Skip to main content

One post tagged with "Jev"

A System One Model for judging relevance with TypeSafe's Jev.

View All Tags

Judging relevance with Jev

· 8 min read
Russ Cam
Founder

There's been a fair bit of interest since the release of Jev, from TypeSafe, a type of model they refer to as a System One model, taking the name from Thinking, Fast and Slow by Daniel Kahneman. System One models are designed to make fast, intuitive, structured decisions, which makes them well suited to a number of search-related tasks, including judging relevance.

Until this release, every AI judge in Releval has been a chat model. Releval sends it a prompt with the query and a result, and asks for a grade and a line of reasoning in return. This works well, and the reasoning is handy when you skim the judgments afterwards. The missing piece is some indication of how sure the model was. A 3 from a model that had no doubt looks exactly like a 3 from one that was much less certain.

The advice in the docs has always been to use AI judges to widen coverage, then sample their work and check it. That's still good advice; the harder part is deciding which judgments to sample.

Releval 1.2.0 adds AI judges that grade with Jev, which helps with exactly that. Rather than generating a grade and a line of reasoning, Jev is asked a question with an ordered set of answers and returns a probability for each one. For relevance judging, that means every judgment comes with a confidence and the distribution of probability across the grades, so the judge tells you where it was unsure, and that's a good place to start sampling. What's more, Jev can be much cheaper and faster than chat-based judges, allowing you to sample more often and get more coverage.