Judging relevance with Jev
There's been a fair bit of interest since the release of Jev, from TypeSafe, a type of model they refer to as a System One model, taking the name from Thinking, Fast and Slow by Daniel Kahneman. System One models are designed to make fast, intuitive, structured decisions, which makes them well suited to a number of search-related tasks, including judging relevance.
Until this release, every AI judge in Releval has been a chat model. Releval sends it a prompt with the query and a result, and asks for a grade and a line of reasoning in return. This works well, and the reasoning is handy when you skim the judgments afterwards. The missing piece is some indication of how sure the model was. A 3 from a model that had no doubt looks exactly like a 3 from one that was much less certain.
The advice in the docs has always been to use AI judges to widen coverage, then sample their work and check it. That's still good advice; the harder part is deciding which judgments to sample.
Releval 1.2.0 adds AI judges that grade with Jev, which helps with exactly that. Rather than generating a grade and a line of reasoning, Jev is asked a question with an ordered set of answers and returns a probability for each one. For relevance judging, that means every judgment comes with a confidence and the distribution of probability across the grades, so the judge tells you where it was unsure, and that's a good place to start sampling. What's more, Jev can be much cheaper and faster than chat-based judges, allowing you to sample more often and get more coverage.
A probability for every gradeโ
A Jev judge is set up like any other AI judge in Releval: a provider, a model and a template. For each result, Releval sends Jev the query, the result's title and its mapped fields, along with one description per grade on the evaluation's scale.
Jev puts a probability on each grade. Releval records the grade as the probability-weighted average, rounded to the nearest grade, and keeps the distribution and a confidence alongside it: 1 when all of the probability sits on one grade, falling as it spreads out.
In the judgments view, each Jev judgment shows that distribution as a bar per grade, with the recorded grade highlighted.
Most of the time it's unremarkable, which is what you want from a judge. Here are two from a small test run against the movie dataset we use for demos, on the graded scale of 0 to 4:
| Query | Result | Grade | Confidence |
|---|---|---|---|
| the princess diaries | The Princess Diaries | 4 | 0.96 |
| lifetime movies about pregnancy | Bee Movie | 0 | 0.98 |
Bee Movie is many things, but it isn't a Lifetime movie about pregnancy ๐
The unsure ones are the interesting onesโ
The judgments at the other end are where the confidence becomes useful. Here's what Jev made of the query "exit room" and the film EXIT:
0 Not relevant โโโโโโโโโโโโโโโโโ 0.34
1 Marginally relevant โโโโโโโโโโโ 0.22
2 Fairly relevant โโโโ 0.07
3 Highly relevant โโโโโ 0.10
4 Perfectly relevant โโโโโโโโโโโโโโ 0.27
Recorded on its own, that's a 2, or "Fairly relevant", but it's the grade Jev assigned the lowest probability of the five.
The distribution tells a more interesting story. Jev split most of its probability between the result being irrelevant and being exactly right, with very little in the middle. Its confidence is 0. Given how ambiguous "exit room" is, that seems reasonable. A chat model would typically give you a grade and a line of reasoning. With Jev, the distribution gives you another signal that this particular judgment is worth looking at.
That gives the sampling advice something to work with:
- Low confidence first. A spread like the one above can mean the grade descriptions overlap for that result, or that the result doesn't contain enough information to place it cleanly.
- Splits between neighbouring grades next. For "a quiet place", Jev spread A Lonely Place to Die across the bottom three grades (0.22, 0.35 and 0.27). When that happens across lots of results, it can be a sign that the distinction between those grades could be sharper.
- Confident surprises too. Confidence describes how certain Jev was in its classification. A confident grade you disagree with is still worth looking at.
Describe situations, not degreesโ
The template for a Jev judge renders the question Jev is asked: some instructions, and one criterion per grade.
The obvious first move is to reuse the scale's labels as the criteria:
"criteria": [
"Not relevant",
"Marginally relevant",
"Fairly relevant",
"Highly relevant",
"Perfectly relevant"
]
It looks right, but in practice it gives Jev fairly little to work with. This is because Jev evaluates each criterion independently against the result. "Fairly relevant", for example, gets much of its meaning from sitting between "Marginally relevant" and "Highly relevant". On its own though, there isn't much concrete information there to match against. Criteria written as degrees tend to spread the probability across grades, which shows up as low confidence across a lot of judgments.
The criteria that work better describe a situation the result is in. Here's the graded scale from the template a new Jev judge starts with:
| Grade | Criterion |
|---|---|
| 0 | The result is unrelated to the query |
| 1 | The result is on the same broad topic as the query but does not help with what it is looking for |
| 2 | The result covers part of what the query is looking for but misses something important |
| 3 | The result satisfies the query, but a more specific or complete result would serve the searcher better |
| 4 | The result is exactly what the query is looking for |
The evaluation run decides the scale rather than the judge, so a Jev template describes the grades for all three scales, branching on the scale with Handlebars.
Releval renders it for every scale when you save the judge, and checks that each produces valid JSON with the expected number of criteria.
The question template docs have the details, including how to add examples to criteria when Jev repeatedly splits two grades you expected it to distinguish more clearly.
Setting it upโ
First up, add a TypeSafe provider with your TypeSafe API key, in Settings under AI Providers, or through the API:
curl -X POST "https://${RELEVAL_HOST}/api/v1/ai-providers" \
-H "Authorization: Bearer ${TOKEN}" \
-H 'Content-Type: application/json' \
-d '{
"type": "typesafe",
"name": "TypeSafe",
"api_key": "${TYPESAFE_API_KEY}"
}'
Next up, create a judge on it. The model can be jev-latest, which moves to each new Jev release, or a pinned version
such as jev-1.13.0. If you're comparing runs judged weeks apart, I'd pin the version so that a new model release doesn't move the grades
underneath the comparison.
Testing the provider or the judge reports which version answered. Then run it against an evaluation run, exactly as you would any other AI judge. Jev grades one result per request and its requests are quick, so throughput is mostly bounded by TypeSafe's rate limits; rate-limited requests are retried.
TypeSafe prices Jev per input token. Since Jev returns probabilities rather than generated text, there are no output tokens to pay for ๐
A few things to knowโ
There are a few differences from a chat-model judge worth knowing before you point it at everything:
- It reads text only. For image judging, a judge using a vision-capable chat model is still a great option.
- The distribution replaces written reasoning. You get probabilities across the grades rather than a generated sentence. I find that useful for deciding what to look at, although it's a different signal from reasoning you can skim.
- It reads literally. TypeSafe's notes on where Jev 1.13 struggles are worth a read. Clear questions and criteria matter, and calculator-like tasks aren't its strength.
- A result can influence its own grade. Jev reads the result as data, but a result written to sound particularly relevant can still pull its grade upwards. That's true of AI judges generally, and another reason to sample their work.
The general advice for AI judges remains the same: use them to widen coverage, and keep people focused on the queries that matter most. What Jev adds is a useful signal for deciding where to look. The confidence and probability distribution make the judgments where the model was uncertain much easier to find.
Where to go nextโ
The TypeSafe provider docs cover setting it up, the question template docs cover writing criteria, and judgments from a Jev judge covers reading the results.
The 1.2.0 release notes have everything else in this release.
I'd be interested to hear how Jev gets on with your queries, particularly where the confidence distributions highlight cases you wouldn't otherwise have looked at.
