Two weeks ago TypeSafe AI released Jev. If you’re reading this, you’ve probably already seen an endless stream of takes about it. I wanted to focus on some practical, applied questions. I’ve written a lot recently about LLM judges, and this is an area where a Jev-like model can be a fit. A model that can ingest unstructured text and output a confidence-weighted judgement much faster (and cheaper) than an LLM sounds like a dream. In this post, I revisit a few of my posts with guidance on LLM judges and explore how that guidance applies to Jev.
Illusion of precision, revisited
In The Illusion of Precision, I cited research showing that while LLMs can provide seemingly precise confidence scores, the scores collapse to a small number of tiers. The recommendation was to ask your judges to provide a score from a small ordinal scale, and to calibrate them on a hand-labeled data set. Jev outputs a probability directly, rather than generating text. So the question is: does that number actually behave like a probability, or is it also a handful of tiers with more precise window-dressing?
I pulled 300 entries from the Nebius SWE-agent data set, and asked two judges (Jev and Sonnet 5) whether the run solved its issue, given the patch the agent submitted (a code diff), and the issue text. Both judges achieved about 77% accuracy, but the numbers in those verdicts were very different. Sonnet returned 32 distinct values, while Jev returned 85. If we compare the reliability curves, we can see that Jev’s scores stay close to the diagonal. This means that Jev’s confidence scores corresponded closely to the actual probability that the run succeeded. Sonnet’s results are much more poorly calibrated, and many are clumped on the lower end.
The advice for small, ordinal scales seems inapplicable to Jev. But while the Jev judge’s calibration was better than Sonnet, its calibration is imperfect, and I only know how it did because I checked. You should still calibrate your Jev judges, but there’s a good chance you’re starting much closer to the truth.
A million token bad habit, revisited
In A Million Token Bad Habit we explored how context is a design decision, and how just because context fits in the window doesn’t mean the model will use it well. I also discussed the importance of location within the context window. To test these on Jev, I sent the same SWE-agent prompts in three ways: the patch diff alone (Patch); the patch code plus the last ten steps of the agent’s run (Tail); and the full trajectory (Full).
The Tail did best on both accuracy and calibration. Full trajectory performed well on accuracy, but was less well-calibrated with probabilities about ten points too low across the board.
I figured this was the junk-drawer effect from the original post: too much irrelevant context in the window. But it wasn’t! To explore this, I used 153 runs that had ten steps or fewer, since for those runs Tail and Full include the same steps with different surrounding context. Cutting the inputs by two-thirds didn’t change anything. Stripping out the agent’s thoughts closed the gap a little bit. But moving the patch and exit status from the end of the state to right after the issue closed most of the gap.
You can see that even when we have the same context, the score jumps nine points solely because of where the patch was in the context window. This is one of the position effects from that post. It appears that for this task, putting the core content of the question at the top of the context window matters most. Changing the structure of context means that judge calibration should be rechecked, too.
Six jobs, or five?
In Should PMs Own Evals, I described the six jobs of an eval program. One of them was calibration:
Calibrate. If you’re using models to evaluate models (LLM-judges, classifiers, etc.), then these measures need to be calibrated themselves because their native outputs may not represent what you expect. Your investment here should be calibrated to your level of risk. Curate smaller, human-labeled sets to check your scorers. Output: scoring models calibrated against a validation set (evals in miniature!)
A decision model like Jev doesn’t eliminate this job, but it may make for a much smaller job! In the previous sections, we saw how our Jev judges started off better calibrated. If you start with a model that’s well-calibrated out of the box, you have much less work to do. But the only way to know for sure is to measure it. And as we saw in the last section, you still have context decisions to make.
Takeaways
Decision models like Jev and OpenAI’s new Decisions API provide an alternative to LLM judges that can be well-calibrated, inexpensive, and fast. I’ve previously argued that you should choose the right model for the job (and that sometimes it’s not an LLM at all), and now we have another in our toolbelt. But a generally well-calibrated judge isn’t necessarily well-calibrated to your problem. And as we saw above, you may find that its performance is still dependent on the design of your context. So there’s still work to do, but these models may change how much time we spend actively tuning and calibrating judges and classifiers vs. just checking them.



