Jev, Regression, and More

Given the recent buzz around Jev, I decided to try it on a few reasoning-intensive regression (RiR) problems. The results were broadly in line with my expectations, Jev is designed for classification, after all. Regardless, I find it worth digging into.

Before getting to the results, let's briefly review what RiR is. RiR refers to natural language regression problems in which predicting a numerical value requires substantial reasoning about the input text. Each instance demands sequential deduction or deeper analysis, beyond identifying surface-level features. The definition is admittedly fuzzy, but it captures an increasingly important class of problems.

Figure 1 places RiR within a hierarchy of natural language regression tasks. I expect most current systems to struggle with Level 3 (RiR), but to fare better at Level 2 (semantic analysis), where fine-grained classification can be an effective approach.

Three levels of natural language regression: feature-based house price prediction, semantic essay grading, and reasoning-intensive math error detection and pairwise RAG comparison.
Figure 1: A hierarchy of natural language regression tasks. Level 1 predicts from supplied features, Level 2 requires shallow semantic understanding of the input, and Level 3 (RiR) requires working through the content before any number can be assigned.

I tested Jev on two RiR tasks: math error detection and pairwise RAG comparison. In the former, we give the model a math problem and an incorrect solution, and ask it to predict the first erroneous step on a scale from 0 to 10. Scores near 0 mean the error occurs early; scores near 10 mean it occurs late. For the latter task, given a query and two answers, we ask the model to predict their relative quality on a scale from −2 to +2. The scale tells us both which answer is preferred (the sign) and how much better it is (the magnitude).

Method

I ran jev zero-shot, with fixed prompts and no task-specific training. Jev doesn't output raw numbers, so each task needed a small wrapper.

For math errors, I split each solution into roughly equal character intervals and asked Jev (via a Choice question) which interval contains the first wrong step. Each interval maps to its midpoint on the 0–10 scale, and the prediction is the probability-weighted average. For pairwise RAG, I used a Score question with five levels, from "substantially worse" to "substantially better," mapped onto −2 to 2. I also tried a Choice over a finer numerical grid.

For comparison, I include baselines from the RiR paper: detailed prompting of GPT-4.1 and GPT-5, and MENTAT, the method proposed in the paper. MENTAT pairs batch-reflective prompt optimization (an LLM reviews its own errors across batches of examples and revises the prompt) with a small neural ensemble that aggregates repeated LLM predictions into a final number.

As in the paper, I report CCC (higher is better) and NMSE (lower is better; always predicting the mean scores 1).

MethodCCC ↑NMSE ↓
Math error detection
Jev — Choice0.65*0.49
GPT-5 — detailed prompt0.690.78
MENTAT — GPT-50.780.42
Pairwise RAG comparison
Jev — Score0.391.35
Jev — Choice0.351.33
GPT-4.1 — detailed prompt0.472.20
GPT-5 — detailed prompt0.312.18
MENTAT — GPT-4.10.520.80

Table 1: CCC and NMSE on the two RiR tasks. The Jev rows are zero-shot, using the wrappers described above; the remaining rows are the corresponding numbers from the RiR paper. *A simple post-hoc calibration (shift and rescaling of Jev’s predictions, fit on training examples and selected on a separate dev set) raises math-error CCC to about 0.70 on the same 750 test examples.

Math Error Detection

For this task, Jev nearly matches detailed GPT-5 prompting on CCC (0.65 vs. 0.69) and beats it by 37% on NMSE.

The NMSE gap doesn't mean Jev is better at finding errors. Averaging over its probabilities pulls predictions toward the middle whenever it's unsure, and that kind of shrinkage is exactly what squared error rewards. You can back this by reasoning over the two metrics: Jev's predictions span at most about 63% of the targets' spread, while the two models have similar correlation with the truth (roughly 0.7). Same signal, different scale.

Jev also had some help. My wrapper handles the position arithmetic, while GPT-5 in the paper had to estimate the fraction itself. In spite of the help, I still find its performance on this task noteworthy!

Pairwise RAG

Jev's CCC lands between the paper's two LLMs: above every GPT-5 configuration, below every GPT-4.1 one. Its NMSE is the best of any zero-shot method, though still above 1, so it's worse than just predicting the mean.

The lower NMSE than GPT-4.1 is misleading, though. Working backward from GPT-4.1's metrics, its correlation with the human scores is at least 0.55, versus 0.41 for Jev. GPT-4.1 carries more real signal; it just spreads its guesses too wide. The finer Choice grid didn't help Jev either, since the difference from Score is within run-to-run noise.

Closing Remarks

There is real interest in small, capable models with low inference costs for classification, and the same likely holds for regression. That's partly why I've kept coming back to the idea of a foundation model for regression for the past year and a half.

Of course, foundation models for numerical prediction are not a new thing. Jev is not the first of its kind in that regard, TabPFN comes to mind, as well as models designed for time-series forecasting. The distinction I'm interested in is the kind of work required to arrive at the prediction, and it goes beyond whether the inputs are tables or text.

Timeline of regression methods from ridge regression in 1970 through LLM-as-a-Judge and RAFT, moving from feature-based to shallow semantic analysis to reasoning intensive.
Figure 2: The regression toolkit over time, from feature-based methods, through (reliable) shallow semantic analysis after transformers, to the reasoning-intensive regime.

I think of RiR as a continuation of the progression from predicting with supplied features, to interpreting the semantics of an input, to performing substantial analysis of what that input implies (Figure 2). In the math task, for example, the numerical rule is straightforward once the first erroneous step is known. The difficult part is establishing which step is actually wrong. Similarly, judging an answer may require resolving a contradiction or recognizing that a seemingly minor omission changes its usefulness. The evidence needed for the score has to emerge from working through the content.

This is what the "reasoning-intensive" part is trying to capture. The boundary is fuzzy, and statistical learning, semantic understanding, and reasoning are not neatly separable. But there is a meaningful question here: how well can a model transfer across numerical prediction tasks where each new example requires this kind of analysis? A model that handles one scoring problem well may still struggle with another whose judgments depend on very different reasoning.

MENTAT was one ad hoc attempt at addressing RiR, but there's also interesting work such as REAL on regression-aware LLM judging. There's plenty left to understand about what works across different tasks. The response to Jev makes me both excited and curious about the demand for models that bring the same practicality to these more involved numerical judgments, and what architectures that exploration might produce.

← Back to home