Open models trail Jev by 30+ points on hard decisions. We’re closing the gap, in the open.
A System One model reads a situation and answers typed questions with calibrated probabilities in a single forward pass. Most of what an agent or a workflow does is a long series of these small decisions. Which team gets this ticket? Does this reply follow the policy? Which tool should run next? Is the user about to cancel?
Jev, from TypeSafe, is the reference for this kind of model, and it is closed. Before vera, the best open alternatives scored 94 / 73 / 34 to 40 on JevBench's easy, original and hard splits, against Jev's 100 / 99 / 74. We build open-weights models that speak Jev's exact wire format and narrow that distance, one benchmark split at a time.
The gap to Jev, by split
JevBench accuracy in percent. The axis starts at 30.
Each hollow marker is the best open model on that split before vera. Laya scores are from the public leaderboard, not our run. reflex-s1 is our run. Full table →
How we close it
Same interface as Jev
Both models answer POST /v1/systemone with Jev's request and response shape. A choice picks one of 2 to 255 options, a score rates on 2 to 10 ordered levels, and a noul returns the probability of yes. Existing Jev clients switch by changing a base URL, so the comparison is like for like.
Scale where it pays
Architecture changes on a 400M encoder moved JevBench by noise. A larger backbone moved it for real: vera-core, built on Qwen3.5-2B, gained 12.6 points over vera-spark on the hard split.
Calibration, not just accuracy
Jev's value is that its probabilities can be trusted. vera-core's overall calibration error on JevBench is 0.098, so a workflow can act alone when the model is sure and hand off to a slower model or a person when it is not.
Measured honestly
We choose checkpoints on our own held-out sets and treat JevBench as a report, not a target. Competitor numbers say whether we ran them or took them from the public leaderboard.
Open weights, built to be fine-tuned
Both models are Apache-2.0 on Hugging Face. On decisions from your own domain, a fine-tuned vera fits your labels and your calibration far better than any generalist base, open or closed.
Where the gap remains
vera-core gets every easy item right and most of the original split. The distance to Jev now sits almost entirely in a few hard families, where the answer depends on reading a long rule or doing several steps of arithmetic.
| Hard family | vera-core | Accuracy |
|---|---|---|
| temporal_numeric | 6/15 | 40% |
| long_policy | 8/19 | 42% |
| probability | 5/10 | 50% |
| ambiguous | 4/7 | 57% |
| judge_hard | 10/17 | 59% |
| trap | 5/8 | 63% |
Get started
Your browser blocked automatic copying. The prompt is selected below: press ⌘C or Ctrl+C.
# pip install vera-s1 import vera agent = vera.load("scar-ai/vera-core") # or "scar-ai/vera-spark" result = agent.predict( "Hi, we were billed twice for March. Please refund the duplicate " "today or we will cancel our plan.", {"department": {"type": "choice", "instructions": "Which department should handle this?", "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages", "other": "everything else"}}, "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"}}) print(result["answers"])
Evaluation
JevBench: measuring the gap
JevBench tests System One decision models across three splits. Easy covers extraction, intent, fact and tool selection. Original adds routing, ordinal scoring, policy checks and adequacy. Hard covers multi-hop reasoning, long policies, temporal and numeric questions, probability, judging, traps and adversarial cases.
Accuracy is in percent, higher is better. ECE is expected calibration error over all 231 items pooled in 10 bins, lower is better. The last column is how many points each model trails Jev on the hard split.
| Model | Params | Easy | Original | Hard | Overall | ECE | Hard gap to Jev |
|---|---|---|---|---|---|---|---|
| vera-core | ~2B | 100 | 93.1 | 58.6 | 77.9 | 0.098 | 15.5 |
| vera-spark | ~630M | 100 | 83.3 | 46.0 | 68.8 | 0.141 | 28.1 |
| Jev reference | – | 100 | 99 | 74.1 | – | – | – |
| Laya public leaderboard, not our run | 421M | 94 | 73 | 34 | – | – | 40.1 |
| Julia-1 our run | 144M | 75.0 | 47.2 | 35.1 | 47.2 | – | 39.0 |
| reflex-s1 our run, fast profile | 23M / 83M | 81.2 | 45.8 | 39.6 | 50.2 | – | 34.5 |
Latency
Median latency per question on one AMD MI300X GPU at batch size one: 29–34 ms for vera-spark and 41–44 ms for vera-core. vera-serve batches concurrent requests.

