REFLEX LABS

Open models trail Jev by 30+ points on hard decisions. We’re closing the gap, in the open.

A System One model reads a situation and answers typed questions with calibrated probabilities in a single forward pass. Most of what an agent or a workflow does is a long series of these small decisions. Which team gets this ticket? Does this reply follow the policy? Which tool should run next? Is the user about to cancel?

Jev, from TypeSafe, is the reference for this kind of model, and it is closed. Before vera, the best open alternatives scored 94 / 73 / 34 to 40 on JevBench's easy, original and hard splits, against Jev's 100 / 99 / 74. We build open-weights models that speak Jev's exact wire format and narrow that distance, one benchmark split at a time.

The gap to Jev, by split

JevBench accuracy in percent. The axis starts at 30.

Laya (easy, original) · reflex-s1 (hard) vera-core Jev Gap that remains

Each hollow marker is the best open model on that split before vera. Laya scores are from the public leaderboard, not our run. reflex-s1 is our run. Full table →

How we close it

Same interface as Jev

Both models answer POST /v1/systemone with Jev's request and response shape. A choice picks one of 2 to 255 options, a score rates on 2 to 10 ordered levels, and a noul returns the probability of yes. Existing Jev clients switch by changing a base URL, so the comparison is like for like.

Scale where it pays

Architecture changes on a 400M encoder moved JevBench by noise. A larger backbone moved it for real: vera-core, built on Qwen3.5-2B, gained 12.6 points over vera-spark on the hard split.

Calibration, not just accuracy

Jev's value is that its probabilities can be trusted. vera-core's overall calibration error on JevBench is 0.098, so a workflow can act alone when the model is sure and hand off to a slower model or a person when it is not.

Measured honestly

We choose checkpoints on our own held-out sets and treat JevBench as a report, not a target. Competitor numbers say whether we ran them or took them from the public leaderboard.

Open weights, built to be fine-tuned

Both models are Apache-2.0 on Hugging Face. On decisions from your own domain, a fine-tuned vera fits your labels and your calibration far better than any generalist base, open or closed.

Where the gap remains

vera-core gets every easy item right and most of the original split. The distance to Jev now sits almost entirely in a few hard families, where the answer depends on reading a long rule or doing several steps of arithmetic.

Hard familyvera-coreAccuracy
temporal_numeric6/1540%
long_policy8/1942%
probability5/1050%
ambiguous4/757%
judge_hard10/1759%
trap5/863%

Get started

# pip install vera-s1
import vera

agent = vera.load("scar-ai/vera-core")   # or "scar-ai/vera-spark"
result = agent.predict(
    "Hi, we were billed twice for March. Please refund the duplicate "
    "today or we will cancel our plan.",
    {"department": {"type": "choice",
                    "instructions": "Which department should handle this?",
                    "criteria": {"billing": "invoices, payments, refunds",
                                 "technical": "bugs, outages",
                                 "other": "everything else"}},
     "churn_risk": {"type": "noul",
                    "instructions": "Does the user threaten to cancel or leave?"}})
print(result["answers"])