Home / What is Jev? An independent take on the System One model

What is Jev? An independent take on the System One model

Analysis·By OpenJev Editorial·Updated 2026-10-03·8 min read

Jev is the first model in a category its maker calls System One: it does not generate a single token of text. You send it a state (structured facts) and a batch of typed questions, and it returns structured answers with calibrated probabilities. Your code reads the answers and branches. That is the whole product.

Where the name comes from — and why it matters

The name is a deliberate nod to Kahneman's Thinking, Fast and Slow. System 1 thinking is fast, automatic, intuitive; System 2 is slow, effortful, deliberate. Generative LLMs are System 2 machines: they are brilliant at producing text, but when you press them into service as a classifier or a gate, you pay for slowness, cost, and a text output you must parse back into structure.

TypeSafe's thesis is that most production AI work is not generation — it is decision. Routing a ticket, scoring a risk, checking whether a step is safe, re-ranking a retrieval: thousands of small, fast, yes/no-ish judgments per minute, machine to machine, no human reading the output. Jev is built for exactly that layer, and nothing else.

Our read: treat "System One" as a decision layer sitting under the generative layer. The interesting architectural shift is not "Jev vs LLM" but "which calls belong in the decision layer and which in the generation layer." Getting that split right is most of the value.

How it is trained: RLCD, a third path

Post-training has had two famous recipes: RLHF (which produced chat assistants) and RLVR (verifiable-reward RL, which produced step-by-step reasoners). TypeSafe's CEO, Diogo Almeida, co-invented RLHF (InstructGPT). Their new recipe is RLCD — RL for Calibrated Decisions — which trains the model to emit a decision plus a calibrated probability rather than text.

Calibration is the load-bearing feature. A model that says "80%" and is right about 80% of the time can be trusted as a control signal. That is what lets you use its output as a threshold in production logic instead of eyeballing a paragraph.

How it runs: one forward pass, no autoregression

Every question in a request is evaluated in parallel and in isolation in a single forward pass. There is no token-by-token decoding. The consequences are practical, not academic:

  • Output tokens are free — you only pay for the input (state + questions).
  • Latency is milliseconds, and it barely moves as you add questions — adding a question is nearly free, which changes how you design calls (see Speculative fan-out).
  • Questions cannot leak into each other — isolation means one tricky question doesn't contaminate another's answer.

What this is not

It is not a chatbot, not a replacement for your LLM, and not a general-purpose "AI brain." It cannot generate text, it is weakest at anything numeric or multi-hop, and it has no built-in defense against prompt injection in the state. If you arrive expecting a smarter GPT, you will be disappointed; if you arrive looking for a cheap, fast, trustworthy judgment primitive to wire into software, this is a genuinely new tool. The rest of this site is a field guide to using it that way.