Pricing & cost math: when Jev beats LLM + JSON
The headline numbers first, then the math. Current version jev-1.13.0 (alias jev-latest): $42 per billion input tokens — $0.042 per million — with output free. Rate limits are 250,000 tokens/sec and 1,200 requests/min (officially "adjusting dynamically"). Context is a 64k total budget; state plus your longest question should stay under 32k (~150k English characters).
Spec sheet
| Item | Spec |
|---|---|
| Price | $42 / Btok ($0.042 / Mtok), input-only billing, output free |
| Rate limits | 250,000 tokens/sec · 1,200 requests/min |
| Context | 64k total; state + longest question ≤ 32k |
| Input | Text only: string / JSON object / array. No image, audio, video |
| Languages | English most accurate; CJK works but less precisely — test on your data |
| Customization | No per-customer fine-tuning; adapt via state content, criteria rules, code-side weights |
The batching effect: why one call beats many
Because questions are evaluated in one forward pass, the number of questions barely changes cost or latency. A community benchmark put it concretely: bundling 13 questions into one call was 12.2× cheaper and 10× faster than 13 separate calls, with identical answers. That single fact reshapes the design: over-query is nearly free, so send every question you might need and let code pick (see Speculative fan-out).
Cost-per-million-judgments (worked example)
The real question is "what does a million judgments cost?" Here is a transparent, assumption-driven comparison. All figures are illustrative estimates — the Jev price is its public rate; the LLM prices are representative of common tiers, not quotes. Run your own numbers with your real token counts.
| Assumption | Jev | Cheap LLM (JSON) | Frontier LLM (JSON) |
|---|---|---|---|
| Price (in / out per Mtok) | $0.042 / free | ~$0.10 / ~$0.40 | ~$3 / ~$15 |
| State + prompt tokens / call | ~300 | ~300 | ~300 |
| Output tokens / call | 0 (free) | ~50 | ~50 |
| Judgments per call (batched) | 13 | ~4 | ~4 |
| Cost per call | ≈ $0.0000126 | ≈ $0.000050 | ≈ $0.00165 |
| Cost per 1M judgments | ≈ $1.3 | ≈ $50 | ≈ $410 |
| Latency per call | milliseconds | ~1–3 s | ~3–30 s |
Reading: at high volume, Jev is roughly an order of magnitude cheaper than a cheap chat model and two to three orders of magnitude cheaper than a frontier model, while being 100–1000× lower latency. The gap widens the more judgments you batch per call.
For the threshold logic that makes these probabilities useful in production, see Confidence-gated routing.