Architecture patterns, with working code
Four recipes cover most of what people build with Jev. The unifying idea: split big judgments into atomic questions, and keep the composition logic in your code. Below each pattern we show the shape, when to use it, and — for the workhorses — actual code.
1. Speculative fan-out
Send every question you might need in a single call — including ones you may end up ignoring — and let code pick what is relevant. Because questions run in parallel and cost pennies, err on the side of more questions. The 13-in-1 benchmark (12.2× cheaper than 13 calls) is the economic reason this is a feature, not a waste.
2. Confidence-gated routing workhorse
Treat confidence as a second decision axis. High → act automatically. Medium → ask for confirmation. Low → escalate to a human. Set the threshold by risk tier: read-only actions can run at 0.5; money or irreversible actions want 0.9+. A third-party test of Norwegian court letters found the model's wrong answers were precisely the ones it reported lowest confidence on — the uncertainty signal is doing real work.
from typesafe import TypeSafeClient
import logging
client = TypeSafeClient() # reads TYPESAFE_API_KEY
log = logging.getLogger("jev.gate")
# risk tier -> minimum confidence to auto-act
THRESHOLDS = {"readonly": 0.50, "write": 0.80, "money": 0.95}
def gate(action_risk: str):
"""Returns (decision, tier) where decision in
{"auto", "confirm", "human"}."""
a = client.system_one(
state=STATE,
questions={
"safe": {"type": "noul",
"instructions": "Is this state safe to act on now?"},
"intent": {"type": "choice",
"instructions": "Primary intent",
"criteria": {"refund": "Asks for money back",
"status": "Asks for status",
"other": "Something else"}},
},
)
p_yes = a.answers["safe"]["noul"]
p_top, top = max(a.answers["intent"]["probabilities"].items(), key=lambda kv: kv[1])
conf = a.answers["intent"]["confidence"]
tier = THRESHOLDS.get(action_risk, 0.80)
# a "not safe" judgment always wins, regardless of tier
if p_yes < 0.50:
return "human", "unsafe"
if conf >= tier:
log.info("AUTO %s conf=%.2f", top, conf)
return "auto", top
if conf >= tier * 0.8:
return "confirm", top
log.warning("ESCALATE %s conf=%.2f", top, conf)
return "human", top
tier per action, not globally. And log everything — the low-confidence escalations are your free evaluation set for tuning thresholds.3. Composite scoring
Break a complex rating into several Score questions (e.g. bug severity × customer sentiment × reproducibility), then combine them with weights in code. The benefit: you tune by editing numbers in your source, not by re-prompting, and each dimension stays interpretable.
4. Intent routing
Choice-classify the user's intent, then route to deterministic code, a specialist LLM, or a human. The payoff is that the large majority of simple requests never touch an expensive model at all — Jev is the cheap triage in front of your LLM budget.