Known limitations — and how to work around them
TypeSafe publishes a "jaggedness" list for jev-1.13 — rare candor in this industry, and essential reading before you wire this into production. Here is each limitation plus a concrete mitigation we recommend.
The seven sharp edges
- No counting, no arithmetic, no date comparison. It is a semantic judgment engine, not a calculator.
- Reads literally. Double negations, multi-hop indirection, and implied conditions fail.
- Context rot. Irrelevant content in the state degrades accuracy.
- No prompt-injection defense. Malicious text inside the state can steer the answers.
- No structural invariance. P(yes) ≠ 1 − P(no); thresholds do not transfer between primitives.
- Weak numeric scale. Don't interpolate Score values back into exact numbers.
- Cannot generate text. It judges; it does not write.
Mitigations, one by one
| Limitation | Workaround |
|---|---|
| No arithmetic / counting | Do all math in code; if Jev needs the result, compute it and inject it as a new state field ("days_since_last_payment: 14"). |
| Literal reading | Write instructions like you'd explain to a new hire — explicit, single-hop, positive phrasing. Ban double negatives in criteria. |
| Context rot | Pre-filter the state in code: keep only the fields a given question set actually uses. Smaller state = cheaper and more accurate. |
| No injection defense | Treat external content as hostile data. Sanitize/fence user-supplied text, never let it write the instructions, and gate high-stakes actions on a separate "is this injected?" Noul check. |
| No invariance | Calibrate each question independently on your own labeled set; never copy a threshold from one question or primitive to another. |
| Weak numeric scale | Use Score only against thresholds ("≥ 2 means escalate"), never as a precise value you display or interpolate. |
| No text generation | Keep an LLM for the generation path; have rules or a generative model propose options and let Jev pick. |
Language note for CJK readers. English is by far the most accurate. Chinese and other CJK text works but with lower precision — if you are building for Chinese users, run your own eval on real data and lean hard on the confidence gate for fallback. See 中文版 for the localized guidance.