Updated 29 September 2026 · by Amy Wilkinson
Tool guide
Jev, explained for growth teams.
Jev is a model that makes decisions and writes nothing. That makes it cheap and fast enough to judge a whole market, and the wrong tool for half the jobs people are trying to give it. Here is how it works, and where we put it in a build.
What it is
A model that only judges.
TypeSafe announced Jev on 15 September 2026, in early access, as the first of what it calls System One Models. You give Jev a situation, which TypeSafe calls the state, and one or more questions. It gives back typed answers with probabilities. TypeSafe's own summary is "unstructured state in, typed probabilistic decisions out."
Because every answer has to match options you defined in advance, Jev cannot return something outside them. It can still pick the wrong one, and TypeSafe says so plainly. What you get is consistency: the same question, applied the same way, to every row you hand it.
Most of the work is writing the question. A vague question gets you a fast, cheap, confident wrong answer, which is the worst kind.
How you ask
Three kinds of question.
You can ask many questions about one state in a single call, and TypeSafe answers them in parallel.
Is this true?
A statement you wrote, answered with a probability. "This company sells to other businesses today."
Which one?
Pick one option from a list you wrote, up to 255 options, with probabilities and a confidence score.
How much?
Rate against ordered levels you defined, up to 10 of them, with probabilities and a confidence score.
Cost, speed and limits
The numbers, from the source.
| Current model | Jev 1.13, jev-1.13.0. The jev-latest alias pointed here when we checked. |
| Price | USD 0.042 per million input tokens. Output is free. |
| Context | 64k tokens per request, and 32k for the state plus the longest question. Text input only. |
| Rate limits | 250,000 tokens a second and 1,200 requests a minute, which TypeSafe says it adjusts as it goes. |
| Speed | TypeSafe quotes 70 to 500 milliseconds end to end. In our own trials the median call took about half a second. |
| Language | English first. Other languages work, in TypeSafe's words, "but not equally well". |
| Data | TypeSafe says it does not train on customer requests or responses. Zero data retention is an enterprise option on its own API, and listed on Cloudflare's model page. |
| Weak at | TypeSafe's own list includes counting and maths, comparing dates, reading instructions too literally, adversarial text and anything that writes. Keep arithmetic and dates in code. |
From TypeSafe's launch post and docs and Cloudflare's model page, checked on 29 September 2026. The speed and price multiples on TypeSafe's homepage come from its own evaluations, which it describes as the higher end of real-world gains.
How to reach it
Six ways in.
api.typesafe.ai/v1/systemone, or try questions in the Playground first.typesafe-sdk, with sync and async clients and Noul, Choice and Score classes. Python 3.10 or newer.@typesafe-ai/sdk on npm, for Node.js 20 or newer.typesafe/jev, called from a Worker or Cloudflare's REST endpoint. This is the route we used for our trials.pydantic-ai-slim[typesafe]. Jev can produce an agent's structured output and choose tool calls, and hand anything that needs words to a language model.langchain-typesafe, with examples for routing between models and gating risky tool calls.What we found
The question mattered more than the model.
We ran our own trials in September 2026, sending public company facts only. Asked to tag 69 companies in our own pipeline by the kind of person we would contact and the kind of trigger behind the timing, Jev matched our hand tags on 69 of 69 and 66 of 69. The whole run cost about a quarter of a cent.
As a tripwire for prompt injection hidden in scraped profile text, it flagged 19 of 20 synthetic attacks. We judged the one miss to be a mislabelled sample.
The result that changed how we use it came from 25 companies an earlier sweep had thrown out. Asked the sweep's own category question, Jev agreed with a careful re-read on 6 of 25. Asked a narrower one, "does this company sell what we sell?", it agreed on 20 of 25, in two seconds. The model was the same in both runs. The question was better.
What it could not do was tell us who would buy. Three written replies cannot validate anything that ranks, and we do not claim it can.
The longer version, with a 998-company test and ten builds we would try next, is in 10 growth jobs I'd hand to a model that only makes decisions.
A classifier does not fix your thinking. It scales it. A person writes and argues about the question; the model applies it to everything at the same standard.
Where it fits
Good jobs, and jobs we never give it.
- A first cut of a scraped or bought list, before any step that costs real money
- Tagging rows against a legend you wrote: trigger type, persona, stage
- Ordering a queue so a person reads the likeliest rows first
- A second reader that flags where it disagrees with the first
- A tripwire for injected or off-topic text before a language model reads it
- Routing a task to the right model or tool
- Approval of any kind
- Exclusion and do-not-contact lists
- Deciding whether two records are the same person
- Counts, thresholds, maths and dates
- Anything that writes text
Questions
Questions about Jev.
What is Jev?
Who makes Jev?
How much does Jev cost?
Can Jev write emails or posts?
Can Jev predict which companies will buy?
Is Jev safe to use with customer data?
Related
Start here
Let's build what runs without you.
Twenty minutes on how growth runs in your company today. If a build is not the answer, we will tell you what is.