
Welcome to AIEdTalks’ Newsletter!
In today's edition:
What Jev actually is, and why it isn't a language model
What "type-safe" guarantees here, and what it very much doesn't
Where it fits in an agent, with code you can run today
The one finding in TypeSafe's own data you can use even if you never touch Jev
Let’s dive in.
This edition took a lot of care to put together. If you have a moment, I'd love your rating below.
Rate today's Newsletter
Who's writing this?
Someone who's spent 18+ years building systems that had to work in production, and learned most lessons the hard way.
Today I build multi-agent systems at a large research lab. Along the way: 35+ patents, 17+ publications, and more failed experiments than I can count.
AIEdTalks is where I share what I'm figuring out, while I'm figuring it out. No hype, no polished theory. Just field notes from someone in the trenches with you.
I'd love to know what you're building. Hit reply. I read every message, and your questions often shape the next edition.

Stop rewriting prompts. Start engineering loops. The Code built The Ultimate Guide to Loop Engineering, giving you the exact system Silicon Valley engineers use. Get it free. Claim your Loop Engineering guide
Today’s Edition
Jev: the model that only returns a type

When I saw "can't hallucinate" in a launch post, I assumed marketing and moved on. Then I saw Vercel had already put it in front of every command their agent runs.
So I went and read the docs properly. The claim is real, and much narrower than it sounds.
A model that can't produce text, can't return an invalid value, and can still be wrong.
Here's a piece of code you have almost certainly written this year. An LLM sits at a decision point in your service — route this ticket, is this comment abusive, should the agent retry or escalate — and you ask it to reply in JSON. Then you write the parser. Then the retry when the JSON is malformed. Then a validator, because the model invented an enum value that isn't in your schema. Then you give up on knowing how sure the model was, because "confidence: 0.9" in the JSON body is a number the model made up.
You have built a typed interface over an untyped remote call, in application code, by hand. Every backend engineer recognizes this shape. It's what talking to a service without an IDL feels like.
On September 15, 2026, a startup called TypeSafe AI shipped a model built specifically for that decision point. It's called Jev, and it cannot produce text at all. You send it some state and a set of typed questions; it sends back typed answers with probabilities, in one round trip, usually inside a couple hundred milliseconds.
This primer explains it from zero: what it is, what the "type-safe" claim actually buys you, where it fits, and where it doesn't. It assumes you know static typing, schema validation and service contracts, and nothing about AI.
1. What Jev actually is
Jev is a hosted, closed-weights model from TypeSafe AI, a San Francisco company founded in 2024. Its CEO, Diogo Almeida, spent four years at OpenAI and is one of the authors on the InstructGPT paper — the work behind RLHF, the technique that made ChatGPT possible.
Three things it is not, because the name misleads:
It is not a type system, compiler or DSL wrapped around an LLM.
It is not a language model. There is no text output. No streaming, no chat, no summaries.
It is not related to Typesafe, the company that used to build Scala and Akka.
What it is: a remote function you call when your code needs a judgment rather than a paragraph. TypeSafe's own phrasing is "unstructured state in, typed probabilistic decisions out." You get exactly three answer types:
Noul — a yes/no question, returned as a probability between 0 and 1
Choice — pick one of up to 255 options you define
Score — rate on a scale you define, 2 to 10 levels
That's the whole surface area. TypeSafe calls this category a "System One model," borrowing Kahneman's fast-thinking framing. The pitch behind it, from Almeida: models have been superhuman at chat for years, so where is all the automation? His answer is that a model that's right 95% of the time but can't tell you when it's in the other 5% can't automate anything.
2. What "type-safe" actually guarantees
This is the part worth being precise about, because the launch marketing says "can't hallucinate" and that is not what's on offer.
The guarantee is structural, at the output boundary, and it is hard. You declare the answer set in the request. The model returns a probability distribution over those options. A value outside your schema is not unlikely, it is unrepresentable — the same way a function returning a Rust enum cannot return a variant that doesn't exist. No parse errors, no invented enum values, no retry loop.
What it does not guarantee is that the answer is right. Jev can confidently pick the wrong option from your list. The best summary of this came from the Hacker News launch thread: it can't emit an invalid type, but it can still emit a wrong valid value.
The mechanism meant to catch that is calibration. Every answer carries a probability, and TypeSafe says those probabilities are trained to match real outcome frequencies. Calibration is a statistical property of aggregates, not a checker. Over a thousand answers at 0.9 confidence, roughly 900 should be right. For any single answer, it tells you nothing certain.
Mapping it to what you know: this is the static-typing guarantee plus a runtime health signal. Your type checker guarantees the shape of the value, never the correctness of your business logic.
3. How it differs from an LLM, mechanically

LLM | Jev | |
Output | Token stream you parse | Typed values plus probabilities |
Decoding | Sequential, one token at a time | One pass, all questions in parallel |
Latency | Seconds | 70–500 ms, most around 100 ms |
Input price | $0.20–$10 per M tokens | $0.042 per M tokens, output free |
Context | Model-dependent | 64k for state plus questions |
The parallel scoring matters more than it sounds. Because the questions don't see each other's answers, asking twenty questions about the same state costs roughly the same wall-clock time as asking one. That is a genuinely different cost curve from an agent that reasons step by step.
The flip side: the questions can't see each other, so Jev cannot chain. Ask "is this a refund request?" and "should we approve it?" in the same call and the second answer is not conditioned on the first. Sequencing is your job, in code.
4. Why you'd care
Three reasons, in order of how much they matter to a production system.
The retry loop disappears. Not "gets rarer" — is structurally impossible. That's a tail-latency source and a class of incident removed by construction.
You get a number you can branch on. The documented pattern is a three-tier cascade: high confidence, act automatically; middle, escalate to an LLM; low, route to a human. That's circuit-breaker thinking applied to model output, and you can't build it on a self-reported confidence score.
The economics change what's worth classifying. At $0.042 per million input tokens, a 300-token support ticket costs about $0.0000126. Per-request, per-log-line, per-event classification becomes affordable where an LLM call wasn't.
5. Where it fits
Real reported uses, all in the first two weeks:
Vercel runs the safety reviewer in its fx agent on Jev — it checks every command before execution. Guillermo Rauch reported it as up to 18x faster at p95 than the LLM it replaced.
Browser Use picks the next browser action and target element with Jev, calling an LLM only when text needs writing. They reported a Zürich-to-London flight search in 7.1 seconds for $0.0039.
LangChain ships a ModelRouterMiddleware that uses Jev to decide cheap model vs. expensive model per request.
Content moderation, ticket triage, relevance filtering, eval graders, and agent-loop gates are the obvious rest.
The shape is consistent: Jev sits at the branch point, an LLM does the writing.
6. How to use it
pip install typesafe-sdk # or npm install @typesafe-ai/sdk
export TYPESAFE_API_KEY=...from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient() # pin "jev-1.13.0" in production, not jev-latest
response = client.system_one(
state=ticket_text,
questions={
"department": Choice(
instructions="Which team should handle this",
criteria={
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions",
},
),
"frustration": Score(
instructions="How frustrated the customer appears",
criteria=["Calm", "Frustrated but civil", "Very angry"],
),
"is_urgent": Noul(instructions="The message conveys time pressure"),
},
)
answer = response.answers["department"]
if answer.confidence >= 0.85:
route_to(answer.choice) # act
elif answer.confidence >= 0.60:
escalate_to_llm(ticket_text) # second opinion
else:
queue_for_human(ticket_text) # give up honestlyDirect signups are currently paused because of demand, but Jev is reachable through Vercel AI Gateway, OpenRouter and Cloudflare Workers AI without a TypeSafe account. Integrations exist for Pydantic AI and LangChain.
The API is an hour's learning. The real work is question design: one judgment per question, explicit criteria, filtered state, and thresholds tuned against your own labeled data.
7. Limitations, honestly
TypeSafe publishes a "jaggedness" page listing its own model's weak spots, which is more than most vendors do. The important ones:
It is not a calculator. It doesn't count reliably and reads dates as text, not as ordered quantities. Keep arithmetic and date logic in code.
Prompt injection still works. The docs say plainly that state is treated as data, not as hostile. An injection can't produce an invalid type, but it can absolutely steer the answer toward one of your valid options. If one option triggers a refund, the type guarantee protects nothing.
Accuracy is mid-tier, not frontier. On TypeSafe's own workflow evals Jev scores 67.8%, roughly level with GPT-5.6 Terra and Luna, and behind GPT-5.6 Sol (74.1%) and Claude Opus 5 (73.1%). Those evals were built by TypeSafe and graded against other models' answers rather than human labels.
Closed and hosted. No weights, no self-hosting, no fine-tuning, single region, no published SLA. The service is two weeks old and has already paused signups.
The latest alias moves. A new release can shift probabilities under a threshold you tuned. Pin the version and log the returned model field.
And one finding from TypeSafe's own benchmark page that you can use whether or not you ever adopt Jev: when they took a single big LLM prompt and split it into a workflow of small typed questions, GPT-5.6 Luna went from 51.9% to 66.8% on the same tasks. Most of the gain came from decomposition, not from the new model. Restructure your prompts before you switch vendors.
Common beginner mistakes
Reading "can't hallucinate" as "can't be wrong." It means "can't emit a value outside your schema."
Trusting a single confidence score. Calibration is an aggregate property. Set thresholds from a labeled sample, not from one impressive demo.
Asking a compound question. "Is this urgent and billing-related?" gives you one muddy probability. Split it; parallel questions are nearly free.
Chaining questions in one call. They can't see each other. Sequence in code.
Sending unfiltered state. Accuracy degrades as irrelevant context piles up, exactly like context rot in an LLM.
Skipping the baseline. Run your current model and Jev side by side on your own data before switching.
Key terms
System One model — TypeSafe's label for a model that returns decisions instead of text. Their category name, not an industry standard.
Noul / Choice / Score — the three answer types: yes/no probability, pick-one, and rating.
Calibration — whether stated confidence matches real accuracy across many predictions.
State — the unstructured input blob (text or JSON) the questions are asked about.
Jaggedness — TypeSafe's term for the documented list of tasks their model is bad at.
The one idea to remember
Jev isn't a smarter model. It's a narrower interface — and the narrowness is the product. By making the answer space closed, it turns an untyped remote call into a typed one, and by returning a probability, it gives you something to branch on when the model is unsure.
Which means the question to ask isn't "should I switch to Jev." It's: at every point where your agent makes a decision, do you know how sure it was? If the answer is no, that's the gap, and you can start closing it today by decomposing your prompts into small typed questions. Jev is one way to run them faster and cheaper.
Your turn
Hit reply and tell me: where in your agent would a typed, calibrated answer help most? I read every reply.
AI is easy to demo. Hard to ship.
Sources
TypeSafe AI: Introducing System One Models & Jev (15 September 2026)
TechCrunch: A new kind of AI model from a ChatGPT inventor (18 September 2026)
Notes on the facts: pricing, latency ranges, the 67.8% eval score and the 51.9%-to-66.8% decomposition result are TypeSafe's own published numbers, measured on evals TypeSafe built and graded against other models rather than human labels. The Vercel and Browser Use figures are developer reports, not controlled benchmarks. Small independent tests confirm Jev is fast and cheap with usable calibration, but put its accuracy level with mid-tier LLMs rather than ahead of frontier models. TypeSafe has published no paper, parameter count or architecture details, so any claim about how Jev works internally is unverified. Model version at the time of writing: jev-1.13.
👋 Before you go
💬 Hit reply. I read every reply.
▶️ Watch it instead. This week's breakdown is on YouTube. Subscribe to the channel for more.
📨 Know an engineer who'd find this useful? Share your referral link and earn rewards.
💡 Topic ideas or sponsorships: Reply to this email.
Until next time,
AIEdTalks team.
P.S. AI is easy to demo. Hard to ship. That's what this newsletter is about.
