Token Drop Podcast · Episode 28
Episode 28 — Jev and “Zero Hallucination”: What It Actually Guarantees, and What It Doesn’t
Also available on YouTube · Spotify · Apple Podcasts
Episode Summary
Sunil opens with a conversation from San Francisco Tech Week, where someone told him Jev “doesn’t hallucinate.” His answer was that it depends on what you mean by hallucination, and Sajjad agrees. Jev is given a fixed set of choices and will never invent a new one, so that kind of hallucination is gone. Whether it picks the correct choice is a separate question. Sajjad gives TypeSafe credit for real progress, and Sam notes that the company’s own launch post says plainly that the 0% hallucination figure is guaranteed by design, not measured, and that Jev can be wrong. The candor is in the body of the post; the headline is what people repeat.
Sam explains how Jev works. It answers three kinds of questions: choice (pick one of up to 255 options you define), score (place something on a scale such as calm, frustrated or angry), and true/false probability. Each answer comes with probabilities and a confidence score. Sunil asks how a model with a confidence score can still be wrong, and Sajjad says that is the heart of the matter: the confidence itself can be wrong. Two launch-week demos show both sides. In one, Jev moved a ping-pong paddle into position in real time. In the other, a version of the trolley problem, it made poor choices between two tracks. Sam lists five places a wrong answer can come from (training data, weights, decoding, output text and the consumer) and explains that Jev removes only two of them. Sunil draws the practical line: a misjudged egg on a sorting line may be tolerable, but a wrong call while flying a plane is not.
Sajjad introduces Kahneman’s Thinking, Fast and Slow: System 1 is fast, automatic perception, and System 2 is slow, deliberate reasoning. Jev is a System 1 model, and reasoning models are closer to System 2. An intelligent platform needs both. Sam explains why Jev is so cheap: it runs only the prefill pass of a transformer, with no token-by-token decode loop, much like a BERT-style encoder with classification heads. The hosts discuss OpenAI’s newly announced Jev-like API and open-source alternatives, and say DaaX is considering a Jev-style model for fast classification and routing inside its pipeline, with ClaimGuard still verifying that answers are grounded. The episode ends with plans for a webinar or whiteboard “chalk talk” to explain the architecture with diagrams.
Chapters
- 0:00 — Tech Week: “Jev doesn’t hallucinate”
- 1:01 — Credit where it’s due
- 2:45 — What’s guaranteed, and what isn’t
- 4:59 — Who supplies the options?
- 5:29 — Jev’s three question types
- 6:42 — If there’s a confidence score, how can it still be wrong?
- 7:08 — Launch-week demos: ping-pong
- 7:52 — Launch-week demos: the trolley problem
- 9:45 — Real-time applications and OpenAI’s egg-sorting example
- 10:52 — Five places a wrong answer can come from
- 12:23 — Errors you can tolerate vs. errors you can’t
- 13:03 — System 1 vs. System 2: Kahneman’s framework
- 16:05 — Jev as System 1, reasoning models as System 2
- 18:45 — OpenAI’s Jev-like API and open-source alternatives
- 19:12 — Using a Jev-style model inside a verification pipeline
- 20:11 — Why Jev is so cheap: prefill without the decode loop
- 23:17 — A webinar is coming
- 25:01 — Reading the fine print on “zero hallucination”
- 26:03 — Speed and cost compared
- 27:32 — Where Jev fits: classify, route, score, and guardrails
- 28:27 — Wrap-up
Full transcript
All opinions expressed are those of the individuals themselves, not necessarily of any organization they work for.
Tech Week: "Jev Doesn't Hallucinate"
Sunil: Okay, Token Drop, Episode 28. I'd like to talk about Jev, because I was at San Francisco Tech Week this week and went to an event today. I was chatting with a guy, and he told me that Jev doesn't hallucinate.
Sunil: I told him that, the way I understand it, it depends on what you mean by hallucination. My understanding of Jev is that it'll have, say, four answers to choose from, and it'll choose one of them. It won't give you a fifth answer, so from that standpoint, it doesn't hallucinate. However, that doesn't mean the answer it chooses is correct. So it all depends on your definition of hallucination. That's the way I understood it, and that's the way I explained it. That's my story and I'm sticking to it. But Sajjad, am I correct or not?
Sajjad: Spot on.
Sunil: Good. I paid attention in our meetings.
Credit Where It's Due
Sajjad: You are spot on. That is indeed correct. But credit should be given where it's due, and I think they've definitely improved on many counts.
Sunil: Improved versus a standard LLM?
Sajjad: Yes, compared to a standard LLM. At least now, when you ask Jev to pick from among four choices, it's not going to pick a fifth or an invalid choice. To that extent, it's been made a little more deterministic. They've chipped away at one of the ways it could hallucinate. But the problem still remains: of the four choices it's required to pick from, did it pick the right one? There, the jury's out, and a lot of tests have shown that it doesn't always. Hallucination in that part of the decision cycle has not been removed yet.
Sajjad: They've also made some real advances. Because this is a single pass through the decision flow, it's much faster. And you can use a smaller model, which translates to cost. It's faster and less expensive, and I think Sam has some metrics. Sam, do you want to pick it up?
What's Guaranteed, and What Isn't
Sam: Yes. This is a very interesting question, Sunil. "Can't hallucinate" is the heart of this episode. It's only true in a narrow sense, as Sajjad said: the output cannot leave the schema. TypeSafe's launch post is explicit that the 0% figure is not empirical. Schema matching is guaranteed, and in the same post, it says Jev can be wrong. So the company is honest in the body and loud in the headline. And the press reaction falls apart after even a moment's consideration.
Sam: So we need to be very careful about what is guaranteed and what is not, and what "can't hallucinate" really means. We've been through this kind of hair-splitting before. Here's what is guaranteed: every answer is one of the options you supply. No type or schema errors, ever. The 0% hallucination figure holds by construction, not by measurement, and the company says so.
Sam: What is not guaranteed is that the chosen option is right. That's very important. It's also not guaranteed that the confidence is calibrated on your data rather than theirs, or that adversarial text inside the input can't sway the judgment. In terms of correctness, a wrong classification is the same failure with a different name.
Sam: On calibration: TypeSafe's calibration evidence comes from its own workflows and synthetic training data. A new language, a new product or a new attacker can quietly break calibration. The only defense is measuring the confidence-versus-accuracy curve on your own traffic, on a regular schedule. That's what Sajjad was talking about.
Who Supplies the Options?
Sunil: Who generates these four options? We keep talking about four options, and Jev chooses one of them. Does Jev generate them, or somebody else?
Sajjad: The user provides them.
Sam: The user. You say what the options are, and it picks from them and gives you a confidence level on each.
Jev's Three Question Types
Sam: If you want to get into how it works, there are three question types and nothing else. The first is called choice: one option from a list you define, up to 255. It returns the pick, each option's probability, and a confidence score.
Sam: The second is called score: a position on levels you describe, like calm, frustrated and angry. It can land between two levels.
Sam: The third is a true/false probability: the probability that a statement is true, from 0 to 1. An answer near 0.5 means unsure, not "medium."
Sam: You can mix all three types in one call. Each question is judged independently, so adding questions barely changes latency, and one answer never leaks into another. That's what they guarantee. I can go deeper on that, but—
If There's a Confidence Score, How Can It Still Be Wrong?
Sunil: Wait, if they have a confidence score, how can it choose the wrong answer?
Sajjad: That's the heart of the matter. The confidence it expresses could itself be incorrect. That's where the hallucination lives.
Sunil: Right.
Launch-Week Demos: Ping-Pong
Sajjad: The day Jev launched, people built lots of very interesting demos. There were good, positive demos showing Jev working and the speed at which decisions could be made, like playing ping-pong. You could see how rapidly the paddle repositioned itself to receive and return the ball. In real time, you were watching Jev decide where to position the paddle. It was fast and good. That was a really good demo.
Launch-Week Demos: The Trolley Problem
Sajjad: I also saw the opposite. Picture a railway lineman whose job is to throw the switch, so the train takes one of two paths at a fork. The lineman makes moral decisions. If there's one person on one track and five people on the other, and an accident is bound to happen whichever way you choose, the lineman makes the morally correct decision, the one that causes the least harm: send the train down the track with one person.
Sajjad: They gave this test to Jev, and Jev made all kinds of incorrect decisions. And it's really just one decision between two options: track A or track B. They would randomly place five strollers on track A and one stroller on track B, or one donkey on track A and five people on track B. They were demonstrating how Jev could make mistakes, and it did.
Sajjad: That is the nub of the matter. Yes, it speeds up decision cycles, and it constrains the choice to the range you want. But the decision itself is questionable. There's no getting away from the fact that there has to be post-decision review and verification.
Real-Time Applications and OpenAI's Egg-Sorting Example
Sunil: It seems like it opens up some real-time applications for AI, like the ping-pong example.
Sajjad: Yes, that was a good example. I saw another one. I think OpenAI announced a Jev-equivalent API today. Their example was eggs on a farm moving along a conveyor belt. A camera makes the decisions: which egg is good quality, which is cracked, which is dirty. Only the good eggs go through, and the dirty, cracked or small ones don't. It's good, it's fast, it's cheap, with very low latency. But those decisions could be wrong, so post-decision verification has to be done even in that case.
Five Places a Wrong Answer Can Come From
Sam: Exactly. If you look at where this fits, there are five places you can get a wrong answer. The first is the training data: what the model saw. The second is the weights: what it learned, including the habit of fluent invention. The third is decoding: how tokens are sampled. The fourth is the output text: the string your code receives. And the fifth is the consumer: whoever acts on it, code or a human. Those are the five places a wrong answer generally comes from.
Sam: What Jev, as a System 1 model, really does is remove two of those stages. There's no decoding loop and no text to validate. What remains is the judgment itself. A wrong pick is still a wrong pick, and everything upstream, the training data and the weights, still shapes it. So, does it really remove hallucination? In a real sense, no. That's the counterargument.
Errors You Can Tolerate vs. Errors You Can't
Sunil: It doesn't, because it can still hallucinate. It just gives you an answer faster, which could be an incorrect answer.
Sam: Exactly.
Sunil: So for applications like egg sorting, it looks like Jev gives you the ability to use AI. If you mischaracterize one egg, maybe it's not the end of the world. Maybe that application can tolerate it. But there are probably a lot of applications that can't tolerate an incorrect answer. Say a real-time application where you're flying a plane, and a wrong answer could cause a catastrophic failure.
System 1 vs. System 2: Kahneman's Framework
Sam: You're absolutely right. That brings up something we were discussing earlier with Sajjad: LLMs as built versus System 1. An LLM as built takes prompt tokens through the transformer stack to the language model head, a softmax over a vocabulary of roughly 100,000 tokens, and produces one token. Whereas with System 1, the state plus the questions—
Sunil: Wait a minute. What is System 1?
Sajjad: There's a famous book by the psychologist Daniel Kahneman called Thinking, Fast and Slow. For the last ten years or more, the AI community has been using that concept. A system that makes a rapid decision with very little input is a System 1 decision. A decision that takes time to think, calculate and reason is a System 2 decision.
Sajjad: Take a human being. If you touch a hot plate, you're not going to calculate whether to lift your hand. You lift it instantly. That's a System 1 decision, driven purely by perception. A System 2 decision is driven by reasoning, calculation and whatever mental processes you choose to apply.
Sajjad: The AI community uses these concepts all the time. In the past, people described the neural network as a System 1 system. It's perceptual, like an image classifier: flash an image and it immediately tells you whether it's a dog or a cat, right or wrong. In Jev's case, it'll instantly tell you whether it's a good egg or a bad egg. That's a System 1 decision.
Sajjad: A System 2 decision is like planning a trip to Boston. You need to work out the route, find the lowest-cost ticket and line up the departure and arrival times with what's comfortable for you. You take time to decide. So in the past, people said LLMs were System 1 systems, and all the work you do with reasoning and knowledge graphs was System 2.
Jev as System 1, Reasoning Models as System 2
Sajjad: Now LLMs are getting to a point where, with what Sam is describing, Jev is treated as System 1, and reasoning models could arguably be treated as System 2.
Sam: Exactly. To answer your question, Sunil: System 1 is fast, automatic and parallel. It's pattern recognition and gut calls. It's confident and sometimes wrong. System 2 is slow, effortful and serial, as opposed to parallel. It's step-by-step reasoning, and it's very expensive. If you look at the latest frontier models, they're pretty slow because they go through step-by-step reasoning, and they're expensive, so System 2 is used sparingly. It takes time to come back with an answer.
Sam: Jev, a classifier that returns probabilities in one pass, is System 1. The knowledge graphs we've talked about, used with LLMs in a neuro-symbolic approach, and what are called reasoning LLMs, with chain of thought, are System 2 kinds of systems. The idea comes from Kahneman's Thinking, Fast and Slow, from 2011. TypeSafe's claim is essentially that a System 1 model can be made more reliable.
Sajjad: And the argument is never one or the other. You probably need both kinds of systems to build an intelligent machine or platform. There are times when you need rapid-fire decisions based on perception, and there are decisions that take time: you need to build a strategy, do some reasoning, do some verification. Both kinds of systems have a role to play, and both are needed.
OpenAI's Jev-Like API and Open-Source Alternatives
Sunil: Since I was at Tech Week, I missed this announcement from OpenAI. You said they announced a Jev equivalent?
Sajjad: They announced a Jev-like API today, so I'm sure they've built a Jev-like model. There are also multiple open-source projects. One of the better-known ones is called Laya. It's open source, so you can download it and play with the model.
Using a Jev-Style Model Inside a Verification Pipeline
Sajjad: We're thinking we might use it for some of our decision work. Given a question and an answer we generated, can we classify whether it's acceptable or not? Should we regenerate the whole answer, resynthesize it claim by claim, or simply abstain from answering? For decisions like those, you could use a Jev-like model, and it would save us money and time. But after the decision, you still have to go back, make sure it was a good decision and do some verification.
Sunil: Because it can hallucinate, like you're saying.
Sajjad: Yes. That's where the potential for hallucination is, and you have to mitigate it.
Why Jev Is So Cheap: Prefill Without the Decode Loop
Sam: And that's why we still need ClaimGuard. Look at what a System 1 model is doing compared with an LLM. How does an LLM produce an answer? Prompt tokens go through the transformer stack. The language model head turns the final hidden state into a softmax over a vocabulary of about 100,000 tokens. One token is sampled and appended to the prompt, and you repeat. Every output token is a separate forward pass that waits for the previous one. That's how normal transformers work. The KV cache makes each step cheaper, but the steps remain serial, and decoding is memory-bandwidth-bound on modern GPUs. We've talked about that a lot.
Sam: The scorer-head approach is different. The state and the questions are formatted into one input. The stack encodes it once, one scorer per question maps that encoding to a distribution over that question's options, and all the scorers run in the same pass. That means adding questions costs almost nothing, so you can have many questions. This is essentially the BERT recipe from 2018: an encoder plus a classification head, with user-defined label sets, on a frontier-scale base. The 255-option cap on choice is suggestive. That's one byte, which smells like a fixed index space in the head.
Sajjad: That's an important point, Sam. Because we've only been talking about four choices, we don't want anyone to go away thinking only four are possible. You can actually have up to 255 different choices.
Sam: Exactly. And the reason it's cheap is that only the prefill pass runs, which is compute-bound and parallel across tokens. There's no decode loop, so there's no serialization. That's why TypeSafe meters input only and calls output too cheap to meter. In the first phase of the transformer, prefill, you can parallelize. In the second phase, decode, you have to go serial, because each step needs the previous output and keeps appending as it works through the cycle.
A Webinar Is Coming
Sunil: Sajjad published a paper on RAG, GraphRAG and neuro-symbolic AI, and people have been asking for something more detailed, maybe a webinar, so they can understand it better. Maybe we can include this whole Jev topic, because your explanation is screaming out for pictures. I only see it through pictures.
Sam: The pictures make more sense, because they show you the exact flow. For example, the difference between an existing LLM and a System 1 model is partly how they're trained. LLMs use RLHF or RLVR, human preferences or automated checkers. With Jev, it's a different form of reinforcement learning focused on calibration, and it's trained on synthetic data only.
Sunil: I'm like the blind man trying to describe the elephant by feeling it. That's how I feel listening to this, because I'm understanding every tenth word. If I saw a picture, I think I'd understand it, so maybe we'll add it to the webinar we've been kicking around.
Sajjad: Yes, Sam has some pictures.
Sam: I can't show them in this podcast. But we can do it and show the entire flow, including where the gate sits in an agentic loop.
Sajjad: You could write a blog post, Sam.
Sunil: Yes, write a blog post too, and then we can turn it into the webinar. I think that would be good.
Reading the Fine Print on "Zero Hallucination"
Sam: Exactly. The one thing I want our listeners to take away is this: even though everyone says there's zero hallucination, you need to understand what's going on behind the scenes. Then you'll see the true meaning of what's being said. The context in which they say hallucination is zero is very narrow, because it's just choosing from a set of options.
Sajjad: Correct.
Sam: And that's not the same as a regular transformer generating text. There's a fundamental difference. It's fair to say it's a marketing term, not a redefinition of hallucination.
Speed and Cost Compared
Sam: If you look at the numbers between traditional frontier LLMs and Jev, frontier models can take anywhere from a few seconds to several minutes end to end. Jev takes about 70 to 500 milliseconds, which is very, very fast.
Sam: On cost, frontier models run anywhere from about 20 cents to $10 per million tokens. If I remember right, Jev is about 4 cents per million input tokens, which is a very low number by comparison.
Sunil: Okay. Sorry, Sam, did you want to say something else?
Sam: Go ahead.
Where Jev Fits: Classify, Route, Score, and Guardrails
Sunil: I think we've covered this. Is there anything else to discuss on Jev?
Sam: The last thing I wanted to say is that existing LLMs have their use cases: chat, copilots, coding agents, verifiable problems. Jev, as Sajjad said, is used to classify, route, score, act as a guardrail and do map-reduce. Those are the applications, mainly classification and routing. That's why we're thinking of using it in our lower layers, and we still need ClaimGuard to ensure the answers are grounded. It's best used for classification and routing tasks.
Wrap-Up
Sunil: You know what I learned today? Next year when I go to Tech Week, I need my wingmen, you guys, to discuss this with people so I can go meet more people. Your explanation would have gone over much better than mine. But anyway, next year.
Sam: We try to make this interesting for our listeners, and hopefully a webinar will do it more justice. I think Sajjad and I can do a webinar on this, walk through a slide deck and present it.
Sajjad: We could do something like a chalk talk. We can meet up in the office at the whiteboard, draw some things and make it a chalk talk.
Sam: Exactly. I know there's a lot of interest. People respond to the articles I publish and ask where Jev fits into the overall picture. Does it replace ClaimGuard? Some even ask what "zero hallucination" means, like you did. There's so much going on that it made sense for us to explain to our listeners what it's really about.
Sunil: So let's plan it. We'll finalize the date and do it at the end of this month. Okay, I think that's it. We'll talk to you later.
Sajjad: Thank you.