Token Drop Podcast · Episode 23
Episode 23 — Astra, Hallucination Rates, and the Myth of the Self-Sufficient LLM
Episode Summary
OpenAI's newest model, Astra, has consumed the AI press this week — impressive benchmark scores, a steep price tag, and Greg Brockman calling it the start of the AGI era. This episode separates the genuine capability gains from the marketing, and lands on a harder question the industry is currently fighting over: what is a “harness,” and can everything around an LLM eventually get absorbed into the model itself — or is there a control plane that simply can't be learned?
Sunil opens with what caught his attention about Astra: pricing of $10 in and $50 out per million tokens, well above the prior state of the art, alongside a million-token context window and benchmark scores that reportedly leapfrog Anthropic's Fable 5.1. Sajjad and Sam walk through the numbers — a 99.5% score on ARC-AGI-2, strong results on OSWorld, and demos converting real estate listing photos into full video walkthroughs.
The back half of the episode is a genuine argument about the word “harness” — a term Sunil finds confusing, since it evokes a horse and buggy rather than anything resembling AI infrastructure. Sajjad defines DaaX's harness as everything that happens before and after an LLM's inference: building context from a knowledge graph upstream, then verifying, rectifying, and regenerating output downstream. Sam introduces a sharper distinction circulating in industry writing — a “cognitive harness” (context, tools, planning) that vendors argue will eventually get absorbed into the model itself, versus a “control harness” (identity, authorization, audit, runtime) that he argues fundamentally cannot be, because a model is a passive file of weights with no way to authenticate itself or hold credentials. Sajjad pushes back hard on the anthropomorphizing language common in the industry: an LLM doesn't learn anything unless someone builds a training loop around it, and conflating that with genuine autonomous capability is scientifically loose.
The conversation closes on a genuinely new wrinkle: Astra no longer exposes its reasoning traces to users, a change Sajjad attributes to competitors mining those traces to train cheaper open-source models. The practical effect is that you get an answer with zero visibility into how it was reached — which, as Sajjad argues, makes independent verification not optional but mandatory, regardless of how good the underlying model claims to be.
Chapters
- 0:00 — Introducing Astra: pricing, context window, and the AGI buzz
- 1:06 — Sajjad's take: leapfrogging Fable 5.1, ARC-AGI-2, and the Zillow demo
- 2:53 — Sam's benchmark rundown: OSWorld, speed, and pricing detail
- 4:20 — Hallucination rates: down, but not to zero
- 5:43 — Was Astra rated “critical” on OpenAI's preparedness scale?
- 6:11 — What does “harness” actually mean?
- 8:06 — DaaX's definition: controlling the LLM before and after inference
- 9:53 — The verification engine: no trust without it
- 10:03 — Cognitive harness vs. control harness
- 17:08 — Is this the first model to use something beyond just an LLM?
- 17:14 — Why an LLM is a passive .bin file, not an active agent
- 20:52 — How hallucination rates are actually measured
- 24:20 — The case against “everything gets absorbed into the model”
- 26:17 — Why explainability can't come from a black box
- 28:39 — Digital twins as the real-world alternative for process optimization
- 29:47 — Pricing vs. Gemini 2.5 Pro, and the coming need for LLM routers
- 30:49 — Astra no longer shares its reasoning traces — and why
- 32:22 — Tokenomics, price elasticity, and whether cost should drive architecture
- 33:47 — Cognitive vs. control plane: what can and can't be absorbed
- 38:41 — Wrap-up: verification is not optional
Full transcript
All opinions expressed are those of the individuals themselves, not necessarily of any company they work for.
Introducing Astra: pricing, context window, and the AGI buzz
Sunil Baliga: Hey, Sajjad, hey Sam, Token Drop time. This week was Astra, and I'd like to talk about it — not Brad Pitt's Astra movie, Astra from OpenAI.
Sunil: Some interesting things I saw about it — it's obviously consumed all the AI press recently, but the pricing really stood out to me. It's $10 in and $50 out. Some of the other models I'd been looking at before, the state of the art was less than $10, maybe $15 in that range. So it's significantly more expensive. They've also increased the context window — it's a million tokens now. A lot of buzz around it — Greg Brockman said it's really the start of the AGI era. What are your thoughts, Sajjad?
Sajjad's take: leapfrogging Fable 5.1, ARC-AGI-2, and the Zillow demo
Sajjad Khazipura: Honestly, it's a really good model. It probably is as good as it gets, because they seem to have leapfrogged Anthropic's Fable 5.1. They check many of the benchmarks out there — especially ARC-AGI-2, where they scored 99.5%. That benchmark was specifically designed to check for early signs of AGI.
Sajjad: It seems like a great model — I saw some of the demos, truly outstanding. One example: Zillow, the house listing marketplace. If you're selling a house, you're listed there with all the pictures. You can provide those pictures to Astra, and it converts that into a walkthrough video of the house. Really impressive demos. Very few people have access to it right now — early, limited access to a few select customers, and everyone's hoping to get their hands on it.
Sajjad: But I think Gary Marcus said it right: it's a good model, but you still need verification.
Sam's benchmark rundown: OSWorld, speed, and pricing detail
Sam Pooni: This is very interesting. What we're seeing here — they're now rolling this out to paid ChatGPT tiers. The numbers get to a point where you stop believing them. Like Sajjad said, 99.5% on ARC-AGI-2 — very significant. There's also a perfect score on something they call ExploitBench.
Sam: The one I care about is OSWorld, and that's 72.6%, finishing in about 4 minutes instead of 75. What they're doing is bringing better and better models right now. The pricing is about $10 per million tokens in and $50 out, and a fast mode doubles the speed at double the price.
Hallucination rates: down, but not to zero
Sajjad: One thing to note, Sam — sorry for interrupting — is the hallucination rates. They've reduced them, but they haven't gone to zero. It's about 4.1%, or roughly 4.2%. And even on Artificial Analysis, the site we keep referring to for hallucination rates — I think they're saying it's roughly 51% for the kinds of questions the LLM should have said it doesn't know, or declined to answer. It still goes ahead and answers, and 51% of those answers are incorrect. That's how they measure hallucination.
Sajjad: So it's quite noteworthy that despite all the training, hallucinations remain. It still has outstanding abilities, but not a lot of people have had the opportunity to test it yet.
Was Astra rated “critical” on OpenAI's preparedness scale?
Sam: The main question I have for Sajjad — did OpenAI rate Astra as critical on its own preparedness scale? If you read the release, they claim Astra as critical on that scale. What does that rating actually change, or is it just a label at that point?
Sunil: No, I did not read that.
Sam: So one of the things I wanted to ask about was this stuff about harnesses.
What does “harness” actually mean?
Sunil: Right, I read some of the stuff about harnesses too. A harness to me evokes the idea of a horse with a harness, pulling a cart behind it. But the harness, as I understood it, is really more of an output gate — to check and make sure. It's not really like a horse-and-buggy kind of thing in my mind. It's the wrong word, I think. What is it harnessing? Why is it a harness?
Sajjad: I'd say a harness on a horse serves two purposes. One, it's a way of using the horse. Two, you're still controlling the horse. If you don't have a harness with a bridle through the nose and those two reins, you're unable to control the horse. You want to be able to control it.
Sajjad: So in this case, the harness is about controlling the LLM. The LLM by itself is a very powerful tool, but it's like a bull in a china shop. You want to be able to harness its energy for doing the right things. So in that sense, I think it's probably the right usage, but it's a very ill-defined term as it concerns LLMs.
DaaX's definition: controlling the LLM before and after inference
Sajjad: For us, the definition of a harness is our ability to verify the outcome of each inference of an LLM, then rectify the output in case it's not right, and iterate through it until we believe the quality bar has been met.
Sajjad: Our harness is a little more elaborate — it's guided by a knowledge graph. The graph models precise knowledge in a fairly precise way, and we use that graph, plus the controls we have, to constitute a harness. But there's no standard definition of what a harness is, so everybody can interpret it as they deem fit.
Sunil: Going back to my chip days, to me it's like an error correction thing — you get some signals off, you double-check if the bits line up with the error code, and if not, you don't pass it along, you say retransmit. I thought that's what the harness was here, but it's a little different, or it's open to interpretation.
Sajjad: I think it is open to interpretation. As Sam was saying, a harness can just be a simple loop with managed context. But for us, the definition of harness is not just managed context — we also manage the outcomes. What good is a harness for us if we let the LLM loose to generate whatever it thinks is appropriate? It's like writing code and shipping it to customers without any testing, without verification. So you've got to have a verification phase.
The verification engine: no trust without it
Sajjad: We have a verification engine, because no matter how good the model is — how can you trust the outcomes unless you verify them?
Sam: Exactly. I think when we talk about the harness part, many people have different definitions of it. You talk about context, goals, roles, rules — that's basically the agent build pack. Normally, in agents, you read agents.md. There's also this whole notion about tools — MCP servers, an AI gateway. Then you have memory — agent memory services, data services, a data intelligence foundation. Then you have skills, specifically playbooks. You have identity, guardrails. So like Sajjad said, the whole definition of harness is to verify.
Cognitive harness vs. control harness
Sam: And the runtime pieces of the harness — the whole supply chain, secrets isolation, routing, bindings — I think that's part of it too. Some people are saying all of this can get absorbed into the LLM eventually, and what remains is HIL — human in the loop. My understanding is: how is that even possible? I really think that's the question we need to address. Is that even possible, and how does it work?
Sam: This is the question for Sajjad: what do you think the harness is — the deterministic boundary around it, the cognitive part, the neuro-symbolic part?
Sajjad: You're right, Sam — I think a harness also needs to include what happens before the LLM runs its inference, and in our definition, it also includes what comes after. Further upstream, we prepare the context from our graph. Some people who don't use a graph prepare it from documents. That gives the LLM greater latitude, and by that definition, you're also allowing it to potentially make errors downstream. So we try to control errors both upstream and downstream — that's our definition of a harness. Given the query, we find the right context: traverse the graph, find the right data points, organize and rank them, augment them with additional information needed, and prepare the package for the LLM to run an inference on. Then post-inference, we do all the verification, rectification, and regeneration. That's the saddle we use to ride the LLM.
Sam: The main interesting thing I found was an article that talked about two types of harnesses. A cognitive harness — context, tools, skills, planning — which is being absorbed, and vendors selling it are on borrowed time. Then a control harness — identity, authorization, audit, limits, runtime — which can't be absorbed, and that's where the whole bet lies.
Is this the first model to use something beyond just an LLM?
Sunil: Let me ask — is this the first time they've had a model that uses something other than just an LLM? You've used vector databases in the past for storing documents, but this seems to be something outside the LLM.
Sajjad: No, no — I think this discussion is going off in a different direction. What is an LLM? What's a neural network at the end of the day? It's a bunch of weights, and then there's the topology of the neurons. That's all. An LLM is not even an active element — it doesn't come to life on its own.
Why an LLM is a passive .bin file, not an active agent
Sajjad: You have to build a process for inferencing, for serving. So for someone to say the LLM will absorb all of that — an LLM is not even an active element, it's a very passive element in that whole stack. You have to build everything around it to make it useful. It's a .bin file — a binary file lying there. On its own, it can't do anything.
Sajjad: So for someone to say the LLM is learning all of this — people tend to anthropomorphize LLMs, according them human abilities. It spices up a conversation, but as scientific people, you have to be true and scientific about it. An LLM will not learn on its own unless somebody puts it in a learning or training loop. And what can it learn? Only from its inferences, if somebody built a reinforcement learning loop around it, or did supervised fine-tuning based on user feedback. One can argue you can improve the LLM's capabilities for that particular application — but how will it absorb identity? How will it absorb all the surrounding infrastructure that's out there? I think people are just being too loose with the word — that all of this is going to get absorbed in there.
Sam: Exactly. It was very different when I heard this for the first time, and I was surprised people put this proposition forward. I said it's not possible — the control part is still there, and you can't do anything about it. How can you absorb that?
The Oracle parallel: land with the database, expand with everything else
Sajjad: One can argue that all the tooling and the harness around an LLM is something the frontier lab companies — OpenAI, Anthropic — might build, and use their LLM as a footprint to upsell everything else. As it happened in the past: Oracle was a database company, then they started upselling all the middleware around the database, then applications, ERP applications, around the database. But the database was the product that would land and then expand their footprint inside an enterprise. To accord anything more than that, I think, is not being scientific.
How hallucination rates are actually measured
Sunil: One of the things you said was really interesting on hallucinations — one benchmark was about 4% hallucinations, not zero, but roughly 4%. The other was about 50%. What benchmark is that — you mentioned something about domain? I didn't quite catch it.
Sajjad: I don't know what benchmarks OpenAI is using when they report bringing hallucination rates down to 4.2% — we don't quite know what benchmarks or datasets they used. But Artificial Analysis, the site we keep referring to, has certain benchmarks, and their definition of a hallucination is: given a question, if the LLM didn't know the answer and still went ahead and responded, what's the rate at which it gets those questions right or wrong? That's their definition of a hallucination rate. Those rates are quite high — some of the bigger models have hallucination rates of up to 96%. According to them, Astra was at about 50 to 51%.
Sam: Actually, if you look at it, that 58% figure is for GPT-4 on legal QA, and 88% for Llama 2, with strict grading criteria.
Sajjad: Correct — if you read that list of metrics, it's pretty scary. There's another leaderboard from a company called Vectara — they used to have a hallucination leaderboard, and for a long time I'd read it and feel good that hallucination rates were coming down into single digits. But more recently I visited their website and saw that the bigger models had hallucination rates back in double digits.
Sunil: We'll put those links in the comments so people can check them out.
The case against “everything gets absorbed into the model”
Sam: I honestly believe there's a lot of loose thinking around these things, and what it's doing to the community is that some people read these blog posts and say, “hallucination — what are you talking about?” There's this whole idea being pitched that everything will get absorbed into the LLM, that it's becoming more super-intelligent, trained on data in such a way that you don't need a symbolic counterpart. But what I'm asking is — if you go to any government organization and they're asking for provenance, for explainable AI, how are you going to get that from the LLM?
Sajjad: It's a black box. It's a bunch of weights, and you're running inferences using those weights. It's quite opaque — there's nothing you can get out of the LLM. If you're denied a credit card and you file a case asking why, there is no way an LLM will answer that question. Those are the kinds of transactions that require explainability, and explainability means transparency in how the decision was arrived at and delivered.
Why explainability can't come from a black box
Sam: Another point I was thinking about last week — you have people working on a business process who want to optimize it over time. Would just a loop help with that? I don't think so. My understanding is: if you want to improve and optimize a business process, you need to execute it in the real world. If you execute it in the real world, there's a cost associated with failure. If you don't want to do that, what's the alternative? It's the digital twin.
Digital twins as the real-world alternative for process optimization
Sunil: That comes to mind, again, with the pricing point I started with — $10 in, $50 out, for a million tokens. Compare that with Gemini 2.5 Pro, which I believe we talked about — roughly $1.25 in, $10 out.
Sunil: So I'm sure that Astra pricing will come down — it's initially being kept high because of limited availability. When it goes mass-market, it'll come down, but I don't think it'll come down to a dollar-and-ten-bucks price; I think it'll stay more expensive. And not every application, not every query, will need Astra. Some may be perfectly fine with Gemini, or ChatGPT-4o Mini, or Llama 2, who knows.
Pricing vs. Gemini 2.5 Pro, and the coming need for LLM routers
Sunil: One challenge I always have as a user is which model to choose. If I go to Anthropic and start using up my credits on their latest model, I'm going to go back to a cheaper one. But how do I know? I think applications will have the same challenge — they'll need some sort of LLM router, to choose between models, maybe start at the cheapest one you guess will be okay and move up from there. These models are very capable, but very expensive, and they may not give you the advantage for your particular application that justifies the cost increase.
Tokenomics, price elasticity, and whether cost should drive architecture
Sam: You're absolutely right — the cost metrics behind these things are going to come down going forward, but generally, when cost goes down, usage props up. So even though the cost went down, you should have gotten a smaller bill, but you don't, because people start using it in different ways, and the overall cost goes up over time. Whenever you reduce the price of a product, people don't buy less of it — they buy more.
Sunil: That's price elasticity of demand — that's the curve.
Sam: Right. So even though the cost of tokens is coming down, I still think people will use a lot more of them. But the bigger question is: should we even build an architecture around tokenomics? Why do we need to take tokenomics into account when we're doing an architecture? I think the fundamental architecture would differ if not for tokenomics.
Sunil: Sorry, what do you mean by tokenomics — the cost-benefit ratio?
Sam: How many tokens are being utilized for reasoning. I don't know how much of what's being developed is based on how many tokens get expended. If you're taking tokenomics into account, are you actually reaching the ideal architecture? Sajjad, you can probably answer that better.
Astra no longer shares its reasoning traces — and why
Sajjad: Thinking back to Astra — one big thing that's happened is that the reasoning tokens, the reasoning output, is not shared with end users anymore. That's a big change. In the past, reasoning traces were available, and the worry these companies had was that those traces could be mined and harvested by lower-cost, open-source LLM providers to train their own models. So this time, the reasoning traces are also not available in Astra. Effectively, you ask Astra a question and get an answer — you don't know how the answer was arrived at. It's up to you to decide whether to trust that answer. Conventional wisdom tells you that you need to run a verification step after that. For those who argue you don't need to — you're taking a risk, flying by the seat of your pants, hoping every answer is right.
Sam: Where I was getting at is why tokenomics becomes very important — because now companies are going to say, we're not going to publish those results. It's like a company that doesn't want to show certain things in their balance sheet — they group them together and just tell you the total. There may be real reasons for them not to provide those reasoning traces, but at the end of the day, it's like saying, “I'm not going to give you the numbers, just look at my overall balance sheet and see if I'm doing well.”
Cognitive vs. control plane: what can and can't be absorbed
Sam: I don't know if that's good or bad. But ultimately, going back to whether the model is absorbing everything — we already said the cognitive parts can get absorbed: planning, routing, interpretation. But there are authoritative parts that remain — identity, durable state, transaction boundaries, audit trails — none of that is going away. So bottom line: what's getting absorbed is the LLM absorbing the agent's control policy, maybe, but not the control plane itself. It cannot. And that's the main thing — the control plane stays where it is, and you can't do anything about that.
Sam: Astra is good, but I'd certainly question the hallucination rates against what kind of harness it was tested with. How do we know what they claim is the truth? Numbers without verification mean nothing to me. I really think the benchmarks may be built in a way that favors their results — they should publish on what basis they're claiming a 4% hallucination mark.
Wrap-up: verification is not optional
Sajjad: What I've started to see is language from the LLM providers themselves saying you need to do your own verification. They don't trust the LLM for everything it says — you need your own verification. So the need for verification isn't going away. While they may have reduced their hallucination rate according to their own claim, according to their own benchmarks — we don't know what those benchmarks are — the point still remains: you need to run some verification on your own.
Sunil: Okay, I think that's it, because we ran over. Thank you both, we'll talk to you later.
Watch on YouTube
Watch on Spotify
Listen on Apple Podcasts