5
Views

It’s been a loud week for frontier AI. Anthropic shipped Claude Fable 5.1 on Tuesday, September 1. Two days later, on Thursday, OpenAI launched GPT-6 Astra, and company president Greg Brockman closed the briefing with a line built for headlines: “Welcome to the AGI era.”

Within days, the story stopped being about AGI and started being about a cache-read rate.

Two Scoreboards, Two Winners

OpenAI’s own launch materials put Astra ahead of Fable 5.1 on nearly every benchmark row it chose to publish. Artificial Analysis, an evaluator with no stake in either company, has the opposite result: Fable 5.1 leads Astra 66 to 61 on its Intelligence Index, and 70 to 67 on its Coding Agent Index — currently the highest Intelligence Index score the firm has recorded for any model.

Some of that gap is a measurement artifact rather than a capability gap. Fable 5.1 ships with more conservative safety filters than its sibling, Claude Mythos 5.1 — the same underlying weights, running with looser restrictions for vetted cybersecurity and life-sciences organizations. On Anthropic’s own Terminal-Bench 4.0 numbers, Mythos scores five points higher than Fable on an identical model, a gap Anthropic attributes to how often a safety classifier interrupts a task rather than any change in raw ability. Fable is also missing from several rows on OpenAI’s table entirely — including two life-sciences evaluations, GeneBench Pro and MedChemBench, where OpenAI reports leading scores — and from OSWorld 2.0, a computer-use benchmark where Anthropic reports its own figure under different task and grading conditions, making a direct comparison impossible.

So: whose table do you trust? A vendor grading its own rival, or an independent lab whose top-line index disagrees with that vendor on both counts.

Different Models, Different Jobs

The more useful read might be that Astra and Fable aren’t really built for the same job anymore.

Astra is tuned for driving software the way a person would. It scores 92.7% on ScreenSpot-Pro, a benchmark for locating and clicking the right element on a screen, up sharply from 76.9% for its predecessor. On BenchCAD, where a model has to reconstruct a 3D part from multiple rendered views and generate working CAD code, Astra hits 95.9% against 84.3% for Fable 5.1 — the widest gap either company has published. OpenAI’s pitch, echoed in Brockman’s framing, is a model that finishes slides, spreadsheets, and technical drawings rather than describing how to make them.

Fable leans the other way, toward raw reasoning. It scores 65.0% on Humanity’s Last Exam with tools, against 57.2% for Astra — one of the few academic benchmarks where the Claude model comes out clearly ahead — and it holds the Artificial Analysis Intelligence Index lead by five points.

The Price Tag Looks Identical. It Isn’t.

Both models charge the same headline rate: $10 per million input tokens, $50 per million output tokens. That symmetry breaks down as soon as a workload starts reusing context. Anthropic cut Fable 5.1’s cache-read price 75% from Fable 5, down to $0.25 per million tokens. Astra’s cached input runs $1.00 per million — four times as much for functionally the same operation. For a single question, the difference is trivial. For an agent that rereads the same files or the same long conversation history hundreds of times over a session, the gap compounds fast, and it’s exactly the kind of cost that doesn’t show up until the invoice arrives.

Yet that same rate card inverts once you look at what a task actually costs to finish rather than what a token costs to read. Artificial Analysis measured Astra completing an average Intelligence Index task for about $1.67, against roughly $3.76 for Fable 5.1 — more than double — because Astra tends to spend far fewer tokens getting to an answer. The cheaper rate card belongs to Fable. The cheaper bill, on this measure, belongs to Astra.

So, Which One Do You Use?

Probably not the one with the higher number on the table you happened to read first.

If the job is operating real software — clicking through a CRM, producing a finished slide deck, generating CAD output, or running agentic science and math workloads — Astra’s numbers are the stronger and more consistently published case, and OpenAI has also pushed it to a “Critical” cybersecurity capability rating under its own Preparedness Framework, the first model the company has classified that way.

If the job is open-ended reasoning, research-style questions without an obvious step-by-step workflow, or long sessions that hammer the same cached context over and over, Fable 5.1’s independent index lead and its cheaper cache rate make a real case.

The AGI framing was always going to generate headlines. The more durable story this week is quieter: two labs measuring two different things, publishing two different scoreboards, and neither one is wrong so much as incomplete on its own.

Article Categories:
Technology

Leave a Reply

Your email address will not be published. Required fields are marked *