AI models

GPT-6 Astra vs Claude Fable 5.1: which one I use for what, and how I prompt both

Two new flagship models landed in the same week of September 2026. Anthropic shipped Claude Fable 5.1 on 1 September. OpenAI shipped GPT-6 Astra on 3 September. Both cost the same on the API, both hold about a million tokens of context, and both labs say theirs is the smartest model in the world. That is not a useful sentence for a business owner.

I run a small AI automation agency. We build support agents, invoice readers and lead follow-up systems for ecommerce brands, real estate offices, clinics and trades businesses. We use both Claude and OpenAI inside those builds, so I have a selfish reason to know which one to reach for. I spent three weeks running the same jobs through both. I also read what other builders were finding on YouTube, X, Substack and the comparison blogs, because my sample is small and theirs adds up.

This post covers what each model is better at, what people disagree about, and how to prompt each one so you stop paying for wasted tokens. The prompts near the end are the ones I actually use, not made up for the post.

GPT-6 Astra and Claude Fable 5.1 side by side, the two flagship models released in the same week of September 2026

The short answer

Pricing for GPT-6 Astra and Claude Fable 5.1: both $10 input and $50 output per million tokens, with Fable 5.1 cached reads at $0.25 against Astra at $1.00

There is no overall winner, and anyone who tells you there is has picked a benchmark that suits them. Astra is the better tool when the job is to drive software, do maths, or finish lots of short tasks cheaply. Fable 5.1 is the better tool when the job is writing a person will read, reading long documents, or running an agent for hours on a codebase. This is my pick for each kind of work.

If your job isI reach forWhy
Emails, proposals, blog posts, anything a human readsFable 5.1Better openings, fewer stock phrases, and it argues back when the brief is wrong
Reading a 300-page contract or a pile of PDFsFable 5.1No surcharge on big requests, and cached reads cost a quarter of Astra's
Filling forms, clicking through a web app, computer useAstraBuilt for it. It wins the computer use benchmarks by a clear margin
Maths, statistics, anything with a right answerAstraIts lead on FrontierMath Tier 4 is the biggest gap in either direction
Quick one-shot tasks by the hundredAstraUses about a third of the output tokens for the same score, so it is cheaper per task
A coding agent running for hours on one repoFable 5.1Holds the goal across a long session and does not pretend to be done
Terminal-heavy automation and scriptsAstraLeads Terminal-Bench on every harness that scored both
Front-end design and unusual UIFable 5.1Even Astra fans concede design to Fable
Dashboards, data visuals, polished SaaS-style pagesAstraWon 35 of 50 blind website builds, mostly on visual polish
Explaining to a client why the AI decided somethingFable 5.1Its reasoning comes back readable. Astra's does not

Price is a tie on paper. Both list at $10 per million input tokens and $50 per million output. The gaps are underneath: Fable 5.1 reads cached context at $0.25 per million against Astra's $1.00, and Astra doubles its input price on requests above 272,000 tokens while Anthropic does not. But Astra writes far less to get to the same answer, so on measured cost per task it usually comes out cheaper. Which of those matters depends on the shape of your workload, not on which logo you like.

Where Astra wins

Illustration of GPT-6 Astra's five effort levels, Low to Max, with Medium marked as the place to start

Astra was built to act, not to chat. OpenAI's own line is that anything you can do on a computer, Astra can do for you. In my testing that is close to true for the boring middle of a workflow: log into a portal, pull a report, paste the numbers into a sheet, send the summary. It clicks, reads the screen, and keeps going through more steps before it stops to ask. For the clinic and real estate clients we serve, that is the part of the job nobody wants.

The maths gap is real and it is the widest gap in the whole comparison. On FrontierMath Tier 4, Astra scores 97.6% against Fable 5.1's 87.8%, according to the DataCamp comparison. If your work is pricing models, statistics or anything where there is one correct number, Astra is the safer default.

The cost story surprised me. Artificial Analysis found that at max effort Astra uses about 27,000 output tokens per task where Fable 5.1 uses around 78,000, for the same score on their Intelligence Index. Their benchmark write-up puts Astra at roughly 40% of Fable's cost per task. So the same rate card can produce a very different bill. If you run hundreds of short jobs a day, that difference is your margin.

Astra also has a fair claim on speed of use. The most common complaint I saw in YouTube comments was not about Astra at all. It was people on the $200 Claude plan hitting the five hour usage limit mid-task and switching to Astra because they could keep working. That is a product decision by Anthropic, not a model weakness, but it matters if you are the one waiting.

One warning. Astra is more likely than Fable to stop and ask a clarifying question, and it has been reported to drift on long unattended runs, touching things you never mentioned. Ruben Hassid's Substack guide makes the same point. I would not leave it running on a client's live system overnight without a tight scope.

Where Fable 5.1 wins

Illustration of Claude Fable 5.1's four effort levels, Low, Medium, High and XHigh, with what each one is for

Writing is the clearest one. When I gave both models the same client brief for a cold email, Fable's opening was the one I would have sent. Improvado ran a similar four-task test and reached the same conclusion: neither model invented a source, both caught the three errors planted in a spreadsheet, and Fable wrote the better opening, although it also wrote about 35% more text. Their full write-up is worth reading if you write for a living. I will say the opposite view exists. One YouTube commenter insisted Claude's short punchy phrasing is itself an AI tell and that Astra sounds more human. I disagree, but I have heard it more than once.

Long documents are the second. Fable 5.1 holds a million tokens with no price change across the window. Astra's window is slightly bigger at 1.05 million, but above 272,000 input tokens it doubles the input price and charges 1.5 times on output for the whole request. If you feed it a full lease, a year of invoices, or a repository, that surcharge lands every single time.

Long agent runs are the third, and this is where the pricing and the behaviour line up. An agent loop rereads the same system prompt, repo map and rules on every turn. Anthropic cut cache read pricing by 75% to $0.25 per million for exactly this reason. Fable also holds the goal better across a long session. Theo from t3.gg, quoted in the MindStudio benchmark analysis, called Astra's raw coding roughly on par with Fable 5 but said front-end polish and "would I actually merge this" confidence still sat with Fable 5.1.

The independent scoreboards mostly lean Fable, which is not what OpenAI's launch table shows. Artificial Analysis had Fable 5.1 at 66 against Astra's 61 on its Intelligence Index at launch, and 70 to 67 on its Coding Agent Index. A later re-run tied them at 53 each after the index was revised. On Humanity's Last Exam with tools, Fable scored 65.0% to Astra's 57.2%. Every table showing Astra ahead almost everywhere came from OpenAI's own launch post.

Design is the one even Astra's loudest fans hand over. The most shared post on X after launch said it was over for Anthropic on speed, cost and debugging, then added that Fable 5.1 is still the best design model in the world. MindStudio's 50-site blind test found Astra won on polish and SaaS-style layouts, while Fable held the edge on audio interfaces and unusual, experimental designs. For Shopify storefronts and landing pages, I still start in Fable and fix in Astra, not the other way round.

What other builders found

I did not want this to be one person's opinion, so here is what the people who tested hardest came back with. They do not agree, which is the point.

WhoWhereWhat they testedTheir call
Nate HerkYouTube, 15 real use casesResearch, sales copy, taxes, a 3D game, browser tasks, a YouTube strategyAstra 10, Fable 5. Fable took research, sales letter copy, the HTML explainer and the second-brain visual. Astra took the browser jobs, taxes and most of the builds
Zapier-sponsored channelYouTube, client work testBoth models on the same agency deliverablesLeaned Astra. Top comment was about Claude's five hour usage cap, not the model
Evidence-first channelYouTube, benchmark reviewIndependent scores against OpenAI's tableMixed. One commenter said Fable declared a big task done after eight hours and had faked three parts
Website build channelYouTube, $50k site promptIdentical premium site briefViewers preferred Fable's functionality and Astra's design, and wanted both fixed by the other
Zach, engineer turned AI builderX threadDay-to-day agent workFable 5.1 if forced to choose. Called Astra frustrating and amazing in the same breath
TekniumX threadCache pricingNervous about Astra's cache reads being 4 to 8 times dearer for agent loops
Ruben HassidSubstackRead both labs' prompting docsUse Astra at Medium effort. Give Fable full control and tell it you are not watching
Artificial AnalysisIndependent benchmarkIntelligence and Coding Agent indicesTie on score. Astra at about 40% of the cost per task
DataCampComparisonEvery published benchmark row plus pricingAstra for computer use, maths and cheap tasks. Fable for reasoning depth and big requests
MindStudio50-site blind buildOne-shot website generationAstra 35 of 50 on human preference. Function was a tie at 48 and 47

A note on Reddit and Quora. I looked for the big threads on r/ClaudeAI, r/OpenAI and Quora before writing this. Search did not return them cleanly, and I am not going to quote posts I could not open. If you have a thread that changed your mind, send it and I will add it here.

The pattern across all of it is the same. Astra wins the moment you ask it to operate something. Fable wins the moment you ask it to think about something for a long time, or to write something you will put your name on. The people who declare a single winner have usually only tested one of those two kinds of work.

How to prompt GPT-6 Astra like a pro

First, find it. In regular ChatGPT the model is called GPT-6 Pro and only Pro, Business and Enterprise plans see it there. On the $20 Plus plan it lives in ChatGPT Work and in Codex, not in the normal chat picker. Engadget's rollout guide has the plan by plan detail. Once you have it, there is an effort slider next to the model name. That slider decides how hard it thinks and how much you pay.

Use Medium effort by default. Almost everyone who has read the docs lands here. Medium is the best mix of cost, speed and intelligence, and High only earns its price on hard research or a build with a lot of verification. Frank Andrade's complete guide ran everything at High, so if you copy his prompts, drop the effort first and see if the result still holds.

Give it a finish line. "Help me analyse this" leaves Astra to decide when it is finished, and it will decide early. "Read these three invoices, flag any line that does not match the purchase order, and give me a table with invoice number, line, amount and reason" tells it what done looks like. You do not need to add "think step by step" any more. Astra plans on its own. What it needs from you is the constraint and the success test.

Decide whether you want it to ask questions. Astra asks for clarification more than Fable does. Sometimes that saves a wasted run. Sometimes it interrupts a job you expected it to finish while you were on a call. If you want it to push through, say so in the prompt: "Do not ask me questions. Make a reasonable assumption, state it at the top, and continue." If the task is risky, say the opposite.

Steer it while it works. You can send a follow-up message mid-task and Astra will change direction without starting over. This is the single habit that saved me the most money. Instead of writing a perfect prompt, I write a good one, watch the first two steps, and correct.

For writing, tell it a standard. Ruben Hassid's favourite trick is the instruction "Use ASD-STE100", which is the simplified technical English spec used in aviation manuals. It strips out the flourishes and makes Astra write short, plain sentences. I use it for anything a client will read.

Scope the computer use. When Astra is driving a browser or an app, name the folders, sites and actions it may touch, and name the ones it may not. The drift people complain about comes from a wide brief, not a bad model.

How to prompt Claude Fable 5.1 like a pro

Fable 5.1 has adaptive thinking that is always on, with an effort setting on top. Alex McFarland's breakdown of the levels is the clearest I found. Low is for quick work on material you supply, like tagging, extraction and formatting. Medium is for documents, spreadsheets and bounded analysis. High is for research, building systems and anything that needs verification. XHigh is for hard debugging and architecture calls. One catch: at Low, Fable is less likely to go and search. If the task depends on current information, raise the effort or tell it in words to verify.

Tell it you are not watching. On long tasks Fable sometimes describes the next step instead of doing it. Anthropic's own guidance, summarised in the Promptessor guide, is to define the request itself as the deliverable and tell it to keep going on reversible work. The line I paste at the top of every unattended run is: "I am not watching and cannot reply. Finish the whole task autonomously without asking. If something is ambiguous, make the safer choice and note it at the end."

Say what done means. "Fix the bug" is an analysis request. "Fix the bug, update the tests, run the suite, fix any failures you caused, and stop only when the requested behaviour works" is a build request. The DDS Hub guide has a dozen examples of this shape. The gap between those two prompts is the gap between a report and a working system.

Put standing context first, and keep it stable. Fable's cheap cache reads only help if the same prefix is reused turn after turn. So the system rules, the repo map, the client's tone guide and the constraints go at the top and never change. The question goes last. Venice's prompt tips explain the mechanics. Rewriting an earlier system prompt mid-conversation can also invalidate its preserved thinking, so treat the history as append-only.

Ask for progress updates. Fable 5.1 says less between tool calls than Fable 5 did. If you are watching an agent run, ask for a one-line update after each phase. If you are not, ask for a summary table at the end: status, files touched, next action, open questions. That table is what you review, not the transcript.

Nudge parallel work. In long loops Fable will sometimes fetch one thing per turn when it could fetch five at once. A single sentence fixes it: "Before you start, list what you need. Request anything independent in parallel."

Let it disagree with you. This is the Fable habit I value most. Tom's Guide tested a single prompt that asked it to interrogate the premise before doing the work, and got a better result than from the work alone. On a client brief I now ask: "Before you write anything, tell me what is wrong with this brief." It usually finds something.

Tune the writing. Fable 5.1 writes denser prose with less formatting than Fable 5. Old prompts full of "no bullet points, no headings" instructions now overshoot and make it terse. Say what you want positively: "short paragraphs, plain words, one idea per sentence", and it lands.

Prompts I actually use

A vague prompt rewritten for GPT-6 Astra and Claude Fable 5.1 with scope, a finish line and an ask or assume rule added

These are lifted from our own builds with the client names removed. Swap the bracketed parts and keep the shape.

Astra, browser task, Medium effort:

Task: Log into [portal], download last month's [report name] as CSV, and paste the totals for [three named columns] into the attached sheet on the tab called Monthly.

You may: open [portal domain], read pages, download the report, edit the attached sheet.
You may not: change any setting in the portal, send anything, or touch any other tab.

Do not ask me questions. If the report is missing, stop and tell me which page you were on.
Done means: the three totals are in the sheet and you show me a screenshot of the row.

Astra, client-facing writing:

Use ASD-STE100.
Write a 120-word email to [name] at [company] confirming [decision]. Include the [date] and the [amount]. No greeting beyond their first name. End with one question.

Fable 5.1, long document review, High effort:

<rules>
You are reviewing a [type of document] for a [type of business]. Quote exactly when you cite a clause. Paraphrase everywhere else and say which you are doing.
</rules>

<document>
[paste the full document here]
</document>

Before you review it, tell me what is wrong with this request. Then list every clause that [the thing I am worried about], with the clause number, the exact words, and one plain sentence on why it matters.

Fable 5.1, unattended coding agent:

I am not watching and cannot reply. Finish the whole task autonomously without asking. If something is ambiguous, make the safer choice and note it at the end.

Goal: [what the feature does when it works].
Scope: only files under [folder]. Do not modify existing API behaviour unless a test proves it is broken.
Done means: the feature works, tests for it exist and pass, the full suite passes, and nothing you did not touch is failing.

Before you start, list what you need and request anything independent in parallel.
At the end give me: status (blocked, ready or done), a table of files touched, next action in one sentence, and open questions as a list.

The common thread is the same in all four. Say what the model may touch, say what done looks like, and say whether it should ask or assume. Every expensive run I have paid for was missing one of those three lines.

What this means if you run a small business

You probably do not need to pick. For a solo owner or a small team, the honest answer is that either subscription will do most of the work you throw at it, and the second one earns its keep only when you have a specific job the first one does badly.

If most of your AI use is writing, reading contracts, summarising calls and thinking through decisions, Claude Pro with Fable 5.1 is the better everyday assistant. Watch the usage cap if you run long jobs. If most of your AI use is operating software, filling forms, working through spreadsheets and doing anything with numbers, ChatGPT with Astra is the better fit, and you will hit fewer walls in a long session.

If you are paying for API usage inside an automation, the model choice is an arithmetic problem, not a taste problem. Count how much of each request is the same context reused turn after turn. If it is most of it, Fable's cache pricing wins. If each request is short and different, Astra's lower token use wins. We run this sum on every build before choosing, and it has gone both ways.

The last thing I would say is that both models are already good enough that the prompt is the bottleneck, not the model. A vague brief to the best model in the world still gets you a vague result. Spend the time on the three lines that matter, and either of them will look like the winner.

Your build starts with the arithmetic.

Bring a week of real examples to the free audit. We come back with the three worth doing first, what each is worth, and a fixed price. If a cheaper tool already covers it, we say so.

Book the free audit →

Sources