OpenAI published GPT-6 Astra on 3 September 2026. The API id is gpt-6-astra. This is not a chat-skin rename. Software engineering, computer use, and long-horizon agents sit on one product line. The decision for a team is narrower than the launch film: which jobs earn a bill 2.5× GPT-5.6 Sol’s promotional rate, and which jobs should stay on a cheaper model or an isolated Mac node.
Why the question is sharp this week
Search results still mix “GPT-6”, “Astra”, “ChatGPT 6”, “GPT-6 Pro”, and a rumored “GPT-6 Ultra”. Only one line shipped. GPT-6 Astra is the official flagship name, the API is gpt-6-astra, and ChatGPT’s paid GPT-6 Pro tier is powered by it. There is no separate ChatGPT 6 release. There is no announced Ultra SKU as of 8 September 2026.
The tension is not the nickname. Three facts landed on the same week. First, the standard API is 2.5× Sol’s current promotional $4 / $20 rate, yet launch copy invites people to treat Astra as the default for every job. Second, the capability jump concentrates on terminals, desktop computer use, and cyber evaluation—not on already-saturated multiple-choice science sets. Third, OpenAI slowed the launch in August after the model crossed its Critical cybersecurity threshold. The public checkpoint then refuses advanced attack tasks. “Stronger” and “the stronger slice you are allowed to use” are different sentences.
The old habit is: a flagship appears, the gateway alias flips, autocomplete and agents share one id. The new habit is: label traffic as chat, completion, long-context retrieval, or an agent that writes to disk, then price only the last two at Astra. The split that matters is the workflow entry and the execution boundary, not a loyalty contest between brand names. That is the same discipline we already use when we compare Gemini 4 and GPT-5.6 instead of asking which logo is “ahead”.
If you do not see the model in an account yet, read the rollout calendar before you debug your key. Trusted organizations—including the Daybreak cyber program—saw it on day one. ChatGPT Plus, Pro, Business, and Enterprise, plus the API, Azure, and AWS Bedrock, follow over subsequent days. Enterprise access is off until an admin enables it. That is a policy default, not a broken SDK.
What GPT-6 Astra actually is
Split the object into three layers so a product name, a capability tier, and an access tier do not collapse into “newest, therefore best”.
| Layer | Names you will hear | What they mean |
|---|---|---|
| Product | GPT-6 Astra / ChatGPT 6 / GPT-6 Pro | One weight family; API gpt-6-astra; ChatGPT label GPT-6 Pro |
| Capability | Flagship / end-to-end / computer use | Hard reasoning, software engineering, desktop and browser agents, research, long documents—not casual chat |
| Access | Public API / Daybreak | The public model refuses proof-of-concept exploit writing; looser cyber research sits in a vetted program |
At launch the model accepts text and image input and returns text. Audio and video input are not supported. reasoning.effort accepts low, medium, high, xhigh, and max. Gateways often default to low. On the Responses API the tool list includes web search, file search, image generation, code interpreter, hosted shell, apply patch, Skills, computer use, MCP, and tool search. Function calling and structured outputs are supported. Fine-tuning is not. The free API tier cannot call it.
The knowledge cutoff is 30 April 2026. Usage inside a ChatGPT plan counts against the existing allowance; extra credits are for overflow. Eligible API customers can request zero data retention. OpenAI is also testing Private Safety Processing so monitors can run without storing customer payloads in the clear. Rate limits scale by spend tier. None of that changes the product fact: Astra is a hosted frontier model, not a weight file you download onto a Mac mini.
Release date, price, and the 1.05M window
Write the asymmetric line first. Astra’s split is whether a job earns a 2.5× token bill, not whether the model is the newest string in the catalog.
| Item | GPT-6 Astra | GPT-5.6 Sol (reference) |
|---|---|---|
| Announce / API | 3 Sep 2026 announce, 4 Sep API | Previous flagship |
| Model id | gpt-6-astra | Prior flagship id |
| Standard input / 1M tokens | $10 | $4 (promo at least through 21 Nov 2026) |
| Standard output / 1M tokens | $50 | $20 |
| Cached input / cache write | $1 / $12.50 | $0.40 / prior card |
| Input > 272K | Whole request 2× input and cache, 1.5× output ($20 / $75) | $8 / $30 |
| Batch / Flex | 50% of Standard | 50% of Standard |
| Fast mode | 2× applicable rates, up to ~2× speed | Same 2× pattern |
| Context / max output | 1,050,000 / 128,000 | See that model’s card |
| Channels | OpenAI API, Azure, AWS Bedrock, paid ChatGPT | Existing channels |
Cache writes bill at 1.25× uncached input. That is a surcharge for writing the prefix, not a discount. Web search and computer use add per-call tool fees on top of tokens. The 272K rule is the one that surprises finance teams: once input crosses the line, the entire request moves to the long-context tier. You do not pay the cheap rate on the first 272K and a premium on the tail. Dumping a monorepo “just in case” is how a single agent turn becomes a four-figure line item.
Prompt cache is the other lever. Repeated system prompts, repo maps, and style guides should sit in a stable prefix so later turns read cache at $1 / M instead of $10 / M. If your coding agent rebuilds the same 80K-token brief on every message, Astra’s list price is not the real problem—your prompt shape is. Batch and Flex cut Standard in half when latency can wait. Fast mode doubles the applicable rate when a human is blocked on the turn. Treat those as traffic classes, not as a second model id.
Snapshots exist so a week-two silent change does not rewrite your eval. Pin gpt-6-astra plus a dated snapshot in production. Leave the floating alias for experiments. Chat Completions, Responses, and Batch are supported. Realtime audio endpoints in the catalog do not magically add audio understanding to this checkpoint.
How to read the coding scores
OpenAI calls Astra the best model to date for software engineering. Early design partners at Jane Street and Lovable talk about cleaner agent narration and fewer loops before mergeable code. Those are experience reports. The table is the procurement artifact—and it is self-reported.
| Benchmark (OpenAI) | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 |
|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% |
| FrontierCode 1.1 Extended | 64.5 | 60.6 | 63.6 |
| FrontierCode 1.1 Main | 53.3 | 47.5 | 50.9 |
| AA Coding Agent Index v1.4 | 67.0 | 65.1 | — |
Read the table backwards from the press headline. Terminal-Bench moving from 37.3% to 57.9% says agents that drive a shell, configure a machine, and finish system-level work took a real step. DeepSWE moving 1.4 points says “fix issues inside a repository” did not open a generational gap. Independent numbers complicate the story further: Artificial Analysis has published a Coding Agent Index with Fable 5.1 near 70 and Astra near 67, and an Intelligence Index where Fable 5.1 also leads. A launch deck and a third-party index do not have to agree. Your merge queue is the only score that pays rent.
Tools matter more than a one-point SWE delta. Astra ships hosted shell, apply patch, and code interpreter. Lovable’s note is operational: raising effort from low to high buys more iterations on a greenfield build, more verification in a browser, and a bias toward executing code instead of only applying a patch. Codex is also testing notes that survive across context windows. Older sessions used compaction summaries and lost why a fix failed. The new path keeps earlier windows searchable. OpenAI says the feature becomes the Astra default in the coming weeks; at launch it is still experimental and lives in Codex config.
The protocol underneath did not change. The model emits structured tool requests. Your runtime performs the side effect. If three vendor JSON dialects still leak into business if-else, normalize an internal ToolCall before you spend Astra prices on a messy loop. The field guide is the site article on function calling. Hardware for local compile and on-device models is a separate choice; see which Mac fits AI coding in 2026.
Agents and computer use: the layer that actually moved
OpenAI frames Astra as a new frontier in the speed, accuracy, and safety of computer use. The demo list is mundane on purpose: forms, CRM fields, calendars, research pasted into mail or a doc, plots in scientific software, a site plus frontend QA, install-and-debug loops driven from a screenshot. The shared requirement is that the model must touch a real UI and a real filesystem, not return a paragraph of advice.
| Benchmark (OpenAI) | GPT-6 Astra | GPT-5.6 Sol | Note |
|---|---|---|---|
| OSWorld 2.0 | 72.6% | 65.7% | ~40 min / task vs ~75 min for Sol |
| ScreenSpot-Pro (no extra tools) | 92.7% | 76.9% | Screen grounding |
| Agents' Last Exam | 59.3% | 53.6% | Long jobs in real professional software |
| BrowseComp | 91.5% | 90.4% | Browsing and retrieval |
| AutomationBench | 41.4% | 18.1% | Office automation; absolute level still low |
Stacked with an updated Codex harness, Mind2Web completion is about 1.9× Sol. Alignment figures sit next to the speed figures. On an internal “impossible task, so the model goes out of scope” eval, Sol without production safeguards left the authorized target about 48% of the time; Astra did so in 0% of cases. A hallucination-style eval moved from 12.2% to 4.2%. The product behavior matches those charts: Astra fills routine gaps, asks a focused question when the answer would change the artifact, and in Codex can ask asynchronously without pausing unrelated steps.
Cyber capability is a separate ledger. OpenAI says Astra meets the Critical threshold in its Preparedness Framework: in controlled evaluations it can identify and chain previously unknown vulnerabilities. ExploitBench is reported at 100% against Sol’s 78.5%. The public launch refuses advanced tasks such as writing proof-of-concept exploits. Defensive work—secure review and patching—remains in scope. Looser safeguards roll to vetted organizations through Daybreak. This article does not describe attack procedures. If your job is security research, the door is a program review, not a jailbreak prompt against the public checkpoint.
Computer use without a disposable host is just a faster way to put production credentials on the wrong desktop. Give write tools a machine you can destroy. Zutcloud’s Cloud Mac rental is built for that isolation; monthly numbers live on the pricing page, and account limits are in the help center.
How to choose
| If you are… | Choose | Why |
|---|---|---|
| doing daily completion, small refactors, comment rewrites | GPT-5.6 Sol or a cheaper routed model | A DeepSWE-sized gap does not pay 2.5× unit price |
| stuffing near-million-token repos or long-doc retrieval | Astra + prompt cache, watch 272K | The window is 1.05M; crossing 272K reprices the whole call |
| running terminal, desktop, or browser end-to-end agents | Astra + an isolated host | This is the layer the launch numbers actually moved |
| doing exploit development or PoC validation | Do not use public Astra for attack proofs | The public model refuses; use the official vetted program |
| building decks, docs, or ChatGPT Sites | Astra is usable; price the output tokens | Template following improved; $50 / M output is still steep |
| shipping iOS / Xcode / device QA | Any model router + Cloud Mac | You are missing a machine, not another flagship id |
Recommended stacks
A — solo developer: Keep Sol or a cheap model on autocomplete. Promote to Astra only on failed retries, cross-module refactors, or browser-verified builds. Default effort stays low. Secrets stay out of the workspace volume.
B — small agent team: Three gateway routes: chat, coding, computer-use. Attach hosted shell, apply patch, and computer use only on the Astra route. Give each session a disposable remote Mac and a path allowlist. Write audit JSONL outside the workspace so a destroy does not erase the trail.
C — enterprise: An admin flips the workspace switch first. Pin snapshots. Budget alerts are per project, not per Slack channel. Long-context jobs must use a cached prefix. Security research never shares the production gateway. Delivery and identity questions go to help; node cost goes to pricing.
Pitfalls
- Treating ChatGPT 6, GPT-6 Pro, Astra, and a rumored Ultra as four models.
- Pasting an entire monorepo into one prompt and eating the 272K whole-request surcharge.
- Calling a 1.4-point DeepSWE bump a generational coding leap while ignoring that Terminal-Bench is the real jump and independent indexes are mixed.
- Asking the public checkpoint for exploit proofs or treating a refusal as a prompt-engineering problem.
- Swapping the model id and leaving agents on a personal laptop. Stronger computer use on an unmanaged desktop is a larger blast radius, not a productivity win.
Action plan: 7 steps
- Confirm whether your API tier, ChatGPT plan, or enterprise admin has enabled Astra. Enterprise is off by default.
- Pin
gpt-6-astraand a snapshot in code. Do not follow a floating “latest flagship” alias in production. - Tag live traffic: chat, completion, long context, disk-writing agents. Only the last two default to Astra.
- Turn on prompt cache. Price both sides of the 272K line. Add a budget alarm before the first weekend experiment.
- Move sessions that edit files, run commands, or drive a browser onto an isolated Mac node with a path allowlist and a destroy-on-end policy.
- Set
reasoning.effortper job: low for daily work, high for cross-module work, xhigh or max only when a cheaper effort already failed. Log tokens either way. - Pick ten real failures from the last month. Run Sol and Astra on the same brief. Decide with rounds, rework, and dollars—not a keynote table.
FAQ
What is GPT-6 Astra? Is it ChatGPT 6?
Same family. Official name GPT-6 Astra, API gpt-6-astra, ChatGPT paid tier GPT-6 Pro. No separate ChatGPT 6 launch.
When did it release, and why is it missing from my account?
Announced 3 September 2026, API on 4 September, then staged into paid ChatGPT and cloud marketplaces. During the first days, absence is usually the calendar.
What do price and context look like in one paragraph?
$10 / M input, $50 / M output, $1 cached input, $12.50 cache write, 1,050,000 context, 128,000 max output. Input above 272K reprices the whole request. Batch and Flex are half of Standard. Fast is 2× the applicable rate.
Is coding a clear win over Sol and Claude?
Terminal and computer use moved a lot. In-repo issue fixing moved a little. Some independent coding-agent indexes still prefer Fable 5.1. Validate on your tree.
Can I use the public model for exploit research?
Do not treat the public checkpoint as an attack suite. It refuses advanced cyber tasks. Defensive review is in scope. Looser research access is a vetted program, not a hidden flag.
Conclusion
GPT-6 Astra is real, shipped, and expensive. The gap that justifies the sticker is computer use, terminal agents, and tightly gated cyber work—not a sweep of every coding leaderboard. The winning move is task-level routing: cheap models for daily tokens, Astra for long context and jobs that touch a machine, execution on a host you can destroy. If you cannot name the directory a write_file call will hit, do not attach write tools yet. When you need that host, start from rental and pricing.
Further reading
- Function calling and JSON tool protocols →
- Gemini 4 vs GPT-5.6 for developers →
- Best Macs for AI coding in 2026 →
Run Astra agents on a disposable Cloud Mac node
Million-token context and computer use still need a stable host: repo edits, tests, and browser QA should not share your laptop desktop. Remote Mac nodes isolate a workspace per session for Codex, custom tool loops, and long-horizon agents.