OpenAI’s GPT-6 Astra Is a Genius That Won’t Stop Talking to Itself

OpenAI’s GPT-6 Astra Is a Genius That Won’t Stop Talking to Itself

GPT-6 Astra arrived on September 3 with numbers that would’ve sounded like science fiction two years ago: 99.9% on ARC-AGI-3, a perfect ExploitBench score, a 1.05 million token context window. OpenAI trained it on more than 100,000 GPUs at the Stargate facility in Texas and called it the most capable model ever shipped. All of that checks out. Two weeks of living with Astra in Codex and ChatGPT Pro has convinced me the benchmarks are the least interesting part of this launch, though. The interesting part is what the model does to your quota, and what that says about the $1.2 trillion funding round now circulating around the company.

A rollout messy enough that Altman apologized

The launch was bumpy. Day-one access went to a short list of organizations, including the Daybreak cybersecurity program, while paying Plus subscribers watched Pro and Enterprise users get in first. Enterprise admins had to manually switch Astra on because it shipped off by default. Sam Altman called the rollout messy and apologized for it, which is more accountability than most companies manage, and broader API and subscriber access followed within days.

Pricing tells its own story. The API costs $10 per million input tokens and $50 per million output tokens, roughly 2.5 times what GPT-5.6 Sol charged. Pro, Business, and Enterprise subscribers get a faster Astra Pro variant, while everyone else draws from existing allowances. Astra also crossed OpenAI’s “Critical” threshold for cybersecurity capability, so the public version now flatly refuses advanced exploit work. We’ve already broken down Astra’s Critical cybersecurity rating, and the scrutiny is warranted given state-linked hacking campaigns aren’t slowing down.

The polling problem nobody priced in

The uncomfortable part showed up in the quota math. Astra burns allowance two to four times faster than Sol on comparable work, so I ran a long-horizon Codex job and pulled the token logs apart.

Read Also:  Meta's Muse Is a Landmark AI Agent, and the Trust Math Doesn't Work

Astra, it turns out, is a restless supervisor. When it orchestrates worker processes, it polls them every 30 seconds, and most of those check-ins return nothing useful. In my logs, idle polling accounted for roughly 68% of the parent conversation’s input tokens. The model spends most of its energy asking its own subordinates whether there’s an update and getting silence back. In the worst run I measured, a five-hour Pro allowance evaporated in 33 minutes.

You can tune around some of it. The latest Codex CLI quietly added Astra support, and stretching polling timeouts in your config claws back a chunk of the waste. OpenAI also introduced new reasoning.effort levels running from low to max, and the prompt packs circulating for this model read less like writing advice and more like parameter documentation. Prompt engineering used to mean words. Now it means knobs.

The capability trade is real. On traditional reasoning and knowledge evals, Astra improves modestly over Sol. Where it pulls away is long-horizon work: sustained multi-hour tasks, large codebases, tool orchestration. Benchmarks like Terminal-Bench and SRE-Bench show the jumps. If your workload is agentic, the premium may pencil out. If you’re using Astra as a chat model, you’re paying 2.5 times more for roughly the same answer, and some teams have paused new Pro subscriptions entirely while they rework their budgets.

A trillion-dollar pitch meeting this friction

Meanwhile the money conversation is accelerating. Annualized revenue has pushed past $40 billion, with roughly a 20% jump after GPT-5.6 Sol shipped, and OpenAI is reportedly shopping a private round at a valuation near $1.2 trillion. An IPO looks unlikely before 2027.

Read Also:  The GTA VI Leak Is a Subscription Service, and Take-Two Is Reading It Wrong

Hold the launch story next to the funding story, because they don’t quite fit. Investors are being sold saturated benchmarks and an AGI-era narrative. Users are getting admin opt-ins, phased access, and subscription allowances a polling loop can drain before lunch. Both are true, and the gap between them is where OpenAI’s execution risk lives. The company is selling something it hasn’t learned to meter yet.

The safety numbers deserve the same skepticism. OpenAI claims 0% scope violations on an internal alignment eval versus 48% for the prior model, which sounds great until you remember the vendor wrote and graded the test. The public refusals on exploit tasks are observable behavior, and those do check out.

My verdict after two weeks is that Astra is legitimately the best agentic model you can use right now, and the quota carnage is mostly a defaults problem OpenAI should have fixed before shipping. Watch two things this month. A patch that tames the polling flips the value story almost overnight. And if that $1.2 trillion round closes at full price, investors are betting OpenAI works out the economics of selling genius by the token before the market does the math itself. That’s the bet. I’m not sure I’d take it at that number.

With ten years in the Industry, I write to provide our readers with the best material and great experience.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *