OpenAI shipped its GPT-5.6 model family and a new agent called ChatGPT Work on Thursday, ending a two-week standoff with Washington that had kept the company's most capable system locked behind a government-vetted preview. The release drops OpenAI into direct combat with Anthropic's Claude Cowork, which reached users in January and has spent the months since converting that head start into enterprise contracts.
ChatGPT Work went live on web and mobile for Pro and Enterprise subscribers, along with Edu accounts. Plus and Business users receive access within days. A rebuilt ChatGPT desktop application is available globally on Windows and Mac, across every plan including the free tier.
The timing was not accidental. Both companies filed confidential IPO paperwork in June, and both are now selling the same pitch to the same buyers.
An Agent Built to Hand Back Finished Files
ChatGPT Work is a mode inside ChatGPT, not a standalone product. Users connect the systems where their work already lives through plugins spanning Slack, Microsoft Teams, Google Drive and SharePoint, plus email, calendars, CRM platforms and project trackers. ChatGPT decides on its own when to reach for a plugin. A user can also point it at a specific application by typing "@" followed by the app name.
What comes back is the part OpenAI wants people to notice. Instead of a chat reply, the agent returns finished spreadsheets, slide decks, documents and shareable web applications. It breaks a goal into steps, works through them without supervision, and stays on a project for hours at a stretch.
Codex sits underneath all of it. OpenAI says more than five million people use its coding tool every week, with over a million now applying it to work that has nothing to do with software. ChatGPT Work takes that engine and hands it to people who have never opened a terminal.
"You can apply the model's ability to code to solve problems across every industry," said Ty Geri, product manager for ChatGPT Work.
Two smaller features shipped alongside it. Sites in ChatGPT entered public beta, turning a project into an interactive site or web application shareable by URL, with OpenAI suggesting live dashboards, launch calendars, internal portals and interactive reports as candidates. Scheduled Tasks handles the repetitive end, letting ChatGPT run an action once, repeat it on a timetable, or watch for changes over time.
Beta testers included staff at Nvidia, Zapier, RingCentral and Virgin Atlantic. Reported use cases ranged across reviewing thousands of sales leads, checking product launch readiness against Jira tickets, comparing airline passenger experiences, and automating event preparation for Nvidia GTC 2026.
Billing works differently here. Because the agent runs long, usage is metered like Codex rather than counted as a standard chat request, so heavier tasks eat more of a plan's allowance. Enterprise and Edu administrators can set spend caps, group limits and individual overrides in the Admin Console.
Three Tiers and a Fourth Price
GPT-5.6 retires OpenAI's old naming mess. The number marks the generation. The name marks a capability tier that advances on its own schedule, replacing the mini and nano suffixes that made the previous lineup difficult to reason about.
Sol is the flagship, priced at $5 per million input tokens and $30 per million output. It carries the deepest reasoning and the only access to the new max reasoning setting. Terra lands at exactly half that, $2.50 and $15, and OpenAI positions it as the migration target for anyone currently running production workloads on GPT-5.5. Luna, the cheapest, costs $1 and $6.
Then there is a fourth line item. Sol Fast serves the same flagship model from Cerebras hardware at up to 750 tokens per second, billed at $12.50 and $75. Speed sold as an explicit paid tier is new for OpenAI.
All three models share a 1.05 million token API context window and a 128,000 token output ceiling. Caching received a rework: explicit cache breakpoints let developers declare boundaries in the prompt, cache reads keep the 90% discount, and cache writes now bill at 1.25 times the uncached input rate. Cached content survives a guaranteed 30 minutes.
Sam Altman told CNBC that GPT-5.6 delivers a 54% improvement in token efficiency on agentic coding compared with GPT-5.5. "Every enterprise now is thinking about spend and the value they're getting," he said.
OpenAI also introduced two compute controls. Max gives Sol additional time to reason through hard problems. Ultra goes further, coordinating four agents in parallel by default, trading heavier token consumption for a better result.
The Benchmarks Cut Both Ways
Sol posts 88.8% on Terminal-Bench 2.1, a test of command-line workflows requiring planning and tool coordination. Ultra mode lifts that to 91.9%, though independent analysis puts the cost of those extra 3.1 points at roughly three times single-agent Sol. Claude Fable 5 scores 86.0% on the same test, per Anthropic's own reporting.
Sol leads the Artificial Analysis Coding Agent Index at 80, sitting 2.8 points above Fable 5. On OSWorld 2.0 it reaches 62.6% while burning 85% fewer output tokens than Claude Opus 4.8.
The picture inverts on repository-level coding. Sol scores 64.6% on SWE-Bench Pro against Claude Mythos 5's 80.3%, a gap of roughly 15 points. One day before the launch, OpenAI published research arguing that around 30% of SWE-Bench Pro tasks are broken and advising model developers to scrutinize results from it. Until a replacement benchmark arrives, the number stands.
Fable 5 also leads on Toolathlon, HealthBench Professional and GDPval-AA v2.
Independent testers have been measured. Simon Willison, who had early access to Sol, described it as competent while noting it had not struck him as better than Anthropic's model on the complex coding tasks he runs.
There is a discrepancy worth flagging. The 53.6 headline score on Agents' Last Exam, cited widely in launch coverage, does not appear in OpenAI's own published table, which lists 52.7%.
Luna, meanwhile, breaks on long context. It drops to 41.3% on the MRCR v2 eight-needle test at both the 256K to 512K range and the 512K to 1M range.
A Desktop Merge and a Browser Retired
The structural news arrived quietly. The standalone Codex application is folding into the new ChatGPT desktop app. The existing desktop client becomes ChatGPT Classic. Codex gains inline editing within diffs, pull request review in a side panel, and multi-repo support in the process. Developers can keep Codex as their default view and even retain the Codex logo as their app icon.
OpenAI also began sunsetting Atlas, the standalone browser it launched less than a year ago. Its agentic browsing work moves into the ChatGPT desktop app's built-in browser and an updated Chrome sidebar extension. Note the phrasing in OpenAI's post: this is the announced beginning of a sunset. Atlas users are being transitioned rather than cut off.
Retiring a browser that was itself positioned as a distribution strategy tells you the company now believes the desktop app is the surface that matters.
Washington Signed Off on This Release
GPT-5.6 was ready weeks ago. On June 25, the White House Office of the National Cyber Director and the Office of Science and Technology Policy asked OpenAI to restrict the rollout to a small set of government-approved partners, roughly 20 organizations. It marked the first time the US government preemptively asked an American AI company to hold back a model before release.
The request traced to an executive order signed earlier that month establishing a voluntary framework under which AI developers offer covered frontier models to the government for up to 30 days of review ahead of release. A source familiar with the discussions said the government intervened because GPT-5.6 possesses what they described as Mythos-like capability, referencing Anthropic's most advanced system.
Altman briefed employees that the government would approve access customer by customer during the preview. He was blunt about it in an internal memo, calling the arrangement "not our preferred long term model." On X, he added that extensive safety testing made sense to him, but that he disliked the idea of the government picking the customers.
Commerce Secretary Howard Lutnick was directly involved, pressing for confirmation that every relevant agency had tested and approved the model before broad release. The Department of Commerce cleared the launch this week following additional testing.
The precedent came from Anthropic. On June 12, an export control directive forced the company to pull Claude Fable 5 and Claude Mythos 5 entirely offline over concerns that foreign nationals could exploit them, specifically their ability to identify software vulnerabilities and carry out multi-step attacks without human direction. Anthropic called the order a misunderstanding. The Department of Commerce lifted the controls on June 30, and access returned on July 1, roughly three weeks later.
A White House official has pushed back on the framing, stating that the administration does not mandate preclearance for model releases.
Enterprise Is the Contested Ground
Anthropic draws roughly 80% of its business from enterprise customers, and the Ramp AI Index published in May showed the company overtaking OpenAI in business adoption for the first time, with 34.4% of firms paying for Anthropic services against OpenAI's 32.3%.
Claude Cowork arrived in January and moved markets. When Anthropic expanded it in late February with 13 new connectors covering Google Workspace, DocuSign, FactSet and WordPress, plus prebuilt plugin templates for HR, finance, engineering and investment banking, IBM shares fell 13.2% while integration partners rallied. Investors read the announcement as a signal about which software categories agents absorb first.
Anthropic published usage data in early July drawn from 1.2 million anonymized Cowork sessions across more than 600,000 organizations. The overwhelming majority of that activity had nothing to do with writing software.
"We expect that every knowledge worker will feel that way about Cowork," said Kate Jensen, Anthropic's Head of Americas, comparing the tool's trajectory to Claude Code, which the company says crossed $1 billion in revenue.
OpenAI's counterattack targets exactly that surface. ChatGPT Work reaches roughly 900 million weekly ChatGPT users, and the company says nearly 100% of its own internal teams, finance and sales included, now run on some combination of Work and Codex. That figure is self-reported.
The Listings Are Watching
Anthropic filed a confidential S-1 with the SEC on June 1 at a private valuation near $965 billion, with reported run-rate revenue above $44 billion and a target of an October Nasdaq debut. OpenAI filed a week later, last valued at $852 billion, with annualized revenue passing $20 billion at the close of 2025.
SpaceX complicated the math. It priced its June IPO at $135, ran above $225 within days, then surrendered roughly 32% of those gains. Bloomberg reported in late June that OpenAI is now weighing a slide into 2027, with Altman holding a $1 trillion floor.
That sequencing carries a consequence most launch coverage skipped. If Anthropic lists first, whatever multiple public investors assign to its revenue in October becomes the denominator every banker plugs into an OpenAI model the following year.
Which makes a pricing announcement about token efficiency something closer to an argument about gross margin, delivered to an audience that will eventually read the filings.
Cost per Task Is the New Benchmark
Analysts have been circling the same point. Neil Shah, vice president for research at Counterpoint, said rising token consumption has produced bill shock across enterprises, forcing organizations to run different models for different workloads and making performance per dollar the metric that decides purchases.
Faisal Kawoosa, co-founder and chief analyst at Techarc, put it more sharply. "The exploratory stage of AI is over," he said.
The routing advice emerging from early analysis reflects that. Terra handles the everyday queue at half of Sol's price while landing within two or three points on most benchmarks. Sol earns its premium on long-horizon agentic work, terminal operations and computer use. Luna wins on cost per unit of work rather than on capability, and collapses on anything requiring long-context recall.
Max Weinbach, an analyst at Creative Strategies, said the smallest model in the family completes tasks about as well as the largest at a fifth of the cost, and that he had not seen small models handle this class of work before.
One line in the system card received less attention than it deserved. All three tiers, not only the flagship, carry OpenAI's High risk classification for cyber and biological capability.
Comments
Join the discussion and share your perspective.