Anthropic Releases Claude Opus 5.5 With Stronger Safety Tests and Lower Price
Around lunchtime in New York on Sept. 22, 2026, Anthropic released Claude Opus 5.5. It was the first model in the Claude 5.5 family. The list price was $4 per million input tokens and $20 per million output tokens.
Anthropic said the model performed near Fable 5.1 on most tasks while costing 40 percent less to run than Opus 5 at default settings on typical workloads. Fable 5.1 listed at $10 per million input tokens and $50 per million output tokens. Cache reads, which reuse context a model has already processed instead of starting over, cost 60 percent less, down to $0.20 per million tokens. Output arrived more than 30 percent faster than Opus 5.
Dianne Penn, head of product management, research and labs, put the release in the language of the bill. "One of the things we're continuing to innovate on is how to make that thinking, how to make the answering more efficient, so it uses less tokens depending on your effort setting," she said. Tokens are the pieces of text a large language model reads and writes; they are what customers pay for, by the million, on the way in and the way out. A large language model is an AI system trained on vast amounts of text so it can generate, summarize, and analyze language.
Opus sits in Anthropic’s high-capacity tier, aimed at professional work in fields such as finance, law, and software development. Earlier releases had chased benchmark records; this one opened on fewer tokens spent and a lower sticker. The $4 and $20 figures were the floor. What the model could actually do at that price was the rest of the claim.
On Terminal-Bench 4.0, which scores whether an agent—a model that takes multi-step actions on its own—can finish complex professional tasks inside a command-line interface, Opus 5.5 hit 66.4 percent. Fable 5.1 scored 55.8 percent. GPT-6 Astra scored 57.9 percent. Early testers had described the prior Opus as "steadier, with far less variance run to run." On FrontierCode, a pass-rate check of whether an agent's code changes would get merged in a real engineering pipeline, Opus 5.5 reached 54.4 percent against Fable 5.1's 50.3 percent. Fable-class work now sat at a price well below Fable's own list.
Opus 5 had launched on July 24, 2026. Fable 5.1 had arrived in early September. Against that tight sequence, the Tuesday release landed inside an industry argument already under way about how fast the frontier should move. Jacob Coxon posted on X on Sept. 8 that he had quit Anthropic. He warned that the labs were "gambling with our lives." The post set off the industry debate.
Dario Amodei, Anthropic's CEO, answered with an essay calling for the industry to pace the frontier—to slow the rate at which labs improve model capabilities so that work on preventing risks could keep up. "We must slow the pace at which we improve the capabilities of AI models," he wrote. "I have become convinced that fully addressing the risks requires even more prudence, not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up." Sam Altman and Elon Musk publicly joined the slowdown call.
Coxon's warning and Amodei's essay shifted the pressure from benchmark scores to the speed of the next release. Whatever Anthropic shipped after the essay would be the company's first model since it had called for the industry to hold back. The company said as much when it put the model out. "Opus 5.5 is our first model since we called for pacing the frontier. As with previous models, it was tested by external evaluators before release, including METR and Frontier Design. On our most comprehensive alignment test, it achieves the strongest score to date." An alignment test checks whether a model behaves as its builders intend—cooperative rather than manipulative or deceptive—instead of scoring raw skill at coding or writing. That strongest score was the hard safety number on the release.
On the company blog, Anthropic tied the moment to a longer build of capacity around systems people would come to rely on. "As AI becomes more capable, public policy should play a larger role in making sure the systems people rely on are safe. That capacity takes time to build, and we've started to put the infrastructure in place to support it." The external reviews by METR and Frontier Design sat inside the claim that the model had been held to the prudence the company had just asked the rest of the industry to practice.
Anthropic raised five-hour usage limits by 20 percent for Pro, Max and Team subscribers. It also added a banked rate limit reset, a feature that would let a customer hold a reset in reserve and spend it later; full details on how it would work were still forthcoming. The extra headroom sat beside the lower token prices as the part of the release that reached the people already paying for the work.
Open-weight rivals had been selling access for less and pulling token share, and both frontier labs were answering customers who wanted more cost-effective models and tighter control of AI spending. OpenAI posted its reply the same afternoon. "Higher usage limits and lower cost give you more flexibility and room to iterate," the company said.
OpenAI rolled out GPT-6 Sol and GPT-6 Luna within hours, cutting prices further. Sol was meant for more complex workloads such as coding. Luna was geared toward high-volume tasks—extracting information, summarizing documents. Both joined the GPT-6 family the same Tuesday. By early afternoon the new tiers were already live.





