Claude Opus 4.6
A release overview of Claude Opus 4.6: context limits, adaptive thinking, agent teams, integrations, benchmarks and launch pricing.
Anthropic released Opus 4.6 on February 5, 2026. This is a launch overview, with the pricing and availability statements below attributed to that announcement.
TL;DR
The launch introduced a one-million-token context beta on the Claude Platform, support for up to 128k output tokens, adaptive thinking, Claude Code agent teams, and office integration updates. The announcement reported strong benchmark results; those are evaluation-specific comparisons.
What Actually Changed
1M Token Context Window
At launch, the one-million-token beta expanded the available context beyond 200k on the Claude Developer Platform. Token capacity is not a fixed word count or proof that every codebase fits, and it does not establish reliable use of every detail.
Anthropic reported 76% on the eight-needle, one-million-token MRCR v2 retrieval test, compared with 18.5% for Sonnet 4.5. That is a specific retrieval improvement, not proof that context degradation is mostly solved across arbitrary long tasks.
128k Output Tokens
You can now get up to 128k tokens in a single response. If you've ever hit that frustrating wall where Claude cuts off mid-code-generation and you have to ask it to continue, this should help significantly.
Adaptive Thinking
Earlier extended-thinking interfaces exposed a token budget as well as an enable/disable setting. Adaptive thinking lets the model decide when deeper reasoning is useful, guided by the supported effort setting.
The Opus 4.6 API launch documented low, medium, high (default), and max effort levels. API effort configuration and Claude Code commands are separate interfaces; use the documentation for the client and model you are running.
Agent Teams in Claude Code
This one is massive for developers. Instead of a single Claude agent working through tasks sequentially, you can now spin up multiple agents that work in parallel and coordinate autonomously. Think of it as having a small engineering team instead of a single developer.
Best use case: tasks that split into independent, read-heavy work like codebase reviews, large refactors, or running analyses across multiple files simultaneously.
Agent teams launched as a Claude Code research preview, separate from the Messages API. Teammates, subagents and their navigation controls are not interchangeable; availability and shortcuts depend on the installed Claude Code version.
Context Compaction
The API compaction beta summarizes and replaces older context near a configured threshold. This can extend a workflow, but summaries can lose detail and other token, usage and runtime limits still apply.
Office Integrations
Claude in PowerPoint
New research preview. Claude now works directly inside PowerPoint as a side panel. It reads your existing layouts, fonts, and slide masters to stay on-brand. No more "create a deck in Claude → download → fix formatting in PowerPoint" workflow. You can now build and iterate directly.
Available for Max, Team, and Enterprise subscribers.
Claude in Excel (Upgraded)
The Excel integration got smarter. It can now ingest messy, unstructured data and figure out the right structure without you having to explain it. Multi-step changes in one pass. It plans before it acts.
Benchmarks
For the numbers people:
- Anthropic reported launch-time leads on Terminal-Bench 2.0, Humanity’s Last Exam and BrowseComp within the comparisons it published.
- Its GDPval-AA comparison put Opus 4.6 about 144 Elo points ahead of the compared GPT-5.2 configuration.
- MRCR v2: 76% versus 18.5% for Sonnet 4.5 on the eight-needle, one-million-token variant.
- Consult the linked system card and benchmark methodology for harnesses, resource limits and uncertainty; these are not permanent global rankings.
Benchmark gains were not uniform. Compare each workload and evaluation configuration rather than treating a headline ranking as a guarantee for software maintenance.
Pricing
The February 5 launch listed $5 per million input tokens and $25 per million output tokens. The then-beta Claude Platform pricing for prompts above 200k was $10/$37.50. These are historical launch terms; check the current rate card before budgeting.
The launch announcement listed US-only inference at 1.1 times token pricing. Availability and current rates need to be checked for the actual deployment.
What It Feels Like
This is the model I'm writing this post with (meta, I know). The difference from Opus 4.5 is noticeable — particularly in how it handles larger codebases and sustained reasoning over long sessions. It doesn't lose the thread as easily. It catches its own mistakes better during code review. And when it encounters ambiguity, it makes better judgment calls instead of asking you to clarify every little thing.
In Anthropic’s launch announcement, Cursor’s Michael Truell described stronger results on long-running tasks and code review in Cursor’s internal testing. That is attributed early-access feedback, not an independent controlled evaluation reported here.
For those of us using Claude Code daily, agent teams are going to be a workflow changer. I'm already thinking about how to restructure my development sessions around parallel agents.
The Bigger Picture
The release brought model and developer-product changes together. Adoption and revenue headlines are different measurements from engineering quality; evaluate the features against representative tasks and operating constraints.
The model is available now on claude.ai, the API (claude-opus-4-6), Amazon Bedrock, and Google Cloud Vertex AI.
Oh, and Anthropic is apparently running a Super Bowl ad on Sunday. We're definitely not in the early adopter phase anymore.