Dynamic Workflows Move to the Server

An agent can now write a program that runs many agents in phases on Anthropic's servers. How a run works, what it costs, and what to check first.

A violet crystal and a three-band plan panel send blue light paths through three columns of glowing cyan cubes in a dark server hall into one white cube.

On October 9, Anthropic added dynamic workflows to Claude Managed Agents, in beta on the Claude Platform. An agent can now write a workflow, which the release notes define as "a program that runs many agents in phases and combines their results", and the server runs it in the background as a workflow run.

The pattern itself is not new. Claude Code has had dynamic workflows since May 28, and they are now generally available there. In Claude Code the script runs on your machine and you can open it, edit it and rerun it. On Managed Agents it runs on Anthropic's servers, the agent decides when to start one, and you pay for it and cap it per session. That changes who holds the controls.

The primary sources are the Claude Platform release notes, the Multiagent orchestration and Workflow runs pages of the Managed Agents docs, the pricing page, and Anthropic's earlier writing on multi-agent systems from 2025 and 2026. Everything below is as of October 9.

What shipped

A Managed Agents agent now has three ways to hand off work. In the docs' words: "With subagents, it delegates tasks itself and reads what each subagent reports. With dynamic workflows, it writes a workflow: a program that runs many agents in the background and combines their results. It can also consult an advisor model for guidance while it does the work itself."

The configuration from the release notes:

{
  "multiagent": {
    "type": "multiagent_20261001",
    "workflows": { "type": "enabled" }
  }
}

With that type, subagents and workflows are both enabled by default, and each can define its own inline agents. The advisor is off until you turn it on. An inline agent "uses the model of the agent that the session runs". To give part of a run a different model, create that agent separately and list it in workflows.predefined_agents, which holds up to 20 agents.

How a run works

You describe the work in a user.message, and the agent decides whether a run is worth starting. There is no separate API call. The docs are explicit that the system prompt is where you steer that choice, and their example is a good template:

"You review contracts. When you're asked to review more than a few contracts, start a workflow run that reads them in parallel and combines the findings. Review one or two contracts yourself, without a run."

Once started, a run has three layers: the run itself, named phases such as "Read the contracts", and session threads where agents work. The program fans agents out in parallel, passes one agent's result to the next, can repeat a step until a review passes, and decides what happens when an agent fails. The key sentence is "An agent's result goes to the program, not to the agent that the session runs." The plan stops living in a conversation and becomes code. Claude Code's docs say it in one sentence: "A workflow moves the plan into code."

While the run works, the main agent is free to talk to you and check on progress. When the run ends, the agent gets a turn to read what it did and answer. Two boundaries matter here. The agent "can't send follow-up messages to a run's threads", so you cannot steer a worker halfway through. Only the agent on the primary thread can start a run, "so runs don't nest".

The limits that shape a run

LimitValue
Threads working at once in one run64. Anthropic notes the API does not guarantee this number.
Agents a workflow starts over the run's life1,000. Past that, the run ends with thread_limit_error.
Run lifetime24 hours by default, or what the agent sets.
Runs open at once in a session10 by default.

The 1,000 cap counts agents, not threads: "The server can run a failed agent again on a new thread, so a run might have more than 1,000 threads." Ordinary subagents keep their own ceiling of 25 child threads at a time per session, and a run's threads do not count toward it.

Throughput is a separate question. A run's model requests count toward your Messages API rate limits "along with your other traffic", so 64 slots do not mean 64 requests in flight if your organization's limits say otherwise.

What it costs

"A run has no price of its own. The tokens its agents use are billed like the session's other tokens, at each model's rates." On top of tokens, Managed Agents charges $0.08 per session-hour, and runtime "accrues only while the session's status is running". A run keeps the session in that state: "expect the session to stay running, even while none of its threads is working". At $0.08 per session-hour, 24 hours in the running state costs $1.92 per session, plus tokens.

For tokens, Anthropic's own numbers on multi-agent work are blunt. In 2025 it wrote that "multi-agent systems use about 15× more tokens than chats". In January 2026 it put the range at "3-10x more tokens than single-agent approaches for equivalent tasks". The Batch API discount does not apply to Managed Agents sessions.

The control is a session budget, and it has two quirks. You must set it when you create the session: "you can't add one to an existing session". When the budget is hit, every open run pauses, but "a run can pass the budget by one request for each working thread". At today's limit that is up to 64 extra requests per run, and more with several runs open.

The details that bite

An interrupt does not stop a run. "Only the agent starts a run. No event you send ends one; archiving the session can." A user.interrupt "stops the agent's turn. It ends no run." To stop a run, you send a message asking the agent to stop it. If your product has a Stop button that only sends an interrupt, the background work keeps going. When you ask the agent to stop a run, confirm that the run has ended by watching for its workflow_run.status_ended event.

"Completed" is not "passed". The docs: "The result doesn't say whether the work passed. A run can end completed even though work on its threads failed, or a thread couldn't be created." Your application needs its own acceptance rule. For a document review that means every document ends with a finding or an explicit failure. Ask the workflow for a coverage report and check it.

The same work can run twice. "A run can create more than one thread for the same piece of work, so make the tools your agents call safe to call twice." All threads also share one sandbox and its files, so give parallel workers separate output paths.

You can't see the program. "You don't see the workflow's code", nor the tool calls that start the run, nor what each thread hands back to the program. The documented workaround is to ask the agent afterwards to "Print the workflow that you started the run with, word for word". In Claude Code the script is a file on disk before it runs. Here you only get the agent's own account of it.

Starting a run skips the permission check. "Permission policies apply to the tools that a run's agents call, not to starting the run." The tools stay gated. The decision to fan out to hundreds of agents does not.

Migrating a coordinator agent can switch both features on. The older coordinator type could not start runs or define its own subagents. Move it to multiagent_20261001 without naming workflows and subagents.inline_agents, and "both take their default and are enabled". The update also replaces the whole block, so a subagent list you leave out comes back empty. Write every setting explicitly in that update. Existing sessions keep their old copy either way.

When it earns its tokens

The docs recommend workflows for "Most work that needs more than one agent: work with many pieces, parallel work, or a long task that should finish sooner", with audits, migrations, deep research and cross-checking as examples. Anthropic also gives criteria for when multiple agents pay off: "when context pollution degrades performance, when tasks can run in parallel" or "when specialization improves tool selection or task focus", because "Outside these situations, the coordination costs typically exceed the benefits." In 2025 it also listed what does not fit: domains "that require all agents to share the same context or involve many dependencies between agents".

My rule of thumb: reach for a run when you can name both the independent units of work and the rule that combines their results. Three hundred contracts and one question pass. A refactor where every change depends on the last one does not. If you cannot write the combining rule in a sentence, test the task with a single agent first.

Before the first run

  • Create the session with a budget, and leave room for the overshoot.
  • Write the run rule into the system prompt, including when not to start one.
  • Name per-model agents in workflows.predefined_agents for bulk phases that don't need your strongest model.
  • Make every tool that writes or sends safe to call twice.
  • Give parallel workers separate output paths in the shared sandbox.
  • Define what "passed" means, and check thread results against it rather than the run's result.
  • Make the Stop button send a message to the agent, then wait for the run's workflow_run.status_ended event.
  • Set workflows and subagents.inline_agents explicitly when you migrate a coordinator agent.
  • After each run, ask the agent to print the workflow and keep it with the output.

Claude Code made the plan a script you can read before it runs. Managed Agents makes it a server-side workload that you budget, observe and stop by asking. For production work, I would start with a small run that includes one failing document, one tool that is called twice, and one cancellation, and see what your application can actually observe.