Claude can now write a program that sends work to other agents, checks what comes back, and combines the results while the main conversation remains available. For anyone who spends hours pulling facts out of product documentation, that is a useful idea. Its value will depend on how quickly it produces a brief that survives a source check, and what that brief costs.
Anthropic added hosted dynamic workflows to Claude Managed Agents in beta on October 9, 2026. Its official release note describes an agent-written program that runs agents in phases and combines their results. Anthropic’s server executes that program in the background as a workflow run.
A good test would give Claude 12 public product documents, require a separate worker to verify each extraction, and ask for a source-linked brief. Run the same job with one agent. Then compare time, total cost, and unsupported claims. A faster answer with invented pricing is a failed answer.
Kingy.ai has checked the launch documentation and pricing. We have not run the hosted beta or a matched single-agent benchmark. The experiment below is a proposed test, and all cost examples are calculations rather than measured results. The release note supplies an announcement date; we could not establish an exact announcement time.
What Anthropic shipped
The distinction from ordinary delegation is who coordinates the work. With subagents, the main agent delegates tasks and reads their reports. With dynamic workflows, a generated program passes context and results between workers and decides what happens next. That leaves the main thread free to communicate with the user and check progress. Anthropic describes this behavior in its multiagent orchestration documentation.
This is a Claude Platform API beta. The documented access header is:
anthropic-beta: managed-agents-2026-04-01
The agent definition can explicitly enable workflows with the following configuration fragment. This is part of an agent definition, not a complete runnable request:
{
"multiagent": {
"type": "multiagent_20261001",
"workflows": { "type": "enabled" },
"subagents": { "type": "disabled" }
}
}
For this proposed experiment, disabling ordinary subagent delegation would keep the workflow arm easier to identify. Workflows and subagents otherwise default to enabled under this multiagent type. The agent still decides when to start a run, so its prompt should specify the circumstances for using one.
Developers also need the surrounding agent, tools, environment, and session setup. Anthropic’s agent setup guide covers those resources and model configuration. Enabling the setting alone does not establish a working document pipeline.
The 12-document test
Product research is a good candidate because the unit of work is clear. One worker can read one document without waiting for the other 11. Verification can inspect the same source separately. The final writer then receives a compact set of checked records instead of having to rediscover every detail inside a long conversation.
Choose 12 official, publicly accessible documents before either run. A practical corpus could contain four products with three documents each: a feature overview, a limits or availability page, and a pricing page. Record every URL, retrieval time, document version where available, and a hash of the saved text. Freeze that material so a page changing between runs cannot decide the winner.
The proposed workflow has three stages:
- Extract. Assign each document its own extraction task. Return atomic facts with document IDs, source URLs, section locations, and short supporting evidence. Preserve currency, billing unit, plan restrictions, dates, and beta labels. A missing value stays unknown.
- Verify. Give a different worker the original document and its extracted records. For every claim, return supported, contradicted, or not established. Check whether a quoted price is monthly or annual, whether a limit belongs to the right plan, and whether “available” means generally available or beta. Flag important facts the extractor omitted.
- Write. Combine the checked records into a brief with links attached to individual claims. Keep unresolved conflicts visible. Exclude rejected claims rather than letting a fluent final writer smooth them into certainty.
Twelve extraction tasks plus 12 verification tasks and one synthesis task give this design 25 logical tasks. That arithmetic describes our proposed workload. Claude could group tasks differently, and retries could create additional threads.
Use a small record format: document ID, claim, source URL, section, evidence, verification status, and reason. The brief should also account for all 12 documents, including any that could not be read. A clean summary of eight documents is incomplete when the assignment was 12.
For a live-fetch version, send the source URLs in the user message. Anthropic’s October 7 release notes say web_fetch needs a URL that appeared in qualifying prior context; a URL present only in an attached document or the system prompt does not qualify. Align the environment and tool host allowlists as explained in the web restrictions guide. For the controlled benchmark, use identical saved documents in both sessions and record any live fetching separately.
Give the single agent the same job
A fair baseline needs the same extraction rules, verification requirements, and final brief format. Ask one agent to process the 12 documents serially, conduct a second source-checking pass, and write the result. Disable workflows, subagents, and advisor use in that arm. Comparing a carefully checked workflow with a one-pass summary would confound extra scrutiny with orchestration.
Keep the model, effort level, available tools, saved documents, output limits, and total spend allowance the same. If the workflow uses a cheaper verifier or a stronger writer, test that separately and label it as a different model mix. Otherwise the outcome cannot tell us whether parallel coordination helped.
Repeat the comparison, alternate which arm runs first, and disclose caching and retries. Separate cold-cache and warm-cache results where possible. Report the spread as well as the median; one lucky run is weak evidence for a latency claim. Predeclare the minimum number of repeated pairs and the pass criteria before seeing outputs.
| Measure | What to record | Why it matters |
|---|---|---|
| Time to final brief | Elapsed time from the task message to the completed, checked brief | Measures the deliverable people need |
| Responsiveness | Time to answer a fixed progress question, tested separately | A responsive conversation can coexist with slow background work |
| Total cost | Session-wide tokens, cache usage, runtime, tools, and retries | Includes the coordinator and every worker |
| Unsupported claims | Count and rate of factual claims not supported by the cited source | Tests whether verification improved the final brief |
| Coverage and failures | Documents processed, required facts omitted, contradictions, failed tasks | Prevents an incomplete or overly cautious brief from looking successful |
This table defines the proposed scorecard. Neither arm has a measured result yet.
Have a reviewer grade the outputs against the saved sources without knowing which arm produced them. Count factual statements individually: a sentence can get the product name right and the eligibility wrong. Also score links that point to a real page but fail to support the attached claim. Fewer unsupported claims means little if the brief avoids all useful details.
The $0.08 rate is only the runtime portion
Anthropic’s Managed Agents pricing lists $0.08 per session-hour in the running state, metered to the millisecond. Tokens use the applicable model rates and caching rules. Web searches cost $10 per 1,000 searches. Idle, rescheduling, and terminated time do not accrue the runtime fee.
At that rate, 10 minutes of running time costs about $0.0133; 30 minutes costs $0.04. Saving 20 minutes would save about $0.0267 in runtime charges. Those are calculated examples for one session, before tokens and tools. A verification pass could easily cost more than the runtime it saves, so the hourly number alone says little about the total bill.
The documented pricing model does not add a separate workflow-specific fee. Every worker’s model consumption still costs money. A short wall-clock run can consume substantial aggregate model work.
For measurement, use the session usage totals. Anthropic says per-thread costs exclude session runtime and are rounded separately; adding them does not exactly reproduce the session total. Its session-level usage.active_seconds counts overlapping thread activity once. Preserve input, output, cache-read, cache-write, and tool counts so someone else can audit the comparison.
The economic result should be expressed as total dollars per acceptable brief. If extra checking raises cost but prevents expensive errors, that may be a useful trade. If one agent already meets the accuracy threshold, parallel workers must justify their additional spend through time saved or better coverage.
Budget and completion need careful handling
Attach a session budget when creating the session. The documented budget amount is an integer number of US cents written as a string: "125" means $1.25. A session that starts without a budget cannot have one added later.
The cap prevents new model requests after the list-cost threshold is reached. Requests already in flight finish. With several workers, spend can pass the cap by one request per active thread. Reserve room for that overshoot, report budget-paused runs as incomplete, and avoid silently raising the cap during the comparison.
Completion also needs more than a progress message. Anthropic’s workflow-run guide says a running workflow can keep the session running even when none of its threads is working. Each created run must end, followed by the main session’s final end_turn idle event, to establish that the work is done.
A workflow ending with completed does not certify that every worker succeeded or that the answer passed a quality check. Inspect the worker events and the output. Interrupting the main turn does not automatically end workflow runs. These details make a visible chat response a poor substitute for a verified final artifact.
A second worker can still agree with a mistake
The verifier needs the original document. If it receives only the extractor’s summary, it can check internal consistency while missing the same omitted condition. Two workers using the same model can also share the same mistaken interpretation.
That is why the acceptance rules should demand evidence for each claim and a separate check of the final brief. A price page might put an annual discount in one section and taxes or usage charges in another. A second worker agreeing with the first is useful process evidence; the source still has to support the statement.
Parallelism is most promising when the documents can be handled independently and the final synthesis is bounded. It is less compelling when one short document contains the whole answer, when every task depends on a shared evolving state, or when coordination produces more intermediate text than useful evidence.
Kingy’s earlier AI loops guide explains the broader practice of generating, checking, and revising work. This hosted beta gives developers another way to execute that pattern. The new coverage question is whether its generated workflow improves a specific deliverable under the same constraints.
A concrete prompt for the demonstration
After configuring the agent, providing the frozen files, and setting the session budget, a workflow-arm prompt could read:
Read these 12 saved public product documents. Start one dynamic workflow with extraction, verification, and synthesis stages. Give each document an extraction task and have a different worker verify its claims against the original file. Preserve plan names, units, dates, prices, restrictions, and source URLs. Mark missing facts as unknown.
Produce a source-linked brief using supported claims, with a separate list of contradictions, unsupported claims, omitted required facts, and unread documents. Do not infer a product capability from marketing language. Do not replace frozen files with newer web content. Return a claim ledger and account for all 12 document IDs. Include errors and retries in the report.
For a video, show the fixed corpus and rules, the workflow’s phase events, the serial baseline, the two final briefs, and the blind source audit. Keep the session-wide usage receipts alongside the latency measurements. That would let viewers judge the result instead of inferring quality from a busy agent timeline.
Before declaring a winner, publish the latency, full bill, document coverage, and source audit together. For this workload, the workflow earns its place if it delivers an acceptable brief sooner at a cost the user is willing to pay.
The Kingy Brief
Get The Kingy Brief.
AI changes, original tests and one practical thing to try. Fridays at 09:00 Vancouver time.
Free · Double opt-in · Unsubscribe anytime
Signup help and newsletter schedule
Signup form provided by Beehiiv. After submitting, check your inbox for "Confirm your subscription to The Kingy Brief" and open its confirmation link. Check Spam or Promotions if you cannot find it.
Fridays at 09:00 Vancouver time: source-checked AI changes, original tests and one practical thing to try. The weekly restart begins October 9, 2026. We skip a week when there is not enough verified material. Free. Unsubscribe anytime.
