One Big Prompt or a Chain of Small Ones?
Should you use one big prompt or chain several small ones?
Start with one prompt and split only when you need to see or control what happens in the middle. Modern models handle multistep reasoning internally, so chaining for its own sake adds cost, latency and failure points. Chain when you need to inspect, log or branch on an intermediate result.
This used to be settled advice in the other direction. Two years ago, breaking a task into small steps was the standard fix for almost any unreliable output.
That advice aged. The models changed underneath it, and a lot of automation built on the old assumption is now more complicated than it needs to be.
What has changed about this question in 2026?
Models now do internally what chaining used to do externally. Anthropic's current prompting documentation puts it directly, saying that "With adaptive thinking and subagent orchestration, Claude handles most multistep reasoning internally." That single sentence invalidates a lot of pipeline design from 2024.
The same page does not dismiss chaining. It says explicit chaining, meaning "breaking a task into sequential API calls," is "still useful when you need to inspect intermediate outputs or enforce a specific pipeline structure."
Read those two statements together and you get the real rule. Chain for control, not for capability. If you are splitting steps because you think the model cannot manage them together, check that assumption. If you are splitting them because you need to log step two before step three runs, that is a good reason.
When does a single prompt win?
When the task is one coherent piece of work and you only care about the final output. Writing a summary, classifying a ticket, extracting fields from a document, drafting a reply. In each case the intermediate reasoning is not something you need to store or act on separately.
It also wins on everything operational. One call means one place for errors, one retry policy, one latency figure and one bill line. Every step you add is another thing that can time out at three in the morning.
Anthropic's guidance on delegation points the same way. Its sample prompt says to work directly rather than delegating "For simple tasks, sequential operations, single-file edits, or tasks where you need to maintain context across steps." That last clause is the one people miss. Splitting a task throws away context, and sometimes the context was the point.
When is chaining genuinely worth it?
When a human or a system needs to see the middle. If a reviewer must approve a draft before it publishes, that is a chain. If you need to store extracted data before the writing step so you can audit it later, that is a chain. If step two branches on step one, that is a chain.
Anthropic names the pattern that earns its keep most often. It describes self correction as the most common chaining approach, where the model generates a draft, then reviews that draft against criteria, then refines based on the review. The docs note that "Each step is a separate API call so you can log, evaluate, or branch at any point."
That last sentence is the whole justification. You are paying for observability. If you are not going to look at the logs or act on the branch, you are paying for nothing.
What does chaining cost you?
Tokens, time and reliability, in that order. Each call resends context, so a four step chain can pay for the same instructions four times. Each call adds its own round trip, so latency stacks. And each call is an independent chance to fail, which means your success rate is the product of four probabilities rather than one.
The reliability maths is the part teams underestimate. A step that works ninety five percent of the time feels fine on its own. Four of them in a row gets you to roughly eighty one percent, and now one run in five needs attention from a person.
There is a quality cost too. A model that writes and then critiques its own work in one pass has the full picture. Split across calls, the reviewing step only sees what you chose to pass it, and the thing you forgot to pass is usually what mattered.
Does prompt caching change the maths?
Substantially, and it is the strongest technical argument for chaining today. Caching means the repeated context in a chain does not cost full price. Anthropic's documentation states that "Cache read tokens are 0.1 times the base input tokens price," with per model exceptions that go lower still.
Writing to the cache carries a premium. The docs say "5-minute cache write tokens are 1.25 times the base input tokens price" and "1-hour cache write tokens are 2 times the base input tokens price." So you pay a little more once to pay much less on every subsequent call.
There is a floor to be aware of. The minimum cacheable prompt is "512 tokens for Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5.5, Claude Opus 5, Claude Sonnet 5.5, Claude Fable 5, and Claude Mythos 5," with higher minimums on some other models. Below that, caching does nothing and the repeated context costs full price.
So a chain sharing a large stable prefix is now affordable in a way it was not before. A chain sharing almost nothing gains nothing. We went into the wider version of this in our notes on keeping AI costs under control.
How does this differ from using subagents?
Subagents are the model choosing to split work, rather than you hard coding the split. Anthropic's docs say that "Claude's latest models orchestrate subagents natively" and can "recognize when tasks would benefit from delegating work to specialized subagents and do so proactively without requiring explicit instruction."
That is a meaningful difference. A chain you design runs the same way every time. Delegation adapts to the task, which is better when the task varies and worse when you need a guaranteed shape.
It can also overshoot. The docs warn to "Watch for overuse," noting one model "has a strong predilection for subagents and may spawn them in situations where a simpler, direct approach would suffice," with the example of spawning subagents for code exploration "when a direct grep call is faster and sufficient." Their suggested fix is to tell the model when delegation is warranted. We compared these shapes in our piece on one agent or many.
How do you decide for a real workflow?
Ask what you need to see, what you need to control, and what you need to keep. Those three questions resolve almost every case without any theory at all. The table below is how we talk it through with a team, working from the job rather than from the model.
| Situation | Better shape | Why |
|---|---|---|
| Classify and reply to a support ticket | Single prompt | Nobody reads the middle step |
| Draft content a person must approve | Chain | The approval is the intermediate step |
| Extract data you will store and audit | Chain | The extracted record has its own value |
| Route to different handling by type | Chain | You branch on step one |
| Rewrite a page in your brand voice | Single prompt | Splitting loses the context that matters |
| Research across many documents | Delegation | Work is parallel and context is isolated |
Notice that none of the answers depend on task difficulty. Difficulty was the old reason to chain, and it is the one that stopped applying.
What mistakes do we see most?
Chains built for a model generation that no longer exists. Someone solved a real problem in 2024 by splitting a task into six steps, the workflow has run ever since, and nobody has asked whether steps two through five are still doing anything. Usually two of them could be deleted with no change in output.
The second mistake is chaining without ever reading the logs. If you accepted the cost and complexity of separate calls in order to inspect intermediate output, and then nobody inspects it, you bought observability and threw away the receipt.
The third is treating the split as permanent. Model capability moves fast enough that pipeline design should be revisited on a schedule, the same way you would revisit a dependency. Our notes on context engineering cover how we think about that drift.
What is our default?
One prompt, with thinking enabled, until something forces a split. Then split at the point where a human or a system needs to act, and nowhere else. Keep the shared context in a cacheable prefix so the extra calls stay cheap.
The broader principle is that complexity should be demanded by a requirement rather than chosen by habit. Every step in a pipeline should be answering a question somebody actually asks.
If you have an automation that grew step by step and nobody is sure which parts still earn their place, we are happy to look through it with you. You can find us at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.