In brief · the one-minute read
- Chained AI steps multiply their errors when nothing checks them: six steps that are each right 95% of the time are all right together about 74% of the time. See it run.
- Chains with real checks do better than a single pass, but only when the check sees something the step couldn't. Why.
- What decides it is a decision's status, not the number of steps. A choice that shapes everything downstream, but that nobody can see, question or change, is a ghost decision.
- One reasonable grouping can produce a well-written, useless article, and rejecting it only reruns the same mistake. Watch it happen.
- The fix is correction boundaries: places where a decision can be challenged and everything built on it updated. Open the decisions with the widest reach and reuse first. The rules and where to start.
The promise: power without the manual
For decades, software made you pick. Simple tools did less. Powerful tools did more, but you had to learn to drive them: taxonomy editors, report builders, rule engines.
AI breaks that trade-off. You describe the result you want and the system runs the tools for you. Jakob Nielsen calls this intent-based outcome specification, the first new UI paradigm in 60 years: you say what you want, not how to get it.
Under the hood, that usually means a chain. One step classifies, the next summarizes, the next analyzes, the next writes. Each step is a specialist, and together they do things no single tool could.
Go deeperThree paradigms of interface, and the one a sealed chain brings back
Nielsen describes three paradigms of user interface [1]. In batch processing, users submitted a whole job at once, and a single error could make the output meaningless. Command-based interaction fixed that by letting users reassess after each command. Intent-based outcome specification has users state the outcome they want and leaves the method to the system. Nielsen notes the cost: when users don't know how a result was produced, it is harder for them to identify or correct problems, and he expects hybrid interfaces that keep graphical controls.
The observation this paper starts from: a chain whose intermediate decisions are hidden and fixed reintroduces batch processing's defining property. The user sees only the final output and can only resubmit, even though the interface looks conversational.
The catch: errors multiply
If every step in a chain builds on the one before, the whole thing is only right when every step is right. So the odds multiply. Run the same chain a hundred times and count:
78 / 100runs came out right. The arithmetic says 74%.
- right
- wrong, and passed on
- built on a wrong step
- caught and sent back
Six steps at 95% each come out right together only about 74% of the time. At 80% per step, it's about 26%. That assumes nothing catches mistakes along the way, which is exactly what a chain looks like when nothing is checking. Put a check between the steps and the picture changes; how much depends on what the check can see.
In the batch-processing era, you handed over a whole job at once, and if it held the slightest error, the output was meaningless. Interactive interfaces fixed that by letting you look after each step and adjust. A sealed AI pipeline quietly brings batch processing back.
Go deeperWhat the research measures
That chained and multi-step AI systems compound errors is well documented:
- Compositional reasoning. Dziri et al. argue that autoregressive models' performance can decline rapidly as task complexity grows, and find that models often solve multi-step problems by pattern matching rather than systematic reasoning [2].
- Agent reliability. τ-bench introduced pass^k, the chance an agent succeeds on all of k repeated trials. Leading models succeeded on under half of tasks, and pass^8 fell below 25% in the retail domain [3].
- Task length. METR measures the length of task agents can complete with 50% reliability, a framing that exists because success falls as tasks get longer [4].
- Multi-agent systems. An analysis of over 1,600 traces across seven multi-agent frameworks found 14 failure modes in three groups: specification and system design, inter-agent misalignment, and task verification. Most were design problems, not model limitations alone [5].
The simulation assumes each step fails independently. Real steps share data and models, so failures correlate, and later steps can sometimes repair earlier ones. It shows the direction of the effect, not a measurement.
The twist: checked chains do better
If that were the whole story, the answer would be easy: use fewer steps. It isn't.
Some of the best AI results come from chains with checking built in. OpenAI's Let's Verify Step by Step trained a model to check every reasoning step instead of only the final answer, and step-by-step checking clearly won on hard math problems [6]. METR found the length of tasks AI agents can finish has been doubling roughly every seven months, driven partly by better reliability and a better ability to adapt to mistakes [4].
So chaining isn't the villain. A well-designed chain can take on bigger work than any single model, because each step is smaller and something can check it. But the checking has to be real:
- Models checking their own work often don't help. They struggle to fix their own reasoning without outside feedback, and sometimes get worse [7].
- AI judges share blind spots. As models get more capable, their mistakes are becoming more alike, and AI judges score models similar to themselves more favorably [8].
- More calls isn't automatically better. Adding model calls can improve results and then hurt them, depending on how hard the questions are [9].
Taken together: chains compound errors when nothing independent can intervene, and catch them when something can.
What decides it: ghost decisions
A chain that compounds errors and a chain that catches them can look identical: same steps, same models. The difference is what happens to each step's decisions.
- Open decisions can be questioned. A judge, a test or a person can challenge a step and send it back.
- Locked decisions are passed downstream as fact. Every later step builds on them, and nothing can push back.
Ghost decision: a choice the system made that shapes everything after it, but that nobody can see, question or change.
Ghost decisions are what turn a self-correcting chain into a compounding one. A small error gets locked in. The next step makes a small decision on top of it, and that gets locked in too. The result drifts further from what the customer needed while staying perfectly consistent with itself: the compounding assumption effect.
A ghost decision doesn't have to come from a model. A default configuration, a hard-coded rule or the product team's own workflow design can make one; a deterministic rule can apply the wrong business definition perfectly. And a check only counts if it can see something the step couldn't. A judge with the same blind spots, or tests written from the same wrong assumption, isn't independent. Sometimes the only one with that outside view is the customer.
Go deeperWhere the idea comes from
The interaction-design side of the problem has a long history, and much of the vocabulary already exists:
| Source | Core idea | Relevance here |
|---|---|---|
| Norman, 1990 [10] | The problem with automation is inappropriate feedback and interaction, not automation itself | A good happy path doesn't make a good product; the exception path is part of the design |
| Woods, 1996 [11] | Automation transforms work rather than removing it; tighter coupling spreads disturbances and complicates diagnosis | Chaining moves effort from operating tools into diagnosing and repairing outputs |
| Green & Blackwell, 1998 [12] | Hidden dependencies, viscosity (resistance to change), progressive evaluation | A ghost decision is a hidden dependency; a sealed chain is viscous |
| Horvitz, 1999 [13] | Mixed initiative: design automation around uncertainty about the user's goals and chances to refine results | An inferred decision is a guess about intent, not permission to proceed |
| Shneiderman, 2020 [14] | Automation and human control are separate dimensions; both can be high | One-click workflows can still support inspection and override |
| Amershi et al., 2019 [15] | Guidelines for human-AI interaction, including efficient correction, global controls and conveying consequences | Correction needs a clear scope and a preview of its effects |
| Wu, Terry & Cai, 2022 [16] | Letting users edit LLM chains and intermediate results improved outcomes and perceived transparency and control | The closest direct precedent for correction boundaries |
| Ehsan et al., 2024 [17] | Seamful design: revealing useful mismatches can support understanding and agency | Exposing a decision is a feature, not an imperfection |
The closest precedent is AI Chains [16]. Its study combined decomposition with editable intermediate steps, so it does not isolate the value of the editable boundary itself, and it was not run on enterprise products at scale.
The proposal, stated plainly. Whether a chain compounds errors or catches them depends less on its number of steps than on the status of its intermediate decisions. Where decisions stay open to an independent check, errors can be caught. Where they are locked and passed downstream as fact, errors compound. This is a synthesis of the work above, not an empirical finding; see the limits.
A ghost decision in the wild
Picture a knowledge-base product that turns support tickets into help-center improvements. One request in, one finished recommendation out. Behind the scenes, it's a four-step chain.
In step one, the AI groups "refund requested," "refund delayed" and "chargeback disputed" into a single topic called Billing. It's a reasonable guess. But this company handles those three with different teams and different policies, and the surge this month is in payouts that arrive late. Try both ways out of it:
Run 1 of the sealed product.
- 1Classify the ticketsGhost decision
“Refund requested”, “refund delayed” and “chargeback disputed” look alike, so they become one topic.
- 2Report on each topicBuilt on itBilling18 → 25 (+39%)
Billing is up 39%. True, and no help: the growth has nowhere to point.Bar: this month · tick: last month
- 3Find coverage gapsBuilt on it
Topic Tickets Articles Gap Billing 25 2 Biggest “Not enough Billing content”: the two existing articles are spread over 25 tickets.
- 4Suggest an articleBuilt on itDraft articleHow to request a refund
Requesting a refund is easy. Open your order, choose Request refund, pick a reason and submit. Most requests are reviewed within two business days…
A help-center article with this title already exists.
The article the sealed chain writes is well made and beside the point: it already exists. Every step did its own job correctly. The report counted the topic it was given, the gap finder found the gap the report implied, the writer wrote. The only mistake was one reasonable grouping that nothing downstream was allowed to question.
Rejecting the article puts the customer in the reroll trap. The next run makes the same grouping and the same mistake. They're fixing the symptom over and over, because the cause is a ghost.
The fix: seamless, not sealed
The answer isn't to bring back every manual step. It's to keep the decisions that matter open.
Correction boundary: a place where a decision can be seen, challenged and changed, and everything that depends on it gets updated.
A correction boundary can be used by a machine or a person:
- A judge or verifier checks the topic grouping against the company's existing help-center structure before reports get built.
- A person sees the Billing topic with its example tickets, splits it in three, and gets a preview: "This recalculates 1 report, reassesses 3 gaps and flags 1 draft for review. Nothing is published."
The happy path stays one smooth flow. The boundaries are just there when someone needs them. Ben Shneiderman makes this point well: automation and human control aren't opposites, and you can have a lot of both [14].
That's the difference between seamless and sealed. A seamless product removes unnecessary work. A sealed one also removes the places where mistakes get caught and customers say how their business is different.
Seven rules for killing ghost decisions
- Make key decisions real records. A topic definition, a metric, a plan: give it a name, a version and a history. You can't check or change what only exists inside a model call.
- Give every check an outside view. A verifier needs something the step didn't have: different data, a different model, the customer's own definitions, or a person [7, 8].
- Show decisions, not machinery. Show the Billing topic and its tickets, not prompts and logs. Seeing every step isn't the same as being able to change anything.
- Let people trace from the result back to the cause. Every output should link to the decisions and evidence behind it. Real records, not a plausible story.
- Make the scope explicit. This ticket, this topic, or every future run? A local fix should never quietly become a global policy [15].
- Update what depends on the change, and protect what's approved. Recompute downstream work, keep edits people already signed off on, and never resend an email just because an analysis reran.
- Match control to consequence. A throwaway draft needs an optional edit. A shared taxonomy needs a preview and permissions. Anything that leaves the building needs approval first.
Where to invest first
Opening up a decision costs something. You have to pull it out of the model call, store it, check it and build a way to change it, which adds engineering time and latency. So start where it pays off: decisions with wide reach (lots of work depends on them) and long reuse (they stick around). Here are the decisions in the same help-center product:
Every report, every gap and every draft, every week.
Consequence can jump the queue. A one-time action that's hard to undo, like emailing customers or deleting data, deserves a strong boundary even if it never repeats.
For each important decision, ask: when this is wrong, who or what notices, where does it get fixed, and what happens to everything built on it? If nobody can answer, you've found a ghost.
When not to bother
- Low-stakes, short-lived steps. Nobody needs to edit how one sentence was phrased in a throwaway draft.
- Decisions nobody can judge. If users lack the expertise to improve a decision, an edit field won't help. A better verifier or an expert reviewer will.
- Honest narrow products. If your product deliberately does one thing, say so. Doing less isn't the problem; pretending to be flexible is.
The goal isn't a settings screen for everything or an approval step at every turn. It's the right check, in the right place, by something that can actually see the mistake.
Before you ship
Chaining AI steps isn't good or bad. Chains with open decisions catch mistakes and take on bigger work. Chains full of ghost decisions multiply mistakes and trap customers in the reroll loop. Before you ship an AI workflow, ask:
- Which decisions shape everything downstream?
- Can a judge, a test or a person challenge each one, with a view the step itself didn't have?
- Can users get from a bad result back to the decision behind it?
- When a decision changes, does the work that depends on it update, without wrecking approved work?
- Have you tested how long it takes to fix a wrong result, not just how fast the first one appears?
Automate the work. Keep the decisions open.
Limits, and how to test this
This is an argument built from other people's evidence, not a study of its own. The parts that would change it are below.
Go deeperLimitations
The proposal is a synthesis, not an empirical finding. It has not been tested directly; the cited studies support its parts in different settings (math reasoning, agent benchmarks, lab studies of chain editing), not the combination in production products. The simulation assumes independent failures, which real chains violate. The worked example and its tickets are made up.
Go deeperOpen questions
- Decomposition cost. Many current systems make intermediate decisions implicitly inside a single model call. How much latency, cost and quality does extracting them as records trade away?
- Discovery. How can teams find their ghost decisions before customers do? Support logs, override frequency and churn reasons are candidate signals.
- Verifier independence. How independent must a machine check be to help, given that model errors are becoming more correlated [8]?
- Who can judge. When users lack the expertise to evaluate a decision, which boundary works best: an edit field, an expert reviewer, or stricter validation?
Go deeperA study that would test it
Compare three interfaces over the same underlying pipeline: a sealed end-to-end experience; one with visible provenance but no source-level edit; and one with provenance, scoped correction and dependent repair. Tasks should include a seeded upstream error, a reasonable but unsuitable business definition, a late-stage error that needs no upstream change, and a case where the default is already right, so the study does not reward editing for its own sake. Participants should be people who own those decisions, such as support leads and knowledge managers.
Measures: time to a validated result, successful source-level corrections, repeated symptom fixes, damage to approved work, and added cost on runs where nothing was wrong. For recurring workflows, check whether corrections persist in later runs.
References
- Nielsen, J. (2023). AI: First New UI Paradigm in 60 Years. Nielsen Norman Group.
- Dziri, N., et al. (2023). Faith and Fate: Limits of Transformers on Compositionality. NeurIPS 2023.
- Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.
- Kwa, T., West, B., et al. (2025). Measuring AI Ability to Complete Long Software Tasks. METR.
- Cemri, M., et al. (2025). Why Do Multi-Agent LLM Systems Fail? NeurIPS 2025.
- Lightman, H., et al. (2023). Let's Verify Step by Step. OpenAI.
- Huang, J., et al. (2024). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024.
- Goel, S., et al. (2025). Great Models Think Alike and this Undermines AI Oversight. ICML 2025.
- Chen, L., et al. (2024). Are More LLM Calls All You Need? Towards the Scaling Properties of Compound AI Systems. NeurIPS 2024.
- Norman, D. A. (1990). The "Problem" with Automation: Inappropriate Feedback and Interaction, Not "Over-Automation". Philosophical Transactions of the Royal Society B, 327, 585–593.
- Woods, D. D. (1996). Decomposing Automation: Apparent Simplicity, Real Complexity. In R. Parasuraman & M. Mouloua (Eds.), Automation and Human Performance, 3–17. Erlbaum.
- Green, T. R. G., & Blackwell, A. F. (1998). Cognitive Dimensions of Information Artefacts: A Tutorial.
- Horvitz, E. (1999). Principles of Mixed-Initiative User Interfaces. CHI '99.
- Shneiderman, B. (2020). Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy.
- Amershi, S., et al. (2019). Guidelines for Human-AI Interaction. CHI 2019.
- Wu, T., Terry, M., & Cai, C. J. (2022). AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. CHI 2022.
- Ehsan, U., Liao, Q. V., Passi, S., Riedl, M. O., & Daumé III, H. (2024). Seamful XAI: Operationalizing Seamful Design in Explainable AI. Proc. ACM HCI, CSCW1.