I put significant effort into reducing generation time in an AI-native workflow I was building. The reasoning felt obvious: waiting is friction, and seeing the result sooner gives a person more time to use it. A blank screen feels like nothing is happening. A finished page feels like progress. So I kept looking for places where the product could move sooner—starting work earlier, passing context more efficiently, and returning visible output without making the person sit through the machinery behind it.
That instinct was not wrong. Speed matters. The mistake was treating the arrival of output as the end of the product problem. Once the product became faster at generating, a different delay became visible. The system could finish its part before the person was ready to make sense of what it had done.
When fast output stopped feeling fast
The product did not generate one isolated object. It moved through a connected chain: product definition, visual direction, assets, and page output. Each stage made choices that shaped the next one. A phrase in the definition could influence the hierarchy of the page. A visual direction could constrain the assets. Those assets could make one layout feel natural and another awkward. By the time the page appeared, many decisions were already embedded in it.
Making that chain run faster created an impressive reveal. It also compressed the explanation of how the result came to be. The person could see a page, but not necessarily the product assumptions, visual commitments, and asset choices nested inside it. What looked like a single output was really a stack of intermediate decisions.
That changed the meaning of review. If the page felt wrong, where should the person intervene? Was the underlying product definition weak? Was the visual direction off? Did an asset distort the intended tone? Or was the final composition simply poor? A fast result did not answer those questions. In some ways, it made them harder to ask, because the polished output encouraged a verdict before it supported an understanding.
I began to see that the person was not reviewing one thing. They were being handed the accumulated consequences of several hidden choices. The product had shortened the time to output while leaving the work of reconstructing those choices to the user. Generation felt fast from the system side. Judgment felt slow from the human side.
Review debt
I think of this condition as review debt: generated output accumulates faster than informed judgment can be applied to it. Like other forms of debt, it can remain invisible for a while. The interface keeps moving. More drafts appear. More options become available. The apparent productivity rises. But every unexamined decision creates an obligation that someone must eventually pay through review, correction, rework, or acceptance of an outcome they do not fully understand.
Review debt appears when the system produces consequences faster than the user can form an informed judgment about them.
The problem is not simply volume. A person can dismiss many low-stakes variations quickly. The harder problem is mixed consequence. A disposable draft, a directional plan, and a factual claim may arrive through the same visual treatment even though they deserve very different kinds of attention. When the product presents all output as equally complete, it forces the user to discover the risk structure for themselves.
Review debt also changes how iteration feels. Producing another version is cheap, so generating again becomes the easiest response to uncertainty. But another version may add another set of decisions without clarifying the first set. The user now compares outputs while still lacking a view into the assumptions that produced them. More choice can create less control.
Generation latency is not decision latency
Generation latency is the time between asking and receiving an output. Decision latency is the time between having an output and reaching a decision the user understands well enough to own. They overlap, but they are not the same measure. A product can have a fast first output and a slow, confusing decision.
This distinction changed what I paid attention to. If I looked only at generation latency, every optimization that made the final page appear sooner looked beneficial. If I looked at decision latency, I had to ask whether the person could locate the consequential choices, inspect the reasoning at the right level, and change direction without discarding everything downstream.
A result can arrive instantly and still strand the user. They may need to reverse-engineer what the system assumed, identify which part is editable, and predict what a revision will disturb. That is waiting too, even if no loading indicator appears. It is cognitive waiting: time spent turning an opaque artifact into something the user can actually decide on.
The useful unit of speed, then, is not output per minute. It is progress toward a sound decision. Sometimes immediate generation helps. Sometimes a visible intermediate artifact helps more. A compact plan that exposes its assumptions may feel slower than an instant finished page, yet lead to a confident choice sooner because the person can intervene before the consequences spread.
An approval button is not control
The obvious answer to an AI review problem is to add an approval step. The system proposes; the person approves; the workflow continues. That creates a checkpoint, but a checkpoint alone does not create agency. If the person cannot see what shaped the proposal, reshape the plan, or identify the consequential choice, approval is mostly a request to accept or reject a bundle.
A binary decision is appropriate when the object itself is understandable and the choice is genuinely binary. It is weak when the output contains several coupled decisions. Rejecting a complete page does not say whether to preserve its structure, change its visual direction, replace an asset, or reconsider its premise. Asking for approval at that point can transfer accountability to the user without giving them a practical way to exercise judgment.
Control requires a legible object of review. The person needs to know what decision they are making now, what remains open later, and what will happen downstream. They also need an edit surface that matches the level of the problem. A directional error needs a way to change direction, not merely a button that produces another fully rendered attempt.
I use these questions as a review budget. The word budget matters because human attention is finite. A product should not spend it evenly across every model action. It should reserve attention for moments where judgment changes the direction, risk, or cost of what follows.
The first question gives the output a job. If nobody can name the decision it supports, the output may be activity rather than progress. The second names the reviewer. Expertise is contextual: the person who can judge tone may not be the person who can verify a factual claim or authorize an expensive action. The third exposes consequence. The fourth tests whether the workflow can safely learn by doing or needs to stop before it commits.
Three levels of consequence
Low-consequence drafts can usually run automatically. Early wording options, rough visual explorations, and disposable variations are useful partly because they are cheap to reject. Their purpose is to widen or clarify the space, not to become a commitment. Interrupting the user after every one would spend attention without buying meaningful safety. The product can generate them, organize them, and let the person review only when a pattern or candidate becomes relevant.
Directional plans need editable review. Product definition and visual direction shape many later outputs, so errors compound. This is where an intermediate artifact earns its place. The person should be able to inspect the plan, see the choices it contains, and alter the part that matters without rewriting the whole thing. Review here is not ceremonial approval. It is a chance to steer before generation turns a questionable direction into a coherent but costly stack of artifacts.
Factual or expensive outputs require explicit confirmation. If an output makes a claim that must be true, triggers a meaningful commitment, or would be costly to undo, silence should not count as consent. The product needs to surface the relevant evidence or uncertainty and ask the qualified person to confirm the decision. The confirmation belongs immediately before the consequence, while the choice is still clear and reversible.
Too much review is also a product failure
There is an uncomfortable trade-off here. Once I noticed review debt, it was tempting to make every intermediate step visible and ask for confirmation everywhere. That would be safe in a narrow sense. It would also recreate manual software: a sequence of forms, checkpoints, and small permissions that makes the user supervise the machine instead of benefit from it.
Too many checkpoints destroy momentum. They make the person repeatedly reload context, even when the next action is low risk and predictable. They also train approval into a reflex. When every screen asks for attention, none of the requests signals importance. The user clicks through the harmless steps and may carry the same habit into the consequential one.
The answer is selective control. Let reversible exploration flow. Pause when a choice sets direction, requires particular expertise, creates an external commitment, or becomes expensive to unwind. At those boundaries, show the smallest artifact that makes the decision legible and editable. Then let the system continue with the context the person has actually endorsed.
This also means the final artifact should not carry the entire burden of explanation. The product can preserve a lightweight trail of important decisions: what was chosen, what remained uncertain, and which later outputs depend on it. The goal is not an exhaustive transcript of model activity. That would produce another kind of review debt. The goal is enough structure for the person to understand how the current result became the current result.
A better definition of fast
I still want the product to feel fast. I still care about the blank screen, the dead moments, and the frustration of waiting for work that could have started sooner. But speed is no longer just a property of generation. It is a property of the whole path from intention to informed commitment.
That path can include automatic drafts, editable plans, and explicit confirmations without becoming slow by default. The difference is whether each review moment has a reason. If an output is easy to reverse, let it run. If it sets direction, make it shapeable. If being wrong carries real consequence, make the decision explicit. Spend the review budget where judgment has leverage.
The original goal was to help people see a result sooner. I have not abandoned it; I have made it more demanding. The goal is not maximum output speed. It is shorter time to a decision the user understands and owns. A product reaches that kind of speed not when it generates before the person can think, but when its generation and their judgment move together.
