Free reading link: Here
The task was done and my work had just started
You ever ask an AI agent for a small change, watch it finish, and open the diff wondering whether you accidentally commissioned a startup?
The summary sounds lovely.
Feature implemented. Tests added. Code cleaned up.
Wait. Cleaned up?
Which code?
Consider a simple example: adding a CSV export button to an admin page. There's already a table. There's already a way to fetch its data. You want users to download what they're allowed to see.
The agent gets to work. When it finishes, the button exists.
So does a new export utility, an extra dependency, a rewritten data-fetching function, and a configuration change you weren't expecting.
Some of that might be justified. You haven't checked yet.
Congratulations. That's your afternoon's new activity.
Every extra change arrives with a question attached.
Does the export respect the same permissions as the table? Does it include all matching records or only the current page? Why did the fetching logic change? Does anything else depend on its previous behavior?
And somewhere between those questions, you discover that the agent added a reusable export framework.
For your one export button.
Apparently, we're expecting a cinematic universe.
The awkward part is that the agent may have done useful work. The feature might be close to finished. It may have handled tedious formatting and test setup you were happy to delegate.
But the handoff still requires attention. You have to separate necessary changes from enthusiastic extras, verify the behavior, and understand enough of the result to maintain it.
The completion message marks a moment in the agent's workflow. Your team still has to decide whether the change belongs in the product.
That distinction gets easy to overlook when files are appearing faster than you can read them.
A small task can quietly become an architectural review. A passing test suite can leave the actual user journey unchecked. A convincing summary can send you searching through files to discover what "improved error handling" involved.
So before we celebrate how quickly the agent finished, there's another question worth asking:
How much work did it leave for the person clicking merge?
The diff got bigger than the problem
A CSV button is easy to describe.
Reviewing everything that appeared around it takes a little longer.
You open the changed files expecting to follow a simple path: button, request, export, download. Instead, you're trying to work out why a shared utility moved and whether the new abstraction replaced something that already existed.
The feature has arrived with luggage.
This is where counting finished tasks can hide the work still sitting on someone's desk. The agent has produced a solution, but the reviewer has to reconstruct the decisions behind it.
Why this dependency? Why change a function used by three other pages?
Some answers may be perfectly reasonable. A proper CSV library, for example, can handle quoting and embedded line breaks that a quick join(",") gets wrong. Rejecting it just because it adds a dependency would be lazy reviewing too.
The question is whether each change earns its place in this task.
A new library chosen for a real formatting requirement has a reason. Renaming an unrelated helper because its old name felt awkward needs a separate conversation.
Put them in the same patch, and the reviewer has to untangle them before judging either.
Think about two possible handoffs for our illustrative export feature.
One reuses the existing authorized query, adds the export behavior, and checks the important cases. The review has a clear route: follow the requested behavior through the changes.
The other also reorganizes the query layer, introduces a general export interface, and adjusts error handling across the admin area.
Now the reviewer has several jobs. Check the export. Understand the new interface. Verify that the query refactor preserved existing behavior. Find out whether another page depended on the old errors.
One ticket. Multiple boss fights.

The second patch could still be the right solution. Maybe the existing query cannot safely support an export. Maybe the shared behavior genuinely needs fixing first.
But that dependency needs to be visible. "I also cleaned this up" doesn't explain why the feature required it.
And file count alone won't settle the question. A careful change spread across several small files can be easier to review than one dense function hiding several behavioral changes.
What matters is how many decisions the reviewer must understand at once.
That's why a useful first pass is simply to group the changes:
Which ones implement the request? Which ones make it possible? Which ones are optional improvements?
The optional work might deserve its own ticket. It might even be excellent work. Separating it gives the team a chance to evaluate it without holding the original feature hostage.
Because once a small request becomes a mixed bag of improvements, "Does this work?" stops being one question.
It becomes a reading assignment.
Every extra file needs an explanation
Open a file you wrote yesterday and you'll probably remember something beyond the code.
Why you avoided that helper. Which approach failed. Why the slightly ugly branch exists.
The diff doesn't contain all of that. Some of it is still in your head.
With agent-generated work, you may encounter the decisions for the first time during review. The implementation is sitting there, fully dressed, while you're still trying to learn its name.
And the final summary can make that introduction suspiciously brief:
"Added robust export handling."
Lovely. What does robust mean here?
For our CSV example, it could mean properly escaping quotes. It could mean exporting thousands of records without loading everything into memory. Or it could mean catching every exception and returning an empty file.
Three very different afternoons.
A useful handoff connects each important decision to a requirement and something you can check.
Compare these two descriptions:
"Updated the data-fetching logic to support exports."
"The table fetches one page at a time. The export needs all matching records, so the export path retrieves additional pages while preserving the existing filters and permission checks."
The second gives the reviewer somewhere to start. You can inspect the pagination, follow the filters, and check where permissions are enforced.
It also exposes a product decision: should this button export every matching record?
Maybe the user expected only the selected rows.
That's a question worth answering before somebody downloads half the database because the button looked friendly.
For unfamiliar changes, three questions do a lot of work:
- What requirement made this necessary?
- What existing behavior could it affect?
- What evidence supports the claimed behavior?
You don't need an essay attached to every import. Focus on decisions with consequences: new dependencies, shared behavior, permissions, stored data, and anything that changes how the feature fails.
Ask the agent for file references and the checks it actually ran. "Tests pass" is much more useful when you know which tests, what they cover, and what remains unchecked.
Then inspect the relevant code. A detailed explanation can help you navigate the patch; it can also confidently describe something the implementation doesn't quite do.
We covered that particular jump scare in the previous article.
There's a maintenance reason for doing this too. Once the patch lands, the team inherits its choices. Someone will eventually need to upgrade the dependency, change the export format, or investigate why a customer's download is incomplete.
That person needs enough context to make the next decision.
Without it, you open another chat and ask an agent to reverse-engineer the last agent's work.
The handoff now has a sequel.
Good delegation leaves a trail you can follow: what changed, why it changed, how it was checked, and where uncertainty remains.
That trail is part of finishing the task.

"While I'm here" is an expensive sentence
There's a particular kind of helpfulness that makes a small task grow teeth.
The agent notices duplicated logic. An awkward name. An old dependency. A helper that could be more general.
All reasonable things to notice.
Then noticing becomes editing, and suddenly your export button is sharing a pull request with a neighborhood renovation.
Developers do this too. We open a file to fix one thing, spot something annoying, and think, "Might as well."
An agent can follow that impulse through several files before you've finished reading its progress update.
The cost shows up when you try to separate the changes.
Imagine a different task: updating a service's health check. While working on the configuration, an agent also refreshes a base image and reorganizes some environment settings.
The health check may be correct. The image update may be sensible. The configuration may look cleaner.
But if the service starts failing after deployment, you now have several plausible causes to investigate. Rolling back the whole patch also removes the health-check improvement you actually needed.
Bundling changes together can make them harder to diagnose and harder to undo independently.
That matters even when every change looked harmless during review.
A rename can obscure the behavioral edit beside it. Formatting can bury the few lines that deserve attention. A dependency update can introduce questions unrelated to the original request.
The patch gets noisier, and the reviewer spends more effort finding the important parts.
None of this means an agent should ignore a problem it discovers. Sometimes the requested feature really does depend on fixing something underneath it.
If our export path bypasses an existing permission check, that belongs in the conversation immediately. Leaving it untouched to keep the diff pretty would be absurd.
The useful distinction is whether the extra work is necessary for the requested behavior.
A prerequisite needs an explanation. An optional improvement can become a separate task.
You can make that boundary explicit in the brief:
Implement the requested behavior using existing project patterns. Keep unrelated refactors and dependency upgrades out of this change. If a broader change is necessary, explain the dependency before expanding the scope. List optional improvements separately.
That instruction won't enforce itself. You still need to inspect the result. But it gives both you and the agent a clearer basis for deciding what belongs.
It also gives useful observations somewhere to go without turning all of them into code.
"Found duplicated formatting logic; left it unchanged and noted it for follow-up" is a perfectly respectable outcome.
Every discovered imperfection does not need to become today's side quest.
You wanted an export button. You're allowed to finish with an export button.
Small tasks make better handoffs
Keeping unrelated work out helps. But the task itself still needs a finish line.
"Add CSV export" leaves several decisions open.
- Export which records?
- Which columns?
- Should filters apply?
- What happens when there's nothing to download?
If those questions stay unanswered, the agent still has to build something. Its assumptions become your review queue.
You can reduce that queue before any code gets written.
For our illustrative admin page, a clearer brief might look like this:
Add CSV export to the orders table.
Expected behavior:
- Export all records matching the current filters,
including records beyond the visible page.
- Include only the columns already shown in the table.
- Preserve the existing server-side access restrictions.
- Disable export when there are no matching records.
Scope:
- Reuse existing query and CSV utilities where suitable.
- Keep unrelated refactors out of this change.
- Explain any necessary dependency addition before making it.
Handoff:
- Summarize the behavior changed.
- List the checks actually run and their results.
- State anything that remains unverified.These are example requirements, not universal rules for export buttons. Another product might deliberately export selected rows or allow empty files.
The point is to make those choices visible while they're still cheap to change.
A good task gives the reviewer something concrete to accept or reject.
You can check whether filters survive the export. You can try records beyond the first page. You can verify that a user cannot retrieve another user's data.
"Make it robust" is considerably harder to test.
Robust against what? Commas? Network failures? Someone uploading the complete works of Shakespeare into the customer-notes field?
Be specific about the behavior that matters.
For a larger task, ask the agent to inspect the relevant code and explain its proposed approach before editing. That gives you a chance to catch a misunderstanding while it's still a paragraph.
You might discover that an export utility already exists. Or that fetching every record through the browser would be a poor fit for the dataset. Those findings can change the plan without leaving a pile of code to unwind.
Once implementation starts, divide substantial work into coherent steps you can verify. Perhaps the server-side export behavior comes first, followed by the interface.
Each step should have a meaningful result. Splitting a change into twenty tiny commits doesn't help much if none makes sense without the other nineteen.
And don't turn every trivial edit into a planning ceremony. Changing a label should not require a council of agents and three architectural diagrams.
The amount of coordination should match the uncertainty.
A useful question is: Can I describe what successful completion looks like, and can I realistically review the result?
If either answer is no, clarify the task or divide it before handing it over.
That's how delegation starts saving attention as well as keystrokes. The agent can move quickly inside a clear boundary, and you can review the result without first discovering what job it decided to do.
Conclusion: done should be easy to verify
Go back to that CSV button.
A useful result lets you follow the change from the user's click to the downloaded file. You can see which records it includes, how access restrictions apply, and what was checked.
Any necessary dependency has a reason. Optional cleanup has its own place. The handoff tells you where uncertainty remains.
There's still review work. But you can begin it without first excavating the task from everything built around it.
That's a much better place to judge whether the agent helped.
The time between sending a prompt and receiving "Done" only covers part of the job. There's also the time spent understanding the patch, correcting mistakes, resolving review questions, and getting the change ready to ship.
If the agent handles tedious implementation and leaves a clear, focused result, that can be a real win. If it produces an impressive pile of code that takes hours to untangle, the completion message arrived before much of the work did.
The handoff belongs in the definition of finished.
For your next task, describe the expected behavior, set a sensible boundary, and ask for the evidence you'll need to review it. Then compare the actual patch with that brief.
Keep the useful discoveries. Give unrelated improvements their own space. Ask questions when a necessary change expands the work.
You're allowed to expect assistance that makes the next step easier.
And if you asked for an export button, it's perfectly reasonable to wonder why you've been handed a framework with a roadmap.
Helpful resources & links
- Google: Small changes How to keep patches focused and easier to review. Useful when your "small fix" starts developing its own architecture department.
- Google: What to look for in a code review A practical guide to reviewing design, functionality, complexity, and tests. Gives you better questions than "Does the code look tidy?"
- Google: Writing good change descriptions Explain what changed and why. Helpful for turning "improved export handling" into a handoff someone can actually understand.
- Anthropic: Claude Code best practices Guidance on clear task context, exploring before implementation, and giving the agent checks it can run. Start with "Give Claude a way to verify its work" and "Explore first, then plan, then code."