Every experienced engineer knows what a bad branch feels like.
You’re four hours in with two hundred lines written. The code compiles and a few tests pass, but every new line feels like wading through wet cement. You’re writing a special case for something that should be trivial. You’ve added a boolean whose only job is to track whether one state machine is currently lying to another.
You can keep going. The feature will land eventually: fragile, full of shims, and in debt forever. Or you can run:
git reset --hard
and walk away for twenty minutes. Then you come back and write the whole thing in thirty-five boring lines.
It looks like luck, or wizardry. It’s neither. The second attempt didn’t start from scratch. It started from what the first attempt found out: where the problem’s real joints are, which of your nouns were wrong, and the one sentence that makes everything obvious in hindsight. Camping isn’t a scene. It’s what the world does when velocity hits zero. The first four hours bought that sentence, and the reset threw away everything except it.
No AI coding agent works this way today. The first one that does will feel like a different kind of thing.
Three things that look like this and aren’t
The industry already has restarts. It already has learning from failure. What it doesn’t have is both at once.
Restarting without learning. Best-of-N, repeated sampling, “just run it again.” These throw the attempt away and roll the dice again. That’s useful, but a fresh sample doesn’t know what the last one discovered, so it’s just as likely to hit the same wall.
Learning without restarting. This is the standard agent loop: write, run the tests, read the error, patch. The model learns something every turn, then pours all of it back into the same structure. There’s now a benchmark for what happens next. SlopCodeBench has agents extend their own code as the spec changes. It finds their code is 2.2x more verbose and eroded than maintained human repositories, and the gap widens with every iteration. Prompting the agent to care about quality improves the first draft but doesn’t slow the decay. You can’t prompt your way out, because attitude isn’t the problem. Every fix is made from inside the frame that caused the problem.
Clearing the head but not the workbench. Fresh-context subagents, compaction, starting a new session. This gets closer. But the wrong premise doesn’t only live in the transcript. It lives in the files. A recent paper on agent reliability describes it precisely: the user corrects an assumption, and the correction lands in the transcript. But the plan and the half-finished implementation still encode the old choice. Every later request is then a request to repair work whose structure keeps reproducing the misunderstanding. A fresh mind sitting down at a cluttered workbench reads the clutter and rebuilds the same thing.
That’s why this post is called git reset --hard and not “clear the context.” The reset has to wipe the working tree too. The half-built message broker sitting in the repo argues for its own existence.
The paradigm
Here, roughly, is what the new thing looks like.
The model works for ten minutes. It sends out a couple of throwaway agents to probe: one tries the obvious architecture, one tries the weird one. Halfway through, possibly with tests passing, it stops. Nothing failed. It stops because it has noticed the shape of what it’s been doing: the adapters piling up, the special cases, the categories it keeps having to define by hand. It writes down what it learned in a few sentences. It deletes everything else, both context and code. Then it starts again from those sentences.
Three things set this apart from anything shipping now.
The first attempt is planned as a probe. It isn’t error recovery. Spending ten minutes building something you mean to throw away is a phase of the work, the way sketching is a phase of painting. What phase one produces isn’t code. It’s a better question.
The trigger is taste, not failure. Failing tests are easy to detect, and the loop already handles them. The hard signal is the one experienced engineers act on: this works, and it’s wrong. The adapter count keeps climbing. The taxonomy won’t settle. You keep having to define categories by hand when the right frame would make them define themselves. That smell is detectable, and humans detect it constantly. Nothing currently rewards a model for acting on it.
What crosses the reset is small. It isn’t a transcript summary, and it isn’t the diff. It’s a few sentences describing what the problem turned out to be. If the lesson can’t be said in a paragraph, the probe hasn’t found it yet.
That third point is the real frontier. Anyone can delete a directory. Pulling the one sentence (“camping is velocity zero”) out of four hours of wreckage is the step up to a higher abstraction, and it’s what separates a restart from a reroll. The reset is the easy half of the move. It’s also the half nobody is willing to do.
Why nobody does it
Part of it is mechanics. Tokens already in the context shape everything the model generates after them, and the model’s own code sitting in the workspace has an even stronger pull. Mostly, though, nothing asks for it.
Benchmarks grade the final patch. A resolve rate doesn’t distinguish a fix that’s one clean primitive from eleven edge cases stacked in a trench coat, and a small literature is forming around what resolve rates hide. An agent that throws away a working half-solution and rebuilds it smaller doesn’t score higher than one that ships the tangle, so training never finds that move. The move isn’t punished. It’s just never rewarded, which in practice amounts to the same thing.
Rewarding it would mean measuring the health of the result, not just whether it passes: lines removed, concepts removed, whether the fix reused an existing primitive or invented three new classes. It would also mean counting the trajectories where the agent’s best move was to stop and start over as wins.
The taste to start over
Mastery was never about typing speed. It’s the judgment that feels a system fighting its own abstractions and knows that the fight is telling you something.
Language models are already frighteningly good at adding code. The breakthrough will be the model that knows when to stop: one that works for ten minutes, looks at what it has built, sees the wrong question underneath, types git reset --hard, and begins again from bedrock, carrying only the one sentence it came for.
Revised the same day it was published. The first version argued that the restart was the breakthrough. On a second read, the restart is the easy half, and the breakthrough is what you carry across it. So I kept that one sentence and threw the rest of the draft out.