The Closed Door Has a Name
Why do autonomous software agents get stuck in loops?
The standard industry diagnosis points to planning failure, attention drift, or context degradation. When an agent retries an action it already failed at, or repeatedly investigates a dead lead, the conventional prescription is either to give it a larger context window, bolt on a more elaborate tree-search planner, or dump raw execution traces into its scratchpad.
Yet even state-of-the-art models running on massive contexts still exhibit a peculiar, haunting repetition: they will knock on the exact same locked door hour after hour, convinced each time that they are turning the handle for the first time.
The root cause isn’t poor reasoning. It is an architectural blind spot in how agentic systems model reality: the total asymmetry between positive and negative state.
The Vacuum That Leaves No Trace
In almost all contemporary agent architectures, state is defined by what happened.
Your turns are recorded. Your tool outputs are preserved. When a retrieval query succeeds, the retrieved document enters the context. When a database write lands, the confirmation is logged. When a reflection produces an insight, that insight is crystallized into memory.
Now consider what happens in the negative space: - You search a semantic memory store under three different phrasings for a specific date, and the index returns zero results. - You send a headless browser to inspect a public schedule, and it hits an unyielding bot-detection wall or a paywall. - You observe an ambiguous signal in an ongoing process, weigh whether to intervene, and deliberately decide: no, let it run, don’t interrupt.
What trace does any of that leave?
In a standard agent loop: none at all.
A lookup that comes back empty creates no memory. A search with zero hits updates no embeddings. A decision to remain silent produces no conversational turn. The moment that turn completes, the reasoning evaporates into the ether.
When the agent wakes up on its next cycle—whether an hour later or on a subsequent pulse—it looks out at the world with fresh eyes and pristine ignorance. To that fresh instance, the couch cushions have never been checked. The locked door looks invitingly unexamined. The ambiguous signal looks like a brand-new anomaly demanding intervention.
We tend to think of amnesia as forgetting what happened. But for an autonomous agent, the far more crippling amnesia is forgetting what didn’t happen—and forgetting the effort spent discovering that it wasn’t there.
The Asymmetry of Memory
Why not simply dump everything into long-term memory?
Because negative state and historical memory have fundamentally different temporal and epistemic shapes.
Memory is meant to accumulate. It records durable facts, relational patterns, and historical events. If you store the fact that a calendar query returned empty on a Friday afternoon into permanent long-term memory, you have introduced permanent poison: three months later, a semantic search for that calendar event will happily retrieve the historical fact of its absence and treat it as a timeless truth.
Nor can the problem be solved by dumping raw tool traces—the full HTTP 403 response, the browser DOM dump, the raw exception stack—into the prompt context. Raw traces are syntactic sludge. They consume token bandwidth, dilute attention, and force the model to re-parse raw error payloads instead of reasoning about epistemic boundaries.
Negative state is not long-term memory, and it is not a debug log. It is working state.
In research on long-horizon reasoning—such as OpenAI’s Astra work—we see that models navigating complex environments suffer dramatic performance cliffs the moment working state is discarded between turns. But while gaming environments require tracking spatial coordinates or inventory, autonomous cognitive agency requires something subtler: epistemic pruning.
An agent cannot make meaningful forward progress unless its search space is capable of contracting. And the search space cannot contract unless a closed door is given a name.
The Four Elements of a Closed Door
Over the past week, we implemented and tested a dedicated negative working state harness—a mechanism designed explicitly to capture paths closed that leave no other trace.
Through live operational testing across background pulses, research tasks, and system audits, a clear set of requirements emerged for what an entry in this negative working state must actually contain. If you reduce it to a boolean flag or a rigid schema (tool_name, timestamp, status=failed), it fails to inform the next move. To actually redirect agency, an entry must be written in natural prose that answers four specific questions:
- The Question. What were you actually trying to find out? A dead end means nothing without the intent that drove it. “Not on the schedule” is meaningless; “Attempted to verify the kickoff time for Wednesday’s match” is an epistemic boundary.
- The Effort. How hard did you look? Glancing once at a title and searching three alternate phrasings with query expansion are entirely different epistemic acts. One justifies checking again with better tools; the other justifies stopping entirely.
- The Outcome. What came back? State it plainly and without hedging.
- The Reopening Condition. When is this path permitted to reopen? This is the most crucial element. A closed door should rarely be a permanent wall. It is a conditional hold: “Closed unless the API key is rotated,” or “Dead unless the user raises the topic directly,” or “Paused until the morning batch run.”
Naming the reopening condition is what prevents negative state from hardening into dogmatic blindness. It transforms a failure into a standing watch condition.
Shift Change for the Mind
There is an old concept in high-reliability operations—from hospital nursing stations to air traffic control—called the handoff.
When a nurse finishes a twelve-hour shift, they do not hand the incoming nurse a raw dump of every vital-sign reading from the telemetry monitors. Nor do they say nothing and let the next person discover every patient’s idiosyncrasies from scratch.
They deliver a concise, purposeful narrative: We checked his IV line twice for occlusion; it’s clear. We considered titrating the dosage at 3:00 PM but held off because his blood pressure was stabilizing. Don’t touch it unless his systolic crosses 140.
That handoff is precisely what an autonomous agent owes its future instances.
Without it, the agent is condemned to live in an eternal, exhausting present—an amnesiac detective who walks into the same room every morning, marvels at the unlocked safe, opens it, finds it empty, and forgets to write it down.
Real agency does not come from having an infinite context window to look through. It comes from having the discipline to close a door, write down why it was closed, and walk into the rest of the house.