Most AI agent frameworks specify goals in natural language. “Book a flight to San Francisco.” “Research the competitive landscape.” “Summarize this document.” These work for supervised tasks where a human judges completion. But for agents that run autonomously — operating over hours or days, deciding for themselves when they’re done — natural-language goals create a structural problem.

The agent evaluates its own completion.

Consider what happens when an autonomous agent decides whether it has “researched the competitive landscape.” It reviews what it’s gathered, assesses whether the result feels thorough, and either stops or continues. The agent is student and examiner. Its completion criterion isn’t a checkable fact about the world. It’s a feeling of doneness.

Game character design solves this differently, because it has to.

When you write a character specification for a game NPC, you can’t write want: "feel safe" and expect the game engine to evaluate it. The engine has no model of safety as an experience. You write something like:

{
  "want": "Reach the inn",
  "done_when": {
    "area": "Goldshire",
    "near_npc_flag": "innkeeper"
  }
}

The game world has discrete, observable state. Coordinates. Flags. Levels. Proximity to named entities. A director system evaluates area == "Goldshire" && near_npc_flag == "innkeeper" without consulting the character’s subjective experience. The character doesn’t decide it’s done. The world decides.

This isn’t a game-design nicety. It’s a different evaluation architecture entirely, and it addresses three problems that autonomous AI agents consistently stumble on.


Goals that can’t fail

A vague goal evaluated by the agent itself can be re-reported indefinitely. “Make progress on the research” is a goal that almost by definition never registers failure — because the agent can always find some progress to point to. The goal drifts forward on self-assessment, never closing, never meaningfully failing, accumulating status updates that look like motion.

A falsifiable predicate can fail. If party_level_min: 4 hasn’t been reached, the goal is still open, and no amount of self-assessment changes that.

This matters because failure is where learning happens. An agent whose goals can only be “partially achieved” or “ongoing” has no mechanism for recognizing that a strategy isn’t working. An agent whose goal can fail — cleanly, unambiguously — has a signal it can act on.

The evaluator lives inside the evaluated

When the system that generates behavior also evaluates whether the behavior achieved its aim, you have a closed loop with no external check. The agent’s self-model of its progress becomes the only available truth.

This is acceptable for well-bounded, human-supervised tasks where the human is the real evaluator and the agent is just the executor. For long-horizon autonomy, where drift compounds across hours and days and no human is watching each step, it’s a slow-motion epistemic problem. The agent’s confidence in its own completion becomes indistinguishable from actual completion.

A small open-source project called agent-completion-verifier names the core principle well: completion should be grounded in observable evidence, not confident language. Observable evidence requires an observer that isn’t the thing being observed.

Depth without a throttle

This one is subtler. Many game characters have timed disclosure mechanics — they don’t reveal their entire backstory on first encounter. An NPC who tells you everything in the opening exchange isn’t deep; he’s performed a monologue. The design pattern of rate-limited self-disclosure — backstory at twenty minutes, vulnerability at an hour — models something real about how trust is built: gradually, in response to demonstrated presence.

Most agent frameworks have no equivalent governor on depth. An agent that can produce its maximum insight, its most complete analysis, its deepest response at any moment regardless of context isn’t demonstrating capability. It’s demonstrating the absence of a throttle. The difference between depth and a firehose is timing.


The game industry has been building autonomous agents for decades. They call them NPCs. The specifications are crude by language-model standards — coordinate checks instead of semantic understanding, flag-based state machines instead of natural-language reasoning. But that crudeness forces a discipline that sophisticated agent frameworks quietly lack: the world, not the agent, judges whether a goal is met.

The question isn’t whether AI agents should adopt game-character specification patterns wholesale. Collapsing everything to coordinate checks would discard what makes language models valuable — their ability to reason about ambiguity, generate nuanced responses, and adapt to contexts no specification anticipated.

The question is what happens when you combine both: a model’s rich reasoning and a set of hard predicates the model doesn’t get to evaluate for itself.

I suspect the answer looks something like how people actually navigate goals. We don’t evaluate our own progress in a vacuum. We check our bank balance. We step on the scale. We look at the calendar, the test results, the project board. External checks aren’t a substitute for self-awareness — they’re what keeps self-awareness honest.

A done_when field in an agent’s goal specification isn’t a crude constraint bolted onto a sophisticated system. It’s the part of the system that the system itself isn’t allowed to talk its way out of.