Developer tool · v0.2.0
A Control System for Long-Running AI Development
An engineering harness for building medium and large software projects with AI agents without losing product intent, scope, or completion discipline.
The agent writes code. The harness preserves state, constrains the task, and decides what must happen before the work can ship.
CONTROL LOOP / 01 FEATURE
- 01Bound scope
- 02AI implements
- 03Machine verifies
- 04Human accepts
- 05Reviewer inspects
Accepted → Ship → Next Feature
01 / Problem
What problem was I trying to solve?
AI coding agents could already generate useful patches. The harder problem was keeping a medium or large project coherent after the first patch—across changing requirements, repeated sessions, reviews, and hundreds of delivery decisions.
I built Yaya Loop as the control system around the agent: a development harness that preserves intent, limits each unit of work, separates completion authority, and leaves a recoverable project state behind.

Read the diagram description
The diagram begins with durable project state: Product records what should exist, Feature records the bounded work still to do, and Coding Rules record how changes should fit the codebase. Product intent is narrowed through a scope gate into one Feature. An AI agent implements that Feature, then automated checks, human acceptance, and an independent fresh-context review evaluate different dimensions of completion. Failed checks or reviews return the same bounded Feature for correction. Only accepted work ships, updates durable state, and allows the next Feature to begin.
02 / Context
Why isn't a bigger prompt enough?
A large prompt can describe the product, the current task, coding rules, and completion criteria—but only inside one finite conversation. The context eventually ends while the repository, unfinished work, and earlier decisions remain.
I wanted the project itself to carry that memory. A new session should read a durable state, not depend on an increasingly fragile summary of an old chat.
The conversation is temporary. The project state must survive it.
03 / Reliability
Why does AI coding become unstable in large projects?
The agent is useful but variable. Over many iterations, requirements can drift, one task can quietly expand, local changes can conflict with system-wide decisions, and old constraints can disappear from working context.
The open-ended engineering work was combining ideas from task decomposition, repository-owned progress, coding rules, automated gates, human judgment, independent review, research, and repeated practice into one flow.
Reliability has to come from the system around the executor—not from assuming every answer will be reliable.
04 / Memory
What state needs to survive between AI sessions?
Product documents preserve what should exist. The Feature ledger records bounded work, dependencies, acceptance criteria, and status. Coding Rules preserve how changes should fit the codebase.
Progress notes and Git history add the operational trail: what has happened, what evidence exists, and what the next session should inspect. Chat history is useful context, but it is not the project database.
Product is What. Feature is Todo. Coding Rules are How.
Repository-owned memory
Three layers survive the chat window
- 01ProductWhat
Versioned product intent, behavior, and constraints.
- 02FeatureTodo
One bounded delivery unit with dependencies, acceptance, and status.
- 03Coding RulesHow
Project and language constraints that shape implementation.
A fresh session reads the same versioned state as the human maintainer.
05 / Scope
Why is Feature the unit of work?
A Product is too broad to authorize as one task, while a loose instruction often has no stable boundary. A Feature is small enough to scope and review, but large enough to describe one meaningful product change.
Each Feature carries its own dependencies, acceptance criteria, status, and evidence. That makes it an independently verifiable delivery unit—and a clear recovery boundary when a run is interrupted.
One bounded delivery loop
Intent moves forward; failed gates return to implementation
- 01
Method
PreflightConfirm branch, state, dependencies, and one selected Feature. - 02
Method
Resource checkLoad product intent, rules, and relevant project evidence. - 03
Agent
StartRecord in-progress state before changing code. - 04
Agent
ImplementChange only the approved scope and self-check the diff. - 05
Machine
Automated verificationRun type, build, tests, and targeted browser contracts.Fail → Implement - 06
Human
Human acceptanceDecide whether the product behavior is right.Fail → Implement - 07
Reviewer
Fresh-context reviewInspect structure and code smells without implementation context.Fail → Implement - 08
Method
DoneRecord completion only after all gates pass. - 09
Method
HandoffPreserve evidence and identify the next safe Feature.
06 / Control
How do I prevent an agent from modifying too much at once?
Before implementation, the loop confirms one active Feature, its dependencies, affected resources, acceptance criteria, and explicit non-goals. The implementation is reviewed against that written boundary rather than against the agent's interpretation alone.
Automated gates and human review can return an over-broad change to the same Feature for correction. This does not guarantee that an agent never crosses the boundary; it makes the boundary visible and enforceable.
07 / Verification
Why aren't tests enough?
Tests answer executable questions: does it type-check, build, preserve known behavior, and satisfy targeted contracts? They cannot decide whether the requested product experience is actually right or whether the change leaves an unhealthy structure behind.
Yaya Loop treats machine verification, product acceptance, and fresh-context engineering review as different gates. A green test run is evidence, not permission to mark the Feature complete.
Three different questions
Completion is a conjunction, not an agent opinion
08 / Authority
Where does human review enter the loop?
The human confirms the Feature and its acceptance criteria before implementation, then evaluates the resulting product behavior after machine checks pass. Public facts and subjective product choices remain human decisions.
The implementation agent cannot promote its own work to done. If acceptance fails, the same bounded Feature returns for correction; only explicit human approval allows the completion workflow to continue.
09 / Recovery
How does the workflow recover after context loss?
A new session does not attempt to recreate the old model context. It rereads the versioned Product, Feature ledger, Coding Rules, current Progress, and relevant Git history.
Those sources reconstruct the current intent, active scope, completed evidence, unresolved decisions, and next safe action. Recovery is a documented project operation—not a claim about the model remembering a previous conversation.
Repository-driven recovery
A new session reconstructs state instead of replaying a chat
Context boundaryPrevious chat context ends
Recovery actorNew AI session
Read versioned project state
- Product
- Feature ledger
- Coding Rules
- Progress
- Git history
Reconstructed stateIntent · active scope · evidence · next safe action
Continue deliberatelyContinue the current Feature or take the next safe action
The recovery path is explicit and inspectable; it does not depend on hidden model memory.
10 / Evidence
Did I actually use this workflow on a real project?
QuackPack became the main stress test: a game project developed through hundreds of bounded Features and a long, multi-branch repository history.
The dated local snapshot below records the exact Feature, commit, code, test, and product-document counts. They show sustained use and project size; they do not prove quality, speed, or commercial success.
Local Git snapshot · 2026-08-27
One harness, sustained project history
- 641
- tracked Features · 616 done
- 2,960
- deduplicated commits across local refs
- 52,688
- tracked core game TypeScript lines
- 18
- Product documents · 5,797 lines
Exact snapshot: 641 Feature records (616 done), 2,960 deduplicated commits across local refs, 52,688 tracked core TypeScript lines, 57,154 tracked root test lines, and 641 Feature detail files. Scale is evidence of sustained use, not a quality or productivity claim.
11 / Limits
What still doesn't work well?
The harness cannot rescue a poorly defined Feature or a weak acceptance test. It still depends on human judgment, disciplined document maintenance, useful verification, and the willingness to stop an Agent that is heading outside the intended boundary.
Different agents behave differently, and the process reduces that variance rather than eliminating it. The ceremony also has a real cost, so I would not use the full workflow for a small script or a short-lived change.
The harness makes important boundaries explicit. It does not make the executor—or the specification—perfect.