← Back to work

Developer tool · v0.2.0

A Control System for Long-Running AI Development

An engineering harness for building medium and large software projects with AI agents without losing product intent, scope, or completion discipline.

The agent writes code. The harness preserves state, constrains the task, and decides what must happen before the work can ship.

Role
Creator and maintainer
Proof
  • 641 tracked Features
  • 2,960 project-history commits
Stack
  • Methodology
  • Markdown
  • Git

CONTROL LOOP / 01 FEATURE

Durable stateProduct · Feature · Rules
  1. 01Bound scope
  2. 02AI implements
  3. 03Machine verifies
  4. 04Human accepts
  5. 05Reviewer inspects

Accepted → Ship → Next Feature

01 / Problem

What problem was I trying to solve?

AI coding agents could already generate useful patches. The harder problem was keeping a medium or large project coherent after the first patch—across changing requirements, repeated sessions, reviews, and hundreds of delivery decisions.

I built Yaya Loop as the control system around the agent: a development harness that preserves intent, limits each unit of work, separates completion authority, and leaves a recoverable project state behind.

Yaya Loop overview showing durable Product, Feature, and Coding Rules state surrounding an AI implementation loop with automated checks, human acceptance, independent review, shipping, and the next Feature.
The agent is the executor; the surrounding harness controls memory, scope, verification, acceptance, review, and the next loop.
Read the diagram description

The diagram begins with durable project state: Product records what should exist, Feature records the bounded work still to do, and Coding Rules record how changes should fit the codebase. Product intent is narrowed through a scope gate into one Feature. An AI agent implements that Feature, then automated checks, human acceptance, and an independent fresh-context review evaluate different dimensions of completion. Failed checks or reviews return the same bounded Feature for correction. Only accepted work ships, updates durable state, and allows the next Feature to begin.

Yaya Loop harness overview

Yaya Loop overview showing durable Product, Feature, and Coding Rules state surrounding an AI implementation loop with automated checks, human acceptance, independent review, shipping, and the next Feature.

02 / Context

Why isn't a bigger prompt enough?

A large prompt can describe the product, the current task, coding rules, and completion criteria—but only inside one finite conversation. The context eventually ends while the repository, unfinished work, and earlier decisions remain.

I wanted the project itself to carry that memory. A new session should read a durable state, not depend on an increasingly fragile summary of an old chat.

The conversation is temporary. The project state must survive it.

03 / Reliability

Why does AI coding become unstable in large projects?

The agent is useful but variable. Over many iterations, requirements can drift, one task can quietly expand, local changes can conflict with system-wide decisions, and old constraints can disappear from working context.

The open-ended engineering work was combining ideas from task decomposition, repository-owned progress, coding rules, automated gates, human judgment, independent review, research, and repeated practice into one flow.

Reliability has to come from the system around the executor—not from assuming every answer will be reliable.

04 / Memory

What state needs to survive between AI sessions?

Product documents preserve what should exist. The Feature ledger records bounded work, dependencies, acceptance criteria, and status. Coding Rules preserve how changes should fit the codebase.

Progress notes and Git history add the operational trail: what has happened, what evidence exists, and what the next session should inspect. Chat history is useful context, but it is not the project database.

Product is What. Feature is Todo. Coding Rules are How.

Repository-owned memory

Three layers survive the chat window

  1. 01
    ProductWhat

    Versioned product intent, behavior, and constraints.

  2. 02
    FeatureTodo

    One bounded delivery unit with dependencies, acceptance, and status.

  3. 03
    Coding RulesHow

    Project and language constraints that shape implementation.

A fresh session reads the same versioned state as the human maintainer.

05 / Scope

Why is Feature the unit of work?

A Product is too broad to authorize as one task, while a loose instruction often has no stable boundary. A Feature is small enough to scope and review, but large enough to describe one meaningful product change.

Each Feature carries its own dependencies, acceptance criteria, status, and evidence. That makes it an independently verifiable delivery unit—and a clear recovery boundary when a run is interrupted.

One bounded delivery loop

Intent moves forward; failed gates return to implementation

  1. 01

    Method

    PreflightConfirm branch, state, dependencies, and one selected Feature.
  2. 02

    Method

    Resource checkLoad product intent, rules, and relevant project evidence.
  3. 03

    Agent

    StartRecord in-progress state before changing code.
  4. 04

    Agent

    ImplementChange only the approved scope and self-check the diff.
  5. 05

    Machine

    Automated verificationRun type, build, tests, and targeted browser contracts.Fail → Implement
  6. 06

    Human

    Human acceptanceDecide whether the product behavior is right.Fail → Implement
  7. 07

    Reviewer

    Fresh-context reviewInspect structure and code smells without implementation context.Fail → Implement
  8. 08

    Method

    DoneRecord completion only after all gates pass.
  9. 09

    Method

    HandoffPreserve evidence and identify the next safe Feature.

06 / Control

How do I prevent an agent from modifying too much at once?

Before implementation, the loop confirms one active Feature, its dependencies, affected resources, acceptance criteria, and explicit non-goals. The implementation is reviewed against that written boundary rather than against the agent's interpretation alone.

Automated gates and human review can return an over-broad change to the same Feature for correction. This does not guarantee that an agent never crosses the boundary; it makes the boundary visible and enforceable.

07 / Verification

Why aren't tests enough?

Tests answer executable questions: does it type-check, build, preserve known behavior, and satisfy targeted contracts? They cannot decide whether the requested product experience is actually right or whether the change leaves an unhealthy structure behind.

Yaya Loop treats machine verification, product acceptance, and fresh-context engineering review as different gates. A green test run is evidence, not permission to mark the Feature complete.

Three different questions

Completion is a conjunction, not an agent opinion

Machine

Does it execute correctly?

Types, build, tests, and targeted contracts.

Human

Is the product behavior right?

Acceptance against the intended experience.

Reviewer

Did the change leave structural debt?

A fresh-context smell and maintainability scan.

Feature may be recorded as done

08 / Authority

Where does human review enter the loop?

The human confirms the Feature and its acceptance criteria before implementation, then evaluates the resulting product behavior after machine checks pass. Public facts and subjective product choices remain human decisions.

The implementation agent cannot promote its own work to done. If acceptance fails, the same bounded Feature returns for correction; only explicit human approval allows the completion workflow to continue.

09 / Recovery

How does the workflow recover after context loss?

A new session does not attempt to recreate the old model context. It rereads the versioned Product, Feature ledger, Coding Rules, current Progress, and relevant Git history.

Those sources reconstruct the current intent, active scope, completed evidence, unresolved decisions, and next safe action. Recovery is a documented project operation—not a claim about the model remembering a previous conversation.

Repository-driven recovery

A new session reconstructs state instead of replaying a chat

Context boundaryPrevious chat context ends

Recovery actorNew AI session

Read versioned project state

  • Product
  • Feature ledger
  • Coding Rules
  • Progress
  • Git history

Reconstructed stateIntent · active scope · evidence · next safe action

Continue deliberatelyContinue the current Feature or take the next safe action

The recovery path is explicit and inspectable; it does not depend on hidden model memory.

10 / Evidence

Did I actually use this workflow on a real project?

QuackPack became the main stress test: a game project developed through hundreds of bounded Features and a long, multi-branch repository history.

The dated local snapshot below records the exact Feature, commit, code, test, and product-document counts. They show sustained use and project size; they do not prove quality, speed, or commercial success.

Local Git snapshot · 2026-08-27

One harness, sustained project history

641
tracked Features · 616 done
2,960
deduplicated commits across local refs
52,688
tracked core game TypeScript lines
18
Product documents · 5,797 lines

Exact snapshot: 641 Feature records (616 done), 2,960 deduplicated commits across local refs, 52,688 tracked core TypeScript lines, 57,154 tracked root test lines, and 641 Feature detail files. Scale is evidence of sustained use, not a quality or productivity claim.

11 / Limits

What still doesn't work well?

The harness cannot rescue a poorly defined Feature or a weak acceptance test. It still depends on human judgment, disciplined document maintenance, useful verification, and the willingness to stop an Agent that is heading outside the intended boundary.

Different agents behave differently, and the process reduces that variance rather than eliminating it. The ceremony also has a real cost, so I would not use the full workflow for a small script or a short-lived change.

The harness makes important boundaries explicit. It does not make the executor—or the specification—perfect.