Thoughts on AI harnesses
"Why spend five minutes doing something by hand when you can spend five hours automating it?" is a common joke about engineers. With the advancement of AI, it is easier to automate things than ever before in our lifetime. Like DHH said, it is like having a genie in your pocket. You can ask it to do almost whatever you want and, within a day or so, it will usually give you a working prototype.
Now the question is: how do we automate the automation? How can we prompt AI automatically, give it the right context, and make it fix its own mistakes instead of requiring our monkey brains to steer it manually using English?
This problem was proposed to me by a Distinguished Engineer named MW when I lived with him for around two months during the convergence of Copilots at the beginning of 2026. We talked about it in a conference room, where he drew a big circle with all feature life cycles as chunk of pies and wrote ENGINEERING in all caps, like a real engineer. The point was that we still had to make the AI work like an engineer instead of simply asking it to fix something and hoping it would not make a mistake.
That sounds obvious, but it leads to a harder question: what does an engineer actually do?
What is software engineering development cycle?
This has a complicated answer because it differs by person. People like Casey Muratori or Jonathan Blow, who are savants of software engineering, will approach a problem differently from most people to a point where they will invent their own programming lanuage.
For many FAANG engineers, it is merely a cycle where people will LARP about the latest technology or tool they have found and explain how it should be incorporated into their gigantic monolotic stack which will takes a year to launch a single hover element in Google doc (my own experiences).
For most people, I like to think of engineering cycle as four steps:
- Define the goal.
- Implement it.
- Test it.
- Polish it.
The first two steps are so common that we usually represent them directly in Jira, Kanban, or whatever board is fashionable this quarter. You are given a ticket, which is the goal, and then you move it into an active state while implementing it. Steps three and four often become review and QA states, but an engineer usually does some version of them before handing the work to anyone else so it is somewhat part of implementation.
These last two steps are also where a surprising amount of programming time goes. A review can take hours or days. When you fix one comment, you might get a merge conflict, uncover another problem, or receive a second review that catches something the first one missed. At most companies, a large part of engineering productivity is consumed by review cycles and getting alignment with other people.
This is somewhat rightful. It is usually harder to safely revert shipped work than it is to implement the initial fix. The cost of a bad change is not the number of lines in the diff. It is the number of people, systems, and assumptions the change can disturb.
Copying the human process
The first instinct is to mimic this process using AI. This is delightful at first because there is no downtime and no need to ping another AI on Teams asking if it has had time to look at your PR. You can spin up another CLI process, give it a fresh reviewer context, and ask it to review the work immediately.
The problem is that it usually never stops.
The AI reviewer will nitpick, warn about states the program can never enter, and suggest increasingly defensive code. The implementation agent then obediently applies those suggestions. After a few rounds you have hundreds of try/catch blocks, hundreds of random tests that should not exist, and abstractions protecting you from imaginary futures. The code becomes larger without becoming safer. Sometimes the final code is worse than the first attempt.
But why?
We copied the visible ceremony of engineering, but not the judgment behind it. A human reviewer knows which concerns matter for this change, which risks are realistic, and when the code is good enough. An AI asked to "find more issues" has no natural reason to stop finding more issues.
Can I get some context, please?

The review process exists partly because we ask people who have more context about an area to look at the code. That might be a code owner who understands the subsystem, or simply a coworker with a fresh view of the first implementation. Either way, almost every PR needs context beyond what is written in the ticket.
That context is different for every PR. If a change touches file A and not file B, the context for file A should usually be preferred. If it touches UI components, the agent should know the UI conventions. If it touches the backend, it should know the service and data constraints. If it changes both, the relevant context may also depend on which phase of the work it is currently doing.
The tempting conclusion is to inject all available context into the entire process. If more context helps review, surely more context will also make implementation better while saving time and tokens later.
This is not necessarily true. You cannot combine implementation and polish into one giant prompt and expect a better result.
Imagine implementing a button that opens a modal containing values from a backend API. During the first implementation pass, the useful context is the API contract, the existing component patterns, and the expected end-to-end behavior. If you also front-load every accessibility rule, internationalization rule, analytics convention, loading-state variation, notification pattern, and modal lifecycle concern, the agent will try to solve all of them before the button even works. It may add aria-labels to elements it later deletes, invent state transitions for impossible conditions, or create an abstraction because three different guidance files each hinted that one might be useful.
The result is not more thoughtful code. It is confused code trying to satisfy too many goals at once.
A better process is staged. First make the smallest end-to-end path work. Then test it. Then inject accessibility context and inspect keyboard and screen-reader behavior. Then inject i18n/a11y context. Then review visual polish and error states. The same model can perform each step, but it should not be asked to think like every specialist at the same time.
Context is not a bag of documentation attached to a task. It is part of the state machine.
From prompts to a workflow

For MAI, we decided to describe as much of this as possible in a data-driven way using YAML. You can think of it as a DAG that is driven by orchetration agents with client side API that will keep track of which prompts were injected, which stage was running, what checks had passed, and where work should go next. We initially ran around 1,000 tasks based on real production work until it got a point where it was able to automate tasks to a production level quality of work. That number is currently around 40,000.
The DAG approach performed better subjectively, especially for UI and UX flows, and objectively through things such as better latency and smaller files. It did not make the model smarter. It made the environment around the model less ambiguous.
The states got more and more complex; defining the goal became its own stage, a task requiring both backend and UI work might implement the backend first, then the UI, and then connect the two. A bug fix might first reproduce the problem using a mock or a real local environment, make the change, and then repeat the reproduction steps to prove that the behavior changed. Over time, we created more skills, prompts, and branches for specific situations. Performance improved for a while, until we had so many skills that some of them were no longer called at all.
This exposed a second-order problem: it is not enough to write good instructions. The harness also has to reliably decide when those instructions are relevant. A skill that is never selected is just documentation in a different folder. A skill selected for every task is noise. Routing context turned out to be as important as writing it.
God does roll dice

If you use AI, you understand what it means for a system to be stochastic. Ask the same question twice and it may produce two different implementations, unless the question is so constrained that it is effectively asking the model to repeat something.
This becomes more obvious with tickets. It does not matter whether a ticket has a deep summary and a long clarification list. If you hand the entire thing to an agent as one initial prompt, there is no guarantee that each requirement will survive into the final code. One run will focus on error handling, another on architecture, and another will somehow spend most of its time renaming variables.
So the problem is not how to make the model deterministic. You probably cannot. The useful question is how to make the system around it deterministic enough that random model behavior cannot quietly become random production behavior.
CI/CD as executable taste
Tests and checks are the most reliable way to enforce this. One simple example was AI using emoji instead of the SVG icon library already used by the product. It did not matter how often we wrote "do not use emoji" in the prompt. Eventually a model would use one anyway.
The reliable solution was to detect icon-like Unicode characters in the diff and fail the check. The agent then had to run that check before it could move to the next stage. The rule stopped being a suggestion in a prompt and became a property of the repository.
The same idea applies to nested try/catch blocks, banned dependencies, generated files, unsupported formatting, missing localization calls, or any other behavior that is both objectively detectable and consistently unwanted. You can think of these as linters, except they are useful because they encode actual unwanted behavior in your codebase instead of forcing a long C++ name into modern art like weird::deeply::nested::namespace::layer
Not every preference should become a test. Subjective checks pretending to be objective will only produce a more frustrating harness. But when a requirement can be checked mechanically, a deterministic failure is better than another paragraph in the system prompt.
Context injection
For requirements that cannot be reduced to a test, the harness should inject context at the stage where it matters and require the agent to account for it before moving on. One basic technique is to create a checklist in the initial task and make each stage update it. Modern CLIs such as OpenCode, Copilot CLI, Codex, and Claude now support versions of this behavior, but earlier we had to build it manually.
For the modal example, the state could look roughly like this:
stages:
- implement:
context: [api-contract, ui-components]
exit: [build-passes, modal-opens]
- test:
context: [test-patterns]
exit: [happy-path-covered, failure-path-covered]
- polish:
context: [accessibility, internationalization]
exit: [keyboard-checked, strings-localized]
- review:
context: [code-ownership, performance]
exit: [required-checks-pass]
The exact YAML is not important. The important part is that context, actions, and exit conditions are attached to stages rather than dumped into one enormous prompt. The harness can see what has been attempted, what remains, and why a task is allowed to move forward.
At this point, we have solved the most basic harness problem: set context, run agents, and gate their progress with tests. But we have not really solved the operational problem because the whole thing still runs on one person's machine. How do we scale it out?
Scaling
The compute part is not especially difficult. You can spin up more Copilot SDK or CLI workers on a server and make them poll a dashboard or queue for tasks. Multiple orchestration servers, which could themselves be driven partly by an LLM, can start workers with an initial prompt, track task state, gate transitions through CI/CD, and inject the context required by each stage.
Adding workers is easy. Giving them a trustworthy world to work in is not.
Each worker needs repository access, credentials with a limited scope, an isolated workspace, a known base commit, and a safe way to publish its result. This is harder than it sounds at a company like Microsoft, but it is still a real problem anywhere else. If an agent can read every repository and deploy every service, you have built an extremely capable security incident. If it cannot access enough, it produces patches that cannot be tested in the environment where they matter.
Then there is observability. It is not enough for a dashboard to say "working." You need to know which stage is running, which commit it started from, what context it received, what tool calls it made, what check failed, whether it is making progress, and how much money it has burned while reconsidering the same test for the twelfth time.
Scaling also makes several social and coordination problems impossible to ignore.
Who gets the credit, and who takes the blame?
Who owns work produced by a harness?
- The person who created the ticket?
- The people who wrote the skills and prompts?
- The people who maintain the harness?
- The code owner who approved the result?
Credit is the less important half of this question. The real issue is accountability. When the output is wrong, somebody still has to decide whether the ticket was underspecified, the context was stale, the model made an unreasonable choice, a deterministic check was missing, or a reviewer approved something they should not have.
Saying "the AI wrote it" is not an ownership model. A useful harness needs a clear human boundary: who requested the change, who is allowed to approve it, and who is responsible after it ships. Without that, automation can make changes faster while making failures harder to assign and therefore harder to learn from.
Race conditions
The world does not pause while an agent works. A person can push a new commit, another agent can edit the same file, a dependency can change, or the base branch can move. This creates questions that a local demo can mostly ignore but a production harness cannot:
- If a person creates a PR without the harness, when and how should automated review start?
- Should review attach to a PR, a branch, or an exact commit SHA?
- If a new commit arrives during review, should the current review be cancelled, resumed, or discarded?
- If two agents make conflicting changes, which one rebases and who judges the reconciliation?
- If an agent dies halfway through a review, can another worker safely resume from its state?
- If a model provider goes down after code is changed but before review finishes, is the task failed, paused, or allowed to fall back to another model?
Commit SHAs help because they make inputs explicit. A review result for commit A should not silently be treated as a review of commit B. But using SHAs does not answer the whole question. The orchestrator still needs cancellation, leases, retries, idempotent transitions, and a rule for invalidating stale results.
Human activity is also part of this race. Designers will always want to vibe code because they need to see the output, and UI/UX remains difficult because animations, spacing, and "this feels wrong" are hard to tokenize. You can inject frame-by-frame data, screenshots, design outlines, or visual diffs, and those can help. They still do not fully capture the feedback a designer gets from interacting with a half-working prototype.
A harness therefore cannot assume it owns the entire development process. It has to reconcile automated work with humans making changes outside of it. In practice, that is a distributed systems problem wearing a code-generation hat.
Conclusion
These are some of the problems I had to solve, at least to a certain degree. We managed to consolidate the work into a version of the harness that is now used in 1JS, the second- or third-largest repository at Microsoft; Epichan, where SuperApp's Autopilot and Codetab UI is being built on; and Canopy, the mobile repository for SuperApp Autopilot and Codetab. Hundreds of developers use it now, which is both amazing and a very fun experiment.
I still wonder what the next version of this idea looks like. Todo lists, context files, and basic agent loops are already built into many AI tools. Models are getting "better" according to benchmarks, although I personally do not feel a dramatic difference in everyday usages. Each model also behaves differently, so a workflow tuned for one model may perform worse with another. Hill climbing on every new model, while also optimizing cost and latency, can become its own endless engineering project.
Part of me thinks we may not need much of this if models become sufficiently capable. Maybe the harness is scaffolding that disappears once the building can hold itself up. But another part of me thinks better models will simply attempt larger tasks, and larger tasks will make context, verification, ownership, and coordination even more important.
The harness is not really there to make the model intelligent. It is there to make an unreliable intelligence participate in a reliable engineering process. Whether that remains a separate product or slowly dissolves into every development tool is the part nobody knows yet.