REVERSE BUILDING: DEFINE SUCCESS, THEN BUILD
PRD first. Gates second. Code last.
Problem: You write code, then figure out if it works. Tests come after implementation. By the time you verify, you've already committed to decisions that might be wrong.
Implementation: I write the PRD first, then define gates that observe the running system. The implementation comes last and changes until those gates pass. I use gateproof to run the checks.
REVERSE BUILD FLOW
==================
1. prd.ts define stories + dependencies
│
2. gates/*.gate.ts observe → act → assert
│
3. bun run prd.ts watch it fail
│
4. server.ts build to satisfy gates
│
5. bun run prd.ts all gates pass
│
└─── done.
The Inversion
Normal development: write code, then write tests to prove it works.
Reverse building: define what "working" means, then write code until the definition is satisfied.
Not TDD. TDD writes unit tests for functions. Reverse building writes stories with gates that verify end-to-end behavior against a running system. The PRD is the source of truth. Gates verify reality.
This Isn't New
Spec-first development is 15+ years old. BDD. Contract testing. Pact. Cucumber. API-first design with OpenAPI. These all exist and work.
What's different here is the audience. These gates aren't primarily for humans running npm test. They're for AI agents running in loops.
When an agent works from a PRD with gates:
- The PRD provides context without the full codebase
- Gate failures give concrete, practical feedback
- The agent can iterate against the same gates without a review between every attempt
- Observability backends capture what actually happened (not just assertions)
If you're a human writing tests for humans, use Playwright or Jest. If you're building harnesses for AI agents that need to observe, act, and verify against running systems, this pattern fits.
Why Not Playwright?
Playwright is excellent for DOM-based E2E tests. Use it for UI flows. But:
- Playwright tests DOM state. Gates observe system logs and backend behavior.
- Playwright assertions are synchronous. Gates poll observability backends over time.
- Playwright is for humans debugging in headed mode. Gates are for agents running headless in CI.
Different tools for different jobs. Gateproof gates work with Playwright — Act.browser() uses it under the hood. But the assertion layer is about observed reality, not DOM snapshots.
A Minimal Example
The dice game is intentionally trivial. A 45-line server doesn't need 4 gate files and 3 action scripts. The overhead is disproportionate on purpose — this is a teaching example, not a production recommendation.
The pattern pays off on larger projects where the PRD has 20+ stories, gates verify cross-service behavior, and an AI agent iterates through failures without human intervention. For real usage, see the Loop and Constraints Not Loops articles.
STEP 1: DEFINE SUCCESS
prd.ts — written before any game code exists
import { definePrd } from "gateproof/prd";
export const prd = definePrd({
stories: [
{
id: "server-responds",
title: "Server responds on port 3000",
gateFile: "./gates/server-responds.gate.ts",
},
{
id: "can-roll",
title: "Roll endpoint returns valid dice value (1-6)",
gateFile: "./gates/can-roll.gate.ts",
dependsOn: ["server-responds"],
},
{
id: "can-bet",
title: "Bet endpoint accepts wager and returns result with score",
gateFile: "./gates/can-bet.gate.ts",
dependsOn: ["can-roll"],
},
{
id: "state-persists",
title: "Game state persists across multiple bets",
gateFile: "./gates/state-persists.gate.ts",
dependsOn: ["can-bet"],
},
] as const,
});
if (import.meta.main) {
const { runPrd } = await import("gateproof/prd");
const result = await runPrd(prd);
if (!result.success) {
console.error(`Failed at: ${result.failedStory?.id}`);
process.exit(1);
}
console.log("All gates passed.");
process.exit(0);
}Four stories. Dependency ordering. server-responds must pass before can-roll runs. state-persists runs last. Run it now — it fails. The game doesn't exist yet. That's the point.
STEP 2: WRITE THE GATES
Each gate uses createHttpObserveResource to poll /api/state for changes, Act.exec to trigger actions, and Assert.custom to verify behavior.
One thing I discovered: Act.exec validates commands for shell safety — no quotes, braces, or metacharacters allowed. Wrap HTTP calls with JSON payloads in script files instead.
scripts/bet.ts — action script that Act.exec can safely call
const res = await fetch("http://localhost:3000/api/bet", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ amount: 10 }),
});
const data = await res.json();
console.log(JSON.stringify(data));gates/can-roll.gate.ts — observe state, trigger roll, assert valid value
import { Gate, Act, Assert } from "gateproof";
import { createHttpObserveResource } from "gateproof";
export async function run() {
const observe = createHttpObserveResource({
url: "http://localhost:3000/api/state",
pollInterval: 500,
});
return await Gate.run({
name: "can-roll",
observe,
act: [
Act.exec("bun run scripts/roll.ts"),
Act.wait(1000),
],
assert: [
Assert.noErrors(),
Assert.custom("valid-dice-value", (logs) => {
return logs.some((log) => {
const body = (log.data as any)?.body;
if (!body) return false;
const roll = Number(body.lastRoll);
return roll >= 1 && roll <= 6;
});
}),
],
stop: { idleMs: 2000, maxMs: 8000 },
});
}gates/state-persists.gate.ts — two bets, verify totalRolls increments
import { Gate, Act, Assert } from "gateproof";
import { createHttpObserveResource } from "gateproof";
export async function run() {
const observe = createHttpObserveResource({
url: "http://localhost:3000/api/state",
pollInterval: 500,
});
return await Gate.run({
name: "state-persists",
observe,
act: [
Act.exec("bun run scripts/bet.ts"),
Act.wait(500),
Act.exec("bun run scripts/bet.ts"),
Act.wait(1500),
],
assert: [
Assert.noErrors(),
Assert.custom("multiple-rolls-tracked", (logs) => {
return logs.some((log) => {
const body = (log.data as any)?.body;
if (!body) return false;
return Number(body.totalRolls) >= 2;
});
}),
],
stop: { idleMs: 2000, maxMs: 15000 },
});
}The HTTP observe resource captures each poll response as a log entry. Response bodies live at log.data.body. Act.wait(1000) after each action gives the observer time to poll the updated state.
STEP 3: BUILD LAST
Now — and only now — write the implementation. The only goal: make the gates pass.
server.ts — Bun HTTP server, no frameworks
const game = {
score: 0,
lastRoll: 0,
totalRolls: 0,
};
function roll(): number {
const value = Math.floor(Math.random() * 6) + 1;
game.lastRoll = value;
game.totalRolls++;
return value;
}
function bet(amount: number) {
const value = roll();
const win = value >= 4;
game.score += win ? amount : -amount;
return {
roll: value,
win,
result: win ? "win" : "lose",
score: game.score,
totalRolls: game.totalRolls,
};
}
const server = Bun.serve({
port: 3000,
async fetch(req) {
const url = new URL(req.url);
if (url.pathname === "/api/roll" && req.method === "POST") {
const value = roll();
return Response.json({ roll: value, totalRolls: game.totalRolls });
}
if (url.pathname === "/api/bet" && req.method === "POST") {
const body = (await req.json()) as { amount: number };
return Response.json(bet(body.amount));
}
if (url.pathname === "/api/state") {
return Response.json({ ...game });
}
return new Response("Not found", { status: 404 });
},
});
console.log(`Dice game running at http://localhost:${server.port}`);Three endpoints. In-memory state. 45 lines. Nothing more than what the gates require.
Run It
--- server-responds: Server responds on port 3000
{
"status": "success",
"durationMs": 380
}
--- can-roll: Roll endpoint returns valid dice value (1-6)
{
"status": "success",
"durationMs": 1344,
"evidence": {
"stagesSeen": ["http"],
"actionsSeen": ["GET http://localhost:3000/api/state"]
}
}
--- can-bet: Bet endpoint accepts wager and returns result with score
{
"status": "success",
"durationMs": 1372
}
--- state-persists: Game state persists across multiple bets
{
"status": "success",
"durationMs": 2423
}
All gates passed.Four stories. Four gates. All pass. The PRD defined success. The gates verified reality. The server was built last.
ROUGH EDGES (HONEST ASSESSMENT)
I've been using this pattern for small, medium, and large production products. I haven't fully given up the wheel, but I drive all development with PRD-first gates as of January 2026. What's still rough:
Polling + wait is timing-dependent. createHttpObserveResource polls on an interval. You need Act.wait() after actions to let the observer capture updated state. This is a race condition waiting to happen. Future versions should support event-driven observation, but today you're managing timing manually.
Shell safety forces workarounds. Act.exec blocks quotes, braces, and pipes. Good for security. Means you can't inline curl -d '{"amount":10}'. Wrapping HTTP calls in script files works, but it's boilerplate. This is early-stage tooling roughness, not intentional design.
Response body nesting is awkward. Your JSON lives at log.data.body, not log.data. The wrapping includes status code and headers, which is useful for debugging but annoying for assertions. Expect the API to evolve.
The payoff is real. Despite the rough edges, writing gates first changes how you build. When the PRD already defines "working," implementation becomes mechanical. No design debates. No scope creep. Build exactly what satisfies the gates.
When to Use It
- API-first projects where endpoints are well-defined
- Agent-driven development where the AI builds to satisfy gates
- CI/CD pipelines that need behavioral verification before deploy
- Microservices where integration contracts matter more than unit tests
THE AI LOOP (WHERE THIS MATTERS)
The real use case isn't humans running bun run prd.ts. It's AI agents in a loop:
while gates fail:
agent reads PRD + failure output
agent modifies code
gates re-run
doneThe agent never sees the full codebase upfront. It gets the PRD (what to build), gate failures (what's wrong), and iterates. Minimal context, concrete feedback, convergence to working code.
I haven't fully automated this for mission-critical loops yet. Current focus is optimizing token usage and iteration count. But the pattern works for guided development where you review between iterations.
When Not to Use It
- Exploratory prototyping — if you're still figuring out the PRD, you can't write gates for it
- UI-heavy features — gates verify backend behavior, not visual layout
- Simple scripts — 4 gate files for a 45-line server is overkill outside demos
- When you need proven tooling — gateproof is new, use Playwright/Jest if stability matters more than the pattern
The Project
Full working example: github.com/acoyfellow/dice-game
dice-game/ ├── prd.ts ├── server.ts ├── gates/ │ ├── server-responds.gate.ts │ ├── can-roll.gate.ts │ ├── can-bet.gate.ts │ └── state-persists.gate.ts ├── scripts/ │ ├── state.ts │ ├── roll.ts │ └── bet.ts └── package.json
The PRD is the source of truth. Gates verify reality. Implementation is the last step.
Write what "done" looks like. Then build until it's true.