# Repo Shape: Experiment → Notebook → Repo
Put this at the start of a session so new or continuing work matches this shape.
**The unit is not the repo. The unit is the experiment, and the notebook of experiments is what compounds.** A repo is what you *extract* from a notebook once the receipts have earned it.
---
## The lifecycle
```text
a claim that can be falsified
-> experiments that test it
-> hand-written receipts: observed / projected / realized
-> a proven/unproven ledger that survives across iterations
-> ONLY THEN, if earned: an extracted repo
```
**Work backwards.** Aim at something audacious enough to be wrong, then work back to the smallest experiment that produces an artifact either way. A `verified-disproved` result is a first-class artifact, not a failure.
Why this beats shipping a clean repo first: a tidy `0.0.1` repo risks nothing. An experiment notebook risks a real claim, keeps the receipts whether the claim lives or dies, and the artifacts compound.
### The notebook shape
A notebook is a directory, iteration-numbered, with an evidence ledger:
```text
README.md the nag -> the bet -> the falsifiable claim -> why a notebook
status.json state, iteration, northStar, constraint, falsifiable,
proven[], unproven[], disproven[], openQuestions[]
experiments/<name>/ self-contained, runnable, one per angle
receipts/ hand-written evidence; observed / projected / realized
```
### The proven/unproven ledger
`status.json` is the heart of the notebook. A claim moves to `proven` only with a linked receipt, and to `disproven` only with a `verified-disproved` receipt. Projection never counts as proof. The ledger is what compounds and what another notebook can cite.
---
## The North Star: build hammers that build hammers
Judge work by what it makes cheaper, safer, or possible next, not by what it does once.
```text
one experiment
-> a reusable receipt schema or primitive
-> a shared vocabulary other notebooks adopt
-> a tool that makes the next experiment faster
-> more, higher-value experiments
```
Each experiment answers three questions:
- **What does this make cheaper or possible for the next experiment?**
- **What evidence does it leave that another notebook can consume directly?**
- **Which hand-rolled substrate does it let a later experiment delete?**
If it produces a one-time output and leaves nothing reusable, it is throughput, not goodput. Record the finding, fold it into a larger system, or retire it.
### Verified now vs. projected vs. realized
Never let a projection wear the costume of an achievement:
```text
verified now the artifact and its evidence exist and run
projected a credible, falsifiable path to future leverage
realized later observed downstream use: imports, linked runs, deleted substrate
```
Words like "exponential", "high leverage", and "compounding" need a named mechanism and a receipt.
---
## The extracted-repo contract (the seven-minute floor)
Once a notebook has earned a repo, a new reader must answer these within seven minutes:
1. **What pain does this remove?** Open with the concrete failure mode.
2. **What is the one-sentence promise?**
3. **How do I run it?** One copyable happy path against the real system.
4. **How does it work?** One compact flow. Name the load-bearing primitive.
5. **How do I know it worked?** An executable check, receipt, replay, or live artifact.
6. **What does it deliberately not do?** Authority, security boundary, non-goals.
7. **Where do I edit behavior?** Policy in one obvious local file.
8. **What does this make possible next?** The leverage claim, with a mechanism.
Items 1-7 are the floor. Item 8 is the North Star made concrete.
---
## The stack
Default to this toolchain unless the project says otherwise:
- **Language**: TypeScript, ESM, ES2022+, `moduleResolution: bundler`, `noEmit`, `strict`. `bun` as package manager and runner, `bun.lock` committed.
- **App shape**: Svelte front end + Hono server in one Worker. Demo sites live under `site/`.
- **Platform**: Cloudflare-native. Workers, Durable Objects (+ `alarm()`), Workflows, Queues, R2, D1, KV, Vectorize, Browser Rendering, Workers AI, AI Gateway, Worker Loaders, Containers, MCP, Access.
- **No daemons**: Cron = heartbeat, Workflows = durable multi-step, Durable Object + `alarm()` = the loop that lives and remembers, Queues = fan-in.
- **Deploy**: `alchemy` (`alchemy.run.ts`) by default, `wrangler` for direct work.
- **Hygiene**: Biome for lint and format. `README.md`, `LICENSE` (MIT), `SECURITY.md`, `CONTRIBUTING.md`, `.gitignore`. Version stays `0.0.1` and `private` unless publishing. Docs follow Diátaxis. No hype.
---
## Default file shape
Delete anything the project does not need.
```text
README.md pain -> promise -> quick start -> proof -> limits -> leverage
SPEC.md / contract stable behavior and invariants
<editable config> modes, phases, policy, or workflow in one file
src/ smallest implementation satisfying the contract
site/ demo surface, if any
test/ deterministic behavioral checks
proof/ or receipts/ generated evidence in a portable shape
alchemy.run.ts deploy path
docs/ deeper material outside the seven-minute path
```
---
## Evidence is a shared contract, not bespoke logs
Receipts should be readable by your other projects. That is where leverage comes from.
- Acceptance gates follow one pattern: observe, act, assert. Pass only on observed correctness. Write a machine-readable receipt.
- A run receipt reads as: intent → execution → evidence → verified constraint cleared → cost, with projection kept apart from realization.
- **The gate is the product. Throughput is the cheap part.** A weak gate that produces confident receipts is worse than no gate. Make the referee strong before you scale volume.
---
## Completion is acceptance, not activity
These are activity signals, not proof: tasks marked done, commits, a passing weak suite, a new test file, the agent saying it finished.
A completion gate names an externally observable claim and links to replayable evidence. The harness may own the *intervention* point. An independent verifier owns *acceptance*.
```yaml
on_attempted_stop:
require:
- phase_tasks_complete
- clean_commit
- tests_pass
- proof_receipt:
claims: [requested_behavior_observed, regression_absent]
on_failure: resume_with: exact_unmet_conditions
on_success: accept: independent_verifier
```
Keep the implementation smaller than the idea it proves.
---
## Liquid primitives, solid invariants
Code, docs, prompts, models, and harnesses are disposable. The spec, the verification loop, the evidence contract, and the auth invariant are solid.
- Prefer plugins, adapters, glue, and small libraries on public platform primitives. Do not wait upstream.
- Delete hand-rolled substrate when an official primitive can carry the load. Keep the end-to-end receipt. Send concrete feedback to the owning team.
- Standalone repos survive only as narrow reusable primitives. Everything else folds into one integrated system or retires.
---
## Two gates
Default to starting an experiment, not a repo. Repos are extracted, not created.
### Gate 1: start an experiment (low bar, do this often)
- a claim that can be **falsified**, in one sentence;
- the smallest experiment that produces an artifact whether the claim lives or dies;
- which existing notebook this extends, or why it needs a new one;
- a credible data, source, auth, and security story;
- 1-3 load-bearing platform primitives, not decorative bindings;
- where the receipt goes and what `observed` it will carry.
If the claim cannot be falsified, it is not an experiment. Sharpen it or drop it.
### Gate 2: extract a repo (high bar, rare)
- **five honest receipts share one schema**;
- at least one `proven` claim with a linked receipt, and one recorded negative result;
- a leverage claim that has begun to **realize** (a second notebook cites the schema, or substrate was deleted);
- a named audience and one public sharing sentence;
- a decision: standalone narrow primitive / fold into a larger system / publish the notebook itself.
If the leverage is still only projected, keep iterating the notebook. Do not extract a repo to manufacture a sense of completion.
---
## Continual review
For a notebook:
- Is the central claim still falsifiable, or has it drifted into a slogan?
- Did this iteration move a claim between `unproven` / `proven` / `disproven`, with a receipt? If nothing moved, what did it produce?
- Are negative results recorded, or quietly dropped?
- Has the leverage realized, or is it still projected? If projected for too long, downgrade the claim.
For an extracted repo:
- Can the promise still be stated in one sentence?
- Does the quick start still exercise the real product?
- Does the proof establish the *current* claim, in a portable shape?
- Can an official primitive delete custom substrate now?
- Should this merge, split, fold in, or retire?
- What can be deleted without weakening the proof?
---
## Security defaults
- Access control before deploy. No public Worker URLs with bindings.
- Secrets before code. The model credential never reaches the browser.
- Dev-only shims stay uncommitted.
---
The unit is the experiment. The notebook is what compounds. The repo is a downstream extraction, earned by receipts. Aim at a claim big enough to be wrong, work backwards to the smallest artifact, and keep the evidence whether the claim lives or dies.
Repo Shape
The unit is the experiment, not the repo. Falsifiable claims, a proven/unproven ledger, receipts, and two gates before a repo is earned.
Operating Docs Updated Oct 8, 2026 ~2.4k tokens
1.2% of 200k
How to use this prompt
- Copy the prompt using the button above
- Paste it into your preferred AI coding assistant
- Adjust any placeholders or context as needed
- Let the agent implement the changes