You edit program.md.
The agent edits the code.
The harness decides.
Tests are frozen and restored before every run. The metric lives in a compiled binary the agent can't reach. Noise is discarded, regressions are rejected.
- one commit per accepted change
- one results.tsv row per experiment — including failures
- stop any time, nothing is thrown away
the loop runs until you stop it
p<0.0001
one overnight run
Real measurements, never illustrations. Validated on go-humanize, mapstructure and google/uuid.
read the case study ↗