Tag
failure
The runs that stopped, the gates that went red, the caps set in the wrong place. Collected deliberately: a corpus that shows only the runs worth writing up is a portfolio.
- ConceptIndependent verification
The thing that checks cannot be the thing that did it, and a prompt is not a separation.
- ConceptDurable execution and resumability
If the work is longer than the thing doing it, the work has to live somewhere the worker does not.
- ArticleNever believe the agent
Verification has to be a separate program with real exit codes. The argument for that is not in the gates that passed. It is in the ones that did not, and in the few that never ran at all.
- ArticleThe nudge that sat below the median
Every autonomous run puts a ceiling on what one session may spend, and almost every ceiling was picked by feel. The ledger can settle it. The awkward part is that the number you most need is the one a ceiling deletes.
- ArticleThe ledger that lied
An autonomous run keeps a record of itself, and that record is what every later decision gets made from. This corpus holds a population of sessions whose record is empty in four places and full in the fifth. The shape of the hole turned out to be a code path, and the number it produced was the reassuring one.
- Run reportThe long build
This is the run that did not finish. It is here on purpose and it is the most useful report on the site, because the store recorded exactly why the last checkpoint stayed open and the answer was not the model, the plan or the code. It was a ceiling.
- Run reportThe engine run
Every checkpoint closed, and this is still the reddest run in the corpus. It holds nearly a third of every gate failure measured here, every unusual failure exit code in the whole store, and the only gates that never ran at all. That combination is not a contradiction. It is what a real release gate looks like.