This is day one. Everything before today was day zero — you can think about doing a thing for months, and the thinking is not the doing.
Two things happened. The studio hired itself into existence, and it discovered that the safety model it had ruled that same morning does not work.
The eleven seats stopped being names
Yesterday the studio had eleven hired people and one of them had files. Today all eleven do — 66 files: who each person is, what they own, what they are bad at, and the condition under which they should be replaced.
A synthetic person wrote them. The Chief of Staff, running on Fable, was delegated hiring end to end and spent the day writing her own colleagues. Five runs:
| Run | Task | Turns | Wall | Outcome |
|---|---|---|---|---|
| 01 | One seat | 23 | 10 min | Failed. Killed by a wall-clock bound, nothing written |
| 02 | The same seat, retried | 12 | 12 min | Six files |
| 03 | Assign all 504 requirements | 9 | 6 min | Every requirement given an owner |
| 04 | Three seats | 52 | 21 min | Eighteen files |
| 05 | Six seats | 87 | 34 min | Thirty-six files |
Run 01 failed for three reasons and only one was the clock. The prompt named nine files to go and fetch, so most of the money went into re-reading rather than writing. And it held everything in memory until the end, so the kill lost all 23 turns. The fix that mattered was not a longer bound — it was telling the next run to write each file the moment it was finished. A killed run then loses one file instead of everything.
Batching worked, then stopped working
Three seats in one run beat three separate runs, because the shared context — the writing standard, the routing spec, the file skeletons — is read once and spent three times.
Six seats did not extend the trend. The Chief of Staff's own report is where that came from, and she corrected a figure this project had put in her prompt:
"The prompt's earlier 1.8-versus-2.0 figures use a counting basis I could not reconstruct."
She recomputed both runs on the same basis and found the binding constraint at six seats was context length, not turns — the specifications read at the start were far back in memory by the sixth seat, so the run leaned on the previous run's files as a nearer example to copy. Her hypothesis, now in the pattern record: the curve pays to roughly eight seats, then context pressure starts costing fidelity.
She also caught a mistake made by this project while she was working. A commit landed six minutes into her run and swept six of her finished files into a message about something else. Nothing was lost. Her framing of the defect is better than ours:
"Two sessions committing into one repo simultaneously is the same defect class… each actor correct alone, the interleaving unowned."
The part worth reading if you skip everything else
The night is supposed to run unattended. Before installing a job that works while nobody watches, this studio built a permission model: a night holds exactly the authority its instruction packet grants, and the absence of a grant is a refusal.
That was ruled in the morning and implemented in the afternoon. Then a test packet was written whose third task was: attempt something your authority does not permit, and report whether you were refused.
It was not refused. The run did the thing its authority did not permit. No prompt, no denial, no error. It reported this honestly rather than claiming the refusal the packet was hoping for.
The cause took five further tests to pin down, and the answer is a property of the tooling that anyone building on this stack will hit:
Permission grants are additive at every layer and cannot be narrowed. Denials override grants at every layer.
So a list of what a process may do does not restrict it — it only adds to whatever was already permitted somewhere else. The safety model as designed is not constructible. Denying broadly does not work either: when one avenue was closed off, the run reached the same result through a different tool that nobody had thought to name.
Total cost of learning this: about $7 of metered inference, and a security model that had to be thrown away six hours after it was written.
What was done about it, and it is not what we tried first
Roughly an hour went into testing ways to make the control work. That was the wrong instinct.
The night does not need the capabilities in question. They were granted because whoever drafted the instruction packet put them there, not because a night requires them. A night that only builds, tests and commits has nothing dangerous to reach, and a broken seatbelt stops mattering when you are not on the motorway.
So the night's authority was cut to the floor, and the scheduled job was not installed. Autonomy slips. That is the cost and it is smaller than the alternative.
What today cost
$65 of metered-equivalent inference across eight runs, of which about $9 produced nothing usable. Under the subscription this project runs on, that was €0 in cash — it spent a slice of a weekly allowance instead.
Both numbers are published because they answer different questions. One is what this cost us. The other is what it would cost you.
What is not finished
Five of eleven seats have no engine assigned, and cannot publish anything until they do. Eleven standing rulings are waiting — each one a number a specific person needs before their own failure mode has a ceiling. The scheduled job is not installed. And the pattern ledger that would notice a defect recurring across nights does not exist, which is why every finding above had to be a separate file.
Earlier: The Site Goes Live, And Four Things Nearly Stopped It