diff --git a/slides/tappaas/one-year-in/index.md b/slides/tappaas/one-year-in/index.md index bf4640d..dc80c61 100644 --- a/slides/tappaas/one-year-in/index.md +++ b/slides/tappaas/one-year-in/index.md @@ -419,7 +419,7 @@ But look at *where* that cost sits: | Paid **once**, by whoever writes it | Paid by **you**, per site | | --- | --- | | the 8 foundation modules | a box and an afternoon | -| 145 test suites that gate every update | `module-manager module add ` | +| 195 test suites that gate every update | `module-manager module add ` | | every install, update and repair script | nothing — patches arrive tested | A hyperscaler amortises a datacentre across a million tenants. @@ -529,7 +529,31 @@ transferable to anyone in the tent using these tools. --- -## Guardrail 1 — the line it may never cross +## Guardrail 1 — nothing starts with a prompt + +By the time the AI is allowed to type, three humans-only artefacts exist: + +1. **An Issue** — a well-defined problem and a proposed fix. No issue, no change. +2. **An ADR**, for anything non-obvious — reviewed by **at least two people**. + Usually me and Erik. 24 of them so far. +3. **A design and implementation plan** — including **how this will be tested**, + written *before* a line is written. + +*Then* the AI takes over — as eight specialist roles: architect, bash, python, +nix, tester, security, infra, PM. + +The model does not decide what to build, or what "done" means. It never has. + + + +--- + +## Guardrail 2 — the line it may never cross > **Never run `git commit` or `git push` — full stop.** > The operator performs ALL commits and pushes themselves. @@ -549,7 +573,7 @@ so the undo button is the thing it must not touch. --- -## Guardrail 2 — the blast radius is designed +## Guardrail 3 — the blast radius is designed | Layer | What it bounds | | --- | --- | @@ -558,7 +582,7 @@ so the undo button is the thing it must not touch. | **Proxmox snapshots** | Minutes-old rollback, per machine | | **PBS + off-site** | Nightly, immutable, pull-based | | **NixOS** | `nixos-rebuild test` before `switch` — a bad config dies at reboot | -| **30,282 lines of tests** | The update does not land unless the service proves it works | +| **34,490 lines of tests** | The update does not land unless the service proves it works | Root access is only terrifying if the system underneath is a snowflake. Mine is disposable by construction. @@ -571,7 +595,7 @@ over-eager agent survivable. Same property, two beneficiaries. --- -## Guardrail 3 — confirm before the irreversible +## Guardrail 4 — confirm before the irreversible The standing rules, as written: @@ -579,8 +603,7 @@ The standing rules, as written: dropping a storage pool, force-pushing `main`/`stable`, wiping `/etc/secrets` - **Fix root causes, not symptoms** — no `--no-verify`, no silenced errors, no bypassed CI to make an install "succeed" -- **Read before you rebuild** — `test` before `switch` -- **Propose tests first**, then run them +- **Read before you rebuild** — `nixos-rebuild test` before `switch` Note what these have in common: they are all rules about **honesty**, not about capability. @@ -592,20 +615,31 @@ makes the red thing turn green by removing the check. Name that explicitly. --- -## Guardrail 4 — I stayed the reviewer +## Guardrail 5 — the tests decide, not the agent -Eight specialist roles — architect, bash, python, nix, tester, security, -infra, PM — with a security review in the path of every change. +Every module ships **two** test levels: -But the load-bearing part is duller than that: +| | When it runs | What it is for | +| --- | --- | --- | +| **quick** | every change, every scheduled update | is this service still itself? | +| **deep** | `test-module.sh --deep` | the full behaviour, too slow for every commit | -- I read the diff -- I write the commit message -- I press the button -- 24 ADRs exist so that *I* still know why the system is shaped this way +And around them: -The day I stop reading diffs, this stops being sovereign -and starts being someone else's system running in my basement. +- the **test plan is written in the ADR**, before the implementation exists +- after implementation, a **full regression** across every module — not just the one touched +- **195 test suites · 34,490 lines · 27% of all source** + +An agent that can edit code can also edit the test that would catch it. +That is exactly why a human writes the test plan first, and why the +regression sweep is the thing that says "done" — not the agent. + + --- @@ -617,7 +651,7 @@ I no longer control every line. I control: - the **architecture** — modules, zones, contracts - the **gate** — commits, pushes, releases -- the **proof** — 145 test suites, 29% of the source, that must be green +- the **proof** — 195 test suites, 27% of the source, that must be green - the **exit** — it is all open source, on my hardware, in my hands That is a real answer, not a comfortable one. Ask me the hard version in Q&A.