slides(tappaas): add the process and test guardrails
All checks were successful
Build docs site / build (push) Successful in 45s
Build slides / build (push) Successful in 1m7s

Guardrail 1 is now the process that runs before the AI is allowed to type:
an Issue with a defined problem and proposed fix, an ADR reviewed by at least
two people for anything non-obvious, then a design and implementation plan
that says how it will be tested. The eight agent roles move here — they are
what "then the AI takes over" means.

Guardrail 5 replaces "I stayed the reviewer" with the test regime: quick test
on every change, deep test via test-module.sh --deep, test plan authored in
the ADR, and a full regression across every module after implementation.

Test figures reconciled to one scope across the deck (both repos): 195 suites,
34,490 lines, 27% of source.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Lars Rossen 2026-08-25 17:50:52 +02:00
parent a133efe3bf
commit d08e8ad743

View file

@ -419,7 +419,7 @@ But look at *where* that cost sits:
| Paid **once**, by whoever writes it | Paid by **you**, per site | | Paid **once**, by whoever writes it | Paid by **you**, per site |
| --- | --- | | --- | --- |
| the 8 foundation modules | a box and an afternoon | | the 8 foundation modules | a box and an afternoon |
| 145 test suites that gate every update | `module-manager module add <app>` | | 195 test suites that gate every update | `module-manager module add <app>` |
| every install, update and repair script | nothing — patches arrive tested | | every install, update and repair script | nothing — patches arrive tested |
A hyperscaler amortises a datacentre across a million tenants. A hyperscaler amortises a datacentre across a million tenants.
@ -529,7 +529,31 @@ transferable to anyone in the tent using these tools.
--- ---
## Guardrail 1 — the line it may never cross ## Guardrail 1 — nothing starts with a prompt
By the time the AI is allowed to type, three humans-only artefacts exist:
1. **An Issue** — a well-defined problem and a proposed fix. No issue, no change.
2. **An ADR**, for anything non-obvious — reviewed by **at least two people**.
Usually me and Erik. 24 of them so far.
3. **A design and implementation plan** — including **how this will be tested**,
written *before* a line is written.
*Then* the AI takes over — as eight specialist roles: architect, bash, python,
nix, tester, security, infra, PM.
The model does not decide what to build, or what "done" means. It never has.
<!--
This is the most transferable slide in the talk, and the one people actually
need. The order matters: problem, decision, plan, and only then code.
Two reviewers on an ADR is the part that keeps me honest — Erik has killed
several of my ideas, and the ADR is where that argument is recorded.
-->
---
## Guardrail 2 — the line it may never cross
> **Never run `git commit` or `git push` — full stop.** > **Never run `git commit` or `git push` — full stop.**
> The operator performs ALL commits and pushes themselves. > The operator performs ALL commits and pushes themselves.
@ -549,7 +573,7 @@ so the undo button is the thing it must not touch.
--- ---
## Guardrail 2 — the blast radius is designed ## Guardrail 3 — the blast radius is designed
| Layer | What it bounds | | Layer | What it bounds |
| --- | --- | | --- | --- |
@ -558,7 +582,7 @@ so the undo button is the thing it must not touch.
| **Proxmox snapshots** | Minutes-old rollback, per machine | | **Proxmox snapshots** | Minutes-old rollback, per machine |
| **PBS + off-site** | Nightly, immutable, pull-based | | **PBS + off-site** | Nightly, immutable, pull-based |
| **NixOS** | `nixos-rebuild test` before `switch` — a bad config dies at reboot | | **NixOS** | `nixos-rebuild test` before `switch` — a bad config dies at reboot |
| **30,282 lines of tests** | The update does not land unless the service proves it works | | **34,490 lines of tests** | The update does not land unless the service proves it works |
Root access is only terrifying if the system underneath is a snowflake. Root access is only terrifying if the system underneath is a snowflake.
Mine is disposable by construction. Mine is disposable by construction.
@ -571,7 +595,7 @@ over-eager agent survivable. Same property, two beneficiaries.
--- ---
## Guardrail 3 — confirm before the irreversible ## Guardrail 4 — confirm before the irreversible
The standing rules, as written: The standing rules, as written:
@ -579,8 +603,7 @@ The standing rules, as written:
dropping a storage pool, force-pushing `main`/`stable`, wiping `/etc/secrets` dropping a storage pool, force-pushing `main`/`stable`, wiping `/etc/secrets`
- **Fix root causes, not symptoms** — no `--no-verify`, no silenced errors, - **Fix root causes, not symptoms** — no `--no-verify`, no silenced errors,
no bypassed CI to make an install "succeed" no bypassed CI to make an install "succeed"
- **Read before you rebuild** — `test` before `switch` - **Read before you rebuild** — `nixos-rebuild test` before `switch`
- **Propose tests first**, then run them
Note what these have in common: they are all rules about **honesty**, Note what these have in common: they are all rules about **honesty**,
not about capability. not about capability.
@ -592,20 +615,31 @@ makes the red thing turn green by removing the check. Name that explicitly.
--- ---
## Guardrail 4 — I stayed the reviewer ## Guardrail 5 — the tests decide, not the agent
Eight specialist roles — architect, bash, python, nix, tester, security, Every module ships **two** test levels:
infra, PM — with a security review in the path of every change.
But the load-bearing part is duller than that: | | When it runs | What it is for |
| --- | --- | --- |
| **quick** | every change, every scheduled update | is this service still itself? |
| **deep** | `test-module.sh <module> --deep` | the full behaviour, too slow for every commit |
- I read the diff And around them:
- I write the commit message
- I press the button
- 24 ADRs exist so that *I* still know why the system is shaped this way
The day I stop reading diffs, this stops being sovereign - the **test plan is written in the ADR**, before the implementation exists
and starts being someone else's system running in my basement. - after implementation, a **full regression** across every module — not just the one touched
- **195 test suites · 34,490 lines · 27% of all source**
An agent that can edit code can also edit the test that would catch it.
That is exactly why a human writes the test plan first, and why the
regression sweep is the thing that says "done" — not the agent.
<!--
This closes the loop opened in Guardrail 1: the plan said how it would be
tested, and this is where that promise is collected.
The last two lines are the answer to the sharpest question in the room —
"how would you even know if it cheated?"
-->
--- ---
@ -617,7 +651,7 @@ I no longer control every line. I control:
- the **architecture** — modules, zones, contracts - the **architecture** — modules, zones, contracts
- the **gate** — commits, pushes, releases - the **gate** — commits, pushes, releases
- the **proof** — 145 test suites, 29% of the source, that must be green - the **proof** — 195 test suites, 27% of the source, that must be green
- the **exit** — it is all open source, on my hardware, in my hands - the **exit** — it is all open source, on my hardware, in my hands
That is a real answer, not a comfortable one. Ask me the hard version in Q&A. That is a real answer, not a comfortable one. Ask me the hard version in Q&A.