From d08e8ad743a24b7ec97ab5013f8d0740718d7ea9 Mon Sep 17 00:00:00 2001 From: Lars Rossen Date: Tue, 25 Aug 2026 17:50:52 +0200 Subject: [PATCH] slides(tappaas): add the process and test guardrails MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Guardrail 1 is now the process that runs before the AI is allowed to type: an Issue with a defined problem and proposed fix, an ADR reviewed by at least two people for anything non-obvious, then a design and implementation plan that says how it will be tested. The eight agent roles move here — they are what "then the AI takes over" means. Guardrail 5 replaces "I stayed the reviewer" with the test regime: quick test on every change, deep test via test-module.sh --deep, test plan authored in the ADR, and a full regression across every module after implementation. Test figures reconciled to one scope across the deck (both repos): 195 suites, 34,490 lines, 27% of source. Co-Authored-By: Claude Opus 5 --- slides/tappaas/one-year-in/index.md | 70 +++++++++++++++++++++-------- 1 file changed, 52 insertions(+), 18 deletions(-) diff --git a/slides/tappaas/one-year-in/index.md b/slides/tappaas/one-year-in/index.md index bf4640d..dc80c61 100644 --- a/slides/tappaas/one-year-in/index.md +++ b/slides/tappaas/one-year-in/index.md @@ -419,7 +419,7 @@ But look at *where* that cost sits: | Paid **once**, by whoever writes it | Paid by **you**, per site | | --- | --- | | the 8 foundation modules | a box and an afternoon | -| 145 test suites that gate every update | `module-manager module add ` | +| 195 test suites that gate every update | `module-manager module add ` | | every install, update and repair script | nothing — patches arrive tested | A hyperscaler amortises a datacentre across a million tenants. @@ -529,7 +529,31 @@ transferable to anyone in the tent using these tools. --- -## Guardrail 1 — the line it may never cross +## Guardrail 1 — nothing starts with a prompt + +By the time the AI is allowed to type, three humans-only artefacts exist: + +1. **An Issue** — a well-defined problem and a proposed fix. No issue, no change. +2. **An ADR**, for anything non-obvious — reviewed by **at least two people**. + Usually me and Erik. 24 of them so far. +3. **A design and implementation plan** — including **how this will be tested**, + written *before* a line is written. + +*Then* the AI takes over — as eight specialist roles: architect, bash, python, +nix, tester, security, infra, PM. + +The model does not decide what to build, or what "done" means. It never has. + + + +--- + +## Guardrail 2 — the line it may never cross > **Never run `git commit` or `git push` — full stop.** > The operator performs ALL commits and pushes themselves. @@ -549,7 +573,7 @@ so the undo button is the thing it must not touch. --- -## Guardrail 2 — the blast radius is designed +## Guardrail 3 — the blast radius is designed | Layer | What it bounds | | --- | --- | @@ -558,7 +582,7 @@ so the undo button is the thing it must not touch. | **Proxmox snapshots** | Minutes-old rollback, per machine | | **PBS + off-site** | Nightly, immutable, pull-based | | **NixOS** | `nixos-rebuild test` before `switch` — a bad config dies at reboot | -| **30,282 lines of tests** | The update does not land unless the service proves it works | +| **34,490 lines of tests** | The update does not land unless the service proves it works | Root access is only terrifying if the system underneath is a snowflake. Mine is disposable by construction. @@ -571,7 +595,7 @@ over-eager agent survivable. Same property, two beneficiaries. --- -## Guardrail 3 — confirm before the irreversible +## Guardrail 4 — confirm before the irreversible The standing rules, as written: @@ -579,8 +603,7 @@ The standing rules, as written: dropping a storage pool, force-pushing `main`/`stable`, wiping `/etc/secrets` - **Fix root causes, not symptoms** — no `--no-verify`, no silenced errors, no bypassed CI to make an install "succeed" -- **Read before you rebuild** — `test` before `switch` -- **Propose tests first**, then run them +- **Read before you rebuild** — `nixos-rebuild test` before `switch` Note what these have in common: they are all rules about **honesty**, not about capability. @@ -592,20 +615,31 @@ makes the red thing turn green by removing the check. Name that explicitly. --- -## Guardrail 4 — I stayed the reviewer +## Guardrail 5 — the tests decide, not the agent -Eight specialist roles — architect, bash, python, nix, tester, security, -infra, PM — with a security review in the path of every change. +Every module ships **two** test levels: -But the load-bearing part is duller than that: +| | When it runs | What it is for | +| --- | --- | --- | +| **quick** | every change, every scheduled update | is this service still itself? | +| **deep** | `test-module.sh --deep` | the full behaviour, too slow for every commit | -- I read the diff -- I write the commit message -- I press the button -- 24 ADRs exist so that *I* still know why the system is shaped this way +And around them: -The day I stop reading diffs, this stops being sovereign -and starts being someone else's system running in my basement. +- the **test plan is written in the ADR**, before the implementation exists +- after implementation, a **full regression** across every module — not just the one touched +- **195 test suites · 34,490 lines · 27% of all source** + +An agent that can edit code can also edit the test that would catch it. +That is exactly why a human writes the test plan first, and why the +regression sweep is the thing that says "done" — not the agent. + + --- @@ -617,7 +651,7 @@ I no longer control every line. I control: - the **architecture** — modules, zones, contracts - the **gate** — commits, pushes, releases -- the **proof** — 145 test suites, 29% of the source, that must be green +- the **proof** — 195 test suites, 27% of the source, that must be green - the **exit** — it is all open source, on my hardware, in my hands That is a real answer, not a comfortable one. Ask me the hard version in Q&A.