slides(tappaas): add the process and test guardrails
Guardrail 1 is now the process that runs before the AI is allowed to type: an Issue with a defined problem and proposed fix, an ADR reviewed by at least two people for anything non-obvious, then a design and implementation plan that says how it will be tested. The eight agent roles move here — they are what "then the AI takes over" means. Guardrail 5 replaces "I stayed the reviewer" with the test regime: quick test on every change, deep test via test-module.sh --deep, test plan authored in the ADR, and a full regression across every module after implementation. Test figures reconciled to one scope across the deck (both repos): 195 suites, 34,490 lines, 27% of source. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
a133efe3bf
commit
d08e8ad743
1 changed files with 52 additions and 18 deletions
|
|
@ -419,7 +419,7 @@ But look at *where* that cost sits:
|
|||
| Paid **once**, by whoever writes it | Paid by **you**, per site |
|
||||
| --- | --- |
|
||||
| the 8 foundation modules | a box and an afternoon |
|
||||
| 145 test suites that gate every update | `module-manager module add <app>` |
|
||||
| 195 test suites that gate every update | `module-manager module add <app>` |
|
||||
| every install, update and repair script | nothing — patches arrive tested |
|
||||
|
||||
A hyperscaler amortises a datacentre across a million tenants.
|
||||
|
|
@ -529,7 +529,31 @@ transferable to anyone in the tent using these tools.
|
|||
|
||||
---
|
||||
|
||||
## Guardrail 1 — the line it may never cross
|
||||
## Guardrail 1 — nothing starts with a prompt
|
||||
|
||||
By the time the AI is allowed to type, three humans-only artefacts exist:
|
||||
|
||||
1. **An Issue** — a well-defined problem and a proposed fix. No issue, no change.
|
||||
2. **An ADR**, for anything non-obvious — reviewed by **at least two people**.
|
||||
Usually me and Erik. 24 of them so far.
|
||||
3. **A design and implementation plan** — including **how this will be tested**,
|
||||
written *before* a line is written.
|
||||
|
||||
*Then* the AI takes over — as eight specialist roles: architect, bash, python,
|
||||
nix, tester, security, infra, PM.
|
||||
|
||||
The model does not decide what to build, or what "done" means. It never has.
|
||||
|
||||
<!--
|
||||
This is the most transferable slide in the talk, and the one people actually
|
||||
need. The order matters: problem, decision, plan, and only then code.
|
||||
Two reviewers on an ADR is the part that keeps me honest — Erik has killed
|
||||
several of my ideas, and the ADR is where that argument is recorded.
|
||||
-->
|
||||
|
||||
---
|
||||
|
||||
## Guardrail 2 — the line it may never cross
|
||||
|
||||
> **Never run `git commit` or `git push` — full stop.**
|
||||
> The operator performs ALL commits and pushes themselves.
|
||||
|
|
@ -549,7 +573,7 @@ so the undo button is the thing it must not touch.
|
|||
|
||||
---
|
||||
|
||||
## Guardrail 2 — the blast radius is designed
|
||||
## Guardrail 3 — the blast radius is designed
|
||||
|
||||
| Layer | What it bounds |
|
||||
| --- | --- |
|
||||
|
|
@ -558,7 +582,7 @@ so the undo button is the thing it must not touch.
|
|||
| **Proxmox snapshots** | Minutes-old rollback, per machine |
|
||||
| **PBS + off-site** | Nightly, immutable, pull-based |
|
||||
| **NixOS** | `nixos-rebuild test` before `switch` — a bad config dies at reboot |
|
||||
| **30,282 lines of tests** | The update does not land unless the service proves it works |
|
||||
| **34,490 lines of tests** | The update does not land unless the service proves it works |
|
||||
|
||||
Root access is only terrifying if the system underneath is a snowflake.
|
||||
Mine is disposable by construction.
|
||||
|
|
@ -571,7 +595,7 @@ over-eager agent survivable. Same property, two beneficiaries.
|
|||
|
||||
---
|
||||
|
||||
## Guardrail 3 — confirm before the irreversible
|
||||
## Guardrail 4 — confirm before the irreversible
|
||||
|
||||
The standing rules, as written:
|
||||
|
||||
|
|
@ -579,8 +603,7 @@ The standing rules, as written:
|
|||
dropping a storage pool, force-pushing `main`/`stable`, wiping `/etc/secrets`
|
||||
- **Fix root causes, not symptoms** — no `--no-verify`, no silenced errors,
|
||||
no bypassed CI to make an install "succeed"
|
||||
- **Read before you rebuild** — `test` before `switch`
|
||||
- **Propose tests first**, then run them
|
||||
- **Read before you rebuild** — `nixos-rebuild test` before `switch`
|
||||
|
||||
Note what these have in common: they are all rules about **honesty**,
|
||||
not about capability.
|
||||
|
|
@ -592,20 +615,31 @@ makes the red thing turn green by removing the check. Name that explicitly.
|
|||
|
||||
---
|
||||
|
||||
## Guardrail 4 — I stayed the reviewer
|
||||
## Guardrail 5 — the tests decide, not the agent
|
||||
|
||||
Eight specialist roles — architect, bash, python, nix, tester, security,
|
||||
infra, PM — with a security review in the path of every change.
|
||||
Every module ships **two** test levels:
|
||||
|
||||
But the load-bearing part is duller than that:
|
||||
| | When it runs | What it is for |
|
||||
| --- | --- | --- |
|
||||
| **quick** | every change, every scheduled update | is this service still itself? |
|
||||
| **deep** | `test-module.sh <module> --deep` | the full behaviour, too slow for every commit |
|
||||
|
||||
- I read the diff
|
||||
- I write the commit message
|
||||
- I press the button
|
||||
- 24 ADRs exist so that *I* still know why the system is shaped this way
|
||||
And around them:
|
||||
|
||||
The day I stop reading diffs, this stops being sovereign
|
||||
and starts being someone else's system running in my basement.
|
||||
- the **test plan is written in the ADR**, before the implementation exists
|
||||
- after implementation, a **full regression** across every module — not just the one touched
|
||||
- **195 test suites · 34,490 lines · 27% of all source**
|
||||
|
||||
An agent that can edit code can also edit the test that would catch it.
|
||||
That is exactly why a human writes the test plan first, and why the
|
||||
regression sweep is the thing that says "done" — not the agent.
|
||||
|
||||
<!--
|
||||
This closes the loop opened in Guardrail 1: the plan said how it would be
|
||||
tested, and this is where that promise is collected.
|
||||
The last two lines are the answer to the sharpest question in the room —
|
||||
"how would you even know if it cheated?"
|
||||
-->
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -617,7 +651,7 @@ I no longer control every line. I control:
|
|||
|
||||
- the **architecture** — modules, zones, contracts
|
||||
- the **gate** — commits, pushes, releases
|
||||
- the **proof** — 145 test suites, 29% of the source, that must be green
|
||||
- the **proof** — 195 test suites, 27% of the source, that must be green
|
||||
- the **exit** — it is all open source, on my hardware, in my hands
|
||||
|
||||
That is a real answer, not a comfortable one. Ask me the hard version in Q&A.
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue