MakerFLOSS/slides/tappaas/one-year-in/index.md
Lars Rossen 485f53f48d
All checks were successful
Build docs site / build (push) Successful in 49s
Build slides / build (push) Successful in 1m8s
slides(tappaas): real hardware/VM diagrams, and count the Community modules
Split the single architecture slide into two: the iron (three nodes, switch,
AP, modem) and the software (foundation band with the AI and productivity
stacks above it). Both read off the running cluster via tappaas-cicd, not
from docs — CPU, RAM, pool topology, link speeds and switch ports are live
values.

The VM layer-cake is inline SVG rather than mermaid: mermaid ignores
`direction LR` inside a subgraph that has a cross-boundary edge, so the bands
came out as a 629x1258 column.

Module slide now counts all three sources — 8 foundation + 12 apps +
21 Community = 41 — and slide 3 is recomputed for today.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 17:33:03 +02:00

29 KiB
Raw Blame History

marp theme class paginate title description
true gaia invert true TAPPaaS one year in — The Good, the Bad, and the Ugly Sommerhack 2026, Taler Teltet — 27 August, 15:30

TAPPaaS — one year in

The Good, the Bad, and the Ugly

Lars Rossen · Sommerhack 2026 · Taler Teltet 27 August, 15:30


Last year, in this tent, I pitched a dream

Everyone needs a cloud in their basement. A Trusted, Automated, fully Private Platform as a Service.

On commodity hardware:

  • your data, your hardware
  • no hidden fees
  • no reliance on closed source
  • no reliance on the Internet

A year later I am back with the messy proof that it is real — and to make the case that you should build one too.


The receipts, up front

First commit 10 May 2025 — "Initial commit"
Commits when I submitted this abstract 1,091
Commits standing here today 1,507
Contributors 3 (I am ~77% of it)
ADRs written 24
Modules — 8 foundation · 12 apps · 21 community 41
Lines of code (bash · TypeScript · Python · Nix) 126,699
Of which test code 34,490 — 27%
Lines of documentation ~39,500

Everything in this talk is in two public repositories. You can check my homework.


What I am actually going to do to you

  1. The Good — it runs. Here is what "it runs" means, and one surprise.
  2. The Bad — what this really cost. The commit graph is not pretty.
  3. The Ugly — the confession. I gave an AI root on every node.
  4. Your turn — why you should build one, and how to start small.

I am not selling anything. There is nothing to buy.


The Good

It runs.


The iron: three boxes, a switch, an AP and a modem

flowchart TB
  n1["<b>tappaas1</b> — ASUS · EPYC 4464P<br/>12c/24t · 64 GB ECC<br/>tanka1 2×4 TB NVMe <b>mirror</b><br/>tankb1 12 TB HDD"]
  n2["<b>tappaas2</b> — Minisforum MS-S1 MAX<br/>Ryzen AI MAX+ 395 · 16c/32t<br/><b>128 GB unified</b> · Radeon 8060S<br/>tanka1 2 TB NVMe"]
  n3["<b>tappaas3</b> — TianBei WTR PRO<br/>Ryzen 7 5825U · 8c/16t · 32 GB<br/>tanka1 1 TB NVMe<br/>tankc1 5 TB HDD — PBS"]
  sw["<b>UniFi USW Pro XG 10 PoE</b><br/>10.0.0.201<br/>trunk: mgmt untagged +<br/>VLAN 200 310 410 420 430 510 610"]
  ap["UniFi Nano HD<br/>access point · port 9"]
  modem["ISP modem<br/>port 5 · VLAN 100 = wan"]
  n1 -- "10G · p12" --> sw
  n2 -- "10G · p2" --> sw
  n3 -- "2.5G · p3" --> sw
  sw --- ap
  sw --- modem

The software: foundation underneath, stacks on top

AI stack — tappaas2, on the GPUvllm-amd · 312LXC · Radeon 8060S128 GB unifiedlitellm · 310gateway · keys4 GBopenwebui · 311the chat window4 GBProductivity — tappaas1nextcloud · 340files · calendar8 GB · 80 GBeuro-office · 343documents8 GBFoundation — every app above rests on thesenetwork · 110OPNsense8 GBtappaas-cicd · 130mothership16 GBidentity · 140Authentik4 GBlogging · 1502 GBunifi-os · 811switch · AP6 GBbackup · PBSon tappaas3tankc1


I promised you eight. There are forty-one.

Where # What is in there
Foundation — TAPPaaS repo 8 cluster network templates tappaas-cicd identity backup logging satellite
Apps — TAPPaaS repo 12 nextcloud euro-office vaultwarden litellm openwebui vllm-amd hass deconz coturn nextcloud-hpb netbird-client windows-server
Community repo — other people 21 immich jellyfin wordpress forgejo mailserver hosting sonos hue synology reolink solaredge alfen unifi shelly-fleet …

Twenty-one of those forty-one are not mine. Erik wrote 13, Andreas 6.

That is the number I did not dare put in the abstract.


"It runs" is a bigger claim than it sounds

Every one of those services, without me:

  • updates itself on a schedule — and the update is gated by a test
  • backs itself up nightly to Proxmox Backup Server, and off-site
  • reports its own health, so I find out before my family does
  • has a certificate that renews
  • has one login — my identity provider, not eight password fields

The interesting engineering is not installing Nextcloud. It is Nextcloud still being there, patched, in eighteen months, unattended.


One idea does all the work: everything is a module

nextcloud/
├── nextcloud.json  # the contract — what it is, needs, provides
├── nextcloud.nix   # the NixOS machine
├── install.sh      # put it there            (once)
├── update.sh       # keep it patched         (on schedule)
├── test.sh         # prove it still works    (gates the update)
└── README.md

The firewall is a module. The backup server is a module. The mothership that installs the modules is a module.

One model to learn, not eight.


dependsOn — the trick the whole thing rests on

{
  "description": "Nextcloud — files, calendar, contacts",
  "vmname": "nextcloud", "vmid": 210,
  "dependsOn": ["cluster:vm", "templates:nixos", "backup:vm",
                "network:proxy", "identity:identity"],
  "config": {
    "cluster:vm": { "cores": 4, "memory": "8192", "diskSize": "64G" },
    "network:proxy": { "proxyPort": 443 }
  }
}

Install order is computed from these declarations. Nobody maintains a list. Add a module, and the platform works out that it needs a VM, an OS, a VLAN, a certificate, a login and a backup job — in that order.


Reduce flexibility. On purpose.

A key design goal is to REDUCE flexibility. There is value in decisions having been taken up front.

Some use cases will not fit TAPPaaS. That is the trade. For what fits, it is dramatically easier.

Zones, VLANs, naming, storage roles, backup policy, identity model — decided. You get to pick the applications, and where they live.


The surprise

My basement grew a brain.


Local AI. Private, offline, mine.

Silicon AMD Ryzen AI MAX+ 395 — Radeon 8060S (Strix Halo)
Memory 128 GB unified — the GPU sees nearly all of it
Largest model tested gpt-oss-120b — 120B parameters
Speed ~50 tok/s at 7B FP16 · ~20 tok/s at 30B 4-bit
API OpenAI-compatible, on my own VLAN
Data leaving the building none

One commodity box. Not a rack, not a hyperscaler, not a monthly bill.


The demo I have been waiting a year to do


Pull the cable out.


Then ask it something.


Why this is the point, not a party trick

Sovereignty is not a checkbox you tick at a vendor.

  • The model runs on hardware you own
  • Your documents are indexed on your VLAN
  • Nobody re-prices it, deprecates it, or reads it
  • It works when the fibre is cut, the account is suspended, or the terms of service change on a Tuesday

litellm in front means apps ask for "a model" — local today, someone else's tomorrow, your choice, revocable.


The Bad

What it actually cost.


The brutal commit graph of a side project

2025-05   99  ██████████  ← the honeymoon
2025-06   21  ██
2025-07   53  █████
2025-08  141  ██████████████  ← Sommerhack 2025
2025-09    6  █  ← life
2025-10    9  █  ← still life
2025-11   41  ████
2025-12   47  █████
2026-01   66  ██████
2026-02  141  ██████████████
2026-03   33  ███
2026-04   19  ██
2026-05  190  ██████████████████
2026-06  351  ██████████████████████████████████  ← something changed
2026-07  164  ████████████████
2026-08   83  ████████  (to the 20th)

This is what a real side project looks like. Not a burndown chart. A heartbeat with two near-death experiences in it.


The two dips are the honest part

September–October 2025: 15 commits in two months.

TODO: what actually happened here — say it plainly, it is the most relatable slide in the deck

The lesson I take from it: a self-hosted platform that needs you every week is not a platform, it is a pet.

The dips are the real test. The services stayed up. Because updates, backups and tests do not need me to be enthusiastic.


June 2026: 351 commits. Something changed.

That is not me getting three times better at typing.

That is the month I leaned all the way into AI-assisted development — which is exactly the confession in part three.

Hold that thought.


The iceberg under every "simple" service

Files Lines
Foundation — the platform 575 114,138
Apps — the things you actually use 144 15,726

For every line in an app module, there are seven lines of platform underneath.

tappaas-cicd alone — the mothership — is 81,942 lines, 72% of the foundation. Inside it: 8 TypeScript managers (36,920 lines) and the controller layer that talks to Proxmox, OPNsense and the switch (29,352).


What it costs in money

Hardware TODO: what you actually spent, be honest, include the mistakes
Electricity TODO: measured W → kr/year at current DK prices
Satellite VPS TODO: kr/month
Domain + DNS TODO:
Licences 0
Versus the hyperscaler equivalent TODO: the comparison, done fairly — include your time at 0

The honest framing: this is not cheaper if you value your time at anything. It is cheaper if you were going to tinker anyway, and it is ownable at any price.


What it costs in things that are not money

  • Learning curve you cannot delegate — Proxmox, NixOS, OPNsense, ZFS, VLANs, PKI, OIDC. Any one of those is a weekend. You need all six.
  • You are the on-call. At 22:00. On holiday.
  • Family SLA. The moment calendars are on it, downtime is a domestic matter, not a technical one.
  • Decision fatigue — which is precisely why the project reduces flexibility.
  • The "why don't you just" tax — from every friend, every time.

TODO: your best war story — the outage that taught you the most


What I got wrong

TODO: pick three, be specific, be unflattering — this slide buys credibility for everything else

Candidates from the graph and the repo:

  • Numbered foundation modules (05-, 10-, 30-…) — retired in the ADR-007 refactor once ordering had to come from dependsOn instead
  • TODO:
  • TODO:

The pattern: everywhere I encoded an ordering or a name as a convention, I later had to make it a declaration.


The Ugly

The confession.


I used AI to build it.

And I gave it root on every node.


Not "AI-assisted autocomplete".

Root. ssh. nixos-rebuild. qm. pvesh. The firewall. The secrets.


Why on earth would you do that

Because the alternative was that it never got built.

  • One retiree, evenings and weekends, six unfamiliar technology stacks
  • 112,000 lines of platform code for thirteen app modules
  • The plumbing is tedious, not clever — the exact shape of work to hand over

The graph does not lie: May 190, June 351. That is what handing over the tedium looks like.

And the honest part: I could not have hand-written the last third of this.


So the real question is not "did you"

It is: what did you fence it with?

The guardrails are not vibes. They are written down, in the repo, loaded on every single session, and they override anything the model would otherwise default to.


Guardrail 1 — the line it may never cross

Never run git commit or git push — full stop. The operator performs ALL commits and pushes themselves. This holds even when a request seems to imply it — "land it", "ship it", "move this to main" — and even when a previous turn involved committing. That is NOT standing authorization.

The AI may change any file on disk. It may not make a change permanent.

Every single line that entered history passed under my eyes.


Guardrail 2 — the blast radius is designed

Layer What it bounds
Modules A mistake lands in one VM, not "the server"
Zones A compromised VM cannot reach what its VLAN forbids
Proxmox snapshots Minutes-old rollback, per machine
PBS + off-site Nightly, immutable, pull-based
NixOS nixos-rebuild test before switch — a bad config dies at reboot
30,282 lines of tests The update does not land unless the service proves it works

Root access is only terrifying if the system underneath is a snowflake. Mine is disposable by construction.


Guardrail 3 — confirm before the irreversible

The standing rules, as written:

  • Confirm before destructive ops — deleting a VM it did not create, dropping a storage pool, force-pushing main/stable, wiping /etc/secrets
  • Fix root causes, not symptoms — no --no-verify, no silenced errors, no bypassed CI to make an install "succeed"
  • Read before you rebuild — test before switch
  • Propose tests first, then run them

Note what these have in common: they are all rules about honesty, not about capability.


Guardrail 4 — I stayed the reviewer

Eight specialist roles — architect, bash, python, nix, tester, security, infra, PM — with a security review in the path of every change.

But the load-bearing part is duller than that:

  • I read the diff
  • I write the commit message
  • I press the button
  • 24 ADRs exist so that I still know why the system is shaped this way

The day I stop reading diffs, this stops being sovereign and starts being someone else's system running in my basement.


Am I still in control? Honestly.

Yes — but "control" moved.

I no longer control every line. I control:

  • the architecture — modules, zones, contracts
  • the gate — commits, pushes, releases
  • the proof — 145 test suites, 29% of the source, that must be green
  • the exit — it is all open source, on my hardware, in my hands

That is a real answer, not a comfortable one. Ask me the hard version in Q&A.


Your turn

Build one too.


Why you, specifically, should

European sovereignty is not a policy problem you can wait out.

It is thousands of small boxes, in basements and back offices, running software nobody can withdraw.

  • Your data has to live somewhere. Somewhere can be here.
  • A skill you own beats a subscription you rent.
  • Every basement cloud makes the next one cheaper to build.

TODO: your one-sentence version of this — say it in your own words, not mine


Start smaller than I did

Step What you get
1. One box, evaluation tier 4 cores / 16 GB / a disk. Nested virt is fine.
2. Proxmox + OPNsense Zones, VLANs, DNS, certificates that renew
3. The mothership tappaas-cicd — the thing that installs the rest
4. One service Vaultwarden. Small, useful, immediately missed.
5. Backup before service two Non-negotiable. Ask me why.

No public IP? A satellite VPS is the escape hatch. No GPU? Skip local AI, keep everything else.


Where to find all of it

  • tappaas.org — docs, work in progress, honest about it
  • codeberg.org/TAPPaaS/TAPPaaS — the code, the 24 ADRs, the commit graph
  • sovereigncomputing.org — the wider argument
  • This deck — slides.makerfloss.eu/tappaas/one-year-in
  • MakerFLOSS — Orange Makerspace, bi-weekly FLOSS jam. Come build one with us.

Contributions welcome. Open an issue before you open a pull request — somebody may already be packaging your app.


Questions

The harder, the better.