The Tree and the Rhizome

Systems models are typically arborescent, or tree-like. There is a root source of truth, and everything else extends from there. Kubernetes has its API server. Ansible has whichever laptop you ran the playbook from. Terraform has a state file with a lock. Even my own deploy-rs setup had a root: one machine with the keys, pushing closures outward to nodes that could only receive. That’s a fine shape when an ops team is the root. For one person with a few computers, it means one of them is special, and the whole system is only as available as that one.

A tree with a single trunk and branching roots beside a rhizome, a tangle of shoots and roots with no center

Tree and rhizome. Drawing by Magda Wojtyra and Marc Ngui, CC BY-SA 4.0, via Wikimedia Commons

A rhizome, like ginger or crabgrass, is a network of connected points. In A Thousand Plateaus, Deleuze and Guattari lay out some of the principles of rhizomatic thought. A rhizome connects different types of things together, with no central pivot. A rhizome may be broken at any point, and it starts up again along its old lines or along new ones. Finally, they insist on the map over the tracing: a tracing copies a structure that was decided in advance, while a map is drawn from whatever actually connected and stays open to redrawing.

But this isn’t a philosophy post. It’s about how the endgame of my fleet configuration ended up taking a rhizomatic shape due to a handful of decisions.

The Fleet Config Monorepo

If you install NixOS on a single computer, you end up with a /etc/nixos/configuration.nix which is the source code for the computer, usually importing an auto-generated hardware-configuration.nix. If you install NixOS on several computers, you end up with several of those. They’ll tend to have some overlap, on everything from timezones to your preferred dev tooling. If you want to upgrade something, you’ll repeat the same process on each computer. If you want to add cross-cutting functionality, well, you get the idea.

The move then is to create a single flake.nix whose schema provides a top-level nixosConfigurations. Your fleet becomes an attrset. Forgiving some pseudo-Nix for the rest of the post:

{
  desktop = { ... };
  laptop = { ... };
}

This of course allows you to abstract out shared segments of their definitions:

let devTools = [ firefox git ]; in
{
  desktop = {
    packages = devTools ++ [ ... ];
  };
  laptop = {
    packages = devTools ++ [ ... ];
  };
}

It’s a pretty neat way to parameterize any number of machines under a single source of truth!

One of the first steps in Rhizome-hood is to introduce a forge.

{
  vps = {
    services.forgejo.enable = true;
  };
}

Its role is to

  1. Host the config repo (among others)
  2. Run CI; in particular, a check.yml that builds all the hosts on every PR. Breaking changes don’t make it in.

I want the forge highly available, so I host it on my low-resource VPS, and my home laptop, which has the RAM for it, serves as the CI runner.

Notice what that does to the shape. The config is no longer administered from outside the fleet by one special machine with the keys. It is hosted, built, and checked by members of the fleet it describes.

Any Node Reaches Any Node

Nebula is a nifty tool from Slack that lets you define a network overlay with certificate-based identities. One well-known node (the “lighthouse”) acts as a rendezvous point, not a relay. Nodes usually talk directly once they find each other, even from behind NAT.

The VPS is the obvious choice for the lighthouse since it has a stable public IP. Making use of the rec keyword for a recursively evaluated attr-set, we can co-configure the nodes like this:

rec {
  vps = {
    services.nebula.networks.rhizome = {
      enable = true;
      isLighthouse = true;
    };
  };
  laptop = {
    services.nebula.networks.rhizome = {
      enable = true;
      lighthouses = [ vps ];  # really the lighthouse's overlay IP
    };
  };
}

Now every node has a stable address on a flat private subnet, no matter where it physically is. The payoff is that the low-resource VPS never has to build anything. It delegates every Nix build to the laptop over the overlay:

{
  vps = {
    nix.settings.max-jobs = 0;  # never build locally
    nix.buildMachines = [{ hostName = laptop; sshUser = "nixremote"; }];
  };
}

The same overlay is how I SSH into the laptop from anywhere, with its sshd closed to everything but the Nebula interface. Other tricks follow, like an nginx reverse proxy on the VPS serving websites that actually run at home.

Automate the Updates

Here’s where we start to close the loop, via 3 different update modes.

  1. Push: Updates to the VPS trigger a CI job wrapping the deploy-rs step.
  2. Pull: Laptops get a systemd timer that checks the forge for main’s tip, then builds and switches to that commit.
  3. Self-Bump: Apps running in the fleet are flake inputs. A green merge triggers a version bump in the config repo.

All the friction in Decoupling My Deployment Model just becomes automated CI steps.

Push

The push to the VPS is the deploy-rs command from the earlier posts, run by CI instead of me. It doesn’t need to happen on every merge, though, so the workflow only fires when the push asks for it. A commit trailer turns out to be the cheapest “deploy this” flag there is:

git commit -m "Tweak nginx" -m "Deploy: vps"

The workflow reads it with git interpret-trailers and runs nix run .#deploy-rs -- .#vps. No second system, no deploy button, just a line in the commit message.

Pull

Laptops don’t accept inbound connections, so instead of being pushed to they poll. A systemd timer asks the forge for the tip of main, and if it isn’t what’s running, switches to exactly that commit:

{
  laptop = {
    systemd.timers.fleet-switch.timerConfig.OnCalendar = "*:0/5";
    systemd.services.fleet-switch.script = ''
      rev=$(git ls-remote $FORGE/fleet.git main | cut -f1)
      [ "$rev" = "$(nixos-version --configuration-revision)" ] && exit 0
      nixos-rebuild switch --flake "git+$FORGE/fleet.git?rev=$rev#laptop"
    '';
  };
}

The comparison works because the flake stamps every build with the commit it came from (system.configurationRevision = self.rev). Pinning the exact revision it observed means a laptop that was asleep for a week catches up to the same commit as everyone else, not to whatever main happens to be mid-build.

Self-Bump

Apps that run on the fleet are flake inputs. Each one’s own CI, after a green merge, runs a single step:

nix run $FORGE/fleet.git#bump -- myapp=$sha

That re-pins the input, runs nix flake check with the new lock (a bump the fleet can’t build never lands), and pushes to main. A small map in the fleet repo says which hosts a moved pin should deploy:

rolling = {
  myapp = [ "laptop" ];
  website = [ "vps" ];
};

That one line and the step above are all the app and the fleet know about each other. The “tab back to the server repo, nix flake update, deploy” loop is gone.

Tool Bumps

The same trick works pointed at upstream instead of my own repos. Claude Code and Codex are pinned independently of nixpkgs, because the model gate is server-side: a CLI older than a new model locks every agent on the fleet out of it the day it ships. So a daily workflow compares the pins with what Anthropic and npm publish, rewrites them, checks, and pushes with a Deploy: vps trailer:

on:
  schedule:
    - cron: "17 6 * * *"
jobs:
  bump:
    runs-on: native
    steps:
      - run: nix run .#bump-tools

These show up in main’s history:

6f84965 bump tools: claude-code 2.1.287 -> 2.1.288
5ddc48f bump tools: claude-code 2.1.286 -> 2.1.287, codex 0.159.3 -> 0.160.0

YOLO Mode: Fleet-Wide Features

With these pieces in place, one PR in one repo can orchestrate cross-cutting features. Every host lives in the same repo, CI builds all of them on every PR, and the merge deploys itself. That makes it a great place to let an agent loose, since a fleet-wide change is a single reviewable diff with a build gate in front of it.

As an example, I asked for an observability system using ntfy. The topic name is the whole credential, so it lives in a secret, and the app on my phone subscribes to it.

{
  vps.fleet.notify.watch.laptop = { host = laptop; port = 22; };
  laptop.fleet.notify.watch.vps = { host = vps; port = 22; };
}

The hosts probe each other, and report each transition, from reachable to silent and back again. Any service unit can also wire in:

systemd.services.backup.onFailure = [ "notify-failed@%n.service" ];

This way, if there’s a rupture in the flows, I’ll know about it quickly.

The Knots

A rhizome can still tie itself in a loop, and that’s where some actual problem-solving went.

A cross-cutting change triggers a re-deploy of the VPS and the laptop. That means the CI runner deploys itself. A switch that restarts the runner kills every job inside it, including the deploy-rs activation of the VPS. Two rules untangle it. The laptop’s pull updater drains the runner before it switches, so no new jobs start mid-deploy. And the updater and CI’s VPS deploy share a single flock, so replacing the runner has to wait for any in-flight deploy to finish, and vice versa.

Secrets are the other knot, since they’re the one thing that can’t be declarative. sops-nix gets most of the way: each file is encrypted for my admin key plus the hosts that consume it, so “move a service to another host” is add a recipient, re-key, commit.

creation_rules:
  - path_regex: secrets/myapp\.yaml$
    key_groups: [{ age: [ *admin, *laptop ] }]

But the Nebula certificates are still signed by hand from a CA directory and copied out of band. Maybe the rhizome has a root after all. Oops!

Summary

In my rhizomatic system,

  • the config is a member of the fleet it describes
  • every node can reach every other
  • main is converged on from three directions
  • every edge reports its own failures

Heterogeneous things connected with no pivot, rupturable anywhere, and a map rather than a tracing. Together they mean I’m basically out of the loop, and the commit log on main reads like the fleet maintaining itself.

As usual, Deleuze and Guattari (illustrated by Marc Ngui and Magda Wojtyra) should help clarify matters.