r/NixOS • u/RocketSeven • 1d ago
What should a NixOS recovery drill prove before you trust the configuration backup?
A version-controlled NixOS configuration is an important recovery asset, but rebuilding the declared system is not the same as restoring the machine. Hardware-specific files, boot setup, secrets, persistent application data, database state, imperatively installed items, and inputs that are no longer available can all sit outside the apparent source of truth. A configuration that still evaluates may also produce a system that cannot reach the network or unlock its storage.
A realistic drill could start from blank media or a disposable VM, recreate partitions and encryption, restore the flake or configuration and pinned inputs, supply secrets through the intended path, rebuild, and then restore stateful data separately. The result would need checks for boot, networking, users and SSH access, mounts, timers, containers or virtual machines, application health, monitoring, and another reboot. The drill should also identify every manual step that was not represented declaratively.
What do you include in a NixOS recovery test? How do you distinguish configuration backup from data and secret backup, and what evidence tells you the machine can be rebuilt after the original disk and cache are both gone?
1
u/zenware 15h ago
Totally depends on the use-case of the system. If it’s user workstations and the files are backed up to network storage, then virtually nothing. If it installed in the first place it’s going to do it again, and if /home dirs are on the network it’s not like they’ll suffer much if any data loss.
If the system is a piece of network hardware, then the same level of rigor and config backup you’d give to e.g. an F5 Load Balancer … and so on.
If the system is hosting services reachable from some network connection then there should be an external monitoring mechanism that tells you if it’s working.
If you’re worried about inputs vanishing (they should still be on your system from whenever you last built so this is virtually a non-issue) then you maintain your own mirrors. (Standard practice for business use of Open Source anyway)
So the question to ask yourself in each case is “What do I need this system to do? How fast do I need to get it working again when something breaks? Does my backup enable that functionality in that time window?” If yes, then you’ve done it. When disaster strikes you’ll be able to continue a known portion of your operations within a known timeframe.
1
u/ppen9u1n 5h ago
Well, my data is on servers/NAS being backed up with Borg. Other hosts are fully declarative and can be provisioned from NixOS anywhere, including disko for partitioning and secrets in sops. This works well as long as you have (ideally, but not an absolute requirement ) a second dev host. I once did a recovery of a laptop after nvme suddenly died in about 20min.
If a server drive were to crash, I’d have to restore in 3 steps: 1) NixOS provision (from self contained flake, trivial), 2) restart nomad cluster (from job specs in git), 3) restore databases from borg.
In case my vault server would also crash, I could restore it after 1), since it’s configured as NixOS service with its own database.
8
u/TomKavees 1d ago
Assuming your system builds are actually declarative, including secrets (e.g. using sops-nix) and network configurations (e.g. you can add wifi ssids and passwords to sops secrets, same with samba/cifs/nfs mounts), you can just rebuild the system in a VM. If it works, you are golden, if not, you need to revise your sydtrm config.
For regular user data (NOT SYSTEM IMAGE), you can use regular backup software like Kopia.