Skip to content

This is the multi-page printable view of this section. .

Return to the regular view of this page.

Design Notes

Architecture decisions, trade-offs, and implementation boundaries behind Barn.

This section explains the architecture and implementation choices in the Barn 0.9.0 candidate.

Start with the product model, then follow the boundaries outward:

  1. Why Barn has no projects — one Inventory, one owner-scoped deployment, and no second source of truth.
  2. Why every node has two NICs — fixed identity for the lab, separate from management egress.
  3. Declarative does not mean destructive — per-node drift with explicit recreate and removal.
  4. A PID is not a virtual machine — QMP identity, process evidence, journals, and bounded recovery.
  5. repo.yaml is intent; catalog.json is evidence — a static image repository whose generated metadata is checked against the actual qcow2 bytes.

Use the documentation for current behavior and the status page for dated verification. These design records explain why those contracts exist; they do not turn a design, build, or local test into release evidence.

1 - Why Barn Has No Projects

Why Barn replaced per-directory project state with one owner-scoped deployment driven by one Pigsty Inventory.
Note

This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.

Barn began with a familiar VM-manager abstraction: a working directory was a project, a hidden marker gave it identity, a registry found projects again, and a host-global lease kept their private networks from colliding. That model can support many independent VM sets. It was also the wrong model for the product Barn was actually becoming.

The target is not a general-purpose hypervisor front end. It is one local, fixed-IP Pigsty lab. The operator already has a complete description of that lab: the Pigsty Inventory. Adding a second VM manifest and a second project identity made every ordinary question harder.

Note

Decision status: current. This record explains the product model. For current filenames, fields, and commands, use the configuration reference.

The old abstraction turned ordinary questions into project-management questions:

  • Which file is authoritative for a node’s name and address?
  • Does moving a directory move the lab or create a new one?
  • What happens when the marker survives but the registry does not?
  • Is a missing directory an abandoned project or an unavailable disk?
  • Which project owns the one host network?

Those are legitimate multi-project questions. Barn chose to stop creating them.

One Inventory is enough

The Inventory handed to Pigsty is also Barn’s desired state. Barn reads a small, documented boundary: host addresses, a few Pigsty-native identity fields, and the vm_* namespace. Validation is strict inside that namespace; the rest of the Inventory stays opaque and passes to Pigsty untouched.

This asymmetric rule matters. A misspelled vm_mem must fail because it would change the machine Barn builds. A new PostgreSQL tuning parameter must not fail merely because the VM layer has never heard of it. The same file can therefore evolve as a Pigsty Inventory without becoming a second Barn format in disguise.

Names follow the same principle. A node uses nodename when present, then a stable Pigsty cluster/sequence derivation, then its address suffix. Barn does not add a parallel vm_name that can disagree with the hostname Pigsty sees.

One owner-scoped state root

Applied state lives under BARN_HOME, normally ~/.barn. There is no marker in the working directory and no per-directory registry. Commands that operate on applied state can run from any directory; commands that propose new desired state discover or receive an Inventory explicitly.

This is an owner-scoped deployment, not a root-enforced machine-wide singleton. Each Unix user has an independent state root. Barn is intended for a trusted development workstation, not hostile-user arbitration on a shared server.

The simplification has practical consequences:

  • moving or renaming the source directory does not move deployment identity;
  • losing the Inventory does not erase the applied state;
  • cleanup follows explicit lifecycle commands; deleting the state directory is not a substitute for stopping VMs or removing managed integrations;
  • image cache, keys, nodes, disks, locks, and deployment state share one inspectable home. Short QMP/pid paths, SSH client integration, and host-global networking have their own managed locations; see Uninstall for complete cleanup.

Configuration absence is not intent

One deployment does not mean the Inventory is disposable. It means Barn can distinguish desired configuration from applied evidence. When a command needs desired state, an existing deployment can supply its applied spec where the command contract allows that fallback. A missing file is never interpreted as a request to remove nodes.

That rule survives every layer of the lifecycle: planning and deletion stay explicit, and recovery preserves ambiguous resources instead of guessing what the user meant.

The trade-off is deliberate

Barn does not support multiple concurrent deployments per user. Projects, registries, and address-level leasing are not hidden future features; they were rejected because they would reintroduce the abstraction the product removed.

If the requirement changes to multi-tenant or multi-host orchestration, that is a different product boundary. For a local Pigsty lab, one Inventory and one deployment make the important things—addresses, ownership, drift, recovery, and cleanup—much easier to explain and prove.

Read next: Why every Barn node has two NICs.

2 - Why Every Barn Node Has Two NICs

Why Barn separates management egress from the fixed-address network used by the host, peers, Ansible, and Pigsty.
Note

This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.

A Pigsty Inventory names machines by stable addresses. PostgreSQL replication, etcd membership, HAProxy backends, VIPs, monitoring targets, and Ansible all assume that 10.10.10.11 continues to mean the same node. A loopback port forward can expose SSH or PostgreSQL to the host, but it cannot provide that network identity to peers.

This is why Barn does not choose between a convenient NAT mode and an advanced fixed-network mode. Every normal node gets both jobs, on separate interfaces.

One interface should not do two incompatible jobs

Interface Addressing Responsibility
management QEMU user-mode NAT, DHCP DNS, default route, outbound internet, and loopback SSH fallback
private, MAC-matched fixed Inventory address host-to-node, node-to-node, Ansible, Pigsty services, and VIP traffic

Barn matches both virtual interfaces by their deterministic MAC addresses and does not rename them. Guest-visible names are whatever the image’s network stack chooses (commonly eth0 or enp0s4); the MAC and address contract, not the display name, determines each interface’s role.

The management interface is deliberately ordinary. It lets a new cloud image reach package repositories before any application exists, without installing NAT rules or a DHCP service on the host.

The private interface has one deterministic RFC1918 address, no default route, and no DNS. Guest setup checks those properties, but since 0.7 private-network checks are optional: a failure is recorded as a private-network warning and does not prevent readiness when management SSH and guest identity are usable. An operation can therefore succeed with limited fixed-IP connectivity. Before deploying a service that needs this interface, inspect the warnings and test the relevant host/peer connection. Repeat up after correcting the cause.

Note

Decision status: current. The topology is part of Barn’s normal lifecycle, not an optional “private mode.” See the design reference for the current platform boundary.

The Inventory owns the address plan

All managed hosts belong to one canonical RFC1918 /24:

Range Meaning
.1 host side of the private network
.2–.8 reserved boundary, including valid L2 VIP space
.9–.254 fixed node addresses

The guest private interface does not request DHCP: cloud-init receives the exact address already declared in the Inventory. On macOS, vmnet’s configured DHCP range ends at .8; managed node addresses begin at .9. Linux uses a bridge with static guest addresses. The VM address contract therefore does not depend on a lease database, and Ansible uses the same address throughout the deployment.

For generated configuration, setup can choose from a small bounded set when the default subnet is already occupied. An explicitly supplied Inventory is never silently rewritten to escape a collision. The whole lab—host interface, nodes, Pigsty addresses, and aliases—must agree on one subnet.

Platform-specific backend, identical guest contract

The guest sees the same topology on every supported host, while Barn follows the host’s native networking owner:

  • macOS: a pinned socket_vmnet service provides the private link. Host mode is the default; shared mode is explicit. QEMU still runs as the user.
  • Linux with active NetworkManager: Barn creates an owned barn0 bridge through nmcli and integrates with active firewalld policy.
  • Linux with systemd-networkd: Barn installs owned units and uses the distribution qemu-bridge-helper so QEMU remains unprivileged.
  • Inactive networkd: activation is allowed only after a pre-mutation scan proves that existing units cannot claim a real host interface.

Supporting both Linux managers is not abstraction for its own sake. Starting networkd on a desktop or RHEL-family host merely because Barn knows how to write .network files can disrupt the host’s real network. The backend must follow the component already in charge.

Host networking is a transaction

A bridge or vmnet daemon outlives one CLI process and crosses a privilege boundary, so Barn treats installation as a reversible host transaction:

  1. inspect routes, interfaces, services, ownership, and existing state;
  2. print the exact plan before mutation;
  3. re-check the preconditions immediately before apply;
  4. write root-owned state describing what Barn created and what existed before it;
  5. prove an unprivileged QEMU attachment;
  6. accept the installation only after readiness checks pass.

A foreign interface that happens to use the same .1/24 is not adopted. A modified owned file is not overwritten. Uninstall refuses while recorded nodes are live and restores only the prestate named by the manifest. Failure is a reason to stop, not permission to delete whatever appears to be in the way.

The accepted trade-off

Guest internet traffic uses QEMU’s user-mode network. That is not the fastest possible forwarding path, but it keeps ordinary egress unprivileged and portable. The traffic that matters to a Pigsty lab—host-to-guest, replication, service calls, and package distribution from a local control node—stays on the private link.

The result is a useful separation of concerns: the management NIC makes a machine easy to bootstrap; the private NIC makes it a stable member of the lab. Neither has to pretend to be the other.

Read next: Declarative does not mean destructive.

3 - Declarative Does Not Mean Destructive

How Barn uses per-node hashes and explicit operations so missing configuration can never authorize deletion.
Note

This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.

“Declarative” is often shortened to “make reality equal the file.” That is a useful slogan until the file is incomplete, the wrong branch is checked out, or one YAML group is temporarily removed. If absence is treated as deletion, an ordinary editing mistake becomes a destructive operation.

Barn uses a narrower rule:

Desired state may authorize creation. It can describe drift. It never authorizes destruction by omission.

The distinction is central to a VM runtime because roots, data disks, SSH keys, and local evidence are not stateless replicas. Recreating them may be correct, but it must be a decision the operator can see.

From Inventory to node identity

Barn does not hash the whole Pigsty Inventory. It first extracts the fields it owns, fills defaults, canonicalizes image selectors and architecture, and builds a canonical resolved spec. Exact Catalog artifact resolution remains separate, so a channel update does not itself change a node’s spec hash. Each node then receives a hash of:

  • the deployment envelope shared by every node, such as subnet, login user, architecture policy, and the deployment-level default image request; and
  • exactly that node’s resolved definition.

Adding a peer therefore does not change an existing node’s hash. Editing an unconsumed Pigsty field—PostgreSQL version, packages, or service policy—does not produce VM drift. The VM layer reacts only to the contract it actually understands.

This also avoids a dangerous half-promise: Barn does not pretend to implement Ansible’s entire variable system. Unknown vm_* keys and conflicting values inside the owned namespace fail. Everything outside the documented boundary is opaque rather than partially interpreted.

Note

Decision status: current. Barn converges additions automatically, but definition changes and removal require explicit commands. See Daily Operations for the command workflow.

The five plan outcomes

barn plan compares desired state, applied deployment state, and committed node state. The result is intentionally small:

Outcome Meaning Apply path
create desired node has no committed state barn up creates it
unchanged definition and runtime still match running peer stays untouched; stopped peer may start
recreate node definition changed explicit barn recreate --force <node>
missing applied node is absent or skipped in the Inventory explicit barn destroy <node> --force, or restore it to the file
envelope drift subnet, login identity, architecture, or runtime policy changed whole-deployment recreate

Plan is read-only. It reports the exact node sets and, in text mode, the command that applies the required explicit transition.

Why up stops at drift

Barn could decide that changing CPU or memory is harmless enough to apply, or that a new image should silently rebuild a root disk. Pre-1.0 intentionally does neither. A changed VM definition is classified as recreate and up returns a typed conflict.

That conservative boundary has two advantages:

  1. all changes that can invalidate Guest state share one visible operation;
  2. Barn can finish every prerequisite check before touching the current node.

The recreate path resolves the selected emulator, acceleration policy, firmware, image bytes, network backend, shares, and persistent-disk contract before destruction. If a foreign emulator is missing or a share is unsafe, the existing VM remains intact.

Why missing nodes block convergence

A node can disappear from desired state for many reasons that do not express deletion intent:

  • the operator opened a reduced Inventory while debugging;
  • a group was renamed or filtered;
  • vm_skip temporarily marks a real or external host;
  • a merge conflict dropped a YAML branch;
  • the configuration file itself is unavailable.

When applied state contains such a node, up stops and names it. The operator must either restore the definition or run the explicit destroy command. This is deliberately more friction than automatic garbage collection—and far less friction than recovering an unintended disk deletion.

Persistent data disks add another boundary. Normal destroy preserves them. Whole-deployment destroy --delete-persistent explicitly includes owned persistent disks, including retained disks, and accepts no node selectors; deleting deployment keys as well requires whole-deployment destroy --purge or purge. Retention is not a backup guarantee: guest bootstrap can reset unrecognized or confirmed damaged test filesystems, even on persistent disks. See Data disks. One confirmation cannot silently grow into broader authority.

Convergence is still incremental

Safety does not mean rebuilding everything. New nodes are created without stopping existing peers. Selected stopped nodes start without recreating running ones. A per-node recreate preserves peers and, when requested by the disk contract, persistent data.

The result is declarative where desired state is strong evidence—creation and comparison—and explicit where the cost is irreversible. Barn does not make the operator manually calculate drift, but it also does not confuse a diff with permission.

Read next: A PID is not a virtual machine.

4 - A PID Is Not a Virtual Machine

Why Barn combines QMP identity, process evidence, typed invocations, and journals before it signals or deletes anything.
Note

This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.

A pidfile answers one question: which integer did a process have when the file was written? It does not prove that the process is still alive, that the PID was not reused, that the executable is QEMU, or that this particular QEMU owns the node an operator wants to stop.

That is not enough authority for SIGKILL, and certainly not enough authority to remove a root disk.

Barn treats identity as a chain of independent evidence. Each link has a different job, and destructive action proceeds only when the required links agree.

QMP is the primary runtime identity

Every VM receives a generated UUID and an expected QEMU name. After launch, Barn connects to the QEMU Machine Protocol socket and asks QEMU for both. The VM is not considered started merely because the process returned or a socket path appeared; QMP must report the expected name and UUID.

The same check guards shutdown. A QMP endpoint with a different name or UUID is not “probably the old VM.” It is a hard identity mismatch, and Barn sends no command through it.

QMP also provides the clean path: request Guest powerdown, wait for the Guest, then ask QEMU to quit if the bounded graceful wait expires. Process signals are fallback tools, not the primary lifecycle API.

Note

Decision status: current. This record explains the fail-closed lifecycle boundary. Operational recovery starts with Troubleshooting, not manual deletion.

Process identity closes the fallback gap

QMP may be unavailable after a crash, a damaged runtime directory, or a half-completed shutdown. For that case Barn records a process tuple:

  • PID;
  • executable path;
  • process start time;
  • SHA-256 of the observed command line.

The complete typed QEMU invocation is stored beside it. Before sending SIGTERM, Barn re-reads the live process and requires the tuple to match. Before escalating to SIGKILL, it captures the tuple again, specifically to close the PID-reuse window created by the bounded TERM wait.

If QMP still answers but a QMP operation fails, Barn does not bypass that live control plane with a signal. If QMP reports another identity, it stops. If the process tuple cannot be verified, it stops. “Unable to prove” is a result, not a reason to weaken the check.

Unreleased 0.9 candidate update: when QMP is unavailable and the recorded PID is positively identified as an unrelated process, Barn treats the old VM as stopped and never signals that unrelated process. This differs from an unreadable or ambiguous identity, which still blocks the operation.

Journals describe work before state exists

Committed node state cannot describe the earliest part of creation: disks and seed media must exist before the VM can start, and the process must start before its identity can be committed. A crash in that interval would otherwise leave artifacts with no trustworthy owner.

Barn writes a mode-0600 prepare journal first. It contains:

  • operation and VM UUIDs;
  • node name and resolved-spec hash;
  • each completed artifact from a fixed kind/path allowlist;
  • the typed QEMU invocation once preparation is complete;
  • the exact node-state path once commit succeeds.

The journal is strict versioned JSON. Unknown fields, invalid UUIDs, unsafe paths, repeated artifacts, wrong modes, symlinks, or a path outside the node directory invalidate it.

Recovery can therefore answer a bounded question: which artifacts did this uncommitted operation create? It is not a request to scan the directory and guess.

Rollback is narrower than cleanup

An offline rollback is allowed only when no committed node state exists and no QMP socket or pidfile from the typed invocation remains. The node directory may contain only the journal and the completed allowlisted artifacts. Rollback then removes the completed artifacts in reverse order, followed by the journal and empty node directory.

Unreleased 0.9 candidate update: a failed first up can be retried after the Inventory is edited. Safe cleanup follows the journal’s completed artifact list rather than requiring the new desired spec to match the failed old spec.

A committed node is never rolled back by a stale prepare journal. A pre-existing disk is never added to the action list. An unexpected file blocks directory removal instead of being swept up as collateral damage.

Normal destroy follows the same philosophy. It proves containment, file type, ownership boundary, process death, and an exact artifact set. Persistent disks live behind their own preservation and purge rules. Directories are removed only after the known files are gone, so an unknown entry turns into an error.

Atomic state makes the evidence durable

State updates use a same-directory temporary file, fsync, atomic rename, and parent-directory fsync; symlink targets are rejected. Deployment and node locks serialize mutations. The goal is not to make crashes impossible—it is to ensure a crash leaves either an old committed fact or a new committed fact, plus a journal for the bounded interval between them.

This design is intentionally conservative. It may ask an operator to inspect an ambiguous resource that a more aggressive tool would delete. For a local database lab, preserving the evidence is the safer failure mode.

Read next: repo.yaml is intent; catalog.json is evidence.

5 - repo.yaml Is Intent; catalog.json Is Evidence

Why Barn separates human-authored image policy from generated metadata that must match the actual qcow2 artifacts.
Note

This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.

A static image repository sounds like a directory of qcow2 files plus a JSON index. The difficult part is deciding which facts a maintainer may write by hand and which facts must be derived from the bytes being published.

If checksums and sizes live in the hand-authored source, they are easy to copy incorrectly. If policy exists only in generated JSON, reviewing a channel change or deprecation requires reading machine output. Barn keeps the two jobs separate.

repo.yaml: what the maintainer means

The source-controlled repo.yaml contains author intent:

  • repository revision and defaults;
  • image families and aliases;
  • movable channels such as stable;
  • exact versions and architectures;
  • boot mode and support status;
  • immutable upstream locations and provenance notes.

It deliberately does not contain generated artifact size, SHA-256, or virtual size. A compact entry can say that d13:stable points to one exact version with amd64 and arm64 variants without pretending to know facts that belong to the files. This policy excerpt illustrates the format; see Image Repositories for a complete working example and the reference for current versions:

defaults: { image: d13, channel: stable, arch: native, boot: uefi }
images:
  d13:
    aliases: [debian13, debian, trixie]
    channels: { stable: "20260810.2566.0" }
    versions:
      "20260810.2566.0":
        status: supported
        variants:
          amd64: {}
          arm64: {}

This is the right layer for review: a pull request can show that a channel moved, a version was deprecated, or a provenance statement changed.

Note

Decision status: current. Repository syntax and client behavior are documented in Images; candidate preparation is a separate image-pipeline contract.

catalog.json: what the repository can prove

catalog.json materializes policy against the local repository. For every variant it records the exact filename, byte count, SHA-256, qcow2 virtual size, boot contract, source user, and immutable upstream provenance.

Filename, byte count, digest, and virtual size are materialized and checked against the artifact. Boot mode, source user, status, and provenance are validated policy copied from repo.yaml; inspection does not independently prove those declarations. Build forces qcow2 parsing, rejects backing files, external data, encryption, and unknown incompatible features, and runs structural checks before atomically replacing the Catalog.

The artifact identity is the tuple (image, exact version, architecture), not the channel that selected it. Files keep readable immutable names:

images/d13-20260810.2566.0-arm64.qcow2

Readable names are an operational feature: an administrator can inspect, mirror, or recover a repository with ordinary filesystem tools. Integrity still comes from generated metadata and verification, not from trusting the name.

Three operations, three responsibilities

The repository CLI keeps observation, generation, and proof separate:

Command Responsibility
barn repo scan report tracked, missing, untracked, or unsafe artifacts without changing anything
barn repo build validate source and artifacts, then atomically materialize the Catalog
barn repo verify rebuild the materialization in memory and require byte-for-byte equality with the published Catalog

build never edits repo.yaml or qcow2 bytes. verify is stronger than “every checksum is valid”: it also proves that no source policy or artifact change was omitted from the generated Catalog.

Publication follows the same direction. Upload immutable image bytes first; publish the Catalog and its matching signature last. Publish that pair together where possible. A client refuses a mismatched pair during a partial upload; the order avoids advertising image bytes that are still in transit.

Selectors may move; artifacts may not

Human configuration needs convenient selectors. d13:stable, el9@9, and el9@9.7 can resolve to newer exact versions as the repository evolves. Numeric prefixes compare dot-separated components as integers, so 9.10 sorts after 9.9.

After resolution, the client persists the exact version, architecture, size, and digest. An already resolved node does not become a different machine because a channel moves. Convenience exists at selection time; immutable identity exists at execution time.

Transport and trust are different questions

Official and plain-HTTP Catalogs require a trusted detached signature. An operator who explicitly selects a local directory or HTTPS repository may use an unsigned Catalog because local ownership or authenticated transport is the explicit trust decision. An implicit compiled default remains in the signed trust domain even if its URL is HTTPS.

Accepted Catalog state is tracked independently per repository. Barn rejects unknown keys, a revision below that repository’s high-water mark, and different bytes at the same revision. An explicit downgrade is visible and scoped to the selected repository; resetting to the embedded Catalog does not erase the anti-rollback record.

Catalog acceptance is only the first half. Every pull still checks byte count, SHA-256, and qcow2 structure. Verified base images become read-only, and node root disks are overlays, so normal VM writes never mutate the trusted base.

Separate trust domains stay separate

Image Catalog keys authorize image policy. Release signing proves the Barn application artifacts and checksum manifest. The two key sets are intentionally independent: permission to publish a VM image must not imply permission to ship a new Barn binary, or vice versa.

The Barn 0.9.0 candidate defaults to https://repo.pigsty.io/barn and exposes --mirror for https://repo.pigsty.cc/barn; --repo remains the explicit custom override. The two official repositories may fall back to each other for image downloads, always verifying the same Catalog size and SHA-256. Custom repositories remain exclusive. Catalog updates still use the selected source; embedded upstream URLs remain provenance and never become an artifact fallback. Source configuration, generated Catalog, uploaded artifacts, signing, and public availability remain separate release gates.

That is the larger design principle: policy should be pleasant to review, but facts about shipped bytes should be generated, reproducible, and independently verifiable.