This is the multi-page printable view of this section. .
Barn blog
-
1: Article
- 2: Design Notes
-
3: Release Notes
Articles, design notes, and release updates for Barn.
1 - Article
Project articles and long-form technical notes about Barn live in this section. For current product behavior, use the documentation.
2 - Design Notes
This section explains the architecture and implementation choices in the Barn 0.9.0 candidate.
Start with the product model, then follow the boundaries outward:
- Why Barn has no projects — one Inventory, one owner-scoped deployment, and no second source of truth.
- Why every node has two NICs — fixed identity for the lab, separate from management egress.
- Declarative does not mean destructive — per-node drift with explicit recreate and removal.
- A PID is not a virtual machine — QMP identity, process evidence, journals, and bounded recovery.
repo.yamlis intent;catalog.jsonis evidence — a static image repository whose generated metadata is checked against the actual qcow2 bytes.
Use the documentation for current behavior and the status page for dated verification. These design records explain why those contracts exist; they do not turn a design, build, or local test into release evidence.
2.1 - Why Barn Has No Projects
This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.
Barn began with a familiar VM-manager abstraction: a working directory was a project, a hidden marker gave it identity, a registry found projects again, and a host-global lease kept their private networks from colliding. That model can support many independent VM sets. It was also the wrong model for the product Barn was actually becoming.
The target is not a general-purpose hypervisor front end. It is one local, fixed-IP Pigsty lab. The operator already has a complete description of that lab: the Pigsty Inventory. Adding a second VM manifest and a second project identity made every ordinary question harder.
Decision status: current. This record explains the product model. For current filenames, fields, and commands, use the configuration reference.
The old abstraction turned ordinary questions into project-management questions:
- Which file is authoritative for a node’s name and address?
- Does moving a directory move the lab or create a new one?
- What happens when the marker survives but the registry does not?
- Is a missing directory an abandoned project or an unavailable disk?
- Which project owns the one host network?
Those are legitimate multi-project questions. Barn chose to stop creating them.
One Inventory is enough
The Inventory handed to Pigsty is also Barn’s desired state. Barn reads a
small, documented boundary: host addresses, a few Pigsty-native identity
fields, and the vm_* namespace. Validation is strict inside that namespace;
the rest of the Inventory stays opaque and passes to Pigsty untouched.
This asymmetric rule matters. A misspelled vm_mem must fail because it would
change the machine Barn builds. A new PostgreSQL tuning parameter must not
fail merely because the VM layer has never heard of it. The same file can
therefore evolve as a Pigsty Inventory without becoming a second Barn
format in disguise.
Names follow the same principle. A node uses nodename when present, then a
stable Pigsty cluster/sequence derivation, then its address suffix. Barn does
not add a parallel vm_name that can disagree with the hostname Pigsty sees.
One owner-scoped state root
Applied state lives under BARN_HOME, normally ~/.barn. There is no
marker in the working directory and no per-directory registry. Commands that
operate on applied state can run from any directory; commands that propose new
desired state discover or receive an Inventory explicitly.
This is an owner-scoped deployment, not a root-enforced machine-wide singleton. Each Unix user has an independent state root. Barn is intended for a trusted development workstation, not hostile-user arbitration on a shared server.
The simplification has practical consequences:
- moving or renaming the source directory does not move deployment identity;
- losing the Inventory does not erase the applied state;
- cleanup follows explicit lifecycle commands; deleting the state directory is not a substitute for stopping VMs or removing managed integrations;
- image cache, keys, nodes, disks, locks, and deployment state share one inspectable home. Short QMP/pid paths, SSH client integration, and host-global networking have their own managed locations; see Uninstall for complete cleanup.
Configuration absence is not intent
One deployment does not mean the Inventory is disposable. It means Barn can distinguish desired configuration from applied evidence. When a command needs desired state, an existing deployment can supply its applied spec where the command contract allows that fallback. A missing file is never interpreted as a request to remove nodes.
That rule survives every layer of the lifecycle: planning and deletion stay explicit, and recovery preserves ambiguous resources instead of guessing what the user meant.
The trade-off is deliberate
Barn does not support multiple concurrent deployments per user. Projects, registries, and address-level leasing are not hidden future features; they were rejected because they would reintroduce the abstraction the product removed.
If the requirement changes to multi-tenant or multi-host orchestration, that is a different product boundary. For a local Pigsty lab, one Inventory and one deployment make the important things—addresses, ownership, drift, recovery, and cleanup—much easier to explain and prove.
Read next: Why every Barn node has two NICs.
2.2 - Why Every Barn Node Has Two NICs
This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.
A Pigsty Inventory names machines by stable addresses. PostgreSQL replication,
etcd membership, HAProxy backends, VIPs, monitoring targets, and Ansible all
assume that 10.10.10.11 continues to mean the same node. A loopback port
forward can expose SSH or PostgreSQL to the host, but it cannot provide that
network identity to peers.
This is why Barn does not choose between a convenient NAT mode and an advanced fixed-network mode. Every normal node gets both jobs, on separate interfaces.
One interface should not do two incompatible jobs
| Interface | Addressing | Responsibility |
|---|---|---|
| management | QEMU user-mode NAT, DHCP | DNS, default route, outbound internet, and loopback SSH fallback |
| private, MAC-matched | fixed Inventory address | host-to-node, node-to-node, Ansible, Pigsty services, and VIP traffic |
Barn matches both virtual interfaces by their deterministic MAC addresses
and does not rename them. Guest-visible names are whatever the image’s network
stack chooses (commonly eth0 or enp0s4); the MAC and address contract, not
the display name, determines each interface’s role.
The management interface is deliberately ordinary. It lets a new cloud image reach package repositories before any application exists, without installing NAT rules or a DHCP service on the host.
The private interface has one deterministic RFC1918 address, no default route,
and no DNS. Guest setup checks those properties, but since 0.7 private-network
checks are optional: a failure is recorded as a private-network warning and
does not prevent readiness when management SSH and guest identity are usable.
An operation can therefore succeed with limited fixed-IP connectivity. Before
deploying a service that needs this interface, inspect the warnings and test
the relevant host/peer connection. Repeat up after correcting the cause.
Decision status: current. The topology is part of Barn’s normal lifecycle, not an optional “private mode.” See the design reference for the current platform boundary.
The Inventory owns the address plan
All managed hosts belong to one canonical RFC1918 /24:
| Range | Meaning |
|---|---|
.1 |
host side of the private network |
.2–.8 |
reserved boundary, including valid L2 VIP space |
.9–.254 |
fixed node addresses |
The guest private interface does not request DHCP: cloud-init receives the
exact address already declared in the Inventory. On macOS, vmnet’s configured
DHCP range ends at .8; managed node addresses begin at .9. Linux uses a
bridge with static guest addresses. The VM address contract therefore does not
depend on a lease database, and Ansible uses the same address throughout the
deployment.
For generated configuration, setup can choose from a small bounded set when the default subnet is already occupied. An explicitly supplied Inventory is never silently rewritten to escape a collision. The whole lab—host interface, nodes, Pigsty addresses, and aliases—must agree on one subnet.
Platform-specific backend, identical guest contract
The guest sees the same topology on every supported host, while Barn follows the host’s native networking owner:
- macOS: a pinned
socket_vmnetservice provides the private link. Host mode is the default; shared mode is explicit. QEMU still runs as the user. - Linux with active NetworkManager: Barn creates an owned
barn0bridge throughnmcliand integrates with active firewalld policy. - Linux with systemd-networkd: Barn installs owned units and uses the
distribution
qemu-bridge-helperso QEMU remains unprivileged. - Inactive networkd: activation is allowed only after a pre-mutation scan proves that existing units cannot claim a real host interface.
Supporting both Linux managers is not abstraction for its own sake. Starting
networkd on a desktop or RHEL-family host merely because Barn knows how to
write .network files can disrupt the host’s real network. The backend must
follow the component already in charge.
Host networking is a transaction
A bridge or vmnet daemon outlives one CLI process and crosses a privilege boundary, so Barn treats installation as a reversible host transaction:
- inspect routes, interfaces, services, ownership, and existing state;
- print the exact plan before mutation;
- re-check the preconditions immediately before apply;
- write root-owned state describing what Barn created and what existed before it;
- prove an unprivileged QEMU attachment;
- accept the installation only after readiness checks pass.
A foreign interface that happens to use the same .1/24 is not adopted. A
modified owned file is not overwritten. Uninstall refuses while recorded nodes
are live and restores only the prestate named by the manifest. Failure is a
reason to stop, not permission to delete whatever appears to be in the way.
The accepted trade-off
Guest internet traffic uses QEMU’s user-mode network. That is not the fastest possible forwarding path, but it keeps ordinary egress unprivileged and portable. The traffic that matters to a Pigsty lab—host-to-guest, replication, service calls, and package distribution from a local control node—stays on the private link.
The result is a useful separation of concerns: the management NIC makes a machine easy to bootstrap; the private NIC makes it a stable member of the lab. Neither has to pretend to be the other.
Read next: Declarative does not mean destructive.
2.3 - Declarative Does Not Mean Destructive
This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.
“Declarative” is often shortened to “make reality equal the file.” That is a useful slogan until the file is incomplete, the wrong branch is checked out, or one YAML group is temporarily removed. If absence is treated as deletion, an ordinary editing mistake becomes a destructive operation.
Barn uses a narrower rule:
Desired state may authorize creation. It can describe drift. It never authorizes destruction by omission.
The distinction is central to a VM runtime because roots, data disks, SSH keys, and local evidence are not stateless replicas. Recreating them may be correct, but it must be a decision the operator can see.
From Inventory to node identity
Barn does not hash the whole Pigsty Inventory. It first extracts the fields it owns, fills defaults, canonicalizes image selectors and architecture, and builds a canonical resolved spec. Exact Catalog artifact resolution remains separate, so a channel update does not itself change a node’s spec hash. Each node then receives a hash of:
- the deployment envelope shared by every node, such as subnet, login user, architecture policy, and the deployment-level default image request; and
- exactly that node’s resolved definition.
Adding a peer therefore does not change an existing node’s hash. Editing an unconsumed Pigsty field—PostgreSQL version, packages, or service policy—does not produce VM drift. The VM layer reacts only to the contract it actually understands.
This also avoids a dangerous half-promise: Barn does not pretend to implement
Ansible’s entire variable system. Unknown vm_* keys and conflicting values
inside the owned namespace fail. Everything outside the documented boundary
is opaque rather than partially interpreted.
Decision status: current. Barn converges additions automatically, but definition changes and removal require explicit commands. See Daily Operations for the command workflow.
The five plan outcomes
barn plan compares desired state, applied deployment state, and committed
node state. The result is intentionally small:
| Outcome | Meaning | Apply path |
|---|---|---|
| create | desired node has no committed state | barn up creates it |
| unchanged | definition and runtime still match | running peer stays untouched; stopped peer may start |
| recreate | node definition changed | explicit barn recreate --force <node> |
| missing | applied node is absent or skipped in the Inventory | explicit barn destroy <node> --force, or restore it to the file |
| envelope drift | subnet, login identity, architecture, or runtime policy changed | whole-deployment recreate |
Plan is read-only. It reports the exact node sets and, in text mode, the command that applies the required explicit transition.
Why up stops at drift
Barn could decide that changing CPU or memory is harmless enough to apply,
or that a new image should silently rebuild a root disk. Pre-1.0 intentionally
does neither. A changed VM definition is classified as recreate and up
returns a typed conflict.
That conservative boundary has two advantages:
- all changes that can invalidate Guest state share one visible operation;
- Barn can finish every prerequisite check before touching the current node.
The recreate path resolves the selected emulator, acceleration policy, firmware, image bytes, network backend, shares, and persistent-disk contract before destruction. If a foreign emulator is missing or a share is unsafe, the existing VM remains intact.
Why missing nodes block convergence
A node can disappear from desired state for many reasons that do not express deletion intent:
- the operator opened a reduced Inventory while debugging;
- a group was renamed or filtered;
vm_skiptemporarily marks a real or external host;- a merge conflict dropped a YAML branch;
- the configuration file itself is unavailable.
When applied state contains such a node, up stops and names it. The operator
must either restore the definition or run the explicit destroy command. This
is deliberately more friction than automatic garbage collection—and far less
friction than recovering an unintended disk deletion.
Persistent data disks add another boundary. Normal destroy preserves them.
Whole-deployment destroy --delete-persistent explicitly includes owned
persistent disks, including retained disks, and accepts no node selectors;
deleting deployment keys as well requires whole-deployment
destroy --purge or purge. Retention is not a backup guarantee: guest
bootstrap can reset unrecognized or confirmed damaged test filesystems, even
on persistent disks. See Data disks.
One confirmation cannot silently grow into broader authority.
Convergence is still incremental
Safety does not mean rebuilding everything. New nodes are created without stopping existing peers. Selected stopped nodes start without recreating running ones. A per-node recreate preserves peers and, when requested by the disk contract, persistent data.
The result is declarative where desired state is strong evidence—creation and comparison—and explicit where the cost is irreversible. Barn does not make the operator manually calculate drift, but it also does not confuse a diff with permission.
Read next: A PID is not a virtual machine.
2.4 - A PID Is Not a Virtual Machine
This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.
A pidfile answers one question: which integer did a process have when the file was written? It does not prove that the process is still alive, that the PID was not reused, that the executable is QEMU, or that this particular QEMU owns the node an operator wants to stop.
That is not enough authority for SIGKILL, and certainly not enough authority
to remove a root disk.
Barn treats identity as a chain of independent evidence. Each link has a different job, and destructive action proceeds only when the required links agree.
QMP is the primary runtime identity
Every VM receives a generated UUID and an expected QEMU name. After launch, Barn connects to the QEMU Machine Protocol socket and asks QEMU for both. The VM is not considered started merely because the process returned or a socket path appeared; QMP must report the expected name and UUID.
The same check guards shutdown. A QMP endpoint with a different name or UUID is not “probably the old VM.” It is a hard identity mismatch, and Barn sends no command through it.
QMP also provides the clean path: request Guest powerdown, wait for the Guest, then ask QEMU to quit if the bounded graceful wait expires. Process signals are fallback tools, not the primary lifecycle API.
Decision status: current. This record explains the fail-closed lifecycle boundary. Operational recovery starts with Troubleshooting, not manual deletion.
Process identity closes the fallback gap
QMP may be unavailable after a crash, a damaged runtime directory, or a half-completed shutdown. For that case Barn records a process tuple:
- PID;
- executable path;
- process start time;
- SHA-256 of the observed command line.
The complete typed QEMU invocation is stored beside it. Before sending
SIGTERM, Barn re-reads the live process and requires the tuple to match.
Before escalating to SIGKILL, it captures the tuple again, specifically to
close the PID-reuse window created by the bounded TERM wait.
If QMP still answers but a QMP operation fails, Barn does not bypass that live control plane with a signal. If QMP reports another identity, it stops. If the process tuple cannot be verified, it stops. “Unable to prove” is a result, not a reason to weaken the check.
Unreleased 0.9 candidate update: when QMP is unavailable and the recorded PID is positively identified as an unrelated process, Barn treats the old VM as stopped and never signals that unrelated process. This differs from an unreadable or ambiguous identity, which still blocks the operation.
Journals describe work before state exists
Committed node state cannot describe the earliest part of creation: disks and seed media must exist before the VM can start, and the process must start before its identity can be committed. A crash in that interval would otherwise leave artifacts with no trustworthy owner.
Barn writes a mode-0600 prepare journal first. It contains:
- operation and VM UUIDs;
- node name and resolved-spec hash;
- each completed artifact from a fixed kind/path allowlist;
- the typed QEMU invocation once preparation is complete;
- the exact node-state path once commit succeeds.
The journal is strict versioned JSON. Unknown fields, invalid UUIDs, unsafe paths, repeated artifacts, wrong modes, symlinks, or a path outside the node directory invalidate it.
Recovery can therefore answer a bounded question: which artifacts did this uncommitted operation create? It is not a request to scan the directory and guess.
Rollback is narrower than cleanup
An offline rollback is allowed only when no committed node state exists and no QMP socket or pidfile from the typed invocation remains. The node directory may contain only the journal and the completed allowlisted artifacts. Rollback then removes the completed artifacts in reverse order, followed by the journal and empty node directory.
Unreleased 0.9 candidate update: a failed first up can be retried after
the Inventory is edited. Safe cleanup follows the journal’s completed artifact
list rather than requiring the new desired spec to match the failed old spec.
A committed node is never rolled back by a stale prepare journal. A pre-existing disk is never added to the action list. An unexpected file blocks directory removal instead of being swept up as collateral damage.
Normal destroy follows the same philosophy. It proves containment, file type, ownership boundary, process death, and an exact artifact set. Persistent disks live behind their own preservation and purge rules. Directories are removed only after the known files are gone, so an unknown entry turns into an error.
Atomic state makes the evidence durable
State updates use a same-directory temporary file, fsync, atomic rename, and
parent-directory fsync; symlink targets are rejected. Deployment and node
locks serialize mutations. The goal is not to make crashes impossible—it is
to ensure a crash leaves either an old committed fact or a new committed fact,
plus a journal for the bounded interval between them.
This design is intentionally conservative. It may ask an operator to inspect an ambiguous resource that a more aggressive tool would delete. For a local database lab, preserving the evidence is the safer failure mode.
Read next: repo.yaml is intent; catalog.json is evidence.
2.5 - repo.yaml Is Intent; catalog.json Is Evidence
This article describes the unreleased Barn 0.9.0 candidate. See Status for current validation and remaining release checks.
A static image repository sounds like a directory of qcow2 files plus a JSON index. The difficult part is deciding which facts a maintainer may write by hand and which facts must be derived from the bytes being published.
If checksums and sizes live in the hand-authored source, they are easy to copy incorrectly. If policy exists only in generated JSON, reviewing a channel change or deprecation requires reading machine output. Barn keeps the two jobs separate.
repo.yaml: what the maintainer means
The source-controlled repo.yaml contains author intent:
- repository revision and defaults;
- image families and aliases;
- movable channels such as
stable; - exact versions and architectures;
- boot mode and support status;
- immutable upstream locations and provenance notes.
It deliberately does not contain generated artifact size, SHA-256, or
virtual size. A compact entry can say that d13:stable points to one exact
version with amd64 and arm64 variants without pretending to know facts that
belong to the files. This policy excerpt illustrates the format;
see Image Repositories for a complete working example
and the reference for current versions:
This is the right layer for review: a pull request can show that a channel moved, a version was deprecated, or a provenance statement changed.
Decision status: current. Repository syntax and client behavior are documented in Images; candidate preparation is a separate image-pipeline contract.
catalog.json: what the repository can prove
catalog.json materializes policy against the local repository. For every
variant it records the exact filename, byte count, SHA-256, qcow2 virtual size,
boot contract, source user, and immutable upstream provenance.
Filename, byte count, digest, and virtual size are materialized and checked
against the artifact. Boot mode, source user, status, and provenance are
validated policy copied from repo.yaml; inspection does not independently
prove those declarations. Build forces qcow2 parsing, rejects backing files,
external data,
encryption, and unknown incompatible features, and runs structural checks
before atomically replacing the Catalog.
The artifact identity is the tuple (image, exact version, architecture), not
the channel that selected it. Files keep readable immutable names:
Readable names are an operational feature: an administrator can inspect, mirror, or recover a repository with ordinary filesystem tools. Integrity still comes from generated metadata and verification, not from trusting the name.
Three operations, three responsibilities
The repository CLI keeps observation, generation, and proof separate:
| Command | Responsibility |
|---|---|
barn repo scan |
report tracked, missing, untracked, or unsafe artifacts without changing anything |
barn repo build |
validate source and artifacts, then atomically materialize the Catalog |
barn repo verify |
rebuild the materialization in memory and require byte-for-byte equality with the published Catalog |
build never edits repo.yaml or qcow2 bytes. verify is stronger than
“every checksum is valid”: it also proves that no source policy or artifact
change was omitted from the generated Catalog.
Publication follows the same direction. Upload immutable image bytes first; publish the Catalog and its matching signature last. Publish that pair together where possible. A client refuses a mismatched pair during a partial upload; the order avoids advertising image bytes that are still in transit.
Selectors may move; artifacts may not
Human configuration needs convenient selectors. d13:stable, el9@9, and
el9@9.7 can resolve to newer exact versions as the repository evolves.
Numeric prefixes compare dot-separated components as integers, so 9.10 sorts
after 9.9.
After resolution, the client persists the exact version, architecture, size, and digest. An already resolved node does not become a different machine because a channel moves. Convenience exists at selection time; immutable identity exists at execution time.
Transport and trust are different questions
Official and plain-HTTP Catalogs require a trusted detached signature. An operator who explicitly selects a local directory or HTTPS repository may use an unsigned Catalog because local ownership or authenticated transport is the explicit trust decision. An implicit compiled default remains in the signed trust domain even if its URL is HTTPS.
Accepted Catalog state is tracked independently per repository. Barn rejects unknown keys, a revision below that repository’s high-water mark, and different bytes at the same revision. An explicit downgrade is visible and scoped to the selected repository; resetting to the embedded Catalog does not erase the anti-rollback record.
Catalog acceptance is only the first half. Every pull still checks byte count, SHA-256, and qcow2 structure. Verified base images become read-only, and node root disks are overlays, so normal VM writes never mutate the trusted base.
Separate trust domains stay separate
Image Catalog keys authorize image policy. Release signing proves the Barn application artifacts and checksum manifest. The two key sets are intentionally independent: permission to publish a VM image must not imply permission to ship a new Barn binary, or vice versa.
The Barn 0.9.0 candidate defaults to https://repo.pigsty.io/barn and exposes
--mirror for https://repo.pigsty.cc/barn; --repo remains the explicit
custom override. The two official repositories may
fall back to each other for image downloads, always verifying the same
Catalog size and SHA-256. Custom repositories remain exclusive. Catalog
updates still use the selected source; embedded upstream URLs remain
provenance and never become an artifact fallback. Source configuration,
generated Catalog, uploaded artifacts, signing,
and public availability remain separate release gates.
That is the larger design principle: policy should be pleasant to review, but facts about shipped bytes should be generated, reproducible, and independently verifiable.
3 - Release Notes
Barn 0.9.0 is an unreleased candidate. Release notes will appear here when a version is published. For now, build from source and see Status for validation and remaining release checks.