High availability: state inventory and design¶
infrabroker is deliberately single-instance today. This document is the design study behind that choice (#145): what actually blocks running two replicas, what each blocker costs, and what a realistic first slice looks like. It is a map, not a commitment — the work is demand-gated, and each blocker below is tracked as its own issue.
The honest summary: multi-replica HA is a re-platforming, not a feature flag. It touches the persistence layer, the session model, the behavior tracker and the audit chain, in three different processes.
What exists today¶
infrabroker runs as three separate processes, each with its own in-memory state and its own audit log:
| Process | Package | Holds |
|---|---|---|
| broker | internal/broker |
live SSH sessions, host/cluster caches |
| control plane | cmd/control-plane, internal/control |
approvals, behavior baselines |
| signer | cmd/signer, internal/signer |
CA key, policy, grants, freezes, rate limits |
state_db (internal/statedb) exists on the signer and the control plane
only; the broker has none. It is a thin opener over modernc.org/sqlite (pure Go,
no CGO) against a local file, in WAL mode, capped at a single connection
because every consumer mutates under its own mutex.
Its architectural note is the thing to internalise before designing HA:
the database is an availability enhancement, not a decision-path component: consumers keep their in-memory state as the source of truth on the hot path and mirror mutations here (write-through).
So even where state is persisted, the in-memory map is what every request reads; the DB is only re-read at startup. That buys restart survival, not sharing — and it is the invariant HA has to invert.
Pointing two replicas at the same SQLite file is not a shortcut: WAL needs
-wal/-shm sidecars and local POSIX locking, and the single-connection design
assumes a single writer process.
State inventory¶
| Subsystem | Where state lives | Persisted? | What breaks with 2 replicas | Difficulty |
|---|---|---|---|---|
| Live sessions (broker) | internal/broker/session.go — sessionManager{sessions map, mu} |
No (broker has no state_db) | Open on A returns a session_id; exec/close routed to B → "unknown or expired session". A session owns a live SSH connection/PTY. |
Hard — not a storage problem |
| Approvals (control plane) | internal/control/approval.go — Registry{items map, mu, db} |
Yes, write-through | Waiter polls A, approver's decision lands on B → the request times out. issuing/consumed are per-process, so four-eyes/consumed-once can double-issue. |
Medium |
| Grants (signer) | internal/signer/grants.go — GrantStore{grants map, mu, db} |
Yes, write-through | A grant added via A is invisible to B's decision path → approve-and-learn is inconsistent by which replica you hit. | Medium |
| Freeze / kill switch (signer) | internal/signer/freeze.go — FreezeStore{frozen map, mu, db} |
Yes, durable + fail-closed | A freeze on A is unknown to B: B keeps signing for a frozen subject, and brokers polling B never get the kill. | Medium-hard (security-critical) |
| Behavior tracker (control plane) | internal/control/behavior.go — BehaviorTracker{subjects map, mu} |
No — no DB variant exists | Each replica sees only its slice of traffic: baselines split-brain, and the per-subject sliding-window rate limit effectively multiplies by replica count. | Hard |
| Rate limiter (signer) | internal/signer/ratelimit.go |
No | N replicas = N× the configured per-caller limit. Weakening, not a correctness break. | Easy |
| Host-key cache (broker) | internal/broker/engine.go — sync.Map, content-addressed |
No | Nothing: deterministic, self-heals per instance. | Trivial |
| Host/cluster caches (broker) | internal/broker/engine.go, refreshed by poller |
No | Nothing: the signer is the source of truth; each replica refetches. | Trivial |
| Compiled command policy (signer) | in-memory, rebuilt from config on reload | No (config-derived) | Nothing, as long as every replica reads the same config. | Trivial |
| Audit chain (all three) | internal/audit/log.go — single-writer O_APPEND file, chain head in memory |
Own file | Each replica writes an independent hash chain, seq restarting at 1. No global order, no single verifiable chain. |
Hard |
Singleton assumptions¶
There is no leader election, no file locking, no "already running" guard anywhere. The background goroutines (session reaper, revocation poll, grant purge, config watchers) are per-process and harmless to run N times — but several enforce global-looking invariants only locally.
One is worth calling out as already HA-shaped: the broker's revocation poll
(GET /v1/revocations, ~10s) is a pull/eventual-consistency propagation model
that survives multi-replica brokers as long as the freeze set it polls is
itself global. Behind an LB in front of multiple signers it silently breaks — so
freeze-sharing is the real dependency, not the poll loop.
Note also that N signer replicas each need CA-key access (shared Key Vault, or the same PEM), which multiplies the trust surface and raises serial allocation.
The four true blockers¶
- Live sessions. The hardest item, and not a storage problem: a session
owns an open SSH connection/PTY that cannot be serialised between processes.
The realistic answer is session affinity — pin
session_idback to the replica that dialed — not shared session state. One-shotExecuteis already stateless and scales freely today. - Freeze / kill switch. Security-critical: a frozen subject must be denied on every replica, which means shared, read-committed freeze state on the sign path — not a local cache.
- Audit chains. Per-process hash chains don't compose into one verifiable history. Needs a strategy decision: per-replica chains stitched by a collector, or an external append-only sink with a single writer/sequencer.
- Behavior tracker. Purely in-memory with no persistence path at all; split-brain baselines and replica-multiplied rate windows.
Medium (mechanical-ish): approvals and grants. Both already do insert-first
write-through against *sql.DB and reload at startup, so the interfaces isolate
the store. The real delta is that each currently treats its in-memory map as the
source of truth on the decision path — HA requires read-through (or subscribing
to invalidations) plus making the consumed/issuing claims DB-transactional
rather than in-process booleans. That inverts the statedb invariant quoted
above, and it is the largest single piece of the work.
Easy / leave alone: host-key and host/cluster caches, compiled policy (all deterministic or config-derived, re-derivable per replica). The rate limiter is per-instance and merely loosens the global bound — accept it, or back it with a shared counter only if a hard cap is contractually required.
Minimum viable HA slice¶
In dependency order, each tracked as its own issue under the Medium-term: HA, durability & secure-by-default milestone:
- Shared transactional backend (e.g. Postgres) for approvals + grants + freezes, replacing local SQLite, with read-through on the decision path and transactional consume/four-eyes. This is the core.
- Broker session affinity
—
session_id-based routing, not shared state. - Distributed behavior tracker — decided: explicit degradation, documented rather than engineered. See "Decided: budgets degrade, and that is the boundary" below.
- HA audit strategy — per-replica chains plus a central verifier, or an external sequencer.
Optional and non-blocking, tracked together: a globally shared rate limiter, and a decision on multi-signer CA custody + serial allocation — decided below.
Decided: budgets degrade, and that is the boundary¶
Two of the items above are closed by a decision rather than by code (#295, #297). Both concern budgets — how much an agent may do — and in both cases the answer is to accept a looser bound under replication and say so, rather than put a shared datastore in the path of every signature.
The reasoning is the same for both: budgets are a detection and damage-limiting layer, not the containment boundary. Containment is the signer's command policy and the approval gate; those are authoritative on every replica because they are derived from the same config, not from accumulated state. A rate limit that admits N× under N replicas is weaker than intended but still bounds a runaway agent; a behaviour baseline that split-brains produces extra escalations, not missed ones. Neither failure mode lets an agent run something policy forbids.
Behaviour tracker (#295). Beyond the split-brain baselines and the N× sliding window already noted above, two facets worth stating because they are not obvious from the table:
Learnteaches only the replica that handled the approval, so an approved-and-learned anomaly stays novel on the other N−1 replicas. The same deviation can be escalated to a human once per replica before it is learned everywhere — noisier, never more permissive.- The sliding window deliberately counts blocked attempts so a flood cannot evade the cap by absorbing rejections. That property is per process too, so under N replicas a flooder's budget is also N×.
Sign-rate limiter (#297). sign_rate_limit_per_min (signer.json) is a
per-CN token bucket in internal/signer/ratelimit.go, keyed on the
authenticated mTLS peer CN and held in memory. N replicas therefore admit up to
N× the configured cap. Size it as desired_total / N, or front the signers with
a global limiter, and treat the config value as a per-replica cap. Backing the
buckets with a shared counter is worth it only if a hard global cap is
contractually required — it would put a network dependency in the path of every
signature, which is the component that must not wobble.
Multi-signer CA custody (#297). Every signer replica needs CA-key access, so custody choice constrains replication:
akv(Azure Key Vault) replicates cleanly — each replica authenticates independently with its own managed identity, and the key never leaves the vault. This is the recommended backend for a replicated signer.agent(ssh-agent: YubiKey PIV / SoftHSM / TPM) is host-local — it is a unix socket on one machine. It does not compose with N replicas without one hardware token per replica (each a separate CA key, which multiplies the trust surface) or a socket-forwarding arrangement that defeats the point. Excellent for a single signer; not a replication story.pemis lab-only regardless.
Certificate serials (#297). randomSerial draws a full 64-bit value from
crypto/rand per issuance (internal/ca/sign.go), and Kubernetes issuances mint
audit serials from the same space. Replication does not change the collision
math: the birthday bound depends on how many certificates exist, not on how many
processes minted them, and no replica identity is mixed in. By that bound,
collision probability stays below 1e-6 up to ~6.1 million certificates, reaches
~0.03% at 10^8 and 50% only at ~5×10^9.
A collision would not break audit correlation — every entry also carries time,
caller, host, session id and outcome, so --serial degrades from a unique key to
an ambiguous one. The materially affected control is the kill switch: freezing a
serial matches by string, so two simultaneously live certificates sharing one
would be closed together. That population is bounded by the 15-minute TTL cap,
not by lifetime issuance — at 10,000 concurrently valid certificates the
probability is ~3e-12. No action needed; partitioning serials per replica would
buy nothing measurable.
Why it stays demand-gated¶
Single-instance is a deliberate choice, not an oversight. state_db already
delivers the property most deployments actually want — restart survival — and
the failure mode of a single broker is a bounded outage, not a security event.
Real HA buys availability at the cost of a shared datastore in the decision path
of every signature, which is a new dependency for the component whose whole job is
to fail closed. Build it when demand is observed, and build the slice above in
order.