HeliosProxy Topology Providers
HeliosProxy Topology Providers
HeliosProxy needs to know which backend node is the current primary so it can route writes correctly and buffer them across a failover. There are two distinct layers to this, and this document keeps them separate because they behave differently:
- The standalone daemon determines the primary from static
[[nodes]]roles plus live health checks (the default), or from an authoritative provider when[topology] provideris set — and reports either at the admin/topologyendpoint. - The
TopologyProviderlibrary abstraction (PrimaryTrackerplus pluggable providers) is the interface the daemon now wires forprovider = "postgres", and is also usable programmatically/embedded.
Every concrete claim below is grounded in src/primary_tracker.rs,
src/admin.rs (compute_topology, TopologyResponse), src/server.rs
(build_primary_tracker, select_primary_until) and the node/config types in
src/config.rs.
Last verified against HeliosProxy 2.0.0.
Layer 1 — How the Standalone Daemon Tracks the Primary
In the running heliosdb-proxy daemon with the default [topology] config, the primary
is not discovered by polling pg_is_in_recovery(). It is the configured [[nodes]]
entry whose role = "primary", that is enabled, and whose health check is currently
passing. Set [topology] provider = "postgres" (feature postgres-topology) to make the
provider authoritative instead — see Layer 2.
Determining the current primary
The write path (select_primary_with_timeout in src/server.rs) looks for
n.role == NodeRole::Primary && n.enabled and a passing health entry. If that node is
healthy, writes go to it. If it is not, the proxy buffers the write, polling health every
100 ms for up to write_timeout_secs (default 30) for a healthy primary to appear; on
timeout it increments the failovers metric and returns NoHealthyNodes. During a
session recovery the same wait runs as select_primary_until, sharing the single
write_timeout_secs deadline with connect/auth, session restore and replay rather than
getting a full window of its own. (See
transaction-replay.md.)
The admin view (compute_topology in src/admin.rs) uses the same rule: currentPrimary
is the address of the first node with role = "primary" (case-insensitive) whose health
entry is healthy = true. None is the correct answer while a failover is in progress and
no primary-role node is healthy.
When a provider is attached ([topology] provider != "static"), compute_topology
and the write path both take the provider’s leader as authoritative, while its lease is
valid. A provider observation is a heartbeat: it refreshes the lease, and authority
expires topology.lease_timeout_secs (default 10) after the last successful observation.
Writes go only to the leased leader (it must be an enabled [[nodes]] entry); while the
provider has no leader or the lease has expired, new writes wait until
write_timeout_secs expires rather than falling back to a configured role. See
Authority leases and fencing limits for what
this does and does not guarantee.
The /topology response
GET /topology returns TopologyResponse (camelCase to map cleanly into the Kubernetes
operator CRD status):
| Field | Meaning |
|---|---|
currentPrimary | Address of the provider’s leader, or (static mode) the first healthy primary-role node; null when neither exists. |
healthyNodes | Count of nodes with a passing health check. |
unhealthyNodes | Count of nodes with a failing health check. |
totalNodes | Number of configured [[nodes]]. |
lastFailoverAt | RFC 3339 timestamp of the last observed primary change; null when none has been observed since boot (currently always null in compute_topology). |
authoritative | Present only when a topology provider is attached and has a leader: {address, epoch, timeline, confirmed, valid, leaseRemainingMs}. epoch increments on every observed leader change; timeline is the leader’s database timeline when the provider reports one (patroni); valid is false once the lease expires (leaseRemainingMs is then absent). |
conflictingPrimariesTotal | Present whenever a topology provider is configured, also while there is no leader: provider polls that found more than one node claiming write authority (resolved by timeline or failed closed). The same count is exported to Prometheus as heliosdb_proxy_topology_conflicting_primaries_total. |
Authority leases and fencing limits
A provider-backed tracker keeps a lease. Every provider observation refreshes the
tracker’s last_refresh; authority is valid only while
now - last_refresh <= topology.lease_timeout_secs. select_primary_until consults
authority_valid() before touching any configured node, so when the provider — or the
control path to it — is unreachable:
- writes keep flowing for at most the lease, so a brief probe gap does not flap traffic;
- after that the tracker drops the stale leader (publishing
PrimaryChanged/Lost) and writes wait, failing withNoHealthyNodesatwrite_timeout_secs— never falling back to a configured role.
The authority epoch is monotonic per process and increments on every observed leader
change, so a tracked address is only ever the provider’s most recent answer. Provider
probes use their own BackendClient connections with the [topology] credentials and
application name (helios-topology), separate from the pooled data path.
What this does not do. The proxy cannot stop a client that connects directly to a
database node, and it cannot prove quorum by itself. Preventing both sides of a partition
from accepting writes requires database-side fencing (Patroni + watchdog, synchronous
replication, pg_promote ordering) or exclusive network/backend control. The lease
bounds the proxy’s own authorization window and refuses stale knowledge; it does not
create zero-loss semantics. A healthy currentPrimary is not a zero-RPO guarantee
for an asynchronous replica that has not acknowledged the last commit — configure
synchronous replication at the database if that is the requirement.
Node configuration and manual control
Nodes are declared with [[nodes]] entries (role = "primary" | "standby" | "replica",
plus host/port/name/enabled). See configuration.md for the
full node schema.
Because the daemon derives the primary from role + health, primary changes are driven by:
- Health checks flipping a node’s
healthyflag (a failing primary drops out;currentPrimarybecomesnulluntil aprimary-role node is healthy again). POST /nodes/{addr}/enable//disable— operator control over which nodes are in rotation.POST /api/chaos— force a node unhealthy (or restore it) to exercise the failover path without external tooling.
To promote a standby in this model, an external HA manager (Patroni, pg_auto_failover,
etc.) performs the promotion, and the proxy’s [[nodes]] roles are updated (config reload
via SIGHUP) or the failed primary node recovers under the same address.
Layer 2 — The TopologyProvider Library Abstraction
src/primary_tracker.rs defines a provider abstraction for automatic primary
tracking. Since 1.9.0 the standalone daemon wires it when [topology] provider is not
static: startup builds a PostgresTopologyProvider over the configured nodes and
spawns PrimaryTracker::run(), and the write path (select_primary_until) consults the
tracker before any configured-role scan. With the default static provider the tracker
is standalone and the historical role+health selection is unchanged.
The PrimaryTracker
PrimaryTracker holds an optional Arc<dyn TopologyProvider> and the current
PrimaryInfo (node_id, address, became_primary_at, is_confirmed, epoch). It can
run in three modes:
- Provider-backed —
PrimaryTracker::with_provider(provider);run()subscribes to the provider’s event stream and updates on eachTopologyEvent. - Standalone —
PrimaryTracker::new_standalone(); the primary is set/cleared explicitly viaset_primary/confirm_primary/clear_primary. - PostgreSQL — pass a
PostgresTopologyProvider(featurepostgres-topology) towith_provider.
The TopologyProvider trait
pub trait TopologyProvider: Send + Sync + 'static { /// Subscribe to topology change events. fn subscribe(&self) -> broadcast::Receiver<TopologyEvent>;
/// Get the current primary node, if one exists. fn get_primary(&self) -> Option<TopologyNodeInfo>;
/// Look up a node by its UUID. fn get_node(&self, id: Uuid) -> Option<TopologyNodeInfo>;}TopologyNodeInfo carries node_id: Uuid, client_addr: String, and is_healthy: bool.
Provider overview
| Provider | Feature Flag | Discovery Method | Latency |
|---|---|---|---|
| PostgreSQL | postgres-topology | Polls pg_is_in_recovery() on each node | Poll interval (default 2 s) |
| HeliosDB | heliosdb-topology | Subscribes to the internal TopologyManager event stream | Event-driven |
| Manual / standalone | (none — always available) | Explicit set_primary() / clear_primary() | Depends on the caller |
PostgreSQL Provider (postgres-topology)
PostgresTopologyProvider (feature postgres-topology) discovers the primary by polling
SELECT pg_is_in_recovery() on each configured node: the node returning false is the
primary; those returning true are standbys/replicas.
How it works
poll_nodes runs each poll interval:
probe_recoveryopens aBackendClientto each node and runsSELECT pg_is_in_recovery().- The first node reporting
falsebecomes the candidate primary. (Taking the first keeps the choice deterministic if a brief split-brain shows two.) - If the primary UUID changed since the previous poll, a
TopologyEvent::PrimaryChanged { old_primary, new_primary }is broadcast. - If a probe errors,
TopologyEvent::HealthChanged { node_id, is_healthy: false }is broadcast for that node.
Construction (programmatic)
The provider is constructed from an explicit Vec<PostgresNode>, not from proxy.toml
[[nodes]] — there is no config wiring that builds a PostgresTopologyProvider in the
daemon. PostgresNode carries node_id, host, port, user, password, database.
use heliosdb_proxy::primary_tracker::{PostgresNode, PostgresTopologyProvider};use std::time::Duration;
let provider = PostgresTopologyProvider::new(nodes) // Vec<PostgresNode> .with_poll_interval(Duration::from_secs(1)) // default is 2s .with_tls_mode(heliosdb_proxy::backend::TlsMode::Prefer);Probe connections are built with a rustls client config from the Mozilla root set;
with_tls_mode sets the TLS policy (default Prefer), connect_timeout is the poll
interval capped at 5 s, and application_name is helios-topology.
Default polling interval
The default poll interval is 2 seconds (Duration::from_secs(2)), tunable with
with_poll_interval. Detection latency is therefore roughly one poll cycle for a lost
primary, plus another cycle after an external HA manager promotes a standby to false.
Compatible HA solutions
Because detection relies solely on pg_is_in_recovery(), the provider works with any HA
solution built on standard PostgreSQL streaming replication — the external manager performs
promotion; the provider detects the resulting change. This includes native streaming
replication, Patroni, pg_auto_failover, Stolon, repmgr, and managed offerings (AWS
RDS/Aurora, Google Cloud SQL, Azure Database for PostgreSQL) that expose replicas answering
pg_is_in_recovery().
Patroni Provider (postgres-topology, provider = "patroni")
Polls a Patroni cluster’s REST API — GET /cluster on each of topology.patroni_endpoints
in order, using the first that answers — and treats the cluster’s leader, as Patroni’s DCS
sees it, as the write primary.
- Only a member with role
leader(or the pre-3.0master) and staterunningcounts. Astandby_leaderleads a standby cluster and is never writable. - The leader’s
host:portmust match a configured[[nodes]]address (host compared case-insensitively). Configure nodes with the same host names or addresses Patroni reports; an unknown leader is not authorized. - Two members reported as running leaders are a conflict: nothing is authorized and the tracker drops its current leader at once.
- When no endpoint answers, the provider reports no primary: the authority lease is not
refreshed and writes stop at the lease boundary (
lease_timeout_secs). - The leader’s timeline is reported as
authoritative.timelineonGET /topology. - Each poll that sees two running leaders counts toward
conflictingPrimariesTotal/heliosdb_proxy_topology_conflicting_primaries_total.
Example:
[topology]provider = "patroni"patroni_endpoints = ["http://pg-a:8008", "http://pg-b:8008", "http://pg-c:8008"]patroni_request_timeout_ms = 2000 # per request; the defaultpoll_interval_secs = 2lease_timeout_secs = 10
# [[nodes]] addresses must match the host:port Patroni reports for each member.Keep lease_timeout_secs below Patroni’s own ttl minus loop_wait so the proxy never
trusts a leader longer than Patroni does (the proxy does not enforce this for you).
Conflicting primaries (postgres provider)
When more than one node answers pg_is_in_recovery() = false — typically an old primary
that kept running after a promotion — the postgres provider reads each node’s timeline
(pg_control_checkpoint().timeline_id) and picks the node on the strictly highest one; a
promotion always starts a new timeline. If the timelines tie, or cannot be read (the probe
role needs permission to call pg_control_checkpoint(): a superuser, or GRANT EXECUTE),
there is no leader and writes fail closed at once. Earlier versions took the first writable
node in probe order.
HeliosDB Provider (heliosdb-topology)
HeliosTopologyProvider<T> (feature heliosdb-topology, in the heliosdb_provider
module) bridges the proxy into HeliosDB’s internal replication TopologyManager instead of
polling. It is defined behind a bridge trait so the standalone proxy can compile without a
hard dependency on the replication crate:
// Implemented by the HeliosDB replication crate.pub trait HeliosTopologyBridge: Send + Sync + 'static { fn subscribe(&self) -> broadcast::Receiver<TopologyEvent>; fn get_primary(&self) -> Option<TopologyNodeInfo>; fn get_node(&self, id: Uuid) -> Option<TopologyNodeInfo>;}
// Adapts the bridge to `TopologyProvider`.pub struct HeliosTopologyProvider<T: HeliosTopologyBridge> { inner: Arc<T>,}Detection is event-driven: when the replication subsystem promotes a standby it emits
TopologyEvent::PrimaryChanged, which PrimaryTracker::run applies immediately. This
provider requires only the feature flag; it is initialized programmatically when the proxy
is built within the HeliosDB workspace.
Manual / Standalone Tracking
With no provider, PrimaryTracker::new_standalone() tracks a primary set entirely through
explicit calls:
use heliosdb_proxy::primary_tracker::PrimaryTracker;
let tracker = PrimaryTracker::new_standalone();
tracker.set_primary(node_id, "pg-primary.local:5432".to_string()); // is_confirmed = falsetracker.confirm_primary(); // is_confirmed = true
// On failover:tracker.clear_primary(); // emits Losttracker.set_primary(new_node_id, "pg-standby.local:5432".to_string());tracker.confirm_primary();This is the mode an external orchestrator would drive: promote out of band, then tell the tracker the new primary’s address.
Primary Lifecycle
PrimaryInfo::is_confirmed encodes a two-phase promotion so writes can begin against a
pending primary before it is fully verified:
set_primary() confirm_primary() ┌──────┐ ─────────────────────▶ ┌─────────┐ ───────────────▶ ┌───────────┐ │ none │ │ pending │ │ confirmed │ └──────┘ ◀───────────────────── └─────────┘ ◀─────────────── └───────────┘ clear_primary() (primary lost / node unhealthy)| State | PrimaryInfo | Meaning |
|---|---|---|
| none | None | No primary known. |
| pending | Some { is_confirmed: false } | Primary set (e.g. mid-switchover), not yet verified. |
| confirmed | Some { is_confirmed: true } | Primary verified and serving. |
Events
TopologyEvent (provider → tracker):
| Event | Trigger |
|---|---|
PrimaryChanged { old_primary, new_primary } | The primary role moved between nodes. |
NodeLeft { node_id } | A node left the cluster. |
HealthChanged { node_id, is_healthy } | A node’s health status changed. |
PrimaryChangeEvent (tracker → subscribers, via PrimaryTracker::subscribe):
| Event | Trigger |
|---|---|
Changed { old, new, address } | A new primary was set (provider event or manual call). |
Lost { old } | The current primary was cleared. |
Confirmed { node_id } | The pending primary was confirmed. |
Custom Topology Providers
Any custom HA source can implement TopologyProvider and be handed to
PrimaryTracker::with_provider. The trait is exactly the three methods above, so a
Consul/Patroni/service-mesh integration only needs to translate its own leader-election
signal into TopologyEvents and answer get_primary / get_node:
struct ConsulTopologyProvider { event_tx: broadcast::Sender<TopologyEvent>, primary: RwLock<Option<TopologyNodeInfo>>, /* consul client, service name, … */}
impl TopologyProvider for ConsulTopologyProvider { fn subscribe(&self) -> broadcast::Receiver<TopologyEvent> { self.event_tx.subscribe() } fn get_primary(&self) -> Option<TopologyNodeInfo> { self.primary.read().clone() } fn get_node(&self, id: Uuid) -> Option<TopologyNodeInfo> { self.primary.read().as_ref().filter(|n| n.node_id == id).cloned() }}
let tracker = PrimaryTracker::with_provider(Arc::new(consul_provider));The crate’s own tests cover a mock provider and a PatroniProvider-shaped custom
implementation, demonstrating the same pattern.
Provider Comparison
| Capability | PostgreSQL | HeliosDB | Manual |
|---|---|---|---|
| Automatic detection | Yes (polling) | Yes (event-driven) | No |
| Detection cadence | Poll interval (default 2 s) | Event-driven | Caller-driven |
| External dependencies | Reachable PG nodes | HeliosDB workspace | External manager |
| Feature flag | postgres-topology | heliosdb-topology | None |
| Wired into standalone daemon | No (programmatic) | No (workspace build) | No (programmatic) |
The standalone daemon’s own primary tracking (Layer 1) is independent of these providers: it uses static
[[nodes]]roles + health checks and reports through/topology.
See Also
- Configuration Reference —
[[nodes]]schema andwrite_timeout_secs. - Transaction Replay — how the primary is used on the write path.
- Admin API Reference —
/topology,/nodes/{addr}/enable|disable,/api/chaos. - Architecture — system overview and module map.