Skip to content

HeliosProxy Topology Providers

HeliosProxy Topology Providers

HeliosProxy needs to know which backend node is the current primary so it can route writes correctly and buffer them across a failover. There are two distinct layers to this, and this document keeps them separate because they behave differently:

  1. The standalone daemon determines the primary from static [[nodes]] roles plus live health checks (the default), or from an authoritative provider when [topology] provider is set — and reports either at the admin /topology endpoint.
  2. The TopologyProvider library abstraction (PrimaryTracker plus pluggable providers) is the interface the daemon now wires for provider = "postgres", and is also usable programmatically/embedded.

Every concrete claim below is grounded in src/primary_tracker.rs, src/admin.rs (compute_topology, TopologyResponse), src/server.rs (build_primary_tracker, select_primary_until) and the node/config types in src/config.rs.

Last verified against HeliosProxy 2.0.0.


Layer 1 — How the Standalone Daemon Tracks the Primary

In the running heliosdb-proxy daemon with the default [topology] config, the primary is not discovered by polling pg_is_in_recovery(). It is the configured [[nodes]] entry whose role = "primary", that is enabled, and whose health check is currently passing. Set [topology] provider = "postgres" (feature postgres-topology) to make the provider authoritative instead — see Layer 2.

Determining the current primary

The write path (select_primary_with_timeout in src/server.rs) looks for n.role == NodeRole::Primary && n.enabled and a passing health entry. If that node is healthy, writes go to it. If it is not, the proxy buffers the write, polling health every 100 ms for up to write_timeout_secs (default 30) for a healthy primary to appear; on timeout it increments the failovers metric and returns NoHealthyNodes. During a session recovery the same wait runs as select_primary_until, sharing the single write_timeout_secs deadline with connect/auth, session restore and replay rather than getting a full window of its own. (See transaction-replay.md.)

The admin view (compute_topology in src/admin.rs) uses the same rule: currentPrimary is the address of the first node with role = "primary" (case-insensitive) whose health entry is healthy = true. None is the correct answer while a failover is in progress and no primary-role node is healthy.

When a provider is attached ([topology] provider != "static"), compute_topology and the write path both take the provider’s leader as authoritative, while its lease is valid. A provider observation is a heartbeat: it refreshes the lease, and authority expires topology.lease_timeout_secs (default 10) after the last successful observation. Writes go only to the leased leader (it must be an enabled [[nodes]] entry); while the provider has no leader or the lease has expired, new writes wait until write_timeout_secs expires rather than falling back to a configured role. See Authority leases and fencing limits for what this does and does not guarantee.

The /topology response

GET /topology returns TopologyResponse (camelCase to map cleanly into the Kubernetes operator CRD status):

FieldMeaning
currentPrimaryAddress of the provider’s leader, or (static mode) the first healthy primary-role node; null when neither exists.
healthyNodesCount of nodes with a passing health check.
unhealthyNodesCount of nodes with a failing health check.
totalNodesNumber of configured [[nodes]].
lastFailoverAtRFC 3339 timestamp of the last observed primary change; null when none has been observed since boot (currently always null in compute_topology).
authoritativePresent only when a topology provider is attached and has a leader: {address, epoch, timeline, confirmed, valid, leaseRemainingMs}. epoch increments on every observed leader change; timeline is the leader’s database timeline when the provider reports one (patroni); valid is false once the lease expires (leaseRemainingMs is then absent).
conflictingPrimariesTotalPresent whenever a topology provider is configured, also while there is no leader: provider polls that found more than one node claiming write authority (resolved by timeline or failed closed). The same count is exported to Prometheus as heliosdb_proxy_topology_conflicting_primaries_total.

Authority leases and fencing limits

A provider-backed tracker keeps a lease. Every provider observation refreshes the tracker’s last_refresh; authority is valid only while now - last_refresh <= topology.lease_timeout_secs. select_primary_until consults authority_valid() before touching any configured node, so when the provider — or the control path to it — is unreachable:

  • writes keep flowing for at most the lease, so a brief probe gap does not flap traffic;
  • after that the tracker drops the stale leader (publishing PrimaryChanged/Lost) and writes wait, failing with NoHealthyNodes at write_timeout_secs — never falling back to a configured role.

The authority epoch is monotonic per process and increments on every observed leader change, so a tracked address is only ever the provider’s most recent answer. Provider probes use their own BackendClient connections with the [topology] credentials and application name (helios-topology), separate from the pooled data path.

What this does not do. The proxy cannot stop a client that connects directly to a database node, and it cannot prove quorum by itself. Preventing both sides of a partition from accepting writes requires database-side fencing (Patroni + watchdog, synchronous replication, pg_promote ordering) or exclusive network/backend control. The lease bounds the proxy’s own authorization window and refuses stale knowledge; it does not create zero-loss semantics. A healthy currentPrimary is not a zero-RPO guarantee for an asynchronous replica that has not acknowledged the last commit — configure synchronous replication at the database if that is the requirement.

Node configuration and manual control

Nodes are declared with [[nodes]] entries (role = "primary" | "standby" | "replica", plus host/port/name/enabled). See configuration.md for the full node schema.

Because the daemon derives the primary from role + health, primary changes are driven by:

  • Health checks flipping a node’s healthy flag (a failing primary drops out; currentPrimary becomes null until a primary-role node is healthy again).
  • POST /nodes/{addr}/enable / /disable — operator control over which nodes are in rotation.
  • POST /api/chaos — force a node unhealthy (or restore it) to exercise the failover path without external tooling.

To promote a standby in this model, an external HA manager (Patroni, pg_auto_failover, etc.) performs the promotion, and the proxy’s [[nodes]] roles are updated (config reload via SIGHUP) or the failed primary node recovers under the same address.


Layer 2 — The TopologyProvider Library Abstraction

src/primary_tracker.rs defines a provider abstraction for automatic primary tracking. Since 1.9.0 the standalone daemon wires it when [topology] provider is not static: startup builds a PostgresTopologyProvider over the configured nodes and spawns PrimaryTracker::run(), and the write path (select_primary_until) consults the tracker before any configured-role scan. With the default static provider the tracker is standalone and the historical role+health selection is unchanged.

The PrimaryTracker

PrimaryTracker holds an optional Arc<dyn TopologyProvider> and the current PrimaryInfo (node_id, address, became_primary_at, is_confirmed, epoch). It can run in three modes:

  1. Provider-backed — PrimaryTracker::with_provider(provider); run() subscribes to the provider’s event stream and updates on each TopologyEvent.
  2. Standalone — PrimaryTracker::new_standalone(); the primary is set/cleared explicitly via set_primary / confirm_primary / clear_primary.
  3. PostgreSQL — pass a PostgresTopologyProvider (feature postgres-topology) to with_provider.

The TopologyProvider trait

pub trait TopologyProvider: Send + Sync + 'static {
/// Subscribe to topology change events.
fn subscribe(&self) -> broadcast::Receiver<TopologyEvent>;
/// Get the current primary node, if one exists.
fn get_primary(&self) -> Option<TopologyNodeInfo>;
/// Look up a node by its UUID.
fn get_node(&self, id: Uuid) -> Option<TopologyNodeInfo>;
}

TopologyNodeInfo carries node_id: Uuid, client_addr: String, and is_healthy: bool.

Provider overview

ProviderFeature FlagDiscovery MethodLatency
PostgreSQLpostgres-topologyPolls pg_is_in_recovery() on each nodePoll interval (default 2 s)
HeliosDBheliosdb-topologySubscribes to the internal TopologyManager event streamEvent-driven
Manual / standalone(none — always available)Explicit set_primary() / clear_primary()Depends on the caller

PostgreSQL Provider (postgres-topology)

PostgresTopologyProvider (feature postgres-topology) discovers the primary by polling SELECT pg_is_in_recovery() on each configured node: the node returning false is the primary; those returning true are standbys/replicas.

How it works

poll_nodes runs each poll interval:

  1. probe_recovery opens a BackendClient to each node and runs SELECT pg_is_in_recovery().
  2. The first node reporting false becomes the candidate primary. (Taking the first keeps the choice deterministic if a brief split-brain shows two.)
  3. If the primary UUID changed since the previous poll, a TopologyEvent::PrimaryChanged { old_primary, new_primary } is broadcast.
  4. If a probe errors, TopologyEvent::HealthChanged { node_id, is_healthy: false } is broadcast for that node.

Construction (programmatic)

The provider is constructed from an explicit Vec<PostgresNode>, not from proxy.toml [[nodes]] — there is no config wiring that builds a PostgresTopologyProvider in the daemon. PostgresNode carries node_id, host, port, user, password, database.

use heliosdb_proxy::primary_tracker::{PostgresNode, PostgresTopologyProvider};
use std::time::Duration;
let provider = PostgresTopologyProvider::new(nodes) // Vec<PostgresNode>
.with_poll_interval(Duration::from_secs(1)) // default is 2s
.with_tls_mode(heliosdb_proxy::backend::TlsMode::Prefer);

Probe connections are built with a rustls client config from the Mozilla root set; with_tls_mode sets the TLS policy (default Prefer), connect_timeout is the poll interval capped at 5 s, and application_name is helios-topology.

Default polling interval

The default poll interval is 2 seconds (Duration::from_secs(2)), tunable with with_poll_interval. Detection latency is therefore roughly one poll cycle for a lost primary, plus another cycle after an external HA manager promotes a standby to false.

Compatible HA solutions

Because detection relies solely on pg_is_in_recovery(), the provider works with any HA solution built on standard PostgreSQL streaming replication — the external manager performs promotion; the provider detects the resulting change. This includes native streaming replication, Patroni, pg_auto_failover, Stolon, repmgr, and managed offerings (AWS RDS/Aurora, Google Cloud SQL, Azure Database for PostgreSQL) that expose replicas answering pg_is_in_recovery().


Patroni Provider (postgres-topology, provider = "patroni")

Polls a Patroni cluster’s REST API — GET /cluster on each of topology.patroni_endpoints in order, using the first that answers — and treats the cluster’s leader, as Patroni’s DCS sees it, as the write primary.

  • Only a member with role leader (or the pre-3.0 master) and state running counts. A standby_leader leads a standby cluster and is never writable.
  • The leader’s host:port must match a configured [[nodes]] address (host compared case-insensitively). Configure nodes with the same host names or addresses Patroni reports; an unknown leader is not authorized.
  • Two members reported as running leaders are a conflict: nothing is authorized and the tracker drops its current leader at once.
  • When no endpoint answers, the provider reports no primary: the authority lease is not refreshed and writes stop at the lease boundary (lease_timeout_secs).
  • The leader’s timeline is reported as authoritative.timeline on GET /topology.
  • Each poll that sees two running leaders counts toward conflictingPrimariesTotal / heliosdb_proxy_topology_conflicting_primaries_total.

Example:

[topology]
provider = "patroni"
patroni_endpoints = ["http://pg-a:8008", "http://pg-b:8008", "http://pg-c:8008"]
patroni_request_timeout_ms = 2000 # per request; the default
poll_interval_secs = 2
lease_timeout_secs = 10
# [[nodes]] addresses must match the host:port Patroni reports for each member.

Keep lease_timeout_secs below Patroni’s own ttl minus loop_wait so the proxy never trusts a leader longer than Patroni does (the proxy does not enforce this for you).

Conflicting primaries (postgres provider)

When more than one node answers pg_is_in_recovery() = false — typically an old primary that kept running after a promotion — the postgres provider reads each node’s timeline (pg_control_checkpoint().timeline_id) and picks the node on the strictly highest one; a promotion always starts a new timeline. If the timelines tie, or cannot be read (the probe role needs permission to call pg_control_checkpoint(): a superuser, or GRANT EXECUTE), there is no leader and writes fail closed at once. Earlier versions took the first writable node in probe order.

HeliosDB Provider (heliosdb-topology)

HeliosTopologyProvider<T> (feature heliosdb-topology, in the heliosdb_provider module) bridges the proxy into HeliosDB’s internal replication TopologyManager instead of polling. It is defined behind a bridge trait so the standalone proxy can compile without a hard dependency on the replication crate:

// Implemented by the HeliosDB replication crate.
pub trait HeliosTopologyBridge: Send + Sync + 'static {
fn subscribe(&self) -> broadcast::Receiver<TopologyEvent>;
fn get_primary(&self) -> Option<TopologyNodeInfo>;
fn get_node(&self, id: Uuid) -> Option<TopologyNodeInfo>;
}
// Adapts the bridge to `TopologyProvider`.
pub struct HeliosTopologyProvider<T: HeliosTopologyBridge> {
inner: Arc<T>,
}

Detection is event-driven: when the replication subsystem promotes a standby it emits TopologyEvent::PrimaryChanged, which PrimaryTracker::run applies immediately. This provider requires only the feature flag; it is initialized programmatically when the proxy is built within the HeliosDB workspace.


Manual / Standalone Tracking

With no provider, PrimaryTracker::new_standalone() tracks a primary set entirely through explicit calls:

use heliosdb_proxy::primary_tracker::PrimaryTracker;
let tracker = PrimaryTracker::new_standalone();
tracker.set_primary(node_id, "pg-primary.local:5432".to_string()); // is_confirmed = false
tracker.confirm_primary(); // is_confirmed = true
// On failover:
tracker.clear_primary(); // emits Lost
tracker.set_primary(new_node_id, "pg-standby.local:5432".to_string());
tracker.confirm_primary();

This is the mode an external orchestrator would drive: promote out of band, then tell the tracker the new primary’s address.


Primary Lifecycle

PrimaryInfo::is_confirmed encodes a two-phase promotion so writes can begin against a pending primary before it is fully verified:

set_primary() confirm_primary()
┌──────┐ ─────────────────────▶ ┌─────────┐ ───────────────▶ ┌───────────┐
│ none │ │ pending │ │ confirmed │
└──────┘ ◀───────────────────── └─────────┘ ◀─────────────── └───────────┘
clear_primary() (primary lost / node unhealthy)
StatePrimaryInfoMeaning
noneNoneNo primary known.
pendingSome { is_confirmed: false }Primary set (e.g. mid-switchover), not yet verified.
confirmedSome { is_confirmed: true }Primary verified and serving.

Events

TopologyEvent (provider → tracker):

EventTrigger
PrimaryChanged { old_primary, new_primary }The primary role moved between nodes.
NodeLeft { node_id }A node left the cluster.
HealthChanged { node_id, is_healthy }A node’s health status changed.

PrimaryChangeEvent (tracker → subscribers, via PrimaryTracker::subscribe):

EventTrigger
Changed { old, new, address }A new primary was set (provider event or manual call).
Lost { old }The current primary was cleared.
Confirmed { node_id }The pending primary was confirmed.

Custom Topology Providers

Any custom HA source can implement TopologyProvider and be handed to PrimaryTracker::with_provider. The trait is exactly the three methods above, so a Consul/Patroni/service-mesh integration only needs to translate its own leader-election signal into TopologyEvents and answer get_primary / get_node:

struct ConsulTopologyProvider {
event_tx: broadcast::Sender<TopologyEvent>,
primary: RwLock<Option<TopologyNodeInfo>>,
/* consul client, service name, … */
}
impl TopologyProvider for ConsulTopologyProvider {
fn subscribe(&self) -> broadcast::Receiver<TopologyEvent> {
self.event_tx.subscribe()
}
fn get_primary(&self) -> Option<TopologyNodeInfo> {
self.primary.read().clone()
}
fn get_node(&self, id: Uuid) -> Option<TopologyNodeInfo> {
self.primary.read().as_ref().filter(|n| n.node_id == id).cloned()
}
}
let tracker = PrimaryTracker::with_provider(Arc::new(consul_provider));

The crate’s own tests cover a mock provider and a PatroniProvider-shaped custom implementation, demonstrating the same pattern.


Provider Comparison

CapabilityPostgreSQLHeliosDBManual
Automatic detectionYes (polling)Yes (event-driven)No
Detection cadencePoll interval (default 2 s)Event-drivenCaller-driven
External dependenciesReachable PG nodesHeliosDB workspaceExternal manager
Feature flagpostgres-topologyheliosdb-topologyNone
Wired into standalone daemonNo (programmatic)No (workspace build)No (programmatic)

The standalone daemon’s own primary tracking (Layer 1) is independent of these providers: it uses static [[nodes]] roles + health checks and reports through /topology.


See Also