Skip to content

Raft Consensus Setup

Raft Consensus Setup — Deploy a 3-Node HeliosDB Cluster


UVP

A single HeliosDB node is fast, but it’s also a single point of failure. HeliosDB Full ships a built-in Raft consensus implementation — leader election, log replication, automatic failover. Point three nodes at each other, set replication_factor: 2, and you get sub-second failover, quorum-protected writes, and live cluster metrics out of the box. No ZooKeeper. No etcd. No external coordinator. Same binary, same protocol, same SQL — just three of them.


Prerequisites

  • Three Linux hosts (or three Docker containers on a shared bridge)
  • HeliosDB Full v7.x+ binary on each
  • Bidirectional TCP between nodes on the cluster port (default 5432)
  • About 15 minutes

1. Cluster Components

The cluster layer consists of three components:

ComponentResponsibility
Consensus Manager (Raft)Leader election, log replication, term management
Health CheckerPeriodic node health checks, alert generation
Failover CoordinatorFailure detection, automatic failover, quorum management

2. Pick Your Three Hosts

Before you write a config file, decide:

  • Node IDs — short, stable, unique strings (node1, node2, node3).
  • Listen addresses — what each node binds to (often 0.0.0.0:5432).
  • Peer addresses — DNS or IPs each node uses to reach the others.
  • Replication factor — how many additional copies of each log entry. RF=2 with 3 nodes means every write is on 2 of 3 before it’s acknowledged.

Quorum math is fixed: floor(N/2) + 1 nodes must be reachable to elect a leader and accept writes. With 3 nodes that’s 2.

Cluster sizeToleratesQuorum
31 failure2
52 failures3
73 failures4

Even-sized clusters are wasteful — a 4-node cluster tolerates the same as 3.


3. The Default Timings

Two knobs matter:

  • heartbeat_interval_ms: 100 — how often the leader pings followers. Smaller = faster failure detection, more network chatter.
  • election_timeout_ms: 300 — how long a follower waits without a heartbeat before starting a new election. Must be greater than heartbeat_interval_ms (typically 3-5x), with jitter so two followers don’t time out simultaneously.

The defaults give you sub-second failover on a healthy LAN. For cross-DC, multiply both by 5-10x to absorb latency variance — see multi-region-active-active.md.


4. Bring Up the Cluster

Configure each node with its own node_id and the other two nodes in peer_addrs. Boot order doesn’t matter — the first two to see each other will elect a leader, and the third will join as a follower.


5. Read the Cluster Metrics

Consensus metrics include:

  • log entries appended
  • log entries committed
  • last applied index
  • election count

Health metrics include:

  • last heartbeat per peer
  • node status (Healthy / Degraded / Unreachable)

These are also exposed via the standard observability surface — see Operations Guide for Prometheus and OpenTelemetry wiring.


6. Trigger a Failover (Drill)

The simplest disaster drill is to kill the leader and watch a follower take over.

Terminal window
# On the leader host
sudo systemctl stop heliosdb

With default timings the new leader is elected in 300-600 ms. Bring the dead node back up — it rejoins as a follower and catches up via log replication.


7. Replication Factor and Durability

replication_factor: 2 means:

  • A write is acknowledged once it’s on the leader plus one follower.
  • You can lose any one node and not lose data.
  • A network partition that isolates the leader from both followers will cause writes to stall (quorum lost) — the leader steps down, no split-brain.

Bumping RF=3 (write to all three before ack) gives stronger durability at the cost of latency — every write waits for the slowest follower.


8. Adding a Fourth Node Later

Live membership changes go through the same Raft log. Bring the new node up with all three existing peers in peer_addrs, and existing nodes will pick it up via gossip + Raft membership change. You don’t need to restart the cluster.


Cross-Region Deployments

Three nodes in one DC tolerates host failure. Three nodes across three DCs tolerates DC failure but pays cross-DC latency on every write. For active-active multi-region with sub-100ms cross-region writes, use the dedicated active-active subsystem — see multi-region-active-active.md.


How to Configure

Cluster topology, node membership, replication factor, quorum, and heartbeat and election timing are not set through a public configuration file or command-line interface. Configuration for this feature is provided during onboarding — contact support@heliosdb.com (or sales@heliosdb.com if you are not yet a customer).

The heliosdb.toml file of the HeliosDB Full server accepts only the documented storage keys (storage.data_dir, storage.memtable_size_mb, storage.compaction_strategy, storage.read_cache_mb and the storage.prefetch_* settings). The server refuses to start if the file contains any other key.

Once the cluster is running, check it from SQL:

SHOW RAFT STATUS;
SHOW CLUSTER NODES;
SHOW CLUSTER HEALTH;

Where Next