Raft Consensus Setup
Raft Consensus Setup — Deploy a 3-Node HeliosDB Cluster
UVP
A single HeliosDB node is fast, but it’s also a single point of failure. HeliosDB Full ships a built-in Raft consensus implementation — leader election, log replication, automatic failover. Point three nodes at each other, set replication_factor: 2, and you get sub-second failover, quorum-protected writes, and live cluster metrics out of the box. No ZooKeeper. No etcd. No external coordinator. Same binary, same protocol, same SQL — just three of them.
Prerequisites
- Three Linux hosts (or three Docker containers on a shared bridge)
- HeliosDB Full v7.x+ binary on each
- Bidirectional TCP between nodes on the cluster port (default 5432)
- About 15 minutes
1. Cluster Components
The cluster layer consists of three components:
| Component | Responsibility |
|---|---|
| Consensus Manager (Raft) | Leader election, log replication, term management |
| Health Checker | Periodic node health checks, alert generation |
| Failover Coordinator | Failure detection, automatic failover, quorum management |
2. Pick Your Three Hosts
Before you write a config file, decide:
- Node IDs — short, stable, unique strings (
node1,node2,node3). - Listen addresses — what each node binds to (often
0.0.0.0:5432). - Peer addresses — DNS or IPs each node uses to reach the others.
- Replication factor — how many additional copies of each log entry. RF=2 with 3 nodes means every write is on 2 of 3 before it’s acknowledged.
Quorum math is fixed: floor(N/2) + 1 nodes must be reachable to elect a leader and accept writes. With 3 nodes that’s 2.
| Cluster size | Tolerates | Quorum |
|---|---|---|
| 3 | 1 failure | 2 |
| 5 | 2 failures | 3 |
| 7 | 3 failures | 4 |
Even-sized clusters are wasteful — a 4-node cluster tolerates the same as 3.
3. The Default Timings
Two knobs matter:
heartbeat_interval_ms: 100— how often the leader pings followers. Smaller = faster failure detection, more network chatter.election_timeout_ms: 300— how long a follower waits without a heartbeat before starting a new election. Must be greater thanheartbeat_interval_ms(typically 3-5x), with jitter so two followers don’t time out simultaneously.
The defaults give you sub-second failover on a healthy LAN. For cross-DC, multiply both by 5-10x to absorb latency variance — see multi-region-active-active.md.
4. Bring Up the Cluster
Configure each node with its own node_id and the other two nodes in peer_addrs. Boot order doesn’t matter — the first two to see each other will elect a leader, and the third will join as a follower.
5. Read the Cluster Metrics
Consensus metrics include:
- log entries appended
- log entries committed
- last applied index
- election count
Health metrics include:
- last heartbeat per peer
- node status (Healthy / Degraded / Unreachable)
These are also exposed via the standard observability surface — see Operations Guide for Prometheus and OpenTelemetry wiring.
6. Trigger a Failover (Drill)
The simplest disaster drill is to kill the leader and watch a follower take over.
# On the leader hostsudo systemctl stop heliosdbWith default timings the new leader is elected in 300-600 ms. Bring the dead node back up — it rejoins as a follower and catches up via log replication.
7. Replication Factor and Durability
replication_factor: 2 means:
- A write is acknowledged once it’s on the leader plus one follower.
- You can lose any one node and not lose data.
- A network partition that isolates the leader from both followers will cause writes to stall (quorum lost) — the leader steps down, no split-brain.
Bumping RF=3 (write to all three before ack) gives stronger durability at the cost of latency — every write waits for the slowest follower.
8. Adding a Fourth Node Later
Live membership changes go through the same Raft log. Bring the new node up with all three existing peers in peer_addrs, and existing nodes will pick it up via gossip + Raft membership change. You don’t need to restart the cluster.
Cross-Region Deployments
Three nodes in one DC tolerates host failure. Three nodes across three DCs tolerates DC failure but pays cross-DC latency on every write. For active-active multi-region with sub-100ms cross-region writes, use the dedicated active-active subsystem — see multi-region-active-active.md.
How to Configure
Cluster topology, node membership, replication factor, quorum, and heartbeat and election timing are not set through a public configuration file or command-line interface. Configuration for this feature is provided during onboarding — contact support@heliosdb.com (or sales@heliosdb.com if you are not yet a customer).
The heliosdb.toml file of the HeliosDB Full server accepts only the documented storage keys (storage.data_dir, storage.memtable_size_mb, storage.compaction_strategy, storage.read_cache_mb and the storage.prefetch_* settings). The server refuses to start if the file contains any other key.
Once the cluster is running, check it from SQL:
SHOW RAFT STATUS;SHOW CLUSTER NODES;SHOW CLUSTER HEALTH;Where Next
- sharding-config.md — scale beyond a single Raft group.
- pitr-recovery.md — point-in-time recovery on top of the Raft log.