HeliosProxy Documentation
HeliosProxy Documentation
Programmable Postgres Data-Plane
What is HeliosProxy?
HeliosProxy is a programmable Postgres data-plane: a PostgreSQL-wire connection router with a real WASM plugin runtime, signed plugin distribution via OCI, a transaction journal with operator-driven replay, a zero-downtime PostgreSQL major-version upgrade orchestrator, and a built-in admin web console.
The current release is 2.0.0 (September 25, 2026) — see Upgrading to 2.0 and the 2.0 announcement. Capabilities are organised into two tiers:
- Connection-Routing Tier — read/write splitting, health checking, circuit breaking, rate limiting, connection pooling (session / transaction / statement), and failover.
- Platform Tier — WASM plugins, query caching, query analytics, multi-tenancy, authentication, query rewriting, GraphQL and MCP gateways, anomaly detection, and edge/geo caching.
Every tier-two capability is a Cargo feature flag, so the binary only carries what you enable — see Feature Flags for the authoritative list, which tracks the [features] table in Cargo.toml flag for flag.
Since 1.0.0 the project’s stated policy is that every shipped feature flag does real work, with intentionally-bounded capabilities documented rather than implied. Where a subsystem is library-only or not yet mounted on the per-query data path, this documentation says so explicitly.
Quick Start
Run with Docker
docker pull ghcr.io/heliosdatabase/hdb-heliosdb-proxy:2.0.0
docker run -d \ --name heliosproxy \ -p 5432:5432 \ -p 9090:9090 \ -v $(pwd)/proxy.toml:/etc/heliosproxy/proxy.toml \ ghcr.io/heliosdatabase/hdb-heliosdb-proxy:2.0.0Published tags follow the release tags (2.0.0, 2.0, latest). Note that the admin API binds loopback by default as of 1.4.0, so reach it from inside the container or publish it deliberately as above.
Install from crates.io
cargo install heliosdb-proxyThe minimum supported Rust version is 1.86.
Configure
Minimal proxy.toml:
listen_address = "0.0.0.0:5432"admin_address = "127.0.0.1:9090"tr_enabled = truetr_mode = "session"
[[nodes]]name = "primary-1"host = "pg-primary.example.com"port = 5432role = "primary"
[[nodes]]name = "standby-1"host = "pg-standby.example.com"port = 5432role = "standby"The sections [pool], [load_balancer], [health] and at least one [[nodes]] entry are required in a configuration file; everything else defaults to off. Config files support ${VAR} and ${VAR:-default} environment substitution (1.4.0), and a ${VAR} with no default and no environment value fails fast at load.
See Configuration for the full reference, and the repo’s config/ directory for complete example files.
Topology
┌─────────────┐ │ HeliosProxy │ ← WASM plugins, replay engine, │ v2.0.0 │ upgrade orchestrator, admin UI └──────┬──────┘ │ PostgreSQL wire protocol ┌────────────────┼────────────────┐ ▼ ▼ ▼ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ Primary │ │ Standby │ │ Replica │ └─────────┘ └─────────┘ └─────────┘The upgrade orchestrator’s validation stage uses a row-count parity check that is portable across PostgreSQL 14–18. Release verification runs against PostgreSQL 18.4.
Documentation Structure
- Architecture — System design, request lifecycle, hook points
- Configuration — Full
proxy.tomlreference, including[limits],[anomaly],[cache],[auth],[tls]and[[hba]] - Feature Flags — Cargo features and module activation
- Admin API — REST endpoints (topology, plugins, plugin KV, anomalies, edge, chaos, migration, replay, metrics)
- Topology Providers — Backend discovery and write authority (static roles, PostgreSQL polling with the timeline rule, Patroni, HeliosDB native)
- Transaction Replay — In-session failover (
tr_mode), the committed-transaction journal ([journal]) and operator replay (POST /api/replay, includingcommitted_history); all in the default build - Deployment Guides — Docker, Kubernetes, standalone
Source Repositories
| Repository | Role | License |
|---|---|---|
HeliosDB-Proxy | Core proxy (this is where most files live) | Apache-2.0 |
HeliosDB-Proxy-Plugins | First-party WASM plugins + helios-plugin CLI | Apache-2.0 |
HeliosDB-Proxy-Operator | Kubernetes operator + Helm chart | Not currently public |
terraform-provider-HeliosDB-Proxy | Terraform provider | Not currently public |
pulumi-HeliosDB-Proxy | Pulumi provider | Not currently public |
Container image: ghcr.io/heliosdatabase/hdb-heliosdb-proxy:2.0.0 (lowercase due to GHCR convention). The project was published under the Dimensigon organisation before 0.5.1; those older paths are no longer valid.
Upgrading to 2.0
HeliosProxy 2.0.0 keeps the 1.x configuration format: a 1.9 proxy.toml loads unchanged.
The major version marks behaviour changes you should review before rolling it out:
- Gateways have a concurrency cap by default.
[graphql_gateway],[mcp]and[http_gateway]admit at mostmax_concurrent_requests = 64requests at once. Beyond that the HTTP and GraphQL gateways answer503withRetry-After, and MCP answers a JSON-RPC error (-32000) fortools/call; health checks and MCPinitialize/ping/tools/listare never gated. Raise the key if you run more concurrent gateway requests.query_timeout_ms(default30000) replaces the fixed 30-second gateway timeout. - Unknown config keys are reported at any nesting depth. A typo inside a section
(
[cache] l1_max_byts) now logsWARN unknown config key 'cache.l1_max_byts' ignoredat startup, and fails startup whenstrict_config = true. Previously only top-level keys were checked, so check your startup log after upgrading. - Conflicting primaries fail closed with
[topology] provider = "postgres". When several nodes are writable, the node on the strictly highest timeline wins; a tie or an unreadable timeline authorizes nothing (1.9 took the first writable node probed). The probe role needs permission to callpg_control_checkpoint(). See Topology Providers. - Query cache (only if you enabled it). Writes on the extended protocol now
invalidate it, a stale entry is refused on every tier (L1 and L3 used to be served
until TTL), and DDL,
TRUNCATE,COPY ... FROM,EXECUTE,CALLandDOinvalidate the whole cache. Expect a lower hit rate on write-heavy workloads in exchange for read-your-commits correctness. The transaction-journal capture now also runs whentr_enabled = falsebut the cache is on. The cache stays off by default; see why. - Library users only: the
TopologyProvidertrait gains defaultedconflicts_total()andleader_timeline()methods (existing implementations keep compiling), andQueryCache::purge_tablesno longer removes L2 entries itself.
New in 2.0 that you may want to turn on: [topology] provider = "patroni"
(Topology Providers),
and the Prometheus counters heliosdb_proxy_topology_conflicting_primaries_total,
heliosdb_proxy_backend_capacity_waits_total and
heliosdb_proxy_backend_capacity_refusals_total.
Coming from 1.7 or earlier, also note the 1.8 change (ha-tr is a no-op; Transaction
Replay is always compiled and tr_enabled is the runtime switch) and that 1.9 added the
[journal], [topology], strict_config and [limits] keys described in
Configuration — all default to the previous behaviour.
What’s New
Highlights since 0.4.0 — see the repo CHANGELOG for the complete history.
- 2.0.0 (2026-09-25) — The HA topology, the query cache and Transaction Replay share one model of what the backend actually committed. The query cache (still opt-in) invalidates on commit on both wire protocols and every tier, with DDL,
TRUNCATE,COPY ... FROM,EXECUTE,CALLandDOinvalidating the whole cache; a Patroni cluster can be the write authority ([topology] provider = "patroni"); conflicting primaries under thepostgresprovider are resolved by the strictly highest timeline or fail closed; gateways gainquery_timeout_msandmax_concurrent_requests; unknown config keys are reported at any nesting depth; a login burst against a backend atmax_connectionsno longer surfaces as a false authentication failure;--tr falseworks. Read the upgrade notes and the announcement. - 1.9.0 (2026-09-22) — Recovery-grade transaction journal: real committed transactions with parameters and outcomes, an optional durable on-disk journal (
[journal] dir, CRC-checked segments, torn-tail truncation,commit/interval/nonefsync), andPOST /api/replaymode: "committed_history". Authoritative topology ([topology] provider = "postgres") with an authority lease (lease_timeout_secs);GET /capabilitiesand opt-instrict_config;[load_balancer] read_strategyhonoured (newpower_of_two); credentialed health probes and measured replica lag (require_known_lag); bounded fair client admission (client_admission_wait_secs); pooled gateway backend connections (gateway_pool_max_idle); byte budgets for the L1/L3 and edge caches;replay_deadline_secs; rustls 0.23.45 (RUSTSEC-2026-0285). - 1.8.0 (2026-09-12) — Transaction Replay ships in the default build; the
ha-trcargo feature is a deprecated no-op andtr_enabledis the runtime switch (falsestops journaling, forcestr_mode = "none", and makesPOST /api/replayreturn503). Dependency security updates (reqwest0.12,wasmtime36,lru0.18); the unusedprometheus/opentelemetrydependencies were dropped.GET /api/migration/statusno longer reportsmigration_readywhile applies are failing. - 1.7.0 — Completes the Transaction Replay series: replay verifies each statement’s response digest against what the client already saw (divergence rolls back with
40001),SERIALIZABLE/REPEATABLE READtransactions are never replayed, an interrupted read is only re-executed when every function it calls is provably side-effect-free (built-in allowlist plus the newtr_read_functions), sessionSETrestore is transactional and savepoint-scoped, exceedingtr_max_session_set_statementsrefuses the failover with08006rather than re-homing with partial state, andwrite_timeout_secsis now one deadline for the entire recovery. New[limits]bounds:tr_max_observation_bytes,backend_response_timeout_secs,max_backend_frame_bytes. - 1.6.1 — Transaction Replay commit-outcome and delivery safety: a possibly-committed statement is never re-executed (
08007asks the client to verify), recovery never appends a second result to a response the client already partly received, the idle backend-watch relay forwards only complete frames, andSETs from a rolled-back transaction are no longer restored. - 1.6.0 —
tr_modenow drives real in-session failover on both query protocols; backend faults return a proper PostgreSQLErrorResponseinstead of a dropped socket; SCRAM/MD5/cleartext backend authentication on redial; new[limits]and[cache]bounds; per-query hot-path performance work. - 1.5.0 —
[limits]and[anomaly]configuration sections;/admin/kv/<plugin>/<key>endpoints for pushing plugin runtime config;/healthz,/livez,/readyzprobe routes; admin dashboard fixed to work withadmin_tokenset. Contains a stored-XSS fix for the 1.4.0 admin dashboard — 1.4.0 operators should upgrade. - 1.4.0 — Admin API binds loopback by default (breaking); environment-variable substitution in config files; edge/geo result-cache mode; MCP bearer auth; gateway and admin request hardening.
- 1.3.0 / 1.3.1 — In-band failure detection, protocol-level health probes,
/api/circuit, idle-connection reaper. - 1.2.0 — Real LDAP search-then-bind authentication (
ldap-auth). - 1.1.0 — Transaction and statement pooling do real work on the data path.
- 1.0.0 — First stable release; every feature flag ships working functionality, with bounded capabilities documented.
- 0.5.0 — Client TLS/mTLS, proxy-terminated SCRAM-SHA-256, pg_hba-style admission rules, MCP agent gateway, HTTP SQL gateway, admin bearer auth, SIGHUP reload and SIGUSR2 binary handoff.
- v0.4.0 — From Connection Router to Programmable Data-Plane — historical release note
- v0.3.1 — Hot-path Performance & Correctness — historical release note
Examples & Demos
End-to-end demo runners live in the core repo’s examples/ directory:
examples/postgres-cluster— full PG-cluster + proxy composeexamples/failover-demo— manual failover with Transaction Replayexamples/multi-tenant— per-tenant pool isolation
The v0.4.0 feature demos remain in demos/v0.4.0/ and are catalogued under Demos.