Skip to content

HeliosProxy Documentation

HeliosProxy Documentation

Programmable Postgres Data-Plane

Version License Container


What is HeliosProxy?

HeliosProxy is a programmable Postgres data-plane: a PostgreSQL-wire connection router with a real WASM plugin runtime, signed plugin distribution via OCI, a transaction journal with operator-driven replay, a zero-downtime PostgreSQL major-version upgrade orchestrator, and a built-in admin web console.

The current release is 2.0.0 (September 25, 2026) — see Upgrading to 2.0 and the 2.0 announcement. Capabilities are organised into two tiers:

  • Connection-Routing Tier — read/write splitting, health checking, circuit breaking, rate limiting, connection pooling (session / transaction / statement), and failover.
  • Platform Tier — WASM plugins, query caching, query analytics, multi-tenancy, authentication, query rewriting, GraphQL and MCP gateways, anomaly detection, and edge/geo caching.

Every tier-two capability is a Cargo feature flag, so the binary only carries what you enable — see Feature Flags for the authoritative list, which tracks the [features] table in Cargo.toml flag for flag.

Since 1.0.0 the project’s stated policy is that every shipped feature flag does real work, with intentionally-bounded capabilities documented rather than implied. Where a subsystem is library-only or not yet mounted on the per-query data path, this documentation says so explicitly.

Quick Start

Run with Docker

Terminal window
docker pull ghcr.io/heliosdatabase/hdb-heliosdb-proxy:2.0.0
docker run -d \
--name heliosproxy \
-p 5432:5432 \
-p 9090:9090 \
-v $(pwd)/proxy.toml:/etc/heliosproxy/proxy.toml \
ghcr.io/heliosdatabase/hdb-heliosdb-proxy:2.0.0

Published tags follow the release tags (2.0.0, 2.0, latest). Note that the admin API binds loopback by default as of 1.4.0, so reach it from inside the container or publish it deliberately as above.

Install from crates.io

Terminal window
cargo install heliosdb-proxy

The minimum supported Rust version is 1.86.

Configure

Minimal proxy.toml:

listen_address = "0.0.0.0:5432"
admin_address = "127.0.0.1:9090"
tr_enabled = true
tr_mode = "session"
[[nodes]]
name = "primary-1"
host = "pg-primary.example.com"
port = 5432
role = "primary"
[[nodes]]
name = "standby-1"
host = "pg-standby.example.com"
port = 5432
role = "standby"

The sections [pool], [load_balancer], [health] and at least one [[nodes]] entry are required in a configuration file; everything else defaults to off. Config files support ${VAR} and ${VAR:-default} environment substitution (1.4.0), and a ${VAR} with no default and no environment value fails fast at load.

See Configuration for the full reference, and the repo’s config/ directory for complete example files.

Topology

┌─────────────┐
│ HeliosProxy │ ← WASM plugins, replay engine,
│ v2.0.0 │ upgrade orchestrator, admin UI
└──────┬──────┘
│ PostgreSQL wire protocol
┌────────────────┼────────────────┐
▼ ▼ ▼
┌─────────┐ ┌─────────┐ ┌─────────┐
│ Primary │ │ Standby │ │ Replica │
└─────────┘ └─────────┘ └─────────┘

The upgrade orchestrator’s validation stage uses a row-count parity check that is portable across PostgreSQL 14–18. Release verification runs against PostgreSQL 18.4.

Documentation Structure

  • Architecture — System design, request lifecycle, hook points
  • Configuration — Full proxy.toml reference, including [limits], [anomaly], [cache], [auth], [tls] and [[hba]]
  • Feature Flags — Cargo features and module activation
  • Admin API — REST endpoints (topology, plugins, plugin KV, anomalies, edge, chaos, migration, replay, metrics)
  • Topology Providers — Backend discovery and write authority (static roles, PostgreSQL polling with the timeline rule, Patroni, HeliosDB native)
  • Transaction Replay — In-session failover (tr_mode), the committed-transaction journal ([journal]) and operator replay (POST /api/replay, including committed_history); all in the default build
  • Deployment Guides — Docker, Kubernetes, standalone

Source Repositories

RepositoryRoleLicense
HeliosDB-ProxyCore proxy (this is where most files live)Apache-2.0
HeliosDB-Proxy-PluginsFirst-party WASM plugins + helios-plugin CLIApache-2.0
HeliosDB-Proxy-OperatorKubernetes operator + Helm chartNot currently public
terraform-provider-HeliosDB-ProxyTerraform providerNot currently public
pulumi-HeliosDB-ProxyPulumi providerNot currently public

Container image: ghcr.io/heliosdatabase/hdb-heliosdb-proxy:2.0.0 (lowercase due to GHCR convention). The project was published under the Dimensigon organisation before 0.5.1; those older paths are no longer valid.

Upgrading to 2.0

HeliosProxy 2.0.0 keeps the 1.x configuration format: a 1.9 proxy.toml loads unchanged. The major version marks behaviour changes you should review before rolling it out:

  • Gateways have a concurrency cap by default. [graphql_gateway], [mcp] and [http_gateway] admit at most max_concurrent_requests = 64 requests at once. Beyond that the HTTP and GraphQL gateways answer 503 with Retry-After, and MCP answers a JSON-RPC error (-32000) for tools/call; health checks and MCP initialize/ping/tools/list are never gated. Raise the key if you run more concurrent gateway requests. query_timeout_ms (default 30000) replaces the fixed 30-second gateway timeout.
  • Unknown config keys are reported at any nesting depth. A typo inside a section ([cache] l1_max_byts) now logs WARN unknown config key 'cache.l1_max_byts' ignored at startup, and fails startup when strict_config = true. Previously only top-level keys were checked, so check your startup log after upgrading.
  • Conflicting primaries fail closed with [topology] provider = "postgres". When several nodes are writable, the node on the strictly highest timeline wins; a tie or an unreadable timeline authorizes nothing (1.9 took the first writable node probed). The probe role needs permission to call pg_control_checkpoint(). See Topology Providers.
  • Query cache (only if you enabled it). Writes on the extended protocol now invalidate it, a stale entry is refused on every tier (L1 and L3 used to be served until TTL), and DDL, TRUNCATE, COPY ... FROM, EXECUTE, CALL and DO invalidate the whole cache. Expect a lower hit rate on write-heavy workloads in exchange for read-your-commits correctness. The transaction-journal capture now also runs when tr_enabled = false but the cache is on. The cache stays off by default; see why.
  • Library users only: the TopologyProvider trait gains defaulted conflicts_total() and leader_timeline() methods (existing implementations keep compiling), and QueryCache::purge_tables no longer removes L2 entries itself.

New in 2.0 that you may want to turn on: [topology] provider = "patroni" (Topology Providers), and the Prometheus counters heliosdb_proxy_topology_conflicting_primaries_total, heliosdb_proxy_backend_capacity_waits_total and heliosdb_proxy_backend_capacity_refusals_total.

Coming from 1.7 or earlier, also note the 1.8 change (ha-tr is a no-op; Transaction Replay is always compiled and tr_enabled is the runtime switch) and that 1.9 added the [journal], [topology], strict_config and [limits] keys described in Configuration — all default to the previous behaviour.

What’s New

Highlights since 0.4.0 — see the repo CHANGELOG for the complete history.

  • 2.0.0 (2026-09-25) — The HA topology, the query cache and Transaction Replay share one model of what the backend actually committed. The query cache (still opt-in) invalidates on commit on both wire protocols and every tier, with DDL, TRUNCATE, COPY ... FROM, EXECUTE, CALL and DO invalidating the whole cache; a Patroni cluster can be the write authority ([topology] provider = "patroni"); conflicting primaries under the postgres provider are resolved by the strictly highest timeline or fail closed; gateways gain query_timeout_ms and max_concurrent_requests; unknown config keys are reported at any nesting depth; a login burst against a backend at max_connections no longer surfaces as a false authentication failure; --tr false works. Read the upgrade notes and the announcement.
  • 1.9.0 (2026-09-22) — Recovery-grade transaction journal: real committed transactions with parameters and outcomes, an optional durable on-disk journal ([journal] dir, CRC-checked segments, torn-tail truncation, commit/interval/none fsync), and POST /api/replay mode: "committed_history". Authoritative topology ([topology] provider = "postgres") with an authority lease (lease_timeout_secs); GET /capabilities and opt-in strict_config; [load_balancer] read_strategy honoured (new power_of_two); credentialed health probes and measured replica lag (require_known_lag); bounded fair client admission (client_admission_wait_secs); pooled gateway backend connections (gateway_pool_max_idle); byte budgets for the L1/L3 and edge caches; replay_deadline_secs; rustls 0.23.45 (RUSTSEC-2026-0285).
  • 1.8.0 (2026-09-12) — Transaction Replay ships in the default build; the ha-tr cargo feature is a deprecated no-op and tr_enabled is the runtime switch (false stops journaling, forces tr_mode = "none", and makes POST /api/replay return 503). Dependency security updates (reqwest 0.12, wasmtime 36, lru 0.18); the unused prometheus/opentelemetry dependencies were dropped. GET /api/migration/status no longer reports migration_ready while applies are failing.
  • 1.7.0 — Completes the Transaction Replay series: replay verifies each statement’s response digest against what the client already saw (divergence rolls back with 40001), SERIALIZABLE/REPEATABLE READ transactions are never replayed, an interrupted read is only re-executed when every function it calls is provably side-effect-free (built-in allowlist plus the new tr_read_functions), session SET restore is transactional and savepoint-scoped, exceeding tr_max_session_set_statements refuses the failover with 08006 rather than re-homing with partial state, and write_timeout_secs is now one deadline for the entire recovery. New [limits] bounds: tr_max_observation_bytes, backend_response_timeout_secs, max_backend_frame_bytes.
  • 1.6.1 — Transaction Replay commit-outcome and delivery safety: a possibly-committed statement is never re-executed (08007 asks the client to verify), recovery never appends a second result to a response the client already partly received, the idle backend-watch relay forwards only complete frames, and SETs from a rolled-back transaction are no longer restored.
  • 1.6.0 — tr_mode now drives real in-session failover on both query protocols; backend faults return a proper PostgreSQL ErrorResponse instead of a dropped socket; SCRAM/MD5/cleartext backend authentication on redial; new [limits] and [cache] bounds; per-query hot-path performance work.
  • 1.5.0 — [limits] and [anomaly] configuration sections; /admin/kv/<plugin>/<key> endpoints for pushing plugin runtime config; /healthz, /livez, /readyz probe routes; admin dashboard fixed to work with admin_token set. Contains a stored-XSS fix for the 1.4.0 admin dashboard — 1.4.0 operators should upgrade.
  • 1.4.0 — Admin API binds loopback by default (breaking); environment-variable substitution in config files; edge/geo result-cache mode; MCP bearer auth; gateway and admin request hardening.
  • 1.3.0 / 1.3.1 — In-band failure detection, protocol-level health probes, /api/circuit, idle-connection reaper.
  • 1.2.0 — Real LDAP search-then-bind authentication (ldap-auth).
  • 1.1.0 — Transaction and statement pooling do real work on the data path.
  • 1.0.0 — First stable release; every feature flag ships working functionality, with bounded capabilities documented.
  • 0.5.0 — Client TLS/mTLS, proxy-terminated SCRAM-SHA-256, pg_hba-style admission rules, MCP agent gateway, HTTP SQL gateway, admin bearer auth, SIGHUP reload and SIGUSR2 binary handoff.
  • v0.4.0 — From Connection Router to Programmable Data-Plane — historical release note
  • v0.3.1 — Hot-path Performance & Correctness — historical release note

Examples & Demos

End-to-end demo runners live in the core repo’s examples/ directory:

The v0.4.0 feature demos remain in demos/v0.4.0/ and are catalogued under Demos.