Skip to content

HeliosDB Observability User Guide

HeliosDB Observability User Guide

Overview

HeliosDB Observability provides zero-configuration automatic instrumentation for distributed tracing, metrics, and monitoring. Automatic tracing is designed to add <5µs of overhead.

Features

  • Zero Configuration: No code changes required
  • <5µs Overhead: operationally hardened performance
  • Automatic Instrumentation: All operations traced automatically
  • OpenTelemetry Compatible: Export to Jaeger, Zipkin, Prometheus
  • Real-time Dashboard: Built-in web UI
  • Smart Alerting: Automatic anomaly detection

Configuration

Environment Variables

Terminal window
# Enable observability
HELIOSDB_OBSERVABILITY_ENABLED=true
# Set service name
HELIOSDB_SERVICE_NAME="my-heliosdb-instance"
# OTLP endpoint
OTEL_EXPORTER_OTLP_ENDPOINT="http://localhost:4317"
# Jaeger endpoint
JAEGER_AGENT_HOST="localhost"
JAEGER_AGENT_PORT="6831"
# Sampling rate (0.0 - 1.0)
HELIOSDB_TRACE_SAMPLE_RATE=1.0
# Log level
RUST_LOG=info

Configuration File

Create observability.toml:

[observability]
enabled = true
service_name = "heliosdb"
[observability.sampling]
# Sample 100% of traces (reduce for high-volume production)
rate = 1.0
[observability.exporters.otlp]
enabled = true
endpoint = "http://localhost:4317"
timeout_seconds = 30
[observability.exporters.jaeger]
enabled = false
agent_host = "localhost"
agent_port = 6831
[observability.dashboard]
enabled = true
host = "0.0.0.0"
port = 9090
[observability.alerts]
enabled = true
email_smtp = "smtp.gmail.com:587"
slack_webhook = "https://hooks.slack.com/services/YOUR/WEBHOOK"

Real-time Dashboard

Dashboard Features

  • Real-time Metrics: Live query latency, throughput, error rates
  • Trace Viewer: Interactive trace timeline with flame graphs
  • System Health: CPU, memory, disk I/O monitoring
  • Alerts: Configure thresholds for automatic notifications
  • Query Analytics: Top queries, slow queries, query patterns

Accessing the Dashboard

Open your browser to http://localhost:9090:

  • Home: System overview and health status
  • Traces: Search and view distributed traces
  • Metrics: Time-series charts for all metrics
  • Alerts: Configure and view alerts
  • Config: Runtime configuration

Integration with Monitoring Tools

Prometheus

Add to prometheus.yml:

scrape_configs:
- job_name: 'heliosdb'
static_configs:
- targets: ['localhost:9091']

Grafana

Import the HeliosDB dashboard:

  1. Open Grafana
  2. Go to Dashboards → Import
  3. Upload the HeliosDB dashboard JSON (heliosdb-dashboard.json)
  4. Select your Prometheus data source

Jaeger

View distributed traces in Jaeger UI:

Terminal window
# Start Jaeger all-in-one
docker run -d --name jaeger \
-p 6831:6831/udp \
-p 16686:16686 \
jaegertracing/all-in-one:latest
# Open Jaeger UI
open http://localhost:16686

OpenTelemetry Collector

Use the OpenTelemetry Collector for advanced pipeline:

otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
processors:
batch:
exporters:
jaeger:
endpoint: jaeger:14250
prometheus:
endpoint: 0.0.0.0:9092
service:
pipelines:
traces:
receivers: [otlp]
processors: [batch]
exporters: [jaeger]
metrics:
receivers: [otlp]
processors: [batch]
exporters: [prometheus]

Alerting

Alert Severities

  • Info: Informational alerts (e.g., deployment events)
  • Warning: Degraded performance, attention needed
  • Error: Service degradation, immediate attention
  • Critical: Service outage, page on-call

Performance Impact

Overhead Benchmarks

Auto-instrumentation overhead (indicative figures, measured on representative fixtures; reproduce on your own hardware):

OperationWithout TracingWith TracingOverhead
Simple SELECT450µs453µs0.7%
Complex JOIN12.5ms12.51ms0.08%
INSERT320µs323µs0.9%
Transaction1.2ms1.203ms0.25%

Average overhead: <5µs per operation

When Disabled

When auto-instrumentation is disabled:

  • Zero overhead: No performance impact
  • Spans are not created
  • No memory allocation for tracing
  • Ideal for performance-critical sections

Best Practices

  1. Production: Sample 10-20% of traces
  2. Staging: Sample 100% for comprehensive testing
  3. Development: Enable verbose logging
  4. Load Testing: Disable for accurate benchmarks

Troubleshooting

Traces Not Appearing

  1. Test connectivity to backend:
    Terminal window
    curl http://localhost:4317/v1/traces

High Overhead

  1. Reduce sampling rate

  2. Disable verbose logging:

    Terminal window
    RUST_LOG=warn

Support