Skip to content

HeliosDB Advanced Observability Dashboard - User Guide

HeliosDB Advanced Observability Dashboard - User Guide

Version: 1.0


Table of Contents

  1. Overview
  2. Architecture
  3. Getting Started
  4. REST API Reference
  5. WebSocket Protocol
  6. Metrics System
  7. Alert Engine
  8. Performance Tuning
  9. Troubleshooting

Overview

The HeliosDB Advanced Observability Dashboard provides comprehensive monitoring and alerting for your HeliosDB clusters with:

  • Real-time metrics visualization (<1s latency)
  • Query performance analytics (P50/P95/P99 latencies)
  • Distributed tracing flamegraphs (integrated with OpenTelemetry)
  • Intelligent alert engine (ML-based anomaly detection)
  • Multi-channel notifications (email, Slack, PagerDuty, webhooks)
  • Prometheus-compatible metrics export
  • 100K+ metrics/sec ingestion capacity

Key Features

FeatureDescriptionPerformance Target
Metrics IngestionTime-series metric collection100K+ metrics/sec
Real-time UpdatesWebSocket streaming<1s latency
Dashboard Load TimeReact SPA loading<2s
API Response TimeREST endpoint latency<50ms
Alert EvaluationRule-based monitoring<100ms
Data RetentionMulti-resolution storage24 hours (configurable)

Architecture

System Components

┌─────────────────────────────────────────────────────────┐
│ React Dashboard UI │
│ (Real-time Charts & Graphs) │
└───────────────┬─────────────────────┬───────────────────┘
│ │
│ HTTP/REST │ WebSocket
│ │
┌───────────────▼─────────────────────▼───────────────────┐
│ actix-web Server (Rust) │
│ ┌─────────────┬──────────────┬───────────────────┐ │
│ │ Metrics API │ Alert Engine │ WebSocket Server │ │
│ └──────┬──────┴──────┬───────┴─────┬─────────────┘ │
│ │ │ │ │
│ ┌──────▼──────┐ ┌───▼────┐ ┌────▼──────┐ │
│ │ Metrics │ │ Alert │ │ Real-time │ │
│ │ Aggregator │ │ Rules │ │ Broker │ │
│ └─────────────┘ └────────┘ └───────────┘ │
└─────────────────────────────────────────────────────────┘
│
│ Metrics Collection
│
┌───────────────▼─────────────────────────────────────────┐
│ HeliosDB Nodes (Query Engine, Storage) │
└─────────────────────────────────────────────────────────┘

Data Flow

  1. Metrics Collection: HeliosDB nodes emit metrics during query execution
  2. Ingestion: Metrics API receives data via REST POST or pub/sub
  3. Aggregation: Time-series data is aggregated into multiple resolutions
  4. Alert Evaluation: Alert engine checks rules against incoming metrics
  5. Real-time Streaming: WebSocket server broadcasts to connected clients
  6. Visualization: React UI renders charts and graphs in real-time

Getting Started

Access the dashboard:


How to Configure

The dashboard is not switched on through a public configuration file. Enabling it, setting its listen address and port, and the tuning and alert-channel settings described on this page are provided during onboarding — contact support@heliosdb.com (or sales@heliosdb.com if you are not yet a customer).

The heliosdb.toml file of the HeliosDB Full server accepts only the documented storage keys (storage.data_dir, storage.memtable_size_mb, storage.compaction_strategy, storage.read_cache_mb and the storage.prefetch_* settings). The server refuses to start if the file contains any other key.


REST API Reference

Metrics Endpoints

POST /api/metrics/query

Query time-series metrics.

Request:

{
"metric": "query_latency_ms",
"start": "2025-11-09T00:00:00Z",
"end": "2025-11-09T23:59:59Z",
"step": 60,
"labels": {
"database": "main"
}
}

Response:

{
"success": true,
"data": [
{
"metric": "query_latency_ms",
"data": [
{"timestamp": 1699488000, "value": 45.2},
{"timestamp": 1699488060, "value": 42.8}
]
}
]
}

GET /api/metrics/prometheus

Export metrics in Prometheus text format.

Response:

# TYPE heliosdb_query_latency_ms histogram
heliosdb_query_latency_ms_bucket{le="10"} 1234
heliosdb_query_latency_ms_bucket{le="50"} 5678
heliosdb_query_latency_ms_sum 123456.0
heliosdb_query_latency_ms_count 10000

Alert Endpoints

GET /api/alerts

Get active alerts.

Response:

{
"success": true,
"data": [
{
"id": "alert_abc123",
"rule_id": "rule_latency",
"name": "High Query Latency",
"severity": "warning",
"value": 1250.5,
"threshold": 1000.0,
"triggered_at": "2025-11-09T12:34:56Z",
"state": "firing"
}
]
}

POST /api/alerts/rules

Create a new alert rule.

Request:

{
"id": "rule_cpu",
"name": "High CPU Usage",
"description": "CPU usage exceeds 80%",
"metric": "cpu_usage_percent",
"condition": "greater_than",
"threshold": 80.0,
"window_secs": 300,
"min_violations": 3,
"severity": "critical",
"channels": ["slack_alerts"],
"enabled": true
}

GET /api/alerts/history?limit=100

Get alert history (up to 1000 alerts).

Stats Endpoint

GET /api/stats

Get dashboard statistics summary.

Response:

{
"success": true,
"data": {
"total_queries": 123456,
"avg_query_latency_ms": 45.2,
"p95_query_latency_ms": 125.8,
"p99_query_latency_ms": 342.1,
"active_connections": 128,
"cache_hit_rate": 0.87,
"active_alerts": 2,
"cluster_status": "healthy",
"timestamp": "2025-11-09T12:34:56Z"
}
}

Cluster Health Endpoint

GET /api/cluster/health

Get cluster health status.

Response:

{
"success": true,
"data": {
"status": "healthy",
"total_nodes": 3,
"healthy_nodes": 3,
"unhealthy_nodes": 0,
"nodes": [
{
"node_id": "node-1",
"status": "online",
"cpu_usage": 45.2,
"memory_usage": 67.8,
"disk_usage": 42.1,
"active_queries": 12,
"uptime_seconds": 123456
}
]
}
}

WebSocket Protocol

Connection

Connect to the WebSocket endpoint:

const ws = new WebSocket('ws://localhost:8080/ws');
ws.onopen = () => {
console.log('Connected to HeliosDB Dashboard');
// Subscribe to metrics
ws.send(JSON.stringify({
type: 'subscribe',
metrics: ['query_latency_ms', 'cpu_usage_percent']
}));
};

Message Format

Client → Server (Subscribe)

{
"type": "subscribe",
"metrics": ["query_latency_ms", "throughput"]
}

Client → Server (Unsubscribe)

{
"type": "unsubscribe",
"metrics": ["throughput"]
}

Server → Client (Metrics Update)

{
"type": "metrics",
"data": [
{
"name": "query_latency_ms",
"value": 45.2,
"timestamp": "2025-11-09T12:34:56Z",
"labels": {"database": "main"},
"metric_type": "histogram"
}
]
}

Metrics System

Supported Metric Types

TypeDescriptionUse Case
CounterMonotonically increasing valueRequest counts, bytes sent
GaugeCurrent value that can go up/downCPU usage, memory usage
HistogramDistribution of valuesQuery latency, response times
SummarySimilar to histogram, pre-calculated quantilesCustom percentiles

Metric Labels

All metrics support labels for multi-dimensional data.

Aggregation Windows

Metrics are automatically aggregated into multiple resolutions for efficient querying:

WindowDurationRetentionUse Case
Raw1s1 hourReal-time monitoring
1-minute60s6 hoursRecent trends
5-minute300s24 hoursDaily patterns
15-minute900s3 daysWeekly trends
1-hour3600s7 daysLong-term analysis
1-day86400s30 daysHistorical data

Alert Engine

Rule Configuration

Alert rules support:

  • Threshold-based alerts (>, <, ==, !=, >=, <=)
  • Windowed evaluation (check multiple violations)
  • Label filtering (alert on specific metric labels)
  • Multi-channel notifications

Notification Channels

Supported channels: Slack (incoming webhook URL and channel), Email (SMTP), and PagerDuty (integration key).


Performance Tuning

Metrics Retention

Adjust retention based on your needs with metrics_retention_hours (for example, 72 for 3 days).

Longer retention = more storage, slower queries.

Alert Check Interval

alert_check_interval sets how often alert rules are checked, in seconds (for example, 5).

Lower interval = faster alerting, higher CPU usage.

WebSocket Ping Interval

ws_ping_interval sets the WebSocket ping interval in seconds (for example, 15).


Troubleshooting

High Memory Usage

Symptom: Dashboard consuming excessive memory Solution: Reduce metrics_retention_hours or increase aggregation pruning frequency

Slow Queries

Symptom: API queries taking >1s Solution: Use appropriate step parameter for downsampling, enable aggregation

Missed Alerts

Symptom: Alerts not triggering Solution: Check min_violations setting, verify metric labels match rule filters

WebSocket Disconnections

Symptom: Real-time updates stop Solution: Check ws_ping_interval, verify network connectivity, review firewall rules


End of User Guide

For additional support, see: