HeliosDB Tier 2/3 AI/ML Features User Guide
HeliosDB Tier 2/3 AI/ML Features User Guide
1. Neural Query Planner
Description
A deep learning-based query optimizer that uses a Transformer encoder + Graph Neural Network architecture for plan generation.
Key Features
- Deep Learning Query Optimization: Transformer + GNN architecture
- Real-time Inference: Sub-5ms latency plan generation
- Learned Cost Model: Neural network-based cost estimation
- Beam Search Exploration: Guided plan search with neural heuristics
- ONNX Export: Production deployment via ONNX
Configuration Options
| Option | Default | Description |
|---|---|---|
inference_batch_size | 64 | Batch size for inference |
beam_width | 5 | Beam search width |
model_type | ”transformer-gnn” | Model architecture |
learning_rate | 0.001 | Training learning rate |
max_plan_depth | 20 | Maximum plan tree depth |
2. Schema AI (Generative Schema Designer)
Description
AI-powered schema design from natural language descriptions. Converts NL requirements to optimized database schemas with automatic normalization.
Key Features
- NL-to-ERD: Natural language to Entity-Relationship Diagrams
- Automatic Normalization: 1NF to BCNF normalization
- Index Recommendations: AI-suggested indexes
- Schema Evolution: Intelligent migration generation
- Multi-dialect Support: PostgreSQL, MySQL, SQLite DDL generation
Schema Generation Options
| Option | Values | Description |
|---|---|---|
target_normalization | 1NF, 2NF, 3NF, BCNF | Normalization level |
generate_indexes | true/false | Auto-generate indexes |
target_dialect | postgresql, mysql, sqlite | SQL dialect |
include_constraints | true/false | Generate constraints |
3. RL-Based Intelligent Cache
Description
Reinforcement learning-based cache eviction using Deep Q-Network (DQN). Learns optimal caching policies from access patterns.
Key Features
- DQN-based Eviction: Deep Q-Network for policy learning
- Contextual Features: Uses query patterns, time-of-day, workload type
- Adaptive Learning: Continuous improvement from production traffic
- Multi-tier Support: Coordinates L1/L2/L3 cache hierarchies
- Workload-aware: Different policies for OLTP vs OLAP
Configuration
| Option | Default | Description |
|---|---|---|
capacity_mb | 1024 | Cache size in MB |
learning_rate | 0.001 | DQN learning rate |
discount_factor | 0.99 | Future reward discount |
exploration_rate | 0.1 | Epsilon for exploration |
update_frequency | 100 | Steps between updates |
4. Multi-Armed Bandit Load Balancer
Description
LinUCB contextual bandit algorithm for intelligent request routing. Balances load while optimizing for latency and throughput.
Key Features
- Contextual Bandits: LinUCB algorithm with context features
- Latency Optimization: Routes to fastest available node
- Adaptive Exploration: Balances exploration vs exploitation
- Health-aware: Considers node health in routing decisions
- Multi-objective: Optimizes latency, throughput, and fairness
5. Anomaly Detection
Description
Multi-algorithm anomaly detection for database metrics, query patterns, and data quality monitoring.
Supported Algorithms
| Algorithm | Use Case | Strengths |
|---|---|---|
| Isolation Forest | General | Fast, handles high dimensions |
| Local Outlier Factor | Density-based | Good for clusters |
| DBSCAN | Clustering | Finds arbitrary shapes |
| One-Class SVM | Novelty detection | Works with limited data |
| LSTM Autoencoder | Time-series | Captures temporal patterns |
| Statistical (Z-score) | Simple metrics | Interpretable, fast |
Algorithm Selection Guide
Query latency monitoring → Isolation Forest or LSTMConnection patterns → DBSCANData quality checks → Statistical (Z-score)Novel query detection → One-Class SVMGeneral monitoring → Ensemble (multiple algorithms)6. Time-Series Forecasting
Description
Comprehensive time-series forecasting for capacity planning, workload prediction, and trend analysis.
Supported Algorithms
| Algorithm | Best For | Accuracy |
|---|---|---|
| ARIMA | Stationary data | High |
| Prophet | Seasonal + holidays | High |
| LSTM | Complex patterns | Very High |
| Exponential Smoothing | Simple trends | Medium |
| Ensemble | General | Highest |
7. AutoML Tuning
Description
Automatic database configuration tuning using Bayesian Optimization and Genetic Algorithms.
Key Features
- Bayesian Optimization: Efficient hyperparameter search
- Genetic Algorithms: Evolves optimal configurations
- Safe Exploration: Constraints to prevent bad configs
- A/B Testing: Validates improvements before rollout
- Workload-aware: Different configs for different workloads
8. Auto-Index
Description
ML-based automatic index recommendation and management based on workload analysis.
Key Features
- Workload Analysis: Learns from query patterns
- Index Recommendations: Suggests optimal indexes
- Impact Prediction: Estimates performance improvement
- Automatic Creation: Creates indexes during low-traffic periods
- Index Consolidation: Removes redundant indexes
9. Probabilistic Data Structures
Description
Memory-efficient probabilistic data structures for approximate queries.
Supported Structures
| Structure | Use Case | Space | Error Rate |
|---|---|---|---|
| Bloom Filter | Membership testing | O(n) bits | Configurable FP |
| Count-Min Sketch | Frequency estimation | O(1) | Configurable |
| HyperLogLog | Cardinality estimation | ~1.5KB | ~2% |
| T-Digest | Percentile estimation | O(compression) | ~1% |
| Cuckoo Filter | Membership + delete | O(n) bits | Configurable |
| MinHash | Similarity estimation | O(k) | 1/√k |
| SimHash | Near-duplicate detection | O(1) | Configurable |
Integration with HeliosDB
Configuration
[ai.neural_planner]enabled = truemodel_path = "models/query_planner.onnx"inference_timeout_ms = 5fallback_to_traditional = true
[ai.schema_ai]enabled = truedefault_normalization = "3NF"
[cache.rl]enabled = truecapacity_mb = 2048learning_enabled = true
[cluster.mab_balancer]enabled = truealpha = 0.5update_frequency = 100
[monitoring.anomaly_detection]enabled = truealgorithm = "ensemble"alert_threshold = 0.8
[monitoring.forecasting]enabled = truealgorithm = "auto"horizon_hours = 24
[tuning.automl]enabled = truemaintenance_window = "02:00-05:00"require_approval = true
[indexing.auto_index]enabled = truemin_improvement_threshold = 0.10max_indexes_per_table = 10Best Practices
1. Start with Defaults
All packages have sensible defaults. Start with defaults and tune based on your workload.
2. Monitor Before Enabling
Monitor your workload characteristics before enabling AI features:
- Query patterns and frequency
- Data distribution
- Peak vs off-peak traffic
3. Use Gradual Rollout
Enable features gradually:
- Start with read-only features (anomaly detection, forecasting)
- Enable learning features (neural planner, RL cache)
- Enable write features (auto-index, automl tuning)
4. Set Safety Constraints
Always configure safety constraints for features that modify behavior:
- Max latency thresholds
- Rollback policies
- Approval requirements
5. Review Recommendations
AI recommendations should be reviewed before automatic application, especially for:
- Index creation
- Configuration changes
- Schema modifications
Troubleshooting
Common Issues
Neural Planner slow inference
- Check model file is loaded (not re-loading per query)
- Reduce beam width if latency exceeds 5ms
- Use ONNX runtime optimizations
RL Cache low hit rate
- Allow more training time (10,000+ accesses)
- Check exploration rate isn’t too high
- Verify feature extraction includes relevant context
Anomaly Detection false positives
- Increase contamination parameter
- Use ensemble mode for higher precision
- Train on longer historical period
AutoML Tuning not improving
- Expand search space ranges
- Increase iteration count
- Check constraint feasibility
Support
- Documentation: https://heliosdb.com/docs/full/
- Support: support@heliosdb.com
- Community: https://discord.gg/yTykuUrFXc
This guide covers HeliosDB v7.1.2 Tier 2/3 AI/ML features.