HeliosDB Vector Search - Production Guide
HeliosDB Vector Search - Production Guide
Overview
HeliosDB Vector Search provides hybrid vector search capabilities combining:
- HNSW (Hierarchical Navigable Small World) - Fast approximate nearest neighbor search
- Multiple Distance Metrics - L2, Cosine, Manhattan, Dot Product, Hamming
- SIMD Optimizations - AVX2 and AVX-512 acceleration (5-10x speedup)
- Hybrid Search - Combine vector similarity with metadata filtering and full-text search
- BM25 Scoring - Industry-standard text ranking algorithm
- Multi-Vector Search - Query with multiple embeddings
- Concurrent Queries - Thread-safe read access for high QPS
Performance Targets
10,000+ QPS for 1M vectors (HNSW with ef=50) 95%+ Recall@10 with default parameters <50ms p95 latency for hybrid search 5-10x SIMD speedup for distance calculations
Distance Metrics
Choosing the Right Metric
| Metric | Use Case | Range | Normalized? |
|---|---|---|---|
| Cosine | Text embeddings, semantic search | [0, 2] | Yes (recommended) |
| L2 (Euclidean) | Image embeddings, general purpose | [0, ∞) | No |
| Manhattan (L1) | Sparse vectors, high dimensions | [0, ∞) | No |
| Dot Product | Already normalized vectors | (-∞, ∞) | No |
| Hamming | Binary vectors, hashing | [0, n] | No |
SIMD Performance
All distance metrics are SIMD-optimized and automatically use AVX-512 if available, else AVX2, else scalar.
Benchmark Results (512-dimensional vectors):
- Scalar: ~500ns per distance calculation
- AVX2: ~80ns per distance calculation (6x speedup)
- AVX-512: ~45ns per distance calculation (11x speedup)
Normalizing Vectors
For cosine similarity, normalize vectors first.
HNSW Parameter Tuning
Key Parameters
-
M (max connections): Controls graph connectivity
- Higher M = better recall, more memory
- Recommended: 16-32
- Range: 4-64
-
ef_construction: Build-time quality
- Higher ef = better graph, slower build
- Recommended: 200-400
- Range: 100-1000
-
ef (search-time): Query recall vs speed
- Higher ef = better recall, slower search
- Recommended: 50-200 for 95%+ recall
- Range: 10-1000
Performance Profiles
- High Recall (Production): M=32,
ef_construction=400,ef=200; expected >98% recall, ~5,000 QPS - Balanced (Recommended): M=16,
ef_construction=200,ef=50; expected >95% recall, ~10,000 QPS - Fast (Low Latency): M=8,
ef_construction=100,ef=20; expected >85% recall, ~25,000 QPS
Memory Usage
Formula: Memory ≈ N × (D × 4 + M × 8) bytes
Where:
- N = number of vectors
- D = vector dimension
- M = max connections parameter
Examples:
- 1M vectors, 128D, M=16: ~1.6 GB
- 10M vectors, 512D, M=16: ~24 GB
- 100M vectors, 768D, M=32: ~360 GB
Hybrid Search Strategies
1. Pre-filtering (Efficient)
Apply metadata filters before vector search.
When to use: High selectivity filters (>10% match rate)
2. Post-filtering (Accurate)
Apply filters after vector search.
When to use: Low selectivity filters (<10% match rate)
3. Score Fusion
Combine vector similarity with text relevance, using either a weighted average of vector, text and metadata scores, or Reciprocal Rank Fusion (RRF), which is better for combining different score ranges.
4. Reranking
Two-stage retrieval for better accuracy: fetch 2x candidates, then rerank them with exact similarity.
BM25 Text Scoring
HeliosDB implements the BM25 algorithm for text relevance.
BM25 Formula:
BM25(D, Q) = Σ IDF(qi) × (f(qi,D) × (k1 + 1)) / (f(qi,D) + k1 × (1 - b + b × |D|/avgdl))Where:
- k1 = 1.5 (term frequency saturation)
- b = 0.75 (length normalization)
- IDF = inverse document frequency
- f(qi,D) = term frequency in document
Concurrent Queries
HNSW index supports concurrent reads.
Performance: Linear scaling up to CPU core count
Filtered Search
Combine vector search with metadata filters.
Performance Impact:
- 10% selectivity: ~2x slower
- 50% selectivity: ~1.3x slower
- 90% selectivity: ~1.1x slower
Multi-Vector Search
Query with multiple embeddings (e.g., multiple text chunks); results include documents that match any of the vectors (max score).
Production Checklist
Index Configuration
- M = 16-32 for balanced performance
- ef_construction = 200-400 for good graph quality
- ef = 50-200 for 95%+ recall at query time
Distance Metric
- Cosine for normalized embeddings (text)
- L2 for non-normalized embeddings (images)
- Normalize vectors before indexing (if using Cosine)
Hybrid Search
- Use pre-filtering for high selectivity (>10%)
- Use post-filtering for low selectivity (<10%)
- Enable BM25 for text relevance
- Tune fusion weights based on use case
Performance
- Benchmark with production data
- Test concurrent query load
- Monitor memory usage (scale with dataset)
- Enable SIMD (AVX2/AVX-512)
Reliability
- Implement index persistence
- Plan for index rebuild strategy
- Monitor recall metrics
- Set up alerting for QPS/latency
Troubleshooting
Low Recall
Problem: Search results missing relevant documents
Solutions:
- Increase
efparameter (50 → 100 → 200) - Increase
ef_constructionfor new index (200 → 400) - Increase
Mparameter (16 → 32) - Verify vector normalization (for Cosine)
- Check for filtering issues (too restrictive)
Low QPS
Problem: Slow query performance
Solutions:
- Decrease
efparameter (200 → 100 → 50) - Decrease
Mparameter (32 → 16 → 8) - Enable CPU SIMD features (AVX2/AVX-512)
- Use pre-filtering instead of post-filtering
- Reduce reranking overhead
- Scale horizontally (distribute index)
High Memory Usage
Problem: Index consuming too much RAM
Solutions:
- Decrease
Mparameter (32 → 16 → 8) - Use lower precision vectors (f32 → f16)
- Implement disk-based storage (mmap)
- Shard index across multiple nodes
- Use IVF-PQ quantization for compression
Filtering Issues
Problem: Filtered search returning too few results
Solutions:
- Use post-filtering instead of pre-filtering
- Increase
kto oversample before filtering - Check filter logic for correctness
- Implement 2-hop traversal (already supported)
- Verify metadata is correctly indexed
Advanced Topics
Distributed Indexing
For billion-scale datasets, shard the index: inserts are routed to the correct shard, and searches query all shards and merge results.
References
- HNSW Paper - Malkov & Yashunin (2018)
- BM25 Algorithm
- SIMD Guide
Support
For issues, questions, or feature requests:
- Support: support@heliosdb.com
- Documentation: heliosdb.com/docs/full
- Community: Discord