Skip to content

HeliosDB Vector Search - Production Guide

HeliosDB Vector Search - Production Guide

Overview

HeliosDB Vector Search provides hybrid vector search capabilities combining:

  • HNSW (Hierarchical Navigable Small World) - Fast approximate nearest neighbor search
  • Multiple Distance Metrics - L2, Cosine, Manhattan, Dot Product, Hamming
  • SIMD Optimizations - AVX2 and AVX-512 acceleration (5-10x speedup)
  • Hybrid Search - Combine vector similarity with metadata filtering and full-text search
  • BM25 Scoring - Industry-standard text ranking algorithm
  • Multi-Vector Search - Query with multiple embeddings
  • Concurrent Queries - Thread-safe read access for high QPS

Performance Targets

10,000+ QPS for 1M vectors (HNSW with ef=50) 95%+ Recall@10 with default parameters <50ms p95 latency for hybrid search 5-10x SIMD speedup for distance calculations

Distance Metrics

Choosing the Right Metric

MetricUse CaseRangeNormalized?
CosineText embeddings, semantic search[0, 2]Yes (recommended)
L2 (Euclidean)Image embeddings, general purpose[0, ∞)No
Manhattan (L1)Sparse vectors, high dimensions[0, ∞)No
Dot ProductAlready normalized vectors(-∞, ∞)No
HammingBinary vectors, hashing[0, n]No

SIMD Performance

All distance metrics are SIMD-optimized and automatically use AVX-512 if available, else AVX2, else scalar.

Benchmark Results (512-dimensional vectors):

  • Scalar: ~500ns per distance calculation
  • AVX2: ~80ns per distance calculation (6x speedup)
  • AVX-512: ~45ns per distance calculation (11x speedup)

Normalizing Vectors

For cosine similarity, normalize vectors first.

HNSW Parameter Tuning

Key Parameters

  1. M (max connections): Controls graph connectivity

    • Higher M = better recall, more memory
    • Recommended: 16-32
    • Range: 4-64
  2. ef_construction: Build-time quality

    • Higher ef = better graph, slower build
    • Recommended: 200-400
    • Range: 100-1000
  3. ef (search-time): Query recall vs speed

    • Higher ef = better recall, slower search
    • Recommended: 50-200 for 95%+ recall
    • Range: 10-1000

Performance Profiles

  • High Recall (Production): M=32, ef_construction=400, ef=200; expected >98% recall, ~5,000 QPS
  • Balanced (Recommended): M=16, ef_construction=200, ef=50; expected >95% recall, ~10,000 QPS
  • Fast (Low Latency): M=8, ef_construction=100, ef=20; expected >85% recall, ~25,000 QPS

Memory Usage

Formula: Memory ≈ N × (D × 4 + M × 8) bytes

Where:

  • N = number of vectors
  • D = vector dimension
  • M = max connections parameter

Examples:

  • 1M vectors, 128D, M=16: ~1.6 GB
  • 10M vectors, 512D, M=16: ~24 GB
  • 100M vectors, 768D, M=32: ~360 GB

Hybrid Search Strategies

1. Pre-filtering (Efficient)

Apply metadata filters before vector search.

When to use: High selectivity filters (>10% match rate)

2. Post-filtering (Accurate)

Apply filters after vector search.

When to use: Low selectivity filters (<10% match rate)

3. Score Fusion

Combine vector similarity with text relevance, using either a weighted average of vector, text and metadata scores, or Reciprocal Rank Fusion (RRF), which is better for combining different score ranges.

4. Reranking

Two-stage retrieval for better accuracy: fetch 2x candidates, then rerank them with exact similarity.

BM25 Text Scoring

HeliosDB implements the BM25 algorithm for text relevance.

BM25 Formula:

BM25(D, Q) = Σ IDF(qi) × (f(qi,D) × (k1 + 1)) / (f(qi,D) + k1 × (1 - b + b × |D|/avgdl))

Where:

  • k1 = 1.5 (term frequency saturation)
  • b = 0.75 (length normalization)
  • IDF = inverse document frequency
  • f(qi,D) = term frequency in document

Concurrent Queries

HNSW index supports concurrent reads.

Performance: Linear scaling up to CPU core count

Combine vector search with metadata filters.

Performance Impact:

  • 10% selectivity: ~2x slower
  • 50% selectivity: ~1.3x slower
  • 90% selectivity: ~1.1x slower

Query with multiple embeddings (e.g., multiple text chunks); results include documents that match any of the vectors (max score).

Production Checklist

Index Configuration

  • M = 16-32 for balanced performance
  • ef_construction = 200-400 for good graph quality
  • ef = 50-200 for 95%+ recall at query time

Distance Metric

  • Cosine for normalized embeddings (text)
  • L2 for non-normalized embeddings (images)
  • Normalize vectors before indexing (if using Cosine)

Hybrid Search

  • Use pre-filtering for high selectivity (>10%)
  • Use post-filtering for low selectivity (<10%)
  • Enable BM25 for text relevance
  • Tune fusion weights based on use case

Performance

  • Benchmark with production data
  • Test concurrent query load
  • Monitor memory usage (scale with dataset)
  • Enable SIMD (AVX2/AVX-512)

Reliability

  • Implement index persistence
  • Plan for index rebuild strategy
  • Monitor recall metrics
  • Set up alerting for QPS/latency

Troubleshooting

Low Recall

Problem: Search results missing relevant documents

Solutions:

  1. Increase ef parameter (50 → 100 → 200)
  2. Increase ef_construction for new index (200 → 400)
  3. Increase M parameter (16 → 32)
  4. Verify vector normalization (for Cosine)
  5. Check for filtering issues (too restrictive)

Low QPS

Problem: Slow query performance

Solutions:

  1. Decrease ef parameter (200 → 100 → 50)
  2. Decrease M parameter (32 → 16 → 8)
  3. Enable CPU SIMD features (AVX2/AVX-512)
  4. Use pre-filtering instead of post-filtering
  5. Reduce reranking overhead
  6. Scale horizontally (distribute index)

High Memory Usage

Problem: Index consuming too much RAM

Solutions:

  1. Decrease M parameter (32 → 16 → 8)
  2. Use lower precision vectors (f32 → f16)
  3. Implement disk-based storage (mmap)
  4. Shard index across multiple nodes
  5. Use IVF-PQ quantization for compression

Filtering Issues

Problem: Filtered search returning too few results

Solutions:

  1. Use post-filtering instead of pre-filtering
  2. Increase k to oversample before filtering
  3. Check filter logic for correctness
  4. Implement 2-hop traversal (already supported)
  5. Verify metadata is correctly indexed

Advanced Topics

Distributed Indexing

For billion-scale datasets, shard the index: inserts are routed to the correct shard, and searches query all shards and merge results.

References

Support

For issues, questions, or feature requests: