Skip to content

HNSW Performance Tuning Guide

HNSW Performance Tuning Guide

Overview

This guide provides detailed performance tuning strategies for achieving production-grade performance with HeliosDB vector search.

Target Metrics

Indicative figures, measured on representative fixtures; reproduce on your own hardware.

  • QPS: 10,000+ queries per second (1M vectors, ef=50)
  • Recall: 95%+ at Recall@10
  • Latency: p50 < 5ms, p95 < 20ms, p99 < 50ms
  • Build Speed: 1,000+ vectors/sec
  • Memory: ~1.6 GB per 1M vectors (128D, M=16)

Parameter Optimization

M (Max Connections)

Controls graph connectivity and search quality.

MRecall@10QPSMemoryBuild SpeedUse Case
485%30k0.7 GB2,500/sMemory-constrained
890%20k1.0 GB2,000/sFast queries
1695%12k1.6 GB1,200/sProduction (balanced)
3298%8k2.8 GB800/sHigh accuracy
6499%5k5.2 GB500/sMaximum recall

Recommendation: Start with M=16, increase to 32 if recall < 95%

ef_construction (Build Quality)

Controls graph quality during index construction.

ef_constructionRecall@10Build TimeUse Case
5088%1.0×Fast development
10092%1.5×Rapid prototyping
20095%2.0×Production (balanced)
40097%3.0×High accuracy
80098%5.0×Maximum quality

Recommendation: Use 200 for production, 400 for mission-critical applications

ef (Search Quality)

Runtime parameter controlling recall vs speed tradeoff.

efRecall@10LatencyQPSUse Case
1075%0.3 ms35kUltra-fast
2085%0.5 ms25kFast search
5095%0.8 ms15kProduction (balanced)
10097%1.2 ms10kHigh recall
20098%2.0 ms6kMaximum recall
50099%5.0 ms3kResearch/evaluation

Recommendation: Start with ef=50, adjust based on recall requirements

SIMD Optimization

CPU Feature Detection

HeliosDB automatically selects the best SIMD implementation (AVX-512, AVX2, or a scalar fallback for all CPUs).

Performance Comparison

512-dimensional vectors, 10,000 distance calculations:

ImplementationTimeSpeedupInstructions
Scalar5.0 ms1.0×~500 per calc
AVX20.8 ms6.2×~80 per calc
AVX-5120.45 ms11.1×~45 per calc

Enabling SIMD

Linux:

Terminal window
# Check CPU features
cat /proc/cpuinfo | grep flags | grep avx512
# Build with native CPU features
RUSTFLAGS="-C target-cpu=native" cargo build --release

Docker:

# Use host CPU features
docker run --cpus=4 --cpu-shares=1024 \
-e RUSTFLAGS="-C target-cpu=native" \
heliosdb/vector-search

Concurrent Query Optimization

Thread Scaling

HNSW provides excellent read concurrency:

ThreadsQPSLatency (p50)Latency (p95)CPU Usage
112,0000.8 ms1.5 ms100%
223,0000.9 ms1.8 ms200%
444,0001.0 ms2.2 ms400%
880,0001.2 ms3.0 ms800%
16140,0001.5 ms4.5 ms1600%

Scaling efficiency: ~90% up to core count

Memory Optimization

Memory Usage Formula

Total Memory = Base + (N × Vector Size) + (N × Graph Memory)
Where:
- Base: ~100 MB (runtime overhead)
- Vector Size: D × 4 bytes (f32)
- Graph Memory: M × 8 bytes (per connection)
Example (1M vectors, 128D, M=16):
= 100 MB + (1M × 128 × 4) + (1M × 16 × 8)
= 100 MB + 512 MB + 128 MB
= 740 MB (actual: ~800 MB with overhead)

Memory-Constrained Strategies

  1. Reduce M parameter: M=8 uses about half the memory of M=16
  2. Disk-based storage (memory-mapped): vectors are loaded on demand from disk
  3. Distributed sharding: for example, shard across 4 nodes, each handling 25% of the data

Filtered Search Optimization

Pre-filtering vs Post-filtering

Pre-filtering (apply filter before vector search):

  • Pros: Fewer distance calculations, faster
  • Cons: May miss results if filter too restrictive
  • Use when: Filter selectivity > 10%

Post-filtering (apply filter after vector search):

  • Pros: Better recall, no missed results
  • Cons: More distance calculations
  • Use when: Filter selectivity < 10%

Performance Impact

Filter SelectivityPre-filter LatencyPost-filter LatencyRecommendation
90% (10% filtered out)0.9 ms1.2 msPre-filter
50% (50% filtered out)1.5 ms2.0 msPre-filter
10% (90% filtered out)8.0 ms3.5 msPost-filter
1% (99% filtered out)50 ms4.0 msPost-filter

Build Optimization

Batch Insertion

Insert vectors in batches for better cache locality.

Build Performance

1M vectors, 128D, M=16, ef_construction=200:

StrategyBuild TimeThroughput
Sequential850s~1,175 vectors/s
Batch (1k)820s~1,220 vectors/s

Query Latency Optimization

Latency Breakdown

For 1M vectors, 128D, M=16, ef=50:

PhaseLatency% Total
Distance calculations0.5 ms62%
Graph traversal0.2 ms25%
Result sorting0.1 ms13%
Total0.8 ms100%

Optimization Strategies

  1. Reduce distance calculations: lower the ef parameter (for example ef=20 is about 3x faster, with about 10% lower recall)
  2. Use SIMD: ensure AVX2/AVX-512 is enabled
  3. Warm up cache: pre-warm with representative queries
  4. Reduce k (top-k results): returning fewer results (for example k=5) is faster than k=100

Recall Optimization

Recall@10 by Configuration

Mef_constructionefRecall@10QPS
81002082%25k
162005095%12k
3240010098%7k
6480020099%4k

Strategies for >99% Recall

  1. Increase build quality: for example M=64, ef_construction=800
  2. Increase search width: for example ef=500
  3. Use exact search for small datasets (<10k vectors): 100% recall, slower
  4. Verify vector normalization (for Cosine)

Hybrid Search Optimization

Score Fusion Tuning

Weight text matching more heavily for text-heavy queries (e-commerce, documents) and vector similarity more heavily for vector-heavy queries (image search, embeddings). For balanced workloads, A/B test the weights.

Reranking Strategy

Two-stage retrieval (fast approximate search, then accurate reranking of a larger candidate set, for example fetch 100 and return the top 10) costs about 1.5x latency for +2-3% recall.

Production Monitoring

Key Metrics

Track query latency, QPS and index memory usage, and alert when latency exceeds your target (for example 50ms).

Troubleshooting Performance Issues

Issue: Low QPS

Symptoms: <5,000 QPS on modern CPU

Checklist:

  • Verify SIMD enabled (RUSTFLAGS="-C target-cpu=native")
  • Check CPU frequency (not throttled)
  • Reduce ef parameter (200 → 50)
  • Reduce M parameter (32 → 16)
  • Verify no disk I/O (index in memory)
  • Check for lock contention (profiling)

Issue: High Latency

Symptoms: p95 > 100ms

Checklist:

  • Reduce ef (100 → 50 → 20)
  • Check index size (>1M vectors may need sharding)
  • Verify no memory swapping
  • Check GC pauses (if using managed language wrapper)
  • Profile hot paths

Issue: Low Recall

Symptoms: Recall@10 < 90%

Checklist:

  • Increase ef (50 → 100 → 200)
  • Increase ef_construction (200 → 400)
  • Increase M (16 → 32)
  • Verify correct distance metric
  • Check vector normalization (for Cosine)
  • Validate test data quality

Issue: High Memory Usage

Symptoms: >3 GB for 1M vectors (128D)

Checklist:

  • Check M parameter (should be ≤32)
  • Verify no memory leaks
  • Check for duplicate indices
  • Consider memory-mapped storage
  • Profile memory allocations

Advanced Tuning

Query-Specific ef

Adjust ef based on query difficulty.

Use a low ef (for example 20) for easy, high-confidence queries and a higher ef (for example 200) for hard, ambiguous queries.

Summary

Production Configuration (1M vectors, 128D): M=16, ef_construction=200, Cosine distance for normalized embeddings, k=10 and ef=50. Expected: 10-15k QPS, 95%+ recall, <5ms p95 latency.

Key Takeaways:

  1. Start with defaults (M=16, ef_construction=200, ef=50)
  2. Enable SIMD for 5-10x speedup
  3. Use pre-filtering for high selectivity (>10%)
  4. Monitor recall, QPS, and latency continuously
  5. A/B test parameter changes in production