Graceful Degradation System
Graceful Degradation System
Overview
The Graceful Degradation system provides intelligent load shedding for HeliosDB under overload conditions. Instead of crashing or rejecting all queries when resources are constrained, the system gracefully degrades performance by automatically adjusting operational parameters.
Architecture
Degradation Modes
The system operates in four distinct modes, automatically transitioning based on system metrics:
-
Normal Mode - Full capacity operation
- All query priorities accepted
- 100% parallelism
- Standard timeouts (1x)
- Full resource allocation (100%)
-
Cautious Mode - Slightly restricted
- All query priorities still accepted
- 80% parallelism (20% reduction)
- Extended timeouts (1.5x)
- 80% resource allocation
-
Restricted Mode - Significantly limited
- Only Normal+ priority queries accepted
- 50% parallelism (50% reduction)
- Extended timeouts (2x)
- 50% resource allocation
-
Survival Mode - Minimal operations only
- Only Critical priority queries accepted
- 20% parallelism (80% reduction)
- Maximum timeouts (3x)
- 30% resource allocation
Mode Transition Thresholds
Transitions occur based on three key metrics:
| Metric | Cautious | Restricted | Survival |
|---|---|---|---|
| CPU Utilization | 80% | 90% | 95% |
| Memory Utilization | 85% | 92% | 97% |
| Queue Depth | 100 | 500 | 1000 |
Transition Rule: The system enters the mode suggested by the worst (highest) metric.
Hysteresis
To prevent mode thrashing (rapid transitions between modes), the system implements hysteresis:
- Hysteresis Factor: 0.85 (configurable)
- Effect: Metrics must improve to 85% of the threshold before downgrading
Example:
- Enter Cautious mode at CPU = 80%
- Can only exit Cautious when CPU < 80% × 0.85 = 68%
Minimum Mode Duration
Prevents rapid mode changes:
- Default: 30 seconds
- System must stay in a mode for at least this duration before transitioning
Integration Points
1. Admission Control
The degradation manager integrates with admission control: queries below the minimum priority for the current mode are rejected.
2. Query Scheduler
The scheduler adjusts parallelism based on degradation mode.
3. Resource Manager
Resource allocation is scaled based on the current mode.
4. Query Execution
Query timeouts are adjusted based on system load.
Configuration
Conservative Thresholds
For systems requiring maximum stability: CPU thresholds of 70% / 85% / 93% and memory thresholds of 75% / 88% / 95% (Cautious / Restricted / Survival), a wider hysteresis band (hysteresis_factor 0.90) and a longer stability period (min_mode_duration_secs 60).
Aggressive Thresholds
For systems that can tolerate higher load: CPU thresholds of 85% / 93% / 97% and memory thresholds of 90% / 95% / 98% (Cautious / Restricted / Survival), a narrower hysteresis band (hysteresis_factor 0.80) and faster recovery (min_mode_duration_secs 15).
Best Practices
1. Set Appropriate Thresholds
- Conservative: Set lower thresholds for critical production systems
- Aggressive: Set higher thresholds for development or test environments
- Calibrate: Monitor system behavior and adjust based on actual load patterns
2. Monitor Mode Transitions
- Log all mode transitions for analysis
- Alert operators when entering Restricted or Survival modes
- Track time spent in each mode
3. Tune Hysteresis
- Too low: Risk of mode thrashing
- Too high: Slow recovery from degraded modes
- Recommended: 0.80 - 0.90 range
4. Query Priority Strategy
- Critical: Health checks, monitoring, essential operations
- High: User-facing queries, important batch jobs
- Normal: Standard workload
- Background: Analytics, reports, maintenance
5. Recovery Monitoring
- Track auto-recovery success rate
- Monitor time to recovery
- Ensure auto-recovery is enabled in production
Metrics and Observability
Key Metrics to Monitor
-
Current Degradation Mode
- Gauge: Current mode (0=Normal, 1=Cautious, 2=Restricted, 3=Survival)
-
Mode Transition Rate
- Counter: Total mode transitions
- Rate: Transitions per hour
-
Time in Mode
- Histogram: Duration in each mode
- Percentage: Time distribution across modes
-
Query Rejection Rate
- Counter: Queries rejected due to degradation
- Rate: Rejections per second
-
Recovery Time
- Histogram: Time spent in degraded modes before recovery
Prometheus Metrics
heliosdb_degradation_mode: current degradation mode (0=Normal, 1=Cautious, 2=Restricted, 3=Survival)heliosdb_degradation_transitions_total: total number of degradation mode transitionsheliosdb_degradation_rejections_total: total queries rejected due to degradation
Troubleshooting
System Stuck in Degraded Mode
Symptoms: System remains in Cautious/Restricted mode despite low load
Possible Causes:
- Hysteresis factor too high
- Minimum mode duration too long
- Auto-recovery disabled
- Load metrics not updating
Solutions:
- Reduce
hysteresis_factor(e.g., from 0.90 to 0.85) - Reduce
min_mode_duration_secs - Enable
enable_auto_recovery - Verify system load updates are happening
Frequent Mode Thrashing
Symptoms: Rapid transitions between modes
Possible Causes:
- Hysteresis factor too low
- Minimum mode duration too short
- Thresholds too close together
- Bursty workload patterns
Solutions:
- Increase
hysteresis_factor(e.g., from 0.80 to 0.90) - Increase
min_mode_duration_secs - Spread threshold values further apart
- Consider workload smoothing
Queries Rejected Too Aggressively
Symptoms: Many queries rejected, but system not actually overloaded
Possible Causes:
- Thresholds set too conservatively
- Priority levels not aligned with workload
Solutions:
- Raise degradation thresholds
- Review and adjust query priorities
- Consider using more gradual degradation levels
Performance Impact
Overhead
- CPU: < 0.1% (periodic metric checks)
- Memory: ~1KB per manager instance
- Latency: < 10μs per query admission check
Benefits
- Availability: Prevents system crashes under overload
- Fairness: Ensures critical queries are processed
- Recovery: Automatic recovery when load decreases
- Observability: Comprehensive metrics for analysis