Skip to content

Graceful Degradation System

Graceful Degradation System

Overview

The Graceful Degradation system provides intelligent load shedding for HeliosDB under overload conditions. Instead of crashing or rejecting all queries when resources are constrained, the system gracefully degrades performance by automatically adjusting operational parameters.

Architecture

Degradation Modes

The system operates in four distinct modes, automatically transitioning based on system metrics:

  1. Normal Mode - Full capacity operation

    • All query priorities accepted
    • 100% parallelism
    • Standard timeouts (1x)
    • Full resource allocation (100%)
  2. Cautious Mode - Slightly restricted

    • All query priorities still accepted
    • 80% parallelism (20% reduction)
    • Extended timeouts (1.5x)
    • 80% resource allocation
  3. Restricted Mode - Significantly limited

    • Only Normal+ priority queries accepted
    • 50% parallelism (50% reduction)
    • Extended timeouts (2x)
    • 50% resource allocation
  4. Survival Mode - Minimal operations only

    • Only Critical priority queries accepted
    • 20% parallelism (80% reduction)
    • Maximum timeouts (3x)
    • 30% resource allocation

Mode Transition Thresholds

Transitions occur based on three key metrics:

MetricCautiousRestrictedSurvival
CPU Utilization80%90%95%
Memory Utilization85%92%97%
Queue Depth1005001000

Transition Rule: The system enters the mode suggested by the worst (highest) metric.

Hysteresis

To prevent mode thrashing (rapid transitions between modes), the system implements hysteresis:

  • Hysteresis Factor: 0.85 (configurable)
  • Effect: Metrics must improve to 85% of the threshold before downgrading

Example:

  • Enter Cautious mode at CPU = 80%
  • Can only exit Cautious when CPU < 80% × 0.85 = 68%

Minimum Mode Duration

Prevents rapid mode changes:

  • Default: 30 seconds
  • System must stay in a mode for at least this duration before transitioning

Integration Points

1. Admission Control

The degradation manager integrates with admission control: queries below the minimum priority for the current mode are rejected.

2. Query Scheduler

The scheduler adjusts parallelism based on degradation mode.

3. Resource Manager

Resource allocation is scaled based on the current mode.

4. Query Execution

Query timeouts are adjusted based on system load.

Configuration

Conservative Thresholds

For systems requiring maximum stability: CPU thresholds of 70% / 85% / 93% and memory thresholds of 75% / 88% / 95% (Cautious / Restricted / Survival), a wider hysteresis band (hysteresis_factor 0.90) and a longer stability period (min_mode_duration_secs 60).

Aggressive Thresholds

For systems that can tolerate higher load: CPU thresholds of 85% / 93% / 97% and memory thresholds of 90% / 95% / 98% (Cautious / Restricted / Survival), a narrower hysteresis band (hysteresis_factor 0.80) and faster recovery (min_mode_duration_secs 15).

Best Practices

1. Set Appropriate Thresholds

  • Conservative: Set lower thresholds for critical production systems
  • Aggressive: Set higher thresholds for development or test environments
  • Calibrate: Monitor system behavior and adjust based on actual load patterns

2. Monitor Mode Transitions

  • Log all mode transitions for analysis
  • Alert operators when entering Restricted or Survival modes
  • Track time spent in each mode

3. Tune Hysteresis

  • Too low: Risk of mode thrashing
  • Too high: Slow recovery from degraded modes
  • Recommended: 0.80 - 0.90 range

4. Query Priority Strategy

  • Critical: Health checks, monitoring, essential operations
  • High: User-facing queries, important batch jobs
  • Normal: Standard workload
  • Background: Analytics, reports, maintenance

5. Recovery Monitoring

  • Track auto-recovery success rate
  • Monitor time to recovery
  • Ensure auto-recovery is enabled in production

Metrics and Observability

Key Metrics to Monitor

  1. Current Degradation Mode

    • Gauge: Current mode (0=Normal, 1=Cautious, 2=Restricted, 3=Survival)
  2. Mode Transition Rate

    • Counter: Total mode transitions
    • Rate: Transitions per hour
  3. Time in Mode

    • Histogram: Duration in each mode
    • Percentage: Time distribution across modes
  4. Query Rejection Rate

    • Counter: Queries rejected due to degradation
    • Rate: Rejections per second
  5. Recovery Time

    • Histogram: Time spent in degraded modes before recovery

Prometheus Metrics

  • heliosdb_degradation_mode: current degradation mode (0=Normal, 1=Cautious, 2=Restricted, 3=Survival)
  • heliosdb_degradation_transitions_total: total number of degradation mode transitions
  • heliosdb_degradation_rejections_total: total queries rejected due to degradation

Troubleshooting

System Stuck in Degraded Mode

Symptoms: System remains in Cautious/Restricted mode despite low load

Possible Causes:

  1. Hysteresis factor too high
  2. Minimum mode duration too long
  3. Auto-recovery disabled
  4. Load metrics not updating

Solutions:

  1. Reduce hysteresis_factor (e.g., from 0.90 to 0.85)
  2. Reduce min_mode_duration_secs
  3. Enable enable_auto_recovery
  4. Verify system load updates are happening

Frequent Mode Thrashing

Symptoms: Rapid transitions between modes

Possible Causes:

  1. Hysteresis factor too low
  2. Minimum mode duration too short
  3. Thresholds too close together
  4. Bursty workload patterns

Solutions:

  1. Increase hysteresis_factor (e.g., from 0.80 to 0.90)
  2. Increase min_mode_duration_secs
  3. Spread threshold values further apart
  4. Consider workload smoothing

Queries Rejected Too Aggressively

Symptoms: Many queries rejected, but system not actually overloaded

Possible Causes:

  1. Thresholds set too conservatively
  2. Priority levels not aligned with workload

Solutions:

  1. Raise degradation thresholds
  2. Review and adjust query priorities
  3. Consider using more gradual degradation levels

Performance Impact

Overhead

  • CPU: < 0.1% (periodic metric checks)
  • Memory: ~1KB per manager instance
  • Latency: < 10μs per query admission check

Benefits

  • Availability: Prevents system crashes under overload
  • Fairness: Ensures critical queries are processed
  • Recovery: Automatic recovery when load decreases
  • Observability: Comprehensive metrics for analysis

References