Paper 1 of 5
Beyond Scalar Flexibility: From Eligible AI Workloads to Dependable Load Relief
Problem
Prior grid studies model data‑center flexibility as a fixed percentage of load, but no public production trace showed how much eligible load persists across event durations or co‑moves across clusters, so the scalar assumption lacks empirical grounding.
Approach
The authors reconstruct hourly power from a 185‑day trace of 155,410 GPUs, building workload‑conditioned GPU power curves and a four‑layer semantic flexibility envelope. They compute immediate eligible curtailment while retaining allocated‑GPU idle power, then evaluate 95%‑available relief for multiple durations using Monte Carlo medians. Scalars are calibrated against the envelope (mean‑calibrated and tail‑calibrated) to assess over‑/under‑statement. Aggregation across 13 clusters is analyzed to measure firmness gains, and the production scheduler’s delay‑based capacity is quantified.
Result
The study finds that 95%‑available relief declines from 2.51 MW for a one‑hour event to 1.95 MW for a 24‑hour event under full eligibility, and that a mean‑calibrated scalar overstates these values by 17% to 47% while a tail‑calibrated scalar understates the one‑hour product by 6% and overstates the 24‑hour product by 17%; aggregating clusters improves firmness but cross‑cluster covariance limits gains.
Why it matters
Grid planners and demand‑response operators should use the duration‑reliability‑portfolio surface instead of a single flexibility percentage to size reliable load‑relief contracts for AI data‑centers.
Method details
- 4,439 hourly power observations reconstructed from a 185‑day trace of 155,410 GPUs
- 300 Monte Carlo model draws expose parameter sensitivity
- Immediate eligible curtailment averages 3.55 MW (12.1% of workload power, 6.35% of median facility power)
- 95%‑available relief: 2.51 MW for one hour, 2.32 MW for four hours, 1.95 MW for 24 hours
- Aggregating 13 clusters raises four‑hour firmness from 0.38 to 0.66
- Newly deferrable arrivals average 0.008 MW and have zero 95%‑available capacity
Numbers
- 55.8 MW fleet time‑averaged Monte Carlo median facility demand
- 3.55 MW immediate eligible curtailment average
- 12.1% of workload power eligible
- 6.35% of median facility power eligible
- 2.51 MW 95%‑available relief for one hour
- 2.32 MW 95%‑available relief for four hours
- 1.95 MW 95%‑available relief for 24 hours
- 0.38 four‑hour firmness for single cluster
- 0.66 four‑hour firmness after aggregating 13 clusters
- 0.008 MW newly deferrable arrivals average
Limitations
The paper only measures historical eligibility, not actual delivered response, and cannot assess sub‑hour dynamics, checkpointing overhead, or control latency.
The production scheduler exposes almost no additional delay-based capacity: newly deferrable arrivals average 0.008 MW and have zero 95%-available capacity.Found in the source text, word for word.
Picked because: Provides a concrete workload‑semantic flexibility analysis of a large GPU fleet with released trace reconstruction and metrics useful for data‑center DevOps automation.