Paper 1 of 5
GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay
Problem
Prior GPU CFR implementations were slower than optimized CPU code because each iteration consists of billions of tiny gather and scatter kernels that finish in microseconds, so kernel launch and framework dispatch overhead dominate runtime.
Approach
GPU-CFR compiles a fixed game once into a static dataflow representation consisting of flat edge and infoset arrays with precomputed indices. It batches operations by tree depth into execution blocks, applies static chance folding and a dual‑lane reach buffer to cut framework operations. The compiled representation is then recorded with CUDA Graph Replay, which replays the entire iteration with a single graph launch. This eliminates per‑iteration host launches while preserving the exact kernel sequence. The method works for any extensive‑form game without changing the CFR update rule.
Result
GPU‑CFR achieves per‑iteration wall‑clock times as low as 0.113 ms (Kuhn) and 0.380 ms (Battleship) on the A100, making it 29.8 to 80.4× faster than the prior GPU baseline (median 44.1×) and 14 to 258× faster than LiteEFG on the four largest games. The compiled CPU path is 2.2 to 51.1× faster than the GPU baseline, with a median speedup of 11.5×.
Why it matters
Researchers and engineers building large extensive‑form game solvers, especially for poker and other imperfect‑information games, can obtain orders‑of‑magnitude speedups on GPUs and CPUs using this static compilation and graph‑replay technique.
Method details
- Compiled float32 solver runs on one NVIDIA A100 80GB PCIe.
- Eight‑game benchmark suite spans 54 to 275,983 infosets.
- Baselines include Kim (2026) sequence‑form CFR+, LiteEFG 0.1.5, and OpenSpiel Python CFR.
- Static chance folding, depth‑level execution blocks, and dual‑lane reach buffer reduce framework operations by up to 18.1x.
- CUDA Graph Replay reduces kernel launches from 87 per iteration to a single graph launch.
- CPU compiled version on 8 threads is 2.2 to 51.1x faster than the A100 baseline.
Numbers
- per‑iteration time (graph) 0.113 ms on Kuhn, fastest system
- speedup over prior GPU baseline 29.8 to 80.4×, median 44.1×
- speedup over LiteEFG up to 258× on largest games
- framework‑operation reduction up to 18.1×
- kernel launches reduced from 87 to 1 per iteration
- CPU compiled (8 threads) 2.2 to 51.1× faster than A100 baseline, median 11.5×
Limitations
The approach cannot capture or accelerate prior GPU baselines that rely on CuPy sparse products, such as Kim (2026), because their iteration cannot be recorded with CUDA graph capture.
graph replay is 258 faster per iteration than LiteEFG.Found in the source text, word for word.
Picked because: Introduces a static dataflow compilation and CUDA graph replay pipeline that yields an 80× speedup for CFR, providing concrete code and performance techniques engineers can adopt for GPU‑accelerated workloads.