Cascade Containment Benchmark BlueChips evaluated N-1-1 contingency analysis on IEEE 140-bus and PEGASE 1,354-bus public power-grid benchmarks. The benchmark covers 14,656 fault scenarios and asks whether a method can contain cascading failure before it propagates. On the PEGASE 1,354-bus network, BlueChips FIBI-Penalised contained 79.1% of cascades while Graph Lasso dropped to 0.8%. A deep learning GCN achieved 14.5% cascade success rate on smaller cases at 101ms per instance and did not scale to the 1,354-bus benchmark. The public result is important because average-case accuracy is not the relevant metric for regional infrastructure failure. The relevant question is whether a worst-case fault can be isolated before it becomes a cascade. Full benchmark summary: benchmarks.html