Networked systems, from power grids to traffic networks and cloud clusters, carry loads across nodes with limited capacity. A node whose load exceeds its capacity fails and sheds its load onto its neighbors, which can trigger a system-wide cascade. We study how to allocate a fixed capacity budget across nodes to resist these cascades under local load redistribution. The problem is difficult because no optimal allocation is known, and the fail-or-survive objective is non-differentiable and piecewise constant, so exact and gradient-based optimization methods do not directly apply. We introduce TANGCO (Topology-Aware Neural Graph-Guided Capacity Optimization), which uses a graph neural network policy trained through the cascade simulator with policy-gradient learning and a heuristic anchor. We evaluate TANGCO on five synthetic graph families and five real networks spanning power, road, air, and Internet topologies. The learned policy improves on the best of four hand-designed heuristics in all 450 synthetic instances and in 40 of 45 real-network conditions, with robustness gains ranging from 1.6% to 246%. The learned policies transfer to unseen graphs within a family and partially across related topologies, and TANGCO$^{pre}$, pre-trained on synthetic graphs, matches per-network training on unseen real networks. Training scales near-linearly with graph size, and TANGCO$^{pre}$ allocates on a new network with no per-target training, matching the deployment cost of a hand-designed heuristic. Free-vector variants without the GNN, stay close to the heuristics, so the graph representation carries the gain beyond numerical search. Finally, analysis of the learned allocations identifies when local risk is sufficient, leads to an improved closed-form heuristic, and reveals the regimes where a topology-aware learned policy remains necessary.
Cyber-physical power systems are vulnerable to cascading failures caused by interdependencies between power and communication infrastructures. Because evaluating large N-k contingency sets with a high-fidelity simulator is computationally expensive, this paper develops a machine-learning surrogate using the previously published Modified Implicative Interdependency Model (MIIM) as the ground-truth cascade simulator. The surrogate predicts contingency severity from leakage-free structural features and derives an association-based component-criticality ranking for resilience screening. On the IEEE 118-bus system, Gradient Boosting achieves a held-out Spearman correlation of 0.849 for contingency-severity ranking. Using five-fold out-of-fold predictions, the resulting component ranking achieves a Spearman correlation of 0.838 with the MIIM-derived ranking and closely approaches the observed cross-sample reproducibility level. Feature-ablation results show that inter-layer dependency features drive most of the surrogate's advantage, while end-to-end screening is approximately 158x faster than direct MIIM evaluation. The results support a two-stage workflow in which the surrogate screens candidate contingencies and components, and MIIM provides selective verification rather than directly identifying optimal hardening actions.