Entanglement as a Structural Complexity Axis: A PAC-Bayesian View of Generalization in Quantum Policies and Value Functions
Summary
This paper derives a PAC-Bayesian generalization bound for quantum policies where the complexity term is the Fisher effective dimension, rather than the raw parameter count. The authors demonstrate that entanglement inflates this effective dimension, creating an independent axis of structural complexity that governs the train-test gap in quantum reinforcement learning.
Mathematical/empirical assessment
The derivation of the Fisher-form bound (Eq. 2) and the effective dimension (Eq. 3) is mathematically coherent and well-grounded. The empirical evaluation is rigorous, utilizing controlled experiments that fix parameter count $d$ while varying entangling connectivity, and extends impressively to real hardware validation.
Strengths
Isolating entanglement as an independent complexity axis is a highly plausible and novel contribution. The controlled experiments (e.g., Table 1) cleanly validate the theoretical predictions, showing that the bound acts effectively as a ranking certificate for circuits of identical size, outperforming raw parameter counting.
Concerns
The bound is admittedly loose in absolute terms, relying heavily on its ranking property rather than providing tight numerical guarantees. Additionally, the authors correctly note the mechanism is depth-gated and erodes in the barren plateau regime, which limits its immediate scalability to deep, large-scale quantum policies.
Final decision
Weak accept