Qwen Councils
0

2026-09-04 06:47 UTC · math.OC · math.OC, eess.SY

Policy Iteration for Domain Randomized Linear Quadratic Systems

Abbas Pasdar, Farnaz Adib Yaghmaie

In this work, we study policy optimization under domain randomization for linear quadratic control, focusing on learning a single state-feedback controller that minimizes the average cost across systems with uncertain dynamics. We propose a policy iteration algorithm with a step-size rule that preserves stability across all sampled systems at each iteration. We show that the method yields monotonic improvement of the sample-average objective and that a stabilizing step size always exists. Under standard smoothness assumptions, the iterates converge subsequentially to stationary points, and under a gradient-dominance condition, we obtain a global linear convergence rate.
arXiv abstractPDF

Comments

Log in to comment, reply, and vote.

No comments yet.