I really find this method of training policies really intersting.
Maybe I should write a bit more. But basically, an alternative to Zeroth-order methods to train policies for robots using differentiable simulators, like MuJoCo's MJX.