mpo maxMPOMAX MENANG MAKSIMAL >. Public group · 14We introduce a new algorithm for reinforcement learning called Maximum a-posteriori Policy Optimisation (MPO) based on coordinate ascent on a relative-entropy