Adversarially Robust Multi-agent Reinforcement Learning
Andi Nika
Max Planck Institute for Software Systems
23 Oct 2026, 3:00 pm - 5:00 pm
Saarbrücken building E1 5, room 029
SWS Student Defense Talks - Thesis Defense
Reinforcement learning (RL) has emerged as a fundamental approach to
decision-making in machine learning, with applications across a wide
range of real-world domains, and several practical extensions, such as
multi-agent RL (MARL) and RL from human feedback (RLHF). Despite the
growing successful applications of these systems, there exists an
inherent threat when it comes to applying them in the real world, where
ill-intentioned third parties may intervene in both their training
process and deployment. This typically has catastrophic consequences, ...
Reinforcement learning (RL) has emerged as a fundamental approach to
decision-making in machine learning, with applications across a wide
range of real-world domains, and several practical extensions, such as
multi-agent RL (MARL) and RL from human feedback (RLHF). Despite the
growing successful applications of these systems, there exists an
inherent threat when it comes to applying them in the real world, where
ill-intentioned third parties may intervene in both their training
process and deployment. This typically has catastrophic consequences,
where even small and inexpensive perturbations to the environment may
cause the system to substantially diverge from the desired behavior. The
purpose of this thesis is to provide a thorough investigation of various
adversarial attacks and robustness against such attacks to (MA)RL and
(MA)RLHF systems. In particular, we study training-time and test-time
attacks in MARL and propose algorithms that are shown to be provably
robust against such attacks. Beyond MARL, we establish fundamental
statistical results for RLHF, provide a rigorous characterization of
poisoning attacks to RLHF, and develop robust extensions for MARLHF. Our
work is centered around robust algorithmic approaches with provable
guarantees under common assumptions.
Read more