Muhammad Huzaifa Β· Sina Mavali Β· Thorsten Eisenhofer
CISPA Helmholtz Center for Information Security
[2026.10] Paper released on arXiv. Code is being refactored and will be released soon.
We show that learned latent communication links can compromise the safety of multi-agent systems built from safety-aligned agents. Even benign link training can increase harmful compliance, and attacks via supervised optimization, data poisoning, or reward-guided optimization amplify it further. Reward-guided updates to the links alone can repair compromised systems.
@article{huzaifa2026safety,
title = {Safety of Latent Communication in Multi-Agent Systems},
author = {Huzaifa, Muhammad and Mavali, Sina and Eisenhofer, Thorsten},
journal = {arXiv preprint arXiv:2609.39788},
year = {2026}
}β Star or watch this repository to be notified when the code is released.