Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
Mason Nakamura, Abhinav Kumar, Saaduddin Mahmud, Sahar Abdelnabi, Shlomo Zilberstein, Eugene Bagdasarian
TL;DR
Terrarium introduces a modular blackboard-based testbed for studying safety, privacy, and security in instruction-augmented LLM-driven MAS solving DCOPs. It formalizes MAS as instruction-augmented DCOPs with a ground-truth objective $F^\star$ and context-dependent policies, enabling controlled evaluation of misalignment, data exfiltration, and DoS-style attacks. The framework emphasizes modularity, configurability, and observability, providing attack surfaces and metrics to rapidly prototype defenses and compare configurations. By enabling diverse problem settings (Meeting Scheduling, Personal Assistant, Smart-Home) and security scenarios, Terrarium advances trustworthy multi-agent systems with a practical, extensible testbed for researchers and practitioners.
Abstract
A multi-agent system (MAS) powered by large language models (LLMs) can automate tedious user tasks such as meeting scheduling that requires inter-agent collaboration. LLMs enable nuanced protocols that account for unstructured private data, user constraints, and preferences. However, this design introduces new risks, including misalignment and attacks by malicious parties that compromise agents or steal user data. In this paper, we propose the Terrarium framework for fine-grained study on safety, privacy, and security in LLM-based MAS. We repurpose the blackboard design, an early approach in multi-agent systems, to create a modular, configurable testbed for multi-agent collaboration. We identify key attack vectors such as misalignment, malicious agents, compromised communication, and data poisoning. We implement three collaborative MAS scenarios with four representative attacks to demonstrate the framework's flexibility. By providing tools to rapidly prototype, evaluate, and iterate on defenses and designs, Terrarium aims to accelerate progress toward trustworthy multi-agent systems.
