Table of Contents
Fetching ...

A Compositional Paradigm for Foundation Models: Towards Smarter Robotic Agents

Luigi Quarantiello, Elia Piccoli, Jack Bell, Malio Li, Giacomo Carfì, Eric Nuertey Coleman, Gerlando Gramaglia, Lanpei Li, Mauro Madeddu, Irene Testa, Vincenzo Lomonaco

TL;DR

Foundation Models deliver broad, cross-modal capabilities but often struggle to adapt to dynamic real-world scenarios without retraining. The paper advocates a paradigm combining Continual Learning and Compositionality to create modular, reusable components that can be dynamically composed to address new tasks with reduced compute. It demonstrates this approach with image-classification adapters (HAM) and robotic-adaptation architectures (WSA) alongside OpenX-Embodiment and Vision-Language-Action models like OpenVLA, reporting improved accuracy and efficiency over baselines. The work outlines a practical path toward smarter, more adaptable robotic agents and highlights benchmarks that support cross-embodiment transfer and continual adaptation in open-world contexts.

Abstract

The birth of Foundation Models brought unprecedented results in a wide range of tasks, from language to vision, to robotic control. These models are able to process huge quantities of data, and can extract and develop rich representations, which can be employed across different domains and modalities. However, they still have issues in adapting to dynamic, real-world scenarios without retraining the entire model from scratch. In this work, we propose the application of Continual Learning and Compositionality principles to foster the development of more flexible, efficient and smart AI solutions.

A Compositional Paradigm for Foundation Models: Towards Smarter Robotic Agents

TL;DR

Foundation Models deliver broad, cross-modal capabilities but often struggle to adapt to dynamic real-world scenarios without retraining. The paper advocates a paradigm combining Continual Learning and Compositionality to create modular, reusable components that can be dynamically composed to address new tasks with reduced compute. It demonstrates this approach with image-classification adapters (HAM) and robotic-adaptation architectures (WSA) alongside OpenX-Embodiment and Vision-Language-Action models like OpenVLA, reporting improved accuracy and efficiency over baselines. The work outlines a practical path toward smarter, more adaptable robotic agents and highlights benchmarks that support cross-embodiment transfer and continual adaptation in open-world contexts.

Abstract

The birth of Foundation Models brought unprecedented results in a wide range of tasks, from language to vision, to robotic control. These models are able to process huge quantities of data, and can extract and develop rich representations, which can be employed across different domains and modalities. However, they still have issues in adapting to dynamic, real-world scenarios without retraining the entire model from scratch. In this work, we propose the application of Continual Learning and Compositionality principles to foster the development of more flexible, efficient and smart AI solutions.
Paper Structure (5 sections, 2 tables)