Table of Contents
Fetching ...

Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots

Haochen Su, Cristian Meo, Francesco Stella, Andrea Peirone, Kai Junge, Josie Hughes

TL;DR

This paper addresses the challenge of deploying Vision-Language-Action models on soft robots to bridge the embodiment gap with rigid manipulators. It introduces a structured finetuning pipeline and an open-source soft-robot demonstration dataset to evaluate cross-embodiment transfer for two state-of-the-art VLA models, OpenVLA-OFT and $π_0$. The experiments reveal that out-of-the-box policies fail on the soft Embuddy due to nonlinear, underactuated dynamics, but targeted finetuning closes the gap and yields performance on par with rigid baselines, with OpenVLA-OFT showing robust, high-frequency control. The work demonstrates the practical potential of combining VLA policies with soft robotics to enable safe, flexible embodied AI in human-centered environments and paves the way for broader task and platform exploration.

Abstract

Robotic systems are increasingly expected to operate in human-centered, unstructured environments where safety, adaptability, and generalization are essential. Vision-Language-Action (VLA) models have been proposed as a language guided generalized control framework for real robots. However, their deployment has been limited to conventional serial link manipulators. Coupled by their rigidity and unpredictability of learning based control, the ability to safely interact with the environment is missing yet critical. In this work, we present the deployment of a VLA model on a soft continuum manipulator to demonstrate autonomous safe human-robot interaction. We present a structured finetuning and deployment pipeline evaluating two state-of-the-art VLA models (OpenVLA-OFT and $π_0$) across representative manipulation tasks, and show while out-of-the-box policies fail due to embodiment mismatch, through targeted finetuning the soft robot performs equally to the rigid counterpart. Our findings highlight the necessity of finetuning for bridging embodiment gaps, and demonstrate that coupling VLA models with soft robots enables safe and flexible embodied AI in human-shared environments.

Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots

TL;DR

This paper addresses the challenge of deploying Vision-Language-Action models on soft robots to bridge the embodiment gap with rigid manipulators. It introduces a structured finetuning pipeline and an open-source soft-robot demonstration dataset to evaluate cross-embodiment transfer for two state-of-the-art VLA models, OpenVLA-OFT and . The experiments reveal that out-of-the-box policies fail on the soft Embuddy due to nonlinear, underactuated dynamics, but targeted finetuning closes the gap and yields performance on par with rigid baselines, with OpenVLA-OFT showing robust, high-frequency control. The work demonstrates the practical potential of combining VLA policies with soft robotics to enable safe, flexible embodied AI in human-centered environments and paves the way for broader task and platform exploration.

Abstract

Robotic systems are increasingly expected to operate in human-centered, unstructured environments where safety, adaptability, and generalization are essential. Vision-Language-Action (VLA) models have been proposed as a language guided generalized control framework for real robots. However, their deployment has been limited to conventional serial link manipulators. Coupled by their rigidity and unpredictability of learning based control, the ability to safely interact with the environment is missing yet critical. In this work, we present the deployment of a VLA model on a soft continuum manipulator to demonstrate autonomous safe human-robot interaction. We present a structured finetuning and deployment pipeline evaluating two state-of-the-art VLA models (OpenVLA-OFT and ) across representative manipulation tasks, and show while out-of-the-box policies fail due to embodiment mismatch, through targeted finetuning the soft robot performs equally to the rigid counterpart. Our findings highlight the necessity of finetuning for bridging embodiment gaps, and demonstrate that coupling VLA models with soft robots enables safe and flexible embodied AI in human-shared environments.
Paper Structure (28 sections, 3 equations, 8 figures, 1 table)

This paper contains 28 sections, 3 equations, 8 figures, 1 table.

Figures (8)

  • Figure 1: Continuum soft robot used for the experimental study. A: The full view of the continuum robot - Embuddy. B: Actuation and structural schematic of Embuddy, indicating tendons, joints, and motors. C: Detailed view of a single section. D: Demonstration setup for the soft and rigid robot.
  • Figure 2: Inference success rate comparisons between OpenVLA-OFT and $\pi_0$ on UR5 and Soft Robot embodiments.
  • Figure 3: Setup and processed image views for UR5 experiments(baseline). Top row: original views, with resolution 640x480; Bottom row: processed views, with resolution 256x256.
  • Figure 4: Setup and processed image views for soft robot experiments. Top row: original views, with resolution 640x480; Bottom row: processed views, with resolution 256x256.
  • Figure 5: The training losses for task 1 and 3 on soft robot with OpenVLA-OFT and $\pi_0$
  • ...and 3 more figures