Table of Contents
Fetching ...

LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Senyu Fei, Siyin Wang, Junhao Shi, Zihao Dai, Jikun Cai, Pengfang Qian, Li Ji, Xinzhe He, Shiduo Zhang, Zhaoye Fei, Jinlan Fu, Jingjing Gong, Xipeng Qiu

TL;DR

This work reveals that state-of-the-art Vision-Language-Action models exhibit brittle robustness under realistic perturbations, challenging the notion that high LIBERO benchmark scores reflect true competency. It systematically analyzes seven single-dimension perturbations across diverse architectures, finds that camera viewpoints and robot initial states drive large performance drops, and uncovers that language inputs are often ignored. The authors formalize compositional generalization gaps and demonstrate significant perturbation interactions, then address these issues by constructing LIBERO-Plus, a large, automated benchmark with multi-dimensional perturbations and difficulty levels. They show that training on generalized, diverse data markedly improves robustness, underscoring the need for evaluation frameworks that capture reliability under realistic variability and for models capable of flexible generalization in embodied tasks.

Abstract

Visual-Language-Action (VLA) models report impressive success rates on robotic manipulation benchmarks, yet these results may mask fundamental weaknesses in robustness. We perform a systematic vulnerability analysis by introducing controlled perturbations across seven dimensions: objects layout, camera viewpoints, robot initial states, language instructions, light conditions, background textures and sensor noise. We comprehensively analyzed multiple state-of-the-art models and revealed consistent brittleness beneath apparent competence. Our analysis exposes critical weaknesses: models exhibit extreme sensitivity to perturbation factors, including camera viewpoints and robot initial states, with performance dropping from 95% to below 30% under modest perturbations. Surprisingly, models are largely insensitive to language variations, with further experiments revealing that models tend to ignore language instructions completely. Our findings challenge the assumption that high benchmark scores equate to true competency and highlight the need for evaluation practices that assess reliability under realistic variation.

LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

TL;DR

This work reveals that state-of-the-art Vision-Language-Action models exhibit brittle robustness under realistic perturbations, challenging the notion that high LIBERO benchmark scores reflect true competency. It systematically analyzes seven single-dimension perturbations across diverse architectures, finds that camera viewpoints and robot initial states drive large performance drops, and uncovers that language inputs are often ignored. The authors formalize compositional generalization gaps and demonstrate significant perturbation interactions, then address these issues by constructing LIBERO-Plus, a large, automated benchmark with multi-dimensional perturbations and difficulty levels. They show that training on generalized, diverse data markedly improves robustness, underscoring the need for evaluation frameworks that capture reliability under realistic variability and for models capable of flexible generalization in embodied tasks.

Abstract

Visual-Language-Action (VLA) models report impressive success rates on robotic manipulation benchmarks, yet these results may mask fundamental weaknesses in robustness. We perform a systematic vulnerability analysis by introducing controlled perturbations across seven dimensions: objects layout, camera viewpoints, robot initial states, language instructions, light conditions, background textures and sensor noise. We comprehensively analyzed multiple state-of-the-art models and revealed consistent brittleness beneath apparent competence. Our analysis exposes critical weaknesses: models exhibit extreme sensitivity to perturbation factors, including camera viewpoints and robot initial states, with performance dropping from 95% to below 30% under modest perturbations. Surprisingly, models are largely insensitive to language variations, with further experiments revealing that models tend to ignore language instructions completely. Our findings challenge the assumption that high benchmark scores equate to true competency and highlight the need for evaluation practices that assess reliability under realistic variation.
Paper Structure (54 sections, 8 equations, 19 figures, 9 tables)

This paper contains 54 sections, 8 equations, 19 figures, 9 tables.

Figures (19)

  • Figure 1: Robustness to object layout perturbations. Comparison of different models under confounding and displacement perturbations, as well as their overall robustness.
  • Figure 2: Illumination robustness and extreme ablation tests. The term Light denotes the condition with light perturbation applied. 3rd Black and All Black represent conditions where only the third-view image is masked and where images from both views are masked, respectively.
  • Figure 3: Accuracy of different models on instruction removed (a) and target modified (b) tasks. Light bars: original success rate with language instruction; (a) dark bars: success rate after removing the instruction; (b) Dark bars: success rate under altered task goal and instruction (task substitution).
  • Figure 4: Heatmap of conditional probabilities under pairwise perturbations. Upper triangular entries represent independence-based products of single-dimension probabilities, while lower triangular entries show actual joint outcomes.
  • Figure 5: Model performance trends across perturbation difficulty levels. The line plots show the success rate of each model as the intensity of four different perturbation dimensions increases.
  • ...and 14 more figures