SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark

Boxi Yu; Yang Cao; Yuzhong Zhang; Liting Lin; Junjielong Xu; Zhiqing Zhong; Qinghua Xu; Guancheng Wang; Jialun Cao; Shing-Chi Cheung; Pinjia He; Lionel Briand

SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark

Boxi Yu, Yang Cao, Yuzhong Zhang, Liting Lin, Junjielong Xu, Zhiqing Zhong, Qinghua Xu, Guancheng Wang, Jialun Cao, Shing-Chi Cheung, Pinjia He, Lionel Briand

TL;DR

SWE-ABS is presented, an adversarial framework that strengthens test suites through a two-stage pipeline: (1) coverage-driven augmentation using program slicing to target untested code regions, and (2) mutation-driven adversarial testing that synthesizes plausible but incorrect patches to expose semantic blind spots.

Abstract

The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we show that this performance is inflated. Our re-evaluation reveals that one in five "solved" patches from the top-30 agents are semantically incorrect, passing only because weak test suites fail to expose their errors. We present SWE-ABS, an adversarial framework that strengthens test suites through a two-stage pipeline: (1) coverage-driven augmentation using program slicing to target untested code regions, and (2) mutation-driven adversarial testing that synthesizes plausible but incorrect patches to expose semantic blind spots. On SWE-Bench Verified (500 instances), SWE-ABS strengthens 50.2% of instances, a 25.1x improvement over prior work, and rejects 19.71% of previously passing patches. As a result, the top agent's score decreases from 78.80% to 62.20%, leading to significant leaderboard reshuffling, with the previous top-ranked agent dropping to fifth place.

SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark

TL;DR

Abstract

Paper Structure (99 sections, 9 equations, 2 figures, 11 tables)

This paper contains 99 sections, 9 equations, 2 figures, 11 tables.

Introduction
Adversarial Benchmark Strengthening.
Results.
Contributions.
Related Work
Benchmarks for Code Generation.
Test Augmentation.
Mutation Testing.
Method
Stage I: Coverage-Driven Test Augmentation
Generating Initial Tests
Decoupling Gold-Patch Dependencies
Program Slicing
Coverage-Guided Augmentation
Stage II: Mutation-Driven Adversarial Strengthening
...and 84 more sections

Figures (2)

Figure 1: Motivating example: Weak tests enable silent failures. (a) A real Django issue (django__django-10973) refactoring requires converting passwords to strings before passing them to subprocess.run via the PGPASSWORD environment variable, because subprocess.run only accepts string-valued environment variables. (b) The original test suite tests only string passwords, enabling semantically incorrect agent patches to pass despite failing to handle non-string inputs. (c) A strengthened test introduces non-string inputs, exposing the missing type conversion and correctly failing the invalid patches.
Figure 2: Overview of the SWE-ABS framework.

SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark

TL;DR

Abstract

SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark

Authors

TL;DR

Abstract

Table of Contents

Figures (2)