Table of Contents
Fetching ...

Do Large Language Models Respect Contracts? Evaluating and Enforcing Contract-Adherence in Code Generation

Soohan Lim, Joonghyuk Hahn, Hyunwoo Park, Sang-Ki Ko, Yo-Sub Han

TL;DR

PACT introduces a contract-aware evaluation framework that extends functional correctness benchmarks by focusing on contract adherence in code generation. It leverages an SMT-based two-stage CVT generation pipeline to produce precise contract-violating inputs and evaluates code generation under Contract Specification (CS) and Example-Augmented Specification (EAS) prompts. The framework provides new metrics (AVC, TS, AAR, AAP) to quantify contract enforcement and reveals a practical trade-off: stronger contract adherence can come at the expense of pure functional correctness. These results highlight the need for multi-objective training and evaluation to build more robust, contract-aware code generation systems with real-world reliability.

Abstract

Prevailing code generation benchmarks, such as HumanEval+ and MBPP+, primarily evaluate large language models (LLMs) with pass@k on functional correctness using well-formed inputs. However, they ignore a crucial aspect of real-world software: adherence to contracts-the preconditions and validity constraints that dictate how ill-formed inputs must be rejected. This critical oversight means that existing benchmarks fail to measure, and models consequently fail to generate, truly robust and reliable code snippets. We introduce PACT, a program assessment and contract-adherence evaluation framework, to bridge this gap. PACT is the first framework designed to systematically evaluate and enhance contract-adherence in LLM-generated code snippets alongside functional correctness. PACT's contributions are threefold: First, it provides a comprehensive test-suite corpus focused on contract violations, extending HumanEval+ and MBPP+. Second, it enables a systematic analysis of code generation under varied prompting conditions. This analysis demonstrates that augmenting prompts with contract-violating test cases significantly enhance a model's ability to respect contracts compared to using contract description alone. Finally, it introduces novel metrics to rigorously quantify contract adherence in both test generation and code generation. By revealing critical errors that conventional benchmarks overlook, PACT provides the rigorous and interpretable metrics to evaluate the robustness of LLM-generated code snippets in both functionality and contract-adherence. Our code and data are available at https://github.com/suhanmen/PACT.

Do Large Language Models Respect Contracts? Evaluating and Enforcing Contract-Adherence in Code Generation

TL;DR

PACT introduces a contract-aware evaluation framework that extends functional correctness benchmarks by focusing on contract adherence in code generation. It leverages an SMT-based two-stage CVT generation pipeline to produce precise contract-violating inputs and evaluates code generation under Contract Specification (CS) and Example-Augmented Specification (EAS) prompts. The framework provides new metrics (AVC, TS, AAR, AAP) to quantify contract enforcement and reveals a practical trade-off: stronger contract adherence can come at the expense of pure functional correctness. These results highlight the need for multi-objective training and evaluation to build more robust, contract-aware code generation systems with real-world reliability.

Abstract

Prevailing code generation benchmarks, such as HumanEval+ and MBPP+, primarily evaluate large language models (LLMs) with pass@k on functional correctness using well-formed inputs. However, they ignore a crucial aspect of real-world software: adherence to contracts-the preconditions and validity constraints that dictate how ill-formed inputs must be rejected. This critical oversight means that existing benchmarks fail to measure, and models consequently fail to generate, truly robust and reliable code snippets. We introduce PACT, a program assessment and contract-adherence evaluation framework, to bridge this gap. PACT is the first framework designed to systematically evaluate and enhance contract-adherence in LLM-generated code snippets alongside functional correctness. PACT's contributions are threefold: First, it provides a comprehensive test-suite corpus focused on contract violations, extending HumanEval+ and MBPP+. Second, it enables a systematic analysis of code generation under varied prompting conditions. This analysis demonstrates that augmenting prompts with contract-violating test cases significantly enhance a model's ability to respect contracts compared to using contract description alone. Finally, it introduces novel metrics to rigorously quantify contract adherence in both test generation and code generation. By revealing critical errors that conventional benchmarks overlook, PACT provides the rigorous and interpretable metrics to evaluate the robustness of LLM-generated code snippets in both functionality and contract-adherence. Our code and data are available at https://github.com/suhanmen/PACT.
Paper Structure (33 sections, 4 equations, 12 figures, 3 tables)

This paper contains 33 sections, 4 equations, 12 figures, 3 tables.

Figures (12)

  • Figure 1: PACT's contract-violating test uncovers an implicit constraint that conventional functional tests miss, proving the need for contract-aware evaluation.
  • Figure 2: Running example of PACT with an Example-Augmented Specification (EAS) prompt, which integrates Contract Specification (CS) and contract-violating test cases (CVTs) to enforce contract-aware code generation.
  • Figure 3: Code and contracts for HumanEval.
  • Figure 4: Code and contracts for MBPP.
  • Figure 5: Code generated by DeepSeek with the contract specification (CS) prompt.
  • ...and 7 more figures