Table of Contents
Fetching ...

Position: Require Frontier AI Labs To Release Small "Analog" Models

Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer, Philip Quirke

TL;DR

The paper addresses the safety–innovation tension in frontier AI by proposing an analog-model mandate: frontier AI labs must publicly release small, distilled analogs trained on the same data and objectives. It supports this with cross-scale transfer evidence showing safety interventions, interpretability tools, and benchmarks discovered in small models generalize to larger systems, aided by representational convergence and smooth scaling laws. The policy details cover definitions, release timelines, licensing, enforcement, and safeguards, along with risk–benefit analyses and precedents from other sectors. The proposed approach aims to accelerate safety research, improve transparency, and preserve competitive incentives, offering a practical path to safer, more open AI development with scalable public benefits.

Abstract

Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovation tradeoff. This paper argues for an alternative regulatory approach that ensures AI safety while actively promoting innovation: mandating that large AI laboratories release small, openly accessible analog models (scaled-down versions) trained similarly to and distilled from their largest proprietary models. Analog models serve as public proxies, allowing broad participation in safety verification, interpretability research, and algorithmic transparency without forcing labs to disclose their full-scale models. Recent research demonstrates that safety and interpretability methods developed using these smaller models generalize effectively to frontier-scale systems. By enabling the wider research community to directly investigate and innovate upon accessible analogs, our policy substantially reduces the regulatory burden and accelerates safety advancements. This mandate promises minimal additional costs, leveraging reusable resources like data and infrastructure, while significantly contributing to the public good. Our hope is not only that this policy be adopted, but that it illustrates a broader principle supporting fundamental research in machine learning: deeper understanding of models relaxes the safety-innovation tradeoff and lets us have more of both.

Position: Require Frontier AI Labs To Release Small "Analog" Models

TL;DR

The paper addresses the safety–innovation tension in frontier AI by proposing an analog-model mandate: frontier AI labs must publicly release small, distilled analogs trained on the same data and objectives. It supports this with cross-scale transfer evidence showing safety interventions, interpretability tools, and benchmarks discovered in small models generalize to larger systems, aided by representational convergence and smooth scaling laws. The policy details cover definitions, release timelines, licensing, enforcement, and safeguards, along with risk–benefit analyses and precedents from other sectors. The proposed approach aims to accelerate safety research, improve transparency, and preserve competitive incentives, offering a practical path to safer, more open AI development with scalable public benefits.

Abstract

Recent proposals for regulating frontier AI models have sparked concerns about the cost of safety regulation, and most such regulations have been shelved due to the safety-innovation tradeoff. This paper argues for an alternative regulatory approach that ensures AI safety while actively promoting innovation: mandating that large AI laboratories release small, openly accessible analog models (scaled-down versions) trained similarly to and distilled from their largest proprietary models. Analog models serve as public proxies, allowing broad participation in safety verification, interpretability research, and algorithmic transparency without forcing labs to disclose their full-scale models. Recent research demonstrates that safety and interpretability methods developed using these smaller models generalize effectively to frontier-scale systems. By enabling the wider research community to directly investigate and innovate upon accessible analogs, our policy substantially reduces the regulatory burden and accelerates safety advancements. This mandate promises minimal additional costs, leveraging reusable resources like data and infrastructure, while significantly contributing to the public good. Our hope is not only that this policy be adopted, but that it illustrates a broader principle supporting fundamental research in machine learning: deeper understanding of models relaxes the safety-innovation tradeoff and lets us have more of both.
Paper Structure (34 sections, 1 figure, 1 table)

This paper contains 34 sections, 1 figure, 1 table.

Figures (1)

  • Figure 1: The Analog-Model Mandate and Its Effect on the Safety–Innovation Frontier. (A) Frontier AI models are distilled into small, openly released “analog” models, enabling broad participation in safety testing, interpretability research, and algorithmic transparency; insights from this open loop are then transferred back to improve the safety of the original large model. (B) By providing a public proxy for each proprietary system, the analog-model mandate (dashed curve) shifts the attainable safety–innovation frontier outward, relaxing the traditional tradeoff (solid curve) between rapid capability development and robust safeguards.