Table of Contents
Fetching ...

COWs and their Hybrids: A Statistical View of Custom Orthogonal Weights

Chad Schafer, Larry Wasserman, Mikael Kuusela

TL;DR

The paper formalizes and extends the COWs framework for separating signal from background in particle physics by relaxing the conditional-independence assumption between discriminant $M$ and control $T$ and representing the joint density as a sum of basis-component densities. It develops estimation tools, identifiability analysis (the Herd), and concrete extensions, including mixtures of copulas, varying-coefficient representations, and least-squares implementations without explicit weights. Key contributions include a rigorous treatment of identifiability, a suite of goodness-of-fit and confidence-band methods, and connections to NMF and copula-based approaches that broaden applicability. The work provides practical algorithms and theoretical insights that enhance signal extraction in high-energy physics analyses and guide future research on identifiability, extensions, and robust inference.

Abstract

A recurring challenge in high energy physics is inference of the signal component from a distribution for which observations are assumed to be a mixture of signal and background events. A standard assumption is that there exists information encoded in a discriminant variable that is effective at separating signal and background. This can be used to assign a signal weight to each event, with these weights used in subsequent analyses of one or more control variables of interest. The custom orthogonal weights (COWs) approach of Dembinski, et al.(2022), a generalization of the sPlot approach of Barlow (1987) and Pivk and Le Diberder (2005), is tailored to address this objective. The problem, and this method, present interesting and novel statistical issues. Here we formalize the assumptions needed and the statistical properties, while also considering extensions and alternative approaches.

COWs and their Hybrids: A Statistical View of Custom Orthogonal Weights

TL;DR

The paper formalizes and extends the COWs framework for separating signal from background in particle physics by relaxing the conditional-independence assumption between discriminant and control and representing the joint density as a sum of basis-component densities. It develops estimation tools, identifiability analysis (the Herd), and concrete extensions, including mixtures of copulas, varying-coefficient representations, and least-squares implementations without explicit weights. Key contributions include a rigorous treatment of identifiability, a suite of goodness-of-fit and confidence-band methods, and connections to NMF and copula-based approaches that broaden applicability. The work provides practical algorithms and theoretical insights that enhance signal extraction in high-energy physics analyses and guide future research on identifiability, extensions, and robust inference.

Abstract

A recurring challenge in high energy physics is inference of the signal component from a distribution for which observations are assumed to be a mixture of signal and background events. A standard assumption is that there exists information encoded in a discriminant variable that is effective at separating signal and background. This can be used to assign a signal weight to each event, with these weights used in subsequent analyses of one or more control variables of interest. The custom orthogonal weights (COWs) approach of Dembinski, et al.(2022), a generalization of the sPlot approach of Barlow (1987) and Pivk and Le Diberder (2005), is tailored to address this objective. The problem, and this method, present interesting and novel statistical issues. Here we formalize the assumptions needed and the statistical properties, while also considering extensions and alternative approaches.
Paper Structure (21 sections, 84 equations, 12 figures)

This paper contains 21 sections, 84 equations, 12 figures.

Figures (12)

  • Figure 1: The Synthetic model used throughout this work. The top two panels show the signal and background distribution in both of the discriminant and control variables. The bottom panel shows a sample drawn from this distribution.
  • Figure 2: Results from the application of sPlot/COWs to the simulated data set shown in Figure \ref{['Syntheticmodel']}. The left panel shows the weight functions $w_{\mathsf{s}}$ and $w_{\mathsf{b}}$, corresponding to the signal and background, respectively. The right panel shows the estimate of $nzh_1(t)$ that results when observations are weighted using $w_{\mathsf{s}}$. The solid line depicts the expected value of these bin counts under the assumed model.
  • Figure 3: Results from the application of mixture weighting approach to the simulated data set shown in Figure \ref{['Syntheticmodel']}. The left panel shows the weight functions $w'_1$ and $w'_2$, corresponding to the signal and background, respectively. The right panel shows the estimate of the signal distribution in the control variable resulting from weighting using $w'_1$. The solid line depicts the expected value of these bin counts under the assumed model.
  • Figure 4: Data simulated under the model with dependence between $M$ and $T$. Note that within a slice of fixed value of $M$, the probability that an observation is signal is independent of $T$.
  • Figure 5: Results when using both COWs and mixture weighting on the data shown in Figure \ref{['fig::copulatoydata']}.
  • ...and 7 more figures