Work in statistics not fitting other categories.
Statistical practice does not automatically follow methodological innovation. Regularization methods, widely advocated to reduce overfitting and stabilize inference, are readily available in modern software, but are not consistently used by data analysts. We investigate this implementation gap in a large-scale empirical study of trust in, and acceptance of, regularization techniques, based on $N = 606$ data analysts. Drawing on measurement frameworks from technology acceptance research, we survey practitioners and embed a randomized experiment to test whether written recommendation of regularization methods increases trust or intended use. We find no evidence of such an effect. Instead, adoption intentions are strongly associated with analysts' perceptions of ease of implementation and practical benefit, such as improved bias control or interpretability. Perceived social norms also emerge as a central driver. These results indicate that uptake of statistical methodology depends less on formal recommendations than on usability, perceived utility, and community practice.
High-dimensional statistical settings ($p \gg n$) pose fundamental challenges for classical inference, largely due to bias introduced by regularized estimators such as the LASSO. To address this, Javanmard and Montanari (2014) propose a debiased estimator that enables valid hypothesis testing and confidence interval construction. This report examines their debiased LASSO framework, which yields asymptotically normal estimators in high-dimensional settings. We present the key theoretical results underlying this approach, specifically, the construction of an optimized debiased estimator that restores asymptotic normality, which enables the computation of valid confidence intervals and $p$-values. To evaluate the claims of Javanmard and Montanari, a subset of the original simulation study and a re-examination of their real-data analysis are presented. Building on this baseline, we extend the empirical analysis to include the desparsified LASSO, a closely related method referenced but not implemented in the original study. The results demonstrate that while the debiased LASSO achieves reliable coverage and controls Type I error, the LASSO projection estimator can offer improved power in low-signal settings without compromising error rates. Our findings highlight a critical practical trade-off: while the LASSO projection estimator demonstrates superior statistical power in an idealized simulated low-signal setting, the estimation procedure employed by Javanmard and Montanari adapts more robustly to complex correlation networks, yielding superior precision and signal detection in real-world genomic data.
Hilbert's sixth problem calls for the axiomatization of physics, particularly the derivation of macroscopic statistical laws from microscopic mechanical principles. A conceptual difficulty arises in classical probability theory: in continuous spaces every individual microstate has probability zero. In this paper, we introduce a probabilistic framework based on Soft Logic and Soft Numbers in which point events possess infinitesimal Soft probabilities rather than the classical zero. We show that Soft probability can be interpreted as an infinitesimal refinement of classical probability and discuss its implications for statistical mechanics and Hilbert's sixth problem. In addition, we show rigorously how to construct a Mobius strip, based on the soft numbers, and we discuss how this Mobius strip representation with soft numbers allows for a deeper understanding of the nature and character of Hilbert's sixth problem.
2603.29786For events $A$ and $B$, we have \[ \mathbb{P}(A\mid B) > \mathbb{P}(A\mid \neg B) \qquad\Longleftrightarrow\qquad \mathbb{P}(B\mid A) > \mathbb{P}(B\mid \neg A) \] whenever all four quantities are defined. In other words, $B$ is evidence for $A$ if and only if $A$ is evidence for $B$. This note gives seven different proofs of this fact -- by cross-multiplication, covariance, coupling parameters, odds ratios, pointwise mutual information, combinatorial double counting, and mixed discrete derivatives -- and develops a surrounding web of interpretations. Once the marginals $\mathbb{P}(A)$ and $\mathbb{P}(B)$ are fixed, a $2\times 2$ table has only one degree of freedom, so every scalar notion of positive association must be governed by the same signed parameter.
2603.28274Statistics 101, 201, and 202 are three open-source interactive web applications built with R \citep{R} and Shiny \citep{shiny} to support the teaching of introductory statistics and probability. The apps help students carry out common statistical computations -- computing probabilities from standard probability distributions, constructing confidence intervals, conducting hypothesis tests, and fitting simple linear regression models -- without requiring prior knowledge of R or any other programming language. Each app provides numerical results, plots rendered with \texttt{ggplot2} \citep{ggplot2}, and inline mathematical derivations typeset with MathJax \citep{cervone2012mathjax}, so that computation and statistical reasoning appear side by side in a single interface. The suite is organised around a broad pedagogical progression: Statistics~101 introduces probability distributions and their properties; Statistics~201 addresses confidence intervals and hypothesis tests; and Statistics~202 covers the simple linear model. All three apps are freely accessible online and their source code is released under a CC-BY-4.0 license.
A framework for probabilistic forecasting of vessel motion is developed and validated for a semisubmersible operating in long period swell. Bayesian statistical methods are applied to predictions of the heave response from a physics model using numerical wave spectra and measured motion data. Model diagnoses motivate an additional level of complexity required for the error structure in the Bayesian model, specifically to account for heteroskedasticity and time-correlated errors. The hybrid model forecasts were evaluated during periods where the heave resonance and cancellation frequencies were excited. The method is demonstrated to be effective for providing reliable quantification of uncertainty and correcting bias in the raw physics model predictions. This justifies its value for improving the efficiency and safety of offshore operations.
Autocalibration is known to be an important requirement for insurance premiums since it guarantees that premium income balances corresponding claims, on average, not only at portfolio level but also inside each group paying similar premiums. Also, fairness has become a major concern because unfair treatment may expose insurers to lawsuits or reputational damage. Translating fairness into conditional mean independence allows actuaries to combine autocalibration and fairness into the multicalibration concept. This paper studies the properties of multicalibration in an insurance context and proposes practical ways to implement it, through local regression or bias correction within groups including credibility adjustments. A case study based on motor insurance data illustrates the relevance of multicalibration in insurance pricing.
2603.15215This article aims to present a unified framework for ranking-based voting rules based on the use of depth functions on permutations, as a counterpart of deepest voting rules on evaluation introduced in Aubin et al. [2022]. It introduces the notion of depth functions, in continuous sets and in permutation sets, the later using the notion of Fr{é}chet means. Deepest voting procedures are then formally defined, and some classical voting rules are expressed as deepest voting procedures, using a large variety of distances on the set of permutations. Links are done between the depth functions mathematical properties and some behaviours of the voting rule, such as Neutrality, Anonymity, Universality, Condorcet winner/loser property and so on.
Sensitivity analysis methods such as the Cornfield inequality and the E-value were developed to assess the robustness of observed associations against unmeasured confounding -- a major challenge in observational studies. However, the calculation and interpretation of these methods can be difficult for clinicians and interdisciplinary researchers. Recent advances in large language models (LLMs) offer accessible tools that could assist sensitivity analyses, but their reliability in this context has not been studied. We assess four widely used LLMs, ChatGPT, Claude, DeepSeek, and Gemini, on their ability to conduct sensitivity analyses using Cornfield inequalities and E-values. We first extract study-specific information (exposures, outcomes, measured confounders, and effect estimates) from four published observational studies in different fields. Using those information, we develop structured prompts to assess the performance of the LLMs in three aspects: (1) accuracy of E-value calculation, (2) qualitative interpretation of robustness to unmeasured confounding, and (3) suggestion of possible unmeasured confounders. To our knowledge, this is the first study to investigate the use of LLMs for sensitivity analysis. The results show that ChatGPT, Claude, and Gemini accurately reproduce the E-values, whereas DeepSeek shows small biases. Qualitative conclusions from all the LLMs align with the magnitude of the E-values and the reported effect sizes, and all models identify biologically and epidemiologically plausible unmeasured confounders. These findings suggest that, when guided by structured prompting, LLMs can effectively assist in evaluating unmeasured confounding, and thereby can support study design and decision-making in observational studies.
Research and Development is the largest budget position in the pharmaceutical industry, with clinical trials being a critical, yet costly and time-consuming component to inform decisions. Beyond drug efficacy, the probability of success and efficiency of research and development are highly dependent on the approaches used for designing, analyzing, and interpreting clinical trials. Deep understanding of statistical methodology and quantitative approaches is therefore essential. Consequently, dedicated methodology groups have emerged in mid-size and large pharmaceutical companies and CROs. Their remit is to lead the conception and implementation of innovative quantitative methodologies in order to improve drug development, often by addressing complexities or offering more efficient designs. To achieve this, they collaborate internally and externally (e.g., with academics, regulators) to identify common challenges and tear down silos in order to invest in methods with the highest impact on efficiency and value to the portfolio. Given the immense financial stakes of drug development -- where delays carry massive implications -- these groups represent a critical strategic investment. However, to realize this business impact, statistical innovations must be rigorously validated and seamlessly integrated. This manuscript explores the setup, remit, and value of dedicated methodology groups, alongside the critical organizational considerations and success factors required to maximize their impact on the speed, efficiency, and probability of success.
Data analyses are often constructed in an imperative manner, where commands representing actions taken on the data are issued sequentially. The publication of these commands, along with the data, is essential to the reproducibility of the analysis by others. However, simply presenting the code and the results of running the code can hide important details about the data analyst's premises, expectations, and assumptions about the data. Understanding this analysis reasoning can be critical to evaluating the quality of an analysis and for suggesting possible improvements. We argue that a formal representation of a data analysis that externalizes its logical construction offers more useful information for statically illustrating an analyst's reasoning. Such a formal representation would allow for the evaluation of some aspects of a data analysis without the need for the data, the visualization of the logical connections leading to a conclusion, and the ability to assess the sensitivity of an analyst's assumptions to unexpected features in the data. In this paper we describe an implementation of this formal representation and how it might be applied to some common data analysis tasks.
Detecting interaction effects (IEs) in meta-regression is challenging, especially when few studies are available and many plausible interactions are considered. In many meta-analyses, interpretability is essential, which limits the use of complex machine learning methods. Tree-based approaches offer a potentially useful compromise, but their role in meta-regression with random effects is not yet well understood. This paper examines how traditional linear and tree-based methods can support variable selection for IEs in random effects meta-regression. We compare test-based and information-criterion-based linear selection procedures with meta-CART approaches. These include fixed effect and random effects trees and their stability-selected ensemble variants. All methods are evaluated using a real-world meta-analytic dataset and a plasmode simulation study. The data-generating process assumes linear IEs and is complemented by settings with nonlinear interactions. Our results show that under strictly linear interactions, linear selection methods perform as expected and achieve superior performance for IE detection. Tree-based methods are more conservative when the number of studies is small, but become competitive as sample size increases, particularly the stability-selected variants. When IEs deviate from strict linearity, even in simple ways, the performance of linear methods deteriorates, whereas tree-based approaches, especially stability-selected fixed effect trees, provide a more robust alternative. Overall, stability-selected random effects trees are useful complementary tools for IE detection in applied meta-regression, particularly for metric covariates. They are well suited for pre-selection and sensitivity analyses, and selection frequency patterns in tree ensembles can help reveal structural patterns in the data.
Linear programming is widely used for decision-making in science, engineering, and operations research, yet in many modern applications the coefficients entering the constraints and objective are not known exactly and must be learned from data. Classical stochastic and robust optimization offer two influential paradigms for handling such uncertainty, but they typically treat the underlying uncertainty description as given and do not directly integrate priors and updated to posteriors guarantees. This paper develops a Bayesian framework for linear programming in which uncertain quantities are modeled probabilistically, updated through observed data, and propagated into optimization through posterior feasibility requirements. We present two complementary computational strategies: a credible-region robustification that converts posterior uncertainty into deterministic protection, and a posterior-scenario approach that uses sampled posterior realizations to construct tractable optimization problems with finite-sample interpretability. We also propose a Monte Carlo certification procedure that provides conservative, data-conditioned assessments of residual infeasibility. Simulation experiments show that the proposed framework substantially improves safety relative to naive plug-in decisions, while a real-data study on single-cell transcriptomic data demonstrates that the approach can produce scientifically interpretable decisions together with explicit uncertainty-aware feasibility diagnostics. The proposed methodology offers a unified bridge between Bayesian learning, optimization under uncertainty, and practical decision certification.
Statistics educators recommend teaching with real data with relevant contexts, but defining relevancy is challenging and varies by student. We investigated whether providing student choice of data context increases engagement through a quasi-experiment in two sections of an introductory probability and statistics course at a large public university (n=65 consenting students). Sections alternated as treatment and control: during their treatment, students chose weekly homework from three similar instructor-provided options varying by data context; during control weeks, they received randomly assigned contexts. We found no significant difference in homework grades between treatment and control conditions. However, thematic analysis revealed students with choice reported enhanced engagement and motivation, greater appreciation for statistics' real-world value, and increased autonomy. Students overwhelmingly preferred contexts relevant to their interests, experiences, daily lives, and career paths-though preferences varied considerably across individuals. Based on these findings, we provide four recommendations for statistics educators: (1) use real data with authentic contexts, (2) select contexts students care about, (3) incorporate variety across data contexts, and (4) consider choice as a pedagogical tool.
2603.03828The philosophical foundations of statistics involve issues in theoretical statistics, such as goals and methods to meet these goals, and interpretation of the meaning of inference using statistics. They are related to the philosophy of science and to the philosophy of probability. We review the core and partly interrelated themes and place them in context.
2603.02372The observation of life on Earth is generally accepted to be uninformative concerning the probability of life on other Earth-like planets, a belief first formalized by Brandon Carter and based on the selection effect of our existence. In a similar way, the Drake equation is either presented as estimate of the total number of active, communicative, extraterrestrial civilizations in our Galaxy ($n^g_{\rm civ}$), i.e. excluding humanity, or humanity is included in the estimate but judged to be an uninformative data point. Daniel Whitmire has recently challenged the Carter abiogenesis argument, claiming the logic behind it is flawed, as the conditional likelihoods used by Carter in Bayes' theorem are not evaluated prior to the occurrence of the evidence of life on Earth, but posterior. Doing so correctly, the anthropic selection effect is removed and the observation of life on Earth is informative after all. Following this argument, we treat the Drake equation as estimate of all technological civilizations in a statistical counting experiment and include the data point of humanity as informative evidence. This allows one to set a pessimistic lower limit on $n^o_{\rm civ}$ for the observable universe, $n^o_{\rm civ} > 0.051$ at 95\% C.L., or $n^g_{\rm civ} > 8\times10^{-13}$ at 95\% C.L. for the Galaxy. In particular, this excludes models that predict $n^o_{\rm civ}\ll 1$ for the observable universe and refines the allowable parameter space for hypotheses like Rare Earth. Our analysis substantially reduces the portion of the Drake equation parameter space that predicts humanity is alone; when applying the lower limit this study finds $P(n^o_{\rm civ}>1 |\, {\rm humanity}) = 97.6\%$, making solitude in the observable universe a disfavored outcome. For the low-end estimate of $n^o_{\rm civ}\! =\! 1$ we calculate a probability of 42\% for the existence of other communicating civilizations.
Keeping pace with rapidly evolving technology is a key challenge in teaching statistics. To equip students with essential skills for the modern workplace, educators must integrate relevant technologies into the statistical curriculum where possible. University-level statistics education has experienced substantial technological change, particularly in the tools and practices that underpin teaching and learning. Statistical programming has become central to many courses, with R widely used and Python increasingly incorporated into statistics and data analytics programmes. Additionally, coding practices, database management, and machine learning now feature within some statistics curricula. Looking ahead, we anticipate a growing emphasis on artificial intelligence (AI), particularly the pedagogical implications of generative AI tools such as ChatGPT. In this article, we explore these technological developments and discuss strategies for their integration into contemporary statistics education.
2602.15581What, if anything, should a frequentist say about a single realized confidence interval (CI) and its chance of having covered the parameter? Jerzy Neyman's original answer was to refuse any nondegenerate probability for coverage ex post and, instead, to "state that the interval covers". In this paper I argue that the usual frequentist machinery already supports a different reading. I treat the coverage event as a Bernoulli random variable, with the nominal level 1-alpha as its design-based success probability, and view "confidence" as a probability forecast for that Bernoulli outcome. Using strictly proper scoring rules, I show that 1-alpha is the unique optimal constant forecast for coverage, both before and after observing the data, and that it remains optimal post-trial in common unbounded, translation-invariant models with pivot-based CIs. When the design yields a theta-free statistic--such as the relative width of the interval in a finite-window uniform model--the conditional coverage given that statistic provides a nonconstant, design-based refinement of 1-alpha that strictly improves predictive performance. Two thought experiments, a Monty Hall-style shell game and the "lost submarine" example of Morey et al. (2016), illustrate how this perspective resolves familiar interpretational puzzles about CIs without appealing to priors or single-case subjective degrees of belief. I conclude with simple "what to do when you see an interval" guidance for applied work and some implications for teaching confidence intervals as tools for forecasting long-run coverage. Keywords: Confidence intervals, coverage probability, proper scoring rules, probabilistic forecasting, frequentist inference Disclaimer: The findings and conclusions in this report are those of the author and do not necessarily represent the official position of the Centers for Disease Control and Prevention
2603.00098The use of profiling evidence in criminal trials is a longstanding controversy in legal epistemology and evidence law theory. Many scholars, even when they oppose its use at trial, still assume that profiling evidence can be probative of guilt. We reject that assumption. Profiling evidence may support a generic hypothesis, but is not evidence that the defendant is guilty of the specific crime of which they are accused. We contrast profiling evidence with case-specific evidence, which speaks more directly to the facts of the case. Our critique departs from others by grounding the argument in a probabilistic analysis of evidentiary value. We also explore the implications of our account for debates about stereotyping.
2602.15562In Neyman's original formulation, a 1-alpha confidence interval procedure is justified by its long-run coverage properties, and a single realized interval is to be described only by the slogan that it either covers the parameter or it does not. On this view, post-data probability statements about the coverage of an individual interval are taken to be conceptually out of bounds. In this paper, I present two kinds of arguments against treating that "either-or" reading as the only legitimate interpretation of confidence. The first is informal, via a set of thought experiments in which the same joint probability model is used to compute both forward-looking and backward-looking probabilities for occurred-but-unobserved events. The second is more formal, recasting the standard confidence-interval construction in terms of infinite sequences of trials and their associated 0/1 coverage indicators. In that representation, the design-level coverage probability 1-alpha and the degenerate conditional probabilities given the full data appear simply as different conditioning levels of the same model. I argue that a strict behavioristic reading that privileges only the latter is in tension with the very mathematical machinery used to define long-run error rates. I then sketch an alternative view of confidence as a predictive probability (or forecast) about the coverage indicator, together with a simple normative rule for when intermediate probabilities for single coverage events should be allowed. Keywords: confidence intervals; coverage probability; frequentist inference; single-case probability; predictive probability; Neyman. Disclaimer: The findings and conclusions in this report are those of the author and do not necessarily represent the official position of the Centers for Disease Control and Prevention.