Distributions of positive signals in pyrosequencing
Yong Kong
TL;DR
This work derives exact distributions for the number of positive signals $r$ in pyrosequencing pyrograms, modeled under fixed r-seq length (FRLM) and fixed flow cycle (FFCM) frameworks using probability generating functions. By solving recurrences and obtaining closed-form GFs, the authors obtain explicit means and variances for $r$ and $f$, show Gaussian limiting behavior, and reveal a robust, approximate relation $ar{r}(f) oughly 2f$ independent of nucleotide probabilities. A key practical result is that simulations and theory align, validating the models for predicting pyrogram signal distributions, which have implications for base-calling thresholds and software design. The work also clarifies how the distributions of $f$, $n$, and $r$ relate, including transitive intuitions and the impact of equal vs unequal nucleotide probabilities on variance, enabling improved understanding of pyrosequencing data generation and analysis.
Abstract
Pyrosequencing is one of the important next-generation sequencing technologies. We derive the distribution of the number of positive signals in pyrograms of this sequencing technology as a function of flow cycle numbers and nucleotide probabilities of the target sequences. As for the distribution of sequence length, we also derive the distribution of positive signals for the fixed flow cycle model. Explicit formulas are derived for the mean and variance of the distributions. A simple result for the mean of the distribution is that the mean number of positive signals in a pyrogram is approximately twice the number of flow cycles, regardless of nucleotide probabilities. The statistical distributions will be useful for instrument and software development for pyrosequencing and other related platforms.
