Probability and Statistics
Probability spaces, distributions, and estimators form the mathematical backbone of the theory of measurement error developed in Measurement, SI Units, and the Theory of Errors and of the statistical mechanics of Statistical Mechanics. The source manuscript reserves a part for this material but leaves it unwritten: it fixes only the division of the subject into a discrete and a continuous theory. That division is retained in the first two sections below, whose content — the descriptive statistics of a distribution and of a sample — is authored rather than ported, the source having left both headings empty. So are the remaining four sections: the central limit theorem, which Part I deferred to here, and then the theory of inference built on it — estimation, hypothesis testing, and the correction that a search over a range of an unknown parameter forces on any significance quoted at a point. Those last three exist because the collider search chapters of Part XI — Quantum Field Theory and the Standard Model test a hypothesis and quote a significance, and a number quoted without the mathematics that defines it is not evidence.
Discrete statistics
A discrete model is one whose possible outcomes can be listed. Everything in it is a sum, and every identity below is an identity between absolutely convergent series, so nothing is needed here beyond the theory of series of Real Analysis. The continuous case of Section 11.2 replaces those sums by integrals and inherits the algebra unchanged. Together the two supply the vocabulary — mean, variance, standard deviation, moments, median, quantile, covariance, correlation — that the rest of this chapter and the error theory of Measurement, SI Units, and the Theory of Errors are written in.
Definitions
A discrete probability space is a pair \((\Omega,\Pr)\) in which the sample space \(\Omega\) is a finite or countably infinite set, whose elements are the possible outcomes, and \(\Pr\) assigns to each \(\omega\in\Omega\) a weight \(\Pr(\set{\omega})\ge0\) with
An event is a subset \(A\subseteq\Omega\), and its probability is \(\Pr(A)=\sum_{\omega\in A}\Pr(\set{\omega})\). A probability is a pure number and carries no SI unit. Rests on Definition 7.45.
The series in Equation (11.1) has non-negative terms, so it converges to the same value in every order and a sum over an unordered countable set is well defined; the same remark licenses every rearrangement made below. The general setting, in which \(\Omega\) need not be countable and the events form a \(\sigma\)-algebra, is not needed anywhere in this chapter and is set out in [Billingsley:1995].
Let \(A,B\subseteq\Omega\) be events and \(A_{1},A_{2},\dots\) pairwise disjoint events. Then
-
\(\Pr(\Omega\setminus A)=1-\Pr(A)\), and \(A\subseteq B\) implies \(\Pr(A)\le\Pr(B)\);
-
\(\Pr(\varnothing)=0\), \(\Pr(\Omega)=1\) and \(0\le\Pr(A)\le1\);
-
\(\Pr(A\cup B)=\Pr(A)+\Pr(B)-\Pr(A\cap B)\);
-
\(\Pr\!\left(\bigcup_{i}A_{i}\right)=\sum_{i}\Pr(A_{i})\).
Rests on Definition 11.1.
Derivation. Derives Proposition 11.2. Every claim is a statement about a series of non-negative terms indexed by a countable set, and such a series may be split and regrouped at will. (i) Splitting the sum Equation (11.1) over \(A\) and its complement gives the first identity; splitting the sum over \(B\) over \(A\) and \(B\setminus A\) gives \(\Pr(B)=\Pr(A)+\Pr(B\setminus A)\ge\Pr(A)\). (ii) The empty sum is \(0\) and the sum over \(\Omega\) is \(1\) by Equation (11.1); the bounds then follow from monotonicity, \(\Pr(\varnothing)\le\Pr(A)\le\Pr(\Omega)\). (iii) Split \(A\cup B\) into \(A\setminus B\), \(B\setminus A\) and \(A\cap B\): the left side is the sum of the three, and the right side counts \(A\cap B\) twice and subtracts it once. (iv) Each outcome of the union lies in exactly one \(A_{i}\), so grouping the terms of the sum over the union by the index \(i\) reproduces the right-hand side.
∎A discrete random variable is a map \(X:\Omega\rightarrow\R\) taking at most countably many values. Its probability mass function is
a non-negative function summing to \(1\) over the values \(X\) attains. If the outcome carries a physical quantity — a count, a length in \(\mathrm{m}\), an energy in \(\mathrm{J}\) — then \(X\) carries that SI unit and \(p_{X}\) does not. Rests on Definition 11.1.
The expectation of a discrete random variable is
defined whenever \(\sum_{\omega}\abs{X(\omega)}\Pr(\set{\omega})<\infty\), in which case \(X\) is called integrable. The expectation carries the SI unit of \(X\). Rests on Definitions 7.45 and 11.3.
Let \(X\) and \(Y\) be integrable random variables on one discrete probability space and let \(a,b,c\in\R\). Then
-
(transfer) for every \(g:\R\rightarrow\R\) with \(\sum_{x}\abs{g(x)}\,p_{X}(x)<\infty\), \(\avg{g(X)}=\sum_{x}g(x)\,p_{X}(x)\); in particular \(\avg{X}=\sum_{x}x\,p_{X}(x)\);
-
(linearity) \(aX+bY\) is integrable and \(\avg{aX+bY}=a\avg{X}+b\avg{Y}\), and \(\avg{c}=c\);
-
(monotonicity) if \(X\le Y\) at every outcome then \(\avg{X}\le\avg{Y}\), and \(\abs{\avg{X}}\le\avg{\abs{X}}\).
Rests on Definition 11.4 and Proposition 11.2.
Derivation. Derives Proposition 11.5. (i) Group the outcomes according to the value of \(X\): the terms of \(\sum_{\omega}g(X(\omega))\Pr(\set{\omega})\) with \(X(\omega)=x\) contribute \(g(x)\sum_{\omega:X(\omega)=x}\Pr(\set{\omega})=g(x)p_{X}(x)\), and the regrouping is legitimate because the series converges absolutely by hypothesis. (ii) Termwise, \(\left(aX+bY\right)(\omega)\Pr(\set{\omega}) =aX(\omega)\Pr(\set{\omega})+bY(\omega)\Pr(\set{\omega})\); both series on the right converge absolutely, so their sum may be split, and \(\abs{aX+bY}\le\abs{a}\abs{X}+\abs{b}\abs{Y}\) gives the integrability. For a constant the sum is \(c\) times Equation (11.1). (iii) Termwise again, since each weight is non-negative; and \(\pm X\le\abs{X}\) gives \(\pm\avg{X}\le\avg{\abs{X}}\).
∎Those three properties are the whole toolkit. Every identity in the rest of this section is a formal consequence of them, which is why the continuous theory of Section 11.2 can take them over wholesale.
Let \(\avg{\abs{X}^{n}}<\infty\). The \(n\)th raw moment of \(X\) is \(\avg{X^{n}}\) and its \(n\)th central moment is \(\mu_{n}=\avg{\left(X-\mu\right)^{n}}\), where \(\mu=\avg{X}\) is the mean. The variance and the standard deviation are
and the skewness and excess kurtosis are \(\gamma_{1}=\mu_{3}/\sigma^{3}\) and \(\gamma_{2}=\mu_{4}/\sigma^{4}-3\). Rests on Proposition 11.5.
If \(X\) carries the SI unit \(u\) then \(\mu\) and \(\sigma\) carry \(u\), the variance carries \(u^{2}\), the \(n\)th moment carries \(u^{n}\), and \(\gamma_{1}\) and \(\gamma_{2}\) are pure numbers — which is what makes the last two comparable between a length distribution and an energy distribution. A finite second moment implies a finite first one, since \(\abs{x}\le\tfrac12\left(1+x^{2}\right)\); the same inequality is used again below to make a covariance exist. The subtracted \(3\) in \(\gamma_{2}\) is a convention fixed by Proposition 11.29, which gives every Gaussian \(\mu_{4}=3\sigma^{4}\) and hence \(\gamma_{2}=0\).
If \(\avg{X^{2}}<\infty\) then
\(\operatorname{var}\left(aX+b\right)=a^{2}\operatorname{var}X\) for constants \(a\) and \(b\), and \(\operatorname{var}X=0\) if and only if \(X=\mu\) with probability one. Rests on Definition 11.6 and Proposition 11.5.
Derivation. Derives Proposition 11.7. Expand the square and use linearity twice: \(\avg{X^{2}-2\mu X+\mu^{2}}=\avg{X^{2}}-2\mu\avg{X}+\mu^{2} =\avg{X^{2}}-\mu^{2}\), since \(\avg{X}=\mu\). For the second claim, \(\avg{aX+b}=a\mu+b\), so \(\left(aX+b\right)-\avg{aX+b}=a\left(X-\mu\right)\) and squaring produces the factor \(a^{2}\). For the third, \(\operatorname{var}X=\sum_{\omega}\left(X(\omega)-\mu\right)^{2} \Pr(\set{\omega})\) is a series of non-negative terms, so it vanishes exactly when every term does, that is when \(X(\omega)=\mu\) at every outcome of non-zero probability.
∎Equation (11.5) is what allows a variance to be accumulated in one pass over the data, keeping only \(\sum X_{i}\) and \(\sum X_{i}^{2}\). It is also a trap in finite-precision arithmetic: when \(\sigma\ll\abs{\mu}\) the two terms on its right are nearly equal and their difference loses most of its significant digits, so a numerical implementation subtracts the mean first and uses Equation (11.4) directly.
For \(\alpha\in(0,1)\), an \(\alpha\)-quantile of \(X\) is any \(q\in\R\) with
A median is a \(\tfrac12\)-quantile. A quantile carries the SI unit of \(X\). Rests on Definition 11.3.
Write \(F(x)=\Pr(X\le x)\). By countable additivity (Proposition 11.2(iv)) \(F(x_{n})\rightarrow F(q)\) whenever \(x_{n}\downarrow q\) and \(F(x_{n})\rightarrow\Pr(X<q)\) whenever \(x_{n}\uparrow q\); also \(F\rightarrow1\) at \(+\infty\) and \(F\rightarrow0\) at \(-\infty\). The set \(\set{x:F(x)\ge\alpha}\) is therefore a non-empty half-line \([q,\infty)\) with \(q\) finite, and that \(q\) satisfies both requirements of Equation (11.6): the first by construction, the second because \(F(x)<\alpha\) for \(x<q\) forces \(\Pr(X<q)\le\alpha\). Uniqueness fails in general: for \(X\) taking the values \(0\) and \(1\) with probability \(\tfrac12\) each, every point of \([0,1]\) is a median.
Let \(X\) be integrable and let \(m\) be a median of \(X\). Then
and if in addition \(\avg{X^{2}}<\infty\) then
with equality exactly at \(c=\mu\). Rests on Definition 11.8 and Proposition 11.7.
Derivation. Derives Proposition 11.10. For Equation (11.8) write \(X-c=\left(X-\mu\right)+\left(\mu-c\right)\) and expand; the cross term is \(2(\mu-c)\avg{X-\mu}=0\).
For Equation (11.7) take first \(c>m\) and put \(d=c-m>0\). Then
Indeed, for \(x\le m<c\) both sides equal \(d\), because \(\abs{x-c}-\abs{x-m}=(c-x)-(m-x)=d\); and for \(x>m\) the right-hand side is \(-d\), while the left-hand side is at least \(-\abs{c-m}=-d\) by the triangle inequality. Put \(x=X\) in Equation (11.9) and take expectations, using monotonicity and linearity (Proposition 11.5):
since \(\Pr(X>m)=1-\Pr(X\le m)\le\tfrac12\) by Equation (11.6) at \(\alpha=\tfrac12\). For \(c<m\) the same argument with \(d=m-c\) and the indicator of \(\set{x<m}\) applies, using \(\Pr(X<m)=1-\Pr(X\ge m)\le\tfrac12\).
∎So the mean is the least-squares measure of location and the median the least-absolute-deviation one. The difference between them is robustness: a single grossly wrong reading displaces the mean without bound, and displaces the median by at most the gap to the next value in the sample.
For \(X\) and \(Y\) with finite second moments,
which exists because \(\abs{uv}\le\tfrac12\left(u^{2}+v^{2}\right)\) makes the product integrable. When both variances are non-zero, the correlation coefficient is
A covariance carries the product of the SI units of \(X\) and \(Y\); a correlation coefficient is a pure number. Rests on Definition 11.6 and Proposition 11.5.
Covariance is symmetric and bilinear: \(\operatorname{cov}(X,X)=\operatorname{var}X\), \(\operatorname{cov}(X,c)=0\) for a constant \(c\), \(\operatorname{cov}(aX+bY,Z)=a\operatorname{cov}(X,Z) +b\operatorname{cov}(Y,Z)\), and \(\operatorname{cov}(X,Y)=\avg{XY}-\avg{X}\avg{Y}\). Consequently, for constants \(a_{1},\dots,a_{n}\),
Rests on Definition 11.11 and Proposition 11.5.
Derivation. Derives Proposition 11.12. The first two claims are Equation (11.4) and the vanishing of \(c-\avg{c}\). For bilinearity, the centred variable of \(aX+bY\) is \(a\left(X-\avg{X}\right)+b\left(Y-\avg{Y}\right)\) by linearity of the expectation, and multiplying by \(Z-\avg{Z}\) and taking expectations splits the result by linearity again; symmetry then gives linearity in the second slot. Expanding Equation (11.10) gives \(\avg{XY}-\avg{X}\avg{Y}-\avg{Y}\avg{X}+\avg{X}\avg{Y} =\avg{XY}-\avg{X}\avg{Y}\). Finally Equation (11.12) is bilinearity applied to \(\operatorname{cov}\!\left(\sum_{i}a_{i}X_{i}, \sum_{j}a_{j}X_{j}\right)\), the diagonal terms giving the variances and the off-diagonal ones pairing up by symmetry.
∎Discrete random variables \(X_{1},\dots,X_{n}\) are independent when
for every choice of values \(x_{1},\dots,x_{n}\). An infinite family is independent when every finite subfamily is. Rests on Definition 11.3.
Derivation. Derives Proposition 11.14. Grouping the outcomes by the value of the pair \((X,Y)\), exactly as in Proposition 11.5(i), gives \(\avg{XY}=\sum_{x,y}xy\Pr(X=x,Y=y)\), and the double series converges absolutely because \(\sum_{x,y}\abs{x}\abs{y}p_{X}(x)p_{Y}(y) =\avg{\abs{X}}\avg{\abs{Y}}<\infty\). Absolute convergence permits summing in either order and splitting the product of the two series (Proposition 7.49), so Equation (11.13) gives \(\avg{XY}=\left(\sum_{x}x\,p_{X}(x)\right) \left(\sum_{y}y\,p_{Y}(y)\right)=\avg{X}\avg{Y}\), and Proposition 11.12 makes the covariance vanish. For the converse, let \(X\) take the values \(-1,0,1\) with probability \(\tfrac13\) each and put \(Y=X^{2}\). Then \(\avg{X}=0\) and \(\avg{XY}=\avg{X^{3}}=0\), so the covariance vanishes; but \(\Pr(X=0,Y=0)=\tfrac13\) while \(p_{X}(0)\,p_{Y}(0)=\tfrac13\cdot\tfrac13\), so Equation (11.13) fails and \(Y\) is in fact a function of \(X\).
∎Zero covariance is a statement about second moments alone, and the example shows how little it excludes. This is exactly why Proposition 11.43(iii) needs full independence rather than mere uncorrelatedness: \(\ee^{\ii tX}\) is a nonlinear function of \(X\).
\(\abs{\rho(X,Y)}\le1\), with \(\rho=\pm1\) if and only if \(Y=aX+b\) with probability one for constants \(a\neq0\) and \(b\); the sign of \(\rho\) is then the sign of \(a\). Rests on Definition 11.11 and Lemma 11.63.
Derivation. Derives Corollary 11.15. Divide Equation (11.50) by the positive number \(\operatorname{var}X\cdot\operatorname{var}Y\) and take square roots. Equality holds exactly under the condition recorded in Lemma 11.63: there are constants \((a_{0},c_{0})\neq(0,0)\) with \(a_{0}\left(X-\avg{X}\right)+c_{0}\left(Y-\avg{Y}\right)=0\) with probability one. Neither coefficient can vanish here, because \(a_{0}=0\) would force \(\operatorname{var}Y=0\) and \(c_{0}=0\) would force \(\operatorname{var}X=0\). Hence \(Y-\avg{Y}=k\left(X-\avg{X}\right)\) with \(k=-a_{0}/c_{0}\neq0\), which is the stated affine relation with \(b=\avg{Y}-k\avg{X}\). For such a \(Y\), Proposition 11.12 gives \(\operatorname{cov}(X,Y)=k\operatorname{var}X\) and \(\operatorname{var}Y=k^{2}\operatorname{var}X\), so \(\rho=k/\abs{k}=\sgn k\).
∎Lemma 11.63 is stated and proved in Section 11.4.3, where it is needed a second time; its derivation uses nothing from this section, so borrowing it here is not circular.
Let \(X_{1},\dots,X_{n}\) be independent with a common distribution and \(n\ge2\). The sample mean, sample variance and sample standard deviation are
and for pairs \((X_{i},Y_{i})\) the sample covariance and sample correlation coefficient are
These are functions of the data alone: they are statistics, estimating the \(\mu\), \(\sigma^{2}\), \(\sigma\), \(\operatorname{cov}(X,Y)\) and \(\rho\) of the distribution that produced them. Rests on Definitions 11.6, 11.11 and 11.13.
With the notation of Equation (11.14) and a common mean \(\mu\) and variance \(\sigma^{2}<\infty\),
while \(\avg{s_{n}}\le\sigma\), with equality only if \(s_{n}\) is constant with probability one. Rests on Definition 11.16, Proposition 11.12 and Proposition 11.14.
Derivation. Derives Proposition 11.17. The first identity is linearity. The second is Equation (11.12) with \(a_{i}=1/n\): the covariances vanish by Proposition 11.14, leaving \(n\cdot n^{-2}\sigma^{2}\). For the third, write \(X_{i}-\bar{X}_{n}=\left(X_{i}-\mu\right)-\left(\bar{X}_{n}-\mu\right)\) and sum the squares; since \(\sum_{i}\left(X_{i}-\mu\right) =n\left(\bar{X}_{n}-\mu\right)\), the cross term is \(-2n\left(\bar{X}_{n}-\mu\right)^{2}\) and
Taking expectations gives \(n\sigma^{2}-n\cdot\sigma^{2}/n =(n-1)\sigma^{2}\), and dividing by \(n-1\) gives \(\avg{s_{n}^{2}}=\sigma^{2}\). Finally \(0\le\operatorname{var}s_{n}=\avg{s_{n}^{2}}-\avg{s_{n}}^{2} =\sigma^{2}-\avg{s_{n}}^{2}\) by Equation (11.5), so \(\avg{s_{n}}\le\sigma\), with equality exactly when \(\operatorname{var}s_{n}\) vanishes.
∎The divisor \(n-1\) is the entire content of the third identity: the deviations in Equation (11.14) are taken from \(\bar{X}_{n}\) rather than from the unknown \(\mu\), and \(\bar{X}_{n}\) is itself pulled towards the data, which is what Equation (11.17) measures. The last clause is a warning worth stating plainly: correcting the variance does not correct the standard deviation. Whatever the distribution, \(s_{n}\) underestimates \(\sigma\) on average, and no divisor can fix that for every distribution at once, since the deficit \(\sigma-\avg{s_{n}}\) depends on the sampling distribution of \(s_{n}\) itself.
Bernoulli: \(X\in\set{0,1}\) with \(\Pr(X=1)=p\). Then \(\avg{X}=p\) and \(\avg{X^{2}}=p\), so \(\operatorname{var}X=p(1-p)\) by Equation (11.5), largest at \(p=\tfrac12\).
Binomial: \(S=\sum_{i=1}^{n}X_{i}\) with the \(X_{i}\) independent Bernoulli variables, so that \(\Pr(S=k)=\binom{n}{k}p^{k}(1-p)^{n-k}\). Linearity gives \(\avg{S}=np\) and Equation (11.12) gives \(\operatorname{var}S=np(1-p)\), neither calculation touching the mass function.
Poisson: \(p(k)=\nu^{k}\ee^{-\nu}/k!\) for \(k=0,1,2,\dots\) and \(\nu>0\), which sums to \(\ee^{-\nu}\ee^{\nu}=1\). Cancelling one factor from \(k/k!=1/(k-1)!\) gives \(\avg{X}=\nu\sum_{j\ge0}\nu^{j}\ee^{-\nu}/j!=\nu\), and the same manoeuvre twice gives \(\avg{X(X-1)}=\nu^{2}\), so \(\avg{X^{2}}=\nu^{2}+\nu\) and \(\operatorname{var}X=\nu\). Mean and variance coincide: that is the signature of a counting process, and the reason a count of \(N\) events is quoted with an uncertainty \(\sqrt{N}\). Rests on Definition 11.6, Proposition 11.12 and Proposition 11.14.
A detector is read out five times over equal live times \(t_{\text{live}}=60\,\mathrm{s}\) and records \(41\), \(52\), \(47\), \(44\) and \(46\) counts. Counts are pure numbers; only the derived rate carries a unit. The sample mean is \(\bar{n}_{5}=230/5=46.0\), the deviations are \(-5,6,1,-2,0\), their squares sum to \(66\), and Equation (11.14) gives \(s_{5}^{2}=66/4=16.5\) and \(s_{5}=4.06\). The standard deviation of the mean is \(s_{5}/\sqrt{5}=1.82\) by Equation (11.16). Sorting the readings as \(41,44,46,47,52\) shows the median to be \(46\), close to the mean here because nothing is grossly discrepant.
Dividing by the live time turns the estimate into a rate,
the division by a constant scaling the standard deviation by that same constant, as Proposition 11.7 requires. Had the five readings come from a Poisson law, Example 11.18 would predict \(\sigma=\sqrt{\nu}\simeq6.8\) for a single reading against the \(s_{5}=4.06\) observed — a difference that five readings cannot resolve, and exactly the kind of question Section 11.5 is built to answer. Rests on Definition 11.16, Proposition 11.17 and Example 11.18.
Continuous statistics
A continuously distributed variable takes each individual value with probability zero, so the mass function of Equation (11.2) is useless and a density takes its place. The sums of Section 11.1 become integrals, and because every identity proved there used only linearity, monotonicity and positivity of the expectation, the algebra survives the change intact. What has to be redone is only what depended on rearranging a series.
Definitions
A real random variable \(X\) is continuously distributed with probability density \(f\) when \(f\ge0\) is Riemann integrable on every bounded interval, the improper integral
converges, and \(\Pr(a\le X\le b)=\int_{a}^{b}f(x)\,\dd x\) for all \(a<b\). If \(X\) carries the SI unit \(u\), then \(f\) carries \(u^{-1}\): only the product \(f(x)\,\dd x\) is a probability, and only a probability is a pure number. Rests on Definition 7.39.
The distribution function of a real random variable \(X\) is \(F_{X}(x)=\Pr(X\le x)\), whether or not \(X\) has a density. When it has one, \(F_{X}(x)=\int_{-\infty}^{x}f(t)\,\dd t\). In general \(F_{X}\) is only non-decreasing, and it may jump: for a variable taking the single value \(c\) with probability one it steps from \(0\) to \(1\) at \(c\). The proposition below is the density case, and continuity is what that case adds. Rests on Definition 11.20.
If \(X\) has a density \(f\), then \(F_{X}\) is non-decreasing and continuous, tends to \(0\) at \(-\infty\) and to \(1\) at \(+\infty\), satisfies \(\Pr(a<X\le b)=F_{X}(b)-F_{X}(a)\) and \(\Pr(X=x)=0\) for every \(x\), and has \(F_{X}'(x)=f(x)\) at every point where \(f\) is continuous. Rests on Definition 11.21 and Theorem 7.42.
Derivation. Derives Proposition 11.22. Monotonicity and the difference formula follow from additivity of the integral over subintervals together with \(f\ge0\); the two limits are Equation (11.19). A Riemann integrable function is bounded on each bounded interval, say by \(M\) near \(x\), so \(\abs{F_{X}(x+h)-F_{X}(x)}\le M\abs{h}\) and \(F_{X}\) is continuous — indeed locally Lipschitz. Hence \(\Pr(X=x)\le F_{X}(x+\varepsilon)-F_{X}(x-\varepsilon)\rightarrow0\). The last claim is the first fundamental theorem of calculus (Theorem 7.42), which needs \(f\) continuous only at the point in question.
∎For \(g:\R\rightarrow\R\) with \(\int\abs{g(x)}f(x)\,\dd x\) convergent,
and the mean, moments, variance, standard deviation, skewness, excess kurtosis, covariance and correlation coefficient are defined by the very formulas of Definitions 11.6 and 11.11, read with this expectation. Quantiles are defined by Equation (11.6); by Proposition 11.22 the two requirements there collapse to \(F_{X}(q)=\alpha\). Rests on Definitions 11.4 and 11.20.
The expectation Equation (11.20) is linear and monotone and satisfies \(\abs{\avg{g(X)}}\le\avg{\abs{g(X)}}\). Consequently Proposition 11.7, Proposition 11.10, Proposition 11.12, Corollary 11.15 and Proposition 11.17 hold verbatim for continuously distributed variables. Rests on Definition 11.23 and Proposition 11.5.
Derivation. Derives Proposition 11.24. Linearity and monotonicity of the integral are the elementary properties of the Darboux construction recorded with Definition 7.39, and applying monotonicity to \(\pm g f\le\abs{g}f\) gives the third property. Reading the proofs of the five results named, each uses the expectation only through those three properties together with pointwise algebraic identities and the pointwise inequality Equation (11.9); none of them rearranges a series. The same holds of Lemma 11.63, whose proof needs only that \(\avg{\left(s\tilde{V}+\tilde{W}\right)^{2}}\ge0\), and which is in any case stated for arbitrary random variables of finite variance. One dependence is not of that kind and has to be replaced rather than reread: Proposition 11.17 also uses the vanishing of the covariance of two independent variables, which here is Proposition 11.26 below and not Proposition 11.14.
∎What does not transfer for free is the product rule for independent variables, since its discrete proof rearranged a double series. It is restated and reproved here.
A pair \((X,Y)\) has joint density \(f_{XY}\ge0\) when \(\Pr\left(a\le X\le b,\;c\le Y\le d\right) =\int_{a}^{b}\!\int_{c}^{d}f_{XY}(x,y)\,\dd y\,\dd x\) for all \(a<b\) and \(c<d\). The two variables are independent when \(f_{XY}(x,y)=f_{X}(x)f_{Y}(y)\) for all \(x\) and \(y\). Rests on Definitions 11.13 and 11.20.
If \(X\) and \(Y\) are independent in the sense of Definition 11.25 and both are integrable, then \(\avg{XY}=\avg{X}\avg{Y}\) and \(\operatorname{cov}(X,Y)=0\). Rests on Definition 11.25 and Proposition 11.24.
Derivation. Derives Proposition 11.26. In the iterated integral \(\avg{XY}=\int\!\left(\int xy\,f_{X}(x)f_{Y}(y)\,\dd y\right)\dd x\) the inner integrand is \(x f_{X}(x)\) times a function of \(y\) alone, so the inner integral equals \(x f_{X}(x)\int y f_{Y}(y)\,\dd y\) — a constant multiple of \(x f_{X}(x)\), the constant being \(\avg{Y}\), which is finite by hypothesis. Integrating in \(x\) gives \(\avg{X}\avg{Y}\). No interchange of integrations is used: the factorisation is of the integrand, not of the order of integration. Proposition 11.12, which transfers by Proposition 11.24, then makes the covariance vanish.
∎Let \(F=F_{X}\) be strictly increasing on an interval carrying all the probability. Then \(U=F(X)\) is uniformly distributed on \([0,1]\), that is \(\Pr(U\le\alpha)=\alpha\) for \(\alpha\in(0,1)\); conversely, if \(U\) is uniformly distributed on \([0,1]\) then \(F^{-1}(U)\) has distribution function \(F\). Rests on Proposition 11.22 and Definition 11.23.
Derivation. Derives Proposition 11.27. \(F\) is continuous by Proposition 11.22 and strictly increasing by hypothesis, so it has a continuous strictly increasing inverse on \((0,1)\) and \(\set{F(X)\le\alpha}=\set{X\le F^{-1}(\alpha)}\). Hence \(\Pr(U\le\alpha)=F\!\left(F^{-1}(\alpha)\right)=\alpha\). Conversely \(\Pr\!\left(F^{-1}(U)\le x\right)=\Pr\!\left(U\le F(x)\right)=F(x)\).
∎The transform is what makes quantiles computable — the \(\alpha\)-quantile is \(F^{-1}(\alpha)\) — and it is the same statement, read through the survival function instead of the distribution function, as the uniformity of a \(p\)-value in Proposition 11.76.
The Gaussian (or normal) distribution \(\mathcal{N}(\mu,\sigma^{2})\) with \(\sigma>0\) has density
The case \(\mu=0\), \(\sigma=1\) is the standard Gaussian, with density \(\varphi\) and distribution function \(\Phi\). Here \(\mu\) and \(\sigma\) carry the SI unit of \(X\), \(f\) carries its reciprocal, and the exponent is a pure number. Rests on Definition 11.20.
Equation (11.21) is a density; its mean is \(\mu\) and its variance \(\sigma^{2}\); its central moments are \(\mu_{2n+1}=0\) and \(\mu_{2n}=(2n-1)!!\,\sigma^{2n}\); and its skewness and excess kurtosis both vanish. Rests on Definitions 11.23 and 11.28.
Derivation. Derives Proposition 11.29. The substitution \(u=(x-\mu)/\sigma\) (Equation (7.27)) reduces every statement to the standard case, with \(\mu_{n}=\sigma^{n}\avg{U^{n}}\) for \(U\) standard. For the normalisation put \(I=\int_{-\infty}^{\infty}\ee^{-u^{2}/2}\,\dd u\) and evaluate \(I^{2}\) as a double integral in polar coordinates, whose Jacobian is \(r\):
so \(I=\sqrt{2\pi}\) and \(\varphi\) integrates to \(1\). This is the one step of the derivation that uses the change-of-variables formula for a double integral.
The odd moments vanish because \(u^{2n+1}\varphi(u)\) is odd and absolutely integrable. For the even ones use \(\varphi'(u)=-u\varphi(u)\) and integrate by parts (Equation (7.28)):
the boundary term vanishing because \(u^{2n-1}\varphi(u)\rightarrow0\) at both ends. With \(\avg{U^{0}}=1\) this gives \(\avg{U^{2n}}=(2n-1)!!\), hence \(\avg{U^{2}}=1\) and \(\avg{U^{4}}=3\). So the mean is \(\mu\), the variance is \(\sigma^{2}\), \(\gamma_{1}=0\) and \(\gamma_{2}=3-3=0\), which is the convention fixed in Definition 11.6.
∎A decay with constant \(\lambda\), carried in \(/\mathrm{s}\), has density \(f(t)=\lambda\ee^{-\lambda t}\) for \(t\ge0\) and \(f(t)=0\) for \(t<0\). It integrates to \(1\), and integrating by parts \(n\) times (Equation (7.28)) gives \(\avg{T^{n}}=n!/\lambda^{n}\). Hence the mean lifetime and the standard deviation are both \(\tau=1/\lambda\), in \(\mathrm{s}\), and
The distribution function is \(F(t)=1-\ee^{-\lambda t}\), so by Definition 11.23 the median is the half-life \(t_{1/2}=\tau\ln 2=0.693\,\tau\), strictly less than the mean — the quantitative form of the positive skewness, and a reminder that for an asymmetric law the two measures of location of Proposition 11.10 answer different questions. Rests on Definition 11.23, Proposition 11.10 and Proposition 11.24.
Six readings of a gauge block on a comparator give, in \(\mathrm{m}\), \(0.0250013\), \(0.0250009\), \(0.0250015\), \(0.0250011\), \(0.0250010\) and \(0.0250014\). Working in units of \(10^{-7}\,\mathrm{m}\) above \(0.025\,\mathrm{m}\), the readings are \(13\), \(9\), \(15\), \(11\), \(10\), \(14\), with sample mean \(12\) and deviations \(1,-3,3,-1,-2,2\) whose squares sum to \(28\). By Equation (11.14),
and the standard deviation of the mean is \(s_{6}/\sqrt{6}=9.7\times 10^{-8}\,\mathrm{m}\). The units track the definitions exactly: a variance in \(\mathrm{m}^{2}\), a standard deviation in \(\mathrm{m}\), and a skewness or a correlation coefficient that would be a pure number. The ordered readings are \(9,10,11,13,14,15\), so any point of \([11,13]\) is a median by Equation (11.6); the customary choice is the midpoint, here \(12\), which coincides with the mean because this sample happens to be symmetric. Rests on Definition 11.16, Proposition 11.17 and Definition 11.8.
The limit theorems that join the two theories are the subject of the next section. The binomial variable of Example 11.18 is a sum of independent identically distributed Bernoulli variables, so Theorem 11.46 applies to it verbatim: a discrete law, suitably centred and scaled, is described in the limit by the continuous Gaussian density Equation (11.21). That is the passage from the first of these two sections to the second, and it is proved rather than asserted in Section 11.3.2.
A distribution need be neither discrete nor continuous. The asymptotic law of the discovery statistic in Equation (11.86) is a half-and-half mixture of a point mass at zero and a chi-squared density, and is of neither kind: its distribution function jumps at the origin and grows continuously above it. The theory that covers every case at once replaces both the sum Equation (11.3) and the integral Equation (11.20) by an integral against a measure [Billingsley:1995]. Every identity proved in these two sections survives that generalisation unchanged, because each rests only on the linearity, monotonicity and positivity of the expectation — which is also why they could be proved once and transferred once, rather than twice.
Measure spaces, null sets and ``almost every''
Both sections above are built on sums and on Riemann integrals, and that is enough for every distribution they treat. It is not enough for one thing the rest of the treatise says repeatedly: that a statement holds for almost every point, or that some exceptional set has zero volume. That phrase has no meaning until sets are assigned sizes, and the assignment has to survive countable operations or the phrase is useless — the whole force of “almost every” is that countably many exceptions may be discarded at once. The definitions below supply exactly that much and nothing more. They are used in Symplectic Geometry of Phase Space for the recurrence theorem, in Nonlinear Dynamics and Chaos for the sets on which a Lyapunov exponent is defined, and in Hilbert Spaces wherever \(L^{2}\) appears.
A \(\sigma\)-algebra on a set \(\Omega\) is a family \(\mathcal{F}\) of subsets of \(\Omega\) such that
-
\(\Omega\in\mathcal{F}\);
-
\(A\in\mathcal{F}\) implies \(\Omega\setminus A\in\mathcal{F}\);
-
\(A_{1},A_{2},\dots\in\mathcal{F}\) implies \(\bigcup_{n\ge1}A_{n}\in\mathcal{F}\).
Its members are the measurable sets, and \((\Omega,\mathcal{F})\) is a measurable space. Complements and countable unions give countable intersections as well, by de Morgan's laws (Proposition 3.37), so \(\mathcal{F}\) is closed under every countable Boolean operation. The Borel \(\sigma\)-algebra of a topological space is the smallest \(\sigma\)-algebra containing the open sets — smallest because an arbitrary intersection of \(\sigma\)-algebras is one. Rests on Definitions 6.1 and 11.1.
A measure on \((\Omega,\mathcal{F})\) is a map \(\mu:\mathcal{F}\longrightarrow[0,\infty]\) with \(\mu(\varnothing)=0\) which is countably additive: for pairwise disjoint \(A_{1},A_{2},\dots\in\mathcal{F}\),
The triple \((\Omega,\mathcal{F},\mu)\) is a measure space, and a probability space when \(\mu(\Omega)=1\), in which case \(\mu\) is written \(\Pr\). A measure carries the SI unit of the quantity it measures — \(\mathrm{m}^{3}\) for a volume in space, the product of the units of the phase-space coordinates for the Liouville measure of Symplectic Geometry of Phase Space — and a probability is a pure number. Rests on Definitions 11.1 and 11.33.
A set \(N\in\mathcal{F}\) is null (or of measure zero) if \(\mu(N)=0\). A statement holds almost everywhere, or for almost every point, if the set where it fails is contained in a null set. In a probability space one says almost surely. Rests on Definition 11.34.
Let \((\Omega,\mathcal{F},\mu)\) be a measure space. Then
-
\(A\subseteq B\) implies \(\mu(A)\le\mu(B)\);
-
\(\mu\left(\bigcup_{n}A_{n}\right)\le\sum_{n}\mu(A_{n})\) for any countable family, disjoint or not;
-
a countable union of null sets is null;
-
consequently, countably many statements each holding almost everywhere hold simultaneously almost everywhere.
Derivation. Derives Proposition 11.36. (i) Write \(B=A\cup(B\setminus A)\), a disjoint union of measurable sets; Equation (11.26) with all later terms empty gives \(\mu(B)=\mu(A)+\mu(B\setminus A)\ge\mu(A)\), the last step because a measure is non-negative.
(ii) Disjointify: put \(B_{1}=A_{1}\) and \(B_{n}=A_{n}\setminus(A_{1}\cup\dots\cup A_{n-1})\), which lie in \(\mathcal{F}\) by Definition 11.33, are pairwise disjoint, have the same union as the \(A_{n}\), and satisfy \(B_{n}\subseteq A_{n}\). Then \(\mu\left(\bigcup_{n}A_{n}\right)=\sum_{n}\mu(B_{n}) \le\sum_{n}\mu(A_{n})\) by Equation (11.26) and (i).
(iii) If every \(\mu(A_{n})=0\) then the right-hand side of (ii) is a series of zeros, so the union has measure at most zero, hence exactly zero.
(iv) Apply (iii) to the union of the countably many exceptional sets: outside it, every one of the statements holds.
∎The definitions above are the whole of what this treatise needs from measure theory, and everything asserted about them is proved in Proposition 11.36. What is not proved anywhere here is that interesting measures exist: that there is a measure on the Borel sets of \(\R^{n}\) assigning to a box the product of its edge lengths — Lebesgue measure — is a construction (Carathéodory's extension of a premeasure from an algebra to the \(\sigma\)-algebra it generates), and it is quoted, from [Billingsley:1995], wherever it is used; so is the Lebesgue integral built on it, which enters this book only through the completeness of \(L^{2}\) in Theorem 12.12 and through the convergence theorems listed in Remark 12.1. The distinction is worth keeping sharp. A statement of the form “the exceptional set has measure zero” needs only Definition 11.35 and Proposition 11.36 to be meaningful and to be combined with others; it needs the construction only when a particular set is to be shown null by computing something. The recurrence theorem of Symplectic Geometry of Phase Space is of the first kind, which is why it can be stated honestly here.
Invariant measures and the ergodic theorem
A dynamical system supplies a measure space with a map of it into itself, and the question that then arises — whether a long time average along one trajectory equals an average over the space — is the question every use of statistical mechanics and of Lyapunov exponents in this book turns on. It is a theorem of probability, not of mechanics, and it belongs here.
Let \((\Omega,\mathcal{F},\mu)\) be a measure space with \(\mu(\Omega)<\infty\). A measurable map \(T:\Omega\longrightarrow\Omega\) preserves \(\mu\), equivalently \(\mu\) is invariant under \(T\), if
A set \(A\) is invariant if \(T^{-1}(A)=A\), and the system \((T,\mu)\) is ergodic if every invariant set is null or has null complement — that is, if the space cannot be split into two parts of positive measure that the motion never mixes. The same definitions apply to a one-parameter family \(\set{T_{t}}_{t\in\R}\), a measure-preserving flow, with Equation (11.27) required of every \(T_{t}\). Rests on Definitions 11.34 and 11.35.
Let \(T\) preserve a finite measure \(\mu\) on \((\Omega,\mathcal{F})\) and let \(f\) be \(\mu\)-integrable. Then the time average
exists for \(\mu\)-almost every \(\omega\), is \(T\)-invariant, and satisfies \(\int\bar{f}\,\dd\mu=\int f\,\dd\mu\). If \((T,\mu)\) is ergodic then \(\bar{f}\) is almost everywhere equal to the constant \(\mu(\Omega)^{-1}\int f\,\dd\mu\): the time average along almost every trajectory equals the space average. The same statements hold for a measure-preserving flow, with the time average \(T^{-1}\int_{0}^{T}f(T_{t}\omega)\,\dd t\). Rests on Definition 11.38 and Proposition 11.36.
Theorem 11.39 is quoted, not proved. It is Birkhoff's theorem [Birkhoff:1931], and its proof rests on a maximal inequality for the partial sums together with the dominated convergence theorem of Lebesgue integration — exactly the theory Remark 11.37 declares this treatise does not develop. Von Neumann's mean ergodic theorem [vonNeumann:1932a], proved a little earlier, gives convergence in \(L^{2}\) rather than pointwise and is a statement about the unitary operator \(f\mapsto f\circ T\) on that space; it is the weaker conclusion and the easier proof, and it is not enough for a statement about individual trajectories. Nothing in this chapter is derived from either. The theorem is stated here so that Nonlinear Dynamics and Chaos, which reports measured Lyapunov exponents, and Statistical Mechanics, which replaces a time average by an ensemble average, can name the hypothesis they need — ergodicity of a specified invariant measure — instead of leaving it implicit. The multiplicative extension that governs the whole Lyapunov spectrum, rather than a single average, is Oseledets' [Oseledets:1968]; it is stated and used, with the same declaration, where the exponents are defined.
Theorem 11.39 says what follows if the system is ergodic; it says nothing about whether any given system is. That is a question about the system, and for mechanical systems it is usually open and sometimes answered in the negative: an integrable system is never ergodic on an energy surface, because each invariant torus of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy is an invariant set of measure zero whose union of positive-measure families fills the surface, and the surviving tori of the KAM theorem in Nonlinear Dynamics and Chaos do the same for a nearly integrable one. The honest form of the statistical-mechanical postulate is therefore an assumption about the measure, not a theorem about the dynamics.
Characteristic functions and the central limit theorem
The theory of measurement error in Measurement, SI Units, and the Theory of Errors rests on one theorem: that a quantity assembled from many small independent contributions is Gaussian, whatever the contributions individually look like. That is what licenses quoting a standard uncertainty at all, and it is proved here.
The proof goes through the Fourier transform of a distribution, which turns the convolution that describes a sum of independent variables into a product. Throughout, \(\avg{\,\cdot\,}\) denotes expectation, as in Measurement, SI Units, and the Theory of Errors.
Characteristic functions
The characteristic function of a real random variable \(X\) is
Unlike the moment generating function \(\avg{\ee^{tX}}\), this exists for every distribution without exception: \(\abs{\ee^{\ii tX}} = 1\), so the expectation is an integral of a bounded function and always converges. That unconditional existence is why the argument below needs no assumption beyond a finite variance.
For any random variables \(X\), \(Y\) and constants \(a,b\in\R\):
-
\(\varphi_{X}(0)=1\) and \(\abs{\varphi_{X}(t)}\le 1\) for all \(t\);
-
\(\varphi_{aX+b}(t) = \ee^{\ii b t}\,\varphi_{X}(a t)\);
-
if \(X\) and \(Y\) are independent, then \(\varphi_{X+Y} = \varphi_{X}\,\varphi_{Y}\).
Rests on Definition 11.42.
Derivation. Derives Proposition 11.43. (i) \(\varphi_X(0)=\avg{1}=1\), and \(\abs{\varphi_X(t)} = \abs{\avg{\ee^{\ii tX}}} \le \avg{\abs{\ee^{\ii tX}}} = 1\). (ii) \(\avg{\ee^{\ii t(aX+b)}} = \ee^{\ii bt}\avg{\ee^{\ii (ta)X}}\). (iii) \(\ee^{\ii t(X+Y)} = \ee^{\ii tX}\ee^{\ii tY}\), and if \(X\) and \(Y\) are independent then \(\avg{f(X)g(Y)} = \avg{f(X)}\avg{g(Y)}\) for all bounded measurable \(f,g\) — which is what independence means, and is strictly stronger than being uncorrelated: the latter is the case \(f=g=\id\) alone and would not suffice here, since \(\ee^{\ii tX}\) is a nonlinear function of \(X\). Both factors are bounded, so both expectations exist.
∎Property (iii) is the whole point: addition of independent variables, which is a convolution of distributions, becomes multiplication of characteristic functions — and no density need exist for that to hold.
If \(\avg{X}=0\) and \(\avg{X^{2}}=\sigma^{2}<\infty\), then
Rests on Definition 11.42.
Derivation. Derives Lemma 11.44. Writing the remainder of \(\ee^{\ii x}\) in integral form by integrating by parts \(n\) times gives, for real \(x\),
Bounding the integrand of that remainder gives the first estimate; the second follows instead from the triangle inequality applied to the order-\((n-1)\) remainder together with the last term of the series. The minimum matters because only the second branch survives when no moment beyond the \(n\)th exists. Taking \(n=2\) and \(x=tX\),
Take expectations, and use \(\abs{\avg{Z}}\le\avg{\abs{Z}}\). Since \(\avg{X}=0\) the left side becomes \(\abs{\varphi_{X}(t) - 1 + \tfrac12 t^{2}\sigma^{2}}\), so
The right-hand side tends to \(0\) as \(t\rightarrow0\). The usual reason given is dominated convergence, which this treatise cannot yet invoke: Real Analysis develops the Riemann integral only, and the Lebesgue theory the dominated-convergence theorem belongs to is not built anywhere in the book. An elementary truncation does the same work. Fix \(K>0\) and split the expectation at \(\abs{X}=K\): where \(\abs{X}\le K\) the first branch of the minimum bounds the integrand by \(\tfrac{1}{6}\abs{t}K^{3}\), and where \(\abs{X}>K\) the second branch bounds it by \(X^{2}\indic{\abs{X}>K}\), so
From Equation (11.31) (take \(n=2\) and \(x=tX\), then split the expectation at \(\abs{X}=K\) and use one branch of the minimum on each piece). Given \(\varepsilon>0\), choose \(K\) so large that the second term is below \(\varepsilon/2\) — possible precisely because \(\avg{X^{2}}\) is finite, a finite integral being one whose tail contribution vanishes [Billingsley:1995] — and then take \(\abs{t}<3\varepsilon/K^{3}\), which puts the first term below \(\varepsilon/2\) as well. That is exactly Equation (11.30).
∎The lemma is where the finite variance enters, and it is the only place it is needed. Note what it does not require: no third moment, no density, no symmetry.
Let \(X_{1},X_{2},\dots\) be random variables with characteristic functions \(\varphi_{n}\). If \(\varphi_{n}(t)\rightarrow\varphi(t)\) for every \(t\in\R\) and \(\varphi\) is continuous at \(t=0\), then \(\varphi\) is the characteristic function of a random variable \(X\) and \(X_{n}\rightarrow X\) in distribution. Rests on Definition 11.42.
This is the bridge from pointwise convergence of transforms back to convergence of the distributions themselves, and it is the deepest analytic statement this section rests on: the proof needs Helly's selection theorem, the Fourier inversion formula and a tightness estimate that is where continuity of the limit at the origin does its work [Billingsley:1995] [Feller:1971]. It is carried out in full in the appendix, so Theorem 11.45 is derived rather than quoted; what the derivation does assume, and all it assumes, are two theorems of Lebesgue theory, named there.
Full derivation in Appendix A.
Derives Theorem 11.45.
The central limit theorem
Let \(X_{1},X_{2},\dots\) be independent, identically distributed random variables with finite mean \(\mu\) and finite variance \(\sigma^{2}\in(0,\infty)\). Then
that is, \(\Pr(Z_{n}\le z)\rightarrow \frac{1}{\sqrt{2\pi}}\int_{-\infty}^{z}\ee^{-u^{2}/2}\,\dd u\) for every \(z\in\R\). Rests on Proposition 11.43, Lemma 11.44 and Theorem 11.45.
Derivation. Derives Theorem 11.46. Put \(Y_{i}=(X_{i}-\mu)/\sigma\), so the \(Y_{i}\) are independent and identically distributed with \(\avg{Y_{i}}=0\) and \(\avg{Y_{i}^{2}}=1\), and \(Z_{n}=n^{-1/2}\sum_{i}Y_{i}\). Write \(\varphi\) for the common characteristic function of the \(Y_{i}\).
By Proposition 11.43(ii) and (iii), the latter extended to \(n\) factors by induction,
Fix \(t\). By Lemma 11.44 applied at argument \(t/\sqrt{n}\),
From Equation (11.30) (the second-order expansion evaluated at argument \(t/\sqrt{n}\)). The remainder \(r_{n}(t)\) is defined by that equation; the convergence is the lemma rescaled. Writing \(\varphi(s) = 1 - s^{2}/2 + R(s)\) with \(R(s)/s^{2}\rightarrow0\), we have \(r_{n}(t) = R(t/\sqrt{n})\) and hence \(n\,r_{n}(t) = t^{2}\,R(s)/s^{2}\) at \(s=t/\sqrt{n}\), which tends to \(0\) for each fixed \(t\) because \(s\rightarrow0\). Now compare the two products using the elementary inequality
which follows by telescoping \(\prod a_{i} - \prod b_{i} = \sum_{i} \left(\prod_{j<i}a_{j}\right) (a_{i}-b_{i})\left(\prod_{j>i}b_{j}\right)\) and bounding each partial product by \(1\). Take \(a_{i}=\varphi(t/\sqrt{n})\), whose modulus is at most \(1\) by Proposition 11.43(i), and \(b_{i}=1-t^{2}/(2n)\), which lies in \([0,1]\) once \(n\ge t^{2}/2\). Then
Since \((1-t^{2}/(2n))^{n}\rightarrow\ee^{-t^{2}/2}\), it follows that
From Equations (11.34), (11.35) and (11.36) (compare the two products factor by factor and let \(n\) grow). The limit is the characteristic function of the standard Gaussian. That is quickest to see without contour shifting: writing \(g(t)=\int\ee^{\ii tu}\ee^{-u^{2}/2}\dd u/\sqrt{2\pi}\) and differentiating under the integral sign, an integration by parts gives \(g'(t) = -t\,g(t)\) with \(g(0)=1\), whence \(g(t)=\ee^{-t^{2}/2}\). It is continuous at \(t=0\). Theorem 11.45 therefore gives \(Z_{n}\rightarrow\mathcal{N}(0,1)\) in distribution.
∎Three qualifications matter for the use made of this result in Measurement, SI Units, and the Theory of Errors.
First, the convergence is in distribution and says nothing about any individual sample. Second, it is a statement about the centre: the approximation is worst in the tails, which is exactly where a coverage factor \(k=3\) is being asked to do work. Third — and this is the hypothesis that actually fails in practice — a finite variance is required. Errors drawn from a heavy-tailed law need not converge to anything Gaussian: the Cauchy distribution has no finite variance — indeed no finite mean — and the average of \(n\) Cauchy samples has the same distribution as a single one, so averaging never narrows it at all and Theorem 2.7's \(\sigma/\sqrt{N}\) has nothing to describe.
For each \(n\) let \(X_{n,1},\dots,X_{n,k_{n}}\) be independent with \(\avg{X_{n,k}}=0\) and \(s_{n}^{2}=\sum_{k}\operatorname{var}X_{n,k}\). If for every \(\varepsilon>0\)
then \(s_{n}^{-1}\sum_{k}X_{n,k}\rightarrow\mathcal{N}(0,1)\) in distribution. Rests on Theorem 11.45 and Lemma 11.44.
Theorem 11.46 is the special case of identically distributed summands. The general form is the one the error theory actually needs, because the disturbances that perturb a measurement are not identically distributed: they are a thermal drift, a vibration, a quantisation step, a reading error.
Condition Equation (11.38) is sufficient, and the implication it does carry is the one wanted here: it forces every individual variance share to vanish. Indeed, splitting the expectation at \(\abs{X_{n,k}}=\varepsilon s_{n}\) gives, for any \(\varepsilon>0\),
so Equation (11.38) implies \(\max_{k}\operatorname{var}X_{n,k}/s_{n}^{2}\rightarrow0\), which is uniform asymptotic negligibility — no single source carries a fixed fraction of the total variance. A dominant source therefore violates the condition.
The converse fails, and it is worth being precise about how, because the temptation is to read Equation (11.38) as a characterisation of “many small contributions”. It is strictly stronger than negligibility: it constrains tail mass, not merely variance shares. Summands equal to \(\pm\sqrt{n}\) with probability \(1/(2n)\) each and \(0\) otherwise have variance share \(1/n\rightarrow0\), yet the Lindeberg sum equals \(1\) for every \(n\) and the limit is compound Poisson, not Gaussian. Nor does failure of the condition force a non-Gaussian limit: adjoining one \(\mathcal{N}(0,n)\) summand to \(n\) independent \(\mathcal{N}(0,1)\) summands breaks Equation (11.38) while leaving the normalised sum exactly \(\mathcal{N}(0,1)\). Feller's companion theorem closes the loop only within the negligible class: if \(\max_{k}\operatorname{var}X_{n,k}/s_{n}^{2} \rightarrow0\) and the normalised sum is asymptotically Gaussian, then Equation (11.38) holds [Feller:1971].
Lindeberg introduced the condition in 1922 [Lindeberg:1922]; the method of that paper — replacing the summands by Gaussians one at a time and controlling the accumulated error — also yields an elementary proof of Theorem 11.46 that avoids characteristic functions altogether. A sufficient condition that is easier to check in practice is Lyapunov's: \(s_{n}^{-(2+\delta)}\sum_{k}\avg{\abs{X_{n,k}}^{2+\delta}}\rightarrow0\) for some \(\delta>0\) [Billingsley:1995] [Feller:1971]; it implies Equation (11.38) in a line, and that line is in the appendix. The characteristic-function argument that establishes Theorem 11.46 generalises to Theorem 11.48, with Lemma 11.44 replaced by an expansion uniform in the summand index and controlled by Equation (11.38); that generalisation is carried out there in full, and Equation (11.39) above is one of its lemmas.
Full derivation in Appendix A.
Derives Theorem 11.48.
The theorem is a limit, and an uncertainty budget needs a rate. The Berry–Esseen inequality supplies one: if \(\rho=\avg{\abs{X-\mu}^{3}}<\infty\), then
with \(\Phi\) the standard normal distribution function and \(C\) an absolute constant, shown finite by Berry [Berry:1941] and, independently and in the same period, by Esseen. It is proved in the appendix along Esseen's route — a smoothing inequality, a bound on the characteristic function of \(Z_{n}\) over \(\abs{t}\lesssim\sqrt{n}\,\sigma^{3}/\rho\), and an integration — which delivers the explicit admissible value \(C=7.23\) [Feller:1971]. Determining the smallest admissible constant is a separate literature and is not attempted there; the sharpest published value for identically distributed summands is \(C\le0.4690\), which is due to Shevtsova alone [Shevtsova:2014].
Two things follow, and neither is comfortable. The bound is uniform in \(z\): it limits the absolute error of the Gaussian approximation equally at the centre and in the tails, so where the probability being approximated is itself small the relative error is large — which is exactly the regime a coverage factor \(k=3\) operates in. And the \(n^{-1/2}\) is slow. Since \(\rho/\sigma^{3}\ge1\) always, the right-hand side of Equation (11.40) is at least \(0.4690/\sqrt{n}\), which for a dozen contributions guarantees only about \(0.14\) in Kolmogorov distance; several hundred are needed before the inequality itself certifies a few per cent. Actual convergence is usually far better than the worst case the bound protects against — but “usually” is not a quantity an uncertainty budget can carry, and the bound is what can be asserted.
Full derivation in Appendix A.
Derives Equation (11.40).
Estimation
The central limit theorem says what a sum of many small contributions looks like. Inference runs the other way: given the numbers a measurement actually produced, what can be said about the parameter that produced them, and how well? This section builds the likelihood function, the estimator that maximises it, the exact lower bound on how well any estimator can do, and the asymptotic law the maximum-likelihood estimator obeys. The last of these is what the next two sections need, and the profile likelihood constructed at the end of this one is the object a collider search actually evaluates.
Two elementary tools are used repeatedly and the earlier sections of this chapter do not yet supply them, so they are derived here.
Let \(Y\ge0\) be a random variable and \(a>0\). Then \(\Pr(Y\ge a)\le\avg{Y}/a\). Consequently, if \(X\) has finite mean \(\mu\) and finite variance \(\sigma^{2}\), then for every \(a>0\)
Rests on Proposition 11.5 and Definition 11.6.
Derivation. Derives Lemma 11.50. Since \(Y\ge0\), \(Y\ge Y\,\indic{Y\ge a}\ge a\,\indic{Y\ge a}\) pointwise; taking expectations gives \(\avg{Y}\ge a\Pr(Y\ge a)\). Apply this to \(Y=(X-\mu)^{2}\) with the level \(a^{2}\): the event \(\abs{X-\mu}\ge a\) is the event \(Y\ge a^{2}\), and \(\avg{Y}=\sigma^{2}\).
∎If \(X_{1},X_{2},\dots\) are independent and identically distributed with mean \(\mu\) and finite variance \(\sigma^{2}\), then the sample mean \(\bar{X}_{n}=n^{-1}\sum_{i=1}^{n}X_{i}\) converges to \(\mu\) in probability: \(\Pr(\abs{\bar{X}_{n}-\mu}\ge a)\rightarrow0\) for every \(a>0\). Rests on Lemma 11.50.
Derivation. Derives Corollary 11.51. \(\avg{\bar{X}_{n}}=\mu\) and, by independence, \(\operatorname{var}\bar{X}_{n}=\sigma^{2}/n\). Then Equation (11.41) gives \(\Pr(\abs{\bar{X}_{n}-\mu}\ge a)\le\sigma^{2}/(na^{2})\rightarrow0\).
∎One further result is needed, a statement about modes of convergence rather than about any model here. It is elementary and is derived in full.
If \(A_{n}\rightarrow A\) in distribution and \(B_{n}\rightarrow b\) in probability with \(b\) a constant, then \(A_{n}+B_{n}\rightarrow A+b\) and \(A_{n}B_{n}\rightarrow bA\) in distribution. If moreover \(b\neq0\), then \(A_{n}/B_{n}\rightarrow A/b\) in distribution. Rests on Definition 11.21.
Derivation of the sum and product clauses. Derives Lemma 11.52. Write \(F\) for the distribution function of \(A\) and \(F_{n}(y)=\Pr(A_{n}\le y)\), both in the general sense of Definition 11.21: no density is assumed here, and \(F\) may jump — as it does for the constant limits met below — which is exactly why the continuity points have to be watched. Convergence in distribution means \(F_{n}(y)\rightarrow F(y)\) at every continuity point \(y\) of \(F\). Two remarks are used repeatedly. First, \(F\) is non-decreasing, so its discontinuities are at most countable — the open intervals \(\left(F(y^{-}),F(y^{+})\right)\) belonging to distinct jumps are disjoint and each contains a rational — and continuity points are therefore dense. Second, \(F(y)\rightarrow1\) as \(y\rightarrow+\infty\) and \(F(y)\rightarrow0\) as \(y\rightarrow-\infty\).
The sum. Let \(x\) be a continuity point of the distribution function of \(A+b\), that is, let \(F\) be continuous at \(x-b\), and let \(\varepsilon>0\). If \(A_{n}+B_{n}\le x\) and \(\abs{B_{n}-b}\le\varepsilon\) then \(A_{n}\le x-B_{n}\le x-b+\varepsilon\); conversely, if \(A_{n}\le x-b-\varepsilon\) and \(\abs{B_{n}-b}\le\varepsilon\) then \(A_{n}+B_{n}\le x\). Passing to probabilities,
with \(\delta_{n}=\Pr\!\left(\abs{B_{n}-b}>\varepsilon\right)\), which tends to \(0\) because \(B_{n}\rightarrow b\) in probability. Restrict \(\varepsilon\) to the values — all but countably many, by the first remark — for which both \(x-b\pm\varepsilon\) are continuity points of \(F\). Taking the lower and upper limits in Equation (11.42) gives
and letting \(\varepsilon\) decrease to zero through admissible values makes both ends tend to \(F(x-b)\), since \(F\) is continuous at \(x-b\). Hence \(\Pr(A_{n}+B_{n}\le x)\rightarrow F(x-b)=\Pr(A+b\le x)\), which is the claim.
Boundedness in probability. Given \(\delta>0\), choose \(K>0\) with \(\pm K\) both continuity points of \(F\), \(F(K)>1-\delta/2\) and \(F(-K)<\delta/2\); the second remark makes this possible and the first makes it possible at continuity points. Since \(\Pr(\abs{A_{n}}>K)\le\left(1-F_{n}(K)\right)+F_{n}(-K)\),
The product. Decompose \(A_{n}B_{n}=b\,A_{n}+A_{n}\left(B_{n}-b\right)\) and treat the two pieces separately. For the second, fix \(\varepsilon>0\) and \(\delta>0\) and take \(K\) as in Equation (11.43). Since \(\abs{A_{n}(B_{n}-b)}>\varepsilon\) forces either \(\abs{A_{n}}>K\) or \(\abs{B_{n}-b}>\varepsilon/K\),
whose second term tends to \(0\); so the upper limit is at most \(\delta\), and \(\delta\) was arbitrary. Thus \(A_{n}(B_{n}-b)\rightarrow0\) in probability.
For the first piece, \(b\,A_{n}\rightarrow b\,A\) in distribution. This is trivial for \(b=0\), both sides being the constant \(0\). For \(b>0\) the events \(\set{bA_{n}\le x}\) and \(\set{A_{n}\le x/b}\) coincide, and \(x\) is a continuity point of the law of \(bA\) exactly when \(x/b\) is one of \(F\). For \(b<0\) they are \(\set{A_{n}\ge x/b}\), so \(\Pr(bA_{n}\le x)=1-\Pr(A_{n}<x/b)\); and at a continuity point \(y\) of \(F\) one has \(F_{n}(y')\le\Pr(A_{n}<y)\le F_{n}(y)\) for every continuity point \(y'<y\), so the limits squeeze \(\Pr(A_{n}<y)\) between \(F(y')\) and \(F(y)\), and \(F(y')\rightarrow F(y)\) as \(y'\) increases to \(y\). Hence \(\Pr(bA_{n}\le x)\rightarrow1-F(x/b)=\Pr(bA\le x)\) at every such \(x\).
Applying the sum clause already proved, with \(bA_{n}\) in place of \(A_{n}\) and the constant \(0\) in place of \(b\), gives \(A_{n}B_{n}\rightarrow bA\) in distribution.
∎The quotient is used below and is not a special case of the other two, so it is worth reducing it to them explicitly rather than leaving the reader to supply the step.
Derivation of the quotient clause from the product clause. Derives Lemma 11.52. Let \(b\neq0\) and \(\varepsilon>0\). The map \(y\mapsto1/y\) is continuous at \(b\), so there is a \(\delta\in(0,\abs{b}/2)\) such that \(\abs{y-b}<\delta\) implies both \(y\neq0\) and \(\abs{1/y-1/b}<\varepsilon\). Hence
the same bound also showing that \(B_{n}\) is nonzero with probability tending to one, so the quotient is defined on an event of probability approaching \(1\). Thus \(1/B_{n}\rightarrow1/b\) in probability — this is the continuous-mapping theorem, in the only case this chapter needs, and its proof is the two lines just given. Applying the product clause to \(A_{n}\) and \(1/B_{n}\) gives \(A_{n}/B_{n}\rightarrow A/b\) in distribution.
∎All three clauses are therefore derived here, and Lemma 11.52 joins the list of results this chapter proves rather than imports. Four results that are stated in this chapter are proved not here but in the appendix, because their derivations are long: Theorem 11.45, Theorem 11.48 and the Berry–Esseen inequality Equation (11.40) of Section 11.3, and Rice's formula Theorem 11.102. What remains genuinely imported is smaller and is worth naming exactly: a handful of theorems of Lebesgue theory — dominated convergence, Fubini–Tonelli — which Real Analysis does not build and which the appendix names at each use; the weak law of large numbers in the form that assumes only integrability of the summand, which Section 11.4.4 needs wherever Corollary 11.51, resting as it does on Chebyshev's inequality, has no finite variance to work with; and, for the multi-parameter form of Theorem 11.88, the Cramér–Wold device and the continuous-mapping theorem in \(\R^{k}\).
The likelihood function
A statistical model for data taking values in a set \(\mathcal{X}\) is a family \(\set{f(\,\cdot\,;\theta):\theta\in\Theta}\) of probability densities (or, for discrete data, probability mass functions) on \(\mathcal{X}\), indexed by a parameter \(\theta\) ranging over a set \(\Theta\subseteq\R^{k}\).
The parameter is what the experiment is trying to learn; the model is the physics, written as a statement about what data each possible value of the parameter would produce. Nothing in the definition is dimensionful, but every component of \(\theta\) in a physical application carries an SI dimension — a mass in \(\mathrm{kg}\), a rate in \(/\mathrm{s}\), a cross-section in \(\mathrm{m}^{2}\) — and the estimators below inherit it. What is dimensionless is everything built from a ratio of likelihoods: the test statistics of Section 11.5, the \(p\)-values and significances read off them, and the trials factor, global \(p\)-value and global significance of Section 11.6. The dimensionful quantities in that last section are the scanned range and the resolution that fix the correction, not the correction itself.
The likelihood function, and with it most of the vocabulary of this section — parameter, statistic, likelihood, sufficiency, efficiency, and the information of Definition 11.60 — is Fisher's, introduced in a single paper [Fisher:1922].
Given observed data \(x\in\mathcal{X}\), the likelihood function is \(L(\theta)=f(x;\theta)\), regarded as a function of \(\theta\) with \(x\) held fixed at its observed value. For an independent sample \(x=(x_{1},\dots,x_{n})\) from a common density,
Rests on Definition 11.53.
\(L\) is not a probability density in \(\theta\). The integral \(\int_{\Theta}L(\theta)\,\dd\theta\) is not \(1\), need not be finite, and changes if \(\theta\) is reparametrised — it has no meaning. Nor is \(L\) determined by the data: multiplying \(f\) by any factor that depends on \(x\) alone leaves the model unchanged and rescales \(L\). What survives both objections is a ratio of likelihoods at two parameter values, and that is why every construction below is built from ratios and never from \(L\) on its own.
A maximum-likelihood estimator (MLE) is any \(\hat{\theta}=\hat{\theta}(x)\) with \(L(\hat{\theta})=\sup_{\theta\in\Theta}L(\theta)\). Rests on Definition 11.54.
Let \(g:\Theta\rightarrow\Theta^{*}\) be any map and define the induced likelihood of \(\tau\in\Theta^{*}\) by
If \(\hat{\theta}\) maximises \(L\), then \(g(\hat{\theta})\) maximises \(L^{*}\). Rests on Definition 11.56.
Derivation. Derives Proposition 11.57. For every \(\tau\) the supremum in Equation (11.46) is taken over a subset of \(\Theta\), so \(L^{*}(\tau)\le\sup_{\Theta}L=L(\hat{\theta})\). On the other hand \(\hat{\theta}\) itself lies in the fibre over \(\tau=g(\hat{\theta})\), so \(L^{*}(g(\hat{\theta}))\ge L(\hat{\theta})\). The two inequalities force \(L^{*}(g(\hat{\theta}))=L(\hat{\theta})=\max_{\tau}L^{*}(\tau)\).
∎Invariance is the property that makes maximum likelihood usable in physics at all: an experiment that estimates a decay rate has thereby estimated the lifetime, with no separate calculation and no ambiguity about which parametrisation was “the right one”. It is also, taken with \(g\) a projection, exactly the construction of the profile likelihood in Section 11.4.5.
A detector records \(n\) events in a fixed exposure, the events being independent and the expected number \(\nu>0\). Then \(f(n;\nu)=\nu^{n}\ee^{-\nu}/n!\), so \(\ell(\nu)=n\ln\nu-\nu-\ln n!\) and \(\dv{\ell}{\nu}=n/\nu-1\), which vanishes at \(\hat{\nu}=n\). The second derivative \(-n/\nu^{2}\) is negative there, so \(\hat{\nu}=n\): the estimate of the expected count is the observed count. By Proposition 11.57 the estimate of \(\sqrt{\nu}\) is \(\sqrt{n}\), and if the exposure is a live time \(t_{\text{live}}\) then the estimate of the rate \(\Gamma=\nu/t_{\text{live}}\) is \(n/t_{\text{live}}\), carried in \(/\mathrm{s}\). Rests on Definition 11.56 and Proposition 11.57.
Score and Fisher information
Everything asymptotic below needs the model to be smooth in \(\theta\) and the data to be informative about it. Those requirements are made explicit once, here, and referred to by name afterwards. Let \(\theta_{0}\) denote the true value.
A model satisfies the regularity conditions at \(\theta_{0}\) when
-
[(R1)] common support: the set \(\set{x:f(x;\theta)>0}\) does not depend on \(\theta\);
-
[(R2)] identifiability: distinct parameter values give distinct distributions, \(\theta\neq\theta'\Rightarrow f(\,\cdot\,;\theta)\neq f(\,\cdot\,;\theta')\);
-
[(R3)] interior point: \(\theta_{0}\) lies in the interior of \(\Theta\);
-
[(R4)] differentiability under the integral: \(\theta\mapsto f(x;\theta)\) is three times continuously differentiable near \(\theta_{0}\), and the first two derivatives may be taken under \(\int\dd x\);
-
[(R5)] non-degenerate information: the quantity \(I_{1}(\theta_{0})\) of Definition 11.60 is finite and nonzero, and there are a neighbourhood of \(\theta_{0}\) and an integrable \(M(x)\) with \(\abs{\pp^{3}\ln f(x;\theta)/\pp\theta^{3}}\le M(x)\) throughout it.
Each of the five fails somewhere in physics, and the failures are not pathologies of the mathematics: (R1) fails for a uniform density on \([0,\theta]\), (R2) for a signal mass when the signal strength is zero, (R3) for a rate or a squared mass constrained to be non-negative and tested at zero. The last two are exactly the cases Section 11.5.4 has to treat.
Both objects of the next definition, and the word “information” attached to the second of them, are Fisher's [Fisher:1922].
For a one-parameter model the score is
and the Fisher information carried by one observation is \(I_{1}(\theta)=\avg{U(\theta;X)^{2}}\), the expectation taken under \(f(\,\cdot\,;\theta)\). For a \(k\)-parameter model the information is the \(k\times k\) matrix with entries \(\left[I_{1}(\theta)\right]_{ab} =\avg{\pdv{\ln f}{\theta^{a}}\pdv{\ln f}{\theta^{b}}}\). Rests on Definition 11.53.
Under (R1) and (R4), \(\avg{U(\theta;X)}=0\), so \(I_{1}(\theta)=\operatorname{var}U(\theta;X)\), and
For an independent sample of size \(n\) the total information is \(I_{n}(\theta)=n\,I_{1}(\theta)\). Rests on Definitions 11.59 and 11.60.
Derivation. Derives Lemma 11.61. Because the support does not move with \(\theta\), the domain of the integral below is fixed, and (R4) licenses differentiating under it:
Differentiate that identity once more with respect to \(\theta\), again under the integral:
In the second integral write \(\pp f/\pp\theta=(\pp\ln f/\pp\theta)f\); it becomes \(\avg{U^{2}}=I_{1}\). Rearranging gives Equation (11.48). Finally, for an independent sample \(\ell=\sum_{i}\ln f(x_{i};\theta)\), so the total score is a sum of \(n\) independent copies of \(U\), each of mean zero; variances of independent variables add, giving \(I_{n}=nI_{1}\).
∎Information is additive in the data and that is the whole reason a measurement improves with running time: doubling the exposure doubles \(I_{n}\) and, by the bound of the next subsection, halves the smallest attainable variance.
For the Poisson model of Example 11.58, \(U(\nu;n)=n/\nu-1\) and \(\pp^{2}\ell/\pp\nu^{2}=-n/\nu^{2}\), so Equation (11.48) gives \(I(\nu)=\avg{n}/\nu^{2}=1/\nu\). The information is therefore largest where the expected count is smallest, but the conclusion to draw from that is the opposite of the obvious one, because the bound it supplies is a bound on an absolute variance. By the Cramér–Rao inequality of the next subsection, Equation (11.53),
A rare process is thus measured well in absolute terms — the standard deviation \(\sqrt{\nu}\) shrinks as the process gets rarer — and badly in fractional terms, the fractional uncertainty \(\nu^{-1/2}\) growing without bound as \(\nu\rightarrow0\). Ten expected events are known to \(\pm3.2\) events, which is about \(32\) per cent; \(10^{6}\) expected events are known to \(\pm10^{3}\) events, which is \(0.1\) per cent. The scale-free comparison is the information about \(\ln\nu\), which by the same calculation is \(I(\ln\nu)=\nu^{2}I(\nu)=\nu\): it is smallest for a rare process, and that is the honest statement of why a rare process needs a long exposure. Rests on Example 11.58, Lemma 11.61 and Theorem 11.64.
The Cramér–Rao bound
The subject of this subsection is the information inequality [Frechet:1943] [Rao:1945], universally called the Cramér–Rao bound; Remark 11.66 records why that name and the priority disagree. It rests on one algebraic fact about second moments, which the treatise states and proves here rather than importing: Hilbert Spaces, where a reader would look for the Cauchy–Schwarz inequality, reserves its headings and proves nothing, and Linear Algebra and Representation Theory likewise leaves the Schwarz inequality a reserved heading with no content.
Let \(V\) and \(W\) be random variables of finite variance. Then
with equality if and only if there are constants \((a,c)\neq(0,0)\) with \(a\left(V-\avg{V}\right)+c\left(W-\avg{W}\right)=0\) with probability one. Rests on Definitions 11.6 and 11.11.
Derivation. Derives Lemma 11.63. Write \(\tilde{V}=V-\avg{V}\) and \(\tilde{W}=W-\avg{W}\), both of mean zero and finite second moment. Their product is integrable, since \(\abs{\tilde{V}\tilde{W}}\le\tfrac12(\tilde{V}^{2}+\tilde{W}^{2})\), so \(\operatorname{cov}(V,W)=\avg{\tilde{V}\tilde{W}}\) is a finite number and
If \(\operatorname{var}V=0\) then \(\tilde{V}=0\) almost surely, both sides of Equation (11.50) vanish, and the stated equality condition holds with \((a,c)=(1,0)\). Otherwise \(\phi\) is a genuine quadratic that is nowhere negative, so its discriminant is non-positive: \(4\operatorname{cov}(V,W)^{2}-4\operatorname{var}V\operatorname{var}W\le0\), which is Equation (11.50). Equality holds exactly when the discriminant vanishes, i.e. when \(\phi\) has a real root \(s_{0}\); and \(\phi(s_{0})=0\) says that the non-negative variable \(\left(s_{0}\tilde{V}+\tilde{W}\right)^{2}\) has zero mean, hence vanishes almost surely, which is the stated condition with \((a,c)=(s_{0},1)\).
∎Let the model satisfy (R1) and (R4), let the information be non-degenerate,
and let \(T(X)\) be an estimator with finite variance and mean \(\avg{T}=\theta+b(\theta)\), where the bias \(b\) is differentiable and \(\pp/\pp\theta\) may be taken under \(\int T f\,\dd x\). Then
In particular an unbiased estimator has \(\operatorname{var}T\ge1/I(\theta)\), which for an independent sample of size \(n\) is \(1/(n I_{1}(\theta))\). Rests on Definition 11.60, Definition 11.59, Lemma 11.61 and Lemma 11.63.
Both halves of Equation (11.52) are load-bearing and neither is cosmetic: finiteness is what lets the Cauchy–Schwarz step be applied to the score at all, since \(\operatorname{var}U=I(\theta)\) must be a number, and non-vanishing is what lets Equation (11.53) divide by it. Where the regularity conditions of Definition 11.59 are in force, Equation (11.52) is the first half of their clause (R5) and nothing new is being assumed; the theorem is stated without (R2), (R3) and (R5) because it needs none of the rest of them.
Derivation. Derives Theorem 11.64. By Lemma 11.61 the score has mean zero, so \(\operatorname{cov}(T,U)=\avg{TU}-\avg{T}\avg{U}=\avg{TU}\). Evaluate that expectation by the same manoeuvre as before, writing \(U f=\pp f/\pp\theta\):
Now apply Lemma 11.63 to \(T\) and \(U\), whose variances are finite by hypothesis and by Equation (11.52) respectively:
which is Equation (11.53).
∎Equality holds in Equation (11.53) for a given \(\theta\) if and only if \(U(\theta;x)=c(\theta)\left(T(x)-\avg{T}\right)\) with probability one, for some \(c(\theta)\). If this holds for all \(\theta\) in an interval then, integrating in \(\theta\),
a one-parameter exponential family with \(T\) as its natural sufficient statistic. Rests on Theorem 11.64 and Lemma 11.63.
Derivation. Derives Corollary 11.65. By Lemma 11.63 equality holds exactly when the two centred variables are proportional with probability one; since \(\operatorname{var}U =I(\theta)>0\) the proportionality can be solved for the score, giving \(U=c(\theta)(T-\avg{T})\) almost surely. Given that for every \(\theta\), integrate \(\pp\ln f/\pp\theta=c(\theta)(T(x)-\avg{T}_{\theta})\) with respect to \(\theta\): the \(x\)-dependence enters only through the factor \(T(x)\) multiplying \(\int c\,\dd\theta\), and the remaining terms are functions of \(\theta\) alone plus a constant of integration \(C(x)\).
∎So the bound is achievable at finite sample size only inside the exponential family; the Poisson model of Example 11.62 is such a case, and indeed \(\operatorname{var}\hat{\nu}=\operatorname{var}n=\nu=1/I(\nu)\), saturating Equation (11.53) exactly. Outside that family the bound is still true and still useful, because Theorem 11.67 shows the MLE attains it asymptotically.
The inequality was obtained independently, within a few years of one another, by Fréchet, Darmois, Rao and Cramér; the name that stuck records only two of the four, and neither of those two was first. The earliest published statement is Fréchet's, in 1943 [Frechet:1943]; Rao's independent derivation, which is the one that introduced the information-geometric reading now standard, appeared in 1945 [Rao:1945]; Cramér's is later still, in his 1946 textbook. This treatise keeps the customary name “Cramér–Rao” because that is what a reader will meet everywhere else, and cites the two papers it actually has — [Frechet:1943] [Rao:1945] — rather than the book, for which it carries no bibliography entry at all. A reader tracing the priority should start from Fréchet, not from either of the names in the title. The statement above is in any case derived in full and stands on its derivation.
Consistency and asymptotic normality
Let \(X_{1},\dots,X_{n}\) be independent and identically distributed with density \(f(\,\cdot\,;\theta_{0})\), and let the regularity conditions (R1)–(R5) hold at \(\theta_{0}\). Then there is a sequence \(\hat{\theta}_{n}\) of solutions of the score equation \(\ell_{n}'(\theta)=0\) such that
-
\(\hat{\theta}_{n}\rightarrow\theta_{0}\) in probability (consistency); and
-
\(\sqrt{n}\left(\hat{\theta}_{n}-\theta_{0}\right) \rightarrow\mathcal{N}\!\left(0,\,I_{1}(\theta_{0})^{-1}\right)\) in distribution (asymptotic normality).
The same holds in \(k\) parameters with \(I_{1}(\theta_{0})^{-1}\) the inverse of the information matrix. Rests on Definition 11.56, Definition 11.59, Lemma 11.61, Theorem 11.46, Corollary 11.51 and Lemma 11.52.
Derivation of (ii), given (i). Derives Theorem 11.67. Write \(\ell_{n}(\theta)=\sum_{i}\ln f(X_{i};\theta)\) and expand its derivative about \(\theta_{0}\) with a Lagrange remainder, which (R4) permits:
with \(\bar{\theta}_{n}\) between \(\theta_{0}\) and \(\hat{\theta}_{n}\). Solve for the deviation and multiply by \(\sqrt{n}\):
Treat the three pieces separately.
The numerator is \(n^{-1/2}\sum_{i}U(\theta_{0};X_{i})\), a normalised sum of independent identically distributed variables of mean zero (Lemma 11.61) and variance \(I_{1}(\theta_{0})\), finite and nonzero by (R5). Theorem 11.46 therefore gives \(n^{-1/2}\ell_{n}'(\theta_{0})\rightarrow \mathcal{N}(0,I_{1}(\theta_{0}))\) in distribution.
The first term of the denominator is \(-n^{-1}\sum_{i}\pp^{2}\ln f(X_{i};\theta_{0})/\pp\theta^{2}\), an average of independent identically distributed variables whose common mean is \(I_{1}(\theta_{0})\) by Equation (11.48) — an identity that presupposes the summand integrable, and asserts nothing about its second moment. Nothing in (R1)–(R5) bounds that second moment, so Corollary 11.51, which is the Chebyshev form and needs a finite variance, does not apply here; the weak law in the form that needs only integrability [Billingsley:1995] gives convergence to \(I_{1}(\theta_{0})\) in probability.
The second term of the denominator is bounded in absolute value by \(\tfrac12\abs{\hat{\theta}_{n}-\theta_{0}}\cdot n^{-1}\sum_{i}M(X_{i})\), using the domination in (R5). The average converges in probability to \(\avg{M}<\infty\), and \(\abs{\hat{\theta}_{n}-\theta_{0}}\rightarrow0\) in probability by (i); the product therefore tends to zero in probability.
The two denominator terms therefore converge in probability, by the sum clause of Lemma 11.52, to the constant \(I_{1}(\theta_{0})\), which is nonzero by (R5). Equation (11.57) is a quotient, not a sum or a product, so the passage to the limit is the quotient clause of Lemma 11.52 — available exactly because \(I_{1}(\theta_{0})\neq0\), and resting on the convergence \(1/B_{n}\rightarrow1/I_{1}(\theta_{0})\) in probability established at Equation (11.44). The ratio therefore converges in distribution to \(\mathcal{N}(0,I_{1}(\theta_{0}))/I_{1}(\theta_{0}) =\mathcal{N}(0,I_{1}(\theta_{0})^{-1})\), since scaling a Gaussian by \(1/I_{1}\) scales its variance by \(1/I_{1}^{2}\).
∎The variance \(1/(nI_{1})\) is precisely the Cramér–Rao bound Equation (11.53) for an unbiased estimator: the MLE is asymptotically efficient in Fisher's sense [Fisher:1922], wasting none of the information in the data in the limit of many observations. That is the whole case for maximum likelihood, and it is an asymptotic case only — at finite \(n\) the MLE can be biased, and Example 11.58 is the exception rather than the rule in being exactly unbiased.
The consistency statement (i) has a short heart and a long technical remainder, and it is worth separating them. The heart is an inequality of Jensen's type, used nowhere else in this treatise and therefore proved here rather than quoted.
Let \(Y>0\) almost surely with \(\avg{Y}<\infty\). Then
with equality if and only if \(Y=\avg{Y}\) with probability one. (The left side may be \(-\infty\), in which case the inequality is trivially true.) Rests on Theorem 7.35.
Derivation. Derives Lemma 11.68. Put \(a=\avg{Y}\), which is finite and strictly positive because \(Y>0\). The logarithm has second derivative \(-1/y^{2}<0\) on \((0,\infty)\), so it is strictly concave there and its graph lies below the tangent at \(a\):
with equality only at \(y=a\). To see Equation (11.59) without appealing to a general theory of convexity, apply the mean value theorem (Theorem 7.35) to \(h(y)=\ln y-\ln a-(y-a)/a\): \(h(a)=0\) and \(h'(y)=1/y-1/a\), which is positive for \(y<a\) and negative for \(y>a\), so \(h\) increases up to \(a\) and decreases after it, attaining its strict maximum \(0\) there. Now put \(y=Y\) in Equation (11.59) and take expectations, which is legitimate because the right-hand side is integrable:
Equality forces the non-positive variable \(\ln Y-\ln a-(Y-a)/a\) to have zero mean, hence to vanish almost surely, which by the strictness in Equation (11.59) means \(Y=a\) almost surely.
∎Apply Lemma 11.68 to the likelihood ratio \(Y=f(X;\theta)/f(X;\theta_{0})\), which is positive almost surely because (R1) makes the support common to all \(\theta\), and whose mean is \(\int f(x;\theta)\,\dd x=1\). For any \(\theta\neq\theta_{0}\),
with equality only if the ratio is constant — and, its mean being \(1\), only if that constant is \(1\), i.e. only if the two distributions coincide, which (R2) excludes. So the expected log-likelihood per observation is strictly maximised at the true value, and by Corollary 11.51 the observed \(n^{-1}[\ell_{n}(\theta)-\ell_{n}(\theta_{0})]\) converges to that negative number for each fixed \(\theta\). Passing from “for each fixed \(\theta\)” to a statement about the global maximiser of the likelihood needs that convergence to be uniform on \(\Theta\), or a compactness argument in its place; that step is Wald's, under conditions weaker than the ones assumed here [Wald:1949], and it is not carried out in this treatise.
Clause (i) of Theorem 11.67 claims less than that, and the less is enough for clause (ii): it asserts a consistent sequence of solutions of the score equation, which the expansion already in hand delivers.
Derivation of (i). Derives Theorem 11.67. Take \(k=1\); the modification for \(k\) parameters is given at the end. Write \(I=I_{1}(\theta_{0})\), which is finite and strictly positive by (R5), and let \(a_{0}>0\) be small enough that \([\theta_{0}-a_{0},\theta_{0}+a_{0}]\) lies both inside \(\Theta\), which (R3) permits, and inside the neighbourhood on which (R5) supplies the dominating function \(M\). For \(0<a\le a_{0}\) and \(\theta=\theta_{0}\pm a\), expand with a Lagrange remainder, as (R4) allows:
with \(\bar{\theta}\) between \(\theta_{0}\) and \(\theta\).
Each of the three coefficients is an average of independent identically distributed terms. The first, \(n^{-1}\sum_{i}U(\theta_{0};X_{i})\), converges in probability to \(0\), the score having mean zero by Lemma 11.61; the second converges in probability to \(-I\) by Equation (11.48); and the third is bounded in absolute value by \(n^{-1}\sum_{i}M(X_{i})\), which converges in probability to \(\bar{M}=\avg{M}<\infty\). Corollary 11.51 gives the first of these, the score having variance \(I_{1}(\theta_{0})\) by (R5). It gives neither of the other two: (R5) bounds the variance of \(\pp^{2}\ln f/\pp\theta^{2}\) no more than it bounds that of \(M\), which it assumes integrable and nothing more, and Equation (11.48) says only that the second derivative has a mean. Both are averages of integrable summands, so for both the weak law in the form that needs only integrability is used instead [Billingsley:1995].
Now shrink the radius, requiring in addition
and let \(E_{n}(a)\) be the event on which all three of
hold. By the previous paragraph \(\Pr(E_{n}(a))\rightarrow1\). On \(E_{n}(a)\), and for either sign of \(\theta-\theta_{0}=\pm a\), the three terms of Equation (11.61) are bounded by \(a^{2}I/16\), by \(-a^{2}I/2+a^{2}I/8\) and, using Equation (11.62), by \(a^{3}(\bar{M}+1)/6<a^{2}I/6\) respectively, so that
From Equations (11.61) and (11.62) (bound the three terms on the event \(E_{n}(a)\) and add). So on \(E_{n}(a)\) the log-likelihood is strictly smaller at both endpoints of \([\theta_{0}-a,\theta_{0}+a]\) than at its midpoint. It is continuous there, so by the extreme value theorem (Theorem 7.24) it attains a maximum on that interval; the maximum is at neither endpoint, hence at an interior point, where the derivative of a differentiable function vanishes. The score equation \(\ell_{n}'(\theta)=0\) therefore has a solution in \((\theta_{0}-a,\theta_{0}+a)\).
Define \(\hat{\theta}_{n}\) to be a solution of the score equation in \([\theta_{0}-a_{0},\theta_{0}+a_{0}]\) nearest to \(\theta_{0}\) when one exists — the solution set is closed, being the zero set of a continuous function, so a nearest point is attained — and any fixed point of \(\Theta\) otherwise. For every \(a\) satisfying Equation (11.62), \(E_{n}(a)\) implies \(\abs{\hat{\theta}_{n}-\theta_{0}}<a\), so \(\Pr(\abs{\hat{\theta}_{n}-\theta_{0}}<a)\rightarrow1\); and since the event grows with \(a\), the same holds for every \(a>0\). That is convergence in probability.
In \(k\) parameters, replace the interval by the closed ball of radius \(a\) about \(\theta_{0}\) and its endpoints by the sphere, expanding along \(\theta=\theta_{0}+a\vect{e}\) with \(\vect{e}\) a unit vector. The linear term is bounded by \(a\) times the length of \(n^{-1}\nabla\ell_{n}\); the quadratic term is \(-\tfrac12a^{2}\,\vect{e}\transpose I\vect{e}\) up to the same error; and the cubic term is bounded as before, with \(M\) dominating every third partial derivative. Read in \(k\) parameters, (R5) makes \(I\) positive definite, and \(\vect{e}\mapsto\vect{e}\transpose I\vect{e}\) is continuous and strictly positive on the unit sphere, which is closed and bounded, so it has a strictly positive minimum \(\lambda\) (Theorem 7.24); putting \(\lambda\) in place of \(I\) throughout gives Equation (11.63) at every point of the sphere, and the argument closes as before with the gradient in place of the derivative.
∎The sequence just constructed solves \(\ell_{n}'(\hat{\theta}_{n})=0\) on the event \(E_{n}(a_{0})\), whose probability tends to one, and is defined arbitrarily on the complement. That is the standard reading of Theorem 11.67(i), and the arbitrariness is immaterial to both clauses: each is a statement about a limit in probability or in distribution, and neither is affected by altering the sequence on events whose probability tends to zero. What is not claimed, here or anywhere in this chapter, is that the root so chosen is the global maximiser of the likelihood; that is the content of Wald's theorem [Wald:1949], and a likelihood with several local maxima can have a consistent root far from its highest peak at any finite \(n\).
Nuisance parameters and the profile likelihood
A real measurement almost never has a one-dimensional parameter. A search for a signal must simultaneously determine the size of the background, the detector efficiency and the energy scale, none of which is of interest and all of which are unknown. Write
with \(\mu\) the parameter of interest and \(\vect{\nu}\) the nuisance parameters.
The profile likelihood is
the double hat marking a maximisation carried out at fixed \(\mu\). The profile likelihood ratio is \(\lambda(\mu)=L_{\mathrm{p}}(\mu)/L(\hat{\mu},\hat{\vect{\nu}})\), where \((\hat{\mu},\hat{\vect{\nu}})\) is the unconstrained MLE. Rests on Definitions 11.54 and 11.56.
\(0<\lambda(\mu)\le1\) for every \(\mu\), with \(\lambda(\hat{\mu})=1\); and \(L_{\mathrm{p}}\) is maximised at \(\hat{\mu}\). Rests on Definition 11.70 and Proposition 11.57.
Derivation. Derives Proposition 11.71. \(L_{\mathrm{p}}\) is the induced likelihood Equation (11.46) for the map \(g(\mu,\vect{\nu})=\mu\), so Proposition 11.57 applies verbatim: \(g\) of the unconstrained maximiser, namely \(\hat{\mu}\), maximises \(L_{\mathrm{p}}\), and the maximum value is \(L(\hat{\mu},\hat{\vect{\nu}})\). The ratio is therefore at most \(1\) and equals \(1\) at \(\hat{\mu}\); it is positive because \(L_{\mathrm{p}}(\mu)\ge L(\mu,\vect{\nu})>0\) for any admissible \(\vect{\nu}\).
∎Partition the information matrix of the full model as
with \(I_{\nu\nu}\) positive definite. Then the asymptotic variance of \(\hat{\mu}\) is \(\left(nI_{\mathrm{p}}\right)^{-1}\) with
with equality if and only if \(I_{\mu\nu}=0\). Rests on Definition 11.60 and Theorem 11.67.
Derivation. Derives Proposition 11.72. By Theorem 11.67 the asymptotic covariance of \(\sqrt{n}(\hat{\theta}-\theta_{0})\) is \(I^{-1}\), so the asymptotic variance of \(\sqrt{n}(\hat{\mu}-\mu_{0})\) is the \(\mu\mu\) entry of \(I^{-1}\). To identify it, solve \(I\,(x,\vect{y})\transpose=(1,\vect{0})\transpose\): the lower block gives \(I_{\nu\mu}x+I_{\nu\nu}\vect{y}=\vect{0}\), hence \(\vect{y}=-I_{\nu\nu}^{-1}I_{\nu\mu}x\), and substituting in the upper block gives \(\left(I_{\mu\mu}-I_{\mu\nu}I_{\nu\nu}^{-1}I_{\nu\mu}\right)x=1\). By construction \(x=\left(I^{-1}\right)_{\mu\mu}\), so \(\left(I^{-1}\right)_{\mu\mu}=I_{\mathrm{p}}^{-1}\). Finally \(I_{\mu\nu}I_{\nu\nu}^{-1}I_{\nu\mu}\ge0\) because \(I_{\nu\nu}^{-1}\) is positive definite, and it vanishes exactly when \(I_{\mu\nu}=0\).
∎The content of Equation (11.67) is that not knowing the background costs precision on the signal, and costs nothing only when the two are information-orthogonal. Designing an analysis so that \(I_{\mu\nu}\) is small — a control region that fixes the background without touching the signal region — is the practical form of that statement.
\(L_{\mathrm{p}}\) is not the likelihood of any probability model for the data: there is generally no density \(h(x;\mu)\) with \(h(x;\mu)=L_{\mathrm{p}}(\mu)\) for all \(x\), because the maximising \(\hat{\hat{\vect{\nu}}}(\mu)\) depends on \(x\) as well as on \(\mu\). The identities of Lemma 11.61 therefore do not hold for it, and its curvature can understate the true uncertainty at finite sample size. What survives is the asymptotic statement proved in Section 11.5.4, which is the only property the search analyses of Part XI — Quantum Field Theory and the Standard Model use [Cowan:2011].
Hypothesis testing
Estimation asks what value a parameter has. Testing asks a coarser and often more consequential question: is a proposed value compatible with the data at all? The apparatus is small — a critical region, its size, its power — and one theorem, due to Neyman and Pearson, says which critical region to use. Everything a modern search does is a corollary of that theorem plus the asymptotic distribution derived in Section 11.5.4.
Tests, size, power and $p$-values
A hypothesis is a subset \(\Theta_{0}\subseteq\Theta\); it is simple if \(\Theta_{0}\) is a single point and composite otherwise. A test is a measurable function \(\psi:\mathcal{X}\rightarrow[0,1]\) giving the probability with which the null hypothesis \(H_{0}:\theta\in\Theta_{0}\) is rejected when \(x\) is observed; a test is non-randomised when \(\psi\) is the indicator of a critical region \(W\subseteq\mathcal{X}\). Its size is
and its power against an alternative \(\theta\in\Theta_{1}\) is \(\beta(\theta)=\avg{\psi}_{\theta}\). Rests on Definition 11.53.
Size is the probability of rejecting a true null — the rate of false discovery claims, fixed in advance by the experimenter. Power is the probability of rejecting a false one, and is not free: it is whatever the data and the chosen critical region deliver.
Let \(t(X)\) be a test statistic, large values of which disfavour a simple \(H_{0}\). The \(p\)-value of an observation \(t_{\mathrm{obs}}\) is
Rests on Definition 11.74.
If \(t(X)\) has a continuous, strictly decreasing survival function \(G(s)=\Pr(t\ge s\mid H_{0})\) on its support, then \(p=G(t(X))\) is uniformly distributed on \([0,1]\) under \(H_{0}\). Consequently the test that rejects when \(p\le\alpha\) has size exactly \(\alpha\). Rests on Definitions 11.74 and 11.75.
Derivation. Derives Proposition 11.76. For \(\alpha\in(0,1)\), \(G\) is invertible on its range, and \(G\) is decreasing, so \(\set{G(t)\le\alpha}=\set{t\ge G^{-1}(\alpha)}\). Hence \(\Pr(p\le\alpha\mid H_{0})=\Pr(t\ge G^{-1}(\alpha)\mid H_{0}) =G(G^{-1}(\alpha))=\alpha\), which is the distribution function of the uniform law on \([0,1]\).
∎That uniformity is what makes a \(p\)-value comparable across experiments: it is a calibrated quantity, and a threshold on it means the same thing whatever the statistic and whatever the model. It is also the source of the most persistent misreading in the subject, which is worth stating without diplomacy.
The \(p\)-value is a probability computed assuming \(H_{0}\) holds. It is \(\Pr(\text{data at least this extreme}\mid H_{0})\). It is not \(\Pr(H_{0}\mid\text{data})\), and the two are different numbers that answer different questions. Nothing in Equation (11.69) even mentions the alternative hypothesis or the plausibility of \(H_{0}\) before the data arrived, and both enter the second quantity:
with \(\pi_{0}\) the prior probability of \(H_{0}\) and \(\pi\) the prior density over the alternative. The gap is not small. Take an observation with \(q=25\), which Section 11.5.5 will show corresponds to \(p=2.87\times 10^{-7}\) — a “five sigma” result.
The comparison needs the direction of one inequality stated, or it proves nothing. What Equation (11.80) supplies is the maximised likelihood ratio \(\ee^{q/2}\), whereas Equation (11.70) contains the likelihood averaged over the prior. Since \(\pi\) is a probability density, integrating to \(1\),
the last two equalities holding whenever the unconstrained maximiser lies in \(\Theta_{1}\), which is the case here because \(q>0\). So replacing the average by the maximum can only inflate the odds against \(H_{0}\): whatever the prior, the posterior odds on the alternative are at most \(\ee^{q/2}(1-\pi_{0})/\pi_{0}\), and \(\Pr(H_{0}\mid x)\) is at least \(\left[1+\ee^{q/2}(1-\pi_{0})/\pi_{0}\right]^{-1}\).
With \(\ee^{q/2}=2.68\times 10^{5}\) and prior odds of \(1\) to \(1000\) against the alternative, the posterior odds are therefore at most \(2.68\times 10^{5}\times10^{-3}=2.68\times 10^{2}\), so \(\Pr(H_{0}\mid x)\ge1/(1+2.68\times 10^{2})=3.7\times 10^{-3}\). The inequality runs the helpful way: the true posterior probability of the null is at least four orders of magnitude larger than the \(p\)-value — larger by a factor of at least \(1.3\times 10^{4}\) — and using a real prior density in place of its maximum can only widen the gap. The two numbers were never the same quantity.
Two corollaries follow. A small \(p\)-value is evidence against \(H_{0}\) only in the presence of an alternative that predicts the observation better; with no such alternative it is a sign that something in the model is wrong, and the something is at least as likely to be the background estimate as the physics. And a \(p\)-value cannot be converted into a probability that a discovery is real without supplying a prior, which is a judgement and must be declared as one.
The Neyman–Pearson lemma
The lemma below is Neyman and Pearson's, from the joint paper of 1933 that created the whole framework of size, power and critical regions used above [Neyman:1933].
Let \(H_{0}\) and \(H_{1}\) be simple, with densities \(f_{0}\) and \(f_{1}\). Fix \(k\ge0\) and let
for any \(\gamma\) taking values in \([0,1]\), and write \(\alpha=\avg{\psi^{*}}_{0}\) for its size. Then every test \(\psi\) with \(\avg{\psi}_{0}\le\alpha\) satisfies \(\avg{\psi}_{1}\le\avg{\psi^{*}}_{1}\): no test of size at most \(\alpha\) is more powerful. Rests on Definition 11.74.
Derivation. Derives Theorem 11.78. Consider the function
It is non-negative everywhere. Where \(f_{1}>kf_{0}\) the second factor is positive and \(\psi^{*}=1\ge\psi\), so the first is non-negative; where \(f_{1}<kf_{0}\) the second factor is negative and \(\psi^{*}=0\le\psi\), so the first is non-positive; where \(f_{1}=kf_{0}\) the second factor vanishes. Integrating,
so that \(\avg{\psi^{*}}_{1}-\avg{\psi}_{1}\ge k\left(\alpha-\avg{\psi}_{0}\right)\). The right-hand side is non-negative because \(k\ge0\) and \(\avg{\psi}_{0}\le\alpha\) by hypothesis.
∎The lemma answers the design question completely for a simple-versus-simple test: order the sample space by the likelihood ratio \(f_{1}/f_{0}\) and reject the largest values. It is the reason every construction that follows is built from a likelihood ratio rather than from anything else. For composite hypotheses no uniformly most powerful test exists in general, and the statistic of Section 11.5.4 — the ratio of maximised likelihoods — is the standard substitute: not optimal by any theorem, but reducing to the Neyman–Pearson statistic when both hypotheses are simple.
The lemma is Neyman and Pearson's, from their joint work of the early 1930s, and the reference is [Neyman:1933]. Two traps attend that citation and are recorded here because both have been fallen into. The Pearson is Egon, not his father Karl, and not either of the two other Pearsons this bibliography carries, who work on silicon and on convection cells. And Crossref stores the title with the issue's article number “IX.” prefixed to it; that numeral is not part of the title and is dropped in the bibliography entry.
The chi-squared distribution
Let \(Z_{1},\dots,Z_{k}\) be independent standard Gaussian variables. The distribution of \(\sum_{i=1}^{k}Z_{i}^{2}\) is the chi-squared distribution with \(k\) degrees of freedom, written \(\chi^{2}_{k}\).
The \(\chi^{2}_{k}\) distribution has density
mean \(k\) and variance \(2k\). Rests on Definition 11.80, Definition 11.42 and Proposition 11.43.
Derivation. Derives Proposition 11.81. Work with the Laplace transform \(M_{k}(s)=\avg{\ee^{-sX}}\), which for a distribution on \((0,\infty)\) exists for all \(s\ge0\). For one degree of freedom, completing the exponent,
using \(\int\ee^{-au^{2}/2}\dd u=\sqrt{2\pi/a}\). The \(k\) squares are independent, so their transforms multiply exactly as characteristic functions do in Proposition 11.43(iii), giving \(M_{k}(s)=(1+2s)^{-k/2}\).
Now transform the claimed density. Substituting \(y=(\tfrac12+s)x\),
the middle integral being \(\Gamma(k/2)\) by definition. The two distributions have the same Laplace transform on \(s>0\); since both are supported on \([0,\infty)\) their transforms extend analytically to \(\operatorname{Re}s>-\tfrac12\) and agree there, in particular on the imaginary axis, where they are the characteristic functions of Definition 11.42. Uniqueness of the characteristic function [Billingsley:1995] [Feller:1971] then forces the distributions to coincide, which is Equation (11.73).
For the moments, differentiate \(M_{k}\) at \(s=0\): \(M_{k}'(s)=-k(1+2s)^{-k/2-1}\) and \(M_{k}''(s)=k(k+2)(1+2s)^{-k/2-2}\), so \(\avg{X}=-M_{k}'(0)=k\) and \(\avg{X^{2}}=M_{k}''(0)=k(k+2)\), whence \(\operatorname{var}X=k(k+2)-k^{2}=2k\).
∎For \(z\ge0\),
Rests on Definition 11.80.
Derivation. Derives Corollary 11.82. \(\chi^{2}_{1}=Z^{2}\) with \(Z\) standard Gaussian, and \(Z^{2}\ge z^{2}\) exactly when \(\abs{Z}\ge z\), an event of probability \(\Pr(Z\ge z)+\Pr(Z\le-z)=2(1-\Phi(z))\) by symmetry.
∎That factor of two is small, obvious, and responsible for more misquoted significances than any other step in the subject. It is tracked explicitly in Section 11.5.5.
If \(\vect{Y}\in\R^{k}\) is Gaussian with mean zero and positive-definite covariance \(\Sigma\), then \(\vect{Y}\transpose\Sigma^{-1}\vect{Y}\sim\chi^{2}_{k}\). More generally, if \(\vect{V}\in\R^{k}\) is Gaussian with mean zero and covariance \(\identity\) and \(P\) is the orthogonal projector onto a subspace of dimension \(r\), then \(\norm{P\vect{V}}^{2}\sim\chi^{2}_{r}\). Rests on Definition 11.80.
Derivation. Derives Lemma 11.83. By the spectral theorem for real symmetric matrices — diagonalise \(\Sigma\) by an orthogonal matrix and take the positive square roots of its eigenvalues, which are positive because \(\Sigma\) is positive definite — \(\Sigma\) has a symmetric positive-definite square root \(\Sigma^{1/2}\). Put \(\vect{V}=\Sigma^{-1/2}\vect{Y}\); a linear image of a Gaussian vector is Gaussian, with covariance \(\Sigma^{-1/2}\Sigma\,\Sigma^{-1/2}=\identity\), so the components of \(\vect{V}\) are independent standard Gaussians. Then \(\vect{Y}\transpose\Sigma^{-1}\vect{Y}=\vect{V}\transpose\vect{V} =\sum_{i}V_{i}^{2}\), which is \(\chi^{2}_{k}\) by Definition 11.80. For the second statement choose an orthonormal basis \(\vect{u}_{1},\dots,\vect{u}_{r}\) of the image of \(P\); the components \(\vect{u}_{j}\cdot\vect{V}\) are jointly Gaussian with \(\avg{(\vect{u}_{i}\cdot\vect{V})(\vect{u}_{j}\cdot\vect{V})} =\vect{u}_{i}\cdot\vect{u}_{j}=\delta_{ij}\), hence independent standard Gaussians, and \(\norm{P\vect{V}}^{2}=\sum_{j=1}^{r} (\vect{u}_{j}\cdot\vect{V})^{2}\).
∎The projector form of Lemma 11.83 has one consequence that metrology uses constantly and that is worth having as a labelled statement rather than rederived at each site.
Let \(x_{1},\dots,x_{n}\) be independent Gaussian measurements of one quantity \(\mu\), with known standard uncertainties \(u_{i}\), and let
be their inverse-variance weighted mean (Theorem 2.10). Then
whatever the value of \(\mu\). In particular \(\avg{\chi^{2}}=n-1\) with variance \(2(n-1)\), and the Birge ratio \(R=\sqrt{\chi^{2}/(n-1)}\) is the factor by which every \(u_{i}\) must be multiplied to make the set internally consistent, in the sense that the \(\chi^{2}\) recomputed with uncertainties \(Ru_{i}\) equals \(n-1\) exactly. Rests on Lemma 11.83, Proposition 11.81 and Theorem 2.10.
Derivation. Derives Corollary 11.84. Put \(V_{i}=(x_{i}-\mu)/u_{i}\), so that \(\vect{V}\in\R^{n}\) is Gaussian with mean zero and covariance \(\identity\). Write \(S=\sum_{k}u_{k}^{-2}\) and let \(\vect{m}\) be the vector with components \(m_{i}=u_{i}^{-1}/\sqrt{S}\), which is a unit vector because \(\sum_{i}m_{i}^{2}=S^{-1}\sum_{i}u_{i}^{-2}=1\). From Equation (11.76),
and therefore
where \(P=\identity-\vect{m}\vect{m}\transpose\) is the orthogonal projector onto the hyperplane orthogonal to \(\vect{m}\), of dimension \(n-1\). Hence \(\chi^{2}=\norm{P\vect{V}}^{2}\), which is \(\chi^{2}_{n-1}\) by the second statement of Lemma 11.83; \(\mu\) has cancelled, as Equation (11.79) shows. The mean and variance are Proposition 11.81 with \(k=n-1\), and the rescaling claim is immediate because \(\chi^{2}\) is homogeneous of degree \(-2\) in the uncertainties.
∎A Birge ratio near unity is consistency, not correctness: it says the scatter of the measurements matches the uncertainties they quote, and it is blind to a bias common to all of them. A ratio well above unity says at least one quoted uncertainty is too small, and gives no information about which. Expanding every uncertainty by \(R\) is therefore a confession rather than a correction, and it is what the CODATA adjustments do where inputs disagree [Mohr:2025]; the worked case is the six modern determinations of the gravitational constant in Experiment: The Cavendish Torsion Balance, where \(R\approx4\). The count of degrees of freedom is \(n-1\) and not \(n\) because one linear combination of the residuals is forced to vanish by the definition of \(\hat{x}\) — that is exactly the content of the projector in Equation (11.79).
The finite-dimensional spectral theorem is proved in this treatise for a self-adjoint operator on a complex inner-product space (Theorem 5.79 in Section 5.6, whose infinite-dimensional counterpart is Theorem 12.59); what is used above is its real orthogonal form, for a real symmetric matrix, which is a standard corollary and is nowhere separately stated. It is quoted here, in the same explicit way as in Expansions of Lie Algebras, and the reader should treat it as an assumption of this lemma rather than as something established earlier in the book. What is actually needed is weaker than the full theorem: only the existence of some invertible \(A\) with \(\Sigma=AA\transpose\), since \(\vect{V}=A^{-1}\vect{Y}\) then has covariance \(\identity\) and \(\vect{Y}\transpose\Sigma^{-1}\vect{Y} =\vect{V}\transpose\vect{V}\); the symmetric square root is one such \(A\), and Gram–Schmidt applied to the inner product defined by \(\Sigma^{-1}\) supplies a triangular one. The same weaker statement is all that Theorem 11.88 uses when it writes \(I^{1/2}\).
The likelihood-ratio statistic and Wilks' theorem
For a null hypothesis \(H_{0}:\theta\in\Theta_{0}\) inside a model with parameter set \(\Theta\), the likelihood-ratio test statistic is
where \(\hat{\theta}_{0}\) is the constrained and \(\hat{\theta}\) the unconstrained maximum-likelihood estimator. For a simple null \(\Theta_{0}=\set{\theta_{0}}\) this is \(q(\theta_{0})=-2\ln\left[L(\theta_{0})/L(\hat{\theta})\right]\); for a parameter of interest with nuisance parameters it is \(q(\mu)=-2\ln\lambda(\mu)\) with \(\lambda\) the profile likelihood ratio of Definition 11.70. Rests on Definitions 11.54, 11.56 and 11.70.
Non-negativity is immediate: the constrained supremum is over a subset, so the ratio is at most \(1\) and its logarithm at most zero. The statistic is also a pure number — the SI dimensions of \(\theta\) cancel in the ratio — which is what allows a single tabulated distribution to serve every measurement.
The asymptotic law of that statistic is Wilks' theorem [Wilks:1938]; its hypotheses are stated in full below because the two ways they fail are exactly the two corrections a collider search has to apply.
Let the model satisfy (R1)–(R5) at the true value \(\theta_{0}\), let the sample be independent and identically distributed of size \(n\), let \(\Theta\subseteq\R^{k}\) and let \(\Theta_{0}\) be, near \(\theta_{0}\), a smooth submanifold of dimension \(k-r\) — equivalently, let \(H_{0}\) impose \(r\) smooth, functionally independent constraints. If \(\theta_{0}\) lies in \(\Theta_{0}\) and in the interior of \(\Theta\), then under \(H_{0}\)
where \(r\) is the number of parameters constrained by the null. Rests on Definition 11.87, Definition 11.59, Definition 11.80, Theorem 11.67, Lemma 11.83 and Lemma 11.52.
Derivation, one parameter and a simple null. Derives Theorem 11.88. Take \(k=r=1\) and \(\Theta_{0}=\set{\theta_{0}}\). Expand \(\ell_{n}\) about the unconstrained MLE \(\hat{\theta}_{n}\), where \(\ell_{n}'\) vanishes:
with \(\bar{\theta}_{n}\) between the two. Multiply by \(-2\):
From Equation (11.82) (multiply by \(-2\) and rescale the deviation by \(\sqrt{n}\)). The first bracket converges in distribution to \(\mathcal{N}(0,I_{1}(\theta_{0})^{-1})\) by Theorem 11.67, and is therefore bounded in probability. The second factor converges in probability to \(I_{1}(\theta_{0})\): it is an average of independent identically distributed terms evaluated at \(\hat{\theta}_{n}\) rather than at \(\theta_{0}\), and the difference is controlled by the domination in (R5) together with the consistency of \(\hat{\theta}_{n}\). The last term is bounded in absolute value by \(\tfrac13\left(n^{-1}\sum_{i}M(X_{i})\right) n\abs{\hat{\theta}_{n}-\theta_{0}}^{3}\), and \(n\abs{\hat{\theta}_{n}-\theta_{0}}^{3} =n^{-1/2}\left[\sqrt{n}\abs{\hat{\theta}_{n}-\theta_{0}}\right]^{3} \rightarrow0\) in probability. By Lemma 11.52,
and \(\sqrt{I_{1}}\,W\) is standard Gaussian, so its square is \(\chi^{2}_{1}\) by Definition 11.80.
∎Derivation, \(k\) parameters and \(r\) constraints. Derives Theorem 11.88. The same expansion in \(k\) dimensions gives, for any \(\theta\) near \(\hat{\theta}_{n}\),
the Hessian term being handled exactly as above and the cubic remainder again vanishing. Since \(\hat{\theta}_{0}\) maximises \(\ell_{n}\) over \(\Theta_{0}\), it minimises \(Q\) there, and \(q=Q(\hat{\theta}_{0})=\min_{\theta\in\Theta_{0}}Q(\theta)\). Near \(\theta_{0}\) the submanifold \(\Theta_{0}\) is its tangent space to the accuracy the expansion retains, \(\Theta_{0}\simeq\theta_{0}+E\) with \(E\) that tangent space, of dimension \(\dim E=k-r\). Writing \(\vect{W}=\sqrt{n}\left(\hat{\theta}_{n}-\theta_{0}\right)\) and \(I=I_{1}(\theta_{0})\),
Substitute \(\vect{V}=I^{1/2}\vect{W}\) and \(S=I^{1/2}E\); the right side is \(\min_{\vect{s}\in S}\norm{\vect{V}-\vect{s}}^{2} =\norm{P\vect{V}}^{2}\) with \(P\) the orthogonal projector onto \(S^{\perp}\), of dimension \(r\). By Theorem 11.67 \(\vect{W}\rightarrow\mathcal{N}(0,I^{-1})\), so \(\vect{V}\) is asymptotically Gaussian with covariance \(I^{1/2}I^{-1}I^{1/2}=\identity\), and Lemma 11.83 gives \(\norm{P\vect{V}}^{2}\sim\chi^{2}_{r}\).
∎Two steps of that derivation are argued here rather than proved: the remainder in the \(k\)-dimensional Taylor expansion must be shown uniform on a neighbourhood of the true parameter shrinking at rate \(n^{-1/2}\), and both maximum-likelihood estimators must be shown to lie in that neighbourhood with probability tending to one, which is what licenses replacing the constraint surface by its tangent space at Equation (11.85). Both are supplied in the appendix, from the same regularity conditions (R1)–(R5) assumed here.
Full derivation in Appendix A.
Derives Theorem 11.88.
For a single parameter of interest \(\mu\) with any number \(m\) of nuisance parameters, the profile statistic \(q(\mu_{0})=-2\ln\lambda(\mu_{0})\) of Definition 11.70 is asymptotically \(\chi^{2}_{1}\) under \(H_{0}:\mu=\mu_{0}\), whatever \(m\) is. Rests on Definition 11.70 and Theorem 11.88.
Derivation. Derives Corollary 11.89. The null constrains one coordinate and leaves the other \(m\) free, so \(\Theta_{0}\) is a submanifold of codimension \(r=1\) and Theorem 11.88 applies with \(r=1\). The nuisance parameters affect the variance of \(\hat{\mu}\), through the Schur complement Equation (11.67), but not the number of degrees of freedom.
∎This is the result a collider search relies on: however many nuisance parameters describe the background and the detector, the significance is read off a one-degree-of-freedom distribution [Cowan:2011]. It is worth saying exactly when it stops being true.
Condition (R3) requires \(\theta_{0}\) to be interior. A signal strength constrained to \(\mu\ge0\) and tested at \(\mu=0\) violates it, and the distribution changes. In the quadratic regime let \(\hat{\mu}_{\text{u}}\) be the unconstrained maximiser, asymptotically Gaussian with mean zero and variance \(\sigma^{2}\) under \(H_{0}\); the maximiser over \(\mu\ge0\) is \(\hat{\mu}=\max(\hat{\mu}_{\text{u}},0)\), because the quadratic log-likelihood opens downward and, with its vertex at a negative abscissa, is decreasing throughout \([0,\infty)\). Then
a half-and-half mixture of a point mass at zero and a chi-squared law, since \(\hat{\mu}_{\text{u}}\le0\) has asymptotic probability \(\tfrac12\) by symmetry. Hence \(\Pr(q\ge u)=\tfrac12\Pr(\chi^{2}_{1}\ge u)\) for \(u>0\): using the uncorrected \(\chi^{2}_{1}\) tail overstates the \(p\)-value by exactly a factor of two, and therefore understates the significance. The boundary correction and the half-and-half mixture are Chernoff's [Chernoff:1954]; the same mixture is used deliberately in Section 11.5.5, and the convention a search actually quotes is fixed in [Cowan:2011]. Rests on Definition 11.59 and Corollary 11.89.
Condition (R2) can fail on \(\Theta_{0}\) itself. Consider a model \(f(x;\mu,m)\) in which \(\mu\) is the strength of a signal and \(m\) its mass. At \(\mu=0\) the density does not depend on \(m\) at all: every value of \(m\) gives the same distribution, the parameter is unidentified, the information matrix is singular in the \(m\) direction, and the expansion behind Theorem 11.88 has nothing to expand in. The statistic \(q(m)\) is then not a single number but a random field over the mass range, and the quantity actually reported — \(\sup_{m}q(m)\) — has no chi-squared distribution and is not covered by any theorem above. This is the look-elsewhere problem, and Section 11.6 treats it. The standard analysis of the unidentified-parameter case is R. B. Davies' [Davies:1977], extended in his sequel of ten years later [Davies:1987]; the two papers share a title and neither supersedes the other. He is not the E. B. Davies of the master-equation literature cited elsewhere in this treatise.
Significance
The significance \(Z\) corresponding to a \(p\)-value is defined by
with \(\Phi\) the standard Gaussian distribution function: \(Z\) is the number of standard deviations beyond which a one-sided Gaussian tail carries probability \(p\). Rests on Definition 11.75.
\(Z\) is a pure number and carries no unit. The thresholds conventionally called evidence and discovery in particle physics are \(Z=3\) and \(Z=5\); the corresponding tail probabilities, computed directly from Equation (11.87), are \(p=1.35\times 10^{-3}\) and \(p=2.87\times 10^{-7}\). The statistical practice of the field, these conventions among it, is reviewed in [Navas:2024], and the profile-likelihood asymptotics that turn a measured statistic into one of these numbers are collected in [Cowan:2011].
Let \(\mu\) be a signal strength with \(H_{0}:\mu=0\), and define the one-sided discovery statistic
which refuses to count a deficit as evidence for a signal. Then under \(H_{0}\), asymptotically, \(\tilde{q}_{0}\sim\tfrac12\delta_{0}+\tfrac12\chi^{2}_{1}\), and for an observed value \(q>0\)
Rests on Definition 11.92, Corollary 11.89, Corollary 11.82 and Equation (11.86).
Derivation. Derives Proposition 11.93. The mixture is the one derived in Equation (11.86): writing \(q_{0}=\hat{\mu}^{2}/\sigma^{2}\) for the two-sided statistic, which is asymptotically \(\chi^{2}_{1}\) by Corollary 11.89, the event \(\tilde{q}_{0}\ge q\) with \(q>0\) is the event \(\set{\hat{\mu}>0}\cap\set{q_{0}\ge q}\). The sign of \(\hat{\mu}\) and the magnitude \(\abs{\hat{\mu}}\) are independent for a Gaussian of mean zero, and the sign is positive with probability \(\tfrac12\), so \(\Pr(\tilde{q}_{0}\ge q)=\tfrac12\Pr(\chi^{2}_{1}\ge q)\). Insert Equation (11.75) with \(z=\sqrt{q}\): the right-hand side is \(\tfrac12\cdot2\left(1-\Phi(\sqrt{q})\right)=1-\Phi(\sqrt{q})\). Comparing with Equation (11.87) gives \(Z=\sqrt{q}\).
∎The identity \(Z=\sqrt{q}\) is the reason a collider search can quote a significance without simulating anything, and it is the convention fixed, together with the asymptotic distributions it rests on, in [Cowan:2011]. The same reference supplies the companion device for the expected significance of a planned measurement: a single artificial “Asimov” data set, in which every observable is set to its expectation under the assumed signal, returns the median \(q\) of an ensemble of simulated experiments directly, so that the sensitivity of a search can be quoted without generating the ensemble at all. The treatise uses only the observed-data statement above; the Asimov construction is named here because the searches of Part XI — Quantum Field Theory and the Standard Model quote sensitivities computed with it.
The identity \(Z=\sqrt{q}\) is a property of the one-sided convention and of nothing else. Used with the two-sided statistic \(q_{0}\), which counts a downward fluctuation as evidence against \(\mu=0\) just as readily as an upward one, the correct \(p\)-value is \(\Pr(\chi^{2}_{1}\ge q)=2\left(1-\Phi(\sqrt{q})\right)\), exactly twice as large, and the significance is \(Z=\Phi^{-1}\!\left(1-2\left(1-\Phi(\sqrt{q})\right)\right)\), which is not \(\sqrt{q}\). At \(q=25\) the two conventions give \(p=2.87\times 10^{-7}\) with \(Z=5.00\), and \(p=5.73\times 10^{-7}\) with \(Z=4.86\). The difference of \(0.14\) in \(Z\) is small enough to look like a rounding disagreement and large enough to move a result across the conventional discovery threshold. State which convention is in force, once, and do not mix them.
An upper limit uses a statistic that is one-sided the other way — large values of \(\mu\) disfavoured — and the same asymptotic distributions apply [Cowan:2011]. The additional difficulty there is that a downward background fluctuation can exclude a signal the experiment had no sensitivity to; the standard remedy is to report the ratio of the two tail probabilities rather than the numerator alone [Read:2002]. That construction belongs with the searches that use it and is not developed here.
The look-elsewhere effect
Everything in Section 11.5 assumes the hypothesis was fixed before the data were examined. A search does not work that way. It scans a range of an unknown parameter — a mass, a lifetime, a frequency — computes a test statistic at every point of the range, and reports the largest value found. The probability that some point of a wide range fluctuates upward is far greater than the probability that a nominated point does, and the difference can be one or two units of significance. This section makes the correction quantitative.
Local and global $p$-values
Let \(M=[m_{1},m_{2}]\) be the scanned range and \(\tilde{q}(m)\) the one-sided statistic Equation (11.88) evaluated with the signal parameter fixed at \(m\). Let \(u=\max_{m\in M}\tilde{q}(m)\) be the observed maximum. The local \(p\)-value is the tail probability at a point fixed in advance,
by Equation (11.89); the global \(p\)-value is
and the trials factor is \(N_{\mathrm{tr}}=p_{\mathrm{glob}}/p_{\mathrm{loc}}\). Rests on Definition 11.75, Equation (11.88) and Equation (11.89).
\(p_{\mathrm{glob}}\ge p_{\mathrm{loc}}\), so \(N_{\mathrm{tr}}\ge1\). Rests on Definition 11.96.
Derivation. Derives Proposition 11.97. For any fixed \(m_{\mathrm{fix}}\in M\) the event \(\set{\tilde{q}(m_{\mathrm{fix}})\ge u}\) is contained in the event \(\set{\max_{m\in M}\tilde{q}(m)\ge u}\); a probability is monotone under inclusion.
∎Suppose the maximum over \(M\) is attained, to the accuracy required, at one of \(N\) points \(m_{1},\dots,m_{N}\). Then
If in addition the \(N\) tests were independent, then
Rests on Definition 11.96.
Derivation. Derives Proposition 11.98. Write \(A_{j}=\set{\tilde{q}(m_{j})\ge u}\). The first claim is Boole's inequality \(\Pr(\bigcup_{j}A_{j})\le\sum_{j}\Pr(A_{j})\), which follows by induction from \(\Pr(A\cup B)=\Pr(A)+\Pr(B)-\Pr(A\cap B)\le\Pr(A)+\Pr(B)\), together with the fact that each \(\Pr(A_{j})\) equals \(p_{\mathrm{loc}}\) because the statistic has the same null distribution at every \(m\). If the \(A_{j}\) are independent then so are their complements, so \(\Pr(\bigcap_{j}A_{j}^{c})=\left(1-p_{\mathrm{loc}}\right)^{N}\); expanding by the binomial theorem gives the series.
∎The customary choice is \(N\simeq\Delta m/\delta m\), the width of the scanned window divided by the experimental resolution, on the reasoning that two hypotheses separated by much less than the resolution are tested by essentially the same data and are not independent trials.
Equation (11.92) is a theorem and Equation (11.93) is not. The union bound holds unconditionally but is weak, and it becomes vacuous as the grid is refined — nothing stops \(N\) from being taken larger, and the bound then exceeds \(1\). The independence assumption is false by construction: \(\tilde{q}(m)\) and \(\tilde{q}(m')\) are computed from the same events and are strongly correlated when \(\abs{m-m'}\) is of order the resolution. The number \(N\) is therefore a judgement about a correlation length, not a quantity the model supplies, and the estimate it gives can err in either direction. The next subsection replaces the judgement by a measurable quantity.
Upcrossings
The honest object is the whole function \(m\mapsto\tilde{q}(m)\), a random field on \(M\) whose null distribution at each point is known but whose maximum is not. What controls the maximum is how often the field crosses a high level on its way up.
For a continuous path \(\tilde{q}(\,\cdot\,)\) on \(M\) and a level \(u\), an upcrossing is a point \(m\in(m_{1},m_{2})\) at which the path passes from below \(u\) to above it. Write \(N_{u}\) for the number of upcrossings in \(M\), a non-negative integer-valued random variable. Rests on Definition 11.96.
The bound below is Davies' [Davies:1977] [Davies:1987]: it is the result that replaces Wilks' theorem when a parameter — here the mass — is present only under the alternative, and it is what the trials-factor practice of a search is built on [Gross:2010].
If the paths of \(\tilde{q}\) are continuous, then for every level \(u\)
that is, \(p_{\mathrm{glob}}\le p_{\mathrm{loc}}+\avg{N_{u}}\). Rests on Definition 11.100, Lemma 11.50 and Theorem 7.23.
Derivation. Derives Theorem 11.101. Suppose \(\max_{m\in M}\tilde{q}(m)\ge u\) and that the left endpoint is below the level, \(\tilde{q}(m_{1})<u\). The path is continuous and reaches a value at least \(u\) somewhere in \(M\), so by the intermediate value theorem (Theorem 7.23) it attains the value \(u\) at some interior point, having come from below; that is an upcrossing, so \(N_{u}\ge1\). Hence
and subadditivity gives \(\Pr(\max\ge u)\le\Pr(\tilde{q}(m_{1})\ge u)+\Pr(N_{u}\ge1)\). Finally \(N_{u}\) is non-negative, so Markov's inequality (Lemma 11.50) with \(a=1\) gives \(\Pr(N_{u}\ge1)\le\avg{N_{u}}\).
∎The bound is exact mathematics and it is also practical, because it replaces an unknown — the distribution of a maximum in the far tail, which no feasible number of simulated experiments can sample — by a mean, which can be estimated at a low level where upcrossings are common. What makes that replacement useful is the way \(\avg{N_{u}}\) depends on the level.
The level dependence below is Gross and Vitells' [Gross:2010], and it descends from Rice's formula for the expected number of level crossings of a smooth Gaussian process. The memoir on random noise is usually cited whole, but the crossing count is specifically section 3.3 of its second instalment [Rice:1945]; the first [Rice:1944] carries the shot effect and the power spectra, and not the formula used here. That formula is derived in the appendix — by a Kac counting integral whose expectation factorises, because a field of constant variance is independent of its own derivative at each point — and everything built on it here follows from it.
Let \(X\) be a Gaussian process on \(M=[m_{1},m_{2}]\) with zero mean, unit variance at every point and continuously differentiable sample paths, and write \(\sigma_{1}(m)^{2}=\operatorname{var}X'(m)\), assumed finite and continuous in \(m\). Then the expected number of upcrossings of a level \(z\in\R\) by \(X\) on \(M\) is
Two features of Equation (11.95) are worth naming, because they are what the next theorem lives on. Constancy of the variance forces \(\operatorname{cov}\!\left(X(m),X'(m)\right) =\tfrac12\,\dd\operatorname{var}X(m)/\dd m=0\), so field and derivative are independent at each fixed \(m\); the level therefore enters only through the Gaussian density of \(X(m)\) at \(z\), the expected positive part of the derivative, \(\avg{X'(m)^{+}}=\sigma_{1}(m)/\sqrt{2\pi}\), being the same at every level. That is the whole reason the right-hand side factorises into a function of \(z\) times a function of the field. For a stationary field \(\sigma_{1}\) is a constant \(\sqrt{\lambda_{2}}\) and Equation (11.95) reduces to Rice's original expression \(\avg{N_{z}}=\left(L/2\pi\right)\sqrt{\lambda_{2}}\,\ee^{-z^{2}/2}\), with \(L\) the length of \(M\).
The appendix carries the derivation, and states there what it must assume: the smoothness hypotheses above in the sharper form the proof uses, the absence of tangential crossings, and the two Lebesgue-theory statements — Tonelli's theorem and dominated convergence — that license the interchange of the expectation with the counting integral and with the limit. Everything the chapter builds on the formula, including the exponential level dependence and the trials factor, is derived in full.
Full derivation in Appendix A.
Derives Theorem 11.102.
Let the one-sided statistic Equation (11.88) be, asymptotically, \(\tilde{q}(m)=X(m)^{2}\,\indic{X(m)>0}\) for a Gaussian field \(X\) satisfying the hypotheses of Theorem 11.102, indexed by the single scanned parameter \(m\). Then for levels \(u,u_{0}>0\) in the range where that asymptotic description holds,
Rests on Definition 11.100 and Theorem 11.102.
Derivation. Derives Theorem 11.103. Fix \(u>0\). An upcrossing of the level \(u\) by \(\tilde{q}\) is exactly an upcrossing of the level \(\sqrt{u}\) by \(X\). Where \(X\le0\) the statistic takes the value \(0\), which is below \(u\), so no crossing of \(u\) occurs there; and on the open set where \(X>0\) one has \(\tilde{q}=X^{2}\), a strictly increasing function of \(X\), so \(\tilde{q}\) passes from below \(u\) to above it precisely where \(X\) passes from below \(\sqrt{u}\) to above it. The two counts therefore agree path by path, \(N_{u}=N_{\sqrt{u}}(X)\), and taking expectations and applying Equation (11.95) at \(z=\sqrt{u}\),
where the integral does not depend on \(u\). Writing the same identity at \(u_{0}\) and dividing cancels the integral and leaves Equation (11.96).
∎Gross and Vitells extend the result to a scan over \(d\) parameters, where the right-hand side of Equation (11.96) acquires a factor \(\left(u/u_{0}\right)^{(d-1)/2}\) [Gross:2010]; the upcrossing count is there replaced by the expected Euler characteristic of the set on which the field exceeds the level, of which it is the case \(d=1\). That extension is reported, not derived here, and nothing in this treatise uses it: every application below is a scan over a single mass.
The procedure this licenses is the one searches actually use: simulate a manageable number of background-only experiments, count upcrossings of a low threshold \(u_{0}\) — \(u_{0}=1\), i.e. a one-sigma level, is a common choice — average to get \(\avg{N_{u_{0}}}\), and extrapolate with Equation (11.96) to the observed level. Nothing has to be simulated at \(5\times 10^{-7}\) probability.
With \(u=Z_{\mathrm{loc}}^{2}\) and \(Z_{\mathrm{loc}}\) large,
so the trials factor grows linearly with the local significance. Rests on Definition 11.96, Theorem 11.101 and Theorem 11.103.
Derivation. Derives Corollary 11.105. Divide Equation (11.94) by \(p_{\mathrm{loc}}\) and use Equations (11.90) and (11.96) for the first form. For the second, the Gaussian tail satisfies
so that \(1-\Phi(z)=\ee^{-z^{2}/2}\left(1+O(z^{-2})\right) /\left(z\sqrt{2\pi}\right)\). To see Equation (11.99), write \(J=\int_{z}^{\infty}\ee^{-t^{2}/2}\dd t\) and integrate by parts with \(\ee^{-t^{2}/2}=t^{-1}\cdot t\,\ee^{-t^{2}/2}\):
The subtracted integral is positive, giving the upper bound, and is at most \(z^{-2}J\), giving \(J(1+z^{-2})\ge\ee^{-z^{2}/2}/z\) and hence the lower bound. Now put \(z=\sqrt{u}=Z_{\mathrm{loc}}\) in the first form of Equation (11.98): the exponentials \(\ee^{-u/2}\) and \(\ee^{-z^{2}/2}\) cancel and what remains is \(\avg{N_{u_{0}}}\ee^{u_{0}/2}\,Z_{\mathrm{loc}}\sqrt{2\pi}\).
∎The same relation appears in the literature with \(\sqrt{\pi/2}\) in place of \(\sqrt{2\pi}\). Both are right, in their own conventions. There the local \(p\)-value is the two-sided tail \(\Pr(\chi^{2}_{1}\ge u)=2\left(1-\Phi(\sqrt{u})\right)\) and \(N_{u}\) counts upcrossings of the two-sided statistic \(q(m)\), which are twice as numerous because the underlying Gaussian field crosses \(+\sqrt{u}\) and \(-\sqrt{u}\) with equal frequency. Both the numerator and the denominator of \(N_{\mathrm{tr}}\) are therefore doubled, and the trials factor — a ratio — is the same number either way, as it must be.
What is fatal is to mix them — but the damage is not exactly a factor of two, because the Davies bound Equation (11.94) carries an additive \(1\) that no convention doubles. Write \(x=\avg{N_{u}}/p_{\mathrm{loc}}\) in the one-sided convention, so that the number a search quotes as its trials factor is \(N_{\mathrm{tr}}=1+x\), the right-hand side of Equation (11.98). A two-sided upcrossing count divided by a one-sided local \(p\)-value doubles \(x\) alone and returns
What the mistake doubles is the excess of the trials factor over one, not the trials factor itself. The overstatement is therefore strictly less than a factor of two, and reaches two only asymptotically, for \(N_{\mathrm{tr}}\gg1\). At the \(N_{\mathrm{tr}}\simeq52\) of Example 11.107 the mixed number is about \(104\), too large by \(1.98\) — close enough to two that the loose statement passes unnoticed. In the opposite regime it is nearly harmless: a genuine \(N_{\mathrm{tr}}=1.05\) is reported as \(1.10\), an overstatement of five per cent, not of a factor of two. The mistake is worth avoiding for its own sake; “exactly twice” is the wrong way to describe it.
A worked example: a bump search over a mass window
A search scans for a narrow resonance in an invariant-mass spectrum over the window
with a mass resolution \(\delta m=1.5\,\mathrm{GeV}/c^{2}=2.67\times 10^{-27}\,\mathrm{kg}\). The statistic \(\tilde{q}(m)\) of Equation (11.88) is evaluated across the window with the background normalisation and shape treated as nuisance parameters, so Corollary 11.89 applies at each fixed \(m\). Its largest value is \(u=25.0\), at some mass inside the window.
Local significance. By Equation (11.89), \(Z_{\mathrm{loc}}=\sqrt{u}=5.00\) and \(p_{\mathrm{loc}}=1-\Phi(5)=2.87\times 10^{-7}\). This is the significance the excess would have if the mass had been nominated in advance, and it is not the significance of the search.
Naive trials factor. Taking \(N=\Delta m/\delta m=40/1.5 \simeq27\) resolution elements and treating them as independent, Equation (11.93) gives \(p_{\mathrm{glob}}\simeq27\times2.87\times 10^{-7}=7.7\times 10^{-6}\), whence \(Z_{\mathrm{glob}}=\Phi^{-1}(1-p_{\mathrm{glob}})=4.32\).
Upcrossings. Suppose background-only simulations at the low threshold \(u_{0}=1.0\) yield an average of \(\avg{N_{u_{0}}}=2.4\) upcrossings per experiment across the window. Then Equation (11.96) extrapolates to
and Equation (11.94) gives
The closed form Equation (11.98) gives \(1+\sqrt{2\pi}\times2.4\times\ee^{0.5}\times5.00=51\), agreeing with the direct evaluation to about three per cent — the difference being the \(O(z^{-2})\) term dropped from the Gaussian tail. Rests on Corollary 11.89, Proposition 11.93, Proposition 11.98, Theorem 11.101 and Theorem 11.103.
| Treatment | $N_{\mathrm{tr}}$ | $p_{\mathrm{glob}}$ | $Z_{\mathrm{glob}}$ |
|---|---|---|---|
| Mass nominated in advance | $1$ | \(2.87\times 10^{-7}\) | \(5.00\) |
| Independent resolution elements, $N=27$ | $27$ | \(7.7\times 10^{-6}\) | \(4.32\) |
| Expected upcrossings, Equation (11.94) | $52$ | \(1.50\times 10^{-5}\) | \(4.17\) |
Three things in Table 11.1 are worth separating. First, the degradation is real and large in probability — a factor of \(52\) — and mild in significance, \(5.00\rightarrow4.17\); that compression is the exponential in Equation (11.87) at work and is why an excess that survives the correction at all tends to survive it comfortably. Second, the naive estimate is not conservative: it happened to be optimistic by a factor of about two here, and which way it errs depends on whether the resolution over-counts or under-counts the correlation length of the field. Third, Equation (11.103) is an upper bound on \(p_{\mathrm{glob}}\), so \(Z_{\mathrm{glob}}=4.17\) is a lower bound on the global significance — the honest direction for a discovery claim.
Nothing in the arithmetic depends on the mass window being the one chosen here, and nothing in it is specific to any particular resonance. What the example fixes is the discipline: a significance quoted for a search must state the range searched, and a significance quoted without it is a local number wearing a global name.