Experiment: The Higgs Boson Discovery

Contents
  1. Historical context and the prediction under test
  2. Apparatus
  3. Procedure
  4. Observations and data
  5. Interpretation
  6. What is not yet measured
  7. Primary references

Tests Phenomena 104.59 and 104.62. Assuming Proposition 104.26, Proposition 104.61 and Definition 11.70.

On 4 July 2012 the ATLAS [Aad:2012] and CMS [Chatrchyan:2012] collaborations independently reported a new boson near \(125\,\mathrm{GeV}\)/\(c^{2}\), each at five standard deviations, in the diphoton and four-lepton channels. This chapter is the experimental closure of Electroweak Unification and the Higgs Boson: the Brout–Englert–Higgs mechanism [Englert:1964] [Higgs:1964] [Guralnik:1964] had been established indirectly for decades by the masses of the \(W\) and \(Z\) [Arnison:1983a] [Arnison:1983b] and by the precision electroweak fits [Schael:2006], but the scalar itself — the one particle of the Standard Model whose mass the theory does not predict — remained unobserved, its allowed mass window closed from below by LEP [Barate:2003] and eroded by the Tevatron [Aaltonen:2012]. The chapter sits at the end of the Standard Model sequence because it completes its particle content; what it does not complete is the theory, and the final sections say so.

Two features of this experiment demand more space than usual. The first is statistics: a discovery claim built on a few hundred events over a large smooth background needs its test statistic, its look-elsewhere correction and its blinding protocol stated as carefully as its detector [Read:2002] [Cowan:2011]. All of that machinery is derived in Probability and Statistics, and Section 113.5.1 applies it to the published numbers rather than gesturing at “five sigma”. The second is that the measurement did not stop in 2012 — spin-parity [Aad:2013] [Chatrchyan:2013], couplings against mass [Aad:2016] [Aad:2022] [Tumasyan:2022] and the total width [Aad:2023] have all been measured since, while the Higgs self-coupling, which is what would actually test the shape of the potential, has not.

Notation 113.1 (Units and the quantities a collider quotes).

Everything here is in SI, in the conventions of Section 104.1: \(\hbar\) and \(c\) appear explicitly, a mass is written as an energy divided by \(c^{2}\), and the electroweak scale is the energy \(E_{v}=v\sqrt{\hbar c}=246.220\,\mathrm{GeV}\) of Equation (104.2). Three collider habits need a translation that is given once and then used silently.

Energy. \(1\,\mathrm{GeV}\) is \(1.602176634\times 10^{-10}\,\mathrm{J}\) exactly, since the elementary charge is exact in the 2019 SI [Mohr:2025]. A mass of \(125.20\,\mathrm{GeV}\)/\(c^{2}\) is therefore \(2.2319\times 10^{-25}\,\mathrm{kg}\).

Cross section. An area. The literature quotes barns (\(10^{-28}\,\mathrm{m}^{2}\)) with SI prefixes, so a picobarn is \(10^{-40}\,\mathrm{m}^{2}\) and a femtobarn is \(10^{-43}\,\mathrm{m}^{2}\). Every cross section below is given in \(\mathrm{m}^{2}\) with the customary figure named in words.

Luminosity. The instantaneous luminosity \(L\) is defined by \(\dd N/\dd t=L\sigma\) and carries \(/\mathrm{m}^{2}/\mathrm{s}\); its time integral carries \(/\mathrm{m}^{2}\). One inverse femtobarn is \(10^{43}\,/\mathrm{m}^{2}\), and the discovery datasets are a few of them.

The letter \(\mu\). Collider practice and field-theory practice collide on this symbol, and five distinct meanings appear in this chapter. To keep them apart: the mass parameter of the scalar potential is written \(\mu\) only inside Equation (113.1) and the curvature statement that follows from it, always with the combination \(\hbar\mu c\) carrying an energy; the signal strength, the ratio of an observed rate to the Standard Model prediction, is written \(\mu\) with a subscript naming the measurement where one is needed (\(\mu_{\text{on}}\), \(\mu_{\text{off}}\), \(\bar{\mu}\)), and the mean pile-up is \(\avg{\mu}\); the renormalization scale is written \(\mu_{R}\) throughout, never bare \(\mu\); and the muon appears only as a particle label, in \(\mu^{\pm}\), \(m_{\mu}\) and \(\kappa_{\mu}\). The SI prefix micro is a different glyph and appears only inside a unit, as in \(18.8\,\mu\mathrm{m}\).

Remark 113.2 (Reading this chapter against the collider literature).

Every experimental paper cited below sets \(\hbar=c=1\), quotes masses in \(\mathrm{GeV}\) with the \(c^{2}\) suppressed, cross sections in barns and luminosities in inverse femtobarns. The dictionary is the one fixed in Remark 104.2 together with the three conversions of Notation 113.1. No derivation in this chapter is carried out with \(\hbar\) or \(c\) suppressed; this remark exists so that a reader can put a published figure beside a formula here and see that they are the same statement.

Historical context and the prediction under test

The Brout–Englert–Higgs mechanism

The problem the mechanism solves is stated in Theorem 104.8: a mass term for a gauge field is not gauge invariant, so a gauge theory of the weak interaction predicts massless carriers, and the carriers are not massless. The obvious repair — write the mass term anyway — destroys renormalizability (Theorem 104.4). The mechanism is the non-obvious repair: leave the Lagrangian invariant and let the vacuum fail to be.

Three papers in 1964 said this independently. Englert and Brout [Englert:1964] computed the vector self-energy directly and found a pole at nonzero momentum; Higgs [Higgs:1964a] [Higgs:1964] gave the abelian model and, in the second paper, noted explicitly that a massive scalar excitation survives; Guralnik, Hagen and Kibble [Guralnik:1964] gave the operator argument for why Goldstone's theorem does not apply. Higgs later wrote the argument out in full [Higgs:1966], and Kibble extended it to a non-abelian group [Kibble:1967], which is the case the Standard Model needs. The physical precedent was already in the literature: Anderson [Anderson:1963] had pointed out that the plasmon in a superconductor is exactly a gauge field that has acquired a mass from a condensate, and that the would-be massless mode is absent there for the same reason. Section 104.3.1 develops the analogy through Nambu's treatment of superconductivity [Nambu:1960].

What has to be evaded is Goldstone's theorem [Goldstone:1961], proved as Theorem 104.17: a spontaneously broken continuous global symmetry produces one massless scalar for every broken generator. Massless scalars are not observed. The evasion, derived as Theorem 104.19 and in the non-abelian case as Theorem 104.22, is that when the broken symmetry is gauged the would-be Goldstone modes are not physical states at all: a gauge transformation removes them from the scalar sector, and they reappear as the longitudinal polarizations that a massless vector does not have and a massive one does.

The bookkeeping is worth doing explicitly, because it is what makes the prediction sharp.

Proposition 113.3 (Degrees of freedom before and after).

The electroweak scalar sector is one complex \(\SU(2)\) doublet, four real fields. Before breaking, the field content \(\SU(2)_{L}\times\U(1)_{Y}\) carries

\[ 4\ \text{(scalar)}+3\times2\ \text{(massless }W^{a}\text{)} +1\times2\ \text{(massless }B\text{)}=12 \]

propagating polarizations. After breaking it carries

\[ 1\ \text{(scalar)}+3\times3\ \text{(massive }W^{\pm},Z\text{)} +1\times2\ \text{(photon)}=12\ep \]

Exactly one physical scalar remains, and its existence is the falsifiable content of the mechanism. Rests on Proposition 26.34, Theorem 104.25 and Equation (104.23).

Derivation. Derives Proposition 113.3. A massless vector field in four spacetime dimensions has two polarizations and a massive one has three; the constraint analysis that states those counts is Remark 26.33 and Proposition 26.34 — where the Proca count is stated and its constraint analysis is itself recorded as owed — and the Proca mass term itself is Equation (104.5). A complex doublet is four real scalar fields. The vacuum Equation (104.29) is invariant under the electric-charge generator and under nothing else in the four-parameter group, so three of the four generators are broken (Theorem 104.25). Goldstone's theorem in its global form would give three massless scalars; gauged, each is instead removed by a gauge choice — the unitary gauge of Equation (104.23) — and supplies the missing longitudinal polarization of one vector. Three vectors gain one polarization each, so the vector sector gains three; the scalar sector loses three. The total is unchanged, which it must be, because a gauge choice cannot create or destroy a physical state. What is left over is the radial excitation of the doublet about the minimum, one real field, and the theory says its mass is \(\sqrt{2\hat{\lambda}}\,E_{v}\) by Equation (104.37).

The counting also fixes what the experiment must find. Not “a scalar”: exactly one, electrically neutral, and — once its mass is measured — with every coupling and every branching fraction already determined. A second scalar, a charged scalar, or a scalar with the wrong couplings would each falsify the minimal mechanism, and none would be an adjustment of it.

Remark 113.4 (What was already evidence before 2012).

The mechanism was not an untested hypothesis in 2012. Three of its consequences had been measured. The \(W\) and \(Z\) exist with the masses the theory predicts from \(E_{v}\) and the mixing angle (Section 104.6.2); the photon is exactly massless, which in Theorem 104.22 is a consequence of the vacuum's residual symmetry and not an input; and the \(\rho\) parameter of Equation (113.6) is unity, which is a statement about the representation of the breaking sector and not merely about the fact of breaking. What was missing was the scalar itself. The distinction matters: a reader who takes 2012 as the year the electroweak theory was confirmed has the history backwards. The theory was confirmed in 1973 and 1983; 2012 completed its particle content.

The one free parameter

The scalar potential is

\begin{equation}\tag{113.1} V\left(\Phi\right)=-\mu^{2}\abs{\Phi}^{2} +\lambda\abs{\Phi}^{4}\ec \end{equation}

with \(\Phi\) of dimension \(\mathrm{J}^{1/2}\,\mathrm{m}^{-1/2}\), \(\mu\) of dimension \(/\mathrm{m}\) and \(\lambda\) of dimension \(/\mathrm{J}/\mathrm{m}\), so that \(V\) is an energy density (Definition 104.24). Two parameters appear, and one combination of them is already measured.

Proposition 113.5 (One number is unknown, and it is the mass).

The vacuum expectation value is fixed by the Fermi constant,

\begin{equation}\tag{113.2} E_{v}=v\sqrt{\hbar c} =\left(\frac{\left(\hbar c\right)^{3}}{\sqrt{2}\,G_{F}} \right)^{1/2}=246.220\,\mathrm{GeV}\ec \end{equation}

and the scalar mass is

\begin{equation}\tag{113.3} m_{H}c^{2}=\sqrt{2\hat{\lambda}}\,E_{v}\ec\qquad \hat{\lambda}=\lambda\hbar c\ep \end{equation}

Since \(\hat{\lambda}\) is not predicted, \(m_{H}\) is the single unknown of the sector; every production cross section and every branching fraction is a function of it alone. Rests on Proposition 104.26, Corollary 104.29 and Proposition 104.61.

Derivation. Derives Proposition 113.5. Equation (113.2) is Proposition 104.26: matching the low-energy limit of \(W\) exchange to Fermi's contact theory cancels the gauge coupling and leaves \(E_{v}^{2}=(\hbar c)^{3}/\sqrt{2}G_{F}\), with \(G_{F}/(\hbar c)^{3}=1.1663787(6)\times 10^{-5}\,/\mathrm{GeV}^{2}\) from the positive-muon lifetime [Tishchenko:2013], that is \(G_{F}=1.43585\times 10^{-62}\,\mathrm{J}\,\mathrm{m}^{3}\). In SI, \(E_{v}=3.94488\times 10^{-8}\,\mathrm{J}\) and the associated length is \(\hbar c/E_{v}=8.0143\times 10^{-19}\,\mathrm{m}\). Equation (113.3) is Corollary 104.29, obtained by expanding Equation (113.1) about the minimum \(\abs{\Phi}=v/\sqrt{2}\) and reading off the coefficient of the quadratic term in the radial fluctuation.

The claim that everything else follows is the content of Proposition 104.61: the coupling of the scalar to any massive particle is that particle's mass term divided by \(v\), so the couplings contain no parameter beyond \(E_{v}\), which is already measured. Fixing \(m_{H}\) therefore fixes every partial width through Equation (104.67) and its gauge analogue, hence every branching fraction (Table 104.4); and it fixes every production cross section, because production is the same set of vertices read backwards, convolved with parton distributions measured elsewhere (Section 116.6.5).

The consequence for the experiment is structural and it is what makes a null result meaningful. A search that scans \(m_{H}\) is testing a one-parameter family of fully specified theories, not fitting a shape to data. At each candidate mass the predicted rate in every channel is a number, so the data either accommodate it or exclude it; there is nothing to adjust. And when a signal does appear, the same rigidity makes the discovery over-determined: the mass measured in the diphoton channel must agree with the mass measured in the four-lepton channel, the ratio of the two rates must be the predicted one, and the production mechanisms must appear in the predicted proportion. Two experiments each measuring two channels give eight numbers where the theory has one.

The wider accounting — how many free parameters the Standard Model has in total, and which measurement fixes each — is tabulated in The Free Parameters of Physics. The scalar mass is one entry of that table, and Section 113.5.2 is the entry that this chapter fills in.

Indirect constraints

Before the scalar was produced it was constrained, because it circulates in loops. The relation between \(M_{W}\), \(M_{Z}\), \(\alpha\) and \(G_{F}\) receives a radiative correction \(\Delta r\), defined by Equation (104.43) and evaluated from the measured masses in Remark 104.33, which finds \(\Delta r=0.036\) and identifies the scalar-mass logarithm as one of the pieces making it up. What matters here is that dependence's form: it is logarithmic,

\begin{equation}\tag{113.4} \frac{\partial\,\Delta r}{\partial\ln\left(m_{H}c^{2}\right)} =\frac{\hat{g}^{2}}{16\pi^{2}}\,c_{H}\ec \end{equation}

with \(\hat{g}\) the dimensionless \(\SU(2)\) coupling of Notation 104.1 and \(c_{H}\) a pure number of order unity that this book does not derive, whereas the dependence on the top mass is quadratic, \(\Delta r|_{t}\propto m_{t}^{2}\) by Equation (104.52). The two sensitivities have very different consequences and the difference is Veltman's screening theorem [Veltman:1977]: a heavy scalar decouples logarithmically, so precision data constrain it weakly, while a heavy top does not decouple at all.

Derivation pending.

The coefficient of the scalar-mass logarithm in the radiative correction \(\Delta r\), written parametrically in this subsection, together with the constant term that accompanies it — which is not a refinement at the measured mass but comparable to the logarithm itself. Obtaining both requires the one-loop \(W\) and \(Z\) self-energies with the scalar circulating, renormalized on shell. The electroweak chapter already records the same logarithm as owed, in the paragraph following its screening-theorem discussion; the two belong together in Appendix A. Nothing in this chapter is evaluated numerically from the coefficient.

Weak is not zero. The LEP and SLC programme measured \(M_{Z}\), the \(Z\) partial widths and a dozen asymmetries to a part in \(10^{3}\) or better [Schael:2006]; combined with \(M_{W}\), \(m_{t}\), \(\alpha\) and \(G_{F}\) the system is over-constrained, and the global fit before 2012 preferred a scalar mass of order \(100\,\mathrm{GeV}\)/\(c^{2}\) with an upper bound in the low hundreds [Baak:2014]. The particle was found at \(125\,\mathrm{GeV}\)/\(c^{2}\), inside that range.

What made the constraint credible in advance was that the same method had already worked once, on a harder case. The quadratic top sensitivity let the LEP data locate the top quark's mass before the Tevatron produced it, and the indirect value agreed with the direct measurement [Abachi:1995] [Abe:1995]. A method that has predicted one unseen particle correctly is evidence about the next one; a method that has not is an extrapolation.

Phenomenon 113.6 (The weak bosons are massive and the photon is not).

The carriers of the charged and neutral weak currents are massive,

\begin{equation}\tag{113.5} m_{Z}c^{2}=91.1880(20)\,\mathrm{GeV}\ec\qquad m_{W}c^{2}=80.369(13)\,\mathrm{GeV}\ec \end{equation}

that is \(m_{Z}=1.62557\times 10^{-25}\,\mathrm{kg}\) and \(m_{W}=1.43271\times 10^{-25}\,\mathrm{kg}\) [Navas:2024], while the photon, which belongs to the same electroweak gauge group, has no measurable mass at all [Navas:2024]. The measured masses moreover satisfy

\begin{equation}\tag{113.6} \rho=\frac{m_{W}^{2}}{m_{Z}^{2}\cos^{2}\theta_{W}}=1 \end{equation}

to within the radiative corrections — a coincidence that a general symmetry-breaking sector has no reason to produce. Rests on Equation (104.36), Theorem 104.41 and Proposition 104.43.

Derivation. Derives Phenomenon 113.6. Let the breaking be produced by a single complex scalar doublet with vacuum expectation value \(v\) (Electroweak Unification and the Higgs Boson). Expanding the square of its covariant derivative about the vacuum generates

\[ m_{W}c^{2}=\tfrac{1}{2}\hat{g}E_{v}\ec\qquad m_{Z}c^{2}=\tfrac{1}{2}\sqrt{\hat{g}^{2}+\hat{g}'^{2}}\,E_{v}\ec \]

by Equation (104.36), with \(\hat{g}\) and \(\hat{g}'\) the dimensionless \(\SU(2)\) and \(\U(1)\) couplings of Notation 104.1, while the orthogonal combination of neutral gauge fields — the photon — gets no mass term at all, because the vacuum is left invariant by the electric-charge generator. The mixing angle is defined by \(\cos\theta_{W}=\hat{g}/\sqrt{\hat{g}^{2}+\hat{g}'^{2}}\), so

\[ \frac{m_{W}c^{2}}{m_{Z}c^{2}\cos\theta_{W}} =\frac{\tfrac{1}{2}\hat{g}E_{v}} {\tfrac{1}{2}\sqrt{\hat{g}^{2}+\hat{g}'^{2}}\,E_{v}} \cdot\frac{\sqrt{\hat{g}^{2}+\hat{g}'^{2}}}{\hat{g}}=1\ec \]

which is Equation (113.6). The relation is a property of the doublet assignment and not of symmetry breaking in general: a scalar in a higher representation gives \(\rho\neq1\), as Theorem 104.41 makes precise. The measured \(\rho\) is therefore evidence about the representation content of the breaking sector, not merely evidence that something breaks the symmetry, and it is what made the search for a single neutral scalar, rather than for some richer sector, the right search to mount.

Numerically, the comparison can only be made in a scheme that defines the mixing angle independently of the mass ratio, and Remark 104.34 is where the schemes are separated. With the on-shell angle, \(\sin^{2}\theta_{W}:=1-M_{W}^{2}/M_{Z}^{2}\), Equation (113.6) is \(\rho=1\) identically by construction and there is nothing to test. Taking instead the \(\overline{\mathrm{MS}}\) angle at the \(Z\) scale, \(\hat{s}^{2}_{Z}=0.23129(4)\) [Navas:2024], and the masses of Equation (113.5),

\[ \rho=\frac{\left(80.369\,\mathrm{GeV}\right)^{2}} {\left(91.1880\,\mathrm{GeV}\right)^{2} \left(1-0.23129\right)}=1.0105\ec \]

a departure from unity of about one per cent — the size of the one-loop correction \(\Delta\rho=9.3\times 10^{-3}\) computed in Proposition 104.43. The relation is not exact and is not supposed to be; what is significant is that the deviation is of the size a loop produces and not of the size a wrong representation would produce.

LEP and Tevatron searches

Two machines searched before the LHC, in two entirely different ways, and between them they set the window the LHC had to cover.

LEP. An \(e^{+}e^{-}\) collider produces the scalar through \(e^{+}e^{-}\to Z^{*}\to ZH\), the same \(HZZ\) vertex read as a production vertex. The event is fully constrained — the initial state has known energy — so the search is a bump hunt in the recoil mass against an identified \(Z\), with essentially no combinatorial background. What limits it is not statistics but kinematics.

Proposition 113.7 (LEP's reach was set by its beam energy).

Associated production \(e^{+}e^{-}\to ZH\) at centre-of-mass energy \(\sqrt{s}\,c\) requires

\begin{equation}\tag{113.7} m_{H}c^{2}\le\sqrt{s}\,c-m_{Z}c^{2}\ec \end{equation}

so LEP's final energy of \(\sqrt{s}\,c=209\,\mathrm{GeV}\) admits at most \(m_{H}c^{2}=117.8\,\mathrm{GeV}\), and the useful reach is a few \(\mathrm{GeV}\) below that because the cross section vanishes at threshold. Rests on Equation (113.5).

Derivation. Derives Proposition 113.7. In the centre-of-mass frame the total energy is \(\sqrt{s}\,c\) and the final state is two particles at rest or better, so \(\sqrt{s}\,c\ge m_{Z}c^{2}+m_{H}c^{2}\), which is Equation (113.7). With \(m_{Z}c^{2}=91.1880(20)\,\mathrm{GeV}\) from Equation (113.5) and \(\sqrt{s}\,c=209\,\mathrm{GeV}\) the bound is \(117.8\,\mathrm{GeV}\). Near threshold the two-body phase space carries the momentum factor \(p^{*}\), and for this process the cross section falls as \(p^{*}\) rather than as \(p^{*3}\) because the \(S\)-wave is allowed; either way it vanishes, so the last few \(\mathrm{GeV}\) of the kinematic range are not usable. The published limit, \(m_{H}c^{2}>114.4\,\mathrm{GeV}\) at \(95\,\mathrm{\%}\) confidence [Barate:2003], sits about \(3.4\,\mathrm{GeV}\) below the kinematic edge, which is the size of that effect.

The LEP combination of ALEPH, DELPHI, L3 and OPAL left a mild excess at the very edge of its sensitivity, near \(115\,\mathrm{GeV}\)/\(c^{2}\), of less than two standard deviations in the four-experiment combination and driven largely by one experiment [Barate:2003]. It did not survive, and the way it was reported is worth recording: the collaborations quoted the limit and the excess together, with the excess described as compatible with a background fluctuation, rather than either suppressing it or promoting it. The combination used the modified frequentist \(CL_{s}\) construction [Read:2002], whose purpose is exactly to prevent a downward background fluctuation from excluding a signal the experiment had no sensitivity to (Remark 11.95).

Tevatron. A \(p\bar{p}\) collider at \(\sqrt{s}\,c=1.96\,\mathrm{TeV}\) cannot use the diphoton channel — the rate is too low for its integrated luminosity — so CDF and D0 attacked the largest branching fraction instead, \(H\to b\bar{b}\), made usable by demanding an associated \(W\) or \(Z\) to suppress the enormous QCD production of \(b\) quarks. The combined analysis of the full Tevatron dataset reported a broad excess in that channel with a maximum local significance of about three standard deviations near \(125\,\mathrm{GeV}\)/\(c^{2}\), falling to roughly two and a half after correcting for the mass range searched [Aaltonen:2012]. The same combination excluded a Standard Model scalar between about \(147\,\mathrm{GeV}\)/\(c^{2}\) and \(179\,\mathrm{GeV}\)/\(c^{2}\).

The Tevatron result is the one place in this story where the associated-production channel with the largest branching ratio was competitive, and it is instructive that it produced evidence rather than a discovery. The excess was broad — the \(b\bar{b}\) mass resolution is of order \(15\,\mathrm{GeV}\)/\(c^{2}\), ten times worse than the diphoton channel — so it constrained a rate over a wide window rather than locating a resonance.

The window entering 2012. Combining the LEP lower bound, the Tevatron exclusion and the LHC's own 2011 exclusions, a Standard Model scalar was confined to roughly \(115\text{–}130\,\mathrm{GeV}\)/\(c^{2}\), with a smaller surviving region at high mass that the electroweak fit already disfavoured. Inside that window the branching fractions of Table 104.4 are changing rapidly: \(WW^{*}\) and \(ZZ^{*}\) are climbing as the thresholds approach, \(b\bar{b}\) is falling, and \(\gamma\gamma\) passes through its maximum. It is the one mass region in which the two high-resolution channels are simultaneously accessible, which is why the LHC was able to settle the question there and would have had a much harder time \(20\,\mathrm{GeV}\)/\(c^{2}\) higher.

Apparatus

Three instruments are involved: one accelerator and two detectors. The accelerator's job is to deliver, in a few years, enough proton–proton collisions that a process occurring once in \(3\times10^{9}\) of them yields a few hundred events. The detectors' job is to recognise those events in real time, discard the rest before they can be written to disk, and measure the survivors well enough that a resonance \(1\,\mathrm{GeV}/c^{2}\) wide can be seen on a continuum background a hundred times larger.

ATLAS and CMS were built to different designs on purpose. Their magnets, their calorimeter technologies and their reconstruction strategies differ in almost every respect, so their systematic errors are largely independent. That is not redundancy: it is the only protection available against a mistake that a single collaboration could make and not detect, and it is why the phrase “two experiments independently” carries weight in Section 113.4.3 that “two analyses of the same data” would not.

The Large Hadron Collider

The machine [Evans:2008] occupies the tunnel dug for LEP, a ring of circumference \(26659\,\mathrm{m}\) lying between \(45\,\mathrm{m}\) and \(170\,\mathrm{m}\) below the surface and tilted by \(1.4\,^\circ\). Reusing the tunnel fixed the circumference and therefore fixed everything else: to reach a beam energy of several \(\mathrm{TeV}\) in a ring of that size the bending field must be far beyond what iron can supply, so the LHC is a superconducting machine, and because two counter-rotating proton beams need opposite bending fields it uses twin apertures inside one cryostat and one yoke rather than two separate rings.

Proposition 113.8 (The dipole field the ring demands).

A relativistic proton of energy \(E\) on a circular arc of radius \(\rho\) requires a field

\begin{equation}\tag{113.8} B=\frac{p}{e\rho}\simeq\frac{E}{ec\rho}\ec \end{equation}

\(p\) being the momentum and \(e\) the elementary charge. The LHC's \(1232\) main dipoles are each \(14.3\,\mathrm{m}\) long, so the total bending length is \(17618\,\mathrm{m}\) and \(\rho=2804\,\mathrm{m}\). The three beam energies that matter here then need

\begin{equation}\tag{113.9} B\left(3.5\,\mathrm{TeV}\right)=4.16\,\mathrm{T}\ec\quad B\left(4\,\mathrm{TeV}\right)=4.76\,\mathrm{T}\ec\quad B\left(7\,\mathrm{TeV}\right)=8.33\,\mathrm{T}\ep \end{equation}

Rests on Equations (40.4) and (40.15).

Derivation. Derives Proposition 113.8. The Lorentz force supplies the centripetal acceleration, \(evB=\gamma m_{p}v^{2}/\rho\), so \(B=\gamma m_{p}v/(e\rho)=p/(e\rho)\) exactly. At \(4\,\mathrm{TeV}\) the proton's Lorentz factor is \(\gamma=E/(m_{p}c^{2})=4000/0.938272=4263\), so \(v/c\) differs from unity by \(1/(2\gamma^{2})=2.8\times 10^{-8}\) and \(p=E/c\) to that accuracy. The bending radius is the total magnetic length divided by \(2\pi\): \(1232\times14.3\,\mathrm{m} =17617.6\,\mathrm{m}\), and \(17617.6\,\mathrm{m}/2\pi =2803.9\,\mathrm{m}\). Then

\[ p=\frac{4\times 10^{12}\,\mathrm{eV}\times 1.602176634\times 10^{-19}\,\mathrm{J}/\mathrm{eV}}{2.99792458\times 10^{8}\,\mathrm{m}/\mathrm{s}} =2.1377\times 10^{-15}\,\mathrm{kg}\,\mathrm{m}/\mathrm{s}\ec \]

and dividing by \(e\rho=1.602176634\times 10^{-19}\,\mathrm{C} \times2803.9\,\mathrm{m}=4.4923\times 10^{-16}\,\mathrm{C}\,\mathrm{m}\) gives \(4.759\,\mathrm{T}\). The other two entries follow by proportionality. An iron-dominated warm dipole saturates near \(2\,\mathrm{T}\); a \(7\,\mathrm{TeV}\) beam in such a ring would need a bending radius above \(11\,\mathrm{km}\) and a circumference near \(100\,\mathrm{km}\), which is the whole argument for superconductivity. The LHC dipoles use niobium–titanium cable operated at \(1.9\,\mathrm{K}\) in superfluid helium, below the \(2.17\,\mathrm{K}\) lambda point, because the critical field of that alloy at \(4.2\,\mathrm{K}\) would not reach \(8\,\mathrm{T}\).

The beams are bunched by \(400\,\mathrm{MHz}\) superconducting radio-frequency cavities, eight per beam delivering about \(2\,\mathrm{MV}\) each. A proton completes the ring at

\begin{equation}\tag{113.10} f_{\text{rev}}=\frac{c}{26659\,\mathrm{m}} =11245\,\mathrm{Hz}\ec \end{equation}

to within the \(2.8\times 10^{-8}\) by which \(v\) falls short of \(c\). In the 2011 and 2012 runs the bunches were spaced by \(50\,\mathrm{ns}\) — twice the design spacing — which puts consecutive bunches \(15.0\,\mathrm{m}\) apart along the beam and gives a crossing rate of \(20\,\mathrm{MHz}\) at each interaction point where bunches are present. With \(1374\) filled bunches in 2012 the actual crossing rate was \(1374\,f_{\text{rev}}=15.4\,\mathrm{MHz}\).

Remark 113.9 (The stored energy is the reason for the caution).

Each 2012 bunch carried about \(1.6\times 10^{11}\) protons, so a beam held \(2.20\times 10^{14}\) protons at \(4\,\mathrm{TeV}\), a stored kinetic energy of

\[ 2.20\times 10^{14}\times4\times 10^{12}\,\mathrm{eV} \times1.602176634\times 10^{-19}\,\mathrm{J}/\mathrm{eV} =1.41\times 10^{8}\,\mathrm{J}\ec \]

\(141\,\mathrm{MJ}\) per beam, comparable to the kinetic energy of a \(4\times 10^{5}\,\mathrm{kg}\) train at \(28\,\mathrm{m}/\mathrm{s}\), which is \(\tfrac{1}{2}\times4\times 10^{5}\,\mathrm{kg} \times(28\,\mathrm{m}/\mathrm{s})^{2}=1.6\times 10^{8}\,\mathrm{J}\). The magnet system stores an order of magnitude more. This is not a detector concern but it dictated the 2011–2012 operating energy: after the 2008 interconnect failure the machine ran at half the design field while the splices were surveyed, which is why the discovery datasets are at \(7\,\mathrm{TeV}\) and \(8\,\mathrm{TeV}\) rather than \(14\,\mathrm{TeV}\).

Proposition 113.10 (Luminosity and the pile-up it forces).

For head-on collisions of \(n_{b}\) bunch pairs with \(N\) protons each, Gaussian transverse profiles of common width \(\sigma\) at the interaction point,

\begin{equation}\tag{113.11} L=\frac{f_{\text{rev}}n_{b}N^{2}}{4\pi\sigma^{2}}\,F\ec\qquad \sigma=\sqrt{\frac{\beta^{*}\epsilon_{n}}{\gamma}}\ec \end{equation}

with \(F\le1\) a geometric reduction from the crossing angle, \(\beta^{*}\) the optical amplitude function at the collision point and \(\epsilon_{n}\) the normalized transverse emittance. The 2012 parameters give \(\sigma=18.8\,\mu\mathrm{m}\) and a peak luminosity of order \(8\times 10^{37}\,/\mathrm{m}^{2}/\mathrm{s}\). The mean number of inelastic collisions per bunch crossing is

\begin{equation}\tag{113.12} \avg{\mu}=\frac{L\sigma_{\text{inel}}}{f_{\text{rev}}n_{b}}\ec \end{equation}

which at that luminosity is about \(36\). Rests on Notation 113.1.

Derivation. Derives Proposition 113.10. The rate of any process is \(\dd N/\dd t=L\sigma\) by the definition of luminosity, and for two Gaussian bunches crossing head-on the overlap integral of two normalized transverse densities of width \(\sigma\) is \(1/(4\pi\sigma^{2})\); multiplying by the number of protons in each bunch and by the rate at which bunch pairs meet gives Equation (113.11). The transverse size follows from the definition of emittance: the geometric emittance is \(\epsilon=\epsilon_{n}/\gamma\) because adiabatic damping shrinks the transverse phase-space area in proportion to the momentum, and \(\sigma^{2}=\beta^{*}\epsilon\) is the definition of the amplitude function. With \(\epsilon_{n}=2.5\times 10^{-6}\,\mathrm{m}\), \(\gamma=4263\) and \(\beta^{*}=0.6\,\mathrm{m}\),

\[ \sigma=\sqrt{0.6\,\mathrm{m}\times \frac{2.5\times 10^{-6}\,\mathrm{m}}{4263}} =1.876\times 10^{-5}\,\mathrm{m}\ep \]

Then

\[ \frac{f_{\text{rev}}n_{b}N^{2}}{4\pi\sigma^{2}} =\frac{11245\,\mathrm{Hz}\times1374\times \left(1.6\times 10^{11}\right)^{2}} {4\pi\left(1.876\times 10^{-5}\,\mathrm{m}\right)^{2}} =8.9\times 10^{37}\,/\mathrm{m}^{2}/\mathrm{s}\ec \]

and the crossing angle of a few hundred microradians, needed to keep the two beams from colliding parasitically elsewhere in the common pipe, supplies \(F\approx0.85\); the measured 2012 peak was \(7.7\times 10^{37}\,/\mathrm{m}^{2}/\mathrm{s}\).

For Equation (113.12), the inelastic proton–proton cross section at \(\sqrt{s}\,c=8\,\mathrm{TeV}\) is about \(7.3\times 10^{-30}\,\mathrm{m}^{2}\), which the collider literature writes as \(73\) millibarns, so the inelastic rate at peak luminosity is \(7.7\times 10^{37}\,/\mathrm{m}^{2}/\mathrm{s}\times 7.3\times 10^{-30}\,\mathrm{m}^{2}=5.6\times 10^{8}\,/\mathrm{s}\). Dividing by the crossing rate \(1.54\times 10^{7}\,/\mathrm{s}\) gives \(36\) collisions per crossing at the start of a fill. Luminosity decays through a fill as the bunches are consumed and heated, so the luminosity-weighted average over the whole 2012 run was about \(21\), against about \(9\) in 2011.

Remark 113.11 (Why pile-up is the defining difficulty).

Twenty simultaneous inelastic collisions is not twenty times more data. All of them occur within \(50\,\mathrm{ns}\) — in fact within the same crossing — and within a luminous region whose longitudinal extent is \(\sigma_{z}\approx5\,\mathrm{cm}\), so they are not separated in time and barely in space. Three specific damages follow, and every element of the analyses in Section 113.3 is shaped by one of them.

First, the event of interest must be assigned to the right vertex out of twenty candidates, and for a final state of two photons there are no tracks from the hard process to point with. Proposition 113.14 shows what a wrong assignment costs.

Second, energy from the other collisions lands in the same calorimeter cells, so any isolation or energy measurement must be corrected for a pedestal that fluctuates event by event. This is in-time pile-up.

Third, the ionization signal in a liquid-argon calorimeter is triangular with a drift time near \(450\,\mathrm{ns}\), nine times the bunch spacing, so the shaped pulse also carries contributions from neighbouring crossings. This is out-of-time pile-up, and it is removed by bipolar shaping whose undershoot cancels the average of the preceding crossings — which works on the average and leaves a fluctuation.

The last machine property that matters is not a machine property at all but a property of the proton. The scalar is produced dominantly by gluon fusion, \(gg\to H\), through a top-quark loop, and that dominance is a statement about parton densities rather than about couplings: at \(\sqrt{s}\,c=8\,\mathrm{TeV}\) a central resonance of mass \(125\,\mathrm{GeV}/c^{2}\) is made from partons carrying momentum fractions near \(x=m_{H}c^{2}/(\sqrt{s}\,c)=0.0156\), and at that \(x\) the gluon distribution is an order of magnitude above any single quark distribution (Section 116.6.5). The distributions themselves were measured in electron–proton scattering (Experiment: Deep Inelastic Scattering) and are usable here only because collinear factorization (Theorem 102.53) makes them universal — a theorem stated there at leading order, with its all-orders proof recorded as owed. Without it, and without the deep inelastic programme that measured its inputs, no LHC cross section could be predicted at all.

ATLAS

ATLAS [Aad:2008] is \(44\,\mathrm{m}\) long and \(25\,\mathrm{m}\) in diameter and weighs about \(7\times 10^{6}\,\mathrm{kg}\). Its distinguishing choices are a modest solenoid around a compact tracker, a sampling electromagnetic calorimeter with unusually fine longitudinal and lateral segmentation, and a very large air-core toroid outside the calorimeters that measures muons independently of the tracker.

Tracking. A superconducting solenoid \(5.3\,\mathrm{m}\) long with a \(2.4\,\mathrm{m}\) bore produces an axial field of \(2\,\mathrm{T}\); it is deliberately thin, about \(0.66\) radiation lengths, because it sits in front of the electromagnetic calorimeter. Inside it are three subsystems covering \(\abs{\eta}<2.5\): a silicon pixel detector of three barrel layers at radii \(50.5\,\mathrm{mm}\), \(88.5\,\mathrm{mm}\) and \(122.5\,\mathrm{mm}\) with \(8.0\times 10^{7}\) channels of \(50\,\mu\mathrm{m}\times400\,\mu\mathrm{m}\); a silicon microstrip tracker of four double barrel layers out to \(514\,\mathrm{mm}\) with \(80\,\mu\mathrm{m}\) pitch; and a transition-radiation tracker of about \(3.5\times 10^{5}\) straw tubes of \(4\,\mathrm{mm}\) diameter, giving roughly \(36\) measurements per track and, from the transition radiation emitted in the interleaved radiator, an electron-versus-pion discriminant.

Electromagnetic calorimetry. Lead absorbers immersed in liquid argon at \(87.3\,\mathrm{K}\), folded into an accordion so that the readout can be brought out from the front and back with no cracks in azimuth. The barrel covers \(\abs{\eta}<1.475\) and the two endcaps \(1.375<\abs{\eta}<3.2\); the total depth is at least \(22\) radiation lengths in the barrel. The segmentation is the point of the design and is worth stating in full:

The design energy resolution is

\begin{equation}\tag{113.13} \frac{\sigma_{E}}{E}=\frac{10\,\mathrm{\%}} {\sqrt{E/1\,\mathrm{GeV}}}\oplus0.7\,\mathrm{\%}\ec \end{equation}

the first term being the sampling fluctuation of a lead–argon stack and the second the constant term set by calibration and mechanical uniformity.

The fine first sampling does two things that the discovery depended on. A neutral pion of \(50\,\mathrm{GeV}\) decays to two photons with a typical opening angle of \(2m_{\pi}c^{2}/E\approx5.4\times 10^{-3}\), which at the calorimeter face is a separation of a few millimetres — resolvable in strips of \(\Delta\eta=0.0031\) and invisible in cells of \(0.025\). And because the first and second samplings measure \(\eta\) at two different radii, the photon's direction can be reconstructed from the calorimeter alone, giving its origin along the beam to about \(15\,\mathrm{mm}\) with no help from the tracker. Proposition 113.14 explains why that number matters.

Hadronic calorimetry and muons. Steel and scintillating tile in the barrel, copper and liquid argon in the endcaps, and a copper–tungsten forward calorimeter reaching \(\abs{\eta}=4.9\), which is what makes a missing-transverse-momentum measurement possible. Outside everything sits the muon spectrometer: eight air-core superconducting coils forming a barrel toroid \(25.3\,\mathrm{m}\) long and \(20.1\,\mathrm{m}\) in outer diameter, plus two endcap toroids, giving a bending integral \(\int B\,\dd l\) between \(1.5\,\mathrm{T}\,\mathrm{m}\) and \(7.5\,\mathrm{T}\,\mathrm{m}\) depending on \(\eta\). Air-core means there is no iron to scatter the muon, so the spectrometer measures momentum on its own to about \(10\,\mathrm{\%}\) at \(1\,\mathrm{TeV}\); the price is a very large and mechanically demanding structure. Precision coordinates come from drift tubes of \(30\,\mathrm{mm}\) diameter filled with argon and carbon dioxide at \(3\times 10^{5}\,\mathrm{Pa}\) — three atmospheres — resolving about \(80\,\mu\mathrm{m}\) per tube.

Trigger. The crossing rate is \(20\,\mathrm{MHz}\) and the recording rate is a few hundred hertz, so the trigger must reject about \(5\times 10^{4}\) events for every one it keeps, in real time and without ever looking at an event twice. A hardware first level, using reduced-granularity calorimeter sums and dedicated muon chambers, made its decision within \(2.5\,\mu\mathrm{s}\) and reduced the rate to about \(70\,\mathrm{kHz}\); two software levels running on a farm of order \(2\times 10^{4}\) processor cores reduced it to about \(400\,\mathrm{Hz}\) recorded in 2012. The diphoton trigger required two electromagnetic clusters above roughly \(35\,\mathrm{GeV}\) and \(25\,\mathrm{GeV}\) of transverse energy. Any event not selected within \(2.5\,\mu\mathrm{s}\) is gone permanently, which is why the trigger menu is part of the physics analysis and not part of the plumbing.

CMS

CMS [Chatrchyan:2008] is \(21.6\,\mathrm{m}\) long, \(14.6\,\mathrm{m}\) in diameter and weighs about \(1.25\times 10^{7}\,\mathrm{kg}\) — half the length of ATLAS and nearly twice the mass. Every one of its choices is the opposite of the corresponding ATLAS choice, and the discovery is stronger for it.

The magnet. A single superconducting solenoid of \(6\,\mathrm{m}\) inner diameter and \(12.5\,\mathrm{m}\) length, operated at \(3.8\,\mathrm{T}\) with a stored energy of \(2.6\,\mathrm{GJ}\). Everything except the muon chambers sits inside it. The high field buys track momentum resolution in a small volume; the cost is that the coil is thick, so the electromagnetic calorimeter must be inside it, which in turn constrains its radius and therefore its cost per unit solid angle. The return flux is carried by an iron yoke of about \(1.8\,\mathrm{T}\) in which the muon chambers are embedded, so the muon system reuses the same magnetic circuit rather than needing one of its own.

Tracking. All silicon, about \(200\,\mathrm{m}^{2}\) of it, the largest such device built at the time. Three pixel barrel layers at radii \(4.4\,\mathrm{cm}\), \(7.3\,\mathrm{cm}\) and \(10.2\,\mathrm{cm}\) carry \(6.6\times 10^{7}\) pixels of \(100\,\mu\mathrm{m}\times150\,\mu\mathrm{m}\); ten microstrip barrel layers reach \(r=1.1\,\mathrm{m}\) with \(9.3\times 10^{6}\) strips. Coverage is \(\abs{\eta}<2.5\).

Electromagnetic calorimetry. This is the subsystem CMS built for the diphoton channel specifically, and the reasoning is explicit in the design report: a homogeneous crystal calorimeter has no sampling fluctuation, so its stochastic term can be several times smaller than a sampling device's, and the diphoton mass resolution is dominated by that term. The calorimeter is \(75\,848\) lead-tungstate (\(\text{PbWO}_{4}\)) scintillating crystals — \(61\,200\) in the barrel for \(\abs{\eta}<1.479\) and \(7324\) in each endcap. The material was chosen for density: \(8280\,\mathrm{kg}/\mathrm{m}^{3}\), a radiation length of \(0.89\,\mathrm{cm}\) and a Molière radius of \(2.2\,\mathrm{cm}\), so a barrel crystal \(230\,\mathrm{mm}\) long is \(25.8\) radiation lengths in a \(22\,\mathrm{mm}\times22\,\mathrm{mm}\) front face. The design resolution measured in a test beam is

\begin{equation}\tag{113.14} \frac{\sigma_{E}}{E}=\frac{2.8\,\mathrm{\%}} {\sqrt{E/1\,\mathrm{GeV}}} \oplus\frac{12\,\mathrm{\%}}{E/1\,\mathrm{GeV}} \oplus0.30\,\mathrm{\%}\ec \end{equation}

a stochastic term nearly four times smaller than Equation (113.13).

That advantage is bought at a price that had to be engineered away. Lead tungstate is a poor scintillator — about \(30\) photons per \(\mathrm{MeV}\) reach the photodetector — so the barrel is read out with avalanche photodiodes of gain \(50\) and the endcaps with vacuum phototriodes, both able to operate inside \(3.8\,\mathrm{T}\) where a photomultiplier cannot. Its light yield also falls by about \(2\,\mathrm{\%}\) per kelvin, so the whole calorimeter is held at \(18\,\mathrm{^\circ\mathrm{C}}\) stabilised to \(0.05\,\mathrm{K}\); and the crystals darken under irradiation and recover between fills, which forces a continuous transparency monitoring with injected laser light. A preshower of two lead converters and two silicon strip planes, three radiation lengths in all with a strip pitch of \(1.9\,\mathrm{mm}\), sits in front of the endcaps and does there what ATLAS's fine strips do in the barrel: separate a \(\pi^{0}\)'s two photons from one.

Reconstruction. CMS reconstructs events by particle flow: rather than reading each subdetector separately, the algorithm builds a list of individual particles by linking tracks to calorimeter clusters to muon segments, and then constructs jets, missing transverse momentum and isolation variables from that list. The method needs a tracker that measures nearly every charged particle and a calorimeter finely segmented enough to associate clusters with tracks, which is exactly what the \(3.8\,\mathrm{T}\) solenoid and the crystal array provide. Its relevance here is isolation: a photon candidate's isolation can be computed from the charged particles associated with the chosen vertex alone, which suppresses the pile-up dependence that a purely calorimetric isolation would carry.

Trigger. The same two-stage structure with different numbers: a hardware level with a \(3.2\,\mu\mathrm{s}\) latency reducing \(20\,\mathrm{MHz}\) to at most \(100\,\mathrm{kHz}\), and a software high-level trigger on a commercial farm reducing that to a few hundred hertz. CMS's diphoton triggers used lower transverse-energy thresholds than ATLAS's, near \(26\,\mathrm{GeV}\) and \(18\,\mathrm{GeV}\), compensated by isolation and shower-shape requirements applied online.

ParameterATLASCMS
Reference[Aad:2008][Chatrchyan:2008]
Length; diameter\(44\,\mathrm{m}\); \(25\,\mathrm{m}\)\(21.6\,\mathrm{m}\); \(14.6\,\mathrm{m}\)
Mass\(7\times 10^{6}\,\mathrm{kg}\)\(1.25\times 10^{7}\,\mathrm{kg}\)
Tracking field\(2\,\mathrm{T}\) solenoid\(3.8\,\mathrm{T}\) solenoid
Muon fieldair-core toroid, $\int B\,\dd l$ from \(1.5\text{–}7.5\,\mathrm{T}\,\mathrm{m}\)\(1.8\,\mathrm{T}\) iron return yoke
Silicon pixels$8.0\times 10^{7}$ of \(50\,\mu\mathrm{m}\)$\times$\(400\,\mu\mathrm{m}\)$6.6\times 10^{7}$ of \(100\,\mu\mathrm{m}\)$\times$\(150\,\mu\mathrm{m}\)
Outer tracking\(3.5\times 10^{5}\) straw tubes, \(4\,\mathrm{mm}\) diameter\(9.3\times 10^{6}\) silicon strips
Electromagnetic calorimeterlead/liquid argon accordion, sampling, $\ge22$ radiation lengths\(75848\) $\text{PbWO}_{4}$ crystals, homogeneous, $25.8$ radiation lengths
Finest lateral cell$\Delta\eta=0.0031$ (first sampling)\(22\,\mathrm{mm}\) crystal face; preshower pitch \(1.9\,\mathrm{mm}\)
Design $\sigma_{E}/E$$10\,\mathrm{\%}/\sqrt{E}\oplus0.7\,\mathrm{\%}$$2.8\,\mathrm{\%}/\sqrt{E}\oplus12\,\mathrm{\%}/E \oplus0.30\,\mathrm{\%}$
Level-1 latency; output\(2.5\,\mu\mathrm{s}\); $\sim70\,\mathrm{kHz}$\(3.2\,\mu\mathrm{s}\); $\le100\,\mathrm{kHz}$
Recorded rate, 2012$\sim400\,\mathrm{Hz}$$\sim400\,\mathrm{Hz}$
Design and construction parameters of the two detectors and of the machine that fed them, as built and as operated in 2011–2012. These are settings and specifications, not measurements, so no uncertainty is attached to them; the measured quantities of this chapter are collected in Table 113.3 and Table 113.5. Energy resolutions are the design or test-beam parametrizations of Equation (113.13) and Equation (113.14), with $E$ in $\mathrm{GeV}$; the resolutions achieved in situ are worse and are discussed in Section 113.4.1.
Machine parameter2011 run2012 run
Circumference\(26659\,\mathrm{m}\)\(26659\,\mathrm{m}\)
Main dipoles$1232\times14.3\,\mathrm{m}$, NbTi at \(1.9\,\mathrm{K}\)as 2011
Bending radius (derived)\(2804\,\mathrm{m}\)\(2804\,\mathrm{m}\)
Beam energy\(3.5\,\mathrm{TeV}\)\(4\,\mathrm{TeV}\)
Centre-of-mass energy $\sqrt{s}\,c$\(7\,\mathrm{TeV}\)\(8\,\mathrm{TeV}\)
Dipole field (derived)\(4.16\,\mathrm{T}\)\(4.76\,\mathrm{T}\)
Bunch spacing\(50\,\mathrm{ns}\) $=15.0\,\mathrm{m}$as 2011
Filled bunches per beamup to \(1380\)up to \(1374\)
Protons per bunch$\approx1.45\times 10^{11}$$\approx1.6\times 10^{11}$
Stored energy per beam (derived)\(1.1\times 10^{8}\,\mathrm{J}\)\(1.4\times 10^{8}\,\mathrm{J}\)
Revolution frequency\(11245\,\mathrm{Hz}\)\(11245\,\mathrm{Hz}\)
Peak luminosity\(3.7\times 10^{37}\,/\mathrm{m}^{2}/\mathrm{s}\)\(7.7\times 10^{37}\,/\mathrm{m}^{2}/\mathrm{s}\)
Mean pile-up$\approx9$$\approx21$
The machine as operated in the two discovery runs [Evans:2008]. These are settings, not measurements, and carry no uncertainty; the two derived rows are computed in Proposition 113.8 and Proposition 113.10 from the geometry and the beam parameters in the rows above them. The mean pile-up is a luminosity-weighted average over the run and is about \(1.7\) times smaller than the value at the start of a fill, because the luminosity decays as the bunches are consumed.

Procedure

Production and decay channels

Four production mechanisms contribute at the LHC, and all four are the vertices of Equation (104.69) read as production rather than decay. In descending order of rate at \(m_{H}c^{2}=125\,\mathrm{GeV}\) and \(\sqrt{s}\,c=8\,\mathrm{TeV}\):

The sum is \(2.21\times 10^{-39}\,\mathrm{m}^{2}\) at \(8\,\mathrm{TeV}\) and \(1.74\times 10^{-39}\,\mathrm{m}^{2}\) at \(7\,\mathrm{TeV}\) — in the customary units \(22.1\) and \(17.4\) picobarns [Aad:2012] [Chatrchyan:2012]. Set against the inelastic proton–proton cross section of \(7.3\times 10^{-30}\,\mathrm{m}^{2}\), the scalar is produced in \(3.0\times 10^{-10}\) of all collisions: one in three thousand million.

Derivation pending.

The Standard Model production cross sections quoted in this section. Obtaining them requires the partonic amplitudes — above all the top-quark triangle that drives gluon fusion, which has no tree-level counterpart — convolved with the measured parton distributions and corrected to next-to-next-to-leading order in the strong coupling, where the corrections are large (a \(K\) factor near two for gluon fusion). The loop amplitude and the resulting cross sections are already recorded as owed in the electroweak chapter's pending derivation of the loop-induced scalar vertices; the convolution and the perturbative corrections belong with them in Appendix A. What is derived in this chapter from those inputs is everything downstream of them: the produced event counts, the channel comparison, the mass resolution and the significance.

Proposition 113.12 (Sensitivity is set by ratio, not by rate).

In the Gaussian limit the significance of a resonance search in a channel with \(s\) expected signal and \(b\) expected background events in the mass window is \(s/\sqrt{b}\), not \(s\). For the ATLAS discovery dataset, the three candidate channels give

\begin{equation}\tag{113.15} \frac{s}{\sqrt{b}}\Big|_{b\bar{b}}\sim7\times 10^{-2}\ec\qquad \frac{s}{\sqrt{b}}\Big|_{\gamma\gamma}\sim2\ec\qquad \frac{s}{\sqrt{b}}\Big|_{4\ell}\sim4\ec \end{equation}

although the produced event counts are in the ratio \(1.2\times 10^{5}:4.9\times 10^{2}:25\). The channel with the largest branching fraction is the least useful of the three. Rests on Propositions 11.93 and 113.5.

Derivation. Derives Proposition 113.12. Take the ATLAS dataset: \(4.8\times 10^{43}\,/\mathrm{m}^{2}\) at \(7\,\mathrm{TeV}\) and \(5.8\times 10^{43}\,/\mathrm{m}^{2}\) at \(8\,\mathrm{TeV}\), that is \(4.8\) and \(5.8\) inverse femtobarns. The number of scalars produced is the product of integrated luminosity and cross section, summed over the two runs:

\[ N_{H}=4.8\times 10^{43}\,/\mathrm{m}^{2} \times1.74\times 10^{-39}\,\mathrm{m}^{2} +5.8\times 10^{43}\,/\mathrm{m}^{2} \times2.21\times 10^{-39}\,\mathrm{m}^{2} =2.12\times 10^{5}\ep \]

Multiplying by the branching fractions of Table 104.4 gives \(1.2\times 10^{5}\) decays to \(b\bar{b}\), \(487\) to \(\gamma\gamma\), and — using \(\mathcal{B}(H\to ZZ^{*})=0.026\) together with \(\mathcal{B}(Z\to e^{+}e^{-})+\mathcal{B}(Z\to\mu^{+}\mu^{-}) =0.0673\), so that \(\mathcal{B}(H\to ZZ^{*}\to4\ell)=0.026\times 0.0673^{2}=1.18\times 10^{-4}\) — \(25\) decays to four charged leptons of the first two generations.

Now the backgrounds.

Bottom pairs. The inclusive \(b\bar{b}\) production cross section at these energies is of order \(2.8\times 10^{-32}\,\mathrm{m}^{2}\), which the collider literature writes as \(280\) microbarns — seven orders of magnitude above the rate at which scalars are produced and decay to \(b\bar{b}\), and four thousandths of the inelastic proton–proton cross section quoted above, which it must be. The same luminosity therefore yields \(1.06\times 10^{44}\,/\mathrm{m}^{2}\times 2.8\times 10^{-32}\,\mathrm{m}^{2}=3\times 10^{12}\) background events. Even at \(100\,\mathrm{\%}\) efficiency, \(s/\sqrt{b}=1.2\times 10^{5}/ \sqrt{3\times 10^{12}}=0.07\). Requiring an associated \(W\) or \(Z\), as the Tevatron did, cuts the background by orders of magnitude but cuts the signal to the associated-production cross section, a twentieth of the total; that trade is what produced three standard deviations at the Tevatron and not five.

Two photons. Photon reconstruction and the selection of Section 113.3.2 retain roughly \(40\,\mathrm{\%}\) of the produced decays, so \(s\approx195\). The background is continuum production of two photons, plus jets misidentified as photons; in a window of a few \(\mathrm{GeV}\) around \(125\,\mathrm{GeV}\) it is of order \(10^{4}\) events in this dataset. Hence \(s/\sqrt{b}\approx195/100\approx2\), and a signal fraction \(s/b\approx2\,\mathrm{\%}\).

Four leptons. Acceptance and lepton efficiency retain roughly \(35\,\mathrm{\%}\), so \(s\approx8.7\); the background in a window of a few \(\mathrm{GeV}\), dominated by the irreducible continuum \(q\bar{q}\to ZZ^{*}\), is of order \(5\) events. Hence \(s/\sqrt{b}\approx4\) — and here \(s/b>1\), so the “background” is not the thing being fought.

The three numbers are Equation (113.15). Two lessons follow, and both shaped the analyses. A large branching fraction is worth nothing if the final state is one the strong interaction produces copiously by itself, which disqualifies \(b\bar{b}\), \(c\bar{c}\) and \(gg\) — three of the four largest entries in Table 104.4. And the two surviving channels are useful for opposite reasons: \(\gamma\gamma\) has many events and a poor signal fraction, \(4\ell\) has few events and an excellent one, so their systematic weaknesses do not overlap either.

Remark 113.13 (The discovery channel has no tree-level vertex).

The scalar is electrically neutral and colourless, so \(H\to\gamma\gamma\) and \(gg\to H\) do not exist at tree level. Both proceed through loops of the heaviest charged or coloured particles: gluon fusion through the top-quark triangle, and the diphoton decay through the top triangle and the \(W\) loop, which interfere destructively with the \(W\) dominating. The discovery therefore rested on a one-loop amplitude at both ends — the dominant production mechanism and the most significant decay channel are each loop-induced — which is an unusual situation and a strong one: any heavy charged or coloured particle coupling to the scalar would change these rates whether or not it can be produced directly. The amplitudes themselves are recorded as a pending derivation in Section 104.7.1 and are not repeated here.

Event selection and background estimation

Photons. A photon candidate is an electromagnetic cluster with no matched track (or, for a converted photon, a matched conversion vertex) whose lateral and longitudinal shower shape is consistent with a single electromagnetic shower rather than with two overlapping ones or with a jet fluctuating to a leading \(\pi^{0}\). About a quarter of photons convert in the tracker material before reaching the calorimeter, which is why both experiments treat converted and unconverted candidates separately. Isolation — the transverse energy in a cone around the candidate, corrected for pile-up — rejects photons produced inside jets. ATLAS required two candidates with transverse energies above \(40\,\mathrm{GeV}\) and \(30\,\mathrm{GeV}\) within \(\abs{\eta}<2.37\) excluding the barrel to endcap transition; CMS scaled its thresholds with the candidate mass, requiring \(E_{T}>m_{\gamma\gamma}c^{2}/3\) and \(m_{\gamma\gamma}c^{2}/4\), which keeps the selection from sculpting a peak into the background.

Leptons. Electrons are clusters with a matched track; muons are tracks matched between the inner detector and the muon system. The four-lepton selection requires two same-flavour opposite-charge pairs, with the pair closer to the \(Z\) mass required to be near it and the second allowed to be far off shell — which is the whole point, since at \(m_{H}c^{2}=125\,\mathrm{GeV}\) the decay \(H\to ZZ\) is below threshold by \(57\,\mathrm{GeV}\) and one boson must be virtual. Lepton transverse-momentum thresholds are low, of order \(7\,\mathrm{GeV}/c\) for the softest, because the four leptons share \(125\,\mathrm{GeV}\) between them.

The measurement that everything else serves is the invariant mass, and its resolution is where the two detector designs cash out.

Proposition 113.14 (What limits the diphoton mass resolution).

For two photons of energies \(E_{1}\), \(E_{2}\) separated by an angle \(\theta\),

\begin{equation}\tag{113.16} \left(m_{\gamma\gamma}c^{2}\right)^{2} =2E_{1}E_{2}\left(1-\cos\theta\right)\ec \end{equation}

and the fractional resolution is

\begin{equation}\tag{113.17} \frac{\sigma_{m}}{m}=\frac{1}{2} \left[\left(\frac{\sigma_{E_{1}}}{E_{1}}\right)^{2} +\left(\frac{\sigma_{E_{2}}}{E_{2}}\right)^{2} +\left(\frac{\abs{\vect{p}_{H}}c}{m_{H}c^{2}}\right)^{2} \sigma_{\theta}^{2}\right]^{1/2}\ec \end{equation}

the angular term being evaluated for the symmetric decay. It vanishes for a scalar produced at rest and grows with the scalar's momentum; with a misassigned collision vertex it dominates. Rests on Equations (113.13) and (113.14).

Derivation. Derives Proposition 113.14. For two massless particles \(\abs{\vect{p}_{i}}c=E_{i}\), so the squared invariant mass of the pair is

\[ \left(m c^{2}\right)^{2} =\left(E_{1}+E_{2}\right)^{2} -\left(\vect{p}_{1}+\vect{p}_{2}\right)^{2}c^{2} =2E_{1}E_{2}-2\vect{p}_{1}\cdot\vect{p}_{2}c^{2} =2E_{1}E_{2}\left(1-\cos\theta\right)\ec \]

which is Equation (113.16). Taking the logarithm and differentiating,

\[ 2\,\frac{\dd m}{m}=\frac{\dd E_{1}}{E_{1}} +\frac{\dd E_{2}}{E_{2}} +\frac{\sin\theta}{1-\cos\theta}\dd\theta =\frac{\dd E_{1}}{E_{1}}+\frac{\dd E_{2}}{E_{2}} +\cot\!\left(\frac{\theta}{2}\right)\dd\theta\ec \]

using \(\sin\theta/(1-\cos\theta)=\cot(\theta/2)\). Adding the three independent errors in quadrature gives Equation (113.17) once \(\cot(\theta/2)\) is expressed kinematically. For the symmetric decay of a parent of energy \(E_{H}\), \(E_{1}=E_{2}=E_{H}/2\), so Equation (113.16) reads \((m_{H}c^{2})^{2}=E_{H}^{2}\sin^{2}(\theta/2)\), whence \(\sin(\theta/2)=m_{H}c^{2}/E_{H}\) and

\[ \cot\!\left(\frac{\theta}{2}\right) =\frac{\sqrt{E_{H}^{2}-\left(m_{H}c^{2}\right)^{2}}}{m_{H}c^{2}} =\frac{\abs{\vect{p}_{H}}c}{m_{H}c^{2}}\ep \]

So the angular sensitivity is exactly the parent's momentum in units of its mass: a scalar at rest decays back to back, \(\theta=\pi\), and its mass is then \(E_{1}+E_{2}\) with no angular information required at all.

Now the numbers, at a typical photon energy of \(60\,\mathrm{GeV}\) and a typical scalar energy of \(200\,\mathrm{GeV}\), for which \(\abs{\vect{p}_{H}}c=156\,\mathrm{GeV}\) and \(\cot(\theta/2)=1.25\).

Energy term. ATLAS's parametrization Equation (113.13) gives \(\sigma_{E}/E=10\,\mathrm{\%}/\sqrt{60}\oplus0.7\,\mathrm{\%} =1.47\,\mathrm{\%}\), so the two energy terms contribute \(1.47\,\mathrm{\%}/\sqrt{2}=1.04\,\mathrm{\%}\), that is \(1.30\,\mathrm{GeV}/c^{2}\). CMS's Equation (113.14) gives \(2.8\,\mathrm{\%}/\sqrt{60}\oplus12\,\mathrm{\%}/60 \oplus0.30\,\mathrm{\%}=0.51\,\mathrm{\%}\) and hence \(0.45\,\mathrm{GeV}/c^{2}\) — a factor of three better, which is the crystal calorimeter's entire justification.

Angular term with the wrong vertex. A photon leaves no track, so the collision vertex must be inferred. If it is taken at random from the twenty available, the error is the longitudinal spread of the luminous region, \(\sigma_{z}\approx5\,\mathrm{cm}\). At a calorimeter radius of \(1.5\,\mathrm{m}\) this is an angular error \(\sigma_{\theta}\approx0.05\,\mathrm{m}/1.5\,\mathrm{m} =0.033\), and \(\tfrac{1}{2}\times1.25\times0.033=2.1\,\mathrm{\%}\), that is \(2.6\,\mathrm{GeV}/c^{2}\). This is twice ATLAS's energy term and six times CMS's: with no vertex information the diphoton channel would not have worked in either detector.

Angular term with the right one. ATLAS's calorimeter pointing — the same shower measured in two samplings at different radii — locates the photon's origin to about \(15\,\mathrm{mm}\) without any tracker information, giving \(\sigma_{\theta}\approx0.010\) and a contribution of \(0.63\,\mathrm{\%}\), or \(0.79\,\mathrm{GeV}/c^{2}\): now subdominant. CMS instead identifies the vertex from the charged tracks recoiling against the diphoton system, and from the tracks of converted photons, reaching a comparable accuracy in the large majority of events.

Combining, the predicted resolutions are \(1.5\,\mathrm{GeV}/c^{2}\) for ATLAS and \(0.9\,\mathrm{GeV}/c^{2}\) for CMS; the values achieved in situ are \(1.6\,\mathrm{GeV}/c^{2}\) to \(1.7\,\mathrm{GeV}/c^{2}\) and \(1.1\,\mathrm{GeV}/c^{2}\) to \(1.5\,\mathrm{GeV}/c^{2}\) respectively [Aad:2012] [Chatrchyan:2012]. The excess over the test-bench numbers is upstream material, intercalibration between \(7.6\times 10^{4}\) channels, and the fact that a converted photon is measured worse than an unconverted one. The design calculation predicts the achieved resolution to within some tens of per cent, which is the level at which such an estimate can be trusted, and it identifies correctly which term dominates in which detector.

Because the resolution and the signal fraction both vary strongly across the detector and across event topologies, neither experiment performed a single inclusive fit. ATLAS divided its diphoton candidates into ten categories in 2011 and fourteen in 2012, by conversion status, by pseudorapidity, and by the component of the diphoton transverse momentum orthogonal to the thrust axis; CMS used a multivariate classifier trained on the expected resolution and predicted signal fraction, plus a separate class tagged by two forward jets for vector-boson fusion. The gain from doing so is not a technicality.

Proposition 113.15 (Why categorizing helps, and by how much).

Let a search be divided into independent categories \(i\) with expected signal \(s_{i}\) and background \(b_{i}\). In the Gaussian limit the combined discovery significance is

\begin{equation}\tag{113.18} Z=\left(\sum_{i}\frac{s_{i}^{2}}{b_{i}}\right)^{1/2} \;\ge\;\frac{\sum_{i}s_{i}}{\left(\sum_{i}b_{i}\right)^{1/2}} =Z_{\text{merged}}\ec \end{equation}

with equality if and only if \(s_{i}/b_{i}\) is the same in every category. Rests on Equation (11.88), Definition 11.60 and Proposition 11.93.

Derivation. Derives Proposition 113.15. Let \(n_{i}\) be Poisson with mean \(\mu s_{i}+b_{i}\) and let \(\mu\) be the signal strength. The profile-likelihood statistic of Equation (11.88) in the Gaussian limit is \(\tilde{q}_{0}=\hat{\mu}^{2}/\sigma_{\mu}^{2}\), and the Fisher information of Definition 11.60 for \(\mu\) at \(\mu=0\) is

\[ \frac{1}{\sigma_{\mu}^{2}} =-\avg{\frac{\pp^{2}\ln L}{\pp\mu^{2}}}\bigg|_{\mu=0} =\sum_{i}\frac{s_{i}^{2}}{b_{i}}\ec \]

since \(\ln L=\sum_{i}\left[n_{i}\ln(\mu s_{i}+b_{i}) -(\mu s_{i}+b_{i})\right]\) up to constants, whose second derivative is \(-\sum_{i}n_{i}s_{i}^{2}/(\mu s_{i}+b_{i})^{2}\), with expectation \(-\sum_{i}s_{i}^{2}/b_{i}\) at \(\mu=0\). For the median dataset under the signal hypothesis \(\hat{\mu}=1\), so \(\tilde{q}_{0}=\sum_{i} s_{i}^{2}/b_{i}\), and \(Z=\sqrt{\tilde{q}_{0}}\) by Proposition 11.93. That is the first equality in Equation (113.18), and the merged case is the special case of a single category.

The inequality is Cauchy–Schwarz. Write \(s_{i}=\left(s_{i}/\sqrt{b_{i}}\right)\sqrt{b_{i}}\); then

\[ \left(\sum_{i}s_{i}\right)^{2} =\left(\sum_{i}\frac{s_{i}}{\sqrt{b_{i}}}\sqrt{b_{i}}\right)^{2} \le\left(\sum_{i}\frac{s_{i}^{2}}{b_{i}}\right) \left(\sum_{i}b_{i}\right)\ec \]

with equality exactly when the vectors \((s_{i}/\sqrt{b_{i}})\) and \((\sqrt{b_{i}})\) are proportional, that is when \(s_{i}/b_{i}\) is constant. Dividing by \(\sum_{i}b_{i}\) gives Equation (113.18).

Example 113.16 (The size of the effect).

Suppose \(100\) signal events sit on \(10000\) background events, so the inclusive significance is \(100/\sqrt{10000}=1.0\). Suppose the events divide into a clean category holding \(30\) signal on \(600\) background and a dirty one holding \(70\) on \(9400\). Then

\[ Z=\left(\frac{900}{600} +\frac{4900}{9400}\right)^{1/2} =\sqrt{2.02}=1.42\ec \]

a gain of \(42\,\mathrm{\%}\) from partitioning the same events into two boxes. Nothing has been added; what has been used is the information that the two boxes have different signal purity, which a single fit throws away. This is why both collaborations spent as much effort on categorization as on the selection itself.

Background estimation. The two channels are treated in opposite ways, and the difference is instructive.

For \(\gamma\gamma\) the background is large, smooth and not calculable: it comes from continuum diphoton production, from photon plus jet with the jet misidentified, and from dijets with two misidentifications, in proportions that no simulation predicts to the accuracy required. It is therefore taken from the data. An analytic function — an exponential, or a polynomial in a Bernstein basis — is fitted to the observed mass spectrum over a wide range simultaneously with a signal peak of fixed shape, so the background normalization and shape are determined by the sidebands and the fit returns the signal yield. The danger of the method is that a wrong functional form can absorb a real signal or manufacture a spurious one, and it is controlled by the spurious-signal test: the chosen function is fitted to a background-only sample far larger than the data, and the fitted “signal” yield it returns is required to be a small fraction of the statistical uncertainty on the real measurement. The residual bias is then carried as a systematic uncertainty. The whole procedure is fixed, function and categories and all, before the signal region is examined.

For \(4\ell\) the background is small and partly calculable. The irreducible part is continuum \(ZZ^{*}\) production, taken from simulation and normalized where the simulation can be checked — on the \(Z\) peak itself, in the same final state, where the rate is thousands of times higher and the physics is the same. The reducible part, \(Z\) plus jets and \(t\bar{t}\) with jets faking leptons, is estimated from control regions in the data defined by inverting the lepton identification or isolation requirements. Both estimates are validated in sidebands before use.

Blinding and the statistical procedure

The statistical machinery is derived in Probability and Statistics and is used here as derived. This subsection states which pieces are in force; Section 113.5.1 applies them to the published numbers.

The test statistic is the profile-likelihood ratio of Definition 11.70. The parameter of interest is a single signal strength \(\mu\), defined so that \(\mu=0\) is the background-only hypothesis and \(\mu=1\) the Standard Model prediction at the mass being tested; every systematic uncertainty — energy scales, resolutions, luminosity, background shape, theoretical cross sections — enters as a nuisance parameter, constrained by an auxiliary measurement and profiled out. The one-sided discovery statistic is \(\tilde{q}_{0}\) of Equation (11.88), which refuses to count a deficit as evidence for a signal, and the significance is \(Z=\sqrt{\tilde{q}_{0}}\) by Proposition 11.93. That identity is what allows a significance to be quoted without simulating an ensemble of experiments, and it rests on Wilks' theorem (Theorem 11.88) applied to the profile statistic (Corollary 11.89); the asymptotic formulae in the form the collaborations use them are [Cowan:2011].

Exclusion uses the same machinery reversed, with the modified frequentist \(CL_{s}\) criterion [Read:2002] — the ratio of the two tail probabilities rather than the numerator alone — so that a downward background fluctuation cannot exclude a signal the experiment had no power to see (Remark 11.95). This is the construction behind the LEP limit of Proposition 113.7 and behind every exclusion plot in the 2012 papers.

Two conventions must be stated because mixing them shifts a significance by more than a tenth of a standard deviation (Remark 11.94). The \(p\)-values quoted below are one-sided, so that \(Z=5\) corresponds to \(p=2.87\times 10^{-7}\); and the significances quoted by the collaborations as “local” are computed at a mass fixed in advance, which no search does. The correction is Section 113.5.1.

Blinding. None of this is worth anything if the analysis was adjusted after the data were seen. Both collaborations therefore fixed the entire chain — trigger, selection, categorization, background parametrization, systematic model and statistical procedure — using simulation and the mass sidebands, with the signal region hidden, and unblinded only afterwards. The reason is not procedural hygiene but arithmetic: the \(p\)-value of Definition 11.75 is the probability, under the null, of a result at least as extreme as the one observed by a test specified in advance. A cut adjusted after seeing the excess makes the number it produces uninterpretable, because the effective number of hypotheses tested is then unknown and unboundable. Blinding is what makes the five-sigma threshold mean the thing it appears to mean, and it is the reason the 2012 result was believed immediately rather than after a decade of argument.

Observations and data

The diphoton channel

What the data show is a narrow peak sitting on a steeply falling smooth continuum. The continuum falls by roughly an order of magnitude between \(m_{\gamma\gamma}c^{2}=100\,\mathrm{GeV}\) and \(160\,\mathrm{GeV}\) and has no structure in it; the peak is one resolution width across, of order \(1.5\,\mathrm{GeV}/c^{2}\) by Proposition 113.14, and stands a few per cent above the continuum in the best categories and well under one per cent in the worst. Nothing about the picture is dramatic; what makes it a measurement is the arithmetic done to it.

ATLAS reported a local significance of \(4.5\) standard deviations near \(m_{H}c^{2}=126.5\,\mathrm{GeV}\) in this channel alone, against \(2.5\) expected for a Standard Model scalar at that mass [Aad:2012]; CMS reported \(4.1\) against \(2.8\) expected near \(125\,\mathrm{GeV}\) [Chatrchyan:2012]. Two things about those pairs of numbers deserve saying plainly, because they are usually skipped.

First, the expected significances — \(2.5\) and \(2.8\) — are close to the crude estimate \(s/\sqrt{b}\approx2\) obtained in Proposition 113.12 from cross section, branching fraction, efficiency and background. That agreement is a check on the whole chain: an experiment whose expected sensitivity disagreed by a factor of two with a back-of-envelope calculation from published inputs would have something wrong with it.

Second, both observed values exceed both expected values. The fitted diphoton signal strength was correspondingly high in both experiments, around \(1.8\) with an uncertainty near \(0.5\) in ATLAS and similarly above unity in CMS. That is an upward fluctuation of about \(1.5\) standard deviations, present in the channel that dominates the combination — exactly the configuration in which an unblinded, after-the-fact analysis would be untrustworthy, and exactly why the selection was frozen beforehand. The later, much larger datasets brought the diphoton signal strength back to unity, which is the expected behaviour of a fluctuation and confirms the 2012 diagnosis.

One structural fact is available from this channel with no further analysis: a resonance that decays to two real photons cannot have spin \(1\). That is the Landau–Yang argument [Landau:1948] [Yang:1950], derived in Theorem 104.60 and again, in the vector form, in the derivation accompanying Phenomenon 113.22. The discovery channel therefore excluded one of the four spin hypotheses on the day it was announced, without any angular measurement.

The four-lepton channel

\(H\to ZZ^{*}\to4\ell\) is the opposite kind of measurement. The final state is fully reconstructed — four charged leptons, all their momenta measured in the tracker — so the mass resolution is set by lepton momentum resolution alone and is of order \(2\,\mathrm{GeV}/c^{2}\), and the background is small enough that the signal fraction exceeds unity in a narrow window. What it does not have is events: the arithmetic of Proposition 113.12 gives about twenty-five produced decays in the ATLAS dataset and fewer than ten surviving selection.

Both experiments observed roughly a dozen candidates in a window \(10\,\mathrm{GeV}/c^{2}\) wide around \(125\,\mathrm{GeV}/c^{2}\), against an expected background of about five, and reported local significances of \(3.6\) standard deviations (ATLAS, against \(2.7\) expected) and \(3.2\) (CMS, against \(3.8\) expected) [Aad:2012] [Chatrchyan:2012]. With numbers this small the Gaussian approximation is doing real work, and the collaborations used the asymptotic formulae [Cowan:2011] whose accuracy at such yields is itself part of what Probability and Statistics establishes; the CMS value fell below its own expectation, which is an ordinary downward fluctuation of a Poisson count.

The complementarity is the point of quoting both channels. A high-rate, high-background channel and a low-rate, low-background one fail in different ways: a mismodelled background shape would move the diphoton result and not the four-lepton one, and a mismeasured lepton efficiency would do the reverse. Their masses agreed to within the resolutions, and their signal strengths agreed with each other and with unity. Two supporting channels, \(H\to WW^{*}\to\ell\nu\ell\nu\) and \(H\to\tau^{+}\tau^{-}\), contributed sensitivity but no mass information at all: the neutrinos in the first and in the tau decays of the second make full reconstruction impossible, so those channels report an excess of events with the expected kinematics and not a resonance. They are evidence for a signal and not evidence for a particle, and both papers treat them as such.

Combined significance, 4 July 2012

Combining channels, ATLAS reported a local significance of \(5.9\) standard deviations at \(m_{H}c^{2}=126.0\,\mathrm{GeV}\) against \(4.9\) expected, and CMS \(5.0\) at \(125.5\,\mathrm{GeV}\) against \(5.8\) expected [Aad:2012] [Chatrchyan:2012]. The masses were

\begin{equation}\tag{113.19} m_{H}c^{2}\Big|_{\text{ATLAS}} =126.0\,\mathrm{GeV}\pm0.4\,\mathrm{GeV} \pm0.4\,\mathrm{GeV}\ec\quad m_{H}c^{2}\Big|_{\text{CMS}} =125.3\,\mathrm{GeV}\pm0.4\,\mathrm{GeV} \pm0.5\,\mathrm{GeV}\ec \end{equation}

the first uncertainty statistical and the second systematic, and the signal strengths relative to the Standard Model prediction were \(\mu=1.4(3)\) and \(\mu=0.87(23)\) respectively.

The two collaborations were blind to each other's results until the seminar of 4 July 2012. That is not a detail of etiquette. Two measurements that agree are evidence in proportion to how independent they are, and these two shared only the accelerator: different magnets, different calorimeter technologies, different reconstruction software, different people, and no exchange of intermediate results. An error common to both would have to originate in the beam or in the theory of the Standard Model, not in either apparatus.

Phenomenon 113.17 (A new boson near 125 GeV).

Two experiments at the Large Hadron Collider, blind to each other's results, reported on the same day a narrow resonance in the diphoton and four-lepton final states, with local significances of \(5.9\) and \(5.0\) standard deviations respectively [Aad:2012] [Chatrchyan:2012]. The combined mass from the two high-resolution channels is

\begin{equation}\tag{113.20} m_{H}c^{2}=125.09(24)\,\mathrm{GeV}\ec \end{equation}

[Aad:2015], and the present world average is \(125.20(11)\,\mathrm{GeV}\), that is \(m_{H}=2.2319\times 10^{-25}\,\mathrm{kg}\) [Navas:2024]. The particle is electrically neutral, decays to two photons and to two \(Z\) bosons, and is produced at a rate compatible with the single scalar that Phenomenon 113.6 requires. Rests on Phenomenon 113.6, Equation (113.19) and Proposition 113.5.

Derivation of the consistency checks. Derives Phenomenon 113.17. Three separate agreements have to hold, and each is an arithmetic statement that could have failed.

The two experiments agree. Adding the errors of Equation (113.19) in quadrature gives \(126.0(6)\,\mathrm{GeV}\) for ATLAS and \(125.3(6)\,\mathrm{GeV}\) for CMS, where \(\sqrt{0.4^{2}+0.4^{2}}=0.57\) and \(\sqrt{0.4^{2}+0.5^{2}}=0.64\). The difference is \(0.7\,\mathrm{GeV}\) with a combined uncertainty \(\sqrt{0.57^{2}+0.64^{2}}=0.86\), that is \(0.8\) standard deviations. Two independent instruments measuring the same mass agreed at once.

The 2012 numbers survive. The joint ATLAS and CMS combination of the full Run-1 data [Aad:2015] gives \(125.09\,\mathrm{GeV}\pm0.21\,\mathrm{GeV}\,(\text{stat}) \pm0.11\,\mathrm{GeV}\,(\text{syst})\), and adding those in quadrature is \(125.09(24)\,\mathrm{GeV}\), which is Equation (113.20). The present world average is \(125.20(11)\,\mathrm{GeV}\) [Navas:2024], higher by \(0.11\,\mathrm{GeV}\).

These two are successive best values, not two measurements. The world average contains the Run-1 combination, so the two are strongly correlated, their errors may not be added in quadrature, and the difference cannot be converted into a number of standard deviations without an estimate of the correlation that neither publication supplies. What can be said without one is the thing worth saying: twelve years and two orders of magnitude in integrated luminosity moved the central value by \(0.11\,\mathrm{GeV}\), less than half the 2015 uncertainty, while more than halving that uncertainty. A shift of that size is what an unbiased measurement being refined looks like; a shift of several \(\mathrm{GeV}\) would have been a retraction.

The rate is the predicted one. The signal strengths \(\mu=1.4(3)\) and \(\mu=0.87(23)\) are ratios of the observed rate to the Standard Model prediction computed from Proposition 113.5 with the mass now fixed. Their weighted mean is

\[ \bar{\mu}=\frac{1.4/0.3^{2} +0.87/0.23^{2}} {1/0.3^{2}+1/0.23^{2}}=1.06\ec\qquad \sigma_{\bar{\mu}}=\left(\frac{1}{0.3^{2}} +\frac{1}{0.23^{2}}\right)^{-1/2}=0.18\ec \]

so \(\bar{\mu}=1.06(18)\), consistent with unity. This is the check that a bump alone would not supply: a narrow resonance at \(125\,\mathrm{GeV}/c^{2}\) produced at a hundredth of the predicted rate, or a hundred times it, would be a new particle and would not be this one.

What the combination does not establish is treated in Section 113.5.1: the quoted significances are local, and the mass was not nominated in advance.

The measured numbers

Table 113.3 collects the quantities of the discovery itself and Table 113.4 what has been measured since; Tables 113.5 and 113.6 collect the coupling test, which is the measurement that distinguishes this scalar from any other. All four sit here, under this rubric, rather than beside the sections that use them, so that a reader looking for the chapter's numbers finds them in one place. Rows that record a run setting carry no uncertainty because they define the conditions rather than measure anything; every row that reports a measured or derived quantity carries one, and a derived row names the equation that produced it.

QuantityValueYearSource
ATLAS integrated luminosity, \(7\,\mathrm{TeV}\) (setting)\(4.8\times 10^{43}\,/\mathrm{m}^{2}\), i.e. \(4.8\) inverse femtobarns2011[Aad:2012]
ATLAS integrated luminosity, \(8\,\mathrm{TeV}\) (setting)\(5.8\times 10^{43}\,/\mathrm{m}^{2}\)2012[Aad:2012]
CMS integrated luminosity, \(7\,\mathrm{TeV}\) (setting)\(5.1\times 10^{43}\,/\mathrm{m}^{2}\)2011[Chatrchyan:2012]
CMS integrated luminosity, \(8\,\mathrm{TeV}\) (setting)\(5.3\times 10^{43}\,/\mathrm{m}^{2}\)2012[Chatrchyan:2012]
Scalars produced in the ATLAS dataset$2.12\times 10^{5}$ (derived from the two rows above and the cross sections of Section 113.3.1; good to the \(10\,\mathrm{\%}\) accuracy of those predictions)2012Proposition 113.12
LEP lower limit, \(95\,\mathrm{\%}\) confidence$m_{H}c^{2}>114.4\,\mathrm{GeV}$2003[Barate:2003]
Tevatron excess in $VH$, $H\to b\bar{b}$local significance $\approx3$, global $\approx2.5$2012[Aaltonen:2012]
ATLAS local significance, $\gamma\gamma$\(4.5\) observed, \(2.5\) expected2012[Aad:2012]
ATLAS local significance, $ZZ^{*}\to4\ell$\(3.6\) observed, \(2.7\) expected2012[Aad:2012]
ATLAS combined local significance\(5.9\) observed, \(4.9\) expected2012[Aad:2012]
ATLAS global significance, \(110\text{–}150\,\mathrm{GeV}\) range\(5.3\)2012[Aad:2012]
ATLAS global significance, \(110\text{–}600\,\mathrm{GeV}\) range\(5.1\)2012[Aad:2012]
CMS local significance, $\gamma\gamma$\(4.1\) observed, \(2.8\) expected2012[Chatrchyan:2012]
CMS local significance, $ZZ^{*}\to4\ell$\(3.2\) observed, \(3.8\) expected2012[Chatrchyan:2012]
CMS combined local significance\(5.0\) observed, \(5.8\) expected2012[Chatrchyan:2012]
The discovery itself: the datasets it rests on, the prior constraints it had to clear, and the significances reported in 2012. Integrated luminosities are given in $/\mathrm{m}^{2}$ with the customary inverse-femtobarn figure alongside (Notation 113.1); one inverse femtobarn is $10^{43}\,/\mathrm{m}^{2}$. Significances are pure numbers. The global significances are the collaboration's own look-elsewhere corrections over the mass range named; Section 113.5.1 reconstructs the trials factors they imply. What was measured afterwards is Table 113.4.
QuantityValueYearSource
ATLAS mass\(126.0\,\mathrm{GeV}\)$/c^{2}$, $\pm0.4$ (stat) $\pm0.4$ (syst); \(126.0(6)\,\mathrm{GeV}\) combined2012[Aad:2012]
CMS mass\(125.3\,\mathrm{GeV}\)$/c^{2}$, $\pm0.4$ (stat) $\pm0.5$ (syst); \(125.3(6)\,\mathrm{GeV}\) combined2012[Chatrchyan:2012]
ATLAS signal strength $\mu$ at \(126\,\mathrm{GeV}\)$/c^{2}$\(1.4(3)\)2012[Aad:2012]
CMS signal strength $\mu$ at \(125.5\,\mathrm{GeV}\)$/c^{2}$\(0.87(23)\)2012[Chatrchyan:2012]
Weighted mean signal strength\(1.06(18)\) (derived)2012Phenomenon 113.17
ATLAS and CMS combined mass\(125.09(24)\,\mathrm{GeV}\)$/c^{2}$2015[Aad:2015]
World-average mass\(125.20(11)\,\mathrm{GeV}\)$/c^{2}$, i.e. \(2.2319(20)\times 10^{-25}\,\mathrm{kg}\)2024[Navas:2024]
Total width, as an energy $\hbar\Gamma_{H}$\(3.7\,\mathrm{MeV}\), $+1.9$, $-1.4$; in SI $5.93\times 10^{-13}\,\mathrm{J}$2024[Navas:2024]
Mean lifetime $\tau=\hbar/(\hbar\Gamma_{H})$\(1.78\times 10^{-22}\,\mathrm{s}\) (derived), with $+1.08$ and $-0.60$ on the same power of ten; decay length $c\tau=5.3\times 10^{-14}\,\mathrm{m}$, with $+3.2$ and $-1.8$ likewise; both errors are the width's $+51\,\mathrm{\%}$, $-38\,\mathrm{\%}$ propagated2024Phenomenon 113.26
Electroweak scale $E_{v}$\(246.220\,\mathrm{GeV}\), i.e. \(3.94488\times 10^{-8}\,\mathrm{J}\) (derived from the muon lifetime; the fractional uncertainty, $3\times 10^{-7}$, is beyond the digits shown)2013[Tishchenko:2013]
What has been measured since. The 2012 masses carry a statistical and a systematic error, quoted separately by the collaborations and combined in quadrature here, with the arithmetic shown in the derivation accompanying Phenomenon 113.17. The two 2012 masses and the later combinations are successive best values of one quantity and not independent measurements to be averaged — the world average already contains the earlier data. Signal strengths are pure numbers. The width, the lifetime and the electroweak scale are derived quantities, and each is given in SI beside its customary form.
PartnerMass [Navas:2024]Predicted reduced couplingMeasured $\kappa$ [Aad:2022] [Tumasyan:2022]Year
$W$\(80.369(13)\,\mathrm{GeV}\)$/c^{2}$; \(1.432708(23)\times 10^{-25}\,\mathrm{kg}\)\(0.326411(53)\)$1$ to about \(10\,\mathrm{\%}\)2024; 2022
$Z$\(91.1880(20)\,\mathrm{GeV}\)$/c^{2}$; \(1.6255738(36)\times 10^{-25}\,\mathrm{kg}\)\(0.3703517(81)\)$1$ to about \(10\,\mathrm{\%}\)2024; 2022
$t$\(172.57(29)\,\mathrm{GeV}\)$/c^{2}$; \(3.07634(52)\times 10^{-25}\,\mathrm{kg}\)\(0.70088(118)\)$1$ to about \(10\,\mathrm{\%}\); $t\bar{t}H$ observed [Sirunyan:2018]2024; 2022
$b$\(4.183(7)\,\mathrm{GeV}\)$/c^{2}$; \(7.4569(12)\times 10^{-27}\,\mathrm{kg}\)\(0.016989(28)\)$1$ to about \(10\,\mathrm{\%}\)2024; 2022
$\tau$\(1.77693(9)\,\mathrm{GeV}\)$/c^{2}$; \(3.1676654(16)\times 10^{-27}\,\mathrm{kg}\)\(0.0072168(4)\)$1$ to about \(10\,\mathrm{\%}\)2024; 2022
$\mu$\(0.1056583755(23)\,\mathrm{GeV}\)$/c^{2}$; \(1.88353163(4)\times 10^{-28}\,\mathrm{kg}\)\(4.291218\times 10^{-4}\)$1$ to about \(20\,\mathrm{\%}\); decay observed at three standard deviations [Sirunyan:2021]2024; 2021
The coupling test. The predicted reduced coupling is $m c^{2}/E_{v}$ for a fermion and $M_{V}c^{2}/E_{v}$ for a weak boson, by Equation (104.70); it is a pure number, it contains no free parameter, and its uncertainty here is propagated from the measured mass alone, since $E_{v}$ is known to three parts in $10^{7}$. Masses are the PDG 2024 values, which are also the masses in $\mathrm{kg}$ shown alongside; each kilogram value carries the same fractional uncertainty as the $\mathrm{GeV}$ value beside it, written in brackets on the last digits quoted. The fourth column records what the ten-year coupling maps report: every measured coupling modifier $\kappa_{i}$, defined so that $\kappa_{i}=1$ reproduces the predicted value, is consistent with unity, and the entry gives the approximate precision, not a central value — the papers cited should be consulted for those and for their correlations. The last column carries the year of each of the two measured entries in the row, the mass first. The partners whose couplings are not yet measured are collected separately in Table 113.6.
PartnerMass [Navas:2024]Predicted reduced couplingStatus [Aad:2022] [Tumasyan:2022]Year
$c$\(1.273(5)\,\mathrm{GeV}\)$/c^{2}$; \(2.2693(89)\times 10^{-27}\,\mathrm{kg}\)\(0.005170(20)\)not measured; only bounded2024; —
$e$, $u$, $d$, $s$below \(0.1\,\mathrm{GeV}\)$/c^{2}$below \(4\times 10^{-4}\)not measured2024; —
Self-coupling $\kappa_{\lambda}$$1$ by Proposition 104.63allowed anywhere from about $-1$ to $+7$—; 2022
What the coupling test does not yet reach. The predicted reduced coupling is defined as in Table 113.5; the difference is that no entry here has a measured $\kappa$. The charm coupling is bounded but not observed, the first-generation and strange couplings are far below present sensitivity, and the self-coupling — the one number that probes the shape of the potential rather than its minimum — is constrained only to a wide interval. The last column carries the year of the mass and of the bound respectively.
Remark 113.18 (Where each number came from, and where it did not).

The masses in Table 113.5 and the world-average scalar mass and width in Table 113.4 are taken from the Particle Data Group's own machine-readable 2024 release, stored with this treatise as evidence [Navas:2024]; the conversions to \(\mathrm{kg}\) and \(\mathrm{J}\) use the exact 2019 SI values of \(e\) and \(c\) [Mohr:2025]. The 2012 significances, signal strengths and per-channel numbers are the collaborations' published values [Aad:2012] [Chatrchyan:2012] and are not recomputed here. The rows marked derived are computed in this chapter from the rows above them, with the arithmetic shown; the produced-scalar count inherits the roughly ten per cent uncertainty of the predicted production cross sections, which is why it is quoted to three figures and not to five. The predicted reduced couplings are exact functions of the masses and \(E_{v}\) and carry no theoretical uncertainty at tree level.

Interpretation

The look-elsewhere effect

Every significance in Table 113.3 labelled local is the answer to a question nobody asked: what is the probability of a fluctuation this large at a mass named in advance? No search names a mass in advance. Both experiments scanned a range and reported the largest excess found in it, and the probability that some point of a range fluctuates upward is larger — here by one to two orders of magnitude — than the probability that a nominated point does. The correction is derived in Section 11.6 and is applied here to the published numbers.

Proposition 113.19 (The trials factors of the 2012 result).

With the one-sided convention of Proposition 11.93, the ATLAS local significance \(Z_{\text{loc}}=5.9\) corresponds to \(p_{\text{loc}}=1.82\times 10^{-9}\), and the two global significances the collaboration quotes correspond to trials factors

\begin{equation}\tag{113.21} N_{\text{tr}}\big|_{110\text{–}150\,\mathrm{GeV}}=31.9\ec \qquad N_{\text{tr}}\big|_{110\text{–}600\,\mathrm{GeV}}=93.4\ep \end{equation}

The first is reproduced to within one per cent by counting resolution elements; the second is not, and the reason is physical. Rests on Definition 11.92, Definition 11.96 and Proposition 113.14.

Derivation. Derives Proposition 113.19. By Definition 11.92, \(p=1-\Phi(Z)\), so \(p_{\text{loc}}=1-\Phi(5.9)=1.8175\times 10^{-9}\), \(1-\Phi(5.3)=5.7901\times 10^{-8}\) and \(1-\Phi(5.1)=1.6983\times 10^{-7}\). The trials factor is the ratio \(N_{\text{tr}}=p_{\text{glob}}/p_{\text{loc}}\) of Definition 11.96, giving \(5.7901\times 10^{-8}/1.8175\times 10^{-9}=31.86\) for the narrow range and \(1.6983\times 10^{-7}/1.8175\times 10^{-9}=93.44\) for the wide one, which is Equation (113.21).

The narrow range, from resolution elements. The naive estimate of Equation (11.93) treats the scan as \(N=\Delta m/\delta m\) independent tests. With \(\Delta m c^{2}=40\,\mathrm{GeV}\) and the diphoton resolution \(\delta m c^{2}=1.7\,\mathrm{GeV}\) of Proposition 113.14 — the channel that dominates the combination — this is \(N=23.5\), whence \(p_{\text{glob}}\simeq23.5\times1.8175\times 10^{-9} =4.27\times 10^{-8}\) and \(Z_{\text{glob}}=\Phi^{-1}(1-p_{\text{glob}})=5.36\). ATLAS quotes \(5.3\). The agreement is closer than the method deserves and should not be over-read — Remark 11.99 says why the independence assumption has no general justification — but it does show that the published correction is the size a reader can reconstruct from the resolution and the window.

The wide range, where the same estimate fails. Repeating it over \(110\text{–}600\,\mathrm{GeV}\) gives \(N=490\,\mathrm{GeV}/1.7\,\mathrm{GeV}=288\) and \(Z_{\text{glob}}=4.88\), whereas ATLAS quotes \(5.1\), corresponding to \(N_{\text{tr}}=93\). The naive count is too large by a factor of three, and the reason is that \(\delta m=1.7\,\mathrm{GeV}/c^{2}\) is not the correlation length of the test statistic above about \(180\,\mathrm{GeV}/c^{2}\). Two effects lengthen it: the channels that carry the sensitivity at high mass (\(ZZ\to4\ell\) with degrading resolution, and \(WW\), which has no mass peak at all) are much coarser than the diphoton channel; and the predicted width of a Standard Model scalar grows rapidly once the \(WW\) and \(ZZ\) thresholds open, reaching tens of \(\mathrm{GeV}\) by \(500\,\mathrm{GeV}/c^{2}\), so neighbouring hypotheses in that region are nearly the same hypothesis. Fewer effectively independent tests, hence a smaller trials factor.

The upcrossing form. The honest machinery is Theorem 11.101 with the level dependence of Theorem 11.103, which for large \(Z_{\text{loc}}\) gives Equation (11.98). That statement is an upper bound — it descends from Davies' bound — and reads \(N_{\text{tr}}\le1+\sqrt{2\pi}\avg{N_{u_{0}}}\ee^{u_{0}/2} Z_{\text{loc}}\) with \(u_{0}=1\), so inverting it about \(\avg{N_{u_{0}}}\) turns it into a lower bound on the upcrossing count and not into a value. Since \(\sqrt{2\pi}\,\ee^{1/2}=4.133\) and \(Z_{\text{loc}}=5.9\), the published trials factors imply mean numbers of one-sigma upcrossings of at least

\[ \avg{N_{1}}\big|_{110\text{–}150\,\mathrm{GeV}} \ge\frac{31.86-1}{4.133\times5.9}=1.27\ec \qquad \avg{N_{1}}\big|_{110\text{–}600\,\mathrm{GeV}}\ge3.79\ep \]

Both floors are of order unity, which is exactly the regime in which the calibrate-and-extrapolate procedure works: a mean upcrossing count near one can be measured from a few hundred simulated background-only experiments, whereas the tail probability it is used to compute, \(5\times 10^{-8}\), could not be sampled by any feasible number of them. And the two floors stand in the ratio \(2.99\), far below the ratio of the two mass ranges, \(12.25\) — which is the quantitative form of the statement that the correlation length grows with mass, provided the bound is comparably tight in both ranges, as it is, both being evaluated at the same \(Z_{\text{loc}}\).

Remark 113.20 (What survives the correction, honestly).

Applied to ATLAS the correction moves \(5.9\) to \(5.3\) over the range where a Standard Model scalar was still allowed, and to \(5.1\) over the full range scanned. Both remain above the conventional threshold. Applied to CMS it does not: repeating the resolution-element estimate with a window \(110\text{–}145\,\mathrm{GeV}\) and a diphoton resolution near \(1.4\,\mathrm{GeV}/c^{2}\) gives \(N=25\), hence \(p_{\text{glob}}\simeq25\times2.87\times 10^{-7}=7.2\times 10^{-6}\) and \(Z_{\text{glob}}\simeq4.3\) — an estimate made here, not a published number, and one that puts the CMS result below five standard deviations globally.

That is the correct state of affairs to record, and it is worth being blunt about it. The 2012 discovery does not rest on either experiment individually crossing an arbitrary threshold. It rests on two independent experiments finding an excess of the same size at the same mass in the same two channels on the same day, with the rate the theory predicts — the three checks in the derivation accompanying Phenomenon 113.17. A joint probability is not the product of two marginal ones when the alternatives are correlated, so no single combined number is quoted here; what can be said is that the agreement of two blind measurements is evidence of a kind that no amount of significance in one of them supplies.

Three further cautions, all of them from Probability and Statistics and all of them routinely mishandled in secondary accounts. A \(p\)-value is the probability of the data given the null, never the probability of the null given the data (Remark 11.77); “a one-in-\(3.5\)-million chance that it is a fluctuation” is a mistranslation of \(Z=5\), not a paraphrase of it. The one-sided and two-sided conventions differ by \(0.14\) in \(Z\) at \(q=25\) (Remark 11.94), which is enough to move a result across the threshold, so the convention must be stated — it is one-sided throughout this chapter. And a significance quoted without the range searched is a local number wearing a global name (Remark 11.108).

The mass

The measured mass is now

\begin{equation}\tag{113.22} m_{H}c^{2}=125.20(11)\,\mathrm{GeV}\ec\qquad m_{H}=2.2319(20)\times 10^{-25}\,\mathrm{kg}\ec \end{equation}

a fractional precision of \(8.8\times 10^{-4}\) [Navas:2024]. Almost all of it comes from the two high-resolution channels, and the systematic error is dominated by one thing: the absolute energy scale of the electromagnetic calorimeters.

That scale is not known a priori. A crystal's light yield, a liquid-argon gap's response and the material in front of both are known to a per cent at best, which would give a mass uncertainty of over \(1\,\mathrm{GeV}/c^{2}\) — ten times the quoted error. What fixes it is a calibration source inside the same detector: the \(Z\to e^{+}e^{-}\) peak, tens of millions of events at \(91.1880(20)\,\mathrm{GeV}/c^{2}\) whose mass is known from LEP to two parts in \(10^{5}\) [Schael:2006] [Navas:2024]. Reconstructing it with the experiment's own electrons and demanding that the reconstructed peak sit at the LEP value calibrates the energy scale in situ, cell by cell and category by category. The residual uncertainty is then the extrapolation from the electron scale at \(45\,\mathrm{GeV}\) to the photon scale at \(60\,\mathrm{GeV}\) — a small step, and it is small because the calibration source and the measurement are in the same instrument and differ by less than a factor of two in energy.

Remark 113.21 (Fixing the mass converts the rest into tests).

\(m_{H}\) was the last free parameter of the Standard Model to be measured. Its value fixes the quartic coupling by Equation (113.3),

\[ \hat{\lambda}=\frac{1}{2} \left(\frac{m_{H}c^{2}}{E_{v}}\right)^{2} =\frac{1}{2}\left(\frac{125.20}{246.220}\right)^{2} =\frac{1}{2}\left(0.50849\right)^{2}=0.12928\ec \]

in agreement with Equation (104.38), and with it the curvature parameter \(\hbar\mu c=m_{H}c^{2}/\sqrt{2} =88.53\,\mathrm{GeV}\). From that moment every other statement the theory makes about the scalar — every branching fraction, every production rate, every self-coupling — is a prediction with no adjustable parameter, and each subsequent measurement is a test rather than a determination. The rest of this section, and the whole of Section 113.6, is the score.

Spin and parity

Four hypotheses were on the table in 2012: \(J^{P}=0^{+}\), the prediction; \(0^{-}\), a pseudoscalar; \(1^{\pm}\); and \(2^{+}\), a graviton-like tensor. One of them was excluded by the discovery channel itself.

Phenomenon 113.22 (The new boson has spin zero and even parity).

The angular distributions of the four-lepton and diphoton final states are those of a \(J^{P}=0^{+}\) resonance; the \(0^{-}\), \(1^{\pm}\) and \(2^{+}\) alternatives are excluded at high confidence by both experiments [Aad:2013] [Chatrchyan:2013]. The exclusion of spin \(1\) requires no angular analysis at all: it follows from the mere existence of the diphoton decay. Rests on Theorem 104.60.

Derivation. Derives Phenomenon 113.22. This is the Landau–Yang argument [Landau:1948] [Yang:1950], given in the vector form here and in the amplitude form as Theorem 104.60. Work in the rest frame of the parent, where the two photons carry momenta \(\pm\vect{k}\) and transverse polarization vectors \(\vect{e}_{1}\) and \(\vect{e}_{2}\), with \(\vect{e}_{1}\cdot\vect{k}=\vect{e}_{2}\cdot\vect{k}=0\). A parent of spin \(1\) is described by a polarization vector \(\vect{J}\), and the decay amplitude must be a rotational scalar, linear in each of \(\vect{e}_{1}\), \(\vect{e}_{2}\) and \(\vect{J}\). The only such scalars available are

\[ \left(\vect{e}_{1}\cdot\vect{e}_{2}\right) \left(\vect{J}\cdot\vect{k}\right)\ec\quad \vect{J}\cdot\left(\vect{e}_{1}\times\vect{e}_{2}\right)\ec\quad \left(\vect{J}\cdot\vect{e}_{1}\right) \left(\vect{k}\cdot\vect{e}_{2}\right) +\left(\vect{J}\cdot\vect{e}_{2}\right) \left(\vect{k}\cdot\vect{e}_{1}\right)\ep \]

The list is complete, and transversality is what makes it so. Since \(\vect{e}_{1}\) and \(\vect{e}_{2}\) are both perpendicular to \(\vect{k}\), the vector \(\vect{e}_{1}\times\vect{e}_{2}\) is parallel to \(\vect{k}\); a further candidate such as \(\left(\vect{J}\cdot\vect{k}\right) \left[\vect{k}\cdot\left(\vect{e}_{1}\times\vect{e}_{2}\right)\right]\) is therefore the second structure above multiplied by \(\abs{\vect{k}}^{2}\), and not an independent one.

The photons are identical bosons, so the amplitude must be unchanged under the exchange \((\vect{e}_{1},\vect{k})\leftrightarrow(\vect{e}_{2},-\vect{k})\). The first structure changes sign, because \(\vect{J}\cdot\vect{k}\) does; the second changes sign, because the cross product does; the third vanishes identically by transversality. No amplitude survives. A spin-\(1\) particle therefore cannot decay into two real photons, whatever its parity, and the observed diphoton channel excludes \(J=1\) on its own.

The argument uses rotational invariance, Bose statistics and transversality and nothing else. It does not use parity, so it kills \(1^{+}\) and \(1^{-}\) together; and it fails for two virtual photons or for a massive vector in the final state, both of which have a longitudinal polarization and therefore a third structure. The angular analysis is needed only to separate \(0^{+}\) from \(0^{-}\) and from the spin-\(2\) hypotheses.

That separation uses the four-lepton final state, which is fully reconstructed and therefore carries its complete angular information: five angles — the production polar angle, the two polar angles of the lepton pairs in their parent frames, and two azimuths, one of them the angle \(\Phi\) between the planes of the two lepton pairs — plus the two lepton-pair masses. The discriminating variable is \(\Phi\). A scalar couples to the two \(Z\) bosons through the structure \(\varepsilon_{1}^{*}\cdot\varepsilon_{2}^{*}\) of Equation (104.69), which produces a distribution containing a \(\cos\Phi\) term; a pseudoscalar can only couple through the antisymmetric structure \(\epsilon_{\mu\nu\rho\sigma}\varepsilon_{1}^{*\mu} \varepsilon_{2}^{*\nu}k_{1}^{\rho}k_{2}^{\sigma}\), which has no \(\cos\Phi\) term and a \(\cos2\Phi\) term of the opposite sign; and a spin-\(2\) state produces polar distributions that neither can. Both collaborations built a likelihood in the seven variables and reported the alternatives disfavoured at high confidence, CMS excluding the pseudoscalar at above \(99\,\mathrm{\%}\) and ATLAS reaching a comparable conclusion [Aad:2013] [Chatrchyan:2013]. The scalar hypothesis was not merely preferred: the alternatives were rejected.

The three structures named in the previous paragraph are quoted, not derived. Landau–Yang is derived above and disposes of \(J=1\) from the diphoton channel alone; the separation of \(0^{+}\) from \(0^{-}\) and from spin \(2\) is the load-bearing half of Phenomenon 113.22 and this chapter does not carry it.

Derivation pending.

The azimuthal distributions that separate \(0^{+}\) from \(0^{-}\) and from the spin-\(2\) hypotheses in \(H\to ZZ^{*}\to4\ell\). What is owed is the decay amplitude for each hypothesis — the scalar contraction of the two boson polarizations, the Levi-Civita contraction that a pseudoscalar is restricted to, and the tensor couplings — projected onto the two lepton pairs and integrated to give the distribution in the angle between their decay planes, exhibiting the \(\cos\Phi\) term the scalar produces, its absence for the pseudoscalar, and the \(\cos2\Phi\) term of opposite sign that the pseudoscalar carries. It belongs in Appendix A beside the vector-boson vertices of the electroweak chapter, whose loop-induced counterparts are already recorded as owed there.

Remark 113.23 (A $CP$-odd admixture is bounded, not excluded).

Excluding a pure \(0^{-}\) state is not the same as establishing that the state is purely \(0^{+}\). What the data constrain is the fraction of the rate arising from a \(CP\)-odd contribution, and the constraint is weak in a specific and instructive way: in the \(HZZ\) channel any \(CP\)-odd coupling would have to be generated by a loop, so it is suppressed by construction and the channel's apparent sensitivity overstates the test. The sharper probe is \(H\to\tau^{+}\tau^{-}\), where a \(CP\)-odd Yukawa coupling enters at tree level and shows up as a phase in the correlation between the tau decay planes; there the mixing angle is currently bounded to a few tens of degrees.

The question is not academic. The Standard Model's \(CP\) violation is quantitatively insufficient to generate the observed baryon asymmetry by many orders of magnitude (Section 111.8.2), and the electroweak transition it predicts is a crossover, which supplies no departure from equilibrium either (Remark 104.65). A \(CP\)-odd component in the scalar sector is one of the few places where new \(CP\) violation could hide and be found with existing machines. It has not been found; it has been bounded.

Couplings against mass

Everything so far would be satisfied by a scalar that had nothing to do with mass generation. The measurement that closes the argument is the pattern of couplings.

Phenomenon 113.24 (Couplings proportional to mass).

The measured couplings of the new boson to the \(W\), the \(Z\), the top, the bottom and the tau, together with the observed \(H\to\mu^{+}\mu^{-}\) decay [Sirunyan:2021] and the observed \(t\bar{t}H\) production [Sirunyan:2018], lie on the single line that the mechanism predicts when they are plotted against the mass of the partner: coupling proportional to mass for fermions and to mass squared for gauge bosons, over more than three decades in mass [Aad:2016] [Aad:2022] [Tumasyan:2022]. Rests on Equation (104.70) and Proposition 104.61.

Derivation. Derives Phenomenon 113.24. In the mechanism a fermion mass comes from a Yukawa term evaluated on the vacuum, \(m_{f}c^{2}=\hat{y}_{f}E_{v}/\sqrt{2}\), while the coupling of the physical scalar to that same fermion is the same \(\hat{y}_{f}\); hence

\begin{equation}\tag{113.23} g_{Hff}=\frac{\hat{y}_{f}}{\sqrt{2}} =\frac{m_{f}c^{2}}{E_{v}}\ec \end{equation}

a pure number. A gauge-boson mass comes instead from the square of the covariant derivative, which on the substitution \(v\to v+H\) multiplies the mass term by \(\left(1+H/v\right)^{2}\) and therefore gives a linear coupling twice the mass term over \(v\),

\begin{equation}\tag{113.24} g_{HVV}=\frac{2\left(M_{V}c^{2}\right)^{2}}{E_{v}}\ec \end{equation}

an energy rather than a pure number. Both are Equation (104.70), derived in Proposition 104.61. Neither contains a free parameter beyond \(E_{v}=246.220\,\mathrm{GeV}\), which is fixed by the muon lifetime.

The variable that puts the two families on one axis is the reduced coupling: \(g_{Hff}\) itself for a fermion and \(\sqrt{g_{HVV}/2E_{v}}=M_{V}c^{2}/E_{v}\) for a vector. In that variable every Standard Model particle lies on one straight line through the origin of slope \(1/E_{v}\), and Table 113.5 evaluates it: from \(4.29\times 10^{-4}\) for the muon to \(0.701\) for the top, a range of \(1630\), each value known to the precision of the corresponding mass and carrying no theoretical uncertainty at tree level.

What is measured is not a coupling but a rate — a product of a production cross section and a branching fraction — so a global fit over many final states is required to disentangle them. The experiments report that fit in terms of modifiers \(\kappa_{i}\) defined so that \(\kappa_{i}=1\) reproduces Equation (113.23) and Equation (113.24), with rates scaling as \(\kappa^{2}\). Every \(\kappa_{i}\) comes out consistent with unity, to roughly a tenth for \(W\), \(Z\), \(t\), \(b\) and \(\tau\) and roughly a fifth for \(\mu\) [Aad:2022] [Tumasyan:2022].

Two features make this a sharper test than a fit to six numbers suggests. First, there is nothing to adjust: a deviation anywhere along the line would falsify the mechanism rather than shift a parameter, because the parameter that would have to shift, \(E_{v}\), is independently fixed by the muon lifetime to seven significant figures and enters every point of the line at once. Second, the \(\gamma\gamma\) and \(gg\) vertices exist only at one loop (Remark 113.13), so their rates test the same couplings in a different combination — the diphoton amplitude is the \(W\) loop interfering destructively with the top loop — which constrains the relative sign of \(\kappa_{W}\) and \(\kappa_{t}\), something no tree-level rate reaches.

Remark 113.25 (Three decades, and where they stop).

The verified range runs from the muon to the top. It does not include the electron, the up, the down or the strange quark, whose reduced couplings are all below \(4\times 10^{-4}\) and none of which has been observed to couple to the scalar at all. The claim that one mechanism gives mass to all the charged fermions is therefore established for the third generation, established at three standard deviations for one second-generation lepton [Sirunyan:2021], and assumed for the first generation and for the light quarks. That is a real gap and Section 113.6.2 states what would close it.

The total width

Phenomenon 113.26 (The resonance is narrow).

The total decay width of the boson lies far below the mass resolution of either detector, so the width of the observed peak is entirely instrumental and no direct line-shape measurement is possible. Comparing \(ZZ\) production far above the resonance with production on it constrains the width to a value consistent with the predicted few megaelectronvolts [Aad:2023], and the world average is \(\hbar\Gamma_{H}=3.7\,\mathrm{MeV}\) with an asymmetric uncertainty of \(+1.9\) and \(-1.4\) [Navas:2024]. Rests on Proposition 113.14 and Equation (104.67).

Derivation. Derives Phenomenon 113.26. Two statements are needed: that the direct measurement is impossible, and that the indirect one works.

The direct measurement. The Standard Model prediction obtained by summing the partial widths of Table 104.4 is \(\hbar\Gamma_{H}=4.1\,\mathrm{MeV}\), that is \(6.57\times 10^{-13}\,\mathrm{J}\), a fraction \(3.3\times 10^{-5}\) of the mass. The best mass resolution achieved is \(1.1\,\mathrm{GeV}/c^{2}\) (Proposition 113.14), larger by a factor of \(270\). The observed peak is therefore a pure instrumental response function and carries no information about \(\Gamma_{H}\) beyond a bound of order \(1\,\mathrm{GeV}\). Equivalently, the mean lifetime is

\[ \tau=\frac{\hbar}{\hbar\Gamma_{H}} =\frac{1.054572\times 10^{-34}\,\mathrm{J}\,\mathrm{s}}{5.93\times 10^{-13}\,\mathrm{J}} =1.78\times 10^{-22}\,\mathrm{s}\ec \]

and the decay length \(c\tau=5.3\times 10^{-14}\,\mathrm{m}\), a twentieth of a picometre: there is no displaced vertex to find either. Both inherit the width's asymmetric error, \(+51\,\mathrm{\%}\) and \(-38\,\mathrm{\%}\), inverted: the width's upper end \(5.6\,\mathrm{MeV}\) gives \(\tau=1.18\times 10^{-22}\,\mathrm{s}\) and its lower end \(2.3\,\mathrm{MeV}\) gives \(2.86\times 10^{-22}\,\mathrm{s}\), so \(\tau\) is known only to a factor of about two either way. Nothing in the argument depends on which end is taken.

The off-shell method. Let \(W\) be the invariant-mass energy of the \(ZZ\) system produced through the scalar in gluon fusion. The differential cross section carries a relativistic Breit–Wigner,

\begin{equation}\tag{113.25} \frac{\dd\sigma}{\dd W^{2}}\propto \frac{\hbar\Gamma_{gg}\,\hbar\Gamma_{ZZ}} {\left[W^{2}-\left(m_{H}c^{2}\right)^{2}\right]^{2} +\left(m_{H}c^{2}\right)^{2} \left(\hbar\Gamma_{H}\right)^{2}}\ec \end{equation}

with \(\hbar\Gamma_{gg}\) and \(\hbar\Gamma_{ZZ}\) the partial widths into the initial and final states, written as energies like every other width in this chapter, so that numerator and denominator carry \(\mathrm{J}^{2}\) and \(\mathrm{J}^{4}\) respectively. Write each coupling as a modifier times its Standard Model value, so \(\hbar\Gamma_{gg}\,\hbar\Gamma_{ZZ} =\kappa_{g}^{2}\kappa_{Z}^{2}\,\hbar\Gamma_{gg}^{\text{SM}}\, \hbar\Gamma_{ZZ}^{\text{SM}}\), and consider two regions.

On the peak, \(\hbar\Gamma_{H}\ll m_{H}c^{2}\) lets the numerator be taken constant across the resonance, and

\[ \int\frac{\dd W^{2}} {\left[W^{2}-\left(m_{H}c^{2}\right)^{2}\right]^{2} +\left(m_{H}c^{2}\right)^{2}\left(\hbar\Gamma_{H}\right)^{2}} =\frac{\pi}{m_{H}c^{2}\,\hbar\Gamma_{H}}\ec \]

the standard Cauchy integral. Hence

\begin{equation}\tag{113.26} \mu_{\text{on}} :=\frac{\sigma_{\text{on}}}{\sigma_{\text{on}}^{\text{SM}}} =\frac{\kappa_{g}^{2}\kappa_{Z}^{2}} {\hbar\Gamma_{H}/\hbar\Gamma_{H}^{\text{SM}}}\ep \end{equation}

Far above the peak, \(W\gg2m_{Z}c^{2}\gg m_{H}c^{2}\), the denominator of Equation (113.25) is \(W^{4}\) up to corrections of order \((m_{H}c^{2}/W)^{2}\) and \((\hbar\Gamma_{H}/W)^{2}\), both utterly negligible: the width has dropped out. Hence

\begin{equation}\tag{113.27} \mu_{\text{off}} :=\frac{\sigma_{\text{off}}}{\sigma_{\text{off}}^{\text{SM}}} =\kappa_{g}^{2}\kappa_{Z}^{2}\ec \end{equation}

and dividing Equation (113.27) by Equation (113.26),

\begin{equation}\tag{113.28} \frac{\hbar\Gamma_{H}}{\hbar\Gamma_{H}^{\text{SM}}} =\frac{\mu_{\text{off}}}{\mu_{\text{on}}}\ep \end{equation}

The couplings cancel and the width is measured by a ratio of two rates, neither of which resolves anything.

The method is usable only because the off-shell rate is not negligible: about a tenth of all \(gg\to H^{*}\to ZZ\) production occurs above \(W=2m_{Z}c^{2}\), because the threshold opening enhances the rate and the top-quark loop amplitude keeps growing until \(W=2m_{t}c^{2}\). ATLAS reported evidence for that off-shell component and used Equation (113.28) to constrain the width to a value consistent with the prediction [Aad:2023].

Remark 113.27 (The assumptions the width bound rests on).

Equation (113.28) is exact given Equation (113.25), and three of its inputs are assumptions rather than measurements. They are stated here because the resulting constraint is often quoted as if it were a direct measurement of a lifetime, and it is not.

First, \(\kappa_{g}\) and \(\kappa_{Z}\) are taken to be the same at \(W=125\,\mathrm{GeV}\) and at \(W\gtrsim400\,\mathrm{GeV}\). Any new particle in the gluon-fusion loop, or any new contribution to the high-mass \(ZZ\) amplitude, breaks that and the extracted width is wrong — and precisely such a particle is what an anomalous width would be evidence for, so the method assumes the absence of the thing it would be used to detect.

Second, the continuum \(gg\to ZZ\) background proceeds through a box diagram with the same initial and final state and therefore interferes with the signal. The interference is destructive and large, and it must be modelled; its higher-order strong corrections are not known to the order the signal's are, and are assumed to be the same.

Third, the narrow-width approximation used in Equation (113.26) is safe here only because \(\hbar\Gamma_{H}/m_{H}c^{2}=3.3\times 10^{-5}\); that part of the derivation is not in doubt.

The honest summary is that the width is constrained, under stated model assumptions, to a value consistent with \(4.1\,\mathrm{MeV}\), and that no assumption-free measurement of it exists or is planned at a hadron collider.

What the discovery does and does not establish

It is worth separating the four claims, because they are established to very different degrees and are routinely run together.

A neutral scalar exists at \(125.20(11)\,\mathrm{GeV}/c^{2}\). Established, by two independent instruments, in two independent channels each, with the mass measured consistently and the rate as predicted. This is as secure as any result in particle physics.

Its quantum numbers are \(J^{P}=0^{+}\). Established: spin \(1\) by Landau–Yang from the existence of the diphoton decay, spin \(0\) over spin \(2\) and even over odd parity by the four-lepton angular analysis [Aad:2013] [Chatrchyan:2013]. A small \(CP\)-odd admixture is bounded, not excluded (Remark 113.23).

Its couplings are proportional to mass. Established over the range from the muon to the top — a factor of \(1630\) — to a precision of roughly ten per cent, with the first generation and the light quarks untested (Remark 113.25).

The potential is the quartic one of Equation (113.1). Not established. Everything above probes the theory at the minimum of the potential. Its shape away from the minimum enters only through the self-couplings, and those are bounded to within a factor of several of the predicted value and not measured (Section 113.6.1). The central postulate of the mechanism — that the vacuum sits at the minimum of a quartic — remains an assumption supported by the mass and the couplings, and not a measurement.

What is not yet measured

The self-coupling and the shape of the potential

Expanding Equation (113.1) in unitary gauge gives, by Proposition 104.63,

\begin{equation}\tag{113.29} V(h)=\frac{1}{2}\left(\frac{m_{H}c}{\hbar}\right)^{2}h^{2} +\frac{\left(m_{H}c\right)^{2}}{2\hbar^{2}v}h^{3} +\frac{\left(m_{H}c\right)^{2}}{8\hbar^{2}v^{2}}h^{4}\ec \end{equation}

so the trilinear and quartic self-couplings are fixed by \(m_{H}\) and \(E_{v}\), both measured. Writing the trilinear coupling as \(\kappa_{\lambda}\) times its Standard Model value, the mechanism predicts \(\kappa_{\lambda}=1\) with nothing to adjust. Nobody has measured it.

The reason is the size of the only channel that carries it. Producing two scalars at once, \(gg\to HH\), proceeds through two interfering amplitudes: a triangle in which a top loop makes one off-shell scalar that then splits, carrying \(\kappa_{\lambda}\) at the splitting vertex, and a box in which two scalars are emitted from the same top loop, carrying no self-coupling at all. The two interfere destructively — a low-energy theorem forces them to cancel at threshold in the limit of a heavy top — so the cross section is a quadratic in \(\kappa_{\lambda}\) with a minimum inside the range of interest rather than a monotone function of it. Three consequences follow, and all three are unhelpful.

First, the rate is tiny: some three orders of magnitude below single production, tens of femtobarns at LHC energies against the tens of picobarns of Section 113.3.1. Second, because the dependence is quadratic with a minimum inside the interesting range, a measured rate maps onto two values of \(\kappa_{\lambda}\), so the constraint is naturally an interval with a possible second solution rather than a value. Third, the destructive interference means the observable loses sensitivity in the neighbourhood of the predicted value rather than gaining it there.

The present constraints [Aad:2022] [Tumasyan:2022] allow \(\kappa_{\lambda}\) over a range of order \(-1\) to \(+7\): consistent with the predicted \(1\), and far too loose to be called a test.

Derivation pending.

The double-production cross section \(\sigma(gg\to HH)\) as a function of \(\kappa_{\lambda}\): the interference of the top-loop triangle carrying the trilinear vertex with the top-loop box that does not, giving the quadratic whose minimum locates the least sensitive value of \(\kappa_{\lambda}\), and the normalization that puts the rate three orders of magnitude below single production. The location of the minimum and the absolute rate are stated in this subsection on the authority of the experimental compilations and are not derived here; the two loop amplitudes belong in Appendix A with the single-production ones already recorded as owed. The published constraint on \(\kappa_{\lambda}\) quoted below is a measurement and needs no derivation.

The high-luminosity programme is expected to reduce that to a band of order tens of per cent around the prediction, which would be the first real measurement of the shape of the potential; that is a projection, not a result, and it is recorded here as one.

Remark 113.28 (What is actually being assumed).

It is worth being precise about what is unmeasured, because the phrase “the Higgs potential” is used for two different things. That the electroweak vacuum breaks the gauge symmetry and sits at some nonzero field value is established — it is what the boson masses of Phenomenon 113.6 measure, and what \(E_{v}\) is. That the function whose minimum it sits at is \(-\mu^{2}\abs{\Phi}^{2}+\lambda\abs{\Phi}^{4}\), with the single coupling \(\lambda\) generating the mass and both self-interactions, is the minimal hypothesis and is unmeasured. Any potential with a minimum in the right place and the right curvature reproduces every result in Section 113.5; only the self-couplings distinguish them. Loop corrections make the point sharper: the object whose minimum determines the vacuum is not the tree-level potential but the effective potential of Coleman and Weinberg [Coleman:1973], whose shape differs (Remark 104.64).

First-generation couplings and rare decays

The scalar has been seen to couple to the top, the bottom, the tau and — at three standard deviations — the muon [Sirunyan:2021]. It has not been seen to couple to the electron, the up quark, the down quark or the strange quark, which between them account for essentially all the mass of ordinary matter that is not binding energy. The claim that one mechanism gives mass to all the fermions is therefore verified for the third generation, for one second-generation lepton, and nowhere else.

Proposition 113.29 (Why the electron coupling is out of reach).

The Standard Model branching fraction to electrons is

\begin{equation}\tag{113.30} \mathcal{B}\left(H\to e^{+}e^{-}\right) =\mathcal{B}\left(H\to\mu^{+}\mu^{-}\right) \left(\frac{m_{e}}{m_{\mu}}\right)^{2} =2.2\times 10^{-4}\times2.34\times 10^{-5}=5.1\times 10^{-9}\ec \end{equation}

so even the full high-luminosity dataset would contain of order one such decay, against a Drell–Yan background of order \(10^{9}\) events in the same mass window. The measurement is not difficult; it is impossible at this machine. Rests on Equations (104.67) and (113.23).

Derivation. Derives Proposition 113.29. By Equation (104.67) the partial width into a fermion pair is proportional to \(\hat{y}_{f}^{2}\) times a phase-space factor which is unity to one part in \(10^{5}\) for both leptons, and \(\hat{y}_{f}\propto m_{f}\) by Equation (113.23). Hence the ratio of branching fractions is \((m_{e}/m_{\mu})^{2}\). With \(m_{e}c^{2}=0.51099895069\,\mathrm{MeV}\) and \(m_{\mu}c^{2}=105.6583755\,\mathrm{MeV}\) [Mohr:2025] [Navas:2024] the ratio is \((4.8363\times 10^{-3})^{2}=2.339\times 10^{-5}\), and \(2.2\times 10^{-4}\times2.339\times 10^{-5}=5.1\times 10^{-9}\), which is Equation (113.30).

The high-luminosity programme aims at \(3\times 10^{46}\,/\mathrm{m}^{2}\), three thousand inverse femtobarns per experiment. At \(\sqrt{s}\,c=13\,\mathrm{TeV}\) the total production cross section is about \(5.5\times 10^{-39}\,\mathrm{m}^{2}\), so \(N_{H}\approx3\times 10^{46}\,/\mathrm{m}^{2}\times 5.5\times 10^{-39}\,\mathrm{m}^{2}=1.7\times 10^{8}\) scalars, of which \(1.7\times 10^{8}\times5.1\times 10^{-9}=0.87\) decay to an electron pair. One event, with an acceptance below unity, on a background of Drell–Yan electron pairs whose cross section in the same mass window exceeds the signal by nine orders of magnitude. No selection recovers that. The light quarks are worse still, because there is no way to tag a light-quark jet as up rather than down, so even the final state cannot be identified.

What is achievable is a bound. An enhanced electron Yukawa coupling, hundreds of times the Standard Model value, would produce a visible \(H\to e^{+}e^{-}\) peak, and searches set limits at that level. A bound three orders of magnitude above the prediction is worth recording as a bound and is not a test of the prediction.

Three other classes of measurement are open and are being pursued.

Loop-level rare decays. \(H\to Z\gamma\) has a predicted branching fraction of \(1.5\times 10^{-3}\) (Table 104.4) and, like \(H\to\gamma\gamma\), no tree-level vertex: it proceeds through \(W\) and top loops in a combination different from the diphoton one, so measuring both constrains the loop content in two independent ways. Evidence for it has been reported by the two collaborations combined, at a significance near the evidence threshold; the treatise's bibliography does not currently hold that publication, and the claim is therefore stated here without a citation rather than attached to the wrong one.

Lepton-flavour-violating decays. \(H\to\mu\tau\) and \(H\to e\tau\) are forbidden in the Standard Model, where the Yukawa matrix and the mass matrix are diagonalized by the same transformation (Theorem 104.38) so that no flavour-changing neutral scalar current exists at tree level. Searches have found nothing and bound the branching fractions at the level of a part in a thousand. This is a null result and it is a meaningful one: it says the scalar's couplings are flavour-diagonal to that accuracy, which the mechanism requires and a general two-doublet sector would not.

Invisible decays. The Standard Model predicts \(\mathcal{B}(H\to ZZ^{*}\to4\nu)\approx1.1\times 10^{-3}\), which is one line of Table 104.4: the branching fraction to \(ZZ^{*}\) is \(0.026\) and each \(Z\) goes to neutrinos with probability \(0.20\), so \(0.026\times0.20^{2}=1.04\times 10^{-3}\), the small remainder coming from the off-shell boson's distorted phase space. Any additional invisible width would signal decays to particles the detector cannot see, so the bound on it — currently of order ten per cent of the total width — constrains a light unseen sector coupling to the scalar. Collecting those searches, together with the direct and indirect ones, belongs to The Dark Sector: Evidence Without Explanation, whose collider subsection is reserved for exactly this material and does not yet contain it.

Vacuum stability

The measured mass has one consequence that reaches far beyond the collider, and it is the most-quoted and least-secure statement in this chapter.

\(\lambda\) is a coupling, so it runs (The Renormalization Group). Its one-loop beta function Equation (104.72) is dominated by the top Yukawa through the term \(-6\hat{y}_{t}^{4}\), negative because it descends from a closed fermion loop, and numerically fifty times larger than the bosonic terms. Proposition 104.66 evaluates it at a renormalization scale \(\mu_{R}=m_{t}c^{2}\) — the scale of Notation 113.1, not the potential's mass parameter and not a signal strength — and finds \(\dd\hat{\lambda}/\dd\ln\mu_{R}=-0.0206\): the quartic coupling falls with scale. Integrating the coupled system upward, with the measured \(m_{H}c^{2}=125.20(11)\,\mathrm{GeV}\) and \(m_{t}c^{2}=172.57(29)\,\mathrm{GeV}\) as inputs, \(\hat{\lambda}\) crosses zero somewhere near \(10^{10}\text{–}10^{11}\,\mathrm{GeV}\) [Degrassi:2012] [Buttazzo:2013]. Above that scale the potential turns over and a deeper minimum exists at large field values: the electroweak vacuum is not the global minimum but a metastable one, with a computed tunnelling lifetime exceeding the age of the universe by hundreds of orders of magnitude.

This chapter records the result and, more importantly, records what it is worth. Remark 104.67 states the three caveats in full and none may be dropped: the conclusion sits on a knife edge in the top mass, where a shift of about \(1\,\mathrm{GeV}\) moves the answer between absolute stability, metastability and instability, and the relation between what a hadron collider reconstructs and the short-distance mass entering the beta function carries a theoretical uncertainty of that size; the extrapolation assumes that nothing couples to the scalar anywhere between \(10^{3}\,\mathrm{GeV}\) and the \(10^{10}\,\mathrm{GeV}\) at which the quartic changes sign — seven decades of energy in which nothing has been looked at, and the extrapolation is only worth as much as that assumption; and the running is a statement about a coupling in a renormalization scheme, not about an observable.

That interval is seven decades wide and not sixteen: \(\log_{10}(10^{10}/10^{3})=7\). Remark 104.67 names the same two endpoints and calls them sixteen, which is an arithmetic slip there; the count used here is the one the endpoints give.

The experimental content of this subsection is therefore narrow and should be stated as such: this chapter measured \(m_{H}\), and Experiment: Deep Inelastic Scattering and its successors measured \(m_{t}\) and \(\alpha_{s}\), and those measurements are the inputs to a calculation whose output is interesting and whose error bar is dominated by an ambiguity in one of its inputs. It is a striking near-coincidence that the Standard Model, extrapolated with no new physics, lands so close to the boundary. It is not evidence for anything.

The related question — why \(E_{v}=246.220\,\mathrm{GeV}\) at all, when nothing in the theory protects a fundamental scalar's mass from corrections scaling with whatever completes it at high energy — is not a question this chapter can address, and it is recorded among the open problems in Section 133.2. The discovery made it sharper rather than softer: before 2012 one could hope the electroweak sector was not built from an elementary scalar; it is.

Primary references

The experiment is two papers, submitted on the same day to the same journal and published back to back. ATLAS [Aad:2012] and CMS [Chatrchyan:2012] each report the diphoton and four-lepton analyses in full, with the supporting \(WW^{*}\), \(\tau\tau\) and \(b\bar{b}\) channels, the categorization, the background models and the statistical treatment. Both are readable end to end and both should be read in preference to any account of them, including this one; the numbers in Table 113.3 attributed to 2012 come from them and are not recomputed here.

The instruments have their own literature and it is where the apparatus section comes from. The ATLAS [Aad:2008] and CMS [Chatrchyan:2008] technical papers in the 2008 LHC machine special issue describe the detectors as built, subsystem by subsystem, with the design resolutions quoted in Equation (113.13) and Equation (113.14); Evans and Bryant [Evans:2008] in the same issue describe the accelerator, and it is the source for every machine parameter in Table 113.2 other than the two derived rows.

The statistical machinery has a small canonical literature and this chapter uses all of it. Cowan, Cranmer, Gross and Vitells [Cowan:2011] fix the asymptotic formulae, the one-sided discovery statistic and the Asimov construction that the collaborations use to quote expected sensitivities; Read [Read:2002] is the \(CL_{s}\) construction behind every exclusion limit quoted here, including LEP's; Wilks [Wilks:1938] is the theorem that makes \(Z=\sqrt{q}\) possible and Gross and Vitells [Gross:2010] the trials-factor practice applied in Proposition 113.19, resting in turn on Davies' bound [Davies:1977] [Davies:1987] and on Rice's level-crossing formula [Rice:1944]. All of it is derived in Probability and Statistics rather than quoted.

The theory the experiment tested was in print half a century before the data. Englert and Brout [Englert:1964], Higgs [Higgs:1964a] [Higgs:1964] [Higgs:1966] and Guralnik, Hagen and Kibble [Guralnik:1964] give the mechanism in 1964, with Kibble's non-abelian extension [Kibble:1967] supplying the case the Standard Model needs and Anderson's plasmon argument [Anderson:1963] the physical precedent; Goldstone [Goldstone:1961] and Nambu [Nambu:1960] give the theorem that has to be evaded and the superconducting analogy that shows how. Glashow [Glashow:1961], Weinberg [Weinberg:1967] and Salam [Salam:1968] assemble the electroweak theory around it, and Electroweak Unification and the Higgs Boson derives the whole of it; Weinberg's second volume [Weinberg:1996] is the standard reference treatment.

The searches that preceded the discovery are the LEP combination of ALEPH, DELPHI, L3 and OPAL [Barate:2003], which set the lower limit and reported the excess at the edge of its sensitivity, and the Tevatron combination of CDF and D0 [Aaltonen:2012], which reported the associated-production excess in \(b\bar{b}\). The indirect constraints are the LEP and SLD precision electroweak measurements [Schael:2006] and the global fits built on them [Baak:2014], together with Veltman's screening theorem [Veltman:1977], which is why the constraint on the scalar mass was logarithmic and therefore weak.

The measurements made since 2012 divide by what they test. The spin-parity analyses are ATLAS [Aad:2013] and CMS [Chatrchyan:2013]; the first joint mass combination is [Aad:2015] and the first joint coupling fit [Aad:2016]; the ten-year coupling maps of the full Run-2 dataset are ATLAS [Aad:2022] and CMS [Tumasyan:2022], which are the source for every \(\kappa\) statement in Tables 113.5 and 113.6 and for the self-coupling range in Section 113.6.1. Note that [Tumasyan:2022] is the CMS ten-year portrait and not a discovery paper, and it is cited here only for couplings. Individual channel observations are \(t\bar{t}H\) production [Sirunyan:2018] and the \(H\to\mu^{+}\mu^{-}\) evidence [Sirunyan:2021]; the off-shell width constraint of Section 113.5.5 is [Aad:2023]. Vacuum stability is Degrassi et al. [Degrassi:2012] and its sequel [Buttazzo:2013].

Two sources are used for numbers rather than for claims. The Particle Data Group's 2024 review [Navas:2024] supplies every world-average mass and width in this chapter, taken from its own machine-readable release and stored with the treatise; CODATA 2022 [Mohr:2025] supplies the exact defining constants used in every conversion between electronvolts, joules and kilograms.

Finally, one gap should be named rather than papered over. The Standard Model production cross sections quoted in Section 113.3.1 are the compilations of the LHC Higgs Cross Section Working Group, which the collaborations cite and which this treatise's bibliography does not hold; they are attributed above to the two discovery papers, which quote and use them. The same applies to the ATLAS and CMS combination reporting evidence for \(H\to Z\gamma\), referred to in Section 113.6.2 without a citation for the same reason.