Cosmic Rays and Astroparticle Physics
Before there were accelerators there was the sky. Every charged particle species known in 1950 — the positron, the muon, the pion and the strange mesons — was first seen in the cosmic radiation, and the highest laboratory-confirmed particle energies still come from it and not from any machine: the atmosphere is struck by nuclei above \(10^{20}\,\mathrm{eV}=16\,\mathrm{J}\), some seven orders of magnitude beyond the beam energy of the collider of Experiment: The Higgs Boson Discovery. This chapter closes Part XI — Quantum Field Theory and the Standard Model by treating that natural beam as what it is — a measured flux, with a spectrum, a composition, an arrival-time structure and an angular distribution, all quoted here with their uncertainties. It covers the discovery of the radiation and the geomagnetic experiments that showed the primaries are charged and predominantly positive; the particles found in it, which feed Experiment: The Positron and Experiment: Time Dilation and Relativistic Kinematics; the primary spectrum with its knee and ankle; extensive air showers and the detection techniques built on them; the astrophysical acceleration mechanisms that are consistent with the data; the Greisen–Zatsepin–Kuz'min cutoff and its observed counterpart; and the neutral messengers — gamma rays and neutrinos — that point back to their sources where charged nuclei cannot.
Its neighbours are the Standard Model chapters it uses as input (Quantum Chromodynamics, Nuclear Forces and Nuclear Structure and Flavour Physics and Neutrinos) and the astrophysics it feeds (Stellar Structure and Nucleosynthesis and Evidence-Based Cosmology). The standing monograph is Gaisser, Engel and Resconi [Gaisser:2016]; current flux and composition numbers are taken from the Particle Data Group review [Navas:2024]. Nothing in this chapter is inferred from a mechanism that has not been observed: where the origin of a component is unknown — as it is above the ankle — the chapter says so.
Units and conventions are those fixed once for the whole part in Section 100.1.1: the metric \(\eta_{\mu\nu}=\diag(+1,-1,-1,-1)\) of Notation 100.1, the mass shell \(p^{2}=m^{2}c^{2}\) of Equation (100.1), and \(\hbar\) and \(c\) written out in every formula and every intermediate step. Nothing below is computed in units where they are set to one; Remark 100.4 gives the dictionary for reading this chapter against the literature, which does. Two conventions of the field need saying at the outset because they are not SI and are used on every page of it.
The rigidity of a particle of momentum \(p\) and charge \(Ze\) is
an SI voltage, measured in volts. It is the natural variable for magnetic problems because a magnetic field bends two particles of the same rigidity along the same path whatever their mass or charge: the Larmor radius of Equation (115.8) depends on \(p\) and \(Ze\) only through \(R\). Rigidities in this chapter are quoted in gigavolts and in volts.
The atmospheric depth \(X\) at a point is the mass of air per unit area lying above it along the direction of travel,
an SI areal mass density in \(\mathrm{kg}/\mathrm{m}^{2}\). The literature universally quotes it in \(\mathrm{g}/\mathrm{cm}^{2}\), and \(1\,\mathrm{g}/\mathrm{cm}^{2}=10\,\mathrm{kg}/\mathrm{m}^{2}\) exactly; both are given here. Its vertical value at sea level follows from hydrostatic equilibrium alone: the whole weight of the column is carried by the surface pressure, so
Every shower in Section 115.4 develops in that calorimeter and in no other.
Discovery of the cosmic radiation
The penetrating radiation
The instrument that found cosmic rays was the gold-leaf electroscope, and understanding what it measures is the whole of the argument. A charged conductor inside a sealed vessel of gas loses its charge because the gas is not a perfect insulator: ionising radiation creates free electron–ion pairs, the field between the conductor and the vessel wall sweeps them apart before they recombine, and the resulting current discharges the leaves. If the field is strong enough that essentially every pair created is collected — the saturation regime — the current no longer depends on the field, and counts ionisation events directly:
with \(q\) the number of ion pairs produced per unit volume per unit time, \(V\) the volume of the vessel and \(e\) the elementary charge. The observable is \(q\), an SI rate density in \(/\mathrm{m}^{3}/\mathrm{s}\), though the period quoted it as ion pairs per cubic centimetre per second and this chapter gives both.
The numbers involved show why the measurement was difficult and why it was nevertheless decisive. A sealed litre vessel at sea level, with every known source of contamination removed, still shows \(q\approx10^{7}\,/\mathrm{m}^{3}/\mathrm{s}\), that is about ten ion pairs per cubic centimetre per second, and by Equation (115.4) this is a current
a femtoampere. No galvanometer of 1910 could see such a current; an electroscope, which integrates it on a capacitance of a few picofarads, turns it into a leaf deflection of a few volts per hour and sees it easily. That is the instrumental fact behind the whole discovery.
By 1900 the residual conductivity of sealed air was known to Wilson, to Elster and Geitel, and to Rutherford, and the standing explanation was gamma radiation from radioactive material in the ground and in the walls of the apparatus. The explanation is testable, because a gamma ray is attenuated exponentially in the mass it traverses. For photons of the few \(\mathrm{MeV}\) that the hardest radium gamma rays carry, the mass attenuation coefficient of air is \(\mu/\rho\approx4\times 10^{-3}\,\mathrm{m}^{2}/\mathrm{kg}\), so at the ground-level density \(\rho\approx1.2\,\mathrm{kg}/\mathrm{m}^{3}\) the linear coefficient is \(\mu\approx5\times 10^{-3}\,/\mathrm{m}\) and the attenuation length is
Wulf therefore carried an electroscope of his own design to the top of the Eiffel Tower, \(300\,\mathrm{m}\) above the ground, where a purely terrestrial source demanded a reduction by
that is to about a fifth of the ground value. What he measured was a reduction to roughly sixty percent [Wulf:1910] — nearly three times the ionisation the terrestrial hypothesis allows. A factor of three is a large discrepancy, but one tower, one instrument and an attenuation coefficient known only to a factor of two could not settle the question, and Wulf did not claim that it had.
Hess settled it by going higher. In seven balloon ascents in 1911 and 1912, the last reaching \(5350\,\mathrm{m}\), he carried three sealed electroscopes of different wall thickness and found that the ionisation first fell with altitude, through a shallow minimum near \(1500\,\mathrm{m}\), and then rose, reaching about twice its sea-level value near \(5000\,\mathrm{m}\) [Hess:1912]. He drew the only available conclusion: a radiation of very great penetrating power enters the atmosphere from above. He also excluded the obvious candidate for a source above, flying at night and during the solar eclipse of 17 April 1912 and finding no reduction — so the Sun is not the source of the penetrating component. Kolhörster extended the measurement to \(9200\,\mathrm{m}\) and found the ionisation about ten times the ground value [Kolhoerster:1913], which fixed the effect beyond argument.
The rate at which a sealed electroscope loses its charge, after every known terrestrial contamination has been excluded, first falls with height above the ground and then rises. In the balloon ascents of 1911–1912 the ionisation passed through a minimum near \(1500\,\mathrm{m}\) and by about \(5000\,\mathrm{m}\) was roughly twice its sea-level value [Hess:1912]; at \(9200\,\mathrm{m}\) it was some ten times the ground value [Kolhoerster:1913]. Flights by night and during a solar eclipse gave the same result, which excludes the Sun as the source [Hess:1912] [Wulf:1910]. Rests on Equation (115.2).
Derivation. Derives Phenomenon 115.2. Let the ionising agent be attenuated exponentially in the mass of absorber traversed, \(I=I_{0}\exp\left(-X/\lambda\right)\), with \(X\) the column density of Equation (115.2) and \(\lambda\) an attenuation length; this is the behaviour of any penetrating radiation with a fixed interaction cross section per unit mass. Two sources are in play and they see opposite overburdens.
For a source in the ground the absorber between source and instrument is the air below the instrument, whose column density increases with height; write it \(\ell(h)\) with \(\ell(0)=0\) and \(\dd\ell/\dd h=\rho(h)>0\). For a source above the atmosphere the absorber is the depth \(X(h)\) of Equation (115.2), which decreases with height, \(\dd X/\dd h=-\rho(h)<0\). The measured rate is the sum,
and differentiating,
The bracket is negative at \(h=0\) whenever the terrestrial term dominates there, and it grows monotonically with \(h\) because the first exponential decreases and the second increases. It therefore has exactly one zero, at the altitude where
and that zero is a minimum. So the observed shape — fall, minimum, rise — is not merely consistent with an external source; it is what an external source superposed on a terrestrial one must produce, and a purely terrestrial source cannot produce it at all, since for \(q_{c}=0\) the bracket is negative everywhere and \(q\) falls monotonically. The two attenuation lengths are then read off the two limbs of the curve, and the upper limb gives a \(\lambda_{c}\) far larger than any then-known gamma radiation possessed, which is the sense in which the radiation is “penetrating”.
∎Phenomenon 115.2 establishes that something enters the atmosphere from outside. It says nothing about what that something is: an exponential absorption curve is produced by a photon beam, an electron beam and a proton beam alike, and the only parameter it yields is \(\lambda_{c}\). This is why the question of the nature of the primaries stayed open for two decades after 1912 and was settled not by absorption but by geomagnetism (Sections 115.1.3 and 115.1.4). It is also worth recording that the rise does not continue indefinitely: above about \(20\,\mathrm{km}\) the ionisation passes through a maximum, the Pfotzer maximum, and falls again, because what the electroscope records at any depth is not the primary but the shower it has made — the subject of Section 115.4.
Naming and the extraterrestrial origin
Millikan and Cameron settled the extraterrestrial origin by a different route and, in doing so, named the subject. Their idea was to use water as the absorber and mountain lakes as the vessel: an instrument lowered to depth \(d\) in a snow-fed lake at altitude \(h\) sits under a total column density
so that one metre of water is worth \(10^{3}\,\mathrm{kg}/\mathrm{m}^{2}=100\,\mathrm{g}/\mathrm{cm}^{2}\), about a tenth of the whole atmosphere by Equation (115.3). Measuring in Muir Lake, at \(3540\,\mathrm{m}\), and in Arrowhead Lake, at \(1530\,\mathrm{m}\), they found that the two ionisation curves, plotted against \(X_{\mathrm{tot}}\) rather than against depth in water, fell on one another [Millikan:1926]. That is exactly what Equation (115.5) predicts for the extraterrestrial term and it is what no terrestrial source can produce, because a terrestrial source has no reason to care about the air above the lake. They called the agent cosmic rays, and the name stuck.
Millikan drew a further conclusion that was wrong, and the way it was wrong is instructive. He read the attenuation curve as that of a gamma-ray beam and identified the photon energies with the binding energies released in the synthesis of helium and heavier elements from hydrogen — the “birth cries of atoms”. The inference from an absorption curve to a photon energy requires a cross section, the cross section assumed was Compton scattering (Radiation and Scattering of Electromagnetic Waves), and the whole chain collapses if the primaries are not photons. Compton argued from the outset that they were charged particles. The dispute could not be settled by absorption measurements of any precision, for the reason given in Remark 115.3: attenuation curves measure \(\lambda\), and \(\lambda\) alone does not identify a species. What settled it was an argument in which the charge appears explicitly, and the Earth supplies the apparatus.
The latitude effect
Clay, making ionisation measurements on the voyages between Java and the Netherlands, found the intensity systematically lower near the equator than at high latitude, by some \(10\,\mathrm{\%}\) [Clay:1927]. Compton organised a worldwide survey with a single standardised ionisation chamber design at sixty-nine stations on every continent, and showed that the variation is ordered not by geographic latitude but by geomagnetic latitude \(\lambda_{m}\), with the intensity flat above about \(\abs{\lambda_{m}}=40\,^\circ\) and falling towards the geomagnetic equator [Compton:1933]. A magnetic ordering can only mean one thing.
The cosmic-ray intensity measured at sea level falls by of order \(10\,\mathrm{\%}\) between geomagnetic latitude \(40\,^\circ\) and the geomagnetic equator, and the ordering parameter is the geomagnetic and not the geographic latitude [Clay:1927] [Compton:1933]. The primaries are therefore electrically charged, and the fall sets the scale of their momenta. Rests on Equation (115.1).
Derivation. Derives Phenomenon 115.4. A photon is undeflected by a magnetic field, so a photon primary produces no latitude effect of any size; the observation of one excludes photons as the dominant component. For a charged primary, the Earth's field acts as a momentum filter, and the scale of the filter follows from a single comparison of lengths.
Model the geomagnetic field as a dipole of moment \(m\), so that in the equatorial plane at distance \(r\) from the centre
the value of \(M\) being fixed by the measured surface field at the magnetic equator, \(B\left(R_{\oplus}\right)\approx3.1\times 10^{-5}\,\mathrm{T}\) with \(R_{\oplus}=6.371\times 10^{6}\,\mathrm{m}\). A particle of momentum \(p\) and charge \(Ze\) moving perpendicular to a field \(B\) turns on a circle of radius
with \(R\) the rigidity of Equation (115.1); the second form shows that \(r_{L}\) depends on the particle only through \(R\), which is the whole reason rigidity is used. A particle can reach the Earth from outside only if its trajectory is not turned round before it arrives, which requires the turning radius in the field it crosses to be at least of order the distance over which that field acts. Setting \(r_{L}=r\) with \(B=M/r^{3}\) gives the rigidity scale
that is \(60\,\mathrm{GV}\). The exact treatment is Störmer's: the motion of a charge in a dipole field has a first integral, and it partitions directions at each point into an allowed cone, from which particles of a given rigidity can arrive, and a forbidden one. The resulting cutoff rigidity for vertical incidence at geomagnetic latitude \(\lambda_{m}\) is [Lemaitre:1933]
which is Equation (115.9) with the numerical coefficient \(\tfrac{1}{4}\) that the exact integration supplies. No particle of rigidity below \(R_{c}\) arrives vertically at that latitude.
The Störmer theory itself: the first integral of the motion of a charge in a magnetic dipole field, the partition of directions at each point into an allowed and a forbidden cone that follows from it, and the integration supplying the coefficient one quarter in the vertical cutoff together with the full azimuthal form used in the east–west subsection below. Belongs in Appendix A; both formulae are quoted here from the source cited beside them, and only their consequences are derived
The latitude effect follows at once. At \(\lambda_{m}=50\,^\circ\) the cutoff is \(14.9\,\mathrm{GV}\times\cos^{4}50\,^\circ =2.5\,\mathrm{GV}\), while at the geomagnetic equator it is the full \(14.9\,\mathrm{GV}\): a detector at high latitude sees the primary spectrum down to a rigidity six times lower than one at the equator, and therefore sees more particles. The size of the effect is then a measurement of the spectrum. If the integral intensity above rigidity \(R\) behaves as \(R^{-\Gamma}\), the ratio of equatorial to polar rates is \(\left(14.9/2.5\right)^{-\Gamma}\), so a \(10\,\mathrm{\%}\) deficit requires \(\Gamma\approx0.06\) for the ionisation-chamber response — that is, the ionisation is dominated by secondaries from primaries well above both cutoffs, which is why the effect is a ten-percent effect and not a factor of six. The qualitative conclusion is independent of all of this: only a charged primary can produce a latitude effect at all.
∎The east–west asymmetry
The latitude effect shows the primaries are charged; it does not show the sign of the charge, because Equation (115.10) depends on \(Z\) only through the rigidity and is the same for positive and negative particles arriving vertically. Rossi saw that the sign appears as soon as one looks away from the vertical: for a given latitude the Störmer cutoff depends on the azimuth of arrival, and it does so with opposite senses for the two signs [Rossi:1930]. The full form of the cutoff, for arrival at zenith angle \(\theta\) and azimuth \(\xi\) measured clockwise from magnetic north — so that arrival from the east is \(\xi=90\,^\circ\) and arrival from the west is \(\xi=270\,^\circ\) — is [Lemaitre:1933]
At the geomagnetic equator and on the horizon, \(\cos\lambda_{m}=1\) and \(\sin\theta=1\), so the bracket carries \(\sin\xi=-1\) for arrival from the west and \(\sin\xi=+1\) for arrival from the east; at the zenith \(\sin\theta=0\) and the bracket is \(\left[1+1\right]^{2}=4\), which returns Equation (115.10). With \(cM/r^{2}=59.8\,\mathrm{GV}\) from Equation (115.9), this gives for a positively charged primary \(R_{c}=10.2\,\mathrm{GV}\) from the west and \(59.6\,\mathrm{GV}\) from the east, against \(14.9\,\mathrm{GV}\) from the zenith; for a negative primary the two horizontal values are exchanged. Since the spectrum falls with rigidity, the lower cutoff is the more intense direction, and the sign of the excess is the sign of the charge.
Near the geomagnetic equator the cosmic-ray intensity measured with a counter telescope at a fixed zenith angle is greater when the telescope points west than when it points east [Johnson:1933] [Alvarez:1933], as had been predicted for a flux of charged particles [Rossi:1930]. The primaries are therefore charged, and predominantly positively charged. Rests on Equation (115.11).
Derivation. Derives Phenomenon 115.5. The sign can be got without Equation (115.11), from the Lorentz force alone. Set up axes at the geomagnetic equator: \(\hat{\vect{z}}\) upward, \(\hat{\vect{x}}\) east, \(\hat{\vect{y}}\) north. There the geomagnetic field is horizontal and points north, \(\vect{B}=B\hat{\vect{y}}\) with \(B\) of order \(3\times 10^{-5}\,\mathrm{T}\). A primary of charge \(q\) arriving vertically has velocity \(\vect{v}=-v\hat{\vect{z}}\), and feels
since \(\hat{\vect{z}}\times\hat{\vect{y}}=-\hat{\vect{x}}\). For \(q>0\) the force is eastward: as the particle descends, its velocity is deflected towards the east, so that it reaches the ground travelling east of vertical — that is, arriving from the west. Positive primaries therefore produce a western excess and negative primaries an eastern one, and the sign of the observed excess fixes the sign of the charge. A neutral primary, which was still the alternative hypothesis in 1933, produces no azimuthal asymmetry whatever. The measurement thus decides between the two hypotheses without any assumption about the energy spectrum or the composition, which is exactly what the absorption curves of Section 115.1.2 could not do.
The size of the asymmetry follows from the cutoffs. Writing the integral intensity above rigidity \(R\) as \(J(>R)\propto R^{-\Gamma}\), the west-to-east ratio at the horizon is
so that even a modest \(\Gamma\) gives a large effect: the observed excess at Mexico City was of order \(10\,\mathrm{\%}\) at moderate zenith angle [Johnson:1933], growing towards the horizon, and it was found independently and with the same sign by Alvarez and Compton [Alvarez:1933]. The primaries are positive. They were later identified as protons and heavier nuclei, the composition measurement being the subject of Section 115.3.4.
∎Solar modulation and geomagnetic transients
The geomagnetic field is not the only magnetic filter between a Galactic source and a detector on Earth. The solar wind carries the solar magnetic field outward through the heliosphere, drawn into an Archimedean spiral by the Sun's rotation, and cosmic rays entering against that outflow must diffuse upstream. The consequence is observed directly: the intensity of low-energy cosmic rays is anticorrelated with the eleven-year sunspot cycle, being highest at solar minimum, and it drops sharply — by a few percent within hours — after a solar disturbance, recovering over the following days. Forbush established the transient decreases from continuously running ionisation chambers [Forbush:1937], and they carry his name; the eleven-year anticorrelation he drew out of the same network over the two decades that followed [Forbush:1954].
The quantitative description is a transport equation for the phase-space density \(f(\vect{r},p,t)\) of cosmic rays, combining outward convection with the wind at speed \(u\), spatial diffusion, and the adiabatic energy loss a particle suffers in an expanding flow. Its spherically symmetric steady state, when diffusion and convection balance, has a first integral of great practical use: the intensity observed at \(1\,\mathrm{au}\) is that of the unmodulated interstellar spectrum evaluated at a higher energy, the particle having paid a toll
on the way in, with \(\Phi\) a modulation potential in volts that absorbs all the transport physics into one number. Measured values of \(\Phi\) range from about \(400\,\mathrm{MV}\) at solar minimum to above \(1\,\mathrm{GV}\) at solar maximum.
Derivation of the force-field intensity. Liouville's theorem, applied along the characteristics of the transport equation, states that the phase-space density is unchanged, so \(f_{\oplus}(p)=f_{\mathrm{LIS}}(p_{\mathrm{LIS}})\) where the subscript denotes the local interstellar value and the two momenta are connected by Equation (115.12). The differential intensity per unit kinetic energy \(T\) is related to \(f\) by the phase-space volume element: the number of particles per unit volume with momentum in \(\dd^{3}p\) is \(f\,\dd^{3}p=4\pi p^{2}f\,\dd p\), and \(J(T):=p^{2}f\) up to constants, since \(\dd T=v\,\dd p\). Hence
the ratio being \(p^{2}/p_{\mathrm{LIS}}^{2}\) written out with the relativistic relation \(p^{2}c^{2}=T\left(T+2mc^{2}\right)\) of Section 40.2. For a proton of \(T=100\,\mathrm{MeV}\) and \(\Phi=500\,\mathrm{MV}\) the factor is severe — the observed intensity is a small fraction of the interstellar one — while for \(T\gg Ze\Phi\) it tends to unity. This is why the spectrum below a few \(\mathrm{GeV}\) per nucleon measured at Earth is not a source spectrum and must not be read as one, and why the acceleration arguments of Section 115.5 are tested above rather than below that energy.
∎The magnetohydrodynamics that produces the spiral field and the turbulent scattering that produces the diffusion coefficient belong to Plasmas and Magnetohydrodynamics, whose headings are reserved and which does not yet carry either; nothing in the present chapter derives them either, and no statement below rests on them. What is used here is only the empirical fact that the modulation exists, that it is measured by neutron monitors as a function of time, and that it is negligible above about \(10\,\mathrm{GeV}\) — which is where every spectral statement below begins.
Cosmic rays as the first particle accelerator
Track detectors
Everything in this section is a statement about a photograph. The instruments that made the photographs are two, and each measures a different pair of quantities.
Wilson's expansion cloud chamber [Wilson:1912] makes a track visible by supersaturating a vapour and letting it condense on the ions left along a charged particle's path. Placed between the poles of a magnet, it measures momentum and sign at once: the projected track is an arc of a circle, and by Equation (115.8) its radius of curvature \(\rho\) gives
in SI a momentum in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\) when \(B\) is in tesla and \(\rho\) in metres. The product \(B\rho\), an SI quantity of dimension \(\mathrm{T}\,\mathrm{m}\), is the historical observable and is quoted in the papers of the period in gauss-centimetres. One gauss-centimetre is \(10^{-4}\,\mathrm{T}\times10^{-2}\,\mathrm{m}=10^{-6}\,\mathrm{T}\,\mathrm{m}\), so \(10^{5}\) of them make \(0.1\,\mathrm{T}\,\mathrm{m}\); Anderson's \(2.1\times 10^{5}\) gauss-centimetres below is \(0.21\,\mathrm{T}\,\mathrm{m}\). The sense of the curvature gives the sign of \(Ze\) once the direction of travel is known, and fixing that direction is the hard part of every one of the discoveries below.
The second instrument is the nuclear emulsion: a thick photographic layer, exposed for weeks at mountain altitude and developed, in which the track is a chain of silver grains. It measures ionisation and range rather than momentum. The grain density along the track follows the specific ionisation, which for a fast particle behaves as
so a heavily ionising track is either slow or multiply charged, and the two cases are separated by following the track to its end: the range of a particle that stops in the emulsion is a single-valued function of its initial energy for each species, and the combination of range and grain density therefore determines mass and charge. The derivation of Equation (115.15) from the electromagnetic interaction of a charge with atomic electrons belongs to Radiation and Scattering of Electromagnetic Waves; it is used here as an experimental calibration curve, which is how the emulsion groups used it.
A single cloud-chamber photograph is one event with one measurement uncertainty, and the uncertainty on a curvature is not small: the fractional momentum error from a sagitta measurement grows linearly with momentum, since a stiffer track is a straighter one. Every discovery in this section was therefore contested at the time and settled by accumulating events, not by the first photograph. Where a number below carries an uncertainty it is the one the discoverers quoted; where it does not, the modern value from [Navas:2024] is given alongside so that the reader can see how good the original measurement was.
The positron
Anderson operated a vertical cloud chamber in a field of \(B=1.5\,\mathrm{T}\), divided across its middle by a lead plate \(6\,\mathrm{mm}\) thick, and photographed a track of positive curvature crossing the plate [Anderson:1933]. The plate is what makes the photograph an argument rather than an ambiguity. A particle loses energy in crossing it, so it is more sharply curved after the crossing than before; the observed curvatures were \(B\rho=0.21\,\mathrm{T}\,\mathrm{m}\) on one side and \(0.077\,\mathrm{T}\,\mathrm{m}\) on the other, which fixes the direction of travel as being from the side of smaller curvature to the side of larger — upward, in this event. With the direction known, the sense of the curvature gives the sign, and the sign is positive.
Cosmic-ray tracks exist that are positively charged and whose specific ionisation is that of a singly charged particle moving at nearly the speed of light, at a momentum at which a proton would be slow and would ionise hundreds of times more heavily [Anderson:1933]. Their mass is therefore of the order of the electron mass and not of the proton mass. The particle is the positron, the antiparticle of the electron. Rests on Equations (115.14) and (115.15).
Derivation. Derives Phenomenon 115.8. Take the segment with \(B\rho=0.21\,\mathrm{T}\,\mathrm{m}\). By Equation (115.14) with \(Z=1\),
so that
Now test the two candidate species against the observed grain density, which was that of a minimum-ionising particle, i.e. one with \(\beta\approx1\) in Equation (115.15).
If the particle is a proton, \(mc^{2}=938.272\,\mathrm{MeV}\) [Navas:2024], then
and Equation (115.15) then demands an ionisation \(1/\beta^{2}=223\) times the minimum. A track ionising two hundred times more heavily than the observed one is not a subtle discrepancy; it is a different photograph. If instead the particle has the electron mass, \(mc^{2}=0.511\,\mathrm{MeV}\), then \(\beta=1-3.3\times 10^{-5}\) and the track is minimum ionising, which is what was seen. Anderson's stated bound from the ionisation was a mass below twenty electron masses; the modern value is exactly one, since the positron and electron masses are equal by the \(CPT\) theorem of Discrete Symmetries and CPT.
A second, independent, check comes from the range. The track was followed for several centimetres of chamber gas. A proton of \(pc=63\,\mathrm{MeV}\) has kinetic energy \(T=E-mc^{2}=2.1\,\mathrm{MeV}\) and, by Equation (115.15), stops within a few centimetres of air; a \(63\,\mathrm{MeV}\) positron crosses the whole chamber and the lead plate as well.
∎The full apparatus, procedure and uncertainty budget of the discovery, together with the Blackett–Occhialini confirmation and the first pair-production photographs, belong to Experiment: The Positron, whose headings for them are reserved and carry no content yet; the field-theoretic account of the antiparticle is that of Quantum Electrodynamics and Renormalization. What matters here is that the first antiparticle ever seen was seen in the cosmic radiation, four years after Dirac's equation and with no accelerator involved.
The muon
By the mid-1930s the cosmic radiation at sea level was known to have two components, distinguished by how much absorber they survive: a soft component absorbed by some \(10\,\mathrm{cm}\) of lead, and a hard or penetrating component that crosses a metre of it. Neddermeyer and Anderson showed that the penetrating particles cannot be electrons [Neddermeyer:1937], and the argument is a comparison of two energy-loss laws.
Derivation: the penetrating particles are not electrons. An electron of energy \(E\) crossing matter loses energy chiefly by bremsstrahlung, at a rate proportional to \(E\) itself,
where \(X\) is the traversed depth of Equation (115.2) and \(X_{0}\) is the radiation length of the material, an SI areal mass density: \(X_{0}=365\,\mathrm{kg}/\mathrm{m}^{2} =36.5\,\mathrm{g}/\mathrm{cm}^{2}\) for air and \(63.7\,\mathrm{kg}/\mathrm{m}^{2}=6.37\,\mathrm{g}/\mathrm{cm}^{2}\) for lead. The exponential is the content of the Bethe–Heitler cross-section [Bethe:1934], whose derivation belongs to the bremsstrahlung section of Radiation and Scattering of Electromagnetic Waves, where the heading is reserved and the calculation is not yet written; what matters here is that the loss is proportional to \(E\) and is therefore catastrophic at high energy.
The Bethe–Heitler bremsstrahlung cross section, the resulting exponential energy-loss law, and the radiation length of air and of lead used throughout this chapter. Belongs in Appendix A or in the bremsstrahlung section of the radiation and scattering chapter, which currently reserves the heading
A metre of lead has depth \(X=11.35\times 10^{3}\,\mathrm{kg}/\mathrm{m}^{3}\times1\,\mathrm{m} =1.135\times 10^{4}\,\mathrm{kg}/\mathrm{m}^{2}\), i.e. \(X/X_{0}=178\) radiation lengths. An electron crossing it would emerge with \(\ee^{-178}\) of its energy: nothing survives, at any incident energy. The observed penetrating particles cross such absorbers losing energy at the far slower, nearly energy-independent rate characteristic of ionisation alone, \(-\dd E/\dd X\approx0.2\,\mathrm{MeV}\,\mathrm{m}^{2}/\mathrm{kg}\), i.e. about \(2\,\mathrm{MeV}\) per \(\mathrm{g}/\mathrm{cm}^{2}\). Since the radiative rate carries a factor \(1/m^{2}\) from the recoil of the radiating particle against the nucleus, a particle heavier than the electron radiates less, and much less: the penetrating component is a particle of mass well above \(m_{e}\) and below \(m_{p}\), because a proton of the observed momenta would ionise far too heavily by Equation (115.15).
∎Street and Stevenson closed the argument by measuring the mass directly, from the curvature and the range of a particle stopping in a cloud chamber, obtaining about \(130\) electron masses [Street:1937]. The modern value is \(m_{\mu}c^{2}=105.6583755(23)\,\mathrm{MeV}\), that is \(m_{\mu}=1.883531627(42)\times 10^{-28}\,\mathrm{kg}\) and \(206.77\,m_{e}\) [Navas:2024] [Mohr:2025], so the first measurement was low by a factor \(1.6\) — an honest reflection of what a single stopping track can do.
What the muon is not took another decade. Yukawa's meson, the quantum of the nuclear force of Section 107.2.1, has a mass in the right range, and the identification was universally assumed. It is wrong, and Conversi, Pancini and Piccioni proved it wrong by stopping negative mesotrons in carbon and finding that they decay with their ordinary lifetime instead of being captured by the nucleus [Conversi:1947].
Derivation: the mesotron is not strongly interacting. A negative particle stopped in matter cascades into a Bohr orbit about a nucleus. The mesic Bohr radius scales as the inverse mass and the inverse nuclear charge,
for carbon, \(Z=6\). The nuclear radius is \(R=r_{0}A^{1/3}=1.2\times 10^{-15}\,\mathrm{m}\times12^{1/3}=2.7\times 10^{-15}\,\mathrm{m}\), so the probability of finding the particle inside the nucleus is of order
If the particle interacts strongly, the rate at which it is absorbed once inside is of order the only rate the strong interaction has, \(c/R=1.1\times 10^{23}\,/\mathrm{s}\). The capture rate is then
a capture time of order \(3\times 10^{-20}\,\mathrm{s}\), against the observed decay rate \(1/\tau_{\mu}=4.6\times 10^{5}\,/\mathrm{s}\). The two differ by fourteen orders of magnitude: a strongly interacting particle stopped in carbon would never be seen to decay. It was seen to decay. Hence the mesotron does not interact strongly and is not Yukawa's meson — the same conclusion reached, with the nuclear-physics apparatus, in Phenomenon 107.18.
∎The muon's other role is as a clock. Rossi and Hall measured the momentum dependence of the surviving fraction between altitudes [Rossi:1941], Frisch and Smith compared the flux on Mount Washington with the flux at sea level [Frisch:1963], and the CERN storage ring measured the dilated lifetime at \(\gamma=29.3\) to a few parts in a thousand [Bailey:1977]. All three are treated as experiments in Experiment: Time Dilation and Relativistic Kinematics (Sections 41.2, 41.3 and 41.4); the cosmic-ray side of the argument is Section 115.8.3 below.
The pion
The particle Yukawa predicted was found in 1947, in emulsion stacks exposed at the Pic du Midi and at Chacaltaya in the Andes, by Lattes, Occhialini and Powell [Lattes:1947]. The signature is a track that ends and gives birth to a second track: a heavily ionising particle comes to rest, and from its end point a second particle sets off, itself coming to rest after a short and — this is the point — always the same range, about \(600\,\mu\mathrm{m}\) of emulsion. A third, much lighter track then leaves the second end point. The chain is \(\pi\to\mu\to e\).
In emulsion the secondary of a stopping cosmic-ray meson has a fixed range, corresponding to a kinetic energy of about \(4\,\mathrm{MeV}\), whatever the history of the primary [Lattes:1947]. The parent therefore decays into exactly two bodies, one of which is unseen, and the parent and the charged secondary are two different particles of different mass. Rests on Equation (40.11).
Derivation. Derives Phenomenon 115.9. Let a particle of mass \(m_{\pi}\) decay at rest into a charged particle of mass \(m_{\mu}\) and \(n\) unseen particles. Conservation of four-momentum, Equation (40.11), gives for the charged secondary
with \(X\) the unseen system. Squaring the four-momentum of \(X\),
so that
If \(n=1\) then \(m_{X}\) is a fixed mass and \(E_{\mu}\) is a single number: the secondary is monoenergetic and its range is constant. If \(n\geq2\) then \(m_{X}\) ranges continuously from \(\sum m_{i}\) upwards, \(E_{\mu}\) is spread over a continuum, and no fixed range can occur. A constant range therefore establishes a two-body decay — which is the same logic that identifies the beta spectrum as three-body in Section 101.1.1, read in the opposite direction.
Putting \(m_{X}=0\) and the modern masses \(m_{\pi^{+}}c^{2}=139.57039\,\mathrm{MeV}\), \(m_{\mu}c^{2}=105.6583755\,\mathrm{MeV}\) [Navas:2024] into Equation (115.17),
in SI \(T_{\mu}=6.60\times 10^{-13}\,\mathrm{J}\), and the range of a \(4.12\,\mathrm{MeV}\) muon in emulsion is indeed about \(600\,\mu\mathrm{m}\). The unseen partner is the muon neutrino, whose own story is Flavour Physics and Neutrinos.
∎The pion is the strongly interacting particle the nuclear force demands, and its role as the lightest quantum of that force is developed in Section 107.2.2; that it is light for a reason — being the pseudo-Goldstone boson of the broken chiral symmetry of Section 102.8 — is the QCD side of the same fact. For the present chapter its importance is different and entirely practical: pion production is the dominant channel by which a primary nucleus converts its energy into a shower, and Section 115.4.2 is built on the branching ratios of pion decay and nothing else.
Strange particles
Rochester and Butler, operating a counter-controlled cloud chamber, photographed tracks that fork in mid-flight with no incoming track at the vertex [Rochester:1947]. A neutral particle had decayed into two charged ones. The mass of the parent follows from the two secondaries by the invariant-mass construction of Equation (40.12),
each \(E_{i}\) and \(\vect{p}_{i}\) coming from the curvature and ionisation of one track, and it came out near \(5\times 10^{8}\,\mathrm{eV}/c^{2}\) — about half the proton mass, and a particle nobody had ordered. The modern value for the neutral kaon is \(m_{K^{0}}c^{2}=497.611(13)\,\mathrm{MeV}\) [Navas:2024].
What made these particles “strange” is a rate, not a mass.
The V particles are produced with cross sections characteristic of the strong interaction, and yet they live of order \(10^{-10}\,\mathrm{s}\) before decaying into ordinary hadrons [Rochester:1947] — some thirteen orders of magnitude longer than the strong-interaction time scale. A quantum number conserved in production and violated in decay is therefore required.
Derivation. Derives Phenomenon 115.10. The only time scale the strong interaction possesses is the light transit time of a hadron. With a hadronic radius \(r_{h}\approx10^{-15}\,\mathrm{m}\),
and any decay proceeding through the strong interaction into kinematically allowed hadronic channels must take a time of that order. The observed lifetime is \(10^{-10}\,\mathrm{s}\), larger by \(3\times 10^{13}\). Since the production is strong — the cross sections are millibarns, not microbarns — the initial and final states are connected by the strong interaction in one direction and not in the other, which is impossible for a single interaction. The resolution is an additive quantum number \(S\), conserved by the strong interaction and violated by the weak one: strange particles are then produced only in pairs of opposite \(S\), so that production conserves \(S\), and each must then wait for a weak decay, whose rate is smaller by the square of the Fermi constant. The classification of the resulting spectrum by \(\SU(3)\) is Section 102.1.1, and the identification of \(S\) as a quark flavour is Section 102.1.3; the weak decay itself is quark mixing, Weak Interactions.
∎After 1953 this physics moved to accelerators, where a controlled beam and a known incident energy do in an afternoon what a mountain laboratory did in a year. What cosmic rays retained is what no accelerator can supply: energies above the beam energy of any machine, which is Section 115.8, and an isotropic natural flux over the whole sky, which is what makes them astronomy.
The primary spectrum and composition
The all-particle spectrum
The primary observable is the differential intensity \(J(E)\): the number of particles per unit area, per unit time, per unit solid angle and per unit energy crossing a surface normal to their arrival direction. Its SI unit is \(/\mathrm{m}^{2}/\mathrm{s}/\mathrm{sr}/\mathrm{J}\); the literature quotes it per \(\mathrm{GeV}\), and \(1\,\mathrm{GeV}=1.602176634\times 10^{-10}\,\mathrm{J}\) converts between them.
Over the range in which direct measurement above the atmosphere is possible, the primary nucleon intensity is described by a single power law [Navas:2024],
valid from a few \(\mathrm{GeV}\) — below which the solar modulation of Equation (115.13) takes over — up to about \(10^{14}\,\mathrm{eV}\). Above that the flux is too small for a satellite and is measured through air showers, and the power law continues, with two changes of index, to the highest energies observed.
The differential intensity of primary cosmic rays is, to a good approximation, a featureless power law
from about \(1\,\mathrm{GeV}\) to above \(10^{20}\,\mathrm{eV}\), across which the flux falls by some thirty-two orders of magnitude. Three integral rates fix the scale: about one particle per square centimetre per second per steradian above \(1\,\mathrm{GeV}\), about four per square metre per second per steradian above \(100\,\mathrm{GeV}\), and of order one per square kilometre per year above \(10^{19}\,\mathrm{eV}\) [Gaisser:2016] [Navas:2024]. No terrestrial accelerator produces anything remotely like this range. Rests on Equation (115.19).
Derivation of the integral rates. Derives Phenomenon 115.11. For a power law \(J(E)=J_{0}\left(E/E_{0}\right)^{-\gamma}\) with \(\gamma>1\) the integral intensity above \(E\) converges and is
Inserting Equation (115.19) with \(E_{0}=1\,\mathrm{GeV}\), \(J_{0}=1.8\times 10^{4}\,/\mathrm{m}^{2}/\mathrm{s}/\mathrm{sr}/\mathrm{GeV}\) and \(\gamma=2.7\),
which is the first anchor, and
which is the second. For the third the index must be changed at the knee of Section 115.3.2. Matching a second power law \(\gamma_{2}=3.1\) continuously at \(E_{k}=3\times 10^{15}\,\mathrm{eV}\) gives, for \(E>E_{k}\),
and at \(E=10^{19}\,\mathrm{eV}=10^{10}\,\mathrm{GeV}\) this evaluates to \(3.3\times 10^{-15}\,/\mathrm{m}^{2}/\mathrm{s}/\mathrm{sr}\). Multiplying by \(10^{6}\,\mathrm{m}^{2}/\mathrm{km}^{2}\) and by \(3.156\times 10^{7}\,\mathrm{s}/\mathrm{yr}\),
so that an array accepting some \(\pi\) steradians records of order one event per square kilometre per year — the third anchor, and the reason the observatories of Section 115.4.3 are the size of a small country. The total dynamic range follows in the same way: \(J\left(10^{20}\,\mathrm{eV}\right)/J\left(1\,\mathrm{GeV}\right) =3.1\times 10^{-32}\), thirty-one and a half orders of magnitude.
∎A quantity falling over thirty-two decades cannot be examined on a linear scale, and even on a logarithmic one a change of index from \(2.7\) to \(3.1\) is invisible to the eye. Every published spectrum is therefore plotted as \(E^{3}J(E)\), or as \(E^{2.7}J(E)\), which turns a pure power law of that index into a horizontal line and turns the knee and ankle into visible kinks. The multiplication is cosmetic and carries a trap: the vertical scale then has SI units \(\mathrm{J}^{2}/\mathrm{m}^{2}/\mathrm{s}/\mathrm{sr}\), and a systematic shift of the energy scale by a factor \(1+\epsilon\) moves every plotted point diagonally — right by \(\epsilon\) and up by \(2\epsilon\), since the abscissa carries one power of the energy and the ordinate the remaining two. What that displacement does to the curve is what matters, and at a fixed abscissa the plotted ordinate is then in error by \(\left(\gamma-1\right)\epsilon\): the common factor \(E^{3}\) cancels between the true and the reconstructed curve, leaving exactly the error Equation (115.40) derives for the integral intensity. Section 115.4.6 treats that effect quantitatively, because it is the dominant uncertainty in every spectral measurement above \(10^{18}\,\mathrm{eV}\).
A featureless power law over eleven decades is the strongest single constraint on the acceleration mechanism. No process with a characteristic energy can produce one: a thermal source produces a Maxwellian with a temperature, a resonance produces a line, and a mechanism with a preferred scale produces a break at that scale. What produces a power law is a process in which the fractional energy gain per step and the escape probability per step are both independent of energy, and Section 115.5.3 shows that a strong shock is exactly such a process.
The knee
Near \(3\times 10^{15}\,\mathrm{eV}\) the spectrum of Equation (115.20) steepens, the index changing from about \(2.7\) to about \(3.1\). The break was first seen in the shower-size spectrum [Kulikov:1959], and was later resolved into successive rigidity-ordered cutoffs of the individual elemental groups, the light component breaking first [Antoni:2005]. Which of the two standing readings — a limit of the Galactic accelerators, or a change in Galactic confinement — is correct is not settled by the data. Rests on Equation (115.20).
Kulikov and Khristiansen found the break not in the energy spectrum but in the shower-size spectrum: the number of showers containing more than \(N\) particles at the ground falls as a power of \(N\), and the power changes near \(N\approx10^{6}\) [Kulikov:1959]. Since \(N\propto E\) for an electromagnetic cascade (Equation (115.29)), a break in \(N\) is a break in \(E\), and the calibration puts it at \(3\times 10^{15}\,\mathrm{eV}=4.8\times 10^{-4}\,\mathrm{J}\).
Derivation: any magnetic origin of the knee is ordered by rigidity. Derives Phenomenon 115.13. Suppose the break is produced by a magnetic process, either at the accelerator (a maximum energy the source can reach) or in transport (a maximum energy the Galaxy can retain). Both are governed by the equation of motion of a charge in a magnetic field, and that equation, written for the momentum,
depends on the species only through \(Ze\). Dividing by \(Ze\) and using Equation (115.1),
so two particles of the same rigidity and the same velocity follow identical trajectories in any magnetic field whatever, regardless of their charge and mass. Any limit imposed by magnetic fields is therefore a limit on rigidity, \(R\leq R_{\max}\), and for an ultrarelativistic nucleus of charge \(Z\) this is a limit on energy
The prediction is sharp: the knee of the proton component sits at \(E_{k}\), that of helium at \(2E_{k}\), and that of iron at \(26E_{k}\), so that the all-particle spectrum steepens gradually as one species after another drops out, and the composition becomes progressively heavier across the knee region. With \(E_{k}=3\times 10^{15}\,\mathrm{eV}\) the iron knee is predicted at \(7.8\times 10^{16}\,\mathrm{eV}\). This is what KASCADE found by unfolding the measured shower-size and muon-number distributions into elemental groups: the light component breaks first, the heavy components later, in the rigidity-ordered sequence Equation (115.22) demands [Antoni:2005].
∎The derivation above is the reason the knee is not yet understood. Both candidate explanations are magnetic and both therefore predict the same rigidity ordering, so the observation that confirmed the ordering cannot discriminate between them. An acceleration limit predicts that the sources stop supplying particles above \(ZeR_{\max}\); a confinement limit predicts that the Galaxy stops retaining them. Distinguishing the two requires an observable that is not a spectrum: the anisotropy expected from a leaking population, or a direct measurement of a source spectrum in gamma rays extending past the knee — which is what the LHAASO sources of Section 115.7.1 begin to supply. It is reported here as open because it is open.
The ankle and the highest energies
Linsley, operating the Volcano Ranch array in New Mexico, recorded in 1962 a shower whose reconstructed primary energy was about \(10^{20}\,\mathrm{eV}=16\,\mathrm{J}\) — the kinetic energy of a well-struck tennis ball, carried by a single nucleus [Linsley:1963]. It remains, sixty years later, within a factor of a few of the highest energy ever measured.
Between the knee and that regime the spectrum steepens once more and then flattens again, near \(4\times 10^{18}\,\mathrm{eV}\), the index falling from about \(3.3\) back to about \(2.6\). The feature is called the ankle and was measured with high statistics at the Pierre Auger Observatory [Abraham:2008]. The standard reading is that it marks the crossing of two populations: a Galactic component dying out and an extragalactic component taking over. The argument that the highest-energy particles cannot be Galactic is a comparison of two lengths and is worth doing in SI.
A proton of \(10^{19}\,\mathrm{eV}\) in the Galactic magnetic field has a gyroradius of several kiloparsecs, larger than the thickness of the Galactic disc by more than an order of magnitude and comparable to the distance from the Sun to the Galactic centre. Such particles cannot be stored in the Galaxy, and a Galactic source population cannot supply an isotropic flux of them. Rests on Equation (115.8).
Derivation. Derives Phenomenon 115.15. The Galactic magnetic field measured through Faraday rotation and synchrotron polarisation has a regular component of order \(B=3\times 10^{-10}\,\mathrm{T}\), that is \(0.3\,\mathrm{nT}\) or three microgauss in the older literature. An ultrarelativistic proton of energy \(E\) has \(p=E/c\), so Equation (115.8) gives
For \(E=10^{19}\,\mathrm{eV}=1.602\,\mathrm{J}\) and \(Z=1\),
using \(1\,\mathrm{pc}=3.0857\times 10^{16}\,\mathrm{m}\). The half-thickness of the Galactic disc is about \(0.15\,\mathrm{kpc}\) and the distance to the Galactic centre about \(8\,\mathrm{kpc}\), so the orbit is twenty-four times the disc thickness and about half the size of the whole Galaxy. At \(10^{20}\,\mathrm{eV}\) the same formula gives \(r_{L}=36\,\mathrm{kpc}\), larger than the Galaxy itself: such a particle crosses the disc essentially in a straight line. A population confined by the Galactic field therefore cannot extend to these energies, and conversely a Galactic source of them would produce a strong anisotropy towards the Galactic plane, which Section 115.6.3 shows is not observed. The particles are extragalactic.
∎Elemental and leptonic composition
Above the modulation region and below the knee the primaries are measured species by species with magnetic spectrometers and calorimeters above the atmosphere. By number they are roughly \(90\,\mathrm{\%}\) protons, \(9\,\mathrm{\%}\) helium and \(1\,\mathrm{\%}\) heavier nuclei, with an electron component of order \(1\,\mathrm{\%}\) of the protons at the same energy. The elemental pattern tracks solar-system abundances closely — with two conspicuous exceptions, and the exceptions are the measurement.
Lithium, beryllium and boron are present in the cosmic radiation at a level of order \(25\,\mathrm{\%}\) of the carbon–nitrogen–oxygen group, whereas in solar-system material their abundance relative to that group is of order \(10^{-6}\). The same holds, less dramatically, for scandium through manganese relative to iron. These nuclei are not made in stars in such quantities; they are made in flight, by the fragmentation of heavier cosmic-ray nuclei on interstellar gas, and their abundance measures the amount of matter traversed. Rests on Equation (115.2).
Derivation of the traversed grammage. Derives Phenomenon 115.16. Let \(N_{C}\) be the number of carbon nuclei in a beam and \(N_{B}\) the number of boron nuclei produced from them by spallation, and let \(X\) be the traversed grammage of Equation (115.2), now measured through interstellar gas rather than air. Over a path element the number of fragmentations is the number of target nuclei per unit area times the cross section,
with \(\bar{m}\) the mean target mass, essentially the proton mass \(1.673\times 10^{-27}\,\mathrm{kg}\) for interstellar hydrogen. As long as the secondary fraction stays small this integrates to
The measured boron-to-carbon ratio at a few \(\mathrm{GeV}\) per nucleon is \(N_{B}/N_{C}\approx0.3\), and the fragmentation cross section is \(\sigma_{C\to B}\approx6\times 10^{-30}\,\mathrm{m}^{2}\), which is \(60\) millibarn in the customary non-SI unit (one barn is \(10^{-28}\,\mathrm{m}^{2}\) exactly), whence
That is the whole of the matter a cosmic ray meets between its source and the Earth. It is a small quantity — less than one percent of the atmosphere of Equation (115.3) — and it is nevertheless enough to make the secondary nuclei, because their solar-system abundance is essentially zero. The corresponding path length in an interstellar medium of number density \(n\approx3\times 10^{5}\,/\mathrm{m}^{3}\) is \(L=X/\left(n\bar{m}\right)=1.7\times 10^{23}\,\mathrm{m}\), some \(5\,\mathrm{Mpc}\) of path — folded, of course, into a volume a few kiloparsecs across, since the trajectory is a random walk and not a straight line. Since \(X\) is by Equations (115.2) and (115.24) the matter traversed along the path, \(L/c=5.6\times 10^{14}\,\mathrm{s}\), that is \(1.8\times 10^{7}\,\mathrm{yr}\), is already the residence time; it is recomputed directly as Equation (115.53) and confirmed there by the independent radioactive clock of Section 115.5.5.
∎The ratio is not constant: AMS-02 measured the boron-to-carbon ratio above a rigidity of \(65\,\mathrm{GV}\) and found it to follow a single power law in rigidity,
with no break [Aguilar:2016]. By Equation (115.24) this says the grammage — and hence the residence time — falls as \(R^{-1/3}\): faster particles leave sooner. That single exponent is the input to every propagation calculation in Section 115.5.5, and the value \(1/3\) is what Kolmogorov turbulence in the Galactic magnetic field would give.
The leptonic composition carries an anomaly that is still unresolved. Secondary positrons are made in the same collisions that make boron — \(p+\text{gas}\to\pi^{+}+\dots\), \(\pi^{+}\to\mu^{+}\to e^{+}\) — so their intensity is that of a secondary and must fall relative to the primary electrons as \(R^{-\delta}\), exactly as the boron does. The positron fraction \(e^{+}/\left(e^{+}+e^{-}\right)\) must therefore decrease monotonically with energy. PAMELA found it rising above about \(10\,\mathrm{GeV}\) [Adriani:2009], and AMS-02 measured the rise precisely over the band \(0.5\,\mathrm{GeV}\) to \(350\,\mathrm{GeV}\), finding a fraction that increases with energy while its slope decreases [Aguilar:2013]. Later AMS-02 data extend the measurement to \(1\,\mathrm{TeV}\) and show the rise flattening and turning over [Aguilar:2019].
A rising positron fraction requires a source of primary positrons. Three classes of candidate are discussed: nearby pulsar wind nebulae, which produce electron–positron pairs and are known to exist; a nearby supernova remnant with an unusually hard secondary production; and the annihilation of a dark-matter particle, which is the subject of The Dark Sector: Evidence Without Explanation. The measurement does not distinguish them. The spectral shape alone cannot: a turnover is produced by a nearby source of finite age and by an annihilation cutoff equally well. What would distinguish them is a measured anisotropy towards a nearby pulsar, and none has been detected at the sensitivity so far reached. The honest statement is that a primary positron source exists and is unidentified; asserting a dark-matter interpretation would be asserting a mechanism that has not been observed, which The Dark Sector: Evidence Without Explanation declines to do and so does this chapter.
Extensive air showers
Auger's coincidence experiment
Above about \(10^{14}\,\mathrm{eV}\) the primary flux is too small for any detector that can be flown, and the atmosphere must be used as the detector. That this is possible at all was discovered by Auger and collaborators, who placed Geiger counters at increasing separations on the roof of a laboratory, at the Jungfraujoch and at the Pic du Midi, and recorded coincidences [Auger:1939]. Coincidences persisted at separations of \(300\,\mathrm{m}\).
Counters separated by distances up to \(300\,\mathrm{m}\) fire simultaneously far more often than uncorrelated counting can produce, and the coincidence rate falls smoothly with separation instead of being independent of it [Auger:1939]. A single primary particle therefore produces a coherent front of secondaries tens of thousands strong spread over a large area at the ground, and the primary energies required are of order \(10^{15}\,\mathrm{eV}\) — six orders of magnitude above anything then measured.
Derivation. Derives Phenomenon 115.18. Two counters with singles rates \(r_{1}\) and \(r_{2}\) and a coincidence resolving time \(\tau\) register accidental coincidences at the rate
since counter \(2\) must fire within \(\pm\tau\) of counter \(1\). For counters of area \(100\,\mathrm{cm}^{2}\) at sea level, where the vertical muon intensity is about \(70\,/\mathrm{m}^{2}/\mathrm{s}/\mathrm{sr}\), the singles rate is \(r\approx1.7\,/\mathrm{s}\); with the \(\tau\approx10\,\mu\mathrm{s}\) of the period, Equation (115.26) gives \(r_{\mathrm{ch}}=5.8\times 10^{-5}\,/\mathrm{s}\), about five accidental coincidences a day.
The decisive point is not the size of this number but its independence of the separation: Equation (115.26) contains no distance at all. The measured coincidence rate, by contrast, falls steeply and smoothly as the counters are drawn apart, and remains above the accidental rate out to \(300\,\mathrm{m}\). A distance-dependent excess over a distance-independent background can only be a correlation, and a correlation over \(300\,\mathrm{m}\) at the speed of light means a single common ancestor.
The energy follows from the particle number. If the front contains \(N\) particles at the ground and each was produced by dividing the primary energy down to the point where production stops — the critical energy \(E_{c}\approx85\,\mathrm{MeV}\) of Section 115.4.2 — then \(E_{0}\gtrsim NE_{c}\). Auger's estimate of \(N\approx10^{6}\) particles gives \(E_{0}\gtrsim8.5\times 10^{13}\,\mathrm{eV}\), and correcting for the fact that the ground lies well past the shower maximum, where most of the particles have already been absorbed, raises this to of order \(10^{15}\,\mathrm{eV}=1.6\times 10^{-4}\,\mathrm{J}\).
∎The shower is thus an amplifier: one particle arriving once per square metre per century is converted by the atmosphere into a signal spread over square kilometres, which a sparse array of detectors covering a \(10^{-5}\) fraction of that area can still sample. Every measurement above \(10^{15}\,\mathrm{eV}\) in this chapter rests on it.
Cascade development
Take the electromagnetic part of the cascade first, since the hadronic part feeds it and cannot be described without it.
A high-energy photon in matter converts to an \(e^{+}e^{-}\) pair; a high-energy electron radiates a bremsstrahlung photon. Both processes have the same characteristic depth, the radiation length \(X_{0}\) of Equation (115.16), which for air is
so that the vertical atmosphere of Equation (115.3) is \(1.0332\times 10^{4}\,\mathrm{kg}/\mathrm{m}^{2}/365\,\mathrm{kg}/\mathrm{m}^{2}=28\) radiation lengths deep — a calorimeter as thick as any built. Below a critical energy \(E_{c}\) the electron loses energy faster by ionisation than by radiation and the multiplication stops; for air \(E_{c}\approx85\,\mathrm{MeV}=1.36\times 10^{-11}\,\mathrm{J}\) [Matthews:2005]. Heitler's model replaces the two processes by one [Heitler:1954]: every particle splits into two, each carrying half the energy, after a fixed splitting length \(d=X_{0}\ln2\).
In the splitting model the number of particles at depth \(X\) and the energy per particle are
the cascade stops at the depth where \(E(X)=E_{c}\), and therefore
Two consequences follow, and both are what is actually measured: the particle number at maximum is linear in the primary energy, and the depth of maximum grows only logarithmically with it, at the rate
per decade of primary energy. Rests on Equations (115.2) and (115.16).
Derives Theorem 115.19. After \(n\) splittings the depth is \(X=nd\) and the number of particles is \(2^{n}\), which is Equation (115.28) with \(n=X/d\). Energy is conserved and shared equally, so \(E(X)N(X)=E_{0}\). Multiplication continues while \(E(X)>E_{c}\) and stops when equality is reached; at that depth \(N=E_{0}/E_{c}\), which is the first of Equation (115.29), and \(2^{X_{\max}/d}=E_{0}/E_{c}\) gives
which is the second. Differentiating with respect to \(\log_{10}E_{0}=\ln E_{0}/\ln10\) gives Equation (115.30).
∎The numbers are informative. For \(E_{0}=10^{15}\,\mathrm{eV}\), Equation (115.29) gives \(N_{\max}=1.2\times 10^{7}\) particles and \(X_{\max}=594\,\mathrm{g}/\mathrm{cm}^{2}\); for \(E_{0}=10^{20}\,\mathrm{eV}\), \(N_{\max}=1.2\times 10^{12}\) and \(X_{\max}=1014\,\mathrm{g}/\mathrm{cm}^{2}\), which is just about the depth of the whole vertical atmosphere — so the highest-energy showers reach maximum near the ground and a detector on a high plateau sees them close to their peak. That is why the observatories of Section 115.4.3 sit at altitude.
Heitler's model is not a calculation of a cascade; it is a book-keeping device that gets the two scaling laws right and the absolute numbers wrong by a fixed factor. Real cascades do not split in two: an electron radiates a photon of any energy fraction with the Bethe–Heitler spectrum, so the particle number at maximum is smaller than \(E_{0}/E_{c}\) by a factor of order \(10\), and the sub-threshold particles are not immediately absorbed. The full treatment is the coupled pair of integro-differential cascade equations, solved by Mellin transform in the depth variable — the technique of Fourier Analysis and Integral Transforms, applied to a physically different problem from the one it is introduced for. What survives unchanged from Theorem 115.19 is \(N_{\max}\propto E_{0}\) and \(X_{\max}\propto\ln E_{0}\), and those two are the entire basis of air-shower energy and composition measurement.
A primary nucleus does not start an electromagnetic cascade; it starts a hadronic one, and the electromagnetic component is a by-product. The interaction length follows from the cross section as \(\lambda_{I}=\bar{m}_{\mathrm{air}}/\sigma_{p\text{-air}}\) with \(\bar{m}_{\mathrm{air}}=14.5\times1.6605\times 10^{-27}\,\mathrm{kg} =2.41\times 10^{-26}\,\mathrm{kg}\) the mean mass of an air nucleus; the inelastic proton–air cross section grows slowly with energy, and near \(10^{15}\,\mathrm{eV}\) it is about \(4.5\times 10^{-29}\,\mathrm{m}^{2}\), that is \(450\) millibarn in the customary unit of Equation (115.24), so
so that by Equation (115.3) the atmosphere is about nineteen interaction lengths deep and a primary interacts near the top with certainty [Gaisser:2016]. Each interaction produces of order \(N_{\mathrm{ch}}\) charged pions and \(N_{\mathrm{ch}}/2\) neutral ones, the isospin content of Quantum Chromodynamics making the three charge states roughly equally populated. The neutral pions decay to two photons in \(8.4\times 10^{-17}\,\mathrm{s}\), far faster than they can interact, and feed the electromagnetic channel; the charged pions either interact again or decay to muons.
Derivation of the pion critical energy. A charged pion of Lorentz factor \(\gamma\) travels a mean distance \(\gamma c\tau_{\pi}\) before decaying, with \(\tau_{\pi}=2.6033\times 10^{-8}\,\mathrm{s}\) and \(c\tau_{\pi}=7.80\,\mathrm{m}\). It travels \(\lambda_{I}/\rho\) before interacting. At an altitude of \(10\,\mathrm{km}\), where most of the shower develops, the air density is \(\rho\approx0.41\,\mathrm{kg}/\mathrm{m}^{3}\) and the interaction length is
Equating the two lengths gives the Lorentz factor at which the two fates are equally likely,
in SI \(3.8\times 10^{-9}\,\mathrm{J}\). Pions above this energy interact and keep the hadronic cascade going; pions below it decay and become muons. The value depends on altitude through \(\rho\) and on energy through \(\lambda_{I}\), and is conventionally taken as \(E_{c}^{\pi}\approx20\,\mathrm{GeV}\) for the shower as a whole [Matthews:2005]; the derivation above shows where that number comes from and how much it can be trusted. Note how sensitive it is to the cross section: the older tabulated interaction length for air, \(90\,\mathrm{g}/\mathrm{cm}^{2}\), which is a few-\(\mathrm{GeV}\) value, would give \(39\,\mathrm{GeV}\) instead and turn a twenty-percent agreement into a factor-of-two puzzle.
∎In the hadronic extension of Theorem 115.19, after \(n\) generations the energy remaining in the hadronic channel is \(E_{0}\left(2/3\right)^{n}\), shared among \(N_{\mathrm{ch}}^{n}\) charged pions; the cascade ends when the pion energy falls to \(E_{c}^{\pi}\), and every surviving charged pion becomes a muon. Hence
so the muon number grows more slowly than linearly with primary energy, while the electromagnetic particle number grows linearly. The ratio of the two is therefore an energy-independent handle on the primary mass. Rests on Theorem 115.19.
Derives Theorem 115.21. Of the \(\tfrac{3}{2}N_{\mathrm{ch}}\) pions made in one interaction, \(N_{\mathrm{ch}}\) are charged and carry \(2/3\) of the energy onward; after \(n\) generations the hadronic energy is \(E_{0}\left(2/3\right)^{n}\) spread over \(N_{\mathrm{ch}}^{n}\) pions, so the energy per pion is \(E_{0}\left(\tfrac{3}{2}N_{\mathrm{ch}}\right)^{-n}\). Setting this equal to \(E_{c}^{\pi}\) determines the number of generations,
and \(N_{\mu}=N_{\mathrm{ch}}^{n}=\exp\left(n\ln N_{\mathrm{ch}}\right)\), which is Equation (115.32). With \(N_{\mathrm{ch}}=10\), \(\beta=\ln10/\ln15=0.850\); the exponent is insensitive to \(N_{\mathrm{ch}}\), moving only from \(0.80\) to \(0.88\) as \(N_{\mathrm{ch}}\) runs from \(5\) to \(20\), which is why it is quoted as a number rather than as a function.
∎A nucleus of mass number \(A\) and total energy \(E_{0}\) behaves, to a first approximation, as \(A\) independent nucleons each of energy \(E_{0}/A\). Then
An iron primary (\(A=56\)) therefore produces \(56^{0.15}=1.83\) times as many muons as a proton of the same total energy, and its shower maximum lies about \(D_{10}\log_{10}56\approx100\,\mathrm{g}/\mathrm{cm}^{2}\) higher in the atmosphere, using the measured hadronic elongation rate \(D_{10}\approx58\,\mathrm{g}/\mathrm{cm}^{2}\) per decade rather than the purely electromagnetic Equation (115.30). Rests on Equations (115.30) and (115.32).
Derives Corollary 115.22. Each of the \(A\) nucleons initiates a sub-shower of energy \(E_{0}/A\); the muon numbers add, giving the first of Equation (115.33) directly from Equation (115.32), and \(1-\beta=0.15\). The depths of maximum do not add: the shower maximum of the superposition is that of a single sub-shower, so \(\avg{X_{\max}}\) is evaluated at \(E_{0}/A\) instead of at \(E_{0}\), and by the definition Equation (115.30) of the elongation rate this shifts it by \(-D_{10}\log_{10}A\).
∎Equation (115.32) is the point at which air-shower physics depends on hadronic interaction models extrapolated an order of magnitude in centre-of-mass energy beyond any collider measurement, and it is where the extrapolation currently fails. Measured muon numbers in showers above \(10^{19}\,\mathrm{eV}\) exceed the predictions of every hadronic model, by a factor of order \(1.3\) to \(1.6\), with the primary mass constrained independently by \(X_{\max}\). The discrepancy is reported here as an unresolved defect of the modelling rather than absorbed into a fitted parameter, and it is one of the reasons a composition inference from muon number alone is not currently trustworthy; Section 115.8.1 returns to it.
Detection techniques
Two techniques dominate, and they measure different things.
A surface array samples the particle density of the shower front where it meets the ground, using either plastic scintillators or sealed tanks of water read by photomultipliers, the latter responding to Cherenkov light (Section 115.4.4) from the traversing particles and so being sensitive to the muonic as well as the electromagnetic component. From the sampled densities a lateral distribution function is fitted, and its value at a fixed distance from the axis — \(1000\,\mathrm{m}\) for a kilometre-spaced array — is used as the energy estimator, because at that distance the fitted density is least sensitive to the shape assumed. The arrival direction comes from the relative timing.
Derivation of the angular resolution of an array. Treat the shower front as a plane advancing at speed \(c\) along the unit vector \(\hat{\vect{n}}\). Two detectors separated by the baseline \(\vect{d}\) are struck at times differing by
for a baseline in the plane of incidence, so that measuring \(\Delta t\) measures \(\theta\). Differentiating, \(\sigma_{\theta}=c\,\sigma_{t}/\left(d\cos\theta\right)\). With a timing precision \(\sigma_{t}=10\,\mathrm{ns}\), achievable with photomultipliers and satellite time transfer, and a baseline \(d=1500\,\mathrm{m}\),
at the zenith. The realised resolution of a large array is about \(1\,^\circ\), an order of magnitude worse, because the front is not a plane but a curved and thick shell whose sampled thickness is tens of nanoseconds; Equation (115.34) sets the floor, not the answer.
∎An air-fluorescence detector does something different: it photographs the shower in the dark. Nitrogen molecules excited by the shower's charged particles fluoresce isotropically in the near ultraviolet, at a yield of a few photons per metre per charged particle, and the light is proportional to the energy deposited. A telescope watching a dark sky therefore records the shower's longitudinal profile \(N(X)\) directly, and the primary energy follows by integration,
which is calorimetry with no hadronic model in it at all. That is the technique's great advantage. It has two costs: the light must traverse kilometres of atmosphere whose transmission must be measured continuously, and Equation (115.35) misses the energy carried into the ground by muons and neutrinos, some \(10\,\mathrm{\%}\) to \(15\,\mathrm{\%}\), which must be added back with a model-dependent correction. The technique was realised in the Fly's Eye detector [Baltrusaitis:1985].
Their weaknesses are complementary. Fluorescence works only on clear moonless nights, a duty cycle of order \(13\,\mathrm{\%}\); an array runs continuously. An array's energy scale depends on shower simulations; a fluorescence detector's does not. The Pierre Auger Observatory therefore does both, covering \(3000\,\mathrm{km}^{2}\) of Argentine plateau with \(1600\) water-Cherenkov tanks on a \(1500\,\mathrm{m}\) triangular grid, overlooked by fluorescence telescopes at four sites, and using the subset of events seen by both — the hybrid events — to calibrate the array's estimator against the calorimetric energy [Aab:2015]. The energy scale of the whole observatory is then that of the fluorescence measurement, with a systematic uncertainty of about \(14\,\mathrm{\%}\), and Section 115.4.6 shows what that costs.
The Cherenkov technique
A charged particle moving through a medium faster than the phase velocity of light in that medium radiates, at a fixed angle to its track. The phenomenon and the cone condition are established in Phenomenon 65.7: with \(n\) the refractive index and \(\beta=v/c\),
which has a solution only if \(\beta>1/n\) [Cherenkov:1937]. The intensity is Frank and Tamm's [Frank:1937]; in SI, for a particle of charge \(Ze\),
an energy per unit path length per unit angular frequency, of SI dimension \(\mathrm{J}\,\mathrm{s}/\mathrm{m}\). Dividing by \(\hbar\omega\) converts it to a photon count, and the factor of \(\omega\) cancels:
with \(\alpha=e^{2}/\left(4\pi\varepsilon_{0}\hbar c\right)\) the dimensionless fine-structure constant of Section 100.1.1. The photon number per unit path is thus flat in frequency, which is why Cherenkov light looks blue: equal numbers per unit \(\omega\) means many more per unit wavelength at the short end. The cone condition Equation (115.36) is derived, by the wavelet construction, in Phenomenon 65.7; the intensity Equation (115.37) is not derived anywhere in this book yet.
The Frank–Tamm intensity formula in SI, obtained from the retarded fields of a uniformly moving charge in a dispersive medium by Fourier transform in space and time. Belongs in Appendix A; this chapter quotes the result and uses only its numerical consequences
Rewriting Equation (115.38) per unit wavelength using \(\omega=2\pi c/\lambda\),
and integrating over the photomultiplier band \(300\,\mathrm{nm}\) to \(600\,\mathrm{nm}\) for \(\beta=1\), \(Z=1\):
-
Water, \(n=1.333\): \(\theta_{c}=41.4\,^\circ\), and \(\dd N/\dd x=3.34\times 10^{4}\,/\mathrm{m}\), that is \(334\) photons per centimetre of track. This is the tank of a surface array and the volume of Experiment: Neutrino Oscillations.
-
Ice, \(n=1.32\): \(\theta_{c}=40.7\,^\circ\) and \(\dd N/\dd x=3.26\times 10^{4}\,/\mathrm{m}\), essentially the same. This is IceCube (Section 115.7.4).
-
Air at sea level, \(n=1.000293\): the cone collapses to \(\theta_{c}=1.39\,^\circ\) and the yield to \(45\,/\mathrm{m}\), but the track lengths are kilometres. The threshold is severe: \(\beta>1/n\) requires
\[ \gamma>\gamma_{\mathrm{th}}=\frac{1}{\sqrt{1-n^{-2}}}=41.3\ec \]so an electron needs \(21\,\mathrm{MeV}\) and a muon \(4.4\,\mathrm{GeV}\). A shower's Cherenkov light is therefore emitted almost entirely by its electromagnetic component and is beamed into a narrow forward cone that illuminates a pool some \(125\,\mathrm{m}\) in radius at the ground.
That narrow cone is the basis of gamma-ray astronomy from the ground. An imaging atmospheric Cherenkov telescope focuses the pool of light onto a camera of photomultipliers and forms an image of the shower; the image of a shower initiated by a gamma ray is a compact, regular ellipse pointing at the source position in the field, while a hadronic shower, which contains many sub-showers with large transverse momenta, gives a broad and irregular image. Hillas showed that a few moments of the light distribution — the ellipse's length, width, orientation and total intensity — separate the two classes with a rejection of better than \(99\) percent [Hillas:1985]. The Whipple telescope used exactly this to detect TeV gamma rays from the Crab Nebula at \(9\sigma\), the first source ever found this way [Weekes:1989], and every ground-based gamma-ray instrument since is a development of it.
Depth of maximum and mass composition
The observable that carries composition information with the least model dependence is \(X_{\max}\), together with its shower-to-shower fluctuation. Both are predicted by Corollary 115.22 and both are measured directly by a fluorescence detector.
The mean depth of shower maximum increases logarithmically with primary energy, and at fixed energy is about \(100\,\mathrm{g}/\mathrm{cm}^{2}=1000\,\mathrm{kg}/\mathrm{m}^{2}\) smaller for an iron primary than for a proton. Its shower-to-shower spread is likewise smaller for heavy primaries. Auger measurements show \(\avg{X_{\max}}\) and its spread consistent with a light, largely protonic composition near the ankle, becoming progressively heavier with increasing energy above it [Aab:2014]. Rests on Equation (115.31) and Corollary 115.22.
Derivation of the fluctuation scaling. Derives Phenomenon 115.25. Two independent sources of fluctuation enter. The first is the depth of the first interaction, which is exponentially distributed with mean \(\lambda_{I}\), so it contributes a standard deviation \(\lambda_{I}\) to \(X_{\max}\). The relevant \(\lambda_{I}\) is the one at the energy of the measurement, and by Equation (115.31) it shrinks as the cross section grows: near \(10^{19}\,\mathrm{eV}\) the inelastic proton–air cross section is about \(5.0\times 10^{-29}\,\mathrm{m}^{2}\), that is \(500\) millibarn, so
The second source is the fluctuation in the subsequent cascade, of comparable size; the two add in quadrature, and \(\left(48^{2}+36^{2}\right)^{1/2}=60\) reproduces the measured proton width of some \(60\,\mathrm{g}/\mathrm{cm}^{2}\).
For a nucleus of mass \(A\), superposition (Corollary 115.22) replaces one shower by \(A\) sub-showers of energy \(E_{0}/A\), each with its own first-interaction depth. The shower maximum of the sum is essentially the average of the \(A\) sub-maxima, and the variance of an average of \(A\) independent variables is \(1/A\) of the individual variance:
For iron \(\sqrt{56}=7.5\), so the rule predicts a collapse from the proton value of some \(60\,\mathrm{g}/\mathrm{cm}^{2}\) to \(8\,\mathrm{g}/\mathrm{cm}^{2}\). The measured iron width is about \(20\,\mathrm{g}/\mathrm{cm}^{2}\): the superposition rule gets the direction and the order of the effect right and underpredicts the width by a factor of about two and a half, which is the same crudeness Remark 115.20 concedes for the model it rests on. The qualitative conclusion — heavy primaries fluctuate much less than light ones — is robust; the number is not. Mean and width are thus two independent composition observables, and they must agree: a mean lying between the proton and iron expectations can be produced either by an intermediate species, which would have a narrow distribution, or by a mixture of protons and iron, which would have a very wide one. Auger sees the mean move heavier above the ankle and the width narrow, which favours a genuinely intermediate composition over a mixture [Aab:2014].
∎The consequence for Section 115.6.2 is direct and uncomfortable. If the highest-energy primaries are not protons, the suppression at the end of the spectrum cannot be read straightforwardly as photopion production on the microwave background, because Equation (115.22) allows the same break to be the maximum rigidity of the sources. That objection is stated where it belongs, in the derivation of Phenomenon 115.42.
Energy reconstruction and its uncertainty
Every spectrum in this chapter is a histogram of reconstructed energies, and two distinct errors act on it. Neither is small, and the statistical machinery for both is that of Probability and Statistics.
Counting. In an energy bin containing \(N\) events the estimator of the rate is the maximum-likelihood estimator of a Poisson mean, which by Example 11.58 is \(N\) itself with standard deviation \(\sqrt{N}\). At the highest energies \(N\) is single digits: a bin with four events carries a \(50\,\mathrm{\%}\) statistical uncertainty, and no amount of care in the reconstruction improves it. This is why the exposure — area times time times solid angle, in \(\mathrm{km}^{2}\,\mathrm{yr}\,\mathrm{sr}\) — is the figure of merit quoted for every observatory.
Energy scale and resolution. These act in opposite directions and must not be confused.
Let the true differential intensity be \(J(E)=AE^{-\gamma}\).
-
If the energy scale carries a systematic multiplicative bias, \(E_{\mathrm{rec}}=\left(1+\epsilon\right)E\), the reconstructed integral intensity above a fixed energy is in error by the fraction
\begin{equation}\tag{115.40} \frac{\Delta J}{J}=\left(\gamma-1\right)\epsilon\ec \end{equation}to first order in \(\epsilon\).
-
If instead the energy is unbiased but smeared, with \(\ln E_{\mathrm{rec}}=\ln E+u\) and \(u\) Gaussian of zero mean and standard deviation \(\sigma\), the reconstructed differential intensity is the true one multiplied by
\begin{equation}\tag{115.41} \frac{J_{\mathrm{rec}}\left(E\right)}{J\left(E\right)} =\exp\!\left[\frac{\left(\gamma-1\right)^{2}\sigma^{2}}{2}\right] >1\ec \end{equation}the same power law with an inflated normalisation. Resolution alone therefore raises a measured flux; it never lowers it.
Rests on Equation (115.21).
Derives Proposition 115.26. For (i), \(J(>E)=AE^{-\left(\gamma-1\right)}/\left(\gamma-1\right)\) by Equation (115.21). A scale bias means the flux reported at \(E\) is the true flux at \(E/\left(1+\epsilon\right)\), so
For (ii), change variable to \(x=\ln E\). The number of events per unit \(x\) is \(J(E)E=Ae^{-\left(\gamma-1\right)x}\), and smearing convolves this with a Gaussian of width \(\sigma\):
by completing the square — this is the moment generating function of a Gaussian evaluated at \(\gamma-1\). Dividing by \(E_{r}\) to return to a differential intensity gives Equation (115.41). The result is exact, not perturbative, and holds for any \(\sigma\).
∎Above the suppression the local index is \(\gamma\approx4.5\). An energy-scale uncertainty of \(\epsilon=14\,\mathrm{\%}\) — the Auger systematic [Aab:2015] — gives by Equation (115.40) a flux uncertainty of \(3.5\times0.14=49\,\mathrm{\%}\): a half-order-of- magnitude band on the vertical axis of every published spectrum, and much larger than the statistical error in the same bins. An energy resolution of \(\sigma=0.20\) in the logarithm gives by Equation (115.41) a flux inflation of \(\exp\left(3.5^{2}\times0.04/2\right)=1.28\), a \(28\,\mathrm{\%}\) overestimate that must be unfolded away before the spectrum is quoted. Both effects grow with \(\gamma\), which is precisely why they matter most where the spectrum is steepest and the physics is most interesting.
The significance quoted for a spectral break — the \(5\sigma\) and \(6\sigma\) of Phenomenon 115.42 — is a likelihood-ratio statistic: the maximised likelihood of a broken power law against that of a single one, converted to a significance by Wilks's theorem (Theorem 11.88) and Proposition 11.93. Two cautions apply and are respected in the published analyses. First, the break energy is a free parameter that is not identified under the null hypothesis, which is the failure mode of Wilks's theorem discussed in Remark 11.91; the correct calibration is by simulation, and it lowers the significance. Second, the systematic of Equation (115.40) is common to all bins and so shifts the whole spectrum rather than scattering it, which means it cannot create a break — a genuinely useful fact, and the reason the suppression survives a \(14\,\mathrm{\%}\) scale uncertainty intact.
Acceleration and origin
The supernova-remnant hypothesis
Baade and Zwicky, in the short paper in which they proposed the neutron star, suggested that supernovae are the source of the cosmic radiation [Baade:1934]; the term super-nova itself was introduced in the companion paper printed immediately before it in the same volume. The proposal was made before any acceleration mechanism was known, on energetic grounds alone, and the energetics remain the strongest argument for it.
The energy density of cosmic rays in the Galaxy is about \(1\,\mathrm{eV}/\mathrm{cm}^{3}=1.60\times 10^{-13}\,\mathrm{J}/\mathrm{m}^{3}\), and they escape on a time scale of about \(10^{7}\,\mathrm{yr}\). Sustaining that density requires a power of order \(6\times 10^{33}\,\mathrm{W}\), which is about \(10\,\mathrm{\%}\) of the kinetic power of Galactic supernovae. No other Galactic population comes within an order of magnitude of supplying it. Rests on Equation (115.24).
Derivation. Derives Phenomenon 115.29. Take the confinement volume to be a disc of radius \(R=15\,\mathrm{kpc}=4.63\times 10^{20}\,\mathrm{m}\) and half-thickness \(h=300\,\mathrm{pc}=9.26\times 10^{18}\,\mathrm{m}\):
so the stored energy is
In a steady state the injected power equals the stored energy divided by the residence time \(\tau_{\mathrm{esc}}\approx10^{7}\,\mathrm{yr} =3.16\times 10^{14}\,\mathrm{s}\) of Section 115.5.5:
The available power is that of the supernovae. A core-collapse supernova delivers of order \(10^{44}\,\mathrm{J}\) to the kinetic energy of its ejecta — three tenths of one percent of the \(3\times 10^{46}\,\mathrm{J}\) of gravitational binding energy released, the remaining ninety-nine and more percent leaving as the neutrino burst of Section 115.7.3 — and the Galactic rate is about two per century:
The required efficiency is therefore
ten percent of the mechanical energy of the blast wave. That is demanding but not absurd, and it is exactly what the shock acceleration of Section 115.5.3 predicts when the accelerated particles are dynamically important. Every alternative Galactic reservoir — stellar winds, pulsar rotational energy, the interstellar magnetic field — falls short of \(P_{\mathrm{CR}}\) by an order of magnitude or more, which is why the supernova hypothesis has survived ninety years without a rival.
∎An energy budget is not a detection, and for most of that ninety years the argument stopped there: supernova remnants were known to be bright in radio synchrotron emission, which proves they accelerate electrons, but the cosmic radiation is overwhelmingly protons and nuclei, and an electron accelerator is not evidence of a proton accelerator. The proof arrived in 2013.
The gamma-ray spectra of the remnants IC 443 and W44, measured between \(60\,\mathrm{MeV}\) and \(2\,\mathrm{GeV}\), show a characteristic break and low-energy turnover that is the kinematic signature of neutral-pion decay, and that no electron process reproduces [Ackermann:2013]. Protons and nuclei are therefore accelerated in these objects, not only electrons. Rests on Phenomenon 115.29.
Derivation of the pion-decay signature. Derives Phenomenon 115.30. A neutral pion at rest decays to two photons each of energy \(m_{\pi^{0}}c^{2}/2=67.5\,\mathrm{MeV}\). Boost it to Lorentz factor \(\gamma_{\pi}\): by the transformation of a null four-momentum, the photon energy in the laboratory is
with \(\theta^{*}\) the emission angle in the pion frame. Because the decay is isotropic in that frame, \(\cos\theta^{*}\) is uniform on \([-1,1]\), so \(E_{\gamma}\) is uniformly distributed between \(\gamma_{\pi}\left(1-\beta_{\pi}\right)m_{\pi}c^{2}/2\) and \(\gamma_{\pi}\left(1+\beta_{\pi}\right)m_{\pi}c^{2}/2\). The distribution is flat, and in \(\log E_{\gamma}\) it is symmetric about \(\log\left(m_{\pi}c^{2}/2\right)\), since the two limits multiply to \(\left(m_{\pi}c^{2}/2\right)^{2}\) whatever \(\gamma_{\pi}\) is. Summing over any pion spectrum preserves that symmetry point. The resulting photon spectrum therefore rises to a maximum at \(E_{\gamma}=67.5\,\mathrm{MeV}\) and falls below it, whatever the parent proton spectrum — the “pion bump”.
No leptonic process produces such a feature. Bremsstrahlung from electrons follows the electron spectrum with no scale in it; inverse-Compton scattering likewise. The observed turnover below \(200\,\mathrm{MeV}\) in IC 443 and W44 is thus a direct measurement of a hadronic population, and it is the observation that converts Phenomenon 115.29 from a budget into evidence [Ackermann:2013].
∎Fermi acceleration
Fermi proposed that cosmic rays gain energy by scattering repeatedly off moving magnetised clouds in the interstellar medium [Fermi:1949]. A magnetic field does no work, so all the energy comes from the motion of the scattering centre, and the sign of the exchange depends on whether the encounter is head-on or overtaking.
A relativistic particle scattering elastically off a randomly moving cloud of speed \(V\ll c\), isotropised in the cloud frame, gains on average the fraction
of its energy per encounter. The gain is second order in \(V/c\), because head-on and overtaking encounters cancel at first order and only the excess of head-on encounters survives. Rests on Definition 40.1 and Equation (40.3).
Derives Theorem 115.31. Write \(\beta_{V}=V/c\) and \(\gamma_{V}=\left(1-\beta_{V}^{2} \right)^{-1/2}\), and let \(\theta\) be the angle in the laboratory between the particle's direction and the cloud's velocity. Transforming the particle's energy into the cloud frame (a boost, as in Section 40.1),
for an ultrarelativistic particle, \(pc=E\). The scattering is elastic in the cloud frame — the field does no work there — so \(E''=E'\) with a new direction \(\theta''\), and transforming back,
Isotropisation in the cloud frame gives \(\avg{\cos\theta''}=0\), so
The laboratory angles are not uniformly sampled: the rate of encounters with a cloud approaching at angle \(\theta\) is proportional to the relative closing speed, \(\propto\left(1-\beta_{V}\cos\theta\right)\), so
Head-on encounters, which gain energy, are commoner than overtaking ones, which lose it. Hence
which is Equation (115.44).
∎Interstellar clouds move at \(V\approx30\,\mathrm{km}/\mathrm{s}\), so \(\beta_{V}=10^{-4}\) and Equation (115.44) gives \(1.3\times 10^{-8}\) per encounter. With a scattering mean free path of order a parsec the encounter interval is \(1\,\mathrm{pc}/c=3.3\,\mathrm{yr}\), so the energy \(e\)-folding time is
some twenty-five times longer than the residence time \(\tau_{\mathrm{esc}}\approx10^{7}\,\mathrm{yr}\) of Section 115.5.5. The particles leave the Galaxy long before they are accelerated. Worse, the predicted spectral index depends on the ratio \(t_{\mathrm{acc}}/\tau_{\mathrm{esc}}\), which is a free parameter with no reason to take the same value in different regions, so the mechanism does not explain why the spectrum is a power law of one universal index. Both objections are removed at a stroke if the scattering centres approach the particle systematically rather than randomly, which is what a shock provides.
Diffusive shock acceleration
A supernova blast wave is a strong shock: a discontinuity across which the fluid speed drops and the density rises, propagating into the interstellar medium at \(u_{1}\approx10^{7}\,\mathrm{m}/\mathrm{s}\). Magnetic irregularities are frozen into the fluid on both sides, so a particle crossing the shock finds the scattering centres on the far side approaching it — whichever way it crosses. Every crossing is head-on, and the first-order cancellation of Theorem 115.31 is gone.
Let a shock have upstream speed \(u_{1}\) and downstream speed \(u_{2}\) in the shock frame, with compression ratio \(r=u_{1}/u_{2}\). A particle that repeatedly crosses and recrosses gains, per complete cycle,
first order in the speeds, and escapes downstream with probability
per cycle. The resulting integral spectrum is a power law
and for a strong shock in a monatomic gas, \(r=4\) and \(\gamma_{s}=2\) — independently of the shock speed, the magnetic field and every other detail. Rests on Theorem 115.31.
Derives Theorem 115.33. The gain. Transform to the frame of the upstream fluid. The downstream fluid approaches it at speed \(u_{1}-u_{2}\), so a particle that crosses into the downstream region and is isotropised there undergoes exactly the calculation of Theorem 115.31 with \(\beta_{V}=\left(u_{1}-u_{2}\right)/c\), but with one difference: the crossing is always head-on, so the average over \(\cos\theta\) is taken over the flux of particles crossing a plane, whose distribution is \(\propto\mu\,\dd\mu\) on \([0,1]\), giving \(\avg{\mu}=\tfrac{2}{3}\). Each of the two crossings that make a cycle contributes \(\beta_{V}\avg{\mu}=\tfrac{2}{3}\beta_{V}\), so \(\xi=\tfrac{4}{3}\beta_{V}\), which is Equation (115.45). The result is first order because the approach is systematic: there are no overtaking encounters to cancel it.
The escape. Downstream the particles are isotropic in the fluid frame and their flux onto the shock from behind is \(n c/4\), the standard one-sided flux of an isotropic gas. They are simultaneously carried away from the shock by the downstream flow at the rate \(nu_{2}\). The fraction lost per cycle is the ratio,
which is Equation (115.46). Note that neither \(\xi\) nor \(P_{\mathrm{esc}}\) depends on the particle's energy. That is the whole reason the spectrum is a power law.
The spectrum. After \(k\) cycles a particle has energy \(E_{k}=E_{0}\left(1+\xi\right)^{k}\) and the number surviving is \(N_{k}=N_{0}\left(1-P_{\mathrm{esc}}\right)^{k}\). Eliminating \(k\),
so \(N(>E)\propto E^{-P_{\mathrm{esc}}/\xi}\) with
which is Equation (115.47); the differential index is one greater.
The compression ratio. Work in the shock frame, where the flow is steady, and impose conservation of mass, momentum and energy across the discontinuity — the Rankine–Hugoniot conditions, which belong to the fluid mechanics of Section 31.10.2:
with the specific enthalpy of an ideal gas \(h=\gamma_{\mathrm{ad}}P/\left[\left(\gamma_{\mathrm{ad}} -1\right)\rho\right]\). A strong shock is one whose upstream pressure is negligible, \(P_{1}\to0\) and hence \(h_{1}\to0\); the momentum condition then gives \(P_{2}=j\left(u_{1}-u_{2}\right)\), and substituting this and \(\rho_{2}=j/u_{2}\) into the energy condition,
Cancelling the common factor \(u_{1}-u_{2}\), which is non-zero for a genuine shock, leaves \(u_{1}+u_{2}=2\gamma_{\mathrm{ad}}u_{2}/ \left(\gamma_{\mathrm{ad}}-1\right)\) and therefore
for a monatomic ionised gas, whence \(\gamma_{s}=2\).
∎The independence of Equation (115.47) from every parameter except \(r\) is what makes the mechanism credible, and it is why the result was found independently and almost simultaneously by several groups [Bell:1978] [Blandford:1978]. It also answers Phenomenon 115.11: a featureless power law over eleven decades requires an energy-independent gain and an energy-independent escape, and a strong shock supplies both.
If sources inject a spectrum \(Q(E)\propto E^{-\gamma_{s}}\) into a Galaxy from which particles escape on a rigidity-dependent time scale \(\tau_{\mathrm{esc}}\propto R^{-\delta}\), the steady-state spectrum observed is
Derives Proposition 115.34. In a steady state the number density at energy \(E\) satisfies injection equals escape: \(Q(E)=n(E)/\tau_{\mathrm{esc}}(E)\), so \(n(E)=Q(E)\tau_{\mathrm{esc}}(E)\), and the intensity is \(J=nc/4\pi\) for an isotropic distribution. Since the particles are ultrarelativistic, \(R=E/\left(Ze\right)\) and the rigidity dependence is an energy dependence at fixed \(Z\).
∎Put the two results together. Theorem 115.33 gives \(\gamma_{s}=2\); the boron-to-carbon measurement Equation (115.25) gives \(\delta=0.333\pm0.014\) [Aguilar:2016]; Equation (115.49) therefore predicts an observed index of \(2.33\pm0.02\). The measured index below the knee is \(2.7\). The discrepancy is \(0.37\) in the exponent, which over the four decades between \(10\,\mathrm{GeV}\) and \(10^{14}\,\mathrm{eV}\) is a factor of \(10^{4\times0.37}=30\) in flux — far outside any uncertainty in either input.
Something in the chain is wrong, and the candidates are known: the test-particle assumption behind Theorem 115.33 fails when the accelerated particles carry a tenth of the shock's energy, as Equation (115.43) requires them to, and their pressure modifies the shock structure so that \(r\) and hence \(\gamma_{s}\) become energy dependent; and the escape time inferred from secondary nuclei is an average over the Galaxy that need not describe the region where the sources are. This chapter does not adjudicate between them. It records that diffusive shock acceleration explains the existence and the universality of the power law, and does not yet explain its index.
Derivation of the maximum energy. Derives Equation (115.51). The mechanism has a limit, and it is the reason the knee exists in the first reading of Remark 115.14. The time to complete one cycle is set by how far the particle diffuses from the shock before returning. Consider the upstream side. The flow sweeps the particles onto the shock at speed \(u_{1}\) while diffusion carries them away from it, and the balance \(u_{1}n=D_{1}\,\dd n/\dd x\) gives a precursor of scale length \(D_{1}/u_{1}\), hence a column \(nD_{1}/u_{1}\) of particles per unit area waiting to recross. They recross at the one-sided isotropic flux \(nc/4\) already used in Equation (115.46), so the mean time spent upstream per crossing is \(\left(nD_{1}/u_{1}\right)/\left(nc/4\right)=4D_{1}/\left(u_{1}c\right)\), and the same argument downstream gives \(4D_{2}/\left(u_{2}c\right)\). One complete cycle therefore takes
and since each cycle multiplies the energy by \(1+\xi\) with the \(\xi\) of Equation (115.45), the time to raise the energy by one \(e\)-folding is \(t_{\mathrm{acc}}=t_{\mathrm{cyc}}/\xi\), that is
Taking the same diffusion coefficient \(D\) on both sides and \(r=4\), so that \(u_{2}=u_{1}/4\) and \(u_{1}-u_{2}=3u_{1}/4\), the bracket is \(5D/u_{1}\) and
the second expression being the smallest diffusion coefficient a magnetic field permits, the Bohm value, obtained when the scattering mean free path equals the gyroradius of Equation (115.8). Setting \(t_{\mathrm{acc}}\) equal to the lifetime \(T\) of the shock,
With \(Z=1\), \(u_{1}=10^{7}\,\mathrm{m}/\mathrm{s}\), \(T=1000\,\mathrm{yr}=3.16\times 10^{10}\,\mathrm{s}\) and the ambient interstellar field \(B=3\times 10^{-10}\,\mathrm{T}\),
a factor of twenty below the knee at \(3\times 10^{15}\,\mathrm{eV}\). The mechanism as stated therefore fails to reach the knee. It succeeds if the field at the shock is amplified above the ambient value by a factor of order \(30\): with \(B=10^{-8}\,\mathrm{T}\), Equation (115.51) gives \(E_{\max}=4.7\times 10^{15}\,\mathrm{eV}\), which does reach it. Such amplification is observed — the X-ray synchrotron filaments at young remnant shocks are thin, and their thinness measures a downstream field of order \(10^{-8}\,\mathrm{T}\) — and the physical mechanism proposed for it is a streaming instability driven by the accelerated particles themselves.
∎The non-resonant streaming instability by which the accelerated particles amplify the magnetic field at the shock, and the saturation level that fixes \({[}B{]}\) in the maximum-energy estimate. Belongs in Appendix A; the estimate in this chapter takes the amplified field from the X-ray filament observations instead of deriving it
The Hillas criterion and candidate sources
Whatever the mechanism, an accelerator cannot confine a particle whose gyroradius exceeds its own size. That single observation, due to Hillas [Hillas:1984], excludes most of the astrophysical zoo at the highest energies and it does so with no assumption about how the acceleration works.
An accelerator of characteristic size \(L\) containing a magnetic field \(B\), whose scattering centres move with speed \(\beta c\), cannot accelerate a nucleus of charge \(Ze\) beyond
an SI energy in joules when \(B\) is in tesla and \(L\) in metres. Equivalently the maximum rigidity of Equation (115.1) is \(R_{\max}=\beta BLc\), in volts. Rests on Equations (115.1), (115.8) and (115.45).
Derives Theorem 115.36. By Equation (115.8) a particle of rigidity \(R\) has a gyroradius \(r_{L}=R/\left(cB\right)\). It remains in the accelerating region only while \(r_{L}\leq L\), i.e. while \(R\leq BLc\); beyond that it leaves in less than one gyration and cannot be turned back. The factor \(\beta\) enters because the energy gain per crossing is first order in the speed of the scattering centres (Equation (115.45)), so a slow shock requires many crossings and loses a corresponding factor. Converting rigidity to energy with \(E=ZeR\) gives Equation (115.52). Dimensionally, \(\left[Ze\right]\left[B\right]\left[L\right]\left[c\right] =\mathrm{C}\cdot\mathrm{T}\cdot\mathrm{m} \cdot\mathrm{m}/\mathrm{s}=\mathrm{J}\), as required.
∎Evaluating Equation (115.52) for \(Z=1\):
| Object | $B$ (\(\mathrm{T}\)) | $L$ (\(\mathrm{m}\)) | $\beta$ | $E_{\max}$ (\(\mathrm{eV}\)) |
|---|---|---|---|---|
| Neutron star | \(10^{8}\) | \(10^{4}\) | $1$ | \(3\times 10^{20}\) |
| Supernova remnant shock | \(10^{-8}\) | \(3.1\times 10^{17}\) | $0.03$ | \(2.8\times 10^{16}\) |
| Active-galaxy jet (parsec scale) | \(10^{-7}\) | \(3.1\times 10^{16}\) | $1$ | \(9\times 10^{17}\) |
| Radio-galaxy lobe | \(10^{-9}\) | \(3.1\times 10^{21}\) | $1$ | \(9\times 10^{20}\) |
| Cluster accretion shock | \(10^{-10}\) | \(3.1\times 10^{22}\) | $0.003$ | \(2.8\times 10^{18}\) |
| Galactic disc (for comparison) | \(3\times 10^{-10}\) | \(4.6\times 10^{20}\) | $1$ | \(4\times 10^{19}\) |
Only two entries reach \(10^{20}\,\mathrm{eV}\) for a proton, and each has a difficulty. The neutron-star value assumes the whole magnetospheric potential is available to one particle, which radiative losses in that field forbid. The radio-galaxy lobe is a plausible site but the nearest such object is tens of megaparsecs away, which the horizon of Section 115.6.1 makes marginal. Note also that Equation (115.52) scales with \(Z\), so an iron nucleus reaches \(26\) times the plotted energy, which is why the composition result of Phenomenon 115.25 changes the shortlist.
This is the honest state of the subject. The Hillas criterion tells us which objects are not excluded; it identifies nothing. The anisotropy of Section 115.6.3 establishes that the highest-energy particles come from outside the Galaxy and no more. Correlations of arrival directions with catalogues of nearby active galaxies and starburst galaxies have been reported at significances between \(3\sigma\) and \(4\sigma\) and have neither strengthened decisively nor gone away with additional exposure. This chapter therefore names candidates and states that none is established.
Galactic propagation and confinement
Between the source and the detector the particles diffuse through the Galactic magnetic field. The simplest description that fits the data is the leaky box: the Galaxy is treated as a container of uniform density from which particles escape with a constant probability per unit time, \(1/\tau_{\mathrm{esc}}\). It has one free function, \(\tau_{\mathrm{esc}}(R)\), and two independent measurements of it.
The first is the grammage of Equation (115.24): the traversed matter is \(X=\bar{m}n\,c\,\tau_{\mathrm{esc}}\) for a particle moving at essentially \(c\) through gas of number density \(n\), so
where the effective density \(n\approx3\times 10^{5}\,/\mathrm{m}^{3}\) (that is \(0.3\,/\mathrm{cm}^{3}\)) is the disc density averaged over a trajectory that spends most of its time in the low-density halo. That average is the weak point of the estimate, and it is why a second, independent measurement matters.
The second is a clock. Beryllium-10 is produced by the same spallation that produces boron, and it is radioactive with a half-life of \(1.39\times 10^{6}\,\mathrm{yr}\), that is a mean life \(\tau_{1/2}/\ln2=2.0\times 10^{6}\,\mathrm{yr}\). If the residence time were much shorter than this, all the \(^{10}\)Be produced would still be present and the measured \(^{10}\)Be\(/^{9}\)Be ratio would equal the production ratio; if it were much longer, essentially none would survive. The observed ratio is suppressed by a factor of several relative to production, which requires
in agreement with Equation (115.53) and independent of the assumed gas density — which is what makes the agreement informative rather than circular. The corresponding diffusion coefficient, for a halo of half-height \(H\approx3\,\mathrm{kpc}=9.3\times 10^{19}\,\mathrm{m}\), is
some thirty orders of magnitude above the self-diffusion coefficient of a gas at room temperature and atmospheric pressure, which is of order \(2\times 10^{-5}\,\mathrm{m}^{2}/\mathrm{s}\), and rising as \(R^{\delta}\) with the \(\delta=0.333\pm0.014\) of Equation (115.25) [Aguilar:2016]. Standard treatments of the whole transport problem, of which the leaky box is the crudest useful limit, are in [Gaisser:2016].
Equation (115.54) is also the end of any hope of pointing with them. A particle that diffuses for \(1.5\times 10^{7}\,\mathrm{yr}\) before leaving the Galaxy has had its direction randomised completely: by Equation (115.23) a \(10^{15}\,\mathrm{eV}\) proton in \(3\times 10^{-10}\,\mathrm{T}\) has a gyroradius of \(0.36\,\mathrm{pc}\), utterly negligible against the kiloparsec scales of the Galaxy, so its arrival direction retains no memory whatever of its source. Only above \(10^{19}\,\mathrm{eV}\), where Phenomenon 115.15 gives gyroradii of kiloparsecs, does any directional information survive — and there the sources are extragalactic and the intervening extragalactic fields are poorly known. This is the reason the subject needs neutral messengers, and Section 115.7 is about them.
The end of the spectrum
The Greisen–Zatsepin–Kuz'min prediction
Within two years of the discovery of the cosmic microwave background, Greisen [Greisen:1966] and, independently, Zatsepin and Kuz'min [Zatsepin:1966] pointed out that it makes intergalactic space opaque to protons above a definite energy. The argument is pure relativistic kinematics applied to a measured photon bath, and it is one of the sharpest predictions in astrophysics: it uses no unknown source, no unknown field, and no free parameter.
A proton propagating through the \(2.72548\,\mathrm{K}\) microwave background of Phenomenon 51.3 loses energy by photopion production, \(p+\gamma\to\Delta^{+}\to p+\pi^{0}\) or \(n+\pi^{+}\), once its energy exceeds a threshold of order \(10^{20}\,\mathrm{eV}\) against a typical background photon and a few times \(10^{19}\,\mathrm{eV}\) against the high-energy tail of the bath. The energy-loss length above the threshold is of order tens of megaparsecs, far smaller than the Hubble radius, so the sources of any observed particle above the threshold must lie within that distance [Greisen:1966] [Zatsepin:1966] [Stecker:1968]. Rests on Phenomenon 51.3 and Equation (40.12).
Derivation of the threshold. Derives Phenomenon 115.40. Let a proton of energy \(E\) and momentum \(p=\sqrt{E^{2}/c^{2} -m_{p}^{2}c^{2}}\) meet a photon of energy \(E_{\gamma}\) head-on. The invariant mass of the pair is given by Equation (40.12):
the limit being the ultrarelativistic one, \(pc\to E\); for a general collision angle \(\theta\) the last term is \(2EE_{\gamma}\left(1-\cos\theta\right)\), maximal head-on. Photopion production \(p+\gamma\to N+\pi\) is kinematically possible once \(\sqrt{s}\geq\left(m_{N}+m_{\pi}\right)c^{2}\), and the lowest such threshold is the \(p\pi^{0}\) channel. Solving Equation (115.55),
The photon bath is measured, not assumed. Its temperature is \(T_{0}=2.72548(57)\,\mathrm{K}\) [Fixsen:2009], so
and the mean energy of a Planck photon is
while their number density is
With \(m_{p}c^{2}=938.272\,\mathrm{MeV}\) and \(m_{\pi^{0}}c^{2}=134.977\,\mathrm{MeV}\) [Navas:2024], Equation (115.56) gives
and the same calculation with the \(\Delta(1232)\) mass in place of \(\left(m_{p}+m_{\pi}\right)\) puts the resonance peak at \(2.5\times 10^{20}\,\mathrm{eV}\).
∎Derivation of the onset against the Planck tail. A threshold computed with the mean photon energy is not the energy at which the effect begins, because the Planck distribution has a tail. At proton energy \(E\) the reaction proceeds on any photon with \(E_{\gamma}\geq\varepsilon_{\mathrm{th}}(E)\), where by Equation (115.56) \(\varepsilon_{\mathrm{th}}=m_{\pi}c^{2}\left(2m_{p}c^{2} +m_{\pi}c^{2}\right)/(4E)\). Writing \(x=\varepsilon_{\mathrm{th}}/\left(k_{B}T_{0}\right)\), the fraction of the bath above that energy is
the approximation being the leading term of the expansion of \(\left(\ee^{u}-1\right)^{-1}\), excellent for \(x\gtrsim3\). At \(E=5\times 10^{19}\,\mathrm{eV}\), \(\varepsilon_{\mathrm{th}}=1.36\times 10^{-3}\,\mathrm{eV}\) and \(x=5.78\), so Equation (115.59) gives \(6.0\,\mathrm{\%}\): six percent of the bath, that is \(2.5\times 10^{7}\,/\mathrm{m}^{3}\) photons, are already above threshold. At \(E=10^{20}\,\mathrm{eV}\), \(x=2.89\) and the fraction is \(37\,\mathrm{\%}\). The onset of the suppression is therefore expected in the range \(4\times 10^{19}\,\mathrm{eV}\) to \(10^{20}\,\mathrm{eV}\), and not at the nominal \(1.07\times 10^{20}\,\mathrm{eV}\) — which is where the observed break in Phenomenon 115.42 lies.
∎Order-of-magnitude derivation of the horizon. The mean free path is \(\ell=1/\left(n_{\mathrm{eff}}\sigma\right)\) with \(n_{\mathrm{eff}}\) the density of above-threshold photons and \(\sigma\) the photopion cross section, which near the \(\Delta\) peak is about \(5\times 10^{-32}\,\mathrm{m}^{2}\), that is \(500\) microbarn. At \(E=10^{20}\,\mathrm{eV}\) the fraction above threshold is \(37\,\mathrm{\%}\), so \(n_{\mathrm{eff}}=1.5\times 10^{8}\,/\mathrm{m}^{3}\) and
Each interaction removes only the fraction \(\kappa\approx0.2\) of the proton's energy — the inelasticity, fixed by the kinematics of Equation (115.17) applied to the \(\Delta\) decay — so the energy-loss length is \(\ell/\kappa\approx20\,\mathrm{Mpc}\). This estimate is crude in two ways that both push it upward: it uses the resonance-peak cross section for every above-threshold photon, and it treats every collision as head-on rather than averaging over angle. The proper calculation, integrating the measured cross section over the Planck spectrum and over angle, gives an energy-loss length of order \(100\,\mathrm{Mpc}\) at \(10^{20}\,\mathrm{eV}\), falling to a few tens of megaparsecs above it [Stecker:1968].
Either number is small in the only comparison that matters. The Hubble radius is
with \(H_{0}\) from [Aghanim:2020]. A particle above the threshold therefore reaches us from at most one or two percent of the radius of the observable universe: the sky at these energies is a local sky, and if a source is identified it will be a nearby one.
∎The photopion energy-loss length as a proper integral of the measured \(p\gamma\) cross section over the Planck spectrum and over collision angle, replacing the order-of-magnitude estimate given here. Belongs in Appendix A
For a nucleus the relevant process is not photopion production but photodisintegration through the giant dipole resonance, which sits at a photon energy of about \(20\,\mathrm{MeV}\) in the nucleus rest frame. A photon of laboratory energy \(E_{\gamma}\) appears in that frame boosted by up to \(2\gamma\), so the threshold is
that is \(1.5\times 10^{19}\,\mathrm{eV}\) per nucleon, and hence \(8\times 10^{20}\,\mathrm{eV}\) for an iron nucleus of \(A=56\). A heavy primary therefore survives to a higher total energy than a proton, and its attenuation, when it sets in, proceeds by shedding nucleons rather than by losing a fifth of its energy at a stroke. The shape of the predicted suppression consequently depends on the composition, which is exactly the loophole Section 115.4.5 opened and Phenomenon 115.42 must respect.
The observed suppression
The flux falls below the extrapolation of Equation (115.20) above a few times \(10^{19}\,\mathrm{eV}\): a break at about \(5.6\times 10^{19}\,\mathrm{eV}\) with a significance near \(5\sigma\) [Abbasi:2008], and a suppression above about \(4\times 10^{19}\,\mathrm{eV}\) at more than \(6\sigma\) [Abraham:2008], measured by two observatories with different techniques. Individual events well above the break exist, including one of \(3.2(9)\times 10^{20}\,\mathrm{eV}\) [Bird:1995]. Rests on Equation (115.20), Phenomenon 115.40 and Equation (115.59).
Derivation and its limits. Derives Phenomenon 115.42. The predicted position is Equation (115.56) softened by the Planck tail, Equation (115.59): an onset between \(4\times 10^{19}\,\mathrm{eV}\) and \(10^{20}\,\mathrm{eV}\). The two measurements bracket that range. The HiRes fluorescence detectors found the spectrum to break at \(5.6\times 10^{19}\,\mathrm{eV}=9.0\,\mathrm{J}\), the break being significant at about five standard deviations against a continued power law [Abbasi:2008]; the Pierre Auger Observatory, with a larger exposure and the hybrid calibration of Section 115.4.3, found the flux above \(4\times 10^{19}\,\mathrm{eV}=6.4\,\mathrm{J}\) to be suppressed relative to the extrapolation at more than six standard deviations [Abraham:2008]. The two techniques share no systematic, which is why the joint statement is stronger than either.
That events exist above the break is not a contradiction; it is required. Attenuation is exponential in distance, not a wall, so a source within \(20\,\mathrm{Mpc}\) can deliver a particle of any energy its accelerator reaches. The Fly's Eye event of \(3.2(9)\times 10^{20}\,\mathrm{eV}=51(14)\,\mathrm{J}\) [Bird:1995] is such a particle, and its existence constrains its own source to lie within the horizon of Phenomenon 115.40 — which is a strong statement, since no candidate accelerator is known within that distance.
The inference is nevertheless not closed, and the reason is Remark 115.41 together with Phenomenon 115.25. If the primaries at these energies are nuclei rather than protons, the same break is produced, equally well, by the maximum rigidity of the sources through Equation (115.22): an accelerator limited to \(R_{\max}\approx5\times 10^{18}\,\mathrm{V}\) delivers protons to \(5\times 10^{18}\,\mathrm{eV}\) and iron to \(1.3\times 10^{20}\,\mathrm{eV}\), and the superposition of the elemental cutoffs mimics a propagation suppression. The measurement therefore establishes that the spectrum is suppressed; it does not by itself establish that the mechanism is the one Greisen, Zatsepin and Kuz'min predicted. Distinguishing them requires the composition above the break, which is where the current data are weakest, and the appearance of the cosmogenic neutrinos that photopion production must also produce, which have not been detected.
∎For a decade the field held two incompatible measurements: AGASA, a scintillator array in Japan, reported a spectrum continuing unbroken past \(10^{20}\,\mathrm{eV}\) with eleven events there, while HiRes reported the break. Both could not be right. The disagreement was resolved not by a new argument but by the energy-scale analysis of Proposition 115.26: a scintillator array infers energy from a simulation of the shower's particle content at the ground, a fluorescence detector measures it calorimetrically by Equation (115.35), and the two scales differed by some \(20\,\mathrm{\%}\). By Equation (115.40) with \(\gamma\approx4.5\), a \(20\,\mathrm{\%}\) scale difference moves the flux by \(70\,\mathrm{\%}\) and moves events across a threshold; the Auger hybrid design, which measures both ways on the same showers, was built to remove exactly this ambiguity, and it did.
Anisotropy at the highest energies
The arrival directions of cosmic rays above \(8\times 10^{18}\,\mathrm{eV}\) are not isotropic: they carry a dipole of amplitude about \(6.5\,\mathrm{\%}\) whose direction lies some \(125\,^\circ\) away from the Galactic centre, with a chance probability below \(10^{-6}\) [Aab:2017]. A dipole pointing away from the Galactic centre is not what a Galactic source population produces; the origin of these particles is extragalactic. Rests on Theorem 11.46 and Definition 11.80.
Derivation: why a few percent is a detection. Derives Phenomenon 115.44. Let \(N\) events arrive with right ascensions \(\varphi_{i}\), and form the first-harmonic amplitudes
Under isotropy each \(\varphi_{i}\) is uniform, so \(\sum\cos\varphi_{i}\) and \(\sum\sin\varphi_{i}\) are sums of \(N\) independent zero-mean terms of variance \(\tfrac{1}{2}\) each; by the central limit theorem (Theorem 11.46) they are Gaussian of variance \(N/2\), and independent. Then \(Na^{2}/2\) and \(Nb^{2}/2\) are independent \(\chi^{2}\) variables of one degree of freedom (Definition 11.80), so \(Nr^{2}/2\) is \(\chi^{2}\) with two, i.e. exponential, and
This is the Rayleigh probability. With the \(N\approx3.2\times 10^{4}\) events recorded above \(8\times 10^{18}\,\mathrm{eV}\) and \(r=0.065\), Equation (115.60) gives \(\exp(-33.8)\), a number so small that it is not the quoted result: the published chance probability is far more conservative, because a three-dimensional dipole must be reconstructed from an exposure that does not cover the whole sky, and the declination component is measured much less well than the right-ascension component. Equation (115.60) shows why a percent-level anisotropy is measurable at all with tens of thousands of events — the sensitivity goes as \(N^{-1/2}\), so \(3\times 10^{4}\) events probe amplitudes near \(1\,\mathrm{\%}\) — and the published \(10^{-6}\) is the honest number after the exposure is accounted for [Aab:2017].
∎It is not an accident that a dipole and not a set of point sources is what is seen. By Equation (115.23) a proton of \(8\times 10^{18}\,\mathrm{eV}\) crossing the Galaxy in a \(3\times 10^{-10}\,\mathrm{T}\) field has a gyroradius of \(2.9\,\mathrm{kpc}\), and the deflection over a path of a few kiloparsecs is therefore tens of degrees; for an iron nucleus, \(Z=26\), it is larger by that factor and the direction is lost entirely. The extragalactic fields add more, and are less well known. The result is that a genuinely anisotropic extragalactic source distribution is smeared into precisely the lowest multipole that survives — a dipole. The absence of a confirmed point source is therefore expected, not anomalous, and no honest reading of the present data identifies one (Remark 115.38).
Neutral messengers
Gamma-ray astronomy
A photon travels in a straight line. Above the atmosphere, a gamma ray is detected by letting it convert to an \(e^{+}e^{-}\) pair in a thin foil, tracking the pair to reconstruct the incoming direction, and stopping it in a calorimeter to measure the energy. That is the Large Area Telescope on the Fermi observatory: a tungsten-and-silicon pair-conversion tracker over a caesium-iodide calorimeter, covering \(20\,\mathrm{MeV}\) to above \(300\,\mathrm{GeV}\) with a field of view of \(2.4\,\mathrm{sr}\), so that it surveys the entire sky every two orbits [Atwood:2009]. Above about \(100\,\mathrm{GeV}\) the flux is too small for any instrument that can be flown, and the ground-based Cherenkov technique of Section 115.4.4 takes over, using the atmosphere as the converter and the calorimeter both.
Two results define the current reach. The MAGIC telescopes detected teraelectronvolt emission from the gamma-ray burst GRB 190114C, establishing that the most violent transients accelerate particles to those energies [Acciari:2019]. And LHAASO, a kilometre-scale array on the Tibetan plateau, detected photons up to \(1.4\,\mathrm{PeV}=2.24\times 10^{-4}\,\mathrm{J}\) from twelve sources within the Galaxy [Cao:2021].
Photons of energy up to \(1.4\,\mathrm{PeV}\) arrive from twelve identified Galactic sources [Cao:2021]. Whatever population produces them accelerates parent particles to at least that energy, and in the hadronic case to more than ten times it — which places accelerators reaching the knee of Phenomenon 115.13 inside the Galaxy. Rests on Phenomenon 115.13 and Equation (115.57).
Derivation of the parent energy, and of the ambiguity. Derives Phenomenon 115.46. Suppose first the emission is hadronic. A proton of energy \(E_{p}\) striking interstellar gas produces pions carrying, on average, some \(20\,\mathrm{\%}\) of its energy each, and a \(\pi^{0}\) decays to two photons sharing its energy, so
and a \(1.4\,\mathrm{PeV}\) photon requires a proton of about \(1.4\times 10^{16}\,\mathrm{eV}\) — above the knee. Suppose instead the emission is leptonic, by inverse-Compton scattering of electrons on the microwave background. In the Thomson regime the scattered photon energy is \(E_{\gamma}\approx\tfrac{4}{3}\gamma_{e}^{2}\avg{E_{\gamma}^{\mathrm{ CMB}}}\), which for \(E_{\gamma}=1.4\,\mathrm{PeV}\) and Equation (115.57) would need \(\gamma_{e}=1.3\times 10^{9}\), i.e. \(E_{e}=0.66\,\mathrm{PeV}\). But at that Lorentz factor the scattering is deep in the Klein–Nishina regime, since \(\gamma_{e}\avg{E_{\gamma}^{\mathrm{CMB}}}=0.8\,\mathrm{MeV}\) exceeds \(m_{e}c^{2}=0.511\,\mathrm{MeV}\); there the electron transfers nearly all its energy in one scattering and the cross section is suppressed, so the required electron energy rises to \(E_{e}\gtrsim1.4\,\mathrm{PeV}\) and the emission becomes inefficient.
Either way the source accelerates particles to at least a petaelectronvolt. Which species it accelerates the gamma-ray spectrum alone cannot say, because both mechanisms produce power-law photon spectra of similar index over a limited band. This is the standing hadronic–leptonic ambiguity, and it is why the detection of neutrinos from a source would be decisive: Section 115.7.4 shows that neutrinos at these energies have no leptonic production channel at all.
∎Atmospheric neutrinos
The showers of Section 115.4.2 are prolific neutrino sources. Every charged pion below the critical energy \(E_{c}^{\pi}\) decays, \(\pi^{+}\to\mu^{+}\nu_{\mu}\), and every muon that decays before reaching the ground gives \(\mu^{+}\to e^{+}\nu_{e}\bar{\nu}_{\mu}\). Counting neutrinos in the full chain,
two muon-flavour neutrinos for every electron-flavour one, a prediction that involves nothing but the decay chain and is therefore reliable to a few percent at low energy. It fails at high energy for a reason already derived: a muon of energy \(E\) has a laboratory decay length \(\gamma c\tau_{\mu}=\left(E/m_{\mu}c^{2}\right)\times659\,\mathrm{m}\), which reaches the \(15\,\mathrm{km}\) of overlying atmosphere at \(E=2.4\,\mathrm{GeV}\); above a few \(\mathrm{GeV}\) most muons reach the ground and are absorbed before decaying, so the electron-flavour component is suppressed and the ratio rises above two.
A large underground water-Cherenkov detector records fewer muon-flavour atmospheric neutrino interactions from below — from neutrinos that have crossed the Earth — than from above, while the electron-flavour rate shows no such zenith dependence. The deficit depends on the ratio of path length to neutrino energy in the manner required by flavour oscillation and in no other [Fukuda:1998]. Rests on Equation (115.61).
Derivation of the oscillation phase in SI units. Derives Phenomenon 115.47. Take two neutrino mass eigenstates of masses \(m_{1}\), \(m_{2}\) and write the flavour states as \(\ket{\nu_{\alpha}}=\cos\theta\ket{\nu_{1}}+\sin\theta\ket{\nu_{2}}\), \(\ket{\nu_{\beta}}=-\sin\theta\ket{\nu_{1}}+\cos\theta\ket{\nu_{2}}\); the unitarity of the mixing matrix is licensed by Linear Algebra and Representation Theory and the physics of the mixing itself belongs to Flavour Physics and Neutrinos. A state created with definite flavour and definite momentum \(p\) evolves with the phase \(\exp\left(-\ii E_{i}t/\hbar\right)\), where
the expansion being valid because \(m_{i}c^{2}\ll pc\) for every neutrino ever detected. After a flight time \(t=L/c\) over a baseline \(L\), the accumulated phase difference is
writing \(E\approx pc\) for the neutrino energy and \(\Delta m^{2}:=m_{2}^{2}-m_{1}^{2}\). Every symbol is SI: \(\hbar c\) has dimension \(\mathrm{J}\,\mathrm{m}\), \(\Delta m^{2}c^{4}\) has dimension \(\mathrm{J}^{2}\), and \(L/E\) has dimension \(\mathrm{m}/\mathrm{J}\), so \(\Delta\Phi\) is a pure number as it must be.
The transition probability follows from the amplitude \(A_{\alpha\alpha}=\cos^{2}\theta\,\ee^{-\ii\Phi_{1}} +\sin^{2}\theta\,\ee^{-\ii\Phi_{2}}\):
so that, defining
the survival and appearance probabilities are
Equation (115.63) is the formula the literature quotes as \(1.267\,\Delta m^{2}L/E\), and the number is derived and not conventional. Put \(\Delta m^{2}c^{4}\) in \(\mathrm{eV}^{2}\), \(L\) in kilometres and \(E\) in \(\mathrm{GeV}\), and use \(\hbar c=1.9732698\times 10^{-7}\,\mathrm{eV}\,\mathrm{m}\):
Derivation of the zenith-angle deficit. Atmospheric neutrinos are produced at an altitude of about \(15\,\mathrm{km}\) over the whole Earth. One arriving from directly overhead has travelled \(L\approx15\,\mathrm{km}\); one arriving from directly below has crossed a diameter, \(L\approx2R_{\oplus}=1.27\times 10^{4}\,\mathrm{km}\). The path length varies by three orders of magnitude with zenith angle, and the neutrino energy does not.
Take the atmospheric mass splitting \(\Delta m^{2}c^{4}\approx2.5\times 10^{-3}\,\mathrm{eV}^{2}\) and \(E=1\,\mathrm{GeV}\). For the downward-going neutrino,
so no oscillation is observable. For the upward-going one,
which is many complete oscillations; averaged over the spread in \(L\) and \(E\) within any experimental bin, \(\avg{\sin^{2}\Phi}=\tfrac{1}{2}\) and
which for maximal mixing is one half. The prediction is therefore specific and asymmetric: the upward-going muon-neutrino rate is halved while the downward-going rate is untouched, and the electron-flavour rate is unaffected at either zenith angle because the corresponding splitting is thirty times smaller and the phase is correspondingly negligible over these baselines. That is exactly the pattern Super-Kamiokande measured [Fukuda:1998], and it is the first compelling evidence that neutrinos have mass.
∎The apparatus, the data and the uncertainty budget of that measurement belong to Experiment: Neutrino Oscillations, and the theory of mixing — three flavours, the PMNS matrix, matter effects — to Flavour Physics and Neutrinos; this chapter supplies the source. It is worth recording that the solar neutrino deficit, measured by Davis's chlorine experiment from 1968 [Davis:1968] [Cleveland:1998] and resolved by the SNO measurement of the neutral-current rate [Ahmad:2002], is the independent half of the same story; solar neutrinos are not cosmic rays and are not treated here, but the two deficits are explained by one mechanism and neither alone would have been believed.
For everything that follows, the atmospheric flux changes role: it stops being the signal and becomes the irreducible background to Section 115.7.4.
Supernova 1987A
On 23 February 1987 a star exploded in the Large Magellanic Cloud, at a distance of about \(50\,\mathrm{kpc}=1.54\times 10^{21}\,\mathrm{m}\). Some three hours before the optical brightening, two underground water-Cherenkov detectors built to look for proton decay recorded bursts of events: eleven in Kamiokande-II over \(13\,\mathrm{s}\) [Hirata:1987] and eight in the IMB detector over \(6\,\mathrm{s}\) [Bionta:1987], with energies between about \(7\,\mathrm{MeV}\) and \(40\,\mathrm{MeV}\). Nineteen events remain, four decades later, the entire observational sample of extragalactic neutrino astronomy below the teraelectronvolt scale, and they carry a remarkable amount of physics.
A core-collapse supernova at \(50\,\mathrm{kpc}\) produced a neutrino burst of duration of order \(10\,\mathrm{s}\) carrying a total energy of order \(3\times 10^{46}\,\mathrm{J}\), detected in coincidence at two independent detectors on opposite sides of the Earth [Hirata:1987] [Bionta:1987]. The energy matches the gravitational binding energy of a neutron star, and the arrival timing bounds both the neutrino speed and the neutrino mass.
Derivation: the energy is the binding energy of a neutron star. Derives Phenomenon 115.48. If a stellar core of mass \(M\) collapses to a uniform sphere of radius \(R\), the gravitational energy released is
the standard self-energy integral. With \(M=1.4M_{\odot} =2.78\times 10^{30}\,\mathrm{kg}\) and \(R=10\,\mathrm{km}\),
which is \(12\,\mathrm{\%}\) of the rest energy \(Mc^{2}=2.50\times 10^{47}\,\mathrm{J}\) of the collapsing core. Essentially none of it emerges as light: the total optical output of the supernova is smaller by three orders of magnitude, and the kinetic energy of the ejecta by two — the \(10^{44}\,\mathrm{J}\) used in Phenomenon 115.29. It emerges as neutrinos, because at the densities of the collapsing core nothing else escapes. The total energy inferred from the nineteen detected events, given the detector masses, the cross section for \(\bar{\nu}_{e}p\to ne^{+}\) and the assumption of equipartition among the six neutrino species, is of order \(3\times 10^{46}\,\mathrm{J}\). The agreement is the confirmation, and it is the only direct observation ever made of the mechanism of a core-collapse supernova.
∎Derivation: the bound on the neutrino speed. The neutrinos preceded the first optical light by about \(\Delta t=3\,\mathrm{h}=1.1\times 10^{4}\,\mathrm{s}\), and that lead is astrophysical — the shock takes hours to reach the stellar surface after the core collapses, whereas neutrinos leave immediately. Taking the lead as an upper bound on any difference in travel time over the flight time
gives
The bound is weak as a fractional velocity but it is obtained over a baseline of a hundred and sixty thousand light years, which no terrestrial measurement can approach, and it is the reason a claim of superluminal neutrino propagation is checked against SN 1987A first.
∎Derivation: the bound on the neutrino mass. A massive neutrino is dispersive. From \(E=\sqrt{p^{2}c^{2}+m^{2}c^{4}}\) the group velocity is
so a neutrino of energy \(E\) arrives later than a massless one by
Two neutrinos of energies \(E_{1}<E_{2}\) emitted simultaneously therefore arrive separated by
The observed burst spanned \(\Delta t\approx13\,\mathrm{s}\) with energies from about \(7.5\,\mathrm{MeV}\) to \(35\,\mathrm{MeV}\). Requiring the dispersion not to exceed the observed spread,
whence
The bound is honest only under an assumption that cannot be checked: that the burst was emitted in a time short compared with \(13\,\mathrm{s}\). Since the emission is expected to last of order that long by itself, the true limit is weaker, and careful analyses that model the emission profile obtain a few electronvolts rather than seventeen. It is quoted here because it is the only kinematic neutrino mass bound ever obtained over an astronomical baseline, and because the method — dispersion over a known distance — is the one used again for photons in Section 115.8.2.
∎Equation (115.67) has been superseded twice over. The kinematic endpoint of tritium beta decay gives \(m_{\nu}c^{2}<0.8\,\mathrm{eV}\) at \(90\,\mathrm{\%}\) confidence [Aker:2022], and the cosmological sum of neutrino masses is bounded well below that [Aghanim:2020]. Oscillation experiments — Equation (115.64) — measure only \(\Delta m^{2}\) and say nothing about the absolute scale, which is why the direct bounds remain necessary. SN 1987A's value today is not its number but its method, and the nineteen events remain the only extragalactic neutrino source ever identified below the teraelectronvolt scale.
High-energy astrophysical neutrinos
A neutrino is neutral, so it points; it is weakly interacting, so it escapes from regions opaque to photons and is not attenuated by the microwave background. The same weakness is the difficulty: a kilometre-scale target is needed. IceCube is a cubic kilometre of Antarctic ice — \(10^{9}\,\mathrm{m}^{3}\), a mass of \(9.2\times 10^{11}\,\mathrm{kg}\) — instrumented with \(5160\) digital optical modules on \(86\) vertical strings between \(1450\,\mathrm{m}\) and \(2450\,\mathrm{m}\) depth, each module a photomultiplier reading the Cherenkov light of Example 115.24 from charged secondaries. The detector concept, and the performance of the nine-string configuration with which it began operating, are described in [Achterberg:2006]; the figures above are those of the completed array.
A cubic kilometre of instrumented Antarctic ice records neutrino events in excess of the atmospheric background: \(28\) events with deposited energies between \(30\,\mathrm{TeV}\) and \(1.2\,\mathrm{PeV}\) against an expectation of about \(10\), a \(4\sigma\) excess [Aartsen:2013]. Neutrinos at these energies are produced only in hadronic collisions, so their detection is direct evidence that protons and nuclei, and not only electrons, are accelerated to such energies somewhere in the universe. Rests on Equation (115.47).
Derivation: why the background falls away above \(100\) teraelectronvolts. Derives Phenomenon 115.50. The atmospheric neutrino flux of Section 115.7.2 is not simply a copy of the primary spectrum. A pion of energy above the critical energy \(E_{c}^{\pi}\approx24\,\mathrm{GeV}\) derived in Section 115.4.2, conventionally rounded to \(20\,\mathrm{GeV}\), interacts before it decays, and the probability that it decays instead is the ratio of the two lengths,
falling as \(1/E\). The neutrino flux is therefore the parent spectrum \(E^{-2.7}\) multiplied by \(E^{-1}\):
one power steeper than the cosmic-ray spectrum that makes it. An astrophysical flux produced by shock acceleration, by contrast, carries the source index of Equation (115.47), near \(E^{-2}\) and certainly harder than \(E^{-3}\), and suffers no such suppression because its pions decay in vacuum. Two power laws differing by more than one unit of index must cross, and they cross near \(100\,\mathrm{TeV}\): below it the sky is atmospheric, above it the atmospheric flux has fallen away and an astrophysical component, if it exists, dominates. That is why the IceCube analysis selects events above \(30\,\mathrm{TeV}\) and why the excess is at the high-energy end of the selection.
A second discriminant is available and is used. An atmospheric neutrino is made in an air shower, and the same shower makes muons that arrive at the detector at the same time. Selecting only events whose interaction vertex is inside the instrumented volume, with no entering track, therefore vetoes atmospheric neutrinos arriving from above against their own accompanying muons — the self-veto — while an astrophysical neutrino, which has no accompanying shower, passes. The veto is what reduces the expected background of the Phenomenon 115.50 sample to about ten events.
∎Derivation: a neutrino detection is evidence of hadronic acceleration. Neutrinos at these energies have exactly one production channel: the weak decay of pions and kaons, which are made only in hadronic collisions. Electrons produce photons — by synchrotron radiation, by bremsstrahlung, by inverse-Compton scattering — and produce no neutrinos at all, because none of those processes has a weak vertex. Detecting a teraelectronvolt neutrino from a direction therefore establishes that protons or nuclei are accelerated there, and resolves the hadronic–leptonic ambiguity of Phenomenon 115.46 in one measurement.
The same chain fixes the expected ratio of neutrinos to gamma rays. In a \(pp\) collision the three pion charge states are produced in roughly equal numbers. The \(\pi^{0}\) gives two photons carrying all its energy; each \(\pi^{\pm}\) gives, through \(\pi^{+}\to\mu^{+}\nu_{\mu}\) and \(\mu^{+}\to e^{+}\nu_{e}\bar{\nu}_{\mu}\), three neutrinos carrying about three quarters of its energy. Hence
a source's all-flavour neutrino luminosity is comparable to its hadronic gamma-ray luminosity, which is what makes a joint detection feasible at all. The flavour composition at the source is \(\nu_{e}:\nu_{\mu}:\nu_{\tau}=1:2:0\) from the same counting, and oscillation over astronomical baselines — where Equation (115.63) gives an enormous \(\Phi\) and every \(\sin^{2}\Phi\) averages to one half — converts this to approximately \(1:1:1\) at Earth. The measured flavour ratio is consistent with that, which is a non-trivial check that the sources are pion sources.
∎Multimessenger astronomy
On 22 September 2017 IceCube recorded a muon-neutrino event of reconstructed energy \(290\,\mathrm{TeV}=4.6\times 10^{-5}\,\mathrm{J}\) and issued an automated alert within a minute. Follow-up by some twenty instruments from radio to gamma rays found the blazar TXS 0506+056 — an active galaxy whose jet points nearly at us — in a gamma-ray flaring state, coincident in direction with the neutrino. The chance-coincidence probability, allowing for the number of blazars and the number of alerts, is at the level of \(3\sigma\) [IceCube:2018]. A search of the archival IceCube data then found, independently, an excess of \(13\pm5\) events from the same direction during 2014–2015 — a period in which the source was not flaring in gamma rays — at \(3.5\sigma\) [Aartsen:2018].
The two observations are statistically independent: different years, different event samples, different analyses, one a directional coincidence with a flare and the other a time-integrated point-source search. Taken together they are the strongest candidate identification of a source of high-energy cosmic rays that exists. Taken separately neither reaches the threshold this book would require to call something observed, and the combination of two \(3\sigma\) results is not a \(6\sigma\) result — significances add in quadrature only for the statistic, not for the trials factor of Section 11.6, and the archival search was made in the direction the first result selected, which is precisely the situation the look-elsewhere correction exists to handle. The honest statement is that TXS 0506+056 is a candidate, that no other source has been identified since, and that the diffuse flux of Phenomenon 115.50 remains unresolved into sources.
The method, however, is established, and its clearest success is elsewhere: the coincidence of a gravitational-wave signal with a short gamma-ray burst, treated in Experiment: Gravitational Waves, in which two independent messengers from one event were combined to identify the source and to bound the propagation speed of gravity. Multimessenger astronomy is that method, and cosmic rays supply two of its four channels.
Cosmic rays as a physics laboratory
Centre-of-mass energies beyond the colliders
A proton of \(10^{20}\,\mathrm{eV}\) striking a nucleon at rest in the atmosphere collides at a centre-of-mass energy \(\sqrt{s}=433\,\mathrm{TeV}\), some thirty-two times the \(13.6\,\mathrm{TeV}\) at which the Large Hadron Collider now runs and fifty-four times the \(8\,\mathrm{TeV}\) of the Higgs discovery datasets of Experiment: The Higgs Boson Discovery. Rests on Equation (40.12).
Derives Proposition 115.52. By Equation (40.12) for a projectile of energy \(E\) on a target of mass \(m_{p}\) at rest,
so
in SI \(6.9\times 10^{-5}\,\mathrm{J}\). The square root is the whole story: a fixed-target collision converts energy to centre-of-mass energy only as \(\sqrt{E}\), which is why colliders were built (Example 40.13), and why even a beam seven orders of magnitude above the LHC's buys only a factor of thirty in \(\sqrt{s}\).
∎That factor is nevertheless real, and two measurements exploit it.
The first is the proton–air cross section, extracted from the tail of the \(X_{\max}\) distribution. A shower whose first interaction happens deep in the atmosphere has a deep maximum, and since the first-interaction depth is exponentially distributed with mean \(\lambda_{I}=\bar{m}_{\mathrm{air}}/\sigma_{p\text{-air}}\), the deep tail of the measured \(X_{\max}\) distribution falls exponentially with a slope that measures \(\lambda_{I}\) and hence \(\sigma_{p\text{-air}}\) directly. Converting to a proton–proton cross section requires a nuclear-overlap calculation, and the results agree with the extrapolation of accelerator measurements [Navas:2024] — a consistency check of the strong interaction three decades in \(\sqrt{s}\) beyond where it was calibrated.
The second is the muon content, and it does not agree.
The muon number predicted by Equation (115.32), with \(N_{\mathrm{ch}}\) and the inelasticity taken from every current hadronic interaction model tuned to LHC data, falls below the measured number in showers above \(10^{19}\,\mathrm{eV}\) by a factor between about \(1.3\) and \(1.6\), with the primary mass fixed independently by \(X_{\max}\) through Phenomenon 115.25. Since \(N_{\mu}\) depends on the mass only through the mild \(A^{1-\beta}\) of Equation (115.33), the discrepancy cannot be absorbed into the composition without contradicting the depth measurement. It is therefore a failure of the hadronic modelling at energies no collider reaches, and it is the one place in this chapter where an extrapolated Standard Model calculation is measurably wrong. It is reported as such.
Tests of Lorentz invariance
The invariance postulates of Minkowski Space and Its Symmetries assert that every species has the same limiting speed. Suppose instead that species \(a\) has maximal attainable velocity \(c_{a}=c\left(1+\delta_{a}\right)\) with \(\abs{\delta_{a}}\ll1\), so that its dispersion relation reads \(E^{2}=p^{2}c_{a}^{2}+m^{2}c_{a}^{4}\). Cosmic rays test this at a sensitivity no laboratory can approach, for a simple reason: the effect of \(\delta\) grows with energy while every terrestrial measurement is made at low energy.
Derivation of the bound from the photopion threshold. Redo the threshold calculation of Equation (115.56) with species-dependent limiting speeds. Writing \(\delta:=\delta_{\pi}-\delta_{p}\), the invariant that must reach \(\left(m_{p}+m_{\pi}\right)^{2}c^{4}\) acquires an extra term of order \(\delta E^{2}\), so the threshold condition becomes
For \(\delta>0\) the reaction is forbidden above a maximum energy rather than allowed above a minimum one, and the photopion suppression would not occur at all. Requiring the second term not to dominate the first at the energy where the suppression is observed, \(E\approx5\times 10^{19}\,\mathrm{eV}\), gives
a bound on a fractional velocity difference at the level of \(10^{-23}\) [Coleman:1999]. For comparison, the best terrestrial Michelson–Morley-type experiments — rotating optical cavities, the modern descendants of [Brillet:1979] — bound anisotropies of the speed of light at the level of \(10^{-17}\) [Herrmann:2009]; the cosmic-ray bound is better by six orders of magnitude, and it is obtained not by building anything but by observing that a predicted suppression occurs.
∎A second and independent channel is the propagation of high-energy photons over cosmological distances. If the photon dispersion relation acquires an energy-dependent correction, photons of different energy emitted simultaneously arrive at different times, and the delay accumulates over the flight path exactly as in Equation (115.66). Gamma-ray bursts, which are both distant and variable on sub-second time scales, are the natural clock, and the observed simultaneity of low- and high-energy photons from them bounds the effect stringently [Abdo:2009].
Both channels give null results, and the resulting bounds are among the strongest constraints on any modification of special relativity. This chapter records that fact and nothing more: no Lorentz-violating framework is developed here, because none has evidence, and the programmes that predict such violations lie outside the scope stated in the front matter. What is in scope is the measurement, and the measurement says the postulates of Minkowski Space and Its Symmetries hold to Equation (115.70).
Muon time dilation as a standing test
Muons produced by cosmic-ray interactions near \(15\,\mathrm{km}\) altitude are detected in quantity at sea level — about \(70\,/\mathrm{m}^{2}/\mathrm{s}/\mathrm{sr}\) from the vertical, or roughly one per square centimetre per minute on a horizontal surface — and the surviving fraction depends on their momentum [Rossi:1941] [Frisch:1963]. The muon lifetime at rest is \(2.197\,\mu\mathrm{s}\) [Navas:2024], and muons circulating in a storage ring at \(\gamma\approx29.3\) are observed to decay with that same proper lifetime, dilated in the laboratory by exactly the factor \(\gamma\) [Bailey:1977]. Rests on Equation (115.3).
Derivation. Derives Phenomenon 115.55. The PDG quotes the muon width as \(\Gamma_{\mu}=2.9959836(30)\times 10^{-19}\,\mathrm{GeV}\) [Navas:2024], which is a lifetime
and the distance a muon covers in one lifetime at a speed arbitrarily close to \(c\) is \(c\tau=658.6\,\mathrm{m}\). Were the laboratory lifetime the same as the proper one, the surviving fraction over \(15\,\mathrm{km}\) of atmosphere would be
and no muons would reach the ground in any detectable number. What the flight consumes is instead the muon's proper time, shorter than the laboratory time by the factor \(\gamma\) (Experiment: Time Dilation and Relativistic Kinematics), so the laboratory decay length is \(\gamma\beta c\tau\); for a typical sea-level muon of a few gigaelectronvolts, \(\gamma\approx20\) — that is \(E=20\,m_{\mu}c^{2}=2.1\,\mathrm{GeV}\), which is also about the energy a muon needs to survive the \(2\,\mathrm{MeV}\) per \(\mathrm{g}/\mathrm{cm}^{2}\) ionisation loss across the atmosphere of Equation (115.3) — and
The observed flux is of this order. The nine and a half orders of magnitude between the two estimates are what make the sea-level muon flux a continuously running test of relativistic kinematics rather than a curiosity, and the momentum dependence of the surviving fraction is the form in which the test was first made quantitative [Rossi:1941].
∎The three measurements that turn this into a quantitative test are sections of Experiment: Time Dilation and Relativistic Kinematics: Rossi and Hall's momentum dependence (Section 41.2), the Frisch–Smith mountain-to-sea-level comparison (Section 41.3) and the CERN storage-ring lifetime (Section 41.4). What this chapter adds is the source: the muons are not supplied by an apparatus, they are supplied by the atmosphere, at a rate fixed by Equation (115.19) and a chain of decays that nothing on Earth controls.
What is settled and what is not
It is worth separating, at the close of the chapter, the statements that rest on measurement from those that do not.
Settled by observation: cosmic rays arrive from outside the atmosphere (Phenomenon 115.2) and are charged and predominantly positive (Phenomena 115.4 and 115.5); the spectrum is a power law over eleven decades with a knee, an ankle and a high-energy suppression (Phenomena 115.11 and 115.13); the flux is suppressed at the very top (Phenomenon 115.42); a single primary makes an extensive air shower whose particle number measures its energy and whose depth of maximum measures its mass (Phenomena 115.18 and 115.25); supernova remnants accelerate protons (Phenomenon 115.30) and can supply the Galactic energy budget (Phenomenon 115.29); petaelectronvolt accelerators exist in the Galaxy (Phenomenon 115.46); the arrival directions above \(8\times 10^{18}\,\mathrm{eV}\) carry a dipole of extragalactic origin (Phenomenon 115.44); an extraterrestrial high-energy neutrino flux exists (Phenomenon 115.50); and atmospheric neutrinos oscillate (Phenomenon 115.47).
Not settled, and stated as such throughout:
-
Which of the two magnetic readings of the knee is correct. Both predict the rigidity ordering that is observed, and nothing in the present data separates them (Remark 115.14).
-
The composition above \(10^{19}\,\mathrm{eV}\), and therefore whether the observed suppression is the propagation effect of Phenomenon 115.40 or the maximum rigidity of the sources (Phenomenon 115.42).
-
The identity of any source of the highest-energy particles. No point source has been confirmed, and the Hillas criterion excludes rather than identifies (Remark 115.38).
-
The origin of the rising positron fraction (Remark 115.17).
-
Why the observed spectral index is \(2.7\) when shock acceleration and the measured escape time predict \(2.33\) (Remark 115.35).
-
Why showers contain more muons than every hadronic model predicts (Remark 115.53).
Six open problems in a subject a century old is not a poor showing, and it is a fair description of where the field stands. The instrument that will settle most of them is exposure: the flux at the top of the spectrum is one particle per square kilometre per century, and every statement about it is limited by how much sky has been watched for how long.