Experiment: The Photoelectric Effect and Compton Scattering

Contents
  1. Hertz and Hallwachs: ultraviolet light discharges a metal (1887–1888)
  2. Lenard: the energy of the electrons ignores the intensity (1902)
  3. Millikan: the slope of the line is $h/e$ (1916)
  4. Compton: the wavelength of scattered X-rays changes (1923)
  5. Bothe and Geiger: the two products appear together (1925)
  6. Grangier, Roger and Aspect: a single quantum is not split (1986)
  7. What the measurements settle

The Photon: Photoelectric and Compton Effects states a hypothesis and derives its consequences: light of frequency \(\nu\) is emitted and absorbed in indivisible amounts \(h\nu\) carrying momentum \(h\nu/c\), from which follow the photoelectric equation, the Compton shift, and the requirement that energy and momentum balance in each single scattering act. It names the experiments that bear on all three and reports none of them. This chapter is that report.

It is organized around one question, and the question is sharper than the textbook treatment usually admits: which of these measurements actually requires the light to be quantized, as opposed to the matter that absorbs it. The answer is unkind to the most famous of them. The photoelectric effect, in every form Hertz, Lenard and Millikan gave it, is reproduced by a quantized atom absorbing from an entirely classical field [Lamb:1969], so the first three sections below establish the equation of Phenomenon 69.2 and the constant in it without establishing the light quantum at all. What carries the weight is momentum and indivisibility: the wavelength shift of Section 74.4, which no driven classical oscillator can produce; the coincidence experiment of Section 74.5, which excludes conservation-on-the-average and with it the last serious counter-theory; and the anticorrelation of Section 74.6, which is the modern experiment that no classical field of any strength or statistics can reproduce. The sections run in that order, which is also chronological, and the closing Section 74.7 states what each is and is not entitled to claim.

Derivation pending.

Experiment: The Photoelectric Effect and Compton Scattering: what is pending is the tabulation of the measured numbers for each experiment, in SI with its uncertainty budget — stopping potential against frequency for each emitting metal, scattered wavelength against scattering angle, coincidence and singles counts against resolving time, and the gated detection probabilities of the beamsplitter experiment. The phenomena stated below carry their derivations inline except where a pending derivation is marked against them individually.

Hertz and Hallwachs: ultraviolet light discharges a metal (1887–1888)

The effect was found by a man who did not want it. Hertz was engaged in the experiment that produced the first electromagnetic waves and confirmed Maxwell's theory; the photoelectric effect entered his apparatus as an interference, was isolated because it would not go away, and was reported as an anomaly whose explanation he could not offer [Hertz:1887]. Hallwachs turned it into a phenomenon the following year with an apparatus a school can build [Hallwachs:1888].

Apparatus

Hertz's arrangement is two spark gaps. An induction coil drives the primary gap, whose oscillatory discharge radiates; a wire loop of adjustable size serves as resonator, and the length of the spark that can be drawn across a micrometer gap in that loop measures the induced electromotive force. The relevant addition is optical, and it is the part that matters here: screens interposed between the two gaps, of glass and of quartz, and a quartz prism with which the light of the primary spark can be dispersed and one region of the spectrum at a time allowed to fall on the micrometer gap. Hallwachs' apparatus is a freshly polished zinc plate mounted on an insulating support and connected to a gold-leaf electroscope, illuminated by the light of a carbon arc.

Procedure

Hertz measured the maximum length of the resonator spark with the micrometer gap exposed to the light of the primary spark and again with it shielded. The diagnostic sequence is the interesting part: enclosing the resonator gap in an opaque box shortens the spark; the effect survives when the box is fitted with a quartz window and disappears when the window is glass; and dispersing the light with a quartz prism locates the active agent beyond the violet end of the visible spectrum. Hallwachs charged his zinc plate negatively, watched the electroscope leaf under illumination, then repeated with the plate charged positively, and finally with the plate initially uncharged.

Observations and data

Illumination of the micrometer gap by ultraviolet light lowers the potential difference at which the spark passes, so that a longer gap sparks at the same drive; glass extinguishes the effect and quartz does not, which places the agent in the ultraviolet [Hertz:1887]. On Hallwachs' plate the sign of the charge decides everything: a negatively charged plate loses its charge under illumination, a positively charged plate does not, and an initially uncharged plate charges up positively [Hallwachs:1888]. The carriers liberated are therefore negative.

[Reserved: the tabulation. For Hertz, the maximum resonator spark length in metres against the screening condition and against the position on the dispersed spectrum, with the reproducibility of the spark-length measurement, which is the whole uncertainty of the experiment. For Hallwachs, the rate of fall of the electroscope leaf against illumination and against the sign and magnitude of the initial charge, converted to a charge loss per second in SI. Neither original was consulted for this chapter and no figures are quoted from memory.]

Phenomenon 74.1 (Ultraviolet light liberates negative charge from a clean metal surface).

Light beyond the violet end of the visible spectrum, falling on a clean metal surface, releases negatively charged carriers from it. The release is blocked by glass and passed by quartz, which identifies the agent as ultraviolet [Hertz:1887]. Because the carriers are negative, an isolated illuminated conductor loses negative charge if it carries any and otherwise acquires a positive charge, while a positively charged conductor is not discharged at all [Hallwachs:1888].

Derivation. Only the sign asymmetry needs an argument, and it is a one-line retarding-field argument that will be used quantitatively in Section 74.2 and Section 74.3. Let the liberated carriers have charge \(-e\) and leave the surface with kinetic energies up to some maximum \(E_{\max}\), and let the isolated conductor sit at potential \(V\) relative to its surroundings. A carrier escaping to infinity must do work \(eV\) against the field of the conductor it has just left more negative — or, for \(V<0\), gains that energy. Hence carriers escape whenever

\begin{equation}\tag{74.1} E_{\max}>eV\ep \end{equation}

If \(V<0\) the condition holds for every carrier and the conductor discharges. If \(V=0\) it holds and the conductor charges positive, which raises \(V\) until Equation (74.1) fails and the process stops by itself — the plate reaches a steady potential \(V=E_{\max}/e\) and no more. If \(V>0\) and already exceeds \(E_{\max}/e\) the carriers cannot leave at all, and a positively charged plate keeps its charge. All three of Hallwachs' observations are the three cases of one inequality, and the same inequality read as a measurement is the stopping-potential method: \(E_{\max}\) is obtained by finding the \(V\) at which emission just ceases.

Nothing here identifies the carriers or requires anything quantum. What Equation (74.1) does establish is that the interesting quantity is \(E_{\max}\) and that it is measurable with an electrometer, which is the whole experimental programme of the next two sections.

Interpretation

In 1887 this is an anomaly and nothing more; there is no theory it contradicts, because there is no theory of it. Its importance is that it makes a quantity measurable. The identification of the carriers came with the corpuscle [Thomson:1897]: their charge-to-mass ratio is that of cathode rays, so the photoelectric current is a current of electrons and the phenomenon is one of the atom's outer structure, treated in Atomic Models and Spectra and, for the metallic case, in Electrons in Solids: Band Theory. It is worth recording that the discovery is incidental in the strict sense: Hertz was verifying Maxwell's equations, the effect degraded his measurements, and he published it because he could not explain it. Epistemology and the Scientific Method takes up what this pattern implies about the difference between a prediction confirmed and an anomaly recorded.

Primary references

[Hertz:1887] [Hallwachs:1888]. The identification of the carriers with the electron rests on [Thomson:1897].

Lenard: the energy of the electrons ignores the intensity (1902)

This is the measurement that made the effect a problem. Hertz and Hallwachs establish that light liberates electrons; Lenard establishes how much energy they carry, and the answer contradicts the only theory of light then available [Lenard:1902].

Apparatus

An evacuated tube containing the illuminated cathode and a collecting electrode, with a quartz window admitting the light — quartz because glass absorbs precisely the wavelengths that work, as Section 74.1 had shown. The source is a carbon arc, mounted so that it can be moved towards and away from the window. A variable potential difference is applied between cathode and collector so that the emitted electrons may be retarded, and the current is read on a quadrant electrometer, the currents being of the order of picoamperes rather than milliamperes. A magnetic field can be applied across the electron path so that the deflection of the beam gives the charge-to-mass ratio of the carriers independently.

Procedure

Three variables are separated. First the retarding potential is raised until the collected current vanishes; by Equation (74.1) that potential is \(E_{\max}/e\) and is the quantity of interest. Then the intensity is changed — by moving the arc towards or away from the window, which changes the irradiance at the cathode by the inverse square of the distance and changes nothing else — and the stopping potential is measured again. Then the spectral composition is changed, by interposing absorbing screens or by changing the source, and the stopping potential measured again. Separately, the time between the first illumination of a shuttered cathode and the first detectable current is looked for.

Observations and data

The photocurrent is proportional to the intensity of the illumination, as any theory would have it. The stopping potential is not: moving the arc changes the current in the ratio the inverse-square law demands and leaves the potential at which the current vanishes unaltered. Changing the spectral composition towards shorter wavelengths raises it. The charge-to-mass ratio of the carriers agrees with that of cathode rays. No delay between illumination and emission is detectable [Lenard:1902].

[Reserved: the tabulation. Stopping potential in volts against source distance — that is, against irradiance in \(\mathrm{W}/\mathrm{m}^{2}\) at the cathode — over the accessible range, showing the constancy; photocurrent in amperes over the same range, showing the proportionality; stopping potential against the spectral band admitted; and the measured charge-to-mass ratio with its uncertainty, against the cathode-ray value. The upper bound on the emission delay, in seconds, is to be given as the resolving time of the electrometer circuit rather than as a measured delay, since no delay was seen.]

Phenomenon 74.2 (The maximum electron energy is fixed by the colour of the light and not by its intensity).

For light of fixed spectral composition, increasing the irradiance at the emitting surface increases the number of photoelectrons per second in proportion and leaves their maximum kinetic energy unchanged; the maximum energy is raised only by shifting the light towards shorter wavelengths [Lenard:1902]. This is item 2 of Phenomenon 69.1 and it is the observation with which no theory treating light as a continuous wave can be reconciled.

Derivation. The two predictions must be set side by side, because the force of the result lies in their disagreement and not in either alone.

Under the hypothesis of The Photon: Photoelectric and Compton Effects the energy available to one electron is the energy of one quantum, so Equation (69.2) gives \(E_{\max}=h\nu-\phi\), in which the irradiance does not appear. Raising the irradiance multiplies the number of quanta crossing the surface per second and therefore the number of electrons ejected per second; it cannot enlarge the parcel any one of them receives. Proportionality of the current to the intensity and constancy of \(E_{\max}\) are then the same statement.

Under any classical field the energy delivered to an electron is set by the work done on it by the field, which is quadratic in the field amplitude and therefore proportional to the irradiance. Doubling the irradiance doubles the rate at which energy arrives at every point of the surface, and there is no mechanism in Maxwell's equations by which an electron can decline the extra energy: \(E_{\max}\) must rise with the intensity. It does not. The classical picture also predicts that \(E_{\max}\) should be insensitive to the frequency, which is the reverse of what is measured.

Phenomenon 74.3 (Emission begins at once, however feeble the light).

The photocurrent appears as soon as the light does, within the resolving time of the measuring circuit, and no accumulation period is observed at any intensity at which the current is detectable at all [Lenard:1902]. This is item 4 of Phenomenon 69.1.

Derivation. The estimate this observation contradicts is worth doing explicitly, because it is the sharpest form of the contradiction and it is arithmetic. Let a classical wave of irradiance \(I\) fall on the surface. An electron can absorb only the energy flowing through its own catchment area; take that area to be atomic, \(A\approx\pi a^{2}\) with \(a\approx10^{-10}\,\mathrm{m}\), so \(A\approx3\times 10^{-20}\,\mathrm{m}^{2}\). To escape it must accumulate the work function, \(\phi\approx2\,\mathrm{eV}\approx3.2\times 10^{-19}\,\mathrm{J}\) for an alkali surface. The waiting time before the first electron can appear is therefore

\begin{equation}\tag{74.2} t\approx\frac{\phi}{IA}\ec \end{equation}

which for a feeble but perfectly ordinary illumination of \(I=10^{-2}\,\mathrm{W}/\mathrm{m}^{2}\) gives \(t\approx10^{3}\,\mathrm{s}\) — of order a quarter of an hour of darkness before the first count. At \(I=1\,\mathrm{W}/\mathrm{m}^{2}\) it is still of order \(10\,\mathrm{s}\). Nothing observed resembles either.

Two objections must be closed, or Equation (74.2) is only a straw man. First, one might let the metal collect over a larger area and funnel the energy to one electron; but a classical field exerts its force locally on each electron, and there is no interaction in classical electrodynamics that concentrates the energy of a wavefront onto a single charge faster than it arrives. Second, one might appeal to electrons that were already energetic before the light arrived; the thermal energy available at room temperature is \(k_{\text{B}}T\approx0.025\,\mathrm{eV}\), two orders of magnitude below \(\phi\), and an appeal to the tail of the distribution predicts a current that depends violently on temperature, which is not observed.

The quantum reading has no waiting time to explain. One quantum is absorbed whole or not at all, so the first electron may leave at the instant the first quantum arrives, and lowering the intensity reduces the rate of emission without introducing any delay in the individual act. The full statement of this argument, including the demonstration that no redistribution within the classical field shortens Equation (74.2), is the pending derivation attached to Phenomenon 69.1 in the theory chapter; what is given here is the order-of-magnitude form that the experiment actually confronts.

Interpretation

Lenard's result is the birth of the problem and, historically, not of its solution: he interpreted the emission as a triggering of electrons already in motion within the atom, the light merely releasing them, and he was still defending that reading long after it had failed. The reading fails on the frequency dependence, which a trigger cannot explain, and it fails on the threshold. But the sceptical structure of Lenard's argument should be preserved rather than smoothed away, because a great deal of what is loosely credited to the photoelectric effect is not established by it: what Phenomenon 74.2 excludes is a classical field delivering energy continuously to a target, and that is not the same as establishing that the field is granular. The energy could be granular because the absorber is. That is exactly the loophole Section 69.2.4 names and Section 74.6 closes.

Primary references

[Lenard:1902].

Millikan: the slope of the line is $h/e$ (1916)

Millikan spent a decade on this measurement in order to disprove the hypothesis it confirms, and the experiment is the better for it: no one has ever had a stronger motive to find the linear law false [Millikan:1916]. What he obtained was the best value of Planck's constant then available by any method, from an experiment in optics and electrometry that touches no cavity and no thermometer.

Apparatus

The difficulty is entirely one of surfaces. An alkali metal exposed to air acquires an oxide layer in minutes, and the layer changes the work function by more than the effect being measured, so the experiment must present a clean surface to the light and keep it clean for the duration of a run. Millikan's answer was to put a machine shop inside the vacuum: a sealed and evacuated glass vessel containing cylinders of sodium, potassium and lithium mounted on a rotatable wheel, together with a cutting knife operated from outside by an electromagnet, so that a fresh surface could be shaved and turned to face the light without ever admitting air. The optical train selects one line at a time from the spectrum of a mercury arc, spanning the visible and the near ultraviolet from about \(550\,\mathrm{nm}\) to about \(250\,\mathrm{nm}\). The emitted electrons are collected in a Faraday cylinder, the current read on a quadrant electrometer, and a known and variable retarding potential is applied between the illuminated surface and the collector.

Procedure

For each metal and each spectral line: shave a fresh surface, admit the line, and record the collected current as the retarding potential is raised, until it vanishes. The potential at which it does is the stopping potential \(V_{0}\) of Equation (69.1). Repeat for each line; plot \(V_{0}\) against the frequency \(\nu\) of the line; and take the slope of the resulting line, which is the measurement. Repeat the whole procedure for a second and a third metal. Separately, determine the contact potential difference between the emitting surface and the collector, which displaces the whole line vertically and must be known before the intercept can be read as a work function — though, as the derivation below shows, not before the slope can be read as \(h/e\).

Observations and data

For every metal the stopping potential is a strictly linear function of frequency over the whole range examined, with no detectable curvature, and the lines obtained for different metals are parallel: the slope is a constant of nature and the material enters only through the intercept. Reading Planck's constant from that slope, Millikan obtained

\begin{equation}\tag{74.3} h=6.57\times 10^{-34}\,\mathrm{J}\,\mathrm{s}\ec \end{equation}

to about \(0.5\,\mathrm{\%}\) [Millikan:1916].

[Reserved: the tabulation. For each metal, the wavelength and frequency of each mercury line used, the measured stopping potential in volts with its uncertainty, and the residual from the fitted straight line — the residuals being the actual evidence for linearity and the thing a summary slope conceals. Then the fitted slope in \(\mathrm{V}\,\mathrm{s}\) with its standard error, the fitted intercept, the separately measured contact potential, and the work function that follows. All in SI.]

Phenomenon 74.4 (The photoelectric slope determines Planck's constant).

The straight line of Equation (69.1) has a slope \(\dd V_{0}/\dd\nu=h/e\) that is the same for every emitting metal and is independent of the work function, of the contact potential of the apparatus, and of the intensity of the light. Millikan established the linearity and the parallelism across the visible and near ultraviolet on alkali surfaces cut fresh in vacuum, and read from the common slope the value Equation (74.3), to about \(0.5\,\mathrm{\%}\) [Millikan:1916] — at the time the best determination of \(h\) by any method, and in agreement with the entirely independent value obtained from the spectrum of cavity radiation [Planck:1901].

Derivation. The claim to be established is not the linearity, which Phenomenon 69.2 already derives, but the robustness of the slope: why an experiment plagued by surface chemistry and by an unknown contact potential nevertheless yields a clean value of \(h/e\).

Let the illuminated surface have work function \(\phi_{\text{E}}\) and the collector \(\phi_{\text{C}}\). The two are in electrical contact through the external circuit, so their Fermi levels coincide and their vacuum levels differ by a fixed contact potential difference \(V_{\text{c}}=\left(\phi_{\text{C}}-\phi_{\text{E}}\right)/e\), whatever the electrometer reads. An electron leaving the surface with kinetic energy \(E\) and crossing to the collector against an applied potential difference \(V\) therefore arrives only if \(E>eV+\phi_{\text{C}}-\phi_{\text{E}}\). The fastest electron just fails when

\begin{equation}\tag{74.4} eV_{0}=E_{\max}-\left(\phi_{\text{C}}-\phi_{\text{E}}\right) =h\nu-\phi_{\text{E}}-\phi_{\text{C}}+\phi_{\text{E}} =h\nu-\phi_{\text{C}}\ep \end{equation}

Two conclusions follow, and they pull in opposite directions.

The intercept is worthless as a measurement of the illuminated surface. By Equation (74.4) it reports the work function of the collector, and any oxide or adsorbed film on that electrode enters the answer; every early attempt to read \(\phi\) off a photoelectric intercept was measuring the wrong electrode.

The slope is untouched. Neither \(\phi_{\text{E}}\) nor \(\phi_{\text{C}}\) depends on the frequency of the light, so differentiating Equation (74.4),

\begin{equation}\tag{74.5} \frac{\dd V_{0}}{\dd\nu}=\frac{h}{e}\ec \end{equation}

in which no property of either metal survives. A contamination that changes the work function displaces the whole line vertically and does not tilt it; that is why the lines for sodium, potassium and lithium come out parallel, and why parallelism is a stronger check on the hypothesis than the value of any one intercept.

What Equation (74.5) measures is the ratio \(h/e\) and not \(h\). In the SI in force since 2019 both constants are fixed by definition [BIPM:2019] [Mohr:2025], so the slope of this line is now an exact number,

\begin{equation}\tag{74.6} \frac{h}{e}=4.135667\times 10^{-15}\,\mathrm{V}\,\mathrm{s}\ec \end{equation}

and Millikan's experiment has become a test of a definition rather than a determination of a constant. Comparing Equation (74.3) with the defined \(h=6.62607015\times 10^{-34}\,\mathrm{J}\,\mathrm{s}\), his value is low by about \(0.8\,\mathrm{\%}\), which is larger than the \(0.5\,\mathrm{\%}\) he quotes. The discrepancy is not in the slope. To convert \(h/e\) into \(h\) one must multiply by a separately measured elementary charge, and the charge Millikan used was his own oil-drop value [Millikan:1913], later found low because of the value then adopted for the viscosity of air. The lesson is generic and belongs to Measurement, SI Units, and the Theory of Errors: an experiment's own precision bounds only the quantity it actually measures, and every constant imported to convert that quantity into another carries its error into the result undiminished.

Derivation pending.

The operational definition of the stopping potential. The measured current does not fall to zero at a sharp potential but approaches it along a tail, because the emitted electrons have a distribution of energies rather than a single maximum, because electrons originate at a range of depths below the surface, and because the collector is not equipotential. What is pending is the derivation of the shape of that current–voltage curve near its foot from the energy distribution of the photoelectrons, and the demonstration that the extrapolation Millikan used recovers the true intercept without bias — the point being that the linearity of the plotted line is a statement about a quantity defined by an extrapolation, and that extrapolation must be shown not to have manufactured it.

Interpretation

The equation is confirmed and its author's hypothesis is not, and Millikan said so. Having verified the linear relation to better than a percent, he continued to regard the light-quantum picture behind it as untenable, and he was not merely being stubborn: the photoelectric equation follows equally from a quantized atom absorbing energy from a classical field, as [Lamb:1969] later showed explicitly and as Section 69.2.4 records. The discreteness that Phenomenon 74.4 demonstrates is a discreteness of energy exchange; the atomic energy levels of Atomic Models and Spectra supply it without any assumption about the field. The measurement's real and undisputed content is metrological: the constant that Planck extracted from a cavity spectrum [Planck:1901] [Planck:1900a], by an argument about oscillators in thermal equilibrium, reappears to half a percent in a table-top electrometer experiment on a shaved sodium surface. Two measurements with nothing in common but the constant is the strongest kind of evidence that the constant is real, and it is why Part VIII — The Transition to Quantum Physics treats \(h\) as a fact about nature before any interpretation of it is settled.

Primary references

[Millikan:1916]. The elementary charge used to convert the measured slope into \(h\) is [Millikan:1913]; the independent value of \(h\) from cavity radiation is [Planck:1901]; the semiclassical account of the same equation is [Lamb:1969].

Compton: the wavelength of scattered X-rays changes (1923)

Here the evidence changes character. Everything before this section concerns the energy the light delivers; this concerns the momentum it carries away, and the momentum of a wave packet is a property of the light itself that no rearrangement of the absorber can supply [Compton:1923].

Apparatus

An X-ray tube with a molybdenum target, whose K\(\alpha\) line at about \(17.5\,\mathrm{keV}\) and \(71\,\mathrm{pm}\) provides a nearly monochromatic incident beam. A block of graphite serves as scatterer: a light element, so that most of its electrons are loosely enough bound to recoil freely. Slits define the incident beam and select the direction in which the scattered radiation is examined, fixing the scattering angle \(\theta\). The scattered radiation enters a Bragg spectrometer — a crystal of known lattice spacing on a divided circle, with an ionization chamber to measure the reflected intensity — so that a wavelength is read as an angle through Bragg's law [Bragg:1913b].

Procedure

Set the scattering angle; scan the spectrometer crystal through its range and record the ionization current against crystal angle, converting angle to wavelength by \(n\lambda=2d\sin\vartheta\); identify the peaks in the resulting spectrum. Repeat at a series of scattering angles from small to large, and repeat with scatterers of different atomic number. The comparison to be made is between the wavelength of each peak and the wavelength of the primary beam measured in the same spectrometer, which cancels the calibration of the crystal spacing from the difference.

Observations and data

At every scattering angle other than zero the scattered spectrum contains two peaks, not one: an unmodified peak at the incident wavelength and a second peak displaced towards longer wavelength. The displacement grows with the scattering angle. It does not depend on the incident wavelength, on the material of the scatterer, or on the intensity of the beam [Compton:1923].

[Reserved: the tabulation. Scattering angle in degrees; measured wavelength of the modified peak in picometres with its uncertainty; measured wavelength of the unmodified peak, as a control that the spectrometer calibration has not drifted; the difference; and the value predicted by Equation (69.4), together with the ratio of measured to predicted shift and its propagated uncertainty. Also the relative height of the two peaks against the atomic number of the scatterer, which is the evidence for the binding interpretation of the unmodified line.]

Phenomenon 74.5 (The shift depends on the scattering angle and on nothing else).

The wavelength displacement of the modified peak is a universal function of the scattering angle alone, given by Equation (69.4) with the Compton wavelength \(\lambda_{\text{C}}=h/m_{\text{e}}c=2.42631\,\mathrm{pm}\) [Mohr:2025], and carries no adjustable parameter and no property of the scattering material [Compton:1923] [Debye:1923]. Classical scattering of a wave by a charge predicts no displacement at any angle, because a charge driven at one frequency re-radiates at that frequency [Jackson:1999]; the observation is therefore not a quantitative refinement of the classical result but a contradiction of it.

Derivation. The relation itself is derived in The Photon: Photoelectric and Compton Effects from the relativistic conservation laws of Relativistic Dynamics applied to a collision between a quantum of energy \(h\nu\) and momentum \(h\nu/c\) and a free electron at rest. What is derived here is the experimental design that follows from it, because the reason this measurement was not made in 1905 is contained in the formula.

Evaluate Equation (69.4) at representative angles. The shift is \(\lambda_{\text{C}}\left(1-\cos\theta\right)\), so it vanishes in the forward direction, equals \(\lambda_{\text{C}}\) at \(90\,^\circ\), and saturates at \(2\lambda_{\text{C}}\) in backscattering; the values are collected in Table 74.1. The absolute shift is a few picometres and no more, whatever one does.

The quantity an instrument must resolve is not the shift but the fractional shift, and that is where the choice of radiation enters:

\begin{equation}\tag{74.7} \frac{\Delta\lambda}{\lambda} =\frac{\lambda_{\text{C}}}{\lambda}\left(1-\cos\theta\right) =\frac{h\nu}{m_{\text{e}}c^{2}}\left(1-\cos\theta\right)\ec \end{equation}

the second form making the physical content plain: the shift is appreciable only when the quantum energy is not negligible against the rest energy of the electron, \(511\,\mathrm{keV}\). For the molybdenum K\(\alpha\) line at \(71\,\mathrm{pm}\), Equation (74.7) at \(90\,^\circ\) gives \(\Delta\lambda/\lambda\approx0.034\), a three-percent change in wavelength — large for a crystal spectrometer, which resolves far better than that. For visible light at \(500\,\mathrm{nm}\) the same formula gives \(\Delta\lambda/\lambda\approx5\times 10^{-6}\), below anything a nineteenth-century spectroscopist could have separated from the width of a line. The Compton effect was not overlooked before 1923; it was inaccessible, and it became accessible when X-ray spectroscopy did [Bragg:1913b].

The same formula explains the peak that does not move. An electron bound tightly enough that the whole atom recoils requires \(m_{\text{e}}\) in Equation (69.4) to be replaced by the atomic mass, which for carbon is some four orders of magnitude larger; the predicted shift falls by the same factor and becomes unobservable. Hence a light element scatters mostly into the modified peak and a heavy one mostly into the unmodified peak, which is a further prediction of the same kinematics and is what the scatterer-dependence measurement tests.

Wavelength shift predicted by Equation (69.4) at representative scattering angles, with the fractional shift for the molybdenum K\(\alpha\) line at \(71\,\mathrm{pm}\) from Equation (74.7). Computed from the formula; these are not measurements.

$\theta$$\Delta\lambda$$\Delta\lambda/\lambda$ at \(71\,\mathrm{pm}\)
\(45\,^\circ\)\(0.711\,\mathrm{pm}\)\(0.010\)
\(90\,^\circ\)\(2.426\,\mathrm{pm}\)\(0.034\)
\(135\,^\circ\)\(4.142\,\mathrm{pm}\)\(0.058\)
\(180\,^\circ\)\(4.853\,\mathrm{pm}\)\(0.068\)

Interpretation

This is the first result in this chapter that a classical field cannot be made to yield. The photoelectric effect concerns what happens inside the absorber, and a quantized absorber suffices; the Compton shift concerns what happens to the radiation, and the radiation leaves with less energy and a changed direction in exactly the proportion that a particle of momentum \(h/\lambda\) would carry away. Two features make the evidence hard to evade. The shift is independent of the scattering material, so it cannot be a property of carbon; and it is independent of the incident intensity, so it cannot be a nonlinear response of the medium.

Three honest qualifications belong here. The kinematics were published independently and within weeks by Debye [Debye:1923], from the same hypothesis and without knowledge of the experiment, and the treatise's citation rule requires that both be named. The kinematics fix the shift but not the rate: the angular distribution requires the relativistic theory of the electron and emerges as the Klein–Nishina cross-section [Klein:1929a], developed from The Dirac Equation, whose low-energy limit is the Thomson result of Radiation and Scattering of Electromagnetic Waves. And — the qualification that mattered most at the time — the experiment measures an average over very many scattering acts. A theory in which energy and momentum are conserved only statistically predicts the same mean shift, and one such theory was on the table within a year [Bohr:1924]. Closing that gap is the business of Section 74.5.

Primary references

[Compton:1923]; the independent derivation of the same kinematics is [Debye:1923]; the crystal-spectrometer technique the measurement rests on is [Bragg:1913b].

Bothe and Geiger: the two products appear together (1925)

Bohr, Kramers and Slater proposed keeping the continuous classical field at the price of strict conservation: energy and momentum would hold only as statistical averages, the atom being coupled to a virtual radiation field that fixes transition probabilities [Bohr:1924]. It is the most serious counter-theory the light quantum ever faced, and its virtue is that it is sharply falsifiable, since it predicts that a scattered quantum and its recoil electron are uncorrelated event by event. This experiment measures that correlation [Bothe:1925].

Apparatus

Two needle counters — point-discharge counters of the kind Geiger had developed — mounted on opposite sides of a chamber filled with hydrogen, which is used as the scatterer because its single loosely bound electron per atom maximizes the modified scattering of Section 74.4 and minimizes absorption. One counter has a thin window and responds to the recoil electrons themselves; the other is arranged to respond to the scattered X-ray quanta, which it registers through the secondary electrons they liberate in a foil. A collimated X-ray beam crosses the chamber between them. Each counter's discharge drives a string electrometer whose fibre is imaged onto a photographic film driven at constant speed, so that the two records are written side by side against a common time axis.

Procedure

Irradiate the hydrogen and let the film run. Develop it, and measure for every discharge in one trace the time interval to the nearest discharge in the other. Coincidences are then defined by a resolving time \(\tau\) — set by how finely two deflections can be separated on the moving film, and here of order \(10^{-4}\,\mathrm{s}\) — and counted. The counting rate of each channel separately is measured on the same records. The number of chance coincidences expected from those singles rates and that resolving time is then computed and subtracted, and the excess is the quantity of interest. The whole experiment consists in showing that the excess is not zero, so the accidental rate must be computed rather than assumed small.

Observations and data

Coincidences between the two counters occur at a rate far above the rate expected from chance alone, with a resolving time of order \(10^{-4}\,\mathrm{s}\) [Bothe:1925]. Every recorded coincidence is therefore a single scattering act whose two products were both detected, and the scattered quantum and the recoil electron are produced together.

[Reserved: the tabulation. Singles counting rate in each channel in \(/\mathrm{s}\); total observation time; number of observed coincidences; the computed accidental number from Equation (74.8) with its Poisson uncertainty; the excess and its significance. Also the control runs that establish that the excess vanishes when it should — with the X-ray beam off, and with an absorber interposed so that only one channel can respond.]

Phenomenon 74.6 (Energy and momentum balance in each individual scattering act).

The scattered quantum and the recoil electron of a Compton scattering appear simultaneously, in the same act, and not merely at rates that agree on the average. Bothe and Geiger recorded coincidences between a counter responding to scattered X-rays and a counter responding to recoil electrons at a rate far exceeding the chance rate computable from the two singles rates and the resolving time [Bothe:1925]. The complementary measurement, in which a cloud chamber records the direction of the recoil electron and the direction of the scattered quantum in the same event and finds them correlated as the kinematics require, is [Compton:1925]. This is Phenomenon 69.6, and it excludes the statistical-only conservation of [Bohr:1924].

Derivation. The two competing predictions are predictions about a number of counts, so the derivation is the calculation of that number under each hypothesis.

Let the two counters fire at rates \(N_{1}\) and \(N_{2}\), and let the apparatus register a coincidence whenever two discharges fall within a resolving time \(\tau\) of each other. Suppose first that the two channels are statistically independent. In an observation of duration \(T\), counter 1 produces \(N_{1}T\) discharges; each of them opens a window of total length \(2\tau\), from \(\tau\) before to \(\tau\) after, within which a discharge of counter 2 will be recorded as coincident. Provided the windows do not overlap, that is \(2\tau N_{2}\ll1\), the expected number of chance coincidences is

\begin{equation}\tag{74.8} N_{\text{acc}}=2\tau N_{1}N_{2}T\ec \end{equation}

with Poisson fluctuation \(\sqrt{N_{\text{acc}}}\). This is the entire prediction of the statistical proposal [Bohr:1924]: on that theory the recoil electrons and the scattered quanta are emitted in unrelated acts, so there is no source of coincidences other than chance and the observed number must agree with Equation (74.8) within its fluctuation.

Under strict conservation each scattering act produces exactly one quantum and exactly one electron. If such acts occur at a rate \(R\) within the region both counters view, and the counters detect their respective products with efficiencies \(\varepsilon_{1}\) and \(\varepsilon_{2}\), the genuine coincidence number is

\begin{equation}\tag{74.9} N_{\text{true}}=\varepsilon_{1}\varepsilon_{2}RT\ec \end{equation}

and the prediction is \(N_{\text{obs}}=N_{\text{true}}+N_{\text{acc}}\), which exceeds Equation (74.8) by an amount that does not fall to zero however the apparatus is arranged.

The design rule follows by dividing the two. The singles rates are themselves proportional to the scattering rate, \(N_{i}=\varepsilon_{i}R\), so

\begin{equation}\tag{74.10} \frac{N_{\text{true}}}{N_{\text{acc}}} =\frac{\varepsilon_{1}\varepsilon_{2}R} {2\tau\varepsilon_{1}\varepsilon_{2}R^{2}} =\frac{1}{2\tau R}\ep \end{equation}

The genuine rate is linear in the beam intensity and the accidental rate is quadratic, so the discrimination improves as the source is weakened and as the resolving time is shortened — the opposite of the usual instinct to increase the signal. This is why the experiment is possible at all with counters of poor efficiency and a resolving time as long as \(10^{-4}\,\mathrm{s}\), and Equation (74.10) is the reason the coincidence method went on to become the basic technique of particle detection in Part XI — Quantum Field Theory and the Standard Model.

Derivation pending.

The angular correlation measured by the cloud-chamber experiment. What is pending is the direction of the recoil electron as a function of the scattering angle of the quantum, obtained from the same conservation laws that give the wavelength shift, together with the angular spread introduced by the initial motion of the struck electron and by multiple scattering of the recoil track in the chamber gas — the latter being what sets the accuracy with which the predicted correlation can be confirmed, and hence what the measurement is really comparing.

Interpretation

The proposal of Section 69.5.1 was withdrawn at once, and the manner of the withdrawal deserves recording: Bohr accepted the result immediately and without attempting to save the theory, which is the behaviour a falsifiable proposal earns. With it went the last scientifically serious attempt to keep a continuous classical electromagnetic field while accounting for the quantum phenomena, and conservation of energy and momentum was re-established as exact rather than statistical — a conclusion that reaches far beyond optics.

What the experiment establishes should be stated precisely, because it is often overstated. It establishes that the two products of a Compton scattering are correlated in time, event by event, and hence that the process is a collision between two objects and not a redistribution among many. It does not by itself establish that the incident radiation consists of localized quanta before the collision: a field that interacts with matter only in whole quanta, but propagates as a wave, survives this test intact. That last distinction — indivisibility of the propagating field itself — is what Section 74.6 measures.

Methodologically this experiment founded the coincidence technique, for which Bothe was later awarded the Nobel Prize, and Equation (74.10) is the design principle every coincidence measurement in this treatise still uses, from the cosmic-ray telescopes of Section 41.2 to the detectors of Part XI — Quantum Field Theory and the Standard Model.

Primary references

[Bothe:1925]; the complementary cloud-chamber measurement of the angular correlation is [Compton:1925]; the theory both of them exclude is [Bohr:1924].

Grangier, Roger and Aspect: a single quantum is not split (1986)

Everything so far is consistent with a classical field that happens to exchange energy in lumps. This experiment is not, and it is included because it is the one measurement in this chapter that no classical field of any strength or statistics can reproduce [Grangier:1986]. The quantity it measures is bounded below for every classical field, by an inequality that follows from nothing but the positivity of intensity, and the measured value lies far beneath the bound.

Apparatus

A beam of calcium atoms, excited by two lasers to a state that decays by a two-photon radiative cascade, so that photons are emitted in pairs: one near \(551\,\mathrm{nm}\) and one near \(423\,\mathrm{nm}\), well separated in colour and therefore separable by filters. One member of each pair is sent to a photomultiplier whose pulse opens an electronic gate of fixed duration — it is not measured further, it merely announces that a partner is on its way. The partner is directed onto a beamsplitter, and each of the two output ports carries its own photomultiplier. Coincidence electronics count, within each gate, the detections at the first output port, at the second, and at both. The same source can alternatively feed a Mach–Zehnder interferometer whose fringe visibility is recorded.

Procedure

Run the source at a rate low enough that the probability of two cascades falling in one gate is small, and accumulate three numbers over \(N\) gates: \(N_{1}\) and \(N_{2}\), the counts at the two output ports, and \(N_{\text{c}}\), the number of gates in which both fired. Form the anticorrelation parameter

\begin{equation}\tag{74.11} \alpha:=\frac{N_{\text{c}}N}{N_{1}N_{2}}\ec \end{equation}

which is the joint detection probability per gate divided by the product of the single detection probabilities. Then replace the beamsplitter arrangement by the interferometer, without changing the source, and record the fringe visibility. The two configurations put the same light to the two incompatible tests — indivisibility and interference — and the point of the experiment is that both are passed.

Observations and data

The measured anticorrelation parameter is

\begin{equation}\tag{74.12} \alpha=0.18\pm0.06\ec \end{equation}

far below the classical lower bound \(\alpha\geq1\) derived below, and the same source in the interferometer produces interference fringes of high visibility [Grangier:1986].

[Reserved: the tabulation. Gate duration in seconds; cascade rate in \(/\mathrm{s}\); the three counts \(N_{1}\), \(N_{2}\) and \(N_{\text{c}}\) with their accumulation time; the accidental contribution computed as in Equation (74.8) from the same gate width; the resulting \(\alpha\) with its statistical uncertainty; and the fringe visibility with its uncertainty. Also the measured dependence of \(\alpha\) on the cascade rate, which is the control that identifies the residual \(\alpha\) as double-cascade contamination rather than as a failure of the prediction.]

Phenomenon 74.7 (No classical field can give $\alpha<1$; the measured value is far below it).

For light described by any classical field whatever — of any intensity, however feeble, and with any statistics — the anticorrelation parameter Equation (74.11) satisfies \(\alpha\geq1\). Light prepared one quantum at a time violates this: the measured value is Equation (74.12), some thirteen standard deviations below the bound [Grangier:1986]. The same light, in an interferometer, nevertheless interferes with high visibility. This is Phenomenon 69.7.

Derivation. Take the classical bound first. Let the field arriving at the beamsplitter during a gate have intensity \(I\), treated as a random variable over an ensemble of gates with any distribution whatever, subject only to \(I\geq0\). Let the beamsplitter transmit a fraction \(T\) and reflect \(R=1-T\), and let each detector respond with a probability proportional to the intensity reaching it, with efficiency \(\eta\). Then the single detection probabilities per gate are

\begin{equation}\tag{74.13} P_{1}=\eta T\avg{I}\ec\qquad P_{2}=\eta R\avg{I}\ec \end{equation}

and, because the two detectors respond independently to the same field realization, the joint probability is

\begin{equation}\tag{74.14} P_{\text{c}}=\eta^{2}TR\avg{I^{2}}\ep \end{equation}

Note where the physics sits: the average of the product is the average of \(I^{2}\), not the product of averages, precisely because a classical wave is present at both outputs at once. Dividing,

\begin{equation}\tag{74.15} \alpha=\frac{P_{\text{c}}}{P_{1}P_{2}} =\frac{\avg{I^{2}}}{\avg{I}^{2}} =1+\frac{\operatorname{Var}\left(I\right)}{\avg{I}^{2}} \geq1\ec \end{equation}

the last step because a variance cannot be negative — equivalently, by the Cauchy–Schwarz inequality applied to \(I\) and the constant function. Equality holds only for a field of perfectly steady intensity; any fluctuation makes \(\alpha\) larger, never smaller. The efficiencies and the splitting ratio have cancelled, so the bound is not a statement about the apparatus. Neither is it a statement about brightness: making the beam weaker scales \(I\) and leaves the ratio in Equation (74.15) untouched. This is the crucial point, and it is what defeats the intuition that a sufficiently dim classical wave would behave like particles.

Now the quantum prediction. A one-photon state incident on port \(a\) of the beamsplitter goes over into a superposition of one photon in the transmitted mode and one photon in the reflected mode. That superposition has no component whatever in which both output modes are occupied, since there is only one excitation to distribute, so the joint detection probability vanishes identically and

\begin{equation}\tag{74.16} \alpha=0 \end{equation}

for an ideal single-photon input — again independently of the splitting ratio and of the detector efficiencies, which appear in \(P_{\text{c}}\) and in \(P_{1}P_{2}\) to the same order and cancel.

The measured Equation (74.12) is neither the classical bound nor the ideal quantum value, and the discrepancy from zero is understood and testable. The heralding is imperfect: if two cascades occur within one gate, the gate contains two photons, which can perfectly well be detected one at each output. The rate of such double-cascade gates grows with the square of the source rate while the single-cascade rate grows linearly, so \(\alpha\) should fall towards zero as the source is weakened — which is why the dependence of \(\alpha\) on the cascade rate is part of the data and not a diagnostic afterthought.

Finally, the quantity in Equation (74.15) is the normalized second-order correlation function at zero delay, and the inequality \(\alpha\geq1\) is the classical bound on it. The same inequality is violated by resonance fluorescence from a single atom, which was the first observation of the effect [Kimble:1977], and its systematic theory belongs to Quantum Optics and the Photon.

Interpretation

Two results are reported together and neither means much without the other. The anticorrelation says that the light is not divided by the beamsplitter: it goes one way or the other, and the classical field picture, in which an amplitude goes both ways and both detectors are illuminated, is excluded by Equation (74.15) at any intensity. The interference says that the very same light, given two paths that are later recombined, produces fringes — which the naive particle picture forbids, since a particle that took one arm cannot know the length of the other. The pair of observations, on one source, is the experimental content of what Interpretations (Evidence-Anchored) discusses, and it should be noted that the discussion rests on data of this kind, not on the photoelectric effect.

The historical placement matters as much as the result. Feeble-light interference had been demonstrated as early as 1909, with an exposure of months at an intensity so low that, on a quantum accounting, the apparatus contained one photon at a time [Taylor:1909] — and it established nothing about granularity, because a weak classical wave interferes exactly as well as a strong one. The complementary error is equally old: photon-counting statistics, and even the correlations between the photocurrents of two detectors viewing one source [HanburyBrown:1956], are reproduced by a classical field of fluctuating intensity, as Equation (74.15) shows for the special case \(\alpha\geq1\). The first experiments to draw the distinction sharply were the tests of the classical field-theoretic predictions for photodetection [Clauser:1974a] and the observation of antibunched fluorescence from a single atom [Kimble:1977]; what [Grangier:1986] adds is a heralded single-photon source, so that the two tests — indivisibility and interference — are applied to the same prepared state. The same hardware, and the same calcium cascade, is what makes the Bell tests of Experiment: Bell Tests possible.

A closing note on the word. “Photon” was coined by Lewis [Lewis:1926] for an entity quite unlike the one this section measures; what the experiment establishes is the indivisibility of a one-excitation state of the quantized field, and the honest definition of that object is the one given in The Photon: Photoelectric and Compton Effects and completed in Part XI — Quantum Field Theory and the Standard Model.

Primary references

[Grangier:1986]. Precursors that bear directly on the same distinction are [Clauser:1974a] [Kimble:1977]; the feeble-light interference that does not establish granularity is [Taylor:1909]; the intensity correlations that a classical field does reproduce are [HanburyBrown:1956].

What the measurements settle

ExperimentQuantity measuredWhat it establishes
Hertz–Hallwachs 1887–1888sign of the carriersultraviolet light liberates electrons from a clean metal; an anomaly, not yet a test of anything
Lenard 1902stopping potentialthe electron energy ignores the intensity and follows the colour; no continuous classical field can deliver energy this way
Millikan 1916slope $\dd V_{0}/\dd\nu$the slope is $h/e$ for every metal, giving $h=6.57\times 10^{-34}\,\mathrm{J}\,\mathrm{s}$; quantized exchange, not quantized light
Compton 1923$\Delta\lambda$ against $\theta$the radiation carries momentum $h/\lambda$; classical scattering predicts no shift at all
Bothe–Geiger 1925coincidence excessenergy and momentum balance in each single act; statistical-only conservation excluded
Grangier et al. 1986$\alpha$ at a beamsplitter$\alpha=0.18\pm0.06$ against the classical bound $\alpha\geq1$; the field itself is indivisible
The measurements reported in this chapter, and the claim each is entitled to. The third column is the point of the chapter: the experiments most often cited as proving that light is granular do not, and the ones that do are less famous.

The chapter has a structure that Table 74.2 makes visible and that the standard textbook order conceals. The first three experiments are decisive against one classical picture — a continuous field pouring energy into a passive absorber — and are silent about the alternative in which the field stays classical and the absorber is quantized. That alternative is not a philosophical reservation: it reproduces the threshold, the linearity in frequency, the independence of intensity and the promptness of the emission [Lamb:1969], which is to say every one of the four observations of Phenomenon 69.1. Anyone who says the photoelectric effect proves the existence of photons is asserting something the photoelectric effect does not contain, and Millikan, who measured it best, said so at the time.

What does the work is the second half. Section 74.4 shows that the radiation carries momentum in the fixed ratio \(1/c\) to its energy and delivers it in a collision, which is a property of the light and not of the target — the shift is the same for graphite, for hydrogen and for paraffin. Section 74.5 shows that the collision is a single event with two products and not an average over many, which removes the only counter-theory that had survived [Bohr:1924]. And Section 74.6 shows that the propagating field is itself indivisible, by violating a bound Equation (74.15) that follows from the positivity of a classical intensity and from nothing else. Only the third of these speaks about the field between the source and the detector; the first two speak about what it does when it arrives.

Two things this chapter does not settle should be named. It does not settle what a photon is — the answer is an excitation of the quantized electromagnetic field, with two helicity states and no rest frame, and it belongs to Part XI — Quantum Field Theory and the Standard Model. And it does not settle the rate of any of these processes, only their kinematics: the angular distribution of Compton scattering requires The Dirac Equation and appears as the Klein–Nishina cross-section [Klein:1929a], and the photoelectric cross-section requires the atomic structure of Atomic Models and Spectra. What is established here, beyond argument and by measurements a century apart, is that light delivers energy in units \(h\nu\), carries momentum \(h\nu/c\), conserves both in every individual act, and cannot be divided at a beamsplitter.