Experiment: The Classical Tests of General Relativity
- Le Verrier and the perihelion of Mercury (1859)
- Dyson, Eddington and Davidson: the eclipse of 1919
- Pound, Rebka and Snider: redshift in a tower (1960–1965)
- Shapiro: the radar time delay (1964–2003)
- The Eötvös lineage and MICROSCOPE (1922–2022)
- Lunar laser ranging and Gravity Probe B (1969–2011)
- Summary of the evidence
The Equivalence Principle and Classical Tests states the classical tests of general relativity as phenomena, and The Einstein Field Equations supplies the field equations from which they follow. Neither shows the measuring. That is the gap this chapter closes: the book has been asserting for five chapters that Mercury's perihelion advances by some 43 seconds of arc per century, that starlight grazing the Sun bends by 1.75 seconds of arc, that a clock at the top of a tower runs fast, and that a radar echo is delayed by passing the Sun — without once exhibiting the apparatus, the raw numbers or the error budget that make those statements evidence rather than doctrine.
Six experiments are reported. They are not six tests of one prediction: they probe logically separable pieces of the theory, and the separation is the point. The redshift measurement and the free-fall null tests constrain the equivalence principle alone, and would be passed by any metric theory of gravity whatever; the light deflection and the radar delay measure the curvature of space, which is where a metric theory first parts company with Newtonian gravity dressed in geometry; the perihelion advance is the first quantity that tests the field equations themselves; and the gyroscope precessions test the behaviour of a locally inertial frame carried through a curved and rotating spacetime. The experiments also span an unusually wide range of evidential quality, from a single photographic campaign of 1919 whose plate selection is still argued over, to a satellite null test at parts in \(10^{15}\). Both are reported here at their true weight.
Le Verrier and the perihelion of Mercury (1859)
Tests Phenomenon 53.1.
This is the oldest of the classical tests and the only one whose anomaly was established before the theory that explains it. Its apparatus is not an instrument but an archive: a century and a half of positional astronomy, reduced by the perturbation theory of Central Forces and Statics until nothing but a residue remained.
Apparatus
Meridian circles and transit instruments of the great European observatories, and above all the record of the transits of Mercury across the solar disc. A transit fixes the instant at which the planet crosses the limb of the Sun to within seconds of time, and therefore constrains the longitude of Mercury's node and perihelion far more sharply than a meridian observation of a small object seen in twilight close to the Sun. Le Verrier's material was the accumulated series of such transits from the end of the seventeenth century to his own day, together with the ordinary meridian observations of the planet [LeVerrier:1859]. The analytical half of the apparatus was the theory of secular perturbations: Lagrange's planetary equations integrated for the perturbations of Mercury by Venus, the Earth, Mars, Jupiter and Saturn.
Procedure
An osculating Keplerian ellipse is fitted to the observations; the computed gravitational perturbations of the other planets are applied; and the observed motion of the perihelion is then compared with the computed one. The comparison must be made in a well-defined reference direction, and the conventional one — the equinox — is itself moving, so the largest term in the budget is not gravitational at all but the general precession of the equinoxes. The procedure is thus a subtraction of three large and well-understood quantities from one large observed quantity, in the hope that the remainder is either zero or interesting. Le Verrier found it was not zero.
Observations and data
The perihelion budget, in the modern determination, stands as follows. All entries are rates in seconds of arc per Julian century, referred to the moving equinox.
The perihelion budget of Mercury, in seconds of arc per Julian century. The first three lines are observation and Newtonian celestial mechanics; the residual on the fourth line is what Le Verrier isolated in 1859, and the fifth is what general relativity supplies with no adjustable parameter.
| Contribution | Rate |
|---|---|
| Observed advance of the perihelion | 5600 |
| General precession of the equinoxes | $-5025$ |
| Perturbations by the other planets | $-531$ |
| Unexplained residual | 43 |
| General relativity, Equation (53.1) | 42.98 |
Le Verrier's own residual was smaller, near 38 seconds of arc per century; the value settled near 43 only after Newcomb's re-reduction of the observations and the improved planetary masses of the later nineteenth century [Newcomb:1895]. Nothing in this chapter turns on the difference: the two agree that a residual exists, of the order of tens of seconds of arc per century, and that it is many times the uncertainty of the observations.
After the precession of the equinoxes and every Newtonian planetary perturbation have been computed and subtracted, the perihelion of Mercury still advances in the direction of the orbital motion at a rate of about 43 seconds of arc per century. The residual was isolated by Le Verrier in 1859 and no distribution of matter consistent with the rest of solar-system astronomy was ever found to account for it [LeVerrier:1859]. General relativity predicts, for a bound orbit of semi-major axis \(a\) and eccentricity \(e\) about a mass \(M\), an advance of
per revolution, with no adjustable parameter [Einstein:1915a]; for Mercury this is 42.98 seconds of arc per century. Rests on Postulate 43.1 and Proposition 45.8.
Derivation. Derives Phenomenon 53.1. Take from Schwarzschild Geometry and Black Holes the orbit equation of a timelike geodesic of the Schwarzschild field in the equatorial plane, written for the inverse radius \(u=1/r\) with \(h=r^{2}\,\dd\varphi/\dd\tau\) the conserved specific angular momentum:
The final term is the whole of the relativistic correction. Its ratio to the term on the left is \(3h^{2}u^{2}/c^{2}\), of order \(v^{2}/c^{2}\) and so of order \(10^{-7}\) for Mercury; it may therefore be treated as a perturbation.
Dropping it leaves the Newtonian conic of Central Forces and Statics,
a closed ellipse whose perihelion does not move. Substituting \(u_{0}\) into the perturbing term gives
The constant piece and the piece in \(\cos 2\varphi\) hidden in \(\cos^{2}\varphi\) drive the oscillator off resonance and produce only bounded corrections to the shape of the orbit. The piece in \(\cos\varphi\) drives it at its own frequency, and it is this term alone that accumulates. Solving
so that to first order
the last step using \(\cos\left(\left(1-\delta\right)\varphi\right) \simeq\cos\varphi+\delta\varphi\sin\varphi\), valid because \(\delta\varphi\ll1\) over a single revolution. The radius therefore returns to its minimum not after \(2\pi\) but after \(\varphi=2\pi/\left(1-\delta\right)\simeq2\pi\left(1+\delta\right)\), and the perihelion advances by
per revolution. Eliminating \(h\) with the Newtonian relation \(h^{2}=GMa\left(1-e^{2}\right)\) yields Equation (53.1).
Inserting \(GM_{\odot}=1.32712\times 10^{20}\,\mathrm{m}^{3}/\mathrm{s}^{2}\), \(a=5.7909\times 10^{10}\,\mathrm{m}\) and \(e=0.20563\) gives \(\Delta\varphi=5.02\times 10^{-7}\) in radian measure per revolution. Mercury's period is \(87.969\) days, so it completes \(415.2\) revolutions in a Julian century, and the accumulated advance is \(2.084\times 10^{-4}\) radian, that is 42.98 seconds of arc per century — the fifth line of Table 53.1.
∎Interpretation
Two things must be said plainly. The first is that this is a postdiction: the number was known for fifty-six years before the theory that produced it, and Einstein computed it in November 1915 knowing what he had to match [Einstein:1915a]. The second is that this costs the result almost nothing, for the reason set out in Epistemology and the Scientific Method: Equation (53.1) contains no free parameter that could have been tuned. The theory was not built to fit 43 seconds of arc; the field equations were fixed by general covariance and the Newtonian limit, and the orbit followed. A postdiction with no adjustable parameter carries the evidential weight of a prediction, and it is a fair test precisely because the alternatives on offer were all parametric — an intramercurial planet of adjustable mass, a solar oblateness of adjustable size, a force law \(r^{-\left(2+\epsilon\right)}\) with adjustable \(\epsilon\) — and each of them, tuned to fit Mercury, then contradicted the other planets or the observed shape of the Sun.
The same effect has since been measured far outside the solar system. The star S2 orbiting the compact object at the Galactic centre shows a Schwarzschild precession of its orbit, detected by near-infrared interferometry and consistent with Equation (53.1) at a gravitational potential some four orders of magnitude deeper than Mercury's [Abuter:2020]; that measurement and its instrument belong to Experiment: Black-Hole Observations.
Primary references
[LeVerrier:1859] for the anomaly, [Einstein:1915a] for the computation that accounts for it, and [Abuter:2020] for the Galactic-centre analogue. The modern budget of Table 53.1 is the standard one summarized by [Will:2014]; the numbers on its first three lines are quoted at second hand and rounded, and only the last line is computed here.
Dyson, Eddington and Davidson: the eclipse of 1919
Tests Phenomenon 53.2.
Light deflection is the first of the classical tests in which the prediction preceded the measurement, and the first that distinguishes general relativity from a theory in which gravity acts on light through the equivalence principle alone. The factor separating them is exactly two, which is what makes a photographic measurement of moderate precision decisive.
Apparatus
Two expeditions of the Greenwich Observatory, to Sobral in northern Brazil and to the island of Príncipe in the Gulf of Guinea [Dyson:1920]. Both used photographic refractors fed by coelostat mirrors, since the instruments had to be shipped and set up in the field and could not be pointed at the Sun directly:
-
at Sobral, an astrographic objective of \(33\,\mathrm{cm}\) aperture and \(3.4\,\mathrm{m}\) focal length, and beside it a smaller lens of \(10\,\mathrm{cm}\) aperture and \(5.8\,\mathrm{m}\) focal length;
-
at Príncipe, a second astrographic objective of the same \(33\,\mathrm{cm}\) pattern.
The original report quotes apertures in inches and focal lengths in feet; they are converted here per the unit axiom of Measurement, SI Units, and the Theory of Errors. Detection is photographic, on glass plates measured afterwards on a screw micrometer at Greenwich.
Procedure
During totality the star field surrounding the eclipsed Sun is photographed. The same field is photographed again from the same instrument at another epoch — months earlier or later, at night, when the Sun is elsewhere in the sky — to give an undeflected reference. The two plates are then compared star by star. What is sought is a systematic radial displacement of each star away from the centre of the Sun, falling off as the reciprocal of the angular distance from it.
The difficulty is that a plate scale which changed between the two epochs would mimic exactly that pattern for a field of limited extent. The reduction therefore solves simultaneously for the scale, the orientation, the plate centre and the deflection constant, and it is the availability of stars at a range of distances from the limb, whose predicted deflections differ, that separates the deflection from the scale. The eclipse of 29 May 1919 was chosen years in advance for two reasons: its totality was among the longest of the century, and the Sun stood in front of the Hyades, so that an unusually rich field of bright stars surrounded it.
Observations and data
The competing predictions for a ray grazing the solar limb are 1.75 seconds of arc from the full theory and 0.87 from the equivalence principle alone; zero is the third possibility, if light does not fall. The measured deflections, reduced to the value at the limb, were
Deflection at the solar limb from the two 1919 stations, in seconds of arc, with the uncertainties as quoted by the authors. The Sobral astrographic plates were set aside in the original reduction; see the text.
| Instrument | Deflection at the limb |
|---|---|
| Sobral, \(10\,\mathrm{cm}\) lens | $1.98\pm0.12$ |
| Príncipe, astrographic | $1.61\pm0.30$ |
| Sobral, astrographic (set aside) | $0.93$ |
| General relativity | $1.75$ |
| Equivalence principle alone | $0.87$ |
The Sobral astrographic plates were rejected in the original reduction on an instrumental ground stated at the time: the coelostat mirror serving that telescope had changed figure in the heat of the day, and the star images on the eclipse plates were visibly out of focus, so that the plate scale could not be transferred reliably from the comparison epoch [Dyson:1920].
Starlight passing close to the Sun is deflected toward it, so that stars seen near the eclipsed limb are displaced radially outward. The two accepted 1919 determinations bracket the general-relativistic value of 1.75 seconds of arc and both exclude the value 0.87 obtained from the equivalence principle alone [Dyson:1920]. For an impact parameter \(b\) with \(b\gg GM/c^{2}\) the theory gives
Very-long-baseline radio interferometry on quasars now confirms Equation (53.3) at the level of one part in \(10^{4}\) [Lebach:1995] [Fomalont:2009][Will:2014], and the same bending on galactic and cluster scales produces the multiple images of gravitational lenses, of which the twin quasar 0957+561 was the first identified [Walsh:1979]. Rests on Postulate 43.1 and Proposition 45.8.
Derivation. Derives Phenomenon 53.2. For a null geodesic the constant term on the right of Equation (53.2) is absent — it came from the rest mass — and the orbit equation of Schwarzschild Geometry and Black Holes reduces to
With the right-hand side dropped the solution is a straight line passing at perpendicular distance \(b\) from the centre,
which indeed gives \(r\sin\varphi=b\). Substituting \(u_{0}\) into the right-hand side,
whose particular solution is
Nothing here is resonant, and the correction stays bounded; what it does is move the asymptotes. The ray comes from infinity and returns to it where \(u=0\). Near \(\varphi=0\) the equation \(u=0\) reads
and near \(\varphi=\pi\), writing \(\varphi=\pi+\epsilon\) and using \(\sin\left(\pi+\epsilon\right)\simeq-\epsilon\), it reads \(\epsilon=+2GM/\left(c^{2}b\right)\). The total angle swept between the two asymptotes therefore exceeds \(\pi\) by
which is Equation (53.3). Half of this comes from the Newtonian-like attraction that the equivalence principle already supplies; the other half comes from the term in \(u^{2}\), that is from the curvature of space, and it is the half that the 1919 plates were looking for. Numerically, with \(GM_{\odot}=1.32712\times 10^{20}\,\mathrm{m}^{3}/\mathrm{s}^{2}\) and \(b=R_{\odot}=6.957\times 10^{8}\,\mathrm{m}\), one finds \(\Delta\vartheta=8.49\times 10^{-6}\) in radian measure, or 1.751 seconds of arc.
∎Interpretation
The logic of the result is not that either station measured 1.75 accurately — neither did, and the quoted uncertainties are large — but that the two accepted determinations lie on opposite sides of 1.75 and that both lie many quoted uncertainties above 0.87. The experiment was designed to separate two hypotheses differing by a factor of two, and it separated them.
The rejection of the Sobral astrographic plates has been argued over ever since, and honesty requires that the argument be reported rather than settled here. The case against the rejection is that a value of 0.93 is close to the half-deflection, that the rejection was decided by people who expected the full value, and that the surviving Príncipe result rests on two usable plates. The case for it is that the defocusing of those plates was an instrumental fact recorded at the time and independent of what they showed, and that later re-measurements of the same Sobral astrographic plates with modern machines returned values consistent with the full deflection [Harvey:1979] — which is the decisive point, because it means the rejection, whatever its statistical propriety, did not bias the conclusion. What settles the matter beyond argument is not the 1919 plates at all but the radio interferometry of the past four decades, which measures the same deflection with a precision four orders of magnitude better and finds Equation (53.3) [Lebach:1995] [Fomalont:2009][Will:2014].
Primary references
[Dyson:1920] for the expeditions and their reduction; [Einstein:1911] for the half-deflection and [Einstein:1915a] for the full value; [Walsh:1979] for the first gravitational lens; [Will:2014] for the modern radio limits, quoted at second hand.
Pound, Rebka and Snider: redshift in a tower (1960–1965)
Tests Phenomenon 53.3.
The gravitational redshift is the one classical test that can be done indoors. It is also the one that tests the least: it follows from the equivalence principle and energy conservation alone, as the derivation in The Equivalence Principle and Classical Tests shows, and no field equation enters. That is precisely why it is worth doing — it isolates the assumption on which everything geometric rests.
Apparatus
The tower of the Jefferson Physical Laboratory at Harvard, giving a height difference of \(22.5\,\mathrm{m}\) between source and absorber. The clock is a nucleus. A source of \(^{57}\)Co decays to \(^{57}\)Fe, which emits a gamma ray of \(14.4\,\mathrm{keV}\); the same transition in a foil of iron enriched in \(^{57}\)Fe absorbs it resonantly. Both source and absorber are bound in a crystal lattice, so that a large fraction of the emissions and absorptions are recoil-free in the sense of Mössbauer [Moessbauer:1958] and the line retains its natural width: the excited state of \(^{57}\)Fe has a mean life of about \(141\,\mathrm{ns}\), corresponding to a fractional linewidth of about \(3\times10^{-13}\). A proportional counter behind the absorber records the transmitted intensity. The gamma path runs inside a mylar tube filled with helium, to keep absorption in air from destroying the count rate over \(22.5\,\mathrm{m}\), and the source is mounted on a magnetic transducer — in effect a loudspeaker cone — which moves it sinusoidally so as to impose a known first-order Doppler shift [Pound:1960] [Pound:1965].
Procedure
The quantity to be detected is a fractional frequency shift of \(gH/c^{2}\), which for \(22.5\,\mathrm{m}\) is \(2.46\times 10^{-15}\): some two orders of magnitude smaller than the natural width of the line being used to measure it. It is detected as a small asymmetry in the transmitted intensity between the two halves of the transducer cycle, and calibrated by the transducer velocity itself: the velocity needed to cancel the gravitational shift is \(v=c\,\Delta\nu/\nu\simeq0.74\,\mu\mathrm{m}/\mathrm{s}\).
The essential trick is a reversal. The whole measurement is performed once with the source at the bottom and the absorber at the top, and once the other way round. The gravitational shift changes sign; almost every systematic error does not. Differencing the two configurations therefore doubles the signal and cancels, to first order, every constant difference between the two nuclear environments — lattice chemistry, hyperfine fields, sample impurities, counter drift. What survives the reversal is a genuine dependence on the direction of propagation in the gravitational field.
One systematic error is not cancelled by the reversal, because it depends on the temperatures of the two samples rather than on their positions: the second-order Doppler shift produced by the thermal vibration of the emitting and absorbing nuclei in their lattices. Its size is derived below. It is comparable to the entire effect for a temperature difference of one kelvin, and the control of source and absorber temperatures is in practice the limiting art of the experiment.
Observations and data
A \(14.4\,\mathrm{keV}\) gamma ray propagating upward through a height \(H=22.5\,\mathrm{m}\) in the Earth's field is received with its frequency lowered by the fraction \(gH/c^{2}=2.46\times 10^{-15}\), and a downward ray is received with its frequency raised by the same fraction. The 1960 measurement gave a ratio of observed to predicted shift of \(1.05\pm0.10\) [Pound:1960], and the 1965 repetition with longer integration and improved control of the sample temperatures gave \(0.9990\pm0.0076\) [Pound:1965]. Rests on Definition 42.2, Phenomenon 42.5 and Phenomenon 38.13.
The derivation of \(gH/c^{2}\) itself is given in The Equivalence Principle and Classical Tests and is not repeated here. What is derived instead is the systematic that limits the measurement, because it is what decides whether the number above can be believed.
Derivation. Derives Phenomenon 53.3. A nucleus bound in a lattice at temperature \(T\) is not at rest: it vibrates, and although its mean velocity vanishes its mean square velocity does not. A moving emitter is seen, to second order in \(v/c\), with its frequency lowered by the time-dilation factor of Lorentz Transformations,
The sign is independent of the direction of the motion, which is why this shift does not average away over a vibration cycle and why it does not reverse when the apparatus is turned upside down. In the classical regime, equipartition over the three vibrational degrees of freedom of a nucleus of mass \(M\) gives \(\tfrac{1}{2}M\avg{v^{2}}=\tfrac{3}{2}k_{\mathrm{B}}T\), hence
For \(^{57}\)Fe, \(M=9.46\times 10^{-26}\,\mathrm{kg}\), and the coefficient is \(-2.4\times 10^{-15}\,/\mathrm{K}\).
That number is the point of the derivation. It is essentially equal to the entire gravitational effect of Phenomenon 53.3, so a systematic temperature difference of one kelvin between source and absorber would mimic or cancel the shift completely. Holding the error from this source below the one per cent finally quoted therefore requires the two samples to agree in temperature to about \(10\,\mathrm{mK}\), since \(0.01\times2.46\times 10^{-15}\) divided by the coefficient above is \(10^{-2}\) kelvin. It is also why the effect survives the reversal test: heating one end of the tower does not care which way the gamma ray is travelling. The measurement is in this sense a thermometry experiment as much as a gravitational one.
∎Interpretation
The redshift exists, has the predicted sign, and has the predicted size to one per cent. Because the derivation of \(gH/c^{2}\) uses only the equivalence principle, kinematics and energy conservation, this result does not test the field equations at all; it tests the assumption that a local laboratory in free fall is indistinguishable from an inertial one, which is the assumption that licenses describing gravity as geometry in the first place. A theory that failed here would not need repairing but abandoning.
The measurement has been superseded many times over and always in the same direction. A hydrogen maser flown on a suborbital rocket verified the shift at the level of \(7\times10^{-5}\) [Vessot:1980]; optical clocks now resolve it over height differences of less than a metre [Chou:2010] and, more recently, across a millimetre-scale atomic sample [Bothwell:2022]; and the same shift has been detected in the spectrum of the star S2 as it passes its closest approach to the Galactic-centre black hole [Abuter:2018]. The satellite-navigation clock budget of The Equivalence Principle and Classical Tests is the same effect run continuously as an engineering necessity, and the flown caesium clocks of Experiment: Time Dilation and Relativistic Kinematics measure it summed with the kinematic dilation.
Primary references
[Pound:1960] [Pound:1965]; the recoil-free emission that makes the measurement possible is [Moessbauer:1958]. The later and more precise measurements quoted in the interpretation are [Vessot:1980] [Chou:2010] [Bothwell:2022] [Abuter:2018].
Shapiro: the radar time delay (1964–2003)
Tests Phenomenon 53.4.
The fourth classical test was not available to Einstein: it requires a transmitter powerful enough to get an echo back from another planet. Shapiro pointed out in 1964 that a radar pulse whose path passes close to the Sun takes measurably longer to return than the same geometric path taken far from it [Shapiro:1964], and the experiment became possible within four years.
Apparatus
For the first measurements, planetary radar: the Haystack antenna, a steerable dish of about \(37\,\mathrm{m}\) aperture working in the centimetre band, and the \(305\,\mathrm{m}\) fixed reflector at Arecibo at metre wavelength, transmitting coded pulse trains at Mercury and Venus and receiving the echo some minutes later [Shapiro:1968]. The timing reference is an atomic frequency standard; the range model is a numerical ephemeris of the solar system.
Later measurements replace the passive planet with a spacecraft transponder, which returns a far stronger and better-defined signal: the Viking landers on Mars, ranged to over several superior conjunctions [Reasenberg:1979], and finally the Cassini spacecraft on its cruise to Saturn, whose radio link was operated simultaneously at two uplink and three downlink frequencies [Bertotti:2003].
Procedure
Range to the target as a function of time through a superior conjunction, when the line of sight passes close to the Sun, and compare the observed round-trip delay with the delay computed from the ephemeris in flat spacetime. The excess grows logarithmically as the ray approaches the limb, peaking at conjunction, and this characteristic shape is what identifies it: it is not degenerate with an error in the planetary distance, which varies on the orbital timescale rather than with the impact parameter.
Two systematics dominate. The first is the solar corona, whose plasma delays a radio signal by an amount proportional to the inverse square of the frequency, and which is both large and variable precisely where the relativistic effect is largest. The first-generation experiments had to model it; the Cassini experiment instead cancelled it, by running the link at several widely separated frequencies and using the known dispersion of a plasma to solve the plasma delay away — which is why Cassini improved on Viking by two orders of magnitude. The second is the topography and rotation of a planetary target, which the spacecraft transponder experiments remove by construction.
Observations and data
A signal travelling between two points at radii \(r_{1}\) and \(r_{2}\) from the Sun, past a closest approach \(b\) with \(b\ll r_{1},r_{2}\), takes longer than the same path in flat spacetime by
for the round trip. For an Earth–Venus circuit grazing the solar limb this is of order \(200\,\mu\mathrm{s}\). The effect was proposed as a fourth test in 1964 [Shapiro:1964] and detected in radar echoes from Mercury and Venus within a few years, in agreement with the prediction at the level of tens of per cent [Shapiro:1968]; ranging to the Viking landers confirmed it to about one part in \(10^{3}\) [Reasenberg:1979]; and the Cassini radio link, which measures the post-Newtonian parameter \(\gamma\) entering Equation (53.5) through an overall factor \(\left(1+\gamma\right)/2\), gave \(\gamma-1=\left(2.1\pm2.3\right)\times10^{-5}\) [Bertotti:2003]. Rests on Postulate 43.1 and Proposition 45.14.
Derivation. Derives Phenomenon 53.4. Write the Schwarzschild metric of Schwarzschild Geometry and Black Holes in isotropic coordinates and keep only first order in \(GM/\left(c^{2}r\right)\):
Isotropic coordinates are the right choice here because their spatial part is conformally flat, so that a coordinate straight line is the unperturbed ray and the whole effect appears in one place. Setting \(\dd s^{2}=0\) for a light ray and writing \(\dd l\) for the Euclidean coordinate length element,
The coordinate speed of light is less than \(c\) in the presence of the mass, and equally so in every direction; the factor two is one part time dilation and one part spatial curvature, exactly as in Equation (53.3).
Integrate along the unperturbed straight line, measuring \(l\) from the point of closest approach, so that \(r=\sqrt{b^{2}+l^{2}}\). The excess over the flat-space time \(\dd l/c\) is, one way,
For \(b\ll l_{i}\) one has \(\operatorname{arcsinh}\left(l/b\right)\simeq\ln\left(2l/b\right)\) and \(l_{i}\simeq r_{i}\), so the bracket is \(\ln\left(4r_{1}r_{2}/b^{2}\right)\); doubling for the return trip gives Equation (53.5).
Numerically, \(4GM_{\odot}/c^{3}=1.97\times 10^{-5}\,\mathrm{s}\). For a circuit from the Earth at \(r_{1}=1.496\times 10^{11}\,\mathrm{m}\) to Venus at \(r_{2}=1.082\times 10^{11}\,\mathrm{m}\) grazing the limb at \(b=6.96\times 10^{8}\,\mathrm{m}\), the logarithm is \(11.8\) and \(\Delta t=2.3\times 10^{-4}\,\mathrm{s}\). Note how weakly the answer depends on the geometry: halving the impact parameter adds only \(2.7\times 10^{-5}\,\mathrm{s}\), which is why the measurement needs a well-sampled conjunction rather than a single well-placed shot.
∎Interpretation
The radar delay is the sharpest of the classical tests, and it is sharp for an unglamorous reason: it measures a time, and time is the quantity physics measures best. The Cassini bound \(\gamma-1=\left(2.1\pm2.3\right)\times10^{-5}\) [Bertotti:2003] is the tightest constraint in the solar system on any metric theory of gravity, and every proposed alternative to general relativity must survive it before anything else is discussed. It is worth noticing what the delay and the deflection have in common: both are governed by the same coefficient \(\gamma\), and both would be exactly half their observed size if space were flat and only time were curved. The two experiments are, at post-Newtonian order, measurements of the same number by entirely different means — one photographic and angular, one electronic and temporal — agreeing to the precision of the weaker.
The same delay appears wherever a signal passes a mass, and it has become a workaday astrophysical tool rather than a test: the Shapiro delay in the timing of a pulsar whose companion passes in front of it is now one of the standard ways of weighing a neutron star, as Compact Stars and Relativistic Astrophysics describes [Demorest:2010].
Primary references
[Shapiro:1964] for the proposal, [Shapiro:1968] for the first radar detection, [Reasenberg:1979] for the Viking ranging and [Bertotti:2003] for the Cassini determination of \(\gamma\).
The Eötvös lineage and MICROSCOPE (1922–2022)
Tests Phenomenon 53.5.
The universality of free fall is the foundation the whole geometric description rests on, and it is the only one of these tests that is a null experiment: what is measured is the absence of a differential acceleration, and the result is a bound. The bound has been driven down by six orders of magnitude in a century, by instruments that have almost nothing in common except the quantity they report.
Apparatus
Two families of instrument, both comparing the free fall of two bodies of different composition.
The torsion balance, in the lineage that descends from the Cavendish instrument of Experiment: The Cavendish Torsion Balance. A horizontal beam is suspended by a fine wire with test bodies of different materials at its ends; a composition-dependent horizontal force appears as a torque about the fibre and is read by an optical lever or autocollimator. In the Eötvös–Pekár–Fekete instrument, operated from 1906 to 1909 and published only in 1922, a beam some tens of centimetres long hung from a platinum-iridium wire inside a triple thermal enclosure, and platinum was compared against copper, water, asbestos, tallow and a magnesium-aluminium alloy among others [Eotvos:1922]. In the modern Eöt-Wash instrument the whole balance sits on a slowly and continuously rotating turntable, and the test bodies are beryllium and titanium [Schlamminger:2008].
The orbiting differential accelerometer. The MICROSCOPE satellite carried two differential electrostatic accelerometers in a sun-synchronous orbit at about \(710\,\mathrm{km}\). Each held two coaxial hollow cylindrical test masses whose centres were made to coincide; electrodes around each mass sensed its displacement and applied the electrostatic force needed to hold it centred, and that force is the measurement. One accelerometer held masses of two different alloys — a platinum-rhodium alloy against a titanium alloy — and the other held two masses of the same platinum-rhodium alloy, serving as a null reference that should show no signal whatever [Touboul:2017] [Touboul:2022].
Procedure
In both families the measured quantity is the Eötvös parameter
and in both the art is to modulate it. A null experiment cannot be done by measuring once, because everything that is not the signal is also a constant: the signal must be made to vary in a known way while the instrument does not.
In the terrestrial torsion balance the modulation is geometric. The attractor is the Earth, or the Sun, or the Galaxy; the source direction rotates with respect to the instrument as the Earth turns, so a violation appears as a torque varying with a diurnal or annual period. Eötvös modulated by hand, turning the beam through \(180^\circ\) in azimuth and differencing the two readings, which reverses the signal but not the instrumental torques. The Eöt-Wash balance modulates continuously, rotating at a period of minutes so that the signal appears as a torque oscillating at the turntable frequency, far above the drift of the fibre and far from every fixed offset.
In orbit the modulation is faster and cleaner. The satellite is spun about an axis perpendicular to the orbital plane, so that the Earth-pointing direction rotates in the instrument frame at the sum of the orbital and spin frequencies. A violation of universality would appear as a differential acceleration between the two test masses at that one known frequency and in that one known phase; everything else — thermal drift, radiation pressure, residual gas — has to conspire to appear at exactly the same frequency and phase to be mistaken for it. The satellite is drag-free, its thrusters cancelling the atmospheric drag so that the test masses are truly in free fall.
Observations and data
Every experiment in the lineage has returned a null result. The bounds on \(\eta\), quoted at the order of magnitude their authors claim:
Bounds on the Eötvös parameter Equation (53.6), by experiment. Every entry is consistent with zero; the column gives the order of magnitude of the limit, not the central value.
| Experiment | Bodies and attractor | Bound on $\abs{\eta}$ |
|---|---|---|
| Eötvös, Pekár and Fekete, published 1922 [Eotvos:1922] | platinum against several, Earth | $10^{-9}$ |
| Roll, Krotkov and Dicke, 1964 [Roll:1964] | aluminium against gold, Sun | $10^{-11}$ |
| Braginsky and Panov, 1972 [Braginsky:1972] | aluminium against platinum, Sun | $10^{-12}$ |
| Eöt-Wash rotating balance, 2008 [Schlamminger:2008] | beryllium against titanium, Earth | $10^{-13}$ |
| MICROSCOPE, first results, 2017 [Touboul:2017] | titanium against platinum alloy, Earth | $10^{-14}$ |
| MICROSCOPE, final results, 2022 [Touboul:2022] | titanium against platinum alloy, Earth | $10^{-15}$ |
No experiment has ever detected a dependence of gravitational acceleration on composition. The Eötvös parameter Equation (53.6) is consistent with zero in every measurement listed in Table 53.3, and the strongest bound now stands at the level of parts in \(10^{15}\), set by comparing titanium and platinum-alloy test masses in orbital free fall [Touboul:2022]. The same experiments constrain composition dependence with the Sun and with the Galactic centre as attractor, so the null result is not an accident of one source [Roll:1964] [Braginsky:1972]. Rests on Definition 42.2, Phenomenon 42.1 and Equation (53.8).
Derivation. Derives Phenomenon 53.5. The bound is established by the torque analysis derived below: Equation (53.8) converts any composition dependence of free fall into a torque about the fibre of magnitude \(\tfrac{1}{2}\,\eta_{AB}\,m\,L\,a_{\mathrm{h}}\), in which every factor multiplying \(\eta_{AB}\) is fixed by the geometry of the instrument and by Equation (53.7). A recorded torque consistent with zero therefore bounds \(\abs{\eta_{AB}}\) directly. The orbital experiments read the same null through the electrostatic suspension force of a differential accelerometer, the driving acceleration being the full orbital \(g\) rather than the horizontal centrifugal residue, and each entry of Table 53.3 is such a null readout divided by the known sensitivity factor of its instrument; no further theory enters.
∎The derivation that a null \(\eta\) is equivalent to the identity of inertial and gravitational mass is given in The Equivalence Principle and Classical Tests. What is derived here is the signal the torsion balance actually looks for, since the geometry is not obvious and it explains why the instrument is sensitive at all.
Derivation. Work in the frame rotating with the Earth, at geographic latitude \(\lambda\). A body at rest there is acted on by the true gravitational attraction \(\vect{g}\), directed very nearly at the centre, and by the centrifugal acceleration of magnitude \(\omega^{2}R\cos\lambda\) directed away from the rotation axis. The centrifugal term is proportional to the inertial mass and the attraction to the gravitational mass, which is what makes the arrangement a test. Resolving the centrifugal acceleration into vertical and horizontal parts, its horizontal component points toward the equator and has magnitude
With \(\omega=7.292\times 10^{-5}\,/\mathrm{s}\) and \(R=6.371\times 10^{6}\,\mathrm{m}\) this is \(1.69\times 10^{-2}\,\mathrm{m}/\mathrm{s}^{2}\) at latitude \(45^\circ\), that is \(1.7\times 10^{-3}\) of \(g\).
Write \(m_{\mathrm{G}}=m_{\mathrm{I}}\left(1+\varepsilon\right)\) for each body, and let the local vertical be defined by the plumb line of a body with \(\varepsilon=0\); that vertical lies along \(\vect{g}+\vect{a}_{\mathrm{cf}}\), so in the plumb frame the horizontal component of \(\vect{g}\) has magnitude \(a_{\mathrm{h}}\) and points away from the equator, exactly cancelling the centrifugal horizontal component. A body with \(\varepsilon\neq0\) has effective acceleration \(\left(1+\varepsilon\right)\vect{g}+\vect{a}_{\mathrm{cf}}\), which differs from the plumb-frame vertical by the horizontal residue \(\varepsilon\,a_{\mathrm{h}}\). Two bodies at the ends of a beam of length \(L\) suspended at its centre and oriented east–west therefore feel horizontal forces along the north–south direction differing by \(m\eta_{AB}a_{\mathrm{h}}\), and since their position vectors from the fibre are \(\pm\left(L/2\right)\) along east, the net torque about the fibre is
The instrument is therefore looking for a torque smaller than the natural scale \(mgL\) by the product of \(\eta\) with \(1.7\times 10^{-3}\) — and \(\eta\) is the quantity being bounded at \(10^{-9}\) or below. That factor \(1.7\times 10^{-3}\) is the whole sensitivity penalty of doing the experiment on a rotating planet rather than in orbit, and it is the reason a satellite gains several orders of magnitude the moment it works at all: in orbit the driving acceleration is the full local \(g\), about \(8\,\mathrm{m}/\mathrm{s}^{2}\) at MICROSCOPE's altitude, rather than \(1.7\times 10^{-2}\,\mathrm{m}/\mathrm{s}^{2}\).
Two further features of Equation (53.8) matter experimentally. The torque vanishes at the equator and at the poles and is maximal near \(45^\circ\) of latitude, which is a signature no instrumental effect shares; and it reverses when the beam is turned through \(180^\circ\) in azimuth, which is the differencing Eötvös performed.
∎Interpretation
The null result is the evidence that licenses geometry. If \(m_{\mathrm{G}}/m_{\mathrm{I}}\) is one universal number, then the trajectory of a freely falling body depends on nothing but where it started and how fast it was going — not on what it is made of — and a family of trajectories with that property can be the geodesics of a connection. If \(\eta\) were nonzero at any level, gravity would be a force with a composition-dependent coupling like every other force, and the geometric description of Geometric Formulation of Gravity would be an approximation valid to whatever precision \(\eta\) had last been bounded at. The bound is therefore not a curiosity but the load-bearing measurement of the entire part.
It should be said what these experiments do not do. They test the weak form of the equivalence principle, for laboratory bodies whose own gravitational binding energy is negligible. Whether a self-gravitating body falls the same way is a separate question, tested by lunar laser ranging in the next section. And a null result is a bound, never a confirmation: what Table 53.3 records is that a century of increasingly ingenious experiments has failed to find a violation, which is the strongest statement an experiment of this kind can make.
Primary references
[Eotvos:1922] [Roll:1964] [Braginsky:1972] and [Schlamminger:2008] [Touboul:2017] [Touboul:2022]. The bounds of Table 53.3 are quoted from those papers at the order of magnitude only; the central values and their separate statistical and systematic uncertainties are given in the sources.
Lunar laser ranging and Gravity Probe B (1969–2011)
Tests Phenomena 42.4 and 53.6.
The last two experiments test what happens to a direction carried through curved spacetime. A gyroscope in orbit, or the orbit of the Moon itself considered as a gyroscope, is parallel-transported along its worldline, and general relativity says that the transported axis does not return to where it started. There are two such effects: the geodetic precession, produced by the curvature due to the central mass, and the very much smaller frame-dragging precession, produced by the rotation of that mass [Lense:1918] [Thirring:1918]. Lunar laser ranging also provides the one test in this chapter of the equivalence principle for a self-gravitating body.
Apparatus
Lunar laser ranging uses retroreflector arrays left on the Moon — by the Apollo 11, 14 and 15 crews, and on the two Soviet Lunokhod rovers — consisting of fused-silica corner cubes, which return light along the direction it came from regardless of the array's orientation. The Earth end is a telescope firing short laser pulses and timing the return of single photons against an atomic clock; the round trip takes about \(2.5\,\mathrm{s}\), and only a few returned photons are recorded per pulse train [Williams:2004] [Murphy:2012].
Gravity Probe B carried four gyroscopes into a polar orbit at about \(642\,\mathrm{km}\) altitude. Each rotor was a sphere of fused quartz \(3.8\,\mathrm{cm}\) in diameter, coated with superconducting niobium, electrostatically suspended in vacuum inside a dewar of superfluid helium at about \(1.8\,\mathrm{K}\), and spun up to tens of revolutions per second. The rotors were figured to within a few tens of nanometres of a perfect sphere. The spin direction was read out magnetically: a spinning superconductor develops a London moment along its spin axis, and the field of that moment was sensed by a SQUID magnetometer. An on-board telescope held the satellite pointed at the guide star IM Pegasi, providing the fixed direction against which the gyroscope drift was measured, and the satellite was drag-free [Everitt:2011].
Procedure
For lunar laser ranging the observable is a time series of Earth–Moon distances accumulated over decades and fitted with a full solar-system and lunar-orbit model. Two relativistic signatures are extracted from it. A violation of the equivalence principle for self-gravitating bodies — the Nordtvedt effect [Nordtvedt:1968] — would make the Earth and the Moon fall toward the Sun at slightly different rates, polarizing the lunar orbit along the Earth–Sun line with a signature periodic at the synodic month. The geodetic precession appears as a slow rotation of the lunar orbit plane, and is separated from Newtonian precessions by its rate and direction.
For Gravity Probe B the observable is the angle between each gyroscope spin axis and the line to IM Pegasi, tracked for the duration of the science mission. The two relativistic drifts are orthogonal by construction, which is why a polar orbit was chosen: the geodetic drift lies in the orbital plane, north–south, and the frame-dragging drift is perpendicular to it, east–west. The guide star has its own proper motion, which was measured independently by very-long-baseline radio interferometry against distant quasars and subtracted.
Observations and data
A gyroscope in a circular orbit of radius \(r\) about a mass \(M\) with angular momentum \(J\) precesses at the geodetic rate
in the plane of the orbit, and, for a polar orbit, at the frame-dragging rate
perpendicular to it [Schiff:1960a]. Gravity Probe B measured both against the guide star IM Pegasi and found, in milliarcseconds per year, \(-6601.8\pm18.3\) for the geodetic drift against a predicted \(-6606.1\), and \(-37.2\pm7.2\) for the frame-dragging drift against a predicted \(-39.2\) [Everitt:2011]. Lunar laser ranging sees the geodetic precession of the lunar orbit at the level of a few parts in \(10^{3}\) and, independently, finds the Nordtvedt parameter consistent with zero at the level of parts in \(10^{4}\), so that self-gravitating bodies fall like test bodies [Williams:2004][Nordtvedt:1968]. Rests on Postulate 43.1 and Theorem 45.30.
Derivation. Derives Phenomenon 53.6. Both rates follow from parallel transport of the spin four-vector in the weak-field metric of a slowly rotating mass — the linearization of the Kerr geometry, which is all the Earth requires. Write, to first order in \(\Phi/c^{2}\) and in the rotation,
with \(\Phi=-GM/r\) and the gravitomagnetic potential
A torque-free gyroscope on a geodesic parallel-transports its spin, \(\dd S^{\mu}/\dd\tau +\Gamma^{\mu}{}_{\nu\rho}u^{\nu}S^{\rho}=0\), with \(S_{\mu}u^{\mu}=0\). To the order required, the Christoffel symbols of Equation (53.11) are \(\Gamma^{i}{}_{00}=\pp_{i}\Phi/c^{2}\), \(\Gamma^{i}{}_{0j} =\tfrac{1}{2}(\pp_{j}h_{i}-\pp_{i}h_{j})\) and \(\Gamma^{i}{}_{jk} =(-\delta_{ij}\pp_{k}\Phi-\delta_{ik}\pp_{j}\Phi +\delta_{jk}\pp_{i}\Phi)/c^{2}\), while the constraint gives \(S^{0}=\vect S\cdot\vect v/c\) at leading order. Assembling the spatial components with \(u^{\mu}\approx(c,\vect v)\),
The coordinate components mix a genuine rotation with the stretching of the coordinate grid; the physical spin is the one referred to the local orthonormal rest frame, \(\hat S^{i}=(1-\Phi/c^{2})S^{i} -\tfrac{1}{2c^{2}}v^{i}(\vect v\cdot\vect S)\), the first factor from the spatial metric in Equation (53.11), the second the low-velocity boost to the gyroscope frame. Differentiating, using \(\dd\vect v/\dd t=-\nabla\Phi\) on the geodesic, and keeping first order, every non-rotational term cancels and what survives is a pure precession,
the second term being \(-\tfrac{c}{2}\nabla\times\vect h\) evaluated from Equation (53.12). For a circular orbit, \(\abs{\vect r\times\vect v}=r^{2}\sqrt{GM/r^{3}}\), so the geodetic magnitude is \(\Omega_{\mathrm{geo}} =\tfrac{3}{2}(GM/c^{2}r)\sqrt{GM/r^{3}}\), which is Equation (53.9), directed along the orbit normal. For a polar orbit the frame-dragging term must be averaged over the orbit: with \(\hat{\vect r}\) sweeping a great circle through the poles, \(\avg{(\vect J\cdot\hat{\vect r})\hat{\vect r}}=\vect J/2\), so \(\avg{\vect\Omega_{\mathrm{fd}}} =(G/c^{2}r^{3})\bigl(\tfrac{3}{2}-1\bigr)\vect J = G\vect J/(2c^{2}r^{3})\), which is Equation (53.10). The same \(\vect\Omega\) applied to the Earth–Moon system as a gyroscope (its orbital angular momentum the spin) gives the geodetic precession of the lunar orbit that lunar laser ranging observes.
∎Granted Equations (53.9) and (53.10), the predicted rates follow from the orbit and from the Earth's mass and spin. Taking \(r=7.013\times 10^{6}\,\mathrm{m}\) for an altitude of \(642\,\mathrm{km}\), and \(GM_{\oplus}=3.986\times 10^{14}\,\mathrm{m}^{3}/\mathrm{s}^{2}\), the orbital angular velocity is \(1.075\times 10^{-3}\,/\mathrm{s}\) and \(GM_{\oplus}/\left(c^{2}r\right)=6.32\times 10^{-10}\), so Equation (53.9) gives \(1.02\times 10^{-12}\,/\mathrm{s}\), which is \(3.22\times 10^{-5}\) radian per year, or about 6640 milliarcseconds per year. With the Earth's spin angular momentum \(J_{\oplus}\simeq5.86\times 10^{33}\,\mathrm{kg}\,\mathrm{m}^{2}/\mathrm{s}\), Equation (53.10) gives \(6.3\times 10^{-15}\,/\mathrm{s}\), about 41 milliarcseconds per year. Both agree with the mission's own predictions to within a few per cent; the residual difference comes from the orbital eccentricity, the Earth's oblateness and the solar geodetic term, all carried by the mission model and none by the two-line estimate made here.
Interpretation
The geodetic result is a clean measurement: a fractional uncertainty of about \(0.3\,\mathrm{\%}\), agreeing with the prediction. The frame-dragging result must be reported more carefully, and this chapter would be dishonest if it were not.
Gravity Probe B was designed to reach an uncertainty below one milliarcsecond per year on both drifts, which would have made the frame-dragging measurement a one-per-cent test. It did not. The delivered uncertainties are \(18.3\) and \(7.2\) milliarcseconds per year, some twenty times the design goal, so the frame-dragging number is a detection at about five times its own uncertainty and a confirmation of the predicted rate at the level of roughly twenty per cent — not one per cent.
The cause is instructive and was not anticipated. The niobium coating on the rotors and the surrounding housing carried patches of varying electric potential, and the interaction of those patches produced classical electrostatic torques on the spinning spheres. Because the rotors' polhode motion — the wobble of the spin axis within the body — drifted slowly in period over the mission, these torques were not constant and could not be removed by a simple fit. The analysis had to build an explicit model of the patch-induced misalignment torque and of the polhode phase for each rotor, and subtract it; the final error budget is dominated not by measurement noise but by the uncertainty of that model [Everitt:2011]. This is worth stating in a treatise that grades its evidence: the geodetic precession is measured, and the frame-dragging precession is detected but only loosely measured, by this experiment.
What lunar laser ranging adds is different in kind. The Earth and the Moon store different fractions of their mass-energy as gravitational binding energy, so they are the two test bodies of an equivalence experiment that the balances of Section 53.5 cannot perform: laboratory masses have no appreciable self-gravity. The null Nordtvedt signal therefore extends Phenomenon 53.5 from the weak to the strong form of the principle, which no other experiment in this chapter touches [Williams:2004].
Primary references
[Everitt:2011] for Gravity Probe B, [Williams:2004] for lunar laser ranging, and [Nordtvedt:1968] for the effect the latter bounds.
Summary of the evidence
| Experiment | Precision reached | What it constrains |
|---|---|---|
| Mercury perihelion, 1859 | a residual of 43 out of 5600 seconds of arc per century | the field equations, at first post-Newtonian order |
| Eclipse of 1919 | enough to resolve a factor of two | the curvature of space; excludes the equivalence principle acting alone |
| Pound–Rebka–Snider | one per cent | the equivalence principle alone; no field equation enters |
| Shapiro delay to Cassini | $\gamma-1$ to $2\times10^{-5}$ | the same curvature coefficient as the deflection, by an independent method |
| Eötvös to MICROSCOPE | $\eta$ below $10^{-15}$ | universality of free fall for laboratory bodies; licenses the geometric description |
| Ranging and Gravity Probe B | \(0.3\,\mathrm{\%}\) geodetic, about \(20\,\mathrm{\%}\) frame-dragging | parallel transport of a direction; the gravitomagnetic sector; the strong form of the principle |
Three remarks close the chapter. The first is that the tests are not redundant. Every row of Table 53.4 could in principle have failed while the others passed, and the third column says how; a theory that reproduced the redshift and the deflection but not the perihelion advance would have been a live possibility in 1919 and is excluded now by radar and by Mercury together.
The second is that the classical tests are all weak-field. The dimensionless potential \(GM/\left(c^{2}r\right)\) reaches at most a few parts in \(10^{6}\), at the solar limb, in every experiment reported here — the Galactic-centre orbit mentioned in Section 53.1 is the exception, and it belongs to Experiment: Black-Hole Observations — and the deviations measured are first-order in it. General relativity is here being tested where it is nearly Newtonian, and passing such tests is necessary but very far from sufficient. The strong-field evidence is elsewhere: in the orbital decay and timing of binary pulsars, in the direct gravitational-wave detections of Experiment: Gravitational Waves, and in the horizon-scale observations of Experiment: Black-Hole Observations.
The third is about the shape of the evidence rather than its content. The best of these measurements are not the most famous ones. The 1919 eclipse made general relativity a public fact and is the weakest result in the chapter; the Cassini radio link is the strongest and is known to almost nobody. That asymmetry is a standing hazard for a treatise organized around evidence, and it is why Table 53.4 reports precision reached rather than historical importance.