Fourier Analysis and Integral Transforms

Contents
  1. Fourier series
  2. The Fourier transform
  3. Convolution and correlation
  4. The Laplace transform
  5. Sampling and the discrete transform
  6. Other integral transforms

Fourier series and transforms, Laplace transforms, convolution and Green's-function methods, with the convergence theorems stated and proved at the depth cap of Part II. The organising idea is a change of basis: a function of time or space is re-expressed as a superposition of harmonics, in which differentiation becomes multiplication and the linear differential equations of physics become algebraic. Fourier introduced the expansion to solve the heat equation [Fourier:1822], and the analysis, synthesis and convergence of such expansions — settled only in 1966 for square-integrable functions [Carleson:1966] — created much of modern analysis on the way.

The chapter feeds the separation-of-variables and Green's-function methods of Partial Differential Equations, the orthonormal-basis language of Hilbert Spaces, the normal-mode analysis of Oscillations and Mechanical Waves, and, through the uncertainty relation between a function and its transform, the wave mechanics of The Postulates of Quantum Mechanics and Matter Waves. Its computational face — sampling and the fast Fourier transform — is what turns the strain record of a gravitational-wave interferometer or the interferogram of a spectrometer into a spectrum, so the experimental chapters consume it too.

The Sturm–Liouville chapter already exhibits the Fourier series and the Fourier integral as the eigenfunction expansion of the one-dimensional Helmholtz operator (Section 9.7.4, Section 9.7.5) and states, without derivation, the orthonormality, the passage to the integral, the derivative rule, the convolution property and the Parseval relation. All five are proved below, in the same normalisation.

Fourier series

Trigonometric expansion of periodic functions

The subject began in a quarrel about a string. In 1747 d'Alembert solved the equation of a stretched string of length \(\ell\) fixed at both ends and found its general motion to be a superposition of two waves travelling in opposite directions at the speed \(c = \sqrt{\tau/\mu}\) — with the tension \(\tau\) in \(\mathrm{N}\) and the linear density \(\mu\) in \(\mathrm{kg}/\mathrm{m}\), so that \(c\) is in \(\mathrm{m}/\mathrm{s}\) [dAlembert:1747]. Six years later Daniel Bernoulli asserted that the same motion is a superposition of the normal modes,

\begin{equation}\tag{17.1} y(x,t) = \sum_{n=1}^{\infty} A_n \sin\frac{n\pi x}{\ell}\, \cos\left(\frac{n\pi c\,t}{\ell} + \varphi_n\right)\ec \end{equation}

and hence — setting \(t = 0\) — that any admissible initial shape of the string is a sine series [Bernoulli:1753]. Euler and d'Alembert objected that the right-hand side of Equation (17.1) is built from analytic functions and could not possibly represent the arbitrary, kinked initial profiles the physical problem allows. The objection was reasonable and wrong, and it was settled in Bernoulli's favour by Fourier, who used precisely such expansions to solve the heat equation and computed the coefficients by the formulas below [Fourier:1822]. What the objectors were right about is that the class of functions represented, and the sense in which the series represents them, both need proof; supplying it occupied a century and is the content of Section 17.1.2.

Definition 17.1 (Fourier coefficients and Fourier series).

Let \(f\) be complex valued, periodic with period \(T > 0\), and absolutely integrable over one period. Write

\begin{equation}\tag{17.2} \omega_1 = \frac{2\pi}{T}\ec\qquad \nu_1 = \frac{1}{T} \end{equation}

for the fundamental angular frequency and the fundamental frequency; if \(t\) is a time in \(\mathrm{s}\) then \(\omega_1\) is in \(\mathrm{rad}/\mathrm{s}\) and \(\nu_1\) in \(\mathrm{Hz}\). The Fourier coefficients of \(f\) are

\begin{equation}\tag{17.3} c_n = \frac{1}{T}\int_{-T/2}^{T/2} f(t)\,\ee^{-\ii n\omega_1 t}\,\dd t\ec \qquad n \in \Z\ec \end{equation}

and the Fourier series of \(f\) is the formal sum

\begin{equation}\tag{17.4} \sum_{n=-\infty}^{\infty} c_n\,\ee^{\ii n\omega_1 t}\ec \end{equation}

whose symmetric partial sums are written \(S_N f(t) = \sum_{n=-N}^{N} c_n \ee^{\ii n\omega_1 t}\).

Equation (17.3) is the analysis step and is defined for every absolutely integrable \(f\), whatever the series does; Equation (17.4) is the synthesis step, and whether it returns \(f\) is the convergence problem. The two are kept separate throughout: it is the failure to separate them that made the eighteenth-century controversy irresolvable.

Lemma 17.2 (Orthogonality of the harmonics).

For \(m, n \in \Z\),

\begin{equation}\tag{17.5} \frac{1}{T}\int_{-T/2}^{T/2} \ee^{\ii m\omega_1 t}\,\overline{\ee^{\ii n\omega_1 t}}\,\dd t = \delta_{mn}\ep \end{equation}

Rests on Definition 17.1.

Proof.

Derives Lemma 17.2. The integrand is \(\ee^{\ii(m-n)\omega_1 t}\). For \(m = n\) it is \(1\) and the average is \(1\). For \(m \neq n\) put \(p = m - n \neq 0\); the antiderivative is \(\ee^{\ii p\omega_1 t}/(\ii p\omega_1)\), so the integral equals

\begin{equation*} \frac{\ee^{\ii p\pi} - \ee^{-\ii p\pi}}{\ii p \omega_1} = \frac{2\ii\sin(p\pi)}{\ii p\omega_1} = 0\ec \end{equation*}

using \(\omega_1 T/2 = \pi\) and \(\sin(p\pi) = 0\) for integer \(p\).

This one computation is the whole mechanism. It is the special case, for the operator \(-\dd^{2}/\dd t^{2}\) on a periodic interval, of the orthogonality of eigenfunctions of a Sturm–Liouville operator belonging to distinct eigenvalues (Section 9.4), and the model for the abstract orthonormal bases of Hilbert Spaces: it is what lets a coefficient be extracted from a sum by a single integral. Equation (17.5) is the relation left as a pending derivation at Equation (9.109).

Theorem 17.3 (Euler–Fourier coefficient formulas).

If a series \(\sum_{n} \gamma_n \ee^{\ii n\omega_1 t}\) converges uniformly on \([-T/2, T/2]\) to a function \(f\), then \(f\) is continuous and periodic, and \(\gamma_n = c_n\) for every \(n\), with \(c_n\) given by Equation (17.3). In particular a function has at most one uniformly convergent trigonometric expansion. Rests on Definition 17.1 and Lemma 17.2.

Proof.

Derives Theorem 17.3. A uniform limit of continuous functions is continuous, and each term is \(T\)-periodic, so \(f\) is. Fix \(m\), multiply the series by \(\ee^{-\ii m\omega_1 t}\) — a bounded factor, so convergence stays uniform — and integrate over one period. Uniform convergence permits the exchange of sum and integral (the tail is bounded by \(T\sup_t\abs{\text{tail}}\)), giving

\begin{equation*} \int_{-T/2}^{T/2} f(t)\,\ee^{-\ii m\omega_1 t}\,\dd t = \sum_{n} \gamma_n \int_{-T/2}^{T/2} \ee^{\ii (n-m)\omega_1 t}\,\dd t = T\gamma_m \end{equation*}

by Lemma 17.2. Divide by \(T\).

Real form.

Splitting the \(n\) and \(-n\) terms of Equation (17.4) with Euler's formula (Proposition 8.4) gives the real trigonometric form

\begin{equation}\tag{17.6} \frac{a_0}{2} + \sum_{n=1}^{\infty} \left[a_n\cos(n\omega_1 t) + b_n\sin(n\omega_1 t)\right]\ec \end{equation}

with

\begin{equation}\tag{17.7} a_n = \frac{2}{T}\int_{-T/2}^{T/2} f(t)\cos(n\omega_1 t)\,\dd t\ec \qquad b_n = \frac{2}{T}\int_{-T/2}^{T/2} f(t)\sin(n\omega_1 t)\,\dd t\ec \end{equation}

and the dictionary between the two sets of coefficients is

\begin{equation}\tag{17.8} a_n = c_n + c_{-n}\ec \qquad b_n = \ii\left(c_n - c_{-n}\right)\ec \qquad c_{\pm n} = \tfrac{1}{2}\left(a_n \mp \ii b_n\right)\ec \quad n \ge 1\ec \end{equation}

with \(a_0 = 2c_0\) twice the mean value of \(f\). For real \(f\) one has \(c_{-n} = \overline{c_n}\), so the real form carries the same information in half as many symbols; the complex form is used below because the harmonics are then eigenfunctions of \(\dd/\dd t\), which is the property the whole subject exploits. Equation (17.7) is Equation (9.118) written with \(T = 2L\), the factor \(2/T\) there appearing as \(1/L\).

The orthonormality relations of the real system follow from Equation (17.5) and the product formulas of trigonometry: for \(m, n \ge 1\),

\begin{equation}\tag{17.9} \frac{2}{T}\int_{-T/2}^{T/2} \cos(m\omega_1 t)\cos(n\omega_1 t)\,\dd t = \delta_{mn}\ec\qquad \frac{2}{T}\int_{-T/2}^{T/2} \sin(m\omega_1 t)\sin(n\omega_1 t)\,\dd t = \delta_{mn}\ec \end{equation}

and the mixed integral vanishes for all \(m,n\). Indeed \(\cos A\cos B = \tfrac12\left[\cos(A-B)+\cos(A+B)\right]\) turns the first into \(\delta_{m,n} + \delta_{m,-n}\) by Lemma 17.2, and the second index coincidence is impossible for positive \(m,n\); the sine relation is the same identity with the opposite sign, and the mixed one is the integral of an odd function over a symmetric interval. This is the derivation left pending at Equations (9.114) to (9.116).

Proposition 17.4 (Least squares and Bessel's inequality).

Let \(f\) be square integrable over a period. Among all trigonometric polynomials \(P(t) = \sum_{\abs{n}\le N}\gamma_n \ee^{\ii n\omega_1 t}\) of degree at most \(N\), the mean-square error \(\frac{1}{T}\int_{-T/2}^{T/2}\abs{f - P}^{2}\dd t\) is smallest exactly when \(\gamma_n = c_n\), that is, when \(P = S_N f\). Moreover

\begin{equation}\tag{17.10} \sum_{n=-N}^{N}\abs{c_n}^{2} \le \frac{1}{T}\int_{-T/2}^{T/2}\abs{f(t)}^{2}\,\dd t \qquad\text{for every } N\ec \end{equation}

so the coefficients are square summable and \(c_n \to 0\) as \(\abs{n}\to\infty\). Rests on Definition 17.1 and Lemma 17.2.

Proof.

Derives Proposition 17.4. Write \(\avg{g,h} = \frac{1}{T}\int_{-T/2}^{T/2} g\bar h\,\dd t\). By Lemma 17.2 the harmonics are orthonormal for \(\avg{\cdot,\cdot}\) and \(\avg{f, \ee^{\ii n\omega_1 t}} = c_n\). Expand:

\begin{equation*} \avg{f - P, f - P} = \avg{f,f} - \sum_{\abs{n}\le N}\left(\gamma_n \overline{c_n} + \overline{\gamma_n} c_n\right) + \sum_{\abs{n}\le N}\abs{\gamma_n}^{2} = \avg{f,f} - \sum_{\abs{n}\le N}\abs{c_n}^{2} + \sum_{\abs{n}\le N}\abs{\gamma_n - c_n}^{2}\ec \end{equation*}

by completing the square term by term. Only the last sum depends on the \(\gamma_n\), and it is minimal — and zero — iff \(\gamma_n = c_n\). Setting \(\gamma_n = c_n\) leaves \(0 \le \avg{f,f} - \sum_{\abs{n}\le N}\abs{c_n}^{2}\), which is Equation (17.10); the bound is uniform in \(N\), so the series of squares converges and its terms tend to zero.

The partial sum is thus the orthogonal projection of \(f\) onto the span of the first \(2N+1\) harmonics — the concrete instance of the projection theorem of Hilbert Spaces. Whether the inequality Equation (17.10) is an equality in the limit is exactly the question of whether the harmonics are a complete orthonormal system; Section 17.2.2 proves that they are.

Example 17.5 (Square wave and sawtooth).

Let \(f\) be the odd square wave of unit amplitude, \(f(t) = \sgn\left(\sin \omega_1 t\right)\). It is odd, so all \(a_n\) vanish, and

\begin{equation}\tag{17.11} b_n = \frac{4}{T}\int_{0}^{T/2}\sin(n\omega_1 t)\,\dd t = \frac{2}{n\pi}\left[1 - (-1)^{n}\right] = \begin{cases} \dfrac{4}{n\pi}\ec & n \text{ odd}\ec\\[4pt] 0\ec & n \text{ even}\ep \end{cases} \end{equation}

Only odd harmonics appear, with amplitudes falling as \(1/n\). For the \(50\,\mathrm{Hz}\) square wave delivered by a simple inverter this places power at \(50\,\mathrm{Hz}\), \(150\,\mathrm{Hz}\), \(250\,\mathrm{Hz}\) and so on, with relative amplitudes \(1\), \(1/3\), \(1/5\); the third harmonic carries \((1/3)^{2} = 11.1\,\mathrm{\%}\) of the fundamental's power. The sawtooth \(f(t) = 2t/T\) on \((-T/2, T/2)\) is also odd, with \(b_n = (-1)^{n+1}\,2/(n\pi)\), again a \(1/n\) envelope. In both cases the \(1/n\) decay is the signature of a jump discontinuity, as Proposition 17.10 will show, and it is what makes these two examples the arena of Section 17.1.3.

Convergence: from Dirichlet to Carleson

Rescaling \(x = \omega_1 t\) turns any period into \(2\pi\) and changes no convergence statement, so this subsection fixes \(T = 2\pi\) and writes \(c_n = \frac{1}{2\pi}\int_{-\pi}^{\pi} f(x)\ee^{-\ii n x}\dd x\).

Lemma 17.6 (Dirichlet kernel).

Define \(D_N(u) = \sum_{n=-N}^{N}\ee^{\ii n u}\). Then, for \(u \notin 2\pi\Z\),

\begin{equation}\tag{17.12} D_N(u) = \frac{\sin\left[\left(N + \tfrac{1}{2}\right)u\right]} {\sin\left(u/2\right)}\ec \end{equation}

\(D_N\) is even and real, \(D_N(0) = 2N+1\), and

\begin{equation}\tag{17.13} \frac{1}{2\pi}\int_{-\pi}^{\pi} D_N(u)\,\dd u = 1\ec\qquad \frac{1}{2\pi}\int_{0}^{\pi} D_N(u)\,\dd u = \frac{1}{2}\ep \end{equation}

Moreover the partial sum is the average of \(f\) against this kernel,

\begin{equation}\tag{17.14} S_N f(x) = \frac{1}{2\pi}\int_{-\pi}^{\pi} f(x - u)\,D_N(u)\,\dd u\ep \end{equation}

Rests on Definition 17.1 and Lemma 17.2.

Proof.

Derives Lemma 17.6. The sum is geometric with ratio \(r = \ee^{\ii u} \neq 1\): \(\sum_{n=-N}^{N} r^{n} = (r^{-N} - r^{N+1})/(1 - r)\). Multiply numerator and denominator by \(\ee^{-\ii u/2}\); the numerator becomes \(\ee^{-\ii(N+1/2)u} - \ee^{\ii(N+1/2)u} = -2\ii\sin[(N+\frac12)u]\) and the denominator \(\ee^{-\ii u/2} - \ee^{\ii u/2} = -2\ii\sin(u/2)\), giving Equation (17.12). Evenness and reality are visible in the closed form; \(D_N(0) = 2N+1\) by counting terms in the sum. Integrating the sum term by term and using Lemma 17.2 leaves only \(n = 0\), which gives the first identity of Equation (17.13), and the second follows by evenness. For Equation (17.14), insert Equation (17.3) into the partial sum and exchange the finite sum with the integral:

\begin{equation*} S_N f(x) = \frac{1}{2\pi}\int_{-\pi}^{\pi} f(s) \sum_{n=-N}^{N}\ee^{\ii n (x - s)}\,\dd s = \frac{1}{2\pi}\int_{-\pi}^{\pi} f(s)\,D_N(x - s)\,\dd s\ec \end{equation*}

and substitute \(u = x - s\), the range returning to \([-\pi,\pi]\) because both \(f\) and \(D_N\) are \(2\pi\)-periodic.

Everything about pointwise convergence is now a statement about the oscillatory kernel Equation (17.12). It has unit average but does not have unit absolute average — that failure is the subject of Lemma 17.12 — and it oscillates faster and faster as \(N\) grows. The instrument for exploiting fast oscillation is the following lemma.

Lemma 17.7 (Riemann–Lebesgue).

Let \(g\) be absolutely integrable on \([a,b]\) and Riemann integrable in the sense of Section 7.7. Then

\begin{equation}\tag{17.15} \lim_{\lambda\to\infty}\int_{a}^{b} g(u)\,\sin(\lambda u)\,\dd u = 0\ec \end{equation}

and likewise with \(\cos(\lambda u)\) or \(\ee^{-\ii\lambda u}\). Rests on Definition 7.39.

Proof.

Derives Lemma 17.7. First let \(g = \indic{u \in [\alpha,\beta]}\) be the indicator of a subinterval. Then

\begin{equation*} \abs{\int_{\alpha}^{\beta}\sin(\lambda u)\,\dd u} = \abs{\frac{\cos(\lambda\alpha) - \cos(\lambda\beta)}{\lambda}} \le \frac{2}{\lambda} \longrightarrow 0\ec \end{equation*}

so the claim holds for indicators and, by linearity, for every step function. Now let \(g\) be Riemann integrable and \(\varepsilon > 0\). By the definition of the Darboux integral (Definition 7.39) there is a partition whose lower step function \(\varphi\) satisfies \(\int_a^b\abs{g - \varphi}\,\dd u < \varepsilon/2\). Then

\begin{equation*} \abs{\int_a^b g\sin(\lambda u)\,\dd u} \le \int_a^b\abs{g - \varphi}\,\dd u + \abs{\int_a^b \varphi\sin(\lambda u)\,\dd u} < \frac{\varepsilon}{2} + \frac{\varepsilon}{2} \end{equation*}

for \(\lambda\) large enough, since the second term tends to zero. As \(\varepsilon\) was arbitrary, Equation (17.15) follows. On an unbounded interval one first truncates: absolute integrability makes the tail contribute less than \(\varepsilon/2\) uniformly in \(\lambda\).

Corollary 17.8 (Localisation).

Whether \(S_N f(x)\) converges, and to what, depends only on the values of \(f\) in an arbitrarily small neighbourhood of \(x\). Two integrable functions that agree on \((x-\delta, x+\delta)\) have Fourier series that converge or diverge together at \(x\), to the same limit when they converge. Rests on Lemmas 17.6 and 17.7.

Proof.

Derives Corollary 17.8. Let \(h = f_1 - f_2\) vanish on \((x-\delta, x+\delta)\). By Equation (17.14),

\begin{equation*} S_N h(x) = \frac{1}{2\pi}\int_{\delta \le \abs{u}\le\pi} \frac{h(x-u)}{\sin(u/2)}\, \sin\left[\left(N+\tfrac12\right)u\right]\dd u\ec \end{equation*}

and on the domain of integration \(\abs{\sin(u/2)} \ge \sin(\delta/2)>0\), so \(h(x-u)/\sin(u/2)\) is absolutely integrable there. Lemma 17.7 sends the integral to zero.

That a global object — the coefficients Equation (17.3) integrate \(f\) over the whole period — should have purely local convergence behaviour is Riemann's localisation principle, and it is the reason the theorem below needs hypotheses only at the point examined.

Theorem 17.9 (Dirichlet).

Let \(f\) be \(2\pi\)-periodic and absolutely integrable, and let \(x\) be a point at which the one-sided limits \(f(x^{\pm})\) exist and the Dini condition

\begin{equation}\tag{17.16} \int_{0}^{\delta} \frac{\abs{f(x+u) - f(x^{+})} + \abs{f(x-u) - f(x^{-})}}{u}\, \dd u < \infty \end{equation}

holds for some \(\delta > 0\). Then

\begin{equation}\tag{17.17} \lim_{N\to\infty} S_N f(x) = \frac{f(x^{+}) + f(x^{-})}{2}\ep \end{equation}

In particular this holds at every point of a piecewise continuously differentiable function, where the series converges to \(f\) at points of continuity and to the midpoint of the jump elsewhere. Rests on Lemma 17.6, Lemma 17.7 and Theorem 7.35.

Proof.

Derives Theorem 17.9. Split Equation (17.14) at \(u = 0\) and use the evenness of \(D_N\) together with the half-period normalisation Equation (17.13):

\begin{equation*} \begin{aligned} S_N f(x) - \frac{f(x^{+}) + f(x^{-})}{2} &= \frac{1}{2\pi}\int_{0}^{\pi} \left[f(x-u) - f(x^{-})\right] D_N(u)\,\dd u\\ &\quad + \frac{1}{2\pi}\int_{0}^{\pi} \left[f(x+u) - f(x^{+})\right] D_N(u)\,\dd u\ep \end{aligned} \end{equation*}

Insert Equation (17.12) and define, on \((0,\pi]\),

\begin{equation*} g_{\pm}(u) = \frac{f(x \pm u) - f(x^{\pm})}{\sin(u/2)}\ec \end{equation*}

so that each term is \(\frac{1}{2\pi}\int_0^\pi g_\pm(u)\sin[(N+\frac12)u]\,\dd u\). It remains to check that \(g_{\pm}\) is absolutely integrable, for then Lemma 17.7 sends both terms to zero. Away from the origin this is clear, since \(\sin(u/2)\) is bounded below there. Near the origin, \(\sin(u/2) \ge u/\pi\) for \(u \in (0,\pi]\) (concavity of the sine on \([0,\pi/2]\)), so

\begin{equation*} \int_{0}^{\delta}\abs{g_{\pm}(u)}\,\dd u \le \pi\int_{0}^{\delta}\frac{\abs{f(x\pm u) - f(x^{\pm})}}{u}\,\dd u < \infty \end{equation*}

by the Dini condition Equation (17.16). Finally, if \(f\) is piecewise \(C^{1}\) then \(\abs{f(x\pm u) - f(x^{\pm})} \le M u\) near \(0\) by the mean value theorem (Theorem 7.35) applied on each side, the integrand of Equation (17.16) is bounded, and the condition holds automatically.

Dirichlet's own hypothesis [Dirichlet:1829] was that \(f\) be bounded with finitely many maxima, minima and discontinuities in a period — a condition of bounded variation rather than of differentiability, which the proof above accommodates with the same estimate. His paper is the first rigorous convergence theorem in analysis, and the first place where the modern habit of stating hypotheses precisely enough to be checked appears in the subject.

Proposition 17.10 (Smoothness and coefficient decay).

Let \(f\) be \(2\pi\)-periodic and \(k-1\) times continuously differentiable, with \(f^{(k)}\) piecewise continuous. Then

\begin{equation}\tag{17.18} c_n\!\left(f^{(k)}\right) = (\ii n)^{k}\,c_n(f)\ec \end{equation}

and consequently \(\abs{c_n} = o\!\left(\abs{n}^{-k}\right)\). Conversely, a jump discontinuity in \(f\) itself forces \(c_n\) to decay no faster than \(1/\abs{n}\), as in Example 17.5. Rests on Definition 17.1, Lemma 17.7 and Equation (7.28).

Proof.

Derives Proposition 17.10. Integrate by parts once (Equation (7.28)):

\begin{equation*} c_n(f') = \frac{1}{2\pi}\int_{-\pi}^{\pi} f'(x)\ee^{-\ii n x}\dd x = \frac{1}{2\pi}\left[f\ee^{-\ii n x}\right]_{-\pi}^{\pi} + \frac{\ii n}{2\pi}\int_{-\pi}^{\pi} f(x)\ee^{-\ii n x}\dd x = \ii n\,c_n(f)\ec \end{equation*}

the boundary term vanishing because \(f\) and \(\ee^{-\ii n x}\) are both \(2\pi\)-periodic. Iterating \(k\) times gives Equation (17.18). Since \(f^{(k)}\) is integrable, its own coefficients tend to zero by Lemma 17.7, whence \(\abs{n}^{k}\abs{c_n(f)} \to 0\). For the converse, suppose \(\abs{c_n}\le C\abs{n}^{-1-\varepsilon}\) for some \(\varepsilon>0\); then \(\sum_n\abs{c_n}<\infty\), the series converges uniformly by the Weierstrass \(M\)-test, and its sum is continuous. A function with a jump — the square wave of Equation (17.11), whose coefficients are exactly of size \(1/n\) — therefore admits no such bound.

Corollary 17.11 (Uniform convergence for $C^{1}$ functions).

If \(f\) is \(2\pi\)-periodic and continuously differentiable, then \(\sum_n \abs{c_n} < \infty\) and the Fourier series converges to \(f\) absolutely and uniformly. Rests on Proposition 17.10, Proposition 17.4 and Theorem 17.9.

Proof.

Derives Corollary 17.11. By Cauchy–Schwarz in the sequence space and Equation (17.18),

\begin{equation*} \sum_{n\neq 0}\abs{c_n} = \sum_{n\neq 0}\frac{1}{\abs{n}}\,\abs{n\,c_n} \le \left(\sum_{n\neq0}\frac{1}{n^{2}}\right)^{1/2} \left(\sum_{n\neq0}\abs{c_n(f')}^{2}\right)^{1/2} \le \frac{\pi}{\sqrt{3}} \left(\frac{1}{2\pi}\int_{-\pi}^{\pi}\abs{f'}^{2}\dd x\right)^{1/2}\ec \end{equation*}

finite, using \(\sum_{n\ge1}n^{-2} = \pi^{2}/6\) and Bessel's inequality Equation (17.10) applied to \(f'\). The Weierstrass \(M\)-test then makes the series uniformly convergent, and Theorem 17.9 identifies its sum with \(f\) at every point.

So smoothness is read off the spectrum and conversely: this is the practical content of Fourier analysis for the solution of differential equations, and it is why the transform methods of Section 10.2.3 work at all.

The failures.

Dirichlet's theorem needs a hypothesis at the point, and no hypothesis of mere continuity will do. The obstruction is measurable, and it is the growth of the kernel's absolute average.

Lemma 17.12 (Divergence of the Lebesgue constants).

Let \(L_N = \frac{1}{2\pi}\int_{-\pi}^{\pi}\abs{D_N(u)}\,\dd u\). Then

\begin{equation}\tag{17.19} L_N \ge \frac{4}{\pi^{2}}\,\ln N \longrightarrow \infty\ep \end{equation}

Rests on Lemma 17.6.

Proof.

Derives Lemma 17.12. By evenness, \(L_N = \frac{1}{\pi}\int_0^{\pi}\abs{D_N(u)}\dd u\), and \(\abs{\sin(u/2)} \le u/2\) on \([0,\pi]\) gives \(\abs{D_N(u)} \ge 2\abs{\sin[(N+\frac12)u]}/u\). Substituting \(v = (N+\frac12)u\),

\begin{equation*} L_N \ge \frac{2}{\pi}\int_{0}^{(N+1/2)\pi}\frac{\abs{\sin v}}{v}\,\dd v \ge \frac{2}{\pi}\sum_{j=1}^{N}\frac{1}{j\pi} \int_{(j-1)\pi}^{j\pi}\abs{\sin v}\,\dd v = \frac{4}{\pi^{2}}\sum_{j=1}^{N}\frac{1}{j} \ge \frac{4}{\pi^{2}}\ln N\ec \end{equation*}

bounding \(1/v \ge 1/(j\pi)\) on the \(j\)-th half-period, using \(\int_{(j-1)\pi}^{j\pi}\abs{\sin v}\dd v = 2\), and comparing the harmonic sum with \(\int_1^{N}\dd j/j\).

Theorem 17.13 (du Bois-Reymond).

There exists a continuous \(2\pi\)-periodic function whose Fourier series diverges at a point [duBoisReymond:1876]. Rests on Lemmas 17.6 and 17.12.

Derivation. Derives Theorem 17.13. By Equation (17.14), \(f \mapsto S_N f(0)\) is a linear functional on the continuous \(2\pi\)-periodic functions with the supremum norm, and its norm is exactly \(L_N\): it is at most \(L_N\) by the triangle inequality, and at least \(L_N\) because a continuous \(f\) with \(\norm{f}_{\infty}\le1\) can be made to follow \(\sgn D_N(-u)\) except on arbitrarily short intervals around the finitely many sign changes. Lemma 17.12 shows this family of functionals is not uniformly bounded. The uniform boundedness principle of functional analysis then supplies an \(f\) for which \(\sup_N \abs{S_N f(0)} = \infty\).

The uniform boundedness principle is the one piece of functional analysis the argument above borrows, and neither it nor the Baire category theorem it rests on is developed in this treatise. It can be dispensed with: the rest of this subsection writes down an explicit continuous function whose Fourier series diverges at the origin, so that Theorem 17.13 depends on nothing unproved. Two elementary estimates come first.

Lemma 17.14 (Uniformly bounded sine sums).

For every integer \(M \ge 1\) and every real \(x\),

\begin{equation}\tag{17.20} \abs{\sum_{k=1}^{M}\frac{\sin(kx)}{k}} \le 1 + 2\pi\ep \end{equation}

The bound is uniform in \(M\) and in \(x\), although the terms are not absolutely summable. Rests on Proposition 8.4.

Proof.

Derives Lemma 17.14. The sum is odd and \(2\pi\)-periodic in \(x\) and vanishes at \(x = 0\), so it suffices to bound it for \(0 < x \le \pi\). Write \(T_m = \sum_{k=1}^{m}\sin(kx)\). Multiplying the finite geometric progression \(\sum_{k=1}^{m}\ee^{\ii k x}\) by \(\ee^{\ii x} - 1\) telescopes it to \(\ee^{\ii(m+1)x} - \ee^{\ii x}\), so for \(\ee^{\ii x} \neq 1\) — the case \(\ee^{\ii x} = 1\) being \(x = 0\), already excluded — that progression sums to \(\ee^{\ii x}\left(\ee^{\ii m x} - 1\right)/\left(\ee^{\ii x} - 1\right)\). Taking imaginary parts (Proposition 8.4),

\begin{equation}\tag{17.21} T_m = \Im\frac{\ee^{\ii x}\left(\ee^{\ii m x} - 1\right)} {\ee^{\ii x} - 1}\ec\qquad\text{so}\qquad \abs{T_m} \le \frac{2}{\abs{\ee^{\ii x} - 1}} = \frac{1}{\sin(x/2)}\ec \end{equation}

using \(\abs{\ee^{\ii x} - 1} = 2\sin(x/2)\) on that range. There \(\sin(x/2) \ge x/\pi\), by concavity of the sine on \([0,\pi/2]\), so \(\abs{T_m} \le \pi/x\) for every \(m\).

Let \(K = \lfloor 1/x \rfloor\). For \(k \le K\) use \(\abs{\sin(kx)} \le kx\), which gives

\begin{equation*} \sum_{k=1}^{\min(K,M)}\frac{\abs{\sin(kx)}}{k} \le K x \le 1\ep \end{equation*}

For \(k > K\), and assuming \(M > K\) since otherwise there is nothing left, summation by parts against \(T_k - T_{k-1} = \sin(kx)\) gives

\begin{equation*} \sum_{k=K+1}^{M}\frac{\sin(kx)}{k} = \frac{T_M}{M} - \frac{T_K}{K+1} + \sum_{k=K+1}^{M-1} T_k\left(\frac{1}{k} - \frac{1}{k+1}\right)\ec \end{equation*}

in which the coefficients \(1/k - 1/(k+1)\) are positive and telescope to \(1/(K+1) - 1/M\). Bounding every \(\abs{T_k}\) by \(\pi/x\), the modulus of the right-hand side is at most

\begin{equation*} \frac{\pi}{x}\left(\frac{1}{M} + \frac{1}{K+1} + \frac{1}{K+1} - \frac{1}{M}\right) = \frac{2\pi}{x\left(K+1\right)} < 2\pi\ec \end{equation*}

the last step because \(K + 1 > 1/x\). Adding the two ranges gives Equation (17.20).

Lemma 17.15 (A polynomial whose partial sum spikes at the origin).

For integers \(1 \le N < n\) put

\begin{equation}\tag{17.22} Q_{N,n}(x) := 2\sin(nx)\sum_{k=1}^{N}\frac{\sin(kx)}{k} = \sum_{k=1}^{N} \frac{\cos\left[(n-k)x\right] - \cos\left[(n+k)x\right]}{k}\ep \end{equation}

Then \(Q_{N,n}\) is a real \(2\pi\)-periodic trigonometric polynomial, and every frequency \(m\) occurring in it satisfies \(n - N \le \abs{m} \le n + N\). Moreover

\begin{equation}\tag{17.23} \abs{Q_{N,n}(x)} \le 2\left(1 + 2\pi\right)\ec\qquad Q_{N,n}(0) = 0\ec\qquad S_{n}Q_{N,n}(0) = \sum_{k=1}^{N}\frac{1}{k} \ge \ln N\ec \end{equation}

the first for every real \(x\). Rests on Lemma 17.14 and Definition 17.1.

Proof.

Derives Lemma 17.15. The second expression in Equation (17.22) is the first, term by term, by \(2\sin(nx)\sin(kx) = \cos\left[(n-k)x\right] - \cos\left[(n+k)x\right]\). It writes \(Q_{N,n}\) as a combination of \(\cos(mx)\) for \(m \in \set{n-N,\dots,n-1}\cup\set{n+1,\dots,n+N}\), all positive because \(N < n\); since \(\cos(mx) = \tfrac12\left(\ee^{\ii m x} + \ee^{-\ii m x}\right)\), the frequencies present are \(\pm m\) for those \(m\), and each satisfies \(n - N \le \abs{m} \le n + N\).

The bound in Equation (17.23) is Equation (17.20) multiplied by \(\abs{2\sin(nx)} \le 2\), and \(Q_{N,n}(0) = 0\) because every \(\sin(kx)\) vanishes there. Finally the symmetric partial sum \(S_{n}\) of Definition 17.1 retains exactly the frequencies with \(\abs{m} \le n\): of the terms just listed it keeps those with \(m = n - k\) and discards those with \(m = n + k\), so

\begin{equation*} S_{n}Q_{N,n}(x) = \sum_{k=1}^{N}\frac{\cos\left[(n-k)x\right]}{k}\ec \end{equation*}

which at \(x = 0\) is the harmonic sum \(\sum_{k=1}^{N}1/k\); comparison with \(\int_{1}^{N+1}\dd t/t\) bounds it below by \(\ln(N+1) > \ln N\).

Proposition 17.16 (An explicit continuous function with a divergent Fourier series).

Put \(N_j = 2^{j^{3}}\) and \(n_j = 2N_j\) for \(j \ge 1\), and

\begin{equation}\tag{17.24} f(x) := \sum_{j=1}^{\infty}\frac{1}{j^{2}}\,Q_{N_j,n_j}(x) \end{equation}

with \(Q_{N,n}\) as in Equation (17.22). Then \(f\) is continuous and \(2\pi\)-periodic, and

\begin{equation}\tag{17.25} S_{n_j}f(0) = \frac{1}{j^{2}}\sum_{k=1}^{N_j}\frac{1}{k} \ge j\ln 2 \longrightarrow \infty\ec \end{equation}

so its Fourier series diverges at the origin. This proves Theorem 17.13 without appeal to the uniform boundedness principle. Rests on Lemma 17.15 and Definition 17.1.

Proof.

Derives Proposition 17.16. Continuity. Each term of Equation (17.24) is a trigonometric polynomial, hence continuous and \(2\pi\)-periodic, and by Equation (17.23) it is bounded in modulus by \(2(1+2\pi)/j^{2}\). Since \(\sum_{j}j^{-2}\) converges, the Weierstrass \(M\)-test makes the convergence uniform, and a uniform limit of continuous periodic functions is continuous and periodic.

The coefficients are the termwise ones. Write \(f_J\) for the \(J\)-th partial sum of Equation (17.24). For any \(m \in \Z\), Equation (17.3) gives

\begin{equation*} \abs{c_m(f) - c_m(f_J)} = \abs{\frac{1}{2\pi}\int_{-\pi}^{\pi} \left[f(x) - f_J(x)\right]\ee^{-\ii m x}\,\dd x} \le \sup_{x}\abs{f - f_J} \longrightarrow 0\ec \end{equation*}

so \(c_m(f) = \sum_{j}j^{-2}c_m\!\left(Q_{N_j,n_j}\right)\). No rearrangement is involved, because the next paragraph shows at most one summand is nonzero for each \(m\).

The blocks occupy disjoint frequency bands. By Lemma 17.15 the \(j\)-th term occupies the frequencies \(m\) with \(N_j \le \abs{m} \le 3N_j\), since \(n_j - N_j = N_j\) and \(n_j + N_j = 3N_j\). For \(i > j\),

\begin{equation*} \frac{N_i}{N_j} \ge \frac{N_{j+1}}{N_j} = 2^{(j+1)^{3} - j^{3}} = 2^{3j^{2} + 3j + 1} \ge 2^{7} > 3\ec \end{equation*}

so the bands are pairwise disjoint and move upward with \(j\).

The partial sum at the origin. Fix \(j\) and apply \(S_{n_j}\), which keeps the frequencies with \(\abs{m} \le n_j = 2N_j\). A block \(i < j\) lies entirely below the cut, since \(3N_i \le 3N_{j-1} < N_j < 2N_j\), and contributes \(i^{-2}Q_{N_i,n_i}(0) = 0\) by Equation (17.23). A block \(i > j\) lies entirely above it, since \(N_i > 3N_j > 2N_j\), and contributes nothing. The block \(i = j\) contributes \(j^{-2}S_{n_j}Q_{N_j,n_j}(0) = j^{-2}\sum_{k=1}^{N_j}k^{-1}\), again by Equation (17.23), and that sum is at least \(\ln N_j = j^{3}\ln 2\). This is Equation (17.25). Since \(\sup_N\abs{S_N f(0)} = \infty\), the series cannot converge at \(x = 0\).

Remark 17.17 (What the explicit example costs and buys).

The soft proof of Theorem 17.13 is shorter and says more: it shows that most continuous functions, in the sense of Baire category, have divergent Fourier series somewhere, whereas Proposition 17.16 exhibits one. What the explicit construction buys is independence: the treatise nowhere proves the uniform boundedness principle, and a chapter that leaned on it would be resting a stated theorem on an unstated one. The two arguments share their mechanism — the harmonic sum \(\sum_{k\le N}1/k\) of Equation (17.23) is the same logarithm that Equation (17.19) produces — which is why the size of the Lebesgue constants is the real content of the theorem either way.

Two further results mark the boundary of the subject and are stated without proof, at the depth cap of Part II. Kolmogorov constructed an absolutely integrable function whose Fourier series diverges at every point [Kolmogorov:1926] — so integrability alone guarantees nothing pointwise. Carleson proved that for square-integrable functions the series converges at almost every point [Carleson:1966], which settles the case relevant to physics, since the states of Hilbert Spaces and the finite-energy signals of Section 17.2.2 are square integrable. The theorem was extended shortly afterwards to \(L^{p}\) for every \(p>1\); the endpoint \(p=1\) is Kolmogorov's counterexample, so the classification is complete.

Remark 17.18 (Two theorems quoted, and not proved).

Carleson's theorem [Carleson:1966] and Kolmogorov's example [Kolmogorov:1926] are the only statements of this chapter that carry no derivation, and the omission is a decision rather than a debt. Carleson's argument, and every later proof of the same theorem, turns on the boundedness of a maximal operator obtained by decomposing the partial-sum operator over dyadic families in frequency and discarding an exceptional set in time; the estimates that control it are the subject matter of a course in harmonic analysis, none of that machinery appears anywhere in this treatise, and introducing it to serve one statement would unbalance Part II. Kolmogorov's counterexample is a construction of the same explicit kind as Proposition 17.16, but far more delicate: the function must be arranged so that the partial sums are unbounded at every point at once. Both are quoted because they fix the boundary of the subject, and a reader needs to know where it lies. Nothing in this chapter, and nothing in the physical parts that consume it, uses either: the working statements are Theorem 17.9, Corollary 17.11, Theorem 17.22 and the mean-square convergence of Section 17.2.2, each proved in full here.

Remark 17.19 (What physics actually needs).

None of the pathology touches the working use of Fourier series. The signals of the experimental chapters are square integrable, so Proposition 17.4 and Section 17.2.2 give convergence in the mean — which is the sense in which energies, spectra and expectation values are computed — and the pieces of them that are piecewise smooth converge pointwise by Theorem 17.9. The counterexamples matter because they say what may not be assumed: that a continuous waveform is reconstructed pointwise by its harmonics is false, and any argument that needs it needs a hypothesis.

The Gibbs phenomenon and summability

At a jump, Theorem 17.9 promises convergence to the midpoint at each fixed point, but says nothing uniform. Something specific and permanent goes wrong nearby.

Theorem 17.20 (Overshoot at a jump).

Let \(f(x) = \sgn(\sin x)\), the odd square wave of Example 17.5 with \(T = 2\pi\), whose jump at the origin has height \(2\). Then the partial sum \(S_{2N+1}f\) attains its first maximum to the right of the origin at \(x_N = \pi/(2N+2)\), and

\begin{equation}\tag{17.26} \lim_{N\to\infty} S_{2N+1}f(x_N) = \frac{2}{\pi}\int_{0}^{\pi}\frac{\sin v}{v}\,\dd v = \frac{2}{\pi}\operatorname{Si}(\pi) = 1.178979\ldots \end{equation}

The partial sums therefore overshoot the limiting value \(1\) by \(0.178979\ldots\), that is by \(8.949\,\mathrm{\%}\) of the jump height, and the overshoot does not diminish with \(N\): it merely migrates toward the discontinuity, since \(x_N \to 0\). Rests on Equation (17.11) and Theorem 7.38.

Derivation. Derives Theorem 17.20. By Equation (17.11) the partial sum through the \((2N+1)\)-th harmonic is

\begin{equation*} S_{2N+1}f(x) = \frac{4}{\pi}\sum_{k=0}^{N}\frac{\sin\left[(2k+1)x\right]}{2k+1}\ec \qquad\text{so}\qquad \dv{}{x}S_{2N+1}f(x) = \frac{4}{\pi}\sum_{k=0}^{N}\cos\left[(2k+1)x\right]\ep \end{equation*}

The finite sum is geometric: \(\sum_{k=0}^{N}\ee^{\ii(2k+1)x} = \ee^{\ii x}\left(1 - \ee^{2\ii(N+1)x}\right)/\left(1-\ee^{2\ii x}\right) = \ee^{\ii(N+1)x}\sin\left[(N+1)x\right]/\sin x\), whose real part is \(\sin\left[2(N+1)x\right]/(2\sin x)\). Hence

\begin{equation}\tag{17.27} \dv{}{x}S_{2N+1}f(x) = \frac{2}{\pi}\,\frac{\sin\left[2(N+1)x\right]}{\sin x}\ep \end{equation}

For \(0 < x < \pi/(2N+2)\) the right-hand side is positive and it changes sign at \(x_N = \pi/(2N+2)\), so that point is the first maximum. Integrating Equation (17.27) from \(0\), where the partial sum vanishes, and substituting \(v = 2(N+1)x\),

\begin{equation*} S_{2N+1}f(x_N) = \frac{2}{\pi}\int_{0}^{x_N}\frac{\sin\left[2(N+1)x\right]}{\sin x}\,\dd x = \frac{2}{\pi}\int_{0}^{\pi}\frac{\sin v} {2(N+1)\sin\left[v/(2(N+1))\right]}\,\dd v\ep \end{equation*}

As \(N \to \infty\) the denominator tends to \(v\) uniformly on \([0,\pi]\) — indeed \(0 \le v - 2(N+1)\sin[v/(2(N+1))] \le v^{3}/(24(N+1)^{2})\) from the alternating Taylor bound for the sine (Theorem 7.38) — so the integrand converges uniformly to \(\sin v/v\), and the limit is Equation (17.26). Numerically \(\operatorname{Si}(\pi) = 1.851937\ldots\), giving \(1.178979\ldots\), and the excess over \(1\) divided by the jump height \(2\) is \(8.949\,\mathrm{\%}\).

Remark 17.21 (The name is wrong, and the record says so).

The overshoot was described, computed and published by Henry Wilbraham in 1848 [Wilbraham:1848], half a century before the exchange in Nature in which Josiah Willard Gibbs — correcting his own earlier letter — gave the value quoted above [Gibbs:1899]. The customary name “Gibbs phenomenon” therefore misattributes the discovery. This treatise keeps the customary name, because a reader searching the literature needs it, and records the misattribution here rather than passing it on silently.

The overshoot is not a defect of the coefficients, which are exactly right, but of the summation method: the sharp truncation \(\sum_{\abs{n}\le N}\) multiplies the spectrum by a rectangular window, and Lemma 17.6 shows this is equivalent to smoothing \(f\) with an oscillating, sign-changing kernel. Change the window and the pathology goes away.

Theorem 17.22 (Fejér).

Let \(\sigma_N f = \frac{1}{N+1}\sum_{n=0}^{N} S_n f\) be the Cesàro average of the partial sums. Then

\begin{equation}\tag{17.28} \sigma_N f(x) = \frac{1}{2\pi}\int_{-\pi}^{\pi} f(x-u)\,K_N(u)\,\dd u\ec \qquad K_N(u) = \frac{1}{N+1} \left(\frac{\sin\left[(N+1)u/2\right]}{\sin(u/2)}\right)^{2}\ec \end{equation}

and if \(f\) is continuous and \(2\pi\)-periodic then \(\sigma_N f \to f\) uniformly [Fejer:1904]. Rests on Lemma 17.6 and Theorem 7.25.

Proof.

Derives Theorem 17.22. Averaging Equation (17.14) gives Equation (17.28) with \(K_N = \frac{1}{N+1}\sum_{n=0}^{N} D_n\). For the closed form multiply by \(\sin^{2}(u/2)\) and use \(\sin\left[(n+\tfrac12)u\right]\sin(u/2) = \tfrac12\left[\cos(nu) - \cos((n+1)u)\right]\); the sum telescopes to \(\tfrac12\left[1 - \cos((N+1)u)\right] = \sin^{2}\left[(N+1)u/2\right]\), which is the assertion.

The closed form exhibits three properties. (i) \(K_N \ge 0\), being a square. (ii) \(\frac{1}{2\pi}\int_{-\pi}^{\pi}K_N = 1\), by averaging Equation (17.13). (iii) For \(\delta \le \abs{u}\le\pi\), \(K_N(u) \le \left[(N+1)\sin^{2}(\delta/2)\right]^{-1}\), which tends to \(0\) uniformly on that set. A family with these three properties is an approximate identity, and the standard argument now applies: given \(\varepsilon>0\), uniform continuity of \(f\) on the compact circle (Theorem 7.25) gives \(\delta\) with \(\abs{f(x-u) - f(x)} < \varepsilon/2\) for \(\abs{u}<\delta\), uniformly in \(x\). Then, using (ii) to write \(f(x) = \frac{1}{2\pi}\int K_N(u) f(x)\,\dd u\),

\begin{equation*} \abs{\sigma_N f(x) - f(x)} \le \frac{1}{2\pi}\int_{\abs{u}<\delta}\!\!\abs{f(x-u)-f(x)}K_N(u)\dd u + \frac{1}{2\pi}\int_{\delta\le\abs{u}\le\pi}\!\!2\norm{f}_{\infty} K_N(u)\,\dd u\ec \end{equation*}

where (i) was used to drop the absolute value on \(K_N\). The first term is below \(\varepsilon/2\) by (ii); the second is below \(\varepsilon/2\) for \(N\) large by (iii). The bound is uniform in \(x\).

Corollary 17.23 (Density, uniqueness, and no overshoot).

(a) Every continuous periodic function is a uniform limit of trigonometric polynomials. (b) If all Fourier coefficients of a continuous periodic \(f\) vanish, then \(f \equiv 0\); hence the coefficients determine the function. (c) \(\min f \le \sigma_N f \le \max f\) pointwise, so the Cesàro means exhibit no overshoot whatever at a jump. Rests on Theorem 17.22.

Proof.

Derives Corollary 17.23. (a) \(\sigma_N f\) is a trigonometric polynomial of degree \(N\). (b) If every \(c_n = 0\) then every \(S_n f = 0\), so \(\sigma_N f = 0\), and Theorem 17.22 gives \(f = \lim\sigma_N f = 0\). (c) \(K_N \ge 0\) with unit average, so Equation (17.28) exhibits \(\sigma_N f(x)\) as a weighted average of values of \(f\), which cannot leave the range of \(f\).

Part (b) is what closes Theorem 17.3 into a genuine one-to-one correspondence, and it is used in Section 17.2.2 to prove completeness. Part (c) is the practical lesson: the overshoot belongs to the rectangular window and not to the data.

Example 17.24 (Ringing in a band-limited reproduction).

A \(1.0\,\mathrm{kHz}\) square wave passed through a system with a sharp cutoff at \(20\,\mathrm{kHz}\) retains harmonics up to the nineteenth, i.e. \(N = 9\) in Theorem 17.20. Its reconstruction still overshoots by \(9\,\mathrm{\%}\) of the step, in a ripple whose first peak sits \(x_N/\omega_1 = 1/\left[4(N+1)\nu_1\right] = 25\,\mu\mathrm{s}\) after the edge. Increasing the cutoff moves the ripple closer to the edge and makes it narrower; it never makes it smaller. Tapering the cutoff instead — which is what a window function does, and what Theorem 17.22 does in its extreme form — trades the ripple for a slower edge. The same trade reappears as spectral leakage in Section 17.5.2 and as the choice of apodisation in every Fourier-transform spectrometer.

Series in several variables

Everything above concerns a function of one variable. A field periodic in space, a motion on a torus, a crystal potential: each is periodic with respect to a lattice in \(\R^{n}\), and the expansion it admits is indexed by the reciprocal lattice. Nothing new has to be proved for it, which is exactly the point of stating it.

Definition 17.25 (Multiple Fourier series).

Let \(\Lambda\subset\R^{n}\) be a lattice with primitive vectors \(\vect{a}_{1},\dots,\vect{a}_{n}\), let \(C\) be a primitive cell of volume \(\abs{C}\), and let \(\Lambda^{*}\) be the reciprocal lattice generated by the \(\vect{b}_{j}\) with \(\vect{a}_{i}\cdot\vect{b}_{j}=2\pi\delta_{ij}\). For \(f\) square integrable on \(C\) and \(\Lambda\)-periodic, the multiple Fourier coefficients are

\begin{equation}\tag{17.29} c_{\vect{G}}=\frac{1}{\abs{C}}\int_{C}f(\vect{x})\, \ee^{-\ii\vect{G}\cdot\vect{x}}\,\dd^{n}x\ec \qquad\vect{G}\in\Lambda^{*}\ec \end{equation}

and the associated series is \(\sum_{\vect{G}\in\Lambda^{*}}c_{\vect{G}}\, \ee^{\ii\vect{G}\cdot\vect{x}}\). If \(f\) carries an SI unit, so does every \(c_{\vect{G}}\); the exponent is a pure number because \(\vect{G}\) carries \(/\mathrm{m}\). Rests on Definition 17.1 and Corollary 17.49.

Proposition 17.26 (The lattice harmonics are an orthonormal basis).

The functions \(e_{\vect{G}}(\vect{x}) =\abs{C}^{-1/2}\ee^{\ii\vect{G}\cdot\vect{x}}\), \(\vect{G}\) running over \(\Lambda^{*}\), form an orthonormal basis of \(L^{2}(C)\). Consequently, for \(\Lambda\)-periodic square-integrable \(f\): the coefficients Equation (17.29) determine \(f\) up to a null set; the series converges to \(f\) in the mean square; and Parseval's identity holds,

\begin{equation}\tag{17.30} \frac{1}{\abs{C}}\int_{C}\abs{f}^{2}\,\dd^{n}x =\sum_{\vect{G}\in\Lambda^{*}}\abs{c_{\vect{G}}}^{2}\ep \end{equation}

If in addition \(\sum_{\vect{G}}\abs{c_{\vect{G}}}<\infty\), the series converges absolutely and uniformly, and its sum is continuous and equal to \(f\). Rests on Definition 17.25, Theorem 17.22 and Theorem 12.30.

Proof.

Derives Proposition 17.26. Change variables to the lattice coordinates \(\vect{x}=u^{i}\vect{a}_{i}\) with \(u^{i}\in[0,1)\), in which \(C\) is the unit cube and \(\vect{G}\cdot\vect{x}=2\pi\,\ell_{i}u^{i}\) for \(\vect{G}=\ell_{i}\vect{b}_{i}\) with integer \(\ell_{i}\). The substitution is linear, so its Jacobian is the constant \(\abs{C}\) and cancels the prefactor of Equation (17.29); no change-of-variables theorem beyond the scaling of a volume by a determinant is used. The integral Equation (17.29) then factorises into \(n\) copies of the one-dimensional integral of Lemma 17.2, so that

\begin{equation}\tag{17.31} \frac{1}{\abs{C}}\int_{C} \ee^{\ii(\vect{G}-\vect{G}')\cdot\vect{x}}\,\dd^{n}x =\prod_{i=1}^{n}\delta_{\ell_{i}\ell_{i}'}\ec \end{equation}

and the family \(\set{e_{\vect{G}}}\) is orthonormal in \(L^{2}(C)\).

For completeness, take the product of \(n\) copies of the Fejér kernel of Equation (17.28), one in each coordinate \(u^{i}\), and let \(\sigma_{N}f\) be the corresponding average of \(f\). The product kernel inherits the three properties established in the proof of Theorem 17.22: it is non-negative, its average over the cube is \(1\), and outside a cube of side \(2\delta\) it is bounded by a quantity tending to zero, because at least one factor is then evaluated at an argument of modulus at least \(\delta\) while the others integrate to at most \(1\). It is therefore an approximate identity, and the argument of that proof, run with the cube in place of the circle and with uniform continuity supplied by Theorem 7.25 on the compact torus, gives \(\sigma_{N}f\to f\) uniformly for continuous \(\Lambda\)-periodic \(f\). Each \(\sigma_{N}f\) is a finite combination of the \(e_{\vect{G}}\), so their span is dense in the continuous periodic functions for the supremum norm and hence, \(C\) having finite volume, in \(L^{2}(C)\) (Remark 17.28). So the orthonormal family is total, which by Theorem 12.30 makes it a basis, and that theorem supplies the mean-square convergence and Equation (17.30); the uniqueness statement is Equation (17.30) applied to a difference.

The last claim is the Weierstrass criterion: each term has modulus \(\abs{c_{\vect{G}}}\), so an absolutely convergent sum of the moduli forces uniform convergence of a series of continuous functions. Its sum is therefore continuous, and it agrees with \(f\) outside a null set because it also converges in the mean square; if \(f\) is itself continuous the two agree at every point.

Remark 17.27 (Where the several-variable case is used).

Three places, and in each the value of Proposition 17.26 is that it costs nothing. A quasiperiodic motion with \(n\) incommensurate frequencies is by construction a function on an \(n\)-torus evaluated along a straight line, so its time dependence is a multiple Fourier series and its power sits on the countable set of integer combinations of the frequencies — which is what Nonlinear Dynamics and Chaos needs when it distinguishes quasiperiodicity from chaos by the presence of a continuous spectral component. A one-electron potential in a crystal is \(\Lambda\)-periodic, and Equation (17.29) defines its structure factors (Electrons in Solids: Band Theory). And Corollary 17.49 is the Poisson-summation counterpart of the same statement, relating a lattice sum to a sum over \(\Lambda^{*}\).

The Fourier transform

Remark 17.28 (Analytical inputs quoted, not proved).

From here on three standard tools of integration theory are used freely and are not developed in this treatise, whose integral is the Riemann–Darboux integral of Section 7.7: the Fubini–Tonelli theorem on exchanging the order of a double integral, the dominated convergence theorem, and the density of smooth rapidly decreasing functions in the absolutely integrable and square-integrable classes. Each is quoted explicitly at the point of use rather than absorbed silently. Everything else in this chapter is proved.

Definition and inversion

Let \(f\) be absolutely integrable on the whole line and, for the moment, supported in \([-T/2, T/2]\). Its Fourier series on that interval, with \(\omega_1 = 2\pi/T\) and \(\omega_n = n\omega_1\), is built from the coefficients Equation (17.3). Introduce

\begin{equation}\tag{17.32} \hat f(\omega) = \frac{1}{\sqrt{2\pi}} \int_{-\infty}^{\infty} f(t)\,\ee^{-\ii\omega t}\,\dd t\ec \end{equation}

so that \(c_n = \sqrt{2\pi}\,\hat f(\omega_n)/T\), and note that \(\sqrt{2\pi}/T = \Delta\omega/\sqrt{2\pi}\) with \(\Delta\omega = \omega_{n+1}-\omega_n = \omega_1\). The synthesis formula Equation (17.4) then reads

\begin{equation}\tag{17.33} f(t) = \frac{1}{\sqrt{2\pi}} \sum_{n=-\infty}^{\infty}\hat f(\omega_n)\,\ee^{\ii\omega_n t}\, \Delta\omega\ec \end{equation}

which is a Riemann sum for \(\frac{1}{\sqrt{2\pi}}\int\hat f(\omega)\ee^{\ii\omega t}\dd\omega\) with mesh \(\Delta\omega = 2\pi/T \to 0\) as the interval is enlarged. This is the passage left as a pending derivation at Equation (9.121); as an argument it is only heuristic, because the summand depends on \(T\) as well as on the mesh, and Theorem 17.32 below supplies the proof. The discrete spectrum of a periodic function becomes a continuous spectrum, and the spacing \(\Delta\nu = 1/T\) of the harmonics — in \(\mathrm{Hz}\) for a signal in time — becomes the resolution of the transform, a fact that returns quantitatively in Section 17.5.2.

Definition 17.29 (Fourier transform; the treatise convention).

For absolutely integrable \(f\) the Fourier transform is Equation (17.32), also written \(\mathcal{F}f = \hat f\), with the inverse transform

\begin{equation}\tag{17.34} \left(\mathcal{F}^{-1}g\right)(t) = \frac{1}{\sqrt{2\pi}} \int_{-\infty}^{\infty} g(\omega)\,\ee^{+\ii\omega t}\,\dd\omega\ep \end{equation}

The symmetric factor \(1/\sqrt{2\pi}\) and the sign \(\ee^{-\ii\omega t}\) in the analysis kernel are fixed for the whole treatise, matching Equation (9.121). For a function of position the pair is written with the wavenumber \(k\) in place of \(\omega\) and \(\ee^{-\ii kx}\) in place of \(\ee^{-\ii\omega t}\); if \(x\) is in \(\mathrm{m}\) then \(k\) is in \(/\mathrm{m}\), and if \(t\) is in \(\mathrm{s}\) then \(\omega\) is in \(\mathrm{rad}/\mathrm{s}\) and \(\nu = \omega/2\pi\) in \(\mathrm{Hz}\). The transform carries the SI unit of \(f\) multiplied by that of \(t\): the transform of a voltage record in \(\mathrm{V}\) is in \(\mathrm{V}\,\mathrm{s}\).

Other normalisations are in use and none is wrong; what is fatal is mixing them, since the factors of \(2\pi\) move between the transform, the inverse and the convolution theorem. Table 17.1 is the dictionary.

ConventionAnalysisSynthesisUsed by
Symmetric, angular frequency (this treatise)$\dfrac{1}{\sqrt{2\pi}}\displaystyle\int f\,\ee^{-\ii\omega t}\dd t$$\dfrac{1}{\sqrt{2\pi}}\displaystyle\int\hat f\, \ee^{\ii\omega t}\dd\omega$mathematical physics
Asymmetric, angular frequency$\displaystyle\int f\,\ee^{-\ii\omega t}\dd t$$\dfrac{1}{2\pi}\displaystyle\int F\,\ee^{\ii\omega t}\dd\omega$field theory, solid state
Ordinary frequency$\displaystyle\int f\,\ee^{-2\pi\ii\nu t}\dd t$$\displaystyle\int\tilde f\,\ee^{2\pi\ii\nu t}\dd\nu$signal processing
The three conventions in common use, and the translation between them. Writing $\hat f$ for the symmetric transform of Equation (17.32), $F$ for the asymmetric one and $\tilde f$ for the ordinary-frequency one, the dictionary is $\hat f(\omega) = F(\omega)/\sqrt{2\pi} = \tilde f(\omega/2\pi)/\sqrt{2\pi}$. Only the first and third are unitary, so only they make Plancherel's theorem (Theorem 17.36) free of stray factors.
Proposition 17.30 (Elementary properties).

Let \(f, g\) be absolutely integrable, \(a \in \R\), \(\lambda \neq 0\). Then \(\hat f\) is bounded, with \(\norm{\hat f}_{\infty}\le\norm{f}_{1}/\sqrt{2\pi}\), uniformly continuous, and \(\hat f(\omega)\to0\) as \(\abs{\omega}\to\infty\). Moreover

\begin{align} \widehat{f(\cdot - a)}(\omega) &= \ee^{-\ii\omega a}\,\hat f(\omega) &&\text{(translation)}\ec\tag{17.35}\\ \widehat{\ee^{\ii\omega_0 \cdot}f}(\omega) &= \hat f(\omega-\omega_0) &&\text{(modulation)}\ec\tag{17.36}\\ \widehat{f(\lambda\,\cdot)}(\omega) &= \frac{1}{\abs{\lambda}}\,\hat f\!\left(\frac{\omega}{\lambda}\right) &&\text{(scaling)}\ec\tag{17.37}\\ \widehat{f'}(\omega) &= \ii\omega\,\hat f(\omega) &&\text{(differentiation)}\ec\tag{17.38}\\ \widehat{t f}(\omega) &= \ii\,\dv{\hat f}{\omega}(\omega) &&\text{(multiplication by }t)\ep\tag{17.39} \end{align}

the last two under the hypotheses that \(f'\), respectively \(tf(t)\), is itself absolutely integrable and that \(f\) vanishes at infinity. Rests on Definition 17.29, Lemma 17.7 and Equation (7.28).

Proof.

Derives Proposition 17.30. The bound is immediate from \(\abs{\hat f}\le\frac{1}{\sqrt{2\pi}}\int\abs{f}\). Uniform continuity: \(\abs{\hat f(\omega+h)-\hat f(\omega)} \le\frac{1}{\sqrt{2\pi}}\int\abs{f(t)}\abs{\ee^{-\ii h t}-1}\dd t\), which tends to zero with \(h\) by dominated convergence, independently of \(\omega\). The decay is Lemma 17.7 extended to the line as noted in its proof. Equation (17.35) and Equation (17.36) are the substitutions \(t\mapsto t+a\) and a regrouping of exponentials; Equation (17.37) is the substitution \(s=\lambda t\), the modulus arising because the limits reverse when \(\lambda<0\). For Equation (17.38), integrate by parts (Equation (7.28)); the boundary term \(\left[f\ee^{-\ii\omega t}\right]_{-\infty}^{\infty}\) vanishes because \(f\to0\). For Equation (17.39), differentiate Equation (17.32) under the integral sign, legitimate by dominated convergence when \(tf(t)\) is absolutely integrable, which brings down a factor \(-\ii t\).

Equation (17.38) is the reason the transform exists as a tool: it converts \(\dd/\dd t\) into multiplication by \(\ii\omega\), hence a constant-coefficient linear differential operator into a polynomial, and a linear partial differential equation into an algebraic one in the transformed variable. That is precisely the manoeuvre of Section 10.2.3, and it is the pending derivative rule Equation (9.122), now proved.

Example 17.31 (The three standard pairs).

(i) Rectangle and sinc. For the rectangular pulse \(f = \indic{\abs{t}\le\tau/2}\) of duration \(\tau\),

\begin{equation}\tag{17.40} \hat f(\omega) = \frac{1}{\sqrt{2\pi}}\int_{-\tau/2}^{\tau/2} \ee^{-\ii\omega t}\,\dd t = \frac{\tau}{\sqrt{2\pi}}\, \operatorname{sinc}\!\left(\frac{\omega\tau}{2}\right)\ec \qquad \operatorname{sinc} x = \frac{\sin x}{x}\ep \end{equation}

The first zero sits at \(\omega = 2\pi/\tau\), i.e. at \(\nu = 1/\tau\): a pulse of duration \(1.0\,\mu\mathrm{s}\) has its first spectral null at \(1.0\,\mathrm{MHz}\).

(ii) Exponential and Lorentzian. For the causal decay \(f(t) = \ee^{-\alpha t}\) for \(t>0\) and \(0\) for \(t<0\), with \(\alpha>0\) in \(/\mathrm{s}\),

\begin{equation}\tag{17.41} \hat f(\omega) = \frac{1}{\sqrt{2\pi}}\,\frac{1}{\alpha + \ii\omega}\ec \qquad \abs{\hat f(\omega)}^{2} = \frac{1}{2\pi}\, \frac{1}{\alpha^{2}+\omega^{2}}\ec \end{equation}

a Lorentzian of full width at half maximum \(\Delta\omega = 2\alpha\), i.e. \(\Delta\nu = \alpha/\pi\). This is the natural line shape of a decaying emitter and the origin of the linewidths measured in Experiment: Precision Spectroscopy and Atomic Clocks.

(iii) Gaussian and Gaussian. For \(f(t) = \ee^{-t^{2}/(2a^{2})}\), Equation (17.38) and Equation (17.39) give \(\ii\omega\hat f = \widehat{f'} = -a^{-2}\widehat{tf} = -\ii a^{-2}\hat f'\), i.e. the first-order equation \(\hat f' = -a^{2}\omega\hat f\) with \(\hat f(0) = a\) from the Gauss integral. Hence

\begin{equation}\tag{17.42} \widehat{\ee^{-t^{2}/(2a^{2})}}(\omega) = a\,\ee^{-a^{2}\omega^{2}/2}\ec \end{equation}

a Gaussian of width \(1/a\) in \(\omega\) for a Gaussian of width \(a\) in \(t\). The Gaussian is thus an eigenfunction of the transform up to scaling, and Theorem 17.43 will show it is the unique minimiser of the bandwidth product. Rests on Definition 17.29 and Proposition 17.30.

Theorem 17.32 (Fourier inversion).

Let \(f\) be absolutely integrable and continuous, and suppose \(\hat f\) is also absolutely integrable. Then for every \(t\)

\begin{equation}\tag{17.43} f(t) = \frac{1}{\sqrt{2\pi}} \int_{-\infty}^{\infty}\hat f(\omega)\,\ee^{\ii\omega t}\,\dd\omega\ep \end{equation}

Rests on Definition 17.29 and Equation (17.42).

Proof.

Derives Theorem 17.32. Fourier stated the integral theorem in 1822 [Fourier:1822]; the proof below is the Gaussian-regularisation argument of Titchmarsh's monograph [Titchmarsh:1937]. For \(\varepsilon>0\) set

\begin{equation*} I_{\varepsilon}(t) = \frac{1}{\sqrt{2\pi}}\int_{-\infty}^{\infty} \hat f(\omega)\,\ee^{-\varepsilon\omega^{2}/2}\, \ee^{\ii\omega t}\,\dd\omega\ep \end{equation*}

Insert Equation (17.32). The double integral converges absolutely, because \(\abs{f(s)}\ee^{-\varepsilon\omega^{2}/2}\) is integrable over the \((s,\omega)\) plane, so Fubini's theorem permits doing the \(\omega\) integral first:

\begin{equation*} I_{\varepsilon}(t) = \int_{-\infty}^{\infty} f(s)\, \left[\frac{1}{2\pi}\int_{-\infty}^{\infty} \ee^{-\varepsilon\omega^{2}/2}\ee^{\ii\omega(t-s)}\dd\omega\right] \dd s = \int_{-\infty}^{\infty} f(s)\,g_{\varepsilon}(t-s)\,\dd s\ec \end{equation*}

where, by the Gaussian pair Equation (17.42) read with \(a^{2} = 1/\varepsilon\),

\begin{equation*} g_{\varepsilon}(u) = \frac{1}{\sqrt{2\pi\varepsilon}}\, \ee^{-u^{2}/(2\varepsilon)}\ec\qquad \int_{-\infty}^{\infty} g_{\varepsilon}(u)\,\dd u = 1\ec\qquad g_{\varepsilon} \ge 0\ep \end{equation*}

The family \(g_{\varepsilon}\) is an approximate identity exactly as in Theorem 17.22: non-negative, of unit mass, and concentrating on \(\abs{u}<\delta\) as \(\varepsilon\to0\), since \(\int_{\abs{u}\ge\delta}g_{\varepsilon} \to 0\). Continuity and boundedness of \(f\) therefore give \(I_{\varepsilon}(t)\to f(t)\). On the other hand \(\abs{\hat f(\omega)\ee^{-\varepsilon\omega^{2}/2}} \le\abs{\hat f(\omega)}\), which is integrable by hypothesis, so dominated convergence gives \(I_{\varepsilon}(t)\to \frac{1}{\sqrt{2\pi}}\int\hat f(\omega)\ee^{\ii\omega t}\dd\omega\). Two limits of one sequence, hence Equation (17.43).

Corollary 17.33 (The transform is injective).

Two continuous absolutely integrable functions with the same transform are equal. Hence a spectrum determines its signal. Rests on Theorem 17.32.

Proof.

Derives Corollary 17.33. Apply Theorem 17.32 to the difference, whose transform vanishes identically and is trivially integrable.

Parseval and Plancherel

Theorem 17.34 (Parseval's identity for series).

Let \(f\) be \(T\)-periodic and square integrable over a period. Then the harmonics are a complete orthonormal system: \(S_N f \to f\) in mean square, and

\begin{equation}\tag{17.44} \frac{1}{T}\int_{-T/2}^{T/2}\abs{f(t)}^{2}\,\dd t = \sum_{n=-\infty}^{\infty}\abs{c_n}^{2} = \frac{\abs{a_0}^{2}}{4} + \frac{1}{2}\sum_{n=1}^{\infty} \left(\abs{a_n}^{2}+\abs{b_n}^{2}\right)\ep \end{equation}

More generally \(\frac{1}{T}\int f\bar g = \sum_n c_n\overline{d_n}\) for the coefficients \(c_n\) of \(f\) and \(d_n\) of \(g\). Rests on Lemma 17.2, Proposition 17.4 and Theorem 17.22.

Proof.

Derives Theorem 17.34. Suppose first that \(f\) is continuous. By Theorem 17.22 the Cesàro means converge uniformly, hence in mean square, and \(\sigma_N f\) is a trigonometric polynomial of degree \(N\); by the least-squares property of Proposition 17.4 the partial sum does at least as well,

\begin{equation*} \frac{1}{T}\int\abs{f - S_N f}^{2} \le \frac{1}{T}\int\abs{f - \sigma_N f}^{2} \le \norm{f - \sigma_N f}_{\infty}^{2}\longrightarrow 0\ep \end{equation*}

Expanding the left-hand side with Lemma 17.2 gives \(\frac{1}{T}\int\abs{f}^{2}-\sum_{\abs{n}\le N}\abs{c_n}^{2}\to0\), which is Equation (17.44). For general square-integrable \(f\), pick continuous \(f_{j}\to f\) in mean square (density, quoted in Remark 17.28); both sides of Equation (17.44) are continuous in that norm — the right side because \(\abs{c_n(f)-c_n(f_j)}\) is controlled by Bessel's inequality Equation (17.10) applied to \(f-f_j\) — so the identity passes to the limit. The bilinear form follows by polarisation, replacing \(f\) by \(f+g\), \(f-g\), \(f+\ii g\), \(f-\ii g\) in turn. The real form is Equation (17.8) substituted term by term.

Parseval stated the identity in 1806 for power series, as a formal consequence of multiplying two series and integrating [Parseval:1806]; its reading as the Pythagorean theorem for an orthonormal system — the length of a vector is the root of the sum of the squares of its components — came later, and is the statement completed in the abstract in Hilbert Spaces. It is the relation left pending at Section 9.4.4.

Lemma 17.35 (Multiplication formula).

If \(f\) and \(g\) are absolutely integrable then

\begin{equation}\tag{17.45} \int_{-\infty}^{\infty}\hat f(\omega)\,g(\omega)\,\dd\omega = \int_{-\infty}^{\infty} f(t)\,\hat g(t)\,\dd t\ep \end{equation}

Rests on Definition 17.29.

Proof.

Derives Lemma 17.35. Both sides equal \(\frac{1}{\sqrt{2\pi}}\iint f(t)g(\omega)\ee^{-\ii\omega t} \dd t\,\dd\omega\), and the double integral converges absolutely since \(\abs{f(t)g(\omega)}\) is integrable on the plane; Fubini's theorem exchanges the order.

Theorem 17.36 (Plancherel).

The Fourier transform, defined by Equation (17.32) on functions that are both absolutely and square integrable, satisfies

\begin{equation}\tag{17.46} \int_{-\infty}^{\infty}\abs{f(t)}^{2}\,\dd t = \int_{-\infty}^{\infty}\abs{\hat f(\omega)}^{2}\,\dd\omega\ec \qquad\text{and}\qquad \int f\bar g\,\dd t = \int \hat f\overline{\hat g}\,\dd\omega\ec \end{equation}

and extends uniquely to a unitary map of the square-integrable functions onto themselves, whose inverse is the extension of Equation (17.34) [Plancherel:1910]. Rests on Definition 17.29 and Theorem 17.32.

Proof.

Derives Theorem 17.36. Take first \(f\) smooth and rapidly decreasing, so that \(f\) and \(\hat f\) are both absolutely integrable and Theorem 17.32 applies. Then, inserting Equation (17.43) and exchanging the order of integration (Fubini, the double integral converging absolutely),

\begin{equation*} \int\abs{f(t)}^{2}\dd t = \int \overline{f(t)}\left[\frac{1}{\sqrt{2\pi}} \int \hat f(\omega)\ee^{\ii\omega t}\dd\omega\right]\dd t = \int\hat f(\omega)\, \overline{\left[\frac{1}{\sqrt{2\pi}}\int f(t)\ee^{-\ii\omega t} \dd t\right]}\,\dd\omega = \int\abs{\hat f(\omega)}^{2}\dd\omega\ep \end{equation*}

So \(\mathcal{F}\) is norm preserving on a dense subspace of the square-integrable functions (density quoted in Remark 17.28). A norm-preserving linear map on a dense subspace of a complete space extends uniquely and continuously: if \(f_j\to f\) then \((f_j)\) is Cauchy, hence so is \((\hat f_j)\) by isometry, hence it converges, and the limit is independent of the approximating sequence because the isometry controls the difference. The extension is again an isometry, and Equation (17.34) extends the same way to a two-sided inverse, since \(\mathcal{F}^{-1}\mathcal{F} = \id\) holds on the dense subspace by Theorem 17.32 and passes to the limit. An invertible isometry of a complex inner-product space is unitary, and the polarisation identity turns the norm statement into the inner-product statement.

The theorem is the relation left pending at Equation (9.123), and it is the single most useful statement in the chapter for physics, because it says that a change to the frequency description conserves the quadratic quantity — energy, probability, power — that the physics is about.

Example 17.37 (Energy spectral density in SI).

Let \(u(t)\), in \(\mathrm{V}\), be the voltage across a resistance \(R\) in \(\mathrm{\Omega}\). The total energy dissipated is

\begin{equation}\tag{17.47} E = \int_{-\infty}^{\infty}\frac{u(t)^{2}}{R}\,\dd t = \frac{1}{R}\int_{-\infty}^{\infty}\abs{\hat u(\omega)}^{2}\, \dd\omega = \frac{2}{R}\int_{0}^{\infty}\abs{\hat u(\omega)}^{2}\,\dd\omega\ec \end{equation}

the last step for real \(u\), where \(\hat u(-\omega) = \overline{\hat u(\omega)}\). Dimensionally \(\hat u\) is in \(\mathrm{V}\,\mathrm{s}\), so \(\abs{\hat u}^{2}/R\) is in \(\mathrm{J}\,\mathrm{s}/\mathrm{rad}\) and its integral over \(\mathrm{rad}/\mathrm{s}\) is in \(\mathrm{J}\): the quantity \(2\abs{\hat u(\omega)}^{2}/R\) is the energy spectral density, telling how many joules sit in each interval of angular frequency. The one-sided form is why laboratory instruments report positive frequencies only, and the factor \(2\) is the commonest bookkeeping error in noise budgets.

Tempered distributions

Equation (17.32) needs \(f\) to be absolutely integrable, and the functions physics most wants to transform are not: a constant, a plane wave \(\ee^{\ii k_{0}x}\), a periodic signal, the Heaviside step. Each has a perfectly definite spectrum in practice — a plane wave is one sharp line — but the integral defining it diverges. Dirac introduced the object that the answer requires, a “function” \(\delta\) with \(\int\delta(x)\varphi(x)\dd x = \varphi(0)\), and used it exactly so [Dirac:1930b]; Schwartz supplied the theory in which the object is legitimate [Schwartz:1950]. The idea is to stop asking what the value of \(\delta\) is at a point and to define it by what it does to test functions.

Definition 17.38 (Schwartz space and tempered distributions).

The Schwartz space \(\mathcal{S}\) consists of the infinitely differentiable \(\varphi : \R\to\C\) such that

\begin{equation}\tag{17.48} p_{m,k}(\varphi) = \sup_{t\in\R}\abs{t}^{m}\abs{\varphi^{(k)}(t)} < \infty \qquad\text{for all } m, k \ge 0\ec \end{equation}

that is, functions decreasing faster than any power together with all their derivatives. A tempered distribution \(T\) is a linear map \(\mathcal{S}\to\C\), written \(\gen{T,\varphi}\), continuous in the sense that \(\abs{\gen{T,\varphi}}\le C\sum_{m,k\le M} p_{m,k}(\varphi)\) for some \(C\) and \(M\). Every function of at most polynomial growth defines one by \(\gen{T_f,\varphi} = \int f\varphi\,\dd t\), and \(\gen{\delta,\varphi} = \varphi(0)\) defines the Dirac distribution.

Proposition 17.39 (The transform preserves $\mathcal{S}$).

If \(\varphi\in\mathcal{S}\) then \(\hat\varphi\in\mathcal{S}\), and \(\mathcal{F}\) is a bijection of \(\mathcal{S}\) with inverse Equation (17.34). Rests on Definition 17.38, Proposition 17.30 and Theorem 17.32.

Proof.

Derives Proposition 17.39. Iterating Equation (17.38) and Equation (17.39), both of which apply because every \(t^{m}\varphi^{(k)}\) is absolutely integrable, gives

\begin{equation}\tag{17.49} \omega^{k}\,\frac{\dd^{m}\hat\varphi}{\dd\omega^{m}}(\omega) = \frac{(-\ii)^{m}}{\ii^{\,k}}\, \widehat{\left[\frac{\dd^{k}}{\dd t^{k}} \left(t^{m}\varphi\right)\right]}(\omega)\ec \end{equation}

each application of Equation (17.39) contributing one factor \(-\ii\) and each application of Equation (17.38) one factor \(\ii\). The right-hand side is the transform of an absolutely integrable function, hence bounded by Proposition 17.30; so every seminorm Equation (17.48) of \(\hat\varphi\) is finite and \(\hat\varphi\in\mathcal{S}\). Since \(\varphi\) and \(\hat\varphi\) are both absolutely integrable, Theorem 17.32 applies and exhibits Equation (17.34) as a two-sided inverse.

Definition 17.40 (Transform and derivative of a distribution).

For \(T\in\mathcal{S}'\) define

\begin{equation}\tag{17.50} \gen{\hat T,\varphi} = \gen{T,\hat\varphi}\ec\qquad \gen{T',\varphi} = -\gen{T,\varphi'}\ec\qquad \varphi\in\mathcal{S}\ep \end{equation}

Both right-hand sides are defined by Proposition 17.39 and by \(\varphi'\in\mathcal{S}\), and both are continuous, so \(\hat T\) and \(T'\) are again tempered distributions. Rests on Definition 17.38 and Proposition 17.39.

The definition is forced, not chosen: for an ordinary absolutely integrable \(f\), Lemma 17.35 says exactly \(\int\hat f\varphi = \int f\hat\varphi\), so Equation (17.50) is the unique extension of the classical transform. The same holds for the derivative, where integration by parts with vanishing boundary terms gives \(\int f'\varphi = -\int f\varphi'\).

Proposition 17.41 (The identities physics uses).

In the sense of Definition 17.40,

\begin{align} \hat\delta &= \frac{1}{\sqrt{2\pi}}\ec &\hat 1 &= \sqrt{2\pi}\,\delta\ec \tag{17.51}\\ \widehat{\ee^{\ii\omega_{0}t}} &= \sqrt{2\pi}\, \delta(\omega - \omega_{0})\ec &\widehat{\cos\omega_{0}t} &= \sqrt{\frac{\pi}{2}} \left[\delta(\omega-\omega_{0}) + \delta(\omega+\omega_{0})\right]\ec \tag{17.52} \end{align}

and, as an identity between distributions in \(t - t'\),

\begin{equation}\tag{17.53} \frac{1}{2\pi}\int_{-\infty}^{\infty} \ee^{\ii\omega(t-t')}\,\dd\omega = \delta(t-t')\ep \end{equation}

Furthermore \(\theta' = \delta\) for the Heaviside step \(\theta\), and \(\widehat{T'} = \ii\omega\hat T\) for every tempered \(T\). Rests on Definition 17.40, Theorem 17.32 and Theorem 7.43.

Proof.

Derives Proposition 17.41. \(\gen{\hat\delta,\varphi} = \gen{\delta,\hat\varphi} = \hat\varphi(0) = \frac{1}{\sqrt{2\pi}}\int\varphi(t)\dd t\), which is the action of the constant \(1/\sqrt{2\pi}\). Next, \(\gen{\hat 1,\varphi} = \int\hat\varphi(\omega)\dd\omega = \sqrt{2\pi}\,\varphi(0)\) by Equation (17.43) evaluated at \(t = 0\), which is the action of \(\sqrt{2\pi}\delta\); this is Equation (17.53) rewritten, since the left side of Equation (17.53) is \(\frac{1}{\sqrt{2\pi}}\hat 1\) evaluated on the difference variable. Equation (17.52) follows by applying the modulation rule Equation (17.36), valid verbatim in the dual pairing, to \(\hat1\), and the cosine case by \(\cos = \frac12(\ee^{\ii\omega_0 t}+\ee^{-\ii\omega_0 t})\). For the step, \(\gen{\theta',\varphi} = -\int_{0}^{\infty}\varphi'(t)\dd t = \varphi(0)\) by the fundamental theorem of calculus (Theorem 7.43) and \(\varphi(\infty)=0\). Finally \(\gen{\widehat{T'},\varphi} = \gen{T',\hat\varphi} = -\gen{T,(\hat\varphi)'} = -\gen{T,\widehat{-\ii t\varphi}} = \gen{\hat T,\ii\omega\varphi}\), using Equation (17.39).

Equation (17.53) is the completeness relation of the continuous plane-wave basis. It is what makes the Green's-function constructions of Section 10.3 legitimate — the fundamental solution of a constant-coefficient operator \(P(\pp_t)\) is \(G(t) = \frac{1}{2\pi}\int \ee^{\ii\omega t}/P(\ii\omega)\,\dd\omega\) precisely because \(P(\pp_t)G = \delta\) reduces to Equation (17.53) — and it is the normalisation of the non-normalisable scattering states of Scattering Theory, whose rigorous home is the rigged triple of Section 12.6.3. Distributional derivatives are likewise what allows a solution of a partial differential equation to have a kink or a shock: that is the weak formulation of Section 10.6, and \(\theta' = \delta\) is its smallest example.

Proposition 17.42 (Sokhotski–Plemelj).

As tempered distributions,

\begin{equation}\tag{17.54} \lim_{\varepsilon\to0^{+}}\frac{1}{x \pm \ii\varepsilon} = \operatorname{PV}\frac{1}{x} \mp \ii\pi\,\delta(x)\ec \end{equation}

where the principal value is \(\gen{\operatorname{PV}\frac1x,\varphi} = \lim_{\eta\to0^{+}}\int_{\abs{x}>\eta}\varphi(x)\,\dd x/x\). Rests on Definition 17.38 and Theorem 17.32.

Proof.

Derives Proposition 17.42. Separate real and imaginary parts, \(\frac{1}{x\pm\ii\varepsilon} = \frac{x}{x^{2}+\varepsilon^{2}} \mp \ii\frac{\varepsilon} {x^{2}+\varepsilon^{2}}\). The imaginary part is \(\pi\) times the Poisson kernel \(\frac{1}{\pi}\frac{\varepsilon}{x^{2}+\varepsilon^{2}}\), which is non-negative, has unit integral — the substitution \(x=\varepsilon s\) gives \(\frac1\pi\int\dd s/(1+s^{2}) = 1\) — and concentrates at the origin, hence is an approximate identity and converges to \(\delta\) as in Theorem 17.32. For the real part, use oddness to write

\begin{equation*} \int\frac{x\,\varphi(x)}{x^{2}+\varepsilon^{2}}\,\dd x = \int_{0}^{\infty}\frac{x\left[\varphi(x)-\varphi(-x)\right]} {x^{2}+\varepsilon^{2}}\,\dd x\ec \end{equation*}

whose integrand is bounded near \(0\) by \(2\norm{\varphi'}_{\infty}x^{2}/(x^{2}+\varepsilon^{2}) \le 2\norm{\varphi'}_{\infty}\), uniformly in \(\varepsilon\); dominated convergence lets \(\varepsilon\to0\) inside, giving \(\int_0^\infty\left[\varphi(x)-\varphi(-x)\right]\dd x/x\), which is the principal value.

Equation (17.54) is the identity behind the causal “\(+\ii\epsilon\)” prescriptions of propagator theory and behind the dispersion relations of Section 17.4.2; it is stated here because it is a statement about distributions and nothing else.

The uncertainty principle

A function and its transform cannot both be sharply localised: this is a theorem about integrals, provable in a page, and it acquires physical content only when the transform variable is given a physical meaning. Both halves are given below, in that order, because conflating them is the commonest confusion in the subject.

Theorem 17.43 (Bandwidth theorem).

Let \(f\) be square integrable with \(\norm{f}_{2}^{2} = \int\abs{f}^{2}\dd t = 1\), and suppose \(tf(t)\) and \(f'\) are square integrable. Define the centroids and dispersions

\begin{equation}\tag{17.55} \avg{t} = \int t\abs{f(t)}^{2}\dd t\ec\quad \sigma_{t}^{2} = \int (t-\avg{t})^{2}\abs{f(t)}^{2}\dd t\ec\quad \sigma_{\omega}^{2} = \int(\omega - \avg{\omega})^{2} \abs{\hat f(\omega)}^{2}\dd\omega\ec \end{equation}

with \(\avg{\omega}\) defined from \(\abs{\hat f}^{2}\) in the same way. Then

\begin{equation}\tag{17.56} \sigma_{t}\,\sigma_{\omega} \ge \frac{1}{2}\ec \end{equation}

with equality if and only if \(f\) is a Gaussian, \(f(t) = C\ee^{\ii\avg{\omega}t}\ee^{-(t-\avg{t})^{2}/(4\sigma_t^{2})}\). Rests on Proposition 17.30, Theorem 17.36 and Equation (17.42).

Derivation. Derives Theorem 17.43. Translating \(t\) and modulating by \(\ee^{-\ii\avg{\omega}t}\) changes neither \(\sigma_t\) nor \(\sigma_\omega\), by Equation (17.35) and Equation (17.36), so assume \(\avg{t} = \avg{\omega} = 0\). By Equation (17.38) and Plancherel Equation (17.46),

\begin{equation*} \sigma_{\omega}^{2} = \int\omega^{2}\abs{\hat f}^{2}\dd\omega = \int\abs{\widehat{f'}}^{2}\dd\omega = \int\abs{f'(t)}^{2}\dd t\ep \end{equation*}

Now integrate \(\dv{}{t}\left(t\abs{f}^{2}\right) = \abs{f}^{2} + t\left(\bar f f' + f\bar f'\right) = \abs{f}^{2} + 2t\,\Re\!\left(\bar f f'\right)\) over the line. The left-hand side integrates to zero, because \(t\abs{f(t)}^{2}\to0\) at infinity when \(tf\) and \(f\) are square integrable, so

\begin{equation*} 1 = \int\abs{f}^{2}\dd t = -2\int t\,\Re\!\left(\bar f(t) f'(t)\right)\dd t \le 2\int\abs{t f(t)}\,\abs{f'(t)}\,\dd t \le 2\,\sigma_{t}\,\sigma_{\omega}\ec \end{equation*}

the last step by the Cauchy–Schwarz inequality applied to \(tf\) and \(f'\). That is Equation (17.56).

Equality requires equality in Cauchy–Schwarz, i.e. \(f'(t) = \lambda\, t\,f(t)\) for a constant \(\lambda\), and equality in \(\Re(\bar ff')\ge-\abs{\bar f f'}\), which forces \(\lambda\) real and negative. Integrating the first-order equation gives \(f(t) = C\ee^{\lambda t^{2}/2}\), a Gaussian; matching \(\sigma_t^{2} = -1/(2\lambda)\) from Equation (17.42) and undoing the translation and modulation yields the stated form.

Corollary 17.44 (Duration and bandwidth in SI).

With \(t\) in \(\mathrm{s}\) and \(\nu = \omega/2\pi\) in \(\mathrm{Hz}\), Equation (17.56) reads

\begin{equation}\tag{17.57} \sigma_{t}\,\sigma_{\nu} \ge \frac{1}{4\pi} \approx 0.0796\ep \end{equation}

A transform-limited optical pulse of \(\sigma_t = 10\,\mathrm{fs}\) therefore has a spectral width of at least \(\sigma_{\nu} = 7.96\,\mathrm{THz}\); at a centre wavelength of \(800\,\mathrm{nm}\), where \(\nu_0 = c/\lambda_0 = 375\,\mathrm{THz}\), that is about \(2\,\mathrm{\%}\) of the carrier frequency, so such a pulse cannot be called monochromatic. Rests on Theorem 17.43 and Equation (17.37).

Proof.

Derives Corollary 17.44. \(\sigma_{\omega} = 2\pi\sigma_{\nu}\) by Equation (17.37) with \(\lambda = 2\pi\); substitute in Equation (17.56). The numbers are \(1/(4\pi\times10^{-14}\,\mathrm{s}) = 7.96\times 10^{12}\,\mathrm{Hz}\) and \(c/\lambda_{0} = 2.998\times 10^{8}\,\mathrm{m}/\mathrm{s}/ 800\times 10^{-9}\,\mathrm{m}\).

Example 17.45 (Two purely classical consequences).

Natural linewidth. An emitter whose amplitude decays as \(\ee^{-t/(2\tau)}\) has, by Equation (17.41) with \(\alpha = 1/(2\tau)\), a Lorentzian spectral profile of full width \(\Delta\nu = 1/(2\pi\tau)\). For an excited state of lifetime \(\tau = 16\,\mathrm{ns}\) — the order of an allowed optical transition — this is \(\Delta\nu = 9.9\,\mathrm{MHz}\), the irreducible floor under every measurement in Experiment: Precision Spectroscopy and Atomic Clocks.

Diffraction. A slit of width \(a\) multiplies an incident plane wave by a rectangle, so by Equation (17.40) the transmitted amplitude as a function of transverse wavenumber \(k_{x}\) is a sinc with first zero at \(k_{x} = 2\pi/a\). Since \(k_{x} = k\sin\theta\) with \(k = 2\pi/\lambda\), the first dark fringe is at \(\sin\theta = \lambda/a\): single-slit diffraction is the rectangle–sinc pair and nothing more. For sodium light, \(\lambda = 589\,\mathrm{nm}\), through \(a = 1.0\,\mathrm{mm}\), \(\theta \approx 5.9\times 10^{-4}\,\mathrm{rad}\). The optics is developed in Electromagnetic Waves and Optics.

Remark 17.46 (The quantum reading, in SI).

Wave mechanics assigns to a particle of momentum \(p\) the wavenumber \(k = p/\hbar\), with \(\hbar = 1.054571817\times 10^{-34}\,\mathrm{J}\,\mathrm{s}\) exactly by the definition of the SI (Matter Waves). If \(\psi(x)\) is a normalised position amplitude and \(\tilde\psi(k)\) its transform, then the momentum amplitude is \(\tilde\psi(p/\hbar)/\sqrt{\hbar}\), normalised in \(p\), and \(\sigma_{p} = \hbar\,\sigma_{k}\). Substituting into Equation (17.56) with the pair \((x,k)\) gives

\begin{equation}\tag{17.58} \sigma_{x}\,\sigma_{p} \ge \frac{\hbar}{2}\ec \end{equation}

which is Kennard's inequality [Kennard:1927], the sharp form of the uncertainty relation that Heisenberg had argued physically from the disturbance caused by a measurement [Heisenberg:1927]. The mathematical content is entirely Theorem 17.43; what physics supplies is the single constant \(\hbar\) converting an inverse length into a momentum. For an electron localised to \(\sigma_{x} = 10^{-10}\,\mathrm{m}\), an atomic diameter, Equation (17.58) gives \(\sigma_{p}\ge5.27\times 10^{-25}\,\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), hence a kinetic-energy scale \(\sigma_p^{2}/(2m_{\mathrm{e}}) \approx 1.5\times 10^{-19}\,\mathrm{J} \approx 0.95\,\mathrm{eV}\) — the electronvolt scale of atomic physics, obtained from a theorem about Fourier integrals. The operator form of the inequality, valid for any pair of non-commuting observables, belongs to The Postulates of Quantum Mechanics. Rests on Theorem 17.43.

Poisson summation

Sampling a function and sampling its transform are not independent operations. The exact relation between them is one formula, and it controls both the lattice sums of solid-state physics and the aliasing of every digitised measurement.

Theorem 17.47 (Poisson summation).

Let \(f\) be continuous with \(\abs{f(t)}\le C(1+\abs{t})^{-1-\eta}\) and \(\abs{\hat f(\omega)}\le C(1+\abs{\omega})^{-1-\eta}\) for some \(\eta, C>0\). Then for every \(T>0\) and every \(t\),

\begin{equation}\tag{17.59} \sum_{n=-\infty}^{\infty} f(t + nT) = \frac{\sqrt{2\pi}}{T}\sum_{m=-\infty}^{\infty} \hat f\!\left(\frac{2\pi m}{T}\right)\, \ee^{2\pi\ii m t/T}\ec \end{equation}

and in particular, at \(t = 0\),

\begin{equation}\tag{17.60} \sum_{n=-\infty}^{\infty} f(nT) = \frac{\sqrt{2\pi}}{T}\sum_{m=-\infty}^{\infty} \hat f\!\left(\frac{2\pi m}{T}\right)\ep \end{equation}

Rests on Definition 17.1, Definition 17.29 and Theorem 17.9.

Proof.

Derives Theorem 17.47. The periodisation \(F(t) = \sum_{n} f(t+nT)\) converges absolutely and uniformly on compact sets by the decay hypothesis (comparison with \(\sum n^{-1-\eta}\)), so \(F\) is continuous, and it is \(T\)-periodic by construction, the sum being over all of \(\Z\). Its Fourier coefficients in the sense of Equation (17.3) are

\begin{equation*} C_m = \frac{1}{T}\int_{0}^{T} F(t)\,\ee^{-2\pi\ii m t/T}\,\dd t = \frac{1}{T}\sum_{n}\int_{0}^{T} f(t+nT)\ee^{-2\pi\ii m t/T}\dd t = \frac{1}{T}\int_{-\infty}^{\infty} f(s)\,\ee^{-2\pi\ii m s/T}\dd s\ec \end{equation*}

where uniform convergence justifies the exchange and the substitution \(s = t+nT\) reassembles the intervals \([nT,(n+1)T]\) into the line, the exponential being unchanged because \(\ee^{-2\pi\ii m n} = 1\). The last integral is \(\sqrt{2\pi}\,\hat f(2\pi m/T)/T\). By hypothesis \(\sum_m\abs{C_m}<\infty\), so the Fourier series of \(F\) converges absolutely and uniformly, and its sum is \(F\) by Theorem 17.9 — or by Corollary 17.23(b) applied to the difference. Writing that statement out is Equation (17.59).

Example 17.48 (Theta identity).

Take \(f(t) = \ee^{-\pi s t^{2}}\) with \(s>0\). By Equation (17.42) with \(a^{2} = 1/(2\pi s)\), \(\hat f(\omega) = (2\pi s)^{-1/2}\ee^{-\omega^{2}/(4\pi s)}\). Put \(T = 1\) in Equation (17.60):

\begin{equation}\tag{17.61} \sum_{n=-\infty}^{\infty}\ee^{-\pi s n^{2}} = \frac{1}{\sqrt{s}}\sum_{m=-\infty}^{\infty}\ee^{-\pi m^{2}/s}\ep \end{equation}

The identity converts a slowly convergent sum at small \(s\) into a rapidly convergent one, which is how such lattice sums are evaluated in practice; it is also the functional equation from which the analytic continuation of the Riemann zeta function is obtained.

Corollary 17.49 (Lattice sums and the reciprocal lattice).

Let \(\Lambda\subset\R^{3}\) be a Bravais lattice with primitive vectors \(\vect{a}_{1},\vect{a}_{2},\vect{a}_{3}\) and cell volume \(V = \vect{a}_{1}\cdot(\vect{a}_{2}\times\vect{a}_{3})\), and let \(\Lambda^{*}\) be the reciprocal lattice generated by the \(\vect{b}_{i}\) with \(\vect{a}_{i}\cdot\vect{b}_{j} = 2\pi\delta_{ij}\). Then, for \(f\) decaying as above in each variable,

\begin{equation}\tag{17.62} \sum_{\vect{R}\in\Lambda} f(\vect{r}+\vect{R}) = \frac{(2\pi)^{3/2}}{V}\sum_{\vect{G}\in\Lambda^{*}} \hat f(\vect{G})\,\ee^{\ii\vect{G}\cdot\vect{r}}\ec \end{equation}

with \(\hat f\) the three-dimensional transform, the convention of Definition 17.29 applied in each Cartesian coordinate. Rests on Theorem 17.47 and Definition 17.29.

Proof.

Derives Corollary 17.49. Change variables to the coordinates \(u_{i}\) defined by \(\vect{r} = \sum_i u_i\vect{a}_i\), in which \(\Lambda\) becomes \(\Z^{3}\) and the Jacobian is \(V\); Theorem 17.47 applies in each of the three variables in turn, and the dual variables assemble into \(\Lambda^{*}\) by the defining relation \(\vect{a}_{i}\cdot\vect{b}_{j} = 2\pi\delta_{ij}\).

For copper, a face-centred cubic crystal with conventional lattice parameter \(a = 0.3615\,\mathrm{nm}\), the reciprocal-lattice scale is \(2\pi/a = 17.4\,/\mathrm{nm}\), which is the size of the Brillouin zone that organises the phonon spectra of Phonons and Lattice Dynamics and the band structure of Electrons in Solids: Band Theory. Equation (17.62) is also the statement that a periodic array of scatterers diffracts only into the reciprocal-lattice directions, which is the Laue condition.

Finally, Equation (17.59) in the frequency variable is the exact statement of what sampling does to a spectrum, and Section 17.5.1 uses it in that form.

Convolution and correlation

The convolution theorem

Definition 17.50 (Convolution).

For \(f, g\) absolutely integrable on \(\R\),

\begin{equation}\tag{17.63} (f * g)(t) = \int_{-\infty}^{\infty} f(s)\,g(t-s)\,\dd s\ep \end{equation}
Proposition 17.51 (Algebra of convolution).

Convolution is bilinear, commutative and associative, it commutes with translations, and

\begin{equation}\tag{17.64} \int_{-\infty}^{\infty}\abs{(f*g)(t)}\,\dd t \le \left(\int\abs{f}\right)\left(\int\abs{g}\right)\ec \end{equation}

so \(f*g\) is again absolutely integrable and is defined for almost every \(t\). Rests on Definition 17.50.

Proof.

Derives Proposition 17.51. Commutativity is the substitution \(s\mapsto t-s\); associativity is a change of variables in the resulting double integral. For Equation (17.64), Tonelli's theorem allows the non-negative double integral to be done in either order:

\begin{equation*} \iint\abs{f(s)}\abs{g(t-s)}\,\dd s\,\dd t = \int\abs{f(s)}\left[\int\abs{g(t-s)}\,\dd t\right]\dd s = \left(\int\abs{f}\right)\left(\int\abs{g}\right)\ec \end{equation*}

finite, which also shows the inner integral in Equation (17.63) converges for almost every \(t\).

Proposition 17.52 (Linear translation-invariant systems).

Let \(\mathcal{A}\) be a linear map on signals that commutes with time translation, \(\mathcal{A}\left[x(\cdot - a)\right] = \left(\mathcal{A}x\right)(\cdot - a)\) for every \(a\), and that is continuous in the sense of Definition 17.38. Then \(\mathcal{A}\) is a convolution,

\begin{equation}\tag{17.65} \left(\mathcal{A}x\right)(t) = \int_{-\infty}^{\infty} h(t-s)\,x(s)\,\dd s = (h*x)(t)\ec\qquad h = \mathcal{A}\delta\ec \end{equation}

with \(h\) the impulse response of the system. Rests on Definition 17.50, Definition 17.38 and Proposition 17.41.

Derivation. Derives Proposition 17.52. By Proposition 17.41 the sifting property writes any input as a superposition of shifted impulses, \(x(t) = \int x(s)\,\delta(t-s)\,\dd s\). Apply \(\mathcal{A}\) and exchange it with the integral, which continuity permits, treating the integral as a limit of finite sums: \(\left(\mathcal{A}x\right)(t) = \int x(s)\, \mathcal{A}\left[\delta(\cdot - s)\right](t)\,\dd s\). Translation invariance identifies \(\mathcal{A}\left[\delta(\cdot-s)\right](t) = h(t-s)\) with \(h = \mathcal{A}\delta\), which is Equation (17.65). The statement made fully rigorous — that every continuous linear map between distribution spaces has a kernel, here forced by translation invariance to depend only on \(t-s\) — is Schwartz's kernel theorem [Schwartz:1950].

Every linear system whose behaviour does not depend on when it is used is therefore a convolution: a filter, an amplifier below saturation, a spectrometer, a lens, the elastic response of a solid. What such a system does to a spectrum is now one line.

Theorem 17.53 (Convolution theorem).

For absolutely integrable \(f, g\),

\begin{equation}\tag{17.66} \widehat{f * g}(\omega) = \sqrt{2\pi}\,\hat f(\omega)\,\hat g(\omega)\ec \end{equation}

and, when both sides are defined,

\begin{equation}\tag{17.67} \widehat{f\,g}(\omega) = \frac{1}{\sqrt{2\pi}}\left(\hat f * \hat g\right)(\omega)\ep \end{equation}

Rests on Definition 17.50, Definition 17.29, Proposition 17.51 and Theorem 17.32.

Proof.

Derives Theorem 17.53. Insert Equation (17.63) into Equation (17.32). The double integral converges absolutely by Equation (17.64), so Fubini's theorem permits exchanging the order:

\begin{equation*} \widehat{f*g}(\omega) = \frac{1}{\sqrt{2\pi}}\int\!\!\int f(s)\,g(t-s)\,\ee^{-\ii\omega t} \,\dd s\,\dd t = \frac{1}{\sqrt{2\pi}}\int f(s)\,\ee^{-\ii\omega s} \left[\int g(u)\,\ee^{-\ii\omega u}\,\dd u\right]\dd s\ec \end{equation*}

substituting \(u = t - s\) in the inner integral and factoring \(\ee^{-\ii\omega t} = \ee^{-\ii\omega s}\ee^{-\ii\omega u}\). The bracket is \(\sqrt{2\pi}\,\hat g(\omega)\) and the remaining integral \(\sqrt{2\pi}\,\hat f(\omega)\), giving Equation (17.66). Equation (17.67) is the same computation run through the inverse transform, or equivalently Equation (17.66) applied to \(\hat f\) and \(\hat g\) and inverted with Theorem 17.32.

This is the pending convolution property of Section 9.7.5, and it is the reason integral transforms solve differential equations. A linear constant-coefficient equation \(P(\pp_t)u = w\) has, by Equation (17.38), \(P(\ii\omega)\hat u = \hat w\), hence \(\hat u = \hat w/P(\ii\omega) = \sqrt{2\pi}\,\hat G\hat w\) with \(\hat G = 1/\left[\sqrt{2\pi}P(\ii\omega)\right]\), and therefore \(u = G * w\): every Green's-function solution of Section 10.3 is a convolution of the source with the fundamental solution, and the Green's function is the impulse response of Equation (17.65). The same statement read backwards is the working fact of experimental spectroscopy.

Example 17.54 (Instrument response and widths in quadrature).

Let \(g_{a}(t) = \left(a\sqrt{2\pi}\right)^{-1}\ee^{-t^{2}/(2a^{2})}\) be the normalised Gaussian of standard deviation \(a\), so that \(\hat g_{a}(\omega) = \left(2\pi\right)^{-1/2}\ee^{-a^{2}\omega^{2}/2}\) by Equation (17.42). Then Equation (17.66) gives

\begin{equation}\tag{17.68} \widehat{g_{a}*g_{b}}(\omega) = \sqrt{2\pi}\,\hat g_{a}\hat g_{b} = \frac{1}{\sqrt{2\pi}}\,\ee^{-(a^{2}+b^{2})\omega^{2}/2} = \hat g_{\sqrt{a^{2}+b^{2}}}(\omega)\ec \end{equation}

so \(g_{a}*g_{b} = g_{\sqrt{a^{2}+b^{2}}}\) by injectivity (Corollary 17.33): Gaussian widths add in quadrature, and so do the full widths at half maximum, which are \(2\sqrt{2\ln 2}\) times the standard deviations. A spectral line of true Doppler width \(1.20\,\mathrm{GHz}\) recorded through an instrument of Gaussian response \(0.90\,\mathrm{GHz}\) is therefore observed with width \(\sqrt{1.20^{2}+0.90^{2}} = 1.50\,\mathrm{GHz}\), and the physical width is recovered by subtracting in quadrature. Every quoted linewidth in Experiment: Precision Spectroscopy and Atomic Clocks is a measured width so corrected.

Remark 17.55 (Deconvolution is ill posed).

Inverting Equation (17.66) suggests recovering the true signal from the record \(y = h*f\) by \(\hat f = \hat y/\left(\sqrt{2\pi}\,\hat h\right)\). The formula is correct and the procedure is unusable as it stands. A physical response \(\hat h\) decays at high frequency — for the Gaussian of Example 17.54, as \(\ee^{-a^{2}\omega^{2}/2}\) — so the division multiplies whatever noise the record carries at frequency \(\omega\) by \(\ee^{+a^{2}\omega^{2}/2}\). An arbitrarily small perturbation of the data therefore produces an arbitrarily large perturbation of the answer: the solution does not depend continuously on the data, which is exactly the third of Hadamard's conditions for a well-posed problem (Partial Differential Equations). Practical deconvolution restores well-posedness by giving something up — truncating the division at the frequency where the signal-to-noise ratio reaches unity, or adding a penalty term that suppresses large high-frequency content. The same arithmetic returns in Section 17.6.3, where the tomographic ramp filter amplifies high frequencies for the same reason.

Correlation and the Wiener–Khinchin theorem

Definition 17.56 (Correlation).

For square-integrable \(f, g\) the cross-correlation and autocorrelation are

\begin{equation}\tag{17.69} R_{fg}(\tau) = \int_{-\infty}^{\infty}\overline{f(t)}\,g(t+\tau)\,\dd t\ec \qquad R_{ff}(\tau) = \int_{-\infty}^{\infty} \overline{f(t)}\,f(t+\tau)\,\dd t\ep \end{equation}

Correlation differs from convolution only by a reflection and a conjugation, \(R_{fg} = \bar f^{-}*g\) with \(f^{-}(t) = f(-t)\), and that one difference is what makes it measure similarity at a lag rather than smoothing.

Theorem 17.57 (Wiener–Khinchin, finite-energy form).

For square-integrable \(f\),

\begin{equation}\tag{17.70} \widehat{R_{ff}}(\omega) = \sqrt{2\pi}\,\abs{\hat f(\omega)}^{2}\ec \end{equation}

so the autocorrelation and the energy spectral density are a Fourier pair. In particular \(R_{ff}(0) = \int\abs{f}^{2}\) recovers Plancherel's theorem, and \(\abs{R_{ff}(\tau)}\le R_{ff}(0)\). Rests on Definition 17.56 and Theorem 17.53.

Proof.

Derives Theorem 17.57. With \(f^{-}(t) = f(-t)\) one has \(\widehat{\overline{f^{-}}}(\omega) = \overline{\hat f(\omega)}\), since conjugating and reflecting the argument each conjugate the analysis kernel. Then \(R_{ff} = \overline{f^{-}}*f\) by inspection of Equation (17.69), and Equation (17.66) gives Equation (17.70). Evaluating the inverse transform at \(\tau = 0\) gives \(\int\abs{f}^{2}=\int\abs{\hat f}^{2}\), and the bound on \(\abs{R_{ff}}\) is Cauchy–Schwarz applied to Equation (17.69).

The physically important version concerns signals that do not decay — noise, thermal fluctuations, a stationary source — which have infinite energy but finite power. The right object is then a statistical one, and the apparatus is that of Probability and Statistics.

Definition 17.58 (Stationary process and power spectral density).

A random process \(X(t)\) is stationary in the wide sense if \(\avg{X(t)}\) is independent of \(t\) — taken to be zero below — and the autocorrelation

\begin{equation}\tag{17.71} R(\tau) = \avg{\overline{X(t)}\,X(t+\tau)} \end{equation}

depends only on the lag \(\tau\), the angle brackets denoting the ensemble average. Its power spectral density is

\begin{equation}\tag{17.72} S(\omega) = \lim_{T\to\infty}\frac{1}{T}\, \avg{\abs{\hat X_{T}(\omega)}^{2}}\ec\qquad X_{T} = X\,\indic{\abs{t}\le T/2}\ep \end{equation}

Rests on Definitions 17.29 and 17.56.

Theorem 17.59 (Wiener–Khinchin).

If \(R\) is absolutely integrable then the limit Equation (17.72) exists and

\begin{equation}\tag{17.73} S(\omega) = \frac{1}{\sqrt{2\pi}}\,\hat R(\omega) = \frac{1}{2\pi}\int_{-\infty}^{\infty} R(\tau)\, \ee^{-\ii\omega\tau}\,\dd\tau\ec \end{equation}

and conversely \(R(\tau) = \int S(\omega)\ee^{\ii\omega\tau}\dd\omega\). The total power is \(R(0) = \int S(\omega)\,\dd\omega\). Rests on Definition 17.58, Definition 17.29 and Theorem 17.32.

Derivation. Derives Theorem 17.59. Expand the definition, using Equation (17.32) on the truncated process:

\begin{equation*} \avg{\abs{\hat X_{T}(\omega)}^{2}} = \frac{1}{2\pi}\int_{-T/2}^{T/2}\!\!\int_{-T/2}^{T/2} \avg{\overline{X(t)}X(t')}\,\ee^{-\ii\omega(t'-t)}\,\dd t\,\dd t' = \frac{1}{2\pi}\int_{-T/2}^{T/2}\!\!\int_{-T/2}^{T/2} R(t'-t)\,\ee^{-\ii\omega(t'-t)}\,\dd t\,\dd t'\ec \end{equation*}

the ensemble average passing inside the integrals by stationarity and integrability. Change variables to \(\tau = t'-t\) and \(t\); for fixed \(\tau\) with \(\abs{\tau}\le T\) the admissible \(t\) form an interval of length \(T-\abs{\tau}\), so

\begin{equation*} \frac{1}{T}\avg{\abs{\hat X_{T}(\omega)}^{2}} = \frac{1}{2\pi}\int_{-T}^{T} \left(1 - \frac{\abs{\tau}}{T}\right) R(\tau)\, \ee^{-\ii\omega\tau}\,\dd\tau\ep \end{equation*}

The triangular factor is bounded by \(1\) and tends to \(1\) pointwise, and \(R\) is absolutely integrable, so dominated convergence gives the limit Equation (17.73) as \(T\to\infty\). The converse statement is Theorem 17.32 applied to \(\hat R\), and setting \(\tau = 0\) gives the power.

Wiener obtained the result as part of a general harmonic analysis of non-decaying functions, in which the correlation replaces the divergent transform [Wiener:1930]; Khinchin gave the probabilistic formulation for stationary processes used above [Khinchin:1934].

Example 17.60 (Noise spectral densities in SI).

For a voltage noise \(V(t)\) in \(\mathrm{V}\), \(R(\tau)\) is in \(\mathrm{V}^{2}\) and, by Equation (17.73), \(S(\omega)\) is in \(\mathrm{V}^{2}\,\mathrm{s}/\mathrm{rad}\). Laboratory practice quotes instead the one-sided density per unit ordinary frequency, \(S_{V}(\nu) = 4\pi S(2\pi\nu)\) in \(\mathrm{V}^{2}/\mathrm{Hz}\), so that \(\avg{V^{2}} = \int_{0}^{\infty}S_{V}(\nu)\,\dd\nu\) with no further factors. The thermal noise of a resistor is white in this measure, with \(S_{V} = 4k_{\mathrm{B}}TR\) — Nyquist's analysis [Nyquist:1928b], derived from the fluctuation–dissipation theorem in Nonequilibrium Thermodynamics and Transport, and quoted here only to fix the units. For \(R = 1.0\,\mathrm{k\Omega}\) at \(T = 300\,\mathrm{K}\), \(S_{V} = 1.66\times 10^{-17}\,\mathrm{V}^{2}/\mathrm{Hz}\), i.e. an amplitude density \(\sqrt{S_{V}} = 4.1\,\mathrm{nV}\) per root hertz, and in a measurement bandwidth of \(100\,\mathrm{kHz}\) the root-mean-square noise is \(\sqrt{S_{V}B} = 1.3\,\mu\mathrm{V}\). The whole calculation is Theorem 17.59 plus arithmetic.

One definite integral is worth isolating, because every estimate of the thermal noise floor of a mechanical resonator in this book reduces to it and each would otherwise redo the same contour.

Proposition 17.61 (Mean-square response of a damped resonator).

Let \(m>0\), \(b>0\), \(k>0\) and let

\begin{equation}\tag{17.74} \chi(\omega)=\frac{1}{k-m\omega^{2}+\ii b\omega} \end{equation}

be the response of the damped oscillator, with \(m\) in \(\mathrm{kg}\), \(b\) in \(\mathrm{kg}/\mathrm{s}\), \(k\) in \(\mathrm{N}/\mathrm{m}\) and \(\chi\) in \(\mathrm{m}/\mathrm{N}\). Then

\begin{equation}\tag{17.75} \int_{-\infty}^{\infty}\abs{\chi(\omega)}^{2}\,\dd\omega =\frac{\pi}{bk}\ec\qquad\text{equivalently}\qquad \int_{0}^{\infty}\abs{\chi(2\pi\nu)}^{2}\,\dd\nu=\frac{1}{4bk}\ep \end{equation}

Consequently, by Proposition 17.52 and Theorem 17.59, a resonator driven by a stationary force whose one-sided spectral density \(S_{F}\) in \(\mathrm{N}^{2}/\mathrm{Hz}\) is constant over the band where \(\chi\) is appreciable has mean-square displacement \(\avg{x^{2}}=S_{F}/(4bk)\). Rests on Theorem 8.24, Proposition 17.52 and Theorem 17.59.

Proof.

Derives Proposition 17.61. Take first \(b^{2}\neq4mk\). For real \(\omega\), \(\abs{\chi(\omega)}^{2}=\chi(\omega)\,\tilde\chi(\omega)\) with \(\tilde\chi(\omega)=(k-m\omega^{2}-\ii b\omega)^{-1}\), and both extend to rational functions of a complex \(\omega\) decaying like \(\abs{\omega}^{-4}\), so the arc at infinity contributes nothing and Theorem 8.24 applies to a contour closed in the upper half-plane. The poles of \(\chi\) are the roots of \(m\omega^{2}-\ii b\omega-k=0\),

\begin{equation}\tag{17.76} z_{\pm}=\frac{\ii b\pm\Delta}{2m}\ec\qquad \Delta=\sqrt{4mk-b^{2}}\ec \end{equation}

and both lie in the upper half-plane: for \(4mk>b^{2}\) each has imaginary part \(b/2m>0\), and for \(4mk<b^{2}\) both are purely imaginary, \(\ii\left[b\pm\sqrt{b^{2}-4mk}\right]/2m\), with positive imaginary part because \(b>\sqrt{b^{2}-4mk}\). The poles of \(\tilde\chi\) are \(-z_{\pm}\), in the lower half-plane, so only \(z_{\pm}\) are enclosed. Writing \(\chi(\omega)=-\left[m(\omega-z_{+})(\omega-z_{-})\right]^{-1}\) and using \(k-mz_{\pm}^{2}=-\ii bz_{\pm}\), which is the equation the roots satisfy, gives \(\tilde\chi(z_{\pm})=-(2\ii bz_{\pm})^{-1}\) and hence, by Proposition 8.23,

\begin{equation}\tag{17.77} \operatorname{Res}_{\omega=z_{+}}\chi\tilde\chi =\frac{1}{2\ii b z_{+}\Delta}\ec\qquad \operatorname{Res}_{\omega=z_{-}}\chi\tilde\chi =-\frac{1}{2\ii b z_{-}\Delta}\ec \end{equation}

since \(z_{+}-z_{-}=\Delta/m\). Their sum is \((2\ii b\Delta)^{-1}(z_{-}-z_{+})/(z_{+}z_{-})\), and the root relations \(z_{+}z_{-}=-k/m\) and \(z_{-}-z_{+}=-\Delta/m\) reduce it to \((2\ii bk)^{-1}\). Multiplying by \(2\pi\ii\) gives the first equality of Equation (17.75); the second follows by substituting \(\omega=2\pi\nu\) and halving, the integrand being even in \(\omega\). Both sides of Equation (17.75) are continuous in \(b\) at \(b^{2}=4mk\), which disposes of the critically damped case.

Example 17.62 (Noise floor of a torsion balance).

A torsion pendulum of stiffness \(k=2.0\times 10^{-8}\,\mathrm{N}\,\mathrm{m}/\mathrm{rad}\) and damping \(b=10^{-11}\,\mathrm{N}\,\mathrm{m}\,\mathrm{s}/\mathrm{rad}\) — the rotational analogues of the constants above, with angle in place of displacement — driven by the thermal torque whose one-sided density is \(S_{\tau}=4k_{\mathrm{B}}Tb\) at \(T=300\,\mathrm{K}\), has mean-square angle \(\avg{\theta^{2}}=S_{\tau}/(4bk)=k_{\mathrm{B}}T/k\) by Equation (17.75). The damping cancels: the answer is \(2.1\times 10^{-13}\,\mathrm{rad}^{2}\), an angular fluctuation of \(0.46\,\mu\mathrm{rad}\), and it is equipartition (Statistical Mechanics) recovered from the spectrum. That the cancellation happens is the content of the proposition; that it must happen is why the fluctuation–dissipation relation takes the form it does in Nonequilibrium Thermodynamics and Transport.

Theorem 17.63 (Matched filter).

Let the record be \(d(t) = s(t) + n(t)\), with \(s\) a known signal and \(n\) stationary noise of strictly positive power spectral density \(S\). Among all linear filters \(q\), the signal-to-noise ratio

\begin{equation}\tag{17.78} \rho^{2} = \frac{\abs{\int \overline{\hat q(\omega)}\, \hat s(\omega)\,\dd\omega}^{2}} {\int S(\omega)\,\abs{\hat q(\omega)}^{2}\,\dd\omega} \end{equation}

of the filtered output at the arrival time is maximised by

\begin{equation}\tag{17.79} \hat q(\omega) = \frac{\hat s(\omega)}{S(\omega)}\ec\qquad\text{giving} \qquad \rho^{2}_{\max} = \int_{-\infty}^{\infty} \frac{\abs{\hat s(\omega)}^{2}}{S(\omega)}\,\dd\omega\ep \end{equation}

Rests on Definition 17.58.

Proof.

Derives Theorem 17.63. Write \(\hat q = \hat u/\sqrt{S}\) and \(\hat\sigma = \hat s/\sqrt{S}\), legitimate since \(S>0\). The numerator of Equation (17.78) is \(\abs{\int\overline{\hat u}\,\hat\sigma\,\dd\omega}^{2}\) and the denominator is \(\int\abs{\hat u}^{2}\dd\omega\), so \(\rho^{2} = \abs{\avg{\hat u,\hat\sigma}}^{2}/\norm{\hat u}^{2} \le \norm{\hat\sigma}^{2}\) by the Cauchy–Schwarz inequality, with equality if and only if \(\hat u\) is a multiple of \(\hat\sigma\), i.e.\ \(\hat q \propto \hat s/S\). Substituting gives \(\rho^{2}_{\max} = \norm{\hat\sigma}^{2}\), which is Equation (17.79).

The optimal filter is therefore the signal itself, weighted inversely by the noise power at each frequency: one trusts the record where the detector is quiet and discounts it where the detector is loud. This is the statistic by which the gravitational-wave strain records of Experiment: Gravitational Waves are searched, and Equation (17.79) is the integral whose value is quoted as the detection significance. The same theorem, read in the frequency domain alone, is why the noise curve of an interferometer — a plot of \(\sqrt{S}\) in \(\mathrm{Hz}^{-1/2}\) against \(\nu\) in \(\mathrm{Hz}\) — is the complete specification of what the instrument can find.

Remark 17.64 (Spectroscopy by autocorrelation).

A two-beam interferometer records, as a function of the path difference \(x\), the intensity \(I(x) \propto R_{EE}(x/c)\) of the autocorrelation of the field. By Theorem 17.57 its Fourier transform is the spectrum. A Fourier-transform spectrometer therefore measures no spectrum directly at all: it measures an interferogram and transforms it, which is why its resolution is set by the maximum path difference, \(\Delta\nu \approx c/(2x_{\max})\) by Corollary 17.44, and why its signal-to-noise advantage over a scanning instrument is that all frequencies are measured at once. The coherence properties this measures are the subject of Quantum Optics and the Photon.

The Laplace transform

Definition, convergence and inversion

The Fourier transform requires decay at both ends of the line. Physical initial-value problems supply neither: a system is switched on at \(t = 0\) and may grow. The remedy is to keep only \(t > 0\) and to insert a convergence factor, which is what Laplace's generating-function transform does [Laplace:1812].

Definition 17.65 (One-sided Laplace transform).

For \(f\) defined on \([0,\infty)\) and \(s\in\C\),

\begin{equation}\tag{17.80} F(s) = \mathcal{L}\left\{f\right\}(s) = \int_{0}^{\infty} f(t)\,\ee^{-st}\,\dd t\ec \end{equation}

whenever the improper integral converges. If \(t\) is in \(\mathrm{s}\) then \(s\) is in \(/\mathrm{s}\) and \(F\) carries the SI unit of \(f\) multiplied by \(\mathrm{s}\).

Theorem 17.66 (Half-plane of convergence).

If Equation (17.80) converges at \(s_{0}\), it converges for every \(s\) with \(\Re s > \Re s_{0}\), uniformly on every sector \(\abs{\arg(s-s_{0})}\le\theta_{0} < \pi/2\). Hence there is an abscissa of convergence \(\sigma_{c}\in[-\infty,+\infty]\) such that the transform exists for \(\Re s > \sigma_{c}\) and fails for \(\Re s < \sigma_{c}\). Rests on Definition 17.65 and Equation (7.28).

Proof.

Derives Theorem 17.66. Let \(g(t) = \int_{0}^{t} f(u)\ee^{-s_{0}u}\,\dd u\), which converges to a limit as \(t\to\infty\) by hypothesis and is therefore bounded, say \(\abs{g}\le M\). Write \(z = s - s_{0}\), with \(\Re z > 0\), and integrate by parts (Equation (7.28)), using \(g' = f\ee^{-s_{0}t}\):

\begin{equation*} \int_{0}^{T} f(t)\ee^{-st}\,\dd t = \int_{0}^{T} g'(t)\,\ee^{-z t}\,\dd t = g(T)\,\ee^{-zT} + z\int_{0}^{T} g(t)\,\ee^{-zt}\,\dd t\ep \end{equation*}

The boundary term is bounded by \(M\ee^{-T\Re z}\to0\), and the remaining integral converges absolutely as \(T\to\infty\) because \(\abs{g\ee^{-zt}}\le M\ee^{-t\Re z}\). Both estimates depend on \(z\) only through \(\abs{z}/\Re z\le 1/\cos\theta_{0}\) and \(\Re z\), whence uniform convergence on the sector. The abscissa is the infimum of \(\Re s\) over the set of convergence, which the first part shows to be a right half-plane.

Proposition 17.67 (Holomorphy).

On the half-plane \(\Re s > \sigma_{a}\) of absolute convergence, \(F\) is holomorphic (Definition 8.5) with

\begin{equation}\tag{17.81} F'(s) = -\int_{0}^{\infty} t\,f(t)\,\ee^{-st}\,\dd t\ec \end{equation}

and more generally \(F^{(n)}(s) = \mathcal{L}\{(-t)^{n}f\}(s)\). Rests on Definition 17.65, Definition 8.5 and Theorem 7.38.

Proof.

Derives Proposition 17.67. Fix \(s\) with \(\Re s = \sigma > \sigma_{a}\) and choose \(\eta>0\) with \(\sigma - \eta > \sigma_{a}\). For \(\abs{h}<\eta/2\),

\begin{equation*} \frac{F(s+h)-F(s)}{h} + \int_{0}^{\infty} t f(t)\ee^{-st}\dd t = \int_{0}^{\infty} f(t)\,\ee^{-st}\, \frac{\ee^{-ht} - 1 + ht}{h}\,\dd t\ec \end{equation*}

and the elementary estimate \(\abs{\ee^{-ht}-1+ht}\le\frac{1}{2}\abs{h}^{2}t^{2}\ee^{\abs{h}t}\), obtained from the Taylor remainder (Theorem 7.38), bounds the integrand by \(\frac{1}{2}\abs{h}\,t^{2}\abs{f(t)}\ee^{-(\sigma-\eta/2)t}\). Since \(t^{2}\ee^{\eta t/2}\) is dominated by \(\ee^{\eta t}\) for large \(t\), the bound is integrable and proportional to \(\abs{h}\), so the difference quotient converges and Equation (17.81) holds. The higher derivatives follow by induction.

Proposition 17.68 (The Laplace transform is a Fourier transform).

Let \(f_{\sigma}(t) = f(t)\,\ee^{-\sigma t}\,\theta(t)\), the causal signal damped by \(\ee^{-\sigma t}\). Then for \(\sigma > \sigma_{a}\)

\begin{equation}\tag{17.82} F(\sigma + \ii\omega) = \sqrt{2\pi}\;\widehat{f_{\sigma}}(\omega)\ep \end{equation}

Rests on Definitions 17.29 and 17.65.

Proof.

Derives Proposition 17.68. Substitute \(s = \sigma+\ii\omega\) in Equation (17.80) and read the result as Equation (17.32) applied to \(f_{\sigma}\), the step function restricting the range to \(t>0\).

So the Laplace transform is the Fourier transform of an exponentially damped causal signal, and its values on any vertical line \(\Re s = \sigma\) are a Fourier spectrum. The advantage over the Fourier transform is that the damping can be chosen to make the integral converge, and that the resulting function of \(s\) is holomorphic, so that its values on one line determine it everywhere in the half-plane by analytic continuation. Inversion follows at once.

Theorem 17.69 (Bromwich inversion integral).

Let \(\gamma > \sigma_{a}\) and suppose \(\omega\mapsto F(\gamma+\ii\omega)\) is absolutely integrable. Then at every point of continuity of \(f\)

\begin{equation}\tag{17.83} f(t) = \frac{1}{2\pi\ii}\int_{\gamma - \ii\infty}^{\gamma + \ii\infty} F(s)\,\ee^{st}\,\dd s\ec\qquad t>0\ec \end{equation}

the integral being taken up the vertical line \(\Re s = \gamma\); and the same integral vanishes for \(t<0\). Rests on Proposition 17.68 and Theorem 17.32.

Proof.

Derives Theorem 17.69. By Equation (17.82) the hypothesis says \(\widehat{f_{\gamma}}\) is absolutely integrable, so Theorem 17.32 applies to \(f_{\gamma}\):

\begin{equation*} f(t)\,\ee^{-\gamma t}\,\theta(t) = \frac{1}{\sqrt{2\pi}}\int_{-\infty}^{\infty} \widehat{f_{\gamma}}(\omega)\,\ee^{\ii\omega t}\,\dd\omega = \frac{1}{2\pi}\int_{-\infty}^{\infty} F(\gamma+\ii\omega)\,\ee^{\ii\omega t}\,\dd\omega\ep \end{equation*}

Multiply by \(\ee^{\gamma t}\) and substitute \(s = \gamma+\ii\omega\), so that \(\dd s = \ii\,\dd\omega\) and the line is traversed upward:

\begin{equation*} f(t)\,\theta(t) = \frac{1}{2\pi}\int_{-\infty}^{\infty} F(\gamma+\ii\omega)\,\ee^{(\gamma+\ii\omega)t}\,\dd\omega = \frac{1}{2\pi\ii}\int_{\gamma-\ii\infty}^{\gamma+\ii\infty} F(s)\,\ee^{st}\,\dd s\ep \end{equation*}

The factor \(\theta(t)\) on the left is the statement that the integral vanishes for \(t<0\).

The contour of Equation (17.83) is Bromwich's, introduced precisely to justify the operational methods of the next subsection [Bromwich:1916]. Evaluating it is a residue computation, and the machinery is that of Section 8.7; the only new ingredient is the estimate that closes the contour.

Lemma 17.70 (Jordan's lemma on the Bromwich contour).

Let \(C_{R}\) be the left-hand semicircle \(s = \gamma + R\ee^{\ii\phi}\), \(\phi\in\left[\pi/2, 3\pi/2\right]\), and let \(M_{R} = \max_{C_{R}}\abs{F}\). Then for \(t>0\)

\begin{equation}\tag{17.84} \abs{\int_{C_{R}} F(s)\,\ee^{st}\,\dd s} \le \frac{\pi\,\ee^{\gamma t}\,M_{R}}{t}\ec \end{equation}

so the semicircular contribution vanishes as \(R\to\infty\) whenever \(M_{R}\to0\). For \(t<0\) the same bound holds on the right-hand semicircle. Rests on Lemma 8.9.

Proof.

Derives Lemma 17.70. On \(C_{R}\), write \(\phi = \pi/2 + \psi\) with \(\psi\in[0,\pi]\); then \(\Re(s-\gamma) = R\cos\phi = -R\sin\psi\) and \(\abs{\ee^{st}} = \ee^{\gamma t}\ee^{-Rt\sin\psi}\). The ML estimate (Lemma 8.9) with \(\abs{\dd s} = R\,\dd\psi\) gives

\begin{equation*} \abs{\int_{C_{R}}F\ee^{st}\dd s} \le \ee^{\gamma t}M_{R}\,R\int_{0}^{\pi}\ee^{-Rt\sin\psi}\,\dd\psi = 2\,\ee^{\gamma t}M_{R}\,R\int_{0}^{\pi/2}\ee^{-Rt\sin\psi}\,\dd\psi\ec \end{equation*}

and Jordan's inequality \(\sin\psi\ge 2\psi/\pi\) on \([0,\pi/2]\) — concavity of the sine there — bounds the last integral by \(\int_{0}^{\pi/2}\ee^{-2Rt\psi/\pi}\dd\psi \le \pi/(2Rt)\), which is Equation (17.84).

Corollary 17.71 (Residue evaluation and causality).

Let \(F\) be holomorphic except for finitely many poles \(s_{1},\dots,s_{k}\), all with \(\Re s_{j} < \gamma\), and let \(F(s)\to0\) uniformly in direction as \(\abs{s}\to\infty\). Then

\begin{equation}\tag{17.85} f(t) = \sum_{j=1}^{k}\operatorname{Res}_{s = s_{j}} \left[F(s)\,\ee^{st}\right] \qquad (t>0)\ec\qquad f(t) = 0 \quad (t<0)\ep \end{equation}

Rests on Theorem 17.69, Lemma 17.70, Theorem 8.24 and Theorem 8.12.

Proof.

Derives Corollary 17.71. For \(t>0\) close the Bromwich line with the left semicircle of Lemma 17.70; the closed contour encircles every pole once counterclockwise, the semicircular contribution vanishes by Equation (17.84), and the residue theorem (Theorem 8.24) evaluates what is left, the \(2\pi\ii\) cancelling the prefactor of Equation (17.83). For \(t<0\) close to the right instead: the enclosed region contains no singularity, so Cauchy's theorem (Theorem 8.12) gives zero.

That the same integral gives zero for \(t<0\) is not a technicality: it is causality, built into the transform by the restriction to \(t>0\), and it is the reason the poles of \(F\) must lie to the left of the contour.

Example 17.72 (Damped sinusoid).

Let \(F(s) = \omega_{0}/\left[(s+\alpha)^{2}+\omega_{0}^{2}\right]\), with \(\alpha,\omega_{0}>0\) in \(/\mathrm{s}\) and \(\mathrm{rad}/\mathrm{s}\). The poles are simple, at \(s_{\pm} = -\alpha\pm\ii\omega_{0}\), and by Equation (8.16) with denominator derivative \(2(s+\alpha)\),

\begin{equation*} \operatorname{Res}_{s=s_{+}}\left[F\ee^{st}\right] = \frac{\omega_{0}\,\ee^{(-\alpha+\ii\omega_{0})t}}{2\ii\omega_{0}} = \frac{\ee^{-\alpha t}\ee^{\ii\omega_{0}t}}{2\ii}\ec \end{equation*}

and the residue at \(s_{-}\) is its complex conjugate with the sign reversed. Adding, \(f(t) = \ee^{-\alpha t}\sin(\omega_{0}t)\) for \(t>0\) — the impulse response of a damped oscillator, recovered from its transform by one contour integral.

Proposition 17.73 (Initial- and final-value theorems).

Let \(f\) be continuous on \([0,\infty)\) with \(f'\) absolutely integrable. Then

\begin{equation}\tag{17.86} \lim_{s\to\infty} s\,F(s) = f(0^{+})\ec\qquad \lim_{s\to0^{+}} s\,F(s) = \lim_{t\to\infty} f(t)\ec \end{equation}

the second whenever the limit on its right exists, the first for \(s\) real and increasing without bound. Rests on Definition 17.65, Equation (7.28) and Theorem 7.43.

Proof.

Derives Proposition 17.73. Integrating Equation (17.80) by parts gives \(\mathcal{L}\{f'\}(s) = sF(s) - f(0^{+})\), so \(sF(s) - f(0^{+}) = \int_{0}^{\infty}f'(t)\ee^{-st}\dd t\). The integrand is dominated by \(\abs{f'}\), which is integrable, and tends to \(0\) pointwise as \(s\to\infty\); dominated convergence gives the first limit. As \(s\to0^{+}\) the same integrand tends pointwise to \(f'(t)\) under the same domination, so the integral tends to \(\int_{0}^{\infty}f' = \lim_{t\to\infty}f(t) - f(0^{+})\) by the fundamental theorem of calculus (Theorem 7.43), which rearranges to the second limit.

Initial-value problems and the operational calculus

Proposition 17.74 (Derivative rule).

If \(f\) is \(n\) times differentiable with \(f^{(n)}\) of exponential order,

\begin{equation}\tag{17.87} \mathcal{L}\left\{f^{(n)}\right\}(s) = s^{n}F(s) - s^{n-1}f(0) - s^{n-2}f'(0) - \cdots - f^{(n-1)}(0)\ep \end{equation}

Rests on Definition 17.65 and Equation (7.28).

Proof.

Derives Proposition 17.74. For \(n=1\), integrate by parts (Equation (7.28)): \(\int_0^\infty f'\ee^{-st}\dd t = \left[f\ee^{-st}\right]_{0}^{\infty} + s\int_0^\infty f\ee^{-st}\dd t = -f(0) + sF(s)\), the boundary term at infinity vanishing for \(\Re s\) large. Induction on \(n\), applying the case \(n=1\) to \(f^{(n-1)}\), gives Equation (17.87).

Equation (17.87) is the whole method: it converts a linear ordinary differential equation with constant coefficients into an algebraic equation, and it swallows the initial data, which enter as the extra polynomial terms rather than as conditions imposed afterwards. This subsumes the constant-coefficient theory of Section 9.1 and discharges the section reserved for the transform at Section 9.13.

Example 17.75 (The driven damped oscillator, in SI).

Consider

\begin{equation}\tag{17.88} m\,\frac{\dd^{2}x}{\dd t^{2}} + b\,\dv{x}{t} + k\,x = f(t)\ec\qquad x(0) = x_{0}\ec\quad \dv{x}{t}(0) = v_{0}\ec \end{equation}

with \(m\) in \(\mathrm{kg}\), \(b\) in \(\mathrm{kg}/\mathrm{s}\), \(k\) in \(\mathrm{N}/\mathrm{m}\) and \(f\) in \(\mathrm{N}\). Transforming with Equation (17.87),

\begin{equation}\tag{17.89} \left(ms^{2}+bs+k\right)X(s) = F(s) + m\left(s x_{0} + v_{0}\right) + b\,x_{0}\ec \end{equation}

an algebraic equation. Define the transfer function

\begin{equation}\tag{17.90} G(s) = \frac{1}{ms^{2}+bs+k}\ec \end{equation}

in \(\mathrm{m}/\mathrm{N}\), whose poles are \(s_{\pm} = -\Gamma \pm \ii\omega_{d}\) with \(\Gamma = b/(2m)\) and \(\omega_{d} = \sqrt{\omega_{0}^{2}-\Gamma^{2}}\), \(\omega_{0} = \sqrt{k/m}\). By Corollary 17.71 the inverse transform of \(G\) is the impulse response \(h(t) = \ee^{-\Gamma t}\sin(\omega_{d}t)/(m\omega_{d})\), and by the convolution theorem in its Laplace form the forced part of the solution is \(x_{f} = h * f\) — Duhamel's principle, the same statement as the causal Green's function of Section 10.3. For \(m = 0.50\,\mathrm{kg}\), \(k = 200\,\mathrm{N}/\mathrm{m}\) and \(b = 2.0\,\mathrm{kg}/\mathrm{s}\) one has \(\omega_{0} = 20\,\mathrm{rad}/\mathrm{s}\), \(\Gamma = 2.0\,/\mathrm{s}\), \(\omega_{d} = 19.9\,\mathrm{rad}/\mathrm{s}\) and a quality factor \(Q = m\omega_{0}/b = 5.0\): the free oscillation decays by a factor \(\ee\) in \(0.50\,\mathrm{s}\), after about \(1.6\) cycles.

Corollary 17.76 (Heaviside's expansion theorem).

Let \(F = P/Q\) with \(P, Q\) polynomials, \(\deg P < \deg Q\), and \(Q\) having only simple zeros \(s_{j}\). Then

\begin{equation}\tag{17.91} f(t) = \sum_{j}\frac{P(s_{j})}{Q'(s_{j})}\,\ee^{s_{j}t}\ec \qquad t>0\ep \end{equation}

Rests on Corollary 17.71 and Equation (8.16).

Proof.

Derives Corollary 17.76. \(F(s)\to0\) as \(\abs{s}\to\infty\) because \(\deg P<\deg Q\), so Corollary 17.71 applies, and each residue is \(P(s_{j})\ee^{s_{j}t}/Q'(s_{j})\) by Equation (8.16).

Equation (17.91) is exactly the rule Heaviside used, and taught, without proof: he treated the differentiation operator \(p = \dd/\dd t\) as an algebraic symbol, wrote the solution of \(Q(p)x = P(p)w\) as \(x = \left[P(p)/Q(p)\right]w\), expanded the rational function in partial fractions, and interpreted each term by rules justified only by their success [Heaviside:1893]. Bromwich's contour integral is what turned the procedure into mathematics [Bromwich:1916], and Corollary 17.76 is the translation.

Proposition 17.77 (Poles and stability).

A linear time-invariant system with rational transfer function \(G\) proper and with poles \(s_{j}\) has an impulse response that decays to zero for every initial condition if and only if \(\Re s_{j} < 0\) for every \(j\); it is bounded but non-decaying if a simple pole sits on the imaginary axis and every other pole is strictly to the left; and it grows without bound otherwise. Rests on Corollary 17.71.

Proof.

Derives Proposition 17.77. By Corollary 17.71 the impulse response is a finite sum of terms \(t^{m_{j}}\ee^{s_{j}t}\) with \(m_{j}\) below the order of the pole, whose modulus is \(t^{m_{j}}\ee^{t\Re s_{j}}\). Such a term tends to zero iff \(\Re s_{j}<0\); it is bounded and non-decaying iff \(\Re s_{j} = 0\) and \(m_{j} = 0\); and it diverges otherwise. Distinct exponentials being linearly independent, no cancellation between terms is possible.

Theorem 17.78 (Causality implies dispersion relations).

Let \(\chi\) be a causal response function — \(\chi(t) = 0\) for \(t<0\) — absolutely integrable, with transform \(\hat\chi\). Then \(\hat\chi\) extends to a function holomorphic in the upper half plane \(\Im\omega>0\), and if \(\hat\chi(\omega)\to0\) as \(\abs{\omega}\to\infty\) there,

\begin{equation}\tag{17.92} \Re\hat\chi(\omega) = \frac{1}{\pi}\operatorname{PV}\!\int_{-\infty}^{\infty} \frac{\Im\hat\chi(\omega')}{\omega'-\omega}\,\dd\omega'\ec \qquad \Im\hat\chi(\omega) = -\frac{1}{\pi}\operatorname{PV}\!\int_{-\infty}^{\infty} \frac{\Re\hat\chi(\omega')}{\omega'-\omega}\,\dd\omega'\ep \end{equation}

Rests on Definition 17.29, Proposition 17.67, Proposition 17.42, Theorem 8.12, Lemma 8.9 and Equation (8.16).

Derivation. Derives Theorem 17.78. For \(\Im\omega>0\) the factor \(\ee^{-\ii\omega t}\) decays as \(\ee^{t\Im\omega}\) on the half-line \(t<0\) — where \(\chi\) vanishes — and as \(\ee^{-t\Im\omega}\) for \(t>0\), so \(\hat\chi(\omega) = \frac{1}{\sqrt{2\pi}}\int_{0}^{\infty} \chi(t)\ee^{-\ii\omega t}\dd t\) converges there and, by the argument of Proposition 17.67, is holomorphic. Now integrate \(\hat\chi(\omega')/(\omega'-\omega)\) for real \(\omega\) around the closed contour consisting of the real axis indented by a small semicircle of radius \(\eta\) above the point \(\omega'=\omega\), closed by a large semicircle in the upper half plane. The integrand is holomorphic inside, so Cauchy's theorem (Theorem 8.12) makes the total vanish. The large semicircle contributes nothing by the decay hypothesis and the ML estimate (Lemma 8.9). The small semicircle, traversed clockwise, contributes \(-\ii\pi\hat\chi(\omega)\) — half a residue, by Equation (8.16) and the parametrisation \(\omega' = \omega+\eta\ee^{\ii\phi}\). What remains of the real axis is the principal value. Hence

\begin{equation*} \operatorname{PV}\!\int_{-\infty}^{\infty} \frac{\hat\chi(\omega')}{\omega'-\omega}\,\dd\omega' = \ii\pi\,\hat\chi(\omega)\ec \end{equation*}

and separating real and imaginary parts gives Equation (17.92). Equivalently, the statement is Equation (17.54) applied to \(1/(\omega'-\omega-\ii\varepsilon)\).

Real and imaginary parts of a causal response are thus not independent: knowing how a medium absorbs at all frequencies determines how it refracts at each, and conversely. The physical instances — the Kramers–Kronig relations for the permittivity and the refractive index — belong to Electrodynamics in Matter; what has been proved here is that they follow from causality and analyticity alone, with no model of the medium.

Sampling and the discrete transform

The sampling theorem

Every measurement that reaches a computer has been sampled: the continuous output of a detector is read at instants \(t_{n} = n\Delta t\) and everything between is discarded. The question of what is lost has an exact answer.

Definition 17.79 (Band-limited function).

\(f\) is band limited to \(\Omega\) if it is square integrable and \(\hat f(\omega) = 0\) for \(\abs{\omega} > \Omega\). In terms of ordinary frequency the band is \(\abs{\nu} \le \nu_{\max} = \Omega/2\pi\), in \(\mathrm{Hz}\). Rests on Definition 17.29.

Theorem 17.80 (Sampling theorem).

Let \(f\) be band limited to \(\Omega\) and let \(\Delta t = \pi/\Omega\). Then \(f\) is determined by its samples at the instants \(n\Delta t\), and is reconstructed from them exactly by the cardinal series

\begin{equation}\tag{17.93} f(t) = \sum_{n=-\infty}^{\infty} f(n\Delta t)\, \operatorname{sinc}\!\left(\frac{\pi\left(t - n\Delta t\right)} {\Delta t}\right)\ec \qquad \operatorname{sinc}x = \frac{\sin x}{x}\ep \end{equation}

Equivalently, in \(\mathrm{Hz}\): a signal containing no frequency above \(\nu_{\max}\) is completely determined by samples taken at any rate \(\nu_{\mathrm{s}} = 1/\Delta t \ge 2\nu_{\max}\), the Nyquist rate. Rests on Definition 17.79, Definition 17.1, Theorem 17.34 and Theorem 17.32.

Proof.

Derives Theorem 17.80. On the interval \([-\Omega,\Omega]\), expand \(\hat f\) in a Fourier series of period \(2\Omega\) in the variable \(\omega\): \(\hat f(\omega) = \sum_{n}\gamma_{n}\ee^{-\ii n\pi\omega/\Omega}\) with \(\gamma_{n} = \frac{1}{2\Omega}\int_{-\Omega}^{\Omega} \hat f(\omega)\ee^{\ii n\pi\omega/\Omega}\dd\omega\), the series converging in mean square by Theorem 17.34. Since \(\hat f\) vanishes outside the band, the inversion formula Equation (17.43) reads \(f(t) = \frac{1}{\sqrt{2\pi}}\int_{-\Omega}^{\Omega} \hat f(\omega)\ee^{\ii\omega t}\dd\omega\); evaluating it at \(t = n\pi/\Omega = n\Delta t\) identifies the coefficients,

\begin{equation}\tag{17.94} f(n\Delta t) = \frac{2\Omega}{\sqrt{2\pi}}\,\gamma_{n}\ec \qquad\text{i.e.}\qquad \gamma_{n} = \frac{\sqrt{2\pi}}{2\Omega}\,f(n\Delta t)\ep \end{equation}

So the samples are the Fourier coefficients of the spectrum, up to a constant, and knowing them is knowing \(\hat f\), hence \(f\). To get Equation (17.93) explicitly, substitute the series for \(\hat f\) into the inversion integral and exchange sum and integral — legitimate because the series converges in mean square on a set of finite measure, so it converges in mean, and \(\ee^{\ii\omega t}\) is bounded:

\begin{equation*} f(t) = \frac{1}{\sqrt{2\pi}}\sum_{n}\gamma_{n} \int_{-\Omega}^{\Omega}\ee^{\ii\omega\left(t-n\Delta t\right)}\dd\omega = \frac{1}{\sqrt{2\pi}}\sum_{n}\gamma_{n}\, \frac{2\sin\left[\Omega\left(t-n\Delta t\right)\right]} {t - n\Delta t}\ep \end{equation*}

Insert Equation (17.94); the constants collapse to \(1/\Omega\) and leave

\begin{equation*} \sum_{n} f(n\,\Delta t)\, \frac{\sin\left[\Omega(t-n\,\Delta t)\right]} {\Omega(t-n\,\Delta t)}\ec \end{equation*}

which is Equation (17.93) because \(\Omega = \pi/\Delta t\). Sampling faster than the Nyquist rate only makes the band a proper subinterval of \([-\Omega,\Omega]\), which changes nothing in the argument.

Whittaker introduced the cardinal series as the interpolation formula distinguished by containing no frequency above the sampling limit [Whittaker:1915]; Nyquist, analysing how many independent symbols a telegraph line of given bandwidth can carry, obtained the rate [Nyquist:1928a]; and Shannon stated and proved the theorem in the form used ever since, as the bridge between a continuous waveform and a finite list of numbers [Shannon:1949].

Corollary 17.81 (Sampling periodises the spectrum; aliasing).

Let \(\omega_{\mathrm{s}} = 2\pi/\Delta t\). For \(f\) satisfying the decay hypotheses of Theorem 17.47,

\begin{equation}\tag{17.95} \sum_{m=-\infty}^{\infty}\hat f\!\left(\omega - m\omega_{\mathrm{s}}\right) = \frac{\Delta t}{\sqrt{2\pi}} \sum_{n=-\infty}^{\infty} f(n\Delta t)\, \ee^{-\ii n\,\omega\,\Delta t}\ep \end{equation}

The right-hand side depends on \(f\) only through its samples; the left is the spectrum with all its translates by multiples of the sampling frequency superposed. If \(f\) is band limited to \(\Omega < \omega_{\mathrm{s}}/2\) the translates do not overlap and \(\hat f\) is recovered by discarding all but the central copy — which is Theorem 17.80 again. If they overlap, a component at \(\nu\) is indistinguishable from components at \(\nu - m\nu_{\mathrm{s}}\) for every integer \(m\): it is aliased. Rests on Theorem 17.47, Theorem 17.32 and Definition 17.79.

Proof.

Derives Corollary 17.81. The left-hand side is periodic in \(\omega\) with period \(\omega_{\mathrm{s}}\), and the periodisation argument of Theorem 17.47 applies verbatim in the frequency variable: its Fourier coefficients are

\begin{equation*} C_{n} = \frac{1}{\omega_{\mathrm{s}}}\int_{-\infty}^{\infty} \hat f(\omega)\,\ee^{-\ii n \Delta t\,\omega}\,\dd\omega = \frac{\sqrt{2\pi}}{\omega_{\mathrm{s}}}\,f(-n\Delta t)\ec \end{equation*}

using \(2\pi/\omega_{\mathrm{s}} = \Delta t\) and the inversion formula Equation (17.43). Summing the series \(\sum_{n}C_{n}\ee^{\ii n\Delta t\,\omega}\) and relabelling \(n \mapsto -n\) gives Equation (17.95), since \(\sqrt{2\pi}/\omega_{\mathrm{s}} = \Delta t/\sqrt{2\pi}\).

Example 17.82 (Aliasing in the laboratory).

Audio is conventionally sampled at \(\nu_{\mathrm{s}} = 44.1\,\mathrm{kHz}\), so the Nyquist frequency is \(22.05\,\mathrm{kHz}\). A \(25.0\,\mathrm{kHz}\) component that reaches the converter is recorded as \(\abs{25.0\,\mathrm{kHz} - 44.1\,\mathrm{kHz}} = 19.1\,\mathrm{kHz}\), squarely inside the audible band and indistinguishable, after sampling, from a real signal there. Nothing downstream can undo this, which is why the analogue low-pass filter that removes such components before the converter is part of the apparatus and must be described as such. The gravitational-wave strain channel of Experiment: Gravitational Waves is recorded at \(16384\,\mathrm{Hz}\), giving a Nyquist frequency of \(8192\,\mathrm{Hz}\), comfortably above the few-hundred-hertz band where the signals sit.

Remark 17.83 (What the theorem does not say).

Three caveats, each of which has cost somebody a measurement. (i) No signal of finite duration is band limited. If \(\hat f\) vanishes outside \([-\Omega,\Omega]\) then \(f(t) = \frac{1}{\sqrt{2\pi}}\int_{-\Omega}^{\Omega} \hat f(\omega)\ee^{\ii\omega t}\dd\omega\) converges for complex \(t\) and is differentiable there, hence entire; were it to vanish on an interval, all its derivatives would vanish at an interior point, its Taylor series (Theorem 8.20) would vanish identically, and \(f\) would be identically zero. Band limitation and finite support are mutually exclusive, so every real record is aliased a little, and the design question is how much. (ii) The reconstruction Equation (17.93) converges slowly: the sinc decays only as \(1/t\), so a truncated cardinal series has errors falling as the reciprocal of the number of samples retained. (iii) An error in the sample times — jitter — is not an error in the sample values and is not corrected by any amount of averaging; its effect is a phase noise proportional to frequency, which is why timing specifications tighten as the band widens.

The discrete Fourier transform and the FFT

Definition 17.84 (Discrete Fourier transform).

For \(x = (x_{0},\dots,x_{N-1})\in\C^{N}\),

\begin{equation}\tag{17.96} X_{k} = \sum_{n=0}^{N-1} x_{n}\,W_{N}^{nk}\ec\qquad W_{N} = \ee^{-2\pi\ii/N}\ec\qquad k = 0,\dots,N-1\ep \end{equation}
Proposition 17.85 (Inversion and Parseval for the DFT).

The roots of unity satisfy

\begin{equation}\tag{17.97} \frac{1}{N}\sum_{n=0}^{N-1} W_{N}^{n(k-k')} = \delta_{kk'} \qquad\text{for } 0\le k,k' \le N-1\ec \end{equation}

hence

\begin{equation}\tag{17.98} x_{n} = \frac{1}{N}\sum_{k=0}^{N-1} X_{k}\,W_{N}^{-nk}\ec \qquad\text{and}\qquad \sum_{n=0}^{N-1}\abs{x_{n}}^{2} = \frac{1}{N}\sum_{k=0}^{N-1}\abs{X_{k}}^{2}\ep \end{equation}

Rests on Definition 17.84.

Proof.

Derives Proposition 17.85. If \(k = k'\) every term of Equation (17.97) is \(1\). If not, \(W_{N}^{k-k'}\neq1\) and the geometric sum is \(\left(1 - W_{N}^{N(k-k')}\right)/\left(1-W_{N}^{k-k'}\right) = 0\), because \(W_{N}^{N} = 1\). This is Lemma 17.2 for a finite cyclic group in place of the circle. Substituting Equation (17.96) into the right side of Equation (17.98) and using Equation (17.97) returns \(x_{n}\); the Parseval identity follows by the same computation applied to \(\sum_{n}\abs{x_n}^2 = \sum_n \bar x_n x_n\).

Proposition 17.86 (The DFT is exact for a band-limited periodic signal).

Let \(f\) be periodic with period \(T_{\mathrm{w}} = N\Delta t\) and have Fourier coefficients \(c_{n}\) vanishing for \(\abs{n}\ge N/2\). Then the DFT of the samples \(x_{n} = f(n\Delta t)\) satisfies \(X_{k} = N c_{k}\) for \(0\le k<N/2\) and \(X_{N-k} = N c_{-k}\) for \(0 < k \le N/2\). No information is lost and no approximation is made. Rests on Definition 17.84, Definition 17.1 and Proposition 17.85.

Proof.

Derives Proposition 17.86. Insert \(f(n\Delta t) = \sum_{\abs{m}<N/2} c_{m}\ee^{2\pi\ii mn/N}\), which is Equation (17.4) at \(t = n\Delta t\) with \(\omega_{1}\Delta t = 2\pi/N\), into Equation (17.96): \(X_{k} = \sum_{m}c_{m}\sum_{n=0}^{N-1}W_{N}^{n(k-m)} = N c_{k \bmod N}\) by Equation (17.97), the index being read modulo \(N\). Negative \(m\) therefore appear at \(N-\abs{m}\), which is the stated indexing.

For a signal that is neither exactly periodic over the record nor exactly band limited — that is, for every real record — the DFT differs from the continuous transform by exactly two effects, both already derived: aliasing, Corollary 17.81, from sampling; and leakage, from the finite record length, since truncating to \(N\) samples multiplies the signal by a rectangular window and hence, by Equation (17.67), convolves the spectrum with the window's transform.

Theorem 17.87 (Cooley–Tukey).

Let \(N = 2^{q}\). Then Equation (17.96) can be evaluated in \(\frac{N}{2}\log_{2}N\) complex multiplications and \(N\log_{2}N\) complex additions, rather than the \(N^{2}\) multiplications of the direct sum. Rests on Definition 17.84.

Derivation. Derives Theorem 17.87. Split the sum into even and odd indices, \(n = 2r\) and \(n = 2r+1\), and use \(W_{N}^{2rk} = W_{N/2}^{rk}\):

\begin{equation}\tag{17.99} X_{k} = \underbrace{\sum_{r=0}^{N/2-1}x_{2r}W_{N/2}^{rk}}_{E_{k}} + W_{N}^{k}\underbrace{\sum_{r=0}^{N/2-1}x_{2r+1}W_{N/2}^{rk}}_{O_{k}}\ec \end{equation}

so \(E\) and \(O\) are the discrete transforms of the even- and odd-indexed halves, each of length \(N/2\). They are periodic in \(k\) with period \(N/2\), and \(W_{N}^{k+N/2} = -W_{N}^{k}\), so the second half of the output costs nothing new:

\begin{equation}\tag{17.100} X_{k} = E_{k} + W_{N}^{k}O_{k}\ec\qquad X_{k+N/2} = E_{k} - W_{N}^{k}O_{k}\ec\qquad 0\le k < N/2\ep \end{equation}

Combining the two half-length transforms therefore costs \(N/2\) complex multiplications and \(N\) additions, whence the recursion \(M(N) = 2M(N/2) + N/2\) with \(M(1) = 0\). Dividing by \(N\), \(M(N)/N = M(N/2)/(N/2) + 1/2\), so \(M(N)/N\) increases by \(1/2\) at each of the \(\log_{2}N\) levels of the recursion, giving \(M(N) = \frac{N}{2}\log_{2}N\); the addition count follows identically.

Example 17.88 (What the saving is worth).

For \(N = 2^{20} = 1048576\) samples — one minute of a channel sampled at \(16384\,\mathrm{Hz}\) — the direct sum needs \(N^{2} \approx 1.1\times 10^{12}\) complex multiplications and the fast algorithm \(\frac{N}{2}\log_{2}N \approx 1.05\times 10^{7}\), a factor of about \(10^{5}\). This is the difference between an hour of computation and a tenth of a second, and it is why every measured spectrum quoted in this treatise, from the angular power spectrum of Experiment: The Cosmic Microwave Background to the strain spectra of Experiment: Gravitational Waves, exists.

The algorithm was published by Cooley and Tukey in 1965 [Cooley:1965] and transformed numerical practice within a few years. It was not new: Gauss had the same doubling recursion in an unpublished note of 1805, written while interpolating the orbit of the asteroid Pallas, and several nineteenth- and twentieth-century authors rediscovered pieces of it — a history reconstructed by Heideman, Johnson and Burrus [Heideman:1984]. The treatise records the attribution as they establish it: the algorithm is Gauss's, its impact is Cooley and Tukey's.

Example 17.89 (Leakage, windowing and resolution).

A record of \(N\) samples at spacing \(\Delta t\) spans \(T_{\mathrm{w}} = N\Delta t\) and resolves frequencies separated by

\begin{equation}\tag{17.101} \Delta\nu = \frac{1}{T_{\mathrm{w}}} = \frac{\nu_{\mathrm{s}}}{N}\ep \end{equation}

With \(N = 2^{16} = 65536\) samples at \(\nu_{\mathrm{s}} = 16384\,\mathrm{Hz}\) the record is \(4.0\,\mathrm{s}\) long and the resolution \(0.25\,\mathrm{Hz}\): no processing of that record can separate two lines closer than that, because Corollary 17.44 forbids it. A sinusoid whose frequency does not fall exactly on a bin is smeared over neighbouring bins by convolution with the transform of the rectangular window, Equation (17.40), whose first side lobe sits at \(-13.3\,\mathrm{dB}\) relative to the peak — the value \(20\log_{10}\abs{\operatorname{sinc}x}\) at the first stationary point \(x \approx 4.4934\) of the sinc, where \(\tan x = x\). A weak line \(20\,\mathrm{dB}\) below a strong neighbour is therefore buried by the neighbour's leakage. Tapering the record with a smooth window suppresses the side lobes at the price of broadening the main lobe: it is Theorem 17.22 again — a non-negative kernel in place of an oscillating one — and the trade between resolution and dynamic range is not avoidable, only chosen.

Other integral transforms

Hankel and Mellin transforms

The Fourier kernel \(\ee^{-\ii\omega t}\) diagonalises translations of the line. Other symmetries call for other kernels: rotations about an axis call for Bessel functions, and dilations of the half-line call for powers. Both transforms below are, at bottom, the Fourier transform in adapted coordinates, and both are treated systematically by Titchmarsh [Titchmarsh:1937] and Morse and Feshbach [Morse:1953].

Proposition 17.90 (Hankel transform as the axially symmetric Fourier transform).

Work in \(\R^{3}\) with cylindrical coordinates \((\rho,\varphi,z)\), and let \(f(\vect{r}) = g(\rho)\,\ee^{\ii m\varphi}\ee^{\ii k_{z}z}\). Then its three-dimensional Fourier transform is supported on the plane \(k_{z}\) fixed and equals

\begin{equation}\tag{17.102} \ii^{-m}\,\ee^{\ii m\varphi_{k}}\, \left(\mathcal{H}_{m}g\right)(k_{\perp})\ec\qquad \left(\mathcal{H}_{m}g\right)(k_{\perp}) = \int_{0}^{\infty} g(\rho)\,J_{m}(k_{\perp}\rho)\,\rho\,\dd\rho\ec \end{equation}

up to the fixed normalisation of the transverse transform, where \(J_{m}\) is the Bessel function of Section 9.8 and \((k_{\perp},\varphi_{k})\) are polar coordinates in the transverse plane of wavevectors. The transform Equation (17.102) is its own inverse:

\begin{equation}\tag{17.103} g(\rho) = \int_{0}^{\infty} \left(\mathcal{H}_{m}g\right)(k_{\perp})\,J_{m}(k_{\perp}\rho)\, k_{\perp}\,\dd k_{\perp}\ep \end{equation}

Rests on Definition 17.29 and Proposition 17.41.

Proof.

Derives Proposition 17.90. The \(z\) dependence separates and produces the constraint on \(k_{z}\) by Equation (17.52). In the transverse plane, write \(\vect{k}_{\perp}\cdot\vect{\rho} = k_{\perp}\rho\cos(\varphi-\varphi_{k})\) and do the angular integral with the Bessel integral representation

\begin{equation*} \int_{0}^{2\pi}\ee^{-\ii k_{\perp}\rho\cos\alpha}\, \ee^{\ii m\alpha}\,\dd\alpha = 2\pi\,\ii^{-m}\,J_{m}(k_{\perp}\rho)\ec \end{equation*}

which is the generating-function identity of Section 9.8.5 evaluated on the unit circle. What survives is the radial integral Equation (17.102). For the inversion, run the same computation on the two-dimensional inversion integral; the angular factor \(\ee^{\ii m\varphi}\) is reproduced and the radial measure \(k_{\perp}\dd k_{\perp}\) appears, giving Equation (17.103).

Example 17.91 (Axially symmetric potential problem).

Let \(\Phi\) satisfy Laplace's equation in the half space \(z>0\) of \(\R^{3}\), with axially symmetric boundary data \(\Phi(\rho,0) = \Phi_{0}(\rho)\) in \(\mathrm{V}\) and \(\Phi\to0\) as \(z\to\infty\). Separating in cylindrical coordinates (Section 10.2) gives solutions \(\ee^{-k z}J_{0}(k\rho)\) for each \(k>0\) in \(/\mathrm{m}\), so

\begin{equation}\tag{17.104} \Phi(\rho,z) = \int_{0}^{\infty} A(k)\,\ee^{-kz}\,J_{0}(k\rho)\, k\,\dd k\ec\qquad A(k) = \left(\mathcal{H}_{0}\Phi_{0}\right)(k)\ec \end{equation}

the coefficient being read off by Equation (17.103) at \(z=0\). The Hankel transform does for a cylindrical boundary what the Fourier transform does for a planar one; the diffraction integrals of Electromagnetic Waves and Optics for a circular aperture are the same transform with \(m=0\).

The Mellin transform.

Where Fourier analysis diagonalises translation, Mellin analysis diagonalises dilation. The multiplicative group of positive reals, with its invariant measure \(\dd x/x\), has characters \(x\mapsto x^{-s}\), and the transform against them is the following.

Definition 17.92 (Mellin transform).

For \(f\) defined on \((0,\infty)\) and \(s\in\C\),

\begin{equation}\tag{17.105} \tilde f(s) = \left(\mathcal{M}f\right)(s) = \int_{0}^{\infty} f(x)\,x^{s-1}\,\dd x\ec \end{equation}

whenever the integral converges absolutely. For \(s = N+1\) with \(N\) a non-negative integer, \(\tilde f(N+1) = \int_{0}^{\infty}f(x)x^{N}\dd x\) is the \(N\)-th moment of \(f\); the Mellin transform is thus the analytic continuation of the sequence of moments to complex order.

Theorem 17.93 (Fundamental strip and holomorphy).

Suppose \(f\) is continuous on \((0,\infty)\) with

\begin{equation}\tag{17.106} \abs{f(x)} \le C\,x^{-a}\ \ (0<x\le1)\ec\qquad \abs{f(x)} \le C\,x^{-b}\ \ (x\ge1)\ec\qquad a < b\ep \end{equation}

Then Equation (17.105) converges absolutely on the open strip \(a < \Re s < b\) and defines a holomorphic function there, bounded on every closed substrip. The strip is called fundamental, and \(\tilde f\) determines \(f\) on it. Rests on Definition 17.92 and Proposition 17.67.

Proof.

Derives Theorem 17.93. Split the integral at \(x = 1\). With \(\sigma = \Re s\), \(\int_{0}^{1}\abs{f}x^{\sigma-1}\dd x \le C\int_{0}^{1}x^{\sigma-a-1}\dd x = C/(\sigma-a)\), finite exactly when \(\sigma>a\); and \(\int_{1}^{\infty}\abs{f}x^{\sigma-1}\dd x \le C\int_{1}^{\infty}x^{\sigma-b-1}\dd x = C/(b-\sigma)\), finite exactly when \(\sigma<b\). Both bounds are uniform on a closed substrip, so the integral converges uniformly there and the argument of Proposition 17.67 — differentiation under the integral sign, the extra factor \(\ln x\) being absorbed by shrinking the substrip — gives holomorphy, with \(\tilde f'(s) = \int_{0}^{\infty}f(x)\ln x\;x^{s-1}\dd x\). Uniqueness follows from Theorem 17.94 below.

Theorem 17.94 (Mellin inversion).

Under the hypotheses of Theorem 17.93, if for some \(c\) in the fundamental strip \(\omega\mapsto\tilde f(c+\ii\omega)\) is absolutely integrable, then at every point of continuity

\begin{equation}\tag{17.107} f(x) = \frac{1}{2\pi\ii}\int_{c-\ii\infty}^{c+\ii\infty} \tilde f(s)\,x^{-s}\,\dd s\ec \end{equation}

the contour being the vertical line \(\Re s = c\), traversed upward. The value of the integral is independent of \(c\) within the strip. Rests on Theorems 8.12, 17.32 and 17.93.

Proof.

Derives Theorem 17.94. Substitute \(x = \ee^{-u}\), so \(\dd x/x = -\dd u\) and the half-line \(0<x<\infty\) becomes the whole \(u\) line reversed. Writing \(\phi(u) = f(\ee^{-u})\) and \(s = c+\ii\omega\),

\begin{equation}\tag{17.108} \tilde f(c+\ii\omega) = \int_{-\infty}^{\infty}\phi(u)\,\ee^{-(c+\ii\omega)u}\,\dd u = \sqrt{2\pi}\,\widehat{\phi_{c}}(\omega)\ec\qquad \phi_{c}(u) = \phi(u)\,\ee^{-cu}\ec \end{equation}

exactly the two-sided analogue of Equation (17.82): the Mellin transform on the line \(\Re s = c\) is the Fourier transform of \(f(\ee^{-u})\ee^{-cu}\). The hypothesis makes \(\widehat{\phi_{c}}\) absolutely integrable, so Theorem 17.32 gives

\begin{equation*} \phi(u)\,\ee^{-cu} = \frac{1}{2\pi}\int_{-\infty}^{\infty} \tilde f(c+\ii\omega)\,\ee^{\ii\omega u}\,\dd\omega\ec \end{equation*}

and multiplying by \(\ee^{cu}\), substituting \(s = c+\ii\omega\) with \(\dd s = \ii\dd\omega\) and returning to \(x = \ee^{-u}\) — so that \(\ee^{su} = x^{-s}\) — yields Equation (17.107). Independence of \(c\) follows from Cauchy's theorem (Theorem 8.12) applied to the rectangle between two admissible lines, the horizontal sides contributing nothing because \(\tilde f\) is bounded on the closed substrip and integrable along the verticals.

Example 17.95 (Two transforms worth knowing).

(i) For \(f(x) = \ee^{-x}\) the bounds Equation (17.106) hold with \(a = 0\) and any \(b\), so the fundamental strip is \(\Re s>0\) and, by the definition of the gamma function,

\begin{equation}\tag{17.109} \mathcal{M}\left\{\ee^{-x}\right\}(s) = \Gamma(s)\ec \qquad \Re s > 0\ep \end{equation}

Its poles at \(s = 0,-1,-2,\dots\) lie to the left of the strip, with residues \((-1)^{n}/n!\); by Proposition 17.96 these poles are the small-\(x\) expansion \(\ee^{-x} = \sum(-x)^{n}/n!\). (ii) For \(f(x) = 1/(1+x)\) the strip is \(0<\Re s<1\) and \(\mathcal{M}f(s) = \pi/\sin(\pi s)\), a function whose poles at every integer encode both the expansion of \(f\) in powers of \(x\) at the origin and its expansion in powers of \(1/x\) at infinity.

Proposition 17.96 (Poles give asymptotics).

Let \(\tilde f\) continue meromorphically to \(c' < \Re s < b\) with finitely many poles \(s_{1},\dots,s_{k}\) there and no pole on \(\Re s = c'\), and suppose \(\tilde f\to0\) uniformly as \(\abs{\Im s}\to\infty\) in the closed strip. Then, as \(x\to0^{+}\),

\begin{equation}\tag{17.110} f(x) = \sum_{j=1}^{k}\operatorname{Res}_{s=s_{j}} \left[\tilde f(s)\,x^{-s}\right] + O\!\left(x^{-c'}\right)\ec \end{equation}

and a simple pole at \(s_{j}\) with residue \(r_{j}\) contributes \(r_{j}x^{-s_{j}}\), a double pole a term in \(x^{-s_{j}}\ln x\). Rests on Theorem 17.94, Theorem 8.24, Lemma 8.9 and Definition 8.22.

Proof.

Derives Proposition 17.96. Apply the residue theorem (Theorem 8.24) to the rectangle with vertical sides \(\Re s = c\) and \(\Re s = c'\) and horizontal sides at \(\Im s = \pm T\). Traversed counterclockwise this runs up the line \(\Re s = c\) and down the line \(\Re s = c'\), so

\begin{equation*} \frac{1}{2\pi\ii}\int_{(c)} - \frac{1}{2\pi\ii}\int_{(c')} = \sum_{j}\operatorname{Res}_{s=s_{j}}\left[\tilde f(s)x^{-s}\right]\ec \end{equation*}

the horizontal sides vanishing as \(T\to\infty\) by the decay hypothesis and the ML estimate (Lemma 8.9). The first term is \(f(x)\) by Equation (17.107); the second is bounded by \(x^{-c'}\int\abs{\tilde f(c'+\ii\omega)}\dd\omega/(2\pi)\), since \(\abs{x^{-s}} = x^{-\Re s}\). The residue of a simple pole is \(r_{j}x^{-s_{j}}\) by Definition 8.22, and of a double pole \(\dd/\dd s\) of \(\tilde f_{2}(s)x^{-s}\), which produces the logarithm.

This is why the Mellin transform is the natural instrument for asymptotics: the analytic structure of \(\tilde f\) to the left of the fundamental strip is, term by term, the expansion of \(f\) near the origin, and the structure to the right is the expansion at infinity.

Multiplicative convolution and the moments

The property that makes the Mellin transform indispensable is that it turns a convolution over the multiplicative group into a product — exactly as Theorem 17.53 does for the additive group.

Definition 17.97 (Multiplicative convolution).

For \(f, g\) on \((0,\infty)\),

\begin{equation}\tag{17.111} \left(f \star g\right)(x) = \int_{0}^{\infty} f(y)\,g\!\left(\frac{x}{y}\right)\, \frac{\dd y}{y}\ep \end{equation}

The measure \(\dd y/y\) is the invariant (Haar) measure of the group \(\left((0,\infty),\times\right)\): it satisfies \(\dd(\lambda y)/(\lambda y) = \dd y/y\) for every \(\lambda>0\), which is what makes Equation (17.111) the exact analogue of Equation (17.63).

Theorem 17.98 (Mellin convolution theorem).

If the fundamental strips of \(\tilde f\) and \(\tilde g\) overlap, then on the common strip

\begin{equation}\tag{17.112} \mathcal{M}\left\{f\star g\right\}(s) = \tilde f(s)\,\tilde g(s)\ep \end{equation}

The transform is commutative and associative, and \(\left(f\star g\right)\star h\) has transform \(\tilde f\tilde g\tilde h\). Rests on Definitions 17.92 and 17.97.

Proof.

Derives Theorem 17.98. Substitute Equation (17.111) into Equation (17.105) and exchange the order of integration, legitimate by Fubini's theorem because on the common strip \(\iint\abs{f(y)}\abs{g(x/y)}x^{\sigma-1}\dd x\,\dd y/y\) is finite by the same computation performed with absolute values:

\begin{equation*} \mathcal{M}\left\{f\star g\right\}(s) = \int_{0}^{\infty}\frac{f(y)}{y} \left[\int_{0}^{\infty}g\!\left(\frac{x}{y}\right)x^{s-1}\dd x\right] \dd y\ep \end{equation*}

In the inner integral put \(x = yu\), so \(\dd x = y\,\dd u\) and \(x^{s-1} = y^{s-1}u^{s-1}\); the bracket becomes \(y^{s}\int_{0}^{\infty}g(u)u^{s-1}\dd u = y^{s}\tilde g(s)\). What remains is \(\tilde g(s)\int_{0}^{\infty}f(y)y^{s-1}\dd y = \tilde f(s)\tilde g(s)\). Commutativity and associativity follow, as for Proposition 17.51, from the substitution \(y\mapsto x/y\) and from Equation (17.112) with uniqueness of the transform.

Corollary 17.99 (Convolution on the unit interval).

Suppose \(f\) and \(g\) vanish for \(x>1\). Then so does \(f\star g\), and for \(0<x\le1\)

\begin{equation}\tag{17.113} \left(f\star g\right)(x) = \int_{x}^{1} f(y)\,g\!\left(\frac{x}{y}\right)\,\frac{\dd y}{y}\ec \end{equation}

while the Mellin transform reduces to the moment integral

\begin{equation}\tag{17.114} \tilde f(N+1) = \int_{0}^{1} f(x)\,x^{N}\,\dd x =: M_{N}[f]\ep \end{equation}

Consequently the moments of a multiplicative convolution are the products of the moments:

\begin{equation}\tag{17.115} M_{N}\!\left[f\star g\right] = M_{N}[f]\;M_{N}[g] \qquad\text{for every } N\ep \end{equation}

Rests on Definition 17.92, Definition 17.97 and Theorem 17.98.

Proof.

Derives Corollary 17.99. In Equation (17.111) the factor \(f(y)\) vanishes unless \(y\le1\) and \(g(x/y)\) vanishes unless \(x/y\le1\), i.e. \(y\ge x\); so the range is \(x\le y\le 1\), which is empty for \(x>1\) and gives Equation (17.113) otherwise. Equation (17.114) is Definition 17.92 at \(s = N+1\) with the support restriction, and Equation (17.115) is Equation (17.112) evaluated at \(s = N+1\), the point lying in the common strip whenever the moments converge.

Equation (17.115) is the statement the rest of this subsection exists for. A convolution over a multiplicative variable — a fraction, a ratio, a scale — is an integral operator whose kernel couples every value of the variable to every other, and it is intractable as it stands. Taking moments replaces it by ordinary multiplication, one independent equation for each \(N\): the moments diagonalise the convolution. The consequence for evolution equations is immediate.

Theorem 17.100 (Moments diagonalise a convolution evolution).

Let \(f(x,\tau)\) be supported in \(0<x\le1\) and satisfy the integro-differential equation

\begin{equation}\tag{17.116} \pdv{f(x,\tau)}{\tau} = \left(P \star f\right)(x,\tau) = \int_{x}^{1} P(y)\,f\!\left(\frac{x}{y},\tau\right)\, \frac{\dd y}{y}\ec \end{equation}

with a fixed kernel \(P\) whose moments \(\tilde P(N+1)\) exist. Then the moments \(M_{N}(\tau) = \int_{0}^{1}f(x,\tau)x^{N}\dd x\) satisfy the decoupled ordinary differential equations

\begin{equation}\tag{17.117} \dv{M_{N}(\tau)}{\tau} = \tilde P(N+1)\,M_{N}(\tau)\ec \qquad\text{whence}\qquad M_{N}(\tau) = M_{N}(0)\,\ee^{\tilde P(N+1)\,\tau}\ec \end{equation}

and \(f(x,\tau)\) is recovered from the moments by the inversion contour Equation (17.107), with \(M_{N} = \tilde f(N+1)\) continued to complex order. Rests on Corollary 17.99 and Theorem 17.94.

Proof.

Derives Theorem 17.100. Take the \(N\)-th moment of Equation (17.116). On the left, differentiation under the integral sign (dominated convergence, the moments being finite on a neighbourhood of \(\tau\)) gives \(\dd M_{N}/\dd\tau\). On the right, Equation (17.115) gives \(\tilde P(N+1)M_{N}(\tau)\). That is Equation (17.117), a linear first-order equation with constant coefficient, whose solution is the exponential. Recovery of \(f\) is Theorem 17.94 applied to the analytic continuation \(\tilde f(s,\tau) = \tilde f(s,0)\ee^{\tilde P(s)\tau}\), which agrees with \(M_{N}\) at \(s = N+1\).

Corollary 17.101 (Coupled channels).

If \(f = (f_{1},\dots,f_{p})\) is a vector of densities on \((0,1]\) obeying \(\pp_{\tau}f_{i} = \sum_{j}P_{ij}\star f_{j}\) with a fixed matrix of kernels, then the vectors of moments \(\vect{M}_{N}(\tau)\) obey

\begin{equation}\tag{17.118} \dv{\vect{M}_{N}}{\tau} = \mathbf{A}_{N}\,\vect{M}_{N}\ec\qquad \left(\mathbf{A}_{N}\right)_{ij} = \tilde P_{ij}(N+1)\ec\qquad \vect{M}_{N}(\tau) = \ee^{\mathbf{A}_{N}\tau}\,\vect{M}_{N}(0)\ep \end{equation}

If \(\mathbf{A}_{N}\) is diagonalisable, with eigenvalues \(\lambda^{(\alpha)}_{N}\) and left eigenvectors \(\vect{u}^{(\alpha)}_{N}\) — those satisfying \(\bigl(\vect{u}^{(\alpha)}_{N}\bigr)\transpose\mathbf{A}_{N} = \lambda^{(\alpha)}_{N}\bigl(\vect{u}^{(\alpha)}_{N}\bigr)\transpose\) — then the scalar combinations \(\vect{u}^{(\alpha)}_{N}\cdot\vect{M}_{N}\), different at each \(N\), evolve as pure exponentials \(\ee^{\lambda^{(\alpha)}_{N}\tau}\) and do not mix. Rests on Theorem 17.100 and Corollary 17.99.

Proof.

Derives Corollary 17.101. Take moments of each component and apply Equation (17.115) to each \(P_{ij}\star f_{j}\); the result is Equation (17.118), a linear system with constant matrix, solved by the matrix exponential. Contracting Equation (17.118) with a left eigenvector gives \(\dd\!\left(\vect{u}^{(\alpha)}_{N}\cdot\vect{M}_{N}\right)/\dd\tau = \lambda^{(\alpha)}_{N}\, \vect{u}^{(\alpha)}_{N}\cdot\vect{M}_{N}\), which decouples the system.

Remark 17.102 (Why this is the useful form).

Three features of Theorem 17.100 deserve emphasis. First, the diagonalisation is exact, not perturbative: no approximation has been made about the kernel \(P\) beyond the existence of its moments. Second, the eigenvalue \(\tilde P(N+1)\) is a number for each \(N\), computed once from the kernel, so an evolution over any interval in \(\tau\) costs one exponential per moment. Third, and least often stated, reconstructing \(f\) from its moments is the ill-conditioned step: the moment problem amplifies small errors in \(M_{N}\) into large errors in \(f\), for the same reason as Remark 17.55, and the stable route is not to invert the moments numerically but to continue \(\tilde f(s,\tau)\) analytically in \(s\) and evaluate the contour integral Equation (17.107) directly. The scale variable \(\tau\) in Equation (17.116) is dimensionless, being the logarithm of a ratio of two quantities of the same SI dimension; the moments \(M_{N}\) carry whatever dimension \(f\,\dd x\) carries and \(N\) affects none of it.

The Radon transform

The last transform in this chapter integrates a function not against a kernel but over a family of submanifolds, and asks whether the function can be recovered from the integrals. Radon settled the question in 1917 for lines in the plane and planes in space [Radon:1917], as pure mathematics; the answer became an instrument half a century later.

Definition 17.103 (Radon transform).

For \(f\) on \(\R^{3}\), integrable and decaying, and for a unit vector \(\vect{n}\) and a real \(p\),

\begin{equation}\tag{17.119} \left(\mathcal{R}f\right)(\vect{n},p) = \int_{\R^{3}} f(\vect{x})\, \delta\!\left(p - \vect{n}\cdot\vect{x}\right)\,\dd^{3}x\ec \end{equation}

the integral of \(f\) over the plane \(\vect{n}\cdot\vect{x} = p\). In a two-dimensional slice \(z = \text{const}\) of the same space the identical formula with \(\dd^{2}x\) gives the integral along the line \(\vect{n}\cdot\vect{x} = p\), which is what a transmission measurement records: for a beam of initial intensity \(I_{0}\) and attenuation coefficient \(\mu(\vect{x})\) in \(/\mathrm{m}\), each element of path removes \(\dd I = -\mu\,I\,\dd\ell\), which integrates along the ray to \(\ln\left(I_{0}/I\right) = \left(\mathcal{R}\mu\right)(\vect{n},p)\) — the Beer–Lambert law, whose left-hand side is dimensionless.

Theorem 17.104 (Projection-slice theorem).

Let \(\hat f\) denote the \(d\)-dimensional Fourier transform of \(f\) and \(\mathcal{F}_{1}\) the one-dimensional transform in the variable \(p\). Then

\begin{equation}\tag{17.120} \mathcal{F}_{1}\left[\left(\mathcal{R}f\right)(\vect{n},\cdot)\right] (\kappa) = (2\pi)^{(d-1)/2}\,\hat f\!\left(\kappa\vect{n}\right)\ec \end{equation}

in the normalisation of Definition 17.29: the one-dimensional transform of one projection is the \(d\)-dimensional transform of \(f\) restricted to the line through the origin in the direction \(\vect{n}\). Rests on Definitions 17.29 and 17.103.

Proof.

Derives Theorem 17.104. Insert Equation (17.119) and do the \(p\) integral first, the delta collapsing it:

\begin{equation*} \frac{1}{\sqrt{2\pi}}\int_{-\infty}^{\infty}\!\!\int_{\R^{d}} f(\vect{x})\,\delta\!\left(p-\vect{n}\cdot\vect{x}\right) \ee^{-\ii\kappa p}\,\dd^{d}x\,\dd p = \frac{1}{\sqrt{2\pi}}\int_{\R^{d}} f(\vect{x})\, \ee^{-\ii\kappa\vect{n}\cdot\vect{x}}\,\dd^{d}x\ec \end{equation*}

which is \(\hat f(\kappa\vect{n})\) times \((2\pi)^{d/2}/\sqrt{2\pi} = (2\pi)^{(d-1)/2}\), the \(d\)-dimensional transform carrying one factor \((2\pi)^{-1/2}\) per coordinate.

So the projections do determine \(f\): they fill the transform space along every ray through the origin, hence everywhere. Inversion is the statement of that fact in a form one can compute with.

Theorem 17.105 (Radon inversion in three-dimensional space).

For \(f\) smooth and rapidly decreasing on \(\R^{3}\),

\begin{equation}\tag{17.121} f(\vect{x}) = -\frac{1}{8\pi^{2}}\int_{S^{2}} \left.\frac{\pp^{2}}{\pp p^{2}} \left(\mathcal{R}f\right)(\vect{n},p)\right|_{p=\vect{n}\cdot\vect{x}} \dd\Omega(\vect{n})\ec \end{equation}

the integral running over the unit sphere of directions. Rests on Theorems 17.32 and 17.104.

Derivation. Derives Theorem 17.105. Write the three-dimensional inversion integral in polar coordinates in \(\vect{k}\), using \(\vect{k} = \kappa\vect{n}\) with \(\kappa\in\R\) and \(\vect{n}\) over the whole sphere, which double-counts and so carries a factor \(\tfrac12\):

\begin{equation*} f(\vect{x}) = \frac{1}{(2\pi)^{3}}\int_{\R^{3}} \left[\int f\,\ee^{-\ii\vect{k}\cdot\vect{y}}\dd^{3}y\right] \ee^{\ii\vect{k}\cdot\vect{x}}\,\dd^{3}k = \frac{1}{2(2\pi)^{3}}\int_{S^{2}}\!\!\int_{-\infty}^{\infty} \kappa^{2}\,g_{\vect{n}}(\kappa)\, \ee^{\ii\kappa\vect{n}\cdot\vect{x}}\,\dd\kappa\,\dd\Omega\ec \end{equation*}

where \(g_{\vect{n}}(\kappa) = \int\left(\mathcal{R}f\right)(\vect{n},p)\ee^{-\ii\kappa p}\dd p\) by Theorem 17.104 in its unnormalised form. Since \(\kappa^{2}\ee^{\ii\kappa p} = -\pp_{p}^{2}\ee^{\ii\kappa p}\), the inner integral is \(-\pp_{p}^{2}\left[2\pi\left(\mathcal{R}f\right)(\vect{n},p)\right]\) evaluated at \(p = \vect{n}\cdot\vect{x}\), by one-dimensional Fourier inversion. Collecting the constants, \(-2\pi/\left[2(2\pi)^{3}\right] = -1/(8\pi^{2})\), gives Equation (17.121).

Theorem 17.106 (Filtered back-projection in a plane slice).

In a two-dimensional slice, with \(\vect{n}(\theta) = (\cos\theta,\sin\theta)\) and projections \(P_{\theta}(p) = \left(\mathcal{R}f\right)(\vect{n}(\theta),p)\),

\begin{equation}\tag{17.122} f(\vect{x}) = \frac{1}{4\pi^{2}}\int_{0}^{\pi}\!\! \int_{-\infty}^{\infty} \left[\int P_{\theta}(p)\,\ee^{-\ii\kappa p}\dd p\right] \abs{\kappa}\,\ee^{\ii\kappa\vect{n}(\theta)\cdot\vect{x}}\, \dd\kappa\,\dd\theta\ep \end{equation}

Reading the formula from the inside out: transform each projection, multiply by the ramp filter \(\abs{\kappa}\), transform back, and smear the result back along the direction it came from (back-projection), summing over directions. Rests on Theorems 17.32 and 17.104.

Derivation. Derives Theorem 17.106. Repeat the previous computation in two dimensions. The inverse transform in polar coordinates \(\vect{k} = \kappa\vect{n}(\theta)\) has area element \(\abs{\kappa}\dd\kappa\,\dd\theta\) with \(\kappa\in\R\) and \(\theta\in[0,\pi)\) — the modulus is the Jacobian and the halved angular range compensates the sign of \(\kappa\) — so

\begin{equation*} f(\vect{x}) = \frac{1}{(2\pi)^{2}}\int_{0}^{\pi}\!\! \int_{-\infty}^{\infty}\hat f_{\mathrm{u}} \left(\kappa\vect{n}(\theta)\right)\, \ee^{\ii\kappa\vect{n}\cdot\vect{x}}\,\abs{\kappa}\, \dd\kappa\,\dd\theta\ec \end{equation*}

with \(\hat f_{\mathrm{u}}\) the unnormalised transform, which Theorem 17.104 identifies with the bracket of Equation (17.122).

The factor \(\abs{\kappa}\) is the whole content of the word “filtered”: plain back-projection without it reconstructs \(f\) blurred by a \(1/r\) kernel, because the polar measure over-weights low frequencies. It is also what makes the reconstruction noise-sensitive, by Remark 17.55: the ramp amplifies exactly the high-frequency components in which detector noise is loudest, so a practical filter is \(\abs{\kappa}\) rolled off above a cutoff, and the choice of cutoff is the trade between resolution and noise.

Corollary 17.107 (How many projections).

To reconstruct features of size \(d\) across an object of diameter \(D\), the detector spacing must satisfy \(\Delta p \le d/2\) by Theorem 17.80 applied to \(p\), and the angular spacing must satisfy \(D\,\Delta\theta/2 \le d/2\), so that adjacent projections do not skip a resolution element at the rim. Hence about \(\pi D/(2d)\) view angles are needed. Rests on Theorem 17.80 and Definition 17.103.

Proof.

Derives Corollary 17.107. The first condition is the Nyquist criterion for the variable \(p\), whose band limit is set by the finest structure present. For the second, two projections at angles differing by \(\Delta\theta\) differ, at radius \(D/2\), by an arc \(\left(D/2\right)\Delta\theta\); requiring that this not exceed the resolution \(d\) gives \(\Delta\theta\le 2d/D\), and the number of views over the half-turn \([0,\pi)\) is \(\pi/\Delta\theta \ge \pi D/(2d)\).

For an object of diameter \(D = 0.30\,\mathrm{m}\) reconstructed to \(d = 1.0\,\mathrm{mm}\), Corollary 17.107 asks for about \(470\) view angles and a detector pitch of \(0.5\,\mathrm{mm}\) — numbers of the order that clinical scanners use, arrived at from the sampling theorem alone. Cormack rediscovered the inversion independently, and, unlike Radon, carried it through to reconstruction algorithms and to a laboratory demonstration on a phantom [Cormack:1963]; computed tomography followed. It remains the most direct vindication in applied mathematics of a transform inversion formula proved for its own sake.

Remark 17.108 (Tomography of a quantum state).

The same inversion reconstructs a quantum state from measurements of rotated quadratures. For a single mode with position \(x\) in \(\mathrm{m}\) and momentum \(p\) in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), the Wigner function

\begin{equation}\tag{17.123} W(x,p) = \frac{1}{\pi\hbar}\int_{-\infty}^{\infty} \overline{\psi(x+y)}\,\psi(x-y)\,\ee^{2\ii p y/\hbar}\,\dd y \end{equation}

[Wigner:1932] has the property that its integral along any line in the \((x,p)\) plane is the probability density of the corresponding rotated quadrature — that is, the measured quadrature distributions are precisely the Radon transform of \(W\). Inverting by Equation (17.122) returns \(W\), hence the state. This is how the states prepared in Quantum Optics and the Photon are characterised, and the \(\hbar\) in Equation (17.123) is what makes the phase-space area element dimensionless: \(\dd x\,\dd p/(2\pi\hbar)\) counts states.

Three sections of this chapter have now met the same pattern from three directions. The Fourier transform diagonalises translation, the Mellin transform dilation, the Radon transform the family of hyperplanes; in each case a hard problem — a differential equation, a convolution over a ratio, an inverse problem — becomes multiplication, and the price is an inversion contour and a discipline about what may be assumed of the data. That is the whole method, and the rest of the treatise uses it.