Ordinary Differential Equations and Sturm–Liouville Theory
- Ordinary differential equations
- Linear systems with constant coefficients
- Linear systems with periodic coefficients
- Orthonormal systems of functions
- The Sturm–Liouville operator
- The semiclassical approximation
- The one-dimensional Helmholtz equation
- The Bessel equation
- The Legendre equation and spherical harmonics
- The Hermite equation
- The Laguerre equation
- Elliptic integrals and the Jacobi elliptic functions
- The Laplace transform
- Summary of the classical families
The linear boundary-value problems of physics — the vibrating string, the temperature of a bar, the electrostatic potential inside a cavity, the stationary states of a quantum particle — all reduce, once the partial differential equation is separated into ordinary ones (Partial Differential Equations), to a single template: a linear second-order ordinary differential equation containing an eigenvalue parameter and a weight function. This is the Sturm–Liouville problem. Its power lies in what all its instances share: the solutions belonging to distinct eigenvalues are orthogonal with respect to the weight, and the full solution set is a basis in which arbitrary admissible functions can be expanded. The present chapter sets up the language of orthonormal function systems, defines the Sturm–Liouville operator, derives the semiclassical approximation that serves every potential for which no closed-form solution exists, and then works through the classical families — Helmholtz (Fourier), Bessel and its modified companion, Legendre and its associated equation with the spherical harmonics, Hermite, Laguerre and Mathieu — recording for each the identification of the Sturm–Liouville data, the solutions, and the orthogonality relations given in the source.
Around that core sit the parts of the general theory of ordinary differential equations that the physical parts of this treatise consume constantly and that no other chapter owns: existence, uniqueness and continuous dependence; the elementary first-order equations solvable by an integrating factor or by separation; linear systems with constant coefficients, the matrix exponential, and the reading of stability off the eigenvalues; Floquet's theorem for periodic coefficients, which is what makes parametric resonance a statement rather than an estimate; and the elliptic integrals and Jacobi elliptic functions, which are what a conservative system with one degree of freedom reduces to when the small-amplitude approximation is dropped.
Ordinary differential equations
An ordinary differential equation (ODE) relates a function of a single real variable to its derivatives; every equation treated in this chapter is an ODE that is linear and of second order. The source outline opens the subject with a section of general definitions — order, linearity, homogeneity, the structure of solution spaces — but that section contains only its heading, so the present section supplies it. None of what follows is special to the Sturm–Liouville problem. It is what entitles every closed form recorded later in this chapter to be called the general solution rather than merely a pair of solutions: a linear homogeneous equation of order \(n\) has a solution set of dimension exactly \(n\), so two solutions of a second-order equation with non-vanishing Wronskian already exhaust it.
Order, linearity, homogeneity
Let \(I \subseteq \R\) be an open interval. An ordinary differential equation of order \(n\) on \(I\) is a relation
between a function \(y\) of the single variable \(x \in I\) and its derivatives up to the \(n\)-th, the \(n\)-th being present. A solution is a function \(y\), \(n\) times differentiable on \(I\), satisfying Equation (9.1) at every point of \(I\).
Equation (9.1) is linear when it can be written
with the coefficients \(a_{0},\ldots,a_{n}\) and the inhomogeneity \(b\) continuous on \(I\) and \(a_{n}\) nowhere zero there. It is homogeneous when \(b\) vanishes identically. The map \(L\) is linear: \(L[\alpha u + \beta v] = \alpha L[u] + \beta L[v]\) for all real constants \(\alpha\), \(\beta\). Rests on Definition 9.1.
Dividing Equation (9.2) by \(a_{n}\) — legitimate because \(a_{n}\) does not vanish — solves the equation for its highest derivative. Any equation so solved, linear or not, becomes a first-order system for the vector \(\vect{Y} = \bigl(y, y', \ldots, y^{(n-1)}\bigr)\) of the function and its first \(n-1\) derivatives,
where \(y^{(n)} = f\bigl(x,y,y',\ldots,y^{(n-1)}\bigr)\) is the equation in normal form. Solutions of the system and of the scalar equation correspond one to one, so every statement below is made for Equation (9.3) and read back for the scalar equation.
Existence and uniqueness
Let \(R = \set{(x,\vect{Y}) : \abs{x - x_{0}} \le a\ec\; \norm{\vect{Y} - \vect{Y}_{0}} \le b}\). A map \(\vect{f} : R \longrightarrow \R^{n}\) is Lipschitz in its second argument, with constant \(L\), if
for every \((x,\vect{U})\) and \((x,\vect{V})\) in \(R\).
A continuously differentiable \(\vect{f}\) satisfies Equation (9.4) on the compact box \(R\), with \(L\) the maximum there of the norm of its derivative with respect to \(\vect{Y}\), by the mean value theorem Theorem 7.35. Mere continuity does not suffice for uniqueness: \(\dd y/\dd x = 3\,y^{2/3}\) with \(y(0) = 0\) is solved both by \(y \equiv 0\) and by \(y = x^{3}\), and \(3y^{2/3}\) is continuous everywhere but Lipschitz nowhere near \(y=0\).
Let \(S \subseteq \R\) and let \(\vect{u}_{j} : S \longrightarrow \R^{n}\), \(j \ge 0\), satisfy \(\norm{\vect{u}_{j}(x)} \le M_{j}\) for every \(x \in S\), with \(\sum_{j\ge0} M_{j}\) convergent. Then \(\sum_{j\ge0}\vect{u}_{j}(x)\) converges absolutely at every \(x\in S\), and its partial sums converge to the sum \(\vect{U}\) uniformly on \(S\): for every \(\varepsilon>0\) there is an \(N\), depending on \(\varepsilon\) alone and not on \(x\), such that
If in addition every \(\vect{u}_{j}\) is continuous on \(S\), then so is \(\vect{U}\). Rests on Proposition 7.47 and Definition 7.20.
Derivation. Derives Lemma 9.6. Fix \(x \in S\). Each component obeys \(\abs{u_{j}^{i}(x)} \le \norm{\vect{u}_{j}(x)} \le M_{j}\), so \(\sum_{j}\abs{u_{j}^{i}(x)}\) converges by comparison with \(\sum_{j}M_{j}\) and the component series converges, both by Proposition 7.47. Call the limit \(\vect{U}(x)\).
For the uniformity, subtract the \(k\)-th partial sum: for \(k \ge N\),
and the right-hand side is the tail of a convergent series of numbers, hence below any prescribed \(\varepsilon\) once \(N\) is large enough. It contains no \(x\), which is exactly Equation (9.5).
Suppose now every \(\vect{u}_{j}\) is continuous and fix \(x_{0}\in S\) and \(\varepsilon>0\). Choose \(N\) with \(\sum_{j>N}M_{j} < \varepsilon/3\) and write \(\vect{S}_{N} = \sum_{j=0}^{N}\vect{u}_{j}\), a finite sum of continuous maps and therefore continuous (Definition 7.20): there is a \(\delta>0\) with \(\norm{\vect{S}_{N}(x)-\vect{S}_{N}(x_{0})} < \varepsilon/3\) for every \(x \in S\) with \(\abs{x-x_{0}}<\delta\). For such an \(x\),
the outer two terms being bounded by \(\sum_{j>N}M_{j}\) by the estimate just proved. Hence \(\vect{U}\) is continuous at \(x_{0}\).
∎Proposition 7.47 is a statement about series of numbers: it decides convergence at each fixed \(x\) separately and says nothing about how fast. What Lemma 9.6 adds is the single \(N\) that serves every \(x\) at once, and it is that uniformity — not the pointwise convergence — which carries continuity across to the limit and licenses the exchange of a limit with an integral. Both halves are stated and proved here because Theorem 9.8 needs them in exactly this form; the same two facts are what the Fourier arguments of Fourier Analysis and Integral Transforms invoke under the name of the \(M\)-test.
Let \(\vect{f}\) be continuous on the box \(R\) of Definition 9.4, bounded there by \(\norm{\vect{f}} \le M\), and Lipschitz in its second argument with constant \(L\). Then the initial value problem
has exactly one solution on \(\abs{x - x_{0}} \le h\), where \(h = \min\set{a,\; b/M}\). Rests on Definition 9.4, Equation (9.3) and Lemma 9.6.
Derivation. Derives Theorem 9.8. A continuous \(\vect{Y}\) solves Equation (9.6) if and only if it solves the integral equation
by the fundamental theorem of calculus (Theorems 7.42 and 7.43): differentiating Equation (9.7) returns Equation (9.6) and evaluating it at \(x_{0}\) returns the initial value, while integrating Equation (9.6) from \(x_{0}\) to \(x\) returns Equation (9.7).
Define the Picard iterates by \(\vect{Y}_{0}(x) = \vect{Y}_{0}\) and
Every iterate stays inside the box on \(\abs{x-x_{0}}\le h\): if \(\norm{\vect{Y}_{k}(s) - \vect{Y}_{0}} \le b\) there, the integrand of Equation (9.8) is defined and bounded by \(M\), so \(\norm{\vect{Y}_{k+1}(x) - \vect{Y}_{0}} \le M\abs{x-x_{0}} \le Mh \le b\); and the hypothesis holds trivially for \(k = 0\).
The successive differences obey
For \(k = 0\) this is the bound just used. Assuming it at \(k-1\) and applying Equation (9.4),
which is Equation (9.9).
Write \(\vect{Y}_{k} = \vect{Y}_{0} + \sum_{j=0}^{k-1}\bigl(\vect{Y}_{j+1} - \vect{Y}_{j}\bigr)\). By Equation (9.9) the terms of that series are bounded in supremum norm on \(\abs{x-x_{0}} \le h\) by \(M L^{j}h^{j+1}/(j+1)!\), whose sum is \(M\bigl(\ee^{Lh}-1\bigr)/L\), a finite number; these are the constants \(M_{j}\) of Lemma 9.6, so the series converges absolutely and uniformly on \(\abs{x-x_{0}}\le h\). Its sum \(\vect{Y}\) is therefore continuous, by the last clause of the same lemma, and obeys \(\norm{\vect{Y}-\vect{Y}_{0}} \le b\). Passing to the limit under the integral sign in Equation (9.8) is legitimate because Equation (9.4) gives \(\norm{\vect{f}(s,\vect{Y}_{k}) - \vect{f}(s,\vect{Y})} \le L\,\norm{\vect{Y}_{k}-\vect{Y}}\), which tends to zero uniformly by Equation (9.5), while the error committed on the integral is at most \(h\) times that supremum. Hence \(\vect{Y}\) solves Equation (9.7).
For uniqueness let \(\vect{Y}\) and \(\tilde{\vect{Y}}\) both solve Equation (9.7) on \(\abs{x-x_{0}} \le h\) and put \(\delta = \sup\norm{\vect{Y}-\tilde{\vect{Y}}}\), finite because both stay in the box. Subtracting the two integral equations and using Equation (9.4) gives \(\norm{\vect{Y}(x)-\tilde{\vect{Y}}(x)} \le L\,\delta\,\abs{x-x_{0}}\); feeding that bound back into the same inequality gives \(L^{2}\delta\abs{x-x_{0}}^{2}/2\), and after \(k\) rounds
which tends to zero as \(k \to \infty\) because the exponential series Equation (7.31) converges. Hence \(\delta = 0\).
∎Let \(A\) be a continuous matrix-valued function and \(\vect{g}\) a continuous vector-valued function on a compact interval \(J\) containing \(x_{0}\). Then \(\dd\vect{Y}/\dd x = A(x)\,\vect{Y} + \vect{g}(x)\) with \(\vect{Y}(x_{0}) = \vect{Y}_{0}\) has exactly one solution, and it is defined on the whole of \(J\). In particular the scalar linear equation Equation (9.2) has exactly one solution on \(I\) for each choice of the initial data \(y(x_{0}),\, y'(x_{0}),\, \ldots,\, y^{(n-1)}(x_{0})\). Rests on Theorem 9.8, Remark 9.3 and Definition 9.2.
Derivation. Derives Corollary 9.9. Here \(\vect{f}(x,\vect{Y}) = A(x)\vect{Y}+\vect{g}(x)\) is defined for every \(\vect{Y}\), so the box of Definition 9.4 may be taken to be \(J\times\R^{n}\) and Equation (9.4) holds globally with \(L = \max_{J}\norm{A}\), the maximum existing because \(J\) is compact and \(A\) continuous (Theorem 7.24). The construction in the proof of Theorem 9.8 then runs on all of \(J\) without the restriction \(h \le b/M\), which was needed only to keep the iterates inside a bounded box: the estimate Equation (9.9) holds on \(J\) with \(M = \max_{J}\norm{\vect{f}(\cdot,\vect{Y}_{0})}\), and the comparison series is again the exponential one. Uniqueness is unchanged. The scalar statement follows through Equation (9.3), whose coefficients \(-a_{i}/a_{n}\) and inhomogeneity \(b/a_{n}\) are continuous precisely because \(a_{n}\) does not vanish.
∎Theorem 9.8 says that the solution through a point exists and is unique. It says nothing about how two solutions through nearby points separate, and that quantitative question is the one every stability argument asks. The answer is a single inequality, and it is worth isolating because it is used far more often than it is named.
Let \(u\) be continuous and non-negative on \([x_{0},x_{1}]\) and suppose that
with constants \(\alpha \ge 0\) and \(\beta > 0\). Then
Rests on Theorem 7.42 and Corollary 7.36.
Derivation. Derives Lemma 9.10. Write \(U(x) = \alpha + \beta\int_{x_{0}}^{x}u(s)\,\dd s\). Since \(u\) is continuous, \(U\) is differentiable with \(U' = \beta u\) (Theorem 7.42), and \(U(x_{0}) = \alpha\). The hypothesis Equation (9.10) is exactly \(u \le U\), so \(U' = \beta u \le \beta U\) and
A function whose derivative is nowhere positive does not increase (Corollary 7.36), so \(\ee^{-\beta(x-x_{0})}U(x) \le U(x_{0}) = \alpha\) throughout, and \(u(x) \le U(x) \le \alpha\,\ee^{\beta(x-x_{0})}\).
∎Let \(\vect{f}\) satisfy the hypotheses of Theorem 9.8, with Lipschitz constant \(L\), and let \(\vect{Y}\) and \(\tilde{\vect{Y}}\) solve \(\dd\vect{Y}/\dd x = \vect{f}(x,\vect{Y})\) on \([x_{0},x_{1}]\) with initial values \(\vect{Y}_{0}\) and \(\tilde{\vect{Y}}_{0}\), both trajectories remaining in the box \(R\). Then
In particular the value of the solution at \(x\) depends continuously on its initial value, and two solutions that agree at one point agree throughout. Rests on Theorem 9.8, Lemma 9.10 and Equation (9.4).
Derivation. Derives Proposition 9.11. Both solutions satisfy the integral equation Equation (9.7) with their own initial values. Subtracting and using Equation (9.4),
This is Equation (9.10) for the continuous non-negative function \(u = \norm{\vect{Y}-\tilde{\vect{Y}}}\) with \(\alpha = \norm{\vect{Y}_{0}-\tilde{\vect{Y}}_{0}}\) and \(\beta = L\), so Lemma 9.10 gives Equation (9.12). Taking \(\alpha = 0\) recovers the uniqueness half of Theorem 9.8.
∎Equation (9.12) is an upper bound and nothing more. It licenses the statement that neighbouring solutions cannot separate faster than \(\ee^{Lx}\) — which is what makes a numerical trajectory meaningful over a finite time, and what turns a linearised decay estimate into a genuine statement about the nonlinear flow in Proposition 9.31. It does not assert that any pair of solutions actually separates at that rate. A system in which some pairs do separate exponentially, and in which the rate is a property of the flow rather than of a Lipschitz constant, is a different subject altogether; the exponents that measure it are defined in Nonlinear Dynamics and Chaos, and Grönwall's inequality is the tool that shows the linear estimate there survives the neglected quadratic terms.
The solution space and the Wronskian
The solutions on \(I\) of the homogeneous linear equation \(L[y] = 0\) of Equation (9.2) form a real vector space of dimension exactly \(n\). The solutions of the inhomogeneous equation \(L[y] = b\) form the affine set \(y_{p} + \ker L\), with \(y_{p}\) any one solution of it. Rests on Definition 9.2 and Corollary 9.9.
Derivation. Derives Theorem 9.13. Linearity of \(L\) makes \(\ker L\) a vector space and makes the difference of two solutions of \(L[y]=b\) a solution of \(L[y]=0\), which is the second statement. Fix \(x_{0} \in I\) and consider the evaluation map
which is linear. It is surjective because Corollary 9.9 produces a solution for every prescribed initial vector, and injective because that solution is unique: a \(y\) in \(\ker L\) with \(E(y) = \vect{0}\) agrees with the zero solution. An isomorphism onto \(\R^{n}\) forces \(\dim \ker L = n\).
∎The Wronskian of \(n\) functions \(y_{1},\ldots,y_{n}\), each \(n-1\) times differentiable on \(I\), is the determinant
If \(y_{1},\ldots,y_{n}\) all solve \(L[y] = 0\), their Wronskian satisfies
In particular \(W\) either vanishes identically on \(I\) or vanishes nowhere. Rests on Definitions 9.2 and 9.14.
Derivation. Derives Proposition 9.15. A determinant is multilinear in its rows, so \(\dd W/\dd x\) is a sum of \(n\) determinants, the \(i\)-th obtained by differentiating the \(i\)-th row of Equation (9.13). For \(i < n\) the differentiated row becomes equal to the row below it, so that determinant has two equal rows and vanishes. Only the last term survives, and in it the bottom row is \(\bigl(y_{1}^{(n)},\ldots,y_{n}^{(n)}\bigr)\). Each solution obeys \(y_{j}^{(n)} = -\sum_{i<n}\left(a_{i}/a_{n}\right)y_{j}^{(i)}\), so that row is a linear combination of the rows of Equation (9.13); subtracting the multiples of the rows carrying \(y^{(i)}\) for \(i \le n-2\), which changes no determinant, leaves \(-a_{n-1}/a_{n}\) times the row \(\bigl(y_{1}^{(n-1)},\ldots,y_{n}^{(n-1)}\bigr)\), that is \(-\left(a_{n-1}/a_{n}\right)W\). The resulting first-order linear equation integrates to the stated exponential, and an exponential never vanishes, so \(W(x)\) and \(W(x_{0})\) vanish together.
∎Let \(y_{1},\ldots,y_{n}\) solve \(L[y] = 0\) on \(I\). They are linearly independent if and only if \(W(x) \neq 0\) at one — equivalently at every — point of \(I\), and in that case every solution of \(L[y]=0\) is a unique linear combination of them. Rests on Theorem 9.13 and Proposition 9.15.
Derivation. Derives Proposition 9.16. If \(\sum_{j}c_{j}y_{j}\) vanishes identically with the \(c_{j}\) not all zero, differentiating \(n-1\) times shows that the columns of Equation (9.13) are linearly dependent at every \(x\), so \(W\) vanishes identically. Conversely, if \(W(x_{0}) = 0\) there is a non-zero vector \(\vect{c}\) annihilating the columns at \(x_{0}\), and \(y = \sum_{j}c_{j}y_{j}\) is then a solution with \(E(y) = \vect{0}\) in the notation of the proof of Theorem 9.13; by uniqueness \(y\) vanishes identically, so the \(y_{j}\) are dependent. Proposition 9.15 makes “at one point” and “at every point” the same statement. Finally \(n\) independent solutions inside an \(n\)-dimensional space (Theorem 9.13) are a basis of it.
∎Every equation treated below is linear, homogeneous and of second order, so its solution space is two-dimensional and exhibiting two solutions with non-vanishing Wronskian settles the general solution completely. That is the status of Equation (9.107), Equation (9.126), Equation (9.147) and Equation (9.158). The Sturm–Liouville equation Equation (9.55) is in normal form wherever \(p\) does not vanish; the endpoints at which \(p\) does vanish — \(x = 0\) for Bessel, \(x = \pm 1\) for Legendre and its associated equation — are exactly the singular points at which Corollary 9.9 does not apply, and they are also the points at which one of the two solutions ceases to be bounded.
Elementary first-order equations
Two first-order equations can be integrated outright, and between them they cover almost every first-order equation the physical parts of this treatise meet: a decay law, the charging of a capacitor, a relaxation towards equilibrium, a body falling against linear or quadratic drag. Theorem 9.8 already guarantees that each has exactly one solution; what follows produces that solution in closed form, so a later chapter can quote a formula instead of exhibiting an answer and verifying it.
Let \(p\) and \(q\) be continuous on an interval \(I\) and let \(x_{0}\in I\). Put \(P(x) = \int_{x_{0}}^{x}p(s)\,\dd s\). Then the solutions on \(I\) of
are exactly the functions
one for each constant \(C = y(x_{0})\). The factor \(\ee^{P}\), which turns the left-hand side of Equation (9.15) into an exact derivative, is the integrating factor. Rests on Corollary 9.9, Theorem 7.42 and Theorem 7.43.
Derivation. Derives Proposition 9.18. \(P\) is differentiable with \(P' = p\) (Theorem 7.42), so for any differentiable \(y\),
and \(\ee^{P}\) never vanishes. Hence \(y\) solves Equation (9.15) if and only if \(\bigl(\ee^{P}y\bigr)' = q\,\ee^{P}\), which by Theorem 7.43 holds if and only if \(\ee^{P(x)}y(x) - \ee^{P(x_{0})}y(x_{0}) = \int_{x_{0}}^{x}q\,\ee^{P}\); since \(P(x_{0}) = 0\) this is Equation (9.16) with \(C = y(x_{0})\). Every \(C\) occurs and no solution is missed, because Corollary 9.9 gives exactly one solution on \(I\) for each initial value.
∎Let \(g\) be continuous on an interval \(I\) and \(h\) continuous and nowhere zero on an interval \(J\), let \(x_{0}\in I\) and let \(y_{0}\) be an interior point of \(J\). Then, near \(x_{0}\), the initial-value problem
has exactly one solution with values in \(J\), and it is obtained by inverting the quadrature
For the autonomous equation \(\dd y/\dd x = F(y)\) with \(F\) continuous and non-vanishing this reads \(x - x_{0} = \int_{y_{0}}^{y}\dd v/F(v)\): the independent variable is recovered from the dependent one by a single integral, and inverting that integral is the whole of the solution. Rests on Proposition 7.32, Theorem 7.42 and Corollary 7.36.
Derivation. Derives Proposition 9.19. Set \(H(y) = \int_{y_{0}}^{y}\dd v/h(v)\) on \(J\) and \(G(x) = \int_{x_{0}}^{x}g(s)\,\dd s\) on \(I\); both are differentiable, with \(H' = 1/h\) and \(G' = g\) (Theorem 7.42). Since \(h\) is continuous and nowhere zero on the interval \(J\) it has one sign there (Theorem 7.23), so \(H'\) has one sign and \(H\) is strictly monotone (Corollary 7.36); it therefore maps \(J\) bijectively onto an interval \(H(J)\) containing \(0 = H(y_{0})\), and its inverse is differentiable with \((H^{-1})' = 1/H'\circ H^{-1} = h\circ H^{-1}\) (Proposition 7.32).
Let \(I_{0}\subseteq I\) be a neighbourhood of \(x_{0}\) on which \(G(x) \in H(J)\); such a neighbourhood exists because \(G\) is continuous and \(G(x_{0}) = H(y_{0}) = 0\) is an interior point of \(H(J)\), \(H\) being a strictly monotone continuous map and \(y_{0}\) interior to \(J\). On \(I_{0}\) define \(y = H^{-1}\circ G\). It takes the value \(H^{-1}(0) = y_{0}\) at \(x_{0}\), and by the chain rule (Proposition 7.31)
so it solves Equation (9.18). Conversely, if \(y\) is any solution with values in \(J\), then \(\dd\bigl[H(y(x)) - G(x)\bigr]/\dd x = y'/h(y) - g = 0\), so \(H(y(x)) - G(x)\) is constant (Corollary 7.36) and equal to its value \(0\) at \(x_{0}\); that is Equation (9.19), and \(H\) being injective determines \(y\) uniquely.
∎Both hypotheses on \(h\) are doing work. If \(h\) vanishes at \(y_{\ast}\) then \(y \equiv y_{\ast}\) is a constant solution, the integral \(\int\dd v/h\) diverges there or ceases to be defined, and uniqueness may fail — Remark 9.5 is exactly this situation with \(h(y) = 3y^{2/3}\). Physically the zeros of \(h\) are the equilibria, and the divergence of the quadrature at one of them is the statement that the equilibrium is reached only asymptotically: a body falling against drag approaches its terminal speed and never attains it, and the integral \(\int^{v_{t}}\dd v/(1 - v/v_{t})\) diverging is that fact in one line.
Linear systems with constant coefficients
The theory so far is scalar. A great deal of physics is not: Euler's equations for a rotating body, a chain of coupled oscillators, a chemical or nuclear decay chain and every linearisation of a nonlinear flow are first-order systems \(\dd\vect{Y}/\dd x = A\vect{Y}\) with \(A\) a constant matrix. Corollary 9.9 already says such a system has exactly one solution for each initial vector; this section produces that solution, describes the shape of every solution, and reads the stability of the equilibrium \(\vect{Y} = \vect{0}\) off the eigenvalues of \(A\). Nothing here is deeper than the series for the exponential and the residue theorem of Complex Analysis, but the conclusions are what every stability argument in the physical parts appeals to.
The matrix exponential
Let \(A\) be an \(n\times n\) matrix with real or complex entries and let \(\norm{\cdot}\) be the matrix norm induced by a norm on the underlying vector space (Definition 5.19),
The matrix exponential is
Rests on Definition 5.19 and Proposition 5.125.
The series Equation (9.21) converges absolutely for every \(A\), and uniformly in \(x\) on every bounded interval when \(A\) is replaced by \(xA\). The map \(x \mapsto \ee^{xA}\) is differentiable with
and \(\ee^{(x+y)A} = \ee^{xA}\ee^{yA}\); in particular \(\ee^{xA}\) is invertible, with inverse \(\ee^{-xA}\). If \(AB = BA\) then \(\ee^{A+B} = \ee^{A}\ee^{B}\). Rests on Definition 9.21, Lemma 9.6 and Theorem 7.51.
Derivation. Derives Proposition 9.22. Submultiplicativity gives \(\norm{x^{k}A^{k}/k!} \le \abs{x}^{k}\norm{A}^{k}/k! \le R^{k}\norm{A}^{k}/k!\) on \(\abs{x}\le R\), and those bounds sum to \(\ee^{R\norm{A}}\); by Lemma 9.6 the series converges absolutely and uniformly there, entry by entry. Each entry is a power series in \(x\) with infinite radius of convergence, so it may be differentiated term by term (Theorem 7.51), and \(\sum_{k\ge1}kx^{k-1}A^{k}/k! = A\sum_{k\ge0}x^{k}A^{k}/k!\), the factor \(A\) coming out on either side since it commutes with its own powers. That is Equation (9.22).
For the addition law with \(AB = BA\), expand the Cauchy product of the two absolutely convergent series (Proposition 7.49) and collect the terms of total degree \(m\): \(\sum_{j+k=m}A^{j}B^{k}/(j!\,k!) = (A+B)^{m}/m!\), the binomial theorem being available precisely because \(A\) and \(B\) commute. Taking \(B = yA\), which commutes with \(A\), gives \(\ee^{(x+y)A} = \ee^{xA}\ee^{yA}\), and \(y = -x\) identifies the inverse.
∎Let \(A\) be a constant \(n\times n\) matrix. The initial-value problem
has exactly one solution, defined for every \(x\in\R\), namely
The inhomogeneous system \(\vect{Y}' = A\vect{Y} + \vect{g}(x)\) with \(\vect{g}\) continuous has the unique solution
Rests on Proposition 9.22 and Corollary 9.9.
Derivation. Derives Theorem 9.23. Equation (9.22) makes Equation (9.24) a solution with the right initial value, and Corollary 9.9 says there is no other and that it lives on the whole line. For the inhomogeneous system, multiply by the integrating factor \(\ee^{-(x-x_{0})A}\) — the matrix analogue of Equation (9.17), legitimate because \(\ee^{-xA}\) commutes with \(A\) — to get \(\bigl(\ee^{-(x-x_{0})A}\vect{Y}\bigr)' = \ee^{-(x-x_{0})A}\vect{g}\), and integrate (Theorem 7.43). Multiplying back by \(\ee^{(x-x_{0})A}\) and using the addition law gives Equation (9.25).
∎Equation (9.22) says that \(x\mapsto\ee^{xA}\) is a one-parameter group of invertible matrices with generator \(A\). Proposition 14.3 proves the converse — every differentiable one-parameter group of linear maps is the exponential of its generator — and is the entry point to the Lie theory of Lie Groups, Lie Algebras, and Fibre Bundles. The two statements are the same computation read in opposite directions, and it is worth seeing that the flow of a linear autonomous system and the exponential map of a matrix group are not two constructions but one.
The shape of the solutions
The exponential is a closed form but not yet an answer: one still wants to know what the solutions look like. They are combinations of \(x^{k}\ee^{\lambda x}\), and the eigenvalues of \(A\) supply the \(\lambda\). The diagonalizable case is immediate; the general case is proved below without any appeal to a normal form for \(A\), by reading the exponential off the resolvent with the residue theorem.
Suppose \(A\) has \(n\) linearly independent eigenvectors \(\vect{v}_{1},\ldots,\vect{v}_{n}\) with eigenvalues \(\lambda_{1},\ldots,\lambda_{n}\). Then every solution of Equation (9.23) is
the constants \(c_{r}\) being the components of \(\vect{Y}_{0}\) in the eigenbasis. If \(A\) is real and \(\lambda = \alpha + \ii\beta\) is a non-real eigenvalue with eigenvector \(\vect{v} = \vect{a}+\ii\vect{b}\), then \(\bar\lambda\) is an eigenvalue with eigenvector \(\bar{\vect{v}}\), and the two real solutions carried by the pair are
Rests on Theorem 9.23, Definition 5.69 and Definition 5.15.
Derivation. Derives Corollary 9.25. For an eigenvector, \(A^{k}\vect{v}_{r} = \lambda_{r}^{k}\vect{v}_{r}\), so term-by-term application of Equation (9.21) gives \(\ee^{xA}\vect{v}_{r} = \ee^{\lambda_{r}x}\vect{v}_{r}\). Expanding \(\vect{Y}_{0} = \sum_{r}c_{r}\vect{v}_{r}\) in the eigenbasis — a basis by hypothesis — and applying Equation (9.24) gives Equation (9.26). For the real form, conjugating \(A\vect{v} = \lambda\vect{v}\) with \(A\) real gives \(A\bar{\vect{v}} = \bar\lambda\bar{\vect{v}}\); the real and imaginary parts of the complex solution \(\ee^{\lambda x}\vect{v}\) are then themselves solutions, because \(A\) is real and the equation is linear, and expanding \(\ee^{(\alpha+\ii\beta)x}(\vect{a}+\ii\vect{b})\) with Euler's formula (Proposition 8.4) exhibits them as Equation (9.27).
∎Let \(A\) be an \(n\times n\) complex matrix, let \(\lambda_{1},\ldots,\lambda_{p}\) be the distinct roots of its characteristic polynomial \(\chi(\lambda) = \det(\lambda\identity - A)\), and let \(m_{j}\) be the multiplicity of \(\lambda_{j}\) as a root. Then every entry of \(\ee^{xA}\) has the form
and the same is true of every component of every solution of Equation (9.23). Every entry of the \(N\)-th power \(A^{N}\) likewise has the form \(\sum_{j}\varpi_{j}(N)\,\lambda_{j}^{N}\) with \(\deg\varpi_{j}\le m_{j}-1\). Rests on Theorems 5.71, 8.24 and 9.23.
Derivation. Derives Theorem 9.26. The resolvent. For \(\abs{\lambda} > \norm{A}\) the series \(\sum_{k\ge0}A^{k}\lambda^{-k-1}\) converges absolutely and uniformly, by Lemma 9.6 with \(M_{k} = \norm{A}^{k}\abs{\lambda}^{-k-1}\) and the geometric series Proposition 7.46; multiplying it by \(\lambda\identity - A\) telescopes to \(\identity\), so its sum is the resolvent
Solving \((\lambda\identity - A)\vect{X} = \vect{e}_{j}\) by Cramer's rule Equation (5.16) expresses the \((i,j)\) entry of \(R(\lambda)\) as a quotient \(D_{ij}(\lambda)/\chi(\lambda)\) in which \(D_{ij}\) is the determinant of a matrix whose entries are affine in \(\lambda\), hence a polynomial of degree at most \(n-1\). So each entry of \(R\) is a rational function of \(\lambda\) whose poles lie among the roots of \(\chi\), that of \(\lambda_{j}\) being of order at most \(m_{j}\).
The contour integral. Let \(\Gamma\) be the positively oriented circle \(\abs{\lambda} = \rho\) with \(\rho > \norm{A}\), so that Equation (9.29) holds on \(\Gamma\) and every root of \(\chi\) lies inside (an eigenvalue \(\lambda_{j}\) has \(\abs{\lambda_{j}} \le \norm{A}\), since an eigenvector would otherwise be lengthened by \(A\) beyond the bound its norm allows). Multiplying Equation (9.29) by \(\ee^{\lambda x}\), which is bounded on \(\Gamma\), preserves the uniform convergence, so the integral may be taken term by term; and \((2\pi\ii)^{-1}\oint_{\Gamma}\ee^{\lambda x}\lambda^{-k-1}\dd\lambda = x^{k}/k!\), the \(k\)-th Taylor coefficient of \(\ee^{\lambda x}\) at \(\lambda = 0\) (Theorem 8.17). Summing,
The residues. By the residue theorem Theorem 8.24 the right-hand side of Equation (9.30) is the sum over \(j\) of the residues of \(\ee^{\lambda x}R(\lambda)\) at \(\lambda_{j}\). Near \(\lambda_{j}\) the entry \((\lambda-\lambda_{j})^{m_{j}}R_{ik}(\lambda)\) is holomorphic and has a Taylor expansion \(\sum_{l\ge0}c_{l}(\lambda-\lambda_{j})^{l}\) (Theorem 8.20), while \(\ee^{\lambda x} = \ee^{\lambda_{j}x} \sum_{i\ge0}x^{i}(\lambda-\lambda_{j})^{i}/i!\). The residue is the coefficient of \((\lambda-\lambda_{j})^{-1}\) in the product (Theorem 8.21 and Definition 8.22), that is
a polynomial in \(x\) of degree at most \(m_{j}-1\) times \(\ee^{\lambda_{j}x}\). Summing over \(j\) gives Equation (9.28), and Equation (9.24) carries it to every solution. Replacing \(\ee^{\lambda x}\) throughout by \(\lambda^{N}\) — also holomorphic, with \((2\pi\ii)^{-1}\oint\lambda^{N-k-1}\dd\lambda = \delta_{k N}\) — gives \(A^{N}\) in place of \(\ee^{xA}\) and the stated form of its entries.
∎Write \(\sigma(A) = \max_{j}\Re\lambda_{j}\). For every \(\sigma > \sigma(A)\) there is a constant \(C\) with
Consequently every solution of Equation (9.23) tends to \(\vect{0}\) as \(x\to+\infty\) if and only if every eigenvalue of \(A\) has strictly negative real part, and some solution is unbounded if some eigenvalue has strictly positive real part. If in addition \(A\) has \(n\) independent eigenvectors, every solution stays bounded on \(x\ge0\) whenever no eigenvalue has positive real part. Similarly \(A^{N}\to0\) as \(N\to\infty\) if and only if every eigenvalue satisfies \(\abs{\lambda_{j}}<1\), and the powers \(A^{N}\) stay bounded whenever every \(\abs{\lambda_{j}}\le1\) and \(A\) has \(n\) independent eigenvectors. Rests on Theorem 9.26 and Corollary 9.25.
Derivation. Derives Corollary 9.27. By Equation (9.28) each entry of \(\ee^{xA}\) is a finite sum of terms \(x^{k}\ee^{\lambda_{j}x}\), whose modulus is \(x^{k}\ee^{x\Re\lambda_{j}}\). For \(\sigma > \sigma(A)\) the quotient \(x^{k}\ee^{x\Re\lambda_{j}}/\ee^{\sigma x} = x^{k}\ee^{-(\sigma-\Re\lambda_{j})x}\) tends to zero as \(x\to+\infty\) and is continuous on \([0,\infty)\), hence bounded there; summing the finitely many bounds gives Equation (9.31). If every \(\Re\lambda_{j}<0\), choose \(\sigma<0\) and every solution decays exponentially. Conversely an eigenvector for \(\lambda\) furnishes the solution \(\ee^{\lambda x}\vect{v}\) of Corollary 9.25, whose norm is \(\ee^{x\Re\lambda}\norm{\vect{v}}\): it does not tend to zero when \(\Re\lambda \ge 0\) and is unbounded when \(\Re\lambda>0\). When \(A\) has an eigenbasis, Equation (9.26) writes every solution as a fixed combination of the \(\ee^{\lambda_{r}x}\vect{v}_{r}\), each of norm \(\ee^{x\Re\lambda_{r}}\norm{\vect{v}_{r}} \le \norm{\vect{v}_{r}}\) for \(x\ge0\) when \(\Re\lambda_{r}\le0\); the triangle inequality bounds the sum. The statements about \(A^{N}\) are the same two arguments on the second half of Theorem 9.26, with \(\abs{\lambda_{j}}^{N}\) in place of \(\ee^{x\Re\lambda_{j}}\) and, in the diagonalizable case, with \(A^{N}\vect{v}_{r} = \lambda_{r}^{N}\vect{v}_{r}\).
∎Let \(a_{0},\ldots,a_{n}\) be constants with \(a_{n}\neq0\), let
be the characteristic polynomial of \(L[y] = \sum_{i}a_{i}\,\dd^{i}y/\dd x^{i}\), and let \(\lambda_{1},\ldots,\lambda_{p}\) be its distinct roots, with multiplicities \(m_{1},\ldots,m_{p}\). Then the \(n\) functions
are a basis of the solution space of \(L[y]=0\). In particular, for \(n=2\) the general solution is \(c_{1}\ee^{\lambda_{1}x}+c_{2}\ee^{\lambda_{2}x}\) when the two roots differ and \((c_{1}+c_{2}x)\ee^{\lambda x}\) when they coincide; and if the \(a_{i}\) are real and the roots are \(\alpha\pm\ii\beta\) with \(\beta\neq0\), the real solutions are spanned by \(\ee^{\alpha x}\cos\beta x\) and \(\ee^{\alpha x}\sin\beta x\). Rests on Theorem 9.13, Theorem 8.19 and Proposition 7.30.
Derivation. Derives Proposition 9.28. Write \(D\) for \(\dd/\dd x\). Polynomials in \(D\) with constant coefficients multiply exactly as polynomials in \(\lambda\) do, because \(D\) commutes with itself and with constants, so the factorization \(P(\lambda) = a_{n}\prod_{j}(\lambda-\lambda_{j})^{m_{j}}\) furnished by the fundamental theorem of algebra (Theorem 8.19) gives the operator identity \(L = a_{n}\prod_{j}(D-\lambda_{j})^{m_{j}}\), in which the factors may be applied in any order.
A single factor acts on the proposed solutions by
(Proposition 7.30), so \((D-\lambda)^{m}\) annihilates \(x^{\kappa}\ee^{\lambda x}\) for every \(\kappa\le m-1\). Applying the factors of \(L\) in an order that puts \((D-\lambda_{j})^{m_{j}}\) last shows every function Equation (9.33) solves \(L[y]=0\).
They are independent. Suppose \(\sum_{j}p_{j}(x)\ee^{\lambda_{j}x}\) vanishes identically, with \(p_{j}\) polynomials; we show every \(p_{j}\) is zero, by induction on the number \(p\) of distinct exponents. For \(p=1\), an exponential never vanishes, so \(p_{1}\) has infinitely many roots and is the zero polynomial. For \(p>1\), multiply by \(\ee^{-\lambda_{p}x}\) and differentiate \(\deg p_{p}+1\) times. The last term dies; each other term \(p_{j}\ee^{\mu_{j}x}\), with \(\mu_{j} = \lambda_{j}-\lambda_{p}\neq0\), becomes \(q_{j}\ee^{\mu_{j}x}\) with \(\deg q_{j} = \deg p_{j}\), since each differentiation multiplies the leading coefficient by \(\mu_{j}\neq0\). The inductive hypothesis makes every \(q_{j}\), hence every \(p_{j}\) with \(j<p\), zero; and then \(p_{p}\) is zero by the case \(p=1\).
There are \(\sum_{j}m_{j} = n\) of them, and the solution space has dimension \(n\) (Theorem 9.13), so they are a basis. The real form of a conjugate pair is obtained as in Corollary 9.25, by taking the real and imaginary parts of \(\ee^{(\alpha+\ii\beta)x}\) (Proposition 8.4); they are independent because their Wronskian is \(\beta\,\ee^{2\alpha x}\neq0\) (Proposition 9.16).
∎Planar systems: classification of the equilibrium
Let \(A\) be a real \(2\times2\) matrix with \(\tau = \tr A\), \(\Delta = \det A\) and discriminant \(D = \tau^{2}-4\Delta\), so that the eigenvalues are
The origin, the only equilibrium of \(\vect{Y}' = A\vect{Y}\) when \(\Delta\neq0\), is classified by \((\tau,\Delta)\) exactly as in Table 9.1. In particular it is asymptotically stable if and only if \(\tau<0\) and \(\Delta>0\), and unstable whenever \(\Delta<0\) or \(\tau>0\). Rests on Corollaries 9.25 and 9.27.
Derivation. Derives Proposition 9.29. The characteristic polynomial of a \(2\times2\) matrix is \(\lambda^{2}-\tau\lambda+\Delta\) (Theorem 5.71), whose roots are Equation (9.34); they multiply to \(\Delta\) and add to \(\tau\). If \(D>0\) the roots are real and distinct: they have opposite signs when \(\Delta<0\) (a saddle, one growing and one decaying direction) and the common sign of \(\tau\) when \(\Delta>0\) (a node). If \(D<0\) they are a conjugate pair \(\tau/2\pm\ii\beta\) with \(\beta = \sqrt{-D}/2 \neq 0\), and Equation (9.27) shows every solution to be a rotation of period \(2\pi/\beta\) scaled by \(\ee^{\tau x/2}\): a focus spiralling in for \(\tau<0\) and out for \(\tau>0\), and for \(\tau=0\) a centre, every orbit of which is closed and periodic. If \(D=0\) the eigenvalue \(\tau/2\) is double and Equation (9.28) permits a factor of \(x\), present exactly when \(A\) is not a multiple of the identity: a degenerate node, or a star node when it is. The stability statements are Corollary 9.27 read through \(\Re\lambda_{\pm}\), both of which are negative precisely when their sum \(\tau\) is negative and their product \(\Delta\) positive.
∎| Condition | Eigenvalues | Equilibrium |
|---|---|---|
| $\Delta<0$ | real, opposite signs | saddle (unstable) |
| $\Delta>0$, $D>0$, $\tau<0$ | real, both negative | stable node |
| $\Delta>0$, $D>0$, $\tau>0$ | real, both positive | unstable node |
| $\Delta>0$, $D<0$, $\tau<0$ | complex, $\Re<0$ | stable focus |
| $\Delta>0$, $D<0$, $\tau>0$ | complex, $\Re>0$ | unstable focus |
| $\Delta>0$, $\tau=0$ | purely imaginary | centre |
| $\Delta>0$, $D=0$ | real, equal | degenerate or star node |
| $\Delta=0$ | one eigenvalue zero | a line of equilibria |
Equilibria of a nonlinear system
A nonlinear autonomous system \(\dd\vect{Y}/\dd x = \vect{F}(\vect{Y})\) with \(\vect{F}(\vect{Y}_{\ast}) = \vect{0}\) is approximated near its equilibrium by the linear system carried by the derivative matrix \(A = D\vect{F}(\vect{Y}_{\ast})\). Two questions follow, and they have different answers. Whether the stability of the equilibrium is decided by \(A\) is settled below, and affirmatively when no eigenvalue sits on the imaginary axis. Whether the whole phase portrait is decided by \(A\) is a deeper statement, and is the one theorem of this chapter that is quoted rather than proved.
An equilibrium \(\vect{Y}_{\ast}\) of \(\dd\vect{Y}/\dd x = \vect{F}(\vect{Y})\), \(\vect{F}\) of class \(C^{1}\), is hyperbolic when no eigenvalue of \(A = D\vect{F}(\vect{Y}_{\ast})\) has zero real part. Rests on Definitions 5.69 and 7.98.
Let \(\vect{F}\) be of class \(C^{1}\) near an equilibrium \(\vect{Y}_{\ast}\) and suppose every eigenvalue of \(A = D\vect{F}(\vect{Y}_{\ast})\) has \(\Re\lambda < 0\). Then \(\vect{Y}_{\ast}\) is asymptotically stable: there are \(\varepsilon>0\), \(C>0\) and \(\mu>0\) such that every solution starting within \(\varepsilon\) of \(\vect{Y}_{\ast}\) satisfies
Rests on Corollary 9.27, Lemma 9.10 and Theorem 9.23.
Derivation. Derives Proposition 9.31. Put \(\vect{Z} = \vect{Y}-\vect{Y}_{\ast}\) and \(\vect{N}(\vect{Z}) = \vect{F}(\vect{Y}_{\ast}+\vect{Z}) - A\vect{Z}\), so that \(\vect{Z}' = A\vect{Z} + \vect{N}(\vect{Z})\). Because \(\vect{F}\) is \(C^{1}\) and \(A\) is its derivative at the equilibrium, \(\norm{\vect{N}(\vect{Z})}/\norm{\vect{Z}} \to 0\) as \(\vect{Z}\to\vect{0}\) (Theorem 7.100): given any \(\eta>0\) there is a \(\delta>0\) with \(\norm{\vect{N}(\vect{Z})}\le\eta\norm{\vect{Z}}\) for \(\norm{\vect{Z}}\le\delta\).
Choose \(\sigma\) with \(\sigma(A)<\sigma<0\) and let \(C\) be the constant of Equation (9.31). Fix \(\eta = -\sigma/(2C)\) and the corresponding \(\delta\), and treat \(\vect{N}\) as an inhomogeneity in Equation (9.25): as long as the solution stays within \(\delta\),
Multiplying by \(\ee^{-\sigma x}\) and writing \(u(x) = \ee^{-\sigma x}\norm{\vect{Z}(x)}\) turns this into \(u(x) \le C\norm{\vect{Z}(0)} + C\eta\int_{0}^{x}u\), so Lemma 9.10 gives \(u(x)\le C\norm{\vect{Z}(0)}\ee^{C\eta x}\), that is
by the choice of \(\eta\). Taking \(\varepsilon = \delta/C\) makes the right-hand side at most \(\delta\) for every \(x\ge0\), so the solution never leaves the ball on which the estimate was assumed and the bound holds for all \(x\ge0\). This is Equation (9.35) with \(\mu = -\sigma/2 > 0\).
∎Let \(\vect{F}\) be of class \(C^{1}\) near a hyperbolic equilibrium \(\vect{Y}_{\ast}\) and let \(A = D\vect{F}(\vect{Y}_{\ast})\). Then there are neighbourhoods \(U\) of \(\vect{Y}_{\ast}\) and \(V\) of \(\vect{0}\) and a homeomorphism \(\varphi : U \longrightarrow V\) carrying the trajectories of \(\vect{Y}' = \vect{F}(\vect{Y})\) in \(U\) onto the trajectories of \(\vect{Z}' = A\vect{Z}\) in \(V\), preserving the direction and the parametrization of time [Hartman:1960]. The nonlinear flow near a hyperbolic equilibrium is therefore a continuous distortion of its linearization, and in particular the equilibrium is unstable whenever some eigenvalue of \(A\) has positive real part. Rests on Definition 9.30, Theorem 9.23 and Definition 6.7.
Theorem 9.32 is the only statement of this chapter about ordinary differential equations that carries no derivation, and the omission is a decision rather than a debt. Its proof constructs the conjugating map as the fixed point of a contraction on a space of bounded continuous maps, after splitting the linearization into its contracting and expanding parts and cutting the nonlinearity off outside a small ball; the technique is that of Theorem 9.8 applied in a function space, but the bookkeeping is a chapter's worth and it serves this one statement.
Three things about the statement are worth fixing in the reader's mind, because each is a place where more is often claimed than is true. The conjugacy is a homeomorphism and in general not differentiable, so it preserves the qualitative phase portrait — which trajectories go where — and not the eigenvalues or any other number computed from the linearization. Hyperbolicity is essential: Proposition 9.29 shows a centre, where \(\Re\lambda = 0\), and an arbitrarily small nonlinear perturbation can turn a centre into a stable or an unstable focus, so no conjugacy can exist there. And the theorem is local; it says nothing about the flow away from the equilibrium.
What the rest of this treatise actually uses is weaker and is proved here. The decay statement is Proposition 9.31, obtained from Equation (9.31) and Grönwall's inequality with no conjugacy anywhere; only the instability half of Theorem 9.32 is genuinely borrowed.
Planar flows and closed orbits
Let \(\vect{F}\) be of class \(C^{1}\) on an open subset of \(\R^{2}\) and let \(K\) be a compact subset of it containing no equilibrium of \(\vect{Y}' = \vect{F}(\vect{Y})\). If a trajectory enters \(K\) and remains in \(K\) for all later times, then it is a closed orbit or it approaches one [Poincare:1882] [Bendixson:1901]. A bounded planar trajectory that does not converge to an equilibrium therefore has periodic behaviour as its only possible fate. Rests on Definition 6.9, Definition 9.1 and Theorem 9.8.
The proof rests on two facts about the plane and on nothing else: a simple closed curve separates the plane into an inside and an outside (the Jordan curve theorem, which this treatise does not prove — Topological and Metric Spaces carries the winding number, Definition 6.20, which is the tool such a proof uses, but stops short of the separation theorem), and successive crossings of a short segment transverse to the flow are therefore monotone along that segment, because a trajectory arc together with the chord joining its ends bounds a region the flow cannot re-enter. It is the dimension that does the work, and the theorem is false in three dimensions: a bounded trajectory there may wander forever without approaching any closed orbit. That failure is not a technicality but the opening of a subject, and it is where Nonlinear Dynamics and Chaos begins.
Let \(\vect{F} = (F_{1},F_{2})\) be of class \(C^{1}\) on a simply connected open \(\Omega\subseteq\R^{2}\) and let \(\varphi\) be a positive \(C^{1}\) function on \(\Omega\). If
has one strict sign throughout \(\Omega\), then \(\vect{Y}' = \vect{F}(\vect{Y})\) has no closed orbit lying in \(\Omega\). Rests on Theorem 7.131 and Definition 6.17.
Derivation. Derives Proposition 9.36. Suppose a closed orbit \(\gamma\) lay in \(\Omega\), and let \(S\) be the region it bounds, contained in \(\Omega\) because \(\Omega\) is simply connected. Green's theorem Theorem 7.131 equates the integral of Equation (9.36) over \(S\) with the circulation \(\oint_{\gamma}\varphi\,(F_{1}\,\dd y_{2} - F_{2}\,\dd y_{1})\). On the orbit \(\dd y_{1} = F_{1}\,\dd x\) and \(\dd y_{2} = F_{2}\,\dd x\), so the integrand is \(\varphi\,(F_{1}F_{2}-F_{2}F_{1})\,\dd x = 0\) and the circulation vanishes. But the area integral of a function of one strict sign over a region of positive area cannot vanish — a contradiction.
∎Linear systems with periodic coefficients
A linear system whose coefficients vary periodically in the independent variable is not solvable in closed form, and yet its stability is decided completely by a single matrix. That is Floquet's theorem, and it is the exact statement behind every parametric-resonance calculation: a swing pumped by shortening the rope at twice its natural frequency, a mass on a vibrating support, a charged particle in the radio-frequency field of an ion trap, the tension modulation of a plucked string. The physical parts of this treatise derive the leading resonance band by a slowly-varying-amplitude argument; what follows is the theorem that argument approximates, and it is what entitles anyone to speak of a tongue structure at all.
The monodromy matrix
Let \(A\) be continuous on \(\R\) with values in the \(n\times n\) matrices. The fundamental matrix of \(\vect{Y}' = A(x)\vect{Y}\) based at \(x_{0}\) is the matrix-valued solution \(\Phi\) of
that is, the matrix whose \(j\)-th column is the solution with initial value \(\vect{e}_{j}\). Every solution is \(\vect{Y}(x) = \Phi(x)\vect{Y}(x_{0})\). Rests on Corollary 9.9 and Theorem 9.13.
The fundamental matrix satisfies
and is therefore invertible at every \(x\). Rests on Definition 9.37 and Proposition 9.15.
Derivation. Derives Proposition 9.38. This is the computation of Proposition 9.15 carried out on the columns of \(\Phi\) instead of on a Wronskian. A determinant is multilinear in its rows, so \(\dd(\det\Phi)/\dd x\) is the sum over \(i\) of the determinant of \(\Phi\) with its \(i\)-th row differentiated. By Equation (9.37) that row is \(\sum_{k}A_{ik}\Phi_{k\cdot}\), a combination of the rows of \(\Phi\); subtracting the multiples of the other rows, which changes no determinant, leaves \(A_{ii}\det\Phi\). Summing over \(i\) gives \(\dd(\det\Phi)/\dd x = (\tr A)\det\Phi\), a scalar linear equation whose solution with \(\det\Phi(x_{0}) = 1\) is Equation (9.38) by Proposition 9.18. An exponential never vanishes.
∎Let \(A\) be continuous and \(T\)-periodic, \(A(x+T) = A(x)\), and let \(\Phi\) be the fundamental matrix based at \(x_{0}=0\). Write \(M = \Phi(T)\), the monodromy matrix, and let \(\rho_{1},\ldots,\rho_{n}\) be its eigenvalues, the characteristic multipliers. Then
-
\(\Phi(x+T) = \Phi(x)\,M\) for every \(x\), and hence \(\Phi(x+NT) = \Phi(x)M^{N}\) for every integer \(N \ge 0\);
-
\(M\) is invertible, with \(\det M = \prod_{i}\rho_{i} = \exp\bigl[\int_{0}^{T}\tr A\bigr]\), so no multiplier is zero;
-
to each multiplier \(\rho\) belongs a solution with \(\vect{Y}(x+T) = \rho\,\vect{Y}(x)\); writing \(\rho = \ee^{\mu T}\) for a characteristic exponent \(\mu\), that solution is
\begin{equation}\tag{9.39} \vect{Y}(x) = \ee^{\mu x}\,\vect{P}(x)\ec \qquad \vect{P}(x+T) = \vect{P}(x)\ec \end{equation}an exponential times a \(T\)-periodic vector function;
-
every solution tends to \(\vect{0}\) as \(x\to+\infty\) if and only if \(\abs{\rho_{i}}<1\) for every \(i\), and some solution is unbounded if \(\abs{\rho_{i}}>1\) for some \(i\).
Rests on Definition 9.37, Proposition 9.38 and Corollary 9.27.
Derivation. Derives Theorem 9.39. (i) Put \(\Psi(x) = \Phi(x+T)\). Then \(\Psi' = A(x+T)\Psi = A(x)\Psi\) by periodicity, so \(\Psi\) solves Equation (9.37) column by column; and \(\Phi(x)M\) solves it too, being a fixed linear combination of solutions. The two agree at \(x=0\), where both equal \(M\), so they agree everywhere by the uniqueness in Corollary 9.9. Iterating gives the statement for \(NT\).
(ii) is Equation (9.38) at \(x=T\) together with the fact that the determinant is the product of the eigenvalues (Theorem 5.71, comparing the constant terms of the characteristic polynomial).
(iii) Let \(M\vect{w} = \rho\vect{w}\) with \(\vect{w}\neq\vect{0}\) and put \(\vect{Y}(x) = \Phi(x)\vect{w}\). By (i), \(\vect{Y}(x+T) = \Phi(x)M\vect{w} = \rho\,\vect{Y}(x)\). Since \(\rho \neq 0\) by (ii), the complex exponential takes the value \(\rho\) (Proposition 8.4: write \(\rho = \abs{\rho}\ee^{\ii\theta}\) and take \(\mu T = \ln\abs{\rho} + \ii\theta\)). Setting \(\vect{P}(x) = \ee^{-\mu x}\vect{Y}(x)\) gives \(\vect{P}(x+T) = \ee^{-\mu T}\ee^{-\mu x}\rho\,\vect{Y}(x) = \vect{P}(x)\), which is Equation (9.39).
(iv) By (i), \(\vect{Y}(x+NT) = \Phi(x)M^{N}\vect{Y}(0)\). The matrix \(\Phi\) is continuous, hence bounded on the compact interval \([0,T]\) (Theorem 7.24), and \(\Phi(x)^{-1}\) is bounded there too by Equation (9.38). So the solution tends to zero, or is bounded, exactly as the sequence \(M^{N}\vect{Y}(0)\) does, and the statements about the powers \(A^{N}\) in Corollary 9.27 settle that by the moduli of the multipliers: \(M^{N}\to0\) when every \(\abs{\rho_{i}}<1\), while an eigenvector for a multiplier of modulus greater than one gives \(M^{N}\vect{w} = \rho^{N}\vect{w}\), unbounded.
∎Floquet's theorem is often stated as the single assertion \(\Phi(x) = P(x)\,\ee^{xB}\) with \(P\) a \(T\)-periodic invertible matrix function and \(B\) constant. That form follows from Theorem 9.39(iii) whenever \(M\) has \(n\) independent eigenvectors: collect them as the columns of \(V\), put \(B = V\diag(\mu_{1},\ldots,\mu_{n})V^{-1}\) and \(P(x) = \Phi(x)\ee^{-xB}\), which is periodic because each column is by (iii). In general it requires a logarithm of the matrix \(M\) itself, which exists because \(M\) is invertible but whose construction needs a normal form for \(M\) that Linear Algebra and Representation Theory does not develop. Nothing below uses the aggregate form: every statement this treatise makes about a periodically forced linear system is a statement about the multipliers, and Theorem 9.39(iii) and (iv) supply those without any logarithm of a matrix. A further caution: \(B\) and \(P\) are generally complex even for a real system, because a negative real multiplier has no real logarithm; the standard repair is to work with period \(2T\), under which the multipliers are squared and become positive.
Hill's equation and the tongues of instability
Let \(Q\) be continuous and \(T\)-periodic and consider
written as a first-order system for \((y,y')\). Then \(\tr A = 0\), so \(\det M = 1\) and the two multipliers satisfy
Consequently, with \(\Delta_{\mathrm{H}} = \tr M\) the discriminant of Equation (9.40):
-
if \(\abs{\Delta_{\mathrm{H}}}<2\) the multipliers are \(\ee^{\pm\ii\theta}\) with \(2\cos\theta = \Delta_{\mathrm{H}}\), distinct and of unit modulus, and every solution is bounded on \(\R\): the equation is stable;
-
if \(\abs{\Delta_{\mathrm{H}}}>2\) the multipliers are real, reciprocal and distinct, one of modulus greater than one, so some solution grows like \(\abs{\rho}^{x/T}\) — exponentially, at the rate \(\mu = T^{-1}\ln\abs{\rho}\): the equation is unstable;
-
if \(\abs{\Delta_{\mathrm{H}}}=2\) the multipliers are both \(+1\) or both \(-1\), and Equation (9.40) has a solution of period \(T\) or of period \(2T\) respectively.
Rests on Theorem 9.39, Proposition 9.38 and Corollary 9.27.
Derivation. Derives Corollary 9.41. In the system for \(\vect{Y} = (y,y')\) the matrix is \(A = \begin{pmatrix} 0 & 1\\ -Q & 0\end{pmatrix}\), of zero trace, so Equation (9.38) gives \(\det M = 1\) and the characteristic polynomial of the \(2\times2\) matrix \(M\) is Equation (9.41) (Proposition 9.29, the same identity read for \(M\)). Its roots are \(\rho_{\pm} = \bigl[\Delta_{\mathrm{H}} \pm\sqrt{\Delta_{\mathrm{H}}^{2}-4}\,\bigr]/2\).
If \(\abs{\Delta_{\mathrm{H}}}<2\) the discriminant is negative, the roots are complex conjugates of product \(1\), hence of modulus \(1\), and distinct; \(M\) then has two independent eigenvectors (Proposition 5.75), so the powers \(M^{N}\) stay bounded by the last clause of Corollary 9.27, and by the argument of Theorem 9.39(iv) every solution is bounded. If \(\abs{\Delta_{\mathrm{H}}}>2\) the roots are real and distinct with product \(1\), so one has modulus greater than \(1\) and Theorem 9.39(iv) gives an unbounded solution; by Equation (9.39) it is \(\ee^{\mu x}\) times a periodic function, with \(\ee^{\mu T} = \rho\). If \(\abs{\Delta_{\mathrm{H}}}=2\) the double root is \(\rho = \Delta_{\mathrm{H}}/2 = \pm1\), and the eigenvector of \(M\) belonging to it furnishes, by Theorem 9.39(iii), a solution with \(y(x+T) = \pm y(x)\).
∎Let \(Q\) depend on two parameters — a mean value \(a\) and a modulation depth \(h\) — continuously. Then \(M\), and with it \(\Delta_{\mathrm{H}}(a,h)\), depends continuously on \((a,h)\), because the solution of an initial-value problem does (Proposition 9.11, applied to the system with parameters carried along). Corollary 9.41 therefore partitions the parameter plane into the open stable set \(\abs{\Delta_{\mathrm{H}}}<2\), the open unstable set \(\abs{\Delta_{\mathrm{H}}}>2\), and the level curves \(\Delta_{\mathrm{H}} = \pm2\) that separate them. On those curves a periodic or antiperiodic solution exists, by Corollary 9.41(iii), which is what makes them locatable in practice: one looks for solutions of period \(T\) and \(2T\) rather than computing \(M\). The unstable set is the union of the regions bounded by those curves, and at \(h=0\) it degenerates — as Proposition 9.43 shows for the Mathieu equation — to a discrete set of points on the \(a\) axis, from which the regions open out as \(h\) grows. That is the origin of the name.
The Mathieu equation
The canonical Hill equation, and the one that the separation of the Helmholtz operator in elliptic-cylinder coordinates produces [Mathieu:1868], is
Its Sturm–Liouville data, in the convention of Definition 9.54, are
so the spectral parameter is \(a\) and the modulation depth \(h\) is a parameter of the operator; much of the literature writes \(q\) for what is here called \(h\), which would collide with the Sturm–Liouville coefficient \(q(x)\). The coefficient has period \(\pi\) in \(x\), so Theorem 9.39 applies with \(T=\pi\). The solutions that are periodic — which by Corollary 9.41(iii) exist exactly on the boundaries of the unstable regions — are the Mathieu functions; the remaining solutions are, by Equation (9.39), an exponential times a periodic function, and it is the sign of the exponent that the experiment sees.
On the line \(h=0\), Equation (9.42) is stable in the sense of Corollary 9.41 at every \(a>0\) with \(\sqrt{a}\) not an integer, and unstable at every \(a\le0\). Consequently the closure of the unstable set meets the ray \(a>0\) exactly at the points
which are the feet from which the instability tongues open out as \(\abs{h}\) grows; the whole ray \(a\le0\) is unstable already at \(h=0\) and is the \(n=0\) region. Rests on Corollary 9.41, Theorem 9.23 and Equation (9.42).
Derivation. Derives Proposition 9.43. At \(h=0\) the equation is \(y''+ay=0\) with constant coefficients, whose fundamental matrix over one period \(T=\pi\) is computed from Theorem 9.23. For \(a>0\) the solutions are \(\cos(\sqrt{a}\,x)\) and \(\sin(\sqrt{a}\,x)/\sqrt{a}\), so
Hence \(\abs{\Delta_{\mathrm{H}}}<2\), the stable case of Corollary 9.41, at every \(a>0\) except where \(\cos(\pi\sqrt{a}) = \pm1\), that is where \(\pi\sqrt{a} = n\pi\), that is at \(a = n^{2}\) with \(n\ge1\); there \(\abs{\Delta_{\mathrm{H}}}=2\) and the second solution grows linearly, so those points belong to the closure of the unstable set. For \(a=0\) the solutions are \(1\) and \(x\), giving \(\Delta_{\mathrm{H}} = 2\) with an unbounded solution, and for \(a<0\) they are real exponentials with \(\Delta_{\mathrm{H}} = 2\cosh(\pi\sqrt{-a}) > 2\), the unstable case outright. So on the line \(h=0\) the unstable set is the ray \(a\le0\) together with the isolated points \(a = n^{2}\), \(n\ge1\), which is the statement.
∎An oscillator of natural angular frequency \(\omega_{0}\) whose restoring coefficient is modulated at angular frequency \(\gamma\) obeys \(\ddot{u} + \omega_{0}^{2}\bigl[1 + \epsilon\cos\gamma t\bigr]u = 0\). Substituting \(x = \gamma t/2\) turns it into Equation (9.42) with
so Equation (9.44) becomes \(\gamma = 2\omega_{0}/n\) with \(n\ge1\): the modulation destabilises the oscillator when it is driven at twice the natural frequency, or at an integer submultiple of that. The \(n=1\) tongue is the widest and is the one observed; the others narrow rapidly with \(n\), which is the quantitative statement that a slowly-varying-amplitude argument cannot make and this theorem can. Oscillations and Mechanical Waves derives the width of the first tongue and the damping threshold, and Experiment: Waves and Acoustics measures the effect on a string.
Orthonormal systems of functions
Orthonormality
Let \(\set{\phi_i}\) be a collection of functions
The functions \(\phi_i\) form an orthonormal set of functions if and only if
For real-valued functions the conjugation in Equation (9.46) is immaterial and the relation reduces to \(\int_I \phi_i\phi_j\,\dd x = \delta_{ij}\), which is the form in which the Legendre and Bessel relations below are stated.
Expansion in series of orthonormal functions
The expansion of a function \(f\) in a series of orthonormal functions \(\set{\phi_i}\) is
Rests on Definition 9.45.
The coefficients of the expansion Equation (9.47) are
Derivation. Derives Proposition 9.48. Multiply Equation (9.47) by \(\overline{\phi_j(x)}\) and integrate over \(I\):
where the orthonormality relation Equation (9.46) collapsed the sum.
∎Expansion of a product
Consider two functions \(f_1\) and \(f_2\), each expanded in the same orthonormal set, together with the expansion of their pointwise product:
where \(c_i^{(1)}\) and \(c_i^{(2)}\) are the expansion coefficients of \(f_1\) and \(f_2\), and \(c_i^{(12)}\) those of the product \(f_1 f_2\).
The structure constants of the orthonormal set \(\set{\phi_i}\) are its triple overlap integrals
Rests on Definition 9.45.
Whenever the three expansions above are valid and the double series below converges absolutely, the coefficients of the product are the values of one fixed bilinear form on the two coefficient sequences,
with \(T_{ijk}\) the structure constants Equation (9.49). Rests on Proposition 9.48 and Definition 9.49.
Derivation. Derives Proposition 9.50. Applying Equation (9.48) to the product and substituting the two series,
and the remaining integral is \(T_{ijk}\) by Equation (9.49).
∎No further reduction is available for a general orthonormal set, and none should be looked for: the numbers \(T_{ijk}\) are independent data attached to the set — they measure exactly how far it is from being closed under multiplication — and Equation (9.50) is the statement that pointwise multiplication, read in coefficients, is the bilinear map they define. What makes a particular set useful is that its structure constants are sparse. The complex exponentials are the extreme case: a single term of the double sum survives and the bilinear map degenerates into a convolution.
Take the complex exponentials \(\phi_n(x) = \ee^{\ii n x}\) on \([-\pi,\pi]\) in the normalisation of Equations (9.110) and (9.111), that is with every integral of Section 9.4 — and in particular the triple overlap Equation (9.49) — read against the normalised measure \(\dd x/(2\pi)\) of Remark 9.91, which is what makes them an orthonormal set at all. Their structure constants are then \(T_{njk} = \delta_{n,\,j+k}\) and Equation (9.50) becomes the discrete convolution
absolutely convergent whenever both coefficient sequences are absolutely summable. Rests on Proposition 9.50 and Definition 9.49.
Derivation. Derives Proposition 9.52. Every integral carries the measure \(\dd x/(2\pi)\) (Remark 9.91), which is also what puts the factor into Equation (9.111), so the triple overlap to be computed is
by Equation (9.109) applied to the integer \(j+k-n\). Inserting this in Equation (9.50) collapses the double sum to the single sum over \(j = m\), \(k = n-m\), which is Equation (9.51). Absolute convergence, and with it the right to rearrange the double series into that single one, is the Cauchy product of two absolutely convergent series (Proposition 7.49).
The other reading of Remark 9.91 gives the same answer, as it must. For the set \(\phi_n = \ee^{\ii n x}/\sqrt{2\pi}\) with plain \(\dd x\) the same integral yields \(T_{njk} = \delta_{n,\,j+k}/\sqrt{2\pi}\) while every coefficient is multiplied by \(\sqrt{2\pi}\), and the three factors cancel in Equation (9.50), leaving Equation (9.51) unchanged.
∎Parseval relation
Let \(f\) be a function whose expansion Equation (9.47) in the orthonormal set \(\set{\phi_i}\) converges to \(f\) in the mean square. The Parseval relation then equates the norm of the function to the norm of its coefficient sequence,
Derivation. Derives Equation (9.52). Mean-square convergence permits the series to be integrated term by term against its own conjugate:
the orthonormality relation Equation (9.46) collapsing the double sum to its diagonal.
∎The hypothesis is not decoration: it is exactly the statement that the orthonormal set is complete for \(f\). For an arbitrary orthonormal set only an inequality holds, and it costs one line: expanding the non-negative quantity \(\int_{I}\bigl|f - \sum_{i\le N}c_i\phi_i\bigr|^{2}\dd x = \int_{I}\abs{f}^{2}\dd x - \sum_{i\le N}\abs{c_i}^{2} \ge 0\), where the cross terms collapse by orthonormality and the definition \(c_i = \int_{I}\bar\phi_i f\,\dd x\), and letting \(N\to\infty\) gives
Bessel's inequality, with equality for every square integrable \(f\) being the definition of completeness. (Proposition 17.4 is this same argument carried out for the trigonometric system, where it also identifies the partial sum as the least-squares approximant; it does not by itself cover an arbitrary orthonormal set, which is why the general statement is given here.) The two systems used in this chapter are complete, and each is proved so: Theorem 17.34 for the trigonometric system of Section 9.7, and Theorem 17.36 for the Fourier integral of Section 9.7.5. The statement in the abstract, for an orthonormal basis of an arbitrary Hilbert space, belongs to Hilbert Spaces.
The Sturm–Liouville operator
Let \(y\) be a function
The Sturm–Liouville differential operator acting on \(y\) is defined as
where \(p\), \(q\) and \(r\) are given real functions and \(\lambda^2\) is a parameter.
Setting the expression Equation (9.54) equal to zero yields the Sturm–Liouville equation. Wherever \(r > 0\) it may be rewritten as
an eigenvalue problem with eigenvalue \(\lambda^2\) and weight function \(r\). Each classical family of special functions treated below is obtained by one specific choice of the data \(\bigl(p, q, \lambda^2, r\bigr)\); the identifications are collected in Table 9.2.
Self-adjointness and orthogonality
The whole force of the Sturm–Liouville form is one identity. Because the second-order part of Equation (9.54) is written as the derivative of \(p\,\dd y/\dd x\) and not as \(p\,\dd^{2}y/\dd x^{2}\), two solutions can be played off against each other under the integral sign, and every orthogonality relation proved in this chapter follows from that single manoeuvre rather than from a separate trick per family.
Write \(L[y] = \dv{}{x}\left[p\,\dv{y}{x}\right] + q\,y\) for the part of Equation (9.54) that does not carry the spectral parameter. For any twice differentiable \(u\) and \(v\),
Rests on Definition 9.54.
Derivation. Derives Lemma 9.56. The terms carrying \(q\) cancel, and by the Leibniz rule Equation (7.16),
in which the two products of first derivatives cancel, leaving Equation (9.56).
∎Let \(u\) and \(v\) be real solutions of the Sturm–Liouville equation Equation (9.55) on \([a,b]\) belonging to different eigenvalues, \(\lambda_{u}^{2} \neq \lambda_{v}^{2}\), and suppose the boundary term vanishes,
Then \(u\) and \(v\) are orthogonal with respect to the weight \(r\),
Rests on Lemma 9.56 and Equation (9.55).
Derivation. Derives Theorem 9.57. By hypothesis \(L[u] = -\lambda_{u}^{2}\,r\,u\) and \(L[v] = -\lambda_{v}^{2}\,r\,v\), so that \(u\,L[v] - v\,L[u] = \bigl(\lambda_{u}^{2}-\lambda_{v}^{2}\bigr)\,r\,u\,v\). Integrating Equation (9.56) over \([a,b]\) therefore gives
which is zero by Equation (9.57). Dividing by the non-zero factor \(\lambda_{u}^{2}-\lambda_{v}^{2}\) gives Equation (9.58).
∎Equation (9.57) is satisfied in each of the situations met below, and they are worth naming because they are the only ones used. Separated boundary conditions: \(u\) and \(v\) each satisfy \(\alpha\,y(a) + \beta\,\dv{y}{x}(a) = 0\) with the same pair \((\alpha,\beta)\), and similarly at \(b\); the bracket then vanishes at each endpoint separately, because \(\bigl(u,\dd u/\dd x\bigr)\) and \(\bigl(v,\dd v/\dd x\bigr)\) are proportional there. This is the Fourier–Bessel problem of Section 9.8.4, where the condition is \(y(a) = 0\). A vanishing coefficient function: \(p(a) = 0\) or \(p(b) = 0\) kills the bracket at that endpoint with no condition on \(y\) beyond boundedness of \(y\) and \(\dd y/\dd x\) there, which is what happens at \(x = 0\) for Bessel and at \(x = \pm 1\) for Legendre and its associated equation. Decay at infinity: on an unbounded interval it suffices that \(p\,y\,\dd y/\dd x\) tend to zero, which the Gaussian and exponential weights of Sections 9.10 and 9.11 guarantee for polynomial solutions. The second mechanism is the one that requires care: \(p\) vanishing is not by itself enough if the solutions are unbounded there, and Remark 9.112 is a worked instance in which the bracket survives.
The semiclassical approximation
The families catalogued in the sections that follow solve the Sturm–Liouville equation exactly, one closed form for each choice of the data \(\bigl(p,q,\lambda^{2},r\bigr)\). Almost no potential met in a laboratory is one of them. What survives when the exact solution does not exist is an asymptotic one: whenever the coefficient multiplying the unknown in a second-order linear equation is large, the equation has a universal approximate solution, an oscillation whose local wavelength and local amplitude are both read off that coefficient. Jeffreys obtained it for the general second-order equation [Jeffreys:1925], and Wentzel, Kramers and Brillouin obtained it independently of him and of each other, within months, for the wave equation of quantum mechanics [Wentzel:1926] [Kramers:1926] [Brillouin:1926] — which is why the method is called WKB by most authors and JWKB by the careful ones. This section derives the approximation, the formulae that carry a solution across the points where it fails, the tunnelling exponent that follows from them — the Gamow factor [Gamow:1928] [Gurney:1928] — and the quantisation condition of the old quantum theory [Bohr:1913] [Sommerfeld:1916], which the same machinery delivers with its half-integer correction attached rather than postulated.
The equation and its Sturm–Liouville data
The equation treated is the stationary wave equation of a particle of mass \(m\) moving on a line in a potential \(V\) [Schroedinger:1926a],
with \(x\) in \(\mathrm{m}\), \(m\) in \(\mathrm{kg}\), \(V\) and \(E\) in \(\mathrm{J}\), and \(\hbar = 1.054571817\times 10^{-34}\,\mathrm{J}\,\mathrm{s}\) exactly (Physical Constants and SI Units). In the template of Definition 9.54 this is a Sturm–Liouville equation with coefficient function \(1\), potential term \(-2mV(x)/\hbar^{2}\), weight \(1\) and eigenvalue \(\lambda^{2} = 2mE/\hbar^{2}\). The potential term and the eigenvalue both carry the SI dimension of an inverse squared length, as they must, being added to a second derivative with respect to \(x\); the coefficient function and the weight are dimensionless.
Dividing Equation (9.59) by \(-\hbar^{2}/(2m)\) puts it in the form used throughout this section,
where \(P(x)\), in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), is the momentum a classical particle of energy \(E\) would have at \(x\). It is real on the classically allowed set \(\set{x : V(x) < E}\) and imaginary on the classically forbidden set \(\set{x : V(x) > E}\), where the real quantity
is used instead. A point \(a\) with \(V(a) = E\), where the two sets meet, is a turning point: classically the particle stops there and reverses.
Throughout this section \(P\) denotes the local momentum Equation (9.60) and not the coefficient function called \(p\) in Definition 9.54, which for this equation is the constant \(1\) and plays no further role. The capital is a deliberate departure from the usual lower-case \(p\) of the semiclassical literature, adopted here only to keep the two meanings apart inside one chapter.
The expansion in powers of $\hbar$
Write the solution as a single exponential,
with \(S\) carrying the SI dimension of an action, \(\mathrm{J}\,\mathrm{s}\), so that the exponent is dimensionless. Nothing is approximated yet: any nowhere-vanishing \(\psi\) can be written this way with a complex \(S\).
The function \(\psi\) of Equation (9.62) solves Equation (9.60) if and only if
Derivation. Derives Proposition 9.60. Differentiating Equation (9.62) twice,
Substituting into Equation (9.60) and cancelling the nowhere-vanishing factor \(\psi\) leaves
and multiplication by \(-\hbar^{2}\) gives Equation (9.63).
∎The substitution has traded a linear equation of second order for a non-linear one of first order in \(\dd S/\dd x\) — a Riccati equation — and that is a bad bargain except for one feature: \(\hbar\) now appears in exactly one place, multiplying the highest derivative. The equation is therefore a singular perturbation problem in \(\hbar\), and its natural treatment is a power series in that constant,
in which \(S_{0}\) has the dimension of an action, \(S_{1}\) is dimensionless, \(S_{2}\) has the dimension of a reciprocal action, and so on. Equation (9.64) is the one place in this treatise where a physical constant is treated as a formal expansion parameter; the whole structure of the method is the ordering of Equation (9.63) by powers of \(\hbar\), so a convention that gave that constant the numerical value one would not simplify the argument but abolish it.
Substituting Equation (9.64) into Equation (9.63) and collecting powers of \(\hbar\) gives, at the first three orders,
Their solutions are
Rests on Proposition 9.60 and Equation (9.64).
Derivation. Derives Proposition 9.61. Insert Equation (9.64) in Equation (9.63). The squared term contributes \((S_{0}')^{2} + 2\hbar S_{0}'S_{1}' + \hbar^{2}\bigl[(S_{1}')^{2} + 2S_{0}'S_{2}'\bigr] + \cdots\), and the second term contributes \(-\ii\hbar S_{0}'' - \ii\hbar^{2}S_{1}'' - \cdots\); the right-hand side \(P^{2}\) is of order \(\hbar^{0}\). Equating coefficients of \(\hbar^{0}\), \(\hbar^{1}\) and \(\hbar^{2}\) gives Equations (9.65) to (9.67). The eikonal equation is solved by quadrature, with the two signs corresponding to the two directions of classical motion, which is the first equality of Equation (9.68). Dividing Equation (9.66) by \(2S_{0}'\) turns it into
and \(\abs{S_{0}'} = P\) by Equation (9.65), so integrating gives the second equality.
∎Truncating Equation (9.64) after the term of order \(\hbar\), the general solution of Equation (9.60) in a classically allowed region free of turning points is
and in a classically forbidden region, with \(\kappa\) of Equation (9.61),
Rests on Proposition 9.61, Equation (9.62) and Equation (9.61).
Derivation. Derives Corollary 9.62. By Equation (9.62) and Equation (9.68),
because \(\ii S_{1} = \ii\cdot\tfrac{\ii}{2}\ln P = -\tfrac12\ln P\), and the second factor is \(P^{-1/2}\). Linear superposition of the two signs gives Equation (9.69). In a forbidden region \(P = \ii\kappa\), the two exponents become real, and absorbing the constant phase of \(P^{-1/2}\) into the coefficients gives Equation (9.70).
∎The amplitude factor is not decoration. It is the exact content of the transport equation, and it says something physical.
For the single travelling wave \(\psi = C_{+}P^{-1/2} \exp\bigl[(\ii/\hbar)\int^{x}P\,\dd x'\bigr]\) of Equation (9.69), the probability current
is independent of \(x\). Rests on Corollary 9.62.
Derivation. Derives Proposition 9.63. Write \(\psi = C_{+}P^{-1/2}\ee^{\ii\theta}\) with \(\theta(x) = \hbar^{-1}\int^{x}P\,\dd x'\), so that \(\dd\theta/\dd x = P/\hbar\). Then
the first term inside the bracket being real because \(P\) is real in the allowed region. Taking the imaginary part and multiplying by \(\hbar/m\) gives Equation (9.71). Had the amplitude been anything other than \(P^{-1/2}\) the bracket would have carried an \(x\)-dependent imaginary part and the current would not have been constant, which for a stationary state on a line is impossible.
∎If \(\psi\) is normalised so that \(\abs{\psi}^{2}\) is a probability per unit length, in \(/\mathrm{m}\), then \(\abs{C_{+}}^{2}\) is in \(\mathrm{kg}/\mathrm{s}\) and \(j\) in \(/\mathrm{s}\): a probability per unit time, as a current must be. Since the classical speed is \(v = P/m\), Equation (9.71) reads \(\abs{\psi}^{2}v = \) constant — the particle is most likely to be found where it moves slowest, which is the classical statement the approximation reproduces.
Where the approximation is valid, and where it is not
For the local momentum of Equation (9.60),
in \(\mathrm{m}\): the de Broglie wavelength [deBroglie:1925] of the local motion, divided by \(2\pi\). Rests on Equation (9.60).
The term of order \(\hbar\) in Equation (9.64) is a small correction to the leading term, in the sense that \(\abs{\hbar\,\dd S_{1}/\dd x} \ll \abs{\dd S_{0}/\dd x}\), if and only if
Both expressions are dimensionless. Rests on Proposition 9.61, Equation (9.64) and Definition 9.64.
Derivation. Derives Proposition 9.65. By Equation (9.68), \(\hbar\,S_{1}' = \ii\hbar P'/(2P)\) and \(S_{0}' = \pm P\), so
the last equality because \(\dd\bar\lambda/\dd x = -\hbar P'/P^{2}\) from Equation (9.72). Differentiating \(P^{2} = 2m(E-V)\) gives \(2PP' = -2mV'\), hence \(P' = -mV'/P\) and \(\hbar\abs{P'}/P^{2} = \hbar m\abs{V'}/P^{3}\).
∎The first form of Equation (9.73) is the statement usually quoted: the local wavelength must change by much less than itself over a distance of one wavelength, so that the wave sees an essentially constant potential over each oscillation. The second form makes the same point in terms of the potential, and shows that the criterion is not a smallness condition on \(V'\) alone — a gentle slope still defeats the approximation where \(P\) is small.
At a turning point \(a\) the momentum vanishes, so \(\bar\lambda\) diverges and Equation (9.73) fails however smooth \(V\) may be. The failure is written on the face of the solution: the amplitudes \(P^{-1/2}\) and \(\kappa^{-1/2}\) in Equations (9.69) and (9.70) become infinite there, whereas the true solution of Equation (9.59) is perfectly finite. Quantitatively, near a simple turning point — one at which \(F = V'(a) \neq 0\) — one has \(P(x)^{2} \approx 2mF(a - x)\), so the second form of Equation (9.73) behaves as \(\abs{a-x}^{-3/2}\) and reaches unity at a distance of order
in \(\mathrm{m}\). Outside a neighbourhood of a few \(\ell\) the semiclassical forms are good; inside it they are worthless, and something else must be put there. That is the business of the next subsection.
Equation (9.64) is an asymptotic expansion. Carrying it to high order does not converge to the exact solution and generally makes matters worse beyond an optimal truncation, in the same way as the perturbation series of the anharmonic oscillator [Bender:1969]. All statements below therefore use the first two terms only, and the accuracy claimed is the one certified by Proposition 9.65.
Connection across a simple turning point
Let \(a\) be a simple turning point, \(V(a) = E\) with \(F = V'(a) \neq 0\), and take \(F > 0\), so that \(x < a\) is classically allowed and \(x > a\) is forbidden. On the scale \(\ell\) of Equation (9.74) the potential may be replaced by its tangent, \(V(x) \approx E + F\,(x-a)\), and Equation (9.60) becomes
From Equation (9.60) (linearise \(V\) about the turning point). Introducing the dimensionless variable \(z = (x-a)/\ell\) turns Equation (9.75) into
the Airy equation, whose solutions Airy introduced to describe the intensity of light near a caustic — the optical turning point, and the same mathematics [Airy:1838]. Its two standard solutions are fixed by a contour integral, and it is their behaviour far from the turning point that the connection formulae are made of.
For \(z \in \C\),
where \(C\) comes in from infinity along the direction \(\arg t = -\pi/3\) and returns to infinity along \(\arg t = +\pi/3\), and
The functions of Definition 9.68 are entire, both solve Equation (9.76), both are real on the real axis, and they are linearly independent. Rests on Definition 9.68 and Equation (9.76).
Derivation. Derives Proposition 9.69. The integrand is entire in \(t\) and decays superexponentially as \(\abs{t}\to\infty\) in each of the three open sectors where \(\Re\bigl(t^{3}\bigr) < 0\), namely those centred on \(\arg t = \pi/3\), \(\arg t = \pi\) and \(\arg t = -\pi/3\). Both ends of \(C\) lie in such sectors, so Equation (9.77) converges and may be differentiated under the integral sign. Since
one has
the bracket vanishing because both ends of \(C\) lie in decay sectors. If \(\omega^{3} = 1\) then \(w(z) = \mathrm{Ai}(\omega z)\) satisfies \(\dd^{2}w/\dd z^{2} = \omega^{2}\,\mathrm{Ai}''(\omega z) = \omega^{3}\,z\,\mathrm{Ai}(\omega z) = z\,w\), so each term of Equation (9.78) solves the equation and so does their sum. Reflection: conjugating Equation (9.77) maps \(C\) onto itself with reversed orientation and sends the integrand to \(\exp\bigl[t^{3}/3 - \bar z t\bigr]\), while \(\overline{1/(2\pi\ii)} = -1/(2\pi\ii)\); the two sign changes cancel and \(\overline{\mathrm{Ai}(z)} = \mathrm{Ai}(\bar z)\). Hence \(\mathrm{Ai}\) is real on the real axis, and for real \(z\) the two terms of Equation (9.78) are complex conjugates of each other, so \(\mathrm{Bi}\) is real there too. Independence follows from Theorem 9.70 below: one solution decays and the other grows as \(z\to+\infty\), so neither is a multiple of the other.
∎As the real variable \(z\) tends to \(+\infty\),
The first of these holds in the whole sector \(\abs{\arg z} < \pi\), with the principal branches of \(z^{3/2}\) and \(z^{1/4}\). Rests on Definition 9.68 and Proposition 9.69.
Derivation. Derives Theorem 9.70. Write the exponent of Equation (9.77) as \(\varphi(t) = t^{3}/3 - z\,t\), so that \(\varphi'(t) = t^{2}-z\) and \(\varphi''(t) = 2t\); the saddle points are \(t = \pm\sqrt{z}\). By Cauchy's theorem and the deformation corollary (Theorem 8.12 and Corollary 8.15) the contour may be moved at will, the integrand being entire, provided its ends stay in the decay sectors. The method of steepest descent chooses the deformation on which \(\Im\varphi\) is constant: the integrand then does not oscillate, its modulus has a sharp maximum at the saddle, and the whole contribution comes from a neighbourhood of it whose width shrinks as \(\abs{\varphi''}^{-1/2}\).
Positive argument. Let \(z = x > 0\). The two saddles are at \(t = \pm\sqrt{x}\), with \(\varphi(\sqrt{x}) = -\tfrac{2}{3}x^{3/2}\) and \(\varphi(-\sqrt{x}) = +\tfrac{2}{3}x^{3/2}\). The ends of \(C\) lie in the two sectors adjacent to the positive real axis, so \(C\) is deformed through the saddle at \(t = +\sqrt{x}\) alone. There \(\varphi'' = 2\sqrt{x}\) is real and positive, so putting \(t = \sqrt{x} + \ii u\) with \(u\) real makes \(\tfrac12\varphi''\,(t-\sqrt{x})^{2} = -\sqrt{x}\,u^{2}\): the vertical line through the saddle is the path of steepest descent. On it \(\dd t = \ii\,\dd u\), and
the first member of Equation (9.79). Nothing in this computation used \(z\) real: the same deformation is available, and the same result holds, for every \(z\) with \(\abs{\arg z} < \pi\).
Negative argument. Let \(z = -x\) with \(x > 0\), so that \(\varphi(t) = t^{3}/3 + x\,t\) and the saddles sit at \(t = \pm\ii\sqrt{x}\) with \(\varphi(\pm\ii\sqrt{x}) = \pm\ii\,\tfrac{2}{3}x^{3/2}\), purely imaginary: neither dominates the other, and both must be kept. The bookkeeping is topological. Call \(V_{1}\), \(V_{2}\), \(V_{3}\) the decay sectors around \(\arg t = \pi/3\), \(\pi\) and \(-\pi/3\); the contour runs from \(V_{3}\) to \(V_{1}\). At a saddle \(s\) the steepest descent direction \(\ee^{\ii\theta}\) is fixed by requiring \(\tfrac12\varphi''(s)\,\ee^{2\ii\theta}\) to be real and negative. At \(s = -\ii\sqrt{x}\), where \(\varphi'' = 2\sqrt{x}\,\ee^{-\ii\pi/2}\), this gives \(\theta = -\pi/4\) or \(\theta = 3\pi/4\), so that descent path joins \(V_{3}\) to \(V_{2}\); at \(s = +\ii\sqrt{x}\), where \(\varphi'' = 2\sqrt{x}\,\ee^{\ii\pi/2}\), it gives \(\theta = \pi/4\) or \(\theta = 5\pi/4\), joining \(V_{2}\) to \(V_{1}\). Their concatenation runs from \(V_{3}\) to \(V_{1}\) and is therefore a legitimate deformation of \(C\). Each saddle contributes \(\ee^{\varphi(s)}\,\ee^{\ii\theta}\sqrt{\pi/\sqrt{x}}\) with \(\theta\) the direction in which it is traversed — \(3\pi/4\) at the lower saddle and \(\pi/4\) at the upper one. Writing \(\zeta = \tfrac{2}{3}x^{3/2}\),
and \(\cos(\zeta-\pi/4) = \sin(\zeta+\pi/4)\), which is the second member of Equation (9.79).
The second solution. Equation (9.78) reduces both members of Equation (9.80) to the result already obtained. For \(z = x > 0\) the two arguments are \(x\,\ee^{\pm2\pi\ii/3}\), inside the sector \(\abs{\arg z} < \pi\) where the first expansion holds; there \(\bigl(x\ee^{\pm2\pi\ii/3}\bigr)^{3/2} = x^{3/2}\ee^{\pm\ii\pi} = -x^{3/2}\) and \(\bigl(x\ee^{\pm2\pi\ii/3}\bigr)^{1/4} = x^{1/4}\ee^{\pm\ii\pi/6}\), so each term of Equation (9.78) equals \(\ee^{\pm\ii\pi/6}\,\ee^{+\frac{2}{3}x^{3/2}} \bigm/ \bigl(2\sqrt{\pi}\,x^{1/4}\,\ee^{\pm\ii\pi/6}\bigr)\): the phases cancel, the two terms are equal, and their sum is the first member of Equation (9.80). For \(z = -x\) the two arguments are \(x\,\ee^{\mp\ii\pi/3}\), again inside the sector, and now \(\bigl(x\ee^{\mp\ii\pi/3}\bigr)^{3/2} = \mp\ii\,x^{3/2}\) and \(\bigl(x\ee^{\mp\ii\pi/3}\bigr)^{1/4} = x^{1/4}\ee^{\mp\ii\pi/12}\), so the two terms are \(\ee^{\pm\ii\pi/4}\,\ee^{\pm\ii\zeta} \bigm/\bigl(2\sqrt{\pi}\,x^{1/4}\bigr)\) and their sum is \(\cos(\zeta+\pi/4)\bigm/\bigl(\sqrt{\pi}\,x^{1/4}\bigr)\), the second member of Equation (9.80).
∎Each result above is the leading term of a full asymptotic expansion, and the expansion is obtained from the same contour. On the steepest-descent path the exponent is \(\varphi(s) + \tfrac12\varphi''(s)(t-s)^{2} + \tfrac16\varphi'''(s)(t-s)^{3}\) exactly, because \(\varphi\) is a cubic and \(\varphi''' = 2\); the quadratic term is the Gaussian already used, and expanding the cubic term and integrating the Gaussian moments one by one produces a series in inverse powers of \(\zeta = \tfrac23 z^{3/2}\), whose first correction to \(\mathrm{Ai}\) is the relative factor \(-5/(72\zeta)\). The series does not converge — it is asymptotic in the sense of Remark 9.67 — but truncating it after any fixed number of terms commits an error of the order of the first term omitted, uniformly on \(\abs{\arg z}\le\pi-\delta\) for each fixed \(\delta > 0\). Only the leading term is used in this chapter, and the connection formulae inherit exactly its accuracy: they are correct to relative order \(\zeta^{-1}\), which is small precisely where the semiclassical forms they connect are themselves good.
With \(a\) a simple turning point and \(F = V'(a) > 0\), the semiclassical solutions on the two sides of \(a\) are related by
the left-hand member holding for \(x > a\) and the right-hand member for \(x < a\). If instead \(F < 0\), so that the allowed region lies to the right of \(a\), the same formulae hold with the roles of the two sides interchanged and every integral taken outward from \(a\). Rests on Corollary 9.62, Equation (9.76), Equation (9.79), Equation (9.80) and Equation (9.74).
Derivation. Derives Theorem 9.72. Work with the linearised equation, for which the exact solution is a combination of \(\mathrm{Ai}(z)\) and \(\mathrm{Bi}(z)\), and evaluate the two sides of Equation (9.81) on it. For \(x > a\), that is \(z > 0\), Equation (9.61) gives \(\kappa(x) = \sqrt{2mF(x-a)} = \sqrt{2mF\ell}\;z^{1/2}\), hence
because \(\ell^{3/2} = \hbar/\sqrt{2mF}\) by Equation (9.74). So on the forbidden side the decaying semiclassical form is \(C\,(2mF\ell)^{-1/4}z^{-1/4}\ee^{-\frac{2}{3}z^{3/2}}\), which by Equation (9.79) is \(2\sqrt{\pi}\,C\,(2mF\ell)^{-1/4}\) times \(\mathrm{Ai}(z)\). For \(x < a\) put \(z = -u\) with \(u > 0\); then \(P(x) = \sqrt{2mF\ell}\;u^{1/2}\) and, by the same computation, \(\hbar^{-1}\int_{x}^{a}P\,\dd x' = \tfrac{2}{3}u^{3/2}\). The second relation of Equation (9.79) therefore gives
which is Equation (9.81). The factor \(2\) is nothing but the ratio of the two prefactors in Equation (9.79), and the phase \(\pi/4\) is the constant carried across by the same pair. Repeating the computation with \(\mathrm{Bi}\) in place of \(\mathrm{Ai}\), whose two prefactors in Equation (9.80) are equal, gives Equation (9.82) with no factor of \(2\). The case \(F<0\) is the reflection \(x \mapsto 2a - x\), which leaves Equation (9.60) invariant.
∎Equation (9.81) may be read reliably only from left to right, and Equation (9.82) only from right to left. The reason is that an exponentially small admixture of the growing solution is invisible beside the decaying one on the forbidden side, yet dominates once traced back to the allowed side; the asymptotic representation of a solution therefore changes discontinuously as the turning point is crossed in the complex plane. This is the Stokes phenomenon, and ignoring it is the commonest error in applications of the method. Every connection made below is therefore labelled with its direction: the quantisation condition of Section 9.6.7 uses only the safe one, and the single delicate step in the tunnelling derivation is flagged where it occurs and its cost accounted for in Remark 9.77.
Turning points of higher order
The linearisation Equation (9.75) presumed \(F = V'(a) \neq 0\). Where the potential meets the energy level tangentially that derivative vanishes too and the Airy problem is empty, so let \(\mu \ge 1\) be the order of contact,
\(F_{\mu}\) carrying the SI unit of an energy divided by the \(\mu\)-th power of a length, so that \(F_{1}\) is the force \(F\) of Section 9.6.4 in \(\mathrm{J}/\mathrm{m}\). The odd and even cases are genuinely different, and only the odd one is a turning point at all.
Let \(\mu\) be odd and \(F_{\mu} > 0\) in Equation (9.83), so that \(x < a\) is classically allowed and \(x > a\) forbidden. Then the decaying solution on the forbidden side connects to
The phase \(\pi/4\) is the same as in Equation (9.81) for every odd \(\mu\); only the amplitude changes, and at \(\mu = 1\) the factor is \(1/\sin(\pi/6) = 2\), which recovers Theorem 9.72. Rests on Theorem 9.72 and Equation (9.83).
Derivation. Derives Proposition 9.74. Put \(s = x-a\) and \(k = \sqrt{2mF_{\mu}}/\hbar\), a positive constant. Near \(a\), Equation (9.60) reads
and for \(s>0\) the substitution \(\zeta = 2k\,s^{q}/(\mu+2)\) with \(q = (\mu+2)/2\), together with \(\psi = \sqrt{s}\,\chi(\zeta)\), turns it by a direct computation into the modified Bessel equation of order \(\nu = 1/(\mu+2)\). A basis of local solutions is therefore \(\sqrt{s}\,I_{\pm\nu}(\zeta)\), and the decaying combination is \(\sqrt{s}\,\bigl[I_{-\nu}(\zeta) - I_{\nu}(\zeta)\bigr]\), proportional to \(\sqrt{s}\,K_{\nu}(\zeta)\). Note that \(\zeta = \hbar^{-1}\int_{a}^{x}\kappa\,\dd x'\) exactly for the local potential, since \(\kappa = \hbar k\,s^{\mu/2}\) there.
Continue to the allowed side by rotating \(s = u\,\ee^{\ii\pi}\) with \(u = a - x > 0\). Then \(\zeta = w\,\ee^{\ii\pi q}\) with \(w = 2k\,u^{q}/(\mu+2) = \hbar^{-1}\int_{x}^{a}P\,\dd x'\). In the power series
the rotation multiplies the \(j\)-th term by \(\ee^{2\pi\ii q j} = \ee^{\ii\pi(\mu+2)j} = (-1)^{j}\), because \(\mu+2\) is odd; and an alternating sign is exactly what turns the modified Bessel series into the ordinary one. Hence \(I_{\sigma}\bigl(w\,\ee^{\ii\pi q}\bigr) = \ee^{\ii\pi q\sigma}\,J_{\sigma}(w)\). With \(q\nu = \tfrac12\) and \(\sqrt{s} = \ee^{\ii\pi/2}\sqrt{u}\) the two basis solutions continue to \(\ee^{\ii\pi/2}\ee^{\pm\ii\pi/2}\sqrt{u}\,J_{\pm\nu}(w)\), that is to \(-\sqrt{u}\,J_{\nu}(w)\) and \(+\sqrt{u}\,J_{-\nu}(w)\), so the decaying solution continues to \(\sqrt{u}\,\bigl[J_{-\nu}(w) + J_{\nu}(w)\bigr]\). This much is exact.
It remains to read off the two amplitudes, and for that one classical fact is imported: the large-argument behaviour of the Bessel functions, \(K_{\nu}(\zeta) \sim \sqrt{\pi/(2\zeta)}\;\ee^{-\zeta}\) and \(J_{\sigma}(w) \sim \sqrt{2/(\pi w)}\, \cos\left(w - \sigma\pi/2 - \pi/4\right)\) [Morse:1953]. The prefactors \(\sqrt{s/\zeta}\) and \(\sqrt{u/w}\) are the same function of \(\abs{s}\) with the same constant in front, and each reproduces the semiclassical amplitude, \(\kappa^{-1/2}\) on one side and \(P^{-1/2}\) on the other, so only the pure numbers matter. On the forbidden side \(I_{-\nu} - I_{\nu} = \left(2\sin\nu\pi/\pi\right)K_{\nu}\) leaves the number \(\left(2\sin\nu\pi/\pi\right)\sqrt{\pi/2} = \sqrt{2/\pi}\,\sin\nu\pi\). On the allowed side
the two phase shifts \(\mp\nu\pi/2\) averaging away and leaving the universal \(-\pi/4\). The ratio of the two numbers is
which with \(\nu = 1/(\mu+2)\) is the factor in Equation (9.84); and \(\cos(w-\pi/4) = \sin(w+\pi/4)\) completes it.
∎For even \(\mu\) the difference \(V-E\) does not change sign at \(a\), so \(a\) is not a turning point at all: the motion is either forbidden on both sides of it or allowed on both, and there are no two regimes to connect. The rotation argument above fails at exactly that point — \(\mu+2\) is then even, \(\ee^{2\pi\ii q j} = +1\), and the modified Bessel series is not turned into the ordinary one. The instance that matters physically is \(\mu = 2\) with \(F_{2} < 0\): the top of a barrier sitting exactly at the energy of the incident particle. There the local equation is the parabolic-cylinder equation rather than Equation (9.76), and the transmission coefficient passes smoothly through the barrier top instead of jumping. That regime is outside the linear connection theory of this section, and Remark 9.77 has already recorded that \(\gamma \lesssim 1\) is precisely where the formulae derived here are not entitled to be used.
The other place where the semiclassical treatment fails for a reason that is not a simple turning point is the origin of the radial equation, where the centrifugal term diverges. That case has its own repair, and it is derived in Section 9.6.10.
Barrier penetration and the Gamow factor
Let \(V > E\) on an interval \((a,b)\) and \(V < E\) outside it, with simple turning points at \(a\) and \(b\). Write
a dimensionless number, the integrand being in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\) and \(\hbar\) in \(\mathrm{J}\,\mathrm{s}\). Then the transmission coefficient — the ratio of transmitted to incident probability current — is
which for a thick barrier, \(\gamma \gg 1\), reduces to
the correction being of relative size \(\ee^{-2\gamma}/2\). The exponent \(2\gamma\) is the Gamow factor. Rests on Corollary 9.62, Theorem 9.72 and Proposition 9.63.
Derivation. Derives Theorem 9.76. Beyond the barrier, for \(x > b\), impose a purely outgoing wave and write it with the phase measured from \(b\),
with \(\Theta(x) = \hbar^{-1}\int_{b}^{x}P\,\dd x' + \pi/4\). At \(b\) the forbidden region lies to the left, so Theorem 9.72 is used in its reflected form. The direction has to be said carefully, because “growing towards \(b\)” and “growing away from \(b\)” name opposite solutions: measured into the barrier, the way Equations (9.81) and (9.82) measure it, the \(\cos\Theta\) part comes from the exponential growing away from \(b\) and the \(\sin\Theta\) part from the one decaying away from it, with the coefficients read off Equation (9.82) and Equation (9.81) respectively. Hence inside the barrier
The first of these two readings is in the safe direction of Remark 9.73; the second is not, and it is the sole source of the order-unity ambiguity recorded in Remark 9.77 — it is the term that carries the coefficient \(\tfrac14\) into Equation (9.86), and it does not affect the exponent.
Now re-express the exponents about the other turning point using \(\int_{x}^{b}\kappa = \hbar\gamma - \int_{a}^{x}\kappa\), which turns the first term into \(\mathcal{A}\ee^{\gamma}\) times the exponential decaying away from \(a\) and the second into \((\ii \mathcal{A}/2)\ee^{-\gamma}\) times the one growing away from \(a\). At \(a\) the allowed region lies to the left, which is exactly the configuration of Theorem 9.72, so with \(\Xi(x) = \hbar^{-1}\int_{x}^{a}P\,\dd x' + \pi/4\),
Writing \(\sin\Xi = (\ee^{\ii\Xi}-\ee^{-\ii\Xi})/(2\ii)\) and \(\cos\Xi = (\ee^{\ii\Xi}+\ee^{-\ii\Xi})/2\) and collecting,
Since \(\dd\Xi/\dd x = -P/\hbar\), the term in \(\ee^{-\ii\Xi}\) advances in phase with increasing \(x\) and is the incident wave, while \(\ee^{\ii\Xi}\) is the reflected one. The amplitude \(P^{-1/2}\) has already made \(\abs{\,\cdot\,}^{2}\) proportional to the current, by Proposition 9.63, so the transmission coefficient is the ratio of the squared coefficients of the transmitted and incident waves,
which is Equation (9.86). Expanding, \(T = \ee^{-2\gamma}\bigl(1+\tfrac14\ee^{-2\gamma}\bigr)^{-2} = \ee^{-2\gamma}\bigl[1 - \tfrac12\ee^{-2\gamma} + \cdots\bigr]\).
∎Only the exponential in Equation (9.87) is a reliable output of the method. The order-unity factor multiplying it depends on how the two turning points are treated, and Equation (9.86) is not exact: the linear connection formulae assume the two turning points are far apart on the scale \(\ell\) of Equation (9.74), and a uniform treatment that does not gives \(T = \bigl(1+\ee^{2\gamma}\bigr)^{-1}\) instead. The place where the two disagree is worth being exact about, because it is the opposite of what one might guess. Both tend to \(\ee^{-2\gamma}\) with the same unit prefactor as \(\gamma\) grows, and they are numerically indistinguishable in the regime a barrier problem usually lives in: their ratio is \(1.009\) at \(\gamma = 2\), and \(1\) to ten significant figures at the \(\gamma = 33.6\) of Example 9.81. They part company only near the top of the barrier, \(\gamma \lesssim 1\), where the ratio reaches \(1.28\) at \(\gamma = 0\) — and that is precisely the regime in which the two turning points are no longer far apart on the scale \(\ell\), so the linear connection formulae were never entitled to be used there. The exponent is the whole physical content, and it is the part the method actually determines.
The exponential is the result Gamow, and independently Gurney and Condon, put to work on the escape of a charged particle from a nucleus [Gamow:1928] [Gurney:1928] [Gurney:1929]; the same expression, with a triangular barrier, is Fowler and Nordheim's account of electron emission from a cold metal in a strong field [Fowler:1928]. What this section supplies is the mathematics: a formula for the exponent, its derivation, and an honest statement of what it does and does not determine.
Bohr–Sommerfeld quantisation and the Maslov index
Let \(V < E\) on an interval \((a,b)\) and \(V > E\) outside it, with simple turning points at \(a\) and \(b\). A solution of Equation (9.59) that decays in both forbidden regions exists only for those energies satisfying
Equivalently, in terms of the action accumulated over one full classical period, with \(h = 2\pi\hbar\) the Planck constant,
Rests on Equation (9.59), Theorem 9.72 and Corollary 9.62.
Derivation. Derives Theorem 9.78. Decay in the left forbidden region, \(x < a\), is the configuration of Theorem 9.72 reflected, and connects to
inside the well. Decay in the right forbidden region, \(x > b\), connects by Equation (9.81) to
These are two expressions for the same function, so they must agree up to sign for every \(x\) in \((a,b)\). Write \(\theta(x) = \hbar^{-1}\int_{a}^{x}P\,\dd x'\) and \(\Phi = \hbar^{-1}\int_{a}^{b}P\,\dd x'\), so that \(\hbar^{-1}\int_{x}^{b}P\,\dd x' = \Phi - \theta\), and expand
For this to be a constant multiple of \(\sin(\theta + \pi/4)\) for all \(\theta\), the coefficient of the independent function \(\cos(\theta+\pi/4)\) must vanish, that is \(\cos\Phi = 0\), whence \(\Phi = (n + \tfrac12)\pi\) with \(n\) a non-negative integer, which is Equation (9.88). The remaining coefficient is then \(-\cos(\Phi+\pi/2) = \sin\Phi = (-1)^{n}\), so \(C' = (-1)^{n}C\) and the two expressions do agree. Doubling and multiplying by \(\hbar\) gives \(\oint P\,\dd x = (2n+1)\pi\hbar = (n+\tfrac12)h\), which is Equation (9.89).
∎The old quantum theory quantised the phase integral as \(\oint P\,\dd x = n h\) [Sommerfeld:1916], generalising Bohr's quantisation of the angular momentum of a circular orbit [Bohr:1913], and had no account of the additional \(h/2\). Theorem 9.78 shows that it is not an empirical patch but an accounting of turning points: each of the two simple turning points contributed a phase \(\pi/4\) through Equation (9.81), and \(2\times\pi/4 = \pi/2\) is exactly the half-integer. The integer \(\mu\) of Equation (9.89) counts those contributions — one per turning point for a one-dimensional well, so \(\mu = 2\) — and is the Maslov index. Keller identified the same index as the count of caustics traversed by the classical trajectory, and thereby extended the condition to systems of more than one degree of freedom, where the quantisation applies to each independent circuit of an invariant torus [Keller:1958]. A turning point at an infinitely steep wall, where the wave function must simply vanish, contributes \(\pi/2\) rather than \(\pi/4\) and so counts twice in \(\mu\).
For \(V(x) = \tfrac12 m\omega^{2}x^{2}\), with \(m\) in \(\mathrm{kg}\) and \(\omega\) in \(\mathrm{rad}/\mathrm{s}\), the turning points of energy \(E\) are at \(x = \pm x_{E}\) with \(x_{E} = \sqrt{2E/(m\omega^{2})}\), and
the integral being the area of an ellipse with semi-axes \(x_{E}\) and \(\sqrt{2mE}\) in the phase plane. Equation (9.89) then gives \(2\pi E/\omega = (n+\tfrac12)\,2\pi\hbar\), that is
in \(\mathrm{J}\). These are the exact eigenvalues of Equation (9.59) for this potential, zero-point energy included — an accident of the harmonic case, not a general property of the approximation, but a reassuring one. For a vibrational mode of angular frequency \(\omega = 5.0\times 10^{14}\,\mathrm{rad}/\mathrm{s}\), typical of a diatomic molecule, the spacing is \(\hbar\omega = 5.3\times 10^{-20}\,\mathrm{J} = 0.33\,\mathrm{eV}\). Rests on Theorem 9.78.
Worked example: a charged particle against a Coulomb barrier
Take a particle of charge \(z e\) and mass \(m\) inside a spherical region of radius \(R\) whose boundary it must cross against the electrostatic repulsion of a residual charge \(Z e\). Outside \(R\) the potential energy is
in \(\mathrm{J}\), with \(e = 1.602176634\times 10^{-19}\,\mathrm{C}\) exactly and \(\varepsilon_{0} = 8.8541878128\times 10^{-12}\,\mathrm{F}/\mathrm{m}\) [Tiesinga:2021]. Take the alpha particle, \(z = 2\) and \(m_{\alpha} = 6.6446573357\times 10^{-27}\,\mathrm{kg}\) [Tiesinga:2021], with \(Z = 82\), \(R = 9.0\,\mathrm{fm}\) and a kinetic energy at infinity of
The charge factor is \(zZe^{2}/(4\pi\varepsilon_{0}) = 3.7836\times 10^{-26}\,\mathrm{J}\,\mathrm{m}\), so the outer turning point, where \(V(b) = E\), sits at
five times the radius of the source region. Only \(b\) is a turning point here: at \(r = R\) the model potential is cut off discontinuously by the nuclear surface, and \(V(R) = 26.24\,\mathrm{MeV}\) is far above \(E\) (Example 9.82), so the momentum does not vanish there and no connection formula can be applied at \(R\). That puts the configuration outside the hypothesis of Theorem 9.76, which asks for simple turning points at both ends of the barrier and performs one connection at each: here the inner end is a hard wall, off which the incident amplitude is reflected rather than matched to an Airy function. What is taken from that theorem is therefore its exponent alone. The barrier integral Equation (9.85) is taken from the edge of the source region to the single turning point \(b\), while the order-unity factor multiplying \(\ee^{-2\gamma}\) — which Equation (9.86) determines only in the two-turning-point case, and there only to the accuracy Remark 9.77 records — is not fixed by anything proved here. Since \(V(r) - E = E\left(b/r - 1\right)\) by Equation (9.92), the exponent Equation (9.85) taken between \(R\) and \(b\) is \(\hbar^{-1}\sqrt{2mE}\int_{R}^{b}\sqrt{b/r - 1}\,\dd r\), and the substitution \(r = b\cos^{2}\vartheta\) evaluates it in closed form: with \(\xi = R/b = 0.1906\),
Numerically \(\sqrt{2mE} = 1.0318\times 10^{-19}\,\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), \(\arccos\sqrt{\xi} = 1.1191\), \(\sqrt{\xi(1-\xi)} = 0.3927\), and the bracket is \(0.7263\), so
dimensionless as required, since \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\times\mathrm{m}\) divided by \(\mathrm{J}\,\mathrm{s}\) is a pure number. The Gamow factor is \(2\gamma = 67.13\), and the exponential of Equation (9.87) gives
about one crossing in \(1.4\times 10^{29}\) attempts — an order-of-magnitude statement, the prefactor being undetermined here for the reason given above. Two checks are worth recording. First, letting \(R \to 0\) replaces the bracket of Equation (9.93) by \(\pi/2\) and gives \(2\gamma \to 2\pi\eta\) with \(\eta = zZe^{2}/(4\pi\varepsilon_{0}\hbar v)\) and \(v = \sqrt{2E/m} = 1.553\times 10^{7}\,\mathrm{m}/\mathrm{s}\); here \(\eta = 23.11\) and \(2\pi\eta = 145.2\), more than twice the actual exponent, so the finite size of the source region is not a detail. Second, with \(c = 299792458\,\mathrm{m}/\mathrm{s}\) exactly, \(v/c = 0.052\), which justifies using the non-relativistic Equation (9.59) at all. Rests on Theorem 9.76.
Return to the data of Example 9.81. At the inner edge \(r = R\) the potential energy is \(V(R) = 26.24\,\mathrm{MeV}\), so \(V(R) - E = 21.24\,\mathrm{MeV}\), \(\kappa(R) = 2.1266\times 10^{-19}\,\mathrm{kg}\,\mathrm{m}/\mathrm{s}\) and \(\bar\lambda(R) = \hbar/\kappa(R) = 0.496\,\mathrm{fm}\). With \(\abs{V'(R)} = zZe^{2}/(4\pi\varepsilon_{0}R^{2}) = 467\,\mathrm{J}/\mathrm{m}\), the criterion Equation (9.73) evaluates to
comfortably small. It grows without bound on approach to \(b\), where \(\abs{V'(b)} = 16.96\,\mathrm{J}/\mathrm{m}\) gives an Airy scale Equation (9.74) of \(\ell = 3.67\,\mathrm{fm}\): the semiclassical forms are worthless over a few femtometres at the outer turning point, of order a tenth of the barrier width \(b - R = 38\,\mathrm{fm}\). That region is precisely the one handled by Theorem 9.72 and not by Equation (9.70). Finally, the free momentum \(\sqrt{2mE}\) corresponds to \(\bar\lambda = 1.02\,\mathrm{fm}\) outside the barrier, small against that same width, which is the qualitative statement the criterion makes quantitative. Rests on Proposition 9.65 and Example 9.81.
The radial equation
In three-dimensional space with a potential depending only on the radius, separation of variables (Partial Differential Equations) writes a stationary state as \(\psi = u(r)\,Y(\theta,\varphi)/r\), with \(Y\) a spherical harmonic of degree \(n\) (Equation (9.163)). The azimuthal index of \(Y\) does not enter the radial equation, which is why it is not displayed here beside the mass \(m\); the radial factor \(u\) obeys
with \(u(0) = 0\). This is Equation (9.59) on the half-line with the centrifugal term added to the potential, so everything derived above applies verbatim once \(V\) is replaced by the bracket. For \(n = 0\) the bracket is \(V\) itself, which is the case used in Example 9.81. For \(n \geq 1\) the replacement is legitimate away from the origin but not at it: the centrifugal term diverges as \(r^{-2}\), the criterion Equation (9.73) fails there however smooth \(V\) is, and the naive substitution gives the wrong energies. The repair, due to Langer, is to replace \(n(n+1)\) by \(\left(n+\tfrac12\right)^{2}\) before applying the approximation, and it is derived next.
The Langer substitution
Fix a reference length \(r_{0}\) in \(\mathrm{m}\) and let \(u\) solve the radial equation Equation (9.95) on \(r>0\). Under the change of independent and dependent variable
the radial equation becomes, exactly and with no approximation,
a problem of the form Equation (9.60) on the whole line, with \(Q\) in \(\mathrm{J}\,\mathrm{s}\) and \(\varrho\) dimensionless. Its semiclassical phase integral, transported back to the radial variable, is
which is the radial phase integral with \(n(n+1)\) replaced by \(\left(n+\tfrac12\right)^{2}\). Rests on Equations (9.60) and (9.95).
Derivation. Derives Theorem 9.83. From Equation (9.96), \(\dd/\dd r = r^{-1}\dd/\dd\varrho\) with \(r = r_{0}\ee^{\varrho}\), so
the cross terms in \(\dd w/\dd\varrho\) cancelling in the second derivative, which is what the factor \(\ee^{\varrho/2}\) in Equation (9.96) was chosen to arrange. Substituting into Equation (9.95) written as
and multiplying through by \(r_{0}^{2}\,\ee^{3\varrho/2}\) gives
because \(r_{0}^{2}\ee^{2\varrho} = r^{2}\). The two constant terms combine as \(n(n+1) + \tfrac14 = \left(n+\tfrac12\right)^{2}\), which is Equation (9.97). For Equation (9.98), \(\dd\varrho = \dd r/r\), so
Both terms of \(Q^{2}\) carry the square of \(\mathrm{J}\,\mathrm{s}\) — the first because \(2m\left(E-V\right)r^{2}\) is a mass times an energy times an area, the second manifestly — so \(Q/\hbar\) is dimensionless, as an exponent per unit of \(\varrho\) must be; the arbitrary \(r_{0}\) has cancelled out of both Equation (9.97) and Equation (9.98), as it must, since shifting it only translates \(\varrho\).
∎Three things are settled at once by Theorem 9.83. First, the half-line \(r>0\), whose endpoint at the origin is singular, becomes the whole line, on which \(Q^{2}\) tends to the finite constant \(-\hbar^{2}\left(n+\tfrac12\right)^{2}\) as \(\varrho \to -\infty\): the coefficient no longer diverges, the criterion Equation (9.73) is satisfied there, and the semiclassical solution \(w \sim \exp\left[\pm\left(n+\tfrac12\right)\varrho\right]\) reproduces through Equation (9.96) the exact behaviours \(u \sim r^{n+1}\) and \(u \sim r^{-n}\) of Equation (9.95) at the origin — which the untransformed approximation does not, since with \(n(n+1)\) in place of \(\left(n+\tfrac12\right)^{2}\) the same computation returns \(u \sim r^{1/2\pm\sqrt{n(n+1)}}\), a pair of irrational exponents that agree with the exact ones only in the limit of large \(n\). Second, the transformation is exact, so nothing has been assumed: the substitution \(n(n+1) \to \left(n+\tfrac12\right)^{2}\) is not a patch applied to the answer but the statement that the approximation had been made in the wrong variable. Third, the turning points of the transformed problem are simple for a potential of the usual kind, so Theorem 9.72 applies to it unchanged and Equation (9.88) may be used with the phase integral Equation (9.98). The Coulomb potential is the standard check [Morse:1953]: with the corrected coefficient the quantisation condition returns the exact hydrogen levels, and with \(n(n+1)\) it does not. Rests on Theorem 9.83.
Most treatments of this material set \(\hbar = 1\) and write the semiclassical exponent as \(\int P\,\dd x\) and the Gamow factor as \(2\int\kappa\,\dd x\). This treatise does not, for the reason given after Equation (9.64): \(\hbar\) is the expansion parameter, and a convention that hides it hides the structure of the method. A reader coming from such a text restores the SI form of every expression above by the single rule that each exponent must be dimensionless: divide any integral of a momentum over a length by \(\hbar\), in \(\mathrm{J}\,\mathrm{s}\), and any energy by \(\hbar\) before comparing it with an angular frequency. Momenta are then in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), lengths in \(\mathrm{m}\) and energies in \(\mathrm{J}\), and no formula in this section needs any other change.
Boundary layers and matched asymptotic expansions
The semiclassical expansion is one half of the theory of a small parameter multiplying the highest derivative. It is the half in which the reduced problem is oscillatory and the small parameter sets a wavelength. The other half is the one in which the reduced problem is of lower order and cannot meet all the boundary conditions: the neglected derivative then reasserts itself in a thin region, a boundary layer, whose thickness is set by the small parameter and inside which the solution turns rapidly to meet the condition the outer solution had to abandon. The technique for handling that case is the same in outline — expand, and match — and is set out here because the physical parts need it and because a reader who has followed Section 9.6 has already done the harder half.
Consider a problem containing a small parameter \(\varepsilon>0\) on a fixed interval. An outer expansion is a series in \(\varepsilon\) solving the problem with \(\varepsilon\) set to zero at leading order, valid where the solution varies on the scale of the interval. An inner expansion is a series solving the problem rewritten in a stretched variable \(\xi = (x-x_{c})/\delta(\varepsilon)\) about a point \(x_{c}\), with \(\delta\to0\); the choice of \(\delta\) for which the rewritten equation retains its highest derivative in a non-trivial balance is the distinguished limit. The two are joined by the matching condition
which says that the inner solution far out and the outer solution close in describe the same thing, and the composite approximation is
the subtraction removing the part they both already carry. Rests on Definitions 7.16 and 9.1.
Let \(a\) and \(b\) be continuous on \([0,1]\) with \(a(x) \ge a_{0} > 0\), and consider
Then (i) the only stretching \(x = x_{c}+\delta\xi\) under which the second-derivative term is retained in a balance with a term at least as large as every neglected one is \(\delta = \varepsilon\); (ii) at that scaling the leading inner equation is \(\dd^{2}Y/\dd\xi^{2} + a(x_{c})\,\dd Y/\dd\xi = 0\), whose solutions are constants and \(\ee^{-a(x_{c})\xi}\); and (iii) that exponential decays away from \(x_{c}\) into the interval only at \(x_{c}=0\). The layer therefore sits at the left endpoint, and the outer solution — the solution of the first-order reduced equation \(a\,y'+b\,y=0\) — must carry the condition at \(x=1\). Rests on Equation (9.101), Theorem 9.23 and Proposition 9.18.
Derivation. Derives Proposition 9.87. Under \(x = x_{c}+\delta\xi\) the three terms of Equation (9.101) carry the factors \(\varepsilon/\delta^{2}\), \(1/\delta\) and \(1\). Suppose the first is retained. If it balances the third, \(\varepsilon/\delta^{2}\sim1\) and \(\delta\sim\sqrt{\varepsilon}\), whereupon the second term is of order \(\varepsilon^{-1/2}\) and dominates both retained terms — so the balance is inconsistent. If it balances the second, \(\varepsilon/\delta^{2}\sim1/\delta\), that is \(\delta\sim\varepsilon\), and the third term, of order \(1\), is smaller by a factor \(\varepsilon\) than the two retained ones: consistent. That is (i), and substituting \(\delta=\varepsilon\) and \(a(x_{c}+\varepsilon\xi)\to a(x_{c})\) gives (ii), whose solutions are read off Theorem 9.23 for the constant-coefficient equation — or directly from Proposition 9.18 applied to \(\dd Y/\dd\xi\).
For (iii): at \(x_{c}=0\) the interval is \(\xi\ge0\) and \(\ee^{-a(0)\xi}\to0\) as \(\xi\to\infty\), so the inner solution settles to a constant and Equation (9.99) can be imposed. At \(x_{c}=1\) the interval is \(\xi\le0\) and \(\ee^{-a(1)\xi}\to\infty\) as \(\xi\to-\infty\): the inner solution grows without bound where it should match, so no matching is possible and there is no layer there. The reduced equation is of first order and therefore accepts exactly one boundary condition, which must then be the one at \(x=1\).
∎Let \(A,B\in\R\) and consider Equation (9.101) with \(a\equiv1\), \(b\equiv1\): \(\varepsilon y'' + y' + y = 0\), \(y(0)=A\), \(y(1)=B\). Its outer solution is \(y_{\mathrm{out}}(x) = B\,\ee^{1-x}\), its inner solution is \(y_{\mathrm{in}}(\xi) = \bigl(A - B\ee\bigr)\ee^{-\xi} + B\ee\), and the composite Equation (9.100) is
There is a constant \(C\), depending only on \(\abs{A}\) and \(\abs{B}\), with
for every \(\varepsilon \in (0,1/8]\). The approximation is therefore uniform: it is correct both inside the layer and outside it. Rests on Proposition 9.87, Proposition 9.28 and Theorem 7.38.
Derivation. Derives Proposition 9.88. The equation has constant coefficients, so by Proposition 9.28 its solutions are combinations of \(\ee^{\lambda_{\pm}x}\) with \(\lambda_{\pm} = \bigl[-1\pm\sqrt{1-4\varepsilon}\bigr]/(2\varepsilon)\), real and distinct for \(\varepsilon<1/4\). Taylor's theorem (Theorem 7.38) applied to \(\sqrt{1-u}\) at \(u=4\varepsilon\) gives \(\abs{\sqrt{1-4\varepsilon}-(1-2\varepsilon)} \le c_{0}\varepsilon^{2}\) on \((0,1/8]\), whence
In particular \(\lambda_{-}\le-1/(2\varepsilon)\), so \(\eta = \ee^{\lambda_{-}}\le\ee^{-1/(2\varepsilon)}\) is smaller than any power of \(\varepsilon\).
Write \(y = C_{+}\ee^{\lambda_{+}x} + C_{-}\ee^{\lambda_{-}x}\). The boundary conditions give \(C_{+}+C_{-}=A\) and \(C_{+}\ee^{\lambda_{+}}+C_{-}\eta = B\), so \(C_{+} = (B-A\eta)/(\ee^{\lambda_{+}}-\eta)\). Since \(\ee^{\lambda_{+}} = \ee^{-1}\ee^{r} = \ee^{-1}\bigl(1+O(\varepsilon) \bigr)\) by Equation (9.104) and \(\eta\) is negligible beside it, \(C_{+} = B\ee + O(\varepsilon)\) and hence \(C_{-} = A - B\ee + O(\varepsilon)\), every implied constant depending only on \(\abs{A}\), \(\abs{B}\) and \(c_{0}\).
Now compare term by term. For the outer term, on \(0\le x\le1\),
using \(\abs{\ee^{rx}-1}\le\abs{r}x\ee^{\abs{r}}\) and Equation (9.104). For the layer term, writing \(\ee^{\lambda_{-}x} = \ee^{-x/\varepsilon}\ee^{sx}\) and using \(\ee^{\lambda_{-}x}\le1\) for \(x\ge0\),
The first term is \(O(\varepsilon)\). In the second, \(\abs{\ee^{sx}-1}\le\abs{s}x\ee^{\abs{s}}\le2\ee^{2}x\) on \([0,1]\), and \(x\,\ee^{-x/\varepsilon}\) attains its maximum \(\varepsilon/\ee\) at \(x=\varepsilon\); so the second term is \(O(\varepsilon)\) too. Adding the two comparisons gives Equation (9.103).
∎Section 9.6 and this subsection treat the same syntactic situation — a small parameter on the highest derivative — and produce different objects, because the sign of the coefficient of the undifferentiated term decides whether the neglected derivative buys oscillation or a layer. Where the reduced problem is oscillatory the rapid variation is spread over the whole interval and the expansion is in the phase, which is Corollary 9.62; where the reduced problem is of lower order the rapid variation is confined to a region of thickness \(\varepsilon\) and the expansion is the two-region one of Definition 9.86. The turning point of Theorem 9.72 is the same matching problem in the first setting: an inner solution — there the Airy function of Definition 9.68 — joined to two outer ones by their common asymptotics.
Nothing above is specific to one dimension or to linear equations. The reduction of the equations of a viscous fluid to Prandtl's boundary-layer equations [Prandtl:1904] is Proposition 9.87 carried out on a partial differential equation, with the Reynolds number in place of \(1/\varepsilon\): the inviscid outer flow cannot satisfy the no-slip condition at a wall, the viscous term is restored in a layer whose thickness the same balance fixes, and Equation (9.99) is what hands the outer pressure gradient to the inner problem. Experiment: Fluid Flow and Turbulence carries that computation and its measurement.
The one-dimensional Helmholtz equation
Sturm–Liouville data
The functions \(p\), \(q\) and \(r\) of the operator are
with \(\lambda^2\) the free eigenvalue parameter.
Homogeneous solution
The homogeneous equation becomes
whose solution is
Using Euler's formula, \(\ee^{\pm\ii t} = \cos t \pm \ii \sin t\), the solution Equation (9.107) may equivalently be written in the real form
Orthonormality
On the interval \([-\pi, \pi]\), and for integer values \(\lambda = n \in \Z\), the solution functions satisfy
Equation (9.109) is Lemma 17.2, proved there by direct integration: the integrand is \(\ee^{\ii(m-n)x}\), which equals \(1\) when \(m = n\) and, when \(m \neq n\), has an antiderivative that returns to its starting value over a full period, so the integral vanishes.
The factor \(1/(2\pi)\) in Equation (9.109) is not cosmetic, and it is worth placing once. Against the plain \(\dd x\) of Equation (9.46) the bare exponentials are merely orthogonal: \(\int_{-\pi}^{\pi}\overline{\ee^{\ii n x}}\;\ee^{\ii n x}\,\dd x = 2\pi\), not \(1\), so \(\set{\ee^{\ii n x}}\) is not an orthonormal set in the sense of Definition 9.45. Two readings repair this, and they are one statement written twice: either keep \(\dd x\) and take the normalised functions \(\phi_n(x) = \ee^{\ii n x}/\sqrt{2\pi}\), or keep the bare exponentials \(\phi_n(x) = \ee^{\ii n x}\) and read every integral of Section 9.4 — Equation (9.46), the coefficient formula Equation (9.48) and the triple overlap Equation (9.49) alike — against the normalised measure \(\dd x/(2\pi)\). The Fourier sections of this chapter take the second reading throughout, which is exactly what puts the \(1/(2\pi)\) into Equation (9.111) and the \(1/L\) into Equation (9.113). The choice cancels out of every expansion, since the factor a coefficient gains is the one its basis function loses; what must not happen is the two readings being mixed inside one calculation.
Fourier series
Expanding a function on \([-\pi,\pi]\) in the complex exponentials — orthonormal in the second sense of Remark 9.91 — gives the Fourier series
an instance of the general coefficient formula Equation (9.48).
Under a change of interval to an arbitrary interval \([a,b]\) of length \(L = b - a\), the series and its coefficients become
Trigonometric form
On the symmetric interval \([-L, L]\) the cosine and sine functions satisfy the orthonormality relations
with \(m, n \geq 1\), and the Fourier series takes the real form
The relations Equations (9.114) to (9.116) are proved at Equation (17.9), where the product formulas of trigonometry reduce each of them to Lemma 17.2: the identity \(\cos A\cos B = \tfrac12[\cos(A-B)+\cos(A+B)]\) turns the first into \(\delta_{mn}\), since the second index coincidence \(m = -n\) is impossible for positive \(m\) and \(n\); the sine relation is the same identity with the opposite sign; and the mixed integral vanishes as that of an odd function over a symmetric interval. The coefficient formulas Equations (9.118) and (9.119) are Equation (17.7) written with \(T = 2L\).
The Fourier transform
Consider a function that can be expanded in the Fourier series Equation (9.112). If \(L \to +\infty\), the sum over the discrete wavenumbers \(k_n\) passes into a Riemann integral, and one obtains
where \(g(k)\), called the Fourier transform of \(f(x)\), is defined by
The Riemann-sum argument sketched above is carried out at Equation (17.33), in the same normalisation: the coefficients of the series on \([-L,L]\) are \(\sqrt{2\pi}\) times the transform sampled at the harmonics \(k_n\), divided by the interval length, and the mesh of the sum is \(\Delta k = 2\pi/L \to 0\). That argument is heuristic, because the summand depends on \(L\) as well as on the mesh; the inversion formula Equation (9.120) is established properly, by Gaussian regularisation, in Theorem 17.32, and the injectivity it implies is Corollary 17.33.
Properties
Transform of derivatives.
If \(f\) is \(n\) times differentiable, \(f^{(n)}\) is absolutely integrable, and every derivative below the top one vanishes at infinity,
then the transform \(g_n\) of the \(n\)-th derivative \(f^{(n)}\) satisfies
The case \(n = 1\) is Equation (17.38) of Proposition 17.30, proved there by a single integration by parts: the boundary term \([f\,\ee^{-\ii k x}]_{-\infty}^{\infty}\) vanishes because \(f\) tends to zero, and what is left is \(\ii k\) times the transform of \(f\). The general \(n\) follows by induction, applying that result to \(f^{(j)}\) in turn — which is exactly why the hypothesis above is imposed on every derivative below the top one and not on \(f\) alone: each step of the induction consumes one boundary term. This is the property that turns a constant-coefficient differential operator into a polynomial and is the reason the transform is a method at all.
Convolution.
The convolution of two functions and the theorem that its transform is the product of theirs are Definition 17.50 and Theorem 17.53: in the normalisation of Equation (9.121) the transform of \(f * g\) is \(\sqrt{2\pi}\) times the product of the transforms, and dually the transform of a product is \((2\pi)^{-1/2}\) times the convolution of the transforms. The derivation is a single application of Fubini's theorem to the double integral, justified by the bound Equation (17.64).
Parseval relation.
The transform preserves the norm,
Equation (9.123) is Plancherel's theorem, Theorem 17.36: the transform of Equation (9.121) is an isometry on the smooth rapidly decreasing functions, by Theorem 17.32 and an exchange of integrals, and extends uniquely to a unitary map of the square-integrable functions onto themselves. The statement for the Fourier series — the completeness of the harmonics, and hence Equation (9.52) for that system — is Theorem 17.34.
The Bessel equation
Sturm–Liouville data
The identification of the operator data is
General solution
The homogeneous equation becomes
whose solution is
with \(J_\nu\) the Bessel functions of the first kind and \(N_\nu\) (also written \(Y_\nu\)) the Neumann functions, the Bessel functions of the second kind [Jackson:1999].
Expanding the derivative in Equation (9.125) and multiplying by \(x\) puts the equation into its familiar form \(x^2 y'' + x y' + (x^2 - \nu^2)\,y = 0\).
For \(\nu \ge 0\) and \(x>0\),
a series convergent for every \(x\) by the ratio test (Proposition 7.48). For non-integer \(\nu\) the same formula with \(\nu\) replaced by \(-\nu\) defines a second, independent solution \(J_{-\nu}\). At integer \(\nu = n\) the reciprocal gamma function vanishes at the non-positive integers, so the first \(n\) terms of that series drop out and \(J_{-n} = (-1)^{n}J_{n}\) is not independent; the second solution is then the Neumann function \(N_{\nu} = \bigl[J_{\nu}\cos\nu\pi - J_{-\nu}\bigr]/\sin\nu\pi\), read at integer order as the limit of that expression.
\(J_{\nu}\) of Equation (9.127) satisfies Equation (9.125). Rests on Definition 9.93 and Equation (9.125).
Derivation. Derives Proposition 9.94. Use the expanded form of Remark 9.92 and substitute Equation (9.127) term by term, which is legitimate inside the disc of convergence (Theorem 7.51). Write \(J_{\nu} = \sum_{j}c_{j}x^{2j+\nu}\) with \(c_{j} = (-1)^{j}\bigl[j!\,\Gamma(j+\nu+1)\,2^{2j+\nu}\bigr]^{-1}\). The operator \(x^{2}\dd^{2}/\dd x^{2} + x\,\dd/\dd x - \nu^{2}\) sends \(x^{2j+\nu}\) to \(\bigl[(2j+\nu)^{2}-\nu^{2}\bigr]x^{2j+\nu} = 4j(j+\nu)\,x^{2j+\nu}\), while the remaining term \(x^{2}J_{\nu}\) shifts the index by one. The coefficient of \(x^{2j+\nu}\) in the whole expression is therefore \(4j(j+\nu)\,c_{j} + c_{j-1}\), and this vanishes: by \(j!\,\Gamma(j+\nu+1) = j\,(j-1)!\,(j+\nu)\,\Gamma(j+\nu)\) and \(2^{2j+\nu} = 4\cdot 2^{2(j-1)+\nu}\),
The term \(j = 0\) has no partner and vanishes on its own, because \(4j(j+\nu) = 0\) there.
∎Recurrence relations
For every order \(\nu\) and every \(x>0\),
Rests on Definition 9.93.
Derivation. Derives Proposition 9.95. Both are term-by-term computations on Equation (9.127). For Equation (9.128),
whose derivative is \(\sum_{j\ge1}(-1)^{j}\,2j\,x^{2j-1} \bigl[j!\,\Gamma(j+\nu+1)2^{2j+\nu}\bigr]^{-1}\); putting \(j = i+1\) and cancelling the factor \(2(i+1)\) against \((i+1)!\) turns this into \(-\sum_{i\ge0}(-1)^{i}x^{2i+1} \bigl[i!\,\Gamma(i+\nu+2)\,2^{2i+\nu+1}\bigr]^{-1}\), which is \(-x^{-\nu}J_{\nu+1}(x)\) by Equation (9.127) at order \(\nu+1\). For Equation (9.129), the derivative of \(x^{\nu}J_{\nu} = \sum_{j}(-1)^{j}x^{2j+2\nu} \bigl[j!\,\Gamma(j+\nu+1)2^{2j+\nu}\bigr]^{-1}\) carries the factor \(2j+2\nu = 2(j+\nu)\), which cancels against \(\Gamma(j+\nu+1) = (j+\nu)\Gamma(j+\nu)\) and lowers the power of two by one, giving \(\sum_{j}(-1)^{j}x^{2j+2\nu-1} \bigl[j!\,\Gamma(j+\nu)2^{2j+\nu-1}\bigr]^{-1} = x^{\nu}J_{\nu-1}(x)\).
∎For every order \(\nu\),
In particular \(\dv{J_{\nu}}{x} = -J_{\nu+1}\) at every zero of \(J_{\nu}\). Rests on Proposition 9.95.
Derivation. Derives Corollary 9.96. Carrying out the differentiations on the left of Equations (9.128) and (9.129) and dividing by \(x^{-\nu}\) and \(x^{\nu}\) respectively gives
Subtracting gives Equation (9.130) and adding gives Equation (9.131). At a zero of \(J_{\nu}\) the first of the two displayed identities reduces to \(\dd J_{\nu}/\dd x = -J_{\nu+1}\).
∎Orthogonality
Equation (9.124) reads Equation (9.125) as a Sturm–Liouville equation in which the order \(\nu\) carries the spectral parameter. The orthogonality relation below belongs to the other reading of the same equation, the one the applications use: fix \(\nu\), and let \(y(x) = J_{\nu}(\alpha x/a)\) on \([0,a]\). Rescaling Equation (9.125) gives
which is Equation (9.54) with \(p(x) = x\), \(q(x) = -\nu^{2}/x\), weight \(r(x) = x\) and eigenvalue \(\lambda^{2} = \alpha^{2}/a^{2}\). It is that spectral parameter, and not \(\nu\), which distinguishes the functions appearing in Equation (9.133).
Let \(\nu \ge 0\) and let \(\alpha_{\nu m}\) denote the \(m\)-th positive zero of \(J_\nu\). On the interval \([0, a]\) the rescaled Bessel functions are orthogonal with weight \(x\):
Rests on Theorem 9.57, Corollary 9.96 and Equation (9.132).
Derivation. Derives Theorem 9.98. Take \(u(x) = J_{\nu}(\lambda x)\) and \(v(x) = J_{\nu}(\mu x)\) with \(\lambda,\mu > 0\), both solutions of Equation (9.132) with \(p(x) = x\), weight \(x\) and eigenvalues \(\lambda^{2}\) and \(\mu^{2}\). Integrating the Lagrange identity Equation (9.56) over \([0,a]\) gives
the prime denoting the derivative of \(J_{\nu}\) with respect to its own argument. The contribution at \(x = 0\) vanishes because \(p(x) = x\) does while \(J_{\nu}\) behaves there as \(x^{\nu}\) by Equation (9.127) — the second mechanism of Remark 9.58. Taking \(\lambda = \alpha_{\nu m}/a\) and \(\mu = \alpha_{\nu n}/a\) with \(m \neq n\) makes both Bessel values on the right vanish while the factor on the left does not, which is the off-diagonal statement; it is Theorem 9.57 with the boundary condition \(y(a)=0\), the first mechanism of Remark 9.58.
For the diagonal, fix \(\lambda = \alpha_{\nu m}/a\), keep \(\mu\) free and use \(J_{\nu}(\lambda a) = 0\) in Equation (9.134):
Let \(\mu \to \lambda\). The numerator vanishes like \(J_{\nu}(\mu a) = (\mu-\lambda)\,a\,J_{\nu}'(\lambda a) + \ldots\) and the denominator like \(-(\mu-\lambda)\,2\lambda\), so the quotient tends to \(\tfrac12 a^{2}\bigl[J_{\nu}'(\lambda a)\bigr]^{2}\). Finally \(J_{\nu}'(\alpha_{\nu m}) = -J_{\nu+1}(\alpha_{\nu m})\) by Corollary 9.96, and squaring removes the sign.
∎Generating function
For integer order the Bessel functions of the first kind are generated by
Equation (9.135) holds for every \(x\) and every complex \(t \neq 0\). Rests on Definition 9.93.
Derivation. Derives Proposition 9.99. Both factors of \(g(x,t) = \ee^{xt/2}\,\ee^{-x/(2t)}\) are entire in their arguments, so the two exponential series may be multiplied out and the resulting double series rearranged absolutely (Proposition 7.49):
the outer index being \(n = j-k\) and the inner sum running over those \(k \ge 0\) with \(k+n \ge 0\). Comparing with Equation (9.127) at integer order, where \(\Gamma(k+n+1) = (k+n)!\) and the reciprocal factorial is zero whenever \(k+n<0\), the inner sum is exactly \(J_{n}(x)\), which is Equation (9.135).
∎The Fourier–Bessel series
Let \(\nu \ge 0\), and let \(f\) on \([0,a]\) admit an expansion
convergent in the mean with respect to the weight \(x\). Then
Rests on Theorem 9.98 and Definition 9.47.
Derivation. Derives Proposition 9.100. Multiply Equation (9.136) by \(J_{\nu}(\alpha_{\nu n}x/a)\,x\) and integrate over \([0,a]\). Mean convergence permits the exchange of sum and integral, and Equation (9.133) leaves the single term \(m = n\) carrying the constant \(a^{2}\left[J_{\nu+1}(\alpha_{\nu n})\right]^{2}/2\). This is the argument of Proposition 9.48 run for a set that is orthogonal but not normalised, the normalising constant being divided out at the end instead of absorbed into the functions.
∎That an expansion Equation (9.136) exists at all — that the system \(\set{J_{\nu}(\alpha_{\nu m}x/a)}_{m\ge1}\) is complete in the square-integrable functions on \([0,a]\) with weight \(x\) — is a statement about the Sturm–Liouville problem of Remark 9.97, not about Bessel functions in particular, and it is not proved here; the general statement belongs to Hilbert Spaces [Courant:1962]. What Proposition 9.100 does establish is that if the expansion exists then it is this one — the same division of labour as in Remark 9.53.
The zeros
The orthogonality relation of Theorem 9.98 and the Fourier–Bessel series of Section 9.8.6 are indexed by the positive zeros of \(J_{\nu}\), and every eigenfrequency of a circular membrane or a cylindrical cavity is one of them. Their existence, their simplicity and their spacing are therefore not decoration.
On \(x>0\) the substitution \(y(x) = x^{-1/2}u(x)\) turns Equation (9.125) into
The first-derivative term is gone, and for large \(x\) the equation is a vanishing perturbation of \(u''+u=0\). Rests on Equation (9.125) and Remark 9.92.
Derivation. Derives Proposition 9.102. Use the expanded form \(x^{2}y''+xy'+(x^{2}-\nu^{2})y = 0\) of Remark 9.92. With \(y = x^{-1/2}u\),
so \(x^{2}y'' = \tfrac34x^{-1/2}u - x^{1/2}u' + x^{3/2}u''\) and \(xy' = -\tfrac12x^{-1/2}u + x^{1/2}u'\). The terms in \(u'\) cancel, and adding \((x^{2}-\nu^{2})x^{-1/2}u\) leaves
which is Equation (9.138) after division by \(x^{3/2}\), since \(\tfrac34-\tfrac12 = \tfrac14\).
∎Let \(\nu\ge0\). Every zero of \(J_{\nu}\) in \(x>0\) is simple, and there are infinitely many of them; writing \(j_{\nu,n}\) for the \(n\)-th in increasing order,
Successive zeros are therefore asymptotically \(\pi\) apart, whatever the order. Rests on Proposition 9.102, Corollary 9.9 and Corollary 9.96.
Derivation. Derives Proposition 9.103. Simplicity. On \(x>0\) the Bessel equation is regular — its leading coefficient \(x^{2}\) does not vanish — so Corollary 9.9 applies: a solution with \(y(x_{\ast}) = y'(x_{\ast}) = 0\) at some \(x_{\ast}>0\) is identically zero. Since \(J_{\nu}\) is not, no positive zero of it can be a zero of its derivative. (Equivalently, by Corollary 9.96, \(J_{\nu}' = -J_{\nu+1}\) at a zero of \(J_{\nu}\), and the two functions have no common positive zero.)
Location. Proposition 9.102 writes \(J_{\nu}(x) = x^{-1/2}u(x)\) with \(u\) solving an equation that tends to \(u''+u=0\): the solutions oscillate with asymptotic half-period \(\pi\), and the classical large-argument form quoted in Section 9.6.5 from [Morse:1953],
makes the phase explicit. The cosine vanishes where \(x - \nu\pi/2 - \pi/4 = \pi/2 + (n-1)\pi\), that is at \(x = (n+\nu/2-\tfrac14)\pi\), and it crosses each of those points with slope of modulus \(1\). A perturbation of size \(O(1/x)\) of a function crossing zero with unit slope displaces the crossing by \(O(1/x)\), and the amplitude \(\sqrt{2/\pi x}\), being non-zero, cannot create or destroy a crossing; so for large \(n\) there is exactly one zero of \(J_{\nu}\) within \(O(1/n)\) of each of those points, which is Equation (9.139), and there are infinitely many of them.
∎Equation (9.140) is not proved in this treatise; it is the same imported fact that Section 9.6.5 uses to fix the connection coefficients at a turning point of higher order, and it is flagged there for the same reason. Only the constant \(-\pi/4\) in the phase depends on it. The \(\pi\) spacing — which is what converts a count of nodal circles into a frequency, and is all that Experiment: Waves and Acoustics needs for the square law of a circular plate — follows from Equation (9.138) alone, since the coefficient there tends to \(1\) and the oscillation of \(u''+u=0\) has half-period exactly \(\pi\).
The modified Bessel functions
Replacing \(x\) by \(\ii x\) in Equation (9.125) reverses the sign of the term that makes the solutions oscillate, and produces the equation obeyed by the radial part of a field that decays rather than propagates — the evanescent interior of a waveguide below cutoff, the screened potential around a charge, the edge-localised part of the deflection of a vibrating plate. This treatise has already used its solutions, in Section 9.6.5; they are defined here.
The modified Bessel equation of order \(\nu\) is
equivalently \(x^{2}y''+xy'-(x^{2}+\nu^{2})y = 0\). Its Sturm–Liouville data in the convention of Definition 9.54 are \(p(x) = x\), \(q(x) = -x\), \(\lambda^{2} = -\nu^{2}\), \(r(x) = 1/x\). The modified Bessel function of the first kind is
the series of Equation (9.127) with the alternating sign removed, so that \(I_{\nu}(x) = \ii^{-\nu}J_{\nu}(\ii x)\). The modified Bessel function of the second kind is
read at integer order as the limit of that expression, exactly as \(N_{\nu}\) is in Definition 9.93. Rests on Definitions 9.54 and 9.93.
\(I_{\nu}\) of Equation (9.142) converges for every \(x\) and solves Equation (9.141). For \(\nu\ge0\) it is positive and strictly increasing on \(x>0\), with \(I_{\nu}(0^{+}) = 0\) for \(\nu>0\) and \(I_{0}(0) = 1\); in particular it has no positive zero, so it carries no oscillation and no eigenvalue condition. Together with \(K_{\nu}\) it spans the solution space of Equation (9.141). Rests on Definition 9.105, Proposition 9.94 and Theorem 9.13.
Derivation. Derives Proposition 9.106. Convergence is the ratio test (Proposition 7.48) exactly as for Equation (9.127), the moduli of the terms being the same. For the equation, run the computation of Proposition 9.94 with \(c_{j} = \bigl[j!\,\Gamma(j+\nu+1)2^{2j+\nu}\bigr]^{-1}\) and with the sign of the term \(x^{2}y\) reversed: the coefficient of \(x^{2j+\nu}\) becomes \(4j(j+\nu)c_{j} - c_{j-1}\), and the identity \(j!\,\Gamma(j+\nu+1) = j\,(j-1)!\,(j+\nu)\,\Gamma(j+\nu)\) used there gives \(4j(j+\nu)c_{j} = c_{j-1}\), so it vanishes; the term \(j=0\) vanishes on its own.
Every coefficient in Equation (9.142) is strictly positive for \(\nu\ge0\), and so is every term for \(x>0\); the same is true of the term-by-term derivative, legitimate inside the radius of convergence (Theorem 7.51). Hence \(I_{\nu}>0\) and \(I_{\nu}'>0\) on \(x>0\), which forbids a zero there. The leading power is \((x/2)^{\nu}/\Gamma(\nu+1)\), giving the stated behaviour at the origin. Finally \(I_{\nu}\) and \(K_{\nu}\) are independent — one is bounded at infinity and the other is not, by Remark 9.107 — so by Theorem 9.13 they span the two-dimensional solution space.
∎The large-argument behaviour
is quoted from [Morse:1953]; the second half of it is the fact already imported in Section 9.6.5, so nothing new is being taken on credit. The qualitative half of Equation (9.144) — that \(I_{\nu}\) grows without bound and \(K_{\nu}\) decays — does not depend on the import: \(I_{\nu}\) is a positive increasing function whose terms already exceed those of \(\ee^{x}/2^{\nu}\) term by term for \(\nu\) an integer, and \(K_{\nu}\) is the solution that Proposition 9.106 leaves over. This is what licenses the reading used in the physical parts: in a superposition \(A\,J_{m}(kr) + B\,I_{m}(kr)\) on a disc, the second term is negligible except within a few \(1/k\) of the rim, so the interior nodal pattern is that of \(J_{m}\) alone.
The Legendre equation and spherical harmonics
The Legendre equation
Sturm–Liouville data
The identification of the operator data is
This is the source's identification and it is the one reproduced in Table 9.2, but with \(\lambda^{2} = 0\) it is not an eigenvalue problem: nothing in it distinguishes one degree from another. Theorem 9.111 therefore reads the same equation the other way round, with \(q = 0\) and \(\lambda^{2} = n(n+1)\), exactly as Remark 9.97 does for the Bessel equation.
Homogeneous solution
The homogeneous equation becomes
whose solution is
with \(P_n\) the Legendre polynomials and \(Q_n\) the Legendre functions of the second kind.
The Legendre polynomials
The Legendre polynomial of degree \(n\) is
\(P_{n}\) of Equation (9.148) is a polynomial of degree \(n\) with leading coefficient \((2n)!/\bigl[2^{n}(n!)^{2}\bigr]\), it satisfies \(P_{n}(1) = 1\), and it solves the Legendre equation Equation (9.146). Rests on Definition 9.108 and Equation (9.146).
Derivation. Derives Proposition 9.109. Write \(v = \left(x^{2}-1\right)^{n}\), a polynomial of degree \(2n\) whose \(n\)-th derivative has degree \(n\); the coefficient of \(x^{n}\) comes from \(\dd^{n}x^{2n}/\dd x^{n} = \bigl[(2n)!/n!\bigr]x^{n}\), which after the division by \(2^{n}n!\) gives the stated leading coefficient. For the value at \(1\), factor \(v = (x-1)^{n}(x+1)^{n}\): in the Leibniz expansion of \(\dd^{n}v/\dd x^{n}\) every term retains at least one factor \(x-1\) except the one in which all \(n\) derivatives fall on \((x-1)^{n}\), so at \(x = 1\) only \(n!\,(x+1)^{n} = n!\,2^{n}\) survives, and division by \(2^{n}n!\) gives \(P_{n}(1) = 1\).
For the equation, start from the identity \(\left(x^{2}-1\right)\dv{v}{x} = 2n\,x\,v\), which is immediate from the definition of \(v\), and differentiate it \(n+1\) times by the Leibniz rule Equation (7.16). On the left only three terms survive, because \(x^{2}-1\) is quadratic:
and on the right only two, because \(x\) is linear: \(2n\,x\,v^{(n+1)} + 2n(n+1)\,v^{(n)}\). Writing \(u = v^{(n)}\) and collecting,
which on multiplication by \(-1\) is \(\bigl[(1-x^{2})u'\bigr]' + n(n+1)u = 0\), that is Equation (9.146). Since \(P_{n}\) is a constant multiple of \(u\), it solves the same linear equation.
∎Every polynomial solution of Equation (9.146) is a constant multiple of \(P_{n}\). Rests on Proposition 9.109.
Derivation. Derives Lemma 9.110. Let \(\Lambda[y] = \bigl[(1-x^{2})y'\bigr]'\). On monomials
so \(\Lambda\) maps the space of polynomials of degree at most \(N\) into itself and is triangular in the monomial basis ordered by degree, with diagonal entries \(-k(k+1)\) for \(k = 0,\ldots,N\). Those entries are pairwise distinct, so each eigenvalue of \(\Lambda\) on that space is simple and its eigenspace is one-dimensional. A polynomial solution of Equation (9.146) of degree at most \(N\) is an eigenvector for the eigenvalue \(-n(n+1)\), and \(P_{n}\) is one such by Proposition 9.109; taking \(N\) at least as large as the degree of the given solution identifies the two up to a constant.
∎Orthogonality
The Legendre polynomials are orthogonal on \([-1, 1]\):
Rests on Theorem 9.57 and Proposition 9.109.
Derivation. Derives Theorem 9.111. For the off-diagonal terms read Equation (9.146) as the eigenvalue problem it is: \(p(x) = 1-x^{2}\), \(q = 0\), weight \(r = 1\) and eigenvalue \(\lambda^{2} = n(n+1)\), distinct for distinct non-negative degrees. The coefficient function \(p\) vanishes at both endpoints while \(P_{m}\), \(P_{n}\) and their derivatives are polynomials and hence bounded there, so Equation (9.57) holds by the second mechanism of Remark 9.58 and Theorem 9.57 gives the vanishing of the integral for \(m \neq n\).
For the diagonal use Equation (9.148) and integrate by parts \(n\) times (Equation (7.28)). At each step the integrated term carries a factor \(\dd^{j}(x^{2}-1)^{n}/\dd x^{j}\) with \(j < n\), which retains a factor \(\left(x^{2}-1\right)\) and so vanishes at \(x = \pm1\). Hence
since the \(2n\)-th derivative of a monic-leading polynomial of degree \(2n\) is the constant \((2n)!\) and \((-1)^{n}(x^{2}-1)^{n} = (1-x^{2})^{n}\). Write \(I_{n} = \int_{-1}^{1}(1-x^{2})^{n}\,\dd x\). Integration by parts gives
the boundary term vanishing and \(x^{2} = 1 - (1-x^{2})\) splitting the integral; hence \(I_{n} = \bigl[2n/(2n+1)\bigr]I_{n-1}\), and with \(I_{0} = 2\),
the closed form following from \(\prod_{k\le n}2k = 2^{n}n!\) and \(\prod_{k\le n}(2k+1) = (2n+1)!/\left(2^{n}n!\right)\). Multiplying,
The source reserves headings for an orthogonality relation of the second-kind solutions \(Q_{n}\), and for one of the general solution Equation (9.147). Neither exists, and the reason is the warning at the end of Remark 9.58: \(p\) vanishing at an endpoint is not enough when the solutions do not stay bounded there. The \(Q_{n}\) diverge logarithmically at \(x = \pm1\) — \(Q_{0}(x) = \tfrac12\ln\left[(1+x)/(1-x)\right]\) and \(Q_{n} = P_{n}Q_{0} - (\text{a polynomial})\) — and a short computation with \(u = Q_{0}\), \(v = Q_{2}\) shows that the surviving part of the bracket is \(\tfrac32\,p\,\dd Q_{0}/\dd x = \tfrac32\) at \(x \to 1\) and \(-\tfrac32\) at \(x \to -1\), so Equation (9.57) equals \(3\) and not \(0\). The identity in the proof of Theorem 9.57 then returns \(\int_{-1}^{1}Q_{0}Q_{2}\,\dd x = -\tfrac12\) rather than zero. Orthogonality on \([-1,1]\) is a property of the solutions that remain bounded at the two singular endpoints, and that requirement is exactly what selects the \(P_{n}\) and makes the Legendre problem a self-adjoint one.
Generating function
For \(\abs{x}\le1\) and \(\abs{t}<1\) the Legendre polynomials are generated by
Rests on Lemma 9.110 and Proposition 9.109.
Derivation. Derives Theorem 9.113. Write \(W = 1-2xt+t^{2}\), so that \(g = W^{-1/2}\). For \(\abs{x}\le1\) the roots of \(W\) in \(t\) are \(\ee^{\pm\ii\theta}\) with \(x = \cos\theta\), both of modulus one, so \(g\) is holomorphic in \(t\) on \(\abs{t}<1\) and has there a Taylor expansion \(g = \sum_{n\ge0}p_{n}(x)\,t^{n}\) (Theorem 8.20). Expanding \(\bigl[1 - t(2x-t)\bigr]^{-1/2}\) by the binomial series and collecting powers of \(t\) shows each \(p_{n}\) to be a polynomial in \(x\) of degree at most \(n\).
Next, \(g\) satisfies the identity
Indeed \(\pp g/\pp x = t\,W^{-3/2}\) and \(\pp g/\pp t = (x-t)\,W^{-3/2}\), so the first bracket differentiates to \(-2xt\,W^{-3/2} + 3\left(1-x^{2}\right)t^{2}W^{-5/2}\) and the second to \(2t(x-t)W^{-3/2} - t^{2}W^{-3/2} + 3t^{2}(x-t)^{2}W^{-5/2}\). The terms in \(W^{-3/2}\) add to \(-3t^{2}W^{-3/2}\), and those in \(W^{-5/2}\) add to \(3t^{2}\bigl[(1-x^{2}) + (x-t)^{2}\bigr]W^{-5/2} = 3t^{2}W\cdot W^{-5/2}\), which cancels the first.
Substituting the Taylor series in that identity and equating coefficients of \(t^{n}\), using \(\pp_{t}\bigl[t^{2}\pp_{t}g\bigr] = \sum_{n}n(n+1)p_{n}t^{n}\), gives
which is Equation (9.146). So \(p_{n}\) is a polynomial solution of the Legendre equation of that degree, hence a constant multiple of \(P_{n}\) by Lemma 9.110. The constant is \(1\) because \(g(1,t) = (1-t)^{-1} = \sum_{n}t^{n}\) gives \(p_{n}(1) = 1\), and \(P_{n}(1) = 1\) by Proposition 9.109.
∎Legendre series and recurrence relations
For \(n \ge 1\),
with \(P_{0}(x) = 1\) and \(P_{1}(x) = x\). Rests on Theorem 9.113.
Derivation. Derives Proposition 9.114. Differentiating Equation (9.150) with respect to \(t\) gives \(W\,\pp g/\pp t = (x-t)\,g\), and with respect to \(x\) gives \(W\,\pp g/\pp x = t\,g\). Substituting the series and reading off the coefficient of \(t^{n}\) turns the first into
which is Equation (9.151), and the second into
Differentiating Equation (9.151) gives \((n+1)P_{n+1}' = (2n+1)P_{n} + (2n+1)xP_{n}' - nP_{n-1}'\), and Equation (9.153) at index \(n+1\) gives \(2xP_{n}' = P_{n+1}' + P_{n-1}' - P_{n}\). Eliminating \(xP_{n}'\) between the two leaves \(\tfrac12 P_{n+1}' = \tfrac12(2n+1)P_{n} + \tfrac12 P_{n-1}'\), which is Equation (9.152). The initial values follow from \(g(x,t) = 1 + xt + O\left(t^{2}\right)\).
∎Let \(f\) on \([-1,1]\) admit an expansion
convergent in the mean. Then
Rests on Theorem 9.111 and Definition 9.47.
Derivation. Derives Proposition 9.115. Multiply Equation (9.154) by \(P_{m}\) and integrate over \([-1,1]\); mean convergence permits the exchange of sum and integral, and Equation (9.149) leaves the single term \(n = m\) with the factor \(2/(2m+1)\), which inverts to Equation (9.155). That an expansion exists for every square-integrable \(f\) — the completeness of the system — is again a Sturm–Liouville theorem and not proved here, exactly as in Remark 9.101.
∎The associated Legendre equation
Sturm–Liouville data
The identification of the operator data is
General solution
The homogeneous equation becomes
whose solution is
with \(P_n^m\) the associated Legendre functions and \(Q_n^m\) the associated Legendre functions of the second kind.
Construction from the Legendre polynomials
For integers \(0 \le m \le n\) and \(\abs{x}\le1\),
so that \(P_{n}^{0} = P_{n}\) and \(P_{n}^{m} = 0\) for \(m>n\), the \(m\)-th derivative of a polynomial of degree \(n\) then vanishing. Rests on Definition 9.108.
If \(y\) solves Equation (9.146) then \(u = \dd^{m}y/\dd x^{m}\) solves
Rests on Equation (9.146).
Derivation. Derives Lemma 9.117. Differentiate \(\left(1-x^{2}\right)y'' - 2xy' + n(n+1)y = 0\) exactly \(m\) times by the Leibniz rule Equation (7.16). The first term contributes \(\left(1-x^{2}\right)y^{(m+2)} - 2mx\,y^{(m+1)} - m(m-1)y^{(m)}\), the second \(-2x\,y^{(m+1)} - 2m\,y^{(m)}\), and the third \(n(n+1)y^{(m)}\); only three and two terms survive respectively because \(1-x^{2}\) is quadratic and \(x\) linear. Collecting, the coefficient of \(y^{(m+1)}\) is \(-2(m+1)x\) and that of \(y^{(m)}\) is \(n(n+1) - m(m-1) - 2m = n(n+1) - m(m+1)\).
∎\(P_{n}^{m}\) of Equation (9.159) solves Equation (9.157). Rests on Definition 9.116 and Lemma 9.117.
Derivation. Derives Proposition 9.118. Put \(\varpi = \left(1-x^{2}\right)^{m/2}\) and \(y = \varpi\,u\) with \(u = \dd^{m}P_{n}/\dd x^{m}\), so that \(\varpi'/\varpi = -mx/\left(1-x^{2}\right)\) and
Substituting \(y = \varpi u\) into the associated equation Equation (9.157), written out as \(\left(1-x^{2}\right)y'' - 2xy' + \bigl[n(n+1) - m^{2}/(1-x^{2})\bigr]y = 0\), and dividing by \(\varpi\) leaves an equation for \(u\) whose coefficient of \(\dd u/\dd x\) is \(2\left(1-x^{2}\right)\varpi'/\varpi - 2x = -2(m+1)x\) and whose coefficient of \(u\) is
that is \(n(n+1) - m(m+1)\). The result is exactly Equation (9.160), which \(u\) satisfies by Lemma 9.117.
∎Orthogonality
At fixed order \(m\), the associated Legendre functions are orthogonal on \([-1,1]\):
Rests on Theorem 9.57, Proposition 9.118 and Theorem 9.111.
Derivation. Derives Theorem 9.119. For \(n \neq n'\) read Equation (9.157) at fixed \(m\) as the eigenvalue problem with \(p(x) = 1-x^{2}\), \(q(x) = -m^{2}/\left(1-x^{2}\right)\), weight \(r = 1\) and eigenvalue \(\lambda^{2} = n(n+1)\). The functions \(P_{n}^{m}\) carry the factor \(\left(1-x^{2}\right)^{m/2}\), so for \(m \ge 1\) both they and \(p\) times their derivatives vanish at \(x = \pm1\), while for \(m = 0\) the argument is that of Theorem 9.111; either way Equation (9.57) holds and Theorem 9.57 applies.
For the diagonal put \(I_{m} = \int_{-1}^{1}\left(1-x^{2}\right)^{m} \bigl[u_{m}(x)\bigr]^{2}\,\dd x\) with \(u_{m} = \dd^{m}P_{n}/\dd x^{m}\), which is the integral to be computed because \(\left(P_{n}^{m}\right)^{2} = \left(1-x^{2}\right)^{m}u_{m}^{2}\). For \(m \ge 1\) write \(u_{m} = \dd u_{m-1}/\dd x\) and integrate by parts (Equation (7.28)); the boundary term carries the factor \(\left(1-x^{2}\right)^{m}\) and vanishes, so
where \(\Theta = \left(1-x^{2}\right)\dd^{2}u_{m-1}/\dd x^{2} - 2m\,x\,\dd u_{m-1}/\dd x\). By Equation (9.160) at order \(m-1\) that combination is \(-\bigl[n(n+1) - m(m-1)\bigr]u_{m-1}\), so
Iterating from \(I_{0} = 2/(2n+1)\), which is Theorem 9.111, gives
which is Equation (9.161).
∎Recurrence relations
For \(1 \le m \le n\),
Rests on Proposition 9.114 and Definition 9.116.
Derivation. Derives Proposition 9.120. Write \(D = \dd/\dd x\). Applying \(D^{m}\) to Equation (9.151) by the Leibniz rule Equation (7.16), in which the factor \(x\) contributes only two terms,
while applying \(D^{m-1}\) to Equation (9.152) gives \(D^{m}P_{n+1} - D^{m}P_{n-1} = (2n+1)\,D^{m-1}P_{n}\). Substituting the second into the first eliminates \(D^{m-1}P_{n}\) and leaves
and multiplying through by \(\left(1-x^{2}\right)^{m/2}\) turns each \(D^{m}P_{k}\) into \(P_{k}^{m}\) by Equation (9.159).
∎Spherical harmonics
The associated Legendre functions in \(\cos\theta\), combined with the complex exponentials of Section 9.7 in the azimuthal angle \(\varphi\), form the spherical harmonics
which are orthonormal on the unit sphere [Jackson:1999],
Separating the angular part of the Laplace operator on the unit sphere as \(Y(\theta,\varphi) = \Theta(\theta)\,\Phi(\varphi)\) forces \(\Phi(\varphi) = \ee^{\ii m\varphi}\) with \(m\) an integer, and \(\Theta(\theta) = P_{n}^{m}(\cos\theta)\) with \(n\) a non-negative integer and \(\abs{m}\le n\). Every other choice of the two separation constants gives a solution that is either multivalued in \(\varphi\) or unbounded at one of the poles. Rests on Equations (9.107) and (9.157).
Derivation. Derives Proposition 9.121. The angular eigenvalue equation is
Substituting the product and dividing by \(\Theta\Phi/\sin^{2}\theta\) separates it: the \(\varphi\) part is \(\dd^{2}\Phi/\dd\varphi^{2} = -m^{2}\Phi\), whose solutions are the exponentials Equation (9.107), and single-valuedness under \(\varphi \mapsto \varphi + 2\pi\) restricts \(m\) to the integers. The \(\theta\) part, in the variable \(x = \cos\theta\) — for which \(\dd/\dd\theta = -\sin\theta\,\dd/\dd x\) and \(\sin^{2}\theta = 1-x^{2}\) — is exactly the associated Legendre equation Equation (9.157) with \(n(n+1)\) replaced by \(\Lambda\). Boundedness at the two poles \(x = \pm1\), the singular endpoints of that equation in the sense of Remark 9.17, selects the solutions \(P_{n}^{m}\): for \(\Lambda = n(n+1)\) with \(n\) a non-negative integer and \(\abs{m}\le n\) these are bounded by Equation (9.159), and the second solution \(Q_{n}^{m}\) diverges there, as does every solution for any other value of \(\Lambda\).
∎With the constant displayed in Equation (9.163) the functions \(Y_{n}^{m}\) satisfy Equation (9.164), and that constant is the only positive one that does so. Rests on Theorem 9.119 and Equation (9.109).
Derivation. Derives Theorem 9.122. The integral factorises. The azimuthal factor is
which is Equation (9.109) on an interval of one period, so every term with \(m \neq m'\) vanishes whatever the polar factor does. At \(m = m'\) the polar factor is, substituting \(x = \cos\theta\) and \(\dd x = -\sin\theta\,\dd\theta\),
by Theorem 9.119. The product of the two factors is \(2\pi \cdot 2(n+m)!\bigl/\bigl[(2n+1)(n-m)!\bigr]\) times \(\delta_{nn'}\delta_{mm'}\), and the squared modulus of the constant in Equation (9.163) is the reciprocal of that number, which gives Equation (9.164). Since the constant enters squared and is required positive, it is unique.
∎Equation (9.159) defines \(P_{n}^{m}\) for \(m \ge 0\) only, and Equation (9.163) is extended to \(-n \le m < 0\) by
a convention — the Condon–Shortley one — chosen so that Equation (9.163) then satisfies \(\overline{Y_{n}^{m}} = (-1)^{m}\,Y_{n}^{-m}\) and so that Theorem 9.122 continues to hold with the same formula, as substituting Equation (9.165) into the two factorials shows. That the \(Y_{n}^{m}\) are moreover complete in the square-integrable functions on the sphere — so that every such function has an expansion in them — is not proved here: it is the statement that the angular Laplace operator has no other eigenfunctions, which belongs with the general spectral theory of Hilbert Spaces [Courant:1962]. The addition theorem, which expresses \(P_{n}(\cos\gamma)\) for the angle \(\gamma\) between two directions as a sum of products of spherical harmonics of those two directions, rests on the rotation behaviour of the space spanned by the \(Y_{n}^{m}\) at fixed \(n\) and is not derived here either; it is used in this treatise only through [Jackson:1999].
The Hermite equation
The source announces the Hermite family with the same systematic headings as the families above — Sturm–Liouville data \(p\), \(q\) and \(r\), general solution, recurrence relations, orthonormality, and Hermite series — but every one of those subsections is empty, so they are supplied here. This family governs the stationary states of the quantum harmonic oscillator, where the Hermite polynomials appear multiplied by a Gaussian weight [CohenTannoudji:1977] [Griffiths:2018].
Sturm–Liouville data
The Hermite equation is
Multiplying it by the integrating factor \(\ee^{-x^{2}}\) puts it in the form Equation (9.54), since \(\dv{}{x}\bigl[\ee^{-x^{2}}\dd y/\dd x\bigr] = \ee^{-x^{2}}\bigl[y'' - 2xy'\bigr]\), with data
on the whole line. The weight is the Gaussian, and the interval is unbounded, so the third mechanism of Remark 9.58 is the one that will apply.
Solutions and generating function
The polynomials Equation (9.168) are generated by
and \(H_{n}\) is a polynomial of degree \(n\) with leading coefficient \(2^{n}\) and parity \((-1)^{n}\). Rests on Definition 9.124.
Derivation. Derives Proposition 9.125. Complete the square: \(2xt-t^{2} = x^{2} - (x-t)^{2}\), so the left of Equation (9.169) is \(\ee^{x^{2}}\,\ee^{-(x-t)^{2}}\). The second factor is a function of \(x-t\) alone, so \(\pp/\pp t\) acting on it equals \(-\pp/\pp x\); hence its \(n\)-th \(t\)-derivative at \(t = 0\) is \((-1)^{n}\dd^{n}\ee^{-x^{2}}/\dd x^{n}\), and the Taylor expansion in \(t\) of a function entire in \(t\) (Theorem 8.20) gives
which is Equation (9.169) with Equation (9.168). Differentiating Equation (9.168) once more, and using \(\dd\ee^{-x^{2}}/\dd x = -2x\,\ee^{-x^{2}}\), gives \(H_{n+1} = 2x\,H_{n} - \dd H_{n}/\dd x\); with \(H_{0} = 1\) this shows by induction that \(H_{n}\) is a polynomial of degree \(n\) whose leading coefficient doubles at each step, hence equals \(2^{n}\). The parity follows from \(\ee^{2(-x)t-t^{2}} = \ee^{2x(-t)-(-t)^{2}}\), which sends \(H_{n}(-x)\) to \((-1)^{n}H_{n}(x)\).
∎For \(n \ge 1\),
and \(H_{n}\) solves the Hermite equation Equation (9.166). Rests on Proposition 9.125 and Equation (9.166).
Derivation. Derives Proposition 9.126. Differentiating Equation (9.169) with respect to \(x\) gives \(2t\,\ee^{2xt-t^{2}} = \sum_{n}\bigl(\dd H_{n}/\dd x\bigr)t^{n}/n!\), and comparing coefficients of \(t^{n}\) gives \(\dd H_{n}/\dd x\bigm/n! = 2H_{n-1}/(n-1)!\), which is Equation (9.170). Combining it with \(H_{n+1} = 2xH_{n} - \dd H_{n}/\dd x\), established in the proof of Proposition 9.125, gives Equation (9.171). For the equation, differentiate that same relation: \(\dd H_{n+1}/\dd x = 2H_{n} + 2x\,\dd H_{n}/\dd x - \dd^{2}H_{n}/\dd x^{2}\), and use Equation (9.170) on the left, which turns it into \(2(n+1)H_{n}\). Cancelling \(2H_{n}\) from both sides leaves \(\dd^{2}H_{n}/\dd x^{2} - 2x\,\dd H_{n}/\dd x + 2n\,H_{n} = 0\).
∎Orthogonality and Hermite series
Rests on Theorem 9.57, Proposition 9.126 and Equation (9.167).
Derivation. Derives Theorem 9.127. For \(m \neq n\) apply Theorem 9.57 with the data Equation (9.167): the eigenvalues \(2m\) and \(2n\) differ, and the boundary term Equation (9.57) vanishes because \(p = \ee^{-x^{2}}\) multiplies polynomials and \(\ee^{-x^{2}}\) beats every polynomial at both ends of the line — the third mechanism of Remark 9.58.
For \(m = n\) insert Equation (9.168) for one factor and integrate by parts \(n\) times (Equation (7.28)), every boundary term vanishing for the same reason:
By Proposition 9.125 the \(n\)-th derivative of \(H_{n}\) is the constant \(2^{n}n!\), and the Gaussian integral \(\int_{-\infty}^{\infty}\ee^{-x^{2}}\dd x = \sqrt{\pi}\) — its square being \(\int\int\ee^{-x^{2}-y^{2}}\dd x\,\dd y = \pi\) in plane polar coordinates — completes Equation (9.172).
∎Let \(f\) admit an expansion \(f(x) = \sum_{n\ge0}a_{n}H_{n}(x)\) convergent in the mean with respect to the weight \(\ee^{-x^{2}}\). Then
Rests on Theorem 9.127 and Definition 9.47.
Derivation. Derives Proposition 9.128. Multiply by \(H_{m}(x)\ee^{-x^{2}}\), integrate over the line, exchange sum and integral by mean convergence, and use Equation (9.172); the surviving term is \(n = m\) and carries the constant \(2^{m}m!\sqrt{\pi}\). Completeness of the system carries the same caveat as Remark 9.101.
∎For the potential \(V(x) = \tfrac12 m_{0}\omega^{2}x^{2}\) of Example 9.80, with \(m_{0}\) in \(\mathrm{kg}\) and \(\omega\) in \(\mathrm{rad}/\mathrm{s}\), the substitution \(\xi = x/\ell_{0}\) with \(\ell_{0} = \sqrt{\hbar/(m_{0}\omega)}\) in \(\mathrm{m}\) turns Equation (9.59) into \(\dd^{2}\psi/\dd\xi^{2} + \bigl[2E/(\hbar\omega) - \xi^{2}\bigr]\psi = 0\), and the substitution \(\psi = \ee^{-\xi^{2}/2}H(\xi)\) turns that into Equation (9.166) for \(H\) with \(2n = 2E/(\hbar\omega) - 1\). Square integrability forces the series solution to terminate, that is \(n\) to be a non-negative integer, which returns the levels Equation (9.91) \(E_{n} = \bigl(n+\tfrac12\bigr)\hbar\omega\) in \(\mathrm{J}\) — there exactly, not semiclassically. The mass is written \(m_{0}\) here only because \(m\) is the order of the associated Legendre functions elsewhere in this chapter.
The Laguerre equation
The source likewise reserves the full set of systematic headings for the Laguerre family — Sturm–Liouville data \(p\), \(q\) and \(r\), general solution, recurrence relations, orthonormality, and Laguerre series — all of them empty, and they are supplied here. The associated Laguerre polynomials furnish the radial eigenfunctions of the hydrogen atom [CohenTannoudji:1977] [Griffiths:2018].
Sturm–Liouville data
The Laguerre equation is
Multiplying by \(\ee^{-x}\) puts it in the form Equation (9.54), since \(\dv{}{x}\bigl[x\ee^{-x}\dd y/\dd x\bigr] = \ee^{-x}\bigl[x y'' + (1-x)y'\bigr]\), with data
on the half-line. Here \(p\) vanishes at the origin and decays at infinity, so the boundary term dies at both ends at once.
Solutions and generating function
The Rodrigues formula Equation (9.176) evaluates to
a polynomial of degree \(n\) with leading coefficient \((-1)^{n}/n!\) and \(L_{n}(0) = 1\), and these polynomials are generated by
Rests on Definition 9.130.
Derivation. Derives Proposition 9.131. By the Leibniz rule Equation (7.16),
and dividing by \(n!\,\ee^{-x}\) and putting \(j = n-k\) gives Equation (9.177); the term \(j=n\) gives the leading coefficient and the term \(j=0\) gives \(L_{n}(0) = 1\). For the generating function expand the exponential and then each factor \((1-t)^{-1-j}\) by the binomial series,
both series being absolutely convergent for \(\abs{t}<1\), so that they may be rearranged (Proposition 7.49). The coefficient of \(t^{n}\) collects \(i = n-j\) and is \(\sum_{j\le n}\binom{n}{j}(-x)^{j}/j!\), which is Equation (9.177).
∎\(L_{n}\) solves Equation (9.174), and for \(n \ge 1\)
Rests on Proposition 9.131 and Equation (9.174).
Derivation. Derives Proposition 9.132. Write \(v = x^{n}\ee^{-x}\), which satisfies \(x\,\dd v/\dd x = (n-x)\,v\). Differentiating that identity \(n+1\) times by the Leibniz rule Equation (7.16), each side producing only two terms because \(x\) and \(n-x\) are linear, gives for \(u = v^{(n)}\)
that is \(x\,u'' + (1+x)\,u' + (n+1)\,u = 0\). Now \(u = n!\,\ee^{-x}L_{n}\) by Equation (9.176), so \(u' = n!\,\ee^{-x}\left(L_{n}' - L_{n}\right)\) and \(u'' = n!\,\ee^{-x}\left(L_{n}'' - 2L_{n}' + L_{n}\right)\); substituting and cancelling \(n!\ee^{-x}\),
which is Equation (9.174). For the recurrence, write \(G\) for the left of Equation (9.178); then \(\ln G = -\ln(1-t) - xt/(1-t)\) gives \(\pp G/\pp t = G\bigl[(1-t)^{-1} - x(1-t)^{-2}\bigr]\), that is \(\left(1-t\right)^{2}\pp G/\pp t = \left(1-t-x\right)G\). Substituting the series and reading off the coefficient of \(t^{n}\) gives \((n+1)L_{n+1} - 2n\,L_{n} + (n-1)L_{n-1} = L_{n} - L_{n-1} - x\,L_{n}\), which rearranges to Equation (9.179).
∎Orthogonality and Laguerre series
The Laguerre polynomials are orthonormal on the half-line with weight \(\ee^{-x}\):
Rests on Theorem 9.57, Proposition 9.131 and Equation (9.175).
Derivation. Derives Theorem 9.133. For \(m \neq n\) apply Theorem 9.57 with the data Equation (9.175): the eigenvalues \(m\) and \(n\) differ, and \(p(x) = x\ee^{-x}\) vanishes at the origin and decays at infinity while the solutions are polynomials, so Equation (9.57) holds.
For \(m = n\) use Equation (9.176) on one factor and integrate by parts \(n\) times (Equation (7.28)). Every boundary term vanishes: at infinity because of \(\ee^{-x}\), and at the origin because each derivative \(\dd^{k}\bigl(x^{n}\ee^{-x}\bigr)/\dd x^{k}\) with \(k<n\) still carries a positive power of \(x\). Hence
By Equation (9.177) the \(n\)-th derivative of \(L_{n}\) is the constant \((-1)^{n}\), and \(\int_{0}^{\infty}x^{n}\ee^{-x}\dd x = \Gamma(n+1) = n!\), so the whole expression is \(\left[(-1)^{n}\right]^{2}n!/n! = 1\).
∎Let \(f\) on \((0,\infty)\) admit an expansion \(f(x) = \sum_{n\ge0}a_{n}L_{n}(x)\) convergent in the mean with respect to the weight \(\ee^{-x}\). Then \(a_{n} = \int_{0}^{\infty}f(x)\,L_{n}(x)\,\ee^{-x}\,\dd x\). Rests on Theorem 9.133 and Definition 9.47.
Derivation. Derives Proposition 9.134. This is Proposition 9.48 verbatim: the system \(\set{L_{n}}\) is already orthonormal for the weight \(\ee^{-x}\) by Equation (9.180), so no normalising constant appears. Completeness carries the caveat of Remark 9.101.
∎The radial eigenfunctions of the hydrogen atom involve not \(L_{n}\) but the associated Laguerre polynomials \(L_{n}^{k} = (-1)^{k}\,\dd^{k}L_{n+k}/\dd x^{k}\), which stand to Definition 9.130 as the associated Legendre functions of Definition 9.116 stand to the Legendre polynomials. They solve \(x\,y'' + (k+1-x)\,y' + n\,y = 0\), which is Equation (9.174) with \(1-x\) replaced by \(k+1-x\) — a Sturm–Liouville problem with \(p(x) = x^{k+1}\ee^{-x}\) and weight \(x^{k}\ee^{-x}\) on the same half-line, so that the boundary term dies at both ends exactly as it does here. Their orthogonality, recurrences and generating function are established by the arguments used above, and they are used in The Hydrogen Atom [CohenTannoudji:1977] [Griffiths:2018].
Elliptic integrals and the Jacobi elliptic functions
Every family treated so far arose from separating a linear partial differential equation. The family of this section arises differently and is needed just as often: whenever a one-degree-of-freedom conservative system is integrated by energy conservation, the result is a quadrature \(t = \int\dd u/\sqrt{f(u)}\), and when \(f\) is a cubic or a quartic that quadrature is an elliptic integral. The pendulum at finite amplitude, the nutation of a heavy symmetric top, the free rotation of a body with three unequal moments of inertia and the shape of an elastic strip under end thrust are all of this kind, and each stops at a quadrature the reader cannot invert without the functions defined here. They are due to Legendre, who reduced every such integral to three normal forms, and to Jacobi and Abel, who saw that the fruitful object is the inverse of the integral.
The three normal forms
For every integer \(n\ge0\),
Rests on Corollary 7.44 and Lemma 7.71.
Derivation. Derives Lemma 9.136. Write \(W_{n}\) for the integral. Integrating by parts (Corollary 7.44) with \(u = \sin^{2n-1}\theta\) and \(\dd v = \sin\theta\,\dd\theta\), the boundary term \(\bigl[-\sin^{2n-1}\theta\cos\theta\bigr]_{0}^{\pi/2}\) vanishes at both ends for \(n\ge1\), leaving
by \(\cos^{2} = 1 - \sin^{2}\) (Lemma 7.71). Hence \(2n\,W_{n} = (2n-1)W_{n-1}\), and with \(W_{0} = \pi/2\) induction gives
the last step because \(\prod_{j\le n}(2j-1) = (2n)!/\bigl(2^{n}n!\bigr)\) and \(\prod_{j\le n}(2j) = 2^{n}n!\).
∎Let \(0 \le k < 1\) be the modulus and \(\varphi\in\R\) the amplitude. The incomplete elliptic integrals of the first, second and third kind are
with \(n\) the characteristic. The complete integrals are their values at \(\varphi = \pi/2\),
Rests on Definition 7.125 and Theorem 7.40.
Three conventions are in use and they are not distinguished by the symbol \(K\). This treatise writes \(K(k)\) with the modulus \(k\), so that \(k^{2}\) appears in Equation (9.182). Much of the mathematical literature, and most computer-algebra systems, write \(K(m)\) with the parameter \(m = k^{2}\); a third convention uses the modular angle \(\alpha\) with \(k = \sin\alpha\). The three disagree by a squaring, so a formula copied across conventions without checking is wrong by an amount that vanishes at small amplitude and is therefore not caught by the obvious limit. Where a result of this treatise is quoted elsewhere — the pendulum period Equation (9.199) is the case that matters — the argument is stated as \(\sin(\theta_{0}/2)\), a modulus.
For \(0\le k<1\),
the series converging for every \(k\) with \(\abs{k}<1\). Moreover \(K\) is strictly increasing on \([0,1)\) and \(K(k)\to+\infty\) as \(k\to1^{-}\). Rests on Definition 9.137, Lemma 9.136 and Lemma 9.6.
Derivation. Derives Proposition 9.139. The binomial series \((1-u)^{-1/2} = \sum_{n\ge0}\bigl[(2n)!/(4^{n}(n!)^{2})\bigr]u^{n}\) converges for \(\abs{u}<1\) (Theorem 7.50, Proposition 7.48: the ratio of successive coefficients is \((2n+1)/(2n+2)\to1\)). Put \(u = k^{2}\sin^{2}\theta\), so that \(0\le u\le k^{2}<1\) uniformly in \(\theta\). The \(n\)-th term is bounded on \([0,\pi/2]\) by \(M_{n} = \bigl[(2n)!/(4^{n}(n!)^{2})\bigr]k^{2n}\), and \(\sum_{n}M_{n} = (1-k^{2})^{-1/2}\) is finite, so by Lemma 9.6 the series converges uniformly on \([0,\pi/2]\) and may be integrated term by term. Each term contributes \(\bigl[(2n)!/(4^{n}(n!)^{2})\bigr]k^{2n}W_{n}\) with \(W_{n}\) the Wallis integral Equation (9.181), and substituting Equation (9.181) produces the square in Equation (9.186). The first coefficients are \(1\), \((1/2)^{2}=1/4\), \((3/8)^{2}=9/64\) and \((5/16)^{2}=25/256\).
Monotonicity is immediate: the integrand of Equation (9.182) increases with \(k\) at every \(\theta\in(0,\pi/2]\). For the divergence, use \(1-k^{2}\sin^{2}\theta = (1-k^{2}) + k^{2}\cos^{2}\theta \le c^{2} + \cos^{2}\theta\) with \(c = \sqrt{1-k^{2}}\), and substitute \(w = \pi/2-\theta\), so \(\cos\theta = \sin w \le w\):
which tends to \(+\infty\) as \(c\to0^{+}\), that is as \(k\to1^{-}\).
∎The Jacobi elliptic functions
Fix \(0\le k<1\). The integrand of Equation (9.182) is positive and \(\pi\)-periodic in \(\theta\), so \(\varphi\mapsto F(\varphi,k)\) is a strictly increasing bijection of \(\R\) onto \(\R\) satisfying \(F(\varphi+\pi,k) = F(\varphi,k) + 2K(k)\). Its inverse is the amplitude \(\operatorname{am}(u,k)\), and the Jacobi elliptic functions are
Rests on Definition 9.137 and Proposition 7.32.
With \(K = K(k)\) and the argument \(k\) suppressed,
and \(\dd(\operatorname{dn}u)/\dd u = -k^{2}\operatorname{sn}u\,\operatorname{cn}u\). Furthermore \(\operatorname{sn}\) is odd and \(\operatorname{cn}\), \(\operatorname{dn}\) are even; \(\operatorname{sn}0 = 0\), \(\operatorname{cn}0 = \operatorname{dn}0 = 1\), \(\operatorname{sn}K = 1\); and
so \(\operatorname{sn}\) and \(\operatorname{cn}\) have period \(4K\) and \(\operatorname{dn}\) has period \(2K\). At \(k=0\) they reduce to \(\sin u\), \(\cos u\) and \(1\). Rests on Definition 9.140, Proposition 7.32 and Lemma 7.71.
Derivation. Derives Proposition 9.141. The first identity in Equation (9.188) is the Pythagorean identity (Lemma 7.71) applied to \(\operatorname{am}\), and the second is Equation (9.187) squared. Since \(\pp F/\pp\varphi = (1-k^{2}\sin^{2}\varphi)^{-1/2}\) is positive and continuous, Proposition 7.32 gives
and the chain rule (Proposition 7.31) turns Equation (9.187) into Equation (9.189). Differentiating \(\operatorname{dn}^{2} = 1-k^{2}\operatorname{sn}^{2}\) and dividing by \(2\operatorname{dn}\) gives the third derivative, \(-k^{2}\operatorname{sn}\operatorname{cn}\).
\(F\) is odd in \(\varphi\) because its integrand is even, so \(\operatorname{am}\) is odd and the parities follow. The values at \(u=0\) are those at \(\varphi=0\), and \(\operatorname{am}(K) = \pi/2\) by Equation (9.185), giving \(\operatorname{sn}K = 1\). Finally \(F(\varphi+\pi) = F(\varphi)+2K\) inverts to \(\operatorname{am}(u+2K) = \operatorname{am}(u)+\pi\), and Equation (9.190) is what \(\sin\), \(\cos\) and \(\sin^{2}\) do under a shift by \(\pi\). At \(k=0\), Equation (9.182) gives \(F(\varphi,0)=\varphi\) and hence \(\operatorname{am}(u,0)=u\).
∎The function \(y(u) = \operatorname{sn}(u,k)\) satisfies
and therefore the second-order equation
with \(y(0)=0\) and \(y'(0)=1\). Conversely Equation (9.193) with those initial data has \(\operatorname{sn}(\cdot,k)\) as its unique solution. Rests on Proposition 9.141 and Theorem 9.8.
Derivation. Derives Proposition 9.142. Squaring Equation (9.189) and eliminating \(\operatorname{cn}\) and \(\operatorname{dn}\) by Equation (9.188) gives Equation (9.192). Differentiating \(\operatorname{sn}' = \operatorname{cn}\operatorname{dn}\) once more and using the three derivative formulas,
which is Equation (9.193). The initial values are those of Proposition 9.141, with \(\operatorname{sn}'(0) = \operatorname{cn}0\,\operatorname{dn}0 = 1\). The right-hand side of Equation (9.193) is a polynomial and therefore locally Lipschitz, so Theorem 9.8 makes the solution unique. Note that Equation (9.192), whose right-hand side has a square root at the turning points, does not by itself determine the solution there; the second-order form does.
∎Quadratures that reduce to elliptic integrals
Let \(f(v) = c\,(v-e_{1})(v-e_{2})(v-e_{3})\) with \(c>0\) and \(e_{1}<e_{2}<e_{3}\), so that \(f>0\) on \((e_{1},e_{2})\). Put
Then for \(u\in[e_{1},e_{2}]\),
so the integral is finite at both endpoints, its value over the whole interval is \(K(k)/\omega\), and the function \(t\mapsto u(t)\) inverting it is
which oscillates between \(e_{1}\) and \(e_{2}\) with period \(2K(k)/\omega\) and satisfies \((\dd u/\dd t)^{2} = f(u)\). Rests on Definition 9.140, Proposition 9.141 and Corollary 7.44.
Derivation. Derives Proposition 9.143. Substitute \(v = e_{1} + (e_{2}-e_{1})\sin^{2}\theta\) (Corollary 7.44), which increases from \(e_{1}\) to \(e_{2}\) as \(\theta\) runs over \([0,\pi/2]\). Then
and, using Equation (9.194), \(v-e_{3} = -(e_{3}-e_{1})\bigl[1-k^{2}\sin^{2}\theta\bigr]\). Hence
whose square root is \(2\omega(e_{2}-e_{1})\sin\theta\cos\theta \sqrt{1-k^{2}\sin^{2}\theta}\) on \([0,\pi/2]\), every factor being non-negative there. Since \(\dd v = 2(e_{2}-e_{1})\sin\theta\cos\theta\,\dd\theta\), everything cancels except
and integrating from \(0\) to \(\varphi\) gives Equation (9.195), finite at both ends because the integrand of Equation (9.182) is bounded for \(k<1\). Taking \(\varphi=\pi/2\) gives \(K(k)/\omega\).
Inverting Equation (9.195) means \(\varphi = \operatorname{am}(\omega t,k)\), hence \(\sin^{2}\varphi = \operatorname{sn}^{2}(\omega t,k)\), which is Equation (9.196). Differentiating it with Equation (9.189), \(\dd u/\dd t = 2\omega(e_{2}-e_{1}) \operatorname{sn}\operatorname{cn}\operatorname{dn}\), whose square is \(4\omega^{2}(e_{2}-e_{1})^{2} \operatorname{sn}^{2}\operatorname{cn}^{2}\operatorname{dn}^{2}\); and the same substitutions that produced \(f(v)\) above, now with \(\sin\theta = \operatorname{sn}\), give \(f(u) = c(e_{2}-e_{1})^{2}(e_{3}-e_{1}) \operatorname{sn}^{2}\operatorname{cn}^{2}\operatorname{dn}^{2}\). The two agree precisely when \(4\omega^{2} = c(e_{3}-e_{1})\), which is Equation (9.194). Finally \(\operatorname{sn}^{2}\) has period \(2K\) by Equation (9.190), so \(u\) has period \(2K/\omega\).
∎Let \(\omega_{0}>0\) and let \(\theta\) solve
with \(\theta(0)=0\) and \(\dd\theta/\dd t(0) = 2\omega_{0}\sin(\theta_{0}/2)\) for some \(\theta_{0}\in(0,\pi)\). Then the motion is periodic between \(-\theta_{0}\) and \(+\theta_{0}\), with
and its period is
which for small amplitude expands as
Rests on Proposition 9.139, Definition 9.140 and Corollary 7.44.
Derivation. Derives Proposition 9.144. Multiplying Equation (9.197) by \(\dd\theta/\dd t\) and integrating once gives the first integral \(\tfrac12\dot\theta^{2} - \omega_{0}^{2}\cos\theta = \text{const}\). The stated initial data fix the constant so that \(\dot\theta\) vanishes at \(\theta = \pm\theta_{0}\), and with \(\cos\theta = 1-2\sin^{2}(\theta/2)\),
The right-hand side is negative beyond \(\abs{\theta} = \theta_{0}\), so the motion is confined there and reverses at the two turning points. Equation (9.197) is unchanged under \(t\mapsto-t\), and also under \(\theta\mapsto-\theta\) because the sine is odd, so the four quarters of the motion are congruent and the period is four times the time from \(\theta=0\) to \(\theta=\theta_{0}\).
On that quarter \(\dot\theta>0\). Substitute \(\sin(\theta/2) = k\sin\varphi\), which carries \(\theta: 0\to\theta_{0}\) into \(\varphi: 0\to\pi/2\) and gives \(\tfrac12\cos(\theta/2)\,\dd\theta = k\cos\varphi\,\dd\varphi\) with \(\cos(\theta/2) = \sqrt{1-k^{2}\sin^{2}\varphi}\), while Equation (9.201) gives \(\dot\theta = 2\omega_{0}k\cos\varphi\). Hence
so \(\omega_{0}t = F(\varphi,k)\), that is \(\varphi = \operatorname{am}(\omega_{0}t,k)\), which is Equation (9.198); and \(\varphi=\pi/2\) at the turning point gives \(T/4 = K(k)/\omega_{0}\), which is Equation (9.199).
For Equation (9.200), expand \(k = \sin(\theta_{0}/2)\) by Taylor's theorem (Theorem 7.38): \(k^{2} = \theta_{0}^{2}/4 - \theta_{0}^{4}/48 + O(\theta_{0}^{6})\) and \(k^{4} = \theta_{0}^{4}/16 + O(\theta_{0}^{6})\). Substituting into Equation (9.186) truncated at \(k^{4}\),
and \(-1/192 + 9/1024 = (-16+27)/3072 = 11/3072\).
∎Proposition 9.143 is the general shape of the answer whenever a conservative system with one degree of freedom is reduced to a quadrature whose radicand is a cubic with three real roots and the motion runs between the lower two. That is the case of the heavy symmetric top, whose nutation angle satisfies exactly such an equation in the variable \(\cos\theta\), so its nutation is Equation (9.196) and its nutation period is the complete integral \(2K(k)/\omega\); the same reduction applies to the free rotation of a body with three distinct moments of inertia. The quartic case is reached from the cubic by the substitution that sends one root to infinity, and the case of two real and two complex roots by a different linear fractional substitution; those reductions are not carried out here, and no statement of this treatise depends on them. Proposition 9.144 is the second standard instance and is used directly by Oscillations and Mechanical Waves and Experiment: The Pendulum: Equation (9.199) is what a pendulum clock's rate depends on, Equation (9.200) is the correction a swinging amplitude introduces, and the two together are the quantitative content of the isochronism experiment. Note that Equation (9.200) is an expansion of an exact formula, not an independent approximation: the exact period is available at every amplitude below \(\pi\), and the series is quoted only because it is what an error budget uses.
The Laplace transform
Beside the Fourier transform of Section 9.7.5, the source reserves a section for the Laplace transform — the integral transform adapted to initial-value problems on the half-line — but leaves it empty. The material is developed in full in Section 17.4, and this section records only what it is for and where each piece of it is proved.
The Fourier transform requires a function to decay at both ends of the line. A physical initial-value problem supplies neither condition: the system is switched on at \(t = 0\) and its response may grow. Restricting the integral to \(t > 0\) and inserting the convergence factor \(\ee^{-st}\) with \(\Re s\) large enough repairs both defects at once, and gives the one-sided Laplace transform of Definition 17.65. Three facts make it a method rather than a definition. It converges on a half-plane \(\Re s > \sigma_c\) and nowhere to the left of it (Theorem 17.66), and it is holomorphic there (Proposition 17.67); on each vertical line it is literally a Fourier transform, of the causal signal damped by \(\ee^{-\sigma t}\) (Proposition 17.68), which is what reduces its inversion to Theorem 17.32 and produces the Bromwich contour integral (Theorem 17.69), evaluable by residues (Corollary 17.71). And it converts differentiation into multiplication while swallowing the initial data, which enter the transformed equation as extra polynomial terms rather than as conditions imposed afterwards (Equation (17.87)).
That last rule is what makes the transform the natural tool for the constant-coefficient equations of Section 9.1: a linear ordinary differential equation becomes an algebraic one, its solution is a rational function of \(s\), and the inverse transform is read off the poles by Heaviside's expansion theorem (Corollary 17.76). The location of those poles is the stability of the system (Proposition 17.77), and the requirement that the response be causal, applied to the same half-plane of holomorphy, yields the Kramers–Kronig dispersion relations (Theorem 17.78).
Summary of the classical families
Table 9.2 collects the Sturm–Liouville data \(\bigl(p, q, \lambda^2, r\bigr)\) of every family treated above, together with the pair of independent solutions. Three of the rows have to be read with Remark 9.97 in mind: the Bessel, Legendre and associated Legendre equations each admit two Sturm–Liouville readings. In the first the order or the degree sits in \(q\) — the source's identification, reproduced in the table — and in the second a scaling parameter or the degree carries the spectral parameter instead. It is the second reading to which the orthogonality relations of Theorems 9.98, 9.111 and 9.119 belong, since orthogonality of eigenfunctions needs distinct eigenvalues to separate them. The Legendre row is where the swap is bluntest, and it is the reason the table alone cannot be read as an eigenvalue problem: it carries \(q = n(n+1)\) and \(\lambda^{2} = 0\), so Equation (9.55) has no spectral parameter left, whereas the proof of Theorem 9.111 reads the same equation with \(q = 0\) and \(\lambda^{2} = n(n+1)\), distinct for distinct degrees.
Two rows of Table 9.2 are of a different character and are included because later parts of the book need them by name. The modified Bessel row (Definition 9.105) is the ordinary Bessel row with the sign of \(q\) reversed, which is exactly the difference between oscillation and monotone growth, and its solutions therefore carry no zeros and no eigenvalue condition (Proposition 9.106). The Mathieu row Equation (9.43) is the only entry whose coefficient \(q\) is periodic rather than polynomial or exponential: its eigenvalue problem is not a boundary-value problem on an interval but the stability question of Corollary 9.41, and its named solutions \(\mathrm{ce}_{n}\) and \(\mathrm{se}_{n}\) are the periodic solutions that exist only on the boundaries of the instability regions (Proposition 9.43). Finally, the elliptic integrals and Jacobi functions of Section 9.12 appear in no row at all, and deliberately: they do not solve a linear Sturm–Liouville equation but invert a quadrature, and Equation (9.193) — the equation \(\operatorname{sn}\) does satisfy — is nonlinear.
| Equation | $p(x)$ | $q(x)$ | $\lambda^2$ | $r(x)$ | Solutions |
|---|---|---|---|---|---|
| Helmholtz (1D) | $1$ | $0$ | $\lambda^2$ | $1$ | $\ee^{\pm\ii\lambda x}$ |
| Bessel | $x$ | $x$ | $-\nu^2$ | $1/x$ | $J_\nu,\ N_\nu$ |
| Modified Bessel | $x$ | $-x$ | $-\nu^2$ | $1/x$ | $I_\nu,\ K_\nu$ |
| Legendre | $1-x^2$ | $n(n+1)$ | $0$ | $1$ | $P_n,\ Q_n$ |
| Associated Legendre | $1-x^2$ | $n(n+1)$ | $-m^2$ | $\dfrac{1}{1-x^2}$ | $P_n^m,\ Q_n^m$ |
| Hermite | $\ee^{-x^2}$ | $0$ | $2n$ | $\ee^{-x^2}$ | $H_n$ |
| Laguerre | $x\,\ee^{-x}$ | $0$ | $n$ | $\ee^{-x}$ | $L_n$ |
| Mathieu | $1$ | $-2h\cos 2x$ | $a$ | $1$ | $\mathrm{ce}_n,\ \mathrm{se}_n$ |