Nonlinear Dynamics and Chaos
Determinism does not imply predictability. The equations of this part are ordinary differential equations whose solutions exist and are unique, and yet—for the generic nonlinear system—two states closer than any measurement can distinguish separate exponentially fast, so that the deterministic flow forgets its initial condition on a finite horizon. This chapter develops the geometry of phase-space flows, fixed points and their bifurcations, limit cycles, and the Lyapunov exponents that quantify sensitivity; it then follows deterministic chaos from Poincaré's three-body analysis to its experimental record, from convecting fluids and driven pendula to Chua-type circuits. The standard modern text is [Strogatz:1994].
The chapter is the counterpoint to the integrable mechanics of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy: there the invariant tori organize phase space, and here the KAM theorem states exactly how much of that order survives perturbation, in the symplectic language of Symplectic Geometry of Phase Space. On the dissipative side it supplies what Fluid Dynamics needs for the transition to turbulence. Chaos is classical determinism at its most honest, and its measured record—period-doubling constants, Lyapunov exponents, fractal dimensions—is as quantitative as any spectroscopy in this book.
Nonlinear dynamics is a branch of mathematics before it is a branch of physics, and this chapter is careful about the difference. Everything that is a computation about a particular system—the amplitude of the van der Pol cycle, the contraction rate of the Lorenz flow, the parameter at which the logistic two-cycle loses stability, the locations of the Kirkwood gaps—is derived here in full. Four general theorems are quoted: Poincaré–Bendixson (Theorem 32.30), Hartman–Grobman (Theorem 32.15), the multiplicative ergodic theorem (Theorem 32.74) and KAM (Theorem 32.51). Each is mathematics that belongs in Ordinary Differential Equations and Sturm–Liouville Theory or in a measure-theoretic chapter that Part II does not yet contain, and each carries a flagged remark saying so where it is used. Three further results—the Hopf bifurcation theorem (Theorem 32.25), the conjugacy of the horseshoe to a shift (Theorem 32.42) and Poincaré's non-integrability theorem (Theorem 32.47)—are likewise quoted, and each says so in its own title. Two further quotations are narrower and are marked where they occur: the existence of the fixed point of the doubling operator, from which Feigenbaum's constants come (Remark 32.88), and the Newhouse–Ruelle–Takens classification of what replaces a destroyed torus (Remark 32.94). Everything quoted is quoted with its primary source where this bibliography has one, and where it does not the attribution is made in words and said to be uncited. Nothing in this chapter's own derivations rests on an unstated result.
The dynamical systems of this subject are almost always written down in dimensionless form—the Lorenz equations, the logistic map, the standard map—and the temptation is to treat the resulting numbers as if they were physical. They are not, and this chapter never quotes one as if it were. Every model below is introduced first in SI variables and then reduced by an explicit scaling, stated in full, so that each dimensionless parameter can be read back as a ratio of measured quantities: the Lorenz \(\sigma\) is the Prandtl number \(\nu/\kappa\) of the fluid, its \(r\) is the Rayleigh number in units of its critical value, the standard map's \(K\) is \(K'T^{2}/I\) built from a kick energy, a kick period and a moment of inertia. Measured quantities—Lyapunov times, cell temperatures, dimensions of attractors reconstructed from data—carry SI units and uncertainties throughout.
Flows in phase space
Dynamical systems, flows and maps
The object of study is a first-order autonomous system of ordinary differential equations. Any higher-order system is brought to this form by taking the derivatives as further coordinates, and any explicitly time-dependent system by taking the time itself as a further coordinate, so nothing is lost by the restriction.
A continuous dynamical system on a phase space \(U\subseteq\R^{n}\) is an autonomous first-order system
with \(\vect{f}\) continuously differentiable. If \(\vect{x}\) carries the SI units of its physical coordinates, \(\vect{f}\) carries those units divided by \(\mathrm{s}\). The flow is the map \(\varphi_{t}(\vect{x}_{0})=\vect{x}(t)\) carrying an initial state to the state \(t\) later. A discrete dynamical system is an iterated map \(\vect{x}_{n+1}=\vect{P}(\vect{x}_{n})\); the two are connected by Definition 32.38 below. Rests on Definitions 9.1 and 13.82.
Wherever solutions of Equation (32.1) exist for all the times involved,
Rests on Definition 32.3 and Theorem 9.8.
Derives Proposition 32.4. Fix \(\vect{x}_{0}\) and set \(\vect{y}(t)=\varphi_{t}(\varphi_{s} (\vect{x}_{0}))\) and \(\vect{z}(t)=\varphi_{t+s}(\vect{x}_{0})\). Both solve Equation (32.1)—the second because the system is autonomous, so a time-translated solution is again a solution—and both take the value \(\varphi_{s}(\vect{x}_{0})\) at \(t=0\). Since \(\vect{f}\) is continuously differentiable it is locally Lipschitz (Definition 9.4), so by Picard–Lindelöf (Theorem 9.8) the initial value problem has exactly one solution and \(\vect{y}=\vect{z}\). The remaining two statements are the cases \(s=0\) and \(s=-t\).
∎Two distinct trajectories of Equation (32.1) never meet, and a single trajectory never crosses itself: a trajectory that returns to a point it has already visited is a closed orbit, traversed periodically forever. Rests on Proposition 32.4 and Theorem 9.8.
Derives Corollary 32.5. If two trajectories shared a point \(\vect{p}\) at times \(t_{1}\) and \(t_{2}\), then by Proposition 32.4 each is the flow of \(\vect{p}\) time-translated, hence they are the same trajectory. If one trajectory has \(\vect{x}(t_{0}+T)=\vect{x}(t_{0})\) with \(T>0\), then \(\varphi_{T}\) fixes \(\vect{x}(t_{0})\) and, by the group property, \(\varphi_{t+T}=\varphi_{t}\circ\varphi_{T}\) agrees with \(\varphi_{t}\) on that point for every \(t\): the solution is \(T\)-periodic.
∎This is why phase portraits in the plane are so rigid, and it is the first hint of the dimension counting that runs through the whole chapter: a bounded, non-self-intersecting curve in a plane has very few places to go.
Let \(V(t)\) be the volume of the image \(\varphi_{t}(V_{0})\) of a region \(V_{0}\) of phase space carried along by the flow Equation (32.1). Then
Rests on Definition 32.3 and Theorem 7.133.
Derives Theorem 32.6. In a time \(\dd t\) a surface element \(\dd\vect{S}\) of the boundary is swept through the displacement \(\vect{f}\,\dd t\), so it contributes \(\vect{f}\cdot\dd\vect{S}\,\dd t\) to the enclosed volume. Summing over the boundary,
and the divergence theorem (Theorem 7.133) converts the surface integral into Equation (32.3).
∎A flow is volume-preserving, or conservative, where \(\vect{\nabla}\cdot\vect{f}=0\), and dissipative where \(\vect{\nabla}\cdot\vect{f}<0\) throughout the region of interest. Rests on Theorem 32.6.
Hamiltonian flows are conservative: that is exactly Liouville's theorem (Theorem 24.21), \(\Div X_{\Ham}=0\), and it is what makes the Hamiltonian chaos of Section 32.4 qualitatively different from the dissipative chaos of Section 32.5. A Hamiltonian system can have no attractor at all, because an attractor would be a set of smaller volume than its basin.
A compact set \(A\) invariant under the flow is an attractor if it has a neighbourhood \(B\) every point of which is carried into \(A\)—\(\varphi_{t}(\vect{x})\to A\) as \(t\to\infty\) for all \(\vect{x}\in B\)—and if \(A\) contains a dense orbit, so that it cannot be split into two smaller attractors. The largest such \(B\) is the basin of attraction of \(A\). Rests on Definitions 6.9 and 32.3.
Three facts together fix where chaos can live. A one-dimensional flow \(\dot{x}=f(x)\) is monotone between its fixed points, so every bounded solution converges to one of them. A two-dimensional flow is confined by Poincaré–Bendixson (Theorem 32.30) to fixed points and closed orbits. Only from three dimensions on can a bounded trajectory avoid both without crossing itself, which by Corollary 32.5 it must. For maps the count is one lower, because a map's orbit is a sequence of points and has nothing to cross: a one-dimensional map can already be chaotic, which is what makes the logistic map of Section 32.6.1 the cheapest laboratory in the subject.
The programme itself—study the geometry of the family of solutions rather than hunt for a formula for one of them—is Poincaré's, set out in the memoir on curves defined by a differential equation [Poincare:1881] and carried into celestial mechanics in [Poincare:1890] [Poincare:1892].
Fixed points and Lyapunov stability
A fixed point (equilibrium, critical point) of Equation (32.1) is a point \(\vect{x}^{*}\) with \(\vect{f}(\vect{x}^{*})=\vect{0}\). It is a trajectory in its own right: \(\varphi_{t}(\vect{x}^{*})=\vect{x}^{*}\) for all \(t\). Rests on Definition 32.3.
A fixed point \(\vect{x}^{*}\) is stable in the sense of Lyapunov if for every \(\varepsilon>0\) there is a \(\delta>0\) such that \(\abs{\vect{x}_{0}-\vect{x}^{*}}<\delta\) implies \(\abs{\varphi_{t}(\vect{x}_{0})-\vect{x}^{*}}<\varepsilon\) for all \(t\ge0\); it is asymptotically stable if in addition \(\varphi_{t}(\vect{x}_{0})\to\vect{x}^{*}\); and unstable otherwise [Lyapunov:1892]. Rests on Definition 32.10.
The two conditions are independent and both are needed. A point can be stable without being asymptotically stable—the centre of a frictionless oscillator, whose neighbours orbit it forever without approaching—and, less obviously, a point can attract every nearby trajectory while failing to be stable, if the approach takes the trajectory far away first.
Write \(\vect{x}=\vect{x}^{*}+\delta\vect{x}\). Taylor's theorem (Theorem 7.38) applied to each component of \(\vect{f}\) gives
in which \(A\) is the Jacobian matrix of \(\vect{f}\) at \(\vect{x}^{*}\), with SI units of \(/\mathrm{s}\). The fixed point is hyperbolic if no eigenvalue of \(A\) has zero real part. Rests on Definition 32.10 and Theorem 7.38.
Let \(A\) be diagonalizable with eigenvalues \(\lambda_{1},\dots, \lambda_{n}\). If \(\Re\lambda_{i}<0\) for every \(i\), then \(\vect{x}^{*}\) is asymptotically stable. If \(\Re\lambda_{i}>0\) for some \(i\), then \(\vect{x}^{*}\) is unstable. Rests on Definition 32.12 and Theorem 5.77.
Derives Theorem 32.13. Consider first the linear system \(\dot{\vect{y}}=A\vect{y}\) alone. Let \(\set{\vect{v}_{i}}\) be a basis of eigenvectors (Definition 5.69), which exists by hypothesis (Theorem 5.77), and expand \(\vect{y}_{0}=\sum_{i}c_{i}\vect{v}_{i}\). Then
satisfies the system—differentiate term by term and use \(A\vect{v}_{i}=\lambda_{i}\vect{v}_{i}\)—and takes the value \(\vect{y}_{0}\) at \(t=0\); by Theorem 9.8 it is the only solution. Its norm is bounded by \(C\max_{i}\ee^{(\Re\lambda_{i})t}\abs{\vect{y}_{0}}\) with \(C\) the condition number of the eigenbasis, so every solution decays to zero when all \(\Re\lambda_{i}<0\), and the component along an eigenvector with \(\Re\lambda_{i}>0\) grows without bound.
The passage from the linear system to Equation (32.4) uses that the neglected terms are \(\mathcal{O}(\abs{\delta\vect{x}} ^{2})\): choosing \(\delta\) small enough that the quadratic remainder is smaller than half the linear decay rate on the ball of radius \(\delta\), the norm is still strictly decreasing there, which gives both conditions of Definition 32.11. The instability statement follows from the same estimate applied on the unstable eigendirection.
∎Theorem 32.13 was stated for diagonalizable \(A\) because that is the case Part II supports. The general case—\(A\) with a non-trivial Jordan block, whose solutions carry polynomial prefactors \(t^{k}\ee^{\lambda t}\)—needs the Jordan normal form, which Linear Algebra and Representation Theory does not contain. Nothing in this chapter uses the degenerate case: every Jacobian evaluated below is diagonalizable at the parameter values quoted.
Near a hyperbolic fixed point the flow of Equation (32.1) is topologically conjugate to the flow of its linearization Equation (32.4): there is a homeomorphism of a neighbourhood of \(\vect{x}^{*}\) onto a neighbourhood of the origin carrying trajectories of the one onto trajectories of the other and preserving both the direction and the parametrization of time [Hartman:1960]; in particular the fixed point is unstable whenever some eigenvalue of the Jacobian has positive real part. This is Theorem 9.32, written in the notation of this chapter. Rests on Theorem 9.32, Definition 32.12 and Definition 6.7.
Theorem 32.15 is a theorem of ordinary differential equations and belongs with the Picard–Lindelöf theory of Ordinary Differential Equations and Sturm–Liouville Theory, which states it as Theorem 9.32 and records there, in Remark 9.33, why it is quoted rather than derived. It is restated here because it is what licenses reading a phase portrait off the eigenvalues, and it is used below only in that qualitative role. Its content is exactly that linearization is trustworthy and no more than that: the conjugacy is a homeomorphism, not a diffeomorphism, so it preserves the topology of the portrait—which trajectories go where—but not rates or angles, and it says nothing at all at a non-hyperbolic point. Every bifurcation in Section 32.1.3 happens precisely at a non-hyperbolic point, which is why bifurcation theory cannot be done by linearization alone.
For \(n=2\) write \(\tau=\tr A\) and \(\Delta=\det A\). The eigenvalues are
and the fixed point is classified as in Table 32.1. Rests on Definition 32.12 and Theorem 5.71.
Derives Proposition 32.17. The characteristic polynomial of a \(2\times2\) matrix is \(\lambda^{2}-\tau\lambda+\Delta\), whose roots (Theorem 5.71) are Equation (32.6). If \(\Delta<0\) the discriminant exceeds \(\tau^{2}\), so the roots are real with opposite signs: a saddle, unstable in one direction and stable in the other. If \(\Delta>0\) and \(\tau^{2}>4\Delta\) the roots are real with the sign of \(\tau\): a node. If \(\Delta>0\) and \(\tau^{2}<4\Delta\) they are a complex conjugate pair with real part \(\tau/2\): a focus, spiralling in for \(\tau<0\) and out for \(\tau>0\). If \(\tau=0\) and \(\Delta>0\) they are purely imaginary and the linearized orbits are closed: a centre. Theorem 32.13 then assigns the stabilities.
∎| Condition | Type | Stability |
|---|---|---|
| $\Delta<0$ | saddle | unstable |
| $\Delta>0$, $\tau^{2}>4\Delta$, $\tau<0$ | node | asymptotically stable |
| $\Delta>0$, $\tau^{2}>4\Delta$, $\tau>0$ | node | unstable |
| $\Delta>0$, $\tau^{2}<4\Delta$, $\tau<0$ | focus | asymptotically stable |
| $\Delta>0$, $\tau^{2}<4\Delta$, $\tau>0$ | focus | unstable |
| $\Delta>0$, $\tau=0$ | centre (linear) | undecided |
Take
The Jacobian at the origin is \(A=\begin{pmatrix}0&-1\\1&0\end{pmatrix}\), with \(\tau=0\) and \(\Delta=1\): a linear centre. But with \(\rho=\tfrac{1}{2}(x^{2}+y^{2})\),
everywhere except the origin, so every trajectory spirals in and the origin is asymptotically stable. Changing the sign of the cubic terms makes it unstable, again without touching \(A\). A centre is not hyperbolic, so Theorem 32.15 does not apply, and this is what its hypothesis is protecting against. Rests on Proposition 32.17 and Theorem 32.15.
Lyapunov's second, or direct, method dispenses with linearization altogether [Lyapunov:1892]. It is the method that survives at a non-hyperbolic point, and in mechanics it is usually the energy.
A continuously differentiable \(V\) on a neighbourhood \(N\) of a fixed point \(\vect{x}^{*}\) is a Lyapunov function if \(V(\vect{x}^{*})=0\), \(V>0\) on \(N\setminus\set{\vect{x}^{*}}\), and
It is strict if the inequality is strict off \(\vect{x}^{*}\). Rests on Definitions 7.97 and 32.10.
If a Lyapunov function exists on \(N\) then \(\vect{x}^{*}\) is stable; if it is strict, \(\vect{x}^{*}\) is asymptotically stable. Rests on Definitions 32.11 and 32.19.
Derives Theorem 32.20. Given \(\varepsilon>0\), choose the ball \(B_{\varepsilon}\subseteq N\) and let \(m=\min\set{V(\vect{x}):\abs{\vect{x}-\vect{x}^{*}}=\varepsilon}\), which is positive because \(V\) is positive and continuous on a compact set (Theorem 7.24). By continuity there is \(\delta>0\) with \(V<m\) on the ball of radius \(\delta\). By Equation (32.7) \(V\) never increases along a trajectory, so a trajectory starting inside that ball can never reach the sphere of radius \(\varepsilon\), on which \(V\ge m\): that is stability. If \(\dot{V}<0\) strictly, then \(V\) decreases monotonically along the trajectory and is bounded below, so it converges to some \(V_{\infty}\ge0\); were \(V_{\infty}>0\), the trajectory would remain in the compact annulus \(\set{V_{\infty}\le V\le V(\vect{x}_{0})}\) on which \(\dot{V}\le-c<0\), forcing \(V\) to fall below \(V_{\infty}\) in finite time. Hence \(V_{\infty}=0\) and \(\vect{x}\to\vect{x}^{*}\).
∎A pendulum of mass \(m\) and length \(\ell\) with linear viscous damping of coefficient \(b\) (SI units \(\mathrm{kg}\,\mathrm{m}^{2}/\mathrm{s}\)) obeys
Take \(V\) to be the mechanical energy measured from the hanging position,
which is positive for \(0<\abs{\theta}<\pi\) and vanishes at the fixed point \((\theta,\dot{\theta})=(0,0)\). Differentiating along Equation (32.8),
so the hanging position is stable by Theorem 32.20. The inequality is not strict—it fails on the whole line \(\dot{\theta}=0\)—so asymptotic stability needs the extra remark that the only trajectory lying entirely in \(\set{\dot{\theta}=0}\) is the fixed point itself, since Equation (32.8) then forces \(\sin\theta=0\). With \(b=0\) the energy is exactly conserved, \(V\) is a first integral, and the fixed point is a centre: the isochronous small oscillation of Oscillations and Mechanical Waves, measured in Experiment: The Pendulum. Rests on Theorem 32.20 and Equation (28.8).
Bifurcations
Let \(\dot{\vect{x}}=\vect{f}(\vect{x};\mu)\) depend on a control parameter \(\mu\). A bifurcation occurs at \(\mu=\mu_{c}\) if the qualitative structure of the phase portrait—the number, type or stability of its invariant sets—differs on the two sides of \(\mu_{c}\). By Theorem 32.15 this can happen only where some fixed point or periodic orbit fails to be hyperbolic. Rests on Definition 32.3 and Theorem 32.15.
The one-dimensional bifurcations come in three normal forms, and the list is short for a reason: expanding \(f(x;\mu)\) about the non-hyperbolic point and discarding what does not matter leaves only these possibilities.
The system \(\dot{x}=\mu-x^{2}\) has two fixed points \(x^{*}=\pm\sqrt{\mu}\) for \(\mu>0\), of which \(+\sqrt{\mu}\) is stable and \(-\sqrt{\mu}\) unstable; they merge at \(\mu=0\) and there is no fixed point for \(\mu<0\). The distance between them grows as \(\sqrt{\mu}\), and the relaxation rate towards the stable one vanishes as \(\sqrt{\mu}\). Rests on Definition 32.22 and Theorem 32.13.
Derives Proposition 32.23. \(\mu-x^{2}=0\) has the two real roots \(\pm\sqrt{\mu}\) for \(\mu>0\), a double root at \(\mu=0\) and none for \(\mu<0\). The linearization is \(f'(x)=-2x\), equal to \(-2\sqrt{\mu}<0\) at \(+\sqrt{\mu}\) and \(+2\sqrt{\mu}>0\) at \(-\sqrt{\mu}\); Theorem 32.13 in one dimension gives the two stabilities and identifies \(2\sqrt{\mu}\) as the relaxation rate, which vanishes at the bifurcation. The slowing down at \(\mu\to0^{+}\) is generic and observable: it is the critical slowing down that announces a saddle-node in a real system before the fixed points actually collide.
∎(i) \(\dot{x}=\mu x-x^{2}\) has fixed points \(x^{*}=0\) and \(x^{*}=\mu\) for all \(\mu\); they exchange stability at \(\mu=0\). (ii) \(\dot{x}=\mu x-x^{3}\) has the single fixed point \(x^{*}=0\), stable, for \(\mu<0\), and for \(\mu>0\) has \(x^{*}=0\) unstable together with a symmetric pair \(x^{*}=\pm\sqrt{\mu}\), both stable (supercritical pitchfork). (iii) \(\dot{x}=\mu x+x^{3}\) has an unstable pair \(x^{*}=\pm\sqrt{-\mu}\) for \(\mu<0\) which annihilates the stable origin at \(\mu=0\), leaving nothing stable beyond (subcritical pitchfork). Rests on Definition 32.22 and Theorem 32.13.
Derives Proposition 32.24. (i) \(x(\mu-x)=0\) gives the two branches; \(f'(x)=\mu-2x\) equals \(\mu\) at the origin and \(-\mu\) at \(x^{*}=\mu\), so the signs are opposite on either side of \(\mu=0\) and the stabilities are exchanged there. (ii) \(x(\mu-x^{2})=0\) gives the origin always and \(\pm\sqrt{\mu}\) for \(\mu>0\); \(f'(x)=\mu-3x^{2}\) equals \(\mu\) at the origin and \(-2\mu<0\) on the outer pair. (iii) is (ii) with \(x^{3}\to-x^{3}\): \(f'=\mu+3x^{2}\) equals \(\mu\) at the origin and \(-2\mu>0\) on the pair, which therefore exists only for \(\mu<0\) and is unstable.
∎The distinction in Proposition 32.24 between the supercritical and subcritical cases is the whole practical content of bifurcation theory in the laboratory. A supercritical transition is continuous and reversible: the new state grows out of the old like \(\sqrt{\mu-\mu_{c}}\), and sweeping the parameter back retraces the same curve. A subcritical transition is discontinuous and hysteretic: the old state loses stability with nothing nearby to fall onto, the system jumps to a distant state, and it stays there when the parameter is swept back past \(\mu_{c}\). Euler buckling (Phenomenon 30.52) is the textbook supercritical pitchfork, its symmetry being the left–right symmetry of the strut; the loss of the convecting rolls in Section 32.5.1 is subcritical, and that is why the Lorenz trajectory has nowhere local to go.
Let \(\dot{\vect{x}}=\vect{f}(\vect{x};\mu)\) in \(\R^{2}\) have a fixed point whose Jacobian has a complex conjugate pair of eigenvalues \(\lambda_{\pm}(\mu)=\alpha(\mu)\pm\ii\omega(\mu)\) with \(\alpha(\mu_{c})=0\), \(\omega(\mu_{c})\neq0\) and \(\alpha'(\mu_{c})\neq0\). Then, generically, a one-parameter family of periodic orbits is born at \(\mu_{c}\), of period tending to \(2\pi/\omega(\mu_{c})\) and of amplitude growing as \(\sqrt{\abs{\mu-\mu_{c}}}\) [Hopf:1942]. Rests on Proposition 32.17 and Definition 32.22.
In polar coordinates the normal form of a Hopf bifurcation is
For \(a>0\) (supercritical) a stable limit cycle of radius \(\rho^{*}=\sqrt{\mu/a}\) exists for \(\mu>0\); for \(a<0\) (subcritical) an unstable cycle of radius \(\sqrt{\mu/a}\) exists for \(\mu<0\) and there is nothing local beyond \(\mu=0\). Rests on Theorem 32.25.
Derives Proposition 32.26. The radial equation is decoupled and is exactly the pitchfork of Proposition 32.24 restricted to \(\rho>0\): its non-zero fixed point is \(\rho^{*}=\sqrt{\mu/a}\) and its linearization there is \(\mu-3a\rho^{*2}=-2\mu\), negative for \(\mu>0\) when \(a>0\). The angular equation then gives a periodic orbit of angular frequency \(\omega+b\mu/a\), which tends to \(\omega\) as \(\mu\to0\). The \(\sqrt{\mu}\) growth of the amplitude in Theorem 32.25 is the \(\sqrt{\mu/a}\) here.
∎Proposition 32.26 is an exact statement about Equation (32.9); Theorem 32.25 is the assertion that a general system satisfying the eigenvalue conditions can be brought to that form by a smooth change of coordinates near the fixed point. That reduction is a theorem of normal-form theory and belongs with the ordinary-differential-equation material of Ordinary Differential Equations and Sturm–Liouville Theory; it is quoted here from [Hopf:1942]. It also requires the implicit function theorem, to follow the fixed point as \(\mu\) varies; that one is available, as Theorem A.286, so the normal-form reduction is the only piece owed. Every Hopf bifurcation located in this chapter is located by the eigenvalue condition \(\alpha(\mu_{c})=0\) directly, computed from the actual Jacobian, so none of the derivations below depends on the reduction.
A system is structurally stable if every sufficiently small smooth perturbation of \(\vect{f}\) has a topologically conjugate phase portrait. Hyperbolic fixed points and hyperbolic limit cycles are structurally stable, by Theorem 32.15 and its analogue for periodic orbits; centres and the exact normal forms above are not, which is why a bifurcation is a codimension-one event—it is seen by sweeping one parameter, and not otherwise. What is plotted against that parameter is the bifurcation diagram: the location and stability of every invariant set as a function of \(\mu\). It is the experimental deliverable of the whole subject. Every measurement in Section 32.7 is a bifurcation diagram in disguise—the Rayleigh number swept in a convection cell, the rotation rate in a Couette cell, the drive amplitude in a circuit—and the universal numbers of Section 32.6.2 are properties of its geometry, not of the system that produced it.
Limit cycles
The Poincaré–Bendixson theorem
A limit cycle is an isolated closed orbit: a periodic trajectory some neighbourhood of which contains no other closed orbit. It is attracting (stable) if all nearby trajectories approach it as \(t\to+\infty\), and repelling if they do so as \(t\to-\infty\). Rests on Corollary 32.5 and Definition 32.8.
The isolation is the whole point, and it is what separates a limit cycle from the continuous family of closed orbits surrounding a centre. A conservative system has families; only a system that dissipates in one region and pumps in another can have an isolated one. Limit cycles were named and first analysed in the second instalment of Poincaré's memoir [Poincare:1882].
Let \(\vect{f}\) be continuously differentiable on an open subset of the plane, and let a trajectory of Equation (32.1) remain, for all \(t\ge t_{0}\), inside a compact region containing no fixed point. Then that trajectory either is a closed orbit or approaches one [Bendixson:1901] [Poincare:1882]. This is Theorem 9.34, written in the notation of this chapter. Rests on Theorem 9.34, Corollary 32.5 and Definition 6.9.
The proof is planar topology—a transversal segment, and the Jordan curve theorem, from which the monotonicity of successive crossings along that segment follows because a trajectory arc together with the chord joining its ends bounds a region the flow cannot re-enter—and it is mathematics that belongs with the ordinary differential equations of Ordinary Differential Equations and Sturm–Liouville Theory, which states it as Theorem 9.34 and sets out in Remark 9.35 exactly which planar facts the proof rests on and which of them this treatise does not carry. It is restated here because the entire architecture of this chapter rests on its corollary: the theorem is what makes the plane too small for chaos, and therefore what sends the subject into three dimensions and into maps.
A bounded trajectory of a continuously differentiable autonomous flow in the plane cannot be chaotic: as \(t\to\infty\) it converges to a fixed point, to a closed orbit, or to a cycle of connecting trajectories between fixed points. Deterministic chaos in a continuous autonomous flow therefore requires at least three phase-space dimensions. Rests on Theorem 32.30 and Remark 32.9.
Derives Corollary 32.32. A bounded trajectory has compact closure. If that closure contains no fixed point, Theorem 32.30 gives convergence to a closed orbit. If it contains fixed points, the limit set consists of those fixed points together with trajectories connecting them, since by Corollary 32.5 nothing else can accumulate without self-intersection in a plane. In neither case do nearby trajectories separate exponentially forever while remaining bounded, which is what Phenomenon 32.72 requires. In three dimensions the argument fails at its first step: a curve may wander indefinitely in a bounded region without crossing itself, because the plane no longer separates it.
∎If \(\vect{\nabla}\cdot\vect{f}\) has one strict sign throughout a simply connected region \(D\) of the plane, then Equation (32.1) has no closed orbit lying entirely in \(D\) [Bendixson:1901]. Rests on Theorem 7.131 and Definition 6.17.
Derives Proposition 32.33. Let \(C\) be a closed orbit in \(D\) bounding the region \(S\). Green's theorem (Theorem 7.131) gives
because \(\dot{x}_{1}=f_{1}\) and \(\dot{x}_{2}=f_{2}\) on the orbit make the integrand vanish identically. The left-hand side cannot be zero if the integrand has one strict sign, so no such \(C\) exists. The criterion is only negative: it excludes cycles, it never produces one—for that one needs a trapping region and Theorem 32.30.
∎A second exclusion principle is topological. The index of a closed curve \(C\) containing no fixed point is the winding number (Definition 6.20) of the vector field \(\vect{f}\) around it; it is an integer, unchanged by continuous deformation of \(C\) (Lemma 6.22), equal to \(+1\) for a node, focus or centre and \(-1\) for a saddle, and it is the sum of the indices of the fixed points enclosed. Since a closed orbit is everywhere tangent to \(\vect{f}\), its own index is \(+1\). Hence every closed orbit encloses at least one fixed point, and the fixed points it encloses have indices summing to \(+1\): a limit cycle cannot surround a single saddle and nothing else, and it cannot lie in a region with no fixed point at all.
Relaxation oscillations: van der Pol
The paradigm of a system with a limit cycle came from radio engineering. A triode circuit in which the effective resistance is negative at small amplitude and positive at large amplitude obeys, for the voltage \(v\) across the tuned circuit,
in which \(\alpha\) has SI units \(/\mathrm{s}\), \(v_{0}\) is the voltage at which the damping changes sign, and \(\omega_{0}\) is the resonant angular frequency of the inductance \(L\) and capacitance \(C\) [vanderPol:1926]. Scaling to the dimensionless amplitude \(u=v/v_{0}\) and the dimensionless time \(s=\omega_{0}t\) puts it in the form used everywhere below,
with the single dimensionless control parameter \(\varepsilon\). Every statement made about Equation (32.11) converts back by \(v=v_{0}u\) and \(t=s/\omega_{0}\).
A triode circuit whose damping changes sign with amplitude settles into a periodic oscillation whose amplitude and period belong to the circuit alone: started from an arbitrarily small disturbance the oscillation grows until it reaches that amplitude, and started from a larger one it decays to the same amplitude [vanderPol:1926]. This is qualitatively unlike the linear oscillator of Oscillations and Mechanical Waves, whose amplitude is set by the initial conditions and is therefore arbitrary, and unlike a driven resonance, since nothing periodic is applied. In the strongly nonlinear regime the waveform is not sinusoidal at all but a relaxation oscillation of slow charging and abrupt discharge, whose period is set by the slow stage and so scales with the nonlinearity rather than with the resonant frequency of the circuit [vanderPol:1926]. Rests on Equation (32.10) and Definition 32.29.
Derivation. Derives Phenomenon 32.35. Weak nonlinearity. For \(\varepsilon\ll1\) write Equation (32.11) as a perturbed harmonic oscillator and follow the energy \(E=\tfrac{1}{2}(\dot{u}^{2}+u^{2})\), where the overdot now means \(\dd/\dd s\). Along a solution,
Over one cycle the solution is nearly \(u=a\cos\psi\) with \(\dot{u}=-a\sin\psi\) and \(\psi\) advancing by \(2\pi\), so averaging Equation (32.12) with \(\avg{\sin^{2}\psi}=\tfrac{1}{2}\) and \(\avg{\cos^{2}\psi\sin^{2}\psi}=\tfrac{1}{8}\) gives
and since \(E=a^{2}/2\) this is an equation for the slowly varying amplitude,
Equation (32.13) is the supercritical pitchfork of Proposition 32.24 in the radial variable: for any \(a>0\) the amplitude moves monotonically to \(a=2\), from below if it starts below and from above if it starts above. That is exactly the observation, and in physical variables the cycle has amplitude \(v=2v_{0}\)—twice the voltage at which the damping changes sign, independent of how the oscillation was started and independent of \(\alpha\). Its angular frequency is \(\omega_{0}\) to leading order.
Strong nonlinearity. For \(\varepsilon\gg1\) use the Liénard variable \(w=\dot{u}/\varepsilon+F(u)\) with \(F(u)=u^{3}/3-u\), which turns Equation (32.11) into the first-order pair
Since \(\varepsilon\) is large, \(u\) moves fast unless \(w\approx F(u)\); the trajectory therefore jumps horizontally onto a branch of the cubic \(w=F(u)\) in a time of order \(\varepsilon^{-1}\) and then crawls along it. A branch is attracting when \(\pp\dot{u}/\pp u=-\varepsilon F'(u)\) is negative, that is when \(F'(u)=u^{2}-1>0\): the two outer branches \(u>1\) and \(u<-1\). On such a branch \(\dot{w}=F'(u)\dot{u}=(u^{2}-1)\dot{u}\) equals \(-u/\varepsilon\), so
On the branch \(u>1\) the trajectory enters at \(u=2\)—where the previous jump deposited it, since \(F(2)=F(-1)=2/3\)—and creeps down to the fold at \(u=1\), at which the branch turns and it jumps across to \(u=-2\), because \(F(-2)=F(1)=-2/3\). The time for one slow stage is therefore
By the symmetry \(u\to-u\), \(w\to-w\) of Equation (32.14) the cycle consists of two such stages and two jumps, and the jumps together with the turning regions at the folds contribute only at order \(\varepsilon^{-1/3}\). Hence
using \(3-2\ln2=1.6137\). The physical period grows in proportion to the nonlinearity \(\alpha\) and falls as the inverse square of the resonant angular frequency—it is not the resonant period \(2\pi/\omega_{0}\) at all, which is the observation stated. The waveform is likewise not sinusoidal: two slow ramps of duration \(\propto\varepsilon\) separated by jumps of duration \(\propto\varepsilon^{-1}\).
∎The derivation above establishes the cycle asymptotically at each end
of the range of \(\varepsilon\). That Equation (32.11) has exactly one
closed orbit, attracting every trajectory but the fixed point at the
origin, and for every \(\varepsilon>0\), is Liénard's theorem
(A. Liénard, Revue générale de l'électricité
23 (1928) 901–912 and 946–954). That paper is not in
this bibliography and the statement is therefore uncited here; the
reader should treat the global uniqueness as quoted, while the
amplitude Equation (32.13) and the period
Equation (32.15) are derived above and stand on their own.
Note that the key entry Lienard:1898 in this bibliography is
Liénard's electrodynamics paper on the potentials of a moving
charge, a different work by the same author, and must not be cited for
this theorem.
A neon-tube relaxation oscillator driven by a sinusoidal voltage does not follow the drive: it locks to a submultiple of the drive frequency, and as a circuit parameter is varied smoothly the observed frequency descends in a staircase of whole-number ratios—van der Pol and van der Mark followed it from the fundamental down to one fortieth. At the transitions between two adjacent steps “often an irregular noise is heard in the telephone receivers before the frequency jumps to the next lower value” [vanderPol:1927]. That noise is not a defect of the apparatus and not thermal: it is the first recorded laboratory encounter with deterministic chaos. Rests on Phenomenon 32.35.
Derivation. Derives Phenomenon 32.37. The staircase follows from the mechanism of the oscillator. A neon tube in parallel with a capacitance \(C\) charged through a resistance \(R\) from a supply \(V_{s}\) has capacitor voltage
after each discharge, where \(V_{e}\) is the extinction voltage. The tube fires when \(v\) reaches the breakdown voltage \(V_{b}\), so the free period is
proportional to \(C\), which is the parameter van der Pol and van der Mark varied. Superposing a small sinusoidal voltage of period \(T_{d}\) modulates the instant at which the threshold is reached: the tube can only fire near a maximum of the drive, so the realized period is an integer multiple \(nT_{d}\), with \(n\) the number of drive periods needed for the ramp to reach \(V_{b}\). As \(C\) is increased, \(T_{0}\) grows continuously by Equation (32.16) while \(n\) can only change by whole numbers, so the observed frequency \(1/(nT_{d})\) falls in steps: the demultiplication. Each step is a locked periodic orbit, stable over a finite interval of \(C\)—the entrainment band, or Arnold tongue, of Section 32.6.3.
The irregular noise sits at the boundaries, where two locking ratios \(n\) and \(n+1\) are both marginal. There the return map of Section 32.3.1 has no stable fixed point of either period, and the sequence of firing phases neither repeats nor converges. The full account of that regime needs the circle map and is given in Section 32.6.3; what is derived here is the staircase, and the location of the noise at its risers.
∎Entrainment of this kind is generic to limit-cycle oscillators and is the nonlinear counterpart of the linear resonance of Oscillations and Mechanical Waves: a linear oscillator responds at the drive frequency and at no other, whereas a limit-cycle oscillator has a frequency of its own and either captures the drive, is captured by it, or does neither.
Poincaré sections and discrete dynamics
The return map
Let \(\Sigma\) be a hypersurface of dimension \(n-1\) in the phase space of Equation (32.1), transverse to the flow (that is, \(\vect{f}\) is nowhere tangent to \(\Sigma\)). The first-return, or Poincaré, map \(\vect{P}:\Sigma\to\Sigma\) sends a point \(\vect{x}\in\Sigma\) to the point at which the trajectory through \(\vect{x}\) next meets \(\Sigma\) in the same sense [Poincare:1892]. A fixed point of \(\vect{P}\) is a periodic orbit of the flow; a period-\(k\) orbit of \(\vect{P}\) is a periodic orbit piercing \(\Sigma\) \(k\) times. Rests on Definition 32.3 and Corollary 32.5.
This is the device that trades a flow for a map one dimension down, and with it the whole theory of iterated maps becomes available for continuous systems. It is also what makes the numerics of the subject tractable and the figures readable: a three-dimensional flow becomes a scatter of points in a plane.
Let \(\vect{x}^{*}\in\Sigma\) be a fixed point of \(\vect{P}\) and let \(\mu_{1},\dots,\mu_{n-1}\) be the eigenvalues of \(D\vect{P} (\vect{x}^{*})\), the characteristic multipliers of the corresponding periodic orbit. If \(\abs{\mu_{i}}<1\) for all \(i\) the orbit is asymptotically stable; if \(\abs{\mu_{i}}>1\) for some \(i\) it is unstable. Rests on Definition 32.38 and Theorem 5.77.
Derives Proposition 32.39. Linearizing \(\vect{P}\) about \(\vect{x}^{*}\) as in Definition 32.12 gives \(\delta\vect{x}_{k+1} =D\vect{P}\,\delta\vect{x}_{k}\) to first order, so \(\delta\vect{x}_{k}=(D\vect{P})^{k}\delta\vect{x}_{0}\). Expanding in an eigenbasis of \(D\vect{P}\) the \(i\)-th component is \(\mu_{i}^{k}\) times its initial value, which tends to zero for \(\abs{\mu_{i}}<1\) and diverges for \(\abs{\mu_{i}}>1\). The passage from the linearized map to the full one repeats the estimate at the end of the proof of Theorem 32.13, with the quadratic remainder dominated on a small enough neighbourhood. A multiplier crossing the unit circle is a bifurcation of the periodic orbit: at \(\mu=+1\) a saddle-node of cycles, at \(\mu=-1\) a period-doubling—the orbit that will drive Section 32.6—and at a complex conjugate pair crossing, the birth of a two-torus.
∎For an autonomous Hamiltonian system with two degrees of freedom, a Poincaré section taken inside a fixed energy surface has a two-dimensional return map that preserves area. Rests on Theorems 22.22 and 24.21.
Derives Proposition 32.40. Take coordinates \((q^{1},p_{1},q^{2},p_{2})\) and let \(\Sigma\) be the surface \(q^{2}=\text{const}\) within \(\Ham=E\), coordinatized by \((q^{1},p_{1})\). The flow is canonical, so by Poincaré's invariant (Theorem 22.22) the integral \(\int\dd q^{a}\,\dd p_{a}\) over a two-surface carried by the flow is conserved. Take a tube of trajectories whose ends are two regions \(S_{0}\) and \(S_{1}=\vect{P}(S_{0})\) of \(\Sigma\). On \(\Sigma\) the coordinate \(q^{2}\) is constant, so \(\dd q^{2}=0\) there and the invariant reduces to \(\int\dd q^{1}\,\dd p_{1}\); on the lateral surface of the tube the same invariant vanishes because the surface is generated by trajectories. Hence \(\int_{S_{0}}\dd q^{1}\,\dd p_{1} =\int_{S_{1}}\dd q^{1}\,\dd p_{1}\): the map preserves area. Liouville's theorem (Theorem 24.21) is the four-dimensional statement of which this is the section.
∎Proposition 32.40 is the reason the surfaces of section of Section 32.4.4 look the way they do: an area-preserving map cannot have an attractor, so the points either lie on invariant curves or wander over an area, and never spiral into anything.
The Smale horseshoe and symbolic dynamics
Chaos in a map has a certified mechanism, and it is geometrical rather than analytical: stretch, fold, and put back.
Let \(Q\) be the unit square. The horseshoe map \(H\) stretches \(Q\) uniformly by a factor \(\lambda>2\) in one direction, contracts it by \(\eta<1/2\) in the other, bends the resulting strip into a horseshoe, and lays it back across \(Q\) so that \(H(Q)\cap Q\) consists of two full horizontal strips [Smale:1967]. The invariant set is
the set of points whose entire forward and backward orbit stays in \(Q\). Rests on Definition 32.3.
\(\Lambda\) is a compact, totally disconnected, perfect set—a Cantor set—and \(H\) restricted to \(\Lambda\) is topologically conjugate to the two-sided shift \(\sigma\) on the space \(\Sigma_{2}\) of bi-infinite binary sequences, \((\sigma s)_{k}=s_{k+1}\), the conjugacy assigning to each point the record of which of the two strips its orbit occupies at each time [Smale:1967]. Rests on Definitions 6.9 and 32.41.
The shift \(\sigma\) on \(\Sigma_{2}\) has exactly \(2^{k}\) points of period \(k\) for every \(k\ge1\), hence a countable infinity of periodic orbits; it has orbits that are dense in \(\Sigma_{2}\); and it has sensitive dependence, two sequences agreeing on \(\abs{j}\le N\) being separated after \(N\) steps. Rests on Theorem 32.42.
Derives Proposition 32.43. A sequence with \(\sigma^{k}s=s\) is determined by the \(k\) entries \(s_{0},\dots,s_{k-1}\) and there are \(2^{k}\) of these, each giving one such sequence. For a dense orbit, list the finite binary words in order of length—\(0\), \(1\), \(00\), \(01\), \(10\), \(11\), \(000,\dots\)—and concatenate them into one half-infinite string, completing it to the left arbitrarily; some forward shift of the result begins with any prescribed finite word, and agreement on a long initial word is exactly closeness in \(\Sigma_{2}\). For sensitivity, take \(s\) and \(s'\) agreeing on \(\abs{j}\le N\) and differing at \(j=N+1\): after \(N+1\) shifts they differ in the zeroth place, which is the coarsest distinction the metric makes. Under the conjugacy of Theorem 32.42 every one of these statements is a statement about \(H\) on \(\Lambda\).
∎The count \(2^{k}\) in Proposition 32.43 is the horseshoe's signature: the number of distinguishable orbit segments of length \(k\) grows like \(\ee^{hk}\) with \(h=\ln2\), and that exponential growth rate of distinguishable histories is the topological entropy. A positive topological entropy is one of the standard definitions of chaos, and for the horseshoe it is exact.
The reason the construction matters physically is that horseshoes are not rare. Wherever the stable and unstable manifolds of a saddle intersect transversally—a transverse homoclinic point—some iterate of the map contains a horseshoe, and therefore all of Proposition 32.43. That is the Smale–Birkhoff homoclinic theorem, quoted here from [Smale:1967]. Its importance for this chapter is that Poincaré had already found such an intersection, and its infinitely tangled consequences, in the restricted three-body problem [Poincare:1890]: the tangle of Section 32.4.1 and the horseshoe are the same object, seen sixty years apart.
Theorem 32.42 calls \(\Lambda\) a Cantor set, and Section 32.5.4 will assign it a dimension. The construction of the Cantor set, its uncountability, and the fact that it has Lebesgue measure zero are results of real analysis and measure theory that Topological and Metric Spaces and Real Analysis do not carry; they are used descriptively here and nothing below depends on more than the counting argument of Proposition 32.43, which is elementary.
Hamiltonian chaos
The three-body problem
Two primaries of masses \(M_{1}\) and \(M_{2}\) move on a circular orbit about their common centre of mass with angular velocity \(\Omega=\left[G(M_{1}+M_{2})/a^{3}\right]^{1/2}\), by Equation (27.51). A third body of negligible mass \(m\) moves in their field. In the frame rotating with the primaries, in which the two are at rest, the Hamiltonian of the third body is
with \(L_{z}=(\vect{r}\times\vect{p})_{z}\) and \(r_{1},r_{2}\) the distances to the primaries. Since Equation (32.18) has no explicit time dependence in that frame, it is conserved; the constant is the Jacobi integral, of SI unit \(\mathrm{J}\). Rests on Equations (22.1) and (27.51).
The system has three degrees of freedom and one known constant of motion. By Liouville–Arnold (Theorem 23.31) three independent constants in involution would make it integrable, and the question Poincaré was set—and answered in the negative—was whether the remaining two exist.
For the restricted three-body problem written as an integrable system perturbed at order \(\varepsilon\) in the mass ratio, there is no first integral, other than a function of \(\Ham\) itself, that is analytic in the phase-space variables and in \(\varepsilon\) and single-valued in the angles [Poincare:1890]. The perturbation series that would construct such an integral diverges, and the geometric reason is the homoclinic tangle: the stable and unstable manifolds of an unstable periodic orbit intersect transversally, and by Corollary 32.5 each intersection forces infinitely many more, so the manifolds fold into a mesh “so complicated that I shall not even try to draw it” [Poincare:1892]. Rests on Theorem 23.31 and Remark 32.44.
An analytic first integral is constant on every trajectory and therefore constant on the closure of any trajectory. In the tangle the closure of a single unstable manifold contains a whole horseshoe (Remark 32.44), whose orbits by Proposition 32.43 come arbitrarily close to every point of a Cantor set of positive topological entropy. An analytic function constant on such a set, and on the infinitely many distinct periodic orbits it contains, is forced to be constant on an open set and hence—by analytic continuation—a function of the energy alone. The contrast with the two-body Kepler problem of Central Forces and Statics is total: there the motion is integrable, its level sets are the closed ellipses of Phenomenon 27.32, and the extra constants (angular momentum and the Laplace–Runge–Lenz vector) exist and are polynomial. Adding one more body destroys them.
The KAM theorem
Take an integrable Hamiltonian in the action–angle variables of Definition 23.28 and perturb it:
with \(\Ham_{1}\) \(2\pi\)-periodic in each angle and \(\varepsilon\) dimensionless. The unperturbed motion lies on the invariant tori of Theorem 23.31, each labelled by its actions and carrying the frequencies \(\omega^{a}(I)=\pp\Ham_{0}/\pp I_{a}\) of Equation (23.52).
Seek a canonical transformation removing the perturbation to first order, generated by \(S(I',\theta)=I'\cdot\theta+\varepsilon S_{1}(I',\theta)\). With the Fourier expansion \(\Ham_{1}=\sum_{\vect{k}\in\Z^{f}}h_{\vect{k}}(I)\, \ee^{\ii\vect{k}\cdot\theta}\), the generator is
which is undefined whenever \(\vect{k}\cdot\vect{\omega}(I)=0\) for some \(\vect{k}\neq\vect{0}\), and arbitrarily large near such a resonance. Rests on Equations (22.18) and (23.52).
Derives Proposition 32.49. The generating function of the second kind (Equation (22.21)) gives \(I_{a}=\pp S/\pp\theta^{a} =I'_{a}+\varepsilon\,\pp S_{1}/\pp\theta^{a}\), so the new Hamiltonian is
using \(\pp\Ham_{0}/\pp I_{a}=\omega^{a}\). Requiring the bracket to be independent of the angles—and taking the angle average of \(\Ham_{1}\), \(h_{\vect{0}}\), into \(\Ham'\)—leaves \(\vect{\omega}\cdot\pp S_{1}/\pp\theta =-\sum_{\vect{k}\neq\vect{0}}h_{\vect{k}}\ee^{\ii\vect{k} \cdot\theta}\). Inserting the Fourier series for \(S_{1}\) turns the derivative into a factor \(\ii\vect{k}\cdot\vect{\omega}\) and solving mode by mode gives Equation (32.20). The denominators are the resonance conditions; since the rationals are dense, every neighbourhood of every torus contains resonant tori, and this is the difficulty that defeated celestial mechanics for two centuries.
∎A frequency vector \(\vect{\omega}\in\R^{f}\) is Diophantine of type \((\gamma,\tau)\) if
with \(\gamma>0\) and \(\tau>f-1\). The condition says the frequencies are badly approximable by rationals, so the denominators in Equation (32.20) shrink only polynomially in \(\abs{\vect{k}}\) while the numerators \(h_{\vect{k}}\) of an analytic \(\Ham_{1}\) shrink exponentially. Rests on Proposition 32.49.
Let \(\Ham_{0}\) be analytic and non-degenerate, in the sense that \(\det\left(\pp^{2}\Ham_{0}/\pp I_{a}\pp I_{b}\right)\neq0\), and let \(\Ham_{1}\) be analytic. Then for every \(\gamma\) there is \(\varepsilon_{0}(\gamma)>0\) such that for \(\abs{\varepsilon}< \varepsilon_{0}\) every unperturbed torus whose frequency vector is Diophantine of type \((\gamma,\tau)\) survives as an invariant torus of the perturbed system, merely deformed, carrying the same frequencies. The surviving tori fill a set whose relative measure tends to \(1\) as \(\varepsilon\to0\), the complement being \(\mathcal{O}(\sqrt{\varepsilon})\) [Kolmogorov:1954] [Arnold:1963] [Moser:1962]. Rests on Definition 32.50 and Theorem 23.31.
Theorem 32.51 carries no derivation in this book, and the omission is a decision rather than a debt—the same treatment Part II gives Carleson's theorem in Remark 17.18. What the proof does is worth stating, because it is what makes the theorem believable and it is also what makes it long. One cannot iterate the first-order step of Proposition 32.49: each pass leaves a remainder of order \(\varepsilon^{2}\), then \(\varepsilon^{3}\), and the small denominators of Equation (32.20) force one to work on a domain shrunk at every pass, so the successive corrections lose analyticity faster than they gain smallness and the series diverges. Kolmogorov's device is to replace the ordinary perturbation series by a Newton iteration [Kolmogorov:1954]: each pass solves the linearized problem exactly and therefore squares the remainder, \(\varepsilon\to\varepsilon^{2}\to\varepsilon^{4}\), which converges fast enough to survive both the shrinking domains and the growth of the denominators permitted by the Diophantine condition Equation (32.21). Arnold supplied the analytic details [Arnold:1963] and Moser removed the analyticity hypothesis in favour of finite differentiability [Moser:1962].
Three ingredients of that argument have no home in this treatise: the Diophantine estimates themselves, the Cauchy estimates on shrinking complex strips that control each pass, and the measure-theoretic accounting that turns “the resonant set is small for each \(\vect{k}\)” into the \(\mathcal{O}(\sqrt{\varepsilon})\) complement quoted in Theorem 32.51—the last of which needs the Lebesgue measure that Part II does not carry. None of them appears anywhere else in this book. What is derived here is the obstruction the theorem overcomes (Proposition 32.49), the condition under which it can be overcome (Definition 32.50), and—in Section 32.4.3—the quantitative estimate of where the tori actually break, which is Chirikov's overlap criterion and is not a corollary of KAM at all. Every number this chapter quotes about the solar system comes from numerical integration [Wisdom:1983] [Laskar:1989], not from KAM, which is perturbative and silent at those perturbation strengths.
Four cautions, each of which is routinely overstated in summaries. First, KAM is perturbative: \(\varepsilon_{0}\) is small, and the theorem says nothing at the perturbation strengths of real celestial mechanics, where the question is settled numerically (Section 32.4.5). Second, the surviving tori form a nowhere dense set of large measure—their complement, the resonant gaps, is dense—so “most initial conditions are regular” and “regular initial conditions form a fat Cantor set riddled with chaotic layers” are both true at once. Third, the theorem is about tori, not about the motion in between: what happens in the gaps is the subject of Section 32.4.3. Fourth, and most importantly, the confinement KAM provides is dimension-dependent. In a system with two degrees of freedom the invariant tori are two-dimensional inside a three-dimensional energy surface, so they separate it and a chaotic orbit is trapped between neighbouring tori forever. With three or more degrees of freedom the tori no longer separate the energy surface, the chaotic layers connect, and slow transport along the resonance web—Arnold diffusion—becomes possible. Its existence is established for constructed examples; its rate in a given physical system is not settled by any theorem and must be measured or integrated. The primary reference for the diffusion example is not in this bibliography, and the statement is therefore made here without one.
Resonance overlap and the standard map
The local model of a perturbed resonance is a rotor given a periodic kick.
A rigid rotor of moment of inertia \(I\) (SI unit \(\mathrm{kg}\,\mathrm{m}^{2}\)), free except for an impulsive torque applied once every \(T\), has the Hamiltonian
with \(L\) the angular momentum (\(\mathrm{J}\,\mathrm{s}\)) and \(K'\) an energy (\(\mathrm{J}\)). Integrating Hamilton's equations across one kick and along the free flight between kicks gives the exact stroboscopic map
which in the dimensionless action \(l_{n}=TL_{n}/I\) becomes the standard map
The map Equation (32.25) is area-preserving—its Jacobian determinant is \(1\), as Proposition 32.40 requires—and \(K\) is the single dimensionless control parameter, built out of the measured kick energy, kick period and moment of inertia.
Each primary resonance of Equation (32.25) sits at \(l=2\pi m\), \(m\in\Z\), and has half-width \(\Delta l=2\sqrt{K}\) in the action. Neighbouring resonances therefore touch when \(4\sqrt{K}=2\pi\), that is at
above which no invariant curve can span the cylinder and the action diffuses without bound [Chirikov:1979]. Rests on Equations (32.19) and (32.25).
Derives Proposition 32.55. Near \(l=0\) and for small \(K\) the map is the stroboscopic image of the pendulum flow \(\dot{\vartheta}=l\), \(\dot{l}=K\sin\vartheta\), whose energy \(\tfrac{1}{2}l^{2}+K\cos\vartheta\) is conserved. Its separatrix passes through \((\vartheta,l)=(0,0)\) with energy \(K\), and at \(\vartheta=\pi\) reaches \(\tfrac{1}{2}l^{2}-K=K\), so the maximum excursion of the separatrix is \(l=2\sqrt{K}\): that is the half-width. The same holds at every \(l=2\pi m\) by the periodicity of the map in \(l\) modulo \(2\pi\). Two neighbouring separatrices, of half-width \(2\sqrt{K}\) each, span the gap \(2\pi\) between resonance centres when \(2\cdot2\sqrt{K}=2\pi\), which is Equation (32.26).
The criterion is a heuristic and is honest about it: it neglects the higher-harmonic resonances that fill the gaps and the deformation of the separatrices by the neighbouring resonance, both of which lower the threshold. Numerically the last invariant curve of Equation (32.25)—the one carrying the golden-mean winding number, the most irrational and therefore the last to break—disappears at \(K_{c}=0.9716\), so Equation (32.26) overestimates by a factor of about \(2.5\) [Chirikov:1979]. That is the accuracy the criterion claims for itself, and it is why the criterion is used for an order of magnitude and never for a number.
∎Just above \(K_{c}\) an invariant curve does not simply vanish: it becomes a cantorus, an invariant Cantor set with gaps, which no longer blocks transport but still obstructs it. Orbits leak through the gaps at a rate that vanishes as \(K\to K_{c}^{+}\), so the diffusion of the action is not the free diffusion of a random walk but is throttled by a hierarchy of partial barriers. This is the mechanism behind the very long timescales of Section 32.4.5: a chaotic orbit can behave regularly for many Lyapunov times before it finds a gap.
The physical realizations of Equation (32.25) are the kicked rotor itself, the periodically driven pendulum of Section 32.7.3, and—the case Chirikov was writing for—any resonance of a nearly integrable system, since Equation (32.22) is what Equation (32.19) reduces to near a single resonance.
Numerical discovery: Hénon–Heiles
A star of mass \(m\) moving in the meridional plane of an axisymmetric galactic potential, expanded to cubic order about the circular orbit, has
with \(\omega\) in \(/\mathrm{s}\) and \(\lambda\) in \(/\mathrm{m}/\mathrm{s}^{2}\) [Henon:1964]. Rescaling by \(x=(\omega^{2}/\lambda)\xi\), \(y=(\omega^{2}/\lambda)\eta\) and \(t=s/\omega\) reduces it to
measured in units of the energy \(E_{\ast}=m\omega^{6}/\lambda^{2}\). Rests on Equation (22.1) and Theorem 7.38.
A slightly perturbed integrable Hamiltonian system does not become wholly chaotic. Surfaces of section computed for the Hénon–Heiles Hamiltonian show, at low energy, closed invariant curves filling the accessible region — the traces of surviving invariant tori — and, as the energy is raised, islands of such curves embedded in a growing sea of points that wander over a two-dimensional area [Henon:1964]. Both kinds of orbit occur at the same energy in the same system, and the chaotic fraction grows continuously with the perturbation: there is no energy at which order gives way to disorder all at once [Kolmogorov:1954] [Arnold:1963] [Moser:1962]. The numerical experiment also settled the astronomical question it was posed for — whether stellar motion in a galactic disc admits a third isolating integral — in the negative for general orbits, while showing that many individual orbits behave as though it existed [Henon:1964]. Rests on Equation (32.27) and Theorem 32.51.
Derivation. Derives Phenomenon 32.58. Three things have to be established: that the system has a finite range of bounded motion in which the question makes sense, that a closed curve on the surface of section is the trace of an invariant torus, and that a scattered area is not.
The escape energy. The cubic term in Equation (32.28) eventually overwhelms the quadratic one, so the potential \(\tilde{V}=\tfrac{1}{2}(\xi^{2}+\eta^{2}) +\xi^{2}\eta-\eta^{3}/3\) is bounded only inside a finite region. Setting \(\vect{\nabla}\tilde{V}=0\) gives \(\xi(1+2\eta)=0\) and \(\eta+\xi^{2}-\eta^{2}=0\). From \(\xi=0\) the second equation gives \(\eta=0\) (the minimum, \(\tilde{V}=0\)) or \(\eta=1\), at which \(\tilde{V}=\tfrac{1}{2}-\tfrac{1}{3}=\tfrac{1}{6}\). From \(\eta=-\tfrac{1}{2}\) it gives \(\xi^{2}=\tfrac{3}{4}\), and there \(\tilde{V}=\tfrac{1}{2}\left(\tfrac{3}{4}+\tfrac{1}{4}\right) -\tfrac{3}{8}+\tfrac{1}{24}=\tfrac{1}{6}\) as well. There are therefore three saddle points at the same height, at the corners of an equilateral triangle, and
Below this the motion is bounded and the question of integrability is well posed; above it, orbits escape through the three channels.
Reading the section. Take the section \(\xi=0\) with \(\pi_{\xi}>0\), coordinatized by \((\eta,\pi_{\eta})\); by Proposition 32.40 the return map preserves area. If the orbit lies on a two-torus—which is what a third isolating integral would guarantee, by Theorem 23.31—then the intersection of that two-torus with the two-dimensional section inside the three-dimensional energy surface is a one-dimensional curve, and the successive points must fall on it. Conversely, points that fill a two-dimensional patch cannot lie on any smooth one-dimensional invariant set, so no third integral exists on that orbit. This is what makes the surface of section a decisive diagnostic rather than a picture: the dimension of the point set is the answer.
The measured outcome. At \(E=\tfrac{1}{12}E_{\ast}\) the computed section is covered by closed curves and the motion looks integrable; at \(E=\tfrac{1}{8}E_{\ast}\) islands of closed curves survive inside a connected region of scattered points; at \(E=\tfrac{1}{6}E_{\ast}\) the scattered region occupies almost the whole accessible area, with a few small islands remaining [Henon:1964]. The relative area of the scattered region rises continuously from zero, which is the statement of the phenomenon, and Theorem 32.51 is the reason it must: the tori that go first are the resonant ones, and the Diophantine ones of Equation (32.21) survive to larger perturbations, so the transition is graded and not sudden. The astronomical conclusion follows directly: no third isolating integral, analytic in the coordinates, exists for general orbits in such a potential, while the surviving tori explain why so many individual stellar orbits behave as though one did.
∎Chaos in the solar system
The asteroid belt is not populated uniformly: it carries deep gaps at the orbital periods commensurate with Jupiter's — the Kirkwood gaps — and the gap at the 3/1 commensurability coincides with a zone in which the eccentricity of a test asteroid does not oscillate but wanders, on a timescale short against the age of the system, until the orbit crosses those of the inner planets and the body is removed [Wisdom:1983]. The planetary orbits themselves are chaotic in the same sense: numerical integration of the averaged equations gives a positive Lyapunov exponent for the inner solar system, with a Lyapunov time of order \(5\times 10^{6}\,\mathrm{yr}\), so that an error of \(15\,\mathrm{m}\) in an initial position grows within about \(10^{8}\,\mathrm{yr}\) to a distance of the order of the Earth–Sun separation, \(1.5\times 10^{11}\,\mathrm{m}\) [Laskar:1989]. The observed regularity of the planets over recorded history is therefore no evidence at all of predictability over the age of the system. Rests on Equation (32.46) and Theorem 32.51.
Derivation. Derives Phenomenon 32.59. Where the gaps are. Kepler's third law Equation (27.51) with the Sun's mass gives, for a heliocentric orbit, \(a=\left(T/1\,\mathrm{yr}\right)^{2/3}\mathrm{au}\). An asteroid in the \(p{:}q\) mean-motion resonance with Jupiter has \(T=(q/p)\,T_{J}\) with \(T_{J}=11.862\,\mathrm{yr}\), so its semi-major axis is fixed by arithmetic alone: Table 32.2 lists the four principal cases against the observed gap positions. The agreement is at the level of the approximation—the derivation used the Sun's mass alone and a circular Jovian orbit—and it identifies the gaps as a resonance phenomenon without yet explaining why a resonance should empty rather than fill.
Why the 3/1 zone empties. At the 3/1 resonance the numerical integration of [Wisdom:1983] shows the eccentricity of a test body not librating but jumping irregularly, in sudden bursts that carry it above \(e=0.3\) on a timescale of order \(10^{6}\,\mathrm{yr}\). That range is the one that matters, because Mars has \(a_{M}=1.5237\,\mathrm{au}\) and \(e_{M}=0.0934\), hence aphelion \(Q_{M}=a_{M}(1+e_{M})=1.666\,\mathrm{au}\); an asteroid at \(a=2.50\,\mathrm{au}\) reaches Mars's orbit when its perihelion \(a(1-e)\) falls to \(Q_{M}\), that is when
A body that reaches such an eccentricity is removed by a close encounter within an interval negligible against the age of the system, and the gap is the record of that removal. The chaotic zone and the observed gap coincide in width, which is the quantitative content of [Wisdom:1983].
The planets. For the planets themselves the exponent is measured, not derived: [Laskar:1989] integrates the averaged secular equations for \(2\times 10^{8}\,\mathrm{yr}\) and finds neighbouring solutions separating exponentially with a Lyapunov time \(\lambda^{-1}\approx5\times 10^{6}\,\mathrm{yr}\). Feeding that into the horizon formula Equation (32.46), an initial uncertainty \(\delta_{0}=15\,\mathrm{m}\) reaches \(\Delta=1.5\times 10^{11}\,\mathrm{m}\) after
which is the growth quoted in [Laskar:1989] to the one significant figure to which the Lyapunov time itself is known. Two conclusions, and only two, follow. Positive \(\lambda\) destroys prediction of the phase of the orbits beyond about \(10^{8}\,\mathrm{yr}\); it does not by itself say that any planet is lost, which is a question about the size of the chaotic zone and not about the exponent, and which the same integrations address separately. The regularity of the planets over the few thousand years of recorded observation is \(10^{-3}\) Lyapunov times and constrains nothing.
∎| Resonance | $T$ (\(\mathrm{yr}\)) | $a$ computed (\(\mathrm{au}\)) | $a$ observed (\(\mathrm{au}\)) |
|---|---|---|---|
| 4:1 | \(2.966\) | \(2.06\) | \(2.06\) |
| 3:1 | \(3.954\) | \(2.50\) | \(2.50\) |
| 5:2 | \(4.745\) | \(2.82\) | \(2.82\) |
| 7:3 | \(5.084\) | \(2.96\) | \(2.96\) |
| 2:1 | \(5.931\) | \(3.28\) | \(3.27\) |
The gap positions in Table 32.2 are arithmetic. What a resonance does to a body sitting in one is dynamics, and the local model is the pendulum of Proposition 32.55 again.
Let a test body of mass \(m_{t}\), negligible beside both other masses, orbit a star of mass \(M\) with semi-major axis \(a\), eccentricity \(e\) and mean motion \(n=\sqrt{GM/a^{3}}\), perturbed by a planet of mass \(m\) on a circular coplanar orbit of mean motion \(n_{J}\), and let it lie near the \(p{:}q\) commensurability \(q\,n=p\,n_{J}\). Write the single resonant harmonic of the perturbing Hamiltonian as
with \(S\) dimensionless and
the resonant angle built from the two mean longitudes and the longitude of pericentre \(\varpi\). Then, while \(\varpi\) may be treated as fixed, \(\sigma\) obeys a pendulum equation whose libration angular frequency is
and whose separatrix has the half-width
in semi-major axis, independent of \(p\) and \(q\) except through \(S\). Furthermore \(S\) vanishes with the eccentricity as \(S=s\,e^{\abs{p-q}}\), with \(s\) a function of \(a/a_{J}\) alone and of order unity. Rests on Proposition 32.55 and Equation (32.19).
Derives Proposition 32.60. The pendulum. Take as the action conjugate to the mean longitude \(\lambda\) the quantity \(\Lambda=m_{t}\sqrt{GMa}\), for which the Kepler Hamiltonian is \(\Ham_{0}=-G^{2}M^{2}m_{t}^{3}/(2\Lambda^{2})\) and
the second by differentiating the first. Hamilton's equation for \(\Lambda\) picks out the \(\lambda\) dependence of Equation (32.31), which enters only through \(\sigma\) and only as \(-q\lambda\):
Differentiating Equation (32.32) twice with \(\varpi\) fixed and \(n_{J}\) constant, \(\dot{\sigma}=p\,n_{J}-q\,n\) and
a pendulum with its stable equilibrium at \(\sigma=\pi\). Its frequency is Equation (32.33), since \(\Lambda=m_{t}na^{2}\) and \(\varepsilon=m_{t}n^{2}a^{2}\mu S\) give \(3q^{2}n\varepsilon/\Lambda=3q^{2}n^{2}\mu S\).
The width. Equation (32.35) conserves \(\tfrac{1}{2}\dot{\sigma}^{2}+\omega_{\text{lib}}^{2}\cos\sigma\). The separatrix passes through the unstable point \(\sigma=0\) with \(\dot{\sigma}=0\), so its constant is \(\omega_{\text{lib}}^{2}\), and at \(\sigma=\pi\) it reaches \(\tfrac{1}{2}\dot{\sigma}^{2}=2\omega_{\text{lib}}^{2}\): the maximum excursion is \(\abs{\dot{\sigma}}=2\omega_{\text{lib}}\). This is the same computation as in Proposition 32.55, on a different pair of variables. Converting to the semi-major axis, \(\dot{\sigma}=-q\,(\dd n/\dd\Lambda)\,\Delta\Lambda =3qn\,\Delta\Lambda/\Lambda\), so
and \(\Lambda\propto\sqrt{a}\) makes \(\Delta a/a=2\Delta\Lambda/\Lambda\), which is Equation (32.34). The factor \(q\) has cancelled.
Why \(S\) carries \(e^{\abs{p-q}}\). The perturbing Hamiltonian depends on the two orbits only through their relative geometry, so it is unchanged by rotating the whole system: \(\lambda\to\lambda+\phi\), \(\lambda_{J}\to\lambda_{J}+\phi\), \(\varpi\to\varpi+\phi\). The combination Equation (32.32) is invariant under exactly that substitution—its three coefficients sum to \(p-q+q-p=0\)—which is what selects the admissible harmonics. Now \(\varpi\) is undefined at \(e=0\), and the perturbation is a smooth function of the position of the test body, hence of the two elements \(e\cos\varpi\) and \(e\sin\varpi\), which are smooth where \(\varpi\) is not. A term carrying \(\ee^{\ii j\varpi}\) can be assembled from those two only as a polynomial of degree at least \(\abs{j}\) in them, so it vanishes at least as \(e^{\abs{j}}\). Here \(j=q-p\), giving \(S=s\,e^{\abs{p-q}}\) with the remaining factor a function of the semi-major-axis ratio alone.
∎Put the numbers in for the 3/1 gap, where \(p=3\), \(q=1\) and the resonance is therefore of second order, \(S=s\,e^{2}\). With \(\mu=m_{J}/M_{\odot}=9.5458\times 10^{-4}\), Equation (32.34) gives
so a body of the typical belt eccentricity \(e\approx0.1\) librates over \(\Delta a\approx0.018\,\mathrm{au}\) about \(a=2.50\,\mathrm{au}\), and the libration period Equation (32.33) is \(2\pi/\omega_{\text{lib}}=74\,\mathrm{yr}/(\sqrt{s}\,e)\), of order \(7\times 10^{2}\,\mathrm{yr}\) there. Two features of those formulas do the work. The width grows with the eccentricity, so a body whose \(e\) is being pumped finds the resonance widening around it rather than escaping from it; and the libration frequency vanishes with \(e\), while the secular precession rate of \(\varpi\)—the very quantity the derivation above froze—is of order \(\mu n\) and does not vanish with \(e\). Their ratio is
which falls below unity for \(e\lesssim0.02\). Below that eccentricity the single-resonance averaging that produced the pendulum is not valid: the resonant and the secular degrees of freedom run at comparable rates, and two comparable resonances that overlap are precisely the configuration Proposition 32.55 identifies as chaotic. The low-eccentricity part of the 3/1 zone is therefore where the pendulum picture must fail, which is the qualitative statement that a body starting on a nearly circular orbit there does not stay on one.
What is not derived here is the rest of [Wisdom:1983]: that the eccentricity actually wanders as far as the Mars-crossing value Equation (32.30) rather than to some smaller ceiling, and that it does so in bursts on a timescale of order \(10^{6}\,\mathrm{yr}\). Those are numerical results, obtained by integrating an algebraic mapping of the resonance over millions of orbits, and this chapter quotes them. The timescale is worth one comment, because it is easy to misread: it is not the Lyapunov time. The only timescale the local model supplies is the libration period, a few hundred years, and a resonance-overlap chaotic layer has a Lyapunov time of that order; the \(10^{6}\,\mathrm{yr}\) on which the eccentricity jumps is three orders of magnitude longer, and the gap between them is the throttled transport of Remark 32.56. A chaotic orbit can look regular for a thousand Lyapunov times before it finds a way through.
Strange attractors
The Lorenz system
Truncating the Boussinesq equations for a fluid layer of depth \(d\) heated from below to one convective mode and two thermal modes gives the three ordinary differential equations
in which the dot is \(\dd/\dd s\) with \(s=\pi^{2}(1+a^{2})\kappa t/d^{2}\) a dimensionless time built from the thermal diffusivity \(\kappa\) (\(\mathrm{m}^{2}/\mathrm{s}\)), \(x\) is proportional to the amplitude of the convective roll, \(y\) and \(z\) to the horizontal and vertical temperature variations, and the three parameters are
with \(\nu\) the kinematic viscosity, \(a\) the dimensionless horizontal wavenumber of the roll, and
the Rayleigh number, built from the acceleration of gravity \(g\), the thermal expansion coefficient \(\alpha_{T}\) (\(/\mathrm{K}\)) and the temperature difference \(\Delta T\) across the layer [Lorenz:1963]. For stress-free boundaries the critical value is \(\mathrm{Ra}_{c}=27\pi^{4}/4=657.5\), attained at \(a^{2}=\tfrac{1}{2}\), whence Lorenz's \(b=8/3\); his other values are \(\sigma=10\) and \(r=28\). Rests on Theorem 32.6 and Definition 32.7.
Full derivation in Appendix A.
Derives Definition 32.62.
The Lorenz System from Boussinesq Convection carries the reduction in full: the Boussinesq equations for a layer of depth \(d\) heated from below, the stream-function form that eliminates the pressure, the linear stability analysis that produces the critical Rayleigh number \(\mathrm{Ra}_{c}=27\pi^{4}/4\) at \(a^{2}=\tfrac{1}{2}\) for stress-free boundaries [Rayleigh:1916], the Galerkin truncation to one velocity mode and two temperature modes, and the scaling that turns what is left into Equation (32.36) with the three parameters Equation (32.37). It is a page of vector calculus and several pages of algebra, which is why it is not reproduced here, and it is what makes each of \(\sigma\), \(r\) and \(b\) a ratio of measured fluid properties rather than a dial: \(\sigma\) is the Prandtl number of the working fluid, \(r\) the Rayleigh number in units of its own critical value, and \(b\) a pure function of the horizontal wavenumber of the roll. The physical content that follows the truncation—the dissipation rate, the fixed points, the onset of the strange attractor—is derived in this chapter and does not depend on the appendix.
Equation (32.36) retains three of the infinitely many modes of a convecting layer, and Lorenz said plainly that the truncation is not quantitatively faithful to convection at \(r=28\), which is far above the range in which a three-mode truncation can be trusted [Lorenz:1963]. What the system is faithful to is the mechanism: a dissipative flow, bounded, with an unstable fixed point and no stable alternative nearby. Its evidential standing rests on Section 32.7.1, where a real convection cell reproduces the universal numbers, and not on any claim that Equation (32.36) describes that cell.
The divergence of Equation (32.36) is the negative constant \(-(\sigma+1+b)\), so every phase volume obeys
which at Lorenz's parameters contracts by a factor \(\ee\) in every \(3/41=0.0732\) of dimensionless time. Rests on Theorem 32.6 and Equation (32.36).
Derives Proposition 32.64. Differentiating each component of Equation (32.36) with respect to its own variable gives \(-\sigma\), \(-1\) and \(-b\), whose sum is the constant \(-(\sigma+1+b)\). Equation (32.3) then reads \(\dot{V}=-(\sigma+1+b)V\), which integrates to Equation (32.39). At \(\sigma=10\) and \(b=8/3\) the rate is \(41/3\) and its reciprocal is \(3/41\).
∎The origin is a fixed point for every \(r\); it is asymptotically stable for \(r<1\) and a saddle for \(r>1\), the transition at \(r=1\) being a supercritical pitchfork that creates the pair
the two steady convecting rolls. For \(\sigma>b+1\) the pair loses stability at
which is \(470/19=24.7368\) at \(\sigma=10\), \(b=8/3\). Rests on Equation (32.36) and Theorem 32.13.
Derives Proposition 32.65. Setting the right-hand sides of Equation (32.36) to zero gives \(y=x\) from the first, then \(x(r-1-z)=0\) from the second and \(x^{2}=bz\) from the third. Either \(x=y=z=0\), or \(z=r-1\) and \(x^{2}=b(r-1)\), which is real only for \(r\ge1\): that is Equation (32.40), and the \(\sqrt{r-1}\) growth of the amplitude identifies the pitchfork of Proposition 32.24(ii), the symmetry being the \((x,y)\to(-x,-y)\) invariance of Equation (32.36).
At the origin the Jacobian is block diagonal, with the entry \(-b\) and the \(2\times2\) block
The block has \(\tau=-(\sigma+1)<0\) and \(\Delta=\sigma(1-r)\), so by Proposition 32.17 it is a stable node for \(r<1\) and a saddle for \(r>1\).
At \(C_{\pm}\), writing \(x_{0}=\pm\sqrt{b(r-1)}\), the Jacobian is
whose trace is \(-(\sigma+1+b)\), whose determinant is \(-2\sigma x_{0}^{2}=-2\sigma b(r-1)\), and whose sum of principal \(2\times2\) minors is \(0+\sigma b+(b+x_{0}^{2})=b(\sigma+r)\). Hence
A Hopf bifurcation needs a purely imaginary pair. Substituting \(\lambda=\ii\omega\) in Equation (32.42) and separating real and imaginary parts gives
and eliminating \(\omega^{2}\) leaves \((\sigma+1+b)(\sigma+r)=2\sigma(r-1)\). Expanding, \(\sigma^{2}+b\sigma+3\sigma+r(1+b-\sigma)=0\), which solves to Equation (32.41); it is positive only for \(\sigma>b+1\). With \(\sigma=10\) and \(b=8/3\), \(r_{H}=10\cdot\tfrac{47}{3}\big/\tfrac{19}{3}=470/19\).
∎Proposition 32.65 places \(r=28\) above \(r_{H}=24.74\), so all three fixed points are unstable there. The Hopf bifurcation at \(r_{H}\) is subcritical: the periodic orbits it involves exist below \(r_{H}\) and are unstable, so they are swallowed by \(C_{\pm}\) rather than emitted by them, in the manner of Proposition 32.24(iii). Above \(r_{H}\) there is therefore no local attractor of any kind, while Proposition 32.64 forces all volumes to zero and a trapping argument keeps trajectories bounded. Something of zero volume and no local structure must attract them: that is the strange attractor of Phenomenon 32.68, and the argument just given is the whole reason to expect one.
Strange attractors and the Ruelle–Takens proposal
An attractor (Definition 32.8) is strange if the dynamics on it has sensitive dependence on initial conditions, that is a positive largest Lyapunov exponent (Definition 32.73). It is a fractal attractor if in addition its dimension (Definition 32.79) is not an integer. The two properties are logically independent, and the term coined by Ruelle and Takens [Ruelle:1971] covers the sets—the common case—that have both. Rests on Definition 32.8 and Phenomenon 32.72.
A dissipative chaotic system settles onto an attracting set of zero phase-space volume which is nevertheless neither a point, nor a closed curve, nor a smooth surface: it is a set of non-integer dimension, measured from experimental and numerical time series by counting the pairs of points lying closer than a given separation [Grassberger:1983]. The Lorenz system is the standard case: its trajectories are drawn onto a bounded set of two sheets on which they never repeat, and the correlation dimension of that set is measured to be a little above two [Grassberger:1983]. Sets of this kind are what Ruelle and Takens named strange attractors, and their geometry is that of the fractals [Ruelle:1971] [Mandelbrot:1982]. Rests on Definition 32.67 and Proposition 32.64.
Derivation. Derives Phenomenon 32.68. The zero-volume half of the statement is Proposition 32.64: every volume decays like \(\ee^{-(\sigma+1+b)s}\), so the attractor can contain no three-dimensional region whatever.
It is equally not a point and not a closed curve. By Remark 32.66 no fixed point and no local periodic orbit is stable at \(r=28\); and on the attractor neighbouring trajectories separate exponentially, by Equation (32.44), while the motion remains bounded, so the set is stretched in one direction and folded back on itself without end. Nor can it be a smooth surface: a two-dimensional invariant surface in a three-dimensional flow would confine trajectories to a plane-like region in which Corollary 32.32 forbids the observed sensitivity.
What remains is fixed quantitatively by the exponents. The Lorenz attractor at \(\sigma=10\), \(b=8/3\), \(r=28\) has computed Lyapunov spectrum \(\lambda_{1}=0.906\), \(\lambda_{2}=0\), \(\lambda_{3}=-14.572\) per unit dimensionless time, whose sum is \(-13.667=-(\sigma+1+b)\) exactly as Proposition 32.76 requires—a check on the numbers that costs nothing and is worth making. The Kaplan–Yorke estimate \(D=2+\lambda_{1}/\abs{\lambda_{3}}=2.06\) then predicts a dimension just above two, and the correlation dimension measured from the trajectory by pair counting is \(2.05 \pm 0.01\) [Grassberger:1983]. The measured value is what settles the statement; the derivation establishes only that it must lie strictly between two and three.
∎Landau's account of the onset of turbulence had each instability add one incommensurate frequency, so that the flow became turbulent by accumulating an unbounded number of them. Ruelle and Takens proposed instead that after a small number of bifurcations the motion settles on an attractor which is strange in the sense of Definition 32.67, with a continuous power spectrum and sensitive dependence [Ruelle:1971]. The two pictures differ in a directly measurable quantity—the shape of the power spectrum—and that is what the Taylor–Couette measurement of Section 32.7.2 was able to decide.
The tractable prototype is the two-dimensional map
with \(\hat{a}=1.4\) and \(\hat{b}=0.3\) [Henon:1976]. Its Jacobian matrix is \(\begin{pmatrix}-2\hat{a}x&1\\\hat{b}&0\end{pmatrix}\), whose determinant is the constant \(-\hat{b}\): the map contracts areas by the exact factor \(0.3\) at every step, whatever the point, which makes it the cleanest available caricature of a dissipative flow. Its largest Lyapunov exponent is \(\lambda_{1}=0.419\) per iterate, so \(\lambda_{2}=\ln\hat{b}-\lambda_{1}=-1.623\) and the Kaplan–Yorke estimate is \(1+\lambda_{1}/\abs{\lambda_{2}}=1.26\); the measured correlation dimension is \(1.21 \pm 0.01\) [Grassberger:1983]. The two differ, as they generally do—the correlation dimension is a lower bound on the information dimension that the Kaplan–Yorke formula estimates—and quoting either as “the” dimension of the attractor without saying which is measured would misreport both. Rests on Definitions 32.38 and 32.67.
The horseshoe of Section 32.3.2 is uniformly hyperbolic: at every point of its invariant set the tangent space splits into a uniformly contracting and a uniformly expanding direction, and that is what makes its symbolic description exact. Real attractors are not uniformly hyperbolic. The Lorenz attractor contains a fixed point, at which the expansion rate differs from that on the rest of the set, and the Hénon attractor has tangencies between stable and unstable manifolds. That the Lorenz equations possess an attractor at all—rather than a very long transient, or a periodic orbit of enormous period—was for four decades a conjecture, settled eventually by a computer-assisted proof whose primary reference is not in this bibliography and which is therefore not cited here. The practical consequence for a reader is worth stating plainly: the numbers quoted in this section are measurements on trajectories, and their status is that of measurements.
Sensitive dependence and Lyapunov exponents
In a chaotic system two initial states differing by an unmeasurably small amount separate at an exponential rate: on the attractor the separation of nearby trajectories grows on average as
with \(\lambda\) — the largest Lyapunov exponent — a property of the attractor and not of the particular trajectory followed [Oseledets:1968]. The effect was established in a numerical model of thermal convection, in which two solutions started from imperceptibly different data were followed until they bore no resemblance to one another [Lorenz:1963], and a positive \(\lambda\) is now extracted routinely from a single measured scalar time series [Wolf:1985]. Determinism is untouched by this; what is destroyed is prediction, on the finite horizon of Equation (32.46). Rests on Equation (32.4) and Theorem 32.74.
Derivation. Derives Phenomenon 32.72. Linearize the flow \(\dot{\vect{x}}=\vect{f}(\vect{x})\) about a reference trajectory: an infinitesimal separation obeys the variational equation
whose solution is \(\delta\vect{x}(t)=M(t)\,\delta\vect{x}(0)\) for a fundamental matrix \(M\) depending on the reference trajectory alone. For a flow with an invariant measure the multiplicative ergodic theorem (Theorem 32.74) guarantees that the limit
exists for almost every initial point and takes one of finitely many values, the largest of which is attained for every \(\delta\vect{x}(0)\) outside a set of lower dimension [Oseledets:1968]. That is Equation (32.44), and it is why \(\lambda\) can be quoted as a number for the system rather than for the run.
The consequence is quantitative. Suppose the initial state is known to within \(\delta_{0}\) and the prediction is useful while the error stays below a tolerance \(\Delta\). Solving \(\delta_{0}\ee^{\lambda T}=\Delta\) gives the horizon
It depends on the measurement precision only through its logarithm: improving the initial data by a factor of ten buys the fixed extra time \(\lambda^{-1}\ln 10\) and no more, however great the effort. A system with \(\lambda>0\) is therefore not one whose equations are unknown — Equation (32.45) used them — but one in which prediction is bought at an exponentially rising price.
∎The Lyapunov exponents of a trajectory are the \(n\) numbers \(\lambda_{1}\ge\lambda_{2}\ge\dots\ge\lambda_{n}\) giving the exponential growth rates of the principal axes of an infinitesimal ball carried by Equation (32.45); equivalently \(\lambda_{i}=\lim_{t\to\infty}t^{-1}\ln\mu_{i}(t)\) with \(\mu_{i}\) the singular values of the fundamental matrix \(M(t)\). They carry SI units of \(/\mathrm{s}\) for a flow and are dimensionless per iterate for a map. A bounded, aperiodic trajectory with \(\lambda_{1}>0\) is called chaotic. Rests on Equation (32.45).
Let \(\varphi_{t}\) preserve a probability measure supported on a compact invariant set, and let \(\vect{f}\) be continuously differentiable. Then for almost every point with respect to that measure the limits in Definition 32.73 exist, are the same for almost all points of an ergodic component, and the tangent space at each point splits into subspaces on which the growth rate is exactly \(\lambda_{i}\) [Oseledets:1968]. Rests on Definitions 32.8 and 32.73.
The multiplicative ergodic theorem is a theorem of ergodic theory, resting on measure theory, on Birkhoff's pointwise ergodic theorem and on the subadditive ergodic theorem. Part II contains none of these: Probability and Statistics is built on discrete and Riemann-integrable probability and says so. The theorem is quoted here because it is what converts “two runs diverged” into “the system has an exponent”, and every numerical exponent quoted in this chapter relies on it for its meaning. The two propositions that follow are independent of it and are proved.
For a flow with an ergodic invariant measure on its attractor,
the average being over the attractor. In particular a Hamiltonian flow has \(\sum_{i}\lambda_{i}=0\), and the Lorenz flow has \(\sum_{i}\lambda_{i}=-(\sigma+1+b)\). Rests on Theorem 32.6 and Definition 32.73.
Derives Proposition 32.76. An infinitesimal ball of radius \(\epsilon\) is carried by Equation (32.45) into an ellipsoid whose principal axes are \(\epsilon\mu_{i}(t)\), so its volume is \(V(t)=V(0)\prod_{i}\mu_{i}(t)\) and
On the other hand Equation (32.3) applied to an infinitesimal volume gives \(\dot{V}/V=\vect{\nabla}\cdot\vect{f}\) evaluated on the trajectory, so \(t^{-1}\ln[V(t)/V(0)]\) is the time average of \(\vect{\nabla}\cdot\vect{f}\) along it, which for an ergodic measure tends to its space average. Equating the two limits gives Equation (32.47). For a Hamiltonian flow the divergence vanishes identically (Theorem 24.21); for Equation (32.36) it is the constant of Proposition 32.64.
∎Let a trajectory of Equation (32.1) remain in a compact invariant set containing no fixed point. Then it has at least one vanishing Lyapunov exponent, associated with displacement along the flow. Rests on Equation (32.45) and Definition 32.73.
Derives Proposition 32.77. Differentiate \(\dot{\vect{x}}=\vect{f}(\vect{x})\) with respect to \(t\): \(\ddot{\vect{x}}=D\vect{f}(\vect{x})\dot{\vect{x}}\), so \(\delta\vect{x}(t)=\vect{f}(\vect{x}(t))\) is itself a solution of the variational equation Equation (32.45). On a compact attractor containing no fixed point, \(\abs{\vect{f}}\) is bounded above and below by positive constants (Theorem 7.24), so \(t^{-1}\ln\abs{\delta\vect{x}(t)}\to0\): the exponent along the flow is zero. The hypothesis excludes a fixed point because \(\abs{\vect{f}}\) must be bounded away from zero; the Lorenz attractor does contain one, at the origin, and the conclusion survives there only because a trajectory spends a vanishing fraction of its time near it, which is a statement about the invariant measure and not about the flow alone. This is why a three-dimensional dissipative flow has the spectrum \((+,0,-)\) with the sum negative—the smallest arrangement that admits chaos—and why a map, which has no such direction, can be chaotic in one dimension.
∎An experiment does not deliver \(D\vect{f}\); it delivers a scalar time series. The standard procedure embeds the series in a higher-dimensional space by delay coordinates, follows a pair of nearby reconstructed points, measures the growth of their separation over a short interval, renormalizes the separation before it leaves the linear regime, and averages the logarithmic growth rates [Wolf:1985]. The output is \(\lambda_{1}\) with an uncertainty dominated by the length of the series and the noise level, and it is on this chain that every experimental claim of a positive exponent in Section 32.7 rests. The embedding step itself—that a scalar observable of a \(d\)-dimensional attractor generically reconstructs it in \(2d+1\) delay coordinates—is a theorem of Takens whose primary reference is not in this bibliography; it is used here as a method and not cited as a result.
Fractal dimension
Let \(N(\epsilon)\) be the least number of boxes of side \(\epsilon\) needed to cover a bounded set \(S\). Its box-counting dimension is
when the limit exists. For a point, a smooth curve and a smooth surface it returns \(0\), \(1\) and \(2\). This is Definition 7.142 specialized twice over: the limit is written as though it exists, where that definition keeps the upper and lower limits apart and calls them \(\overline{\dim}_{\mathrm{B}}\) and \(\underline{\dim}_{\mathrm{B}}\); and the covering family is boxes of a fixed side rather than arbitrary sets of diameter at most \(\delta\). The two covering conventions give the same value whenever either limit exists, since a box of side \(\epsilon\) has diameter \(\epsilon\sqrt{3}\) and a set of diameter \(\delta\) fits in a box of side \(\delta\), so the two counts differ by a factor that is constant in \(\epsilon\) and drops out of the logarithmic ratio. Proposition 7.143 records that the Hausdorff dimension of a set never exceeds the lower box dimension, so a box count bounds the finer quantity from above and is not a substitute for it. Rests on Definitions 6.5 and 7.142.
Remove the open middle third of \([0,1]\), then the middle thirds of the two remaining intervals, and so on. At stage \(k\) the set is covered by \(N=2^{k}\) intervals of length \(\epsilon=3^{-k}\), so Equation (32.48) gives
a dimension strictly between that of a point and that of a line. The horseshoe's invariant set Equation (32.17) is the product of two such sets, with \(2\) and \(3\) replaced by the map's stretching and contraction factors. Rests on Definition 32.79 and Theorem 32.42.
For \(N\) points \(\vect{x}_{i}\) sampled from a trajectory, let
be the fraction of pairs closer than \(\epsilon\), with \(\Theta\) the step function. Where \(C(\epsilon)\propto\epsilon^{D_{2}}\) over a range of \(\epsilon\), the exponent \(D_{2}\) is the correlation dimension [Grassberger:1983]. Rests on Definition 32.79.
Equation (32.49) is what makes fractal dimension an experimental quantity rather than a mathematical one: it needs only a list of points, it is computed by counting pairs, and it applies without change to a numerical trajectory and to an attractor reconstructed from a measured scalar series by the embedding of Remark 32.78. The measured values are collected in Table 32.3.
| Attractor | $D_{2}$ measured [Grassberger:1983] | Kaplan–Yorke estimate |
|---|---|---|
| Lorenz, $\sigma=10$, $b=8/3$, $r=28$ | \(2.05 \pm 0.01\) | \(2.06\) |
| Hénon, $\hat{a}=1.4$, $\hat{b}=0.3$ | \(1.21 \pm 0.01\) | \(1.26\) |
The Kaplan–Yorke estimate used in Table 32.3 is
with \(j\) the largest index for which the partial sum is still non-negative: the dimension at which the volume growth rate changes sign. It is a conjecture, not a theorem, and its primary reference is not in this bibliography; it is quoted here as an estimate and never as a measured value.
More fundamentally, the canonical definition of a fractal dimension is Hausdorff's, built on an outer measure. Section 7.12 constructs it (Definition 7.140) without any Lebesgue theory, and Proposition 7.143 relates it to the box count: the Hausdorff dimension never exceeds the lower box-counting dimension, and the gap can be total — the rationals of \([0,1]\) have Hausdorff dimension \(0\) and box dimension \(1\). This chapter works with Equation (32.48) and Equation (32.49) because they are what a finite data set can actually return; what they return is therefore an upper bound on the Hausdorff dimension of the attractor, never a determination of it. The general fractal vocabulary is Mandelbrot's [Mandelbrot:1982]. Together with the Lyapunov spectrum, the dimension is the second quantitative fingerprint of a chaotic attractor, and the two are the numbers that an experiment actually returns.
Universal routes to chaos
The logistic map
A population \(N_{n}\) in generation \(n\), growing at fractional rate \(\rho\) per generation and limited by a carrying capacity \(K\), obeys the difference equation \(N_{n+1}=N_{n}\left[1+\rho\left(1-N_{n}/K\right)\right]\). Setting \(x_{n}=\rho N_{n}/\left[(1+\rho)K\right]\) and \(r=1+\rho\) turns it into the logistic map
which maps \([0,1]\) into itself for \(0<r\le4\). Both \(x\) and \(r\) are dimensionless, being a population in units of a population and a rate per generation [May:1976]. Rests on Definition 32.3.
Equation (32.50) has fixed points \(x^{*}=0\), with multiplier \(r\), and \(x^{*}=1-1/r\), with multiplier \(2-r\). The origin is stable for \(r<1\); the second fixed point exists and is stable for \(1<r<3\); at \(r_{1}=3\) its multiplier passes through \(-1\) and a stable two-cycle is born, which itself loses stability at
Rests on Equation (32.50) and Proposition 32.39.
Derives Proposition 32.84. Solving \(rx(1-x)=x\) gives \(x^{*}=0\) and \(x^{*}=1-1/r\). Since \(f_{r}'(x)=r(1-2x)\), the multipliers are \(f_{r}'(0)=r\) and \(f_{r}'(1-1/r)=r\left(1-2+2/r\right)=2-r\); Proposition 32.39 makes each fixed point stable exactly while its multiplier lies in \((-1,1)\), which gives \(r<1\) and \(1<r<3\). At \(r=1\) the multiplier is \(+1\) and the two branches exchange stability: the transcritical bifurcation of Proposition 32.24(i). At \(r=3\) it is \(-1\), which is a period doubling.
For the two-cycle \(\set{p,q}\), add and multiply the two conditions \(q=rp(1-p)\) and \(p=rq(1-q)\). Subtracting gives \(p+q=(r+1)/r\), and substituting back gives \(pq=(r+1)/r^{2}\). The multiplier of the cycle is the product of the derivatives,
using the two symmetric functions just found. Setting this to \(+1\) returns \(r=3\), the birth of the cycle; setting it to \(-1\) gives \(r^{2}-2r-5=0\), whose positive root is Equation (32.51). That is the second period doubling, and the process repeats.
∎The doublings continue: \(r_{3}=3.544090\), \(r_{4}=3.564407\), \(r_{5}=3.568759\), and the sequence accumulates geometrically at
beyond which the map has orbits of every period and no stable cycle of any period over most of the remaining interval [May:1976]. The chaotic region is not solid: it is shot through with windows of stable periodic behaviour, the widest being the period-three window opening at \(r=1+\sqrt{8}=3.828427\), where the third iterate \(f_{r}^{3}\) becomes tangent to the diagonal in a saddle-node bifurcation (Proposition 32.23) creating a stable and an unstable three-cycle at once. The algebra of that tangency is not reproduced here and the value is quoted [May:1976].
The existence of a three-cycle is not a curiosity. Li and Yorke proved that a continuous map of an interval possessing a periodic point of period three has periodic points of every period, and an uncountable set of points whose orbits neither converge to any periodic orbit nor to each other [Li:1975]; the paper is where the word “chaos” entered the mathematical literature. The theorem is quoted, not proved here.
The change of variable \(x=\sin^{2}(\pi\vartheta)\) conjugates \(f_{4}\) to the doubling map \(\vartheta\mapsto2\vartheta\bmod1\). Hence \(x_{n}=\sin^{2}\left(2^{n}\pi\vartheta_{0}\right)\), the orbit is periodic exactly when \(\vartheta_{0}\) is rational, and the Lyapunov exponent is \(\lambda=\ln2\) per iterate for almost every \(\vartheta_{0}\). Rests on Equation (32.50) and Definition 32.73.
Derives Proposition 32.86. With \(x=\sin^{2}(\pi\vartheta)\),
which is the same function of \(2\vartheta\): the conjugacy holds and iterating gives \(x_{n}\). Writing \(\vartheta_{0}\) in binary, the doubling map is the shift on its digits, so the orbit is eventually periodic precisely for rational \(\vartheta_{0}\)—these form a set of zero length—and for every other starting point the orbit never repeats. The derivative of the doubling map is \(2\) at every point, so by Definition 32.73 the exponent in the \(\vartheta\) variable is \(\ln2\) exactly; the conjugacy is smooth with nowhere-vanishing derivative except at the two endpoints, which the orbit visits with probability zero, so the exponent is unchanged in \(x\). This is the cheapest complete example of chaos in the book: an explicit solution, an explicit positive exponent, and sensitivity that is nothing more than the loss of one binary digit of \(\vartheta_{0}\) per step—which is also exactly the symbolic dynamics of Proposition 32.43 with the two symbols the two binary digits.
∎Feigenbaum universality
As a control parameter is swept, a wide class of systems reaches chaos through an infinite cascade of period-doubling bifurcations that accumulates at a finite parameter value. The parameter intervals between successive doublings shrink geometrically, and the ratio
is the same number for the logistic map, for a driven pendulum and for a convecting fluid, provided only that the return map of the system has a smooth quadratic maximum [Feigenbaum:1978] [May:1976]. This is the observed fact that a dimensionless constant of a difference equation is measurable in a laboratory: in a mercury Rayleigh–Bénard cell the cascade was followed through several doublings and \(\delta\) recovered to within a few percent [Libchaber:1982]. Rests on Propositions 32.39 and 32.84.
Derivation. Derives Phenomenon 32.87. Three things have to be shown: that the doublings recur without end, that their parameter intervals shrink geometrically so that the cascade accumulates at a finite parameter value, and that the ratio is a property of the cascade rather than of the map. The first two are derived here; the third is derived only as far as identifying what would have to be true, and the statement that it is true is quoted—see Remark 32.88.
The doublings recur. Proposition 32.84 carries the first two exactly: the fixed point \(1-1/r\) loses stability at \(r_{1}=3\) when its multiplier passes through \(-1\), and the two-cycle born there loses stability the same way at \(r_{2}=1+\sqrt{6}\). The mechanism is self-reproducing. Let \(\set{p_{1},\dots,p_{2^{n}}}\) be a stable \(2^{n}\)-cycle at a parameter just below \(r_{n}\). The map \(f_{r}^{2^{n}}\) fixes each \(p_{i}\), and the chain rule makes its derivative there the product \(\prod_{j}f_{r}'(p_{j})\) over the whole cycle—the same number at every \(p_{i}\), since the factors are the same and multiply in any order—which is the multiplier that Proposition 32.39 grades and which reaches \(-1\) at \(r_{n}\). Take a small interval \(I_{i}\) about \(p_{i}\), so small that \(f_{r}^{2^{n}}\) maps it into itself; on it that iterate is smooth with negative slope at \(p_{i}\). It also has a critical point there: the orbit of \(I_{i}\) visits, once per period, the interval containing the single critical point \(x=1/2\) of \(f_{r}\), so by the chain rule \(\left(f_{r}^{2^{n}}\right)'\) vanishes at the point of \(I_{i}\) whose orbit lands on \(1/2\). That critical point is quadratic, because \(f_{r}''(1/2)=-2r\neq0\) and the chain rule multiplies it by factors that do not vanish. So \(f_{r}^{2^{n}}\) restricted to \(I_{i}\), reflected and rescaled to unit size, is again a unimodal map with a quadratic maximum and a multiplier sweeping downwards through \(-1\): exactly the situation at \(r_{1}\), one level up. It doubles again, and the argument applies at every level.
The intervals shrink geometrically. Write \(\mathcal{R}\) for the operation just described—take \(f_{r}\), iterate twice, restrict, reflect and rescale—and let \(r\mapsto\mathcal{R}(r)\) denote the induced map on the parameter, defined by requiring the rescaled map at \(\mathcal{R}(r)\) to look like the original at \(r\). By construction it carries the \(n\)th doubling to the \((n-1)\)th, so \(\mathcal{R}(r_{n})=r_{n-1}\), and it fixes the accumulation point, \(\mathcal{R}(r_{\infty})=r_{\infty}\), because a map at \(r_{\infty}\) has already doubled infinitely often and rescaling removes one doubling from an infinite supply. If \(\mathcal{R}\) is differentiable at \(r_{\infty}\) with derivative \(\delta\), then linearizing about the fixed point in the manner of Definition 32.12 gives \(r_{n-1}-r_{\infty}\approx\delta\,(r_{n}-r_{\infty})\), hence
which is Equation (32.53). Two consequences follow at once. The doubling parameters form a geometric sequence, so their sum converges and the cascade accumulates at a finite \(r_{\infty}\) after infinitely many doublings—a cascade that did not shrink geometrically would run off the end of the parameter range instead. And \(\delta>1\), since the doublings crowd together rather than spread apart, so \(r_{\infty}\) is a repelling fixed point of \(\mathcal{R}\) in the parameter.
The arithmetic checks. Using the exact \(r_{1}=3\) and \(r_{2}=1+\sqrt{6}=3.449490\) of Proposition 32.84 and the numerically located \(r_{3}=3.544090\), \(r_{4}=3.564407\), \(r_{5}=3.568759\) of Remark 32.85, the successive ratios in Equation (32.54) are
converging on \(\delta=4.6692\) from alternate sides. Summing the remaining geometric tail from \(r_{5}\),
against the directly computed \(r_{\infty}=3.5699456\) of Equation (32.52): agreement in the seventh digit, which is the quantitative test that the convergence really is geometric with that ratio and not merely fast.
Why the ratio is not a property of the logistic map. Nothing in the first step used the quadratic form of Equation (32.50) beyond two facts: that the map has one critical point in the interval, and that the critical point is quadratic. Any system whose return map (Definition 32.38) has those two properties enters the same construction at the first level, and after one application of \(\mathcal{R}\) it is described by a rescaled map of the same class. So the cascade of every such system is governed by the same operation \(\mathcal{R}\), and if \(\mathcal{R}\) has a single fixed point attracting that whole class, with one repelling direction of multiplier \(\delta\), then every member inherits that \(\delta\) regardless of what it is made of—a difference equation, a driven pendulum, or a convecting layer of mercury. That is the content of the observation, and Remark 32.90 states the operator whose fixed point it is. Its existence is what is quoted [Feigenbaum:1978].
∎The boundary is sharp and worth marking, in the manner of Remark 17.18. Derived above: that the doublings recur, that they accumulate geometrically at a finite parameter value, that the ratio is asymptotically the multiplier of a repelling fixed point of the doubling operation, and that the class of systems entering the construction is fixed by two local properties of the return map and by nothing else. Quoted: that the doubling operator \(\mathcal{T}\) of Remark 32.90 actually has a fixed point \(g\) in a suitable space of unimodal maps, that its linearization there has exactly one eigenvalue outside the unit circle, and that the eigenvalue is \(4.669201609102990\ldots\). Those three statements are a theorem of analysis, first established by a computer-assisted proof whose primary reference is not in this bibliography; the numbers themselves are quoted from [Feigenbaum:1978]. No measured quantity in this chapter depends on the quoted half: the comparison in Section 32.7.1 is between a measured ratio and a number computed from the logistic map's own cascade, and Equation (32.54) is what connects them.
Two numbers govern the cascade [Feigenbaum:1978]. The first is \(\delta\) of Equation (32.53), the rate at which the parameter intervals shrink,
and the second is the rate at which the attractor itself rescales: successive period-doubled orbits reproduce the previous pattern reduced by
in the state variable. Both are mathematical constants of the doubling operator, so they carry no uncertainty of their own; what carries uncertainty is their measurement in a physical system. For the logistic map the finite-order ratios approach \(\delta\) from above—the first three, computed from the values in Proposition 32.84 and Remark 32.85, are \(4.7514\), \(4.6562\) and \(4.6684\)—so an experiment that resolves only the first few doublings is not limited by that convergence but by its own resolution. In the mercury cell of Section 32.7.1 the measured value is \(\delta=4.4 \pm 0.1\) [Libchaber:1982], about \(6\,\mathrm{\%}\) below the exact constant; the deficit is the resolution of the higher doublings, whose parameter intervals shrink by the factor \(\delta\) at each step and soon fall below the stability of the control.
The mechanism is a renormalization group, in exactly the sense of The Renormalization Group. Define the doubling operator \((\mathcal{T}f)(x)=-\alpha\,f\bigl(f(-x/\alpha)\bigr)\), which composes a map with itself and rescales so that the result is again normalized: it takes the description of a system at one level of the cascade to the description at the next. If \(\mathcal{T}\) has a fixed point \(g\), then \(g\) describes a system that looks the same after every doubling, and every map in the basin of \(g\) flows to it under repeated doubling and therefore inherits its numbers. Linearizing \(\mathcal{T}\) about \(g\), the fixed point turns out to have exactly one eigenvalue outside the unit circle, and that eigenvalue is \(\delta\); \(\alpha\) is the scaling built into the operator. What is universal is thus the unstable direction of a fixed point in a space of maps, and what a physical system contributes is only which point of that space it starts from—which is why the only hypothesis needed is a smooth quadratic maximum. The existence of \(g\) is not elementary; it was established by a computer-assisted proof whose reference is not in this bibliography, and this chapter quotes the constants rather than deriving them.
Quasiperiodicity and intermittency
Period doubling is one of three routes by which a system is observed to become chaotic, and an experiment must distinguish them.
The phase of a limit-cycle oscillator sampled at the period of an external drive obeys, to leading order, the circle map
in which \(\Omega\) is the ratio of the natural frequency of the oscillator to the drive frequency and \(K\) the dimensionless coupling. The winding number \(W=\lim_{n\to\infty}(\vartheta_{n}-\vartheta_{0})/n\) is the observed frequency ratio. Rests on Phenomenon 32.37 and Definition 32.38.
For \(K<1\) the winding number of Equation (32.55) is a continuous, non-decreasing function of \(\Omega\) that is constant on an interval around every rational value—a devil's staircase. The intervals are the mode-locked states; plotted in the \((\Omega,K)\) plane they are the Arnold tongues, widening from zero width at \(K=0\) until at \(K=1\) they fill the axis and the irrational winding numbers form a set of zero length. That is the staircase heard by van der Pol and van der Mark (Phenomenon 32.37), and the critical line \(K=1\) is where the last quasiperiodic motion is destroyed. In the underlying flow, quasiperiodic motion with two incommensurate frequencies is motion on an invariant two-torus; mode locking is the capture of that motion onto a closed orbit on the torus, and the destruction of the torus at \(K=1\) is the second route to chaos, the one proposed by Ruelle and Takens [Ruelle:1971].
The line \(K=1\) deserves its own statement, because everything the quasiperiodic route claims turns on what happens as it is crossed.
For the map Equation (32.55),
Hence: for \(K<1\) the map is an orientation-preserving diffeomorphism of the circle; at \(K=1\) its derivative vanishes at the single point \(\vartheta=0\), where the second derivative vanishes as well, so the critical point is an inflection of cubic type; and for \(K>1\) the derivative is negative on an interval, the map is not injective, and it carries two quadratic critical points. In the last case \(\abs{\dd\vartheta_{n+1}/\dd\vartheta_{n}}\) reaches \(1+K>2\) at \(\vartheta=\tfrac{1}{2}\), so the map both stretches and folds. Rests on Definitions 32.38 and 32.91.
Derives Proposition 32.93. Differentiating Equation (32.55) gives Equation (32.56) at once. Its minimum over the circle is \(1-K\), attained at \(\vartheta=0\), and its maximum is \(1+K\), attained at \(\vartheta=\tfrac{1}{2}\). So the derivative is strictly positive everywhere exactly when \(K<1\), and a smooth circle map with everywhere positive derivative is a diffeomorphism: it is a local diffeomorphism, it is strictly increasing as a map of the line commuting with the translation by \(1\), and it therefore descends to a bijection of the circle preserving orientation.
At \(K=1\) the derivative vanishes at \(\vartheta=0\) and nowhere else. The second derivative of Equation (32.55) is \(2\pi K\sin(2\pi\vartheta)\), which also vanishes at \(\vartheta=0\), while the third derivative there is \(4\pi^{2}K\neq0\). The critical point is therefore cubic, not quadratic—which is why the critical line \(K=1\) has universal numbers of its own and not those of Section 32.6.2, whose whole hypothesis was a quadratic maximum.
For \(K>1\) the derivative is negative wherever \(\cos(2\pi\vartheta)>1/K\), an interval about \(\vartheta=0\), and positive elsewhere. A map that decreases on one interval and increases on another takes some values twice, so it is not injective; and its two turning points \(\vartheta_{\pm}=\pm(2\pi)^{-1}\arccos(1/K)\) are quadratic, since there \(\sin(2\pi\vartheta_{\pm})=\mp\sqrt{1-K^{-2}}\neq0\) makes the second derivative nonzero. The bound on the derivative is read off the maximum computed above.
∎The consequences run in both directions. Below \(K=1\) the return map is a circle diffeomorphism, and a circle diffeomorphism cannot be chaotic: its orbits are ordered on the circle for all time, so nearby points cannot separate and later re-approach in the way Phenomenon 32.72 requires. The invariant circle of the section—the two-torus of the flow—survives, and the motion on it is either mode-locked or quasiperiodic, exactly the dichotomy of Remark 32.92. Above \(K=1\) that protection is gone. Proposition 32.93 supplies precisely the two ingredients Section 32.3.2 identified as sufficient: stretching, by a factor up to \(1+K\), and folding, at the two quadratic turning points. Two points with different histories can now share an image, so the past is not recoverable from the present, and the invariant set of the section is no longer a smooth circle—there is no smooth invariant two-torus left to carry the motion. That is the torus breakdown, in the one model where it can be seen explicitly.
Two things are quoted rather than proved. That an orbit of a circle diffeomorphism has a well-defined winding number independent of the starting point, and that an irrational winding number forces the map to be semi-conjugate to the rigid rotation by it, are Poincaré's rotation-number theorem and Denjoy's theorem; both are ordinary differential equations and circle dynamics, and belong with the material of Ordinary Differential Equations and Sturm–Liouville Theory, which does not carry them. And the general classification—that after three generic bifurcations from a fixed point an invariant three-torus is typically replaced by an attractor with a continuous spectrum—is the Newhouse–Ruelle–Takens theorem, quoted here from [Ruelle:1971]. Its proof is a genericity argument in a space of vector fields, machinery that appears nowhere else in this treatise, and Remark 17.18 is the precedent for quoting such a thing rather than owing it. Nothing measured in Section 32.7.2 depends on the quoted half: what the experiment tests is the presence of a continuous spectral component after a small number of transitions, and that is decided by the argument given in Phenomenon 32.97.
Let a stable and an unstable fixed point of a return map annihilate in a saddle-node bifurcation (Proposition 32.23) at a parameter \(\mu_{c}\), and let the map beyond the bifurcation be \(x_{n+1}=x_{n}+\epsilon+ax_{n}^{2}\) with \(\epsilon\propto\mu-\mu_{c}>0\) and \(a>0\). Then the orbit spends a mean number of iterations
in the narrow channel where the fixed points used to be, producing long stretches of nearly periodic behaviour interrupted by bursts [Pomeau:1980]. Rests on Proposition 32.23 and Definition 32.38.
Derives Proposition 32.95. In the channel the increment per iterate is small, so the difference equation may be replaced by the differential equation \(\dd x/\dd n=\epsilon+ax^{2}\). Separating variables and integrating across the channel,
the limits being taken to the edges of the channel, where the quadratic approximation fails but the integrand is already negligible. That is Equation (32.57). The signature is a distinguishing one: a system approaching chaos by intermittency shows periodic behaviour interrupted by bursts whose mean spacing diverges as an inverse square root of the distance from threshold, and no subharmonics at all.
∎An experiment sweeping a control parameter therefore has three signatures to look for, and they are mutually exclusive in the power spectrum. Period doubling: a sequence of new lines at \(f/2\), \(f/4\), \(f/8\), at parameter intervals shrinking by Equation (32.53). Quasiperiodicity: a second line incommensurate with the first, then broadband. Intermittency: no new lines at all, but bursts interrupting a periodic signal, with the laminar length scaling as Equation (32.57). Each has been observed. The measurements in Section 32.7 are the record of the first two.
Experimental chaos
Rayleigh–Bénard convection
The convection cell is the workhorse of experimental chaos because its control parameter, the Rayleigh number Equation (32.38), is set by a temperature difference and can be held and swept with great stability. The background—Bénard's cells and Rayleigh's stability calculation—belongs to Section 31.8.2; for a real cell with rigid upper and lower boundaries the critical value is \(\mathrm{Ra}_{c}=1708\), not the \(657.5\) of the stress-free idealization used in Definition 32.62.
Libchaber, Laroche and Fauve used liquid mercury, whose Prandtl number \(\sigma=\nu/\kappa\) is about \(0.025\)—from \(\nu\approx1.1\times 10^{-7}\,\mathrm{m}^{2}/\mathrm{s}\) and \(\kappa\approx4.4\times 10^{-6}\,\mathrm{m}^{2}/\mathrm{s}\)—and is therefore two orders of magnitude below that of water. A small Prandtl number makes the convecting rolls oscillate rather than remain steady at modest \(\mathrm{Ra}\), which is exactly what is wanted: the oscillation supplies the periodic orbit whose doubling is to be followed. The cell was small enough to contain only a few rolls, so that the number of active degrees of freedom stayed low, and a horizontal magnetic field was applied to damp the transverse motions that would otherwise introduce further frequencies. The temperature was recorded by a bolometer inside the fluid, and the power spectrum of that signal was computed as \(\mathrm{Ra}\) was raised [Libchaber:1982].
What the spectrum showed was the cascade of Remark 32.96: the fundamental oscillation, then a line at half its frequency, then at a quarter, then at an eighth, each appearing at a value of \(\mathrm{Ra}\) closer to the last than the previous interval by a factor near \(\delta\). From the ratios of the successive intervals the experiment returns
against the exact \(4.669\ldots\) of Remark 32.89 [Libchaber:1982]. The agreement is at the \(6\,\mathrm{\%}\) level, and the shortfall is a resolution effect: each successive interval is smaller than the last by \(\delta\), so the fourth doubling requires holding the temperature difference to a part in \(\delta^{3}\approx100\) of the first interval.
That is the whole evidential point of this chapter. A constant computed from a one-line difference equation with no physical content whatever appears, to a few percent, in the temperature record of a centimetre of mercury—because both systems are governed by the same fixed point of the doubling operator of Remark 32.90, and by nothing else they share.
Taylor–Couette flow
Fluid between concentric cylinders, the inner one rotating, becomes unstable to a stack of toroidal vortices above a critical rotation rate; Taylor computed the threshold and measured it in the same paper, obtaining what was then the first quantitative agreement between a hydrodynamic stability calculation and an experiment [Taylor:1923]. In the narrow-gap limit the criterion is a critical Taylor number of \(1708\), the same number as the rigid-boundary Rayleigh–Bénard threshold, the two problems having the same linear operator.
Gollub and Swinney took that apparatus into the chaotic regime and measured, by laser-Doppler velocimetry, the radial velocity at a point in the fluid as the rotation rate was increased through and well beyond the Taylor threshold [Gollub:1975]. Velocimetry is what makes the experiment decisive: it delivers a time series long enough and clean enough for a power spectrum, in a flow that a visualization would only show to be complicated.
In a fluid confined between rotating cylinders, the power spectrum of the velocity acquires, as the rotation rate is raised, first one sharp frequency, then a second incommensurate with it, and then — after only a small number of such steps — a broadband continuous component that grows until it dominates [Gollub:1975]. The measurement excludes Landau's picture of turbulence as the accumulation of an unbounded sequence of incommensurate frequencies, which would show a spectrum of ever more discrete lines and no continuum, and supports instead a route in which a low-dimensional attractor with a continuous spectrum appears after a few bifurcations [Ruelle:1971]. The underlying instability of the rotating flow is Taylor's [Taylor:1923]. Rests on Remark 32.69 and Definition 32.67.
Derivation. Derives Phenomenon 32.97. What has to be derived is why the shape of the spectrum decides between the two pictures at all, since both predict complicated motion.
Landau's flow after \(n\) instabilities is quasiperiodic with \(n\) incommensurate frequencies: by construction the velocity at a point is
with \(V\) continuous and \(2\pi\)-periodic in each of its arguments. Such a \(V\) has a multiple Fourier series, so
and all the power sits on the set \(\set{\vect{k}\cdot\vect{\omega}:\vect{k}\in\Z^{n}}\), which is countable for every finite \(n\). A countable set carries no continuous spectral component, whatever its density: the spectrum of a quasiperiodic signal consists of lines and nothing else, and this is true however large \(n\) becomes. Landau's route therefore predicts a spectrum that grows ever denser in lines and never acquires a continuum, at any stage.
A continuous component in the spectrum requires an autocorrelation that decays. If \(v\) has an integrable autocorrelation then its transform is a continuous function tending to zero at large frequency, by the Riemann–Lebesgue lemma (Lemma 17.7); conversely Equation (32.59) has an autocorrelation \(\sum_{\vect{k}}\abs{c_{\vect{k}}}^{2} \ee^{\ii\vect{k}\cdot\vect{\omega}\tau}\) which is itself quasiperiodic and never decays. So the appearance of a broadband component is precisely the statement that correlations are being lost—which is what a positive Lyapunov exponent does, by Equation (32.44), and what motion on a torus cannot do.
The measurement therefore reads as follows. Sharp lines appearing one at a time are consistent with both pictures. The appearance of a broadband component, after a number of transitions that can be counted on one hand, is consistent with only one of them. Gollub and Swinney observed exactly that: one frequency at the Taylor transition, a second incommensurate frequency, and then, at a rotation rate of order ten times the critical one, a continuous component that could not be resolved into further lines [Gollub:1975]. The exact multiple depends on the radius ratio and aspect ratio of the cell and is not a universal number; the number of transitions before the continuum appears is the observable that matters, and it is small.
∎The derivation above settles the spectral half of the reading: a quasiperiodic flow of any number of frequencies has a pure point spectrum, so a continuous component cannot be Landau's accumulation. The other half—how the invariant two-torus is destroyed once the coupling is strong enough—is Proposition 32.93, whose loss of invertibility at \(K=1\) produces the stretching and folding that a torus cannot support, and Remark 32.94 states exactly which part of the general classification is quoted from [Ruelle:1971] and which is derived. Nothing Gollub and Swinney measured depends on the quoted part.
Circuits and driven oscillators
Electronics is where chaos was first heard, and it remains where it is most easily produced. A circuit has no boundary layers, no aspect ratio and no thermal inertia; its state variables are two voltages and a current, its parameters are components with catalogue values, and its time series is available directly on an oscilloscope.
The first instance is van der Pol and van der Mark's neon-tube oscillator of Phenomenon 32.37, in which the irregular noise heard between locking steps was reported in 1927 and understood half a century later [vanderPol:1927]. The modern bench standard is Chua's circuit, an inductor, two capacitors, a linear resistor and one nonlinear resistor with a piecewise-linear characteristic. In SI variables its equations are
with \(v_{1},v_{2}\) in \(\mathrm{V}\), \(i_{L}\) in \(\mathrm{A}\), \(C_{1},C_{2}\) in \(\mathrm{F}\), \(L\) in \(\mathrm{H}\), \(R\) in \(\mathrm{\Omega}\), and the nonlinear element's characteristic
whose slopes \(G_{a},G_{b}\) are in \(\mathrm{S}\) and whose breakpoint \(E\) is in \(\mathrm{V}\). The system is three-dimensional and autonomous, the minimum Remark 32.9 permits, and it is dissipative wherever the total conductance is positive. Equations (32.60), (32.61) and (32.62) were shown numerically to possess a chaotic attractor by Matsumoto [Matsumoto:1984], and the same attractor is produced on a breadboard in an undergraduate afternoon; sweeping \(R\) carries the circuit through a period-doubling cascade of the kind measured in Section 32.7.1.
The mechanical counterpart is the periodically driven, damped pendulum,
with the drive torque \(\Gamma\) in \(\mathrm{N}\,\mathrm{m}\). Writing the drive phase as a third variable makes Equation (32.64) an autonomous three-dimensional system, and taking a Poincaré section once per drive period (Definition 32.38) reduces it to a two-dimensional map; in the limit of weak damping and impulsive driving that map is the standard map Equation (32.25). The same apparatus that measures the isochronism of Experiment: The Pendulum at small amplitude becomes, driven hard, one of the cleanest demonstrations of a strange attractor available.
In every one of these systems the quantitative claim—that the motion is chaotic and not merely complicated—rests on the same measurement chain: record a scalar time series, embed it, extract \(\lambda_{1}\) by the algorithm of [Wolf:1985], and estimate \(D_{2}\) by pair counting Equation (32.49). A positive exponent and a non-integer dimension, measured from data with stated uncertainties, are what distinguish deterministic chaos from noise. That distinction—and not any of the mathematics above—is what makes chaos an experimental subject.