Hamiltonian Mechanics
The Lagrangian formulation of Lagrangian Mechanics describes a system with \(f\) degrees of freedom by the generalized coordinates \(q^{a}\) and their velocities \(\dot{q}^{a}\). The Hamiltonian formulation trades the velocities for the generalized momenta \(p_{a}=\pp\Lag/\pp\dot{q}^{a}\), and in doing so moves the dynamics onto phase space, the \(2f\)-dimensional space of the pairs \((q^{a},p_{a})\). The reward for the exchange is structural: the equations of motion become a symmetric first-order system, the transformations that preserve their form—the canonical transformations—constitute a far larger class than the point transformations of Lagrangian mechanics, and the whole dynamics is expressed through a single algebraic operation, the Poisson bracket [Goldstein:2002] [Landau:1976].
The index conventions of Notation 21.1 remain in force; additionally \(c,d,e=1,\ldots,2f\) label phase-space components. Poisson brackets are written with braces, \(\pb{u}{v}\), and Lagrange brackets with square brackets, \(\left[u,v\right]\), a subscript indicating the coordinate pair in which they are computed.
From the Lagrangian to the Hamiltonian
The Legendre transformation of the Lagrangian
The generalized momentum conjugate to the coordinate \(q^{a}\) was defined in Definition 21.34 as \(p_{a}=\pp\Lag/\pp\dot{q}^{a}\). It is the variable that replaces the generalized velocity in the description that follows.
The Hamiltonian of a system is the Legendre transformation, in the generalized velocities \(\dot{q}^{a}\), of the Lagrangian of the system:
the generalized velocities being expressed in terms of \((q,p,t)\) through the inversion of Equation (21.43). Rests on Definitions 21.31 and 21.34.
The definition presupposes that the inversion can be carried out, and that is a genuine hypothesis rather than a formality: it is the exact point at which the two formulations can fail to be equivalent.
Let \(\Lag\) be of class \(C^{2}\) and let the velocity Hessian
be non-singular at a point of the tangent bundle. Then Equation (21.43) can be solved near that point for the velocities, \(\dot{q}^{a}=\dot{q}^{a}(q,p,t)\), by a \(C^{1}\) function, and Equation (22.1) defines \(\Ham\) there. Rests on Equation (21.43) and Corollary A.292.
Derives Proposition 22.3. At fixed \((q,t)\) the map \(\dot{q}\longmapsto p\) of Equation (21.43) is \(C^{1}\), and its Jacobian matrix is
symmetric by Proposition 7.105. Where \(\det W\neq0\) the inverse function theorem Corollary A.292 supplies a \(C^{1}\) local inverse, and the dependence on the parameters \((q,t)\) is inherited from that of \(\Lag\).
∎For the Lagrangians of Section 21.7.3, \(\Lag=\tfrac12 g_{ab}(q)\dot{q}^{a}\dot{q}^{b}-V(q)\), the Hessian is the kinetic metric itself, \(W_{ab}=g_{ab}\), which is positive definite, so the inversion is global in the velocities and gives \(\dot{q}^{a}=g^{ab}p_{b}\). The condition \(\det W\neq0\) is also the strict form of Legendre's necessary condition Theorem 16.51 for the variational problem the Lagrangian poses, which is why the same convexity governs minimality and the passage to the Hamiltonian. When it fails — as it does for every gauge theory, where the Lagrangian does not depend on some of the velocities at all — the momenta are not independent functions of the velocities, the relations among them are constraints, and the formalism of this chapter must be rebuilt: that is the subject of Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism. Rests on Proposition 22.3 and Definition 21.66.
\(\Ham\) carries the unit of energy, \(\mathrm{J}\), whatever the units of the individual \(q^{a}\). Each product \(\dot{q}^{a}p_{a}\) in Equation (22.1) is an energy, and necessarily so: by Equation (21.43) the unit of \(p_{a}\) is that of \(\Lag\) divided by that of \(\dot{q}^{a}\), so the two units cancel term by term. For a Cartesian coordinate \(p_{a}\) is a linear momentum, \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\); for an angle it is an angular momentum, \(\mathrm{J}\,\mathrm{s}\). The one combination whose unit never depends on the choice of coordinates is \(p_{a}\,\dd q^{a}\), which is an action, \(\mathrm{J}\,\mathrm{s}\) — the unit of \(\hbar\), and the reason the symplectic structure built from it in Symplectic Geometry of Phase Space is measured in the same units as the quantum of action (Remark 5.120). Rests on Equations (21.43) and (22.1).
Hamilton's equations
The equations that follow are Hamilton's, from the general method in dynamics of [Hamilton:1834]; he obtained them as the characteristic-function formulation whose partial differential equation is the subject of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy, and the canonical pair is what remains when that function is not solved for.
The equations of motion of a system with Hamiltonian \(\Ham(q,p,t)\) are
together with the relation
Here the active forces are taken to be conservative, so that the Lagrangian obeys Equation (21.39). Rests on Definition 22.2, Equation (21.39) and Equation (21.43).
Derives Theorem 22.6. Differentiate the Lagrangian as a function of its own arguments \((q,\dot{q},t)\):
the last step by the definition Equation (21.43) of the momentum. Differentiating Equation (22.1) instead, and subtracting,
From Equations (22.1) and (22.6) (the two terms in \(\dd\dot{q}^{a}\) cancel). The cancellation of \(p_{a}\,\dd\dot{q}^{a}\) is the whole content of the Legendre transformation: no differential of a velocity survives in Equation (22.7), so \(\Ham\) is genuinely a function of \((q,p,t)\) and of nothing else — which is what Proposition 22.3 licenses us to assume.
Comparing Equation (22.7) with the differential of \(\Ham\) read off its own arguments,
and using that \(\dd q^{a}\), \(\dd p_{a}\) and \(\dd t\) are independent, gives three identities:
The first is Equation (22.3) and the third is Equation (22.5). Neither uses the equations of motion: both are properties of the transformation alone.
The dynamics enters only now. By the Euler–Lagrange equations Equation (21.39) and Equation (21.43),
the last step by the second identity of Equation (22.8). This is Equation (22.4).
∎The asymmetry exposed by the proof is worth keeping in view. Equation (22.3) holds for any \(\Ham\) obtained by the Legendre transformation, whether or not the motion satisfies the equations of motion; it merely says that the transformation is invertible and inverts the way a Legendre transformation does. Equation (22.4) is Newton's second law in disguise. The elegance of the pair is therefore partly notational — but only partly, since it is exactly this symmetry of appearance that makes the canonical transformations of Section 22.2 available, and those have no Lagrangian counterpart. Rests on Theorem 22.6.
If some active forces are not derivable from a potential, the Euler–Lagrange equations carry the generalized force \(Q_{a}\) of Equation (21.38) on their right-hand side, and the last line of the proof gives instead
Equation (22.3) being untouched, since it does not use the equations of motion. Everything in this chapter that rests on the pair Equations (22.3) and (22.4) — conservation of \(\Ham\), the Poisson bracket form of the evolution, Liouville's theorem — fails in exactly the way Equation (22.9) predicts when \(Q_{a}\neq0\). Rests on Theorem 22.6 and Equation (21.38).
Where the Lagrange equations are \(f\) second-order equations for the \(q^{a}(t)\), Hamilton's equations are \(2f\) first-order equations for the phase-space trajectory \((q^{a}(t),p_{a}(t))\).
The Hamiltonian and the energy
Along the motion,
Derives Proposition 22.9. \(\Ham\) depends on time both through the trajectory and explicitly, so by the chain rule Proposition 7.104,
substituting Equations (22.3) and (22.4). The first two terms are the same sum with opposite signs and cancel.
∎If the active forces of a system are conservative and the Hamiltonian of the system does not depend explicitly on time, then the Hamiltonian is constant. Rests on Proposition 22.9.
Derives Theorem 22.10. The hypothesis on the forces is what puts Equation (22.4) in force and hence Equation (22.10); the hypothesis \(\pp\Ham/\pp t=0\) then makes its right-hand side vanish, so \(\dd\Ham/\dd t=0\) along every motion and \(\Ham\) is constant on each trajectory (Corollary 7.36). By Equation (22.5) the condition \(\pp\Ham/\pp t=0\) is the same as \(\pp\Lag/\pp t=0\), so the hypothesis may be tested on whichever of the two functions is in hand.
∎If the kinetic energy \(T\) does not depend explicitly on time, then the Hamiltonian is the total energy of the system,
Rests on Equation (22.1) and Proposition 21.38.
Derives Proposition 22.11. Under the stated hypothesis Proposition 21.38 gives \(E=\dot{q}^{a}p_{a}-\Lag\), and the right-hand side is Equation (22.1) verbatim. The two quantities are therefore the same function, expressed once in the velocities and once in the momenta.
∎“\(T\) does not depend explicitly on time” is the wording of Proposition 21.38 and is a shorthand. What the identification \(\Ham=T+V\) actually needs is that \(T\) be a homogeneous quadratic form in the generalized velocities and \(V\) be independent of them — which follows from Equation (21.74) when the constraints are scleronomous (Remark 21.8), and fails when they are not. The step is Euler's theorem on homogeneous functions: if \(T\) is homogeneous of degree \(k\) in the \(\dot{q}^{a}\), then \(\dot{q}^{a}\,\pp T/\pp\dot{q}^{a}=kT\), so for \(k=2\),
A rheonomous constraint makes \(T\) the sum of terms of degrees \(2\), \(1\) and \(0\) in the velocities, the theorem then gives \(\dot{q}^{a}p_{a}=2T_{2}+T_{1}\), and \(\Ham=T_{2}-T_{0}+V\) is a conserved quantity that is not the energy — the rotating-frame case of Rigid Bodies and Rotating Frames, where the difference is the centrifugal term. Conservation of \(\Ham\) and conservation of \(E\) are thus two different statements that coincide only under the hypothesis above. Rests on Proposition 22.11 and Equation (21.74).
Euler's theorem on homogeneous functions, used in Equation (22.12), is a statement of several-variable differential calculus and belongs in Real Analysis beside the chain rule Proposition 7.104, which is all its one-line proof requires. It is not stated there, and it is quoted here rather than proved, so that the gap is visible rather than papered over. Nothing else in this chapter depends on it: Proposition 22.11 itself rests only on Proposition 21.38, and Theorem 22.10 does not use it at all. Rests on Propositions 7.104 and 22.11.
Canonical transformations
Definition and transformation conditions
A transformation of coordinates of phase space is a transformation of the form
It is assumed invertible and twice continuously differentiable, so that \((q,p)\) may equally be regarded as functions of \((Q,P,t)\). Rests on Definition 22.2.
A canonical transformation is a transformation of phase-space coordinates that leaves Hamilton's equations invariant in form; that is, one for which
Here \(S[q,p]=\int_{t_{a}}^{t_{b}}\left(p_{a}\dot{q}^{a}-\Ham\right)\dd t\) is the action of Definition 21.45 written in phase-space variables, and \(S[Q,P]=\int_{t_{a}}^{t_{b}}\left(P_{a}\dot{Q}^{a}-K\right)\dd t\) is the same functional built from the new variables and from a new Hamiltonian \(K\). Rests on Definition 22.14, Theorem 22.6 and Definition 21.45.
Equation (22.13) lets \(Q^{a}\) depend on the momenta, which a point transformation Equation (21.63) of Lagrangian mechanics cannot. The gain is not cosmetic: a canonical transformation may exchange coordinates and momenta outright — \(Q=p\), \(P=-q\) satisfies every condition below — so the distinction between a “position” and a “momentum” has no invariant meaning in phase space. That is what makes it possible to look for a transformation in which every new coordinate is cyclic, which is the programme of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy. Rests on Definitions 21.60 and 22.15.
A transformation Equations (22.13) and (22.14) is canonical if and only if
where the derivatives on the left-hand sides are taken in the old variables and those on the right-hand sides in the new. That is: on the left, the variables held fixed are the remaining \((q,p)\); on the right, the remaining \((Q,P)\). Rests on Definition 22.15, Equation (22.13) and Equation (22.14).
The conditions are \(4f^{2}\) scalar equations relating derivatives taken in two different sets of variables, and the shortest route to them runs through the generating function of Section 22.2.2. The derivation is therefore deferred to Section 22.2.3, once that function is available.
Generating functions
Postulating a canonical transformation leaves an arbitrary function \(F\) at our disposal: since the variational statements Equation (22.15) agree, the two actions may differ by the integral of a total time derivative,
The function \(F\) is called the generating function of the transformation. Once it is known, the transformation from \((q,p)\) to \((Q,P)\) is determined uniquely. To obtain the transformation one must take
which involves the \(4f\) phase-space variables, of which only \(2f\) are independent, the transformation itself relating the two sets; the dependent variables are eliminated by passing to one of the following mixed forms.
In each case \(K\) denotes the Hamiltonian that governs the transformed variables \((Q,P)\).
Let a canonical transformation admit a generating function depending on one old and one new variable of each conjugate pair, the two being independent. Then the transformation and the new Hamiltonian are given by the following relations.
-
[Transformation \((q,Q)\rightarrow(p,P)\).]
\begin{align} F_{1}&=F_{1}(q,Q,t)\ec& p_{a}&=\pdv{F_{1}}{q^{a}}\ec& P_{a}&=-\pdv{F_{1}}{Q^{a}}\ec& K&=\Ham+\pdv{F_{1}}{t}\ep\tag{22.20} \end{align} -
[Transformation \((q,P)\rightarrow(p,Q)\).]
\begin{align} F_{2}&=F_{2}(q,P,t)\ec& p_{a}&=\pdv{F_{2}}{q^{a}}\ec& Q^{a}&=\pdv{F_{2}}{P_{a}}\ec& K&=\Ham+\pdv{F_{2}}{t}\ep\tag{22.21} \end{align} -
[Transformation \((p,Q)\rightarrow(q,P)\).]
\begin{align} F_{3}&=F_{3}(p,Q,t)\ec& q^{a}&=-\pdv{F_{3}}{p_{a}}\ec& P_{a}&=-\pdv{F_{3}}{Q^{a}}\ec& K&=\Ham+\pdv{F_{3}}{t}\ep\tag{22.22} \end{align}
Rests on Equation (22.18) and Definition 22.15.
Derives Proposition 22.18. Write out Equation (22.18) with the two actions of Definition 22.15. Since the equality of the two variational problems must hold for every comparison path between the same endpoints, and not merely for solutions, the integrands agree identically on phase space:
Take first \(F=F_{1}(q,Q,t)\), which is legitimate when \((q,Q)\) are \(2f\) independent variables. Then
and substituting into Equation (22.23) and collecting terms,
At a given instant the \(2f\) quantities \(\dot{q}^{a}\) and \(\dot{Q}^{a}\) may be prescribed arbitrarily and independently, because \((q,Q)\) are independent coordinates and the relation holds along every path. An affine expression that vanishes for all values of its variables has vanishing coefficients, which gives the three relations of Equation (22.20).
The remaining forms follow by Legendre transformation in one pair of variables, exactly as \(\Ham\) followed from \(\Lag\). Put \(F_{2}=F_{1}+P_{a}Q^{a}\); then
using \(\dd F_{1}=p_{a}\dd q^{a}-P_{a}\dd Q^{a}+(\pp F_{1}/\pp t)\dd t\) from Equation (22.20). No \(\dd Q^{a}\) survives, so \(F_{2}\) is a function of \((q,P,t)\), and reading off its partial derivatives gives Equation (22.21) — the last of them because \(\pp F_{2}/\pp t=\pp F_{1}/\pp t\). Put instead \(F_{3}=F_{1}-p_{a}q^{a}\); the same computation gives \(\dd F_{3}=-q^{a}\dd p_{a}-P_{a}\dd Q^{a}+(\pp F_{1}/\pp t)\dd t\), which is Equation (22.22). The fourth member of the family, \(F_{4}=F_{1}+P_{a}Q^{a}-p_{a}q^{a}\), is obtained by performing both Legendre transformations and satisfies \(q^{a}=-\pp F_{4}/\pp p_{a}\), \(Q^{a}=\pp F_{4}/\pp P_{a}\), \(K=\Ham+\pp F_{4}/\pp t\).
∎Not every canonical transformation admits every mixed form, and some admit none of the four. The identity transformation \(Q^{a}=q^{a}\), \(P_{a}=p_{a}\) is generated by \(F_{2}=q^{a}P_{a}\) but has no \(F_{1}\) at all, since \(q\) and \(Q\) are not then independent; the exchange \(Q^{a}=p_{a}\), \(P_{a}=-q^{a}\) has \(F_{1}=q^{a}Q^{a}\) but no \(F_{2}\). The requirement in each case is the independence assumed in Proposition 22.18, which by Corollary A.292 is the non-vanishing of the corresponding Jacobian — \(\det\left(\pp^{2}F_{1}/\pp q^{a}\pp Q^{b}\right)\neq0\) for the first form, and correspondingly for the others. Nothing in the results of this chapter depends on a generating function existing in a particular form: Lemma 22.20 below is proved from Equation (22.18) itself, which asks only that some \(F\) exist. Rests on Proposition 22.18 and Corollary A.292.
Derivation of the transformation conditions
Assembling the \(2f\) new coordinates into a single column, as Equation (22.55) does for the old ones, the transformation Equations (22.13) and (22.14) has at each instant the Jacobian matrix
Everything below is a statement about \(M\) and the symplectic matrix \(J\) of Equation (22.56), whose defining properties \(J\transpose=-J\), \(J^{2}=-\identity_{2f}\) and \(J^{-1}=-J\) are recorded in Equation (5.173).
The Jacobian Equation (22.24) of a canonical transformation satisfies, at every point and at every instant,
that is, \(M\in\Sp(2f,\R)\). Conversely, a transformation not involving the time explicitly whose Jacobian satisfies Equation (22.25) is canonical. Rests on Equation (22.18), Definition 5.135 and Proposition 7.105.
Derives Lemma 22.20. Necessity. Regard \((q,p,t)\) as the independent variables, so that \(Q^{a}\), \(P_{a}\) and the generating function \(F\) of Equation (22.18) are functions of them. Expanding \(\dot{Q}^{a}\) and \(\dd F/\dd t\) by the chain rule Proposition 7.104 in Equation (22.23) and collecting the coefficients of the independently prescribable \(\dot{q}^{b}\) and \(\dot{p}_{b}\), as in the proof of Proposition 22.18, gives
together with \(K=\Ham+P_{a}\,\pp Q^{a}/\pp t+\pp F/\pp t\), which plays no part here.
Now impose on \(F\) the equality of mixed second partial derivatives (Proposition 7.105). Differentiating Equation (22.26) with respect to \(p_{c}\) and Equation (22.27) with respect to \(q^{b}\), the terms carrying second derivatives of \(Q^{a}\) cancel between the two — they are the same by Proposition 7.105 applied to \(Q^{a}\) — and what is left is
Differentiating Equation (22.26) with respect to \(q^{c}\) and comparing with the same expression with \(b\) and \(c\) exchanged gives, after the same cancellation, that \(A\transpose C\) is symmetric; doing the same with Equation (22.27) gives that \(B\transpose D\) is symmetric. Written out in blocks,
whose four blocks are, in order, \(0\) by the symmetry of \(A\transpose C\), \(\identity_{f}\) by Equation (22.28), \(-\identity_{f}\) because that block is minus the transpose of the one above it, and \(0\) by the symmetry of \(B\transpose D\). That is Equation (22.25).
Sufficiency for a transformation free of the time. With \(\zeta=\zeta(\eta)\) and \(K(\zeta)=\Ham(\eta(\zeta))\), the chain rule gives \(\pp\Ham/\pp\eta=M\transpose\,\pp K/\pp\zeta\), so by the matrix form Equation (22.57) of Hamilton's equations, verified in Proposition 22.41 below,
Transposing and inverting Equation (22.25) gives \(MJM\transpose=J\) as well, so \(\dot{\zeta}=J\,\pp K/\pp\zeta\): Hamilton's equations keep their form, for every \(\Ham\). The corresponding statement for explicitly time-dependent transformations, and the geometric reading of Equation (22.25) as the preservation of a differential form, are Theorem 24.18.
∎Derives Proposition 22.17. By Lemma 22.20 a transformation is canonical only if \(M\transpose J M=J\), which since \(J^{-1}=-J\) is the same as
the last equality by multiplying out the blocks of Equation (22.24) with those of Equation (22.56). But \(M^{-1}\) is the Jacobian of the inverse transformation,
whose entries are exactly the right-hand sides of Equations (22.16) and (22.17). Reading Equation (22.30) block by block therefore gives \(\pp p_{b}/\pp P_{a}=A_{ab}=\pp Q^{a}/\pp q^{b}\), \(\pp q^{b}/\pp P_{a}=-B_{ab}=-\pp Q^{a}/\pp p_{b}\), \(\pp p_{b}/\pp Q^{a}=-C_{ab}=-\pp P_{a}/\pp q^{b}\) and \(\pp q^{b}/\pp Q^{a}=D_{ab}=\pp P_{a}/\pp p_{b}\), which are the four conditions. Conversely, the four conditions assert Equation (22.30); multiplying by \(M\) on the left gives \(\identity_{2f}=-M J M\transpose J\), and multiplying by \(J^{-1}=-J\) on the right returns \(M J M\transpose=J\), hence Equation (22.25), and Lemma 22.20 closes the equivalence.
∎The four families of Equations (22.16) and (22.17) are \(4f^{2}\) scalar equations, but Equation (22.25) shows that they are not independent: \(M\transpose J M\) is antisymmetric for every \(M\), so requiring it to equal \(J\) imposes only \(f(2f-1)\) conditions. The dimension of \(\Sp(2f,\R)\) is accordingly \(4f^{2}-f(2f-1)=f(2f+1)\) (Proposition 5.139), which for one degree of freedom is \(3\): the canonical transformations of a two-dimensional phase space are, pointwise, the area- and orientation-preserving linear maps. Rests on Propositions 5.139 and 22.17.
Invariant integrals
The integral invariants are due to Poincaré, who introduced them in volume III (1899) of Les méthodes nouvelles de la mécanique céleste [Poincare:1892].
Let \(S\) be a two-dimensional surface of phase space, parametrized by \((u,v)\) over a domain \(D\), so that
Then for a canonical transformation from \((q,p)\) to \((Q,P)\), applied at one instant,
Rests on Lemma 22.20 and Definition 7.125.
Derives Theorem 22.22. Write \(\eta\) for the column of Equation (22.55) and \(\zeta\) for the same column built from \((Q,P)\). Each Jacobian determinant in Equation (22.31) is a difference of two products, and summing over \(a\) assembles them into the single matrix product
as multiplying out the blocks of Equation (22.56) shows. By the chain rule Proposition 7.104, the same surface described in the new variables has \(\pp\zeta/\pp u=M\,\pp\eta/\pp u\) and likewise for \(v\), with \(M\) the Jacobian Equation (22.24). Hence
by Equation (22.25). The two integrands of Equation (22.32) agree at every point of \(D\), so the integrals agree.
∎Equation (22.31) defines \(I\) through a chosen parametrization of \(S\), and the value is unchanged by any other regular parametrization of the same oriented surface: the integrand Equation (22.33) is bilinear and antisymmetric in \(\pp\eta/\pp u\) and \(\pp\eta/\pp v\), so a change of parameters multiplies it by the Jacobian determinant of that change, which the transformation of the area element \(\dd u\,\dd v\) then cancels. The cancellation is the change-of-variables theorem for a multiple integral, which Real Analysis does not state. It is quoted twice in this chapter and nowhere else — here, and for the transformation of the volume element in Corollary 22.24 — and in neither place is it needed for the statement itself: the two integrands of Theorem 22.22 are equal point by point over one and the same \(D\), and Equation (22.34) is the assertion \(\det M=1\) about a determinant. Rests on Theorem 22.22.
A canonical transformation has \(\det M=1\), so the phase-space volume element is invariant,
In particular the flow generated by a Hamiltonian preserves phase-space volume, since the map carrying the state at time \(t_{a}\) to the state at time \(t_{b}\) is itself a canonical transformation. Rests on Lemma 22.20 and Theorem 5.138.
Derives Corollary 22.24. By Lemma 22.20 the Jacobian \(M\) lies in \(\Sp(2f,\R)\), and every element of the symplectic group has determinant \(+1\) (Theorem 5.138); the volume element transforms by \(\abs{\det M}\), whence Equation (22.34).
For the flow, let \(M(t)\) be the Jacobian of the map carrying the state at time \(t_{a}\) to the state at time \(t\), and write \(\Ham''_{cd}=\pp^{2}\Ham/\pp\eta^{c}\pp\eta^{d}\), symmetric by Proposition 7.105. Differentiating Equation (22.57) with respect to the initial data gives \(\dot{M}=J\,\Ham''M\), and therefore
using \(J\transpose=-J\), \(J^{2}=-\identity_{2f}\) and the symmetry of \(\Ham''\). At \(t=t_{a}\) the map is the identity, for which \(M\transpose J M=J\) holds; hence it holds for all \(t\), and the flow is a canonical transformation by Lemma 22.20.
∎Corollary 22.24 is Liouville's theorem in coordinates. Its coordinate-free form — that the Hamiltonian flow preserves the \(2f\)-form built from the symplectic structure, of which the invariance of the volume element is one consequence among several — is Theorem 24.21, and the physical use to which it is put, as the theorem that singles out the phase-space measure of statistical mechanics, is Remark 24.22. The chain Equations (22.32) and (22.34) continues: Poincaré's invariant is the first of a family of \(f\) integral invariants, of dimensions \(2,4,\ldots,2f\), the last of which is the volume. They are the exterior powers of one object (Equation (24.13)). Rests on Corollary 22.24 and Theorem 24.21.
The Lagrange bracket
The Lagrange bracket of two phase-space functions \(u\) and \(v\) with respect to the coordinates \((q,p)\) is
the sum over \(a\) of the Jacobian determinants \(\pp(q^{a},p_{a})/\pp(u,v)\). Rests on Definition 22.14 and Proposition 7.104.
Comparing Equation (22.35) with Equation (22.33), the integrand of Poincaré's invariant is the Lagrange bracket of the two surface parameters, and in matrix form
From Equations (22.33) and (22.35) (the same sum of Jacobian determinants, written as a matrix product). This is the form in which its properties are read off.
-
Invariance. The Lagrange bracket is invariant under canonical transformations:
\begin{equation}\tag{22.37} \left[u,v\right]_{q,p}=\left[u,v\right]_{Q,P}\ep \end{equation} -
Antisymmetry.
\begin{equation}\tag{22.38} \left[u,v\right]_{q,p}=-\left[v,u\right]_{q,p}\ep \end{equation} -
Fundamental Lagrange brackets.
\begin{equation}\tag{22.39} \left[q^{a},q^{b}\right]_{q,p}=0\ec\qquad \left[p_{a},p_{b}\right]_{q,p}=0\ec\qquad \left[q^{a},p_{b}\right]_{q,p}=\delta^{a}_{b}\ep \end{equation}
Rests on Equation (22.36) and Lemma 22.20.
Derives Proposition 22.27. (1) Invariance. By Equation (22.36) written in the new variables and the chain rule \(\pp\zeta/\pp u=M\,\pp\eta/\pp u\),
the last step by Equation (22.25). This is Equation (22.37), and it is the same computation as the one that proved Theorem 22.22 — as it must be, the invariant being the integral of the bracket.
(2) Antisymmetry. Exchanging \(u\) and \(v\) in Equation (22.35) exchanges the two products and so reverses the sign, which is Equation (22.38); equivalently, \(J\transpose=-J\) in Equation (22.36).
(3) Fundamental brackets. Take the parameters to be two of the phase-space coordinates themselves, so that the derivatives in Equation (22.35) are Kronecker deltas or zero: \(\pp q^{c}/\pp q^{a}=\delta^{c}_{a}\), \(\pp p_{c}/\pp q^{a}=0\), \(\pp q^{c}/\pp p_{a}=0\), \(\pp p_{c}/\pp p_{a}=\delta_{ca}\). Then
which is Equation (22.39).
∎As computed, Equation (22.39) is a triviality: it says only that the coordinates are coordinates. The statement with content is the one obtained by combining it with Equation (22.37), namely
for a canonical transformation — and these \(f(2f-1)\) independent equations are Equation (22.25) again, written in yet another notation. The same remark applies to the Poisson brackets of Equation (22.44) below, and is why Theorem 24.18 can list the three as equivalent tests for a transformation to be canonical. Rests on Proposition 22.27 and Lemma 22.20.
The Poisson bracket
The Poisson bracket of two phase-space functions \(u\) and \(v\) with respect to the coordinates \((q,p)\) is
Where the Lagrange bracket differentiates the coordinates with respect to the functions, the Poisson bracket differentiates the functions with respect to the coordinates. In the matrix notation of Equation (22.36) the two read
where \(\pp u/\pp\eta\) is the column of the \(2f\) partial derivatives of \(u\). The contrast is the whole of Proposition 22.32 below.
-
Antisymmetry.
\begin{equation}\tag{22.43} \pb{u}{v}_{q,p}=-\pb{v}{u}_{q,p}\ep \end{equation} -
Fundamental Poisson brackets.
\begin{equation}\tag{22.44} \pb{q^{a}}{q^{b}}_{q,p}=0\ec\qquad \pb{p_{a}}{p_{b}}_{q,p}=0\ec\qquad \pb{q^{a}}{p_{b}}_{q,p}=\delta^{a}_{b}\ep \end{equation} -
Invariance. The Poisson bracket is invariant under canonical transformations:
\begin{equation}\tag{22.45} \pb{u}{v}_{q,p}=\pb{u}{v}_{Q,P}\ep \end{equation} -
Jacobi identity. For \(u\), \(v\), and \(w\) functions on phase space,
\begin{equation}\tag{22.46} \pb{u}{\pb{v}{w}}+\pb{v}{\pb{w}{u}}+\pb{w}{\pb{u}{v}}=0\ep \end{equation}
Rests on Equation (22.42), Lemma 22.20 and Proposition 7.105.
Derives Proposition 22.30. (1) Antisymmetry. Exchanging \(u\) and \(v\) in Equation (22.41) exchanges the two products, giving Equation (22.43).
(2) Fundamental brackets. As in the Lagrange case, take the two arguments to be coordinates, so the derivatives are Kronecker deltas or zero:
which is Equation (22.44).
(3) Invariance. By the chain rule Proposition 7.104, a function's gradients in the two sets of variables are related by \(\pp u/\pp\eta=M\transpose\,\pp u/\pp\zeta\), so \(\pp u/\pp\zeta=\left(M^{-1}\right)\transpose\,\pp u/\pp\eta\). Hence, by Equation (22.42) in the new variables,
because Equation (22.25) is equivalent to \(MJM\transpose=J\) and hence, multiplying by \(M^{-1}\) on the left and by \(\left(M^{-1}\right)\transpose\) on the right, to \(M^{-1}J\left(M^{-1}\right)\transpose=J\). This is Equation (22.45). The invariance is what allows the subscript on a Poisson bracket to be dropped, and it will be from here on.
(4) Jacobi identity. For fixed \(u\) the map \(D_{u}:w\longmapsto\pb{u}{w}\) is, by Equation (22.41), a linear first-order differential operator in \(w\): it involves the first derivatives of \(w\) and no higher ones. Now inspect the left-hand side of Equation (22.46), which we call \(\Sigma\). Each of its three terms is a Poisson bracket one of whose arguments is itself a Poisson bracket, so each term is a sum of products of first derivatives of two of the functions with a second derivative of the third — exactly one second-derivative factor in every term.
Collect the terms carrying second derivatives of \(w\). They can arise only from the first two terms of \(\Sigma\), and by antisymmetry those two are
the commutator of two first-order operators. In such a commutator the second-order parts cancel: writing \(D_{u}w=B^{e}\pp_{e}w\) and \(D_{v}w=A^{d}\pp_{d}w\) with coefficients built from the first derivatives of \(u\) and \(v\) alone, the second-order part of \(D_{u}D_{v}w\) is \(B^{e}A^{d}\,\pp_{e}\pp_{d}w\) and that of \(D_{v}D_{u}w\) is \(A^{e}B^{d}\,\pp_{e}\pp_{d}w\), and the two contractions are equal because \(\pp_{e}\pp_{d}w\) is symmetric in \(e\) and \(d\) (Proposition 7.105). So \(\left(D_{u}D_{v}-D_{v}D_{u}\right)w\) is again first order in \(w\). The third term \(\pb{w}{\pb{u}{v}}\) carries only first derivatives of \(w\) to begin with. Hence \(\Sigma\) contains no second derivative of \(w\) at all. Since \(\Sigma\) is unchanged by cyclic permutation of \(u\), \(v\), \(w\), it contains no second derivative of \(u\) and none of \(v\) either. But every term of \(\Sigma\) carries exactly one second derivative, of one of the three functions. Therefore every term cancels and \(\Sigma=0\), which is Equation (22.46).
∎Properties (1) and (4) say precisely that the real vector space of smooth functions on phase space, equipped with \(\pb{\cdot}{\cdot}\), is a Lie algebra in the sense of Definition 5.127 — an infinite-dimensional one. The bracket obeys in addition the Leibniz rule \(\pb{u}{vw}=\pb{u}{v}w+v\pb{u}{w}\), immediate from Equation (22.41), so each \(D_{u}=\pb{u}{\cdot}\) is a derivation of the algebra of observables (Definition 5.128). Both facts are structural rather than dynamical: no Hamiltonian has been named. It is because they hold that the correspondence with the commutator of quantum operators of The Poisson Algebra and the Canonical Bridge to Quantum Mechanics can be a correspondence of algebras and not merely an analogy of formulas. Rests on Proposition 22.30, Definition 5.127 and Definition 5.128.
Let \(u_{1},\ldots,u_{2f}\), with \(u_{c}=u_{c}(q^{a},p_{a})\), be a set of mutually independent functions on phase space. Then
Writing \(\Lambda_{cd}=\left[u_{c},u_{d}\right]_{q,p}\) and \(\Pi_{ce}=\pb{u_{c}}{u_{e}}_{q,p}\) for the two matrices, both antisymmetric, Equation (22.47) says \(\Lambda\transpose\Pi=\identity_{2f}\); each matrix therefore determines the other, and
Rests on Equation (22.42) and Corollary A.292.
Derives Proposition 22.32. Let \(T\) be the \(2f\times2f\) matrix of the derivatives \(T^{g}{}_{c}=\pp\eta^{g}/\pp u_{c}\). The \(u_{c}\) being \(2f\) mutually independent functions of the \(2f\) variables \(\eta\), the map \(\eta\longmapsto u\) has a local inverse (Corollary A.292) and \(T\) is invertible, with \(\left(T^{-1}\right)_{c}{}^{g}=\pp u_{c}/\pp\eta^{g}\). The two matrices of Equation (22.42) are then
Multiplying them,
using \(J^{2}=-\identity_{2f}\) from Equation (5.173). Since \(\Lambda\transpose=-\Lambda\),
which is Equation (22.47), and Equation (22.48) follows on multiplying by \(\left(\Lambda\transpose\right)^{-1}\).
∎Neither bracket determines a transformation on its own, but each is a complete test. Taking the \(u_{c}\) to be the new coordinates \((Q,P)\) themselves, Equation (22.49) reads \(\Lambda=\left(M^{-1}\right)\transpose J\,M^{-1}\) and \(\Pi=MJM\transpose\); the transformation is canonical when either equals \(J\), and Equation (22.47) shows that one follows from the other. This is why the fundamental Lagrange brackets and the fundamental Poisson brackets are interchangeable as tests (Remark 22.28), a fact used without comment in most of the literature. Rests on Proposition 22.32 and Lemma 22.20.
Motion in phase space
For an arbitrary phase-space function \(U=U(q,p,t)\), along the motion,
Derives Theorem 22.34. By the chain rule Proposition 7.104 along a trajectory, and then by Hamilton's equations Equations (22.3) and (22.4),
and the first two terms are Equation (22.41) with \(v=\Ham\).
∎If the function \(U\) does not depend explicitly on time, then
Rests on Theorem 22.34.
Derives Corollary 22.35. Set \(\pp U/\pp t=0\) in Equation (22.50).
∎Equation (22.50) is the most compact statement of Hamiltonian mechanics available: it computes the rate of change of any observable from the bracket alone, with \(\Ham\) singled out only as the observable that generates the flow. Taking \(U=\Ham\) recovers Equation (22.10), since \(\pb{\Ham}{\Ham}=0\) by Equation (22.43). It is also the classical statement of which the Heisenberg equation of motion is the quantum counterpart, \(\pb{\cdot}{\cdot}\longmapsto\comm{\cdot}{\cdot}/\left(\ii\hbar\right)\), and the correspondence — with the exact sense in which it holds and the precise sense in which it fails — is The Poisson Algebra and the Canonical Bridge to Quantum Mechanics. The factor \(\ii\hbar\) is forced by units: by Remark 22.5 a Poisson bracket of two observables carries their units divided by an action, and \(\hbar\) is the action that supplies it. Rests on Theorem 22.34 and Remark 22.5.
In particular, the equations of motion themselves take the bracket form:
From Equations (22.44) and (22.51) (taking \(U=q^{a}\) and \(U=p_{a}\), whose brackets with \(\Ham\) reduce to a single term by Equation (22.44)).
If the Poisson bracket of a function that does not depend explicitly on time with the Hamiltonian is zero, then the function is a constant of motion. Rests on Corollary 22.35.
Derives Proposition 22.37. By Equation (22.51), \(\dd U/\dd t=\pb{U}{\Ham}=0\) along every motion, so \(U\) is constant on each trajectory (Corollary 7.36). The converse holds as well: a \(U\) free of explicit time dependence that is constant along every motion has \(\pb{U}{\Ham}=0\) everywhere, since the trajectory through a given point may be chosen to pass through it in any direction the equations allow.
∎The Poisson bracket of two constants of motion is itself a constant of motion. Rests on Equation (22.46) and Theorem 22.34.
Derives Theorem 22.38. Let \(u\) and \(v\) be constants of motion, so that by Equation (22.50)
Differentiating Equation (22.41) with respect to \(t\) at fixed \((q,p)\) gives the Leibniz rule \(\pp\pb{u}{v}/\pp t=\pb{\pp u/\pp t}{v}+\pb{u}{\pp v/\pp t}\). Next, the Jacobi identity Equation (22.46) with \(w=\Ham\), rearranged by antisymmetry Equation (22.43), gives
the last step by Equation (22.54). Adding the two results and applying Equation (22.50) to the function \(\pb{u}{v}\),
since \(\pb{v}{\pp u/\pp t}=-\pb{\pp u/\pp t}{v}\). In particular, when neither constant depends explicitly on time the argument reduces to \(\pb{\pb{u}{v}}{\Ham}=0\) and Proposition 22.37.
∎The theorem is a genuine tool — the three components of angular momentum close on themselves under the bracket, and knowing two of them to be conserved delivers the third — but it is not a machine for producing integrals. Most often \(\pb{u}{v}\) turns out to be a constant already known, a function of \(u\) and \(v\), or simply zero. It says that the constants of motion of a system form a Lie subalgebra of the algebra of Remark 22.31, and the interest lies in which algebra that is: for the Kepler problem it is larger than the rotation algebra, which is the statement that the Laplace–Runge–Lenz vector is conserved (Central Forces and Statics). Rests on Theorem 22.38 and Remark 22.31.
Symplectic formulation
The pairs \((q^{a},p_{a})\) may be assembled into a single object, and Hamilton's equations into a single matrix equation. Define the coordinate vector \(\eta^{c}\), \(c=1,\ldots,2f\), and the corresponding vector of derivatives of the Hamiltonian by
where the first \(f\) components run over the coordinates and the last \(f\) over the momenta.
The symplectic matrix is the \(2f\times2f\) matrix
It is the matrix of Equation (5.173), the standard non-degenerate antisymmetric form of Proposition 5.119, at \(2n=2f\); the transformations preserving it are the symplectic group \(\Sp(2f,\R)\) of Definition 5.135.
Rests on Equation (5.173) and Definition 5.135.
Hamilton's equations Equations (22.3) and (22.4) are equivalent to
Derives Proposition 22.41. Split the column Equation (22.55) into its two blocks of \(f\) components and multiply out, using Equation (22.56):
while \(\dot{\eta}\) is the column with blocks \(\dot{q}^{a}\) and \(\dot{p}_{a}\). Equating the two columns block by block gives Equation (22.3) from the upper block and Equation (22.4) from the lower, and conversely. The whole effect of \(J\) is to interchange the two blocks and change the sign of one of them, which is exactly the asymmetry between Equations (22.3) and (22.4).
∎Rewriting two equations as one buys nothing by itself. What it buys is that every statement about canonical transformations becomes a statement about the single matrix \(J\): Equation (22.25) is the condition, Equation (22.42) are the brackets, and the proofs in Section 22.2, Section 22.3 and Section 22.3.2 above are three lines each because of it. The price is that \(\eta\) mixes quantities of different physical dimension — Remark 22.5 — so \(J\) is not a tensor on phase space in any metric sense and \(\eta\) is not a vector; the object that is coordinate-free is the two-form whose matrix \(J\) is, and that is the subject of Symplectic Geometry of Phase Space. Rests on Proposition 22.41 and Remark 22.5.
Symplectic form of the transformation condition
The heading reserved at this point in the source outline asks for the symplectic characterization of canonical transformations: the condition on the Jacobian matrix of the transformation that is equivalent to the \(4f^{2}\) direct conditions of Proposition 22.17. It has already been established, because it was needed before the notation of this section was introduced. It is Lemma 22.20: the single matrix identity \(M\transpose J\,M=J\), stating that the Jacobian Equation (22.24) lies in \(\Sp(2f,\R)\) at every point and at every instant.
Four apparently different tests for a transformation to be canonical have now been shown to be one test, and it is worth listing them together.
-
The action principles agree, Equation (22.15) — the definition.
-
The Jacobian satisfies Equation (22.25).
-
The fundamental Lagrange brackets Equation (22.40), equivalently the fundamental Poisson brackets, are preserved.
That (1) implies (3) is the necessity half of Lemma 22.20; that (3) implies (1) for transformations free of the time is its sufficiency half, and in general Theorem 24.18; the equivalence of (2) and (3) is Proposition 22.17; and (4) is (3) read block by block, Remark 22.28. Only (3) survives unchanged into the coordinate-free language of Symplectic Geometry of Phase Space, where it becomes the statement that the transformation pulls the symplectic form back to itself.
Beyond the canonical formalism
Three bodies of theory grow directly out of this chapter, and each is given a chapter of its own rather than a section here.
Hamilton–Jacobi theory
Pushed to its limit, the search for a convenient canonical transformation asks for one that makes every new coordinate and momentum constant. The generating function that achieves it satisfies a single first-order partial differential equation, and its level surfaces propagate through configuration space exactly as optical wavefronts propagate through a medium. Hamilton–Jacobi Theory and the Optical–Mechanical Analogy develops the equation, the action–angle variables it produces for bounded motion, and the optical–mechanical analogy that leads out of classical mechanics altogether.
Symplectic geometry
The matrix \(J\) of Equation (22.56) is the component form of a closed nondegenerate two-form on phase space, and the Poisson bracket, the canonical transformations, the invariant integrals of Section 22.3 and the conservation laws are all statements about it. Symplectic Geometry of Phase Space develops that geometry.
Symmetries
The symmetry content of the canonical formalism—the correspondence between one-parameter families of canonical transformations, their generators, and the constants of motion—follows from the bracket alone, and with it the correspondence rule that carries the whole formalism into quantum mechanics. The Poisson Algebra and the Canonical Bridge to Quantum Mechanics develops both, and Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism treats the systems for which the Legendre transformation of Definition 22.2 fails to be invertible in the first place—which is to say every gauge theory.