DANIEL MARIN

Deriving Einstein's Field Equations from First Principles

A self-contained derivation from six postulates: definitions, the variational calculation in full, the Newtonian limit, and a survey of theories that relax each postulate

27 min

The Einstein field equations can be obtained from a small number of postulates about the structure of spacetime together with a variational principle. This article carries out that derivation in full. Sec. 0 fixes conventions and defines the objects used throughout. Sec. 1 states the six postulates. Sec. 2 constructs the connection and the curvature tensors that follow from them, and Sec. 3 determines the admissible form of the action. The variation is performed in Sec. 4 and the coupling constant is fixed by the Newtonian limit in Sec. 5. Sec. 6 gives equivalent forms of the equations and Sec. 7 verifies internal consistency. Sec. 8 collects every assumption made and Sec. 9 surveys the theories obtained by relaxing them.

Equations are numbered by section and referred to by number. Familiarity with special relativity and multivariable calculus is assumed; no differential geometry is presupposed.

0. Conventions and definitions

The following conventions are used without further comment. Roughly half of the signs appearing below depend on them.

  • Signature. The metric has signature (,+,+,+)(-,+,+,+), so that timelike intervals satisfy ds2<0ds^2 \lt 0.
  • Indices. Greek indices μ,ν,ρ,\mu,\nu,\rho,\dots range over 0,1,2,30,1,2,3, with x0=ctx^0 = ct. Latin indices i,j,ki,j,k range over the spatial values 1,2,31,2,3.
  • Summation. An index appearing once as a superscript and once as a subscript within a term is summed over its range: AμBμμ=03AμBμA^\mu B_\mu \equiv \sum_{\mu=0}^{3} A^\mu B_\mu.
  • Curvature sign. The Riemann tensor is defined by [μ,ν]Vρ=R σμνρVσ[\nabla_\mu, \nabla_\nu] V^\rho = R^\rho_{\ \sigma\mu\nu} V^\sigma, the convention of Misner, Thorne and Wheeler.
  • Symmetrization. A(μν)=12(Aμν+Aνμ)A_{(\mu\nu)} = \tfrac{1}{2}(A_{\mu\nu} + A_{\nu\mu}) and A[μν]=12(AμνAνμ)A_{[\mu\nu]} = \tfrac{1}{2}(A_{\mu\nu} - A_{\nu\mu}).
  • Units. The constants cc and GG are retained explicitly, since the object of Sec. 5 is to determine the combination in which they appear.

A tensor of type (p,q)(p,q) at a point is an object with pp contravariant and qq covariant components transforming under a change of coordinates xμxμx^\mu \to x'^\mu as

T  ν1νqμ1μp=xμ1xα1xβ1xν1T  β1βqα1αp(0.1)T'^{\mu_1 \dots \mu_p}_{\ \ \nu_1 \dots \nu_q} = \frac{\partial x'^{\mu_1}}{\partial x^{\alpha_1}} \cdots \frac{\partial x^{\beta_1}}{\partial x'^{\nu_1}} \cdots T^{\alpha_1 \dots \alpha_p}_{\ \ \beta_1 \dots \beta_q} \tag{0.1}

Two consequences of (0.1) are used repeatedly. First, an equality between tensors that holds in one coordinate system holds in all of them. Second, a tensor whose components vanish at a point in one coordinate system vanishes at that point in every coordinate system. The second is applied four times in Sec. 2 and Sec. 4 to promote a relation established in a convenient frame to a covariant statement.

1. The postulates

P1. Spacetime is a four-dimensional smooth manifold MM. Each point admits a neighborhood homeomorphic to an open subset of R4\mathbb{R}^4, with smooth transition functions between overlapping charts. The dimension is an input to the theory.

P2. MM carries a Lorentzian metric gμνg_{\mu\nu}, and the metric determines the readings of ideal clocks and rods. The metric is a symmetric non-degenerate tensor field of type (0,2)(0,2) and signature (,+,+,+)(-,+,+,+). It determines the line element

ds2=gμνdxμdxν(1.1)ds^2 = g_{\mu\nu}\, dx^\mu dx^\nu \tag{1.1}

and the proper time elapsed along a timelike worldline γ\gamma is

τ=1cγgμνdxμdxν(1.2)\tau = \frac{1}{c}\int_\gamma \sqrt{-g_{\mu\nu}\,dx^\mu dx^\nu} \tag{1.2}

Equation (1.2) supplies the operational content of the metric. Non-degeneracy guarantees the existence of the inverse metric gμνg^{\mu\nu} satisfying gμαgαν=δνμg^{\mu\alpha}g_{\alpha\nu} = \delta^\mu_\nu, by means of which indices are raised and lowered.

P3. General covariance. The laws of physics take the same form in all coordinate systems; equivalently, the action is a scalar functional invariant under diffeomorphisms xμxμ+ξμ(x)x^\mu \to x^\mu + \xi^\mu(x). Coordinates are labels and carry no physical content.

P4. The Einstein equivalence principle. At every point pMp \in M there exist coordinates, called locally inertial or normal coordinates at pp, in which

gμν(p)=ημν=diag(1,1,1,1),λgμν(p)=0(1.3)g_{\mu\nu}(p) = \eta_{\mu\nu} = \mathrm{diag}(-1,1,1,1), \qquad \partial_\lambda g_{\mu\nu}(p) = 0 \tag{1.3}

and in which all non-gravitational physics reduces to special relativity. A freely falling observer therefore cannot detect a gravitational field by any local experiment. Two consequences are used below: matter couples to gμνg_{\mu\nu} and to no other gravitational field (minimal coupling, Sec. 3.4), and the connection is metric-compatible (Sec. 2.2).

P5. The field equations follow from a stationary action. There exists a functional S[g,ψ]S[g,\psi] of the metric and the matter fields ψ\psi whose stationary points under variations vanishing on the boundary are the physical configurations. The metric is the only gravitational field.

P6. Locality and second order. The Lagrangian density is a local scalar constructed from gμνg_{\mu\nu} and finitely many of its derivatives, and the field equations contain derivatives of gμνg_{\mu\nu} of order no higher than two.

P6 requires comment, as it is the postulate least often stated explicitly and the one most frequently relaxed. Second-order equations render the initial-value problem well posed in the standard form: the metric and its first time derivative on a spacelike hypersurface determine the solution. Field equations of higher order generically possess an Ostrogradsky instability, that is, a Hamiltonian unbounded from below.1 The theories of Sec. 9.1 accept this cost in exchange for improved ultraviolet behavior.

2. Consequences of the postulates

2.1 The covariant derivative

The partial derivative of a vector field is not a tensor. Differentiating (0.1) for a type (1,0)(1,0) tensor produces a term proportional to 2x/xx\partial^2 x' / \partial x \partial x, which spoils the transformation law. One therefore introduces the covariant derivative

νVμ=νVμ+ΓνλμVλ,νωμ=νωμΓνμλωλ(2.1)\nabla_\nu V^\mu = \partial_\nu V^\mu + \Gamma^\mu_{\nu\lambda} V^\lambda, \qquad \nabla_\nu \omega_\mu = \partial_\nu \omega_\mu - \Gamma^\lambda_{\nu\mu} \omega_\lambda \tag{2.1}

where the connection coefficients Γμνλ\Gamma^\lambda_{\mu\nu} are defined to transform inhomogeneously in precisely the manner that cancels the unwanted term. They are consequently not the components of a tensor. On scalars μf=μf\nabla_\mu f = \partial_\mu f, and (2.1) extends to tensors of arbitrary type by the Leibniz rule, with one +Γ+\Gamma term for each contravariant index and one Γ-\Gamma term for each covariant index.

A vector field VμV^\mu is parallel-transported along a curve with tangent uμu^\mu if uννVμ=0u^\nu \nabla_\nu V^\mu = 0. A geodesic is a curve that parallel-transports its own tangent, uννuμ=0u^\nu \nabla_\nu u^\mu = 0; in coordinates,

d2xμdτ2+Γαβμdxαdτdxβdτ=0(2.2)\frac{d^2 x^\mu}{d\tau^2} + \Gamma^\mu_{\alpha\beta}\frac{dx^\alpha}{d\tau}\frac{dx^\beta}{d\tau} = 0 \tag{2.2}

Equation (2.2) also follows from extremizing the proper time (1.2) of a free test particle, whose action is Spt=mc2 ⁣dτS_{\rm pt} = -mc^2\!\int d\tau. It is used in Sec. 5.1 to identify the Newtonian potential.

2.2 Determination of the connection

Two conditions determine Γμνλ\Gamma^\lambda_{\mu\nu} uniquely.

Lemma 2.1 (Metric compatibility). P4 implies λgμν=0\nabla_\lambda g_{\mu\nu} = 0.

Proof. In normal coordinates at a point pp the first derivatives of the metric vanish by (1.3). The connection coefficients are constructed from those first derivatives, so they too vanish at pp, and hence λgμν(p)=λgμν(p)=0\nabla_\lambda g_{\mu\nu}(p) = \partial_\lambda g_{\mu\nu}(p) = 0 in this frame. Since λgμν\nabla_\lambda g_{\mu\nu} is a tensor, its vanishing at pp in one frame implies its vanishing at pp in every frame. The point pp was arbitrary. \blacksquare

Geometrically, Lemma 2.1 states that lengths and angles are preserved under parallel transport, as (1.2) requires if the metric is to be measured by transported clocks and rods.

The second condition is the vanishing of the torsion tensor T μνλ=2Γ[μν]λT^\lambda_{\ \mu\nu} = 2\Gamma^\lambda_{[\mu\nu]}, equivalently Γμνλ=Γνμλ\Gamma^\lambda_{\mu\nu} = \Gamma^\lambda_{\nu\mu}. The antisymmetric part of the connection is a tensor, and P1–P6 do not require it to vanish; its vanishing is an independent assumption, recorded in Sec. 8 and relaxed in Sec. 9.4.

Proposition 2.2 (Levi-Civita). A torsion-free metric-compatible connection exists and is unique, with coefficients

Γμνλ=12gλρ(μgνρ+νgμρρgμν)(2.3)\Gamma^\lambda_{\mu\nu} = \frac{1}{2}g^{\lambda\rho}\left(\partial_\mu g_{\nu\rho} + \partial_\nu g_{\mu\rho} - \partial_\rho g_{\mu\nu}\right) \tag{2.3}

Proof. Expanding ρgμν=0\nabla_\rho g_{\mu\nu} = 0 and permuting the free indices cyclically gives

ρgμν=Γρμλgλν+Γρνλgμλμgνρ=Γμνλgλρ+Γμρλgνλνgρμ=Γνρλgλμ+Γνμλgρλ(2.4)\begin{aligned} \partial_\rho g_{\mu\nu} &= \Gamma^\lambda_{\rho\mu}g_{\lambda\nu} + \Gamma^\lambda_{\rho\nu}g_{\mu\lambda} \\ \partial_\mu g_{\nu\rho} &= \Gamma^\lambda_{\mu\nu}g_{\lambda\rho} + \Gamma^\lambda_{\mu\rho}g_{\nu\lambda} \\ \partial_\nu g_{\rho\mu} &= \Gamma^\lambda_{\nu\rho}g_{\lambda\mu} + \Gamma^\lambda_{\nu\mu}g_{\rho\lambda} \end{aligned} \tag{2.4}

Adding the second and third equations of (2.4) and subtracting the first, six of the eight terms on the right cancel in pairs by the symmetry Γμνλ=Γνμλ\Gamma^\lambda_{\mu\nu} = \Gamma^\lambda_{\nu\mu}, leaving 2Γμνλgλρ2\Gamma^\lambda_{\mu\nu}g_{\lambda\rho}. Contraction with gρσg^{\rho\sigma} gives (2.3). Uniqueness follows because the construction inverts the two conditions with no freedom remaining. \blacksquare

The geometry is therefore completely determined by the metric.

2.3 The curvature tensor

Since Γ\Gamma vanishes at any preassigned point in normal coordinates, no tensor constructed from the first derivatives of the metric alone can represent the gravitational field. The invariant content lies in the second derivatives, and is expressed by the failure of covariant derivatives to commute. Evaluating [μ,ν]Vρ[\nabla_\mu,\nabla_\nu]V^\rho from (2.1), the terms containing derivatives of VρV^\rho cancel by the symmetry of Γ\Gamma, and the remainder is algebraic in VρV^\rho:

[μ,ν]Vρ=R σμνρVσ(2.5)[\nabla_\mu, \nabla_\nu] V^\rho = R^\rho_{\ \sigma\mu\nu} V^\sigma \tag{2.5}

with

R σμνρ=μΓνσρνΓμσρ+ΓμλρΓνσλΓνλρΓμσλ(2.6)R^\rho_{\ \sigma\mu\nu} = \partial_\mu \Gamma^\rho_{\nu\sigma} - \partial_\nu \Gamma^\rho_{\mu\sigma} + \Gamma^\rho_{\mu\lambda}\Gamma^\lambda_{\nu\sigma} - \Gamma^\rho_{\nu\lambda}\Gamma^\lambda_{\mu\sigma} \tag{2.6}

The left-hand side of (2.5) is a tensor for arbitrary VσV^\sigma, and therefore so is R σμνρR^\rho_{\ \sigma\mu\nu}, notwithstanding its construction from the non-tensorial quantities (2.3). This is the Riemann curvature tensor. It admits an equivalent characterization as the change in a vector parallel-transported around an infinitesimal closed loop, and R σμνρ=0R^\rho_{\ \sigma\mu\nu} = 0 throughout a region is the necessary and sufficient condition for that region to be flat.

With the first index lowered, Rρσμν=gρλR σμνλR_{\rho\sigma\mu\nu} = g_{\rho\lambda}R^\lambda_{\ \sigma\mu\nu}, the tensor obeys

Rρσμν=Rσρμν,Rρσμν=Rρσνμ,Rρσμν=Rμνρσ(2.7)R_{\rho\sigma\mu\nu} = -R_{\sigma\rho\mu\nu}, \qquad R_{\rho\sigma\mu\nu} = -R_{\rho\sigma\nu\mu}, \qquad R_{\rho\sigma\mu\nu} = R_{\mu\nu\rho\sigma} \tag{2.7}

together with the first Bianchi identity Rρ[σμν]=0R_{\rho[\sigma\mu\nu]} = 0. In nn dimensions the relations (2.7) leave n2(n21)/12n^2(n^2-1)/12 algebraically independent components, which for n=4n = 4 is twenty, as against the ten components of the metric.

2.4 Contractions

The symmetries (2.7) admit essentially one contraction to a tensor of rank two. Contracting the first and third indices defines the Ricci tensor

Rμν=R μλνλ(2.8)R_{\mu\nu} = R^\lambda_{\ \mu\lambda\nu} \tag{2.8}

which is symmetric by the pair-exchange symmetry in (2.7). Contraction of the first with the second, or of the third with the fourth, vanishes by antisymmetry; contraction of the first with the fourth yields Rμν-R_{\mu\nu}. A further contraction defines the Ricci scalar

R=gμνRμν(2.9)R = g^{\mu\nu}R_{\mu\nu} \tag{2.9}

By the same argument, RR is the only scalar constructible from the metric and its first two derivatives that is linear in the second derivatives. This fact determines the form of the action in Sec. 3.2.

2.5 The Bianchi identities

Differentiating (2.6) in normal coordinates at a point, where Γ\Gamma vanishes and the quadratic terms do not contribute, and antisymmetrizing, gives the second Bianchi identity

λR σμνρ+μR σνλρ+νR σλμρ=0(2.10)\nabla_\lambda R^\rho_{\ \sigma\mu\nu} + \nabla_\mu R^\rho_{\ \sigma\nu\lambda} + \nabla_\nu R^\rho_{\ \sigma\lambda\mu} = 0 \tag{2.10}

Both sides are tensors, so (2.10) holds in every frame and at every point. It is an identity of the geometry, valid for an arbitrary metric, and independent of any field equation.

Contracting ρ\rho with μ\mu in (2.10) and using R σνμμ=RσνR^\mu_{\ \sigma\nu\mu} = -R_{\sigma\nu} in the third term,

λRσννRσλ+μR σνλμ=0(2.11)\nabla_\lambda R_{\sigma\nu} - \nabla_\nu R_{\sigma\lambda} + \nabla_\mu R^\mu_{\ \sigma\nu\lambda} = 0 \tag{2.11}

Contracting (2.11) with gσνg^{\sigma\nu}, the first term gives λR\nabla_\lambda R, the second σRσλ-\nabla^\sigma R_{\sigma\lambda}, and the third, on applying the symmetries (2.7), a further μRμλ-\nabla^\mu R_{\mu\lambda}. Hence the contracted Bianchi identity

μRμν=12νR(2.12)\nabla^\mu R_{\mu\nu} = \frac{1}{2}\nabla_\nu R \tag{2.12}

Defining the Einstein tensor

Gμν=Rμν12gμνR(2.13)G_{\mu\nu} = R_{\mu\nu} - \frac{1}{2}g_{\mu\nu}R \tag{2.13}

identity (2.12) is equivalent to μGμν=0\nabla^\mu G_{\mu\nu} = 0. The tensor (2.13) is thus symmetric, constructed from the metric and its first two derivatives, linear in the second derivatives, and identically divergence-free. These four properties characterize it, by the theorem quoted in Sec. 3.2.

3. The action

3.1 The invariant volume element

The coordinate volume d4xd^4x is not invariant, acquiring the Jacobian determinant under a change of coordinates. Writing gdet(gμν)g \equiv \det(g_{\mu\nu}), which transforms with the square of the inverse Jacobian and is negative for Lorentzian signature, the combination

d4xg(3.1)d^4x\,\sqrt{-g} \tag{3.1}

is invariant. By P3 every admissible action is therefore of the form d4xgL\int d^4x\,\sqrt{-g}\,\mathcal{L} with L\mathcal{L} a scalar.

3.2 The admissible scalars

By P6 the scalar L\mathcal{L} is constructed from the metric and its derivatives and must yield field equations of second order. The candidates are enumerated as follows.

  1. A constant is admissible and cannot be excluded. It contributes the cosmological term.
  2. No non-constant scalar can be built from gμνg_{\mu\nu} and λgμν\partial_\lambda g_{\mu\nu} alone, since normal coordinates set the first derivatives to zero at any preassigned point, whereas a scalar cannot vanish at a point in one frame and not in another.
  3. At second order in derivatives the only available tensor is the Riemann tensor, and by Sec. 2.4 the only scalar linear in it is RR. Its variation yields second-order equations.
  4. Scalars quadratic in the curvature, such as R2R^2, RμνRμνR_{\mu\nu}R^{\mu\nu} and RρσμνRρσμνR_{\rho\sigma\mu\nu}R^{\rho\sigma\mu\nu}, are permitted by P1–P5 but yield fourth-order equations and are excluded by P6. The Gauss-Bonnet combination G=R24RμνRμν+RρσμνRρσμν\mathcal{G} = R^2 - 4R_{\mu\nu}R^{\mu\nu} + R_{\rho\sigma\mu\nu}R^{\rho\sigma\mu\nu} is an exception, being in four dimensions a topological density whose variation vanishes identically; it contributes nothing here and is dynamically relevant only in five or more dimensions.

The enumeration is made precise by the following result, which will not be proved here.

Theorem 3.1 (Lovelock, 1971). In four dimensions, the most general symmetric, divergence-free tensor of rank two constructed from the metric and its first two derivatives is

aGμν+bgμν(3.2)a\,G_{\mu\nu} + b\,g_{\mu\nu} \tag{3.2}

for constants aa and bb.

Given P1–P6, Theorem 3.1 exhausts the possible field equations, leaving two constants to be fixed: aa by the Newtonian limit and bb by observation.

3.3 The Einstein-Hilbert action

SEH=12κd4xg(R2Λ)(3.3)S_{EH} = \frac{1}{2\kappa} \int d^4x\, \sqrt{-g}\,(R - 2\Lambda) \tag{3.3}

The constant κ\kappa is left undetermined; no argument internal to the derivation fixes it, and its value is obtained in Sec. 5 by comparison with Newtonian gravity. The constant Λ\Lambda is the cosmological constant.

3.4 The matter action

By P4 matter couples to the metric alone. The corresponding prescription, minimal coupling, is to take the Lagrangian of a field in special relativity and replace ημνgμν\eta_{\mu\nu} \to g_{\mu\nu}, μμ\partial_\mu \to \nabla_\mu, and d4xd4xgd^4x \to d^4x\sqrt{-g}:

SM=d4xg  LM(gμν,ψ,ψ)(3.4)S_M = \int d^4x\, \sqrt{-g}\;\mathcal{L}_M(g_{\mu\nu}, \psi, \nabla\psi) \tag{3.4}

No term coupling the curvature directly to matter, such as Rϕ2R\phi^2, is included. P1–P6 do not strictly forbid such a term; the equivalence principle excludes it only in the limit in which the radius of curvature is large compared with the scale of the experiment. This is an assumption, recorded in Sec. 8.

The total action is

S=SEH+SM(3.5)S = S_{EH} + S_M \tag{3.5}

and by P5 the field equations are δS/δgμν=0\delta S / \delta g^{\mu\nu} = 0.

4. Variation of the action

The variation is taken with respect to the inverse metric gμνg^{\mu\nu}; variation with respect to gμνg_{\mu\nu} yields the same equations with an overall change of sign.

4.1 Variation of the inverse metric

Varying gμαgαν=δνμg^{\mu\alpha}g_{\alpha\nu} = \delta^\mu_\nu, whose right-hand side is constant,

δgμαgαν+gμαδgαν=0δgμν=gμαgνβδgαβ(4.1)\delta g^{\mu\alpha}\,g_{\alpha\nu} + g^{\mu\alpha}\,\delta g_{\alpha\nu} = 0 \qquad\Longrightarrow\qquad \delta g^{\mu\nu} = -g^{\mu\alpha}g^{\nu\beta}\,\delta g_{\alpha\beta} \tag{4.1}

Contracting (4.1) with the metric gives the relation

gμνδgμν=gμνδgμν(4.2)g_{\mu\nu}\,\delta g^{\mu\nu} = -g^{\mu\nu}\,\delta g_{\mu\nu} \tag{4.2}

used at several points below.

4.2 Variation of the volume element

For an invertible matrix MM the identity lndetM=trlnM\ln\det M = \mathrm{tr}\ln M gives on differentiation δdetM=detM  tr(M1δM)\delta \det M = \det M\;\mathrm{tr}(M^{-1}\delta M), known as Jacobi's formula. Applied to the metric,

δg=ggμνδgμν(4.3)\delta g = g\,g^{\mu\nu}\,\delta g_{\mu\nu} \tag{4.3}

whence

δg=δg2g=ggμνδgμν2g=12g  gμνδgμν(4.4)\delta\sqrt{-g} = \frac{-\,\delta g}{2\sqrt{-g}} = \frac{-g\,g^{\mu\nu}\delta g_{\mu\nu}}{2\sqrt{-g}} = \frac{1}{2}\sqrt{-g}\;g^{\mu\nu}\,\delta g_{\mu\nu} \tag{4.4}

and, on substituting (4.2),

δg=12g  gμνδgμν(4.5)\delta \sqrt{-g} = -\frac{1}{2}\sqrt{-g}\;g_{\mu\nu}\,\delta g^{\mu\nu} \tag{4.5}

The term 12gμνR-\tfrac{1}{2}g_{\mu\nu}R in the Einstein tensor originates entirely in (4.5).

4.3 Variation of the connection

Although Γμνλ\Gamma^\lambda_{\mu\nu} is not a tensor, the difference of two connections is, the inhomogeneous parts of their transformation laws being identical. The variation δΓμνλ\delta\Gamma^\lambda_{\mu\nu} is therefore a tensor and admits a covariant expression. Evaluating it in normal coordinates at a point, where Γ=0\Gamma = 0 and covariant derivatives reduce to partial derivatives,

δΓμνλ=12gλρ(μδgνρ+νδgμρρδgμν)(4.6)\delta\Gamma^\lambda_{\mu\nu} = \frac{1}{2}g^{\lambda\rho}\left(\nabla_\mu \delta g_{\nu\rho} + \nabla_\nu \delta g_{\mu\rho} - \nabla_\rho \delta g_{\mu\nu}\right) \tag{4.6}

Both sides of (4.6) are tensors agreeing at an arbitrary point in one frame, and therefore agree everywhere in every frame.

4.4 The Palatini identity

Varying (2.6) in normal coordinates, where the quadratic terms do not contribute at the point, gives δR σμνρ=μδΓνσρνδΓμσρ\delta R^\rho_{\ \sigma\mu\nu} = \partial_\mu \delta\Gamma^\rho_{\nu\sigma} - \partial_\nu \delta\Gamma^\rho_{\mu\sigma}. Every object in this relation is now a tensor, so it may be written covariantly:

δR σμνρ=μδΓνσρνδΓμσρ(4.7)\delta R^\rho_{\ \sigma\mu\nu} = \nabla_\mu \delta\Gamma^\rho_{\nu\sigma} - \nabla_\nu \delta\Gamma^\rho_{\mu\sigma} \tag{4.7}

Contracting ρ\rho with μ\mu,

δRμν=λδΓμνλνδΓλμλ(4.8)\delta R_{\mu\nu} = \nabla_\lambda \delta\Gamma^\lambda_{\mu\nu} - \nabla_\nu \delta\Gamma^\lambda_{\lambda\mu} \tag{4.8}

Equation (4.8) expresses the variation of the Ricci tensor, which contains second derivatives of δgμν\delta g_{\mu\nu}, as a difference of covariant divergences.

4.5 The boundary term

From (2.9),

δR=Rμνδgμν+gμνδRμν(4.9)\delta R = R_{\mu\nu}\,\delta g^{\mu\nu} + g^{\mu\nu}\,\delta R_{\mu\nu} \tag{4.9}

In the second term gμνg^{\mu\nu} may be taken through the covariant derivatives in (4.8), by Lemma 2.1, giving

gμνδRμν=λvλ,vλ=gμνδΓμνλgλνδΓμνμ(4.10)g^{\mu\nu}\,\delta R_{\mu\nu} = \nabla_\lambda v^\lambda, \qquad v^\lambda = g^{\mu\nu}\delta\Gamma^\lambda_{\mu\nu} - g^{\lambda\nu}\delta\Gamma^\mu_{\mu\nu} \tag{4.10}

or equivalently vλ=gμνλδgμνμδgλμv^\lambda = g_{\mu\nu}\nabla^\lambda \delta g^{\mu\nu} - \nabla_\mu \delta g^{\lambda\mu}. For any vector field, gλvλ=λ(gvλ)\sqrt{-g}\,\nabla_\lambda v^\lambda = \partial_\lambda(\sqrt{-g}\,v^\lambda), so by the divergence theorem the contribution of (4.10) to the variation is a surface integral,

Md4x  g  λvλ=MdΣλ  vλ(4.11)\int_M d^4x\;\sqrt{-g}\;\nabla_\lambda v^\lambda = \oint_{\partial M} d\Sigma_\lambda\; v^\lambda \tag{4.11}

The vanishing of (4.11) requires both δgμν\delta g^{\mu\nu} and its normal derivative to vanish on M\partial M, which over-determines the variational problem. The standard remedy is to add to (3.3) the Gibbons-Hawking-York surface term

SGHY=1κMd3x  h  K(4.12)S_{GHY} = \frac{1}{\kappa}\oint_{\partial M} d^3x\; \sqrt{|h|}\;K \tag{4.12}

in which hh is the induced metric on M\partial M and KK the trace of its extrinsic curvature; the variation of (4.12) cancels the normal-derivative contribution exactly.2 For variations of compact support the surface term is absent, and it is omitted in what follows.

4.6 The gravitational contribution

Combining (4.5), (4.9) and (4.10), and varying the cosmological term with (4.5),

δSEH=12κd4x  g(Rμν12gμνR+Λgμν)δgμν(4.13)\delta S_{EH} = \frac{1}{2\kappa}\int d^4x\;\sqrt{-g}\left(R_{\mu\nu} - \frac{1}{2}g_{\mu\nu}R + \Lambda g_{\mu\nu}\right)\delta g^{\mu\nu} \tag{4.13}

The cosmological term enters with +gμν+g_{\mu\nu} because δ ⁣(g(2Λ))=+Λggμνδgμν\delta\!\left(\sqrt{-g}\,(-2\Lambda)\right) = +\Lambda\sqrt{-g}\,g_{\mu\nu}\,\delta g^{\mu\nu}.

4.7 The energy-momentum tensor

The energy-momentum tensor is defined as the response of the matter action to a variation of the metric,

Tμν2g  δ ⁣(gLM)δgμνδSM=12d4xg  Tμνδgμν(4.14)T_{\mu\nu} \equiv -\frac{2}{\sqrt{-g}}\;\frac{\delta\!\left(\sqrt{-g}\,\mathcal{L}_M\right)}{\delta g^{\mu\nu}} \qquad\Longleftrightarrow\qquad \delta S_M = -\frac{1}{2}\int d^4x\, \sqrt{-g}\;T_{\mu\nu}\,\delta g^{\mu\nu} \tag{4.14}

and is symmetric by construction. Definition (4.14) is justified by its agreement with the expressions obtained in special relativity. As an illustration, consider a real scalar field with

LM=12gμνμϕνϕV(ϕ)(4.15)\mathcal{L}_M = -\frac{1}{2}g^{\mu\nu}\partial_\mu\phi\,\partial_\nu\phi - V(\phi) \tag{4.15}

whose only metric dependence is the explicit inverse metric and the volume element. Then

δ ⁣(gLM)=g(12μϕνϕ)δgμν12g  gμνLMδgμν(4.16)\delta\!\left(\sqrt{-g}\,\mathcal{L}_M\right) = \sqrt{-g}\left(-\frac{1}{2}\partial_\mu\phi\,\partial_\nu\phi\right)\delta g^{\mu\nu} - \frac{1}{2}\sqrt{-g}\;g_{\mu\nu}\mathcal{L}_M\,\delta g^{\mu\nu} \tag{4.16}

so that Tμν=μϕνϕ+gμνLMT_{\mu\nu} = \partial_\mu\phi\,\partial_\nu\phi + g_{\mu\nu}\mathcal{L}_M. In flat space with c=1c = 1 its time-time component is ϕ˙2(12ϕ˙2+V)=12ϕ˙2+V\dot\phi^2 - \left(-\tfrac{1}{2}\dot\phi^2 + V\right) = \tfrac{1}{2}\dot\phi^2 + V, the sum of the kinetic and potential energy densities, in agreement with the special-relativistic result.

4.8 The field equations

Requiring δS=δSEH+δSM=0\delta S = \delta S_{EH} + \delta S_M = 0 for arbitrary δgμν\delta g^{\mu\nu} annihilates the integrand of (4.13) and (4.14) pointwise,

12κ(Rμν12gμνR+Λgμν)12Tμν=0(4.17)\frac{1}{2\kappa}\left(R_{\mu\nu} - \frac{1}{2}g_{\mu\nu}R + \Lambda g_{\mu\nu}\right) - \frac{1}{2}T_{\mu\nu} = 0 \tag{4.17}

that is,

Rμν12gμνR+Λgμν=κTμν(4.18)R_{\mu\nu} - \frac{1}{2}g_{\mu\nu}R + \Lambda g_{\mu\nu} = \kappa\,T_{\mu\nu} \tag{4.18}

These are ten coupled nonlinear partial differential equations of second order for the ten independent components of gμνg_{\mu\nu}.

5. The Newtonian limit and the value of κ\kappa

The constant κ\kappa is determined by requiring that (4.18) reproduce Newtonian gravitation in the regime in which the latter is valid. That regime is defined by three approximations: the field is weak, gμν=ημν+hμνg_{\mu\nu} = \eta_{\mu\nu} + h_{\mu\nu} with hμν1|h_{\mu\nu}| \ll 1 and all terms beyond first order in hμνh_{\mu\nu} neglected; the field is static, 0hμν=0\partial_0 h_{\mu\nu} = 0; and the sources move slowly, dxi/dtc|dx^i/dt| \ll c. The cosmological term is negligible at the relevant scales and is set to zero.

Step 1. Identification of the potential. For a slowly moving particle dx0/dτcdx^0/d\tau \approx c while the spatial components of dxμ/dτdx^\mu/d\tau are negligible, so that (2.2) reduces to

d2xidt2c2Γ00i(5.1)\frac{d^2x^i}{dt^2} \approx -c^2\,\Gamma^i_{00} \tag{5.1}

For a static field, (2.3) gives to first order

Γ00i=12giλ(20gλ0λg00)=12δijjh00(5.2)\Gamma^i_{00} = \frac{1}{2}g^{i\lambda}\left(2\partial_0 g_{\lambda 0} - \partial_\lambda g_{00}\right) = -\frac{1}{2}\delta^{ij}\partial_j h_{00} \tag{5.2}

Comparison of (5.1) and (5.2) with the Newtonian equation of motion d2xi/dt2=iΦd^2x^i/dt^2 = -\partial_i\Phi yields

h00=2Φc2,that isg00(1+2Φc2)(5.3)h_{00} = -\frac{2\Phi}{c^2}, \qquad\text{that is}\qquad g_{00} \approx -\left(1 + \frac{2\Phi}{c^2}\right) \tag{5.3}

The Newtonian potential is thus a component of the metric; gravitational time dilation is the same statement applied to (1.2).

Step 2. Linearized curvature. To first order in hμνh_{\mu\nu} the quadratic terms of (2.6) are of second order and may be dropped, so that RμνλΓμνλνΓλμλR_{\mu\nu} \approx \partial_\lambda\Gamma^\lambda_{\mu\nu} - \partial_\nu\Gamma^\lambda_{\lambda\mu}. For a static field the second term contributes nothing to the time-time component, and with (5.2),

R00iΓ00i=122h00=1c22Φ(5.4)R_{00} \approx \partial_i \Gamma^i_{00} = -\frac{1}{2}\nabla^2 h_{00} = \frac{1}{c^2}\nabla^2\Phi \tag{5.4}

Step 3. The source. For pressureless dust of rest-mass density ρ\rho at rest, Tμν=ρuμuνT_{\mu\nu} = \rho\,u_\mu u_\nu with uμ(c,0,0,0)u^\mu \approx (c,0,0,0), giving T00ρc2T_{00} \approx \rho c^2 and T=gμνTμνρc2T = g^{\mu\nu}T_{\mu\nu} \approx -\rho c^2.

Step 4. Comparison. The time-time component of the trace-reversed equation (6.1) below, with Λ=0\Lambda = 0, reads

1c22Φ=κ(ρc212(1)(ρc2))=12κρc2(5.5)\frac{1}{c^2}\nabla^2 \Phi = \kappa\left(\rho c^2 - \frac{1}{2}(-1)(-\rho c^2)\right) = \frac{1}{2}\kappa \rho c^2 \tag{5.5}

so that 2Φ=12κρc4\nabla^2\Phi = \tfrac{1}{2}\kappa\rho c^4. Comparison with Poisson's equation 2Φ=4πGρ\nabla^2 \Phi = 4\pi G\rho gives

κ=8πGc4(5.6)\kappa = \frac{8\pi G}{c^4} \tag{5.6}

The factor 8π8\pi is a consequence of the factor 4π4\pi in Poisson's equation, which in turn follows from the normalization of Newton's law of gravitation. With (5.6), the field equations (4.18) take their standard form

Rμν12gμνR+Λgμν=8πGc4Tμν(5.7)R_{\mu\nu} - \frac{1}{2}g_{\mu\nu}R + \Lambda g_{\mu\nu} = \frac{8\pi G}{c^4}\,T_{\mu\nu} \tag{5.7}

6. Equivalent forms

6.1 Trace-reversed form

Contracting (5.7) with gμνg^{\mu\nu} and using gμνgμν=δμμ=4g^{\mu\nu}g_{\mu\nu} = \delta^\mu_\mu = 4 gives R+4Λ=κT-R + 4\Lambda = \kappa T, whence R=4ΛκTR = 4\Lambda - \kappa T. Substituting back,

Rμν=κ(Tμν12gμνT)+Λgμν(6.1)R_{\mu\nu} = \kappa\left(T_{\mu\nu} - \frac{1}{2}g_{\mu\nu}T\right) + \Lambda g_{\mu\nu} \tag{6.1}

For Λ=0\Lambda = 0 in the absence of matter this reduces to the vacuum field equations

Rμν=0(6.2)R_{\mu\nu} = 0 \tag{6.2}

A solution of (6.2) is Ricci-flat but need not be flat. Of the twenty independent components of the Riemann tensor, (6.2) constrains only the ten contained in the Ricci tensor; the remaining ten, comprising the Weyl tensor, are unconstrained and propagate. Gravitational radiation in vacuum is a solution of (6.2).

6.2 The cosmological term as vacuum energy

Transposing the cosmological term in (5.7) to the right-hand side,

Gμν=8πGc4(TμνΛc48πGgμν)(6.3)G_{\mu\nu} = \frac{8\pi G}{c^4}\left(T_{\mu\nu} - \frac{\Lambda c^4}{8\pi G}\,g_{\mu\nu}\right) \tag{6.3}

which is the stress tensor of a perfect fluid of energy density ρΛc2=Λc4/8πG\rho_\Lambda c^2 = \Lambda c^4/8\pi G and pressure p=ρΛc2p = -\rho_\Lambda c^2. The field equations do not distinguish between a geometrical constant and a vacuum energy density. The discrepancy between the value of Λ\Lambda inferred from cosmological observation and the vacuum energy estimated from quantum field theory, which exceeds it by many tens of orders of magnitude, is the cosmological constant problem and remains unresolved.

7. Consistency conditions

7.1 Conservation of energy and momentum

Taking the divergence of (5.7) and applying the contracted Bianchi identity (2.12), which annihilates the left-hand side identically, gives

μTμν=0(7.1)\nabla^\mu T_{\mu\nu} = 0 \tag{7.1}

Local conservation of energy and momentum is therefore not an independent postulate but a consequence of the field equations; a matter model violating (7.1) is inconsistent with them.

The same conclusion follows from P3 alone, independently of the field equations. Under an infinitesimal diffeomorphism generated by ξμ\xi^\mu the metric changes by its Lie derivative,

δgμν=(μξν+νξμ)(7.2)\delta g^{\mu\nu} = -\left(\nabla^\mu\xi^\nu + \nabla^\nu\xi^\mu\right) \tag{7.2}

General covariance requires SMS_M to be invariant, so that for matter fields satisfying their own equations of motion, using (4.14), (7.2) and the symmetry of TμνT_{\mu\nu}, and integrating by parts,

0=δSM=d4xg  Tμνμξν=d4xg  (μTμν)ξν(7.3)0 = \delta S_M = \int d^4x\,\sqrt{-g}\;T_{\mu\nu}\nabla^\mu\xi^\nu = -\int d^4x\, \sqrt{-g}\;\left(\nabla^\mu T_{\mu\nu}\right)\xi^\nu \tag{7.3}

Since ξν\xi^\nu is arbitrary, (7.1) follows. Equation (7.1) is thus the Noether identity associated with general covariance, of which the contracted Bianchi identity is the gravitational counterpart.

7.2 Counting of degrees of freedom

Equations (5.7) number ten, as do the components of the metric, but the four contracted Bianchi identities imply that four combinations of them contain no second time derivatives. These four are constraints on the initial data rather than evolution equations. Independently, P3 implies that any solution may be transformed by a diffeomorphism, corresponding to four arbitrary functions of gauge. Subtracting both from the ten components leaves two propagating degrees of freedom, corresponding to the two polarization states of a gravitational wave.

8. Summary of assumptions

The following table lists every assumption employed above, the point at which it enters, and the theories obtained by relaxing it. The last column is expanded in Sec. 9.

AssumptionEnters atRelaxed by
Four dimensions(2.7); Sec. 3.2, item 4; Sec. 6.1Kaluza-Klein theory, string theory, Lovelock gravity in D>4D \gt 4
Global Lorentzian metricP2Euclidean quantum gravity, causal set theory
Smooth manifold structureP1Loop quantum gravity, causal dynamical triangulations
Metric compatibilityLemma 2.1; (4.10)Metric-affine and Weyl geometries, symmetric teleparallel gravity
Vanishing torsionSec. 2.2; (2.5)Einstein-Cartan theory, Poincaré gauge theory
Metric is the only gravitational fieldP5Brans-Dicke, Horndeski, TeVeS, massive gravity
Second-order field equationsP6; Sec. 3.2, item 4Quadratic gravity, f(R)f(R) gravity, Hořava-Lifshitz gravity
Action principle and localityP5, P6Non-local gravity, thermodynamic derivations
Minimal couplingSec. 3.4Non-minimally coupled scalar fields, scalar-tensor theories
Vanishing boundary terms(4.11)Manifolds with boundary; black hole thermodynamics
Λ\Lambda constant(3.3), (4.13)Unimodular gravity, quintessence
Classical field theoryThroughoutCandidate quantum theories of gravity

9. Theories relaxing the postulates

9.1 Quadratic gravity

Adding αR2+βRμνRμν\alpha R^2 + \beta R_{\mu\nu}R^{\mu\nu} to (3.3) relaxes P6. Stelle showed that the resulting theory is renormalizable in perturbation theory, which general relativity is not. The fourth-order field equations, however, propagate a massive spin-2 mode of negative kinetic energy, so that unitarity is not manifest; whether this ghost is physical or an artifact of the perturbative treatment remains a matter of debate. Independently of that question, quadratic terms are unavoidable as corrections: treated as an effective field theory, general relativity generates them at one loop with coefficients suppressed by the Planck scale, where they are innocuous.

9.2 f(R)f(R) gravity

Replacing RR in (3.3) by a function f(R)f(R) likewise yields fourth-order equations, but the additional propagating mode is a scalar of positive kinetic energy rather than a ghost, and the theory is dynamically equivalent to a scalar-tensor theory. It has been studied extensively as a means of producing accelerated cosmological expansion geometrically rather than by a dark energy component, and is strongly constrained by solar-system tests.

9.3 Lovelock and Gauss-Bonnet gravity

Retaining second-order field equations but relaxing the restriction to four dimensions, the Gauss-Bonnet term of Sec. 3.2 ceases to be topological and contributes to the equations of motion, and Theorem 3.1 admits an infinite series of higher-order invariants. The four-dimensional case is thus a degenerate member of a larger family.

9.4 Einstein-Cartan theory

Relaxing the vanishing of torsion while retaining P1–P6 gives a theory in which intrinsic spin acts as a source for torsion. The torsion field is non-propagating and algebraically determined by the spin density, so the theory coincides with general relativity except at densities at which spin alignment is appreciable.

9.5 Scalar-tensor and Horndeski theories

Relaxing the requirement in P5 that the metric be the only gravitational field, one introduces a scalar ϕ\phi coupled to the curvature. The Brans-Dicke theory, motivated by Mach's principle, is the original example, and the Horndeski construction gives the most general scalar-tensor theory with second-order field equations. The near-simultaneous arrival of gravitational waves and light from the neutron star merger GW170817, after a propagation time of order 10810^8 years, constrains the speed of gravitational waves to equal cc to high precision and excludes those members of the family predicting otherwise.

9.6 Massive gravity

Assigning a mass to the graviton breaks diffeomorphism invariance, or requires its restoration by the introduction of Stückelberg fields. Generic mass terms propagate a sixth mode, the Boulware-Deser ghost; the construction of de Rham, Gabadadze and Tolley eliminates it by a particular tuning of the potential.

9.7 Unimodular gravity

Restricting the variation of the action to those variations preserving g\sqrt{-g} removes the trace of the field equations from the variational principle. The cosmological constant then arises as a constant of integration rather than as a coupling in the action. The classical predictions are otherwise identical to those of general relativity, so the theory reformulates the cosmological constant problem without resolving it.

9.8 Thermodynamic derivations

Jacobson showed that imposing the Clausius relation δQ=TdS\delta Q = T\,dS on every local Rindler horizon, with entropy proportional to horizon area, yields (5.7) as an equation of state. This construction does not relax a single postulate but reinterprets the entire framework, suggesting that the metric is a coarse-grained thermodynamic variable and that the quantization of gμνg_{\mu\nu} is analogous to the quantization of sound waves in a solid.

10. Further development

The derivation above yields the field equations and nothing further. Their content is obtained by solving them in particular cases.

The natural first application is the vacuum equation (6.2) under the assumption of spherical symmetry and staticity. The metric then contains two unknown functions of the radial coordinate, the field equations reduce to ordinary differential equations, and integration yields the Schwarzschild solution together with Birkhoff's theorem, which states that it is the unique spherically symmetric vacuum solution. The classical tests, namely the precession of the perihelion of Mercury and the deflection of light by the Sun, are then perturbative calculations on that background, and the horizon at r=2GM/c2r = 2GM/c^2 appears as a property of the solution.

References

  • Misner, C. W., Thorne, K. S., and Wheeler, J. A., Gravitation, Freeman (1973). Sign conventions and the variational derivation.
  • Wald, R. M., General Relativity, University of Chicago Press (1984). Appendix E for the boundary term.
  • Weinberg, S., Gravitation and Cosmology, Wiley (1972). The Newtonian limit in detail.
  • Carroll, S. M., Spacetime and Geometry, Addison-Wesley (2004). Chapters 3 and 4.
  • Landau, L. D., and Lifshitz, E. M., The Classical Theory of Fields, 4th ed., Pergamon (1975).
  • Lovelock, D., "The Einstein tensor and its generalizations," J. Math. Phys. 12, 498 (1971). Theorem 3.1.
  • York, J. W., Phys. Rev. Lett. 28, 1082 (1972); Gibbons, G. W., and Hawking, S. W., Phys. Rev. D 15, 2752 (1977). The surface term (4.12).
  • Stelle, K. S., "Renormalization of higher-derivative quantum gravity," Phys. Rev. D 16, 953 (1977). Sec. 9.1.
  • Donoghue, J. F., "General relativity as an effective field theory," Phys. Rev. D 50, 3874 (1994). Sec. 9.1.
  • Sotiriou, T. P., and Faraoni, V., "f(R)f(R) theories of gravity," Rev. Mod. Phys. 82, 451 (2010). Sec. 9.2.
  • Hehl, F. W., von der Heyde, P., Kerlick, G. D., and Nester, J. M., "General relativity with spin and torsion," Rev. Mod. Phys. 48, 393 (1976). Sec. 9.4.
  • Brans, C., and Dicke, R. H., Phys. Rev. 124, 925 (1961); Horndeski, G. W., Int. J. Theor. Phys. 10, 363 (1974). Sec. 9.5.
  • Boulware, D. G., and Deser, S., Phys. Rev. D 6, 3368 (1972); de Rham, C., Gabadadze, G., and Tolley, A. J., Phys. Rev. Lett. 106, 231101 (2011). Sec. 9.6.
  • Jacobson, T., "Thermodynamics of spacetime: the Einstein equation of state," Phys. Rev. Lett. 75, 1260 (1995). Sec. 9.8.

Footnotes

  1. The instability is due to Ostrogradsky (1850). A Lagrangian depending non-degenerately on derivatives higher than the first yields a Hamiltonian linear in one of the canonical momenta, and therefore unbounded below.

  2. The surface term contributes the whole of the Bekenstein-Hawking entropy in the Euclidean evaluation of the black hole partition function, and cannot be neglected in that context.