Chapter 4

The Form of the Evolution Equation

With the diffraction of electrons we saw that a wave is in some way associated with the electron. If there is a wave, there must be an equation that describes it: this is the question Schrödinger asked himself. That equation describes the time evolution of the system, that is, how the state changes in time. In this card we shall look for the general form of the evolution equation.

Probabilistic interpretation.

A wave associated with a particle can be imagined in more than one way. Here we assume that the wave gives the probability of finding the particle at a point in space. We then describe the wave with a complex number at each point of space. In other areas of physics complex numbers are a computational convenience for dealing with waves, and it is therefore natural to use them for our own purpose as well, which is to find the equation of time evolution of what will be a wave.

According to this probabilistic interpretation we postulate that a material particle, from the point of view of Quantum Mechanics, is a physical system on which position measurements can be made. When we perform such a measurement, the result is in general not certain, but random. So we have a probability distribution p(x)p(x) for the variable x. We postulate the principle of the complex function: the probabilities are obtained as the squared moduli of certain complex numbers, called probability amplitudes. So we have a certain complex function ψ(x)\psi(x) and the probability is given by p(x)=ψ(x)2p(x)=|\psi(x)|^2.

The function ψ(x)\psi(x) is a distribution of probability amplitudes and characterises the state of the particle; this function can be seen as a vector of infinitely many complex numbers, and we will represent it with the ket-vector symbol ψψ(x)|\psi\rangle\equiv\psi(x).

When the state of the particle evolves, the function ψ changes with time, so we obtain a function that depends on space and time ψ(x,t)\psi(x,t), which we can represent with a time-dependent ket ψt|\psi t\rangle.

We postulate the principle of linear superposition: two states can be combined, each with a complex coefficient, and the combination is still a state; the state evolves continuously in time and, under evolution, each term evolves as it would on its own, with the same coefficient, and the subsequent state is the sum of the evolved terms. It follows that the time evolution is described by a linear law ψt=U(t0t)ψt0|\psi t\rangle=U\left(t_0\to t\right)|\psi t_0\rangle. Here ψt0|\psi t_0\rangle is the vector associated with the function ψ(x,t0)\psi(x,t_0) representing the initial state; ψt|\psi t\rangle is the vector associated with the function ψ(x,t)\psi(x,t), and U(t0t)U\left(t_0\to t\right) is the time-evolution matrix.

At this point the problem arises of determining the time-evolution matrix U(t0t)U\left(t_0\to t\right). We will solve this problem later; for now we want to anticipate the following qualitative result: the final equation we will obtain will closely resemble the wave equation, and the solutions ψ(x,t)\psi(x,t) will have the appearance of wave packets. So the wave that accompanies a material particle, of which we spoke in the previous section, is nothing other than the vector of probability amplitudes ψ(x,t)\psi(x,t) associated with the position variable.

For clarity we recap the important points we have introduced:

A material particle is characterised by the position variable, that is, by a coordinate which we will briefly denote by the symbol x.

The state of a material particle at an instant t is represented by the distribution of probability amplitudes ψ(x,t)\psi(x,t) of the variable x, which we denote by the ket symbol ψt|\psi t\rangle.

The evolution of the state is described by the following equation:

ψt=U(t0t)ψt0|\psi t\rangle=U(t_0\to t)|\psi t_0\rangle

where U(t0t)U\left(t_0\to t\right) is a matrix of ∞×∞ components.

Time-evolution equation in differential form.

The time-evolution equation we have written allows a finite jump between the instants t0 and t. However, it is more convenient to consider an infinitesimal time jump tt+dtt\to t+dt. In this case we have the equation

ψ(t+dt)=U(tt+dt)ψt|\psi(t+dt)\rangle=U(t\to t+dt)|\psi t\rangle

subtracting ψt|\psi t\rangle from both sides and dividing by dt we obtain

ψ(t+dt)ψtdt=U(tt+dt)ψtψtdtddtψt=U(tt+dt)1dtψt\begin{aligned} \frac{|\psi(t+dt)\rangle-|\psi t\rangle}{dt} & =\frac{U(t\to t+dt)|\psi t\rangle-|\psi t\rangle}{dt}\Leftrightarrow \\ \frac{d}{dt}|\psi t\rangle & =\frac{U(t\to t+dt)-1}{dt}|\psi t\rangle \end{aligned}

To simplify the right-hand side we can write the matrix U(t0t)U\left(t_0\to t\right) with a first-order approximation

U(tt+dt)=U(tt)Term for dt=0+A(t)dtFirst-order increment+o(dt2)Higher-order infinitesimalU(t\to t+dt)=\underbrace{U(t\to t)}_{\text{Term for }dt=0}+\underbrace{A(t)\cdot dt}_{\text{First-order increment}}+\underbrace{o(dt^2)}_{\text{Higher-order infinitesimal}}

where A(t)A(t) is the derivative

limdt0U(tt+dt)U(tt)dt\lim_{dt\to0}\frac{U(t\to t+dt)-U(t\to t)}{dt}

Substituting this first-order approximation we have

U(tt+dt)1dt=U(tt)+A(t)dt+o(dt2)1dtobserving that U(tt)=1 we have=A(t)dt+o(dt2)dt=A(t)\begin{aligned} \frac{U(t\to t+dt)-1}{dt} & =\frac{U(t\to t)+A(t)dt+o(dt^2)-1}{dt} \\ & \quad\text{observing that }U(t\to t)=1\text{ we have} \\ & =\frac{A(t)dt+o(dt^2)}{dt} \\ & =A(t) \end{aligned}

So in the end we obtain the equation

ddtψt=A(t)ψt\frac{d}{dt}|\psi t\rangle=A(t)|\psi t\rangle

The problem of determining U(t0t)U\left(t_0\to t\right) has become the problem of determining the derivative A(t)A(t).

Before proceeding, we must dwell on some mathematical topics.

Algebra of operators.

Vectors and matrices of infinite dimension

In this card we have introduced vectors and matrices of infinite dimension. In finite-dimensional cases a vector is represented by an n-tuple of components v=(v1vn)|v\rangle=\left(v_1\dots v_n\right); in infinite-dimensional cases, instead, we can represent a vector by a function v=v(x)|v\rangle=v(x), where the variable x takes the role of the indices. Analogously, an infinite-dimensional matrix is represented by a function of two variables A=A(x,x)A=A(x,x').

The product of a matrix and a vector, for example, can be written by means of an integral

u=Avu(x)=+A(x,x)v(x)dx\begin{aligned} |u\rangle & =A|v\rangle\Leftrightarrow \\ u(x) & =\int_{-\infty}^{+\infty}A(x,x')v(x')\:dx' \end{aligned}

In an analogous way one can write other types of products between matrices or between vectors.

Adjoint matrix and Hermitian, anti-Hermitian and unitary matrices.

Suppose we have a ket vector u|u\rangle given by the product of a matrix A and another vector v|v\rangle

u=Av|u\rangle=A|v\rangle

and suppose we have to determine the conjugate bra u\langle u|. Consider for example the two-dimensional case

u=(u1u2)u=(u1,u2)=(u1u2)t=ut\begin{aligned} |u\rangle & =\left(\begin{gathered} u_1 \\ u_2 \end{gathered}\right)\Rightarrow \\ \langle u| & =(u_1^*,u_2^*) \\ & ={\left(\begin{gathered} u_1 \\ u_2 \end{gathered}\right)}^{*t} \\ & ={|u\rangle}^{*t} \end{aligned}

So the bra u\langle u| is obtained by conjugating ()({}^*) and transposing (t)({}^t) the ket u|u\rangle

u=ut=(Av)t=vtAt=vAt\begin{aligned} \langle u| & =|u\rangle^{*t} \\ & ={\left(A|v\rangle\right)}^{*t} \\ & =|v\rangle^{*t}A^{*t} \\ & =\langle v|A^{*t} \end{aligned}

In the end we can write

u=Avu=vAt|u\rangle=A|v\rangle\Leftrightarrow\langle u|=\langle v|A^{*t}

where the matrix AtA^{*t} is obtained by conjugating and transposing the matrix A.

By definition we will say that the matrix AtA^{*t} is the adjoint of the matrix A, and we will denote it with the symbol A+A^+

AtA+A^{*t}\equiv A^+

For example, for a two-dimensional matrix

A=(a11a12a21a22)A+=At=(a11a21a12a22)\begin{aligned} A & =\begin{pmatrix} a_{11} & a_{12} \\ a_{21} & a_{22} \end{pmatrix}\Rightarrow \\ A^+ & =A^{*t} \\ & =\begin{pmatrix} a_{11}^* & a_{21}^* \\ a_{12}^* & a_{22}^* \end{pmatrix} \end{aligned}

On the basis of this definition we can write that the conjugate bra of the ket AuA|u\rangle is uA+\langle u|A^+.

By definition, if a matrix is equal to its adjoint A=A+A=A^+, then it is said to be Hermitian; if instead the matrix is equal to its adjoint with the sign changed A=A+A=-A^+, then it is said to be anti-Hermitian. A matrix such that AA+=A+A=IAA^+=A^+A=I, where I is the identity matrix, is said to be unitary.

The operation of taking the adjoint matrix, in the field of matrices, takes the role of the operation of taking the complex conjugate in the field of complex numbers. So Hermitian matrices take the role of purely real numbers, while anti-Hermitian matrices take the role of purely imaginary numbers. Unitary matrices, finally, take the role of numbers of unit modulus.

Hermitian, anti-Hermitian and unitary matrices have very interesting properties that we will study later, and they are of considerable importance in Quantum Mechanics.

Relation between matrices and operators.

In the case of finite-dimensional vectors we know, from the study of linear algebra, that any linear operator vR  w|v\rangle\overset{R\;}{\to}|w\rangle can be written as a matrix product w=Rv|w\rangle=R|v\rangle, where R is a particular matrix associated with the operator in question.

This property remains valid also in the case of infinite-dimensional vectors, but it is much more delicate from the mathematical point of view.

Consider for example the derivative operator, which we will denote by D. The operator D transforms a vector v|v\rangle associated with the function v(x)v(x) into the vector DvD|v\rangle associated with the function dv(x)/dxdv(x)/dx

v(x) D ddxv(x)v(x)\xrightarrow{\ D\ }\frac{d}{dx}v(x)

We ask: what is the matrix, that is, the function of two variables, that represents the derivative operator? The function we seek is the derivative of the Dirac δ

ddxδ(xx)-\frac{d}{dx'}\delta(x'-x)

indeed, performing the matrix product and integrating by parts, we have

+dδ(xx)dxv(x)dx=δ(xx)v(x)+++δ(xx)dv(x)dxdx\int_{-\infty}^{+\infty}-\frac{d\delta(x'-x)}{dx'}v(x')\:dx'=-\delta(x'-x)v(x')\Big|_{-\infty}^{+\infty}+\int_{-\infty}^{+\infty}\delta(x'-x)\frac{dv(x')}{dx'}\:dx'

The boundary term vanishes because δ vanishes at infinity; in the remaining integral the δ selects the value of the derivative at the point x:

+δ(xx)dv(x)dxdx=dv(x)dx\int_{-\infty}^{+\infty}\delta(x'-x)\frac{dv(x')}{dx'}\:dx'=\frac{dv(x)}{dx}

The operator Ddδ(xx)/dxD\leftrightarrow-d\delta(x'-x)/dx' is anti-Hermitian, indeed

(dδ(xx)dx)t=(dδ(xx)dx)=dδ(xx)dx=dδ(xx)dx\begin{aligned} {\left(-\frac{d\delta(x'-x)}{dx'}\right)}^{*t} & ={\left(-\frac{d\delta(x'-x)}{dx'}\right)}^\dagger \\ & =-\frac{d\delta(x-x')}{dx} \\ & =\frac{d\delta(x'-x)}{dx'} \end{aligned}

that is, by conjugating and transposing the matrix associated with D one obtains a matrix of opposite sign D+=DD^+=-D.

In general one prefers to use the operator K=iDK=iD, which is Hermitian, indeed

K+=(iD)+=iD+=(i)(D)=iD=KK^+=(iD)^+=i^*D^+=(-i)(-D)=iD=K

From now on we will speak indifferently of operators or of matrices. A derivative operator will in general be represented by the derivative operation rather than by the associated matrix; in any case it is important to know that to any linear operator there is always associated a definite matrix, whether finite- or infinite-dimensional.

Functions of operators, or functions of matrices.

If we take an operator A and apply it twice we obtain the operator AA=A2AA=A^2; in the same way we can obtain A3A^3, A4A^4, etc.

If we take the inverse operator A1A^{-1} and apply it twice we define the operator A1A1=A2A^{-1}A^{-1}=A^{-2}; in the same way we obtain A3A^{-3}, A4A^{-4}, etc.

If we consider a polynomial

p(x)=cnxn++c1x+c0+c1x1++cmxmp(x)=c_nx^n+\dots+c_1x+c_0+c_{-1}x^{-1}+\dots+c_{-m}x^{-m}

from it we can build the operator

p(A)=cnAn++c1A+c0I+c1A1++cmAmp(A)=c_nA^n+\dots+c_1A+c_0I+c_{-1}A^{-1}+\dots+c_{-m}A^{-m}

In general, from a function f(x)f(x) that can be expanded in a power series

f(x)=n=+cnxnf(x)=\sum_{n=-\infty}^{+\infty}c_nx^n

we can build the operator f(A)f(A)

f(A)=n=+cnAnf(A)=\sum_{n=-\infty}^{+\infty}c_nA^n

(one sets A0=IA^0=I, where I is the identity operator)

Let us now return to the study of the time-evolution equation.

Conservation of the scalar product.

Finally, we postulate the principle of conservation of the total: as long as the system is not observed, the sum of the squared moduli over all positions does not change during the evolution. From this principle, together with linearity, there follows an important property of the time-evolution process: during this process the scalar products between different kets remain constant. Now we will see what this fact implies for the matrix A that appears in the equation

ddtψt=A(t)ψt\frac{d}{dt}|\psi t\rangle=A(t)|\psi t\rangle

Consider two kets αt|\alpha t\rangle and βt|\beta t\rangle; let us compute the evolution of these kets at the instant t+dtt+dt

α(t+dt)=αt+ddtαtdt+o(dt2)=(1+A(t)dt)αt+o(dt2)\begin{aligned} |\alpha(t+dt)\rangle & =|\alpha t\rangle+\frac{d}{dt}|\alpha t\rangle dt+o(dt^2) \\ & =(1+A(t)dt)|\alpha t\rangle+o(dt^2) \end{aligned}
β(t+dt)=βt+ddtβtdt+o(dt2)=(1+A(t)dt)βt+o(dt2)\begin{aligned} |\beta(t+dt)\rangle & =|\beta t\rangle+\frac{d}{dt}|\beta t\rangle dt+o(dt^2) \\ & =(1+A(t)dt)|\beta t\rangle+o(dt^2) \end{aligned}

By the property of conservation of the scalar product, it must be

α(t+dt)β(t+dt)=αtβt\langle\alpha(t+dt)|\beta(t+dt)\rangle=\langle\alpha t|\beta t\rangle

substituting the formulas found for α(t+dt)|\alpha(t+dt)\rangle and β(t+dt)|\beta(t+dt)\rangle we have

αt(1+A+(t)dt)(1+A(t)dt)βt+o(dt2)=αtβtαtβt+αt(A+(t)+A(t))dtβt+αtA+(t)dt2A(t)βt+o(dt2)=αtβtαt(A+(t)+A(t))dtβt+αtA+(t)dt2A(t)βt+o(dt2)=0\begin{aligned} \langle\alpha t|(1+A^+(t)dt)(1+A(t)dt)|\beta t\rangle+o(dt^2) & =\langle\alpha t|\beta t\rangle\Leftrightarrow \\ \langle\alpha t|\beta t\rangle+\langle\alpha t|(A^+(t)+A(t))dt|\beta t\rangle+\langle\alpha t|A^+(t)dt^2A(t)|\beta t\rangle+o(dt^2) & =\langle\alpha t|\beta t\rangle\Leftrightarrow \\ \langle\alpha t|(A^+(t)+A(t))dt|\beta t\rangle+\langle\alpha t|A^+(t)dt^2A(t)|\beta t\rangle+o(dt^2) & =0 \end{aligned}

dividing by dt and taking the limit as dt→0 we have

αt(A+(t)+A(t))βt=0\langle\alpha t|\left(A^+(t)+A(t)\right)|\beta t\rangle=0

This equation is valid for every αt|\alpha t\rangle and βt|\beta t\rangle, so it must be

A+(t)+A(t)=0A+(t)=A(t)\begin{aligned} A^+(t)+A(t) & =0\Leftrightarrow \\ A^+(t) & =-A(t) \end{aligned}

So the matrix A is anti-Hermitian. We replace the matrix A with the matrix H(t)=iA(t)H\left(t\right)=iA\left(t\right), where ii is the imaginary unit.

The matrix H is called the Hamiltonian and is Hermitian, indeed

H+=(iA)+=iA+=(i)(A)=iA=HH^+={\left(iA\right)}^+=i^*A^+=(-i)(-A)=iA=H

The evolution equation written in terms of the matrix H appears thus

iddtψt=H(t)ψti\frac{d}{dt}|\psi t\rangle=H(t)|\psi t\rangle

This equation is called the Schrödinger equation.

Italiano English