Chapter 3

Complex Numbers and State Vectors

In the previous chapter the electrons drew diffraction rings, and to explain them we had to associate a wave with the beam. To write the equation of that wave we need a system of numbers: here we see which one to choose, and we gather the rules of calculation we shall use from the next chapter onwards.

Which numbers we use

There are several kinds of numbers with characteristics similar to the real numbers. What we need is a normed division algebra over the real numbers.

What the name says

Algebra means that the numbers can be added and multiplied. Of addition we want what holds for the real numbers: the order of the terms does not matter, there is a zero, every number has its opposite. Of multiplication we want it to distribute over addition, and real factors to be movable out of it: multiplying xx by twice uu must give twice xuxu. We also want the number 11 to exist: a number that, multiplied by any other, leaves it as it is.

Division means that one can divide by any number other than zero. Solving ax=bax=b means dividing both sides by aa: if for some number other than zero the division could not be carried out, that equation would be left without a solution. This is not a schoolroom case. In the arithmetic of the twelve-hour clock 262\cdot 6 makes 1212, which on the clock is zero, although neither 22 nor 66 is zero. To ask that division always be possible is to ask that this shall not happen.

Normed means that every number a has a modulus, a non-negative real number which we write a|a|, and that the modulus of a product is the product of the moduli; “norm” is the name mathematicians give to the modulus. Among the real numbers the modulus is the absolute value and the rule already holds: ab=ab|ab|=|a|\,|b|. When a number is made of several real components we take as modulus the ordinary distance from the origin, a12++an2\sqrt{a_1^2+\cdots+a_n^2}: this is what is already done with the complex numbers, where a+ib=a2+b2|a+ib|=\sqrt{a^2+b^2}. It is the property that narrows the field the most.

Over the real numbers means that every number of the algebra is made of one or more real components, and that the coefficients by which we multiply it are real numbers.

How many possibilities do we have?

How many real numbers it takes to write a number of the algebra: this is what we call the dimension. A real number has dimension 1. A complex number a+iba+ib has dimension 2, because giving it requires the two reals aa and bb. If there were numbers made of three reals, their algebra would have dimension 3.

One would expect to be able to build one in every dimension. It is not so: the mathematician Adolf Hurwitz proved that the algebras of this kind are only four:

AlgebraDimensionWhat stops holding
real numbers1
complex numbers2the ordering
quaternions4commutativity of multiplication
octonions8associativity of multiplication

The third column is to be read cumulatively. The four starting properties — addition, multiplication, division, modulus — hold in all four algebras: that is what makes them normed division algebras. But each time we go up in dimension a property that still held in the previous algebra is lost, and the losses carry over: the real numbers have everything; the complex numbers have everything except the ordering; the quaternions have lost commutativity as well; the octonions associativity too.

Let us now see what each loss means.

The ordering. Of two real numbers we can always say which is the greater, and the relation agrees with the operations: adding the same quantity to two numbers does not change their order, and the product of two positive numbers is positive. Among complex numbers no such ordering exists: asking whether 2+3i2+3i is greater or smaller than 3+2i3+2i makes no sense. This is a loss we can afford, because the numbers we read off our instruments remain real, and those can still be ordered.

In exchange for the loss there is a gain, and it is worth saying at once because further on it will be decisive: over the complex field every nonconstant polynomial equation has at least one solution. Over the reals this is not so — x2+1=0x^2+1=0 has none — and it was precisely in order to solve it that ii was born. This property is called algebraic closure, the complex numbers are the smallest extension of the reals that possesses it, and among the four algebras they are the only ones that have it.

Commutativity. The quaternions have a real part and three imaginary units, ii, jj, kk: they are written a+bi+cj+dka+bi+cj+dk, and Hamilton found them in 1843. Their product depends on the order of the factors: ij=kij=k, but ji=kji=-k. Anyone using them must state each time from which side they multiply, and every rule splits into a right-hand and a left-hand version. They are not a curiosity: they are the usual way of representing rotations in space, and they are used in computer graphics, in robotics and in the attitude control of satellites. There non-commutativity is an advantage, because rotations do not commute either: two rotations performed in a different order lead to different positions.

Associativity. The octonions have eight components, one real and seven imaginary. In them even the grouping of factors changes: (ab)c(ab)c and a(bc)a(bc) may differ, and steps we usually take without thinking must be redone from scratch. Outside mathematics they have no established use: they appear in some attempts at theoretical physics, but nothing settled.

To sum up: the complex numbers are the last algebra in which multiplication remains commutative and associative, the only property we lose there is the ordering, which we do not need, and in exchange we gain algebraic closure.

Waves point us towards complex numbers

In the chapter on the diffraction of electrons we saw that a beam of electrons must be associated with a wave. To describe a wave the amplitude is not enough: two waves of the same amplitude can reinforce or cancel each other depending on their relative phase, and it is precisely these cancellations that form the rings on the screen. Every wave therefore carries two pieces of information, the amplitude and the phase, and the second counts as much as the first.

A complex number carries exactly two pieces of information: the modulus and the angle that the point a+iba+ib makes with the real axis, called the argument. Calling them AA and ϕ\phi,

a+ib=A(cosϕ+isinϕ)=Aeiϕ,a+ib=A(\cos\phi+i\sin\phi)=Ae^{i\phi},

and the last equality is Euler’s formula. We write the real wave

Acos(kxωt+ϕ)A\cos(kx-\omega t+\phi)

as the real part of

Aeiϕei(kxωt).Ae^{i\phi}e^{i(kx-\omega t)}.

The single factor AeiϕAe^{i\phi} keeps amplitude and phase together. And the operations we need become products: shifting the phase means multiplying by a number of modulus one; differentiating with respect to time, for a wave of definite frequency, means multiplying by iω-i\omega. Above all, adding two waves becomes adding two complex numbers, and the sum takes the relative phase into account by itself.

Taking the real part is legitimate because the equations we write are linear and have real coefficients. None of this concerns quantum waves only: it holds for sound, for light, for alternating currents.

A single electromagnetic field

Complex numbers are not indispensable to classical electromagnetism, but they show their usefulness there too. Maxwell’s equations in vacuum, with the charges and currents in place, are

 ⁣ ⁣E=ρε0, ⁣ ⁣B=0, ⁣× ⁣E=Bt, ⁣× ⁣B=μ0J+1c2Et.\begin{aligned} \nabla\!\cdot\!\mathbf E&=\frac{\rho}{\varepsilon_0}, & \nabla\!\cdot\!\mathbf B&=0,\\ \nabla\!\times\!\mathbf E&=-\frac{\partial\mathbf B}{\partial t}, & \nabla\!\times\!\mathbf B&=\mu_0\mathbf J+\frac{1}{c^2}\frac{\partial\mathbf E}{\partial t}. \end{aligned}

where ρ\rho is the charge density, J\mathbf J the current density, cc the speed of light in vacuum, and the two constants of free space are related by ε0μ0c2=1\varepsilon_0\mu_0c^2=1.

Let us set

F=E+icB.\mathbf F=\mathbf E+ic\mathbf B.

The factor cc gives E\mathbf E and cBc\mathbf B the same dimensions. The real part of F\mathbf F is the electric field, the imaginary part divided by cc is the magnetic field. The two divergence equations become a single complex equation,

 ⁣ ⁣F=ρε0,\nabla\!\cdot\!\mathbf F=\frac{\rho}{\varepsilon_0},

while the two evolution equations become

iFt=c ⁣× ⁣Fiε0J.i\frac{\partial\mathbf F}{\partial t}=c\,\nabla\!\times\!\mathbf F-\frac{i}{\varepsilon_0}\mathbf J.

To verify this, it is enough to substitute the definition of F\mathbf F and separate the two parts. In the first equation the real part gives Gauss’s law and the imaginary part says that the magnetic field has no sources; in the second the real part gives Faraday’s law and the imaginary part the Ampère-Maxwell law. The complex vector F\mathbf F is called the Riemann-Silberstein vector. The four equations have become two, sources and all, and the electric and the magnetic field sit in a single field.

Here the measurable fields remain real: the complex number does not change electromagnetism, it shortens the way we write it. The same advantage is found in alternating-current circuits, where resistance and reactance form the complex impedance Z=R+iXZ=R+iX.

The eigenvalue problem

There is one last reason, and in our treatment it is the most concrete. We shall see that to every observable quantity there corresponds a matrix, and that the values that quantity can take are the eigenvalues of that matrix. An eigenvalue is a number λ\lambda for which there exists a non-zero vector that the matrix merely multiplies by λ\lambda, without changing its direction; to find them one solves a polynomial equation.

Over the real field a polynomial equation may have no solutions. The matrix that rotates the plane by a right angle has no real eigenvalue, and it is clear why: under a rotation of ninety degrees no direction stays where it was. If we worked with real numbers alone we would have to accept that certain quantities have no possible value at all.

Over the complex field this never happens: it is the algebraic closure we anticipated, also known as the fundamental theorem of algebra — every nonconstant polynomial has at least one root. This is where that gain is collected: by choosing the complex numbers we make sure that the eigenvalue problem always has a solution.

To sum up: among the four possible algebras the complex numbers are the last in which multiplication remains commutative and associative; they hold the amplitude and the phase of a wave in a single quantity; they gather the electric and the magnetic field into one; they guarantee that the eigenvalue problem has a solution.

Vectors, bras and kets

Given a complex number c=a+ib we will denote its complex conjugate by the symbol c*=a-ib (c* is read “c star”). The product cc*=a2+b2 is a positive real number and is called the squared modulus of c.

Let us now consider a column vector (α1α2)\left(\begin{gathered} \alpha_1 \\ \alpha_2 \end{gathered}\right), where α1\alpha_1 and α2\alpha_2 are two complex numbers: every component of the vector is a complex number, and every complex number is in turn made of two real numbers. We start from two components so that the examples stay simple; the components can be as many as we like, and for a wave on a line they are infinitely many. We can define the dual of this column vector by the row vector (α1,α2)(\alpha_1^*,\alpha_2^*). In this way the product between the vector and its dual

(α1,α2)(α1α2)=α1α1+α2α2=α12+α22\begin{aligned} (\alpha_1^*,\alpha_2^*)\left(\begin{gathered} \alpha_1 \\ \alpha_2 \end{gathered}\right) & =\alpha_1^*\alpha_1+\alpha_2^*\alpha_2 \\ & =|\alpha_1|^2+|\alpha_2|^2 \end{aligned}

gives a positive real number that is called the squared modulus of the vector.

If we have a physical state α, we will denote the vector associated with it by the symbol α|\alpha\rangle

α=(α1α2)|\alpha\rangle=\left(\begin{gathered} \alpha_1 \\ \alpha_2 \end{gathered}\right)

and we will denote its dual by the symbol α\langle\alpha|

α=(α1,α2)\langle\alpha|=\left(\alpha_1^*,\alpha_2^*\right)

The symbols   \langle\;| and   |\;\rangle are called bra and ket respectively, because joining these names gives the word “bracket”, which in English means parenthesis. We introduced them because they let us distinguish between rows and columns, and make very clear the reading of expressions containing matrix products.

For example we can have

αβ=(α1,α2)(β1β2)=α1β1+α2β2\begin{aligned} \langle\alpha|\beta\rangle & =\left(\alpha_1^*,\alpha_2^*\right)\left(\begin{gathered} \beta_1 \\ \beta_2 \end{gathered}\right) \\ & =\alpha_1^*\beta_1+\alpha_2^*\beta_2 \end{aligned}

or we can have

αβ=(α1α2)(β1,β2)=(α1β1α1β2α2β1α2β2)\begin{aligned} |\alpha\rangle\langle\beta| & =\left(\begin{gathered} \alpha_1 \\ \alpha_2 \end{gathered}\right)\left(\beta_1^*,\beta_2^*\right) \\ & =\begin{pmatrix} \alpha_1\beta_1^* & \alpha_1\beta_2^* \\ \alpha_2\beta_1^* & \alpha_2\beta_2^* \end{pmatrix} \end{aligned}

These two products are completely different: in one case we have a row vector times a column vector and obtain a number, in the other case we have a column vector times a row vector and obtain a matrix.

Scalar product.

Given two vectors α|\alpha\rangle and β|\beta\rangle, the product

αβ=(α1,α2)(β1β2)=α1β1+α2β2\begin{aligned} \langle\alpha|\beta\rangle & =\left(\alpha_1^*,\alpha_2^*\right)\left(\begin{gathered} \beta_1 \\ \beta_2 \end{gathered}\right) \\ & =\alpha_1^*\beta_1+\alpha_2^*\beta_2 \end{aligned}

is called the scalar product between α|\alpha\rangle and β|\beta\rangle.

Let us recall that in the real field the scalar product between two real vectors α|\alpha\rangle and b|b\rangle is defined without the complex conjugate:

ab=(a1,a2)(b1b2)=a1b1+a2b2\begin{aligned} \langle a|b\rangle & =\left(a_1,a_2\right)\left(\begin{gathered} b_1 \\ b_2 \end{gathered}\right) \\ & =a_1b_1+a_2b_2 \end{aligned}

Let us observe, first of all, that the two definitions are compatible, because the complex conjugate of a real number is the number itself a*=a. So the definition given in the complex field reduces to the one given in the real field if the vectors to which it is applied are real.

However a doubt may remain: why in the complex field do we not define the scalar product simply without the complex conjugate

αβ=(α1,α2)(β1β2)=α1β1+α2β2\langle\alpha|\beta\rangle=(\alpha_1,\alpha_2)\left(\begin{gathered} \beta_1 \\ \beta_2 \end{gathered}\right)=\alpha_1\beta_1+\alpha_2\beta_2 ?

The answer is very simple: this definition does not enjoy some very important properties. For example, with this definition the product of a vector with itself αα\langle\alpha|\alpha\rangle is not necessarily a real number, so the concept of modulus is lost. Moreover, in some cases the product αα\langle\alpha|\alpha\rangle could be zero without the vector α|\alpha\rangle being zero, for example if

α=(1i)|\alpha\rangle=\left(\begin{gathered} 1 \\ i \end{gathered}\right)

we have

αα=(1,i)(1i)=1+i2=11=0\begin{aligned} \langle\alpha|\alpha\rangle & =(1,i)\left(\begin{gathered} 1 \\ i \end{gathered}\right) \\ & =1+i^2 \\ & =1-1 \\ & =0 \end{aligned}

whereas with the correct definition

αα=(1,i)(1i)=1i2=1+1=2\begin{aligned} \langle\alpha|\alpha\rangle & =(1,-i)\left(\begin{gathered} 1 \\ i \end{gathered}\right) \\ & =1-i^2 \\ & =1+1 \\ & =2 \end{aligned}

To conclude this digression let us see the fundamental properties of the scalar product in a complex space:

αα is real and is 0αα=0 if and only if α is the null vectorαβ=βααβ1+β2=αβ1+αβ2\begin{aligned} & \langle\alpha|\alpha\rangle\text{ is real and is }\geq0 \\ & \langle\alpha|\alpha\rangle=0\text{ if and only if }|\alpha\rangle\text{ is the null vector} \\ \langle\alpha|\beta\rangle & =\langle\beta|\alpha\rangle^* \\ \langle\alpha|\beta_1+\beta_2\rangle & =\langle\alpha|\beta_1\rangle+\langle\alpha|\beta_2\rangle \end{aligned}

Decomposition of vectors.

by definition we will say that two vectors α|\alpha\rangle and β|\beta\rangle are orthogonal if their scalar product is zero

αβ=0\langle\alpha|\beta\rangle=0 ⇔ they are orthogonal

Theorem: if n non-zero vectors α1αn|\alpha_1\rangle\cdots|\alpha_n\rangle are pairwise orthogonal

αjαk=0 jk\langle\alpha_j|\alpha_k\rangle=0\ \forall j\neq k

then these n vectors are linearly independent, that is none of them can be obtained as a linear combination of the others.

Proof: by contradiction, suppose that one vector, for example the first, is a combination of the others

α1=λ2α2++λnαn|\alpha_1\rangle=\lambda_2|\alpha_2\rangle+\cdots+\lambda_n|\alpha_n\rangle

we multiply scalarly on the right by α1\langle\alpha_1|

α1α1=λ2α1α2++λnα1αn\langle\alpha_1|\alpha_1\rangle=\lambda_2\langle\alpha_1|\alpha_2\rangle+\dots+\lambda_n\langle\alpha_1|\alpha_n\rangle

since the vectors are by hypothesis orthogonal, all the terms on the right-hand side are zero, hence

α1α1=0α1\langle\alpha_1|\alpha_1\rangle=0\Rightarrow|\alpha_1\rangle is null

but this is absurd because the vectors are all non-zero.

Now suppose we have an n-dimensional vector space and we must choose a basis for it. As is known, one can choose any n-tuple of linearly independent vectors, but it is very convenient to choose these vectors so that they are orthogonal to each other and have unit modulus αkαk=1\langle\alpha_k|\alpha_k\rangle=1. Indeed, doing so it becomes very simple to compute the components of a generic vector v|v\rangle. Let us see how:

We take a vector v|v\rangle and a basis of orthogonal vectors of unit modulus

α1αn|\alpha_1\rangle\cdots|\alpha_n\rangle

v|v\rangle can be written as a linear combination of the vectors α1αn|\alpha_1\rangle\cdots|\alpha_n\rangle

v=c1α1++cnαn|v\rangle=c_1|\alpha_1\rangle+\cdots+c_n|\alpha_n\rangle

to compute the coefficient ck we multiply this expression on the left by αk\langle\alpha_k|

αkv=c1αkα1++ckαkαk++cnαkαn\langle\alpha_k|v\rangle=c_1\langle\alpha_k|\alpha_1\rangle+\cdots+c_k\langle\alpha_k|\alpha_k\rangle+\cdots+c_n\langle\alpha_k|\alpha_n\rangle

since the basis vectors are pairwise orthogonal we have

αkv=ckαkαk\langle\alpha_k|v\rangle=c_k\langle\alpha_k|\alpha_k\rangle

since αk|\alpha_k\rangle has unit modulus we have

αkv=ck\langle\alpha_k|v\rangle=c_k

So we are left with the formula to compute ck

ck=αkvc_k=\langle\alpha_k|v\rangle

the component of v|v\rangle along the basis vector αk|\alpha_k\rangle is given by the scalar product between these two vectors.

These are the rules we shall use: the scalar product, orthogonality and the component of a vector along a basis vector. The algebra of complex vector spaces has many other developments, which we shall introduce when they are needed.

Italiano English