Chapter 3
Complex Numbers and State Vectors
In the previous chapter the electrons drew diffraction rings, and to explain them we had to associate a wave with the beam. To write the equation of that wave we need a system of numbers: here we see which one to choose, and we gather the rules of calculation we shall use from the next chapter onwards.
Which numbers we use
There are several kinds of numbers with characteristics similar to the real numbers. What we need is a normed division algebra over the real numbers.
What the name says
Algebra means that the numbers can be added and multiplied. Of addition we want what holds for the real numbers: the order of the terms does not matter, there is a zero, every number has its opposite. Of multiplication we want it to distribute over addition, and real factors to be movable out of it: multiplying by twice must give twice . We also want the number to exist: a number that, multiplied by any other, leaves it as it is.
Division means that one can divide by any number other than zero. Solving means dividing both sides by : if for some number other than zero the division could not be carried out, that equation would be left without a solution. This is not a schoolroom case. In the arithmetic of the twelve-hour clock makes , which on the clock is zero, although neither nor is zero. To ask that division always be possible is to ask that this shall not happen.
Normed means that every number a has a modulus, a non-negative real number which we write , and that the modulus of a product is the product of the moduli; “norm” is the name mathematicians give to the modulus. Among the real numbers the modulus is the absolute value and the rule already holds: . When a number is made of several real components we take as modulus the ordinary distance from the origin, : this is what is already done with the complex numbers, where . It is the property that narrows the field the most.
Over the real numbers means that every number of the algebra is made of one or more real components, and that the coefficients by which we multiply it are real numbers.
How many possibilities do we have?
How many real numbers it takes to write a number of the algebra: this is what we call the dimension. A real number has dimension 1. A complex number has dimension 2, because giving it requires the two reals and . If there were numbers made of three reals, their algebra would have dimension 3.
One would expect to be able to build one in every dimension. It is not so: the mathematician Adolf Hurwitz proved that the algebras of this kind are only four:
| Algebra | Dimension | What stops holding |
|---|---|---|
| real numbers | 1 | — |
| complex numbers | 2 | the ordering |
| quaternions | 4 | commutativity of multiplication |
| octonions | 8 | associativity of multiplication |
The third column is to be read cumulatively. The four starting properties — addition, multiplication, division, modulus — hold in all four algebras: that is what makes them normed division algebras. But each time we go up in dimension a property that still held in the previous algebra is lost, and the losses carry over: the real numbers have everything; the complex numbers have everything except the ordering; the quaternions have lost commutativity as well; the octonions associativity too.
Let us now see what each loss means.
The ordering. Of two real numbers we can always say which is the greater, and the relation agrees with the operations: adding the same quantity to two numbers does not change their order, and the product of two positive numbers is positive. Among complex numbers no such ordering exists: asking whether is greater or smaller than makes no sense. This is a loss we can afford, because the numbers we read off our instruments remain real, and those can still be ordered.
In exchange for the loss there is a gain, and it is worth saying at once because further on it will be decisive: over the complex field every nonconstant polynomial equation has at least one solution. Over the reals this is not so — has none — and it was precisely in order to solve it that was born. This property is called algebraic closure, the complex numbers are the smallest extension of the reals that possesses it, and among the four algebras they are the only ones that have it.
Commutativity. The quaternions have a real part and three imaginary units, , , : they are written , and Hamilton found them in 1843. Their product depends on the order of the factors: , but . Anyone using them must state each time from which side they multiply, and every rule splits into a right-hand and a left-hand version. They are not a curiosity: they are the usual way of representing rotations in space, and they are used in computer graphics, in robotics and in the attitude control of satellites. There non-commutativity is an advantage, because rotations do not commute either: two rotations performed in a different order lead to different positions.
Associativity. The octonions have eight components, one real and seven imaginary. In them even the grouping of factors changes: and may differ, and steps we usually take without thinking must be redone from scratch. Outside mathematics they have no established use: they appear in some attempts at theoretical physics, but nothing settled.
To sum up: the complex numbers are the last algebra in which multiplication remains commutative and associative, the only property we lose there is the ordering, which we do not need, and in exchange we gain algebraic closure.
Waves point us towards complex numbers
In the chapter on the diffraction of electrons we saw that a beam of electrons must be associated with a wave. To describe a wave the amplitude is not enough: two waves of the same amplitude can reinforce or cancel each other depending on their relative phase, and it is precisely these cancellations that form the rings on the screen. Every wave therefore carries two pieces of information, the amplitude and the phase, and the second counts as much as the first.
A complex number carries exactly two pieces of information: the modulus and the angle that the point makes with the real axis, called the argument. Calling them and ,
and the last equality is Euler’s formula. We write the real wave
as the real part of
The single factor keeps amplitude and phase together. And the operations we need become products: shifting the phase means multiplying by a number of modulus one; differentiating with respect to time, for a wave of definite frequency, means multiplying by . Above all, adding two waves becomes adding two complex numbers, and the sum takes the relative phase into account by itself.
Taking the real part is legitimate because the equations we write are linear and have real coefficients. None of this concerns quantum waves only: it holds for sound, for light, for alternating currents.
A single electromagnetic field
Complex numbers are not indispensable to classical electromagnetism, but they show their usefulness there too. Maxwell’s equations in vacuum, with the charges and currents in place, are
where is the charge density, the current density, the speed of light in vacuum, and the two constants of free space are related by .
Let us set
The factor gives and the same dimensions. The real part of is the electric field, the imaginary part divided by is the magnetic field. The two divergence equations become a single complex equation,
while the two evolution equations become
To verify this, it is enough to substitute the definition of and separate the two parts. In the first equation the real part gives Gauss’s law and the imaginary part says that the magnetic field has no sources; in the second the real part gives Faraday’s law and the imaginary part the Ampère-Maxwell law. The complex vector is called the Riemann-Silberstein vector. The four equations have become two, sources and all, and the electric and the magnetic field sit in a single field.
Here the measurable fields remain real: the complex number does not change electromagnetism, it shortens the way we write it. The same advantage is found in alternating-current circuits, where resistance and reactance form the complex impedance .
The eigenvalue problem
There is one last reason, and in our treatment it is the most concrete. We shall see that to every observable quantity there corresponds a matrix, and that the values that quantity can take are the eigenvalues of that matrix. An eigenvalue is a number for which there exists a non-zero vector that the matrix merely multiplies by , without changing its direction; to find them one solves a polynomial equation.
Over the real field a polynomial equation may have no solutions. The matrix that rotates the plane by a right angle has no real eigenvalue, and it is clear why: under a rotation of ninety degrees no direction stays where it was. If we worked with real numbers alone we would have to accept that certain quantities have no possible value at all.
Over the complex field this never happens: it is the algebraic closure we anticipated, also known as the fundamental theorem of algebra — every nonconstant polynomial has at least one root. This is where that gain is collected: by choosing the complex numbers we make sure that the eigenvalue problem always has a solution.
To sum up: among the four possible algebras the complex numbers are the last in which multiplication remains commutative and associative; they hold the amplitude and the phase of a wave in a single quantity; they gather the electric and the magnetic field into one; they guarantee that the eigenvalue problem has a solution.
Vectors, bras and kets
Given a complex number c=a+ib we will denote its complex conjugate by the symbol c*=a-ib (c* is read “c star”). The product cc*=a2+b2 is a positive real number and is called the squared modulus of c.
Let us now consider a column vector , where and are two complex numbers: every component of the vector is a complex number, and every complex number is in turn made of two real numbers. We start from two components so that the examples stay simple; the components can be as many as we like, and for a wave on a line they are infinitely many. We can define the dual of this column vector by the row vector . In this way the product between the vector and its dual
gives a positive real number that is called the squared modulus of the vector.
If we have a physical state α, we will denote the vector associated with it by the symbol
and we will denote its dual by the symbol
The symbols and are called bra and ket respectively, because joining these names gives the word “bracket”, which in English means parenthesis. We introduced them because they let us distinguish between rows and columns, and make very clear the reading of expressions containing matrix products.
For example we can have
or we can have
These two products are completely different: in one case we have a row vector times a column vector and obtain a number, in the other case we have a column vector times a row vector and obtain a matrix.
Scalar product.
Given two vectors and , the product
is called the scalar product between and .
Let us recall that in the real field the scalar product between two real vectors and is defined without the complex conjugate:
Let us observe, first of all, that the two definitions are compatible, because the complex conjugate of a real number is the number itself a*=a. So the definition given in the complex field reduces to the one given in the real field if the vectors to which it is applied are real.
However a doubt may remain: why in the complex field do we not define the scalar product simply without the complex conjugate
?
The answer is very simple: this definition does not enjoy some very important properties. For example, with this definition the product of a vector with itself is not necessarily a real number, so the concept of modulus is lost. Moreover, in some cases the product could be zero without the vector being zero, for example if
we have
whereas with the correct definition
To conclude this digression let us see the fundamental properties of the scalar product in a complex space:
Decomposition of vectors.
by definition we will say that two vectors and are orthogonal if their scalar product is zero
Theorem: if n non-zero vectors are pairwise orthogonal
then these n vectors are linearly independent, that is none of them can be obtained as a linear combination of the others.
Proof: by contradiction, suppose that one vector, for example the first, is a combination of the others
we multiply scalarly on the right by
since the vectors are by hypothesis orthogonal, all the terms on the right-hand side are zero, hence
but this is absurd because the vectors are all non-zero.
Now suppose we have an n-dimensional vector space and we must choose a basis for it. As is known, one can choose any n-tuple of linearly independent vectors, but it is very convenient to choose these vectors so that they are orthogonal to each other and have unit modulus . Indeed, doing so it becomes very simple to compute the components of a generic vector . Let us see how:
We take a vector and a basis of orthogonal vectors of unit modulus
can be written as a linear combination of the vectors
to compute the coefficient ck we multiply this expression on the left by
since the basis vectors are pairwise orthogonal we have
since has unit modulus we have
So we are left with the formula to compute ck
the component of along the basis vector is given by the scalar product between these two vectors.
These are the rules we shall use: the scalar product, orthogonality and the component of a vector along a basis vector. The algebra of complex vector spaces has many other developments, which we shall introduce when they are needed.