Chapter 2
Cascaded Stern–Gerlach Experiments
In this card we will describe a series of experiments carried out by arranging several Stern–Gerlach machines in cascade. The results of these experiments will lead us to introduce the fundamental features of Quantum Mechanics and, in the final part, we will give a first simplified formulation of this theory.
Experiment number 1.
(See fig. 1) We have two Stern–Gerlach machines in a row, one with a vertical magnetic axis, the other inclined with respect to the vertical by an angle ϑ.
The first machine splits the atomic beam into two parts; we will write that one part has magnetic moment m0=+k and the other m0=-k. The subscript zero in the symbol m0 represents the inclination angle of the machine; for example, if the angle is 45° we will write m45. The symbol k represents the absolute value of the measured magnetic moment which, as we saw in the previous card, is always the same in every experiment.
Figure 2 shows a diagram representing the arrangement of the machines and the path of the atoms. At the output of the first machine the beam with m0=-k is blocked and only the beam with m0=+k is let through; we will therefore say that the first machine selects atoms with m0=+k.
The beam with m0=+k is passed into the second machine and here it splits again into two parts, mϑ=+k and mϑ=-k. The two fractions are in general not equal and it is seen experimentally that they are proportional to the terms and .
For example, if the second machine is also placed vertically (fig. 3 ϑ=0°) then the beams mϑ=+k and mϑ=-k have intensities proportional to the terms 1 and 0, that is the beam mϑ=+k has all the intensity and the beam mϑ=-k has zero intensity. If instead the second machine is placed at ninety degrees (fig. 4 ϑ=90°) then the beams m90=+k and m90=-k have intensities proportional to the terms ½ and ½, that is they are equal.
It is interesting to observe that already these first results cannot be explained by the concepts of classical mechanics.
If we also rotate the first machine (fig. 5) we observe that the results do not change substantially and, as was easy to imagine, the two outgoing beams have intensities proportional to the terms and .
This means that what matters is only the relative inclination between the two machines and not the absolute inclination with respect to the plumb line.
▶Try the interactive simulator · Experiment 1→
Experiment number 2.
(See fig. 6) We have three machines. Let us consider the beam coming out of machine two; this beam is split by machine three into two parts of intensities proportional to the terms and . If the inclination of the first machine is varied, we observe that the intensity of the beam coming out of the second machine varies, but the way in which the third machine splits this beam does not vary, that is the ratios between the intensities of the outgoing beams and the intensity of the beam entering the third machine do not vary. This experiment leads us to conclude that the atoms coming out of machine two are in a physical state that is independent of the orientation of machine one. We will indicate this physical state by saying that the atoms are in the state mϑ2=+k.
▶Try the interactive simulator · Experiment 2→
Probabilistic interpretation of the results.
In this section we will give a possible interpretation of these first experiences that we have described.
Let us consider the first experiment. We can think that machine one places the atoms in a particular physical state, which we will denote by m0=+k. If a measurement of m0 is performed on these atoms, orienting machine two vertically, the value m0=+k is obtained with certainty. If instead a measurement of mϑ (with ϑ≠0) is performed by inclining the machine, one obtains that each single atom has a probability of giving mϑ=+k, and a probability of giving mϑ=-k.
In other words, our interpretation reads as follows:
if, on an atom in the state m0=+k, the quantity mϑ (projection of the magnetic moment in the direction ϑ) is measured, two possible outcomes occur: mϑ=+k, with probability , and mϑ=-k with probability .
According to this interpretation, the beam coming out of the first machine is not composed of a fraction of atoms with mϑ=+k and another with mϑ=-k, but consists of atoms that are all in the state m0=+k and each of them, when mϑ is measured, has a certain probability of giving mϑ=+k and another of giving mϑ=-k.
This way of seeing things is perhaps not the only possible one, but it is certainly very difficult to imagine another view consistent with all the facts we have described and will describe.
At this point we have outlined a first important feature of Quantum Mechanics, the one linked to probability. When a measurement is performed on a physical system several outcomes are possible, in our case two, but there may be more or even infinitely many. Each possible outcome has a certain probability of occurring, and Quantum Mechanics will allow us to compute these probabilities.
Superposition of states.
Figure 7 shows that with two Stern–Gerlach machines it is possible to recombine two beams that were previously separated by a first machine. The experiments of this section are based on the superposition of two beams.
Experiment number 3.
(See fig. 8) The first machine places the atoms in the state m0=+k, the second splits the beam into two parts having m90=+k and m90=-k, and the last two recombine the beam.
Three measurements of m0 are performed (fig. 9).
In case 1 we measure m0 on atoms in the state m90=+k, so we have the two possibilities m0=+k and m0=-k with probabilities 0.5 and 0.5.
In case 2 we measure m0 on atoms in the state m90=-k; we again have the two possibilities m0=+k with probability 0.5 and m0=-k with probability 0.5.
In case 3 we measure m0 on atoms that are in a state which is a superposition of the states m90=+k and m90=-k. The result is that we have the two possibilities m0=+k with probability 1 and m0=-k with probability 0, exactly as if there had been no splitting and recombination.
To summarise, we have:
The thing to note is that by superposing two states that each have probability 0.5 of giving m0=-k one obtains a state that has zero probability of having m0=-k,
0.5+0.5=0!
So by superposing the states the probabilities do not add. We ask whether there is some other mathematical entity linked to the physical states that, when the states are superposed, combines according to a simple law, perhaps a law of addition.
▶Try the interactive simulator · Experiment 3→
Experiment number 4.
(See fig. 10) This experiment is almost identical to the previous one, with the only difference that the beam with m90=-k travels a longer path than the beam with m90=+k. This can be achieved with the system represented schematically in figure 11. The lower beam passes through a region where the field gradient is more intense, so it curves more and travels a longer path.
Let us see what the results are as the phase shift ∆ between the two paths varies.
If the phase shift is zero, at the output we have the state m0=+k exactly as in the previous experiment, but if the phase shift is different from zero we have a different state. If ∆ becomes equal to a certain characteristic length λ, at the output we again have m0=+k; even if ∆ is a multiple of λ, ∆=2λ,3λ, etc., we still have the state m0=+k. If ∆ is equal to λ/2, at the output we have the state m0=-k!
In general, for the probabilities we have the following formulas:
where we have introduced the variable φ=2π∆/λ.
Now let us make the following mathematical calculations. Let us take two numbers in the complex field, for example the number one twice, 1 and 1, and represent them as vectors.
Now we phase-shift the second number by multiplying it by the phasor eiφ, obtaining the pair 1 and eiφ.
We add the result thus obtained and take its squared modulus. We obtain the result:
Now we repeat everything starting from the two numbers 1 and –1.
We phase-shift the second number by φ, obtaining 1 and -eiφ.
We add and take the squared modulus. We obtain the result:
At this point it seems fairly clear that the probabilities obtained experimentally can be seen as the square of the sum of two complex numbers. The factor 4 present in the results does not matter because it can be eliminated by normalising the probability distribution.
Now we can organise the ideas as follows.
With the state m90=+k we associate two complex numbers, the number 1 taken from the first calculation and another number 1 taken from the second calculation. With the state m90=-k we associate the second number 1 of the first calculation and the number –1 of the second calculation.
when we phase-shift the state m90=-k we multiply its two numbers by eiφ
when we superpose the states, we add the corresponding pairs of numbers and the resulting state is represented by the sum
The squared moduli of the numbers found give the probability distribution of obtaining the values m0=+k and m0=-k when a measurement of m0 is performed.
These numbers refer to a particular physical quantity of the system, in our case m0. Each single complex number is associated with one of the possible values that the physical quantity can take, and the squared modulus of these numbers gives the probability that, on measuring, the associated value is obtained.
We observe that the pairs of complex numbers representing the state of our system form a vector space, because they can be added and can be multiplied by a number.
The experiments we have described led us to think that, to represent physical states, it is convenient to use complex vectors, because these entities add simply when we superpose the states.
The idea of using complex vectors to represent states will prove very fruitful; indeed all the laws of Quantum Mechanics are written as simple linear equations in terms of these vectors. Before proceeding with the theory it is convenient to devote a section to the algebra of complex vector spaces.
▶Try the interactive simulator · Experiment 4→
A review of the algebra of complex vector spaces.
Given a complex number c=a+ib we will denote its complex conjugate by the symbol c*=a-ib (c* is read “c star”). The product cc*=a2+b2 is a positive real number and is called the squared modulus of c.
Let us now consider a column vector ; we can define its dual by the row vector . In this way the product between the vector and its dual
gives a positive real number that is called the squared modulus of the vector.
If we have a physical state α, we will denote the vector associated with it by the symbol
and we will denote its dual by the symbol
The symbols and are called bra and ket respectively, because joining these names gives the word “bracket”, which in English means parenthesis. We introduced them because they let us distinguish between rows and columns, and make very clear the reading of expressions containing matrix products.
For example we can have
or we can have
These two products are completely different: in one case we have a row vector times a column vector and obtain a number, in the other case we have a column vector times a row vector and obtain a matrix.
Scalar product.
Given two vectors and , the product
is called the scalar product between and .
Let us recall that in the real field the scalar product between two real vectors and is defined without the complex conjugate:
Let us observe, first of all, that the two definitions are compatible, because the complex conjugate of a real number is the number itself a*=a. So the definition given in the complex field reduces to the one given in the real field if the vectors to which it is applied are real.
However a doubt may remain: why in the complex field do we not define the scalar product simply without the complex conjugate
?
The answer is very simple: this definition does not enjoy some very important properties. For example, with this definition the product of a vector with itself is not necessarily a real number, so the concept of modulus is lost. Moreover, in some cases the product could be zero without the vector being zero, for example if
we have
whereas with the correct definition
To conclude this digression let us see the fundamental properties of the scalar product in a complex space:
Decomposition of vectors.
by definition we will say that two vectors and are orthogonal if their scalar product is zero
Theorem: if n non-zero vectors are pairwise orthogonal
then these n vectors are linearly independent, that is none of them can be obtained as a linear combination of the others.
Proof: by contradiction, suppose that one vector, for example the first, is a combination of the others
we multiply scalarly on the right by
since the vectors are by hypothesis orthogonal, all the terms on the right-hand side are zero, hence
but this is absurd because the vectors are all non-zero.
Now suppose we have an n-dimensional vector space and we must choose a basis for it. As is known, one can choose any n-tuple of linearly independent vectors, but it is very convenient to choose these vectors so that they are orthogonal to each other and have unit modulus . Indeed, doing so it becomes very simple to compute the components of a generic vector . Let us see how:
We take a vector and a basis of orthogonal vectors of unit modulus
can be written as a linear combination of the vectors
to compute the coefficient ck we multiply this expression on the left by
since the basis vectors are pairwise orthogonal we have
since has unit modulus we have
So we are left with the formula to compute ck
the component of along the basis vector is given by the scalar product between these two vectors.
The algebra of complex vector spaces has many other developments that we will deal with later. For now we have wished to introduce the scalar product because it is indispensable for formulating a first simplified version of Quantum Mechanics.
Let us return to our simple system.
The state m0=+k is characterised by the fact that if a measurement of m0 is performed the value m0=+k is obtained with certainty, that is the probability of obtaining m0=+k is one and the probability of obtaining m0=-k is zero.
So this state can be represented by the ket vector
In this choice we have a certain degree of freedom; indeed the probabilities are the squared modulus of the components. So the vectors or are also fine, because . This degree of freedom must be seen as, in classical mechanics, the choice of the reference frame is seen: it is completely indifferent which choice we make, so it is convenient to make the most convenient one.
For the state m0=-k an analogous discussion holds
We observe that the vectors we have chosen are orthogonal and have unit modulus
Let us now consider a state α obtained from a combination of the states m0=+k and m0=-k
We can say that this state has properties intermediate to those of the component states, in the sense that, if m0 is measured, the two possible values m0=+k and m0=-k are obtained with probabilities proportional to the terms and
To obtain the actual probabilities we must normalise the distribution
If we multiply the two numbers α+ and α- by the same complex number c
we have the distribution
normalising, we obtain exactly the same probabilities as before
This is consistent with the fact that in Quantum Mechanics the vectors
represent the same combination of the basis states and therefore the same physical state. Choosing one vector or the other is completely indifferent.
The most convenient choice is to take normalised vectors, that is vectors such that . In this way the probabilities are obtained directly from the squared moduli, without having to normalise
So from now on we represent states with unit-modulus vectors. The coefficients of a unit-modulus state vector are given a special name: they are called “probability amplitudes”. α+ is the probability amplitude of obtaining m0=+k and is the probability of obtaining m0=+k, and analogously for α- and for .
Let us now consider two states mϑ=+k and mφ=+k, and consider the vectors that represent them
the quantities αϑ+, αϑ- and αφ+, αφ- are, respectively, the probability amplitudes of finding m0=+k and m0=-k when the system is in the states mϑ=+k and mφ=-k.
But what is the probability amplitude of finding mϑ=+k by measuring mϑ on atoms in the state mφ=+k?
The sought probability amplitude equals the product
that is, it equals the component of the vector along the vector .
The probability, as usual, is given by the squared modulus
This rule was easy to guess, but it cannot be proved on the basis of the things we had already said. It is a fundamental principle of Quantum Mechanics, and has general validity:
suppose we have a system in a state α and we want to know the probability that, measuring a quantity g, a particular value is obtained.
First we must take the normalised ket associated with the state α.
Then we must take the normalised ket associated with the state in which, if g is measured, the value is obtained with certainty.
Finally we must compute the component of the ket along the ket
thus obtaining the probability amplitude.
The probability is given by the squared modulus
Now we are able to choose the right components of the vectors and for each angle ϑ.
For the probabilities relative to we have the following experimental results
which, for example, may derive from the vectors
In this case too we have a certain degree of freedom, because we cannot directly measure the probability amplitudes, but only the probabilities. The simplest choice is the following
For the state we can choose
The minus sign in front of is necessary because the two vectors must be orthogonal . Indeed we know that the probability of obtaining mϑ=+k in the state mϑ=-k is zero.
At this point, to verify the consistency of the things we have said, let us try to compute the probability of finding mϑ=+k in a state mφ=+k.
Analogously one can compute
In perfect agreement with the cascaded Stern–Gerlach thought experiments we have described.
As a last step it remains for us to study the time evolution of the state vectors.
Time evolution.
Given a system, suppose we know the initial state ket and the “ambient conditions”. We ask what the state ket will be at a generic instant t.
We can organise an experiment as follows:
we take a beam of atoms whose state we know, we pass it through a box containing various Stern–Gerlach machines and, on exit after the time needed to cross the box, we observe the state of the atoms (see fig. 12).
When setting up the box, however, we must be careful. We are studying a simplified system, in which the atoms are characterised by their magnetic moment alone, with no other property such as position. In the earlier experiments we did use septa to block certain beams, but those septa were there to prepare the initial state; everything downstream of them was then observed without further interference. We do the same here. Upstream we may position septa — say, to select the beam with ϑ=0 — in order to set the initial state; from then on we simply let the system evolve, with the magnetic field acting on the spin as the only external influence. Inside the box, however, we cannot place septa to block the beams during this evolution: any atom that struck one of them would be in a state no longer described by its magnetic moment alone but also by its position — which would take us beyond the simple picture we are studying.
We can consider an experiment we have already carried out (Experiment number 4), whose figure we reproduce.
Suppose we can rotate the first machine by an angle ϑ, so at the input we have atoms in the state . When these atoms enter the box the state is split into two components
where the terms c+ and c- are given by the scalar products
The beam is phase-shifted by a factor eiφ and at the output we have the composite state
where U is the matrix
This matrix is called the time-evolution matrix between the instants t0 and t, and refers to the particular environment, that is to the apparatus that constitutes the experiment.
In general, for any phenomenon there exists a matrix such that, given , one computes with the product .
An important property of the time-evolution matrix is that it preserves scalar products. That is, if we take two state vectors and , and take the vectors obtained from their evolution
the scalar products between the first pair and between the second pair are equal
We can verify this property for the particular case of experiment 4, but first let us look at a detail of algebra:
If we take a vector , its dual is . If we multiply by a number c we have , and the dual will be . Hence the dual of is .
Now we compute and show that the result is equal to .
performing the scalar product and eliminating the null terms
To conclude, let us summarise the fundamental principles we have arrived at: we have not proved them but introduced them — they are suggested by the thought experiments we have carried out.
Principles of Quantum Mechanics (simplified version).
Complex vector. The state of the system is described by a vector with one component for each of the possible values of the observed quantity. For every value that the measurement of the quantity g can give there exists a state, written as the ket , in which that measurement yields with certainty; taking these kets as a basis, the component is the scalar product and is called the probability amplitude. For a system in the state , the probability of obtaining the value is proportional to the squared modulus of the corresponding amplitude:
Linear superposition. Two states can be combined, each with a complex coefficient, and the combination is still a state. The state evolves continuously in time and, under evolution, each term evolves as it would on its own, with the same coefficient, and the subsequent state is the sum of the evolved terms:
Conservation of the total. As long as the system is not observed, the sum of the squared moduli of all the components does not change during the evolution.
The matrix that appears in the second principle depends on the system and on the “ambient conditions”. It is called the time-evolution matrix and, by the principles just stated, leaves scalar products invariant
What we have summarised is a simplified version of the fundamental principles. In general, to describe the state of a system it is not sufficient to use a single observable quantity, but it will be necessary to use more than one. In what follows we will show how the theory is generalised to more complex cases, and we will address the important problem of determining the time-evolution matrix for the various systems we will study. In particular, when the observed quantity is position the possible values become infinitely many and the vector of complex numbers turns into a function.
These are the three principles stated in the introduction: there in the particular form of a particle moving in space, here in general form — and the “function” of the introduction is here a vector, because the possible values are finite in number. The introduction adds a fourth to them, agreement with Newton’s mechanics: it would serve no purpose here, because there is not yet any motion in space to compare with classical mechanics.