Background The Dirac δ and its derivative
The is defined by the following equality, valid for every sufficiently regular function :
It is this equality that defines it: the is characterised by what it produces under the integral sign, not by the values it takes point by point.
No ordinary function can satisfy it: it would have to vanish everywhere except at one point and still have integral one. The belongs to a wider class of objects, the distributions.
If a mental image helps, one can think of the as the limit of functions that grow ever narrower and ever taller, with area always equal to one — Gaussians whose width tends to zero, for instance. That limit does not exist as a function, but the rule above does, and it is the only thing we ever use.
The derivative. This too is defined by its effect under the integral, and the definition is chosen so that integration by parts keeps holding. Moving the derivative from the onto the function, the boundary term vanishes — the is zero away from the origin — and what remains is:
In words: the derivative of the , applied to a function, returns minus the derivative of that function, evaluated at the point. The minus sign comes from the integration by parts, not from a convention.
Why it is needed in the chapter. The derivative operator, written as an infinite-dimensional matrix , has precisely as its kernel. The computation in the chapter is the check of this fact, and it is a direct application of the formula above.
A caveat. The manipulations performed with the — differentiating it, integrating it by parts — are legitimate because they are definitions, not theorems about ordinary functions. The framework that makes them rigorous is the theory of distributions, due to Laurent Schwartz; there is no need to know it here: the two lines of definition suffice.