|
Archive:
Subtopics:
Comments disabled |
Tue, 11 Aug 2026
There are two kinds of theorems
In mathematical study there are two kinds of theorems, which serve very different purposes. Math instruction follows the same pattern. Students are often very puzzled by this, and rightly so, because it's never explained, or at least I've never seen it explained. There is this crucial, critical piece of mathematical methodology which is never made explicit, students just have to figure it out on their own, and many of them never do. When we do mathematics, we construct a simplified model of some phenomenon. For example, Euclidean geometry is a simplified model of how shapes and lines actually work. In formal geometry, things are simple: lines have no thickness, and three or more lines might all intersect at the exact same point. There are perfect circles, where every point is the exact same distance from the center, and there are perfect rectangles with perfectly straight sides and perfectly equal angles. Real shapes aren't like this. Nobody can draw an infinitely thin line. Nobody has ever seen a geometrically perfect circle or rectangle. Three lines, however carefully drawn, will always intersect in three different places. That's okay! The point of geometry is to construct a simplified model that is easier to deal with. When we're setting up a mathamtical model, we start by describing its basic objects, like points and lines, and with axioms and postulates, what properties we intend the objects to have. For example, Euclid has:
Having done that, we state and prove the theorems of the first kind. We're not studying the actual phenomenon yet. We're not yet trying to learn anything new about shapes and circles. Instead, we're investigating the model itself:
So for example Euclid starts by proving extremely simple theorems. For example, propositions 4 and 5:
Obviously, yes, anyone can see that! We didn't need to develop a whole mathematical theory in order to discover that vertical angles were equal. Everyone already knew that, long before Euclid. The point of proving this theorem, the real discovery, is: our simple model is strong enough to demonstrate that vertical angles are equal. The theory didn't explicitly include anything about vertical angles, but the vertical alngle theorem was latent in the model anyway. Consider the opposite situation, where we couldn't prove that vertical angles were equal. Or worse, what if the model allowed us to construct a pair of unequal vertical angles? Would this tell us something about vertical angles? Obviously not. Vertical angles are equal, regardless of what the theory does or doesn't prove. It would, though, tell us something about the theory, namely that it wasn't fit for purpose, and we'd better try a different model. This is a common pattern in all mathematics education and indeed in all mathematics, but I've rarely seen it called out as such, not even with a passing remark like “here's why we're doing this”. In nearly every undergraduate class I've ever been in, someone was puzzled about why we were doing this. Why is Euclid proving all these theorems that are obvious? In Euclid the mode goes back and forth: Euclid will do some model-verifying theorems, then move on to interesting-result theorems, then back for a while to introduce something new to the model, then forward again to prove interesting theorems about the thing he introduced. The first transition happens around proposition 32 or 36 or so. Up until that point there were a lot of proposition like this:
and this:
But then very soon after, the theorems start to have a different flavor:
That is:
This is actually an interesting fact about parallelograms, and not intuitively obvious. Even though the two parallelograms are not at all the same shape, they have equal areas, since they lie between the same parallels and !!AB=CD!!. (Interactive version) The same issue comes up in many different contexts. We develop the theory of the Peano numbers, define addition, and prove that addition is commutative, Was that because we didn't know how to do addition? No. We already knew that addition was commutative. The point of the theorem is to show that Peano arithmetic knows that addition is commutative. But I have more than once seen instructors demonstrate the proof, via a double induction, and then finish with a remark like “therefore, addition is commutative!” The students, to their credit, were suspicious of this. They knew something wasn't quite right, even if they weren't sure what. The right announcement would have been something like “therefore, the Peano axioms aren't complete rubbish!” Or: We explore Dedekind cuts, we define a model for the real numbers as cuts of rationals, and a construction that we claim characterizes addition. And then we prove a batch of theorems that are intended to show things we already know about addition, not because we want to know whether addition is commutative (news flash: it is) but to show that it's plausible that our construction really does characterize addition. One of these, that the addition operation we defined on cuts, which looks nothing like the addition we defined on rational numbers, actually agrees with it when the cuts themselves correspond to rationals. Another, that if !!a < b!! then !!a+c < b+c!!. Taking a look at Rudin Principles of Mathematical Analysis, I see that the first fifteen or so pages are like this, theorems like !!\lvert z\rvert = \lvert \bar z \rvert!!, which is a basic property of the fundamental notions !!\lvert z\rvert!! and !!\bar z!!. And then the mode starts to shift, first a little bit, with
Okay, that's a triangle inequality again… and suddenly, seemingly out of nowhere, something not at all obvious: Theorem 1.35, the Cauchy-Schwartz inequality for !!\Bbb C^1!!: $$ \left\lvert\sum_{j=1}^n a_j\bar b_j\right\rvert ^2 ≤ \sum_{j=1}^n\lvert a_j\rvert^2 \sum_{j=1}^n\lvert b_j\rvert^2 $$ A completely different kind of theorem, not a basic property of anything. Remember the whole point of the process: We wanteded to model some object of study, we built a model, we proved a lot of theorems to lend plausibility to our model, to verify that the model wasn't broken, to confirm that the model captures the properties of interest. And then came time to use the model, and we started to prove theorems that told us new things about the original object of study. Does Rudin announce this shift? Of course not, Rudin never announces anythingġ. (Usually he mutters, and sometimes if you are especially unlucky he fixes you with a glare that dares you to question the remark he throws away in an undertone.) But Rudin is Rudin, and nobody else seems to announce this shift either. Almost always, it's passed over, usually without remark, even in gentler textbooks that give more attention to pedagogical matters. Sometimes the shift is sudden, sometimes gradual, but it's almost never pointed out. In advanced study, that's okay, because advanced students should be expected to recognize the pattern. But why do we expect high schoolers and undergraduates to understand this without explanation? In summary:
[Other articles in category /math] permanent link |