Matrices as Transformations
Loading learning experience...
Lecture transcript
Read the narration for Matrices as Transformations
From data vectors x to learned transformations A
Dr. Lena Hartmann: Neural nets look mysterious until you notice the same operation repeating everywhere: take a vector of numbers and transform it. Today we are going to discover the geometric meaning of matrix times vector, and you will be able to look at y equals A x and say: that is a rotation, a stretch, or a shear. We will start from the vector idea from last time, then move to pictures in 2D, and end by connecting it back to weight matrices in AI.
Kai: So last time was vectors as data. Now a matrix is like the thing that changes the data vector?
Dr. Lena Hartmann: Exactly. In L01, a single vector represented one data item, like an embedding or a feature list. Now we add a machine that takes that vector as input and produces a new vector as output.
Dr. Lena Hartmann: Driving question for today: when you see A times x, what does that transformation do to x in geometric terms?
Why AI cares: a linear layer is just y equals W x
Dr. Lena Hartmann: Here is the workhorse equation of modern AI: input vector x goes through weights W and becomes output vector y. If you understand the geometry of this, a lot of neural network behavior becomes less magical.
Dr. Lena Hartmann: This equation is saying: a matrix takes a vector and produces a new vector, like a controlled remix of features.
Kai: When people say the model learns weights, they literally mean it learns this matrix?
Dr. Lena Hartmann: Yes. Training adjusts W so that these transformations make useful intermediate representations. Next we will make W feel like a geometric action, not just a table of numbers.
First geometric handle: watch what A does to the basis
Dr. Lena Hartmann: In 2D, a clean way to understand a matrix is to track what it does to the two basic direction vectors: the unit step in the x direction and the unit step in the y direction. Once you know where those two directions end up, you can predict the transformation of every vector.
Dr. Lena Hartmann: This equation shows the key idea: the first column is where the first basis vector goes, and the second column is where the second basis vector goes.
Kai: So columns are not random numbers. They are literally the new axes after the transformation?
Dr. Lena Hartmann: Exactly. And any vector with coordinates x one and x two becomes x one times column one plus x two times column two. That is why matrix times vector is a weighted combination of columns.
Worked example: one matrix, one vector, one clear geometric effect
Dr. Lena Hartmann: Let us do a concrete multiply and then interpret it as a transformation. Before we compute anything, try to predict what will happen to the two coordinates when we apply this matrix to this vector.
Kai: I think the second output coordinate will stay 2, because it looks like it just copies the second input. For the first coordinate, I predict it becomes 2 times 1 plus 1 times 2, so the output should be the vector 4, 2.
Dr. Lena Hartmann: Now we compute and check that prediction. The first output entry is y one equals 2 times 1 plus 1 times 2, which is 4. The second output entry is y two equals 0 times 1 plus 1 times 2, which is 2. So A times x really is 4, 2, and the second coordinate stayed the same.
Kai: So the matrix is kind of pushing points sideways, but not really moving them up or down?
A matrix moves shapes: unit square to parallelogram
Dr. Lena Hartmann: Vectors are great, but shapes make the transformation pop. If we transform four corner points, the whole square comes along for the ride.
Dr. Lena Hartmann: We will keep using the same matrix A from Slide 4. Start with the unit square corners. These are the simplest test points because they are built from the basis directions.
Kai: And after applying A, those corners become a slanted shape, right? That is the shear we just computed with that same matrix.
Dr. Lena Hartmann: Exactly. And one extra AI-relevant detail: the area scaling of this parallelogram is controlled by the determinant of A, which later shows up in change of variables and probability densities.
Another shape test: unit circle to ellipse
Dr. Lena Hartmann: The unit square shows shear nicely. The unit circle shows stretching nicely: a circle becomes an ellipse under most matrices.
Dr. Lena Hartmann: This line is the key: all vectors of length one do not stay length one after a general matrix. Some directions get stretched more than others.
Kai: Is that why some directions in training feel like they explode or vanish? Like gradients get huge in some directions?
Dr. Lena Hartmann: Yes. Preview only: people often describe singular values as the maximum and minimum stretch over unit directions, but that is not the formal definition yet. The precise statement requires an optimization over all unit vectors and it connects to the matrix A transpose A; we will do that carefully later. For now, keep the picture: circle in, ellipse out.
Formal definition: linear transformations preserve addition and scaling
Dr. Lena Hartmann: Now that we have a picture, we can say what makes a transformation linear: it respects adding vectors and scaling vectors.
Dr. Lena Hartmann: This equation encodes both rules at once: transform the combination equals the same combination of transformed vectors.
Kai: So linear means no bending. It is just rotate, stretch, shear, and combinations of those?
Dr. Lena Hartmann: Exactly. And the big payoff is: every linear transformation can be represented as multiplication by some matrix A, once you choose coordinates.
Composing transformations: matrix multiplication is "do one, then the next"
Dr. Lena Hartmann: Neural networks stack layers. Mathematically, when you do one linear transformation and then another, you are composing transforms, and matrix multiplication is the notation for that composition.
Dr. Lena Hartmann: This equation shows the rule: applying B first and then A equals one combined matrix, A times B, acting on x.
Kai: The order still messes with me. AB means B happens first?
Dr. Lena Hartmann: Yes. You read it right to left because x is on the right. In code, it is like nested function calls: A of B of x.
Undoing a transformation: inverse matrices and "information loss"
Dr. Lena Hartmann: Sometimes we want to reverse a transformation. That is only possible if the matrix did not squash space into a lower dimension.
Dr. Lena Hartmann: This identity means the inverse perfectly undoes the matrix, returning every vector to where it started.
Kai: So if the transformation collapses a plane onto a line, there is no way back, because different inputs look the same after A?
Dr. Lena Hartmann: Exactly. In the square case, one quick test is the determinant: nonzero means invertible, zero means some direction got crushed to zero area.
AI connection: attention is built from learned matrix projections
Dr. Lena Hartmann: Let us connect the geometry back to something you have seen in model diagrams: attention. The key link is that attention starts by applying learned matrix transformations to embeddings.
Dr. Lena Hartmann: One quick note on notation: here X is typically n by d with token vectors stacked as rows, so we write Q equals X times W sub Q, and similarly for K and V. Earlier we often used column vectors like y equals W times x; both conventions are the same idea, just related by transpose and shape bookkeeping.
Kai: So the model is not just comparing original embeddings. It is comparing transformed versions that it learned to be useful?
Dr. Lena Hartmann: Yes. Think of it as the network learning the right coordinate system for similarity and information flow. That is matrices as transformations in production.
What to remember when you see A x
Dr. Lena Hartmann: Let us compress everything into a few mental shortcuts you can use immediately when reading ML papers or code.
Dr. Lena Hartmann: First: treat a matrix as an action, not as a grid to memorize. It moves vectors and shapes.
Kai: Second: I can predict the movement by looking at where the basis vectors go, which is the columns.
Dr. Lena Hartmann: And third: multiplication means composition, so deep nets are chains of transformations. Next time we will leverage this to solve systems and talk about when information is lost.
Exit ticket: columns-as-images test
Dr. Lena Hartmann: Quick check: the core idea is that a matrix is a transformation, and its columns tell you where the basic unit directions end up.
Dr. Lena Hartmann: Here is the matrix A we will use. Do not multiply everything yet, just remember what the columns mean.
Kai: So A times e one should just pick out the first column, and A times e two should pick out the second column.
Dr. Lena Hartmann: Exactly. So the answers are the two columns: A times e one is [1, 3] and A times e two is [2, 4]. Reflection-wise, notice how often you call a linear layer in code: every one of those calls is doing this kind of basis mixing in a much higher dimension.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!