Systems of Linear Equations
Loading learning experience...
Lecture transcript
Read the narration for Systems of Linear Equations
Last time: matrices move vectors; today: recover the vector
Kai: So last time, we treated a matrix like a function that takes in a vector and spits out a new vector, right?
Dr. Lena Hartmann: Exactly. Now we reverse the direction: we are given the output and we want the input that created it. That is the heart of solving systems of equations.
Dr. Lena Hartmann: In AI, this shows up as fitting parameters from data. Sometimes the system has an exact solution, and sometimes it is overdetermined, so we solve the best approximate version with least squares.
A system is one idea: match a transformation to a target
Dr. Lena Hartmann: This says: apply the matrix transformation A to the unknown vector x, and you must land exactly on b.
Kai: So x is literally just the list of numbers I am trying to solve for?
Dr. Lena Hartmann: Yes. Each row of A forms one dot product with x, and that dot product has to equal the corresponding entry of b. Many equations, one compact object.
In $2D$: solving means finding where lines intersect
Dr. Lena Hartmann: Here are two linear equations in x and y. Think of them as two constraints that x and y must satisfy at the same time.
Kai: So if I plot both lines, I should see the answer as their intersection point?
Dr. Lena Hartmann: Exactly. If they cross once, you have a unique solution. If they never cross, no solution. If they overlap, infinitely many solutions.
Worked example: Gaussian elimination is just organized algebra
Dr. Lena Hartmann: This is the same system, but written so we can do the same operations we do in algebra: add a multiple of one equation to another.
Kai: Why do we keep the second row as-is and change the first row?
Dr. Lena Hartmann: Either row could be used as the pivot, but here we keep the second row unchanged and use it to cancel the x term in the first row.
Dr. Lena Hartmann: We are trying to eliminate x from the first equation so it only talks about y. Once you see three y equals three, you get y equals one, and plugging back gives x equals two.
Let me show you in code: solve $A\mathbf{x}=\mathbf{b}$ in NumPy
Kai: I think x should be [2, 1], and A at x should come back as [5, 1]. That check feels like a unit test for linear algebra: solve, then assert the forward pass matches b.
Dr. Lena Hartmann: Yes — when you run it, numpy will print x as about [2. 1.], and the check A at x will be about [5. 1.]. And later, when data is noisy, you will not get perfect equality. Then we switch from exact solving to best-fit solving, which is least squares.
Row operations: change the look, not the solution set
Dr. Lena Hartmann: The goal is to transform the augmented matrix into an upper-triangular form, so solving becomes simple back-substitution.
Kai: These are exactly the tricks I used in algebra, just written as operations on rows.
Dr. Lena Hartmann: Right. Pivots are the columns where a variable gets pinned down. If a variable never gets a pivot, it becomes a free variable, which is where infinitely many solutions come from.
The matrix view: dimensions tell you what can happen
Dr. Lena Hartmann: This line also encodes dimensions: A has m equations and n unknowns. That m versus n is the first clue about unique, none, or many solutions.
Kai: Is square the nice case where we can hope for exactly one answer?
Dr. Lena Hartmann: Often, yes. With square systems you might get a unique solve if the matrix is invertible. With non-square systems, you usually get either many exact solutions or no exact solution, so you move to best-fit methods.
When does a system have a solution at all?
Dr. Lena Hartmann: This says: when you append b as an extra column, the amount of independent information should not increase. If it increases, b is asking for something impossible.
Kai: So the rank jumps when b forces a contradiction, like two parallel lines that never meet.
Dr. Lena Hartmann: And if the system is consistent but you do not have enough pivots to pin down every variable, you get a family of solutions parameterized by free variables.
Let me show you in code: intersection vs no intersection
Kai: This makes it click: solving is literally searching for agreement between constraints.
Dr. Lena Hartmann: Exactly. And in higher dimensions you cannot easily plot, but the same logic is still there: do the constraints intersect or not?
AI connection: most training is an overdetermined linear system
Dr. Lena Hartmann: Instead of demanding perfect equality, we pick x that makes A x as close as possible to b in squared error. This is linear regression in one line.
Dr. Lena Hartmann: You can also minimize this same kind of objective by gradient descent, just scaled up to huge nonlinear models. Linear systems are the clean starting point.
Checkpoint: what you should be able to do now
Dr. Lena Hartmann: If you can build A and b correctly, you have already done most of the modeling work: coefficients go into A, constants go into b.
Kai: And the geometry story is my mental model: intersection once, never, or everywhere.
Dr. Lena Hartmann: Exactly. And always verify: compute A times your solution and compare to b, or check the residual size for least squares.
Exit ticket: solve, then interpret
Dr. Lena Hartmann: Add the equations to eliminate y: you get two x equals six, so x equals three. Then plug back into x plus y equals four to get y equals one.
Kai: In regression, b is the targets y. In calibration, b could be measured outputs I want my model to match.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!