Vectors: The Language of Data
Loading learning experience...
Lecture transcript
Read the narration for Vectors: The Language of Data
Data needs a language
Dr. Lena Hartmann: Every AI system you use, from search to image generation to chatbots, has to turn messy human information into numbers before it can reason with it. The stakes are practical: if we understand those numbers, we understand the first layer of how AI sees the world.
Dr. Lena Hartmann: This first point says that images, text, and audio all become numbers. Not because the world is naturally a spreadsheet, but because computation needs a precise representation.
Dr. Lena Hartmann: Our destination is the idea of a vector: an ordered package of numbers that can represent a data point, a direction, a feature set, or an AI embedding.
Kai: So we are not starting with abstract symbols first. We are starting with how real data becomes something a model can compute with?
A vector is a data point with structure
Dr. Lena Hartmann: Let me show you the idea before we name every symbol. A vector is not just a random list; it is a list where position matters.
Dr. Lena Hartmann: This first phrase, ordered numbers, is the key. If the first number means height and the second means weight, swapping them changes the meaning completely.
Dr. Lena Hartmann: Each coordinate is like one measured feature. In a simple dataset, one coordinate might mean price, another might mean size, and another might mean rating.
Kai: So dimension just means how many feature slots the data point has?
First contact in Python
Dr. Lena Hartmann: Before we look at formal notation, let me show you in code. This is the most honest first definition of a vector for a programmer.
Dr. Lena Hartmann: The code creates two NumPy arrays, house and apartment. Each array has three entries, so each one is a three-dimensional vector representing one home. Notice how the order of numbers is important: square feet, then bedrooms, then bathrooms. What do you expect the output of house - apartment to be, and what would it represent?
Kai: I think it will subtract each value individually, so the output will show the difference in square footage, bedrooms, and bathrooms. Like, the house is 400 square feet larger, has 1 more bedroom, and 1 more bathroom.
Dr. Lena Hartmann: Excellent! Your prediction is spot on. Because these are vectors, specifically NumPy arrays, the subtraction happens element-wise. So it truly shows the difference in each feature: 400 square feet, 1 bedroom, and 1 bathroom. This element-wise operation is fundamental to how vectors are handled computationally.
Geometry: a vector as an arrow
Dr. Lena Hartmann: Now let us switch from data tables to geometry. In two dimensions, a vector can be drawn as an arrow.
Dr. Lena Hartmann: The tail starts at the origin. That gives us a common reference point, like saying every arrow begins at zero zero.
Dr. Lena Hartmann: The head lands at the coordinate pair. If the vector is three four, move three steps right and four steps up.
Kai: So the same object is both data in Python and an arrow in space?
Plot it with Matplotlib
Dr. Lena Hartmann: Now that we've discussed vectors conceptually, let's see how we can visualize them using a programming library like Matplotlib. This visual representation can be a powerful thinking tool.
Dr. Lena Hartmann: The code uses a numpy array for the vector's coordinates. The matplotlib dot pyplot dot quiver function then draws an arrow from the origin. Notice how vector v at index zero provides the x-coordinate, which is three, and vector v at index one provides the y-coordinate, which is four, for the arrow's head.
Kai: I like that. The vector is not just stored; I can literally see where it points.
Now formalize the object
Dr. Lena Hartmann: After exploring vectors through Python code and as geometric arrows, it's crucial that we formalize their mathematical definition. This allows us to work with vectors precisely, in any number of dimensions, and communicate complex mathematical ideas unambiguously.
Dr. Lena Hartmann: On screen, we now see the standard mathematical notation for a vector. It's typically represented as a column matrix, with its components stacked vertically. The boldface 'v' denotes the vector itself.
Dr. Lena Hartmann: The bold 'v' on the slide clarifies that it represents the entire vector, encompassing all its constituent parts.
Kai: Okay, so the bold 'v' is essentially the label for the whole vector object, that entire column of numbers.
Dr. Lena Hartmann: Each individual entry within the vector, such as v sub i, is known as a coordinate or component. Each component gives the vector's signed coordinate along a particular axis or feature direction.
Dr. Lena Hartmann: And finally, 'n' signifies the dimension of the vector, indicating how many components or coordinates it has. So, an n-dimensional vector has 'n' individual values.
Addition combines changes
Dr. Lena Hartmann: Vector addition is a fundamental operation that allows us to combine multiple changes or movements represented by vectors. When we add vectors, we are essentially finding the combined effect of these individual changes.
Dr. Lena Hartmann: Crucially, vectors must have the same dimension. You can't add a 2D vector to a 3D vector, as there would be no corresponding component. This ensures data compatibility.
Kai: So if my house vector had three features, but my apartment vector only had two, I couldn't add them to compare the total change?
Dr. Lena Hartmann: Exactly. Dimension compatibility is key. The result is a coordinate-wise combination where each component is the sum of its corresponding parts.
Dr. Lena Hartmann: Geometrically, vector addition is 'tip-to-tail' motion. As shown in the diagram, place the tail of the second vector at the tip of the first. The sum vector then connects the initial start to the final end point.
Scaling changes size
Dr. Lena Hartmann: Scalar multiplication changes a vector's magnitude, or length, and sometimes its direction. This is a core operation for manipulating vectors.
Kai: So, scalar multiplication essentially resizes the vector?
Dr. Lena Hartmann: Exactly. A scalar 'c' multiplies each component of the vector 'v', creating a new vector scaled proportionally.
Dr. Lena Hartmann: If scalar 'c' is greater than one, the vector stretches, becoming longer while maintaining its direction.
Kai: So, a 'c' value like 2 would make the vector twice as long?
Dr. Lena Hartmann: Precisely. If 'c' is positive between zero and one, the vector shrinks, becoming shorter with direction unchanged.
Kai: And if 'c' is negative, it reverses direction?
Dr. Lena Hartmann: That's right. If 'c' is negative, the vector's direction reverses. Its length still scales by the absolute value of 'c'.
Dr. Lena Hartmann: The diagram on the right shows these effects. We see vector 'v', '2v' demonstrating a stretch, and negative 'v' showing reversal of direction. Scalar multiplication is fundamental for manipulating vector size and orientation.
Length is information
Dr. Lena Hartmann: Beyond direction, a vector's length often carries crucial information, representing things like stronger forces or more significant data points.
Kai: So, length isn't just a number; it's part of the vector's message?
Dr. Lena Hartmann: Mathematically, we quantify a vector's length using its norm. This formula, similar to the Pythagorean theorem, calculates the Euclidean length of vector v as the square root of the sum of its squared components.
Dr. Lena Hartmann: The norm, also called the vector's magnitude, tells us 'how much' of something there is, without considering its direction.
Kai: So, magnitude gives us a sense of the vector's 'power' or 'impact,' regardless of direction?
Dr. Lena Hartmann: For a two-dimensional vector v with components 3 and 4, its magnitude is 5, a classic 3-4-5 right triangle, illustrating the norm formula.
Kai: That makes sense. The 3-4-5 triangle clarifies how the formula works.
Dr. Lena Hartmann: In AI, vector length can represent signal strength, image feature intensity, or model confidence. A larger norm implies more significant information.
Dr. Lena Hartmann: Understanding a vector's length is fundamental for extracting meaningful insights from data, from physics to AI.
Compute the basics
Dr. Lena Hartmann: Now we trust but verify. NumPy gives us a quick way to test the operations we just described.
Dr. Lena Hartmann: This code adds two vectors, scales one vector, and uses numpy linalg norm to compute its length. Given what we just learned about vector addition, scalar multiplication, and the norm, what results do you expect to see printed for sum vector, scaled vector, and length v?
Kai: Based on our discussions, I predict 'u + v' will be [5, 5], because 2 plus 3 is 5 and 1 plus 4 is 5. '2v' should be [6, 8] because two times three is six and two times four is eight. And the length of 'v' should be 5, from the 3-4-5 triangle we just discussed.
Dr. Lena Hartmann: Excellent! Your predictions are spot on. NumPy performs these operations exactly as we formalized them: coordinate-wise for addition and scaling, and using the Euclidean norm formula for length. This confirms our understanding and shows how seamlessly these mathematical concepts translate into code.
Vectors compare data
Dr. Lena Hartmann: Vectors become powerful when we compare them. A model rarely cares about one data point in isolation.
Dr. Lena Hartmann: If each user is represented by a vector, recommendation systems can compare users by their preferences.
Dr. Lena Hartmann: If each document is represented by a vector, search systems can compare your query to millions of documents.
Kai: And image embeddings are vectors too, so similar images should land near each other.
The dot product measures alignment
Dr. Lena Hartmann: The dot product is a fundamental operation in linear algebra that helps us understand the relationship, or 'alignment', between two vectors. It's particularly useful for determining how much one vector goes in the direction of another.
Dr. Lena Hartmann: Mathematically, the dot product of two vectors, u and v, is calculated by multiplying their corresponding components and then summing these products, as you see here.
Kai: Hold on. That is just multiply the matching entries and add them up. Where does an angle come into it?
Dr. Lena Hartmann: Good, that is exactly the right thing to be suspicious about. The angle is hiding in the algebra. Write out the squared length of u minus v in two ways, once with the law of cosines from geometry and once coordinate by coordinate, and the two agree only if u dot v equals length u times length v times cosine theta. We will not grind through that today. What matters is the consequence: the dot product is the two lengths multiplied by how aligned the directions are, so its sign tells you the angle.
Dr. Lena Hartmann: While this algebraic definition shows us how to compute it, the true power of the dot product lies in its geometric interpretation. For any two non-zero vectors u and v, their dot product is also equal to the product of their magnitudes and the cosine of the angle between them. This elegant relationship, u dot v equals the magnitude of u times the magnitude of v times the cosine of theta, is why the dot product measures alignment. Specifically, when cosine of theta is 1, they point in the exact same direction, and when it's minus 1, they are exactly opposite.
Dr. Lena Hartmann: This cosine relationship means that if the dot product is positive, the angle between the vectors is less than 90 degrees. This suggests the vectors are generally pointing in the same direction, or at least have a strong positive correlation in their orientation.
Dr. Lena Hartmann: If the dot product is zero, it means the cosine of the angle between them is zero. This happens when the angle is exactly 90 degrees, indicating the vectors are orthogonal or perpendicular. It's a key indicator that they have no directional relationship or influence on each other.
Kai: So if the dot product is zero, the two vectors have nothing in common?
Dr. Lena Hartmann: In terms of direction, yes. Zero means neither one leans along the other at all, which is why we call them orthogonal. They can still both be large; they simply point in unrelated directions.
Dr. Lena Hartmann: And finally, if the dot product is negative, the angle between the vectors is greater than 90 degrees. This tells us the vectors are pointing in generally opposite directions, or have a negative correlation in their orientation.
Dr. Lena Hartmann: So, if you had two vectors, one representing someone's interest in 'dogs' and another their interest in 'cats', and their dot product was a large negative number, what would that suggest about their interests?
Kai: It would probably mean they like dogs but dislike cats, or vice-versa, indicating a strong opposition in their preferences, right?
Dr. Lena Hartmann: Exactly! That's a great way to think about how dot products reveal underlying relationships in data.
Cosine similarity ignores size
Dr. Lena Hartmann: We'll now explore cosine similarity. It measures the orientation of two vectors, ignoring their magnitude. This is useful when only direction matters.
Dr. Lena Hartmann: As shown, it's the dot product of u and v, divided by their magnitudes. This normalization removes length's influence from the similarity score.
Kai: So, the division by magnitudes makes vector lengths irrelevant, focusing on the angle. Is that correct?
Dr. Lena Hartmann: Exactly. The result is between minus one and one. One means vectors point in the same direction, indicating maximum similarity.
Kai: And a score of zero means they're totally unrelated, like perpendicular, right?
Dr. Lena Hartmann: Correct. Zero means orthogonal or perpendicular, meaning they have no directional relationship. We call this 'unrelated direction'.
Dr. Lena Hartmann: Finally, minus one indicates vectors point in completely opposite directions, showing maximum dissimilarity in orientation.
Similarity search in code
Dr. Lena Hartmann: Vector search is a fundamental operation in many applications. Let's look at how we can implement a basic version of it in code.
Dr. Lena Hartmann: This cosine similarity function calculates the similarity between two vectors based on their dot product and magnitudes. We're comparing a query vector against two document vectors: doc AI and doc cooking. What do you expect the cosine similarity scores to be for 'query versus AI' and 'query versus cooking', and what would those scores tell us about their relevance?
Kai: Given the query's similarity to doc AI in its first two components, I expect 'query versus AI' to be high, close to 1. Doc cooking is very different, with a high value only in its third component, so 'query versus cooking' should be low, probably close to 0.
Dr. Lena Hartmann: Precisely. The output shows the query is very similar to doc AI, with a score around zero point nine nine. Conversely, its similarity to doc cooking is very low, around zero point one two. This is how AI systems quickly find relevant information by comparing vector directions.
Embeddings: meaning as position
Dr. Lena Hartmann: Now we can name one of the most important AI uses of vectors: embeddings.
Dr. Lena Hartmann: A word embedding turns a word into a vector. The coordinates are learned from data rather than hand chosen by a human.
Dr. Lena Hartmann: A sentence embedding does the same thing for a whole sentence, placing the sentence somewhere in a high-dimensional space.
Kai: So related ideas are not just linked by labels. They are close in vector space.
Dr. Lena Hartmann: It is worth saying why that works, because it is not a coincidence. An embedding model is trained on huge amounts of text with a simple objective: words that show up in similar contexts should get similar vectors. Cat and kitten appear in the same kinds of sentences, so training pulls their vectors together; car appears elsewhere, so it drifts away. The geometry is the residue of that training, not something anyone placed by hand.
PyTorch says the same thing
Dr. Lena Hartmann: If you work with modern AI code, you will see these operations in PyTorch as well as NumPy.
Dr. Lena Hartmann: This code uses torch dot nn dot functional dot cosine similarity. You'll notice the unsqueeze calls, which add a batch dimension. This is a common pattern in neural network code where operations are often performed on batches of data, even if it's just a single item for now.
Kai: That matches what I have seen in documentation: tensors are basically vectors and arrays with extra structure.
Dr. Lena Hartmann: Yes, exactly right, Kai! Your observation is key. PyTorch tensors are essentially multi-dimensional arrays, just like NumPy arrays, but they come with additional capabilities optimized for deep learning, such as G P U acceleration and automatic differentiation. So, while the underlying mathematical concept of a vector remains the same, PyTorch wraps it in a data structure that's powerful for AI.
What changes in high dimensions?
Dr. Lena Hartmann: Jumping from two dimensions to thousands sounds scary. But remember, the core math and operations we've learned don't break down, even if our intuition does.
Dr. Lena Hartmann: In two dimensions, we easily draw a vector as an arrow on paper, from the origin to its coordinates. It's very intuitive.
Dr. Lena Hartmann: In three dimensions, we can still imagine physical space, like moving through a room. But even drawing it accurately on a 2D screen becomes challenging.
Kai: So does that mean in really high dimensions, like for an AI embedding, the vectors lose their 'arrow' meaning? Or is it just that I can't see it?
Dr. Lena Hartmann: That's an excellent question, Kai. The 'arrow' meaning doesn't disappear; it's simply our human ability to visualize that hits a limit. Vectors still have direction and magnitude; we just can't see them. The coordinate rules still work perfectly, and we rely on the math to tell us about their relationships in that vast space.
Three mental models to keep
Dr. Lena Hartmann: Let us compress the lesson into three mental models. You will use all three, depending on the problem.
Dr. Lena Hartmann: As a data row, a vector is an ordered list of features. This is the spreadsheet view.
Dr. Lena Hartmann: As an arrow, a vector has direction and length. This is the geometry view.
Kai: And as a meaning location, a vector lets AI compare words, images, documents, and users.
Exit ticket
Dr. Lena Hartmann: Before we finish, try this quick check. It tests the basic symbol and the geometric meaning at the same time.
Dr. Lena Hartmann: For the practice item, the vector has coordinates six and eight. Take a moment to calculate the norm yourself. Then, we'll check it together. To find the norm, you square the coordinates to get thirty six and sixty four, add them to get one hundred, then take the square root. The norm is ten.
Kai: For the reflection, I would say the same vector can be all three depending on context: features, arrow, or location. For the check, I'd say using vectors for semantic search, like when comparing a query to documents.
Dr. Lena Hartmann: Precisely, Kai! You've perfectly captured the versatility of vectors: features, arrows, or locations, all depending on context. Your semantic search example, comparing queries and documents with vectors, is a fantastic illustration of their power in intelligent systems.
Keep going: the next lesson builds on this one.
Thanks for spending time with Thinking in Math. Leave your questions in the YouTube comments and we'll tackle them in future videos.