Hook
More breakout videos from this creator.
Every token that comes into the LLM is a number. And LLM uses these numbers to calculate the next token. The only thing here it doesn't use this ones. Today we'll talk about vectors and how information is represented in a model. My name is Ivan, I'm a computer science teacher, founder of easybits.tech, and I help adults to keep up with tech. Let's talk vectors. We know that when text comes into a model, we convert into sequence of tokens. And every token is basically a number. We can open websites like Tiktokenizer and see real numbers there. But in reality, these numbers, they are just labels. It's literally means that this is token no. 574. And once the sequence of these labels or as they called token IDs coming into a model, it actually falls into something that is called embedding matrix. And conceptually, this embedding matrix is a huge table where we have the ID of the token and its vector representation. So what is vector? Essentially, vector is the array of numbers, so kind of ordered sequence. And you may ask, why we need multiple numbers if we already have this one? Because with one number, we are quite limited. We need multiple numbers to have a flexibility and represent different properties of the word. See, when model producing the text, we are not working with words as set of characters. We actually work also with the meaning of the word and with the relations of words between each other. And we need to encode quite a bit of information. Let's take one very simple example. Imagine we have two words: cat and plane. And we want to encode them with vectors. And we see that the values, they are quite different because cat is very different from a plane. And if we want now add dog here, we understand that values will be quite close to the cat, not the same ones because it's a different species. But it's will still be pretty far from a plane. But if we will decide to add, for example, bird, bird is already a little bit closer to a plane because it also can fly and it have wings. So you can see that different words have different relations in different characteristics. And the amount of numbers in our vector is called dimension. So here in our oversimplified example, we have three-dimensional vector. But in the reality, we have something like 4096 dimensional vector. And once we processed all our tokens through this embedding matrix, for each of them, we have a separate vector. And after that, we kind of stack these vectors together. And we have something that is called in math, matrix. And this matrix with some additional information is actually what goes into LLM for calculation. This set of factors, it's not going through just one mathematical operation. It actually goes through multiple layers. But we will talk about it in the next episodes. So follow me to not miss it.