Transcripts

Hands-On AI 2 Transcript

Please be advised that this transcript is AI-generated and may not be word-for-word. Time codes refer to the approximate times in the ad-free version of the show.
 

Mikah Sargent [00:00:00]:
All right. Remember how last week I was typing into the Ollama app and I was getting answers about a torque wrench? Well, these are the files that are responsible for those answers, a 6.59 gigabyte file. And I thought, hey, let's open it up and give it a look. And this text here that I was talking to, oh look, there's the word Facebook, was responsible for giving us that answer. I can scroll through here and see all of this code, all of this text that actually is equivalent to numbers, and it's what's responsible for us being able to talk and get responses from The model. There's no copy of the internet in here. There's no list of facts. It's just numbers, and they appear as letters, but there are billions of them.

Mikah Sargent [00:01:06]:
So how, how in the world do billions of numbers write a paragraph about torque wrenches? Well, that is the question we are answering today. I'm Micah Sargent, and this is Hands-on AI. Podcasts you love. From people you trust. This is TWiT. Welcome or welcome back to Hands-on AI. I hope you enjoyed the first episode. Today we are digging in even more, and today we're talking about, well, what an LLM is, how it works, and why you can type things in and get responses from a bunch of numbers.

Mikah Sargent [00:01:55]:
So in order to do that, I thought we'd take a look at the model that we were working with last week, which of course was Quen 3.59B. And look, there are a bunch of numbers and a bunch of parameters that all mean that we have a model that could run on our laptop. But let's kind of dig in and talk about what makes up this model. So we'll head over to macOS and take a look. All right, here we are on macOS. And the first thing that I wanna show you is, well, A, how to run some code on the terminal. So I've got the Terminal app open on the Mac and I'm going to type in O-L-L-A-M-A, that's Ollama, show. And then what we're going to do is take a look at our specific model that we're running.

Mikah Sargent [00:02:44]:
This is the model we downloaded last week. That's Quen 3.5, And that is 9B for 9 billion parameters. Boom. Now, when we look at this, we can see 9.7 billion parameters is how many we're working with. Also the context and some other information that makes up this model. And we can dig into it even more than we did last week, if you can believe it. So what do all of these different numbers and letters mean? Well, we can take a look at that. Now, first and foremost, what makes up a model? Well, you have your parameters, and in this case, 9.7 billion.

Mikah Sargent [00:03:33]:
You remember that last week I called these parameters sort of knobs, and I'm sticking with that. They're little adjustments that can be made to what the model is made of. Every one of these knobs is a number, and every one got its value during training. We're going to talk about training soon, but it's important to understand that that training happens— it's an odd thing to say, but training happens during training. Those knobs are set during training. It's not as if someone goes in and makes adjustments to all 9 billion knobs themselves. Of course, we also have quantization. We talked about that.

Mikah Sargent [00:04:16]:
That's the Q4KM part, and that is the bitrate Another term that we're going to kind of stick to, how tightly those numbers got squeezed after they were set. Then there's also the embedding. This is the embedding length. So in this case, 4,096. And we'll talk about that number in a minute. So we'll come back to it. Just kind of keep that in your mind. 4,096.

Mikah Sargent [00:04:47]:
And of course, Well, there's the 5.6 gigabyte file. It shows up differently depending on your system. In this case, mine was 6.59 gigabytes, but folks may see 5.6 gigabytes depending on if they got that extra amount that was part of a vision model. So I wondered myself, how in the world Do 9 billion numbers actually fit into 5.6 gigabytes? That's where the squeezing happens, because if you think about it, it's 16-bit precision. The 9 billion numbers should take about 18 gigabytes of storage. This file, of course, is 5.6 gigabytes. So it does this through the quantization that we talked about before. If you're wondering what the heck is quantization, well, it's the rounding that happens.

Mikah Sargent [00:05:53]:
At 16 bits, each number can be any of— this is nerdy, but 65,000 values. Each number can be any of about 65,000 values. So it's like a ruler that has 65,000 little tick marks on it, right? Right? And at 4 bits, there are only 16 tick marks, and then every number will snap to the nearest one. So that's where, when we talked about it last week, that 4 bits number that I mentioned comes from. Each group of numbers snaps to those sections and gets its own little ruler. That is sized to fit that specific group. So those 16 marks land wherever the numbers actually are, and the file will store that ruler once per group and then a 4-bit tick number for each of the knobs that we talked about before. Now, this is where things get magical and are the unique qualities of each of the different models, because a few of the parts will actually be kept at higher precision.

Mikah Sargent [00:07:08]:
And then the rulers there will take up more space. So when that happens, that's why the average ends up coming out to 5 instead of 4. But when you're doing that rounding, that quantization, well, the cost is that every number is now just a little bit off. But the good thing is with this much math working all at once, The model mostly holds up anyway. So you do get a file about a third of the size that you would get with just a little bit of a hit to quality. So that leads to the 5 bits per each individual knob. Pretty nice, right? And again, being able to kind of adjust different parts of the different different knobs and make them more precise or less precise is one of the things that can make for a unique model. And if you count all of it, it averages out to, again, an experience that gives you the results that are at least seemingly accurate.

Mikah Sargent [00:08:18]:
But we'll talk about that soon. Now, again, with this specific model, this Quen that we downloaded, it is able to take in images, and that's why it is a little bit bigger at that 6.59 or nearly 6.6 gigabytes. But with just the text part of the model, that's what makes it smaller. Now I want to talk about scale here because this is pretty incredible when it comes to this model. The previous generation, of Quen was trained on— trained, which we'll talk about too— 36 trillion tokens. And tokens are basically pieces of text. We'll get into that as well, if you can believe it. But understand that, you know, Quen doesn't have, for this model, an exact number that it is provided.

Mikah Sargent [00:09:17]:
But the point is, when you are looking at this from the specific math point of view, and you're thinking, how could a person possibly turn all these knobs? No, they're not. It's the model finding patterns in, like, during training that results in those knobs getting turned. And that is what keeps them in place and results in the outcomes that we see. Like its ability to answer our question. So we just left off talking about, well, what makes up a model. And now it's time to talk about how a model actually reads, how it understands, and how it processes. And it does so with a term that you may have heard before. The term is attention.

Mikah Sargent [00:10:14]:
Attention. Tokens. And what's great is, well, we've got a wonderful tool that is going to help us figure out tokens, and that tool comes from OpenAI. OpenAI has a tool called Tokenizer, and what it does is it tells you what are tokens, how Words are made up of tokens and what they do. And before a model is able to do anything with your words, it has to chop them up into these little pieces. Those pieces are tokens. And in fact, if you look at our show artwork, we sort of made a reference to the fact that each of those words, Hands on AI, and the bits and pieces that make it up are parts of tokens. Now, tokens aren't just words.

Mikah Sargent [00:11:06]:
They can be a whole word, they can be a part of a word, they can be a space and a word together, they can be punctuation, they can even be a single letter. So let's give this a shot. First, we're going to try a sentence like Hands on AI, if I can type this correctly, is my favorite show. Because of course, that's how you all feel, right? Now check this out. We can see there are 8 tokens that make up this this sentence. 32 characters, but that's for us to think about and for computers to think about. The large language models are using tokens as part of this, and the different colors that you can see here, those are representative of each of the tokens. So notice that hands is one token, the hyphen and the on are another, the space and AI are another, The space and I is are another.

Mikah Sargent [00:12:07]:
The space and M-Y are another. The space and favorite are another. The space and show are another. And lastly, the period is another. That is one example. But let's do some more. Let's go ahead and type in the word, uh, let's go with unbelievable. If I can spell that while talking.

Mikah Sargent [00:12:30]:
Check this out. Unbelievable is an unbelievable word. Because it's made up of multiple tokens. 1, 2, 3 tokens make up the word unbelievable, and then there's a little space afterward because I put that there. That's 4 tokens with the word unbelievable. Let's do a number. 1, 2, 3, 4, 5, 6, 7. And check this out.

Mikah Sargent [00:12:56]:
You may think that each of those numbers is going to be a token. No. 1, 2, 3 is a token. 4, 5, 6 is a token. 7 is a token. And it's not just English. It also can be another language as well. And so when you do this in other languages, you may get different results.

Mikah Sargent [00:13:18]:
So I have the word hands-on, or the sentence hands-on AI is my favorite show in German. I'm not going to try to pronounce that. Let's see how it compares though. Well, hands, same. On, same. Space AI, same. Is, same. My, same.

Mikah Sargent [00:13:38]:
And well, technically favorite is the same, but then things get a little weird at the end because podcast and period all are spread, are all their own tokens. So pod is its own token, cast is its own token, and period is its own token. So depending on the language that you're using, you may get a different result. And that too is unbelievable. Long or unusual words and long numbers, they break into more pieces. And of course, different languages can break into different places as well. So the rule of thumb changes based on what you're working with. But when it comes to the English language, in that case, tokens in English, it's about 1 token for every 4 characters.

Mikah Sargent [00:14:34]:
So about 3/4 of a word. So if we're looking at 100 tokens, that's about 75 words. The next time you see that rating for tokens per second, you get an understanding of how quickly the model is able to process what you're saying. Of course, that, that I showed you was the OpenAI tokenizer, and every model has its tokens set a little differently. Even Quen has its own, but of course it works the same way. So it works the same way in the sense of breaking things up by token. Let's take a look at how our local model that we've been working with thus far breaks things up. So we're going to head back to the terminal and we'll downsize all of these windows so we can get to the terminal.

Mikah Sargent [00:15:30]:
We'll open up a new terminal window here and let's stretch it out and go ahead and we're going to type in ollama because that's the tool that we use, show, quen 3.5 9B, that's the model, and verbose. This will give us a readout of some of the information that is, that makes up this model, including the tokens. We need to look for the line that says tokenizer ggml_tokens. So we'll scroll up, For the most part, it's alphabetical, but then it breaks its own little rules here. And we can find tokenizer GGML tokens and we'll see 248,317 tokens after showing us the first few. The first few tokens that make up the Qwen model are an exclamation point, A quotation mark, and a pound sign, or the octothorpe, plus 248,317 more tokens. And I remind you that tokens are bits of words, not just the entire word. Really kind of mind-boggling that 248,320 tokens.

Mikah Sargent [00:17:07]:
Tokens are what we, what it has to work with, but that's the extent of it. So when we look at that math, it means that Quen's vocabulary is about 250,000 tokens. It's, it's the vocabulary for it is not exactly equivalent to words. It's a quarter of a million pieces that it can put together to make that possible. And remember I told you to remember— remember how I told you to remember 4,096? Well, that is where this comes in. Every single one of those tokens, every single one of those 248,320 tokens gets its own starting list of 4,096 words. Numbers. You can kind of think of it as sort of like the model's first impression of that piece of text, like what it sees when it first looks at that individual piece of text.

Mikah Sargent [00:18:16]:
It could be that exclamation point, it could be the quotation mark, it could be the octothorpe, it could be something else. And it uses this to Figure out what it needs to provide a response to you. And we'll get even more into that, if you can believe it. Now, that, if we start to look at this math, means that the starting table, so sort of the tokens that are there and the potential numbers that can be used as part of these tokens, the starting table is about a billion. Billion numbers. So more than a 9th of the model's total knobs. And that's only where it starts. Because depending on where a word, and more importantly, the piece of the word exists in a given string of text, that same word can mean different things.

Mikah Sargent [00:19:23]:
So it keeps working those numbers over and over again in context. That's what context is. And that is why context— if you've ever used one of the Frontier models and it talks about compacting and it talks about how much context it has left, that's why that stuff starts to add up so quickly, because every additional bit of context means even more understanding of how the probability of what comes next can change. Do you remember when last week, if you tuned into the show last week, and I certainly hope you did, I asked Quen to count the letters for R in strawberry, the number of Rs in strawberry? And you may have heard this test run before, and you may be going, why in the world would that be so difficult for a model to do? It's because of the way that tokens work. It's not looking at a word as an individual set of characters. It's looking at a word as chunks. It sees chunks of data. It doesn't see letters.

Mikah Sargent [00:20:38]:
And so those chunks may or may not have an R in it. And that is where it can get confused. It got it right, which was great. It had though to spell the word out letter By letter first. That's the only way it was able to chunk it properly so it could figure out how many letters were actually an R. It's, it, it kind of counteracts, or for me at least, it makes me go, oh. These systems are not thinking in the same way. That we do.

Mikah Sargent [00:21:20]:
But at the same time, they're kind of reading the same way that we do. Because when I look at a sentence and I look at words, I'm not at this point, as, as someone who can read, looking at individual letters to understand what it is that I'm reading. I'm looking at chunks. It's tokens, baby, all the way down. Now we get into, if you can believe it, that all sets us up for how a large language model is able to give us an answer. Its one job is to just guess the next piece. So it's really like, this is, this is all of what we do here at Hands-On AI, if you can believe it. It needs to score every possible token to figure out everything that it's written so far.

Mikah Sargent [00:22:30]:
It then does these calculations to determine A score for every possible next token. What does that mean? Yeah, really complicated. So you put in a prompt, you're asking a question, uh, how many Rs in strawberry? It's not actually understanding what you're asking it. It's taking what you're asking. And it's also taking the context of everything it's written so far. And then it just does a bunch of calculations, and the score determines how it responds. If you can believe it, it's almost a quarter of a million guesses happening at once. And that's again why these things have to run off, uh, in, in some server, uh, room somewhere in most cases to be able to do this.

Mikah Sargent [00:23:27]:
And it's also why a 6 gigabyte file On our local Mac is kind of mind-boggling. That's a lot of calculations it's doing all at once. So it runs all these calculations, it scores them, and then it just picks one. And then it adds it to the end. And then it starts over. So when you were watching everything stream in last week, when you were watching the model get suspicious, it seemed, And talk about what it was trying to figure out. All of that thinking, all of the stuff that it was doing, it's all just calculations that are running over and over and over and over again to try and give the best possible answer. And it's often in this area, in this magical area, it seems, that the, the, the outcomes can change and can be improved.

Mikah Sargent [00:24:25]:
But I wanna kind of dig into something because we can actually take a closer look at how models are scoring things. Okay? You don't have to just have the first answer. And so I've worked out a way to show you some of the ways that these models are scoring the next possible answer. And by that, I don't even mean the next possible answer. I do mean the next possible token as a response to a question that is asked. But in this case, it's really not even thinking about the fact that it's a question. It's just taking what is being said and running calculations to give us an answer. So let's head over to macOS and to the terminal again.

Mikah Sargent [00:25:17]:
So we can take a look. First and foremost, what we're doing is running some odds calculations. And so again, this is going to give us an understanding of how Quen is rating the possible answers that we have in the background. So the first question, all of this code that I'm going to paste in here It's the prompt. The important thing is I'm asking it, I'm saying to it, the capital of France is, then its job is to just tell me what the next 5 possible answers are and rate them based on percentages. So that's what this command does. I'm going to hit enter. And we'll see what it has to say.

Mikah Sargent [00:26:15]:
So there's a 65.5% rating for Paris. Remember that the prompt I gave it was the capital of France is. 65.5% rating for the word and the token Paris. The next are A. So the capital of France is A. That only gets a 4.6% rating. The next one is the capital of France is known. That's a 2.8% rating.

Mikah Sargent [00:26:44]:
The next is the capital of France is London. That's a 2.2% rating. And the last one is the capital of France is not. That's a 1.7% rating. Now you're probably starting to get an idea of how training makes an impact, right? Because if it has all of this data, that comes into its training. Over time, those knobs get shifted and it picks up this understanding that most of the time when the tokens that make up the word or the phrase, the capital of France is, are presented, most of the time, 65.5% of the time, according to its own training, the next response is Paris. And God help us. We hope that's correct.

Mikah Sargent [00:27:33]:
Now let's do something that does not have a true answer. My favorite food is, and let's see what it has seen the most of. Well, it's not as high of a rating. 10.7% pizza, 2.8% chicken, 2.8% pasta, 2.6% uh, or the letter A, and 2.5% sushi. Hard to believe that sushi rates so high, but it's there. So with a non-true answer, I'm not surprised that pizza ranks so highly. I think back to my elementary school days and how everyone said their favorite food was pizza. But that gives you an idea.

Mikah Sargent [00:28:26]:
That this is just a matter of plausible answers, plausible responses. I hate to even say answers because these aren't really answers. It's just probability of what would come out next. But it ends up getting a little bit more complex than just what the token rating is. And that is why we talk about non-deterministic responses, meaning responses that may not be the same every time. See, the model doesn't always take the top guess. It kind of works with, with weighted dice. And that is part of the Training that takes place with the model that can make a difference in what it comes out with.

Mikah Sargent [00:29:26]:
It's the setting that controls how much those dice rolls are going to choose or favor the top guess. And what is that rating called? There's actually a term for it. It's called temperature. Temperature. Now with the Ollama build of Quen, so the Ollama version of this model of Quen, the temperature defaults to 1. And if you remember earlier when we were looking in the terminal, if you did this on your own computer, you would've seen a setting that says the temperature is 1. But I wanna show you something interesting. Okay? So we're gonna head back to the terminal and take a look at how temperature can have an impact On what we see and how the weighted dice make a difference in giving us different answers.

Mikah Sargent [00:30:23]:
So let's once again start with a new terminal session here. We'll close out of these. Whoops, no, we'll keep that open. And all of my windows look the same, don't they? We will type in First and foremost, Ollama run quen-3.59b. So that's the model we've been working with this whole time. And now we've got the model running. This is the same thing as running it in the Ollama app, but we're just doing it on the terminal this time. And what I'm doing first and foremost is typing in /set.

Mikah Sargent [00:31:07]:
No think. And what this does is it keeps the model from doing that thinking that we saw last time, that thing that led to it going, oh, it could be try— the, the, the user might be trying to trick me. The user might be asking questions. And it gives us an answer pretty quick. Now I'm gonna type clear, and that sort of clears the session context. It means that the model doesn't have anything previously to rely on. And I wanna do a simple sentence. Suggest one name for my new fish.

Mikah Sargent [00:31:45]:
And then I'm gonna say just the name. So that way it knows exactly what I'm asking for. Bubbles. Then I'm going to type in clear because again, that clears the context. And then I'm going to say the same thing again. This time it said Barnaby. Now we're gonna type in clear one more time, removing the context. And now once more, and this time it said Bubbles.

Mikah Sargent [00:32:10]:
So twice it said Bubbles, once it said Barnaby, and that's with the current settings. But now we're going to do something interesting. We're going to choose /set parameter. Temperature 0. Remember, 1 is the temperature where the weighted dice plays more of a factor, or it is where the model may not favor the most common or most chosen answer. But with 0, what it's doing is making a choice that is more likely to be the, the result that has the most occurrences. So I've set that. Now we're going to clear one more time, and then we are going to suggest the name.

Mikah Sargent [00:33:18]:
We've done that once. We'll do clear, suggest the name, do that once. We'll do clear and then suggest the name. Now, because we've set the temperature to 0, it's choosing the top-rated answer every time. And what is that? It's bubbles. 3 times I asked it, 3 times it told me bubbles. That Is what temperature is all about. So temperature, when it's set to 1, you've got the weighted dice where different names or different outcomes are likely.

Mikah Sargent [00:34:00]:
Well, again, honestly, weighted dice, think about it more as, um, It's sort of non-weighted dice. When a temperature is set to 1, right, it is saying, let's make it more random. But when it's set to 0, that is where the weighted dice come in because it's going to choose the top pick every time. So same prompt, exact same settings, meaning the context hasn't changed. You're going to expect a much more repeatable answer. And in this case, bubbles, bubbles, bubbles was the answer. And that frankly is one reason that a chatbot can answer differently when you ask the same question twice. It's not that the model changed its mind, it doesn't have one.

Mikah Sargent [00:34:59]:
It's just that a lot of the time that dice roll Is what makes the difference. And a lot of the time you don't have control over the temperature. So going back to the knobs, remember I told you that we don't directly go in and boop, boop, boop, adjust those knobs. We gotta ask ourselves, where do those knobs come from? Who sets the 9 billion knobs? It's the training. When the training run happens, it does so. It starts with the knob set to basically random numbers. And if you were to talk to a model before any of this pre-training, well, the model is going to just respond with a bunch of random tokens, not even random words, random tokens. So Pre-training comes in to show it text, and by text we mean mountains and mountains and mountains of text.

Mikah Sargent [00:36:06]:
It predicts the next tokens based on the text that is coming in. It compares its guesses with the real text that it's provided, and then it makes slight adjustments to the numbers, the knobs that it has, so that it gets a little better. And it does this over and over and over and over again across trillions, remember, of tokens. That is what pre-training is. So what you get in the end, well, it's a model that's really good at finishing documents. It's a model that if you give it the start of a recipe, then it's going to probably give you the rest of the recipe. It's looked at so many recipes that it knows what's going to come next. If you ask it a question, depending on the training, it might answer the question or it might just write out more questions.

Mikah Sargent [00:37:05]:
And that's part of where the thinking comes in, right? That is where it may have been trained on a page full of questions, but it's post-training that really helps to narrow things in. It's what really helps to sort of shape shave off the edges and take something that, you know, finishes the words in a sentence or accidentally asks a bunch more questions and actually makes it the helpful models that we see today. Post-training, that can include practicing on example conversations. It definitely includes getting feedback from human beings and other AI models on its answers. And that is what makes it actually act like an assistant instead of just being sort of a sentence finisher, even though it's technically still a sentence finisher, a token finisher. Quen does have both types. There's a base model of Quen that you can use, and that one has no post-training at all. Now I want to give you, because this blew my mind, a sort of sense of scale.

Mikah Sargent [00:38:27]:
Back in 2023, Andrej Karpathy, who's of course one of the founding members of OpenAI, talked about the rough numbers for, it was Meta's at the time, Llama 2. Which had 70 billion tokens. In order to train that model, Andrej said it was 10 terabytes of text. That's just text. If you've ever saved a text file that's like, you know, thousands and thousands of words and saw how small that file was, you know that to make 10 terabytes of text is a whole heck of a lot. So 10 terabytes of text, 6,000 GPUs, 12 days of training, roughly $2 million. And all of that work, all of that training boiled down into a 140-gigabyte file. Andrej compared it to lossy compression like a JPEG.

Mikah Sargent [00:39:39]:
I would argue it's quite a bit of lossy compression for For sure. Now, when the training is done, when the model has been pre-trained and post-trained, the models that we're using, it kind of goes in and puts locks on all of the knobs. Those knobs, they don't move anymore. So you, when you're talking to an AI, You aren't actually training it. Chatting with the AI system does not retrain the base model. And that's where people sometimes get confused. But typically what's happening in those instances is the model is simply pulling in context From its own memory and adding it to the tokens that it has and finding new patterns that would— new patterns of tokens that would make for a much more satisfying for us because it's trained to be more satisfying outcome. So it's not as if you are able to go in and Adjust those knobs left and right unless you're working with open source, open weight models, and you can make some adjustments there.

Mikah Sargent [00:41:14]:
So when one of these cloud systems remembers you, it's a different trick other than you training the basic model. It's sort of something that gets layered on top, and different systems do it in different ways. This is the part where people go, okay, cool. So we've got this system and really all it's doing is running a bunch of calculations, ridiculous calculations, incredible calculations based on thousands and thousands of gigabytes of data and training and time and money. But still, it's just doing some predictions. So why does it sound so sure? Tokens in, one guess at a time on what should come next, and then weighted dice to decide if it chooses the top answer or something else. And those weighted dice, of course, being weighted based on the training. It's interesting because one might think that there's some sort of step that is going to look for an answer somewhere and use that as a means of determining if what it's saying is true.

Mikah Sargent [00:42:44]:
But what sounds plausible Can be true, but it can also be false, right? It— the model doesn't know the difference. It just knows what is most likely to be the next outcome based on what it's been given. So if a lot of people, if a lot of text and a lot— it's— this is the thing too. It's not exactly that. a bunch of humans have said this because the training is not necessarily just that. But if its patterns in its training suggest that the next answer is, the next token is this, and maybe that's 3 tokens that make up the word, then that's the response that it gives. And sometimes, again, it doesn't choose that top one. So you may remember last week I asked for the week's news, right? And it told me it didn't have access and told me instead that there were some great sources to check out.

Mikah Sargent [00:43:49]:
That's the outcome we want, right? That's what we want. But it doesn't always work that way, and some models aren't trained to do so. So there are times where you can kind of push a model to make a difference. So Let's say you had a fresh chat, you had your Wi-Fi off, and you asked it, what was the first episode of the TWiT podcast, Hands on AI, about? And the model, of course, came out months before the first episode. So without a source in the chat or a search tool, it's not able to look the episode up. It may or may not have an answer for us. And so I was curious to find out what that might be. So let's see what it says.

Mikah Sargent [00:44:41]:
Let me actually head over to macOS and we'll take a look. Again, this is a model that has some training and we're gonna clear out this session. Actually, we'll start a new session because this one currently has the temperature set to— 80. 0. So we'll do ollama run quen-3.59b, and then we will set the model to nothink, and then we'll ask that question. What was the first episode of The TWiT podcast, Hands on AI, about? The first episode of TWiT Hands on AI, which debuted in February 2023, featured host Leo Laporte interviewing Dr. Oren Etzioni, the CEO of Allen Institute for AI. The primary topic of the conversation was the launch and mission of the Hands on AI podcast series itself.

Mikah Sargent [00:45:58]:
Specifically, they discussed: the origin story, the goal of the show, the series structure. Following this introductory episode, the show went on to feature interviews with various experts such as Drew Houston, the founder of Dropbox, discussing privacy-preserving machine learning, and other leaders from companies like Google and Amazon focusing on specific AI tools. By goodness and by golly, that sure does sound like a very real thing. So much so that I'm going to look that up to make sure that this show did not exist back in the day. And that's the point, right? That's the point. Is that it can say things that feel so true and so sure. In this case, I'm not seeing any show from back then that Leo hosted that was called Hands on AI. But without that thinking, without that extra step, it was able to give us this very sure answer and not get suspicious seemingly about what I was asking.

Mikah Sargent [00:47:11]:
See, part of that thinking helps the model to come up with a better response. You kind of want it to say it doesn't know, not something that sounds so very true given all of the facts that it has there. Post-training helps a model to not do that. OpenAI has published some research on this saying that, Look, the fact is models get rewarded for guessing. And so if you get rewarded for guessing, then you're more likely to guess. It's a multiple choice test, which means that you could get lucky and you could get it right. And the point of these models, the thing that is reinforcing their behavior is for them to get it right. So not having an answer Is something that in many cases is not seen as the right answer.

Mikah Sargent [00:48:13]:
So what can you do? Well, one thing that I like to say to my agents is, you know, some form of, I don't have the answer to this, in quotes, is a perfectly acceptable response. I will tell it that it's okay to not know if it doesn't know. That's only one possibility though. It, it, it sometimes that works, sometimes it doesn't. And there are times where that matters more. So giving it the source is something that can help. If you give it the source like that email summary as we did, it worked because the email was right there in the prompt, but it's still worth checking. Also, if the prompt is more obscure or if the prompt is more recent, well, it's going to have to do more guessing if it doesn't have access to the information with a local model.

Mikah Sargent [00:49:09]:
So that isn't— wasn't in its training. It didn't have access and it couldn't add that to its context to give us a better answer. And just keep in mind that just because a model is incredibly confident does not mean that it's correct. A confident tone Be it from a human being or from an AI chatbot, does not tell you anything about its veracity. It just tells you how the sentence sounds. Folks, again, I have to go back to this because it blows my mind. I'm really not trying to patronize. I just think it's incredible.

Mikah Sargent [00:49:50]:
What is in that file that we looked at at the beginning of it? 9 billion numbers, a vocabulary of nearly a quarter million tokens, and then the training that gives it its job to guess the next piece of the puzzle. Not the next word, the next piece. That is the core loop behind every word that we saw it write today. Now, next week, of course, you know that for today's show, everything that we did was all about the words, right? And understanding what makes up a large language model and how it operates. So it was all about text in, tokens in, text out, tokens out. You may remember for the first episode I mentioned Muse, Meta's new AI agent. Well, next week we're going to put it to work. It's an incredibly popular model among more people, and that means that we should take a look.

Mikah Sargent [00:50:56]:
You know, it's an agent that doesn't just answer you, it actually goes out and does the thing, the thing that you're asking it to do. So I'm gonna give it some simple jobs, some hard jobs, and even some weird tasks so we can kind of see what it's capable of doing. In the meantime, if you'd like a little bit of homework, go ahead and run ollama show on any model that you've downloaded and then find its parameters and its quantization. Now you know what those mean, or perhaps you already did, but you can get a better look at your specific models. Then paste your own name or a sentence that you often use into the tokenizer to find out how many pieces You're seen as when it comes to an AI model, in that case from OpenAI. In the meantime, send your AI questions to handsonai@twit.tv and you just might hear yours answered on the show. Please do hit subscribe or follow wherever you're watching or listening. Head to twit.tv/HOAI.

Mikah Sargent [00:52:03]:
And if you'd like this show ad-free, well, that's clubtwit, twit.tv/clubtwit. I am and will continue to be Micah Sargent. Don't just wonder about AI, get your hands on it. Thanks for tuning in. I'll see you again next week.

 

All Transcripts posts