BASUDEV_OS

2026-08-22

Understanding The Term 'LLM'

Understanding The Term 'LLM'
If you are interested in the field of AI/ML, You probably have learned or heared about the term LLM which stands for Large Language Model. In this Blog, We will be understanding the term LLM in detail. **Large Language Model** Breaking words from this terms we can generate a rough idea regarding what it is. **Large :** It refers to the scale of the model which include its parameter,training data. **Language :** It means patterns, relationship, structure, information in human written text. **Model** : It is a mathematical terms also known as a system which generates output from input. **INTORDUCTION TO LLM** To make it easy imagine LLM as a person who has read the lot of things. He has not memorized but knows about patterns, relationship among them and has become capable to complete the sentence from half sentence based on probability. Getting deeper LLM is usually a Neural Network based on Transformer architecture which is trained on massive text . It holds a capacity to predicts based on probability distribution given a previous sequence of token like p(token n| token 1,...... token n-1). It is trained with billions of paramteter( weights). It requires no labelling to predict next text but can predicts based on token probability without requiring human to manually label. Before going deeper into this lets understand term Token, Token can be a whole word or can be part of words, character or can even be a space. Lets dive into things happen when you give a prompt to AI models like claude, chatgpt. * **TOKENIZATION** In this raw text splits into token which is a small units so AI can process it easily. like for input " I Love Maths." Its get divided in tokenization process as \["I" , "love" , "Maths" "."] Next important things happens in tokenization process is each token is mapped to token ID which is a number. Like for i love maths. I->45 Love ->271 Cats-> 1203 this depends entirely on model's vocabulary. * **EMBEDDING** This is a way of converting token ID from tokenization process to a vector of number. Flowsheet till now ```markdown TEXT ↓ ↓ Token ↓ ↓ Token ID ↓ ↓ Embedding ``` Embedding is mandatory as it helps to find relation and pattern between tokens like Embeddings for Cat, Dog, Cow has similar pattern as they fall in same category Animal while cat,math has less similar pattern like Embedding for cat and dog "cat" → \[0.21, -0.73, 0.45, 0.18, ...] "dog" → \[0.19, -0.69, 0.48, 0.22, ...] * **SELF ATTENTION** This process allows each token to look at other previous token which help to understand the context and helps to find which is important for new token. Lets take an example for a statement " I like Python because It is needed for AI/ML." Here to understand what 'it' means model look to other token of text and find it means Python not other word here. #### Attention Score It needs to understand what attention score means and how it helps in predicting another token. For every token model creates Vector with 3 things Query(Q): What infomation I am looking for Key(K) : What information I already have Value(V) : What information to pass ahead. Using this attention Score is calculated * **FEED FORWARD PROCESS** It process information which is gathered by self attention. * **RESIDUAL CONNECTION** This keeps input representation while adding transformed output. * **LAYER NORM** It keeps resulting information Well Scaled. * **OUTPUT LAYER** Converts information learned by transformer into a final prediction.Transformer Process the input and creates a final representation which is taken by Output Layer and produces score for next token like for Capital City of Nepal is ``` TOKEN SCORE SOFTMAX KATHMANDU 8.5 95% BUTWAL 2.1 2% . . . . . . . . . . . . LONDON 0.1 0.1% ``` Softmax converts raw Score into probability. ![TRANSFORMER](https://miro.medium.com/v2/resize:fit:1200/1*eNYtdGpIaGwd8KWCwLUQ9w.png "TRANSFORMER BLOCK") * **Autoagressive decoding** : This picks next token. This can be done by either 1. **Greedy Decoding**: Choose with highest probability like ``` TOKEN PROBABILITY Butwal 41% Pokhara 30% Kathmandu 21% ``` Here In Greedy Decoding it chooses Butwal 1. **SAMPLING** Here next token is randomly choosed according to probabilities like not only choosing highest probability but choosing least also but there is less probability in choosing from least probable event but high chance of choosing among high probability.This result in more variation. 1. **TEMPERATURE** This is another Key aspect which decides which token is to be choosen If temperature is Low--------> Prefers High probability token High---------> Prefers Low probability token Thus high temperature brings more variation. 1. **TOP K SAMPLING** Top K means consider only top k token For Example ``` IF k=3 TOKEN PROBABILITY School 77% Library 20% Resturant 1% Stadium 0.5% Here while choosing next token it considers only top 3 and ignores rest and next token is from first 3 ``` TOP P SAMPLING This keep enough of most likely token until total probability reach P For example ``` IF P=0.90 TOKEN PROBABILITY SCHOOL 70% LIBRARY 9% HOME 11% RESTURANT 8% Here If we combine probability of first 3 it becomes exactly 90 percent so it stops and ignores 4 th as p is 0.90 which means 90 percent ``` * **APPEND** After choosing token, it add that token to existing text or para. * **REPEAT** Model does same thing again until it decides to stop. This is how it gives response for given text . This is every activity that happens inside LLM Model from giving prompt to generating a response.