Hook

Agar main aap se poochun ke crocodile, lizard aur cat in teenon creatures mein se sabse zyada similar kaun si hain? Toh aap kahoge crocodile aur lizard similar hain, cat inse different hai. Isko agar main inki visually represent karun ek vector ke upar toh main is tarah se karunga. In dono ko main similar saath rakhunga aur cat ko main thoda sa side pe rakhunga. Agar main aap se poochun ab snake ko batao ke snake in mein se compare karo kiske zyada similar hai? Toh aap kahoge snake iske zyada inse bhi similar nahi hai aur iske saath bhi similar nahi hai. Toh snake ko aap us vector ke upar agar place karna chahoge toh aap nahi kar paoge. Uske liye aapko alag vector create karna padega aur udhar aap snake ko jo hai woh rakhoge. Isi tarah main parrot ke baare mein poochunga, parrot ke baare mein toh parrot bhi inse different creature hai. Toh uske liye main na isko is vector pe place kar sakta hun na is, iske liye mujhe kya karna padega? Mujhe different vector create karna padega. Isi tarah agar main crow ke baare mein aap se poochunga toh crow jo hai woh parrot ke similar tha toh maine isko crow parrot wale vector ke upar jo hai woh place kar diya. Theek hai? In dono ke andar nahi kiya. Toh reason yeh hai ki ab kya hoga ke ab jitni bhi zyada creatures honge, jitni bhi zyada animal honge unko main ek ek karke is vector ke andar, is vector store ke andar main place karta jaunga. Agar toh woh similar honge, theek hai toh unko main ek hi vector ke andar add karta jaunga. Agar woh different honge toh main unke liye alag alag vector jo hai woh create karta jaunga. Is tarah hum thousands of vector jo hai woh create kar lenge. Theek hai? Toh yeh saara process karne ka fayda kya ho raha hai? Isse yeh ho raha hai ki hum jo hai woh similarity find out kar sakte hain. Ab yahan pe ab dekho ke in dono creatures ke darmiyan jo hai woh less distance hai. Less distance ka matlab yeh hai ki less more similarity. Means ki yeh dono aapas mein zyada similar hain aur yahan pe agar main aapko dikhaun toh yeh snake aur cat inke darmiyan maine jo hai woh line laga ke likha hua hai more distance. Jinka distance zyada hoga iska matlab yeh hai ki unke darmiyan jo hai woh less similarity hogi. Yahi aapki embedding hai. Embeddings kya karti hain ki woh aapka jo data set hota hai usko woh different vectors ke andar jo hai woh represent kar deti hain. Har vector ke andar jo data hota hai woh closely related to each other hota hai. Theek hai? Woh more similar hota hai. Unke darmiyan distance jo hai woh kam hota hai aur jo dusre vectors ke upar data hota hai woh pehle wale vector se jo hai woh different hota hai. Theek hai? Toh yeh karne ka fayda kya hai? Embeddings toh humein pata lag gayi lekin embedding ka fayda kya hota hai? Aapne kabhi notice kiya hoga ki jab bhi aap chat GPT se query karti ho toh woh aapko usi ke related kyun answer deta hai? Aap usse poochoge ki crocodile ke baare mein mujhe batao toh woh aapko hamesha crocodile ke baare mein hi batayega. Woh cat ke baare mein nahi batayega, woh snake ke baare mein nahi batayega. Reason yahi hai ki jab aap usse query karte ho toh woh kya karta hai woh just randomly data search nahi karta hai. Woh jis data set ke upar train hua hota hai woh just us mein randomly se search nahi karta. Woh kya karta hai ki woh us vector ke andar jaata hai jiske vector ke andar crocodile ka data pada hota hai taake woh udhar se jo hai woh similar data jo hai woh fetch out kar paaye. Kyunki humne baat ki hai ki jis vector ke andar crocodile hoga iska matlab yeh hai ki us vector ke andar crocodile ke related hi information hoga. Theek hai? Yahi toh vectors banane ka fayda, yahi toh embeddings ka fayda hai. Toh jab bhi chat GPT ko aap koi bhi query karoge ki yaar crocodile ke baare mein mujhe batao toh hamesha woh crocodile ki information isi wajah se deta hai correct ki woh jo hai woh jis vector ke andar search karta hai udhar jo hai woh saara data crocodile ka hi hota hai. Theek hai? Woh kabhi bhi aapko cat ki information nahi dega, kabhi bhi aapko snake ki information nahi dega. Toh jab hamare LLM train ho rahe hote hain toh woh kya karte hain woh basically jo data set hota hai un data set ka isi tarah embeddings create kar lete hain. Theek hai? Jis tarah se agar main aapko dikhaun yahan pe main agar aapko scroll karke dikhaunga toh agar main inko word ke andar difference kar dun inka main snake ka word likh dun aur crocodile likh dun, cat dekhen toh yeh suppose karo yeh hamara data set hai aur jab bhi LLM train hota hai toh woh is data set ko in vectors ke andar represent kar deta hai. Is vectors ke andar jo hai woh convert kar deta hai. Toh yeh basically hamare jo hain woh embeddings create ho jaati hain. Jab yeh embedding create ho jaati hai toh jab bhi LLM humein response deta hai toh woh inhi vector store ke andar aata hai. In embedding ke andar search out karta hai aur most similar result aapko find out karke de deta hai. Ab aap usse parrot ka poochoge toh woh aap se parrot ka jo hai woh parrot wale vector ke andar jayega aur udhar se jo hai woh aapko accurate information fetch karke laake de dega. Theek hai? Toh tabhi aapka LLM jo hai woh yahan pe maine likha hua hai ki this is the only reason LLM generate accurate answers. Yahi answers jo hai woh accurate is wajah se deta hai ki woh jo embeddings create karta hai woh un vectors ke andar similar data pada hota hai. Toh this is the only reason aapka LLM jo hai woh itna accurate answer kyun deta hai.
Their other posts in the index, biggest breakout first.