Hook
More breakout videos from this creator.
OpenAI says their new model GPT-6 Astra has 2,055,844,907,407 parameters. Hypothesis -- not an official specification. Start with one learned value. AI with Prof. Whitaker. Parameters. Unfortunately, mathematics. X1w1 + X2w2 + X3w3 + b. A short course in learned numbers. First, a correction. A = B. FACT. FILE. [ ]. MEMORY. </>. CODE. ≠ PARAMETER. A parameter is a learned number. A mathematical function. Input. Computation. Prediction. TEXT. VECTORS. TRANSFORM. TRANSFORM. PROBABILITIES. MAT 72%. FLOOR 14%. CHAIR 6%. Illustrative; other tokens omitted. One tiny neuron. Σ neuron. Three inputs. One computation. Weights. And bias. X1 w1. X2 w2. X3 w3. b. Σ neuron. The x values are inputs. The w values and b are learned. Numbers change influence. X1 w1 = +1.00. X2 w2 = +1.00. X3 w3 = +1.00. b. Σ neuron. Hold x1 = x2 = x3 = 1, b = 0. Weighted sum = +3.00. Illustrative input and weight values. Same inputs. Different influence. X1 w1 = +2.14. X2 w2 = +1.00. X3 w3 = +1.00. b. Σ neuron. Hold x1 = x2 = x3 = 1, b = 0. Weighted sum = +4.14. Illustrative input and weight values. Same inputs. Different influence. X1 w1 = +2.40. X2 w2 = +1.00. X3 w3 = +1.00. b. Σ neuron. Hold x1 = x2 = x3 = 1, b = 0. Weighted sum = +4.40. Illustrative input and weight values. Same inputs. Different influence. X1 w1 = +2.40. X2 w2 = +0.50. X3 w3 = +1.00. b. Σ neuron. Hold x1 = x2 = x3 = 1, b = 0. Weighted sum = +3.90. Illustrative input and weight values. Same inputs. Different influence. X1 w1 = +2.40. X2 w2 = +0.01. X3 w3 = +1.00. b. Σ neuron. Hold x1 = x2 = x3 = 1, b = 0. Weighted sum = +3.41. Illustrative input and weight values. Same inputs. Different influence. X1 w1 = +2.40. X2 w2 = +0.01. X3 w3 = -1.30. b. Σ neuron. Hold x1 = x2 = x3 = 1, b = 0. Weighted sum = +1.11. Illustrative input and weight values. Same inputs. Different influence. Not hand-tuned. w1 w2 w3 w4 w5 w6 w7 w8 w9 w10 w11 w12 w13 w14 w15 w16 w17 w18 w19 w20 w21 w22 w23 w24 w25 w26 w27 w28. NOT MANUALLY PROGRAMMED. Thankfully. We have weekends. From weights to matrices. W1 +0.310. W2. W3. One reusable idea, at a much larger scale. W1 +0.310 -0.820 +0.140 +1.070 +0.030. -0.410 +0.006 -0.028 +0.013 +0.225. +0.715 -0.102 +0.333 +0.009 -0.074. WEIGHT MATRIX. vector X W new vector. Inside a transformer. TOKEN → EMBEDDING. ATTENTION. Learned values throughout the computation. TOKEN → EMBEDDING. ATTENTION. FEED-FORWARD. OUTPUT SCORES. MAT 72%. FLOOR 14%. CHAIR 6%. Illustrative; other tokens omitted. Learned values throughout the computation. 70 billion... what? 70B PARAMETERS. learned numerical values. +0.310. Values used in computation. +0.310 -0.820 +0.140 +1.070 +0.030. -0.410 +0.006 -0.028 +0.013 +0.225. +0.715 -0.102 +0.333 +0.009 -0.074. Values used in computation. The wrong mental model. WRONG MENTAL MODEL. #8,482,193. There is no one-fact-per-slot map. #8,482,193 Earth orbits Sun. #15,821,992 Python uses indentation. Distributed influence. PARIS / FRANCE. Conceptual illustration, not a mechanistic map. Many values. Overlapping effects. GRAMMAR. Conceptual illustration, not a mechanistic map. Many values. Overlapping effects. GRAMMAR. FACTUAL ASSOCIATIONS. Conceptual illustration, not a mechanistic map. Many values. Overlapping effects. STYLE. CONCEPTS. Conceptual illustration, not a mechanistic map. Many values. Overlapping effects. STYLE. RELATIONSHIPS. CONCEPTS. PATTERNS. Conceptual illustration, not a mechanistic map. Many values. Overlapping effects. Where do the numbers come from? +0.310 -0.820 +0.140 +1.070 +0.030. -0.410 +0.006 -0.028 +0.013 +0.225. +0.715 -0.102 +0.333 +0.009 -0.074. TRAINING. Training. Start almost at random. +0.012. INITIAL WEIGHTS. Small values. An untrained function. +0.012 -0.033 +0.006 +0.043 +0.001. -0.016 +0.000 -0.001 +0.001 +0.009. +0.029 -0.004 +0.013 +0.000 -0.003. +0.001 +0.017 -0.021 +0.008 +0.027. INITIAL WEIGHTS. Then: predict text. Small values. An untrained function. A prediction. The capital of France is... Other tokens: 30% (not shown). The training data supplies the target. London 35%. Paris 20%. Rome 15%. Other tokens: 30% (not shown). TARGET: PARIS. The training data supplies the target. Measure the error. PREDICTION. TARGET: PARIS. COMPARE. Natural-log cross-entropy. A number the training system can minimize. LOSS = 1.61. -ln(0.20) = 1.61. Natural-log cross-entropy. A number the training system can minimize. Follow the gradient. WEIGHTS. LAYERS. LOSS. +0.310 -0.820 +0.140. Backprop computes gradients. The optimizer updates weights. GRADIENTS. +0.310 -0.820 +0.140. +0.312 -0.817 +0.139. OPTIMIZER: SMALL UPDATE. Backprop computes gradients. The optimizer updates weights. Repeat. PREDICTION. LOSS. BACKPROP. GRADIENTS. UPDATE. Many small updates. PREDICTION. STEP 1. Illustrative iteration counter. Many small updates. PREDICTION. STEP 10,000. Illustrative iteration counter. Many small updates. PREDICTION. STEP 1,000,000. Illustrative iteration counter. Many small updates. Training changes numbers. MEMORY SLOT #4,000,000. PARIS. +0.310 -0.820 +0.140 +1.070 +0.030. -0.410 +0.006 -0.028 +0.013 +0.225. +0.715 -0.102 +0.333 +0.009 -0.074. SMALL NUMERICAL UPDATES. Paris 20% → Paris ↑. Better predictions through small adjustments. +0.312 -0.817 +0.139 +1.072 +0.033. -0.411 +0.008 -0.025 +0.012 +0.227. +0.718 -0.103 +0.335 +0.012 -0.075. SMALL NUMERICAL UPDATES. Paris 20% → Paris ↑. Better predictions through small adjustments. Patterns take shape. Distributed patterns, not separate filing cabinets. GRAMMAR. Distributed patterns, not separate filing cabinets. GRAMMAR. FACTUAL ASSOCIATIONS. Distributed patterns, not separate filing cabinets. STYLE. CONCEPTS. Distributed patterns, not separate filing cabinets. STYLE. RELATIONSHIPS. CONCEPTS. PATTERNS. Distributed patterns, not separate filing cabinets. Capacity. FEWER CONTROLS. MORE CONTROLS. Illustration of flexibility, not a performance test. More adjustable degrees of freedom. Bigger ≠ better. 500B > 100B. Parameter count is not a performance score. 500B ≠ 100B. ARCHITECTURE. DATA. OPTIMIZATION. COMPUTE. POST-TRAINING. INFERENCE. CONTEXT. COUNT ≠ PERFORMANCE. Parameter count is not a performance score. Mixture of Experts. TOKEN. ROUTER. EXPERT A. EXPERT B. EXPERT C. EXPERT D. EXPERT E. EXPERT F. EXPERT G. EXPERT H. ILLUSTRATIVE EXAMPLE. Some models route tokens through selected components. Total ≠ active. 500B TOTAL. ILLUSTRATIVE EXAMPLE. Wonderful. Another number. 500B TOTAL. TOTAL PARAMETERS ≠ ACTIVE PARAMETERS. ILLUSTRATIVE EXAMPLE. Wonderful. Another number. Numbers need storage. Before runtime overhead, KV cache, etc. Raw weights only. Runtime needs more memory. 70B parameters × 16 bits. 16 bits = 2 bytes. 70B × 2 bytes ≈ 140 GB of weights. Before runtime overhead, KV cache, etc. Raw weights only. Runtime needs more memory. Less precision. PARAMETERS = 70B. 16-BIT. 1 0 1 0 1 0 1 0. 1 0 1 0 1 0 1 0. 4-BIT. 1 0 1 0. fewer precision levels. QUANTIZATION. The number of parameters stays the same. Same count. Less storage. 70B 16-bit ≈ 140 GB. SAME COUNT: 70B. Weight storage approximation. 70B 4-bit ≈ 35 GB. Return to the number. 70,000,000,000 PARAMETERS. FACT? FACT? FACT? FACT? FACT? FACT? TRANSFORM → PREDICT. Numbers shaping a mathematical function. +0.310 -0.820 +0.140 +1.070 +0.030. -0.410 +0.006 -0.028 +0.013 +0.225. +0.715 -0.102 +0.333 +0.009 -0.074. TRANSFORM → PREDICT. Numbers shaping a mathematical function. Learned numbers. LEARNED NUMBERS. SHAPE COMPUTATION. SHAPE BEHAVIOR. Parameter count ≠ fact count. A rough measure of learned capacity. Parameter count ≠ fact count. PARAMETER COUNT ≠ FACT COUNT. A rough measure of learned capacity. Parameter count ≠ fact count. A great many knobs. +0.310 -0.820 +0.140 +1.070 +0.030. -0.410 +0.006 -0.028 +0.013 +0.225. +0.715 -0.102 +0.333 +0.009 -0.074. +0.016 +0.415 -0.521 +0.197 +0.682. A GREAT MANY KNOBS. AI with Prof. Whitaker.