Hook

Everyone in machine learning glazes Relu like it's the goat activation function,nbut there's one place sigmoid just cooked it,nand it's quite literally in your face.nSo quick recap for context.nEveryone ditched the sigmoid function for Relu because Relu's derivative is a step function,nso there was no vanishing gradients and it was really fast.nThe sigmoid function, on the other hand,nwas a smooth S curve, and the derivative is a bell curve.nSigmoid: Vanishing Gradient ProblemnBasically forced nodes to stay on even if they weren't contributing to training.nNow here is your first plot twist.nFace recognition models encode your face as a vector in high dimensional space.nTwo components to a vector,nit's magnitude and its direction.nMagnitude was utterly uselessnbecause you could have a photo of you be absolutely high resolution,nwhich gives you a large magnitude,nbut you can have a blurry photo of you,nwhich makes a shorter magnitude.nEven though they're both the photo of you,nthey have different magnitudes,nwhich causes noise. So S facenand this whole family of modelsnL2 normalizes everything.nEvery face embedding gets projected onto a unit hyper sphere,nbasically think of a 3d unit circle in trig :)nradius 1 always.nNow that the magnitude is gone,nthe only thing that matters is the angle between the vectors.nIf they are the same person,nthe vectors point in the same direction.nAnd I'm not saying that these angles are absolute,nby the way. They're mainly a range so you could have two vectors that are withinnlike 5 degrees of each other,nand that's enough.nNow you've got all your face and embeddings living on the surface of a hyper sphere,nand your job is push the same identity vectors closer together angularlynand push the different identity vectors further apart.nIt's also known as a minimizingnintrclass angle and maximizingninterclass angle.nIt sounds clean, but real trainingndata is incredibly messy.nYou'll have mislabeled images,nmotion blur, low resolutionnselfies.nAnd very strict optimizers will justnoverfit to the noise.nThey'll drag a blurry photo towardnthe wrong identity clusternwith full confidence. So in 2021,nzhong et al dropped SFace,nwhich is a sigmoid constrainednhyper sphere lossnHoly buzzword, the fix was to notnoptimize those angular distancesnall the way,nbut do it moderately with nuance.nAnd the things that provided thatnnuance were two sigmoid rescalenfunctions,none throttling intraclass and onenthrottling interclass optimization.nRemember the bell curvenderivative that everybodynclowned?nThat is literally the whole point.nIt naturally slows down whennembeddingsnare already close to where theynshould be on the sphere,nand ramps up when they startndrifting.nThis was very smooth, verynadaptive,nand critically, noise resistant.nReLU's step function derivative,non the other hand, would havenjust said,npush harder, but sigmoid said,nactually, we're kind of good here.nSo the end move was to strip outnthe magnitude of the vector.nSo Only the direction mattered.nAnd then use then"weak" activation function tonmoderate how hard you chasenthat direction.nThe Machine LearningnCommunity spent years sayingnsigmoid was too soft,nblah, blah, blah.nBut face recognition said wenneed soft and we need smooth.nAnd guess what? It won.nS face outperformed arcface,nCos face, blah, blah, blah.nMega face, the underdognactivation function won thenbenchmark.
Their other posts in the index, biggest breakout first.