cjzdaily
6 days ago
https://x.com/amitiitbhu/status/2082489293852008549?s=12
X (formerly Twitter)
Amit Shekhar (@amitiitbhu) on X
- Math behind Attention- Q, K, and V
- Math behind √dₖ Scaling Factor in Attention
- Math Behind Backpropagation
- Math Behind Gradient Descent
- Math Behind Cross-Entropy Loss
- Math Behind RoPE (Rotary Position Embedding)
- RMSNorm (Root Mean Square Layer…
Home
Powered by
BroadcastChannel
&
Sepia