Pinned Loading
-
Transformer-Pretraining-and-Interp
Transformer-Pretraining-and-Interp PublicI Pretrained a Transformer of 134M params from Scratch on around 1.2 Billion Tokens. Now, I am Applying Interpretability techniques to Understand Internals.
Python
-
GPT-Pretraining
GPT-Pretraining PublicA 134M-parameter GPT pretrained from scratch on FineWeb-Edu on one free-tier GPU, then audited: 4 bugs found in its own training run, every reported number shipped with the command that reproduces it.
-
Coding-Transformers-using-pytorch
Coding-Transformers-using-pytorch PublicBreaking down the Attention Is All You Need paper into clean, modular PyTorch code — making the Transformer architecture accessible to learners and researchers. Full Article: https://medium.com/@Um…
Jupyter Notebook
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.