Pinned Loading
-
Transformer-Pretraining-and-Interp
Transformer-Pretraining-and-Interp PublicI Pretrained a Transformer of 134M params from Scratch on around 1.2 Billion Tokens. Now, I am Applying Interpretability techniques to Understand Internals.
Python
-
GPT-Pretraining
GPT-Pretraining PublicA 134M-parameter GPT written from scratch in PyTorch, no nn.Transformer and no HuggingFace, pretrained on a 7B tokens corpus of FineWeb-Edu, where it consumed 1.2B tokens on one free Kaggle P100. H…
-
Coding-Transformers-using-pytorch
Coding-Transformers-using-pytorch PublicBreaking down the Attention Is All You Need paper into clean, modular PyTorch code — making the Transformer architecture accessible to learners and researchers. Full Article: https://medium.com/@Um…
Jupyter Notebook
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.