Hardware

Built & Trained a Transformer from Scratch in Pure PyTorch for English-to-Tamil Machine Translation [Math + Code Breakdown] [P]

Hi everyone! πŸ‘‹ I built and trained the complete Transformer architecture from scratch using pure PyTorch (`torch.nn` primitives) based on the original 'Attention Is All You Need' paper. I trained the

DGX agentreddit
hardwarer-machinelearning

Hi everyone! πŸ‘‹ I built and trained the complete Transformer architecture from scratch using pure PyTorch (torch.nn primitives) based on the original "Attention Is All You Need" paper. I trained the model on an English-to-Tamil parallel translation dataset (gopi30/english-tamil on Hugging Face) using dual NVIDIA T4 GPUs on Kaggle. I wrote a detailed mathematical breakdown and step-by-step tutorial covering every equation, tensor shape transformation, and PyTorch block. Full Blog Post: https://imrancoder786.github.io/blog-post.html?post=transformer-from-scratch GitHub Repository: https://github.com/imrancoder786/ML_FROM_SCRATCH/tree/main/Transformer_from_scratch I’d love to hear your feedback, suggestions, or any questions on the code/math! I’d love to hear your feedback, suggestions, or any questions on the code/math! submitted by /u/imrancoder [link] [comments]

Source: r/MachineLearning | 2026-07-27

Loading related sources…