Hi everyone! 👋 I built and trained the complete Transformer architecture from scratch using pure PyTorch (`torch.nn` primitives) based on the original "Attention Is All You Need" paper. I trained the model on an English-to-Tamil parallel translation dataset (`gopi30/english-tamil` on Hugging Face) u