FastAPI service running PyTorch inference to process NMT and export attention tensors.
237
Attention-Seeker is a comprehensive Transformer Model trained for Translation. It is a from-scratch implementation of the Transformer architecture as proposed in "Attention Is All You Need."
Unlike standard implementations that utilize high-level abstractions, this project is a ground-up reconstruction of the core architecture. This includes the manual implementation of Multi-Head Attention, Positional Encodings, and specialized masking logic.
Content type
Image
Digest
sha256:c908eb894…
Size
3.1 GB
Last updated
6 months ago
docker pull kinjal1234/transformer-backend