Yotta Labs: Python 3.11 + AWS Neuron SDK + Mini-SGLang for Trainium & Inferentia2 inference.
620
Official Yotta Labs container for running LLM inference and training on AWS Trainium (trn1/trn2) and Inferentia2 (inf2) hardware.
Includes Mini-SGLang — Yotta Labs' lightweight, high-performance LLM inference server built for Neuron hardware, featuring Radix Cache (KV cache reuse), chunked prefill, and tensor parallelism.
| Component | Details |
|---|---|
| Python | 3.11 |
| Mini-SGLang | Yotta Labs LLM inference server for Neuron |
| torch-neuronx | 2.x (Neuron pip repo) |
| neuronx-cc | 2.x (Neuron compiler) |
| transformers-neuronx | Latest |
| neuronx-distributed | Latest |
| Hugging Face | transformers, datasets, huggingface_hub, accelerate |
| JupyterLab | Accessible on port 8888 |
| SSH | Password + public key auth |
The Neuron kernel driver must be installed on the EC2 host before running this container:
# Amazon Linux 2023
sudo tee /etc/yum.repos.d/neuron.repo > /dev/null <<'EOF'
[neuron]
name=Neuron YUM Repository
baseurl=https://yum.repos.neuron.amazonaws.com
enabled=1
EOF
sudo rpm --import https://yum.repos.neuron.amazonaws.com/GPG-PUB-KEY-AMAZON-AWS-NEURON.PUB
sudo yum install -y aws-neuronx-dkms
Mini-SGLang is pre-installed at /opt/mini-sglang-neuron. Start the OpenAI-compatible API server:
docker exec -it aws-neuron bash
cd /opt/mini-sglang-neuron
bash server.sh
Or run a model directly:
docker exec -it aws-neuron bash /opt/mini-sglang-neuron/run_minisgl.sh
# Neuron devices
docker exec aws-neuron neuron-ls
# torch-neuronx version
docker exec aws-neuron python3.11 -c "import torch_neuronx; print(torch_neuronx.__version__)"
# Mini-SGLang
docker exec aws-neuron python3.11 -c "import minisgl; print('Mini-SGLang ready')"
Content type
Image
Digest
sha256:4e0e57634…
Size
5.6 GB
Last updated
7 months ago
docker pull yottalabsai/aws-neuron:py3.11-ubuntu22.04-2026032501