Train a ChatGPT clone from scratch base on karpathy/nanochat. PyTorch + CUDA. 8×H100, 4hrs, $100.
1.0K
Train your own ChatGPT clone from scratch in 4 hours for $100
A pre-configured Docker image for nanochat - a minimal, full-stack LLM implementation designed for educational purposes and rapid experimentation.
To test the pre-trained model, you can check out our online demo
| Component | Description |
|---|---|
| Base Stack | PyTorch 2.8.0 + CUDA 12.8.1 |
| Package Manager | uv (fast Python package installer) |
| Tokenizer | Pre-built rustbpe for efficient BPE encoding |
| Codebase | Complete nanochat with all dependencies |
| Initial Data | 8 training shards (~800MB) pre-downloaded |
| Startup Script | Interactive guide with system info and commands |
# 1. Launch pod with 8×H100 GPUs
# 2. Container starts automatically with helpful info displayed
# 3. Start training (recommended: use screen session)
screen -L -Logfile speedrun.log -S speedrun bash speedrun.sh
# 4. Wait ~4 hours for complete pipeline
# 5. Chat with your trained model
python -m scripts.chat_web
# Access at http://YOUR_POD_IP:8000
Tokenization → Pretraining → Finetuning → Evaluation → Inference
After the 4-hour pipeline completes:
| Output | Size | Description |
|---|---|---|
report.md | ~50KB | Complete training metrics & evaluation results |
tok_checkpoints/ | ~1MB | Trained BPE tokenizer (vocab size 65,536) |
chatsft_checkpoints/ | ~2GB | Final SFT model (ready for inference) |
base_checkpoints/ | ~2GB | Pretrained base model (561M params) |
mid_checkpoints/ | ~2GB | Midtraining checkpoints |
| Requirement | Specification |
|---|---|
| GPUs | 8×H100 (80GB VRAM each) or 8×A100 |
| Disk Space | ~60GB (24GB data + 6GB checkpoints + overhead) |
| Network | Internet connection for dataset download |
| Time | ~4 hours for full pipeline |
| Cost | ~$100 on 8×H100 @ $24/hr |
Enable Wandb Logging:
wandb login
WANDB_RUN=my_run bash speedrun.sh
Download Trained Model:
tar -czf nanochat_model.tar.gz -C ~/.cache/nanochat \
tok_checkpoints/ chatsft_checkpoints/
Troubleshoot OOM:
# Reduce batch size if running out of memory
torchrun --standalone --nproc_per_node=8 -m scripts.base_train -- --device_batch_size=16
Built with ❤️ by Andrej Karpathy | Powered by RunPod
Content type
Image
Digest
sha256:30bf00c56…
Size
14.7 GB
Last updated
12 months ago
docker pull axiilay/nanochat