Qwen Image 2.1 Nunchaku: Ada INT4 and RTX 5090 native NVFP4 research workers
164
Built with Qwen. Research and evaluation only. License and attribution ship in /app/LICENSE, /app/NOTICE and /app/THIRD_PARTY_NOTICES.md. This is an independent custom Qwen Image 2.1 implementation, not an official Qwen/Nunchaku release or a near-lossless claim.
| Image tag | Target | Status |
|---|---|---|
v1-fp4, 20260921-r128-fp4 | RTX 5090 / Blackwell, NVFP4 rank 128, resident pipeline | Published at the same verified digest; all 25 RTX 5090 cloud requests and uploads passed |
v1-int4, 20260921-r128 | RTX 4070 Ti SUPER / RTX 4090 Ada, INT4 rank 128, offloaded pipeline | Earlier published/tested release; historical evidence |
Linux/amd64, Torch 2.8.0, CUDA 12.8 and generic Nunchaku 1.2.1. Both variants use an NF4 text encoder and BF16 VAE. Select FP4 for RTX 5090; do not relabel or load the INT4 transformer as FP4.
The FP4 image selects qwen-image-2.1-fp4-r128.safetensors from model revision 868fc8b147de20f2c6e6835b694b2c4d9de9b5a6: 4,876,425,112 bytes, SHA256 80122444361bf8bd0d974abf1eb69b839b52b8447ab1d9f42a31f859046049d8. The corresponding INT4 consolidated file is 4,655,402,112 bytes, SHA256 ce0792ce69b0c037bebd17d93174584ac6dfe84f77fa20dabddf07e6522b41af. FP4's auxiliary processor/scheduler/encoder/VAE/model-index components are pinned separately to immutable revision 2b1c7e324511916189b1635b6622b40d54ae73f5. Models are baked and offline at runtime; no HF token is required for inference. Build/runtime source is in the separate local POC bundle, not the model-only repository.
At 1024×1024, generation medians across eight prompts were 4.755 / 7.399 seconds for native NVFP4 versus 10.181 / 16.064 seconds for the original BF16 transformer at 25 / 40 steps. Both pipelines were resident on the same RTX 5090 and shared the NF4 text encoder, BF16 VAE, seeds and settings. The 24-job suites include generation, one-reference edits and a two-reference case; maximum sampled board memory was 19,192 MiB FP4 / 28,038 MiB BF16. Prompt/reference caches and compilation were disabled. These inference timings exclude serialization, queue, startup and network/upload; they are not endpoint response times. This is not a near-lossless claim, and a same-seed FP4 repeat changed visibly.
RunPod uses the default command, one GPU and concurrency 1. FP4 defaults to QWEN_OFFLOAD=0; home/INT4 defaults remain offloaded. Set RUNPOD_LOG_LEVEL=INFO. The worker waits for actual model readiness, then returns output.images and baked checkpoint identity:
{"input":{"prompt":"An editorial photograph of two potters shaping a bowl in a sunlit studio","width":1024,"height":1024,"steps":40,"CFGScale":1,"seed":42,"outputFormat":"PNG"}}
Add up to two referenceImages for editing and a presigned PUT uploadUrl for direct delivery; otherwise output is base64. Current limits are 256–1024 pixels per dimension, multiples of 32, 1–60 steps and CFG exactly 1. Forty steps is the upstream starting recommendation. The inner request timeout is 240 seconds, intended for a 300-second endpoint execution ceiling. Queue/startup/upload time must not be confused with model inference timing.
The published FP4 image digest is sha256:e7270890d2d05cb80c0501c2f33923961e02ef56b0ba2706547277fae9bd9512. Dedicated endpoint 38yo5rzdni3zf1 is configured for RTX 5090 only, using template hukn4fot5e and minimum zero / maximum one worker. All 25 actual cloud requests completed with the expected FP4 precision, rank, model revision and transformer hash. Their presigned uploads were retrieved and verified as 1024×1024 PNGs. Earlier successful RTX 4090 cloud generation/edit jobs and image digests remain historical records. MediaRouter remains disabled and undeployed.
All 25 actual cloud requests completed with the expected FP4 precision, rank, model revision and transformer hash. Their presigned uploads were retrieved and verified as 1024×1024 PNGs.
| Workload | Steps | n | Engine | SDK execution | Client wall |
|---|---|---|---|---|---|
| Generation | 25 | 8 | 4.757 s | 6.938 s | 8.752 s |
| Generation | 40 | 8 | 7.400 s | 9.581 s | 11.246 s |
| Single-reference edit | 25 | 3 | 6.383 s | 10.081 s | 11.666 s |
| Single-reference edit | 40 | 3 | 9.743 s | 13.076 s | 14.968 s |
| Two-reference case | 25 | 1 | 8.115 s | 11.793 s | 14.010 s |
| Two-reference case | 40 | 1 | 12.183 s | 16.167 s | 18.318 s |
These are medians over the 24 sequential requests after the separately reported initial cold request; the two-reference rows each contain one case. Two worker IDs were observed, so this is not a strictly warm, single-worker benchmark. Subsequent queue delays ranged from 0.411 to 22.127 seconds. Default serving caches were enabled. SDK execution includes serving and upload overhead; client wall also includes submission, queueing and polling. The first 25-step request waited 349.152 seconds in the queue, then took 7.503 seconds SDK execution (5.222 seconds engine) and 357.567 seconds client wall. These timings are separate from the controlled resident comparison above and do not establish image-quality equivalence.
Content type
Image
Digest
sha256:e7270890d…
Size
14.2 GB
Last updated
3 days ago
docker pull mesmerlord/mesmer-image21-runpod:20260921-r128-fp4