Sign inSign up

mesmerlord/mesmer-image21-runpod

By mesmerlord

Updated 3 days ago

Qwen Image 2.1 Nunchaku: Ada INT4 and RTX 5090 native NVFP4 research workers

Image
0

164

mesmerlord/mesmer-image21-runpod repository overview

Mesmer Image 21 Nunchaku

Built with Qwen. Research and evaluation only. License and attribution ship in /app/LICENSE, /app/NOTICE and /app/THIRD_PARTY_NOTICES.md. This is an independent custom Qwen Image 2.1 implementation, not an official Qwen/Nunchaku release or a near-lossless claim.

Image tagTargetStatus
v1-fp4, 20260921-r128-fp4RTX 5090 / Blackwell, NVFP4 rank 128, resident pipelinePublished at the same verified digest; all 25 RTX 5090 cloud requests and uploads passed
v1-int4, 20260921-r128RTX 4070 Ti SUPER / RTX 4090 Ada, INT4 rank 128, offloaded pipelineEarlier published/tested release; historical evidence

Linux/amd64, Torch 2.8.0, CUDA 12.8 and generic Nunchaku 1.2.1. Both variants use an NF4 text encoder and BF16 VAE. Select FP4 for RTX 5090; do not relabel or load the INT4 transformer as FP4.

The FP4 image selects qwen-image-2.1-fp4-r128.safetensors from model revision 868fc8b147de20f2c6e6835b694b2c4d9de9b5a6: 4,876,425,112 bytes, SHA256 80122444361bf8bd0d974abf1eb69b839b52b8447ab1d9f42a31f859046049d8. The corresponding INT4 consolidated file is 4,655,402,112 bytes, SHA256 ce0792ce69b0c037bebd17d93174584ac6dfe84f77fa20dabddf07e6522b41af. FP4's auxiliary processor/scheduler/encoder/VAE/model-index components are pinned separately to immutable revision 2b1c7e324511916189b1635b6622b40d54ae73f5. Models are baked and offline at runtime; no HF token is required for inference. Build/runtime source is in the separate local POC bundle, not the model-only repository.

Controlled RTX 5090 measurement

At 1024×1024, generation medians across eight prompts were 4.755 / 7.399 seconds for native NVFP4 versus 10.181 / 16.064 seconds for the original BF16 transformer at 25 / 40 steps. Both pipelines were resident on the same RTX 5090 and shared the NF4 text encoder, BF16 VAE, seeds and settings. The 24-job suites include generation, one-reference edits and a two-reference case; maximum sampled board memory was 19,192 MiB FP4 / 28,038 MiB BF16. Prompt/reference caches and compilation were disabled. These inference timings exclude serialization, queue, startup and network/upload; they are not endpoint response times. This is not a near-lossless claim, and a same-seed FP4 repeat changed visibly.

RunPod uses the default command, one GPU and concurrency 1. FP4 defaults to QWEN_OFFLOAD=0; home/INT4 defaults remain offloaded. Set RUNPOD_LOG_LEVEL=INFO. The worker waits for actual model readiness, then returns output.images and baked checkpoint identity:

{"input":{"prompt":"An editorial photograph of two potters shaping a bowl in a sunlit studio","width":1024,"height":1024,"steps":40,"CFGScale":1,"seed":42,"outputFormat":"PNG"}}

Add up to two referenceImages for editing and a presigned PUT uploadUrl for direct delivery; otherwise output is base64. Current limits are 256–1024 pixels per dimension, multiples of 32, 1–60 steps and CFG exactly 1. Forty steps is the upstream starting recommendation. The inner request timeout is 240 seconds, intended for a 300-second endpoint execution ceiling. Queue/startup/upload time must not be confused with model inference timing.

The published FP4 image digest is sha256:e7270890d2d05cb80c0501c2f33923961e02ef56b0ba2706547277fae9bd9512. Dedicated endpoint 38yo5rzdni3zf1 is configured for RTX 5090 only, using template hukn4fot5e and minimum zero / maximum one worker. All 25 actual cloud requests completed with the expected FP4 precision, rank, model revision and transformer hash. Their presigned uploads were retrieved and verified as 1024×1024 PNGs. Earlier successful RTX 4090 cloud generation/edit jobs and image digests remain historical records. MediaRouter remains disabled and undeployed.

Actual RunPod RTX 5090 verification

All 25 actual cloud requests completed with the expected FP4 precision, rank, model revision and transformer hash. Their presigned uploads were retrieved and verified as 1024×1024 PNGs.

WorkloadStepsnEngineSDK executionClient wall
Generation2584.757 s6.938 s8.752 s
Generation4087.400 s9.581 s11.246 s
Single-reference edit2536.383 s10.081 s11.666 s
Single-reference edit4039.743 s13.076 s14.968 s
Two-reference case2518.115 s11.793 s14.010 s
Two-reference case40112.183 s16.167 s18.318 s

These are medians over the 24 sequential requests after the separately reported initial cold request; the two-reference rows each contain one case. Two worker IDs were observed, so this is not a strictly warm, single-worker benchmark. Subsequent queue delays ranged from 0.411 to 22.127 seconds. Default serving caches were enabled. SDK execution includes serving and upload overhead; client wall also includes submission, queueing and polling. The first 25-step request waited 349.152 seconds in the queue, then took 7.503 seconds SDK execution (5.222 seconds engine) and 357.567 seconds client wall. These timings are separate from the controlled resident comparison above and do not establish image-quality equivalence.

Tag summary

Content type

Image

Digest

sha256:e7270890d

Size

14.2 GB

Last updated

3 days ago

docker pull mesmerlord/mesmer-image21-runpod:20260921-r128-fp4