(1)
Image generation model, uses a base latent diffusion model plus a refiner.
8m
50K+
8
SmolVLM: lightweight multimodal model for video, image, and text analysis, optimized for devices
11m
10K+
2
SmolVLM: lightweight multimodal model for video, image, and text analysis, optimized for devices.
12m
10K+
4
Image generation model, uses a base latent diffusion model plus a refiner.
8m
1.5K
SmolVLM: lightweight multimodal model for video, image, and text analysis, optimized for devices
8m
1.6K
SmolVLM: lightweight multimodal model for video, image, and text analysis, optimized for devices
12m
1.1K
9B multimodal model with vision, speech, and full-duplex streaming for text, image, video, audio
7m
1.9K