Sign inSign up

ai/granite-4.0-h-tiny

Verified Publisher

By Docker

•Updated 12 months ago

7B long-context instruct model with RL alignment, IF, tool use, and enterprise optimization.

Model
4

10K+

ai/granite-4.0-h-tiny repository overview

⁠Granite-4.0-h-Tiny

logo

⁠Description

Granite-4.0-H-Tiny is a 7B parameter long-context instruct model finetuned from Granite-4.0-H-Tiny-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging. Granite 4.0 instruct models feature improved instruction following (IF) and tool-calling capabilities, making them more effective in enterprise applications.

⁠Characteristics

AttributeDetails
ProviderGranite Team, IBM
Architecturegranitehybrid
Cutoff dateNot disclosed
LanguagesEnglish, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, Chinese (extensible via finetuning)
Tool calling✅
Input modalitiesText
Output modalitiesText
LicenseApache 2.0

⁠Available model variants

Model variantParametersQuantizationContext windowVRAM¹Size
ai/granite-4.0-h-tiny:7B

ai/granite-4.0-h-tiny:7B-Q4_K_M

ai/granite-4.0-h-tiny:latest
64x994MMOSTLY_Q4_K_M1M tokens4.41 GiB3.94 GB

¹: VRAM estimated based on model characteristics.

latest → 7B

⁠Use this AI model with Docker Model Runner

docker model run ai/granite-4.0-h-tiny

⁠Considerations

  • Optimized for instruction following, tool/function calling, and long-context (up to 128K tokens) scenarios.
  • Strong generalist capabilities: summarization, classification, extraction, QA/RAG, coding, function-calling, and multilingual dialogue.
  • Multilingual: best performance in English; a few-shot approach or light finetuning can help close gaps for other languages.
  • Safety & reliability: despite alignment, the model can still produce inaccurate or biased outputs—apply domain-specific evaluation and guardrails.
  • Infrastructure note: trained on NVIDIA GB200 NVL72 at CoreWeave; use acceleration libraries (e.g., accelerate, optimized attention/KV cache settings) for efficient inference.

⁠Benchmark performance

CategoryMetricGranite-4.0-h-Tiny
General Tasks
MMLU (5-shot)68.65
MMLU-Pro (5-shot, CoT)44.94
BBH (3-shot, CoT)66.34
AGI EVAL (0-shot, CoT)62.15
GPQA (0-shot, CoT)32.59
Alignment Tasks
AlpacaEval 2.030.61
IFEval (Instruct, Strict)84.78
IFEval (Prompt, Strict)78.10
IFEval (Average)81.44
ArenaHard35.75
Math Tasks
GSM8K (8-shot)84.69
GSM8K Symbolic (8-shot)81.10
Minerva Math (0-shot, CoT)69.64
DeepMind Math (0-shot, CoT)49.92
Code Tasks
HumanEval (pass@1)83.00
HumanEval+ (pass@1)76.00
MBPP (pass@1)80.00
MBPP+ (pass@1)69.00
CRUXEval-O (pass@1)39.63
BigCodeBench (pass@1)41.06
Tool Calling Tasks
BFCL v357.65
Multilingual Tasks
MULTIPLE (pass@1)55.83
MMMLU (5-shot)61.87
INCLUDE (5-shot)53.12
MGSM (8-shot)45.36
Safety
SALAD-Bench97.77
AttaQ86.61

Tag summary

Content type

Model

Digest

sha256:c7b6e5774…

Size

3.9 GB

Last updated

12 months ago

docker model pull ai/granite-4.0-h-tiny

This week's pulls

Pulls:

31

Sep 14 to Sep 20