Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 14 additions & 1 deletion docusaurus.config.ts
Original file line number Diff line number Diff line change
Expand Up @@ -76,7 +76,10 @@ const config: Config = {
'@docusaurus/plugin-client-redirects',
{
// Pre-wired empty: add entries here if any URLs change after the IBM.com cutover.
redirects: [],
redirects: [
// The legacy /models/granite URL always lands on the latest generation.
{from: '/models/granite', to: '/models/granite4-2'},
],
} satisfies RedirectOptions,
],
[
Expand All @@ -99,6 +102,16 @@ const config: Config = {
{name: 'keywords', content: 'IBM Granite, AI, foundation models, LLM'},
{name: 'description', content: 'IBM Granite documentation — models, serving guides, cookbooks'},
],
announcementBar: {
// Bump this id if the message changes and you want dismissals reset
// (only relevant while isCloseable is true).
id: 'site-sunset-2026',
content:
'⚠️ This documentation site is no longer being updated. For the latest Granite documentation and models, see <a target="_blank" rel="noopener noreferrer" href="https://huggingface.co/ibm-granite">Hugging Face</a> and <a target="_blank" rel="noopener noreferrer" href="https://github.com/ibm-granite">GitHub</a>.',
backgroundColor: '#fcf4d6',
textColor: '#1c1200',
isCloseable: false,
},
navbar: {
title: '',
logo: {
Expand Down
File renamed without changes.
267 changes: 267 additions & 0 deletions granite/docs/models/granite4-2.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,267 @@
---
title: "Granite 4.2"
description: "Dense reasoning language model family in 3B, 8B, and 30B sizes with built-in chain-of-thought thinking, flexible thinking modes, and reasoning-augmented tool calling."
sidebar_custom_props:
icon: "brain"
---

<CardGroup cols={2}>
<Card
title="Model Collection"
icon="arrow-right"
href="https://huggingface.co/collections/ibm-granite/granite-42-language-models"
>
View the full Granite 4.2 collection on Hugging Face
</Card>
<Card
title="GitHub Organization"
icon="github"
href="https://github.com/ibm-granite/granite-4.2-language-models"
>
Granite models and documentation
</Card>
</CardGroup>

## Overview

**Granite 4.2** is a family of dense reasoning language models available in three sizes: 3B, 8B, and 30B parameters. Granite 4.2 introduces native reasoning (thinking) capabilities, allowing models to perform step-by-step chain-of-thought reasoning before producing final answers. This significantly improves performance on complex math, coding, multi-step logic, and agentic tool-calling tasks.

### Model Variants

- **granite-4.2-3b**: Compact reasoning model optimized for edge deployment and resource-constrained environments
- **granite-4.2-8b**: Balanced reasoning model for general-purpose enterprise applications
- **granite-4.2-30b**: Flagship reasoning model for complex reasoning and specialized tasks

All models natively support a 128K context window (with long-context extension to 512K on the 30B model) and are released under the Apache 2.0 license with cryptographic signatures, ISO certification, and full transparency disclosures, enabling unrestricted commercial and academic use.

### Key Capabilities

**Built-in Reasoning**: Granite 4.2 features native chain-of-thought reasoning inside `<think>...</think>` tags, significantly improving performance on math, coding, and complex multi-step problems.

**Flexible Thinking Modes**: Seamlessly switch between full thinking (default), non-thinking, and low-effort modes within a single model, allowing users to balance depth vs. latency on a per-query basis.

**Reasoning-Augmented Tool Calling**: The model reasons about which tools to invoke and why before making the call, producing more accurate function calls for agentic workflows.

**Multilingual Dialog**: Granite 4.2 is tested across English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese.

## Getting Started

First, install the required libraries:

<CodeGroup>

```bash Install
pip install torch torchvision torchaudio
pip install accelerate
pip install transformers
```

</CodeGroup>

### Generation

This is a simple example of how to use the Granite-4.2-30B model in thinking mode:

<CodeGroup>

```python Python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = "ibm-granite/granite-4.2-30b"
tokenizer = AutoTokenizer.from_pretrained(model_path)
# drop device_map if running on CPU
model = AutoModelForCausalLM.from_pretrained(model_path, device_map="cuda", torch_dtype=torch.bfloat16)
model.eval()

# change input text as desired
messages = [
{ "role": "user", "content": "How many r's are in the word 'strawberry'?" },
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

# generate output tokens
output = model.generate(**inputs, max_new_tokens=8192, temperature=1.0, top_p=0.95, do_sample=True)

# decode output tokens into text
print(tokenizer.decode(output[0][inputs.input_ids.shape[-1]:], skip_special_tokens=False))
```

</CodeGroup>

Expected output:

<CodeGroup>

```text Output
<think>
First, I need to write out the word: s t r a w b e r r y.
Now, I count the number of 'r' letters: one at position 3, one at position 8, one at position 9.
Total r's = 3.
</think>
There are **3** r's in the word "strawberry".<|im_end|>
```

</CodeGroup>

> **Generation parameters:** Use `temperature=1.0` and `top_p=0.95` across **all tasks and serving backends**, including general chat, reasoning, and tool calling.

### Thinking Modes

Granite 4.2 supports three thinking modes, selected via chat-template parameters:

| Mode | Template Parameters | Behavior |
|:-----|:-------------------|:---------|
| **Thinking** (default) | `enable_thinking=True` | Full chain-of-thought reasoning inside `<think>...</think>` |
| **Non-thinking** | `enable_thinking=False` | Direct answer with no reasoning overhead |
| **Low-effort** | `enable_thinking=True, low_effort=True` | Brief reasoning for simpler queries |

#### Non-Thinking Mode

Disable reasoning for a direct answer with no chain-of-thought overhead:

<CodeGroup>

```python Python
messages = [
{"role": "user", "content": "What is the capital of France?"},
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

output = model.generate(**inputs, max_new_tokens=2048, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(output[0][inputs.input_ids.shape[-1]:], skip_special_tokens=False))
```

</CodeGroup>

Expected output:

<CodeGroup>

```text Output
<think></think>The capital of France is Paris.<|im_end|>
```

</CodeGroup>

### Tool Calling

Granite 4.2 supports tool calling with integrated reasoning — the model thinks about which tool to call and why before making the call. Define a list of tools using OpenAI's function definition schema:

<CodeGroup>

```python Python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = "ibm-granite/granite-4.2-30b"
tokenizer = AutoTokenizer.from_pretrained(model_path)
# drop device_map if running on CPU
model = AutoModelForCausalLM.from_pretrained(model_path, device_map="cuda", torch_dtype=torch.bfloat16)
model.eval()

tools = [
{
"type": "function",
"function": {
"name": "get_current_weather",
"description": "Get the current weather for a specified city.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "Name of the city"
}
},
"required": ["city"]
}
}
}
]

# change input text as desired
messages = [
{ "role": "user", "content": "What's the weather like in Boston right now?" },
]

text = tokenizer.apply_chat_template(messages, tokenize=False, tools=tools,
add_generation_prompt=True, enable_thinking=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

# generate output tokens
output = model.generate(**inputs, max_new_tokens=4096, temperature=1.0, top_p=0.95, do_sample=True)

# decode output tokens into text
print(tokenizer.decode(output[0][inputs.input_ids.shape[-1]:], skip_special_tokens=False))
```

</CodeGroup>

Expected output:

<CodeGroup>

```text Output
<think>
The user is asking for the weather in Boston right now. There's a function called
get_current_weather that takes a city parameter. I need to call that with the city set to Boston.
</think>
<tool_call>
<function=get_current_weather>
<parameter=city>
Boston
</parameter>
</function>
</tool_call>
<|im_end|>
```

</CodeGroup>

### Serving with vLLM

Granite 4.2 is optimized for deployment with [vLLM](https://github.com/vllm-project/vllm) (v0.20+).

> **Reasoning parser:** Use the custom `granite_thinking_parser`, available in each model's Hugging Face repository. The models also work with the built-in `nemotron_v3` parser, but `granite_thinking_parser` provides better formatting of reasoning output.
> **Tool calling parser:** Use `qwen3_coder`.

<CodeGroup>

```bash Serve
vllm serve ibm-granite/granite-4.2-30b \
--served-model-name granite-4.2-30b \
--dtype bfloat16 \
--max-model-len 131072 \
--reasoning-parser granite_thinking_parser \
--reasoning-parser-plugin ./granite_thinking_parser.py \
--tool-call-parser qwen3_coder \
--enable-auto-tool-choice
```

</CodeGroup>

The model then exposes an OpenAI-compatible API on `http://localhost:8000/v1`, which integrates with popular agentic coding harnesses such as OpenCode, Pi, and OpenHands out of the box.

<CodeGroup>

```python Python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")

response = client.chat.completions.create(
model="granite-4.2-30b",
messages=[{"role": "user", "content": "Explain the Riemann hypothesis in simple terms."}],
temperature=1.0,
top_p=0.95,
max_tokens=8192,
)

print(response.choices[0].message.content)
```

</CodeGroup>
2 changes: 1 addition & 1 deletion granite/docs/models/guardian.mdx
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
title: "Granite Guardian 4.1"
title: "Granite Guardian"
description: "Risk detection models and LoRA adapters for evaluating AI safety, RAG groundedness, and agentic function calling — with support for custom bring-your-own criteria (BYOC)."
sidebar_custom_props:
icon: "shield"
Expand Down
15 changes: 7 additions & 8 deletions granite/docs/models/speech.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
title: "Granite Speech"
description: "Compact multilingual models for automatic speech recognition (ASR) and translation (AST) across English, French, German, Spanish, Portuguese, and Japanese."
description: "Compact speech models for automatic speech recognition (ASR)."
sidebar_custom_props:
icon: "mic"
---
Expand All @@ -14,9 +14,9 @@ sidebar_custom_props:
View the full Granite Speech collection on Hugging Face
</Card>
<Card
title="Speech Demo"
title="Streaming Audio Demo"
icon="rocket"
href="https://huggingface.co/spaces/ibm-granite/granite-speech"
href="https://huggingface.co/spaces/ibm-granite/granite-speech-streaming-webgpu"
>
Try Granite Speech in action
</Card>
Expand All @@ -38,21 +38,20 @@ sidebar_custom_props:

## Overview

The **Granite Speech 4.1** model family provides compact and efficient speech-language models for multilingual automatic speech recognition (ASR) and automatic speech translation (AST), supporting English, French, German, Spanish, Portuguese, and Japanese. All models are trained on 174,000 hours of audio from public corpora and tailored synthetic datasets.
The Granite Speech model family provides compact and efficient speech-language models for multilingual automatic speech recognition (ASR) and automatic speech translation (AST), supporting English, French, German, Spanish, Portuguese, and Japanese.

### Model Variants

The Granite Speech 4.1 suite includes three specialized variants:
The Granite Speech suite includes four specialized variants:

- **[granite-speech-5.0-470m-turboctc](https://huggingface.co/ibm-granite/granite-speech-5.0-470m-turboctc)**: compact English ASR model with very high inference speed that is well suited for deployment on laptops, smartphones and other edge devices
- **[granite-speech-4.1-2b](https://huggingface.co/ibm-granite/granite-speech-4.1-2b)**: Balanced ASR and AST capabilities with improved punctuation and capitalization across all supported languages
- **[granite-speech-4.1-2b-plus](https://huggingface.co/ibm-granite/granite-speech-4.1-2b-plus)**: Speech-to-text model with speaker-attributed ASR, timestamps, and keyword-prompted ASR for enhanced recognition of names, acronyms, and technical jargon
- **[granite-speech-4.1-2b-nar](https://huggingface.co/ibm-granite/granite-speech-4.1-2b-nar)**: Non-autoregressive variant ([NLE architecture](https://arxiv.org/abs/2603.08397)) optimized for fast and accurate ASR with significantly lower latency

### Performance

Granite Speech 4.1 models deliver industry-leading performance on the [OpenASR Leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard): **granite-speech-4.1-2b ranks #1** for accuracy, while **granite-speech-4.1-2b-nar places #3** with exceptional speed through its non-autoregressive architecture.

Granite Speech is released under the Apache 2.0 license, making it freely available for both research and commercial purposes, with full transparency into its training data.
Granite Speech models deliver industry-leading performance on the [OpenASR Leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard) and the FFASR Leaderboard. Granite Speech is released under the Apache 2.0 license, making it freely available for both research and commercial purposes, with full transparency into its training data.

[Granite Speech Paper](https://arxiv.org/abs/2505.08699)

Expand Down
2 changes: 1 addition & 1 deletion granite/docs/run/granite-with-vllm-containerized.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,6 @@ Now run the container with the added parameters:
docker run --runtime nvidia --gpus all -v ~/.cache/huggingface:/root/.cache/huggingface -p 8000:8000 vllm/vllm-openai:latest --model ibm-granite/granite-4.0-h-small --tool-call-parser granite4 --enable-auto-tool-choice
```

Once the container is up, you can start running requests using the OpenAI API. Refer to the documentation on OpenAI API [tool calling](/models/granite#tool-calling) for examples.
Once the container is up, you can start running requests using the OpenAI API. Refer to the documentation on OpenAI API [tool calling](/models/granite4-0#tool-calling) for examples.

To run vLLM with the Granite 3 models and tool calling, use the additional parameters specified in the vLLM documentation [here](https://docs.vllm.ai/en/stable/features/tool_calling/#ibm-granite) as part of the `docker run` command share in Section 2.
2 changes: 1 addition & 1 deletion granite/docs/use-cases/docling-rag.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ sidebar_custom_props:
icon: "magnifying-glass"
---

In this tutorial, you will create an AI-powered document retrieval system with [Docling](https://github.com/DS4SD/docling), [LangChain](https://github.com/langchain-ai/langchain), and [Granite 3.1](/models/granite). The tutorial is designed to enable you to gain proficiency in document processing and chunking, integrate vector databases to enhance retrieval capabilities, and utilize RAG to perform efficient and accurate data retrieval for real-world applications.
In this tutorial, you will create an AI-powered document retrieval system with [Docling](https://github.com/DS4SD/docling), [LangChain](https://github.com/langchain-ai/langchain), and [Granite 3.1](/models/granite4-0). The tutorial is designed to enable you to gain proficiency in document processing and chunking, integrate vector databases to enhance retrieval capabilities, and utilize RAG to perform efficient and accurate data retrieval for real-world applications.

<Card
title="Granite Docling RAG notebook"
Expand Down
3 changes: 2 additions & 1 deletion sidebars.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,9 @@ const sidebars: SidebarsConfig = {
type: 'category',
label: 'Models',
items: [
'models/granite4-2',
'models/granite4-1',
'models/granite',
'models/granite4-0',
'models/docling',
'models/vision',
'models/speech',
Expand Down
3 changes: 2 additions & 1 deletion src/components/Icon.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ import useBaseUrl from '@docusaurus/useBaseUrl';
import {FontAwesomeIcon} from '@fortawesome/react-fontawesome';
import type {IconDefinition, SizeProp} from '@fortawesome/fontawesome-svg-core';
import {
faArrowRight, faBicycle, faBolt, faBook, faBox, faBriefcase, faBug,
faArrowRight, faBicycle, faBolt, faBook, faBox, faBrain, faBriefcase, faBug,
faChartLine, faChurch, faCircleExclamation, faCircleQuestion, faCloud,
faCode, faCookieBite, faCopy, faCube, faDesktop, faDownload, faEye, faFileCode,
faImages, faLaptop, faMagnifyingGlass, faMicrophone, faNewspaper,
Expand All @@ -26,6 +26,7 @@ const FA_MAP: Record<string, IconDefinition> = {
'bolt': faBolt,
'book': faBook,
'box': faBox,
'brain': faBrain,
'briefcase': faBriefcase,
'chart-line': faChartLine,
'church': faChurch,
Expand Down
Loading
Loading