Mistral Small 4
mistral · mistral/mistral-small-2603
Fast Mistral production model for chat, extraction, and cost-sensitive agents
≈171k Chinese characters of context · Reads images · Step-by-step reasoning · Calls tools / agents · Structured JSON output · Open weights, self-hostable
Read a 200,000-Chinese-character book and write a 2,000-Chinese-character summary: about $0.05
An order-of-magnitude estimate at 1.5 tokens per Chinese character and the vendor's own rate — not a quote. Full method: the glossary on the AI Models page.
Specs
| Context | 256,000(≈171k Chinese characters) |
|---|---|
| Max output | 256,000 |
| Structured output | Yes |
| Open weights | Yes |
| Released | 2026-03-16 |
Source: models.dev snapshot 2026-08-29
Who serves it, at what price
| Provider | In /M | Out /M | Cache read /M |
|---|---|---|---|
| Mistralfirst-party | $0.15 | $0.6 | — |
| Cortecs | $0.143 | $0.568 | $0.014 |
| Requesty | $0.165 | $0.66 | $0.165 |
| Venice AI | $0.1875 | $0.75 | — |
Rows tagged “first-party” are the model vendor's own pricing; untagged rows are resale channels and may differ.
When to choose it
This model is on our watch list but has no written assessment yet — we do not write what we cannot support. The specs and prices above still stand.
Leaderboard
Not yet on any leaderboard we track. That does not mean the model performs poorly — it means we have not taken this snapshot for it yet, and we do not publish a score we have not measured.