GLM-Image
zhipuai · zhipuai/glm-image
GLM-Image is an image generation model adopts a hybrid autoregressive + diffusion decoder architecture. In general image generation quality, GLM‑Image aligns with mainstream latent diffusion approaches, but it shows significant advantages in text-rendering and knowledge‑intensive generation scenarios. It performs especially well in tasks requiring precise semantic understanding and complex information expression, while maintaining strong capabilities in high‑fidelity and fine‑grained detail generation. In addition to text‑to‑image generation, GLM‑Image also supports a rich set of image‑to‑image tasks including image editing, style transfer, identity‑preserving generation, and multi‑subject consistency.
≈7k Chinese characters of context · Reads images · Open weights, self-hostable
Read a 200,000-Chinese-character book and write a 2,000-Chinese-character summary: about No public rate
An order-of-magnitude estimate at 1.5 tokens per Chinese character and the vendor's own rate — not a quote. Full method: the glossary on the AI Models page.
Specs
| Context | 10,240(≈7k Chinese characters) |
|---|---|
| Max output | — |
| Structured output | Unknown (not declared upstream) |
| Open weights | Yes |
| Released | 2026-01-19 |
Source: models.dev snapshot 2026-09-20
Who serves it, at what price
No channel in the snapshot currently serves this model — the vendor may have retired it, or it may never have had a public API. The specs above are the vendor's published figures and still stand.
When to choose it
We have not assessed this model. It comes from the full models.dev snapshot — the specs are the vendor's published figures, and a price is the vendor's own only on rows tagged first-party; the rest are resale channels. The fit is simply not something anyone here has written.
Leaderboard
Not yet on any leaderboard we track. That does not mean the model performs poorly — it means we have not taken this snapshot for it yet, and we do not publish a score we have not measured.