GLM
15 articles tagged with this topic
GLM Beats DeepSeek on Two GPUs — Chinese Open-Source Stops Compromising
On two NVIDIA DGX Sparks, GLM-5.3 Flash beat DeepSeek V4 Flash on HumanEval (97% vs 94.5%). GLM ran 30% slower with one-quarter the context.
Alibaba and Zhipu Bet on Small Models — Local AI Faces Choice Overload
Qwen Flash and GLM Flash launched together, leaving local users with choice overload. China's open-source LLMs shift from parameter wars to same-tier
Zhipu GLM Runs Locally on Mac — And This Matters More Than It Looks
ds4 (co-maintained by Redis creator antirez) added Zhipu GLM Flash support this week, running on 128GB M4 Max. A concrete step for local Chinese LLMs.
DeepSeek as Brain, Qwen as Assistant — China's Open-Source LLMs Split by Role
3 Chinese open-source LLMs tested on 4 DGX Spark cards. GLM cut for slowness. DeepSeek leads, Qwen supports — multi-model orchestration replaces singl
Zhipu GLM Flash Beats Qwen — Size Doesn't Cut It
GLM Flash beat Qwen on three of four Reddit benchmarks. Qwen only led on graduate-level science by half a point — despite GLM being the larger model.
Zhipu Puts GLM-5.3-Flash on Hugging Face — China's Open-Weight Push Continues
Zhipu posts GLM-5.3-Flash to Hugging Face, betting on speed and low cost. As open-weight models multiply, can pay-per-call APIs survive?
GLM 5.3 Flash Stays a Rumor; Zhipu's Next Model Path Is Unclear
Zhipu's GLM 5.3 and a rumored "Flash" variant fuel speculation, but weights are unreleased. Rumors hint at open-source competition, not launch facts.
EvoX Swarm Mode Lifts Accuracy from 26% to 71%, Putting Architecture in Focus
EvoX, EvoMap's desktop Agent, uses task splitting and context isolation to lift accuracy to 71%, suggesting architecture matters as much as model capa
Zhipu GLM's 'Hilarious' Thinking Goes Viral as Chinese Open LLMs Race on Inner Monologue
Zhipu GLM 5.3's 'hilarious' thinking hit r/LocalLLaMA. The meme masks Chinese open LLMs selling transparent reasoning as a post-R1 differentiator.
llama.cpp Fork Turns Retired AMD Server GPUs into AI Inference Rigs
Reddit user milpster and Zhipu AI's GLM team release a llama.cpp fork optimized for AMD GFX906, letting Mi50, Mi60, and Radeon VII run LLMs locally.
Codename Ox Alpha Appears on Reddit — Z.ai's Next-Gen GLM Model Leaked Early
Ox Alpha surfaces on Reddit's LocalLLaMA, claiming to be Zhipu's next-gen GLM with 1M token context and image/video input.
After a 6x cost spike, LLM cost control enters the engineering era
LLM costs now depend on output length, reasoning, and tool-call loops. Per-request and per-task budgets are shifting from cost tricks to reliability e
5 AI Bombshells in One Day — Only 1 Hits Your Side Hustle
Friday: 5 AI bombshells — OpenAI $40B, Apple+Alibaba, GLM-5.3, Cursor×SpaceX, Gemini 3.7 Flash. The 1 that actually matters for indie builders.
Tim Dettmers Teases New Quantization Method: 7 tok/s on One Box — Industry Wary
Tim Dettmers claims GLM 5.3 hits 7 tok/s on a single DGX Spark. If true, enterprise LLM hardware costs halve — Reddit says: wait for benchmarks.
GLM 5.1 Dominates Open-Source Code Arena: China's Programming AI Inflection Point
Zhipu's GLM 5.1 topping open-source code rankings signals that low-cost programming AI is within reach, with software outsourcing and IT service prici