Back to home

Open-Source LLM

12 articles tagged with this topic

OrnithQwen3

Ornith Runs 35B Coding Model on 8GB VRAM: Local AI's Sweet Spot Arrives

A Reddit user ran Ornith-1.5-35B-A3B on an 8GB RTX 3070 laptop at ~32 tokens/sec, completing agentic coding tasks end-to-end. Consumer hardware is now

1d ago2 min read
Z.AIZhipu

Zhipu Open-Sources Ox Alpha vs DeepSeek — Open Source Becomes China's LLM Default

Z.AI confirms Ox Alpha is GLM's next gen, open-sourced tonight. After DeepSeek, a second Chinese AI firm bets open source for ecosystem.

3d ago2 min read
DeepSeekApple

DeepSeek Model Hits 25 token/s on $7K Mac — Local AI Catches the Cloud

DeepSeek's latest model hits 25.8 tokens/s locally on an M2 Ultra Mac — smaller than official quant. Chinese open-source LLMs are now viable.

6d ago2 min read
QwenApple Silicon

Qwen Runs 45 Tokens/Second Locally — Apple Silicon's Silent Win for Open LLMs

Reddit user benchmarked Qwen 3.8B at 45+ tok/s on Apple silicon — a record on consumer hardware. Local LLM inference shifts from geek toy to viable op

Aug 212 min read
QwenAlibaba

Qwen 27B Crushed to 1 Bit — Runs on 8GB VRAM, Output Is Brain-Dead

Reddit user crushed Qwen 27B to 1-bit, ran it on an 8GB laptop — output was gibberish. We dig into what this reveals about local AI limits.

Aug 202 min read
ZhipuGLM-5

GLM-5.3 Hits Artificial Analysis — China's First Open-Source Flagship Vetted

Zhipu's GLM-5.3 completes Artificial Analysis benchmarks — first Chinese open-source model to land in the global top tier with an independent score.

Aug 192 min read
QwenAlibaba Cloud

Qwen 35B Pre-Launch Hype Validates China's Open-Source LLM Roadmap

Reddit's local AI community awaits Qwen 3.8 35B A3B — likely a MoE model (35B total / 3B active), extending China's "small but strong" open-source pla

Aug 182 min read
QwenAlibaba

Qwen Pulls a Local-Deploy Favorite — Is Alibaba's Open-Source Cadence Shifting?

Qwen dropped a popular MoE model from GitHub, alarming r/LocalLLaMA. China's top open-weight LLM is moving to on-demand over completeness.

Aug 162 min read
QwenLocal Deployment

16GB GPU Hits 'Performance Cliff' Running Qwen — Local LLM Bar Is Higher Than You Think

A Reddit user found Qwen3-27B's KV cache precision tweak on a 16GB GPU cratered speed from 9 to 1.5 tokens/sec. Local LLM deployment is far harder tha

Aug 152 min read
AlibabaQwen

Alibaba Open-Sources Qwen Again — Can Closed-Source AI's Price Moat Hold?

Alibaba's new Qwen matches GPT-4o and Claude 3.5 Sonnet benchmarks. When open-source is good enough and free, must-buy-closed logic gets rewritten.

Aug 152 min read
QwenAlibaba

Qwen 27B Runs Full Web App on Dual 3090s — Open-Source Catches Closed Models

Qwen 27B ships a full retro web game on dual RTX 3090s — China open-source hits usable perf on consumer GPUs, giving firms a local coding option.

Aug 152 min read
MetaMuse Glimmer

Meta Open-Sources 30B Agent — Your Life Data Is the Real Price

Meta open-sources its 30B Muse Glimmer Agent under Apache 2.0. The catch: it pitches "deep access to personal life" — a red flag for any ad business.

Aug 122 min read