Open-source LLM
13 articles tagged with this topic
700 Lines of C Code Run Google's Latest LLM — Solo Project Beats llama.cpp
Open-source gemma4.c runs Google Gemma 4 E2B in 700 lines of C, hitting 25.9 tok/s on a regular CPU and beating llama.cpp.
Qwen Pushed on Code Autocomplete — AI Tools Still Far From Real Workflows
A Reddit dev asked if Qwen open-source models support FIM code autocomplete like Copilot — exposing the engineering gap from benchmarks to daily use.
Zhipu's GLM Just Got Faster Locally — The Self-Hosted LLM Bar Drops Again
Zhipu's GLM-4.5-Air (106B total / 12B active) unlocked MTP in llama.cpp, validated on RTX 3090. A signal for data-localization-focused firms.
Single RTX 5090 Runs Qwen 27B at 262K Context — Local LLM Threshold Crushed
NVIDIA's NVFP4 quantization compresses Qwen3.8-27B to 19GB on a single RTX 5090, running full 262K context at 77 tokens/sec — local LLM threshold fall
Qwen 27B Local Coding Test: Two AI Coding Agents, Surprisingly Wide Gap
On an RTX 3090, a developer ran two AI coding Agents on Qwen 27B. PI Agent won on resources and context, showing open-source LLMs can run Agent tasks.
Qwen 27B Ties Claude Opus on AIME 2026; Open-Source LLMs Trail GPT by 3 Points
Alibaba's Qwen 27B scored 96.7% on AIME 2026 (29/30), tying Claude Opus 4.6 but trailing GPT-5.6's 99.9% by 3 points. FP8 runs 2.7x faster.
Ornith 1.5 Drops Three Sizes at Once — Indie Devs Now Ship 397B MoE
Developer tarruda dropped Ornith 1.5 on r/LocalLLaMA: 9B dense, 35B MoE, and 397B MoE in one release. The 397B scale is now reachable for indie develo
Alibaba Qwen3.8-27B runs 128K on dual 4090s — local LLM cost hits 6-figure RMB
Alibaba's Qwen3.8-27B runs 128K on dual RTX 4090s at ~$15K hardware cost. Near-GPT-grade local LLM deployment now within SMB reach.
RTX 4090 Runs 27B Model Locally — On-Prem AI Hardware Cost Curve Breaks Through
A developer runs 27B Qwen on a single RTX 4090, hitting 250K-350K tokens context. Hardware bar for local LLMs slips below SME expectations.
Qwen abliteration strips 99% to 6% refusal with only 1.3 test point loss
Open-source abliteration strips Qwen's refusal with only 1.3 point test loss — making "open-source AI safety thickness" a concrete numbers question.
Kimi K3 Lands in llama.cpp — Another Chinese Model Goes Local
llama.cpp received a PR this week adding Kimi K3 support, meaning the new Chinese model could run locally like Llama and Qwen — giving users a cloud-i
Qwen 3.8 Spotted on GitHub — Alibaba's Open-Source Cadence Outpaces Rivals
Traces of Qwen 3.8 35B-A3B surfaced in Alibaba Tongyi's ms-swift framework on GitHub, hinting at the next open-source drop from a leading model family
Qwen 3.8 Launches with 5 Bugs—Community Devs Ship a Universal Fix in One Week
Alibaba's Qwen 3.8 launched with adjustable reasoning depth but shipped with 5 critical chat template bugs. Community dev froggeric released a univers