Back to home

Open-source LLM

13 articles tagged with this topic

GoogleGemma

700 Lines of C Code Run Google's Latest LLM — Solo Project Beats llama.cpp

Open-source gemma4.c runs Google Gemma 4 E2B in 700 lines of C, hitting 25.9 tok/s on a regular CPU and beating llama.cpp.

1d ago2 min read
QwenOpen-source LLM

Qwen Pushed on Code Autocomplete — AI Tools Still Far From Real Workflows

A Reddit dev asked if Qwen open-source models support FIM code autocomplete like Copilot — exposing the engineering gap from benchmarks to daily use.

4d ago2 min read
Zhipu AIGLM-4.5-Air

Zhipu's GLM Just Got Faster Locally — The Self-Hosted LLM Bar Drops Again

Zhipu's GLM-4.5-Air (106B total / 12B active) unlocked MTP in llama.cpp, validated on RTX 3090. A signal for data-localization-focused firms.

6d ago2 min read
QwenRTX 5090

Single RTX 5090 Runs Qwen 27B at 262K Context — Local LLM Threshold Crushed

NVIDIA's NVFP4 quantization compresses Qwen3.8-27B to 19GB on a single RTX 5090, running full 262K context at 77 tokens/sec — local LLM threshold fall

Aug 222 min read
QwenAlibaba

Qwen 27B Local Coding Test: Two AI Coding Agents, Surprisingly Wide Gap

On an RTX 3090, a developer ran two AI coding Agents on Qwen 27B. PI Agent won on resources and context, showing open-source LLMs can run Agent tasks.

Aug 222 min read
QwenAlibaba

Qwen 27B Ties Claude Opus on AIME 2026; Open-Source LLMs Trail GPT by 3 Points

Alibaba's Qwen 27B scored 96.7% on AIME 2026 (29/30), tying Claude Opus 4.6 but trailing GPT-5.6's 99.9% by 3 points. FP8 runs 2.7x faster.

Aug 202 min read
OrnithOpen-source LLM

Ornith 1.5 Drops Three Sizes at Once — Indie Devs Now Ship 397B MoE

Developer tarruda dropped Ornith 1.5 on r/LocalLLaMA: 9B dense, 35B MoE, and 397B MoE in one release. The 397B scale is now reachable for indie develo

Aug 192 min read
QwenAlibaba Tongyi Qianwen

Alibaba Qwen3.8-27B runs 128K on dual 4090s — local LLM cost hits 6-figure RMB

Alibaba's Qwen3.8-27B runs 128K on dual RTX 4090s at ~$15K hardware cost. Near-GPT-grade local LLM deployment now within SMB reach.

Aug 182 min read
QwenRTX 4090

RTX 4090 Runs 27B Model Locally — On-Prem AI Hardware Cost Curve Breaks Through

A developer runs 27B Qwen on a single RTX 4090, hitting 250K-350K tokens context. Hardware bar for local LLMs slips below SME expectations.

Aug 162 min read
QwenAlibaba

Qwen abliteration strips 99% to 6% refusal with only 1.3 test point loss

Open-source abliteration strips Qwen's refusal with only 1.3 point test loss — making "open-source AI safety thickness" a concrete numbers question.

Aug 162 min read
Moonshot AIKimi

Kimi K3 Lands in llama.cpp — Another Chinese Model Goes Local

llama.cpp received a PR this week adding Kimi K3 support, meaning the new Chinese model could run locally like Llama and Qwen — giving users a cloud-i

Aug 152 min read
QwenAlibaba Tongyi

Qwen 3.8 Spotted on GitHub — Alibaba's Open-Source Cadence Outpaces Rivals

Traces of Qwen 3.8 35B-A3B surfaced in Alibaba Tongyi's ms-swift framework on GitHub, hinting at the next open-source drop from a leading model family

Aug 152 min read
QwenTongyi Qianwen

Qwen 3.8 Launches with 5 Bugs—Community Devs Ship a Universal Fix in One Week

Alibaba's Qwen 3.8 launched with adjustable reasoning depth but shipped with 5 critical chat template bugs. Community dev froggeric released a univers

Aug 142 min read