Back to home

Unsloth

19 articles tagged with this topic

UnslothHuggingFace

HF 收购疑云下,Unsloth 让普通显卡跑得动大模型 — 开源生态的真正护城河

Under HF acquisition rumors, Unsloth lets 24GB GPUs run 70B models. Small open-source teams—not platforms—will decide if AI stays affordable.

Aug 272 min read
QwenUnsloth

Reddit Pushes Alibaba 35B Re-Quantization — Open-Source Compute Drops a Notch

A Reddit request urges re-quantizing Alibaba's Qwen 35B with Unsloth's new UD 3.0 — open-source AI compute keeps getting cheaper.

Aug 272 min read
QwenUnsloth

Qwen Flash Drops, Unsloth Adopts Day Zero — Open-Source AI's Pace Just Changed

Qwen Flash launched; Unsloth adopted it same day. The release-to-consumer-hardware gap shrank to under 24 hours. Enterprise IT should take note.

Aug 252 min read
QwenUnsloth

$500 GPU ships the first real PR — local AI coding exits demo

An indie dev ran Qwen 27B on a 4060Ti (~3,000 RMB) and shipped a human-reviewed PR — open-source local AI's first full engineering cycle.

Aug 252 min read
QwenUnsloth

Qwen 27B Compressed to 1-2 Bits — Local AI Memory Savings, Quality Wavers

Reddit's jojohai quantized Qwen3.8-27B to Q1-Q2 with MTP baked into weights. Lower memory than external MTP, but the model "gets stupid" off thinking

Aug 232 min read
QwenAlibaba Tongyi

Qwen 27B Runs 80-Step Agent on a Single GPU — Local Models Can Now Do Real Work

A Reddit test caught our eye: Qwen 27B on a consumer GPU made 80 autonomous tool calls from one prompt. The "cloud-only Agent" default is crumbling.

Aug 202 min read
UnslothQwen

Unsloth squeezes Qwen onto 8GB laptops — local LLMs now run on almost anything

Unsloth's Dynamic v3.0 lets Qwen models run on 8GB laptops, with 1-bit versions keeping 77% accuracy — easing local enterprise LLM deployment.

Aug 192 min read
QwenAlibaba Tongyi Qianwen

Qwen 3.8 Open-Source: 27B on RTX 4090 — Closed-Source Flagship Moat Loosens

Qwen 3.8 open-weights: 27B beats Claude Opus 4.6 Max on coding/agent benchmarks. 4-bit fits 24GB. New option for enterprise IT, dev tools, hardware.

Aug 172 min read
UnslothMuse-Glimmer

Unsloth Releases 30B Open-Source Model — Local LLMs Step Out of the Geek Bubble

Unsloth ships a GGUF-quantized Muse-Glimmer-30B, letting consumer laptops run a 30-billion-parameter LLM locally. On-prem AI is shifting from hobbyist

Aug 102 min read
KimiMoonshot

Kimi K3 Trimmed 33% for English-Only — Local LLMs Are Getting Quietly Affordable

Reddit user hellohazime trimmed Moonshot's Kimi K3 from 711GB to 478GB by stripping multilingual weights. A 2-bit variant reportedly beats its predece

Aug 92 min read
MLXLocalLLaMA

MLX 4bit Quantization Showdown: Which Compression Format Actually Wins on Apple Silicon?

A Reddit thread compares four 4bit quantization schemes for running Qwen3.6 on Apple silicon. We break down what each format trades off — and why it m

Aug 82 min read
KimiKimi K3

Kimi K3 May Hit 2.8T Parameters as China’s AI Race Shifts to Deployment

A rumored 2.8T-parameter Kimi K3 suggests China’s model race is shifting from launches to post-training capability and deployment cost.

Jul 162 min read
IBMGranite

IBM Open-Sources Granite 4.1: 21 Quantized Versions Prove Bottleneck Isn't Size

IBM open-sources Granite 4.1. A 21-version quantization test shows no quality difference: small models' bottleneck is base capability, not compression

May 52 min read
MistralUnsloth

Mistral Local GGUF Bug Fixed — Open Source QA Gaps Are Bigger Than You Think

Mistral Medium 3.5 GGUF files corrupted, community-fixed. Reveals open source QA gap: APIs tested, local formats not—impacts enterprise deployments.

May 22 min read
MistralUnsloth

Mistral 3.5 Inference Bug Fixed by Open-Source Team — LLM Delivery QA Flashing Red

Unsloth fixed a Mistral Medium 3.5 inference bug from a core config error, exposing absent QA in commercial LLMs. Beware the "community beta" business

May 22 min read
QwenUnsloth

Qwen3.6-27B Quantized Fits Single Consumer GPU: Local Deployment Sweet Spot

Unsloth Q5-quantized Qwen3.6-27B runs stably on a single RTX 5090 across 19 rounds. Mid-size model local deployment is hitting the cost-capability swe

May 12 min read
UnslothQwen3.6

Qwen3.6 GGUF Benchmarks

Un sloth claims top KLD-vs-disk-space performance for Qwen3.6-35B-A3B quants in 21 of 22 pareto frontier comparisons.

Apr 172 min read
llama.cppQwen3

GPoUr with ~12gb vram and a 3080 getting 40tg/s on qwen3.6 35BA3B w/ 260k ctx

A llama.cpp fork with turbo3 KV cache quantization achieves ~40 tok/s on Qwen3-35 B-A3B with only 12GB VRAM.

Apr 162 min read
UnslothMiniMax-M2.7

Unsloth Releases Full GGUF Quant Suite for MiniMax M2.7

Unsloth uploads 22 GGUF quantizations of MiniMax M2.7, ranging from 1-bit (60.7 GB) to BF16 (457 GB).

Apr 122 min read