Unsloth
19 articles tagged with this topic
HF 收购疑云下,Unsloth 让普通显卡跑得动大模型 — 开源生态的真正护城河
Under HF acquisition rumors, Unsloth lets 24GB GPUs run 70B models. Small open-source teams—not platforms—will decide if AI stays affordable.
Reddit Pushes Alibaba 35B Re-Quantization — Open-Source Compute Drops a Notch
A Reddit request urges re-quantizing Alibaba's Qwen 35B with Unsloth's new UD 3.0 — open-source AI compute keeps getting cheaper.
Qwen Flash Drops, Unsloth Adopts Day Zero — Open-Source AI's Pace Just Changed
Qwen Flash launched; Unsloth adopted it same day. The release-to-consumer-hardware gap shrank to under 24 hours. Enterprise IT should take note.
$500 GPU ships the first real PR — local AI coding exits demo
An indie dev ran Qwen 27B on a 4060Ti (~3,000 RMB) and shipped a human-reviewed PR — open-source local AI's first full engineering cycle.
Qwen 27B Compressed to 1-2 Bits — Local AI Memory Savings, Quality Wavers
Reddit's jojohai quantized Qwen3.8-27B to Q1-Q2 with MTP baked into weights. Lower memory than external MTP, but the model "gets stupid" off thinking
Qwen 27B Runs 80-Step Agent on a Single GPU — Local Models Can Now Do Real Work
A Reddit test caught our eye: Qwen 27B on a consumer GPU made 80 autonomous tool calls from one prompt. The "cloud-only Agent" default is crumbling.
Unsloth squeezes Qwen onto 8GB laptops — local LLMs now run on almost anything
Unsloth's Dynamic v3.0 lets Qwen models run on 8GB laptops, with 1-bit versions keeping 77% accuracy — easing local enterprise LLM deployment.
Qwen 3.8 Open-Source: 27B on RTX 4090 — Closed-Source Flagship Moat Loosens
Qwen 3.8 open-weights: 27B beats Claude Opus 4.6 Max on coding/agent benchmarks. 4-bit fits 24GB. New option for enterprise IT, dev tools, hardware.
Unsloth Releases 30B Open-Source Model — Local LLMs Step Out of the Geek Bubble
Unsloth ships a GGUF-quantized Muse-Glimmer-30B, letting consumer laptops run a 30-billion-parameter LLM locally. On-prem AI is shifting from hobbyist
Kimi K3 Trimmed 33% for English-Only — Local LLMs Are Getting Quietly Affordable
Reddit user hellohazime trimmed Moonshot's Kimi K3 from 711GB to 478GB by stripping multilingual weights. A 2-bit variant reportedly beats its predece
MLX 4bit Quantization Showdown: Which Compression Format Actually Wins on Apple Silicon?
A Reddit thread compares four 4bit quantization schemes for running Qwen3.6 on Apple silicon. We break down what each format trades off — and why it m
Kimi K3 May Hit 2.8T Parameters as China’s AI Race Shifts to Deployment
A rumored 2.8T-parameter Kimi K3 suggests China’s model race is shifting from launches to post-training capability and deployment cost.
IBM Open-Sources Granite 4.1: 21 Quantized Versions Prove Bottleneck Isn't Size
IBM open-sources Granite 4.1. A 21-version quantization test shows no quality difference: small models' bottleneck is base capability, not compression
Mistral Local GGUF Bug Fixed — Open Source QA Gaps Are Bigger Than You Think
Mistral Medium 3.5 GGUF files corrupted, community-fixed. Reveals open source QA gap: APIs tested, local formats not—impacts enterprise deployments.
Mistral 3.5 Inference Bug Fixed by Open-Source Team — LLM Delivery QA Flashing Red
Unsloth fixed a Mistral Medium 3.5 inference bug from a core config error, exposing absent QA in commercial LLMs. Beware the "community beta" business
Qwen3.6-27B Quantized Fits Single Consumer GPU: Local Deployment Sweet Spot
Unsloth Q5-quantized Qwen3.6-27B runs stably on a single RTX 5090 across 19 rounds. Mid-size model local deployment is hitting the cost-capability swe
Qwen3.6 GGUF Benchmarks
Un sloth claims top KLD-vs-disk-space performance for Qwen3.6-35B-A3B quants in 21 of 22 pareto frontier comparisons.
GPoUr with ~12gb vram and a 3080 getting 40tg/s on qwen3.6 35BA3B w/ 260k ctx
A llama.cpp fork with turbo3 KV cache quantization achieves ~40 tok/s on Qwen3-35 B-A3B with only 12GB VRAM.
Unsloth Releases Full GGUF Quant Suite for MiniMax M2.7
Unsloth uploads 22 GGUF quantizations of MiniMax M2.7, ranging from 1-bit (60.7 GB) to BF16 (457 GB).