inclusionAI
5 articles tagged with this topic
4060Ti Replaces Gemma-12B — Big Models' Side Hustle Is Being Redistributed
A Reddit user replaced Gemma-12B with Ling Tiny on a 4060Ti, calling the speed "amazing." Auxiliary AI tasks are migrating to consumer hardware.
$249 Box Runs 128K LLM at 33 tok/s — Edge AI Starts Prying at Cloud
$249 Jetson Orin Nano runs inclusionAI's Ling-3.0-tiny (7.9B, 128K context) at 33 tok/s — enterprise AI hardware could drop from rack to desktop scale
Ling-3 Tiny Outscores Qwen3.5 9B — Our Open-Source Attention Is Too Concentrated
Reddit found Ling-3 Tiny (1-2B params) outperforming Qwen3.5 9B on reasoning benchmarks. Not an isolated case — strong mid-tier open-source labs are g
124B model ran stably 7 min on desktop — local AI crosses usability threshold
Reddit test: 124B Ling model held 35.7 tok/s for 15K+ tokens on a single NVIDIA DGX Spark — local LLMs shifting from geek toy to enterprise option.
Two Flags Nearly Double Small Model Throughput — But the Hidden Compatibility Trap Matters More
InclusionAI's Ling-3.0-flash INT4 hits 38.7 tok/s on DGX Spark with two config tweaks — but default vLLM silently breaks V3 architecture, producing fl