Back to home
AI compute
2 articles tagged with this topic
QwenAlibaba
Qwen Makes Thinking Depth Adjustable — Alibaba Lets LLMs Allocate Compute On Demand
Alibaba's Qwen now lets users adjust 'thinking depth'—quick answers for easy questions, more reasoning for hard ones. LLMs shift from on/off switch to
15h ago2 min read
local deploymentinference acceleration
Running 122B LLM at 198 tok/s on Local Deployment: Countdown to Cloud AI Rental Providers' Doom
Consumer GPUs matching enterprise inference speeds locally signals cloud AI rental's disruption—should bosses renew contracts or build in-house?
Apr 102 min read