Not available in English yet
你的客户可能不再打开App — OpenAI的AI手机提前一年来了
Related Reading
From ai_news
Engineer Pushes Qwen to the Limit: 260K Tokens Is Local AI's Hard Ceiling
Engineer pushed Qwen to extremes: context over 100K tokens drops generation 75%. Long-context remains local AI's hard ceiling—proof enterprises can't
Million-Token Context on Two 5090s — Amateur Dev Shatters Enterprise AI Myth
Reddit developer NInfer hits 1.04M token context on consumer RTX 5090s at 119 tok/s — 2.8x faster than vLLM with a 27B Qwen model.
llama.cpp Has 50 PRs Pending — Local AI No Longer Needs a High-End GPU
Open-source llama.cpp has 50+ performance PRs pending merge, some claiming 3x CPU inference speedup. Local LLM deployment is shedding its dependence o
DeepSeek Hits 67 token/s on Two $9K Mini Boxes — Local LLM Floor Is Caving In
Reddit user hit 67-84 token/s on DeepSeek V4 Flash with a 1M-token context window on two ~$9K NVIDIA DGX Sparks. The local-LLM cost barrier is collaps
Tenstorrent Runs Qwen3.7-27B — Non-NVIDIA AI Chip Breaks Commercial Ice
Tenstorrent user shares inference data on QuietBox 2 running Qwen3.7-27B. First near-commercial benchmark from the non-NVIDIA camp, but still far from
Financial LLM Benchmark Exposed: Each Model Wears Its Own Gear—Is That Fair?
Reddit dissected Ling's financial LLM benchmark: tests used different reasoning, agents, tools. Rankings measure setups, not models—a buyer alert.