Back to home
AMD Strix Halo
2 articles tagged with this topic
LocalLLaMAllama.cpp
AI Inference Goes Specialist: New Engines Drop Generality to Max Out Hardware
A wave of narrow inference engines optimized for one model + one chip is emerging. AI deployment is shifting from "one engine for all" to "one engine
Oct 32 min read
GufoHalogen
Claimed 70 tokens/sec fails real test — AI benchmark inflation strikes again
Gufo claims 70 tokens/sec; real tests show ~38, 13% slower than Halogen. Benchmark task lets speculative decoding cheat. Another AI inflation case.
Oct 12 min read