Back to home
speculative-decoding
5 articles tagged with this topic
DeepSeekspeculative-decoding
Speculative decoding is becoming standard — open-source LLMs now predict ahead
Reddit users spotted speculative decoding working on local GPUs—AI instantly outputting phrases via MTP. The local inference cost curve is being quiet
Aug 292 min read
DeepSeekAMD
DeepSeek V4 at 22-28 Tokens/sec on a Mini PC — Local LLMs May Finally Be Usable
We see DeepSeek V4 Flash hit 22-28 tokens/sec on AMD Strix Halo. Chinese open-source LLMs shifting from 'big' to local — but hardware costs stay high.
Aug 182 min read
inclusionAILing-3.0-flash
Two Flags Nearly Double Small Model Throughput — But the Hidden Compatibility Trap Matters More
InclusionAI's Ling-3.0-flash INT4 hits 38.7 tok/s on DGX Spark with two config tweaks — but default vLLM silently breaks V3 architecture, producing fl
Aug 92 min read
AWS-Trainium2vLL M
Speculative Decoding on AWS Trainium2 Cuts LLM Lat ency Up to 3x
AWS benchmarks show speculative decoding with vLLM on Trainium2 reduces inter -token latency up to 3x for decode-heavy workloads.
Apr 152 min read
MLXQwen3.5
DFlash speculative decoding on Apple Silicon: 4.1x on Qwen3.5-9B, now open source (MLX, M5 Max)
Open-source DFlash achiev es 4.13x speedup on Qwen3.5-9B using MLX on M5 Max with 89.4% token acceptance rate.
Apr 132 min read