Back to home
MTP
4 articles tagged with this topic
OrnithMTP
Ornith 1.5 Shipped with an Untrained MTP Head — Open-Source AI's QC Problem
Ornith 1.5 shipped with an untrained MTP head. Not a bug — a symptom of QA gaps in open-source AI that any cost-cutting enterprise should heed.
Aug 202 min read
llama.cppMTP
llama.cpp Adds Adaptive MTP — One Fewer Knob for Local LLM Users
llama.cpp adds adaptive MTP — the model picks its own token depth. Coding up to 2x faster, prose ~3% slower. A shift from manual tuning to self-tuning
Aug 172 min read
QwenRTX 3090
Consumer GPU Hits 100K Context: Local LLM Hardware Thresholds Drop Fast
We see an RTX 3090 run a 27B model, 100K context, 50 tokens/s via quant+MTP+KV compression. Consumer inference now rivals last year's enterprise setup
May 72 min read
llama.cppMTP
llama.cpp MTP Hits Beta: Local LLM Inference Speed Gap Narrowing
llama.cpp MTP beta supports Qwen3.5. With tensor parallelism maturing, the local-cloud inference speed gap is narrowing, making local LLM deployment m
May 42 min read