1 article tagged with this topic
MoE offload support in Freetoken lets 24GB consumer GPUs attempt 70B models — previously a 48GB minimum. Another small step for edge AI.