1 article tagged with this topic
Developer tomtsai28 compressed Gemma-2B inference into 5KB of pure x86-64 assembly, hitting 4.6 tokens/sec on an old i5 desktop. Zero PyTorch, zero C