llama.cpplocal LLM
llama.cpp Adds 'Lazy Loading': 100B Models Run Locally, Hardware Bar Drops
llama.cpp adds TENSOR_READ_LAZY: weights load on-demand instead of fully residing in memory. Enables larger models on consumer hardware—a real shift f
2d ago·2 min read