Liquid AI released LFM2.5-VL-3B this week: only 3.1B parameters and roughly 2GB in size, already running locally on the iPhone 17. But the number we think matters more is this one — on the ScreenSpot-v2 desktop benchmark, it jumped from 6 in the previous generation to 78.7.
What this is
Liquid AI is a US startup out of MIT, focused on efficient small models. LFM2.5-VL-3B is a vision-language model (VL — can see images and read text), with only 3.1B parameters and around 2GB of weights, small enough to fit in a typical phone's memory and run independently, with no cloud dependency.
Compared to the previous generation, the key improvement isn't "can it recognize what's in the image" but "can it point to where it is." The original poster's test: hand it a photo of a Minecraft Steve toy and ask, and instead of just answering "this is Steve," the model described Steve's appearance part by part — a direct demonstration of strengthened localization. The cost is 2 minutes 31 seconds of inference time, still too slow on a phone.
Industry view
The optimistic side sees a clear trend: models small enough to run locally mean sensitive images (production lines, documents, customer sites) can get visual analysis without upload — directly good news for manufacturing, healthcare, and government scenarios.
But critics flag three issues. First, a 2+ minute wait is essentially unusable on a phone, far from "production-ready." Second, cloud large models (GPT-4V, Gemini, Qwen-VL) still lead on vision capability, and there's no industry consensus on whether local small models are "good enough" or "left behind." Third, the original poster is himself the founder of the app running this model — there's an inherent conflict of interest in the test conclusions.
Impact on regular people
For enterprise IT: data-sensitive industries (manufacturing inspection, retail store audits, preliminary medical imaging screening) can start evaluating local vision solutions — images stay on-premises, never reach the cloud. This is the most direct landing scenario for this wave of device-side (running locally on the endpoint device) models.
For professionals: no short-term change to daily work — nobody wants to wait 2 minutes for a description. If inference can be compressed to under 10 seconds in the future, local vision assistants become a practical tool during business trips and on-site visits.
For consumer market: these capabilities will gradually appear in photo-identification and document-scanning apps, but "can run" doesn't mean "works well" — the average user's experience won't improve noticeably in the short term.