What this is
Over the past few months, a researcher reviewed dozens of papers on "LLM Agents" — AI assistants that autonomously operate mobile apps — and reached a sobering conclusion: the industry currently has no reliable way to measure whether phone-based AI assistants can truly work like a human. He compiled the relevant papers into a reading list (sourced from alphaxiv.org) and surfaced three recurring patterns. First, the vast majority of evaluations run on emulators, making real-device data nearly impossible to obtain. Second, the "deployment-hard" metrics — battery drain, heat, thermals — are almost never measured. Third, everyday tasks are scattered across different benchmarks, languages, and apps, with no globally adopted "real-user task set."
Industry view
The mainstream narrative is optimistic — phone Agents are widely treated as the next platform entry point, and Google, Apple, Samsung, and Chinese phone OEMs are all doubling down. But this paper list is pouring cold water on that story: the entire industry sits in a state of "emulator confidence, real-world blindness." What should give us pause is that high emulator scores cannot predict power draw on an actual phone, will not surface the failure when a banking app redesigns its UI and the Agent stalls, and certainly will not capture the multitasking scenario of a Chinese user juggling WeChat while switching screens. Our recommendation: any enterprise seriously considering procuring "phone-based AI assistants" in 2025 should first move the vendor's demo onto ten real devices and run it for a week. Benchmark scores can wait.
Impact on regular people
For enterprise IT: when procuring mobile AI tools, don't be fooled by demo videos and benchmark numbers — demand that vendors deliver test reports built on real devices, real tasks, and real battery consumption.
For working professionals: don't yet treat "letting AI operate your phone" as stable productivity; right now it's closer to a flash-of-brilliance toy that breaks down when it matters most.
For the consumer market: phone makers will aggressively push "AI assistants" as their new selling point, but user口碑 will expose the problems faster than any lab dataset — the cadence will look a lot like the early smart-speaker era.