Researchers have presented a paper on using small multimodal LLMs to understand sequences of user interactions on the web and on mobile devices on device, rather than sending data to a server. The work, titled “Small Models, Big Results: Achieving Superior Intent Extraction Through Decomposition”, was presented at EMNLP 2025.
The paper says that helpful agents need to understand what a user is doing, or trying to do, across current and previous tasks so they can anticipate next actions. It gives an example: if a user previously searched for music festivals across Europe and is now looking for a flight to London, the agent could offer to find festivals in London on those specific dates.
According to the source, large multimodal LLMs are already quite good at understanding user intent from a user interface trajectory, but using them usually requires sending information to a server. The paper says this can be slow, costly, and may expose sensitive information.
The approach in the paper breaks intent understanding into two stages: first summarizing each screen separately, then extracting an intent from the sequence of generated summaries. The source says this makes the task more tractable for small models.
The paper also formalizes metrics for evaluating model performance and says the approach yields results comparable to much larger models. The source says this suggests potential for on-device applications and builds on previous work from the team on user intent understanding.
Source: research.google.
Companies can share verified announcements through Newz9’s international press release submission page.

