Large-Scale User Behavior Analysis in Multimodal AI-Assisted Manual Task Execution
| Source: arXiv AI
Tags: conversational AI, multimodal AI, user behavior, HCI, task assistants, UX research
A large-scale real-world study of thousands of users interacting with multimodal AI task assistants (cooking/DIY) in-the-wild reveals key gaps between controlled lab findings and actual deployment — with concrete design guidelines for user interaction flows, intent patterns, and satisfaction drivers.
Details
Conversational Task Assistants (CTAs) support users in complex real-world tasks like cooking and DIY through voice, text, image, and video. Prior research focused on controlled lab settings, leaving limited understanding of how people actually use CTAs at scale.\n\nThis study presents findings from thousands of real-world users interacting with a deployed CTA system in-the-wild. The analysis covers four dimensions: user-CTA interaction flows, user intent patterns, conversational traits, and behavioral factors associated with user satisfaction. The findings reveal concrete design opportunities in interaction design and task engagement, concluding with actionable design guidelines.\n\nThe paper was accepted at EPIA 2026 (European Conference on AI Applications). Methodology uses in-the-wild behavioral data analysis rather than controlled experiments — a meaningful ecological validity advantage.\n\nThe study does not name the specific CTA product, which limits context. Findings are behavioral patterns applicable broadly to multimodal copilot design. Practitioners building AI assistants, enterprise copilots, or multimodal task support systems will find the design guidelines directly applicable.