Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.CV) · July 17, 2026

Vision-Language Assistant for Emotional Reactions to Risky Driving

Harine Choi, Eun Hak Lee, Zhengzhong Tu

Vision-language models have become good at describing what is happening on a road and reasoning about it. What they almost never consider is how the moment feels to the person in the driver's seat, which is odd, because that is the part determining whether anyone acts on the warning.

Keep Yelling Assistant watches for high-risk manoeuvres in real time, a sudden cut-in being the example given, and answers with an emotionally expressive response instead of a flat alert. A vision module handles the perception, and a language model produces the reaction, shaped to the individual driver's preferences.

The intuition is one every passenger knows. A friend who gasps when a car swerves in front of you conveys the danger faster than any chime on the dashboard, and how much you want that friend to react is very much a matter of personal taste. Tailoring the response to preference is an admission that the same warning does not suit everybody.

From the arXiv (cs.CV) abstract

This study introduces a vision-language pipeline that detects risky driving behaviors and generates emotionally expressive responses to support driver awareness and comfort. Although vision-language models have advanced perception and reasoning in autonomous driving, existing systems rarely consider the emotional dimension or real-world user experience. Keep Yelling Assistant (KYA) detects high-risk driving maneuvers in real time, such as sudden cut-ins. It then produces…


More Artificial Intelligence papers