Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · July 17, 2026

HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

Bhavana Verma, Priyanka Meel, Dinesh Kumar Vishwakarma

Sarcasm rarely lives in the words alone. A caption reading "what a beautiful day" over a photo of torrential rain means the opposite of what it says, and neither the text nor the image carries that meaning by itself. It exists in the mismatch. The same is true of much online bullying, where an innocuous phrase turns cruel next to a particular picture.

Most multimodal systems handle this by fusing features or running cross-modal attention, which mixes the two streams and hopes the contradiction survives the blending. The authors argue that this misses inconsistencies that live at different levels of representation, some in fine detail, some in overall meaning.

HCIG models the incongruity itself as a hierarchical graph rather than treating it as a by-product of fusion. Making the mismatch the object of study is the right instinct here, because in sarcasm the mismatch is not noise around the signal, it is the signal.

From the arXiv (cs.AI) abstract

Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multimodal approaches primarily rely on feature fusion or cross-modal attention, which may not effectively capture hierarchical semantic inconsistencies across different levels of representation. To address this limitation, this paper proposes HCIG (Hierarchical…


More Artificial Intelligence papers