Bhavana Verma, Priyanka Meel, Dinesh Kumar Vishwakarma
Sarcasm rarely lives in the words alone. A caption reading "what a beautiful day" over a photo of torrential rain means the opposite of what it says, and neither the text nor the image carries that meaning by itself. It exists in the mismatch. The same is true of much online bullying, where an innocuous phrase turns cruel next to a particular picture.
Most multimodal systems handle this by fusing features or running cross-modal attention, which mixes the two streams and hopes the contradiction survives the blending. The authors argue that this misses inconsistencies that live at different levels of representation, some in fine detail, some in overall meaning.
HCIG models the incongruity itself as a hierarchical graph rather than treating it as a by-product of fusion. Making the mismatch the object of study is the right instinct here, because in sarcasm the mismatch is not noise around the signal, it is the signal.
Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multimodal approaches primarily rely on feature fusion or cross-modal attention, which may not effectively capture hierarchical semantic inconsistencies across different levels of representation. To address this limitation, this paper proposes HCIG (Hierarchical…
Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening
arXiv (cs.CV) · July 16, 2026Weakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based Diffusion
arXiv (cs.CV) · July 16, 2026Video = World + Event Stream
arXiv (cs.AI) · July 16, 2026Man, Machine, and Masterpiece: Artistic Ownership in the AI Era