1.湖南工商大学计算机学院,湖南省长沙市410205
2.统计学习与智能计算湖南省重点实验室,湖南省长沙市410205
3.湖南工商大学数字媒体工程与人文学院,湖南省长沙市410205
[ "刘耀" ]
收稿:2025-02-13,
修回:2025-04-20,
录用:2025-04-28,
移动端阅览
黄少年, 彭永涛, 文沛然, 等. 融合注意力协同和对比学习的跨模态突发事件识别方法[J/OL]. 智能科学与技术学报, 2025.
HUANG Shaonian, PENG Yongtao, WEN Peiran, et al. A Cross-Modal Emergency Recognition Method Integrating Attentional Collaboration and Contrastive Learning[J/OL]. Chinese Journal of Intelligent Science and Technology, 2025.
针对跨模态突发事件识别中图像复杂、文本有限以及模态间误导信息干扰等问题,提出了一种融合注意力协同和对比学习的跨模态融合方法。首先,通过构建融合多支路自注意力的图像特征编码器和融合外部知识的文本特征编码器,分别提取图像的深层语义特征和文本的上下文关联信息,以克服单一模态信息表达的局限性。其次,为有效捕捉图像-文本特征间的复杂关联,设计跨模态注意力协同网络增强模态间特征的交互一致性,减少误导信息的传递。最后,引入对比学习机制,采用联合监督策略优化图像-文本特征的贡献,实现模态间信息的动态平衡。在公开数据集CrisisMMD和自建数据集上开展的实验证明了提出方法的优越性。
To address the challenges of image complexity
limited textual information
and inter-modal misleading data in cross-modal emergency recognition
this paper proposes a novel fusion recognition method that integrates attentional collaboration and contrastive learning. First
a multi-branch self-attention mechanism is integrated into an image feature encoder
while a text feature encoder is augmented with external knowledge. These encoders are designed to extract rich semantic features from images and contextual information from text
respectively
for overcoming the limitations of single-modal representations. Next
a cross-modal attentional collaboration network is introduced to capture intricate relationships between image-text features
enhancing inter-modal consistency and mitigating the influence of misleading information. Finally
a contrastive learning mechanism
that is optimized through a joint supervision strategy
is employed to dynamically balance the contributions of image and text features. Extensive experiments conducted on the public CrisisMMD dataset and the self-constructed dataset demonstrate the superior performance of the proposed method.
管泽礼 , 杜军平 , 薛哲 等 . 基于强化联邦GNN的个性化公共安全突发事件检测 [J ] . 软件学报 , 2024 , 35 ( 4 ): 1774 - 1789 .
Guan Zeli , Du Junping , Xue Zhe , et al . Personalized public Security Emergency Detection based on enhanced Federal GNN [J ] . Journal of Software , 2024 , 35 ( 4 ): 1774 - 1789 .
Abhina Kumar , Jyoti Prakash Singh , Nripendra P . Rana, et al. Multi-channel convolutional neural network for the identification of eyewitness tweets of disaster[J ] . Information Systems Frontiers , 2023 , 25 ( 4 ): 1589 - 1604 .
周红磊 , 张海涛 , 栾宇 等 . 基于文本-图像增强的突发事件识别及分类方法研究 [J ] . 情报理论与实践 , 2024 , 47 ( 4 ): 181 - 188 .
Zhou Honglei , ZHANG Haitao , Luan Yu et al . Research on emergency recognition and classification based on text-image enhancement [J ] . Information Theory & Practice , 2024 , 47 ( 4 ): 181 - 188 .
Wang Zhihong , Guo Yi , Wang Jiahui . Empower Chinese event detection with improved atrous convolution neural networks [J ] . Neural Computing and Applications , 2021 , 33 ( 11 ): 5805 - 5820 .
Li Jiankai , Wang Yuhong , Li Weixin . MGMP: Multimodal graph message propagation network for event detection [C ] // International Conference on Multimedia Modeling . Cham : Springer International Publishing , 2022 : 141 - 153 .
陈宏 , 钱胜胜 , 李章明 等 . 基于多模态掩码Transformer网络的社会事件分类 [J ] . 北京航空航天大学学报 , 2024 , 50 ( 2 ): 579 - 587 .
Chen Hong , QIAN Shengsheng , Li Zhangming , et al . Social event classification based on Multimodal Mask Transformer Network [J ] . Journal of Beijing University of Aeronautics and Astronautics , 2024 , 50 ( 2 ): 579 - 87 .
Rani Koshy , Sivasankar Elango . Multimodal tweet classification in disaster response systems using transformer-based bidirectional attention model [J ] . Neural Computing and Applications , 2023 , 35 ( 2 ): 1607 - 1627 .
Muhammad Mujahid , Khadija Kanwal , Furqan Rustam , et al . Arabic ChatGPT tweets classification using RoBERTa and BERT ensemble model [J ] . ACM Transactions on Asian and Low-Resource Language Information Processing , 2023 , 22 ( 8 ): 1 - 23 .
Prakash Babu Yandrapati , R. Eswari . Classifying informative tweets using feature enhanced pre-trained language model [J ] . Social Network Analysis and Mining , 2024 , 14 ( 1 ): 48 .
Zhou Han , Yin Hongpeng , Zheng Hengyi , et al . A survey on multi-modal social event detection [J ] . Knowledge-Based Systems , 2020 , 195 : 105695 .
Mathilde Brousmiche , Jean Rouat ,, Stéphane Dupont . Multimodal attentive fusion network for audio-visual event recognition [J ] . Information Fusion , 2022 , 85 : 52 - 59 .
Abhinav Kumar , Jyoti Prakash Singh , Yogesh K . Dwivedi, et al. A deep multi-modal neural network for informative Twitter content classification during emergencies[J ] . Annals of Operations Research , 2022 : 1 - 32 .
Iustin Sirbu , Tiberiu Sosea , Cornelia Caragea , et al . Multimodal semi-supervised learning for disaster tweet classification [C ] // Proceedings of the 29th international conference on computational linguistics . 2022 : 2711 - 2723 .
Sreenivasulu Madichetty , Sridevi M . A novel method for identifying the damage assessment tweets during disaster [J ] . Future Generation Computer Systems , 2021 , 116 : 440 - 454 .
Sreenivasulu Madichetty , Sridevi M . A stacked convolutional neural network for detecting the resource tweets during a disaster [J ] . Multimedia tools and applications , 2021 , 80 ( 3 ): 3927 - 3949 .
Abhishek Upadhyay , Yogesh Kumar Meena , Ganpat Singh Chauhan . SatCoBiLSTM: Self-attention based hybrid deep learning framework for crisis event detection in social media [J ] . Expert Systems with Applications , 2024 , 249 : 123604 .
Firoj Alam , Ferda Ofli , Muhammad Imran . Processing social media images by combining human and machine computing during crises [J ] . International Journal of Human–Computer Interaction , 2018 , 34 ( 4 ): 311 - 327 .
Daryl B. Valdez , Rey Anthony G . Godmalin. A deep learning approach of recognizing natural disasters on images using convolutional neural network and transfer learning [C ] // Proceedings of the international conference on artificial intelligence and its applications . 2021 : 1 - 7 .
Christos Kyrkou , heocharis Theocharides . Deep-Learning-Based Aerial Image Classification for Emergency Response Applications Using Unmanned Aerial Vehicles [C ] // CVPR workshops . 2019 : 517 - 525 .
Ganesh Nalluru , Rahul Pandey , Hemant Purohit . Relevancy classification of multimodal social media streams for emergency services [C ] // 2019 IEEE International Conference on Smart Computing (SMARTCOMP) . IEEE , 2019 : 121 - 125 .
Douwe Kiela , Suvrat Bhooshan , Hamed Firooz , et al . Supervised multimodal bitransformers for classifying images and text [J ] . arXiv preprint arXiv: 1909.02950 , 2019 .
Valentin Vielzeu , Alexis Lechervy , Stephane Pateux , et al . Centralnet: a multilayer approach for multimodal fusion [C ] // Proceedings of the European Conference on Computer Vision (ECCV) Workshops . 2018 : 575 - 589 .
John Arevalo , Thamar Solorio , Manuel Montes-y-Gómez , et al . Gated multimodal networks [J ] . Neural Computing and Applications , 2020 , 32 : 10209 - 10228 .
Liunian Harold Li , Mark Yatskar , Da Yin , et al . Visualbert: A simple and performant baseline for vision and language [J ] . arXiv preprint arXiv: 1908.03557 , 2019 .
Zhicheng Huang , Zhaoyang Zeng , Bei Liu , et al . Pixel-bert: Aligning image pixels with text by deep multi-modal transformers [J ] . arXiv preprint arXiv: 2004.00849 , 2020 .
Wonjae Kim , Bokyung Son , Ildoo Kim . Vilt: Vision-and-language transformer without convolution or region supervision [C ] // International conference on machine learning . PMLR , 2021 : 5583 - 5594 .
Akira Fukui , Dong HukPark , Daylen Yang , et al . Multimodal compact bilinear pooling for visual question answering and visual grounding [J ] . arXiv preprint arXiv: 1606.01847 , 2016 .
Wang Jing , Yang Shuo , Zhao Hui . Crisis event summary generative model based on hierarchical multimodal fusion [J ] . Pattern Recognition , 2023 , 144 : 109890 .
Lv Jiandong , Wang Xingang , Shao Cuiling . AMAE: Adversarial multimodal auto-encoder for crisis-related tweet analysis [J ] . Computing , 2023 , 105 ( 1 ): 13 - 28 .
Mahdi Abavisani , Wu Liwei , Hu Shengli , et al . Multimodal categorization of crisis events in social media [C ] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2020 : 14679 - 14689 .
Wu Xuehua , Mao Jin , Xie Hao , et al . Identifying humanitarian information for emergency response by modeling the correlation and independence between text and images [J ] . Information Processing & Management , 2022 , 59 ( 4 ): 102977 .
Shubham Gupta , Nandini Saini , Suman Kundu , et al . CrisisKan: Knowledge-infused and explainable multimodal attention network for crisis event classification [C ] // European Conference on Information Retrieval . Cham : Springer Nature Switzerland , 2024 : 18 - 33 .
Mao Yudong , Jiang Qiuping , Cong Runmin , et al . Cross-modality fusion and progressive integration network for saliency prediction on stereoscopic 3D images [J ] . IEEE Transactions on Multimedia , 2021 , 24 : 2435 - 2448 .
Kevin Clark , Minh-Thang Luong , Quoc V . Le, et al. Electra: Pre-training text encoders as discriminators rather than generators[J ] . arXiv preprint arXiv: 2003.10555 , 2020 .
Firoj Alam , Ferda Ofli , Muhammad Imran . Crisismmd: Multimodal twitter datasets from natural disasters [C ] // Proceedings of the international AAAI conference on web and social media . 2018 , 12(1) .
Jacob Devlin , Ming-Wei Chang , Kenton Lee , et al . Bert: Pre-training of deep bidirectional transformers for language understanding [C ] // Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) . 2019 : 4171 - 4186 .
Mariham Rezk , Noureldin Elmadany , Radwa K . Hamad, et al. Categorizing crises from social media feeds via multimodal channel attention[J ] . IEEE Access , 2023 , 11 : 72037 - 72049 .
Liang Tao , Lin Guosheng , Wan Mingyang , et al . Expanding large pre-trained unimodal models with multimodal information injection for image-text multimodal classification [C ] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2022 : 15492 - 15501 .
Onkar Susladkar , Gayatri Deshmukh , Dhruv Makwana , et al . Gafnet: A global fourier self attention based novel network for multi-modal downstream tasks [C ] // Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 2023 : 5242 - 5251 .
0
浏览量
2
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621
