IMPACT-Agent: Intelligent Multimodal Prompt-Driven Agent for Collaborative Teaming

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

The future of disaster response, search and rescue (SAR) technology is rapidly evolving with the integration of Artificial Intelligence (AI) and intelligent unmanned platforms. Manned-Unmanned Teaming (MUM-T) is revolutionizing on-site disaster structure and emergency rescue strategies by synchronizing manned and unmanned vehicles to support operations. This paper investigates the potential applications of Multimodal Large Language Models (MLLMs) in intelligent MUM-T systems. We present a system capable of using aerial imagery to identify and locate survivors and hazards/risks, and to generate safe path plans for unmanned ground vehicles (UGVs). The main contributions are: 1) the implementation of a Few-Shot In-Context Learning (ICL) method for MLLMs, which eliminates the need for fine-tuning or re-training the AI model, and 2) the application of the Chain-of-Thought (CoT) prompting technique. The ICL method reduces the time required for data gathering and model training, while the CoT reasoning technique improves the path planning generated by MLLMs. Our experiments, using aerial imagery, demonstrate that the proposed system achieves up to 86% mean Average Precision (mAP) in survivor and hazard identification through 5-shot ICL. Furthermore, applying the CoT technique doubled the UGV's path generation success rate from 20% to 40%. The proposed MLLM-based system effectively performed open-vocabulary object detection without additional training. This research highlights the potential of MLLMs for diverse and complex disaster relief missions in intelligent MUM-T systems, presenting an innovative approach to enhance capabilities and mission success while reducing human risk and operational costs. A demonstration video of our system is available at https://youtu.be/vReaSMDWu8o?si=RluQsqS4O_j5lVcS

키워드

DisastersHazardsArtificial intelligenceDecision makingCollaborationTrainingLarge language modelsNavigationMulti-agent systemsToy manufacturing industrySearch and rescue (SAR)manned-unmanned teaming (MUM-T)observeorientdecideact (OODA) loopmultimodal large language models (MLLMs)few-shot in-context learning (ICL)Chain-of-Thought (CoT) promptinghuman-machine collaboration
제목
IMPACT-Agent: Intelligent Multimodal Prompt-Driven Agent for Collaborative Teaming
저자
Cho, TaewanYou, HyunwooChoi, Andrew Jaeyong
DOI
10.1109/ACCESS.2026.3674148
발행일
2026-03
유형
Article
저널명
IEEE Access
14
페이지
50575 ~ 50583

파일 다운로드