상세 보기
End-to-end 자율주행을 위한 경량 VLM: Q-Former 학습 유무에 따른 성능 분석
- 조재현;
- 안종현
SCOPUS
0초록
Recent advances in vision-language models (VLMs) have enabled integrated visual perception and language understanding for autonomous driving. However, the high computational cost of large-scale VLMs limits their practical deployment in real-time systems. In this study, we revisit the common assumption that jointly training the Q-Former improves adaptability in End-to-End autonomous driving. Through empirical analysis, we demonstrate that training the Q-Former can, in some cases, hinder performance by introducing representational bottlenecks or disrupting pre-trained visual semantics. Instead, we show that applying Low-Rank Adaptation (LoRA) to the language model alone achieves better accuracy with significantly fewer trainable parameters. Our findings highlight the importance of careful architectural tuning and suggest that simpler adaptation strategies may be more effective for real-world VLM deployment in autonomous driving.
키워드
- 제목
- End-to-end 자율주행을 위한 경량 VLM: Q-Former 학습 유무에 따른 성능 분석
- 제목 (타언어)
- Lightweight Vision-language Models for End-to-end Autonomous Driving: An Analysis of Performance With and Without Q-former Training
- 저자
- 조재현; 안종현
- 발행일
- 2025-11
- 유형
- Y
- 저널명
- 제어.로봇.시스템학회 논문지
- 권
- 31
- 호
- 11
- 페이지
- 1314 ~ 1318