NeXt-DETR: A scalable and efficient transformer-based detector for resource-constrained systems

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

2

초록

While Transformer-based object detectors, particularly DETR variants, have demonstrated strong benchmark performance, their high computational costs during both training and inference, coupled with rigid architectural designs, often hinder real-world deployment in resource-constrained systems. We propose NeXt-DETR to address these challenges: a scalable and efficient detection framework optimized for diverse hardware from edge devices to servers. Its memory-efficient engineering enables stable training with small batch sizes, an efficiency rooted in two key architectural innovations. First, the framework employs a ConvNeXt-v2 backbone that deliberately eliminates batch normalization, ensuring training stability under micro-batch conditions. Second, we introduce FuNeXt, a novel module that optimizes the encoder by integrating large-kernel depthwise convolutions with an efficient attention branch. This lightweight multi-scale fusion block efficiently expands the encoder's receptive field to enhance feature representation with minimal computational overhead. The NeXt-DETR family spans from a 5.4 M parameter Atto model to a 111 M Base model. Notably, the 8.3 M parameter Femto variant achieves 46.4 AP on MS COCO, virtually matching RT-DETR-R18 while using 59 % fewer parameters and 61 % less computation. The Base model achieves 55.4 AP , outperforming RT-DETR-R101 by +1.1 AP with comparable complexity. Real-world latency tests further validate its efficiency: all variants run under 60ms on an H100 GPU, while on an RTX 4070, the Atto model achieves 58ms, demonstrating scalability across hardware. Ultimately, NeXt-DETR's superior accuracy-efficiency trade-off, architectural flexibility, and practical features-stable, batch-normalization-free training and fast multi-platform inference-bridge the gap between academic research and deployment-ready systems.

키워드

Transformer-based object detectionLightweight deep learning modelsEdge and resource-constrained deploymentMulti-scale feature fusionScalable model architecture
제목
NeXt-DETR: A scalable and efficient transformer-based detector for resource-constrained systems
저자
Choi, Chan-YoungLee, Sang-Woong
DOI
10.1016/j.ins.2025.122913
발행일
2026-04
유형
Article
저널명
Information Sciences
731