BitLoRA: Quantization-Compatible Adapter Tuning for 1.58-bit LLM in Federated On-Device AI-Agent

Citations

WEB OF SCIENCE

0

초록

An increasing curiosity in bespoke AI methodologies has revealed the prospects and hurdles that extensive Large language model (LLM) encounter in contexts where resources are constrained and privacy is of paramount importance. Despite the impressive adaptability and generalization shown by LLM, the substantial memory demands and dependence on centralized frameworks introduce challenges for deployment in federated On-Device situations, where both data security and operational effectiveness are paramount. This investigation presents Bit-LoRA, a quantization-compatible adapter tuning framework that amalgamates 1.58 bit BitNet quantization with Parameter-Efficient Fine-Tuning methodologies. BitLoRA guarantees effective model adaptation while preserving stringent data confidentiality by solely transmitting lightweight adapter updates within a Federated Learning milieu. Comprehensive empirical assessments have demonstrated that BitLoRA consistently surpasses the benchmark model concerning accuracy while diminishing GPU memory usage by 85% or more. These findings indicate that BitLoRA not only facilitates substantial fine-tuning of LLM within a constrained computational budget but also furnishes a scalable foundation for the establishment of privacy, resource efficiency, and personalized AI-Agents. This proposed structure creates opportunities for the durable merger of LLM into federal and On-Device frameworks, bridging the gap between advanced model skills and essential deployment considerations.

키워드

BitLoRATernary QuantizationFederated LearningOn-Device AIPersonalized AI-Agent
제목
BitLoRA: Quantization-Compatible Adapter Tuning for 1.58-bit LLM in Federated On-Device AI-Agent
저자
Song, InSeoLee, KangYoon
DOI
10.1016/j.eswa.2026.131397
발행일
2026-05
유형
Article
저널명
Expert Systems with Applications
311