Optimizing cryptocurrency trades with twin delayed DDPG: Adaptive multi-factor reward function with diverse data sources

Citations

WEB OF SCIENCE

1

초록

Cryptocurrency trading presents significant opportunities but also poses unique challenges due to extreme market volatility, complex price dynamics, and the influence of diverse data sources such as social sentiment and blockchain activity. Traditional trading models struggle to adapt to rapid price fluctuations, often relying on simplistic buy-sell-hold strategies that fail to optimize trade quantities and risk exposure. To address these challenges, this study proposes a Twin Delayed Deep Deterministic Policy Gradient (TD3)-based trading model that leverages historical price data, sentiment analysis from cryptocurrency-related tweets, and on-chain metrics for Bitcoin, Litecoin, and Ethereum. Unlike existing approaches, TD3 mitigates overestimation bias, stabilizes policy updates, and improves decision-making in large action spaces. To enhance profitability and risk control, we introduce an Adaptive Multi-Factor Reward Function (AMRF), incorporating profit maximization, risk management, active trading, and confidence-based adjustments. The model's performance is evaluated using Return on Investment (ROI) and Sharpe Ratio (SR), achieving 38.7% ROI and 2.438 SR for Bitcoin, 25.71% ROI and 2.609 SR for Litecoin, and 30.04% ROI and 2.632 SR for Ethereum. Results demonstrate that the proposed TD3-based strategy outperforms conventional methods, offering a more robust and adaptive framework for cryptocurrency trading.

키워드

TD3CryptocurrencyTradingReinforcement learningBITCOINDIVERSIFICATIONEXCHANGERETURN
제목
Optimizing cryptocurrency trades with twin delayed DDPG: Adaptive multi-factor reward function with diverse data sources
저자
Otabek, SattarovChoi, Jaeyoung
DOI
10.1016/j.eswa.2026.131527
발행일
2026-06
유형
Article
저널명
Expert Systems with Applications
313