Text-Conditioned Diffusion Model for 3D Point Cloud Generation

초록

Current point-cloud diffusion models can generate overall shapes well, but they struggle to capture detailed semantic features and specific shape attributes described in captions. To improve this, we present a text-guided 3D point cloud generation method that uses a CLIP-conditioned diffusion model. Our approach adds natural-language information to the denoising network through token-level cross-attention and sentence-level semantic feature modulation. To prevent the model from focusing too much on less important words like articles and prepositions, we use a simple stop-word-aware attention reweighting strategy. This keeps the full token sequence but lowers the influence of stop words. We train and test our model on 3D chair and table data paired with matching natural-language descriptions. Our experiments look at overall generation quality, how well the shapes match the text, token-level attention patterns, and results from different attention setups. Quantitatively, the proposed model achieved MMD-CD scores of 6.80/6.33, COV-CD scores of 49.75/43.75, 1-NN-CD scores of 77.99/71.75, and JSD values of 9.6/13.2 on the Chair/Table categories, respectively. In the chair-category ablation study, stop-word-aware attention reweighting improved COV-CD from 37.18% to 49.75% and reduced JSD from 12.2 to 9.6. The results also show that stop-word-aware reweighting helps the model focus on key words and improves semantic coverage. Still, our model has some limits in grounding fine-grained attributes and generalizing to more object categories.

키워드

Deep Learning3D Point Cloud GenerationDiffusion ModelCLIPText-Conditioned Generation딥러닝3D 포인트 클라우드 생성확산 모델CLIP텍스트 조건부 생성
제목
Text-Conditioned Diffusion Model for 3D Point Cloud Generation
저자
장현준조정찬
DOI
10.23019/kingpc.22.3.202606.007
발행일
2026-06
유형
Y
저널명
한국차세대컴퓨팅학회 논문지
22
3
페이지
101 ~ 116