상세 보기
Voice-to-Mesh: A Mixed Reality Pipeline for 3D Content Generation Using Text-to-3D Model
- 김예진;
- 강민준;
- 김수현;
- 정윤현
초록
Recent advances in text-to-3D generation have enabled the creation of 3D assets from natural language prompts. However, Mixed Reality (MR) authoring workflows still rely on manually prepared 3D assets, making it difficult to generate and utilize new objects. In addition, generated mesh assets often vary in scale, object origin, orientation, and coordinate conventions, requiring additional normalization and alignment for direct placement in MR environments. This additional processing interrupts the continuity of the content creation workflow. To address these limitations, we present an end-to-end pipeline that generates 3D assets from multilingual voice commands and automatically normalizes the generated assets for use in MR environments. Speech-to-Text recognition and Large Language Model-based parsing convert natural language into structured commands. The proposed pipeline enables generated mesh assets to be immediately utilized, stored, and reused within the MR environments. Through a storytelling scenario, we demonstrate the feasibility of interactive MR content creation.
키워드
- 제목
- Voice-to-Mesh: A Mixed Reality Pipeline for 3D Content Generation Using Text-to-3D Model
- 저자
- 김예진; 강민준; 김수현; 정윤현
- 발행일
- 2026-06
- 유형
- Y
- 저널명
- Journal of Digital Media & Culture Technology
- 권
- 6
- 호
- 1
- 페이지
- 69 ~ 76