이미지 기반 모델이 볼륨 렌더링을 이해할까?

Do Vision Foundation Models Understand Volume Renderings?

초록

Vision Foundation Models (VFMs) are pre-trained on large-scale natural image datasets and possess strong capabilities for capturing both morphological structures and complex contextual patterns in images. In contrast, Direct Volume Rendering (DVR) images contain more intricate information than natural images, as they project 3D volumetric data onto a 2D plane while incorporating transparency accumulation and the optical properties defined by a transfer function. This study investigates whether VFMs pre-trained solely on natural images can also extract meaningful features from DVR images. To this end, we analyze the image encoders of three representative VFMs—DINO, CLIP, and SAM—which differ in their training objectives and input modalities. By visualizing the features extracted from volume-rendered images using these models, we compare their capabilities in terms of structural expressiveness discrimination. Furthermore, by examining how each model’s training methodology influences its feature visualization outcomes, we provide an analysis of the characteristic differences that arise from their respective pre-training strategies.

키워드

Direct Volume RenderingVision Foundation ModelFeature Extraction직접 볼륨 렌더링시각 기반 모델특징 추출
제목
이미지 기반 모델이 볼륨 렌더링을 이해할까?
제목 (타언어)
Do Vision Foundation Models Understand Volume Renderings?
저자
강지윤안하일정윤현
DOI
10.29056/jdaem.2026.06.03
발행일
2026-06
유형
Y
저널명
디지털예술공학멀티미디어논문지
13
2
페이지
171 ~ 178