2026-07-22

9 papers from arXiv

← Back to history

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Zhengyu Zou, Hao Li, Kuixuan Jiao, Liu Liu, Tingyang Xiao et al.

Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic rec...

cs.CVPR

Anatomy-Aware 3D Mesh Refinement of Pericardium Segmentations on Computed Tomography

Andreas W. Aspe, Jonas Jalili Loft, Michael Huy Cuong Pham, Andreas Ohrt Johansen, Jørgen Tobias Kühl et al.

Accurate delineation of the pericardium in a cardiac CT scan is essential for quantifying epicardial adipose tissue, yet it remains one of the most challenging structures to segment due to its poor contrast boundaries. Instead of solely relying on image gradients, our framework leverages the anatomical context of surrounding anatomical structures to guide the segmentation. This work introduces a n...

cs.CVPR

Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding

Junlin Chang, Longhao Zou, Rui Li

Fine-tuning pre-trained point-cloud backbones typically updates all parameters, resulting in substantial computation and memory overhead. More importantly, modern point backbones rely on aggressive tokenization and downsampling, which yields compact global tokens but irreversibly discards fine-grained local geometry, an inherent bottleneck for parameter-efficient adaptation. Consequently, existing...

cs.CVECCV

FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility

Yihalem Yimolal Tiruneh, Muhammad Salman Ali, Uyoung Jeong, Muneeb A. Khan, MD Khalequzzaman Chowdhury Sayem et al.

Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital content creation. Existing approaches typically couple parametric body models with neural rendering or 3D Gaussian splatting and optimize all body regions jointly from short videos, which often degrades fidelity in the visible areas. To overcome t...

cs.CVECCV

CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement

Zhihao Yang, Zhiyu Xiang, Peng Xu, Tianyu Pu, Kai Wang et al.

V2X collaborative object detection features overcoming the limitations of single-vehicle systems by aggregating environmental features from multiple collaborative agents. However, existing mainstream V2X perception methods mainly focus on 2D BEV object detection. When 3D detection task is concerned, inferior results are obtained because they ignore the 3D spatial misalignment caused by differing h...

cs.CVcs.AIECCV

Mi-Memory: A Lifecycle Memory Framework for Personal AI

Xule Liu, Hanlin Teng, Chao Li, Yanan Ni, Shuo Lu et al.

Personal AI is moving beyond chat-only interaction toward continuous services that span phones, cars, homes, wearables, cameras, and tools. In this setting, memory cannot remain a cache of prior conversations. It should serve as a continuity and governance substrate: preserving durable user state, grounding answers in multimodal and device evidence, supporting correction and forgetting, bounding p...

cs.AIPR

OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining

Rian Touchent, Éric de la Clergerie

We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models. Our approach has three stages: random walks through ontology graphs capture hierarchical and causal relations between medical codes, a large language model reformulates these walks into fluent textbook-style prose, and the resulting text is used to train ModernCamemBERT, a 14...

cs.AIPR

Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning

Stella Ho, Joel Villalobos, Joseph West, Jingyang Liu, Weijie Qi et al.

ECoG-based visual semantic decoding enables inference of semantic interpretation of visual perception from complex, noisy brain activity. This study examines the feasibility of visual semantic decoding using an end-to-end deep learning framework using electrocorticography (ECoG). Specifically, the decoding task is to predict visual categories from video stimuli using time-series neural inputs. A p...

cs.LGq-bio.NCPR

Reliability-Aware 3D Geometric Injection for Universal Person Re-identification

Bohan Su, Jiashuo Wang, Fangyi Liu, Mang Ye

Universal person re-identification (ReID) aims to retrieve pedestrian identities across diverse real-world scenarios, including severe occlusions, clothing changes, and cross-modality shifts, within a unified model. However, existing 2D representations fundamentally struggle with spatial ambiguities due to a lack of depth and topological awareness, while naively introducing monocular 3D priors oft...

cs.CVECCV