2026-08-13

10 papers from arXiv

← Back to history

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

Yuyang Yin, Zixiang Li, Longxuan Deng, Hongkai Li, Shifang Zhao et al.

Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-temporal dynamics. Yet existing generative methods rely on simple prompts to jointly control all of these factors through one-shot image or video synthesis, offering weak controllability and limited support ...

cs.CVPR

GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors

Jiazheng Liu, Hang Li, Jiawei Zhang, Jiahe Li, Xiaohan Yu et al.

Generative models like Diffusion Models and Flow Matching have demonstrated remarkable capabilities in synthesizing high-fidelity driving videos, but are severely constrained by high inference latency due to the requirement of extensive sampling steps. We argue that this inefficiency stems from the prevailing reliance on a standard Gaussian source distribution, where consecutive frames are initial...

cs.CVECCV

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

Yung-Hsu Yang, Luigi Piccinelli, Samuel Rota Bulò, Sunghwan Hong, Denis Rozumny et al.

Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This motivates monocular 3D detection, which avoids additional constraints, yet it faces a major obstacle: from a single image, depth, and especially absolute scale, are underconstrained. As a result, the prevailing pattern of d...

cs.CVECCV

ADEPT: A Unified Framework for Deep Learning Test Adequacy

Yidi Kao, Shawn Burnham, Tommi Rose Fahy, Ali Ghanbari

Over the past decade, many test adequacy metrics have been proposed for deep learning that characterize test dataset adequacy from different perspectives, e.g., neuron activation behavior, latent feature coverage, decision-boundary exploration, etc. However, these metrics are typically released as independent research prototypes with substantially different installation and preprocessing requireme...

cs.SEcs.LGPR

Attractor Image-Based Deep Learning of Arterial Pulse Waves for Age Classification

Sara Vardanega, Patrick Segers, Philip Aston, Ernst Rietzschel, Jordi Alastruey et al.

Arterial pulse waveform morphology evolves with age, reflecting structural and functional changes in the cardiovascular system. Thus, vascular age is a valuable surrogate marker of cardiovascular health, and premature vascular ageing can indicate increased disease risk. Pulse wave analysis could support risk stratification in otherwise asymptomatic adults. We transformed pulse wave time-series dat...

cs.LGPR

Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP

Nikolette Pedersen, Regitze Sydendal, Veronika Cheplygina, Théo Sourget

Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence. However, recent work reveals that CLIP-based models remain vulnerable to shortcuts. We investigate how real-world shortcuts manifest across different layers of the medical CLIP-based model, MedCLIP, and its vision encoder,...

cs.CVcs.LGPR

TESLA: Taylor Expansion of Sinusoidal Learnable Activations

Daehwa Ko, Jaehyeon Kim, Seunghyun Ham, Jay Hoon Jung

The parity problem--deciding whether the number of ones in a binary vector is odd or even--remains challenging for standard neural networks due to linear inseparability and the need for global interactions. We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order component...

cs.LGPR

Dual Anchors, Do It Better: Hierarchical Group Merging for Zero-Shot Anomaly Detection

Jimin Roh, DongKyu Kim, Suk-Ju Kang

Zero-shot anomaly detection (ZSAD) aims to identify anomalies in unseen domains, a setting that is particularly critical for industrial and medical applications where domain shifts are prevalent. However, most CLIP-based ZSAD methods anchor semantics solely on the text modality, making performance highly sensitive to prompt design and leading to weak visual grounding. To mitigate these limitations...

cs.CVCVPR

Warping Earth Observations for better ice labeling in the Marginal Marginal Ice Zone

Tom Kelly, Martin S. J. Rogers

Multimodal satellite imagery provides complementary information for Earth Observation, but accurately combining heterogeneous sensors remains challenging in dynamic environments. Fast-changing regions, such as the Antarctic marginal ice zone, cannot fully exploit multimodal information from different satellite sensors because surface features move between image acquisitions. This spatial and tempo...

cs.CVECCV

BoltNet: An Ultra-Lightweight Convolutional Network for On-Device Plant Species Identification

Daniel Rossi, Guido Borghi, Roberto Vezzani

Automated plant species identification from citizen-science imagery is an established, demanding fine-grained recognition problem: large taxonomic label spaces, visually similar species, and long-tailed observations require real model capacity, while field use constrains memory, latency, and power. Model size is only part of the deployment cost: intermediate activations held in memory during infer...

cs.CVECCV