10752 条条目 · 106 个活跃源
2026年9月26日
04:00
arXiv cs.CV

Match4Annotate: Cross-Video Annotation Transfer in Ultrasound via Implicit Feature Flow-Guided Matching

04:00
arXiv cs.CV

Band-Attention Modulation Network for Robust Face Forgery Detection

04:00
arXiv cs.CV

Cross-Task Generalization in Handwriting-Based Alzheimer's Screening via Vision Language Adaptation

04:00
arXiv cs.CV

Interpretable Similarity of Synthetic Image Utility

04:00
arXiv cs.CV

OncoVision: Integrating Mammography and Clinical Data through Attention-Driven Multimodal AI for Enhanced Breast Cancer Diagnosis

04:00
arXiv cs.CV

Comparing YOLOv11 and YOLOv8 for instance segmentation of occluded and non-occluded immature green fruits in complex orchard environment

04:00
arXiv cs.CV

LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping

04:00
arXiv cs.CV

MDE-VIO: Enhancing Visual-Inertial Odometry Using Learned Depth Priors

04:00
arXiv cs.CV

WaterClear-GS: Optical-Aware Gaussian Splatting for Underwater Reconstruction and Restoration

04:00
arXiv cs.CV

Context-aware Skin Cancer Epithelial Cell Classification with Scalable Graph Transformers

04:00
arXiv cs.CV

One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation

04:00
arXiv cs.CV

GeoBlur: Epipolar Geometry Estimation from a Single Motion-Blurred Image

04:00
arXiv cs.CV

SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction

04:00
arXiv cs.CV

OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding

04:00
arXiv cs.CV

Open-access model for detecting openly dumped dispersed municipal solid waste from crowdsourced UAV imagery in Sub-Saharan Africa

04:00
arXiv cs.CV

Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion

04:00
arXiv cs.CV

BARRIER: Bounded Activation Regions for Robust Information Erasure

04:00
arXiv cs.CV

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

04:00
arXiv cs.CV

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

04:00
arXiv cs.CV

COMPASS: Fusion-Matched Supervision for Missing-Modality Human Sensing

04:00
arXiv cs.CV

Toward a Foundation Plug-and-Play Prior for Computed Tomography Reconstruction via a Multimodal Diffusion Model

04:00
arXiv cs.CV

SAMI3D-DW: Interactive Segmentation of Any 3D Medical Images

04:00
arXiv cs.CV

Virtual Encoders in Multimodal Transformers

04:00
arXiv cs.CV

AIR: Analytic Imbalance Rectifier for Continual Learning

04:00
arXiv cs.CV

Bridging the Inter-Domain Gap through Low-Level Features for Cross-Modal Medical Image Segmentation

04:00
arXiv cs.CV

PARTE: Plane-Assisted Robust Transformation Estimation for Point Cloud Registration

04:00
arXiv cs.CV

RotVLA: Rotational Latent Action for Vision-Language-Action Model

04:00
arXiv cs.CV

When Search Becomes Memory: Accelerating Robot Design Discovery with Self-Evolving Skills

04:00
arXiv cs.CV

GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective

04:00
arXiv cs.CV

Learning to Navigate with Minimal Parameters: Decomposing Visual Navigation Through Closed-Form Geometric Interfaces

04:00
arXiv cs.CV

ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion

2026年9月25日
04:00
arXiv cs.CV

M-plicits:基于嵌套多尺度残差的神经隐式曲面

提出嵌套多尺度残差神经隐式曲面,窄带监督抗噪,支持实时渲染,参数少且性能优。

04:00
arXiv cs.CV

CinematicVQA:面向大型视觉语言模型的电影语法推理基准测试

首个电影语法推理基准CinematicVQA,用电影场景图评测视觉语言模型:能描述画面却难识技法,思维链反致性能下降。

04:00
arXiv cs.CV

💓Heartian:生理感知的可重光照高斯头部化身

Heartian以心脏周期调制高斯头部化身肤色反照率,嵌入可恢复rPPG信号,心率MAE仅0.29 bpm,重建质量几乎无损。

04:00
arXiv cs.CV

UltraBench 2:迈向超声视觉基础模型的稳健评估

UltraBench 2 提供覆盖广泛解剖与任务的标准化超声评测,视觉基础模型专用预训练分类领先,通用模型分割追平。

04:00
arXiv cs.CV

面向高光谱图像分类的令牌聚类与语义序列Mamba

提出STMamba,通过令牌聚类构建语义序列,融合空间-光谱Mamba,提升高光谱图像分类并超越SOTA。

04:00
arXiv cs.CV

PePESeg3D:感知先验增强3D高斯泼溅的多尺度分割

将感知先验注入几何重建与对比学习,弥补多尺度掩码不完整,实现高精度3D高斯多尺度分割。

04:00
arXiv cs.CV

GeoNLI——卫星影像的自然语言解释器

提出融合SAM变体与多模态大模型的统一流水线,集成多模型投票,完成遥感描述、VQA与定位任务。

04:00
arXiv cs.CV

DrGait:面向可解释临床步态分析的生物力学基础视觉推理

DrGait免训练框架让VLM转为临床规划,经分诊-验证-综合调用几何工具,减少幻觉,产出可审计诊断报告。

04:00
arXiv cs.CV

小巧却助益:面向低视力者的空间感知后训练

提出 500M 紧凑视觉语言模型,经蒸馏与 GRPO 后训练增强视障导航的空间与危险感知,支持手机离线运行。

04:00
arXiv cs.CV

DeltaWAM:面向双臂操作的增量世界动作模型

提出DeltaWAM,联合预测视觉增量与动作,并用流式增量记忆提升双臂操作成功率、降低计算与延迟。

04:00
arXiv cs.CV

动作表示中的方向-尺度分解:重新思考视觉-语言-动作模型的词元化对象

DSD将动作增量分解为方向与尺度再词元化,提升离散词元VLA成功率,并缓解混合数据训练的性能下降。

04:00
arXiv cs.CV

MoVISA:面向视频目标分割的多令牌推理

MoVISA 用多分割令牌跨帧表示目标,实现语言与时空掩码细粒度对齐,MeViS 基准 J&F 提升 13.2%。

04:00
arXiv cs.CV

M²PFN:面向阿尔茨海默病可泛化多模态上下文学习的端到端解耦对齐

提出M²PFN,将TabPFN扩展为多模态阿尔茨海默病诊断模型,端到端解耦对齐,在ADNI及外部队列上泛化优于基线。

04:00
arXiv cs.CV

看起来相同,答案却不同:面向鲁棒视觉语言推理的翻转方向引导

提出免训练方法FlipDir,以低秩翻转子空间引导解码抑制VLM答案翻转,构建VisFlip基准。

04:00
arXiv cs.CV

面向语言引导医学图像分割的多模态路由与区域细化

提出MRSeg:以图像-文本对路由特征适配,区域桥细化,仅7.11M参数即达领先精度。

04:00
arXiv cs.CV

MEVL-STP:面向任意形状场景文本检测与识别的多编码器与视觉语言模型

提出多编码器分割与视觉语言模型识别两阶段法,生成多边形掩码,无需合成预训练,CTW1500端到端H均值85.86%。

04:00
arXiv cs.CV

PlenoCI:面向视角依赖感知变化分类的全光特征

从3DGS解析推导全光导数特征PlenoCI,忽略朗伯纹理、捕获视角依赖行为,用于变化分类,误报低两个数量级。

04:00
arXiv cs.CV

HelloWorld:迈向生成式驾驶世界模型的实际应用

提出2B驾驶世界模型HelloWorld,实现多传感器可控生成与高效交互式仿真。

04:00
arXiv cs.CV

利用多模态大语言模型中的目标知识实现鲁棒的少样本分割

提出MK-FSS,用MLLM挖掘查询图像的空间与语义知识,经双记忆融合与跨模态提示增强少样本分割鲁棒性。