10851 条条目 · 106 个活跃源
2026年9月11日
04:00
arXiv cs.CV

SegCol Challenge: Semantic Segmentation for Tools and Fold Edges in Colonoscopy data

04:00
arXiv cs.CV

Divergence-Based Similarity Function for Multi-View Contrastive Learning

04:00
arXiv cs.CV

SSS: Semi-Supervised SAM-2 with Efficient Prompting for Medical Imaging Segmentation

04:00
arXiv cs.CV

MedGEN-Bench: A Contextually Entangled Benchmark for Open-ended Multimodal Medical Generation

04:00
arXiv cs.CV

Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation

04:00
arXiv cs.CV

Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge

04:00
arXiv cs.CV

Adaptive Dual-Constrained Line Aggregation for Cross-Paradigm Line Segment Detection

04:00
arXiv cs.CV

Hologram Representation via Quadratic Phase Gaussian Splatting

04:00
arXiv cs.CV

DirectSwap: Paired, Mask-Free Video Head Swapping with Full-Reference Evaluation

04:00
arXiv cs.CV

Gaussian Belief Propagation Network for Depth Completion

04:00
arXiv cs.CV

Discriminative Span as a Predictor of Synthetic Data Utility via Classifier Reconstruction

04:00
arXiv cs.CV

InstantHDR: Single-forward Gaussian Splatting Initialization for HDR 3D Reconstruction

04:00
arXiv cs.CV

V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval

04:00
arXiv cs.CV

CLIP-RD: Relational Distillation for Efficient CLIP Knowledge Distillation

04:00
arXiv cs.CV

Leveraging Avatar Fingerprinting: A Multi-Generator Photorealistic Talking-Head Public Database and Benchmark

04:00
arXiv cs.CV

Automated multi-class wound assessment using dedicated instance segmentation models for boundary detection and classification

04:00
arXiv cs.CV

Towards Automated Solar Panel Integrity: Hybrid Deep Feature Extraction for Advanced Surface Defect Identification

04:00
arXiv cs.CV

Reconstruction of a 3D wireframe from a single line drawing via generative depth estimation

04:00
arXiv cs.CV

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

04:00
arXiv cs.CV

Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters

04:00
arXiv cs.CV

SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models

04:00
arXiv cs.CV

Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model

04:00
arXiv cs.CV

Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning

04:00
arXiv cs.CV

SegKAN: High-Resolution Medical Image Segmentation with Long-Distance Dependencies

04:00
arXiv cs.CV

A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction Models

04:00
arXiv cs.CV

RAIDAL: Redundancy-Aware Information Density Active Learning for CTC-Based Continuous Sign Language Recognition

04:00
arXiv cs.CV

CGSM: Concept-Guided Segmentation Model for Precise Pulmonary Lesion Delineation

04:00
arXiv cs.CV

Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision

04:00
arXiv cs.CV

Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage

04:00
arXiv cs.CV

PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images

04:00
arXiv cs.CV

FastMap: Real-Time Semantic Map Completion via Bitwise Masked Modeling

04:00
arXiv cs.CV

DefVINS: Visual-Inertial Odometry for Deformable Scenes

04:00
arXiv cs.CV

CertDW: Towards Certified Dataset Ownership Verification via Conformal Calibration

04:00
arXiv cs.CV

Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation

04:00
arXiv cs.CV

Motus2: A Self-Evolving General World Model for Dexterous Manipulation

04:00
arXiv cs.CV

Domain Elastic Transform: Bayesian Function Registration for High-Dimensional Scientific Data

04:00
arXiv cs.CV

Representation learning of human cortical folding to reveal long lasting neurodevelopmental signatures

04:00
arXiv cs.CV

Diagnosing and Dynamically Filtering Occupancy World Models for Active Mapping

2026年9月10日
04:00
arXiv cs.CV

OmniPoint:任意相机的通用单目度量点云

提出统一框架OmniPoint,以解耦光线-距离表示和双向增强,实现多相机单目度量三维重建,零样本性能领先。

04:00
arXiv cs.CV

问题关键证据渐进丢失下选择性视觉推理的证据顺序校准

研究VLM关键证据渐进缺失下置信可靠性,提出证据顺序校准,降低违规,但未证明通用置信优势。

04:00
arXiv cs.CV

在潜在空间中融合单模态与视觉-语言表示用于多标签胸部X射线分类

融合单模态与视觉语言表示,经潜在空间精炼后混合融合,提升胸部X光多标签分类,AUROC达0.840。

04:00
arXiv cs.CV

信息分布在协同头中漂移时,多模态大语言模型会产生幻觉

提出HEAL:解耦头部信息分布,向协同头注入动态校准因子,有效抑制多模态大模型幻觉。

04:00
arXiv cs.CV

M2LG-DG:面向跨站点重性抑郁障碍分类的多模态局部-全局域泛化框架

提出M2LG-DG源域多模态局部-全局框架,融合双流rs-fMRI编码与跨站点对比学习,提升跨站点MDD分类性能。

04:00
arXiv cs.CV

DensePol:面向基于学习的偏振视觉的密集角度偏振数据集

提出密集角度RGB-偏振数据集DensePol,含2018组图像、180个分析器角度,显著提升偏振预测与表面法线估计。

04:00
arXiv cs.CV

无镜头视线并非默认隐私安全:跨披露面的身份泄露审计

无镜头近眼传感虽不可解读,攻击者仍能恢复身份:模拟管道识别率96.7%,各披露面均泄露,隐私须实测而非凭外观。

04:00
arXiv cs.CV

AgenticGen:面向广告的奖励引导智能体视频生成

提出奖励引导智能体框架 AgenticGen,将广告视频生成拆为策略选择与草稿生成,借在线反馈优化,TikTok 上显著提升点击与转化。

04:00
arXiv cs.CV

活态图书馆:将档案馆藏转化为对话式知识系统——来自西奥多·罗斯福总统图书馆的经验

罗斯福图书馆将30万条档案数字化,经AI处理与专家校审,构建支撑数字人"与TR对话"展项的对话式知识系统。

04:00
arXiv cs.CV

Video-MOPD:面向视频理解的多教师同策略蒸馏

该模型经三领域强化学习与多教师同策略蒸馏,融合视频时序定位、理解与STEM推理,在多项视频基准达同规模最优。

04:00
arXiv cs.CV

VANTAGE-Bench:评估视觉语言模型中的基础设施AI差距

VANTAGE-Bench基准揭示视觉语言模型在基础设施AI上的缺口,时间定位与稠密描述最弱。

04:00
arXiv cs.CV

MotionBlind:探究视频大语言模型中运动理解的错觉

Video-LLM实为运动盲:新基准显示多数模型仅达随机水平,仅Gemini3.1 Pro整体通过。