全部OpenAIHugging FaceArXiv AIGoogle AIVentureBeat AIMarkTechPostPragmatic EngineerThe GradientOne Useful ThingCrunchbase NewsTechCrunchGoogle DeepMindMicrosoft ResearchLlamaIndex BlogarXiv cs.AIarXiv cs.LGarXiv cs.CLarXiv cs.CVNature Machine IntelligenceLilian WengAndrej KarpathySebastian RuderBAIR BlogAgile Lab EngineeringGoogle AI BlogMeta NewsroomMIT Tech Review AITechCrunch AIWiredFull-Stack AI EngineerAgentplexa16zSequoia CapitalY CombinatorElad GilTomasz TunguzNot BoringThe GeneralistSimon WillisonSimon Willison NewsletterLatent SpaceImport AIInterconnectsHamel HusainDAIR.AIEnterprise AI GovernanceStratecherySemiAnalysisBenedict EvansX @OpenAIX @claudeaiX @AnthropicAIX @karpathyX @samaX @demishassabisX @DarioAmodeiX @elonmuskX @miramuratiX @ilyasutX @gdbX @AravSrinivasX @arthurmenschX @ClementDelangueX @alexandr_wangX @mustafasuleymanX @ylecunX @geoffreyhintonX @AndrewYNgX @fcholletX @polynoamialX @_jasonweiX @DrJimFanX @rasbtX @lilianwengX @OriolVinyalsMLX @tri_daoX @percyliangX @ch402X @janleikeX @swyxX @simonwX @jeremyphowardX @OfficialLoganKX @svpinoX @SuhailX @emollickX @_akhaliqX @natolambertX @rowancheungX @bilawalsidhuX @Francis_YAO_X @ShunyuYao14X @tqchenmlX @haozhangmlX @GoogleDeepMindX @huggingfaceX @MistralAIX @perplexity_aiX @NVIDIAAIX @xaiX @victor207755822X @deepseek_aiX @bchernyX @Alibaba_QwenX @Kimi_MoonshotX @dotey
10851 条条目 · 106 个活跃源
2026年9月11日
04:00
arXiv cs.CV
SCINTILLA-SNN: A Spiking Multi-Scale Selective Aggregation Network for Perineural Invasion Prediction
04:00
arXiv cs.CV
From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
04:00
arXiv cs.CV
Fast and Accurate Monomodal 3D High Resolution Deep Registration of Drosophila Larval Brain Volumes
04:00
arXiv cs.CV
SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-Views
04:00
arXiv cs.CV
Improving Faint Object Detection for Space Situational Awareness with Variational Autoencoders
04:00
arXiv cs.CV
Order-Aware 2.5D Multiple Instance Learning for Preoperative MRI-Based Perineural Invasion Risk Assessment in Intrahepatic Cholangiocarcinoma
04:00
arXiv cs.CV
GRIPNet: Gaussian Radial Intensity Prior Guided Architecture for Pulmonary Nodule Detection in CT
04:00
arXiv cs.CV
Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models
04:00
arXiv cs.CV
R4Tun: LLM-guided adaptive segmental tunnel lining segmentation in point clouds
04:00
arXiv cs.CV
Predictive Multi-Landmark OCT Tracking for Increased Motion Robustness
04:00
arXiv cs.CV
Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification
04:00
arXiv cs.CV
DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging
04:00
arXiv cs.CV
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation
04:00
arXiv cs.CV
Brain-PACE: A Deep Siamese MRI Framework for Modelling Longitudinal Brain Acceleration
04:00
arXiv cs.CV
Multi-Modal Controlled Coherent Motion Generation
04:00
arXiv cs.CV
BruNet: A Cross-Domain Transfer Framework for Bruise Segmentation
04:00
arXiv cs.CV
BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable Registration
04:00
arXiv cs.CV
Pre- and Post-Treatment Brain Metastases Segmentation Using nnU-Net with Post-Processing for BraTS 2026
04:00
arXiv cs.CV
Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs
04:00
arXiv cs.CV
Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation
04:00
arXiv cs.CV
UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from Ultrasound
04:00
arXiv cs.CV
LoopVAE: Recurrent Depth Across Scales for Visual Tokenization
04:00
arXiv cs.CV
Prototype Matters: Modality-unified Prototype Self-distillation for Unsupervised Visible-infrared Person Re-identification
04:00
arXiv cs.CV
Learn the Solid, Not the File: Canonical Inputs for Neural Networks on CAD Boundary Representations
04:00
arXiv cs.CV
Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates
04:00
arXiv cs.CV
World in World: Explore the World with World Models
04:00
arXiv cs.CV
A Comparative Evaluation of Pre-trained Convolutional Neural Networks for Melanoma Detection
04:00
arXiv cs.CV
OmniKVQuant: KV Cache Quantization for Omni-LLMs
04:00
arXiv cs.CV
LangStreet: Persistent Language Fields for Anchor-Decoded Street Gaussians
04:00
arXiv cs.CV
Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need
04:00
arXiv cs.CV
MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous Modalities
04:00
arXiv cs.CV
Multimodal Taxonomic Conditioning for Generative Plankton Imagery
04:00
arXiv cs.CV
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
04:00
arXiv cs.CV
Self-Supervised Cardiac Phase Detection via Single-Parameter Latent Orbits
04:00
arXiv cs.CV
Spectral Adapters for Segment Anything Model-based Segmentation of Colorectal Liver Metastases in Computed Tomography
04:00
arXiv cs.CV
Single-Stream Multi-Feature Fusion with Temporal Robustness for Gait Emotion Recognition
04:00
arXiv cs.CV
Language-Augmented Semantic Priors for B-Spline Surface Fitting
04:00
arXiv cs.CV
MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View Images
04:00
arXiv cs.CV
Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image Generators
04:00
arXiv cs.CV
Exponential Pixelating Integral transform with dual fractal features for enhanced chest X-ray abnormality detection
04:00
arXiv cs.CV
Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
04:00
arXiv cs.CV
Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling
04:00
arXiv cs.CV
3D Point Splatting for mmWave Radar Novel View Synthesis
04:00
arXiv cs.CV
Scale-Aware 3D Deep Learning for Robust Brain Metastasis Detection in Multimodal MRI
04:00
arXiv cs.CV
SenseNova-U1.5: Towards Native Unified Visual Intelligence
04:00
arXiv cs.CV
HuRo: Robotizing Human Videos for Scalable VLA Pretraining
04:00
arXiv cs.CV
Seamless Whole Slide Label-Free Virtual Staining
04:00
arXiv cs.CV
IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies
04:00
arXiv cs.CV
Breaking the Central Bias: Spatially Partitioned Experts for Coordinate-Based Neuroevolution
04:00
arXiv cs.CV