5294 条条目 · 106 个活跃源
2026年9月12日
04:00
arXiv cs.AI

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

04:00
arXiv cs.AI

Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language

04:00
arXiv cs.AI

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

04:00
arXiv cs.AI

Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

04:00
arXiv cs.AI

Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents

04:00
arXiv cs.AI

Breaking Predictions Is Not Enough: Specified-Foil Counterfactuals for Temporal Graphs

04:00
arXiv cs.AI

MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG

04:00
arXiv cs.AI

DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat

04:00
arXiv cs.AI

Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation

04:00
arXiv cs.AI

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

04:00
arXiv cs.AI

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

04:00
arXiv cs.AI

When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting

04:00
arXiv cs.AI

Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce

04:00
arXiv cs.AI

Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model

04:00
arXiv cs.AI

Memory Compression for High-Fanout Agent Sandboxes

04:00
arXiv cs.AI

Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

04:00
arXiv cs.AI

An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning

04:00
arXiv cs.AI

RouteRepair: Instance-Level Failure Diagnosis and Targeted Repair in LLM-Based Automated Heuristic Design for Routing Optimization

04:00
arXiv cs.AI

Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning

04:00
arXiv cs.AI

Flexible and Interpretable Accent Distance Measurements

04:00
arXiv cs.AI

Characterizing Job Power Elasticity for Power-Flexible AI Training

04:00
arXiv cs.AI

Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting

04:00
arXiv cs.AI

MAPLE: Memory-Augmented Planning with Language and Evolution

04:00
arXiv cs.AI

Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents

04:00
arXiv cs.AI

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

04:00
arXiv cs.AI

Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models

04:00
arXiv cs.AI

What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

04:00
arXiv cs.AI

No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vulnerabilities in MCP Servers

04:00
arXiv cs.AI

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

04:00
arXiv cs.AI

Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training

04:00
arXiv cs.AI

Role differentiation as ignition of a collective information engine: Structuration in Agent Populations

04:00
arXiv cs.AI

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

04:00
arXiv cs.AI

DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation Attacks

04:00
arXiv cs.AI

A Mathematical Theory of Pragmatic Information

04:00
arXiv cs.AI

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

04:00
arXiv cs.AI

Less can be More: What Aspects of Speech Drive End-of-Turn Detection

04:00
arXiv cs.AI

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

04:00
arXiv cs.AI

terms.txt: A Consent and Compensation Protocol for Agentic Web Access

04:00
arXiv cs.AI

Exploring Second-Order Pattern Recognition in Speaker Recognition

04:00
arXiv cs.AI

X-RACE: XAI-assisted Recurrent neural network Attribution for Channel Estimation

04:00
arXiv cs.AI

AI Soccer Analyst: Stage-Aware and Verifiable Human-AI Collaboration for Soccer Data Analysis

04:00
arXiv cs.AI

Buyer Artificial Intelligence-Enabled Environmental Governance and Supplier Environmental Controversies: An Organizational Information Processing and Signaling

04:00
arXiv cs.AI

Agent-Integrated Software: Interaction Contracts and Continuous Assurance

04:00
arXiv cs.AI

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

04:00
arXiv cs.AI

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

04:00
arXiv cs.AI

Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks

04:00
arXiv cs.AI

Warrant Theory

04:00
arXiv cs.AI

Physics-Informed Neural Networks to Infer the Perpendicular Energy Conductivity in the Scrape-Off Layer of Stellarator Devices

04:00
arXiv cs.AI

ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

04:00
arXiv cs.AI

Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms