AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
29 May 2026

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

Model ReleasesDGX agent

arXiv:2605.29360v1 Announce Type: new Abstract: Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that t

Mitigating Hallucination in Vision-Language Models through Barrier-Regulated Adaptive Closed-form Steering

Model ReleasesDGX agent

arXiv:2605.29881v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) often hallucinate objects that are not present in the input image, largely because visual grounding weakens as de

Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2605.29862v1 Announce Type: cross Abstract: AI-driven respiratory sound classification (RSC) is promising for automated pulmonary disease detection, yet multi-site deployment is hindered by inte

Model Fusion via Retrofitting

SafetyDGX agent

arXiv:2507.00037v2 Announce Type: replace-cross Abstract: Model fusion seeks to combine independently trained neural networks into a single model without retraining, but is complicated by representati

Modeling Hierarchical Thinking in Large Reasoning Models

SafetyDGX agent

arXiv:2510.22437v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) solve complex tasks by generating long Chain-of-Thought (CoT) sequences; however, the emergent dynamics governing reas

Modularizing Educational LLM-Agency for Fostering Responsible Learning Assistance

SafetyDGX agent

arXiv:2605.30187v1 Announce Type: new Abstract: The widespread adoption of AI chatbots in education will drastically change learning, making responsible deployment a critical concern. While large lang

Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation

ResearchDGX agent

arXiv:2605.29873v1 Announce Type: new Abstract: Key-Value (KV) cache remains a major bottleneck for deploying Large Language Models (LLMs) in long-generation tasks. Prior work often applies uniform co

MOO: A Multi-view Oriented Observations Dataset for Viewpoint Analysis in Cattle Re-Identification

ApplicationsDGX agent

arXiv:2603.04314v2 Announce Type: replace-cross Abstract: Animal re-identification (ReID) faces critical challenges due to viewpoint variations, particularly in Aerial-Ground (AG-ReID) settings where

MOOSE-Copilot: A Web-Based Interactive Assistant for Unified Exploratory and Fine-Grained Scientific Hypothesis Discovery

AgentsDGX agent

arXiv:2605.29475v1 Announce Type: cross Abstract: Large language models (LLMs) show remarkable potential in scientific hypothesis discovery. However, existing approaches face two critical limitations:

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing

Model ReleasesDGX agent

arXiv:2605.22100v2 Announce Type: replace Abstract: Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information sys

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

Model ReleasesDGX agent

arXiv:2605.29738v1 Announce Type: cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual

Multi-Level Barriers to Generative AI Adoption Across Disciplines and Professional Roles in Higher Education

ApplicationsDGX agent

arXiv:2603.27052v2 Announce Type: replace-cross Abstract: Generative Artificial Intelligence (GenAI) is rapidly reshaping higher education, yet barriers to its adoption across different disciplines an

Multi-Resolution End-to-End Deep Neural Network for Optimizing Latency-Accuracy Tradeoff in Autonomous Driving

SafetyDGX agent

arXiv:2605.29138v1 Announce Type: cross Abstract: Latency-accuracy tradeoffs are fundamental in real-time applications of deep neural networks (DNNs) for cyber-physical systems. In autonomous driving,

MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization

Model ReleasesDGX agent

arXiv:2605.29951v1 Announce Type: new Abstract: Understanding how harm emerges from interaction between otherwise benign image-text pairs requires intent-aware cross-modal reasoning beyond surface-lev

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

Model ReleasesDGX agent

arXiv:2605.29300v1 Announce Type: cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses ar

NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs

Model ReleasesDGX agent

arXiv:2605.29716v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising non-autoregressive generative paradigm. Given the prohibitive computational cost of

Network Optimization Aspects of Autonomous Vehicles: Challenges and Future Directions

AgentsDGX agent

arXiv:2605.29518v1 Announce Type: cross Abstract: Global megatrends, such as urbanization, population growth, and emerging network solutions are accelerating the development of the Connected and Auton

Neural Network Verification using Partial Multi-Neuron Relaxation

SafetyDGX agent

arXiv:2605.30155v1 Announce Type: cross Abstract: The increasing integration of deep neural networks in critical systems has spawned a theoretical and practical interest in formally guaranteeing safet

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

Model ReleasesDGX agent

arXiv:2605.29685v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social

No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval

Model ReleasesDGX agent

arXiv:2605.30120v1 Announce Type: cross Abstract: Multi-vector retrieval (MVR) models, exemplified by ColBERT, have established new benchmarks in retrieval accuracy by preserving fine-grained token-le

No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand

AgentsDGX agent

arXiv:2605.28836v1 Announce Type: cross Abstract: The Plain Writing Act in the United States requires government documents to be accessible in clear and simple language that the general public can eas

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

Model ReleasesDGX agent

arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default languag

Obfuscation Rules for Detecting and Detoxifying Korean Toxicity

ApplicationsDGX agent

arXiv:2510.10961v3 Announce Type: replace-cross Abstract: As language models become increasingly deployed in online environments, toxicity detection and detoxification have received growing attention.

OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning

ResearchDGX agent

arXiv:2605.29657v1 Announce Type: cross Abstract: Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and

Offline Reinforcement Learning with Generative Trajectory Policies

SafetyDGX agent

arXiv:2510.11499v2 Announce Type: replace-cross Abstract: Generative models have emerged as a powerful class of policies for offline reinforcement learning (RL) due to their ability to capture complex

OISD: On-Policy Internal Self-Distillation of Language Models

Model ReleasesDGX agent

arXiv:2605.29089v1 Announce Type: cross Abstract: Recent reinforcement learning (RL) post-training approaches primarily optimize the final output policy using sparse outcome-level rewards, while large

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

ResearchDGX agent

arXiv:2602.12304v4 Announce Type: replace-cross Abstract: Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual pro

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

Model ReleasesDGX agent

arXiv:2605.29833v1 Announce Type: new Abstract: As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdi

OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources

Model ReleasesDGX agent

arXiv:2605.29250v1 Announce Type: cross Abstract: Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graph

On Distributional Reinforcement Learning in Chaotic Dynamical Systems

AgentsDGX agent

arXiv:2605.30160v1 Announce Type: cross Abstract: Chaotic dynamical systems pose a fundamental challenge for Reinforcement Learning (RL): exponential sensitivity to initial conditions induces high-var

On Language Generation in the Limit with Bounded Memory

ResearchDGX agent

arXiv:2605.30324v1 Announce Type: cross Abstract: We study language generation in the limit under bounded memory. In this task, a learner observes examples from an unknown target language one at a tim

On the Geometry of Games and their Solvers

SafetyDGX agent

arXiv:2605.29919v1 Announce Type: new Abstract: A central challenge in game theory and learning systems such as GANs is understanding which algorithms can efficiently compute equilibria across the het

On the Optimizer Dependence of Neural Scaling Laws

ResearchDGX agent

arXiv:2605.29387v1 Announce Type: cross Abstract: The scaling exponent alpha in neural scaling laws L(N) propto N^{-alpha} is commonly treated as a fixed constant set by architecture and data. We pres

Online Fair Division with Additional Information

SafetyDGX agent

arXiv:2505.24503v3 Announce Type: replace-cross Abstract: We study the problem of fairly allocating indivisible goods to agents in an online setting, where goods arrive sequentially and must be alloca

OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories

Model ReleasesDGX agent

arXiv:2605.29253v1 Announce Type: new Abstract: Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambi

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

Model ReleasesDGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verification

ResearchDGX agent

arXiv:2605.29556v1 Announce Type: new Abstract: Building mathematical optimization models is critical in operations research (OR), while it requires substantial human expertise. Recent advancements ha

OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation

Model ReleasesDGX agent

arXiv:2605.29829v1 Announce Type: new Abstract: Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient par

Orthogonal Concept Erasure for Diffusion Models

Model ReleasesDGX agent

arXiv:2605.28902v1 Announce Type: new Abstract: Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face signifi

Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies

Model ReleasesDGX agent

arXiv:2605.30148v1 Announce Type: cross Abstract: Evolution Strategies (ES) has recently emerged as a competitive alternative to reinforcement learning (RL) for large language model (LLM) fine-tuning,

P^2RAG: Efficient Privacy-Preserving RAG Service Supporting Arbitrary Top-k Retrieval

ApplicationsDGX agent

arXiv:2603.14778v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) enables large language models to use external knowledge, but outsourcing the RAG service raises privacy c

Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents

SafetyDGX agent

arXiv:2605.29174v1 Announce Type: new Abstract: DeFi investment agents, systems that use AI for autonomous on-chain trading, have attained over USD 3 billion in combined token valuations since late 20

Parallax: Parameterized Local Linear Attention for Language Modeling

Model ReleasesDGX agent

arXiv:2605.29157v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remain

ParaTool: Shifting Tool Representations from Context to Parameters

Model ReleasesDGX agent

arXiv:2605.29561v1 Announce Type: new Abstract: Tool calling extends large language models (LLMs) by enabling grounded interaction with external executable interfaces, thereby supporting environment-c

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding

ApplicationsDGX agent

arXiv:2605.30126v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) map visual inputs into dense token sequences, imposing a quadratic computational bottleneck for inference. Elasti

PassNet: Scaling Large Language Models for Graph Compiler Pass Generation

ApplicationsDGX agent

arXiv:2605.29357v1 Announce Type: new Abstract: Modern tensor compilers such as TorchInductor deliver substantial speedups on mainstream models, yet face a systematic performance ceiling on long-tail

Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit

ApplicationsDGX agent

arXiv:2605.30207v1 Announce Type: new Abstract: The same prompt -- 'best CRM software' -- reaches AI assistants from buyers in widely different contexts: a solo founder, an enterprise VP, a UK SMB own

PersonaAgent: Bridging Memory and Action for Personalized LLM Agents

SafetyDGX agent

arXiv:2506.06254v2 Announce Type: replace Abstract: Large Language Model (LLM) empowered agents have recently emerged as advanced paradigms that exhibit impressive capabilities in a wide range of doma

Personalized Turn-Level User Conversation Satisfaction Benchmark

Model ReleasesDGX agent

arXiv:2605.29711v1 Announce Type: cross Abstract: User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user

PhoneWorld: Scaling Phone-Use Agent Environments

Model ReleasesDGX agent

arXiv:2605.29486v1 Announce Type: cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Ex

PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions

AgentsDGX agent

arXiv:2605.30268v1 Announce Type: cross Abstract: We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target obje

Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software

Model ReleasesDGX agent

arXiv:2605.30353v1 Announce Type: new Abstract: Are AI agents tools, co-authors, or researchers? We present a quantified case study (N=1): a physicist supervising an AI coding agent (Claude Code, Sonn

Planning with the Views via Scene Self-Exploration

Model ReleasesDGX agent

arXiv:2605.29563v1 Announce Type: new Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1)understandin

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.29299v1 Announce Type: cross Abstract: Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cos

PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers

Model ReleasesDGX agent

arXiv:2605.30094v1 Announce Type: new Abstract: Poker is a landmark challenge for artificial intelligence. The dominant approach relies on equilibrium solvers built on counterfactual regret minimizati

Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning

ApplicationsDGX agent

arXiv:2506.08354v2 Announce Type: replace-cross Abstract: This position paper argues that text embedding research should move beyond surface meaning and embrace implicit semantics as a central modelin

Practitioner Beliefs and Behaviors in AI-Enhanced Education: DOT Framework Survey Evidence

SafetyDGX agent

arXiv:2605.29041v1 Announce Type: new Abstract: This study reports findings from a cross-sectional survey (n = 72) of higher education practitioners examining beliefs, behaviors, and institutional con

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

Model ReleasesDGX agent

arXiv:2605.29815v1 Announce Type: new Abstract: The growing number of submitted papers has motivated the exploration of Large Language Models (LLMs) as a means to support and augment the peer review p

Predicting Causal Effects from Natural Language Queries using Structured Representations

Model ReleasesDGX agent

arXiv:2605.29631v1 Announce Type: cross Abstract: Randomized controlled trials are a cornerstone of medicine and the social sciences as they enable reliable estimates of causal effects. However, they

PrismFlow: Residual Dynamics for Flow Matching in Time-Series Generation

TutorialsDGX agent

arXiv:2605.28867v1 Announce Type: cross Abstract: Generating high-quality time-series data is challenging because real-world signals often exhibit multimodal patterns and multiscale dynamics, includin

← Previous
1…192193194195196…358
Next →