AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
28 May 2026

Anomaly as Non-Conformity via Training-Free Graph Laplacian Energy Minimization

ResearchDGX agent

arXiv:2605.28428v1 Announce Type: cross Abstract: Detecting subtle visual anomalies in images remains challenging, particularly when only normal samples are available a priori. Such unsupervised anoma

Apple Intelligence Foundation Language Models

Model ReleasesDGX agent

arXiv:2407.21075v2 Announce Type: replace Abstract: We present foundation language models developed to power Apple Intelligence features, including a ~3 billion parameter model designed to run efficie

Architecture-driven Shift: towards a lightweight selector for capturing the trends of logit shift

ApplicationsDGX agent

arXiv:2605.27469v1 Announce Type: cross Abstract: Continual Learning (CL) is a practical paradigm to utilize power of deep pre-trained neural networks, but which pre-trained model has a better ability


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration

ResearchDGX agent

arXiv:2605.27752v1 Announce Type: new Abstract: LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence. These signals are sometimes

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications

Model ReleasesDGX agent

arXiv:2605.27472v1 Announce Type: cross Abstract: Assertion-based verification (ABV) is a cornerstone of modern hardware design, yet manually translating design intent into formal SystemVerilog Assert

ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference

Model ReleasesDGX agent

arXiv:2505.19342v2 Announce Type: replace-cross Abstract: Multi-device inference can reduce Transformer latency by parallelizing computation. However, existing methods require high inter-device bandwi

AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

Model ReleasesDGX agent

arXiv:2605.27995v1 Announce Type: new Abstract: Large language model (LLM)-based agents have shown strong capabilities in using external tools to solve complex tasks. However, existing evaluations oft

Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies

ResearchDGX agent

arXiv:2512.01970v3 Announce Type: replace Abstract: Does Reinforcement Learning (RL) merely amplify existing skills, or synthesize novel skills? We investigate this question through the lens of Comple

Auditable Decision Models with Learned Abstention and Real-Time Steering

SafetyDGX agent

arXiv:2605.27768v1 Announce Type: new Abstract: Production AI systems often operate with incomplete, conflicting, or insufficient evidence. Forced classifiers collapse such cases into action labels, w

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation

AgentsDGX agent

arXiv:2605.28655v1 Announce Type: new Abstract: Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision. AI agents can automate parts

Backdoor Attacks on Fault Detection and Localization in Cyber-Physical Systems

Local AiDGX agent

arXiv:2605.27674v1 Announce Type: cross Abstract: Cyber-Physical Systems (CPS) integrate sensing, communication, computation, and control to support critical infrastructure, including smart grids, ind

Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspective

ResearchDGX agent

arXiv:2605.27476v1 Announce Type: cross Abstract: We characterize the pre-softmax attention matrix mathbf{QK^op} in transformers as an associative memory matrix encoding pairwise associations between

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

Model ReleasesDGX agent

arXiv:2605.28642v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment par

Bayesian Gated Non-Negative Contrastive Learning

SafetyDGX agent

arXiv:2605.28441v1 Announce Type: cross Abstract: While Contrastive Learning (CL) has revolutionized self-supervised representation learning, its latent representations remain highly entangled and opa

Behavioural Analysis of Alignment Faking

SafetyDGX agent

arXiv:2605.27681v1 Announce Type: new Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deploym

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

Model ReleasesDGX agent

arXiv:2605.28508v1 Announce Type: new Abstract: Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape us

Benchmarking Fairness in Spiking Neural Networks: Data Bias, Spurious Features, and Hardware Effects

Model ReleasesDGX agent

arXiv:2605.27407v1 Announce Type: cross Abstract: Evaluating fairness in Spiking Neural Networks (SNNs) demands rigorous benchmarks that reflect real-world complexities, yet existing assessments remai

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

Model ReleasesDGX agent

arXiv:2605.27492v1 Announce Type: cross Abstract: LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

Model ReleasesDGX agent

arXiv:2605.28183v1 Announce Type: cross Abstract: We introduce the BenGER (Benchmark for German Law) dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The BenGER d

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

Model ReleasesDGX agent

arXiv:2605.28301v1 Announce Type: new Abstract: Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics

Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI

Model ReleasesDGX agent

arXiv:2605.28707v1 Announce Type: new Abstract: Critical decision-making in socially consequential spaces is increasingly involving AI systems at varying capacities. Yet, despite the ubiquity of auton

Beyond Binary: Sim-to-Real Dexterous Manipulation with Physics-Grounded Contact Representation

SafetyDGX agent

arXiv:2605.28812v1 Announce Type: cross Abstract: A primary bottleneck in contact-rich manipulation is the difficulty of collecting real-world data. Sim-to-real reinforcement learning offers a scalabl

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

Model ReleasesDGX agent

arXiv:2502.05242v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain uncl

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2509.23074v3 Announce Type: replace-cross Abstract: In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark lea

Beyond Motion Primitives: Behavioral Activity Recognition from Head-Mounted IMU

Model ReleasesDGX agent

arXiv:2605.27464v1 Announce Type: cross Abstract: AR smart glasses need continuous behavioral context to offer proactive assistance, yet their most practical always-on sensor, the head-mounted Inertia

BiasEdit: A Training-Free Bias-Detect-and-Edit Framework for Learning Fair Visual Classifiers

SafetyDGX agent

arXiv:2605.28450v1 Announce Type: cross Abstract: Visual data from the Web power image classifiers, which often underpin many web services, such as recommendation and content moderation. However, the

BioELX: Cross-lingual Biomedical Entity Linking via Alias-based Retrieval and LLM Ranking

Model ReleasesDGX agent

arXiv:2605.27380v1 Announce Type: cross Abstract: Cross-lingual biomedical entity linking (BEL) maps mentions in any language to unique identifiers in a biomedical knowledge base (KB), supporting clin

BIRDNet: Mining and Encoding Boolean Implication Knowledge Graphs as Interpretable Deep Neural Networks

ResearchDGX agent

arXiv:2605.28739v1 Announce Type: cross Abstract: Tabular data in knowledge-rich domains often carries a latent prior in the form of Boolean implication relationships (BIRs) between pairs of features.

BIRDS: Characterizing and Understanding Biodiversity Impact of Large Language Model Serving

ResearchDGX agent

arXiv:2605.27480v1 Announce Type: cross Abstract: Large language model (LLM) serving creates environmental impacts beyond carbon and water, including ecosystem damage through biodiversity-related path

BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models

Model ReleasesDGX agent

arXiv:2605.28067v1 Announce Type: new Abstract: The remarkable generation quality of modern diffusion models often comes at the cost of massive parameter counts, which necessitate server-side inferenc

Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking

SafetyDGX agent

arXiv:2605.28632v1 Announce Type: cross Abstract: Cryptographic watermarking is a leading defense for attributing text generated by large language models (LLMs). Existing schemes, including KGW, Unigr

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information

SafetyDGX agent

arXiv:2605.28070v1 Announce Type: new Abstract: We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

Model ReleasesDGX agent

arXiv:2605.27383v1 Announce Type: cross Abstract: Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However,

BuddyBench: A Privacy-Constrained Multi-Task Benchmark for Pediatric Social-Communication Personalization

Model ReleasesDGX agent

arXiv:2605.28089v1 Announce Type: new Abstract: BuddyBench introduces a privacy-constrained multi-task benchmark for pediatric social-communication personalization. Unlike existing neurodevelopmental

C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning

TutorialsDGX agent

arXiv:2605.27860v1 Announce Type: new Abstract: Retrieval-augmented generation combined with reinforcement learning has shown promise for grounding large language models in trustworthy medical evidenc

Calibrating Conservatism for Scalable Oversight

AgentsDGX agent

arXiv:2605.28807v1 Announce Type: new Abstract: Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain mea

CaMBRAIN: Real-time, Continuous EEG Inference with Causal State Space Models

ResearchDGX agent

arXiv:2605.28792v1 Announce Type: new Abstract: Electroencephalography (EEG) is a critical, non-invasive method to monitor electrical brain activity. EEGs can span anywhere from a couple seconds to mu

Can I Have Your Order? Monte-Carlo Tree Search for Slot Filling Ordering in Diffusion Language Models

ResearchDGX agent

arXiv:2602.12586v2 Announce Type: replace Abstract: While plan-and-infill decoding in Masked Diffusion Models (MDMs) shows promise for mathematical and code reasoning, performance remains highly sensi

Can Quantum Federated Learning Withstand Circuit-Level Backdoors?

ResearchDGX agent

arXiv:2605.27416v1 Announce Type: cross Abstract: Quantum Federated Learning (QFL) inherits the core vulnerability of federated optimization to malicious clients, while also introducing an attack surf

Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought

Model ReleasesDGX agent

arXiv:2605.27764v1 Announce Type: cross Abstract: Recent segmentation models couple large language models (LLMs) with mask decoders to ground complex language expressions into masks, yet their instruc

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

SafetyDGX agent

arXiv:2603.22335v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior dis

Checking Fact with Better Retrieval: Dynamic Contrastive Learning for Evidence Retrieval

TutorialsDGX agent

arXiv:2605.27449v1 Announce Type: cross Abstract: In the field of multimodal fact checking, the accuracy of retrieving evidence from different modalities has a significant impact on the downstream cla

ChildEval: When large language models meet children's personalities

Model ReleasesDGX agent

arXiv:2605.27805v1 Announce Type: cross Abstract: While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-spec

CircuitLM: A Multi-Agent LLM-Aided Design Framework for Generating Circuit Schematics from Natural Language Prompts

AgentsDGX agent

arXiv:2601.04505v3 Announce Type: replace Abstract: Generating accurate circuit schematics from high-level natural language descriptions remains a persistent challenge in electronic design automation

CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text

Model ReleasesDGX agent

arXiv:2605.27700v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while contain

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models

Local AiDGX agent

arXiv:2605.28115v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face severe memory and latency bottlenecks due to high-resolution visual tokens. While current token reduction methods the

CLANE: Continual Learning of Actions on Neuromorphic Hardware from Event Cameras

Local AiDGX agent

arXiv:2605.28387v1 Announce Type: cross Abstract: Recognizing and continuously learning novel human actions without forgetting prior classes is a requirement for emerging AR/VR and robotics applicatio

Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings

ResearchDGX agent

arXiv:2605.28034v1 Announce Type: new Abstract: Clark Hash is a small method for storing neural embeddings in less space. It normalizes each database vector, applies a deterministic sparse signed John

Clinical Validation of the Melanoscope AI Mobile Dermoscopy Clinical Decision Support System

ResearchDGX agent

arXiv:2605.27561v1 Announce Type: cross Abstract: Introduction. Early detection of malignant skin lesions is critical for prognosis, yet dermatologist shortages in Russian regions limit screening cove

Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code

Model ReleasesDGX agent

arXiv:2603.24631v2 Announce Type: replace-cross Abstract: Code agents resolve 65-70% of SWE-bench Verified issues, but Pass@1 cannot tell us why the rest fail, and, as we show, capable-model failures

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

SafetyDGX agent

arXiv:2602.15198v2 Announce Type: replace-cross Abstract: Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperativ

Comparative Analysis of Liquid Neural Networks and LSTM for Sequential Pattern Recognition: Robustness, Efficiency, and Clinical Utility

Model ReleasesDGX agent

arXiv:2605.27467v1 Announce Type: cross Abstract: Traditional Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) units operate on discrete time steps, often failing to capture the flui

Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback

Model ReleasesDGX agent

arXiv:2605.28010v1 Announce Type: new Abstract: Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. H

Constrained Auto-Bidding via Generative Response Modeling

ResearchDGX agent

arXiv:2605.27811v1 Announce Type: new Abstract: Auto-bidding systems aim to maximize advertiser value over long horizons under budget constraints and ratio targets such as cost-per-acquisition, yet fu

Continual Model Routing in Evolving Model Hubs

Model ReleasesDGX agent

arXiv:2605.28577v1 Announce Type: new Abstract: AI model hubs provide access to a rapidly growing collection of powerful pre-trained models, enabling off-the-shelf mixture-of-experts systems with diff

COOP^2: Defining, Observing, and Repairing Cooperation in LLM Multi-Agent Systems

AgentsDGX agent

arXiv:2603.00349v2 Announce Type: replace Abstract: Many complex tasks require extended effort, diverse capabilities, or coordinated actions beyond what a single agent can provide. However, simply add

CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning

ResearchDGX agent

arXiv:2605.28742v1 Announce Type: new Abstract: Language models can use verifiable rewards to improve at a wide variety of reasoning tasks. However, both parametric (e.g. RLVR) and non-parametric (e.g

COTTA: Context-Aware Transfer Adaptation for Trajectory Prediction in Autonomous Driving

SafetyDGX agent

arXiv:2604.00402v2 Announce Type: replace-cross Abstract: Developing robust models to accurately predict the trajectories of surrounding agents is fundamental to autonomous driving safety. However, mo

CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders

SafetyDGX agent

arXiv:2604.01604v2 Announce Type: replace Abstract: While modern LLMs are aligned to refuse harmful requests, it is essential to understand the underlying mechanistic basis of this refusal behavior fo

Cross-Entropy Games and Frost Training

SafetyDGX agent

arXiv:2605.27701v1 Announce Type: new Abstract: We present Frost Training, a method for improving Monte Carlo-based policy optimization for a large family of LLM-as-a-judge tasks called Cross-Entropy

← Previous
1…196197198199200…358
Next →