AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
19 May 2026

Automated Root-Cause Subclassification and No-Code Fix Generation for Invalid Bug Reports

Model ReleasesDGX agent

arXiv:2605.17561v1 Announce Type: cross Abstract: Issues faced when using software are reported in the form of bug reports. However, many bug reports are invalid, meaning they do not require code chan

Automatic Generation of High-Performance RL Environments

SafetyDGX agent

arXiv:2603.12145v2 Announce Type: replace-cross Abstract: Translating complex reinforcement learning (RL) environments into high-performance implementations has traditionally required months of specia

Automatic Unsupervised Ensemble Outlier Model Selection--Extended Version

ApplicationsDGX agent

arXiv:2605.16567v1 Announce Type: cross Abstract: Unsupervised outlier detection is attractive because it eliminates the need for labeled data. Moreover, forming multi-model ensembles can improve dete


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

SafetyDGX agent

arXiv:2605.17602v1 Announce Type: new Abstract: Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images acc

Avoiding Structural Failure Modes in Tabular Fair SSL: Online Primal-Dual Allocation under Confidence Gating

SafetyDGX agent

arXiv:2605.16446v1 Announce Type: cross Abstract: Semi-supervised learning (SSL) enables prediction with limited labels, but high-stakes tabular applications (medical, credit, recidivism) require stat

Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models

AgentsDGX agent

arXiv:2605.16725v1 Announce Type: new Abstract: Executable world models can be read, edited, executed, and reused for planning, but only if the program captures the environment's transition law rather

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

Model ReleasesDGX agent

arXiv:2605.17971v1 Announce Type: cross Abstract: Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuri

BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting

Model ReleasesDGX agent

arXiv:2605.17937v1 Announce Type: cross Abstract: Quantitative backtesting is essential for evaluating trading strategies but remains hampered by high technical barriers and limited scalability. While

Balancing Knowledge Distillation for Imbalance Learning with Bilevel Optimization

ResearchDGX agent

arXiv:2605.17839v1 Announce Type: cross Abstract: Knowledge distillation transfers knowledge from a high capacity teacher to a compact student using a mixture of hard and soft losses. On imbalanced da

Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity

Model ReleasesDGX agent

arXiv:2510.00304v3 Announce Type: replace-cross Abstract: Deep learning models excel in stationary data but struggle in non-stationary environments due to a phenomenon known as loss of plasticity (LoP

Bayesian-Monte Carlo Schedule Updating for Construction Digital Twins: A Probabilistic Framework for Dynamic Project Forecasting

Model ReleasesDGX agent

arXiv:2605.17608v1 Announce Type: cross Abstract: Construction projects frequently experience schedule delays and forecasting uncertainty due to variability in labor productivity, material availabilit

Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models

Model ReleasesDGX agent

arXiv:2510.16727v2 Announce Type: replace-cross Abstract: Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that

Benchmarking Mythos-Linked Bug Rediscovery

Model ReleasesDGX agent

arXiv:2605.17416v1 Announce Type: cross Abstract: Anthropic's April 2026 Mythos materials combine benchmark claims with concrete bug-finding stories across OpenBSD, FreeBSD, Linux, FFmpeg, and browser

BESplit: Bias-Compensated Split Federated Learning with Evidential Aggregation

Model ReleasesDGX agent

arXiv:2605.17508v1 Announce Type: cross Abstract: Split Federated Learning (SFL) enables privacy-preserving collaborative training by partitioning models between clients and a server. However, under n

Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs

Model ReleasesDGX agent

arXiv:2602.09805v2 Announce Type: replace-cross Abstract: As reasoning LLMs increasingly trade tokens for accuracy through deliberation, search, and self-correction, a single accuracy score can no lon

Beyond Accuracy: Robustness, Interpretability and Expressiveness of EEG Foundation Models

ResearchDGX agent

arXiv:2605.17562v1 Announce Type: cross Abstract: EEG foundation models (EEG-FMs) have been evaluated predominantly on clean, in-distribution accuracy, leaving their robustness, interpretability and r

Beyond Catalogue Counts: the Dataset Visibility Asymmetry in Low-Resource Multilingual NLP

ApplicationsDGX agent

arXiv:2605.17442v1 Announce Type: cross Abstract: Multilingual NLP often relies on dataset counts from centralized catalogues to characterize which languages are resource-rich or resource-poor. Howeve

Beyond Compliance: How AI Could Help Creative Writers by Refusing Them

SafetyDGX agent

arXiv:2605.16272v1 Announce Type: cross Abstract: Mainstream creativity support design prioritizes compliant AI for seamless writing interactions, but concerns over inappropriate AI reliance highlight

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training

ResearchDGX agent

arXiv:2509.03403v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves final-answer accuracy on reasoning tasks, but it does not reliably improve reas

Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation

ResearchDGX agent

arXiv:2605.17174v1 Announce Type: cross Abstract: Reinforcement Learning (RL) is an important paradigm for aligning Diffusion Language Models (DLMs) toward functional correctness in code generation. H

Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech

ApplicationsDGX agent

arXiv:2605.16280v1 Announce Type: cross Abstract: Automating legal reasoning forces a choice between imperfect alternatives: symbolic systems offer transparency but struggle with ambiguity, whereas ne

Beyond Inference-Time Search: Reinforcement Learning Synthesizes Reusable Solvers

Local AiDGX agent

arXiv:2605.18374v1 Announce Type: cross Abstract: Large language models (LLMs) typically approach combinatorial optimization as an inference-time procedure, solving each instance separately through sa

Beyond Linear Superposition: Discovering Climate Features in AI Weather Models with KAN-SAE

ResearchDGX agent

arXiv:2605.17493v1 Announce Type: cross Abstract: Deep learning weather prediction models achieve remarkable predictive skill yet remain largely opaque: we know little about how they represent physica

Beyond Morphology: Quantifying the Diagnostic Power of Color Features in Cancer Classification

ResearchDGX agent

arXiv:2605.18522v1 Announce Type: cross Abstract: In histopathology, human experts primarily rely on color as a means of enhancing contrast to interpret tissue morphology, whereas machine vision model

Beyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon Planning

SafetyDGX agent

arXiv:2508.03018v2 Announce Type: replace Abstract: Large Language Reasoning Models have demonstrated remarkable success on static tasks, yet their application to multi-round agentic planning in inter

Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2601.16527v2 Announce Type: replace-cross Abstract: Multimodal LLMs are powerful but prone to object hallucinations, which describe non-existent entities and harm reliability. While recent unlea

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks

Local AiDGX agent

arXiv:2605.18194v1 Announce Type: new Abstract: While Multi-Modal Large Language Models (MLLMs) demonstrate impressive capabilities in general reasoning, their embodied spatial intelligence remains ha

BioProAgent: Neuro-Symbolic Grounding for Constrained Scientific Planning

Model ReleasesDGX agent

arXiv:2603.00876v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant reasoning capabilities in scientific discovery but struggle to bridge the gap to physical

BLAgent: Agentic RAG for File-Level Bug Localization

Local AiDGX agent

arXiv:2605.17965v1 Announce Type: cross Abstract: Bug localization remains a key bottleneck in downstream software maintenance tasks, including root cause analysis, triage, and automated program repai

BlendedNet++: A dataset and benchmark for field-resolved aerodynamics and inverse design of blended wing body aircraft

Model ReleasesDGX agent

arXiv:2512.03280v2 Announce Type: replace-cross Abstract: The conceptual design of Blended Wing Body (BWB) aircraft is often constrained by the high computational cost of resolving complex aerodynamic

Body-Grounded Perspective Formation and Conative Attunement in Artificial Agents

SafetyDGX agent

arXiv:2605.16728v1 Announce Type: new Abstract: This paper proposes a minimal architecture for body-grounded perspective formation in artificial agents. Extending prior work, the model introduces an i

BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks

Model ReleasesDGX agent

arXiv:2605.17000v1 Announce Type: cross Abstract: Optimization of LLM training and inference configurations, such as hyperparameters, data mixtures, and prompts, is critical to performance, but it is

Brain Vascular Age Prediction Using Cerebral Blood Flow Velocity and Machine Learning Algorithms

ResearchDGX agent

arXiv:2605.16969v1 Announce Type: new Abstract: Defining vascular age in terms of physiological function has become one focal point of the extensive studies to categorize and track chronological age.

Breaking the accuracy-resource dilemma: a lightweight adaptive video inference enhancement

ResearchDGX agent

arXiv:2601.14568v2 Announce Type: replace-cross Abstract: Existing video inference (VI) enhancement methods typically aim to improve performance by scaling up model sizes and employing sophisticated n

Bridging the Version Gap: Multi-version Training Improves ICD Code Prediction, Especially for Rare Codes

ResearchDGX agent

arXiv:2605.17755v1 Announce Type: cross Abstract: Clinical coding maps clinical documentation to standardized medical codes, an essential yet time-consuming administrative task that could benefit from

Building Reliable Arithmetic Multipliers Under NBTI Aging and Process Variations

SafetyDGX agent

arXiv:2605.18444v1 Announce Type: cross Abstract: Hardware aging poses a significant challenge for integrated circuits (ICs), leading to performance degradation and eventual failure. In this work, we

Byzantine-Resilient Federated Learning via QUBO-Based Client Selection on Quantum Annealers

ResearchDGX agent

arXiv:2605.16438v1 Announce Type: cross Abstract: Federated Learning (FL) trains a global model across decentralized clients while preserving data privacy, but at scale it is vulnerable to malicious u

Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents

AgentsDGX agent

arXiv:2602.16699v3 Announce Type: replace-cross Abstract: LLM agents are deployed in environments where they must interact to acquire information. In these scenarios, the agent must reason about inher

CAM-Bench: A Benchmark for Computational and Applied Mathematics in Lean

Model ReleasesDGX agent

arXiv:2605.17255v1 Announce Type: new Abstract: Formal theorem-proving benchmarks enable mechanically verifiable evaluation of mathematical reasoning in large language models. However, existing benchm

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection

ResearchDGX agent

arXiv:2605.17133v1 Announce Type: cross Abstract: The rapid advancement of Deepfake technologies and video manipulation tools poses a critical challenge to multimedia forensics, judicial evidence inte

Can Heterogeneous Language Models Be Fused?

Model ReleasesDGX agent

arXiv:2604.01674v2 Announce Type: replace Abstract: Model merging aims to integrate multiple expert models into a single model that inherits their complementary strengths without incurring the inferen

Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment

AgentsDGX agent

arXiv:2603.23638v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly tested on complex tasks, but their ability to allocate scarce resources over long horizons remain

Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks

ResearchDGX agent

arXiv:2510.01782v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) should refuse to answer questions beyond their knowledge. This capability, which we term knowledge-aware refusal,

Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench

Model ReleasesDGX agent

arXiv:2605.17079v1 Announce Type: cross Abstract: LLMs are increasingly used as ``digital consumers'' to simulate public opinion, pre-test marketing decisions, and anticipate audience response. Howeve

CANSURF: An ASV-View Can Dataset and Benchmark for Detection and Tracking of Surface-Level Debris

Model ReleasesDGX agent

arXiv:2605.16774v1 Announce Type: cross Abstract: Surface-level marine debris remains a practical bottleneck for autonomous clean-up, where small, reflective targets (e.g., aluminum cans) must be dete

Capturing LLM Capabilities via Evidence-Calibrated Query Clustering

ResearchDGX agent

arXiv:2605.17110v1 Announce Type: new Abstract: Query clustering organizes queries into groups that reflect shared latent capability demands, enabling capability-aware LLM evaluation. Existing cluster

CarbonScaling: Extending Neural Scaling Laws for Carbon Footprint in Large Language Models

Model ReleasesDGX agent

arXiv:2508.06524v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly follow neural scaling laws that tie performance gains to rapidly expanding computational budgets, ra

CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

Model ReleasesDGX agent

arXiv:2605.17176v1 Announce Type: new Abstract: Emotion understanding is a core capability for LLMs to interact effectively with humans, yet existing evaluation paradigms rely on discrete emotion labe

CasualSynth: Generating Structurally Sound Synthetic Data

Model ReleasesDGX agent

arXiv:2605.17528v1 Announce Type: cross Abstract: Large Language Models (LLMs) generate realistic synthetic data but offer no guarantee that their outputs respect the causal mechanisms governing the t

CATA: Continual Machine Unlearning via Conflict-Averse Task Arithmetic

ResearchDGX agent

arXiv:2605.18610v1 Announce Type: cross Abstract: Vision-language models (VLMs) have shown remarkable ability in aligning visual and textual representations, enabling a wide range of multimodal applic

CatalyticMLLM: A Graph-Text Multimodal Large Language Model for Catalytic Materials

SafetyDGX agent

arXiv:2605.17254v1 Announce Type: new Abstract: Property prediction and inverse structural design of catalytic materials are typically modeled as two independent tasks: the former predicts target prop

Catastrophic Overfitting, Entropy Gap and Participation Ratio: A Noiseless l^p Norm Solution for Fast Adversarial Training

ResearchDGX agent

arXiv:2505.02360v2 Announce Type: replace-cross Abstract: Adversarial training is a cornerstone of robust deep learning, but fast methods like the Fast Gradient Sign Method (FGSM) often suffer from Ca

Causal Intervention-Based Memory Selection for Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2605.17641v1 Announce Type: new Abstract: Long-horizon LLM agents rely on persistent memory to support interactions across sessions, yet existing memory systems often retrieve context using sema

Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows

Model ReleasesDGX agent

arXiv:2605.18327v1 Announce Type: new Abstract: AI agents deployed into SRE workflows currently derive their understanding of environment state from raw observability telemetry at query time, paying a

CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning

TutorialsDGX agent

arXiv:2605.16416v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have achieved strong performance on general multimodal reasoning, yet remain challenged in integrating nonlocal visual i

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings

TutorialsDGX agent

arXiv:2605.17370v1 Announce Type: new Abstract: Cognitive behavioural therapy is widely used to help patients understand and manage psychological distress. It is often delivered through spoken convers

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference

HardwareDGX agent

arXiv:2605.17164v1 Announce Type: cross Abstract: Deploying large-scale LLM training and inference with optimal performance is exceptionally challenging due to a complex design space of parallelism st

ChartDesign: Towards LLM Designer of Data Visualization

SafetyDGX agent

arXiv:2605.16274v1 Announce Type: cross Abstract: Charts are the dominant medium for visualizing data, discovering patterns and trends, and communicating data driven insights, yet designing them still

CheckSupport: A Local LLM-Powered Tool for Automated Manuscript Submission Checklist Selection and Completion

Local AiDGX agent

arXiv:2605.16377v1 Announce Type: cross Abstract: Transparent and standardized reporting is essential for reproducible scientific research, yet adherence to reporting guidelines remains inconsistent b

ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding

SafetyDGX agent

arXiv:2605.17214v1 Announce Type: new Abstract: While Large Language Models (LLMs) have revolutionized scientific text processing, they exhibit a significant capability gap when interpreting chemical

← Previous
1…233234235236237…358
Next →