AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
11 May 2026

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight

Model ReleasesDGX agent

arXiv:2605.07021v1 Announce Type: new Abstract: Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors do not surface until reasoning concludes. To addr

Belief Memory: Agent Memory Under Partial Observability

AgentsDGX agent

arXiv:2605.05583v2 Announce Type: replace Abstract: LLM agents that operate over long context depend on external memory to accumulate knowledge over time. However, existing methods typically store eac

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs

Model ReleasesDGX agent

arXiv:2605.07731v1 Announce Type: cross Abstract: This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE)


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Benchmarking World-Model Learning with Environment-Level Queries

Model ReleasesDGX agent

arXiv:2510.19788v4 Announce Type: replace Abstract: World models are central to building AI agents capable of flexible reasoning and planning. Yet current evaluations (i) test only properties measurab

Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

SafetyDGX agent

arXiv:2605.07806v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved fro

Beyond Factor Aggregation: Gauge-Aware Low-Rank Server Representations for Federated LoRA

Model ReleasesDGX agent

arXiv:2605.06733v1 Announce Type: cross Abstract: Federated LoRA enables parameter-efficient adaptation of large language models under decentralized data and limited client resources.However, directly

Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation

Model ReleasesDGX agent

arXiv:2605.07111v1 Announce Type: cross Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides the representational pla

Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph

SafetyDGX agent

arXiv:2605.08037v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) aligns language models using pairwise preference comparisons, offering a simple and effective alternative to Rein

Beyond Retrieval: A Multitask Benchmark and Model for Code Search

Model ReleasesDGX agent

arXiv:2605.04615v2 Announce Type: replace-cross Abstract: Code search has usually been evaluated as first-stage retrieval, even though production systems rely on broader pipelines with reranking and d

Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies

SafetyDGX agent

arXiv:2602.23811v4 Announce Type: replace-cross Abstract: We investigate the theoretical aspects of offline reinforcement learning (RL) under general function approximation. While prior works (e.g., X

Beyond the Black Box: Interpretability of Agentic AI Tool Use

Model ReleasesDGX agent

arXiv:2605.06890v1 Announce Type: new Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagn

BGM-IV: an AI-powered Bayesian generative modeling approach for instrumental variable analysis

Model ReleasesDGX agent

arXiv:2605.07029v1 Announce Type: cross Abstract: Instrumental-variable (IV) regression enables causal estimation under endogeneity, but modern IV problems often involve nonlinear structural effects a

Bifurcation Models: Learning Set-Valued Solution Maps with Weight-Tied Dynamics

TutorialsDGX agent

arXiv:2605.07277v1 Announce Type: cross Abstract: Many scientific and combinatorial problems admit multiple correct solutions, not a single label. Standard supervised learning resolves this ambiguity

BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation

Model ReleasesDGX agent

arXiv:2605.07306v1 Announce Type: cross Abstract: Biological laboratory automation can reduce repetitive manual work and improve reproducibility, but reliable embodied execution in wet-lab environment

Bounded Fitting for Expressive Description Logics

ResearchDGX agent

arXiv:2605.07452v1 Announce Type: new Abstract: Bounded fitting is an attractive paradigm for learning logical formulas from labeled data examples that offers PAC-style generalization guarantees and c

Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing

Model ReleasesDGX agent

arXiv:2605.06936v1 Announce Type: cross Abstract: LLM-based agents are increasingly applied to the 'last mile' of Electronic Design Automation (EDA): repairing residual sign-off Design Rule Check (DRC

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation

Model ReleasesDGX agent

arXiv:2605.08057v1 Announce Type: cross Abstract: While recent advancements in inference-time learning have improved LLM reasoning on Text-to-SQL tasks, current solutions still struggle to perform wel

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning

Model ReleasesDGX agent

arXiv:2605.07251v1 Announce Type: new Abstract: Large Language Models (LLMs) have become increasingly capable as tool-using agents, with benchmarks spanning diverse general agentic tasks. Yet rigorous

Can You Break RLVER? Probing Adversarial Robustness of RL-Trained Empathetic Agents

Model ReleasesDGX agent

arXiv:2605.07138v1 Announce Type: new Abstract: Reinforcement learning from verifiable emotion rewards RLVER has produced language models with strong empathetic performance, evaluated on benchmarks th

CASCADE: Case-Based Continual Adaptation for Large Language Models During Deployment

AgentsDGX agent

arXiv:2605.06702v1 Announce Type: new Abstract: Large language models (LLMs) have become a central foundation of modern artificial intelligence, yet their lifecycle remains constrained by a rigid sepa

CASCADE: Context-Aware Relaxation for Speculative Image Decoding

ResearchDGX agent

arXiv:2605.07230v1 Announce Type: cross Abstract: Autoregressive generation is a powerful approach for high-fidelity image synthesis, but it remains computationally demanding and slow even on the most

Causal EpiNets: Precision-corrected Bounds on Individual Treatment Effects using Epistemic Neural Networks

SafetyDGX agent

arXiv:2605.07065v1 Announce Type: cross Abstract: Individual treatment effects are not point-identified from data. The Probability of Necessity and Sufficiency (PNS) circumvents this limitation by cha

Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models

HardwareDGX agent

arXiv:2605.07194v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a small synthetic set that preserves downstream training utility. While most existing method

CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers

Model ReleasesDGX agent

arXiv:2605.07905v1 Announce Type: cross Abstract: Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness.

Cognitive Agent Compilation for Explicit Problem Solver Modeling

SafetyDGX agent

arXiv:2605.07040v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used for tutoring, feedback generation, and content creation, but their broad pretraining makes them hard to c

Conditional generation of antibody sequences with classifier-guided germline-absorbing discrete diffusion

SafetyDGX agent

arXiv:2605.06720v1 Announce Type: cross Abstract: Antibody therapeutics are among the most successful modern medicines, yet computationally designing antibodies with desirable binding and developabili

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable

SafetyDGX agent

arXiv:2605.07353v1 Announce Type: new Abstract: Large reasoning models often reach correct answers through flawed intermediate steps, creating a gap between final accuracy and reasoning reliability. E

Contextual Invertible World Models: A Neuro-Symbolic Agentic Framework for Colorectal Cancer Drug Response

AgentsDGX agent

arXiv:2603.02274v2 Announce Type: replace-cross Abstract: Precision oncology is currently limited by the small-N, large-P paradox, where high-dimensional genomic data is abundant but pharmacological r

Continually Evolving Skill Knowledge in Vision Language Action Model

Model ReleasesDGX agent

arXiv:2511.18085v4 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains chal

CRAFT: Forgetting-Aware Intervention-Based Adaptation for Continual Learning

ResearchDGX agent

arXiv:2605.05732v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can acquire new capabilities through fine-tuning, but continual adaptation often leads to catastrophic forgetting

Cross-Attention and Encoder-Decoder Transformers: A Logical Characterization

ResearchDGX agent

arXiv:2605.07705v1 Announce Type: cross Abstract: We give a novel logical characterization of encoder-decoder transformers, the foundational architecture for LLMs that also sees use in various setting

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs

Model ReleasesDGX agent

arXiv:2605.06115v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs), trained primarily on English-centric data, frequently generate culturally inappropriate or misaligned resp

CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2601.03728v3 Announce Type: replace-cross Abstract: Composed Image Retrieval (CIR) enables users to search for target images using both a reference image and manipulation text, offering substant

CSR: Infinite-Horizon Real-Time Policies with Massive Cached State Representations

Model ReleasesDGX agent

arXiv:2605.07325v1 Announce Type: cross Abstract: Deploying massive large language models (LLMs) as continuous cognitive engines for robotics is bottlenecked by the time-to-first-token (TTFT) latency

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences

SafetyDGX agent

arXiv:2605.07724v1 Announce Type: cross Abstract: Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios

Model ReleasesDGX agent

arXiv:2605.07830v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenom

DCGL: Dual-Channel Graph Learning with Large Language Models for Knowledge-Aware Recommendation

SafetyDGX agent

arXiv:2605.07314v1 Announce Type: cross Abstract: Knowledge Graphs (KGs) have proven highly effective for recommendation systems by capturing latent item relationships, while recent integration of Lar

Decentralized Time-Varying Optimization for Streaming Data via Temporal Weighting

SafetyDGX agent

arXiv:2605.06971v1 Announce Type: cross Abstract: Classical optimization theory largely focuses on fixed objective functions, whereas many modern learning systems operate in dynamic environments where

Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks

Model ReleasesDGX agent

arXiv:2603.04676v2 Announce Type: replace-cross Abstract: Multi-image reasoning remains a significant challenge for vision-language models (VLMs). We investigate a previously overlooked phenomenon: du

Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation

ResearchDGX agent

arXiv:2605.07694v1 Announce Type: cross Abstract: Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which com

Detecting Distillation Data from Reasoning Models

Model ReleasesDGX agent

arXiv:2510.04850v3 Announce Type: replace-cross Abstract: Reasoning distillation has emerged as a prevailing paradigm for transferring reasoning capabilities from large reasoning models to small langu

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks

ApplicationsDGX agent

arXiv:2506.13351v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) training of large language models (LLMs) on unverifiable tasks is challenging even when a reasonable-quality refer

Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs

SafetyDGX agent

arXiv:2602.22831v2 Announce Type: replace-cross Abstract: Moral benchmarks for LLMs typically score models on context-free prompts, implicitly treating the measured choice rate as stable. We test this

Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points

AgentsDGX agent

arXiv:2412.11194v2 Announce Type: replace-cross Abstract: Security vulnerabilities in software can have severe consequences; however, manual vulnerability detection is costly and does not scale, espec

Discovering Learning-Friendly Generation Orders for Sequential Computation

ApplicationsDGX agent

arXiv:2506.23875v4 Announce Type: replace-cross Abstract: Sequential computation via autoregressive generation can make difficult tasks learnable, but the generation order of intermediate states stron

Discovering Multiagent Learning Algorithms with Large Language Models

SafetyDGX agent

arXiv:2602.16928v3 Announce Type: replace-cross Abstract: Much of the advancement in Multi-Agent Reinforcement Learning (MARL) for imperfect-information games has historically depended on the manual,

Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation

Model ReleasesDGX agent

arXiv:2605.07323v1 Announce Type: new Abstract: Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing symbolic regres

Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport

ResearchDGX agent

arXiv:2605.06785v1 Announce Type: cross Abstract: Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We prop

Divide and Conquer: Object Co-occurrence Helps Mitigate Simplicity Bias in OOD Detection

Model ReleasesDGX agent

arXiv:2605.07821v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models. Existing methods mostly focus on regular entangle

Do Joint Audio-Video Generation Models Understand Physics?

Model ReleasesDGX agent

arXiv:2605.07061v1 Announce Type: cross Abstract: Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visu

Does Your Neural Network Extrapolate? Feature Engineering as Identifiability Bias for OOD Generalization

SafetyDGX agent

arXiv:2605.07483v1 Announce Type: cross Abstract: Successful deep neural networks discover salient features of data. We show when and why they fail to learn out-of-distribution (OOD)-relevant represen

Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas

Model ReleasesDGX agent

arXiv:2605.06673v1 Announce Type: cross Abstract: Aggregate metacognitive quality scores mask within-model variation across MMLU benchmark domains. We administered 1,500 MMLU items (250 per domain, un

Don't Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment

SafetyDGX agent

arXiv:2605.06885v1 Announce Type: cross Abstract: Diffusion language models (DLMs) have recently demonstrated capabilities that complement standard autoregressive (AR) models, particularly in non-sequ

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation

HardwareDGX agent

arXiv:2605.07985v1 Announce Type: cross Abstract: Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, s

DPG-CD: Depth-Prior-Guided Cross-Modal Joint 2D-3D Change Detection

ResearchDGX agent

arXiv:2605.07151v1 Announce Type: cross Abstract: Urban spatial evolution is manifested not only through horizontal expansion but also through vertical structural changes. Consequently, jointly captur

Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training

SafetyDGX agent

arXiv:2605.07063v1 Announce Type: cross Abstract: Data selection methods address a critical challenge in LLM post-training: effectively leveraging scarce, high-fidelity target data alongside abundant

Drawing Lines in Psychological Space: What K-means Clustering Reveals in Simulated and Real Psychometric Data

ResearchDGX agent

arXiv:2605.06989v1 Announce Type: cross Abstract: K-means clustering is widely used in psychological and psychometric research to identify profiles, subgroups, and potential typologies, yet its classi

Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow

SafetyDGX agent

arXiv:2605.07727v1 Announce Type: cross Abstract: We propose Drifting Field Policy (DFP), a non-ODE one-step generative policy built on the drifting model paradigm. We frame the policy update as a rev

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

Model ReleasesDGX agent

arXiv:2605.07699v1 Announce Type: cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently a

DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization

Model ReleasesDGX agent

arXiv:2512.14263v2 Announce Type: replace-cross Abstract: Preferential Bayesian Optimization (PBO) aims to find a decision-maker's most preferred solution in as few pairwise comparisons as possible. E

← Previous
1…278279280281282…358
Next →