AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
22 Apr 2026

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors

SafetyDGX agent

arXiv:2603.15956v2 Announce Type: replace-cross Abstract: Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data. While human demonstrations (

Explicit Trait Inference for Multi-Agent Coordination

AgentsDGX agent

arXiv:2604.19278v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) show promise on complex tasks but remain prone to coordination failures such as goal drift, error cascades, and misa

Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck

SafetyDGX agent

arXiv:2601.12499v2 Announce Type: replace Abstract: Despite scaling to massive context windows, Large Language Models (LLMs) struggle with multi-hop reasoning due to inherent position bias, which caus


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Fairness Audits of Institutional Risk Models in Deployed ML Pipelines

SafetyDGX agent

arXiv:2604.19468v1 Announce Type: cross Abstract: Fairness audits of institutional risk models are critical for understanding how deployed machine learning pipelines allocate resources. Drawing on mul

FASE : A Fairness-Aware Spatiotemporal Event Graph Framework for Predictive Policing

SafetyDGX agent

arXiv:2604.18644v1 Announce Type: cross Abstract: Predictive policing systems that allocate patrol resources based solely on predicted crime risk can unintentionally amplify racial disparities through

FASTER: Value-Guided Sampling for Fast RL

SafetyDGX agent

arXiv:2604.19730v1 Announce Type: cross Abstract: Some of the most performant reinforcement learning algorithms today can be prohibitively expensive as they use test-time scaling methods such as sampl

FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware Fusion

Model ReleasesDGX agent

arXiv:2604.19015v1 Announce Type: cross Abstract: Federated fine-tuning of Large Language Models (LLMs) is obstructed by a trilemma of challenges: protecting LLMs intellectual property (IP), ensuring

Fine-Tuning Code Language Models to Detect Cross-Language Bugs

ResearchDGX agent

arXiv:2507.21954v2 Announce Type: replace-cross Abstract: Multilingual programming, which involves using multiple programming languages (PLs) in a single project, is increasingly common due to its ben

Fine-tuning DeepSeek-OCR-2 for Molecular Structure Recognition

Model ReleasesDGX agent

arXiv:2604.03476v2 Announce Type: replace-cross Abstract: Optical Chemical Structure Recognition (OCSR) is critical for converting 2D molecular diagrams from printed literature into machine-readable f

Fine-Tuning Small Reasoning Models for Quantum Field Theory

Model ReleasesDGX agent

arXiv:2604.18936v1 Announce Type: cross Abstract: Despite the growing application of Large Language Models (LLMs) to theoretical physics, there is little academic exploration into how domain-specific

Formally Verified Patent Analysis via Dependent Type Theory: Machine-Checkable Certificates from a Hybrid AI + Lean 4 Pipeline

ApplicationsDGX agent

arXiv:2604.18882v1 Announce Type: new Abstract: We present a formally verified framework for patent analysis as a hybrid AI + Lean 4 pipeline. The DAG-coverage core (Algorithm 1b) is fully machine-ver

Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents

Model ReleasesDGX agent

arXiv:2604.19457v1 Announce Type: new Abstract: Long-horizon enterprise agents make high-stakes decisions (loan underwriting, claims adjudication, clinical review, prior authorization) under lossy mem

From Craft to Kernel: A Governance-First Execution Architecture and Semantic ISA for Agentic Computers

AgentsDGX agent

arXiv:2604.18652v1 Announce Type: cross Abstract: The transition of agentic AI from brittle prototypes to production systems is stalled by a pervasive crisis of craft. We suggest that the prevailing o

From Experience to Skill: Multi-Agent Generative Engine Optimization via Reusable Strategy Learning

Model ReleasesDGX agent

arXiv:2604.19516v1 Announce Type: new Abstract: Generative engines (GEs) are reshaping information access by replacing ranked links with citation-grounded answers, yet current Generative Engine Optimi

From Natural Language to Executable Narsese: A Neuro-Symbolic Benchmark and Pipeline for Reasoning with NARS

Model ReleasesDGX agent

arXiv:2604.18873v1 Announce Type: new Abstract: Large language models (LLMs) are highly capable at language generation, but they remain unreliable when reasoning requires explicit symbolic structure,

GAIN: Multiplicative Modulation for Domain Adaptation

ResearchDGX agent

arXiv:2604.04516v2 Announce Type: replace-cross Abstract: Adapting LLMs to new domains causes forgetting because standard methods (e.g., full fine-tuning, LoRA) inject new directions into the weight s

GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations

ResearchDGX agent

arXiv:2503.16683v2 Announce Type: replace-cross Abstract: Vision Transformer (ViT) has been widely used in computer vision tasks with excellent results by providing representations for a whole image o

Gated Memory Policy

Model ReleasesDGX agent

arXiv:2604.18933v1 Announce Type: cross Abstract: Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks that depend

Generalization at the Edge of Stability

ResearchDGX agent

arXiv:2604.19740v1 Announce Type: cross Abstract: Training modern neural networks often relies on large learning rates, operating at the edge of stability, where the optimization dynamics exhibit osci

GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines

Model ReleasesDGX agent

arXiv:2508.06226v2 Announce Type: replace Abstract: Geometry problem solving (GPS) poses significant challenges for Multimodal Large Language Models (MLLMs) in diagram comprehension, knowledge applica

Geometric Decoupling: Diagnosing the Structural Instability of Latent

ResearchDGX agent

arXiv:2604.18804v1 Announce Type: cross Abstract: Latent Diffusion Models (LDMs) achieve high-fidelity synthesis but suffer from latent space brittleness, causing discontinuous semantic jumps during e

GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes

ResearchDGX agent

arXiv:2604.19411v1 Announce Type: cross Abstract: Understanding road scenes in a geometrically consistent, scene-centric representation is crucial for planning and mapping. We present GOLD-BEV, a fram

Gradient-Based Program Synthesis with Neurally Interpreted Languages

TutorialsDGX agent

arXiv:2604.18907v1 Announce Type: cross Abstract: A central challenge in program induction has long been the trade-off between symbolic and neural approaches. Symbolic methods offer compositional gene

GRAIL:Learning to Interact with Large Knowledge Graphs for Retrieval Augmented Reasoning

SafetyDGX agent

arXiv:2508.05498v2 Announce Type: replace Abstract: Large Language Models (LLMs) integrated with Retrieval-Augmented Generation (RAG) techniques have exhibited remarkable performance across a wide ran

Graph Data Augmentation with Contrastive Learning on Covariate Distribution Shift

ApplicationsDGX agent

arXiv:2512.00716v2 Announce Type: replace-cross Abstract: Covariate distribution shift occurs when certain structural features present in the test set are absent from the training set. It is a common

GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models

Model ReleasesDGX agent

arXiv:2604.19398v1 Announce Type: new Abstract: Large language models (LLMs) are expensive to serve because model parameters, attention computation, and KV caches impose substantial memory and latency

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models

Model ReleasesDGX agent

arXiv:2604.19300v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models

Handling and Interpreting Missing Modalities in Patient Clinical Trajectories via Autoregressive Sequence Modeling

ApplicationsDGX agent

arXiv:2604.18753v1 Announce Type: cross Abstract: An active challenge in developing multimodal machine learning (ML) models for healthcare is handling missing modalities during training and deployment

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams

Model ReleasesDGX agent

arXiv:2604.18901v1 Announce Type: cross Abstract: Harmful intent is geometrically recoverable from large language model residual streams: as a linear direction in most layers, and as angular deviation

Has Automated Essay Scoring Reached Sufficient Accuracy? Deriving Achievable QWK Ceilings from Classical Test Theory

Model ReleasesDGX agent

arXiv:2604.19131v1 Announce Type: new Abstract: Automated essay scoring (AES) is commonly evaluated on public benchmarks using quadratic weighted kappa (QWK). However, because benchmark labels are ass

HELM: Harness-Enhanced Long-horizon Memory for Vision-Language-Action Manipulation

Model ReleasesDGX agent

arXiv:2604.18791v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models fail systematically on long-horizon manipulation tasks despite strong short-horizon performance. We show that this

Hierarchically Robust Zero-shot Vision-language Models

SafetyDGX agent

arXiv:2604.18867v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) can perform zero-shot classification but are susceptible to adversarial attacks. While robust fine-tuning improves their

How Adversarial Environments Mislead Agentic AI?

Model ReleasesDGX agent

arXiv:2604.18874v1 Announce Type: new Abstract: Tool-integrated agents are deployed on the premise that external tools ground their outputs in reality. Yet this very reliance creates a critical attack

How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning

TutorialsDGX agent

arXiv:2604.19149v1 Announce Type: cross Abstract: Thinking LLMs produce reasoning traces before answering. Prior activation steering work mainly targets on shaping these traces. It remains less unders

How does the optimizer implicitly bias the model merging loss landscape?

SafetyDGX agent

arXiv:2510.04686v2 Announce Type: replace-cross Abstract: Model merging combines independent solutions with different capabilities into a single one while maintaining the same inference cost. Two popu

How to Teach Large Multimodal Models New Skills

SafetyDGX agent

arXiv:2510.08564v2 Announce Type: replace Abstract: How can we teach large multimodal models (LMMs) new skills without erasing prior abilities? We study sequential fine-tuning on five target skills wh

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

Model ReleasesDGX agent

arXiv:2604.19406v1 Announce Type: cross Abstract: Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, alt

Human-Guided Harm Recovery for Computer Use Agents

Model ReleasesDGX agent

arXiv:2604.18847v1 Announce Type: new Abstract: As LM agents gain the ability to execute actions on real computer systems, we need ways to not only prevent harmful actions at scale but also effectivel

Human-Machine Co-Boosted Bug Report Identification with Mutualistic Neural Active Learning

ApplicationsDGX agent

arXiv:2604.18862v1 Announce Type: cross Abstract: Bug reports, encompassing a wide range of bug types, are crucial for maintaining software quality. However, the increasing complexity and volume of bu

Idle is the New Sleep: Configuration-Aware Alternative to Powering Off FPGA-Based DL Accelerators During Inactivity

ResearchDGX agent

arXiv:2407.12027v2 Announce Type: replace-cross Abstract: In the rapidly evolving Internet of Things (IoT) domain, we concentrate on enhancing energy efficiency in Deep Learning accelerators on FPGA-b

Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI

ResearchDGX agent

arXiv:2604.19578v1 Announce Type: cross Abstract: With the rapid advancement of Large Language Models (LLMs), the academic community has faced unprecedented disruptions, particularly in the realm of a

Improved Anomaly Detection in Medical Images via Mean Shift Density Enhancement

ResearchDGX agent

arXiv:2604.19191v1 Announce Type: cross Abstract: Anomaly detection in medical imaging is essential for identifying rare pathological conditions, particularly when annotated abnormal samples are limit

IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text

Model ReleasesDGX agent

arXiv:2604.19298v1 Announce Type: cross Abstract: We introduce IndiaFinBench, to our knowledge the first publicly available evaluation benchmark for assessing large language model (LLM) performance on

Inductive Subgraphs as Shortcuts: Causal Disentanglement for Heterophilic Graph Learning

Local AiDGX agent

arXiv:2604.19186v1 Announce Type: cross Abstract: Heterophily is a prevalent property of real-world graphs and is well known to impair the performance of homophilic Graph Neural Networks (GNNs). Prior

Industrial Surface Defect Detection via Diffusion Generation and Asymmetric Student-Teacher Network

Local AiDGX agent

arXiv:2604.19240v1 Announce Type: new Abstract: Industrial surface defect detection often suffers from limited defect samples, severe long-tailed distributions, and difficulties in accurately localizi

InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation

Model ReleasesDGX agent

arXiv:2509.21080v2 Announce Type: replace-cross Abstract: Advancements in Large language models (LLMs) have enabled a variety of downstream applications like story and interview script generation. How

Integrating Anomaly Detection into Agentic AI for Proactive Risk Management in Human Activity

SafetyDGX agent

arXiv:2604.19538v1 Announce Type: new Abstract: Agentic AI, with goal-directed, proactive, and autonomous decision-making capabilities, offers a compelling opportunity to address movement-related risk

Intentional Updates for Streaming Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.19033v1 Announce Type: cross Abstract: In gradient-based learning, a step size chosen in parameter units does not produce a predictable per-step change in function output. This often leads

Investigating the structure of emotions by analyzing similarity and association of emotion words

ResearchDGX agent

arXiv:2602.06430v2 Announce Type: replace-cross Abstract: In the field of natural language processing, some studies have attempted sentiment analysis on text by handling emotions as explanatory or res

Knowledge-Guided Time-Varying Causal Inference for Arctic Sea Ice Dynamics

SafetyDGX agent

arXiv:2601.17647v2 Announce Type: replace-cross Abstract: Quantifying the causal relationship between sea ice thickness and sea surface height (SSH) is essential for understanding the mechanisms drivi

Large Language Models Exhibit Normative Conformity

SafetyDGX agent

arXiv:2604.19301v1 Announce Type: new Abstract: The conformity bias exhibited by large language models (LLMs) can pose a significant challenge to decision-making in LLM-based multi-agent systems (LLM-

LASER: Learning Active Sensing for Continuum Field Reconstruction

SafetyDGX agent

arXiv:2604.19355v1 Announce Type: cross Abstract: High-fidelity measurements of continuum physical fields are essential for scientific discovery and engineering design but remain challenging under spa

LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation

HardwareDGX agent

arXiv:2604.19167v1 Announce Type: cross Abstract: Deploying large language models (LLMs) in resource-constrained environments is hindered by heavy computational and memory requirements. We present LBL

Learning Evolution via Optimization Knowledge Adaptation

ResearchDGX agent

arXiv:2501.02200v2 Announce Type: replace-cross Abstract: The iterative search process of evolutionary algorithms (EAs) encapsulates optimization knowledge within historical populations and fitness ev

Learning Hybrid-Control Policies for High-Precision In-Contact Manipulation Under Uncertainty

SafetyDGX agent

arXiv:2604.19677v1 Announce Type: cross Abstract: Reinforcement learning-based control policies have been frequently demonstrated to be more effective than analytical techniques for many manipulation

Learning Lifted Action Models from Unsupervised Visual Traces

TutorialsDGX agent

arXiv:2604.19043v1 Announce Type: new Abstract: Efficient construction of models capturing the preconditions and effects of actions is essential for applying AI planning in real-world domains. Extensi

LePREC: Reasoning as Classification over Structured Factors for Assessing Relevance of Legal Issues

Model ReleasesDGX agent

arXiv:2604.19464v1 Announce Type: cross Abstract: More than half of the global population struggles to meet their civil justice needs due to limited legal resources. While Large Language Models (LLMs)

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.18803v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed in settings where reliable visual grounding carries operational consequences, yet their behavi

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control

Local AiDGX agent

arXiv:2604.19018v1 Announce Type: cross Abstract: Inference-time LLM alignment methods, particularly activation steering, offer an alternative to fine-tuning by directly modifying activations during g

Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs

SafetyDGX agent

arXiv:2604.19292v1 Announce Type: cross Abstract: Multilingual large language models (LLMs) have minimized the fluency gap between languages. This advancement, however, exposes models to the risk of b

← Previous
1…317318319320321…354
Next →