AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
12 May 2026

DataMaster: Towards Autonomous Data Engineering for Machine Learning

AgentsDGX agent

arXiv:2605.10906v1 Announce Type: cross Abstract: As model families, training recipes, and compute budgets become increasingly standardized, further gains in machine learning systems depend increasing

Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents

SafetyDGX agent

arXiv:2605.08717v1 Announce Type: cross Abstract: Software engineering agents are increasingly deployed in evaluable engineering environments, yet post-failure recovery remains costly, manual, and ad

Decoupling Endpoint and Semantic Transition Learning for Zero-Shot Composed Image Retrieval

SafetyDGX agent

arXiv:2605.08389v1 Announce Type: cross Abstract: Zero-shot composed image retrieval (ZS-CIR) retrieves a target image from a reference image and a text modification without human-annotated CIR triple


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Deep Arguing

TutorialsDGX agent

arXiv:2605.10569v1 Announce Type: new Abstract: Deep learning has become the dominant approach for creating high capacity, scalable models across diverse data modalities. However, because these models

DeepRefine: Agent-Compiled Knowledge Refinement via Reinforcement Learning

AgentsDGX agent

arXiv:2605.10488v1 Announce Type: cross Abstract: Agent-compiled knowledge bases provide persistent external knowledge for large language model (LLM) agents in open-ended, knowledge-intensive downstre

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

Model ReleasesDGX agent

arXiv:2605.09679v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, e

Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents

ResearchDGX agent

arXiv:2605.08442v1 Announce Type: cross Abstract: Persistent memory attacks against LLM agents achieve high attack success rates against open-source models. In these attacks, malicious instructions in

Detect, Localize, and Explain: Interactive Hierarchical Log Anomaly Analytics with LLM Augmentation

Model ReleasesDGX agent

arXiv:2605.09222v1 Announce Type: cross Abstract: Logs are ubiquitous in modern systems. Unfortunately, their unstructured nature in flat sequences limits understanding of execution behaviors, hinderi

Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

Model ReleasesDGX agent

arXiv:2604.01151v2 Announce Type: replace Abstract: As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human o

Deterministic Decomposition of Stochastic Generative Dynamics

ResearchDGX agent

arXiv:2605.08794v1 Announce Type: cross Abstract: Modern generative models can be understood as probability transport from a simple base distribution to a target data distribution. Deterministic trans

Developing a foundation model for high-resolution remote sensing data of the Netherlands

TutorialsDGX agent

arXiv:2605.10184v1 Announce Type: cross Abstract: We develop a foundation model using 1.2m high resolution satellite images of the Netherlands. By combining a Convolutional Neural Network and a Vision

Diagnosing Spectral Ceilings in Equivariant Neural Force Fields

Model ReleasesDGX agent

arXiv:2605.08286v1 Announce Type: cross Abstract: We introduce a spectral-injection diagnostic for measuring which angular frequencies a trained equivariant force-field backbone preserves: inject a co

DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules

Model ReleasesDGX agent

arXiv:2605.08614v1 Announce Type: new Abstract: Monitoring complex industrial assets relies on engineer-authored symbolic rules that trigger based on sensor conditions and prompt technicians to perfor

diffGHOST: Diffusion based Generative Hedged Oblivious Synthetic Trajectories

TutorialsDGX agent

arXiv:2605.10647v1 Announce Type: new Abstract: Trajectories are nowadays valuable information for a wide range of applications. However they are also inherently sensitive, as they contain highly pers

Digital Image Forgery Detection Using Transfer Learning

ApplicationsDGX agent

arXiv:2605.08167v1 Announce Type: cross Abstract: The increasing availability of advanced image editing tools has led to a significant rise in manipulated digital content, posing serious challenges fo

Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT

ApplicationsDGX agent

arXiv:2605.09719v1 Announce Type: cross Abstract: Large-scale 3D vision-language models (VLMs) like LLaVA-3D offer strong spatial reasoning but are difficult to deploy due to high computational costs.

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

Model ReleasesDGX agent

arXiv:2605.08462v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge in Large Language Models (LLMs), particularly in context-grounded settings such as RAG and agentic AI sys

Do Linear Probes Generalize Better in Persona Coordinates?

SafetyDGX agent

arXiv:2605.09391v1 Announce Type: new Abstract: It is becoming increasingly necessary to have monitors check for harmful behaviors during language model interactions, but text-only monitoring has not

Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas

TutorialsDGX agent

arXiv:2605.09159v1 Announce Type: new Abstract: Recent work shows that large language models (LLMs) encode behavioural traits ('personas') as linear directions in activation space, often called 'perso

Do multimodal models imagine electric sheep?

ResearchDGX agent

arXiv:2605.09693v1 Announce Type: cross Abstract: Yes. We find that large multimodal models develop mental imagery when solving spatial puzzles, and they do imagine sheep when solving sheep puzzles. W

Do not copy and paste! Rewriting strategies for code retrieval

Model ReleasesDGX agent

arXiv:2605.08299v1 Announce Type: cross Abstract: Embedding-based code retrieval often suffers when encoders overfit to surface syntax. Prior work mitigates this by using LLMs to rephrase queries and

Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation

Model ReleasesDGX agent

arXiv:2605.09315v1 Announce Type: new Abstract: Recent advances in LLM agents enable systems that autonomously refine workflows, accumulate reusable skills, self-train their underlying models, and mai

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.08747v1 Announce Type: new Abstract: Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call te

Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces

Model ReleasesDGX agent

arXiv:2605.09497v1 Announce Type: new Abstract: Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Ex

DP-LAC: Lightweight Adaptive Clipping for Differentially Private Federated Fine-tuning of Language Models

Local AiDGX agent

arXiv:2605.10272v1 Announce Type: cross Abstract: Federated learning (FL) enables the collaborative training of large-scale language models (LLMs) across edge devices while keeping user data on-device

Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs

SafetyDGX agent

arXiv:2605.10281v1 Announce Type: cross Abstract: Generating realistic drum audio directly from symbolic representations is a challenging task at the intersection of music perception and machine learn

Dsat: A Native SAT Solver for Discrete Logic

ResearchDGX agent

arXiv:2605.09347v1 Announce Type: new Abstract: Discrete variables are common in many applications, such as probabilistic reasoning, planning and explainable AI. When symbolic reasoning techniques are

DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments

Model ReleasesDGX agent

arXiv:2503.06047v2 Announce Type: replace Abstract: Large language model (LLM)-based agents are increasingly applied to complex strategic environments that demand long-horizon reasoning, multi-agent i

DUALFloodGNN: Physics-informed Graph Neural Network for Operational Flood Modeling

Local AiDGX agent

arXiv:2512.23964v2 Announce Type: replace-cross Abstract: Flood models inform strategic disaster management by simulating the spatiotemporal hydrodynamics of flooding. While physics-based numerical fl

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards

Model ReleasesDGX agent

arXiv:2605.08441v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) generates hundreds of thousands of tokens per training step, with rollout generation dominating

DuetFair: Coupling Inter- and Intra-Subgroup Robustness for Fair Medical Image Segmentation

SafetyDGX agent

arXiv:2605.10521v1 Announce Type: cross Abstract: Medical image segmentation models can perform unevenly across subgroups. Most existing fairness methods focus on improving average subgroup performanc

Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning

ApplicationsDGX agent

arXiv:2605.10765v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, yet real-world deployment often requires continual cap

Dynamic Linear Coregionalization for Realistic Synthetic Multivariate Time Series

ResearchDGX agent

arXiv:2604.05064v2 Announce Type: replace-cross Abstract: Synthetic data is essential for training foundation models for time series (FMTS), but most generators assume static correlations, and are typ

Dynamics-Aligned Shared Hypernetworks for Contextual RL under Discontinuous Shifts

Model ReleasesDGX agent

arXiv:2602.06550v2 Announce Type: replace-cross Abstract: Zero-shot generalization in contextual reinforcement learning remains a core challenge, particularly when the context is latent and must be in

DynGhost: Temporally-Modelled Transformer for Dynamic Ghost Imaging with Quantum Detectors

ResearchDGX agent

arXiv:2605.10185v1 Announce Type: cross Abstract: Ghost imaging reconstructs spatial information from a single-pixel bucket detector by correlating structured illumination patterns with scalar intensi

E-TCAV: Formalizing Penultimate Proxies for Efficient Concept Based Interpretability

SafetyDGX agent

arXiv:2605.10261v1 Announce Type: new Abstract: TCAV (Testing with Concept Activation Vectors) is an interpretability method that assesses the alignment between the internal representations of a train

Echo-LoRA: Parameter-Efficient Fine-Tuning via Cross-Layer Representation Injection

Model ReleasesDGX agent

arXiv:2605.08177v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) has become a practical route for adapting large language models to downstream tasks, with LoRA-style methods be

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection

ApplicationsDGX agent

arXiv:2510.19414v2 Announce Type: replace-cross Abstract: The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and ident

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies

Model ReleasesDGX agent

arXiv:2602.09514v3 Announce Type: replace-cross Abstract: Long-horizon planning is widely recognized as a core capability of autonomous LLM-based agents; however, current evaluation frameworks suffer

EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation

Model ReleasesDGX agent

arXiv:2605.09378v1 Announce Type: cross Abstract: Long-horizon video generation has advanced in visual quality, yet existing methods still struggle to maintain knowledge consistency and coherent pedag

Effective Explanations Support Planning Under Uncertainty

SafetyDGX agent

arXiv:2605.08406v1 Announce Type: cross Abstract: Explaining how to get from A to B can be challenging. It requires mentally simulating what the listener will do based on what they are told. To captur

Efficient Ensemble Selection from Binary and Pairwise Feedback

Model ReleasesDGX agent

arXiv:2605.09588v1 Announce Type: cross Abstract: Organizations increasingly deploy multiple AI systems across task domains, but selecting a small, high-performing ensemble can require costly model ca

Efficient Estimation of Kernel Surrogate Models for Task Attribution

TutorialsDGX agent

arXiv:2602.03783v2 Announce Type: replace-cross Abstract: Modern AI agents such as large language models are trained on diverse tasks -- translation, code generation, mathematical reasoning, and text

Efficient LLM Collaboration via Planning

Local AiDGX agent

arXiv:2506.11578v4 Announce Type: replace Abstract: Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achie

Efficient Prompt Learning for Traffic Forecasting

ApplicationsDGX agent

arXiv:2605.08273v1 Announce Type: cross Abstract: Accurate traffic prediction is essential for optimizing transportation systems, enhancing resource allocation, and improving overall urban administrat

EGL-SCA: Structural Credit Assignment for Co-Evolving Instructions and Tools in Graph Reasoning Agents

SafetyDGX agent

arXiv:2605.10366v1 Announce Type: new Abstract: Graph reasoning agents operating from natural-language inputs must solve a coupled problem: they must reconstruct a structured graph instance from text,

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

Model ReleasesDGX agent

arXiv:2605.09874v1 Announce Type: cross Abstract: Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more

Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts

ApplicationsDGX agent

arXiv:2509.21892v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models typically fix the number of activated experts k at both training and inference. However, real-world deployment

ELF: Embedded Language Flows

ResearchDGX agent

arXiv:2605.10938v1 Announce Type: cross Abstract: Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their

Elite Polarization in European Parliamentary Speeches: a Novel Measurement Approach Using Large Language Models

ApplicationsDGX agent

arXiv:2507.06658v2 Announce Type: replace-cross Abstract: Theories of democratic stability, populism, and party-system crisis often point to a form of polarization that comparative research rarely mea

Embeddings for Preferences, Not Semantics

ResearchDGX agent

arXiv:2605.08360v1 Announce Type: new Abstract: Modern AI is opening the door to collective decision-making in which participants express their views as free-form text rather than voting on a fixed se

EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents

Model ReleasesDGX agent

arXiv:2605.10332v1 Announce Type: new Abstract: Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since embodied enviro

Emergence of Physical Intelligence via Controllable Information Production

SafetyDGX agent

arXiv:2601.22449v2 Announce Type: replace Abstract: Intrinsic Motivation (IM) aims to train agents without external rewards, enabling useful behavior to emerge from the agent's interaction with its en

Emergent Semantic Role Understanding in Language Models

TutorialsDGX agent

arXiv:2605.09187v1 Announce Type: new Abstract: Understanding how linguistic structure emerges in language models is central to interpreting what these systems learn from data and how much supervision

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning

SafetyDGX agent

arXiv:2605.09395v1 Announce Type: new Abstract: In this paper, we propose the first VLnderline{extbf{M}} nderline{extbf{a}}gentic nderline{extbf{r}}easoning framework for few-nderline{extbf{s}}hot mul

Empty SPACE: Cross-Attention Sparsity for Concept Erasure in Diffusion Models

ResearchDGX agent

arXiv:2605.10198v1 Announce Type: cross Abstract: Erasing specific concepts from text-to-image diffusion models is essential for avoiding the generation of copyrighted and explicit content. Closed-for

EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.09826v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to track others epistemic state, makes humans efficient collaborators. AI agents need the same capacity in multi agent

enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways

ApplicationsDGX agent

arXiv:2604.16838v2 Announce Type: replace-cross Abstract: We present enclawed, a hard-fork hardening framework built on the OpenClaw AI assistant gateway. enclawed targets deployments that need attest

Energy Consumption of Dataframe Libraries for End-to-End Deep Learning Pipelines:A Comparative Analysis

HardwareDGX agent

arXiv:2511.08644v3 Announce Type: replace-cross Abstract: This paper presents a detailed comparative analysis of the performance of three major Python data manipulation libraries - Pandas, Polars, and

Engineering Robustness into Personal Agents with the AI Workflow Store

AgentsDGX agent

arXiv:2605.10907v1 Announce Type: cross Abstract: The dominant paradigm for AI agents is an 'on-the-fly' loop in which agents synthesize plans and execute actions within seconds or minutes in response

← Previous
1…265266267268269…358
Next →