AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
26 May 2026

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

Model ReleasesDGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

Beyond Generative Priors: Minority Sampling with JEPA-Guided Diffusion

ApplicationsDGX agent

arXiv:2605.24631v1 Announce Type: cross Abstract: Minority sampling aims to generate low-density instances on a data manifold and is of central importance in applications such as medical diagnosis, an

Beyond Inference-Only Deployment: Comparing Weight-Based Consolidation Against Cascading Compaction

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.24657v1 Announce Type: new Abstract: Major LLM platforms deploy models in an inference-only configuration: the model serves requests but never updates per-user weights. Users must repeatedl

Beyond Killer Robots: General AI Attitudes and Public Support for Military AI in Nine Countries

SafetyDGX agent

arXiv:2605.25196v1 Announce Type: cross Abstract: AI-enabled military systems are a fixture of modern military conflict. Applications vary from autonomous drones for surveillance and attack to AI-supp

Beyond Predefined Learning Objects: A Thinking-Learning Interaction Model for Up-to-Date Autonomous Robot Learning

AgentsDGX agent

arXiv:2605.23987v1 Announce Type: new Abstract: Autonomous robots operating in open and changing environments cannot always rely on predefined inputs, outputs, and action routines. Although existing l

Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching

Model ReleasesDGX agent

arXiv:2605.25558v1 Announce Type: new Abstract: Optimizing the trade-off among predictive performance and computational cost is a central focus in the deployment of Large Language Models (LLMs). Curre

Beyond Static Uncertainty: Modeling Temporal Uncertainty Dynamics for Probabilistic Time Series Forecasting

ApplicationsDGX agent

arXiv:2603.24254v2 Announce Type: replace-cross Abstract: Real-world time series exhibit temporally structured uncertainty: volatility clusters in turbulent regimes, dissipates in stable periods, and

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models

Model ReleasesDGX agent

arXiv:2605.26100v1 Announce Type: cross Abstract: Code review is a critical practice in software engineering, yet the growing scale and frequency of code patches in modern projects, together with the

Beyond the Aggregation Dilemma: Prior-Retaining Decoupled Learning for Multimodal Graphs

HardwareDGX agent

arXiv:2605.24684v1 Announce Type: cross Abstract: Multimodal Attributed Graph Learning (MAGL) integrates intrinsic node attributes with structural topology via graph aggregation. However, as pretraine

Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling

ResearchDGX agent

arXiv:2605.25143v1 Announce Type: new Abstract: Test-time scaling improves language model reasoning by spending additional compute to explore multiple solution trajectories. The key challenge is to ma

Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training

SafetyDGX agent

arXiv:2505.20110v3 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) excel at sampling diverse, high-reward objects. In many practical applications where active reward querie

Bilevel Optimization of Synthetic Trajectories for Multi-Turn LLM Fine-Tuning

ResearchDGX agent

arXiv:2605.24743v1 Announce Type: cross Abstract: While LLMs excel at single-turn generation, they struggle with long-horizon, multi-turn interactions. Offline reinforcement learning (RL) offers a sca

Binding Visual Features Point by Point

ResearchDGX agent

arXiv:2605.25427v1 Announce Type: cross Abstract: Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, inclu

BODHI: Precise OS Kernel Specification Inference

Model ReleasesDGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

Boosting Inference with Guided Reasoning: Stochastic Exploration for Recursive Models

Local AiDGX agent

arXiv:2605.25230v1 Announce Type: new Abstract: Recent work on recursive architectures has shown that tiny neural networks can be surprisingly powerful on structured reasoning tasks. The trick is to m

BoxLitE: A Faithful Knowledge Base Embedding Based on Convex Optimization

TutorialsDGX agent

arXiv:2605.23937v1 Announce Type: new Abstract: Knowledge base (KB) embeddings aim at combining the capability of classical knowledge graph embeddings to generalize the information present in facts, t

Breaking the Chains of Probability: Neutrosophic Logic as a New Framework for Epistemic Uncertainty in Large Language Models

ResearchDGX agent

arXiv:2605.24053v1 Announce Type: new Abstract: Large Language Models (LLMs) are predominantly governed by probabilistic frameworks in which the sum of outcome probabilities is constrained to unity. T

Bridging Evolutionary Algorithms and Reinforcement Learning: A Comprehensive Survey on Hybrid Algorithms

ResearchDGX agent

arXiv:2401.11963v5 Announce Type: replace-cross Abstract: Evolutionary Reinforcement Learning (ERL), which integrates Evolutionary Algorithms (EAs) and Reinforcement Learning (RL) for optimization, ha

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion

SafetyDGX agent

arXiv:2605.24975v1 Announce Type: cross Abstract: Proximal Policy Optimization (PPO) has become the de facto standard for training legged robots, thanks to its robustness and scalability in massively

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference

ResearchDGX agent

arXiv:2511.16449v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown great potential for embodied AI by integrating visual perception, language understanding, and a

By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode

ApplicationsDGX agent

arXiv:2605.25186v1 Announce Type: cross Abstract: Formalizing legal provisions promises machine-accessible law and automated legal reasoning, and recent LLMs make it tempting to generate such formaliz

Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.25920v1 Announce Type: cross Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint

CARL-CXR: Continual Adapter-Based Routing for Task-Unknown Chest Radiograph Classification

ResearchDGX agent

arXiv:2602.15811v2 Announce Type: replace-cross Abstract: Clinical deployment of chest radiograph classifiers requires models that can be updated as new datasets become available without retraining on

Cascade-KDE: Robust Time-Series Restoration under Out-of-Distribution Impulse Corruptions

Model ReleasesDGX agent

arXiv:2605.24055v1 Announce Type: cross Abstract: Real-world time-series data in industrial sensing, healthcare, and energy systems is often corrupted by a mixture of Gaussian noise and occasional lar

Catching MRI outliers: unsupervised detection and localization of MRI artefacts and clinical anomalies using deep learning

Local AiDGX agent

arXiv:2605.24609v1 Announce Type: cross Abstract: Artificial intelligence is increasingly integrated into radiotherapy workflows, yet such pipelines remain vulnerable to out-of-distribution image data

Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning

ResearchDGX agent

arXiv:2605.23925v1 Announce Type: cross Abstract: Intelligent tutoring systems increasingly provide automated feedback on student work, but robust feedback requires assessing reasoning, not only final

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express

Model ReleasesDGX agent

arXiv:2605.25891v1 Announce Type: cross Abstract: We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

Model ReleasesDGX agent

arXiv:2605.26029v1 Announce Type: new Abstract: We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates bo

CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures

AgentsDGX agent

arXiv:2605.25338v1 Announce Type: cross Abstract: Large language model (LLM) agents frequently fail on multi-step tasks involving reasoning, tool use, and environment interaction. While such failures

Certified Robustness from Approximate Gaussian Mixture Structures in Pretrained Latent Spaces

SafetyDGX agent

arXiv:2605.25352v1 Announce Type: cross Abstract: Deep learning models are vulnerable to adversarial perturbations, raising important concerns for safety-critical deployment. Empirical defenses can ac

Chain-of-Thought Hijacking

Model ReleasesDGX agent

arXiv:2510.26418v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reas

Channel-wise Vector Quantization

ResearchDGX agent

arXiv:2605.26089v1 Announce Type: cross Abstract: We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens. Unlik

ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale

Model ReleasesDGX agent

arXiv:2605.24305v1 Announce Type: cross Abstract: Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason

Characterizing Linear Alignment Across Language Models

SafetyDGX agent

arXiv:2603.18908v4 Announce Type: replace Abstract: Language models increasingly appear to learn similar representations, despite differences in training objectives, architectures, and data modalities

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference

Model ReleasesDGX agent

arXiv:2510.02361v2 Announce Type: replace-cross Abstract: Transformer-based large models excel in natural language processing and computer vision, but face severe computational inefficiencies due to t

CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities

Model ReleasesDGX agent

arXiv:2605.26036v1 Announce Type: new Abstract: Urban representation learning encodes complex urban environments into general-purpose embeddings for diverse downstream tasks and emerging urban foundat

Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation

ResearchDGX agent

arXiv:2605.25831v1 Announce Type: cross Abstract: Large language models (LLMs) define a distribution over text, which can be viewed as a probabilistic representation of uncertainty: sampling K respons

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

Model ReleasesDGX agent

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Y

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

ResearchDGX agent

arXiv:2506.17629v2 Announce Type: replace-cross Abstract: Embodied Visual Reasoning (EVR) seeks to follow complex, free-form instructions based on egocentric video, enabling semantic understanding and

Clustering as Reasoning: A k-Means Interpretation of Chain-of-Thought Graph Learning

SafetyDGX agent

arXiv:2605.24867v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has shown promise in enhancing the reasoning capabilities of large language models (LLMs) on text-attributed graphs (TA

Coarse-to-Fine Domain Incremental Learning with Attentive Distillation for Mining Footprint Segmentation in Multispectral Imagery

ResearchDGX agent

arXiv:2605.24460v1 Announce Type: cross Abstract: Automatically mapping and segmenting global mining footprints using remote sensing and deep learning is critical for monitoring the socio-environmenta

Code2UML: Agentic LLMs with context engineering for scalable software visualization

Model ReleasesDGX agent

arXiv:2605.24453v1 Announce Type: cross Abstract: Large Language Model (LLM)-based code analysis tools are adopted to automate software documentation tasks. However, the scalability of these approache

CODESKILL: Learning Self-Evolving Skills for Coding Agents

SafetyDGX agent

arXiv:2605.25430v1 Announce Type: new Abstract: Coding agents produce rich trajectories while solving software-engineering tasks. To enable agent self-evolution, these trajectories can be distilled in

CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation

Model ReleasesDGX agent

arXiv:2605.25378v1 Announce Type: cross Abstract: Customized image editing aims to equip pre-trained diffusion models with specific visual effects using limited paired data, typically via Low-Rank Ada

Committed SAE-Feature Traces for Audited-Session Substitution Detection in Hosted LLMs

Model ReleasesDGX agent

arXiv:2604.18179v2 Announce Type: replace-cross Abstract: Hosted-LLM providers have a silent-substitution incentive: advertise a stronger model while serving cheaper replies. Probe-after-return scheme

Complement Submodular Information Measures for Balanced and Robust Data Selection

Model ReleasesDGX agent

arXiv:2605.24779v1 Announce Type: cross Abstract: Submodular optimization has become a fundamental paradigm for data selection, retrieval, summarization, and representation learning due to its ability

Concept Drift Adaptation Using Self-Supervised and Reinforcement Learning In Android Malware Detection

SafetyDGX agent

arXiv:2605.24294v1 Announce Type: cross Abstract: Android malware detectors often degrade after deployment because of concept drift, while full retraining at each maintenance step is costly. We propos

Concept Unlearning via Cross-Attention Activation Projection for Diffusion Models

Model ReleasesDGX agent

arXiv:2605.25765v1 Announce Type: cross Abstract: Concept unlearning aims to erase a target concept from a pretrained text-to-image diffusion model without retraining. Closed-form methods are attracti

ConceptM^3oE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology

ApplicationsDGX agent

arXiv:2605.24399v1 Announce Type: new Abstract: Healthcare models are transitioning from unimodal prediction toward multimodal reasoning over heterogeneous diagnostic inputs. In computational patholog

Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding

ResearchDGX agent

arXiv:2605.26067v1 Announce Type: cross Abstract: Conditionally positive definite (CPD) kernels are defined with respect to a function class F. It is well known that such a kernel K is associated with

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM

Local AiDGX agent

arXiv:2605.24786v1 Announce Type: cross Abstract: Long-horizon LLM inference turns the key--value (KV) cache into the dominant GPU memory consumer and makes per-token attention increasingly expensive.

Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

ResearchDGX agent

arXiv:2605.26045v1 Announce Type: cross Abstract: Activation oracles aim to make the activations of other models legible to humans and yield promising results compared to white-box interpretability te

Confidence Calibration in Large Language Models

ResearchDGX agent

arXiv:2605.23909v1 Announce Type: new Abstract: We investigate the calibration of large language models' (LLMs') confidence across diverse tasks. The results of our preregistered study show that the c

Constraint-Anchored Attribution: Feasibility-Certified Counterfactuals and Bonferroni-PAC Sufficient Subsets for Neural CO Policies

ResearchDGX agent

arXiv:2605.25235v1 Announce Type: cross Abstract: We give an attribution method for neural combinatorial-optimisation (CO) policies that (i) decomposes a decision by constraint families via LP-relaxat

Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis

ResearchDGX agent

arXiv:2605.25354v1 Announce Type: new Abstract: While LLMs excel at reasoning over prompts using static pretrained knowledge, they struggle significantly with context learning-the ability to dynamical

Context-Instrumental Data Distillation for Kubernetes Manifest Generation: Method and Experimental Evaluation

Model ReleasesDGX agent

arXiv:2605.25835v1 Announce Type: cross Abstract: This paper examines the specialization of Small Language Models (SLMs) with up to 4 billion parameters for generating artifacts in domain-specific lan

Context: Proactive Goal-Directed Intelligence via Composable Sandboxed Programs, Declarative Wiring, and Structured Interaction

ResearchDGX agent

arXiv:2605.23928v1 Announce Type: new Abstract: We present Context, the intelligence layer of the Magarshak Architecture, which replaces reactive query-response chatbots with proactive goal-directed a

Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

SafetyDGX agent

arXiv:2602.08499v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an effective paradigm for improving the reasoning capabilities of large language mode

Continual Speaker Identity Unlearning with Minimal Interference

Model ReleasesDGX agent

arXiv:2605.25962v1 Announce Type: cross Abstract: Machine unlearning removes designated concepts or knowledge from pre-trained models. Recent work has extended this paradigm to speaker identity unlear

Continuous-Depth Field Theory for Transformer Patching and Mechanistic Interpretability

Local AiDGX agent

arXiv:2605.25225v1 Announce Type: cross Abstract: Mechanistic interpretability often uses activation patching, causal tracing, path patching, and steering directions to reveal behaviorally meaningful

← Previous
1…211212213214215…358
Next →