AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
27 May 2026

Reliable Extraction of Clinical Follow-Up Instructions: A Hybrid Neural-Symbolic Pipeline

Model ReleasesDGX agent

arXiv:2605.26560v1 Announce Type: cross Abstract: Objective. Outpatient notes carry follow-up instructions pairing actions with future times ('MRI brain in two weeks'). Extracting (action, date) pairs

ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference

Model ReleasesDGX agent

arXiv:2605.27081v1 Announce Type: cross Abstract: Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining hi

RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.26177v1 Announce Type: cross Abstract: Code agents are currently having skillful performance on repository-level software engineering benchmarks, but it remains unclear whether success on e

Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models

ResearchDGX agent

arXiv:2605.26661v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection has emerged as a popular technique to enhance the reliability of machine learning models by identifying unexpected

Rethinking the Trust Region in LLM Reinforcement Learning

SafetyDGX agent

arXiv:2602.04879v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) ser

Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective

SafetyDGX agent

arXiv:2605.26441v1 Announce Type: cross Abstract: This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposa

ReVEL: Multi-Turn Reflective LLM-Guided Heuristic Evolution via Structured Performance Feedback

Local AiDGX agent

arXiv:2604.04940v2 Announce Type: replace Abstract: Designing effective heuristics for NP-hard combinatorial optimization problems remains challenging and often requires substantial domain expertise.

Risk Averse Alert Prioritization for IDS Using Subnormal Gaussian Fuzzy Models

Model ReleasesDGX agent

arXiv:2605.27299v1 Announce Type: cross Abstract: Modern intrusion detection systems generate thousands of alerts daily, but alert fatigue severely limits security operations effectiveness due to too

Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompting Attacks

ApplicationsDGX agent

arXiv:2506.03627v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable performance across various tasks by effectively utilizing a prompting strategy. Howe

Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling

ResearchDGX agent

arXiv:2605.26702v1 Announce Type: cross Abstract: Reliable watermarking of panoramic imagery is fundamentally challenged by arbitrary 3D rotations. As panoramas are defined on the sphere, they natural

RulePlanner: All-in-One Reinforcement Learner for Unifying Design Rules in 3D Floorplanning

ApplicationsDGX agent

arXiv:2601.22476v2 Announce Type: replace-cross Abstract: Floorplanning determines the coordinate and shape of each module in Integrated Circuits. With the scaling of technology nodes, in floorplannin

Scalable GANs with Transformers

ResearchDGX agent

arXiv:2509.24935v2 Announce Type: replace-cross Abstract: Scalability has driven recent advances in generative modeling, yet its principles remain underexplored for adversarial learning. We investigat

Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

Model ReleasesDGX agent

arXiv:2605.27134v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking,

Scaling GraphLLM with Bilevel-Optimized Sparse Querying

Model ReleasesDGX agent

arXiv:2602.09038v2 Announce Type: replace-cross Abstract: LLMs have recently shown strong potential in enhancing node-level tasks on text-attributed graphs (TAGs) by providing explanation features. Ho

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

Model ReleasesDGX agent

arXiv:2605.26340v1 Announce Type: new Abstract: Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetect

SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge Graphs

Model ReleasesDGX agent

arXiv:2512.04868v2 Announce Type: replace-cross Abstract: Knowledge-based conversational question answering (KBCQA) confronts persistent challenges in resolving coreference, modeling contextual depend

Searching the Internet for Challenging Benchmarks at Scale

ResearchDGX agent

arXiv:2509.26619v3 Announce Type: replace-cross Abstract: Many static benchmarks are beginning to saturate: as models rapidly improve, they achieve near-perfect scores on fixed test sets, leaving litt

Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation

SafetyDGX agent

arXiv:2510.19420v2 Announce Type: replace-cross Abstract: Multi-Agent Systems (MAS) have become a prevalent paradigm for Large Language Model (LLM) applications. However, the complex multi-agent desig

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability

Model ReleasesDGX agent

arXiv:2605.26788v1 Announce Type: cross Abstract: Large language models (LLMs) achieve impressive performance when a task is fully specified in a single turn, yet the same models lose up to 39% of tha

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

Model ReleasesDGX agent

arXiv:2601.07737v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in mainstream visual understanding tasks, but their ability

Self-Cascaded Diffusion Models for Arbitrary-Scale Image Super-Resolution

ResearchDGX agent

arXiv:2506.07813v2 Announce Type: replace-cross Abstract: Arbitrary-scale image super-resolution aims to upsample images to any desired resolution, offering greater flexibility than traditional fixed-

Self-Improvement Imitation with Biologically Guided Search for Protein Design Under Oracle Budgets

SafetyDGX agent

arXiv:2605.26690v1 Announce Type: cross Abstract: Protein sequence optimization under tight oracle budgets requires methods that explore vast combinatorial spaces while making each evaluation informat

Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning

AgentsDGX agent

arXiv:2510.06843v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have exhibited impressive capabilities across diverse application domains. Recent work has explored Multi-LLM Age

Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection

SafetyDGX agent

arXiv:2605.27155v1 Announce Type: cross Abstract: Testing object detectors in safety-critical domains requires semantically meaningful probes beyond pixel-level corruptions. We present SemProbe, a too

Semigroup Consistency as a Diagnostic for Learned Physics Simulators

AgentsDGX agent

arXiv:2605.26324v1 Announce Type: cross Abstract: Learned physics simulators are often evaluated by one-step or short-horizon prediction error, but these metrics can miss failures in temporal composit

SetupX: Can LLM Agents Learn from Past Failures in Functionality-Correct Code Repository Setup?

AgentsDGX agent

arXiv:2605.26186v1 Announce Type: cross Abstract: Functionality-correct repository setup aims to configure execution environments (e.g., dependencies, build scripts) to successfully execute a reposito

Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs

ResearchDGX agent

arXiv:2601.04275v2 Announce Type: replace-cross Abstract: Machine unlearning aims to selectively remove the influence of specific training samples to satisfy privacy regulations such as the GDPR's 'Ri

SIA: Self Improving AI with Harness & Weight Updates

HardwareDGX agent

arXiv:2605.27276v1 Announce Type: new Abstract: Humans are the bottleneck in building and improving AI. Both the models and the agents that wrap them are written, tuned, and corrected by people. The l

SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation

SafetyDGX agent

arXiv:2605.26704v1 Announce Type: cross Abstract: Epidemic forecasting faces a fundamental challenge: human behavior dynamically responds to disease spread, creating feedback loops that induce distrib

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training

SafetyDGX agent

arXiv:2605.26606v1 Announce Type: cross Abstract: Reinforcement learning (RL) is the dominant paradigm for post-training large language models. However, in the online, on-policy setting, rollout gener

Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context Prefilling

Model ReleasesDGX agent

arXiv:2604.18103v2 Announce Type: replace Abstract: Prefilling computational costs pose a significant bottleneck for Large Language Models (LLMs) and Large Multimodal Models (LMMs) in long-context set

Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models

ResearchDGX agent

arXiv:2605.26733v1 Announce Type: cross Abstract: Looped Language Models (LoopLMs) enable efficient latent reasoning through depth recurrence, yet exhibit unreliable test-time scaling behavior: perfor

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.27140v1 Announce Type: new Abstract: Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hin

Strategies for Guiding LLMs to Use Software Design Patterns: A Case of Singleton

Model ReleasesDGX agent

arXiv:2605.26898v1 Announce Type: cross Abstract: Large Language Models (LLMs) can generate functional source code from natural-language prompts, but often fail to consistently follow higher-level arc

StreamSplit: Continuous Audio Representation Learning via Uncertainty-Guided Adaptive Splitting

Local AiDGX agent

arXiv:2605.26523v1 Announce Type: cross Abstract: Large-batch Contrastive Learning (CL), the foundation of modern representation learning, is fundamentally incompatible with the volatile resource cons

Structure-Adaptive Conformal Inference for Large-Scale Out-of-Distribution Testing

ResearchDGX agent

arXiv:2605.26429v1 Announce Type: cross Abstract: This paper addresses structured out-of-distribution (OOD) testing in high-stakes machine learning applications. Traditional conformal methods rely on

SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking

ResearchDGX agent

arXiv:2511.04711v2 Announce Type: replace-cross Abstract: Large-scale vision-language models, especially CLIP, have demonstrated remarkable performance across diverse downstream tasks. Soft prompts, a

TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews

Model ReleasesDGX agent

arXiv:2605.26911v1 Announce Type: new Abstract: LLM-generated peer reviews are increasingly common at major venues, yet their deficiencies are hard to detect because they are uniformly fluent and well

Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2

ResearchDGX agent

arXiv:2605.26628v1 Announce Type: new Abstract: This report describes Tail-Aware HiFloat4, our submission to the low-bit text-to-video generation quantization challenge. Our method adapts the public V

Targeted Remasking: Replacing Token Editing with Token-to-Mask Refinement in Discrete Diffusion Language Models

ResearchDGX agent

arXiv:2605.26436v1 Announce Type: cross Abstract: Discrete masked diffusion language models such as LLaDA generate text through iterative denoising, where mask tokens are progressively replaced with p

The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

SafetyDGX agent

arXiv:2601.07085v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding mi

The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context

ResearchDGX agent

arXiv:2605.26778v1 Announce Type: new Abstract: Retrieval-augmented generation promises to ground language model outputs in external evidence, yet the field has no reliable way to verify whether retri

The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?

Model ReleasesDGX agent

arXiv:2605.27176v1 Announce Type: new Abstract: Knowledge graphs (KGs) can provide structured scientific context to language models, but it remains unclear which graph facts actually shape the generat

The Kalman Evolve: Closing the Gap in Kalman Filtering via Interpretable Algorithm Discovery

ApplicationsDGX agent

arXiv:2605.26830v1 Announce Type: cross Abstract: State estimation is a fundamental problem in control and signal processing, for which the Kalman Filter provides an optimal solution under linear dyna

The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models

SafetyDGX agent

arXiv:2605.26670v1 Announce Type: cross Abstract: Sequential editing of structured knowledge in large language models allows targeted factual updates without retraining, yet existing methods often rel

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

AgentsDGX agent

arXiv:2605.26494v1 Announce Type: new Abstract: We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum

The Necessity of a Unified Framework for LLM-Based Agent Evaluation

AgentsDGX agent

arXiv:2602.03238v2 Announce Type: replace Abstract: With the advent of Large Language Models (LLMs), general-purpose agents have seen fundamental advancements. However, evaluating these agents present

The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP

SafetyDGX agent

arXiv:2605.26415v1 Announce Type: cross Abstract: Deploying Vision-Language Models on resource-constrained hardware typically requires INT8 quantization, but in joint-embedding architectures such as C

The Sensation Modulating Network:Haltability as the architectural ground for object-directed phenomenology

AgentsDGX agent

arXiv:2605.26856v1 Announce Type: cross Abstract: Cognitive science remains split between cognitivism - which accounts for recursion and language but cannot ground formal symbols in meaning - and 4E a

The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection

AgentsDGX agent

arXiv:2605.26872v1 Announce Type: cross Abstract: LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current p

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

SafetyDGX agent

arXiv:2505.11063v3 Announce Type: replace Abstract: LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly sh

Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V

Model ReleasesDGX agent

arXiv:2605.27003v1 Announce Type: cross Abstract: W4A4 quantization of large video diffusion Transformers offers substantial memory savings but is hindered by two main challenges: sparse large-magnitu

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets

Model ReleasesDGX agent

arXiv:2605.26165v1 Announce Type: cross Abstract: Agentic RAG systems that equip language models with dozens to hundreds of tool definitions face a critical resource conflict: tool schemas consume the

Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation

ResearchDGX agent

arXiv:2605.26958v1 Announce Type: cross Abstract: Reinforcement learning in open-ended long-form generation is challenging because reliable reference answers and automatic metrics are often unavailabl

Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records

Model ReleasesDGX agent

arXiv:2605.26463v1 Announce Type: cross Abstract: Data consistency between unstructured clinical notes and structured tables in Electronic Health Records (EHRs) is essential for patient safety and cli

Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

HardwareDGX agent

arXiv:2605.26720v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong empirical gains as self-evolving agents for CUDA kernel generation, driven by feedback-conditioned planni

Towards Generalization-Oriented Models for Vehicle Routing Problems with Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2605.26776v1 Announce Type: cross Abstract: In recent years, Deep Reinforcement Learning (DRL) has achieved substantial progress on Vehicle Routing Problems (VRPs). However, existing DRL-based m

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

Model ReleasesDGX agent

arXiv:2601.05899v2 Announce Type: replace Abstract: Recent breakthroughs in Large Language Models (LLMs) have positioned them as a promising paradigm for agents, with long-term planning and decision-m

Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry

Model ReleasesDGX agent

arXiv:2605.27071v1 Announce Type: new Abstract: Key knowledge for steel-industry volatile organic compounds (VOCs) governance is scattered across unstructured scientific literature, making it difficul

Tracing Computation Density in LLMs

ResearchDGX agent

arXiv:2605.27033v1 Announce Type: cross Abstract: Transformer-based large language models (LLMs) are comprised of billions of parameters arranged in deep and wide computational graphs, but it is not c

← Previous
1…208209210211212…358
Next →