AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
24 Apr 2026

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

Model ReleasesDGX agent

arXiv:2604.21700v1 Announce Type: cross Abstract: The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studie

Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power

SafetyDGX agent

arXiv:2604.21043v1 Announce Type: cross Abstract: This paper examines the strategic use of language in contemporary artificial intelligence (AI) discourse, focusing on the widespread adoption of metap

Strategic Scaling of Test-Time Compute: A Bandit Learning Approach

TutorialsDGX agent

arXiv:2506.12721v2 Announce Type: replace Abstract: Scaling test-time compute has emerged as an effective strategy for improving the performance of large language models. However, existing methods typ


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Stream2LLM: Overlap Context Streaming and Prefill for Reduced Time-to-First-Token (TTFT)

HardwareDGX agent

arXiv:2604.16395v2 Announce Type: replace-cross Abstract: Context retrieval systems for LLM inference face a critical challenge: high retrieval latency creates a fundamental tension between waiting fo

StructMem: Structured Memory for Long-Horizon Behavior in LLMs

ResearchDGX agent

arXiv:2604.21748v1 Announce Type: cross Abstract: Long-term conversational agents need memory systems that capture relationships between events, not merely isolated facts, to support temporal reasonin

Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework

AgentsDGX agent

arXiv:2604.21090v1 Announce Type: cross Abstract: AI governance programmes increasingly rely on natural language prompts to constrain and direct AI agent behaviour. These prompts function as executabl

Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2603.21697v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) extend text-only LLMs with visual reasoning, but also introduce new safety failure modes under visual

StyleVAR: Controllable Image Style Transfer via Visual Autoregressive Modeling

SafetyDGX agent

arXiv:2604.21052v1 Announce Type: cross Abstract: We build on the Visual Autoregressive Modeling (VAR) framework and formulate style transfer as conditional discrete sequence modeling in a learned lat

Supervised Learning Has a Necessary Geometric Blind Spot: Theory, Consequences, and Minimal Repair

SafetyDGX agent

arXiv:2604.21395v1 Announce Type: cross Abstract: We prove that empirical risk minimisation (ERM) imposes a necessary geometric constraint on learned representations: any encoder that minimises superv

Survey on Evaluation of LLM-based Agents

SafetyDGX agent

arXiv:2503.16416v2 Announce Type: replace Abstract: LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environ

Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning

Model ReleasesDGX agent

arXiv:2604.21346v1 Announce Type: new Abstract: Vision--language models (VLMs) often fail on abstract visual reasoning benchmarks such as Bongard problems, raising the question of whether the main bot

SyMTRS: Benchmark Multi-Task Synthetic Dataset for Depth, Domain Adaptation and Super-Resolution in Aerial Imagery

Model ReleasesDGX agent

arXiv:2604.21801v1 Announce Type: cross Abstract: Recent advances in deep learning for remote sensing rely heavily on large annotated datasets, yet acquiring high-quality ground truth for geometric, r

Synthetic Data in Education: Empirical Insights from Traditional Resampling and Deep Generative Models

Model ReleasesDGX agent

arXiv:2604.21031v1 Announce Type: cross Abstract: Synthetic data generation offers promise for addressing data scarcity and privacy concerns in educational technology, yet practitioners lack empirical

TAPO-Description Logic for Information Behavior: Refined OBoxes, Inference, and Categorical Semantics

ResearchDGX agent

arXiv:2604.21172v1 Announce Type: cross Abstract: This paper develops a refined version of TAPO-description logic for the analysis of information behavior. The framework is treated not as a single hom

Task-specific Subnetwork Discovery in Reinforcement Learning for Autonomous Underwater Navigation

SafetyDGX agent

arXiv:2604.21640v1 Announce Type: cross Abstract: Autonomous underwater vehicles are required to perform multiple tasks adaptively and in an explainable manner under dynamic, uncertain conditions and

Ternary Memristive Logic: Hardware for Reasoning Realized via Domain Algebra

ResearchDGX agent

arXiv:2604.20891v1 Announce Type: cross Abstract: Memristive crossbars store numerical weights needing aggregation and decoding; a single junction means nothing alone. This paper presents a fundamenta

The AI Criminal Mastermind

AgentsDGX agent

arXiv:2604.20868v1 Announce Type: cross Abstract: In this paper, I evaluate the risks of an AI criminal mastermind, an AI agent capable of planning, coordinating, and committing a crime through the on

The CriticalSet problem: Identifying Critical Contributors in Bipartite Dependency Networks

ResearchDGX agent

arXiv:2604.21537v1 Announce Type: new Abstract: Identifying critical nodes in complex networks is a fundamental task in graph mining. Yet, methods addressing an all-or-nothing coverage mechanics in a

The Economics of p(doom): Scenarios of Existential Risk and Economic Growth in the Age of Transformative AI

SafetyDGX agent

arXiv:2503.07341v2 Announce Type: replace-cross Abstract: Recent advances in artificial intelligence (AI) have led to a wide range of predictions about its long-term impact on humanity. A central focu

The Effect of Idea Elaboration on the Automatic Assessment of Idea Originality

SafetyDGX agent

arXiv:2604.20569v1 Announce Type: cross Abstract: Automatic systems are increasingly used to assess the originality of responses in creative tasks. They offer a potential solution to key limitations o

The First Challenge on Remote Sensing Infrared Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

Model ReleasesDGX agent

arXiv:2604.21312v1 Announce Type: cross Abstract: This paper presents the NTIRE 2026 Remote Sensing Infrared Image Super-Resolution (x4) Challenge, one of the associated challenges of NTIRE 2026. The

The Last Harness You'll Ever Build

AgentsDGX agent

arXiv:2604.21003v1 Announce Type: new Abstract: AI agents are increasingly deployed on complex, domain-specific workflows -- navigating enterprise web applications that require dozens of clicks and fo

The Path Not Taken: Duality in Reasoning about Program Execution

Model ReleasesDGX agent

arXiv:2604.20917v1 Announce Type: cross Abstract: Large language models (LLMs) have shown remarkable capabilities across diverse coding tasks. However, their adoption requires a true understanding of

Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry

Model ReleasesDGX agent

arXiv:2604.20983v1 Announce Type: cross Abstract: Vision evaluations are typically done through multi-step processes. In most contemporary fields, experts analyze images using structured, evidence-bas

Thinking with Reasoning Skills: Fewer Tokens, More Accuracy

TutorialsDGX agent

arXiv:2604.21764v1 Announce Type: new Abstract: Reasoning LLMs often spend substantial tokens on long intermediate reasoning traces (e.g., chain-of-thought) when solving new problems. We propose to su

Time, Causality, and Observability Failures in Distributed AI Inference Systems

SafetyDGX agent

arXiv:2604.21361v1 Announce Type: new Abstract: Distributed AI inference pipelines rely heavily on timestamp-based observability to understand system behavior. This work demonstrates that even small c

TingIS: Real-time Risk Event Discovery from Noisy Customer Incidents at Enterprise Scale

ApplicationsDGX agent

arXiv:2604.21889v1 Announce Type: cross Abstract: Real-time detection and mitigation of technical anomalies are critical for large-scale cloud-native services, where even minutes of downtime can resul

To See the Unseen: on the Generalization Ability of Transformers in Symbolic Reasoning

Model ReleasesDGX agent

arXiv:2604.21632v1 Announce Type: new Abstract: We investigate the ability of decoder-only transformer models to perform abstract symbolic reasoning; specifically solving propositional logic reasoning

Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows

Model ReleasesDGX agent

arXiv:2604.21816v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has become a common interface for connecting large language model (LLM) agents to external tools, but its reliance on s

Towards Multimodal Active Learning: Efficient Learning with Limited Paired Data

Model ReleasesDGX agent

arXiv:2510.03247v2 Announce Type: replace-cross Abstract: Active learning (AL) is a principled strategy to reduce annotation cost in data-hungry deep learning. However, existing AL algorithms focus al

TraceScope: Interactive URL Triage via Decoupled Checklist Adjudication

SafetyDGX agent

arXiv:2604.21840v1 Announce Type: cross Abstract: Modern phishing campaigns increasingly evade snapshot-based URL classifiers using interaction gates (e.g., checkbox/slider challenges), delayed conten

Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models

Model ReleasesDGX agent

arXiv:2604.21860v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This pape

TRAVELFRAUDBENCH: A Configurable Evaluation Framework for GNN Fraud Ring Detection in Travel Networks

Model ReleasesDGX agent

arXiv:2604.21093v1 Announce Type: cross Abstract: We introduce TravelFraudBench (TFG), a configurable benchmark for evaluating graph neural networks (GNNs) on fraud ring detection in travel platform g

Trust but Verify: Introducing DAVinCI -- A Framework for Dual Attribution and Verification in Claim Inference for Language Models

ApplicationsDGX agent

arXiv:2604.21193v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable fluency and versatility across a wide range of NLP tasks, yet they remain prone to factual ina

Trust-SSL: Additive-Residual Selective Invariance for Robust Aerial Self-Supervised Learning

SafetyDGX agent

arXiv:2604.21349v1 Announce Type: cross Abstract: Self-supervised learning (SSL) is a standard approach for representation learning in aerial imagery. Existing methods enforce invariance between augme

Trustworthy Clinical Decision Support Using Meta-Predicates and Domain-Specific Languages

Model ReleasesDGX agent

arXiv:2604.21263v1 Announce Type: new Abstract: extbf{Background:} Regulatory frameworks for AI in healthcare, including the EU AI Act and FDA guidance on AI/ML-based medical devices, require clinical

UbiQVision: Quantifying Uncertainty in XAI for Image Recognition

ResearchDGX agent

arXiv:2512.20288v2 Announce Type: replace-cross Abstract: Recent advances in deep learning have led to its widespread adoption across diverse domains, including medical imaging. This progress is drive

Unbiased Prevalence Estimation with Multicalibrated LLMs

SafetyDGX agent

arXiv:2604.21549v1 Announce Type: new Abstract: Estimating the prevalence of a category in a population using imperfect measurement devices (diagnostic tests, classifiers, or large language models) is

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning

Model ReleasesDGX agent

arXiv:2604.21327v1 Announce Type: cross Abstract: Test-time reinforcement learning (TTRL) always adapts models at inference time via pseudo-labeling, leaving it vulnerable to spurious optimization sig

Using ASP(Q) to Handle Inconsistent Prioritized Data

ResearchDGX agent

arXiv:2604.21603v1 Announce Type: cross Abstract: We explore the use of answer set programming (ASP) and its extension with quantifiers, ASP(Q), for inconsistency-tolerant querying of prioritized data

Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue

Model ReleasesDGX agent

arXiv:2604.21144v1 Announce Type: cross Abstract: Situated dialogue requires speakers to maintain a reliable representation of shared context rather than reasoning only over isolated utterances. Curre

Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models

SafetyDGX agent

arXiv:2604.20995v1 Announce Type: new Abstract: Alignment faking, where a model behaves aligned with developer policy when monitored but reverts to its own preferences when unobserved, is a concerning

VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution

Model ReleasesDGX agent

arXiv:2604.21450v1 Announce Type: cross Abstract: Recent advancements in visual autoregressive models (VAR) have demonstrated their effectiveness in image generation, highlighting their potential for

VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought

Model ReleasesDGX agent

arXiv:2604.21396v1 Announce Type: cross Abstract: The advancement of Large Vision-Language Models (LVLMs) requires precise local region-based reasoning that faithfully grounds the model's logic in act

VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models

SafetyDGX agent

arXiv:2604.03956v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models are emerging as embodied foundation models for robotic manipulation, but their deployment introduces a new

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation

Model ReleasesDGX agent

arXiv:2604.21375v1 Announce Type: cross Abstract: Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repet

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping

ResearchDGX agent

arXiv:2511.13587v2 Announce Type: replace-cross Abstract: Visual autoregressive (AR) generation models have demonstrated strong potential for image generation, yet their next-token-prediction paradigm

Watts-per-Intelligence Part II: Algorithmic Catalysis

ResearchDGX agent

arXiv:2604.20897v1 Announce Type: cross Abstract: We develop a thermodynamic theory of algorithmic catalysis within the watts-per-intelligence framework, identifying reusable computational structures

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs

Model ReleasesDGX agent

arXiv:2604.21911v1 Announce Type: cross Abstract: Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA

Model ReleasesDGX agent

arXiv:2511.01458v2 Announce Type: replace-cross Abstract: Safety and reliability are critical for deploying visual question answering (VQA) systems in surgery, where incorrect or ambiguous responses c

Who Defines 'Best'? Towards Interactive, User-Defined Evaluation of LLM Leaderboards

Model ReleasesDGX agent

arXiv:2604.21769v1 Announce Type: new Abstract: LLM leaderboards are widely used to compare models and guide deployment decisions. However, leaderboard rankings are shaped by evaluation priorities set

Who Defines Fairness? Target-Based Prompting for Demographic Representation in Generative Models

SafetyDGX agent

arXiv:2604.21036v1 Announce Type: new Abstract: Text-to-image(T2I) models like Stable Diffusion and DALL-E have made generative AI widely accessible, yet recent studies reveal that these systems often

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

SafetyDGX agent

arXiv:2604.21751v1 Announce Type: cross Abstract: LLMs have been showing limitations when it comes to cultural coverage and competence, and in some cases show regional biases such as amplifying Wester

Why Do Language Model Agents Whistleblow?

SafetyDGX agent

arXiv:2511.17085v3 Announce Type: replace-cross Abstract: The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds

Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2604.08016v2 Announce Type: replace Abstract: Regardless of its foundational role in human discovery and sense-making, abductive reasoning--the inference of the most plausible explanation for an

Words that make SENSE: Sensorimotor Norms in Learned Lexical Token Representations

ResearchDGX agent

arXiv:2602.00469v2 Announce Type: replace-cross Abstract: While word embeddings derive meaning from co-occurrence patterns, human language understanding is grounded in sensory and motor experience. We

Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model

Model ReleasesDGX agent

arXiv:2604.21223v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their ability to generate human-like text has ra

23 Apr 2026

A Field Guide to Decision Making

AgentsDGX agent

arXiv:2604.20669v1 Announce Type: cross Abstract: High-consequence decision making demands peak performance from individuals in positions of responsibility. Such executive authority bears the obligati

A Multi-Plant Machine Learning Framework for Emission Prediction, Forecasting, and Control in Cement Manufacturing

ApplicationsDGX agent

arXiv:2604.19903v1 Announce Type: cross Abstract: Cement production is among the largest contributors to industrial air pollution, emitting ~3 Mt NOx/year. The industry-standard mitigation approach, s

A Survey of Scaling in Large Language Model Reasoning

SafetyDGX agent

arXiv:2504.02181v2 Announce Type: replace Abstract: The rapid advancements in large Language models (LLMs) have significantly enhanced their reasoning capabilities, driven by various strategies such a

← Previous
1…310311312313314…354
Next →