AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
22 Apr 2026

Towards Understanding the Robustness of Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2604.18756v1 Announce Type: cross Abstract: Large Language Models (LLMs) remain vulnerable to optimization-based jailbreak attacks that exploit internal gradient structure. While Sparse Autoenco

TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards

SafetyDGX agent

arXiv:2512.07761v3 Announce Type: replace Abstract: Large language models have seen widespread adoption, yet they remain vulnerable to multi-turn jailbreak attacks, threatening their safe deployment.

TurboEvolve: Towards Fast and Robust LLM-Driven Program Evolution

ResearchDGX agent

arXiv:2604.18607v1 Announce Type: cross Abstract: LLM-driven program evolution can discover high-quality programs, but its cost and run-to-run variance hinder reliable progress. We propose TurboEvolve


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Two-dimensional early exit optimisation of LLM inference

Model ReleasesDGX agent

arXiv:2604.18592v1 Announce Type: cross Abstract: We introduce a two-dimensional (2D) early exit strategy that coordinates layer-wise and sentence-wise exiting for classification tasks in large langua

UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction

ApplicationsDGX agent

arXiv:2604.19221v1 Announce Type: new Abstract: Full-duplex speech interaction, as the most natural and intuitive mode of human communication, is driving artificial intelligence toward more human-like

Uncertainty Quantification in Detection Transformers: Object-Level Calibration and Image-Level Reliability

Local AiDGX agent

arXiv:2412.01782v4 Announce Type: replace-cross Abstract: DETR and its variants have emerged as promising architectures for object detection, offering an end-to-end prediction pipeline. In practice, h

Understanding LLM Performance Degradation in Multi-Instance Processing: The Roles of Instance Count and Context Length

ResearchDGX agent

arXiv:2603.22608v2 Announce Type: replace Abstract: Users often rely on Large Language Models (LLMs) for processing multiple documents or performing analysis over a number of instances. For example, a

Unifying Controller Design for Stabilizing Nonlinear Systems with Norm-Bounded Control Inputs

ResearchDGX agent

arXiv:2403.03030v2 Announce Type: replace-cross Abstract: This paper revisits a classical challenge in the design of stabilizing controllers for nonlinear systems with a norm-bounded input constraint.

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling

Model ReleasesDGX agent

arXiv:2604.19734v1 Announce Type: cross Abstract: Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative,

Unlocking the Edge deployment and ondevice acceleration of multi-LoRA enabled one-for-all foundational LLM

Model ReleasesDGX agent

arXiv:2604.18655v1 Announce Type: cross Abstract: Deploying large language models (LLMs) on smartphones poses significant engineering challenges due to stringent constraints on memory, latency, and ru

User Simulation in the Era of Generative AI: User Modeling, Synthetic Data Generation, and System Evaluation

SafetyDGX agent

arXiv:2501.04410v2 Announce Type: replace Abstract: User simulation is an emerging interdisciplinary topic with multiple critical applications in the era of Generative AI. It involves creating an inte

VideoAgent: Personalized Synthesis of Scientific Videos

Model ReleasesDGX agent

arXiv:2509.11253v2 Announce Type: replace Abstract: The technical complexity of research papers often limits their reach, necessitating more accessible formats like scientific videos to disseminate ke

ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios

Model ReleasesDGX agent

arXiv:2601.08620v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) pipelines must address challenges beyond simple single-document retrieval, such as interpreting visual elements

Visual Reasoning Agent: Robust Vision Systems in Remote Sensing via Inference-Time Scaling

Model ReleasesDGX agent

arXiv:2509.16343v2 Announce Type: replace-cross Abstract: Building robust vision systems for high-stakes domains such as remote sensing requires stronger visual reasoning than what single-pass inferen

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2604.19728v1 Announce Type: cross Abstract: We present VLA Foundry, an open-source framework that unifies LLM, VLM, and VLA training in a single codebase. Most open-source VLA efforts specialize

WebUncertainty: Dual-Level Uncertainty Driven Planning and Reasoning For Autonomous Web Agent

AgentsDGX agent

arXiv:2604.17821v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have empowered autonomous web agents to execute natural language instructions directly on real-w

When Graph Structure Becomes a Liability: A Critical Re-Evaluation of Graph Neural Networks for Bitcoin Fraud Detection under Temporal Distribution Shift

ResearchDGX agent

arXiv:2604.19514v1 Announce Type: cross Abstract: The consensus that GCN, GraphSAGE, GAT, and EvolveGCN outperform feature-only baselines on the Elliptic Bitcoin Dataset is widely cited but has not be

Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs

ResearchDGX agent

arXiv:2604.18880v1 Announce Type: cross Abstract: LLMs frequently generate fictitious yet convincing citations, often expressing high confidence even when the underlying reference is wrong. We study t

Who Benefits from AI? Self-Selection, Skill Gap, and the Hidden Costs of AI Feedback

TutorialsDGX agent

arXiv:2409.18660v2 Announce Type: replace-cross Abstract: Feedback from artificial intelligence (AI) is increasingly easy to access and research has already established that people learn from it. But

Who Shapes Brazil's Vaccine Debate? Semi-Supervised Modeling of Stance and Polarization in YouTube's Media Ecosystem

ResearchDGX agent

arXiv:2604.18586v1 Announce Type: cross Abstract: Vaccination remains a cornerstone of global public health, yet the COVID-19 pandemic exposed how online misinformation, political polarization, and de

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation

Model ReleasesDGX agent

arXiv:2604.02368v4 Announce Type: replace Abstract: As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficienc

20 Apr 2026

(1D) Ordered Tokens Enable Efficient Test-Time Search

ResearchDGX agent

arXiv:2604.15453v1 Announce Type: cross Abstract: Tokenization is a key component of autoregressive (AR) generative models, converting raw data into more manageable units for modeling. Commonly, token

1S-DAug: One-Shot Data Augmentation for Robust Few-Shot Generalization

Model ReleasesDGX agent

arXiv:2602.00114v4 Announce Type: replace-cross Abstract: Few-shot learning (FSL) challenges model generalization to novel classes based on just a few shots of labeled examples, a testbed where tradit

A Comparative Study on the Impact of Traditional Learning and Interactive Learning on Students' Academic Performance and Emotional Well-Being

TutorialsDGX agent

arXiv:2604.15335v1 Announce Type: cross Abstract: The growing adoption of interactive learning tools in higher education offers new opportunities to enhance student performance and well-being. This st

A PennyLane-Centric Dataset to Enhance LLM-based Quantum Code Generation using RAG

Model ReleasesDGX agent

arXiv:2503.02497v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) offer powerful capabilities in code generation, natural language understanding, and domain-specific reasoning. Th

A Q-learning-based QoS-aware multipath routing protocol in IoMT-based wireless body area network

ApplicationsDGX agent

arXiv:2604.15489v1 Announce Type: cross Abstract: The Internet of Medical Things (IoMT) enables intelligent healthcare services but faces challenges such as dynamic topology, energy constraints, and d

A Two-Stage, Object-Centric Deep Learning Framework for Robust Exam Cheating Detection

Local AiDGX agent

arXiv:2604.16234v1 Announce Type: cross Abstract: Academic integrity continues to face the persistent challenge of examination cheating. Traditional invigilation relies on human observation, which is

Aerial Multi-Functional RIS in Fluid Antennas-Aided Full-Duplex Networks: A Self-Optimized Hybrid Deep Reinforcement Learning Approach

SafetyDGX agent

arXiv:2604.14309v2 Announce Type: replace-cross Abstract: To address high data traffic demands of sixth-generation (6G) networks, this paper proposes a novel architecture that integrates autonomous ae

AgentV-RL: Scaling Reward Modeling with Agentic Verifier

AgentsDGX agent

arXiv:2604.16004v1 Announce Type: cross Abstract: Verifiers have been demonstrated to enhance LLM reasoning via test-time scaling (TTS). Yet, they face significant challenges in complex domains. Error

AI Agents and Hard Choices

SafetyDGX agent

arXiv:2504.15304v2 Announce Type: replace Abstract: Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a t

AI-assisted Protocol Information Extraction For Improved Accuracy and Efficiency in Clinical Trial Workflows

ApplicationsDGX agent

arXiv:2602.00052v2 Announce Type: replace-cross Abstract: Increasing clinical trial protocol complexity, amendments, and challenges around knowledge management create significant burden for trial team

AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection

SafetyDGX agent

arXiv:2604.16207v1 Announce Type: cross Abstract: As forgery types continue to emerge consistently, Incremental Face Forgery Detection (IFFD) has become a crucial paradigm. However, existing methods t

AISysRev -- LLM-based Tool for Title-abstract Screening

Model ReleasesDGX agent

arXiv:2510.06708v3 Announce Type: replace-cross Abstract: Conducting systematic reviews is laborious. In the screening or study selection phase, the number of papers can be overwhelming. Recent resear

Analyzing Chain of Thought (CoT) Approaches in Control Flow Code Deobfuscation Tasks

TutorialsDGX agent

arXiv:2604.15390v1 Announce Type: cross Abstract: Code deobfuscation is the task of recovering a readable version of a program while preserving its original behavior. In practice, this often requires

Anthropomorphism and Trust in Human-Large Language Model interactions

ResearchDGX agent

arXiv:2604.15316v1 Announce Type: cross Abstract: With large language models (LLMs) becoming increasingly prevalent in daily life, so too has the tendency to attribute to them human-like minds and emo

Applied Explainability for Large Language Models: A Comparative Study

ApplicationsDGX agent

arXiv:2604.15371v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong performance across many natural language processing tasks, yet their decision processes remain difficult t

ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence

Model ReleasesDGX agent

arXiv:2603.24621v2 Announce Type: replace Abstract: We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents

ArrayTac: A Closed-loop Piezoelectric Tactile Platform for Continuously Tunable Rendering of Shape, Stiffness, and Friction

ResearchDGX agent

arXiv:2603.13829v2 Announce Type: replace-cross Abstract: Human touch depends on the integration of shape, stiffness, and friction, yet existing tactile displays cannot render these cues together as c

AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units

Model ReleasesDGX agent

arXiv:2601.07160v2 Announce Type: replace Abstract: To meet the ever-increasing demand for computational efficiency, Neural Processing Units (NPUs) have become critical in modern AI infrastructure. Ho

ASMR-Bench: Auditing for Sabotage in ML Research

Model ReleasesDGX agent

arXiv:2604.16286v1 Announce Type: new Abstract: As AI systems are increasingly used to conduct research autonomously, misaligned systems could introduce subtle flaws that produce misleading results wh

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing

ResearchDGX agent

arXiv:2604.16056v1 Announce Type: cross Abstract: Text-based speech editing aims to modify specific segments while preserving speaker identity and acoustic context. Existing methods rely on task-speci

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency

Model ReleasesDGX agent

arXiv:2604.16158v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex tasks. Yet ensuring that the reasoning trace both

AutoFed: Personalized Federated Traffic Prediction via Adaptive Prompt

Model ReleasesDGX agent

arXiv:2512.24625v2 Announce Type: replace-cross Abstract: Accurate traffic prediction is essential for Intelligent Transportation Systems, including ride-hailing, urban road planning, and vehicle flee

Automatic Combination of Sample Selection Strategies for Few-Shot Learning

ResearchDGX agent

arXiv:2402.03038v2 Announce Type: replace-cross Abstract: In few-shot learning, the selection of samples has a significant impact on the performance of the model. While effective sample selection stra

Automating Crash Diagram Generation Using Vision-Language Models: A Case Study on Multi-Lane Roundabouts

Model ReleasesDGX agent

arXiv:2604.15332v1 Announce Type: cross Abstract: Crash diagrams are essential tools in transportation safety analysis, yet their manual preparation remains time-consuming and prone to human variabili

BAGEL: Benchmarking Animal Knowledge Expertise in Language Models

Model ReleasesDGX agent

arXiv:2604.16241v1 Announce Type: cross Abstract: Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language model

Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI

Model ReleasesDGX agent

arXiv:2604.15808v1 Announce Type: cross Abstract: Spatial reasoning and visual grounding are core capabilities for vision-language models (VLMs), yet most medical VLMs produce predictions without tran

Beyond Distribution Sharpening: The Importance of Task Rewards

Model ReleasesDGX agent

arXiv:2604.16259v1 Announce Type: cross Abstract: Frontier models have demonstrated exceptional capabilities following the integration of task-reward-based reinforcement learning (RL) into their train

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants

Model ReleasesDGX agent

arXiv:2510.24328v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal co

Beyond Passive Viewing: A Pilot Study of a Hybrid Learning Platform Augmenting Video Lectures with Conversational AI

ApplicationsDGX agent

arXiv:2604.15334v1 Announce Type: cross Abstract: The exponential growth of AI education has brought millions of learners to online platforms, yet this massive scale has simultaneously exposed critica

Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning

Local AiDGX agent

arXiv:2604.15414v1 Announce Type: cross Abstract: Continual reinforcement learning must balance retention with adaptation, yet many methods still rely on single-model preservation, committing to one e

Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations

ResearchDGX agent

arXiv:2604.16217v1 Announce Type: cross Abstract: Large language models are increasingly deployed in settings where reliability matters, yet output-level uncertainty signals such as token probabilitie

Bilevel Optimization of Agent Skills via Monte Carlo Tree Search

AgentsDGX agent

arXiv:2604.15709v1 Announce Type: new Abstract: Agent exttt{skills} are structured collections of instructions, tools, and supporting resources that help large language model (LLM) agents perform part

BioHiCL: Hierarchical Multi-Label Contrastive Learning for Biomedical Retrieval with MeSH Labels

ResearchDGX agent

arXiv:2604.15591v1 Announce Type: cross Abstract: Effective biomedical information retrieval requires modeling domain semantics and hierarchical relationships among biomedical texts. Existing biomedic

Bridging the phenotype-target gap for molecular generation via multi-objective reinforcement learning

TutorialsDGX agent

arXiv:2509.21010v2 Announce Type: replace-cross Abstract: The de novo generation of drug-like molecules capable of inducing desirable phenotypic changes is receiving increasing attention. However, pre

Bureaucratic Silences: What the Canadian AI Register Reveals, Omits, and Obscures

ResearchDGX agent

arXiv:2604.15514v1 Announce Type: new Abstract: In November 2025, the Government of Canada operationalized its commitment to transparency by releasing its first Federal AI Register. In this paper, we

Can LLMs Understand the Impact of Trauma? Costs and Benefits of LLMs Coding the Interviews of Firearm Violence Survivors

ResearchDGX agent

arXiv:2604.16132v1 Announce Type: cross Abstract: Firearm violence is a pressing public health issue, yet research into survivors' lived experiences remains underfunded and difficult to scale. Qualita

Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations

AgentsDGX agent

arXiv:2602.05523v2 Announce Type: replace-cross Abstract: Agentic large language models (LLMs) are increasingly evaluated on cybersecurity tasks using capture-the-flag (CTF) benchmarks, yet existing p

Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs

ResearchDGX agent

arXiv:2604.16060v1 Announce Type: cross Abstract: Multimodal Reasoning Models (MRMs) leveraging Chain-of-Thought (CoT) based thinking have revolutionized mathematical and logical problem-solving. Howe

Characterising LLM-Generated Competency Questions: a Cross-Domain Empirical Study using Open and Closed Models

Model ReleasesDGX agent

arXiv:2604.16258v1 Announce Type: new Abstract: Competency Questions (CQs) are a cornerstone of requirement elicitation in ontology engineering. CQs represent requirements as a set of natural language

← Previous
1…320321322323324…354
Next →