AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
27 May 2026

Hi-SAM: A Hierarchical Structure-Aware Multi-modal Framework for Large-Scale Recommendation

SafetyDGX agent

arXiv:2602.11799v2 Announce Type: replace Abstract: Multi-modal recommendation has gained traction as items possess rich attributes like text and images. Semantic ID-based approaches effectively discr

High-Quality Synthetic Financial Time-Series using a GAN-Diffusion Framework

ResearchDGX agent

arXiv:2605.27113v1 Announce Type: cross Abstract: In recent years, financial institutions and firms have increasingly adopted synthetic data to address data scarcity and to generate counterfactual mar

HiSpec: Hierarchical Speculative Decoding for LLMs

ResearchDGX agent

arXiv:2510.01336v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates LLM inference by using a smaller draft model to speculate tokens that a larger target model verifies. Verific


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

How Chain-of-Thought Works? Tracing Information Flow from Decoding, Projection, and Activation

Model ReleasesDGX agent

arXiv:2507.20758v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting significantly enhances model reasoning, yet its internal mechanisms remain poorly understood. We analyze CoT's oper

How Reliable are LLMs for Reasoning on the Re-ranking task?

SafetyDGX agent

arXiv:2508.18444v2 Announce Type: replace-cross Abstract: With the improving semantic understanding capability of Large Language Models (LLMs), they exhibit a greater awareness and alignment with huma

How to Square Tensor Networks and Circuits Without Squaring Them

TutorialsDGX agent

arXiv:2512.17090v2 Announce Type: replace-cross Abstract: Squared tensor networks (TNs) and their extension as computational graphs--squared circuits--have been used as expressive distribution estimat

HRVConformer: Neonatal Hypoxic-Ischemic Encephalopathy Classification from the Heart Rate signals

Local AiDGX agent

arXiv:2605.26190v1 Announce Type: cross Abstract: This paper presents the HRVConformer, a novel deep learning architecture for the classification of hypoxic-ischemic encephalopathy (HIE) using the ins

HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML

Model ReleasesDGX agent

arXiv:2605.26807v1 Announce Type: cross Abstract: LLMs can now produce full HTML pages, but many of those pages are only superficially correct: they render once, then fail under scroll, hover, click,

ICCU: In-Context Continual Unlearning via Pattern-Induced Refusal Rules

ApplicationsDGX agent

arXiv:2605.27138v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of specific data from trained language models. In real-world deployments, unlearning requests often arri

ICICLE: Expanding Retrieval with In-Context Documents

ResearchDGX agent

arXiv:2605.26902v1 Announce Type: cross Abstract: Generative retrieval (GR) maps queries directly to document identifiers (docids) using parametric knowledge, However, this design makes corpus expansi

Implementation of Big Data Analytics for Diabetes Management: Needs Assessment in the Rwanda Healthcare System

ApplicationsDGX agent

arXiv:2605.26786v1 Announce Type: cross Abstract: Diabetes is a chronic metabolic disease that can lead to serious health problems if not diagnosed and managed early. Big Data Analytics (BDA) and mach

Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction

ResearchDGX agent

arXiv:2510.03352v3 Announce Type: replace-cross Abstract: Diffusion models have been used as priors for solving inverse problems. However, existing approaches typically overlook side information that

InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization

Model ReleasesDGX agent

arXiv:2605.26175v1 Announce Type: cross Abstract: Low-bit activation quantization remains a major bottleneck in efficient large language model (LLM) deployment. The difficulty is not only that activat

Innovation: An Almost Characterization of Hallucination

ResearchDGX agent

arXiv:2605.26808v1 Announce Type: cross Abstract: Hallucination is a central limitation of large language models (LLMs), and substantial effort has been devoted to understanding and mitigating it. Tow

Intelligent Detection and Mitigation of Carpet-Bombing DDoS Attacks in SDN Using Retrieval-Augmented Generation and Large Language Models

Model ReleasesDGX agent

arXiv:2605.26307v1 Announce Type: cross Abstract: Software-Defined Networking (SDN) provides flexible and programmable network management; however, its centralized control architecture remains highly

Intelligent Offloading in Vehicular Edge Computing: A Comprehensive Review of Deep Reinforcement Learning Approaches and Architectures

SafetyDGX agent

arXiv:2502.06963v3 Announce Type: replace-cross Abstract: The increasing complexity of Intelligent Transportation Systems (ITS) has led to significant interest in computational offloading to external

InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward

Model ReleasesDGX agent

arXiv:2605.26520v1 Announce Type: cross Abstract: While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow an

Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory

Local AiDGX agent

arXiv:2605.26252v1 Announce Type: new Abstract: Long-running AI agents need persistent memory. Memory supports learning across sessions, reduces repeated context injection, and enables auditing of pas

It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

SafetyDGX agent

arXiv:2605.27288v1 Announce Type: cross Abstract: Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behav

It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

Model ReleasesDGX agent

arXiv:2605.26731v1 Announce Type: new Abstract: A prevalent assumption in LLM agent deployment holds that more structured harnesses universally improve reliability, and that higher-capability models n

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models

SafetyDGX agent

arXiv:2605.26409v1 Announce Type: cross Abstract: Evaluating and mitigating a generative system's susceptibility to jailbreak attacks is critical to its safe deployment. Given the number of deployable

JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

HardwareDGX agent

arXiv:2605.26636v1 Announce Type: cross Abstract: We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention

JobBench: Aligning Agent Work With Human Will

Model ReleasesDGX agent

arXiv:2605.26329v1 Announce Type: new Abstract: Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluat

JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors

Model ReleasesDGX agent

arXiv:2605.26955v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed to users around the world, they are integrated into everyday tasks across diverse cultural c

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

Model ReleasesDGX agent

arXiv:2511.14993v3 Announce Type: replace-cross Abstract: This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis.

Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations

Model ReleasesDGX agent

arXiv:2605.26874v1 Announce Type: cross Abstract: LLM-based agents for industrial asset operations show limited accuracy when reasoning over flat document stores. AssetOpsBench (KDD 2026) establishes

L2Rec: Towards Dual-View Understanding of LLMs for Personalized Recommendation

Model ReleasesDGX agent

arXiv:2605.26717v1 Announce Type: cross Abstract: Adapting large language models (LLMs) for personalized recommendation requires aligning their general-purpose capabilities with user-specific preferen

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments

AgentsDGX agent

arXiv:2605.27209v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have facilitated the widespread deployment of LLMs as interactive agents capable of reasoning, planning,

Learning When to Think While Listening in Large Audio-Language Models

Model ReleasesDGX agent

arXiv:2605.27190v1 Announce Type: cross Abstract: Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reas

LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems

ResearchDGX agent

arXiv:2512.01556v3 Announce Type: replace Abstract: Foundation models often generate unreliable answers, while heuristic uncertainty estimators fail to fully distinguish correct from incorrect outputs

Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Spatial-Relation Data

ResearchDGX agent

arXiv:2601.12809v2 Announce Type: replace-cross Abstract: Spatial understanding remains a key challenge in vision-language models. Yet it is still unclear whether such understanding is truly acquired,

LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation

ApplicationsDGX agent

arXiv:2605.26956v1 Announce Type: new Abstract: Entity linking is a key component of many downstream NLP systems, yet existing approaches are often tied to the specific target knowledge bases and doma

Less is More: Early Stopping Rollout for On-Policy Distillation

SafetyDGX agent

arXiv:2605.27028v1 Announce Type: cross Abstract: On-policy distillation has recently emerged as a promising alternative to standard sequence-level imitation, training a student by scoring its own rol

Lessons from Penetration Tests on Large-Scale Agent Systems

AgentsDGX agent

arXiv:2605.27042v1 Announce Type: cross Abstract: As AI systems gain increasing autonomy and execution capability, the number of discovered security vulnerabilities continues to rise. However, many of

Linear and Neural Dueling Bandits with Delayed Feedback

SafetyDGX agent

arXiv:2605.26554v1 Announce Type: cross Abstract: Contextual dueling bandits form a cornerstone of preference-based decision-making, with critical applications in recommender systems and large languag

LiPUP-MA: A Residential Experience-centric Multi-Agent Framework for Living-in-the-loop Participatory Urban Planning

AgentsDGX agent

arXiv:2412.20505v2 Announce Type: replace Abstract: Participatory Urban Planning (PUP) is increasingly supported by LLM-based agents, yet existing methods largely rely on static preference elicitation

LitSeg: Narrative-Aware Document Segmentation for Literary RAG

ResearchDGX agent

arXiv:2605.27156v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge, particularly for long-tail domains suc

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

Model ReleasesDGX agent

arXiv:2605.26781v1 Announce Type: new Abstract: Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors

LLMs versus the Halting Problem: Characterizing Program Termination Reasoning

Model ReleasesDGX agent

arXiv:2601.18987v5 Announce Type: replace-cross Abstract: Determining whether a program terminates is a central problem in computer science. Turing's Halting Problem established termination as undecid

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

ResearchDGX agent

arXiv:2605.27365v1 Announce Type: cross Abstract: Vision-language models (VLMs) commonly formulate visual grounding and detection as a coordinate-token generation problem, serializing each 2D box into

Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)

ResearchDGX agent

arXiv:2605.27268v1 Announce Type: cross Abstract: Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabularies. W

LUCoS: Latent Unsupervised Context Selection for Tabular Foundation Models

ResearchDGX agent

arXiv:2605.27254v1 Announce Type: cross Abstract: Selecting which instances to label is a key challenge in low-label tabular learning. For recent Tabular Foundation Models such as TabPFN, context sele

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

Model ReleasesDGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

Maat: The Agentic Legal Research Assistant for Competition Protection

Model ReleasesDGX agent

arXiv:2605.27331v1 Announce Type: new Abstract: Competition law experts conducting legal research must review extensive volumes of cases, decisions, and judicial reports to identify precedents and ass

Managing Uncertainty in LLM-Generated Procedural Knowledge for Virtual Laboratory Planning

ResearchDGX agent

arXiv:2605.26333v1 Announce Type: new Abstract: Educational virtual laboratories can make experimental training more scala-ble, adaptive, and accessible, especially when students have limited access t

Many Logics, One Methodology: A Plea for Logical Pluralism in Formalised Reasoning (preprint)

ResearchDGX agent

arXiv:2605.27246v1 Announce Type: cross Abstract: This position statement looks back on two decades of work on shallow embeddings of non-classical logics in classical higher-order logic (HOL), a line

MatFormBench: A Benchmarking Evaluation Framework for Target-Driven Materials Formulation

TutorialsDGX agent

arXiv:2605.26741v1 Announce Type: cross Abstract: Inverse design of materials has significantly advanced target-driven formulation optimization, yet existing materials machine learning benchmarks rema

Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training

ResearchDGX agent

arXiv:2605.26189v1 Announce Type: cross Abstract: Quantization-aware training (QAT) with low-bit floating-point formats enables efficient LLM deployment, yet introduces subtle failure modes invisible

Measuring Prediction Uncertainty in Neural Cellular Automata

SafetyDGX agent

arXiv:2605.26726v1 Announce Type: cross Abstract: Neural cellular automata (NCA) provide a lightweight alternative to encoder-decoder segmentation networks. However, it can be difficult to decide when

MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis

AgentsDGX agent

arXiv:2603.01131v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise in clinical diagnosis but remain limited by unreliable report generation, weak evidence ground

MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning

ResearchDGX agent

arXiv:2605.26567v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode evidence-based decision logic that clinicians apply by evaluating patient variables, conditional criteria, an

MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation

Model ReleasesDGX agent

arXiv:2605.26621v1 Announce Type: cross Abstract: Volumetric Reasoning Segmentation (VRS) aims to segment a target region in a 3D medical scan from a free-form clinical query, where the referent is of

MemFail: Stress-Testing Failure Modes of LLM Memory Systems

Model ReleasesDGX agent

arXiv:2605.26667v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external memory systems to remain consistent across long-horizon interactions, but little empiric

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

SafetyDGX agent

arXiv:2605.26154v1 Announce Type: cross Abstract: LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents

Message-Passing State-Space Models: Improving Graph Learning with Modern Sequence Modeling

ResearchDGX agent

arXiv:2505.18728v2 Announce Type: replace-cross Abstract: The recent success of State-Space Models (SSMs) in sequence modeling has motivated their adaptation to graph learning, giving rise to Graph St

MetaSICL: Adapting Audiroty LLM via Meta Speech In-Context Learning

ResearchDGX agent

arXiv:2601.18904v2 Announce Type: replace-cross Abstract: Auditory Large Language Models (LLMs) have demonstrated strong performance across a wide range of speech and audio understanding tasks. Nevert

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

AgentsDGX agent

arXiv:2605.26691v1 Announce Type: new Abstract: Medical AI agents increasingly use external tools for diagnosis, treatment recommendation, and evidence retrieval, yet most existing approaches assume t

MiRD: Reliable Set-Valued Prediction for Open-Ended Question Answering via Miscoverage Risk Decomposition

ResearchDGX agent

arXiv:2605.27091v1 Announce Type: cross Abstract: Reliable set-valued prediction provides a principled way to mitigate hallucinations in open-ended question answering (QA), yet existing conformal appr

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

Model ReleasesDGX agent

arXiv:2605.26546v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems foc

MobileMoE: Scaling On-Device Mixture of Experts

Model ReleasesDGX agent

arXiv:2605.27358v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales

← Previous
1…206207208209210…358
Next →