AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlog
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
Model Releases

InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization

DGX agent

arXiv:2605.26175v1 Announce Type: cross Abstract: Low-bit activation quantization remains a major bottleneck in efficient large language model (LLM) deployment. The difficulty is not only that activat

model-releasesarxiv-cs-ai
27 May 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Innovation: An Almost Characterization of Hallucination

DGX agent

arXiv:2605.26808v1 Announce Type: cross Abstract: Hallucination is a central limitation of large language models (LLMs), and substantial effort has been devoted to understanding and mitigating it. Tow

researcharxiv-cs-ai
27 May 2026
Model Releases

Intelligent Detection and Mitigation of Carpet-Bombing DDoS Attacks in SDN Using Retrieval-Augmented Generation and Large Language Models

DGX agent

arXiv:2605.26307v1 Announce Type: cross Abstract: Software-Defined Networking (SDN) provides flexible and programmable network management; however, its centralized control architecture remains highly

model-releasesarxiv-cs-ai
27 May 2026
Safety

Intelligent Offloading in Vehicular Edge Computing: A Comprehensive Review of Deep Reinforcement Learning Approaches and Architectures

DGX agent

arXiv:2502.06963v3 Announce Type: replace-cross Abstract: The increasing complexity of Intelligent Transportation Systems (ITS) has led to significant interest in computational offloading to external

safetyarxiv-cs-ai
27 May 2026
Model Releases

InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward

DGX agent

arXiv:2605.26520v1 Announce Type: cross Abstract: While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow an

model-releasesarxiv-cs-ai
27 May 2026
Local Ai

Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory

DGX agent

arXiv:2605.26252v1 Announce Type: new Abstract: Long-running AI agents need persistent memory. Memory supports learning across sessions, reduces repeated context injection, and enables auditing of pas

local-aiarxiv-cs-ai
27 May 2026
Safety

It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

DGX agent

arXiv:2605.27288v1 Announce Type: cross Abstract: Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behav

safetyarxiv-cs-ai
27 May 2026
Model Releases

It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

DGX agent

arXiv:2605.26731v1 Announce Type: new Abstract: A prevalent assumption in LLM agent deployment holds that more structured harnesses universally improve reliability, and that higher-capability models n

model-releasesarxiv-cs-ai
27 May 2026
Safety

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models

DGX agent

arXiv:2605.26409v1 Announce Type: cross Abstract: Evaluating and mitigating a generative system's susceptibility to jailbreak attacks is critical to its safe deployment. Given the number of deployable

safetyarxiv-cs-ai
27 May 2026
Hardware

JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

DGX agent

arXiv:2605.26636v1 Announce Type: cross Abstract: We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention

hardwarearxiv-cs-ai
27 May 2026
Model Releases

JobBench: Aligning Agent Work With Human Will

DGX agent

arXiv:2605.26329v1 Announce Type: new Abstract: Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluat

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors

DGX agent

arXiv:2605.26955v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed to users around the world, they are integrated into everyday tasks across diverse cultural c

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

DGX agent

arXiv:2511.14993v3 Announce Type: replace-cross Abstract: This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis.

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations

DGX agent

arXiv:2605.26874v1 Announce Type: cross Abstract: LLM-based agents for industrial asset operations show limited accuracy when reasoning over flat document stores. AssetOpsBench (KDD 2026) establishes

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

L2Rec: Towards Dual-View Understanding of LLMs for Personalized Recommendation

DGX agent

arXiv:2605.26717v1 Announce Type: cross Abstract: Adapting large language models (LLMs) for personalized recommendation requires aligning their general-purpose capabilities with user-specific preferen

model-releasesarxiv-cs-ai
27 May 2026
Agents

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments

DGX agent

arXiv:2605.27209v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have facilitated the widespread deployment of LLMs as interactive agents capable of reasoning, planning,

agentsarxiv-cs-ai
27 May 2026
Model Releases

Learning When to Think While Listening in Large Audio-Language Models

DGX agent

arXiv:2605.27190v1 Announce Type: cross Abstract: Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reas

model-releasesarxiv-cs-ai
27 May 2026
Research

LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems

DGX agent

arXiv:2512.01556v3 Announce Type: replace Abstract: Foundation models often generate unreliable answers, while heuristic uncertainty estimators fail to fully distinguish correct from incorrect outputs

researcharxiv-cs-ai
27 May 2026
Research

Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Spatial-Relation Data

DGX agent

arXiv:2601.12809v2 Announce Type: replace-cross Abstract: Spatial understanding remains a key challenge in vision-language models. Yet it is still unclear whether such understanding is truly acquired,

researcharxiv-cs-ai
27 May 2026
Applications

LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation

DGX agent

arXiv:2605.26956v1 Announce Type: new Abstract: Entity linking is a key component of many downstream NLP systems, yet existing approaches are often tied to the specific target knowledge bases and doma

applicationsarxiv-cs-ai
27 May 2026
Safety

Less is More: Early Stopping Rollout for On-Policy Distillation

DGX agent

arXiv:2605.27028v1 Announce Type: cross Abstract: On-policy distillation has recently emerged as a promising alternative to standard sequence-level imitation, training a student by scoring its own rol

safetyarxiv-cs-ai
27 May 2026
Agents

Lessons from Penetration Tests on Large-Scale Agent Systems

DGX agent

arXiv:2605.27042v1 Announce Type: cross Abstract: As AI systems gain increasing autonomy and execution capability, the number of discovered security vulnerabilities continues to rise. However, many of

agentsarxiv-cs-ai
27 May 2026
Safety

Linear and Neural Dueling Bandits with Delayed Feedback

DGX agent

arXiv:2605.26554v1 Announce Type: cross Abstract: Contextual dueling bandits form a cornerstone of preference-based decision-making, with critical applications in recommender systems and large languag

safetyarxiv-cs-ai
27 May 2026
Agents

LiPUP-MA: A Residential Experience-centric Multi-Agent Framework for Living-in-the-loop Participatory Urban Planning

DGX agent

arXiv:2412.20505v2 Announce Type: replace Abstract: Participatory Urban Planning (PUP) is increasingly supported by LLM-based agents, yet existing methods largely rely on static preference elicitation

agentsarxiv-cs-ai
27 May 2026
Research

LitSeg: Narrative-Aware Document Segmentation for Literary RAG

DGX agent

arXiv:2605.27156v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge, particularly for long-tail domains suc

researcharxiv-cs-ai
27 May 2026
Model Releases

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

DGX agent

arXiv:2605.26781v1 Announce Type: new Abstract: Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

LLMs versus the Halting Problem: Characterizing Program Termination Reasoning

DGX agent

arXiv:2601.18987v5 Announce Type: replace-cross Abstract: Determining whether a program terminates is a central problem in computer science. Turing's Halting Problem established termination as undecid

model-releasesarxiv-cs-ai
27 May 2026
Research

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

DGX agent

arXiv:2605.27365v1 Announce Type: cross Abstract: Vision-language models (VLMs) commonly formulate visual grounding and detection as a coordinate-token generation problem, serializing each 2D box into

researcharxiv-cs-ai
27 May 2026
Research

Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)

DGX agent

arXiv:2605.27268v1 Announce Type: cross Abstract: Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabularies. W

researcharxiv-cs-ai
27 May 2026
Research

LUCoS: Latent Unsupervised Context Selection for Tabular Foundation Models

DGX agent

arXiv:2605.27254v1 Announce Type: cross Abstract: Selecting which instances to label is a key challenge in low-label tabular learning. For recent Tabular Foundation Models such as TabPFN, context sele

researcharxiv-cs-ai
27 May 2026
Model Releases

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

DGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Maat: The Agentic Legal Research Assistant for Competition Protection

DGX agent

arXiv:2605.27331v1 Announce Type: new Abstract: Competition law experts conducting legal research must review extensive volumes of cases, decisions, and judicial reports to identify precedents and ass

model-releasesarxiv-cs-ai
27 May 2026
Research

Managing Uncertainty in LLM-Generated Procedural Knowledge for Virtual Laboratory Planning

DGX agent

arXiv:2605.26333v1 Announce Type: new Abstract: Educational virtual laboratories can make experimental training more scala-ble, adaptive, and accessible, especially when students have limited access t

researcharxiv-cs-ai
27 May 2026
Research

Many Logics, One Methodology: A Plea for Logical Pluralism in Formalised Reasoning (preprint)

DGX agent

arXiv:2605.27246v1 Announce Type: cross Abstract: This position statement looks back on two decades of work on shallow embeddings of non-classical logics in classical higher-order logic (HOL), a line

researcharxiv-cs-ai
27 May 2026
Tutorials

MatFormBench: A Benchmarking Evaluation Framework for Target-Driven Materials Formulation

DGX agent

arXiv:2605.26741v1 Announce Type: cross Abstract: Inverse design of materials has significantly advanced target-driven formulation optimization, yet existing materials machine learning benchmarks rema

tutorialsarxiv-cs-ai
27 May 2026
Research

Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training

DGX agent

arXiv:2605.26189v1 Announce Type: cross Abstract: Quantization-aware training (QAT) with low-bit floating-point formats enables efficient LLM deployment, yet introduces subtle failure modes invisible

researcharxiv-cs-ai
27 May 2026
Safety

Measuring Prediction Uncertainty in Neural Cellular Automata

DGX agent

arXiv:2605.26726v1 Announce Type: cross Abstract: Neural cellular automata (NCA) provide a lightweight alternative to encoder-decoder segmentation networks. However, it can be difficult to decide when

safetyarxiv-cs-ai
27 May 2026
Agents

MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis

DGX agent

arXiv:2603.01131v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise in clinical diagnosis but remain limited by unreliable report generation, weak evidence ground

agentsarxiv-cs-ai
27 May 2026
Research

MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning

DGX agent

arXiv:2605.26567v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode evidence-based decision logic that clinicians apply by evaluating patient variables, conditional criteria, an

researcharxiv-cs-ai
27 May 2026
Model Releases

MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation

DGX agent

arXiv:2605.26621v1 Announce Type: cross Abstract: Volumetric Reasoning Segmentation (VRS) aims to segment a target region in a 3D medical scan from a free-form clinical query, where the referent is of

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MemFail: Stress-Testing Failure Modes of LLM Memory Systems

DGX agent

arXiv:2605.26667v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external memory systems to remain consistent across long-horizon interactions, but little empiric

model-releasesarxiv-cs-ai
27 May 2026
Safety

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

DGX agent

arXiv:2605.26154v1 Announce Type: cross Abstract: LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents

safetyarxiv-cs-ai
27 May 2026
Research

Message-Passing State-Space Models: Improving Graph Learning with Modern Sequence Modeling

DGX agent

arXiv:2505.18728v2 Announce Type: replace-cross Abstract: The recent success of State-Space Models (SSMs) in sequence modeling has motivated their adaptation to graph learning, giving rise to Graph St

researcharxiv-cs-ai
27 May 2026
Research

MetaSICL: Adapting Audiroty LLM via Meta Speech In-Context Learning

DGX agent

arXiv:2601.18904v2 Announce Type: replace-cross Abstract: Auditory Large Language Models (LLMs) have demonstrated strong performance across a wide range of speech and audio understanding tasks. Nevert

researcharxiv-cs-ai
27 May 2026
Agents

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

DGX agent

arXiv:2605.26691v1 Announce Type: new Abstract: Medical AI agents increasingly use external tools for diagnosis, treatment recommendation, and evidence retrieval, yet most existing approaches assume t

agentsarxiv-cs-ai
27 May 2026
Research

MiRD: Reliable Set-Valued Prediction for Open-Ended Question Answering via Miscoverage Risk Decomposition

DGX agent

arXiv:2605.27091v1 Announce Type: cross Abstract: Reliable set-valued prediction provides a principled way to mitigate hallucinations in open-ended question answering (QA), yet existing conformal appr

researcharxiv-cs-ai
27 May 2026
Model Releases

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

DGX agent

arXiv:2605.26546v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems foc

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MobileMoE: Scaling On-Device Mixture of Experts

DGX agent

arXiv:2605.27358v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales

model-releasesarxiv-cs-ai
27 May 2026
← Previous
1…258259260261262…448
Next →