AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,040 results
Model Releases

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

DGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

model-releasesarxiv-cs-ai
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Memory Architectures for Multi-Turn Text-to-SQL: A Benchmark and Empirical Study

DGX agent

arXiv:2605.26394v1 Announce Type: new Abstract: Multi-turn Text-to-SQL is central to enterprise analytics yet remains predominantly evaluated in single-turn settings. We introduce EnterpriseMem-Bench,

model-releasesarxiv-cs-cl
27 May 2026
Tutorials

ORCA: An End-to-End Interactive Copilot for Optimized Root Cause Analysis

DGX agent

arXiv:2605.27022v1 Announce Type: new Abstract: Causal analysis is a crucial task in many domains, including manufacturing, social science, and medicine. However, despite recent progress, the conceptu

tutorialsarxiv-cs-ai
27 May 2026
Research

Probabilistic Recurrent Intention Switching Model

DGX agent

arXiv:2605.26998v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) recovers reward functions from observed behavior, yet traditional methods assume a single stationary reward that ca

researcharxiv-cs-lg
27 May 2026
Hardware

Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU

DGX agent

arXiv:2605.26118v1 Announce Type: cross Abstract: Porting deep learning algorithms to new hardware accelerators requires developers to repeatedly apply the same low-level optimizations -- quantization

hardwarearxiv-cs-ai
27 May 2026
Safety

AI-Assisted Systematization for Evaluating GenAI Systems

DGX agent

arXiv:2605.26001v1 Announce Type: cross Abstract: Evaluating generative AI (GenAI) systems is challenging because many targets of evaluation are broad, contested concepts, such as 'reasoning,' 'fairne

safetyarxiv-cs-ai
26 May 2026
Model Releases

An Interactive Paradigm for Deep Research

DGX agent

arXiv:2605.24266v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled deep research systems that synthesize comprehensive, report-style answers to open-ended q

model-releasesarxiv-cs-ai
26 May 2026
Research

ATWL: A Formal Language for Representing, Comparing, and Reusing Visual Analytics Workflows

DGX agent

arXiv:2605.25489v1 Announce Type: new Abstract: Visual analytics (VA) workflows are inherently complex, involving data transformation, feature engineering, visual representation, and human interpretat

researcharxiv-cs-ai
26 May 2026
Model Releases

Benchmarking and Learning Real-World Customer Service Dialogue

DGX agent

arXiv:2510.22143v3 Announce Type: replace Abstract: Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements

model-releasesarxiv-cs-cl
26 May 2026
Safety

Causal methods for LLM development and evaluation

DGX agent

arXiv:2605.25998v1 Announce Type: new Abstract: Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and

safetyarxiv-cs-lg
26 May 2026
Model Releases

Distilling Game Code World Model Generation into Lightweight Large Language Models

DGX agent

arXiv:2605.24375v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown great ability in generating executable code from natural language, opening the possibility of automatically cons

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction

DGX agent

arXiv:2605.24686v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly integrated into emotionally sensitive domains, the structural integrity of their emotional intelligence

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

FOUND-IT: Foundation-model-first Task-driven 3D Scene Graphs with Granularity on Demand

DGX agent

arXiv:2605.25371v1 Announce Type: new Abstract: We present the first approach to build hierarchical task-driven 3D scene graphs of arbitrary indoor or outdoor environments using an uncalibrated monocu

model-releasesarxiv-cs-ro
26 May 2026
Safety

Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation

DGX agent

arXiv:2605.25170v1 Announce Type: cross Abstract: Training data for olfaction is scattered through disparate, non-standardized datasets that limit the ability to build representative world models. Olf

safetyarxiv-cs-ai
26 May 2026
Model Releases

Hylos: Operability Contracts for Model-Native Spatial Intelligence

DGX agent

arXiv:2605.24728v1 Announce Type: new Abstract: Foundation models can increasingly describe, reconstruct, and generate 3D objects, assemblies, scenes, and environments, but visually plausible spatial

model-releasesarxiv-cs-ai
26 May 2026
Safety

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition

DGX agent

arXiv:2605.24005v1 Announce Type: new Abstract: The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous

safetyarxiv-cs-ai
26 May 2026
Research

Logic-Guided Socially-aware Robot Navigation World Model

DGX agent

arXiv:2510.23509v2 Announce Type: replace Abstract: Social robot navigation increasingly relies on large language models for reasoning, path planning, and enabling movement in dynamic human spaces. Ho

researcharxiv-cs-ro
26 May 2026
Research

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments

DGX agent

arXiv:2605.24020v1 Announce Type: cross Abstract: Advancements at the intersection of computer vision and natural language processing are crucial for applications like assistive tech, multimedia query

researcharxiv-cs-ai
26 May 2026
Model Releases

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction

DGX agent

arXiv:2605.24562v1 Announce Type: cross Abstract: Pedestrian intention and trajectory prediction are critical for the safe deployment of autonomous driving systems, directly influencing navigation dec

model-releasesarxiv-cs-ai
26 May 2026
Research

Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content

DGX agent

arXiv:2605.24421v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as analyst assistants in security operations centers (SOCs), where they ingest log and alert data t

researcharxiv-cs-lg
26 May 2026
Safety

Right-Sizing Communication and Recommendation Set Size in AI-Assisted Search

DGX agent

arXiv:2605.23944v1 Announce Type: new Abstract: We model the interaction between a user and an AI driven recommendation system. The user initiates the process by conveying preference information throu

safetyarxiv-cs-ai
26 May 2026
Model Releases

SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation

DGX agent

arXiv:2605.25984v1 Announce Type: cross Abstract: Ensuring safe and contextually appropriate behaviour in Large Language Models (LLMs) remains a critical challenge for real-world deployment. We presen

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions

DGX agent

arXiv:2605.25073v1 Announce Type: cross Abstract: Background: Fine-tuning is central to adapting pre-trained Large Language Models (LLMs) to downstream tasks, but its reliance on training data, parame

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Teaching large language models to reason like expert diagnosticians

DGX agent

arXiv:2509.12194v2 Announce Type: replace Abstract: Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

The Time is Here for Just-in-Time Systems: Challenges and Opportunities

DGX agent

arXiv:2605.24096v1 Announce Type: cross Abstract: Core systems like key-value stores have historically taken years to build, and are designed to be general so as to amortize cost across deployments, p

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction

DGX agent

arXiv:2605.25133v1 Announce Type: new Abstract: Reliably knowing when a language model is correct is almost as important as being correct. We introduce prover-verifier deliberation (PVD), an inference

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Truthful Online Preference Aggregation for LLM Fine-Tuning in Mobile Crowdsourcing

DGX agent

arXiv:2605.24052v1 Announce Type: cross Abstract: To better serve users' demands in mobile applications (e.g., navigation), mobile crowdsourcing platforms can iteratively align large language model (L

model-releasesarxiv-cs-ai
26 May 2026
Tutorials

Two-Sided Time-Independent Regret for Matching Markets with Limited Interviews

DGX agent

arXiv:2602.12224v2 Announce Type: replace-cross Abstract: Two-sided matching platforms rely on preferences from both sides, yet participants can evaluate only a small fraction of potential partners. I

tutorialsarxiv-cs-ai
26 May 2026
Research

Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence

DGX agent

arXiv:2509.23573v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to help security analysts manage the surge of cyber threats, automating tasks from vulnerab

researcharxiv-cs-ai
26 May 2026
Safety

Unifying Value Alignment and Assignment in Cross-Domain Offline Reinforcement Learning with Heterogeneous Datasets

DGX agent

arXiv:2605.24862v1 Announce Type: new Abstract: Cross-domain offline reinforcement learning (RL) aims to learn a policy in the target domain with a limited target domain dataset and a source domain da

safetyarxiv-cs-lg
26 May 2026
Model Releases

What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA

DGX agent

arXiv:2605.25988v1 Announce Type: new Abstract: Medical RAG needs evidence-grounded claims, so plugging a claim-level NLI checker into retrieval-augmented RL is intuitive. extbf{We find that the check

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

DGX agent

arXiv:2605.23243v1 Announce Type: cross Abstract: We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

DGX agent

arXiv:2605.23281v1 Announce Type: new Abstract: Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across div

model-releasesarxiv-cs-cv
25 May 2026
Safety

Direct Dynamic Retargeting for Humanoid Imitation Learning from Videos

DGX agent

arXiv:2605.23762v1 Announce Type: new Abstract: Imitation Learning from monocular video demonstrations provides a scalable approach for teaching complex skills to humanoid robots. However, translating

safetyarxiv-cs-ro
25 May 2026
Research

DRL-Driven Edge-Aware Utility Optimization for Multi-Slice 6G Networks

DGX agent

arXiv:2605.23056v1 Announce Type: cross Abstract: Virtual Reality (VR) services delivered over 6G networks demand ultra-low latency and high bandwidth to ensure seamless user experiences. This paper p

researcharxiv-cs-ai
25 May 2026
Model Releases

Evaluating Large Language Models in a Complex Hidden Role Game

DGX agent

arXiv:2605.22826v1 Announce Type: cross Abstract: Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments.

model-releasesarxiv-cs-ai
25 May 2026
Safety

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving

DGX agent

arXiv:2605.23163v1 Announce Type: new Abstract: End-to-end autonomous driving via Vision-Language-Action (VLA) models demands a precarious balance between high-fidelity trajectory planning and efficie

safetyarxiv-cs-cl
25 May 2026
Model Releases

How Far Are We from Generating Missing Modalities with Foundation Models?

DGX agent

arXiv:2506.03530v3 Announce Type: replace-cross Abstract: Multimodal foundation models have demonstrated impressive capabilities across diverse tasks. However, their potential as plug-and-play solutio

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

DGX agent

arXiv:2605.23170v1 Announce Type: cross Abstract: Position-controlled evaluation is standard for retrieval tasks such as Needle-in-a-Haystack and RULER, but mainstream reasoning benchmarks do not cont

model-releasesarxiv-cs-ai
25 May 2026
Safety

SeedER: Seed-and-Expand Retrieval from Knowledge Graphs

DGX agent

arXiv:2605.23753v1 Announce Type: new Abstract: Knowledge graphs (KGs) offer a rich representation for relational knowledge, but their irregular structure makes retrieval challenging: ego-graph expans

safetyarxiv-cs-lg
25 May 2026
Model Releases

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

DGX agent

arXiv:2605.22841v1 Announce Type: cross Abstract: What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty c

model-releasesarxiv-cs-ai
25 May 2026
Applications

The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems

DGX agent

arXiv:2605.23024v1 Announce Type: new Abstract: Large language models now write software, draft legal documents, and produce clinical notes, yet fundamental limits, from Turing and Arrow to the No Fre

applicationsarxiv-cs-ai
25 May 2026
Research

USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots

DGX agent

arXiv:2510.07869v4 Announce Type: replace Abstract: Underwater environments pose unique challenges for robotic navigation and manipulation. While existing research has primarily focused on task-specif

researcharxiv-cs-ro
25 May 2026
Safety

CCLab: Adversarial Testing of Learning- and Non-Learning-Based Congestion Controllers

DGX agent

arXiv:2605.21915v1 Announce Type: cross Abstract: Congestion controllers (CCs) are critical to network performance, and yet their robustness under adverse conditions remains insufficiently understood.

safetyarxiv-cs-lg
23 May 2026
Model Releases

Chebyshev Policies and the Mountain Car Problem: Reinforcement Learning for Low-Dimensional Control Tasks

DGX agent

arXiv:2605.22305v1 Announce Type: new Abstract: We analytically solve the Mountain Car problem, a canonical benchmark in RL, and derive an optimal control solution, closing a gap after 36 years. This

model-releasesarxiv-cs-lg
23 May 2026
Safety

Long-term Fairness with Selective Labels

DGX agent

arXiv:2605.22291v1 Announce Type: new Abstract: Long-term fairness algorithms aim to satisfy fairness beyond static and short-term notions by accounting for the dynamics between decision-making polici

safetyarxiv-cs-lg
23 May 2026
Safety

Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games

DGX agent

arXiv:2602.10894v2 Announce Type: replace Abstract: Two-player games such as board games have long been used as traditional benchmarks for reinforcement learning. This work revisits a policy optimizat

safetyarxiv-cs-lg
23 May 2026
Model Releases

Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability

DGX agent

arXiv:2605.22142v1 Announce Type: new Abstract: Reinforcement learning under partial observability requires deciding what information to retain, yet most memory-based approaches do not explicitly mode

model-releasesarxiv-cs-lg
23 May 2026
← Previous
1…214215216217218…230
Next →