AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “openai”

GridTimelineEvolution
227 results
Model Releases

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

DGX agent

arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, h

model-releasesarxiv-cs-cl
26 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs

DGX agent

arXiv:2507.10593v3 Announce Type: replace-cross Abstract: Every LLM tool call is structurally an RPC -- a function name, JSON arguments, and a serialized result -- yet each protocol (native Python, MC

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering

DGX agent

arXiv:2605.23497v1 Announce Type: new Abstract: Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds

model-releasesarxiv-cs-cl
25 May 2026
Safety

Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control

DGX agent

arXiv:2605.23415v1 Announce Type: cross Abstract: Reinforcement learning has long struggled with poor sample efficiency. One promising approach to mitigate this problem is leveraging group-invariant M

safetyarxiv-cs-ai
25 May 2026
Local Ai

A Mechanistic Explanatory Strategy for XAI

DGX agent

arXiv:2411.01332v5 Announce Type: replace Abstract: Despite significant advancements in XAI, scholars note a persistent lack of solid conceptual foundations and integration with broader scientific dis

local-aiarxiv-cs-lg
23 May 2026
Safety

Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction

DGX agent

arXiv:2605.21653v1 Announce Type: cross Abstract: AI text detectors amplify a pretrained typicality axis; they do not construct an AI-vs-human boundary. On raw encoders before any task supervision, pr

safetyarxiv-cs-cl
22 May 2026
Model Releases

Code Researcher: Deep Research Agent for Large Systems Code and Commit History

DGX agent

arXiv:2506.11060v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code rema

model-releasesarxiv-cs-ai
22 May 2026
Agents

Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling

DGX agent

arXiv:2605.21470v1 Announce Type: new Abstract: Computer-use agents (CUA) automate tasks specified with natural language such as 'order the cheapest item from Taco Bell' by generating sequences of cal

agentsarxiv-cs-lg
21 May 2026
Safety

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents

DGX agent

arXiv:2605.19932v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over long and recurring external contexts, like document corpora and code repositories. Across in

safetyarxiv-cs-ai
20 May 2026
Model Releases

A Machine With Human-Like Memory Systems

DGX agent

arXiv:2204.01611v3 Announce Type: replace Abstract: Inspired by the cognitive science theory, we explicitly model an agent with both semantic and episodic memory systems, and show that it is better th

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Can LLMs Generate and Solve Linguistic Olympiad Puzzles?

DGX agent

arXiv:2509.21820v2 Announce Type: replace Abstract: In this paper, we introduce a combination of novel and exciting tasks: the solution and generation of linguistic puzzles. We focus on puzzles used i

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows

DGX agent

arXiv:2605.18327v1 Announce Type: new Abstract: AI agents deployed into SRE workflows currently derive their understanding of environment state from raw observability telemetry at query time, paying a

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Episodic-Semantic Memory Architecture for Long-Horizon Scientific Agents

DGX agent

arXiv:2605.17625v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into persistent scientific collaborators, context window saturation has emerged as a critical bottleneck. Scienti

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps

DGX agent

arXiv:2605.17554v1 Announce Type: new Abstract: Frontier deep research agents (DRAs) plan a research task, synthesize across documents, and return a structured deliverable on demand. They are being de

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

EvilGenie: A Reward Hacking Benchmark

DGX agent

arXiv:2511.21654v2 Announce Type: replace Abstract: We introduce EvilGenie, a benchmark for reward hacking in programming settings. We source problems from LiveCodeBench and create an environment in w

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

Fidelity Probes for Specification--Code Alignment

DGX agent

arXiv:2605.17246v1 Announce Type: cross Abstract: We introduce fidelity probes: natural-language questions generated from a reference artifact with code-derived ground-truth answers, answered from a c

model-releasesarxiv-cs-ai
19 May 2026
Agents

Transfer Learning for Customized Car Racing Environments

DGX agent

arXiv:2605.17928v1 Announce Type: cross Abstract: Transfer Learning, a technique where a model/agent can use the knowledge/expertise that it gained from one task and exploit that to solve another clos

agentsarxiv-cs-lg
19 May 2026
Model Releases

Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification

DGX agent

arXiv:2605.17691v1 Announce Type: cross Abstract: Automating the classification of negative treatment in legal precedent is a critical yet nuanced NLP task where misclassification carries significant

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

DGX agent

arXiv:2510.16252v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on

model-releasesarxiv-cs-cl
19 May 2026
Safety

When Vision Speaks for Sound

DGX agent

arXiv:2605.16403v1 Announce Type: new Abstract: Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual c

safetyarxiv-cs-cv
19 May 2026
Research

An LLM-RAG Approach for Healthy Eating Index-Informed Personalized Food Recommendations

DGX agent

arXiv:2605.15213v1 Announce Type: cross Abstract: Diet quality is a leading determinant of chronic disease risk. Advances in artificial intelligence (AI) have enabled food recommendation systems to ad

researcharxiv-cs-ai
18 May 2026
Model Releases

Agentic Design of Compositional Descriptors via Autoresearch for Materials Science Applications

DGX agent

arXiv:2605.14671v1 Announce Type: cross Abstract: Autoresearch offers a flexible paradigm for automating scientific tasks, in which an AI agent proposes, implements, evaluates, and refines candidate s

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models

DGX agent

arXiv:2605.14635v1 Announce Type: cross Abstract: This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language

model-releasesarxiv-cs-ai
15 May 2026
Safety

Progent: Securing AI Agents with Privilege Control

DGX agent

arXiv:2504.11703v3 Announce Type: replace-cross Abstract: AI agents interact with external environments through tool calls, exposing them to attacks like indirect prompt injection that can trigger una

safetyarxiv-cs-ai
15 May 2026
Model Releases

From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning

DGX agent

arXiv:2603.26839v2 Announce Type: replace-cross Abstract: How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce ex

model-releasesarxiv-cs-cv
14 May 2026
Research

Just Ask for a Table: A Thirty-Token User Prompt Defeats Sponsored Recommendations in Twelve LLMs

DGX agent

arXiv:2605.12772v1 Announce Type: new Abstract: Wu et al. (2026) showed that most frontier large language models (LLMs) recommend a sponsored, roughly twice-as-expensive flight when their system promp

researcharxiv-cs-cv
14 May 2026
Model Releases

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

DGX agent

arXiv:2605.11086v1 Announce Type: cross Abstract: AI agents are rapidly gaining capabilities that could significantly reshape cybersecurity, making rigorous evaluation urgent. A critical capability is

model-releasesarxiv-cs-lg
13 May 2026
Safety

The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested

DGX agent

arXiv:2605.11496v1 Announce Type: cross Abstract: Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and

safetyarxiv-cs-lg
13 May 2026
Model Releases

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks

DGX agent

arXiv:2605.09611v1 Announce Type: new Abstract: This preprint presents an empirical analysis of byte-exact chunk-level deduplication in Retrieval-Augmented Generation (RAG) pipelines. We measure conte

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

CHAINTRIX: A multi-pipeline LLM-augmented framework for automated smart-contract security auditing

DGX agent

arXiv:2605.09350v1 Announce Type: new Abstract: Smart-contract exploits have caused billions of USD in cumulative losses, yet audits remain expensive and slow. Automated tools have emerged to close th

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction

DGX agent

arXiv:2605.09973v1 Announce Type: cross Abstract: Reliable detection of personally identifiable information (PII) is increasingly important across modern data-processing systems, yet the task remains

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models

DGX agent

arXiv:2510.08592v3 Announce Type: replace-cross Abstract: Test-Time Scaling (TTS) improves LLM reasoning by exploring multiple candidate responses and then operating over this set to find the best out

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark

DGX agent

arXiv:2410.14702v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstra

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology

DGX agent

arXiv:2605.10761v1 Announce Type: new Abstract: Cancer screening is a reasoning task. A radiologist observes findings, compares them to prior scans, integrates clinical context, and reaches a diagnost

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas

DGX agent

arXiv:2605.06673v1 Announce Type: cross Abstract: Aggregate metacognitive quality scores mask within-model variation across MMLU benchmark domains. We administered 1,500 MMLU items (250 per domain, un

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG

DGX agent

arXiv:2605.07273v1 Announce Type: cross Abstract: Multimodal RAG systems increasingly rely on vision-language retrievers to ground visual queries in external textual evidence. Existing adversarial stu

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

LLM-Based Agents for Competitive Landscape Mapping in Drug Asset Due Diligence

DGX agent

arXiv:2508.16571v4 Announce Type: replace Abstract: In this paper, we describe and benchmark a competitor-discovery component used within an agentic AI system for fast drug asset due diligence. A comp

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents

DGX agent

arXiv:2605.03952v1 Announce Type: cross Abstract: Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The chal

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting

DGX agent

arXiv:2605.04426v1 Announce Type: new Abstract: We introduce Telegraph English (TE), a prompt-compression protocol that rewrites natural language into a symbol-rich, formally-structured dialect. Where

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use

DGX agent

arXiv:2605.02964v1 Announce Type: new Abstract: Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomou

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration

DGX agent

arXiv:2605.01970v2 Announce Type: cross Abstract: Memory systems enable otherwise-stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. We characte

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Benchmarking Retrieval Strategies for Biomedical Retrieval-Augmented Generation: A Controlled Empirical Study

DGX agent

arXiv:2605.02520v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) offers a well-established path to grounding large language model (LLM) outputs in external knowledge, yet the quest

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Interpretable experiential learning based on state history and global feedback

DGX agent

arXiv:2605.00940v1 Announce Type: new Abstract: A new interpretable experiential learning model based on state history and global feedback is presented. It is capable of learning a behavioral model re

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces

DGX agent

arXiv:2605.02801v1 Announce Type: new Abstract: As large language model (LLM) agents evolve from isolated tool users into coordinated teams, reinforcement learning (RL) must optimize not only individu

model-releasesarxiv-cs-cl
5 May 2026
Safety

Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts

DGX agent

arXiv:2510.22628v2 Announce Type: replace-cross Abstract: This paper presents a real-time modular defense system named Sentra-Guard. The system detects and mitigates jailbreak and prompt injection att

safetyarxiv-cs-ai
5 May 2026
Safety

Training Non-Differentiable Networks via Optimal Transport

DGX agent

arXiv:2605.01928v1 Announce Type: new Abstract: Neural networks increasingly embed non-differentiable components (spiking neurons, quantized layers, discrete routing, blackbox simulators, etc.) where

safetyarxiv-cs-lg
5 May 2026
Safety

Ultrasound Vision-Language Alignment via Contrastive Learning

DGX agent

arXiv:2605.02126v1 Announce Type: new Abstract: Ultrasound foundation models have achieved strong performance on structured prediction tasks but remain exclusively vision-based, limiting zero-shot and

safetyarxiv-cs-cv
5 May 2026
Model Releases

When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models

DGX agent

arXiv:2605.02363v1 Announce Type: new Abstract: Deployed language models must produce outputs that are both correct and format-compliant. We study this structured-output reliability gap using two math

model-releasesarxiv-cs-cl
5 May 2026
← Previous
12345
Next →