AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

DGX agent

arXiv:2605.17453v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses im

model-releasesarxiv-cs-cl
19 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation

DGX agent

arXiv:2605.17140v1 Announce Type: cross Abstract: Brain tumor diagnosis is largely dependent on Magnetic Resonance Imaging (MRI) evaluation, which requires radiologists to synthesize thousands of imag

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Verify-Gated Completion as Admission Control in a Governed Multi-Agent Runtime: A Bounded Architecture Case Study

DGX agent

arXiv:2605.17998v1 Announce Type: cross Abstract: As multi-agent systems move from short interactions to tool-using workflows with specialized roles and persistent state, completion becomes a runtime-

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Few-Shot Large Language Models for Actionable Triage Categorization of Online Patient Inquiries

DGX agent

arXiv:2605.15680v1 Announce Type: new Abstract: Online patient inquiries are often informal, incomplete, and written before professional assessment, yet they must still be routed to an appropriate lev

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models

DGX agent

arXiv:2605.15589v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in the mental health domain, yet it remains unclear how well they capture related biomedical knowledg

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints

DGX agent

arXiv:2504.11320v3 Announce Type: replace-cross Abstract: Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day. Each request requires toke

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels

DGX agent

arXiv:2605.15208v1 Announce Type: cross Abstract: Large Language Models are routinely compressed via post-training quantization to reduce inference costs and memory footprint for cloud and edge deploy

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

CLOVER: Closed-Loop Value Estimation & Ranking for End-to-End Autonomous Driving Planning

DGX agent

arXiv:2605.15120v1 Announce Type: cross Abstract: End-to-end autonomous driving planners are commonly trained by imitating a single logged trajectory, yet evaluated by rule-based planning metrics that

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Descriptor: Distance-Annotated Traffic Perception Question Answering (DTPQA)

DGX agent

arXiv:2511.13397v2 Announce Type: replace-cross Abstract: The remarkable progress of Vision-Language Models (VLMs) on a variety of tasks has raised interest in their application to automated driving.

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Fusion-fission forecasts when AI will shift to undesirable behavior

DGX agent

arXiv:2605.14218v1 Announce Type: new Abstract: The key problem facing ChatGPT-like AI's use across society is that its behavior can shift, unnoticed, from desirable to undesirable -- encouraging self

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

DGX agent

arXiv:2605.14237v1 Announce Type: new Abstract: Deploying AI agents for repetitive periodic tasks exposes a critical tension: Large Language Models (LLMs) offer unmatched flexibility in tool orchestra

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

NodeSynth: Socially Aligned Synthetic Data for AI Evaluation

DGX agent

arXiv:2605.14381v1 Announce Type: cross Abstract: Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, thes

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

DGX agent

arXiv:2605.15118v1 Announce Type: cross Abstract: We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4imes6 Target imes Technique mat

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

The Moltbook Observatory Archive: an incremental dataset of agent-only social network activity

DGX agent

arXiv:2605.13860v1 Announce Type: cross Abstract: Moltbook is a social media platform in which posts and comments are authored exclusively by autonomous AI agents. We present the Moltbook Observatory

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Vision-Based Runtime Monitoring under Varying Specifications using Semantic Latent Representations

DGX agent

arXiv:2605.13923v1 Announce Type: cross Abstract: We study certified runtime monitoring of past-time signal temporal logic (ptSTL) from visual observations under partial observability. The monitor mus

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Action Emergence from Streaming Intent

DGX agent

arXiv:2605.12622v1 Announce Type: cross Abstract: We formalize action emergence as a target capability for end-to-end autonomous driving: the ability to generate physically feasible, semantically appr

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

Connecting the Dots: A Machine Learning Ready Dataset for Ionospheric Forecasting Models

DGX agent

arXiv:2511.15743v2 Announce Type: replace Abstract: Operational forecasting of the ionosphere remains a critical space weather challenge due to sparse observations, complex coupling across geospatial

model-releasesarxiv-cs-lg
14 May 2026
Model Releases

Negation Neglect: When models fail to learn negations in training

DGX agent

arXiv:2605.13829v1 Announce Type: cross Abstract: We introduce Negation Neglect, where finetuning LLMs on documents that flag a claim as false makes them believe the claim is true. For example, models

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

DGX agent

arXiv:2605.11398v1 Announce Type: cross Abstract: We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations.

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor

DGX agent

arXiv:2601.05752v3 Announce Type: replace Abstract: We introduce AutoMonitor-Bench, the first benchmark designed to systematically evaluate the reliability of LLM-based misbehavior monitors across div

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Crash Assessment via Mesh-Based Graph Neural Networks and Physics-Aware Attention

DGX agent

arXiv:2605.11784v1 Announce Type: cross Abstract: Full-vehicle crash simulations are computationally expensive, limiting their use in iterative design exploration. This work investigates learned hybri

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Multi-Task Representation Learning for Conservative Linear Bandits

DGX agent

arXiv:2605.12176v1 Announce Type: new Abstract: This paper presents the Constrained Multi-Task Representation Learning (CMTRL) framework for linear bandits. We consider T linear bandit tasks in a d di

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models

DGX agent

arXiv:2605.11887v1 Announce Type: new Abstract: Large language models have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque, li

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Robust Promptable Video Object Segmentation

DGX agent

arXiv:2605.12006v1 Announce Type: new Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

Urban Risk-Aware Navigation via VQA-Based Event Maps for People with Low Vision

DGX agent

arXiv:2605.11782v1 Announce Type: new Abstract: Visual impairment affects hundreds of millions of people worldwide, severely limiting their ability to navigate urban environments safely and independen

model-releasesarxiv-cs-cv
13 May 2026
Model Releases

Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling

DGX agent

arXiv:2604.08178v2 Announce Type: replace Abstract: In classical Reinforcement Learning from Human Feedback (RLHF), Reward Models (RMs) serve as the fundamental signal provider for model alignment. As

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour

DGX agent

arXiv:2506.12090v2 Announce Type: replace Abstract: This paper introduces ChatbotManip, a novel dataset for studying manipulation in Chatbots. It contains simulated generated conversations between a c

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Decomposing and Steering Functional Metacognition in Large Language Models

DGX agent

arXiv:2605.08942v1 Announce Type: new Abstract: Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

GONE: Structural Knowledge Unlearning via Neighborhood-Expanded Distribution Shaping

DGX agent

arXiv:2603.12275v1 Announce Type: cross Abstract: Unlearning knowledge is a pressing and challenging task in Large Language Models (LLMs) because of their unprecedented capability to memorize and dige

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs

DGX agent

arXiv:2508.20325v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly integral to various domains, their potential to generate harmful responses has prompted si

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving

DGX agent

arXiv:2605.09972v1 Announce Type: cross Abstract: End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving

model-releasesarxiv-cs-cv
12 May 2026
Local Ai

Interactive Critique-Revision Training for Reliable Structured LLM Generation

DGX agent

arXiv:2605.08327v1 Announce Type: cross Abstract: In structured decision-making workflows such as form filling, compliance checking, and maintenance reporting, LLM outputs must be locally correct, glo

local-aiarxiv-cs-ai
12 May 2026
Model Releases

Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models

DGX agent

arXiv:2510.08592v3 Announce Type: replace-cross Abstract: Test-Time Scaling (TTS) improves LLM reasoning by exploring multiple candidate responses and then operating over this set to find the best out

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

DGX agent

arXiv:2605.08445v1 Announce Type: new Abstract: AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

DGX agent

arXiv:2605.10002v1 Announce Type: new Abstract: Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet inco

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

DGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

DGX agent

arXiv:2605.10639v1 Announce Type: new Abstract: The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluati

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Phoenix-VL 1.5 Medium Technical Report

DGX agent

arXiv:2605.10391v1 Announce Type: cross Abstract: We introduce Phoenix-VL 1.5 Medium, a 123B-parameter natively multimodal and multilingual foundation model, adapted to regional languages and the Sing

model-releasesarxiv-cs-ai
12 May 2026
Local Ai

Playing games with knowledge: AI-Induced delusions need game theoretic interventions

DGX agent

arXiv:2605.08409v1 Announce Type: new Abstract: Conversational AI has a fundamental flaw as a knowledge interface: sycophantic chatbots induce epistemic entrenchment and delusional belief spirals even

local-aiarxiv-cs-ai
12 May 2026
Model Releases

Position: AI Security Policy Should Target Systems, Not Models

DGX agent

arXiv:2605.09504v1 Announce Type: cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared memory, paral

model-releasesarxiv-cs-ai
12 May 2026
Local Ai

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents

DGX agent

arXiv:2605.08468v1 Announce Type: cross Abstract: Local LLM-based coding agents increasingly work in settings where correctness is earned through execution feedback, persistent state, and bounded repa

local-aiarxiv-cs-ai
12 May 2026
Model Releases

Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality

DGX agent

arXiv:2605.10142v1 Announce Type: cross Abstract: Artificial intelligence models are increasingly scaled to improve predictive accuracy, yet it remains unclear whether scale improves the quality of po

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs

DGX agent

arXiv:2509.02372v3 Announce Type: replace-cross Abstract: Large Language Models have become critical to modern software development, but their reliance on uncurated web-scale datasets for training int

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

DGX agent

arXiv:2604.02608v2 Announce Type: replace Abstract: Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both stee

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

The Echo Amplifies the Knowledge: Somatic Marker Analogues in Language Models via Emotion Vector Re-Injection

DGX agent

arXiv:2605.08611v1 Announce Type: new Abstract: Current language model memory systems store what happened but not how it felt. This distinction -- between semantic memory (knowing about a past event)

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

DGX agent

arXiv:2605.08737v1 Announce Type: cross Abstract: On-policy distillation (OPD) is widely used for LLM post-training. When pushed with a reward-extrapolation coefficient lambda > 1, the student can lif

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering

DGX agent

arXiv:2601.20164v2 Announce Type: replace-cross Abstract: Prior work suggests that language models, while trained on next token prediction, show implicit planning behavior: they may select the next to

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Beyond the Black Box: Interpretability of Agentic AI Tool Use

DGX agent

arXiv:2605.06890v1 Announce Type: new Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagn

model-releasesarxiv-cs-ai
11 May 2026
← Previous
1…249250251252253…255
Next →