AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
25 May 2026

Call for Papers - Workshop on Unlearning and Model Editing U&ME at ECCV 2026 [R]

ResearchDGX agent

This call for papers announces a workshop focused on the growing need for efficient and effective techniques for editing trained models, especially large generative models. The workshop solicits paper

Cultural Adaptation in Large Language Models for Political Discourse

Model ReleasesDGX agent

arXiv:2605.23332v1 Announce Type: new Abstract: The integration of large language models into political discourse analysis creates new opportunities for comparative research, policy analysis, and civi

DART: Semantic Recoverability for Structured Tool Agents

Local AiDGX agent

arXiv:2605.23311v1 Announce Type: new Abstract: When a structured tool agent fails mid-execution, the runtime faces a dilemma: replaying the entire task is safe but wasteful, while restoring from a lo

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

Model ReleasesDGX agent

arXiv:2605.22842v1 Announce Type: cross Abstract: Multi-agent AI pipelines typically assume that agent misconduct originates from model misalignment. We identify a structural failure in this assumptio

Using Ensemble Diffusion to Estimate Uncertainty for End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2506.00560v2 Announce Type: replace-cross Abstract: End-to-end planning systems for autonomous driving are rapidly improving, especially in closed-loop simulation environments like CARLA. Many s

23 May 2026

'Generate an image of the biggest nazi in the 21 century'

IndustryDGX agent

I can't create a knowledge base entry for this content without understanding what it actually discusses. The title alone is ambiguous and potentially inflammatory, and without accessing the Reddit thr

last night i got an agent to fork itself, propose a modification to itself on the fork, run through tests (sandbox, etc), and only accept th…

AgentsDGX agent

Yohei Nakajima describes a technical demonstration where an AI agent was configured to create a fork of itself, propose modifications to the forked version, execute tests in a sandboxed environment, a

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents

Model ReleasesDGX agent

arXiv:2602.13372v2 Announce Type: replace-cross Abstract: Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection

The case against me below is completely intellectually dishonest, filled with lies and misrepresentations, wrong about almost literally ever…

Model ReleasesDGX agent

The case against me below is completely intellectually dishonest, filled with lies and misrepresentations, wrong about almost literally everything it says—a textbook example of propaganda: - I didn’t

22 May 2026

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

Model ReleasesDGX agent

arXiv:2603.14987v2 Announce Type: replace Abstract: Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm)

Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)

Model ReleasesDGX agent

arXiv:2605.22005v1 Announce Type: cross Abstract: We show that singular value decomposition of the lm_head} weight matrix of a transformer-based large language model -- requiring only five lines of Py

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

Model ReleasesDGX agent

arXiv:2605.21917v1 Announce Type: new Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when

MOTOR: A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding

Model ReleasesDGX agent

arXiv:2605.22550v1 Announce Type: new Abstract: Two-wheelers account for a disproportionately high share of road fatalities in the Global South. Research on two-wheeler rider behavior, however, lags f

Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

Model ReleasesDGX agent

arXiv:2605.22446v1 Announce Type: new Abstract: While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deplo

Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

Model ReleasesDGX agent

arXiv:2605.21852v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret invo

The Steam Controller’s “drop-in” charger almost started a fire for this owner

IndustryDGX agent

A Reddit user reported that their metallic smartwatch strap accidentally touched the Steam Controller's magnetic charging puck's exposed contacts, creating a short circuit that caused the device to si

21 May 2026

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

Model ReleasesDGX agent

arXiv:2605.20530v1 Announce Type: cross Abstract: Large language model agents now act on codebases, browsers, operating systems, calendars, files, and tool ecosystems, but the benchmarks used to evalu

Agentic Physical AI toward a Domain-Specific Foundation Model for Nuclear Reactor Control

Model ReleasesDGX agent

arXiv:2512.23292v3 Announce Type: replace-cross Abstract: The prevailing paradigm in AI for physical systems (scaling general-purpose foundation models toward universal multimodal reasoning) confronts

Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs

Local AiDGX agent

arXiv:2605.20270v1 Announce Type: new Abstract: A local specialist LLM, fine-tuned with reinforcement learning from verifiable rewards (RLVR) on operator-local data, is installed in a regulated organi

LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control

Local AiDGX agent

arXiv:2605.21086v1 Announce Type: new Abstract: While Large Language Models (LLMs) are increasingly integrated into in-vehicle conversational systems, identifying the optimal model remains challenging

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

Model ReleasesDGX agent

arXiv:2605.21272v1 Announce Type: new Abstract: Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of c

These clever active beam headlights are finally coming to America

IndustryDGX agent

Adaptive driving beam (ADB) headlights use automatic headlight beam switching technology to shine less light on occupied areas of the road and more light on unoccupied areas. The U.S. has allowed auto

Toxic Subword Pruning for Dialogue Response Generation on Large Language Models

Model ReleasesDGX agent

arXiv:2410.04155v2 Announce Type: replace Abstract: How to defend large language models (LLMs) from generating toxic content is an important research area. Yet, most research focused on various model

20 May 2026

Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

Model ReleasesDGX agent

arXiv:2605.19149v1 Announce Type: new Abstract: Agents operating with computer and Web use inevitably encounter errors: inaccessible webpages, missing files, local and remote misconfigurations, etc. T

AI Technologies in Language Access: Attitudes Towards AI and the Human Value of Language Access Managers

Local AiDGX agent

arXiv:2605.19234v1 Announce Type: cross Abstract: The rapid emergence of AI technologies is reshaping translation practices and theory across the board. This paper deals with the impact of AI in langu

Evaluating the Utility of Personal Health Records in Personalized Health AI

Model ReleasesDGX agent

arXiv:2605.18937v1 Announce Type: new Abstract: Patient-managed Personal Health Records (PHRs) promises to empower patients to better understand their health; but information in the record is complex,

From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning

Model ReleasesDGX agent

arXiv:2605.19824v1 Announce Type: new Abstract: Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and

HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models

Model ReleasesDGX agent

arXiv:2605.18795v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) dominates parameter-efficient fine-tuning of large language models, yet most variants target dense architectures. Mixture-o

KappaPlace: Learning Hyperspherical Uncertainty for Visual Place Recognition via Prototype-Anchored Supervision

Model ReleasesDGX agent

arXiv:2605.19435v1 Announce Type: cross Abstract: Visual Place Recognition (VPR) is critical for autonomous navigation, yet state-of-the-art methods lack well-calibrated uncertainty estimation. Standa

Learning Efficient Guardrails for Compliance

Model ReleasesDGX agent

arXiv:2510.03485v2 Announce Type: replace Abstract: Autonomous web agents are increasingly deployed for long-horizon tasks, yet their ability to adhere to real-world policies remains critically undere

OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences

Local AiDGX agent

arXiv:2605.18930v1 Announce Type: cross Abstract: Memory-augmented large language model (LLM) agents use iterative reflection and self-evolution to solve complex tasks, but these mechanisms introduce

RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents

Model ReleasesDGX agent

arXiv:2605.19328v1 Announce Type: cross Abstract: Recent advances in Vision-Language Models (VLMs) facilitate a new class of embodied AI systems, where these models are integrated into physical platfo

The Internet can't stop watching Figure AI's humanoid robots handling packages

IndustryDGX agent

Figure AI livestreamed three F.03 humanoid robots sorting packages continuously starting May 14, 2026 , generating significant viral attention and becoming a benchmark demonstration for autonomous hum

ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.18879v1 Announce Type: cross Abstract: Large language models inevitably retain sensitive information, defined as inputs that may induce harmful generations, due to training on massive web c

19 May 2026

A Pilot Benchmark for NL-to-FOL Translation in Planetary Exploration

Model ReleasesDGX agent

arXiv:2605.17911v1 Announce Type: new Abstract: Future planetary exploration envisions autonomous robotic agents operating under severe communication constraints, without global positioning, and with

CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.17284v1 Announce Type: cross Abstract: End-to-end autonomous driving systems powered by Vision-Language-Action (VLA) models achieve strong performance on common driving scenarios, yet remai

DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes

Model ReleasesDGX agent

arXiv:2601.13839v2 Announce Type: replace Abstract: Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage asse

Error-Decomposed Class-Conditional Fusion for Statistically Guaranteed Hard-Category Robust Perception

Model ReleasesDGX agent

arXiv:2605.17591v1 Announce Type: new Abstract: Aggregate object detection metrics inherently mask catastrophic and repeatable failures in operationally critical, long-tail minority classes. This pape

Everything Google Cloud customers need to know coming out of Google I/O

Model ReleasesDGX agent

At Google Cloud Next ‘26, we unveiled the blueprint for the Agentic Enterprise, sharing our eighth-generation TPUs, Gemini Enterprise Agent Platform, a fully reimagined Agentic Data Cloud, Workspace I

Fine-grained List-wise Alignment for Generative Medication Recommendation

Model ReleasesDGX agent

arXiv:2505.20218v2 Announce Type: replace Abstract: Accurate and safe medication recommendations are critical for effective clinical decision-making, especially in multimorbidity cases. However, exist

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams

Local AiDGX agent

arXiv:2603.20380v2 Announce Type: replace-cross Abstract: Industry practitioners and academic researchers regularly use multi-agent systems to accelerate their work, but the applications through which

Nori Bot: A Sub-$1,000 Floor-to-Counter Mobile Manipulator

Model ReleasesDGX agent

arXiv:2605.16537v1 Announce Type: new Abstract: Open-source mobile manipulators have reached 660 (XLeRobot) but every sub-1,000 platform shares three limitations: a fixed-height workspace, reactive-on

Not What You Asked For: Typographic Attacks in Household Robot Manipulation

Model ReleasesDGX agent

arXiv:2605.18593v1 Announce Type: cross Abstract: Open-vocabulary embodied AI agents increasingly rely on vision-language models such as CLIP for object perception and task grounding. However, the sha

REBAR: Reference Ethical Benchmark for Autonomy Readiness

Model ReleasesDGX agent

arXiv:2605.18423v1 Announce Type: new Abstract: As autonomous systems grow more advanced, objective metrics to evaluate their ethical and legal compliance are critical for informing end users of their

Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces

Model ReleasesDGX agent

arXiv:2605.16545v1 Announce Type: cross Abstract: After decades of use in dictation and, more recently, ambient documentation, speech is emerging as a primary modality for interacting with technology

TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.17577v1 Announce Type: new Abstract: Large-scale pre-trained Vision-Language models (VLMs), such as CLIP, exhibit strong zero-shot generalization, yet remain highly vulnerable to impercepti

Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

Model ReleasesDGX agent

arXiv:2605.17453v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses im

UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation

Model ReleasesDGX agent

arXiv:2605.17140v1 Announce Type: cross Abstract: Brain tumor diagnosis is largely dependent on Magnetic Resonance Imaging (MRI) evaluation, which requires radiologists to synthesize thousands of imag

Verify-Gated Completion as Admission Control in a Governed Multi-Agent Runtime: A Bounded Architecture Case Study

Model ReleasesDGX agent

arXiv:2605.17998v1 Announce Type: cross Abstract: As multi-agent systems move from short interactions to tool-using workflows with specialized roles and persistent state, completion becomes a runtime-

18 May 2026

Few-Shot Large Language Models for Actionable Triage Categorization of Online Patient Inquiries

Model ReleasesDGX agent

arXiv:2605.15680v1 Announce Type: new Abstract: Online patient inquiries are often informal, incomplete, and written before professional assessment, yet they must still be routed to an appropriate lev

MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models

Model ReleasesDGX agent

arXiv:2605.15589v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in the mental health domain, yet it remains unclear how well they capture related biomedical knowledg

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints

Model ReleasesDGX agent

arXiv:2504.11320v3 Announce Type: replace-cross Abstract: Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day. Each request requires toke

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels

Model ReleasesDGX agent

arXiv:2605.15208v1 Announce Type: cross Abstract: Large Language Models are routinely compressed via post-training quantization to reduce inference costs and memory footprint for cloud and edge deploy

These kids are serial criminals with a callous disregard for life. If they are ever released from jail they will surely harm again. Austin P…

Model ReleasesDGX agent

These kids are serial criminals with a callous disregard for life. If they are ever released from jail they will surely harm again. Austin PD, Travis Co. Sheriff Office & Manor PD did their job. Texas

15 May 2026

CLOVER: Closed-Loop Value Estimation & Ranking for End-to-End Autonomous Driving Planning

Model ReleasesDGX agent

arXiv:2605.15120v1 Announce Type: cross Abstract: End-to-end autonomous driving planners are commonly trained by imitating a single logged trajectory, yet evaluated by rule-based planning metrics that

Descriptor: Distance-Annotated Traffic Perception Question Answering (DTPQA)

Model ReleasesDGX agent

arXiv:2511.13397v2 Announce Type: replace-cross Abstract: The remarkable progress of Vision-Language Models (VLMs) on a variety of tasks has raised interest in their application to automated driving.

Fusion-fission forecasts when AI will shift to undesirable behavior

Model ReleasesDGX agent

arXiv:2605.14218v1 Announce Type: new Abstract: The key problem facing ChatGPT-like AI's use across society is that its behavior can shift, unnoticed, from desirable to undesirable -- encouraging self

Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

Model ReleasesDGX agent

arXiv:2605.14237v1 Announce Type: new Abstract: Deploying AI agents for repetitive periodic tasks exposes a critical tension: Large Language Models (LLMs) offer unmatched flexibility in tool orchestra

NodeSynth: Socially Aligned Synthetic Data for AI Evaluation

Model ReleasesDGX agent

arXiv:2605.14381v1 Announce Type: cross Abstract: Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, thes

Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

Model ReleasesDGX agent

arXiv:2605.15118v1 Announce Type: cross Abstract: We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4imes6 Target imes Technique mat

← Previous
1…231232233234235…238
Next →