AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
28 Apr 2026

Beyond Context: Large Language Models' Failure to Grasp Users' Intent

Model ReleasesDGX agent

arXiv:2512.21110v3 Announce Type: replace Abstract: Current Large Language Models (LLMs) safety approaches focus on explicitly harmful content while overlooking a critical vulnerability: the inability

Context-Aware Hospitalization Forecasting Evaluations for Decision Support using LLMs

SafetyDGX agent

arXiv:2604.23949v1 Announce Type: new Abstract: Medical and public health experts must make real-time resource decisions, such as expanding hospital bed capacity, based on projected hospitalization tr

Designing escalation criteria for international AI incident response: criteria, triggers, and thresholds

SafetyDGX agent

arXiv:2604.23183v1 Announce Type: cross Abstract: AI incident reporting requirements are emerging in regulation and policy, yet no operational criteria exist for determining when a detected AI inciden

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage

SafetyDGX agent

arXiv:2511.22177v2 Announce Type: replace-cross Abstract: Most post-training methods for text-to-image samplers focus on model weights: either fine-tuning the backbone for alignment or distilling it f

From Stateless Queries to Autonomous Actions: A Layered Security Framework for Agentic AI Systems

SafetyDGX agent

arXiv:2604.23338v1 Announce Type: cross Abstract: Agentic AI systems face security challenges that stateless large language models do not. They plan across extended horizons, maintain persistent memor

Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents

SafetyDGX agent

arXiv:2604.24686v1 Announce Type: new Abstract: Autonomous AI agents can remain fully authorized and still become unsafe as behavior drifts, adversaries adapt, and decision patterns shift without any

LLM-Auction: Generative Auction towards LLM-Native Advertising

SafetyDGX agent

arXiv:2512.10551v2 Announce Type: replace-cross Abstract: The commercialization of LLM applications is the next frontier in online advertising, with LLM-native advertising emerging as a promising para

Polychromic Objectives for Reinforcement Learning

SafetyDGX agent

arXiv:2509.25424v5 Announce Type: replace-cross Abstract: Reinforcement learning fine-tuning (RLFT) is a dominant paradigm for improving pretrained policies for downstream tasks. These pretrained poli

Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs

SafetyDGX agent

arXiv:2604.04177v2 Announce Type: replace Abstract: As large language models (LLMs) are increasing integrated into fact-checking pipelines, formal logic is often proposed as a rigorous means by which

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation

SafetyDGX agent

arXiv:2604.23099v1 Announce Type: cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models

Protecting the Trace: A Principled Black-Box Approach Against Distillation Attacks

SafetyDGX agent

arXiv:2604.23238v1 Announce Type: cross Abstract: Frontier models push the boundaries of what is learnable at extreme computational costs, yet distillation via sampling reasoning traces exposes closed

Safe Navigation in Unknown and Cluttered Environments via Direction-Aware Convex Free-Region Generation

SafetyDGX agent

arXiv:2604.23648v1 Announce Type: new Abstract: Convex free regions provide a structured and optimization-friendly representation of collision-free space for robot navigation in unknown and cluttered

Security Considerations for Multi-agent Systems

SafetyDGX agent

arXiv:2603.09002v2 Announce Type: replace-cross Abstract: Multi-agent artificial intelligence systems or MAS are systems of autonomous agents that exercise delegated tool authority, share persistent m

Sliding Mode Control for Safe Trajectory Tracking with Moving Obstacles Avoidance: Experimental Validation on Planar Robots

SafetyDGX agent

arXiv:2604.24518v1 Announce Type: cross Abstract: This paper presents a unified control framework for robust trajectory tracking and moving obstacle avoidance applicable to a broad class of mobile rob

Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning

SafetyDGX agent

arXiv:2511.01490v2 Announce Type: replace Abstract: As synthetic data becomes widely used in language model development, understanding its impact on model behavior is crucial. This paper investigates

Training a General Purpose Automated Red Teaming Model

SafetyDGX agent

arXiv:2604.23067v1 Announce Type: cross Abstract: Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing fo

Value Alignment Tax: Measuring Value Trade-offs in LLM Alignment

SafetyDGX agent

arXiv:2602.12134v2 Announce Type: replace Abstract: Existing work on value alignment typically characterizes value relations statically, ignoring how alignment interventions, such as prompting, fine-t

27 Apr 2026

Adaptive vs. Static Robot-to-Human Handover: A Study on Orientation and Approach Direction

SafetyDGX agent

arXiv:2604.22378v1 Announce Type: new Abstract: Robot-to-human handovers often rely on static, open-loop strategies (or, at best, approaches that adapt only the position), which generally do not consi

AgentBound: Securing Execution Boundaries of AI Agents

SafetyDGX agent

arXiv:2510.21236v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have evolved into AI agents that interact with external tools and environments to perform complex tasks. The Mode

Controllable Spoken Dialogue Generation: An LLM-Driven Grading System for K-12 Non-Native English Learners

SafetyDGX agent

arXiv:2604.22542v1 Announce Type: cross Abstract: Large language models (LLMs) often fail to meet the pedagogical needs of K-12 English learners in non-native contexts due to a proficiency mismatch. T

Don't try to build a self-improving AI agent without evals. You are just wasting time and compute. An agent can't improve from traces it can…

SafetyDGX agent

Don't try to build a self-improving AI agent without evals. You are just wasting time and compute. An agent can't improve from traces it can't evaluate. This is why it's exciting to see @FutureAGI_ go

Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework

SafetyDGX agent

arXiv:2604.22119v1 Announce Type: new Abstract: As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve their own ob

Sam Altman cannot be trusted. • The OpenAI board fired him because he was not always honest with them. They said he should not control power…

SafetyDGX agent

Sam Altman cannot be trusted. • The OpenAI board fired him because he was not always honest with them. They said he should not control powerful AI. • A major report talked to over 100 people and saw s

24 Apr 2026

Bounding the Black Box: A Statistical Certification Framework for AI Risk Regulation

SafetyDGX agent

arXiv:2604.21854v1 Announce Type: new Abstract: Artificial intelligence now decides who receives a loan, who is flagged for criminal investigation, and whether an autonomous vehicle brakes in time. Go

Enabling and Inhibitory Pathways of University Students' Willingness to Disclose AI Use: A Cognition-Affect-Conation Perspective

SafetyDGX agent

arXiv:2604.21733v1 Announce Type: new Abstract: The increasing integration of artificial intelligence (AI) in higher education has raised important questions regarding students' transparency in report

HARBOR: Automated Harness Optimization

SafetyDGX agent

arXiv:2604.20938v1 Announce Type: cross Abstract: Long-horizon language-model agents are dominated, in lines of code and in operational complexity, not by their underlying model but by the harness tha

Lawmakers and lobbyists say the Trump administration has lobbied against legislation that would regulate AI in at least six Republican-led states (Amrith Ramkumar/Wall Street Journal)

SafetyDGX agent

Amrith Ramkumar / Wall Street Journal: Lawmakers and lobbyists say the Trump administration has lobbied against legislation that would regulate AI in at least six Republican-led states — ‘I am disappo

Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review

SafetyDGX agent

arXiv:2603.18740v2 Announce Type: replace-cross Abstract: Automated Code Review (ACR) systems integrating Large Language Models (LLMs) are increasingly adopted in software development workflows, rangi

Mind the Prompt: Self-adaptive Generation of Task Plan Explanations via LLMs

SafetyDGX agent

arXiv:2604.21092v1 Announce Type: new Abstract: Integrating Large Language Models (LLMs) into complex software systems enables the generation of human-understandable explanations of opaque AI processe

Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

SafetyDGX agent

arXiv:2603.12845v2 Announce Type: replace Abstract: Predicting enzyme kinetic parameters quantifies how efficiently an enzyme catalyzes a specific substrate under defined biochemical conditions. Canon

Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build …

SafetyDGX agent

Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build upon them and use them to evaluate the monitorability of the

23 Apr 2026

AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation

SafetyDGX agent

arXiv:2604.20134v1 Announce Type: cross Abstract: Security Operations Centers (SOCs) increasingly encounter difficulties in correlating heterogeneous alerts, interpreting multi-stage attack progressio

Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents

SafetyDGX agent

arXiv:2604.20572v1 Announce Type: new Abstract: Online lifelong learning enables agents to accumulate experience across interactions and continually improve on long-horizon tasks. However, existing me

FLOSS: Federated Learning with Opt-Out and Straggler Support

SafetyDGX agent

arXiv:2507.23115v2 Announce Type: replace-cross Abstract: Previous work on data privacy in federated learning systems focuses on privacy-preserving operations for data from users who have agreed to sh

LLMs Can Get 'Brain Rot': A Pilot Study on Twitter/X

SafetyDGX agent

arXiv:2510.13928v2 Announce Type: replace-cross Abstract: We propose and test the LLM Brain Rot Hypothesis: continual exposure to junk web text induces lasting cognitive decline in large language mode

MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment

SafetyDGX agent

arXiv:2604.20685v1 Announce Type: new Abstract: Aligning large language models (LLMs) to desirable human values requires balancing multiple, potentially conflicting objectives such as helpfulness, tru

SAMix: Calibrated and Accurate Continual Learning via Sphere-Adaptive Mixup and Neural Collapse

SafetyDGX agent

arXiv:2510.15751v2 Announce Type: replace Abstract: While most continual learning methods focus on mitigating forgetting and improving accuracy, they often overlook the critical aspect of network cali

Verification of Machine Unlearning is Fragile

SafetyDGX agent

arXiv:2408.00929v2 Announce Type: replace Abstract: As privacy concerns escalate in the realm of machine learning, data owners now have the option to utilize machine unlearning to remove their data fr

22 Apr 2026

BAPO: Boundary-Aware Policy Optimization for Reliable Agentic Search

SafetyDGX agent

arXiv:2601.11037v2 Announce Type: replace Abstract: RL-based agentic search enables LLMs to solve complex questions via dynamic planning and external search. While this approach significantly enhances

Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications

SafetyDGX agent

arXiv:2604.19281v1 Announce Type: cross Abstract: The use of Large Language Models (LLMs) to support patients in addressing medical questions is becoming increasingly prevalent. However, most of the m

Counting Worlds Branching Time Semantics for post-hoc Bias Mitigation in generative AI

SafetyDGX agent

arXiv:2604.19431v1 Announce Type: cross Abstract: Generative AI systems are known to amplify biases present in their training data. While several inference-time mitigation strategies have been propose

How to Teach Large Multimodal Models New Skills

SafetyDGX agent

arXiv:2510.08564v2 Announce Type: replace Abstract: How can we teach large multimodal models (LMMs) new skills without erasing prior abilities? We study sequential fine-tuning on five target skills wh

Hybrid Task and Motion Planning with Reactive Collision Handling for Multi-Robot Disassembly of Complex Products: Application to EV Batteries

SafetyDGX agent

arXiv:2509.21020v2 Announce Type: replace Abstract: This paper addresses the problem of multi-robot coordination for complex manipulation task sequences. We present a vision-driven task-and-motion pla

LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit

SafetyDGX agent

arXiv:2604.19117v1 Announce Type: new Abstract: When a language model agrees with a user's false belief, is it failing to detect the error, or noticing and agreeing anyway? We show the latter. Across

Multimodal embodiment-aware navigation transformer

SafetyDGX agent

arXiv:2604.19267v1 Announce Type: new Abstract: Goal-conditioned navigation models for ground robots trained using supervised learning show promising zero-shot transfer, but their collision-avoidance

Neuromorphic Continual Learning for Sequential Deployment of Nuclear Plant Monitoring Systems

SafetyDGX agent

arXiv:2604.18611v1 Announce Type: cross Abstract: Anomaly detection in nuclear industrial control systems (ICS) requires continuous, energy-efficient monitoring across multiple subsystems that are oft

One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization

SafetyDGX agent

arXiv:2601.18572v2 Announce Type: replace Abstract: Personalization of LLMs by sociodemographic subgroup often improves user experience, but can also introduce or amplify biases and unfair outcomes ac

Our newsroom AI policy

SafetyDGX agent

Ars Technica's newsroom AI policy forbids unlabeled AI material in reported stories and requires human confirmation of every quotation's accuracy. The publication does not permit the publication of AI

Safe Continual Reinforcement Learning in Non-stationary Environments

Model ReleasesDGX agent

arXiv:2604.19737v1 Announce Type: new Abstract: Reinforcement learning (RL) offers a compelling data-driven paradigm for synthesizing controllers for complex systems when accurate physical models are

SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention

SafetyDGX agent

arXiv:2506.14387v3 Announce Type: replace Abstract: Adapting LLMs with new knowledge is increasingly important, but standard fine-tuning often erodes aligned epistemic abstention: the ability to ackno

The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation

SafetyDGX agent

arXiv:2604.19064v1 Announce Type: new Abstract: In vision-and-language navigation (VLN), self-improvement from policy-induced experience, using only standard VLN action supervision, critically depends

VIGIL: An Extensible System for Real-Time Detection and Mitigation of Cognitive Bias Triggers

SafetyDGX agent

arXiv:2604.03261v2 Announce Type: replace Abstract: The rise of generative AI is posing increasing risks to online information integrity and civic discourse. Most concretely, such risks can materialis

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving

SafetyDGX agent

arXiv:2411.18275v2 Announce Type: replace Abstract: Vision-language models (VLMs) have significantly advanced autonomous driving (AD) by enhancing reasoning capabilities. However, these models remain

Weakly supervised framework for wildlife detection and counting in challenging Arctic environments: a case study on caribou (Rangifer tarandus)

SafetyDGX agent

arXiv:2601.18891v3 Announce Type: replace Abstract: Caribou across the Arctic has declined in recent decades, motivating scalable and accurate monitoring approaches to guide evidence-based conservatio

21 Apr 2026

A Hamilton-Jacobi Reachability-Guided Search Framework for Efficient and Safe Indoor Planar Robot Navigation

SafetyDGX agent

arXiv:2604.17679v1 Announce Type: new Abstract: Autonomous navigation requires planning to reach a goal safely and efficiently in complex and potentially dynamic environments. Graph search-based algor

I agree with Connor that most folks concerned about AI have been far too coy about extinction risk, and that ControlAI is one of few excepti…

SafetyDGX agent

I agree with Connor that most folks concerned about AI have been far too coy about extinction risk, and that ControlAI is one of few exceptions. The world needs more earnest communication efforts. We

Navigating the Conceptual Multiverse

SafetyDGX agent

arXiv:2604.17815v1 Announce Type: cross Abstract: When language models answer open-ended problems, they implicitly make hidden decisions that shape their outputs, leaving users with uncontextualized a

One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment

SafetyDGX agent

arXiv:2601.18731v2 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individu

Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding

Model ReleasesDGX agent

arXiv:2604.17132v1 Announce Type: new Abstract: Safety-aligned large language models (LLMs) often generate refusal responses to harmless queries due to the over-refusal problem. However, existing meth

PowerCLIP: Powerset Alignment for Contrastive Pre-Training

SafetyDGX agent

arXiv:2511.23170v5 Announce Type: replace Abstract: Contrastive vision-language pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-languag

← Previous
1…3031323334…240
Next →