AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
Safety

Security Considerations for Multi-agent Systems

DGX agent

arXiv:2603.09002v2 Announce Type: replace-cross Abstract: Multi-agent artificial intelligence systems or MAS are systems of autonomous agents that exercise delegated tool authority, share persistent m

safetyarxiv-cs-ai
28 Apr 2026
Safety

Sliding Mode Control for Safe Trajectory Tracking with Moving Obstacles Avoidance: Experimental Validation on Planar Robots

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2604.24518v1 Announce Type: cross Abstract: This paper presents a unified control framework for robust trajectory tracking and moving obstacle avoidance applicable to a broad class of mobile rob

safetyarxiv-cs-ro
28 Apr 2026
Safety

Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning

DGX agent

arXiv:2511.01490v2 Announce Type: replace Abstract: As synthetic data becomes widely used in language model development, understanding its impact on model behavior is crucial. This paper investigates

safetyarxiv-cs-cl
28 Apr 2026
Safety

Training a General Purpose Automated Red Teaming Model

DGX agent

arXiv:2604.23067v1 Announce Type: cross Abstract: Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing fo

safetyarxiv-cs-cl
28 Apr 2026
Safety

Value Alignment Tax: Measuring Value Trade-offs in LLM Alignment

DGX agent

arXiv:2602.12134v2 Announce Type: replace Abstract: Existing work on value alignment typically characterizes value relations statically, ignoring how alignment interventions, such as prompting, fine-t

safetyarxiv-cs-ai
28 Apr 2026
Safety

Adaptive vs. Static Robot-to-Human Handover: A Study on Orientation and Approach Direction

DGX agent

arXiv:2604.22378v1 Announce Type: new Abstract: Robot-to-human handovers often rely on static, open-loop strategies (or, at best, approaches that adapt only the position), which generally do not consi

safetyarxiv-cs-ro
27 Apr 2026
Safety

AgentBound: Securing Execution Boundaries of AI Agents

DGX agent

arXiv:2510.21236v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have evolved into AI agents that interact with external tools and environments to perform complex tasks. The Mode

safetyarxiv-cs-ai
27 Apr 2026
Safety

Controllable Spoken Dialogue Generation: An LLM-Driven Grading System for K-12 Non-Native English Learners

DGX agent

arXiv:2604.22542v1 Announce Type: cross Abstract: Large language models (LLMs) often fail to meet the pedagogical needs of K-12 English learners in non-native contexts due to a proficiency mismatch. T

safetyarxiv-cs-ai
27 Apr 2026
Safety

Don't try to build a self-improving AI agent without evals. You are just wasting time and compute. An agent can't improve from traces it can…

DGX agent

Don't try to build a self-improving AI agent without evals. You are just wasting time and compute. An agent can't improve from traces it can't evaluate. This is why it's exciting to see @FutureAGI_ go

safetydair-ai--x
27 Apr 2026
Safety

Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework

DGX agent

arXiv:2604.22119v1 Announce Type: new Abstract: As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve their own ob

safetyarxiv-cs-ai
27 Apr 2026
Safety

Sam Altman cannot be trusted. • The OpenAI board fired him because he was not always honest with them. They said he should not control power…

DGX agent

Sam Altman cannot be trusted. • The OpenAI board fired him because he was not always honest with them. They said he should not control powerful AI. • A major report talked to over 100 people and saw s

safetyelon-musk--x
27 Apr 2026
Safety

Bounding the Black Box: A Statistical Certification Framework for AI Risk Regulation

DGX agent

arXiv:2604.21854v1 Announce Type: new Abstract: Artificial intelligence now decides who receives a loan, who is flagged for criminal investigation, and whether an autonomous vehicle brakes in time. Go

safetyarxiv-cs-ai
24 Apr 2026
Safety

Enabling and Inhibitory Pathways of University Students' Willingness to Disclose AI Use: A Cognition-Affect-Conation Perspective

DGX agent

arXiv:2604.21733v1 Announce Type: new Abstract: The increasing integration of artificial intelligence (AI) in higher education has raised important questions regarding students' transparency in report

safetyarxiv-cs-ai
24 Apr 2026
Safety

HARBOR: Automated Harness Optimization

DGX agent

arXiv:2604.20938v1 Announce Type: cross Abstract: Long-horizon language-model agents are dominated, in lines of code and in operational complexity, not by their underlying model but by the harness tha

safetyarxiv-cs-ai
24 Apr 2026
Safety

Lawmakers and lobbyists say the Trump administration has lobbied against legislation that would regulate AI in at least six Republican-led states (Amrith Ramkumar/Wall Street Journal)

DGX agent

Amrith Ramkumar / Wall Street Journal: Lawmakers and lobbyists say the Trump administration has lobbied against legislation that would regulate AI in at least six Republican-led states — ‘I am disappo

safetytechmeme
24 Apr 2026
Safety

Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review

DGX agent

arXiv:2603.18740v2 Announce Type: replace-cross Abstract: Automated Code Review (ACR) systems integrating Large Language Models (LLMs) are increasingly adopted in software development workflows, rangi

safetyarxiv-cs-ai
24 Apr 2026
Safety

Mind the Prompt: Self-adaptive Generation of Task Plan Explanations via LLMs

DGX agent

arXiv:2604.21092v1 Announce Type: new Abstract: Integrating Large Language Models (LLMs) into complex software systems enables the generation of human-understandable explanations of opaque AI processe

safetyarxiv-cs-ai
24 Apr 2026
Safety

Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

DGX agent

arXiv:2603.12845v2 Announce Type: replace Abstract: Predicting enzyme kinetic parameters quantifies how efficiently an enzyme catalyzes a specific substrate under defined biochemical conditions. Canon

safetyarxiv-cs-cv
24 Apr 2026
Safety

Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build …

DGX agent

Today we open sourced many of OpenAI's monitorability evaluations. We hope that the research community and other model developers can build upon them and use them to evaluate the monitorability of the

safetysam-altman--x
24 Apr 2026
Safety

AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation

DGX agent

arXiv:2604.20134v1 Announce Type: cross Abstract: Security Operations Centers (SOCs) increasingly encounter difficulties in correlating heterogeneous alerts, interpreting multi-stage attack progressio

safetyarxiv-cs-ai
23 Apr 2026
Safety

Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents

DGX agent

arXiv:2604.20572v1 Announce Type: new Abstract: Online lifelong learning enables agents to accumulate experience across interactions and continually improve on long-horizon tasks. However, existing me

safetyarxiv-cs-cl
23 Apr 2026
Safety

FLOSS: Federated Learning with Opt-Out and Straggler Support

DGX agent

arXiv:2507.23115v2 Announce Type: replace-cross Abstract: Previous work on data privacy in federated learning systems focuses on privacy-preserving operations for data from users who have agreed to sh

safetyarxiv-cs-ai
23 Apr 2026
Safety

LLMs Can Get 'Brain Rot': A Pilot Study on Twitter/X

DGX agent

arXiv:2510.13928v2 Announce Type: replace-cross Abstract: We propose and test the LLM Brain Rot Hypothesis: continual exposure to junk web text induces lasting cognitive decline in large language mode

safetyarxiv-cs-ai
23 Apr 2026
Safety

MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment

DGX agent

arXiv:2604.20685v1 Announce Type: new Abstract: Aligning large language models (LLMs) to desirable human values requires balancing multiple, potentially conflicting objectives such as helpfulness, tru

safetyarxiv-cs-lg
23 Apr 2026
Safety

SAMix: Calibrated and Accurate Continual Learning via Sphere-Adaptive Mixup and Neural Collapse

DGX agent

arXiv:2510.15751v2 Announce Type: replace Abstract: While most continual learning methods focus on mitigating forgetting and improving accuracy, they often overlook the critical aspect of network cali

safetyarxiv-cs-lg
23 Apr 2026
Safety

Verification of Machine Unlearning is Fragile

DGX agent

arXiv:2408.00929v2 Announce Type: replace Abstract: As privacy concerns escalate in the realm of machine learning, data owners now have the option to utilize machine unlearning to remove their data fr

safetyarxiv-cs-lg
23 Apr 2026
Safety

BAPO: Boundary-Aware Policy Optimization for Reliable Agentic Search

DGX agent

arXiv:2601.11037v2 Announce Type: replace Abstract: RL-based agentic search enables LLMs to solve complex questions via dynamic planning and external search. While this approach significantly enhances

safetyarxiv-cs-ai
22 Apr 2026
Safety

Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications

DGX agent

arXiv:2604.19281v1 Announce Type: cross Abstract: The use of Large Language Models (LLMs) to support patients in addressing medical questions is becoming increasingly prevalent. However, most of the m

safetyarxiv-cs-ai
22 Apr 2026
Safety

Counting Worlds Branching Time Semantics for post-hoc Bias Mitigation in generative AI

DGX agent

arXiv:2604.19431v1 Announce Type: cross Abstract: Generative AI systems are known to amplify biases present in their training data. While several inference-time mitigation strategies have been propose

safetyarxiv-cs-ai
22 Apr 2026
Safety

How to Teach Large Multimodal Models New Skills

DGX agent

arXiv:2510.08564v2 Announce Type: replace Abstract: How can we teach large multimodal models (LMMs) new skills without erasing prior abilities? We study sequential fine-tuning on five target skills wh

safetyarxiv-cs-ai
22 Apr 2026
Safety

Hybrid Task and Motion Planning with Reactive Collision Handling for Multi-Robot Disassembly of Complex Products: Application to EV Batteries

DGX agent

arXiv:2509.21020v2 Announce Type: replace Abstract: This paper addresses the problem of multi-robot coordination for complex manipulation task sequences. We present a vision-driven task-and-motion pla

safetyarxiv-cs-ro
22 Apr 2026
Safety

LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit

DGX agent

arXiv:2604.19117v1 Announce Type: new Abstract: When a language model agrees with a user's false belief, is it failing to detect the error, or noticing and agreeing anyway? We show the latter. Across

safetyarxiv-cs-lg
22 Apr 2026
Safety

Multimodal embodiment-aware navigation transformer

DGX agent

arXiv:2604.19267v1 Announce Type: new Abstract: Goal-conditioned navigation models for ground robots trained using supervised learning show promising zero-shot transfer, but their collision-avoidance

safetyarxiv-cs-ro
22 Apr 2026
Safety

Neuromorphic Continual Learning for Sequential Deployment of Nuclear Plant Monitoring Systems

DGX agent

arXiv:2604.18611v1 Announce Type: cross Abstract: Anomaly detection in nuclear industrial control systems (ICS) requires continuous, energy-efficient monitoring across multiple subsystems that are oft

safetyarxiv-cs-ai
22 Apr 2026
Safety

One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization

DGX agent

arXiv:2601.18572v2 Announce Type: replace Abstract: Personalization of LLMs by sociodemographic subgroup often improves user experience, but can also introduce or amplify biases and unfair outcomes ac

safetyarxiv-cs-cl
22 Apr 2026
Safety

Our newsroom AI policy

DGX agent

Ars Technica's newsroom AI policy forbids unlabeled AI material in reported stories and requires human confirmation of every quotation's accuracy. The publication does not permit the publication of AI

safetyars-technica
22 Apr 2026
Model Releases

Safe Continual Reinforcement Learning in Non-stationary Environments

DGX agent

arXiv:2604.19737v1 Announce Type: new Abstract: Reinforcement learning (RL) offers a compelling data-driven paradigm for synthesizing controllers for complex systems when accurate physical models are

model-releasesarxiv-cs-lg
22 Apr 2026
Safety

SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention

DGX agent

arXiv:2506.14387v3 Announce Type: replace Abstract: Adapting LLMs with new knowledge is increasingly important, but standard fine-tuning often erodes aligned epistemic abstention: the ability to ackno

safetyarxiv-cs-ai
22 Apr 2026
Safety

The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation

DGX agent

arXiv:2604.19064v1 Announce Type: new Abstract: In vision-and-language navigation (VLN), self-improvement from policy-induced experience, using only standard VLN action supervision, critically depends

safetyarxiv-cs-cv
22 Apr 2026
Safety

VIGIL: An Extensible System for Real-Time Detection and Mitigation of Cognitive Bias Triggers

DGX agent

arXiv:2604.03261v2 Announce Type: replace Abstract: The rise of generative AI is posing increasing risks to online information integrity and civic discourse. Most concretely, such risks can materialis

safetyarxiv-cs-cl
22 Apr 2026
Safety

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving

DGX agent

arXiv:2411.18275v2 Announce Type: replace Abstract: Vision-language models (VLMs) have significantly advanced autonomous driving (AD) by enhancing reasoning capabilities. However, these models remain

safetyarxiv-cs-cv
22 Apr 2026
Safety

Weakly supervised framework for wildlife detection and counting in challenging Arctic environments: a case study on caribou (Rangifer tarandus)

DGX agent

arXiv:2601.18891v3 Announce Type: replace Abstract: Caribou across the Arctic has declined in recent decades, motivating scalable and accurate monitoring approaches to guide evidence-based conservatio

safetyarxiv-cs-cv
22 Apr 2026
Safety

A Hamilton-Jacobi Reachability-Guided Search Framework for Efficient and Safe Indoor Planar Robot Navigation

DGX agent

arXiv:2604.17679v1 Announce Type: new Abstract: Autonomous navigation requires planning to reach a goal safely and efficiently in complex and potentially dynamic environments. Graph search-based algor

safetyarxiv-cs-ro
21 Apr 2026
Safety

I agree with Connor that most folks concerned about AI have been far too coy about extinction risk, and that ControlAI is one of few excepti…

DGX agent

I agree with Connor that most folks concerned about AI have been far too coy about extinction risk, and that ControlAI is one of few exceptions. The world needs more earnest communication efforts. We

safetyconnor-leahy--x
21 Apr 2026
Safety

Navigating the Conceptual Multiverse

DGX agent

arXiv:2604.17815v1 Announce Type: cross Abstract: When language models answer open-ended problems, they implicitly make hidden decisions that shape their outputs, leaving users with uncontextualized a

safetyarxiv-cs-cl
21 Apr 2026
Safety

One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment

DGX agent

arXiv:2601.18731v2 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individu

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding

DGX agent

arXiv:2604.17132v1 Announce Type: new Abstract: Safety-aligned large language models (LLMs) often generate refusal responses to harmless queries due to the over-refusal problem. However, existing meth

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

PowerCLIP: Powerset Alignment for Contrastive Pre-Training

DGX agent

arXiv:2511.23170v5 Announce Type: replace Abstract: Contrastive vision-language pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-languag

safetyarxiv-cs-cv
21 Apr 2026
← Previous
1…3839404142…299
Next →