AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
30 Jun 2026

Pose-Based Fall Detection System: Efficient Monitoring on Standard CPUs

SafetyDGX agent

arXiv:2503.19501v2 Announce Type: replace-cross Abstract: Falls among elderly residents in assisted living homes pose significant health risks, often leading to injuries and a decreased quality of lif

Propagation of~Interval Belief Structures and~Imprecise Copulas for~Neural Network Verification

SafetyDGX agent

arXiv:2606.30105v1 Announce Type: new Abstract: Quantitative verification of neural networks requires reasoning about probabilities under substantial uncertainty in both input distributions and their

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

Model ReleasesDGX agent

arXiv:2606.29887v1 Announce Type: new Abstract: In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies,

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group Discovery

SafetyDGX agent

arXiv:2606.29403v1 Announce Type: cross Abstract: Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in s

“Sometimes I wonder, maybe everything is conscious, or nothing is conscious. I prefer the former; it just seems more fun.” – Elon Musk

SafetyDGX agent

“Sometimes I wonder, maybe everything is conscious, or nothing is conscious. I prefer the former; it just seems more fun.” – Elon Musk Media Neuralink has solved through-dura electrode implantation! T

Test-Time Detoxification without Training or Learning Anything

SafetyDGX agent

arXiv:2602.02498v2 Announce Type: replace-cross Abstract: Large language models can produce toxic or inappropriate text even for benign inputs, creating risks when deployed at scale. Detoxification is

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

SafetyDGX agent

arXiv:2510.06096v3 Announce Type: replace-cross Abstract: The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a gr

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems

SafetyDGX agent

arXiv:2606.28425v1 Announce Type: cross Abstract: Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The natural defen

TrajRS: Towards Certified Robustness in Pedestrian Trajectory Prediction

SafetyDGX agent

arXiv:2606.28716v1 Announce Type: new Abstract: The robustness of trajectory prediction models is crucial for developing safe autonomous driving systems. Adversarial attacks on trajectory prediction c

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control

SafetyDGX agent

arXiv:2606.28760v1 Announce Type: new Abstract: Social robot navigation (SRN) requires more than geometric path planning; it demands understanding human intentions, social norms, and contextual cues t

29 Jun 2026

Artificial Intelligence Index Report 2026

SafetyDGX agent

arXiv:2606.15708v2 Announce Type: replace Abstract: Welcome to the ninth edition of the AI Index report. As AI continues to advance rapidly, the question becomes whether the systems built around it ca

CacheMPC: Certified Cached Model Predictive Control for Quadruped Locomotion

SafetyDGX agent

arXiv:2606.28300v1 Announce Type: new Abstract: Model Predictive Control (MPC) is the standard predictive layer in hierarchical quadruped controllers, but the per-cycle QP solve limits the update rate

CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association

SafetyDGX agent

arXiv:2606.28179v1 Announce Type: cross Abstract: Identifying robust associations between cardiac imaging phenotypes and clinical diseases is fundamental to population-scale cardiovascular research an

Drifting in the Future: Stabilizing Path Following Drifting on High-Latency Vehicle Systems

SafetyDGX agent

arXiv:2606.27914v1 Announce Type: new Abstract: Autonomously controlling and handling a vehicle at and beyond its stability limit is a mathematically and computationally demanding task. Prior demonstr

Halt Fast! Early Stopping for Certified Robustness

SafetyDGX agent

arXiv:2606.27694v1 Announce Type: cross Abstract: Randomized Smoothing (RS) provides rigorous robustness guarantees for neural networks without architectural constraints, yet its adoption is limited b

hia-gat: A Heterogeneous Interaction-Aware Graph Attention Network For Frame-Level Traffic Conflict Risk Prediction On Freeways

SafetyDGX agent

arXiv:2606.27577v1 Announce Type: cross Abstract: This paper formulates frame-level freeway risk assessment as a multi-agent scene graph-level binary classification problem, where each video or trajec

Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs

SafetyDGX agent

arXiv:2601.21233v2 Announce Type: replace Abstract: Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, and self-d

MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation

SafetyDGX agent

arXiv:2510.10271v2 Announce Type: replace-cross Abstract: Unlike regular tokens derived from existing text corpora, special tokens are artificially created to annotate structured conversations during

OverFlowLight: Real-Time Gridlock Prevention and Traffic Signal Optimization for Urban Intersections

SafetyDGX agent

arXiv:2606.27381v1 Announce Type: cross Abstract: Queue overflow, a severe consequence of urban traffic congestion, occurs when vehicle queues exceed intersection capacity, obstructing upstream traffi

Position: The Term 'Machine Unlearning' Is Overused in LLMs

SafetyDGX agent

arXiv:2606.27379v1 Announce Type: cross Abstract: Large language models increasingly face demands to 'forget' training data, knowledge, or behaviors due to regulatory deletion obligations, copyright/l

PRISON: Unmasking the Criminal Potential of Large Language Models

SafetyDGX agent

arXiv:2506.16150v4 Announce Type: replace-cross Abstract: As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research overlooked th

Regularized Reward-Punishment Reinforcement Learning

SafetyDGX agent

arXiv:2606.28152v1 Announce Type: new Abstract: We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR,

ReWorld: Learning Better Representations for World Action Models

SafetyDGX agent

arXiv:2606.27504v1 Announce Type: new Abstract: World Action Models (WAMs) model future environment evolution under action conditioning, offering a scalable paradigm for autonomous driving. However, e

27 Jun 2026

not dunking on Dario here bc Mythos is to GPT2 what a lion is to a bumblebee, but it’s insane that GPT-2, and successive models, have had ~0…

SafetyDGX agent

not dunking on Dario here bc Mythos is to GPT2 what a lion is to a bumblebee, but it’s insane that GPT-2, and successive models, have had ~0% impact on causing bad universal outcomes & so SHOULD HAVE

26 Jun 2026

100% correct. When I was at the Del Rio, TX Haitian bridge camp in 2021, many of the Haitians told me they had been living in Chile & Brazil…

SafetyDGX agent

100% correct. When I was at the Del Rio, TX Haitian bridge camp in 2021, many of the Haitians told me they had been living in Chile & Brazil for years before coming to the US illegally for economic (n

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

SafetyDGX agent

arXiv:2606.26396v1 Announce Type: new Abstract: Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-wo

Bridging Performance and Generalization in Reinforcement Learning for Agile Flight

SafetyDGX agent

arXiv:2606.27348v1 Announce Type: new Abstract: Autonomous drone racing is a fundamentally challenging regime for autonomous aerial robots, requiring time-optimal control while operating under persist

LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction

SafetyDGX agent

arXiv:2606.26661v1 Announce Type: cross Abstract: Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios. While existin

Radical AI Interpretability

SafetyDGX agent

arXiv:2606.26523v1 Announce Type: new Abstract: We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanis

Reducing Conversational Escalation in Large Language Model Dialogue with Nonviolent Communication Constraints

SafetyDGX agent

arXiv:2606.26106v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in emotionally charged situations involving interpersonal conflict, frustration, and distress. Whil

Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling

SafetyDGX agent

arXiv:2606.26922v1 Announce Type: cross Abstract: Continuous driver monitoring in automated vehicles requires low-latency inference while avoiding unsafe decisions under uncertain driver states. Large

Sources: Meta lobbyists are urging California lawmakers to exempt social media platforms from legislation that would increase penalties in child-harm cases (Tyler Katzenberger/Politico)

SafetyDGX agent

Tyler Katzenberger / Politico: Sources: Meta lobbyists are urging California lawmakers to exempt social media platforms from legislation that would increase penalties in child-harm cases — Meta's plea

THIS. The reason the US Government seems out of its depth now is because people Andreessen and Sacks steered them wrong before and the USG g…

SafetyDGX agent

THIS. The reason the US Government seems out of its depth now is because people Andreessen and Sacks steered them wrong before and the USG got caught out, utterly unprepared. It seems both foolish and

Topology-Informed Neural Networks for Flood Detection in Optical and Synthetic Aperture Radar Imagery

SafetyDGX agent

arXiv:2606.26204v1 Announce Type: new Abstract: Floods frequently impact regions around the world. Rapid and accurate flood detection is crucial for emergency response and timely mitigation of human a

25 Jun 2026

Adaptive-Horizon Conflict-Based Search for Closed-Loop Multi-Agent Path Finding

SafetyDGX agent

arXiv:2602.12024v3 Announce Type: replace Abstract: Multi-Agent Path Finding (MAPF) is a core coordination problem for large robot fleets in automated warehouses and logistics. Existing approaches are

Bias-Controlled Primal-Dual Natural Actor-Critic: Optimal Rates for Constrained Multi-Objective Average-Reward RL

SafetyDGX agent

arXiv:2606.25012v1 Announce Type: new Abstract: Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisf

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

Model ReleasesDGX agent

arXiv:2606.19887v2 Announce Type: replace-cross Abstract: Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance vio

Generative AI for Safe and Photorealistic Drone Light Shows

SafetyDGX agent

arXiv:2606.25458v1 Announce Type: new Abstract: Drone light shows are redefining aerial entertainment, yet their widespread adoption is bottlenecked by labor-intensive, manual animation. While generat

GUI agent: Guided Exploration of User-Sensitive Screens

SafetyDGX agent

arXiv:2606.25705v1 Announce Type: new Abstract: LLM agents are increasingly being used to automate tasks for users within an open GUI environment. They inevitably encounter screens containing user-sen

Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions

SafetyDGX agent

arXiv:2606.25396v1 Announce Type: new Abstract: AI companions powered by large language models increasingly interact with cognition-developing users, including children and adolescents, creating risks

Narrative Feature or Structured Feature? A Study of Large Language Models to Identify Cancer Patients at Risk of Heart Failure

SafetyDGX agent

arXiv:2403.11425v4 Announce Type: replace-cross Abstract: Cancer treatments are known to introduce cardiotoxicity, negatively impacting outcomes and survivorship. Identifying cancer patients at risk o

Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant

SafetyDGX agent

arXiv:2606.25181v1 Announce Type: cross Abstract: Early identification of speech sound errors in children is often limited by access to specialists, motivating lightweight screening tools that can ope

Reliability-Asymmetric Spacecraft Autonomy: Co-Designing a Capable Learned GNC Stack with a Verified, Adaptation-Aware Runtime Shield

SafetyDGX agent

arXiv:2606.25366v1 Announce Type: new Abstract: Deep-space missions need onboard autonomy that is both capable and certifiable. Rule-based autonomy is certifiable but brittle, while learned autonomy i

Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See

SafetyDGX agent

arXiv:2606.25127v1 Announce Type: new Abstract: We investigate how reward design shapes the internal attention patterns of reinforcement learning agents trained for autonomous driving. Using three Per

Statistically Valid Hyperparameter Selection: From Tuning to Guarantees

SafetyDGX agent

arXiv:2606.25601v1 Announce Type: cross Abstract: Hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom suc

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

SafetyDGX agent

arXiv:2601.16529v3 Announce Type: replace Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based gui

24 Jun 2026

Bilevel Data Curation for LLM Fine-tuning: Offline Selection and Online Self-Refining Generation

SafetyDGX agent

arXiv:2511.21056v2 Announce Type: replace-cross Abstract: Supervised fine-tuning (SFT) datasets are critical to the downstream performance of large language models, yet they often contain low-quality

Critique of Agent Model

SafetyDGX agent

arXiv:2606.23991v1 Announce Type: new Abstract: What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding agents'', ``AI co-scientists'', and

DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model

SafetyDGX agent

arXiv:2606.24051v1 Announce Type: new Abstract: Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follow

FT-WBC: Learning Fault-Tolerant Whole-Body Control for Legged Loco-Manipulation

SafetyDGX agent

arXiv:2606.24466v1 Announce Type: new Abstract: Legged manipulators combine the mobility of legged platforms with the manipulation capability of robotic arms. However, arm-induced Center-of-Mass shift

PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation

SafetyDGX agent

arXiv:2606.24081v1 Announce Type: cross Abstract: As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. More important

Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection

SafetyDGX agent

arXiv:2606.24575v1 Announce Type: new Abstract: Modern finance relies heavily on complex machine learning models to find patterns in the stock market. However, as these AI models get more complicated,

Selective Capability Unlearning in End-to-End Spoken Language Understanding

SafetyDGX agent

arXiv:2606.24063v1 Announce Type: cross Abstract: Modern spoken language understanding (SLU) systems are increasingly deployed in real-world settings, where specific functionalities may need to be rem

When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models

SafetyDGX agent

arXiv:2606.22974v2 Announce Type: replace Abstract: Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LL

23 Jun 2026

A Generative Model for Closed-Loop Microsimulation of Signalized Intersections

SafetyDGX agent

arXiv:2606.23588v1 Announce Type: new Abstract: Traffic microsimulators rely on hand-crafted behavior models that reproduce aggregate flow but miss the heterogeneous interactions between vehicles at s

Adversarial observations in probabilistic State-Space Models for robust Reinforcement Learning

SafetyDGX agent

arXiv:2606.20880v1 Announce Type: cross Abstract: Decision-making under partial or adversarial observability requires accurate inference of the environment's latent state and its associated uncertaint

BadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous Driving

SafetyDGX agent

arXiv:2606.21172v1 Announce Type: new Abstract: Video world models are increasingly used in autonomous driving to forecast future scene evolution and provide future-aware spatio-temporal representatio

Continuous-Time Probabilistic Correctors for Uncertainty-Aware Physics-Based Spacecraft Trajectory Forecasting

SafetyDGX agent

arXiv:2606.21021v1 Announce Type: new Abstract: Long-horizon spacecraft trajectory forecasting suffers from error accumulation due to the absence of corrective observations in the forecast regime, mak

D2HDMap: Non-visible Driveline Map Prior for Online Vectorized HD Map Prediction

SafetyDGX agent

arXiv:2606.20725v1 Announce Type: new Abstract: Accurate, up-to-date representations of road structures are critical for the safe operation of autonomous vehicles. Existing systems rely either on cost

Helping build shared standards for advanced AI

SafetyDGX agent

OpenAI discusses its efforts to contribute to the development of shared industry standards and best practices for advanced artificial intelligence systems. The article likely covers OpenAI's involveme

← Previous
1…3839404142…240
Next →