AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
Safety

Contour Errors: Ego-Centric Matching for 3D Multi-Object Tracking Performance Evaluation

DGX agent

arXiv:2506.04122v4 Announce Type: replace Abstract: Open-loop performance evaluation of 3D multi-object tracking in autonomous driving requires matching criteria that effectively penalize translationa

safetyarxiv-cs-cv
23 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models

DGX agent

arXiv:2607.19528v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. However, the ma

safetyarxiv-cs-ai
23 Jul 2026
Safety

Effort-Based Criticality Metrics for Evaluating 3D Perception Errors in Autonomous Driving

DGX agent

arXiv:2603.28029v2 Announce Type: replace-cross Abstract: Criticality metrics such as time-to-collision (TTC) quantify collision urgency but do not distinguish the operational consequences of false-po

safetyarxiv-cs-ro
23 Jul 2026
Safety

EGRNet: A Lightweight Semantic Segmentation Network with Edge-Gated Refinement and Adversarial Sensing

DGX agent

arXiv:2607.19617v1 Announce Type: new Abstract: As autonomous systems and smart cities continue to evolve, the demand for efficient and robust scene understanding becomes increasingly critical. Semant

safetyarxiv-cs-cv
23 Jul 2026
Safety

It’s important that outside agencies and independent bodies hold frontier AI companies accountable for the actions of their models. This is …

DGX agent

It’s important that outside agencies and independent bodies hold frontier AI companies accountable for the actions of their models. This is not a “whoops” situation, it’s a deliberate policy choice. W

safetygary-marcus--x
23 Jul 2026
Safety

Membership Inference Attacks for Unseen Classes

DGX agent

arXiv:2506.06488v3 Announce Type: replace Abstract: A key tool in developing safe AI models is data auditing, i.e., using statistical tools to determine whether harmful content may have been used in t

safetyarxiv-cs-lg
23 Jul 2026
Safety

More than 300 million people turn to ChatGPT with health-related questions each week—and we’re continuing to improve how our models respond.…

DGX agent

More than 300 million people turn to ChatGPT with health-related questions each week—and we’re continuing to improve how our models respond. We work with hundreds of physicians around the world to mea

safetyopenai--x
23 Jul 2026
Safety

Rater State Bias in RLHF Preference Data: An Audit Framework

DGX agent

arXiv:2607.16195v2 Announce Type: replace Abstract: We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compa

safetyarxiv-cs-ai
23 Jul 2026
Safety

The Ethics of Autonomous AI Agents for Offensive Security

DGX agent

arXiv:2607.20255v1 Announce Type: cross Abstract: LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly scoped, and o

safetyarxiv-cs-ai
23 Jul 2026
Safety

The Mechanism Matters: When Knowledge Graphs Help Reinforcement Learning

DGX agent

arXiv:2607.19616v1 Announce Type: new Abstract: Knowledge graphs (KGs) are widely used to inject prior knowledge into reinforcement learning (RL), yet the literature is dominated by single-domain, pos

safetyarxiv-cs-lg
23 Jul 2026
Safety

From the original HuggingFace report, before they knew it was OpenAI. Wild stuff. Things are going to get weird “When we started the log ana…

DGX agent

From the original HuggingFace report, before they knew it was OpenAI. Wild stuff. Things are going to get weird “When we started the log analysis, we first used frontier models behind commercial APIs.

safetyclem-delangue--x
22 Jul 2026
Model Releases

A Fireside Chat with Cat and Thariq from the Claude Code team

DGX agent

Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, c

model-releasessimon-willison
21 Jul 2026
Safety

ok the Hugging Face breach writeup is one of the more honest post-mortems I've read in a while, and there's a detail buried in it that's way…

DGX agent

ok the Hugging Face breach writeup is one of the more honest post-mortems I've read in a while, and there's a detail buried in it that's way more interesting than 'AI agent hacked us.' Their IR team t

safetyclem-delangue--x
21 Jul 2026
Safety

A VAE-Driven Multi-Task Satellite-Aided Semantic Communication Framework for 6G-Enabled Connected Autonomous Vehicles

DGX agent

arXiv:2607.13494v1 Announce Type: new Abstract: The development of smart transportation systems and the introduction of 6G wireless communication technologies have significantly changed vehicle networ

safetyarxiv-cs-lg
16 Jul 2026
Safety

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management

DGX agent

arXiv:2607.13239v1 Announce Type: new Abstract: Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation management center

safetyarxiv-cs-ai
16 Jul 2026
Safety

Layered Risk Mapping for Autonomous Patient Transport in Expeditionary Medical Facilities

DGX agent

arXiv:2607.13497v1 Announce Type: new Abstract: In expeditionary medical facilities, routine patient transport imposes a compounding burden of personal protective equipment consumption, staff diversio

safetyarxiv-cs-ro
16 Jul 2026
Safety

Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities

DGX agent

arXiv:2607.13596v1 Announce Type: cross Abstract: When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledgi

safetyarxiv-cs-ai
16 Jul 2026
Safety

STITCHER: Constrained Trajectory Planning in Complex Environments with Real-Time Motion Primitive Search

DGX agent

arXiv:2510.14893v4 Announce Type: replace Abstract: Autonomous high-speed navigation through large, complex environments requires real-time generation of agile trajectories that are dynamically feasib

safetyarxiv-cs-ro
16 Jul 2026
Safety

Temperature Scaling Is Not Enough: Calibration Gaps Under Human Label Distributions

DGX agent

arXiv:2607.13423v1 Announce Type: new Abstract: Temperature scaling is the dominant post-hoc calibration method in modern deep learning. Its theoretical justification rests on an assumption that is ra

safetyarxiv-cs-lg
16 Jul 2026
Safety

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

DGX agent

arXiv:2607.13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at

safetyarxiv-cs-ai
16 Jul 2026
Safety

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

DGX agent

arXiv:2607.12273v1 Announce Type: cross Abstract: As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, w

safetyarxiv-cs-ai
15 Jul 2026
Safety

G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis

DGX agent

arXiv:2607.11892v1 Announce Type: cross Abstract: Human-factor event diagnosis is essential for learning from operational events in nuclear power plants, yet its quality depends strongly on expert int

safetyarxiv-cs-ai
15 Jul 2026
Safety

Git-Assistant: Planning-Based Support for Updating Git Repositories

DGX agent

arXiv:2607.09224v2 Announce Type: replace-cross Abstract: Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners. Re

safetyarxiv-cs-ai
15 Jul 2026
Model Releases

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

DGX agent

arXiv:2607.12792v1 Announce Type: cross Abstract: Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, however, are se

model-releasesarxiv-cs-ai
15 Jul 2026
Safety

Streamlining stereo differentiable rendering for marker-free real-time tracking of surgical robots

DGX agent

arXiv:2607.12604v1 Announce Type: new Abstract: Purpose: Marker-based tracking of surgical robots is occlusion-prone in cluttered operating rooms. We evaluate stereo differentiable rendering for marke

safetyarxiv-cs-ro
15 Jul 2026
Safety

Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap

DGX agent

arXiv:2607.12113v1 Announce Type: cross Abstract: One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around fi

safetyarxiv-cs-ai
15 Jul 2026
Safety

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

DGX agent

arXiv:2607.12790v1 Announce Type: new Abstract: Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable e

safetyarxiv-cs-ai
15 Jul 2026
Safety

The Anglo-Scottish Enlightenment – the real antidote to Rousseau and Voltaire The French Enlightenment and the Anglo-Scottish Enlightenment …

DGX agent

The Anglo-Scottish Enlightenment – the real antidote to Rousseau and Voltaire The French Enlightenment and the Anglo-Scottish Enlightenment happened simultaneously, in the same century, reading the sa

safetyelon-musk--x
11 Jul 2026
Safety

Detecting Ladder Logic Bombs in IEC 61131-3 PLC Programs using ESBMC-PLC+: A Formal Verification Approach with Trigger Synthesis

DGX agent

arXiv:2607.08417v1 Announce Type: new Abstract: A Ladder Logic Bomb (LLB) is malicious control logic in a Programmable Logic Controller (PLC) program that lies dormant until a trigger activates a payl

safetyarxiv-cs-cl
10 Jul 2026
Safety

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

DGX agent

arXiv:2607.07859v1 Announce Type: new Abstract: Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While h

safetyarxiv-cs-ai
10 Jul 2026
Safety

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

DGX agent

arXiv:2607.08028v1 Announce Type: new Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization

safetyarxiv-cs-ai
10 Jul 2026
Safety

In preliminary findings, the EU Commission said Facebook's and Instagram's 'addictive design' violates the DSA, telling Meta to make changes or risk hefty fines (Adam Satariano/New York Times)

DGX agent

Adam Satariano / New York Times: In preliminary findings, the EU Commission said Facebook's and Instagram's “addictive design” violates the DSA, telling Meta to make changes or risk hefty fines — Euro

safetytechmeme
10 Jul 2026
Safety

In vivo feasibility study of humanoid robots in surgery

DGX agent

arXiv:2607.07972v1 Announce Type: new Abstract: Recent advances in actuation, control and learning have rapidly pushed humanoid robots from a distant vision towards near-term real-world deployment. He

safetyarxiv-cs-ro
10 Jul 2026
Safety

It Takes a MAESTRO To Prune Bad Experts

DGX agent

arXiv:2607.08601v1 Announce Type: new Abstract: Sparsely-activated Mixture-of-Experts (MoE) language models achieve remarkable inference efficiency by activating only a small fraction of parameters pe

safetyarxiv-cs-cl
10 Jul 2026
Safety

Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

DGX agent

arXiv:2607.07903v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing appro

safetyarxiv-cs-ai
10 Jul 2026
Safety

Agentic Data Environments

DGX agent

arXiv:2607.07397v1 Announce Type: new Abstract: Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. Th

safetyarxiv-cs-ai
9 Jul 2026
Safety

Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

DGX agent

arXiv:2607.07370v1 Announce Type: cross Abstract: In embodied intelligence systems, the motion controller serves as the critical bridge between semantic reasoning and physical execution. Humanoid cont

safetyarxiv-cs-ai
9 Jul 2026
Safety

Manual, Joystick, or Haptic Control? An In Vitro Comparison of Navigation Strategies for Robotic Interventional Neuroradiology Procedures

DGX agent

arXiv:2607.07253v1 Announce Type: new Abstract: Objective: To evaluate robotic controller interfaces for interventional neuroradiology procedures in-vitro incorporating a force-sensing platform to ass

safetyarxiv-cs-ro
9 Jul 2026
Safety

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning

DGX agent

arXiv:2607.07316v1 Announce Type: new Abstract: This article offers a comprehensive overview of mechanistic interpretability, an emerging field that seeks to reverse-engineer the internal algorithms o

safetyarxiv-cs-lg
9 Jul 2026
Safety

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors

DGX agent

arXiv:2607.07368v1 Announce Type: cross Abstract: AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent. AI Control usually studies a single

safetyarxiv-cs-ai
9 Jul 2026
Safety

NonTextual Target Attack

DGX agent

arXiv:2510.02999v5 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with

safetyarxiv-cs-ai
9 Jul 2026
Safety

R^3: Advertisement Compliance Rectification via Group-Relative Experience Extractor and Curriculum Reinforcement

DGX agent

arXiv:2607.07318v1 Announce Type: new Abstract: Rigorous content moderation is crucial for online advertising but leads to millions of daily rejections. This scale renders manual rectification infeasi

safetyarxiv-cs-cl
9 Jul 2026
Safety

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

DGX agent

arXiv:2607.07663v1 Announce Type: new Abstract: AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data t

safetyarxiv-cs-ai
9 Jul 2026
Safety

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

DGX agent

arXiv:2607.07436v1 Announce Type: new Abstract: A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the stru

safetyarxiv-cs-ai
9 Jul 2026
Safety

The Power of Backdoor Absorption in Community Training

DGX agent

arXiv:2607.06643v1 Announce Type: cross Abstract: Backdoor attacks severely threaten large-scale AI models. When model owners delegate training to external compute providers within a decentralized tra

safetyarxiv-cs-lg
9 Jul 2026
Safety

Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer

DGX agent

arXiv:2510.24108v2 Announce Type: replace-cross Abstract: Human demonstrations are widely considered the cornerstone of end-to-end (E2E) autonomous driving despite human demonstration's scarcity for l

safetyarxiv-cs-cv
9 Jul 2026
Local Ai

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

DGX agent

arXiv:2607.06120v1 Announce Type: new Abstract: Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploi

local-aiarxiv-cs-cv
8 Jul 2026
Model Releases

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

DGX agent

arXiv:2607.06196v1 Announce Type: new Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional law

model-releasesarxiv-cs-cl
8 Jul 2026
← Previous
1…4546474849…300
Next →