AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
Safety

AeroSpectra Sentinel: An Auditable LLM Prompt-Chaining Decision-Support Workflow for Acute Asthma Risk Assessment from Respiratory Sounds and Clinical Signals

DGX agent

arXiv:2606.08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing,

safetyarxiv-cs-ai
9 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

alienating your customers before you IPO is maybe not the best idea, @AnthropicAI

DGX agent

alienating your customers before you IPO is maybe not the best idea, @AnthropicAI Brilliant idea! Next up: Apple randomly reboots your Mac if you're building competing tech, Gmail silently edits your

safetygary-marcus--x
9 Jun 2026
Safety

Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care

DGX agent

arXiv:2606.08982v1 Announce Type: new Abstract: Baichuan-M4 is Baichuan Intelligence's clinical-grade medical large model, designed for continuous care rather than single-turn medical question answeri

safetyarxiv-cs-ai
9 Jun 2026
Safety

Beyond Accuracy: Interpreting Topic Representation in Suicide Ideation Detection Models

DGX agent

arXiv:2606.07714v1 Announce Type: cross Abstract: Suicide ideation detection models are typically evaluated using aggregate performance metrics, yet little is known about how they internally represent

safetyarxiv-cs-ai
9 Jun 2026
Safety

Beyond Linear Activation Steering: Invertible Latent Transformations for Controlling LLM Behavior

DGX agent

arXiv:2606.08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation

safetyarxiv-cs-lg
9 Jun 2026
Safety

Brilliant idea! Next up: Apple randomly reboots your Mac if you're building competing tech, Gmail silently edits your email if you mention r…

DGX agent

Brilliant idea! Next up: Apple randomly reboots your Mac if you're building competing tech, Gmail silently edits your email if you mention rival platforms, and Tesla Autopilot swerves if it detects yo

safetyjeremy-howard--x
9 Jun 2026
Safety

Distilling LLM Reasoning into an Interpretable Policy Tree for Human-AI Collaboration

DGX agent

arXiv:2606.08596v1 Announce Type: new Abstract: Constructing efficient and reliable policies to assist humans is indispensable for human-AI collaboration. Existing methods mainly follow two lines of w

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Distilling Safe LLM Systems via Soft Prompts for On Device Settings

DGX agent

arXiv:2606.09388v1 Announce Type: new Abstract: Deploying safe large language models (LLMs) on resource-constrained edge devices presents a critical challenge: while dual-model systems combining LLMs

model-releasesarxiv-cs-lg
9 Jun 2026
Safety

Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO

DGX agent

arXiv:2606.09701v1 Announce Type: cross Abstract: AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel att

safetyarxiv-cs-ai
9 Jun 2026
Safety

Memetic Capture: A Pluralistic Policy Framework for Governing AI-Driven Cultural Disempowerment

DGX agent

arXiv:2606.07802v1 Announce Type: cross Abstract: Culture is the most insidious vector of gradual human disempowerment by AI: unlike economic or political displacement, cultural displacement attacks t

safetyarxiv-cs-ai
9 Jun 2026
Safety

Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies

DGX agent

arXiv:2606.08312v1 Announce Type: new Abstract: In this work we study offline reinforcement learning (RL) under temporally extended task constraints expressed in Linear Temporal Logic over finite trac

safetyarxiv-cs-ai
9 Jun 2026
Safety

Overcoming the Regulatory Bottleneck via Agent-to-Agent Protocols: A Nuclear Case Study

DGX agent

arXiv:2606.07866v1 Announce Type: new Abstract: Regulatory review of advanced nuclear reactor designs routinely spans more than three years and consumes hundreds of millions of dollars in combined reg

safetyarxiv-cs-ai
9 Jun 2026
Safety

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

DGX agent

arXiv:2606.07612v1 Announce Type: cross Abstract: We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for

safetyarxiv-cs-ai
9 Jun 2026
Safety

Revisiting the shutdown problem

DGX agent

arXiv:2606.08296v1 Announce Type: new Abstract: A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut d

safetyarxiv-cs-ai
9 Jun 2026
Safety

RPO-PDT: Demonstrating Role-Play-Based Knowledge Adaptation for Student Support Dialogue (Demonstration System)

DGX agent

arXiv:2606.09255v1 Announce Type: new Abstract: We present RPO-PDT: a retrieval-grounded, role-play-based dialogue system for adaptive student support in higher education. RPO-PDT is: (1) able to prov

safetyarxiv-cs-ro
9 Jun 2026
Safety

SAD-Flower: Flow Matching for Safe, Admissible, and Dynamically Consistent Planning

DGX agent

arXiv:2511.05355v3 Announce Type: replace Abstract: Flow matching (FM) has shown promising results in data-driven planning. However, it inherently lacks formal guarantees for ensuring state and action

safetyarxiv-cs-lg
9 Jun 2026
Model Releases

SafeRun: Enabling Determinism in LLM Planning for Running

DGX agent

arXiv:2606.09027v1 Announce Type: cross Abstract: Large Language Models enable flexible natural-language planning but remain unreliable in determinism-critical domains due to their probabilistic natur

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Stain-Aware Wavelet Regularization for Instant Adversarial Purification in Histopathology

DGX agent

arXiv:2606.08745v1 Announce Type: new Abstract: Deep learning has become prevalent in computational pathology pipelines that support tasks such as cancer screening and digital pathology analysis. Howe

safetyarxiv-cs-cv
9 Jun 2026
Safety

Toward autocorrection of chemical process flowsheets using large language models

DGX agent

arXiv:2312.02873v2 Announce Type: replace-cross Abstract: The process engineering domain widely uses Process Flow Diagrams (PFDs) and Process and Instrumentation Diagrams (P&IDs) to represent process

safetyarxiv-cs-ai
9 Jun 2026
Safety

TRACER: Token ReAssignment for Concept ERasure in Generative Recommendation

DGX agent

arXiv:2606.07688v1 Announce Type: cross Abstract: Generative recommendation formulates next-item prediction as autoregressive generation over semantic ID (SID) sequences derived from users' historical

safetyarxiv-cs-ai
9 Jun 2026
Safety

Transforming Police-Car Swerving for Mitigating Isolated Stop-and-Go Traffic Waves: A Practice-Oriented Jam-Absorption Driving Strategy

DGX agent

arXiv:2602.10234v3 Announce Type: replace-cross Abstract: Stop-and-go traffic waves, a major form of freeway congestion, impose severe and persistent adverse impacts, including reduced traffic efficie

safetyarxiv-cs-ai
9 Jun 2026
Safety

Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video

DGX agent

arXiv:2606.08828v1 Announce Type: new Abstract: Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains cha

safetyarxiv-cs-ro
9 Jun 2026
Safety

An Abstract Architecture for Explainable Autonomy in Hazardous Environments

DGX agent

arXiv:2606.07211v1 Announce Type: cross Abstract: Autonomous robotic systems are being proposed for use in hazardous environments, often to reduce the risks to human workers. In the immediate future,

safetyarxiv-cs-ai
8 Jun 2026
Safety

Bounded-Abstention Pairwise Learning to Rank

DGX agent

arXiv:2505.23437v2 Announce Type: replace-cross Abstract: Ranking systems influence decision-making in high-stakes domains like health, education, and employment, where they can have substantial econo

safetyarxiv-cs-ai
8 Jun 2026
Safety

CARVE-Q: Quantum-Proposed, Classically Certified Interactive Driving Repair

DGX agent

arXiv:2606.06531v1 Announce Type: new Abstract: The critical question after a correct driving veto is not only whether a maneuver is unsafe, but whether the blocked interaction admits a lawful, audita

safetyarxiv-cs-ai
8 Jun 2026
Safety

check out Marcus on AI

DGX agent

Gary Marcus, a prominent AI researcher and critic, shared a post on X directing followers to learn more about his perspectives on artificial intelligence. The post likely promotes his work, writings,

safetygary-marcus--x
8 Jun 2026
Safety

Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing

DGX agent

This Import AI newsletter issue covers three main topics: societal implications of reward hacking (optimizing for measurable metrics at the expense of intended goals), new reinforcement learning data

safetyimport-ai
8 Jun 2026
Safety

Position: Don't Just 'Fix it in Post': A Science of AI Must Study Training Dynamics

DGX agent

arXiv:2606.06533v1 Announce Type: new Abstract: What would it mean to have a scientific understanding of AI? Models are not static objects: they are snapshots of time-evolving processes shaped by data

safetyarxiv-cs-ai
8 Jun 2026
Safety

UK PM Keir Starmer says tech companies must introduce 'device controls' that stop kids from sending and receiving nude images or face laws forcing them to do so (Reuters)

DGX agent

Reuters: UK PM Keir Starmer says tech companies must introduce “device controls” that stop kids from sending and receiving nude images or face laws forcing them to do so — Big tech firms operating in

safetytechmeme
8 Jun 2026
Safety

Which title is better? Seven [Lies/Myths/Bitter Truths] About AI

DGX agent

Gary Marcus critiques common misconceptions about AI, presenting seven corrective perspectives on prevailing beliefs about artificial intelligence's capabilities and limitations. The post likely addre

safetygary-marcus--x
8 Jun 2026
Safety

Consistency Training Along the Transformer Stack

DGX agent

arXiv:2606.05817v1 Announce Type: cross Abstract: Consistency training encourages models to behave similarly across different contexts, and has shown promise for reducing misalignment. We broaden the

safetyarxiv-cs-ai
6 Jun 2026
Safety

If leading AI companies are indeed approaching the point of recursive self-improvement, a coordinated, verifiable, and universally applied p…

DGX agent

If leading AI companies are indeed approaching the point of recursive self-improvement, a coordinated, verifiable, and universally applied pause is probably the only responsible solution to mitigate s

safetyyoshua-bengio--x
6 Jun 2026
Safety

Risk Assessment of Autonomous Driving: Integrating Technical Failures, Ethical Dilemmas, and Policy Frameworks

DGX agent

arXiv:2606.06396v1 Announce Type: new Abstract: Autonomous driving technology has the potential to reduce the large number of road traffic accidents caused by human error each year, but it also brings

safetyarxiv-cs-ai
6 Jun 2026
Safety

Towards World Models in Biomedical Research

DGX agent

arXiv:2606.05925v1 Announce Type: new Abstract: A central goal of biomedicine is to understand, predict and ultimately control the dynamic mechanisms by which biological systems respond to perturbatio

safetyarxiv-cs-ai
6 Jun 2026
Safety

Unsupervised Pattern Analysis in Japanese Veterinary Toxicology: A Regulatory-Compliant Framework for Cross-Species Risk Assessment

DGX agent

arXiv:2606.06207v1 Announce Type: new Abstract: Veterinary pharmacovigilance systems are essential for monitoring adverse drug events (ADEs), yet existing approaches often fail to capture region-speci

safetyarxiv-cs-ai
6 Jun 2026
Safety

Willing but Unable: Separating Refusal from Capability in Code LLMs via Abliteration

DGX agent

arXiv:2606.05396v1 Announce Type: cross Abstract: Producing a labeled vulnerable code at scale is a recurring obstacle for learning-based vulnerability detection: mined corpora carry substantial label

safetyarxiv-cs-ai
6 Jun 2026
Safety

Alignment Risks from Capability-Seeking RL Training

DGX agent

arXiv:2602.12124v2 Announce Type: replace-cross Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capab

safetyarxiv-cs-cl
5 Jun 2026
Safety

Forgive or forget: Understanding the context of hate in audio retrieval systems

DGX agent

arXiv:2606.05857v1 Announce Type: new Abstract: Handling toxic retrieval in text-to-audio systems is challenging due to contextual dependencies. Existing strategies (e.g., rephrasing, summarization) r

safetyarxiv-cs-cl
5 Jun 2026
Safety

HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

DGX agent

arXiv:2606.06493v1 Announce Type: new Abstract: For a humanoid robot to be deployed in the real world, the choice of command space (i.e., the interface between task planning and whole-body control) is

safetyarxiv-cs-ro
5 Jun 2026
Safety

Literally the only reason to bailout OpenAI.

DGX agent

Gary Marcus argues for a specific rationale supporting a potential OpenAI bailout, though the exact reasoning is not detailed in the available information. Given Marcus's background as an AI researche

safetygary-marcus--x
5 Jun 2026
Safety

What's in a Name? Morphological Shortcuts by LLMs in Pharmacology

DGX agent

arXiv:2606.05616v1 Announce Type: new Abstract: The morphological form of a word can often give cues to its meaning, but purely relying on these mappings can lead to overgeneralization in high-stakes

safetyarxiv-cs-cl
5 Jun 2026
Safety

A Pathology Foundation Model for Gastric Cancer with Real-World Validation

DGX agent

arXiv:2606.04792v1 Announce Type: new Abstract: Gastric cancer remains a major cause of cancer mortality, yet its histological and molecular heterogeneity complicates diagnosis and risk stratification

safetyarxiv-cs-cv
4 Jun 2026
Safety

Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control

DGX agent

arXiv:2606.04775v1 Announce Type: cross Abstract: Text-to-video (T2V) models trained on large-scale web data can generate undesired content, motivating interventions that reduce harmful outputs withou

safetyarxiv-cs-ai
4 Jun 2026
Safety

Certified Neural Approximations of Nonlinear Dynamics

DGX agent

arXiv:2505.15497v3 Announce Type: replace Abstract: Neural networks hold great potential to act as approximate models of nonlinear dynamical systems, with the resulting neural approximations enabling

safetyarxiv-cs-lg
4 Jun 2026
Safety

Formal Semantics for Agentic Tool Protocols: A Process Calculus Approach

DGX agent

arXiv:2603.24747v2 Announce Type: replace Abstract: The emergence of large language model agents capable of invoking external tools has created urgent need for formal verification of agent protocols.

safetyarxiv-cs-ai
4 Jun 2026
Safety

Learning Empirically Admissible Neural Heuristics for Combinatorial Search

DGX agent

arXiv:2606.04860v1 Announce Type: cross Abstract: Finding optimal solution paths for combinatorial puzzles like the Rubik's Cube, sliding tile puzzles, and Lights Out remains a classical challenge in

safetyarxiv-cs-ai
4 Jun 2026
Safety

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

DGX agent

arXiv:2606.04027v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safe

safetyarxiv-cs-ai
4 Jun 2026
Safety

Off-Distribution Voices: Fanfiction Subgenres as Universal Vernacular Jailbreaks for Aligned LLMs

DGX agent

arXiv:2606.04483v1 Announce Type: new Abstract: Existing jailbreaks against aligned LLMs are discrete artifacts whose surface forms are easy to fingerprint and patch. We argue that the real failure mo

safetyarxiv-cs-cl
4 Jun 2026
← Previous
1…5051525354…300
Next →