AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,707 results
Safety

Connected Dependability Cage: Run-Time Function and Anomaly Monitoring for the Development and Operation of Safe Automated Vehicles

DGX agent

arXiv:2604.27728v1 Announce Type: new Abstract: The advancement of automated vehicles introduces complex safety challenges, particularly in dynamic and unpredictable environments where AI-enabled perc

safetyarxiv-cs-ro
1 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk

DGX agent

arXiv:2601.22993v3 Announce Type: replace Abstract: We introduce the Value-at-Risk Constrained Policy Optimization algorithm (VaR-CPO), a sample efficient and conservative method designed to optimize

safetyarxiv-cs-lg
1 May 2026
Safety

Consumer Attitudes Towards AI in Digital Health: A Mixed-Methods Survey in Australia

DGX agent

arXiv:2604.27744v1 Announce Type: new Abstract: AI applications are increasingly being introduced into digital health. While technical performance has advanced rapidly, successful deployment mainly de

safetyarxiv-cs-ai
1 May 2026
Safety

Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations

DGX agent

arXiv:2604.27372v1 Announce Type: cross Abstract: This paper investigates the continuous-time counterpart of the Q-function for entropy-regularized mean-field control (MFC) with controlled common nois

safetyarxiv-cs-lg
1 May 2026
Safety

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning…

DGX agent

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning fixes get bolted on at post-training. By then, the patterns

safetydair-ai--x
1 May 2026
Safety

Cost-Aware Learning

DGX agent

arXiv:2604.28020v1 Announce Type: new Abstract: We consider the problem of Cost-Aware Learning, where sampling different component functions of a finite-sum objective incurs different costs. The objec

safetyarxiv-cs-lg
1 May 2026
Safety

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability

DGX agent

arXiv:2602.17469v2 Announce Type: replace Abstract: Recent advances in multilingual representation learning aim to bridge the performance gap between high- and low-resource languages, yet their abilit

safetyarxiv-cs-cl
1 May 2026
Safety

Cross-Subject Generalization for EEG Decoding: A Survey of Deep Learning Methods

DGX agent

arXiv:2604.27033v1 Announce Type: new Abstract: Deep learning for cross-subject EEG decoding is hindered by high inter-subject variability, which introduces a severe domain shift between training and

safetyarxiv-cs-lg
1 May 2026
Safety

Customer service in the age of AI has been become truly horrible. Excruciatingly bad.

DGX agent

Customer service in the age of AI has been become truly horrible. Excruciatingly bad. Hey @DHLCanadaHelp, your customer service really and truly sucks. You claimed to try to deliver a package, but did

safetygary-marcus--x
1 May 2026
Safety

D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery

DGX agent

arXiv:2604.27977v1 Announce Type: new Abstract: Despite recent progress in language models and agents for scientific data-driven discovery, further advancing their capabilities is held back by the abs

safetyarxiv-cs-ai
1 May 2026
Safety

Debiasing Reward Models via Causally Motivated Inference-Time Intervention

DGX agent

arXiv:2604.27495v1 Announce Type: cross Abstract: Reward models (RMs) play a central role in aligning large language models (LLMs) with human preferences. However, RMs are often sensitive to spurious

safetyarxiv-cs-ai
1 May 2026
Safety

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards

DGX agent

arXiv:2603.09117v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) significantly enhances large language models (LLMs) reasoning but severely suffers from

safetyarxiv-cs-ai
1 May 2026
Safety

Design Structure Matrix Modularization with Large Language Models

DGX agent

arXiv:2604.28018v1 Announce Type: cross Abstract: Design Structure Matrix (DSM) modularization, the task of partitioning system elements into cohesive modules, is a fundamental combinatorial challenge

safetyarxiv-cs-ai
1 May 2026
Safety

Designing Ethical Learning for Agentic AI: Toegye Yi Hwang's Ethical Emotion Regulation Framework

DGX agent

arXiv:2604.26958v1 Announce Type: cross Abstract: Agentic AI systems capable of autonomous goal setting and proactive intervention introduce new challenges for regulating moral-emotional processes in

safetyarxiv-cs-ai
1 May 2026
Safety

Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture

DGX agent

arXiv:2604.27045v1 Announce Type: cross Abstract: As Large Language Model (LLM) agents transition from single-session tools to persistent systems managing longitudinal healthcare journeys, their memor

safetyarxiv-cs-ai
1 May 2026
Safety

Distributional Alignment Games for Answer-Level Fine-Tuning

DGX agent

arXiv:2604.27166v1 Announce Type: new Abstract: We focus on the problem of Answer-Level Fine-Tuning (ALFT), where the goal is to optimize a language model based on the correctness or properties of its

safetyarxiv-cs-lg
1 May 2026
Safety

DOT-Sim: Differentiable Optical Tactile Simulation with Precise Real-to-Sim Physical Calibration

DGX agent

arXiv:2604.27367v1 Announce Type: cross Abstract: Simulating optical tactile sensors presents significant challenges due to their high deformability and intricate optical properties. To address these

safetyarxiv-cs-cv
1 May 2026
Safety

Dreaming Across Towns: Semantic Rollout and Town-Adversarial Regularization for Zero-Shot Held-Out-Town Fixed-Route Driving in CARLA

DGX agent

arXiv:2604.27994v1 Announce Type: new Abstract: Learned driving agents often degrade when deployed in unseen environments. This paper studies a deliberately bounded instance of that problem in the CAR

safetyarxiv-cs-ro
1 May 2026
Safety

Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry

DGX agent

arXiv:2604.27019v1 Announce Type: cross Abstract: Safety-aligned language models must refuse harmful requests without collapsing into broad over-refusal, but the training-time mechanisms behind this t

safetyarxiv-cs-cl
1 May 2026
Safety

Efficient Preimage Approximation for Neural Network Certification

DGX agent

arXiv:2505.22798v3 Announce Type: replace-cross Abstract: The growing reliance on artificial intelligence in safety- and security-critical applications is raising concerns about the robustness of neur

safetyarxiv-cs-ai
1 May 2026
Safety

Elon Musk, Sam Altman, the future of humanity, and … goblins. Terrific interview @theinformation with @rocketalignment https://www.youtube.c…

DGX agent

This post references an interview featuring Elon Musk and Sam Altman discussing AI safety, existential risks, and the future of humanity, with an unconventional or humorous element involving goblins.

safetygary-marcus--x
1 May 2026
Safety

Exploration Hacking: Can LLMs Learn to Resist RL Training?

DGX agent

arXiv:2604.28182v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become essential to the post-training of large language models (LLMs) for reasoning, agentic capabilities and alignmen

safetyarxiv-cs-cl
1 May 2026
Safety

Exploring Applications of Transfer-State Large Language Models: Cognitive Profiling and Socratic AI Tutoring

DGX agent

arXiv:2604.27454v1 Announce Type: new Abstract: Large language models (LLMs) sometimes exhibit qualitative shifts in response style under sustained self-referential dialogue conditions (Berg et al., 2

safetyarxiv-cs-cl
1 May 2026
Safety

EXPO: Stable Reinforcement Learning with Expressive Policies

DGX agent

arXiv:2507.07986v3 Announce Type: replace-cross Abstract: We study the problem of training and fine-tuning expressive policies with online reinforcement learning (RL) given an offline dataset. Trainin

safetyarxiv-cs-ai
1 May 2026
Safety

Fairness for distribution network operations and planning

DGX agent

arXiv:2604.27669v1 Announce Type: new Abstract: The incorporation of fairness into the distribution network (DN) planning and operation has become a key goal of recent studies. The cost of implementin

safetyarxiv-cs-ai
1 May 2026
Safety

Focus Session: Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification

DGX agent

arXiv:2604.27807v1 Announce Type: new Abstract: The design of embedded safety-critical systems such as those used in next-generation automotive and autonomous platforms, is increasingly challenged by

safetyarxiv-cs-ai
1 May 2026
Safety

FP-IRL: Fokker--Planck Inverse Reinforcement Learning -- A Physics-Constrained Approach to Markov Decision Processes

DGX agent

arXiv:2306.10407v3 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is a powerful paradigm for uncovering the incentive structure that drives agent behavior, by inferring an

safetyarxiv-cs-ai
1 May 2026
Safety

Frequency-Aware Semantic Fusion with Gated Injection for AI-generated Image Detection

DGX agent

arXiv:2604.27875v1 Announce Type: new Abstract: AI-generated images are becoming increasingly realistic and diverse, posing significant challenges for generalizable detection. While Vision Foundation

safetyarxiv-cs-cv
1 May 2026
Safety

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback

DGX agent

arXiv:2502.07645v3 Announce Type: replace Abstract: Behavior cloning (BC) optimizes policies by treating human demonstrations as pointwise action labels. While effective with accurate action labels, t

safetyarxiv-cs-ro
1 May 2026
Safety

From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems

DGX agent

arXiv:2604.27267v1 Announce Type: cross Abstract: As large language models are integrated into autonomous robotic systems for task planning and control, compromised inputs or unsafe model outputs can

safetyarxiv-cs-ai
1 May 2026
Safety

From surveillance to signalling: escalation channels as environmental controls for agentic AI

DGX agent

arXiv:2510.05192v2 Announce Type: replace-cross Abstract: When AI agents operating with access to sensitive information encounter a conflict between completing an assigned task and following rules or

safetyarxiv-cs-ai
1 May 2026
Safety

GAVEL: Towards Rule-Based Safety Through Activation Monitoring

DGX agent

arXiv:2601.19768v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be appare

safetyarxiv-cs-ai
1 May 2026
Safety

GSDrive: Reinforcing Driving Policies by Multi-mode Trajectory Probing with 3D Gaussian Splatting Environment

DGX agent

arXiv:2604.28111v1 Announce Type: new Abstract: End-to-end (E2E) autonomous driving presents a promising approach for translating perceptual inputs directly into driving actions. However, prohibitive

safetyarxiv-cs-ro
1 May 2026
Safety

Hey @DHLCanadaHelp, your customer service really and truly sucks. You claimed to try to deliver a package, but didn’t actually contact me, d…

DGX agent

Hey @DHLCanadaHelp, your customer service really and truly sucks. You claimed to try to deliver a package, but didn’t actually contact me, didn’t leave a service card, your automated software won’t le

safetygary-marcus--x
1 May 2026
Safety

How Hard Is Continuous Clustering? Lower Bounds from the Existential Theory of the Reals

DGX agent

arXiv:2604.26972v1 Announce Type: cross Abstract: This paper studies the computational difficulty of clustering problems that are defined directly on a continuous probability density. Rather than work

safetyarxiv-cs-lg
1 May 2026
Safety

How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

DGX agent

arXiv:2604.27147v1 Announce Type: cross Abstract: In generative modeling, we often wish to produce samples that maximize a user-specified reward such as aesthetic quality or alignment with human prefe

safetyarxiv-cs-ai
1 May 2026
Safety

I dunno. The competition is really tight. But Zuck certainly a top contender! Who’s your “favorite”?

DGX agent

This appears to be a casual social media post by AI researcher Gary Marcus discussing competitive dynamics in the AI field, with a conversational reference to Mark Zuckerberg as a notable figure or 't

safetygary-marcus--x
1 May 2026
Safety

I have many beefs with Dario and don’t trust him or his hype — but he has certainly eaten OpenAI’s lunch, despite their immense initial lead…

DGX agent

I have many beefs with Dario and don’t trust him or his hype — but he has certainly eaten OpenAI’s lunch, despite their immense initial lead. Maybe “clown” isn’t the right word here. I think Jensen ag

safetygary-marcus--x
1 May 2026
Safety

I just had my first ride with HOVR, a Canadian alternative to Uber. They are cheaper and pay their drivers more. They already have over 3000…

DGX agent

I just had my first ride with HOVR, a Canadian alternative to Uber. They are cheaper and pay their drivers more. They already have over 3000 drivers in the Greater Toronto Area. You can download HOVR

safetygeoffrey-hinton--x
1 May 2026
Safety

I should also add that one of the strongest alignment actions that OpenAI did was to name their product chatgpt with gpt 5.5 medium, names s…

DGX agent

I should also add that one of the strongest alignment actions that OpenAI did was to name their product chatgpt with gpt 5.5 medium, names so uninspired that nobody could see it as a friend. Unlike Cl

safetyethan-mollick--x
1 May 2026
Safety

If you think you are going to get alignment out of LLMs you are sadly mistaken. If you live in a society in which people are rolling out LLM…

DGX agent

If you think you are going to get alignment out of LLMs you are sadly mistaken. If you live in a society in which people are rolling out LLMs at massive scale, without a robust solution to alignment (

safetygary-marcus--x
1 May 2026
Safety

Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks

DGX agent

arXiv:2505.13230v3 Announce Type: replace Abstract: Scaling laws in deep learning -- empirical power-law relationships linking model performance to resource growth -- have emerged as simple yet striki

safetyarxiv-cs-lg
1 May 2026
Safety

In-context Learning vs. Instruction Tuning: The Case of Small and Multilingual Language Models

DGX agent

arXiv:2503.01611v3 Announce Type: replace Abstract: Instruction following is a critical ability for Large Language Models to perform downstream tasks. The standard approach to instruction tuning has r

safetyarxiv-cs-cl
1 May 2026
Safety

Intern-Atlas: A Methodological Evolution Graph as Research Infrastructure for AI Scientists

DGX agent

arXiv:2604.28158v1 Announce Type: new Abstract: Existing research infrastructure is fundamentally document-centric, providing citation links between papers but lacking explicit representations of meth

safetyarxiv-cs-ai
1 May 2026
Safety

Jensen is one the smartest and most far seeing folks the world. 'If an AI scientist warns people that AI is going to permeate across radiolo…

DGX agent

Jensen is one the smartest and most far seeing folks the world. 'If an AI scientist warns people that AI is going to permeate across radiology and radiologists are going to get wiped out, it might see

safetyclem-delangue--x
1 May 2026
Safety

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning

DGX agent

arXiv:2604.28005v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have increasingly relied on reinforcement learning (RL) to improve their reasoning capabilities. Three a

safetyarxiv-cs-lg
1 May 2026
Safety

Knowledge Graph Representations for LLM-Based Policy Compliance Reasoning

DGX agent

arXiv:2604.27713v1 Announce Type: new Abstract: The risks posed by AI features are increasing as they are rapidly integrated into software applications. In response, regulations and standards for safe

safetyarxiv-cs-ai
1 May 2026
Safety

LA-Pose: Latent Action Pretraining Meets Pose Estimation

DGX agent

arXiv:2604.27448v1 Announce Type: new Abstract: This paper revisits camera pose estimation through the lens of self-supervised pretraining, focusing on inverse-dynamics pretraining as a scalable alter

safetyarxiv-cs-cv
1 May 2026
← Previous
1…212213214215216…265
Next →