AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,704 results
1 May 2026

I should also add that one of the strongest alignment actions that OpenAI did was to name their product chatgpt with gpt 5.5 medium, names s…

SafetyDGX agent

I should also add that one of the strongest alignment actions that OpenAI did was to name their product chatgpt with gpt 5.5 medium, names so uninspired that nobody could see it as a friend. Unlike Cl

If you think you are going to get alignment out of LLMs you are sadly mistaken. If you live in a society in which people are rolling out LLM…

SafetyDGX agent

If you think you are going to get alignment out of LLMs you are sadly mistaken. If you live in a society in which people are rolling out LLMs at massive scale, without a robust solution to alignment (

Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks

Safety

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2505.13230v3 Announce Type: replace Abstract: Scaling laws in deep learning -- empirical power-law relationships linking model performance to resource growth -- have emerged as simple yet striki

In-context Learning vs. Instruction Tuning: The Case of Small and Multilingual Language Models

SafetyDGX agent

arXiv:2503.01611v3 Announce Type: replace Abstract: Instruction following is a critical ability for Large Language Models to perform downstream tasks. The standard approach to instruction tuning has r

Intern-Atlas: A Methodological Evolution Graph as Research Infrastructure for AI Scientists

SafetyDGX agent

arXiv:2604.28158v1 Announce Type: new Abstract: Existing research infrastructure is fundamentally document-centric, providing citation links between papers but lacking explicit representations of meth

Jensen is one the smartest and most far seeing folks the world. 'If an AI scientist warns people that AI is going to permeate across radiolo…

SafetyDGX agent

Jensen is one the smartest and most far seeing folks the world. 'If an AI scientist warns people that AI is going to permeate across radiology and radiologists are going to get wiped out, it might see

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning

SafetyDGX agent

arXiv:2604.28005v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have increasingly relied on reinforcement learning (RL) to improve their reasoning capabilities. Three a

Knowledge Graph Representations for LLM-Based Policy Compliance Reasoning

SafetyDGX agent

arXiv:2604.27713v1 Announce Type: new Abstract: The risks posed by AI features are increasing as they are rapidly integrated into software applications. In response, regulations and standards for safe

LA-Pose: Latent Action Pretraining Meets Pose Estimation

SafetyDGX agent

arXiv:2604.27448v1 Announce Type: new Abstract: This paper revisits camera pose estimation through the lens of self-supervised pretraining, focusing on inverse-dynamics pretraining as a scalable alter

Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning

SafetyDGX agent

arXiv:2604.27998v1 Announce Type: cross Abstract: Latent reasoning offers a more efficient alternative to explicit reasoning by compressing intermediate reasoning into continuous representations and s

Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care

SafetyDGX agent

arXiv:2604.28010v1 Announce Type: cross Abstract: We reframe clinician overrides of clinical AI recommendations as implicit preference data - the same signal structure exploited by reinforcement learn

Learning Rate Transfer in Normalized Transformers

SafetyDGX agent

arXiv:2604.27077v1 Announce Type: cross Abstract: The Normalized Transformer, or nGPT (arXiv:2410.01131) achieves impressive training speedups and does not require weight decay or learning rate warmup

Learning Tactile-Aware Quadrupedal Loco-Manipulation Policies

SafetyDGX agent

arXiv:2604.27224v1 Announce Type: new Abstract: Quadrupedal loco-manipulation is commonly built on visual perception and proprioception. Yet reliable contact-rich manipulation remains difficult: visio

Learning-to-Explain through 20Q Gaming: An Explainable Recommender for Cybersecurity Education

SafetyDGX agent

arXiv:2604.26964v1 Announce Type: cross Abstract: The growing sophistication of contemporary cyber threats necessitates a more effective and adaptive approach to cybersecurity training. Intuitive and

Learning When to Remember: Risk-Sensitive Contextual Bandits for Abstention-Aware Memory Retrieval in LLM-Based Coding Agents

SafetyDGX agent

arXiv:2604.27283v1 Announce Type: cross Abstract: Large language model (LLM)-based coding agents increasingly rely on external memory to reuse prior debugging experience, repair traces, and repository

Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention

SafetyDGX agent

arXiv:2604.27712v1 Announce Type: cross Abstract: Scene-text image captioning requires fusing three information streams -- visual features, OCR-detected text, and linguistic knowledge -- to generate d

LLM Biases

SafetyDGX agent

arXiv:2604.26960v1 Announce Type: cross Abstract: Transformer-based agentic AI is rapidly being deployed on major platforms to help users shop, watch, and navigate content with less effort. While thes

Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior

SafetyDGX agent

arXiv:2604.27624v1 Announce Type: cross Abstract: Large Language Models (LLMs) can strongly shape social discourse, yet datasets investigating how LLM outputs vary across controlled social and context

“Marcus’ specific point about coding is structurally important: a model that produces code which compiles and passes the tests it was given …

SafetyDGX agent

“Marcus’ specific point about coding is structurally important: a model that produces code which compiles and passes the tests it was given is not the same as a model that produces correct, secure, ma

Mechanized Foundations of Structural Governance: Machine-Checked Proofs for Governed Intelligence

SafetyDGX agent

arXiv:2604.27289v1 Announce Type: new Abstract: We present five results in the theory of structural governance for cognitive workflow systems. Three are mechanized in Coq 8.19 using the Interaction Tr

Meta is basically Black Mirror incarnate.

SafetyDGX agent

Meta is basically Black Mirror incarnate. This is a confusingly written piece, but the upshot is that Meta's smart glasses record even when you don't want them to, and that Meta's data analysis teams

METASYMBO: Multi-Agent Language-Guided Metamaterial Discovery via Symbolic Latent Evolution

SafetyDGX agent

arXiv:2604.27300v1 Announce Type: new Abstract: Metamaterial discovery seeks microstructured materials whose geometry induces targeted mechanical behavior. Existing inverse-design methods can efficien

MIFair: A Mutual-Information Framework for Intersectionality and Multiclass Fairness

SafetyDGX agent

arXiv:2604.28030v1 Announce Type: cross Abstract: Fairness in machine learning remains challenging due to its ethical complexity, the absence of a universal definition, and the need for context-specif

Mind the Gap: Structure-Aware Consistency in Preference Learning

SafetyDGX agent

arXiv:2604.27733v1 Announce Type: new Abstract: Preference learning has become the foundation of aligning Large Language Models (LLMs) with human intent. Popular methods, such as Direct Preference Opt

Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO

SafetyDGX agent

arXiv:2603.21016v2 Announce Type: replace-cross Abstract: Large language models (LLMs) used for multiple-choice and pairwise evaluation tasks often exhibit selection bias due to non-semantic factors l

MotuBrain: An Advanced World Action Model for Robot Control

SafetyDGX agent

arXiv:2604.27792v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models achieve strong semantic generalization but often lack fine-grained modeling of world dynamics. Recent work explores

MSR:Hybrid Field Modeling for CT-MRI Rigid-Deformable Registration of the Cervical Spine with an Annotated Dataset

SafetyDGX agent

arXiv:2604.27654v1 Announce Type: new Abstract: Accurate CT-MRI registration of the cervical spine is essential for preoperative planning because this region is anatomically complex,highly variable,an

OmniRobotHome: A Multi-Camera Platform for Real-Time Multiadic Human-Robot Interaction

SafetyDGX agent

arXiv:2604.28197v1 Announce Type: cross Abstract: Human-robot collaboration has been studied primarily in dyadic or sequential settings. However, real homes require multiadic collaboration, where mult

One of the things I hate the most about this site is the consistent lack of nuance. That’s why everything is an argument, and progress here …

SafetyDGX agent

Gary Marcus critiques social media platforms for lacking nuance in discourse, which he identifies as a root cause of persistent arguments and stalled progress on the site. The post reflects concerns a

Online semi-supervised perception: Real-time learning without explicit feedback

SafetyDGX agent

arXiv:2604.27562v1 Announce Type: new Abstract: This paper proposes an algorithm for real-time learning without explicit feedback. The algorithm combines the ideas of semi-supervised learning on graph

OpAgent: Operator Agent for Web Navigation

SafetyDGX agent

arXiv:2602.13559v2 Announce Type: replace Abstract: To fulfill user instructions, autonomous web agents must contend with the inherent complexity and volatile nature of real-world websites. Convention

OpenAI o1 System Card

SafetyDGX agent

arXiv:2412.16720v2 Announce Type: replace Abstract: The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provi

PALCAS: A Priority-Aware Intelligent Lane Change Advisory System for Autonomous Vehicles using Federated Reinforcement Learning

SafetyDGX agent

arXiv:2604.27118v1 Announce Type: cross Abstract: We present a priority-aware intelligent lane change advisory system based on multi-agent federated reinforcement learning, namely PALCAS, for autonomo

Performance-Driven QUBO for Recommender Systems on Quantum Annealers

SafetyDGX agent

arXiv:2410.15272v3 Announce Type: replace-cross Abstract: Quantum annealers offer a promising hardware platform for solving combinatorial optimization problems, especially those formulated as Quadrati

Policy-Grounded Safety Evaluation of 20 Large Language Models

SafetyDGX agent

arXiv:2507.14719v2 Announce Type: replace Abstract: As large language models (LLMs) become increasingly integrated into real-world applications, scalable and rigorous safety evaluation is essential. T

Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor

SafetyDGX agent

arXiv:2604.27633v1 Announce Type: new Abstract: Large language models (LLMs) are commonly evaluated for political bias based on their responses to fixed questionnaires, which typically place frontier

Preserving Temporal Dynamics in Time Series Generation

SafetyDGX agent

arXiv:2604.27182v1 Announce Type: cross Abstract: Time-series data augmentation plays a crucial role in regression-oriented forecasting tasks, where limited data restricts the performance of deep lear

Real-Time GPU-Accelerated Monte Carlo Evaluation of Safety-Critical AEB Systems Under Uncertainty

SafetyDGX agent

arXiv:2604.27193v1 Announce Type: new Abstract: Automatic Emergency Braking (AEB) systems represent a safety-critical national interest, with the National Highway Traffic Safety Administration (NHTSA)

Residual Gaussian Splatting for Ultra Sparse-View CBCT Reconstruction

SafetyDGX agent

arXiv:2604.27552v1 Announce Type: new Abstract: While 3D Gaussian splatting (3DGS) offers explicit and efficient scene representations for cone-beam computed tomography reconstruction, conventional ph

RHyVE: Competence-Aware Verification and Phase-Aware Deployment for LLM-Generated Reward Hypotheses

SafetyDGX agent

arXiv:2604.28056v1 Announce Type: new Abstract: Large language models (LLMs) make reward design in reinforcement learning substantially more scalable, but generated rewards are not automatically relia

Robot Learning from Human Videos: A Survey

SafetyDGX agent

arXiv:2604.27621v1 Announce Type: cross Abstract: A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of

Safe Bilevel Delegation (SBD): A Formal Framework for Runtime Delegation Safety in Multi-Agent Systems

SafetyDGX agent

arXiv:2604.27358v1 Announce Type: new Abstract: As large language model (LLM) agents are deployed in high-stakes environments, the question of how safely to delegate subtasks to specialized sub-agents

Sam Altman is a master at insincerity, Author @GaryMarcus claims. 'He is a master of projecting insincerity, a master at telling the room wh…

SafetyDGX agent

Sam Altman is a master at insincerity, Author @GaryMarcus claims. 'He is a master of projecting insincerity, a master at telling the room what it wants to hear and not always truthful'. “Altman, sitti

Sample-efficient evidence estimation of score based priors for model selection

SafetyDGX agent

arXiv:2602.20549v2 Announce Type: replace-cross Abstract: The choice of prior is central to solving ill-posed imaging inverse problems, making it essential to select one consistent with the measuremen

SCOOP: A pro-AI dark money group backed by a powerful super PAC funded by execs tied to Palantir and OpenAI, has been secretly paying influe…

SafetyDGX agent

SCOOP: A pro-AI dark money group backed by a powerful super PAC funded by execs tied to Palantir and OpenAI, has been secretly paying influencers to push pro-AI, anti-China propaganda on TikTok and IG

Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception

SafetyDGX agent

arXiv:2604.28048v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used as proxies for human perception in urban analysis, yet it remains unclear whether persona prompting p

Stating that women do not have penises is conservative. Stating that the scientific method is superior to ancestral tribal dances for seekin…

SafetyDGX agent

Stating that women do not have penises is conservative. Stating that the scientific method is superior to ancestral tribal dances for seeking truth is conservative. Supporting a rational immigration p

Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification

SafetyDGX agent

arXiv:2602.16516v2 Announce Type: replace Abstract: This paper introduces ParlaCAP, a large-scale dataset for analyzing parliamentary agenda setting across Europe, and proposes a cost-effective method

Taxon: Hierarchical Tax Code Prediction with Semantically Aligned LLM Expert Guidance

SafetyDGX agent

arXiv:2601.08418v2 Announce Type: replace-cross Abstract: Tax code prediction is a crucial yet underexplored task in automating invoicing and compliance management for large-scale e-commerce platforms

Test Before You Deploy: Governing Updates in the LLM Supply Chain

SafetyDGX agent

arXiv:2604.27789v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used as core dependencies in software systems. However, the hosted LLM services evolve continuously thro

Test-Time Distillation for Continual Model Adaptation

SafetyDGX agent

arXiv:2506.02671v3 Announce Type: replace Abstract: Deep neural networks often suffer performance degradation upon deployment due to distribution shifts. Continual Test-Time Adaptation (CTTA) aims to

The Effects of Visual Priming on Cooperative Behavior in Vision-Language Models

SafetyDGX agent

arXiv:2604.27953v1 Announce Type: new Abstract: As Vision-Language Models (VLMs) become increasingly integrated into decision-making systems, it is essential to understand how visual inputs influence

The Field of Safe Motion: Operationalizing Affordances in the Field of Safe Travel Using Reachability Analysis

SafetyDGX agent

arXiv:2604.27168v1 Announce Type: new Abstract: We present the Field of Safe Motion (FSM), a quantitative safety model for determining whether a driver maintains a collision-free escape route, or 'out

The Likelihood Ratio Wall: Structural Limits on Accurate Risk Assessment for Rare Violence

SafetyDGX agent

arXiv:2604.27282v1 Announce Type: cross Abstract: Pretrial risk assessment tools are used on over one million U.S. defendants each year, yet their use for predicting rare violent re-offense faces a ba

The Two Boundaries: Why Behavioral AI Governance Fails Structurally

SafetyDGX agent

arXiv:2604.27292v1 Announce Type: new Abstract: Every system that performs effects has two boundaries: what it can do (expressiveness) and what governance covers (governance). In nearly all deployed A

Tokenmaxxing is stupid. Change my mind?

SafetyDGX agent

Gary Marcus critiques the AI industry's focus on scaling model parameters and training data (tokenmaxxing) as an inefficient approach to advancing AI capabilities. He argues that simply increasing tok

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance

SafetyDGX agent

arXiv:2601.20239v4 Announce Type: replace Abstract: Fine-grained and contact-rich manipulation remain challenging for robots, largely due to the underutilization of tactile feedback. To address this,

Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles

SafetyDGX agent

arXiv:2604.28087v1 Announce Type: cross Abstract: Rule-based systems remain central in safety-critical domains but often struggle with scalability, brittleness, and goal misspecification. These limita

TRUST: A Framework for Decentralized AI Service v.0.1

SafetyDGX agent

arXiv:2604.27132v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) and Multi-Agent Systems (MAS) in high-stakes domains demand reliable verification, yet centralized approaches suffer four

Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues

SafetyDGX agent

arXiv:2506.05412v3 Announce Type: replace-cross Abstract: Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer ga

← Previous
1…170171172173174…212
Next →