AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,357 results
Safety

D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery

DGX agent

arXiv:2604.27977v1 Announce Type: new Abstract: Despite recent progress in language models and agents for scientific data-driven discovery, further advancing their capabilities is held back by the abs

safetyarxiv-cs-ai
1 May 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Debiasing Reward Models via Causally Motivated Inference-Time Intervention

DGX agent

arXiv:2604.27495v1 Announce Type: cross Abstract: Reward models (RMs) play a central role in aligning large language models (LLMs) with human preferences. However, RMs are often sensitive to spurious

safetyarxiv-cs-ai
1 May 2026
Safety

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards

DGX agent

arXiv:2603.09117v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) significantly enhances large language models (LLMs) reasoning but severely suffers from

safetyarxiv-cs-ai
1 May 2026
Safety

Design Structure Matrix Modularization with Large Language Models

DGX agent

arXiv:2604.28018v1 Announce Type: cross Abstract: Design Structure Matrix (DSM) modularization, the task of partitioning system elements into cohesive modules, is a fundamental combinatorial challenge

safetyarxiv-cs-ai
1 May 2026
Safety

Designing Ethical Learning for Agentic AI: Toegye Yi Hwang's Ethical Emotion Regulation Framework

DGX agent

arXiv:2604.26958v1 Announce Type: cross Abstract: Agentic AI systems capable of autonomous goal setting and proactive intervention introduce new challenges for regulating moral-emotional processes in

safetyarxiv-cs-ai
1 May 2026
Safety

Distributional Alignment Games for Answer-Level Fine-Tuning

DGX agent

arXiv:2604.27166v1 Announce Type: new Abstract: We focus on the problem of Answer-Level Fine-Tuning (ALFT), where the goal is to optimize a language model based on the correctness or properties of its

safetyarxiv-cs-lg
1 May 2026
Safety

DOT-Sim: Differentiable Optical Tactile Simulation with Precise Real-to-Sim Physical Calibration

DGX agent

arXiv:2604.27367v1 Announce Type: cross Abstract: Simulating optical tactile sensors presents significant challenges due to their high deformability and intricate optical properties. To address these

safetyarxiv-cs-cv
1 May 2026
Safety

Exploration Hacking: Can LLMs Learn to Resist RL Training?

DGX agent

arXiv:2604.28182v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become essential to the post-training of large language models (LLMs) for reasoning, agentic capabilities and alignmen

safetyarxiv-cs-cl
1 May 2026
Safety

Exploring Applications of Transfer-State Large Language Models: Cognitive Profiling and Socratic AI Tutoring

DGX agent

arXiv:2604.27454v1 Announce Type: new Abstract: Large language models (LLMs) sometimes exhibit qualitative shifts in response style under sustained self-referential dialogue conditions (Berg et al., 2

safetyarxiv-cs-cl
1 May 2026
Safety

EXPO: Stable Reinforcement Learning with Expressive Policies

DGX agent

arXiv:2507.07986v3 Announce Type: replace-cross Abstract: We study the problem of training and fine-tuning expressive policies with online reinforcement learning (RL) given an offline dataset. Trainin

safetyarxiv-cs-ai
1 May 2026
Safety

Fairness for distribution network operations and planning

DGX agent

arXiv:2604.27669v1 Announce Type: new Abstract: The incorporation of fairness into the distribution network (DN) planning and operation has become a key goal of recent studies. The cost of implementin

safetyarxiv-cs-ai
1 May 2026
Safety

FP-IRL: Fokker--Planck Inverse Reinforcement Learning -- A Physics-Constrained Approach to Markov Decision Processes

DGX agent

arXiv:2306.10407v3 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is a powerful paradigm for uncovering the incentive structure that drives agent behavior, by inferring an

safetyarxiv-cs-ai
1 May 2026
Safety

Frequency-Aware Semantic Fusion with Gated Injection for AI-generated Image Detection

DGX agent

arXiv:2604.27875v1 Announce Type: new Abstract: AI-generated images are becoming increasingly realistic and diverse, posing significant challenges for generalizable detection. While Vision Foundation

safetyarxiv-cs-cv
1 May 2026
Safety

From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback

DGX agent

arXiv:2502.07645v3 Announce Type: replace Abstract: Behavior cloning (BC) optimizes policies by treating human demonstrations as pointwise action labels. While effective with accurate action labels, t

safetyarxiv-cs-ro
1 May 2026
Safety

GSDrive: Reinforcing Driving Policies by Multi-mode Trajectory Probing with 3D Gaussian Splatting Environment

DGX agent

arXiv:2604.28111v1 Announce Type: new Abstract: End-to-end (E2E) autonomous driving presents a promising approach for translating perceptual inputs directly into driving actions. However, prohibitive

safetyarxiv-cs-ro
1 May 2026
Safety

Hey @DHLCanadaHelp, your customer service really and truly sucks. You claimed to try to deliver a package, but didn’t actually contact me, d…

DGX agent

Hey @DHLCanadaHelp, your customer service really and truly sucks. You claimed to try to deliver a package, but didn’t actually contact me, didn’t leave a service card, your automated software won’t le

safetygary-marcus--x
1 May 2026
Safety

How Hard Is Continuous Clustering? Lower Bounds from the Existential Theory of the Reals

DGX agent

arXiv:2604.26972v1 Announce Type: cross Abstract: This paper studies the computational difficulty of clustering problems that are defined directly on a continuous probability density. Rather than work

safetyarxiv-cs-lg
1 May 2026
Safety

How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance

DGX agent

arXiv:2604.27147v1 Announce Type: cross Abstract: In generative modeling, we often wish to produce samples that maximize a user-specified reward such as aesthetic quality or alignment with human prefe

safetyarxiv-cs-ai
1 May 2026
Safety

I dunno. The competition is really tight. But Zuck certainly a top contender! Who’s your “favorite”?

DGX agent

This appears to be a casual social media post by AI researcher Gary Marcus discussing competitive dynamics in the AI field, with a conversational reference to Mark Zuckerberg as a notable figure or 't

safetygary-marcus--x
1 May 2026
Safety

I have many beefs with Dario and don’t trust him or his hype — but he has certainly eaten OpenAI’s lunch, despite their immense initial lead…

DGX agent

I have many beefs with Dario and don’t trust him or his hype — but he has certainly eaten OpenAI’s lunch, despite their immense initial lead. Maybe “clown” isn’t the right word here. I think Jensen ag

safetygary-marcus--x
1 May 2026
Safety

I just had my first ride with HOVR, a Canadian alternative to Uber. They are cheaper and pay their drivers more. They already have over 3000…

DGX agent

I just had my first ride with HOVR, a Canadian alternative to Uber. They are cheaper and pay their drivers more. They already have over 3000 drivers in the Greater Toronto Area. You can download HOVR

safetygeoffrey-hinton--x
1 May 2026
Safety

I should also add that one of the strongest alignment actions that OpenAI did was to name their product chatgpt with gpt 5.5 medium, names s…

DGX agent

I should also add that one of the strongest alignment actions that OpenAI did was to name their product chatgpt with gpt 5.5 medium, names so uninspired that nobody could see it as a friend. Unlike Cl

safetyethan-mollick--x
1 May 2026
Safety

If you think you are going to get alignment out of LLMs you are sadly mistaken. If you live in a society in which people are rolling out LLM…

DGX agent

If you think you are going to get alignment out of LLMs you are sadly mistaken. If you live in a society in which people are rolling out LLMs at massive scale, without a robust solution to alignment (

safetygary-marcus--x
1 May 2026
Safety

Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks

DGX agent

arXiv:2505.13230v3 Announce Type: replace Abstract: Scaling laws in deep learning -- empirical power-law relationships linking model performance to resource growth -- have emerged as simple yet striki

safetyarxiv-cs-lg
1 May 2026
Safety

In-context Learning vs. Instruction Tuning: The Case of Small and Multilingual Language Models

DGX agent

arXiv:2503.01611v3 Announce Type: replace Abstract: Instruction following is a critical ability for Large Language Models to perform downstream tasks. The standard approach to instruction tuning has r

safetyarxiv-cs-cl
1 May 2026
Safety

Intern-Atlas: A Methodological Evolution Graph as Research Infrastructure for AI Scientists

DGX agent

arXiv:2604.28158v1 Announce Type: new Abstract: Existing research infrastructure is fundamentally document-centric, providing citation links between papers but lacking explicit representations of meth

safetyarxiv-cs-ai
1 May 2026
Safety

Jensen is one the smartest and most far seeing folks the world. 'If an AI scientist warns people that AI is going to permeate across radiolo…

DGX agent

Jensen is one the smartest and most far seeing folks the world. 'If an AI scientist warns people that AI is going to permeate across radiology and radiologists are going to get wiped out, it might see

safetyclem-delangue--x
1 May 2026
Safety

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning

DGX agent

arXiv:2604.28005v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have increasingly relied on reinforcement learning (RL) to improve their reasoning capabilities. Three a

safetyarxiv-cs-lg
1 May 2026
Safety

Knowledge Graph Representations for LLM-Based Policy Compliance Reasoning

DGX agent

arXiv:2604.27713v1 Announce Type: new Abstract: The risks posed by AI features are increasing as they are rapidly integrated into software applications. In response, regulations and standards for safe

safetyarxiv-cs-ai
1 May 2026
Safety

LA-Pose: Latent Action Pretraining Meets Pose Estimation

DGX agent

arXiv:2604.27448v1 Announce Type: new Abstract: This paper revisits camera pose estimation through the lens of self-supervised pretraining, focusing on inverse-dynamics pretraining as a scalable alter

safetyarxiv-cs-cv
1 May 2026
Safety

Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning

DGX agent

arXiv:2604.27998v1 Announce Type: cross Abstract: Latent reasoning offers a more efficient alternative to explicit reasoning by compressing intermediate reasoning into continuous representations and s

safetyarxiv-cs-cl
1 May 2026
Safety

Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care

DGX agent

arXiv:2604.28010v1 Announce Type: cross Abstract: We reframe clinician overrides of clinical AI recommendations as implicit preference data - the same signal structure exploited by reinforcement learn

safetyarxiv-cs-ai
1 May 2026
Safety

Learning Rate Transfer in Normalized Transformers

DGX agent

arXiv:2604.27077v1 Announce Type: cross Abstract: The Normalized Transformer, or nGPT (arXiv:2410.01131) achieves impressive training speedups and does not require weight decay or learning rate warmup

safetyarxiv-cs-ai
1 May 2026
Safety

Learning Tactile-Aware Quadrupedal Loco-Manipulation Policies

DGX agent

arXiv:2604.27224v1 Announce Type: new Abstract: Quadrupedal loco-manipulation is commonly built on visual perception and proprioception. Yet reliable contact-rich manipulation remains difficult: visio

safetyarxiv-cs-ro
1 May 2026
Safety

Learning-to-Explain through 20Q Gaming: An Explainable Recommender for Cybersecurity Education

DGX agent

arXiv:2604.26964v1 Announce Type: cross Abstract: The growing sophistication of contemporary cyber threats necessitates a more effective and adaptive approach to cybersecurity training. Intuitive and

safetyarxiv-cs-ai
1 May 2026
Safety

Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention

DGX agent

arXiv:2604.27712v1 Announce Type: cross Abstract: Scene-text image captioning requires fusing three information streams -- visual features, OCR-detected text, and linguistic knowledge -- to generate d

safetyarxiv-cs-cl
1 May 2026
Safety

LLM Biases

DGX agent

arXiv:2604.26960v1 Announce Type: cross Abstract: Transformer-based agentic AI is rapidly being deployed on major platforms to help users shop, watch, and navigate content with less effort. While thes

safetyarxiv-cs-ai
1 May 2026
Safety

Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior

DGX agent

arXiv:2604.27624v1 Announce Type: cross Abstract: Large Language Models (LLMs) can strongly shape social discourse, yet datasets investigating how LLM outputs vary across controlled social and context

safetyarxiv-cs-ai
1 May 2026
Safety

“Marcus’ specific point about coding is structurally important: a model that produces code which compiles and passes the tests it was given …

DGX agent

“Marcus’ specific point about coding is structurally important: a model that produces code which compiles and passes the tests it was given is not the same as a model that produces correct, secure, ma

safetygary-marcus--x
1 May 2026
Safety

Meta is basically Black Mirror incarnate.

DGX agent

Meta is basically Black Mirror incarnate. This is a confusingly written piece, but the upshot is that Meta's smart glasses record even when you don't want them to, and that Meta's data analysis teams

safetygary-marcus--x
1 May 2026
Safety

METASYMBO: Multi-Agent Language-Guided Metamaterial Discovery via Symbolic Latent Evolution

DGX agent

arXiv:2604.27300v1 Announce Type: new Abstract: Metamaterial discovery seeks microstructured materials whose geometry induces targeted mechanical behavior. Existing inverse-design methods can efficien

safetyarxiv-cs-ai
1 May 2026
Safety

MIFair: A Mutual-Information Framework for Intersectionality and Multiclass Fairness

DGX agent

arXiv:2604.28030v1 Announce Type: cross Abstract: Fairness in machine learning remains challenging due to its ethical complexity, the absence of a universal definition, and the need for context-specif

safetyarxiv-cs-ai
1 May 2026
Safety

Mind the Gap: Structure-Aware Consistency in Preference Learning

DGX agent

arXiv:2604.27733v1 Announce Type: new Abstract: Preference learning has become the foundation of aligning Large Language Models (LLMs) with human intent. Popular methods, such as Direct Preference Opt

safetyarxiv-cs-lg
1 May 2026
Safety

Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO

DGX agent

arXiv:2603.21016v2 Announce Type: replace-cross Abstract: Large language models (LLMs) used for multiple-choice and pairwise evaluation tasks often exhibit selection bias due to non-semantic factors l

safetyarxiv-cs-ai
1 May 2026
Safety

MotuBrain: An Advanced World Action Model for Robot Control

DGX agent

arXiv:2604.27792v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models achieve strong semantic generalization but often lack fine-grained modeling of world dynamics. Recent work explores

safetyarxiv-cs-ro
1 May 2026
Safety

MSR:Hybrid Field Modeling for CT-MRI Rigid-Deformable Registration of the Cervical Spine with an Annotated Dataset

DGX agent

arXiv:2604.27654v1 Announce Type: new Abstract: Accurate CT-MRI registration of the cervical spine is essential for preoperative planning because this region is anatomically complex,highly variable,an

safetyarxiv-cs-cv
1 May 2026
Safety

One of the things I hate the most about this site is the consistent lack of nuance. That’s why everything is an argument, and progress here …

DGX agent

Gary Marcus critiques social media platforms for lacking nuance in discourse, which he identifies as a root cause of persistent arguments and stalled progress on the site. The post reflects concerns a

safetygary-marcus--x
1 May 2026
Safety

Online semi-supervised perception: Real-time learning without explicit feedback

DGX agent

arXiv:2604.27562v1 Announce Type: new Abstract: This paper proposes an algorithm for real-time learning without explicit feedback. The algorithm combines the ideas of semi-supervised learning on graph

safetyarxiv-cs-lg
1 May 2026
← Previous
1…237238239240241…300
Next →