AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,600 results
15 Apr 2026

I spent some time trying to distill all the complex factors impacting open models -- economics, capabilities, distribution, policy, etc. -- …

SafetyDGX agent

I spent some time trying to distill all the complex factors impacting open models -- economics, capabilities, distribution, policy, etc. -- into a clear list of beliefs. Here they are in full. 1. It’s

Incentivizing High-Quality Human Annotations with Golden Questions

SafetyDGX agent

arXiv:2505.19134v2 Announce Type: replace-cross Abstract: Human-annotated data plays a vital role in training large language models (LLMs), such as supervised fine-tuning and human preference alignmen

Information-Geometric Decomposition of Generalization Error in Unsupervised Learning

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.12340v1 Announce Type: cross Abstract: We decompose the Kullback--Leibler generalization error (GE) -- the expected KL divergence from the data distribution to the trained model -- of unsup

InsightFlow: LLM-Driven Synthesis of Patient Narratives for Mental Health into Causal Models

SafetyDGX agent

arXiv:2604.12721v1 Announce Type: new Abstract: Clinical case formulation organizes patient symptoms and psychosocial factors into causal models, often using the 5P framework. However, constructing su

Labeled TrustSet Guided: Batch Active Learning with Reinforcement Learning

SafetyDGX agent

arXiv:2604.12303v1 Announce Type: new Abstract: Batch active learning (BAL) is a crucial technique for reducing labeling costs and improving data efficiency in training large-scale deep learning model

LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries

SafetyDGX agent

arXiv:2601.10398v3 Announce Type: replace Abstract: In LLM-based text-to-SQL systems, unanswerable and underspecified user queries may generate not only incorrect text but also executable programs tha

Learning step-level dynamic soaring in shear flow

SafetyDGX agent

arXiv:2604.12413v1 Announce Type: cross Abstract: Dynamic soaring enables sustained flight by extracting energy from wind shear, yet it is commonly understood as a cycle-level maneuver that assumes st

Learning Versatile Humanoid Manipulation with Touch Dreaming

SafetyDGX agent

arXiv:2604.13015v1 Announce Type: new Abstract: Humanoid robots promise general-purpose assistance, yet real-world humanoid loco-manipulation remains challenging because it requires whole-body stabili

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

SafetyDGX agent

arXiv:2604.13010v1 Announce Type: cross Abstract: On-policy distillation (OPD) has emerged as an efficient post-training paradigm for large language models. However, standard OPD requires a live teach

LiveMoments: Reselected Key Photo Restoration in Live Photos via Reference-guided Diffusion

SafetyDGX agent

arXiv:2604.12286v1 Announce Type: new Abstract: Live Photo captures both a high-quality key photo and a short video clip to preserve the precious dynamics around the captured moment. While users may c

LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software

SafetyDGX agent

arXiv:2604.12994v1 Announce Type: cross Abstract: Logical vulnerabilities in software stem from flaws in program logic rather than memory safety, which can lead to critical security failures. Although

Man and machine: artificial intelligence and judicial decision making

SafetyDGX agent

arXiv:2603.19042v4 Announce Type: replace Abstract: The integration of artificial intelligence (AI) technologies into judicial decision-making, particularly in pretrial, sentencing, and parole context

Meet Dynamic Individual Preferences: Resolving Conflicting Human Value with Paired Fine-Tuning

SafetyDGX agent

arXiv:2604.12479v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have significantly improved the alignment of models with general human preferences. However, a major cha

Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation

SafetyDGX agent

arXiv:2604.12277v1 Announce Type: new Abstract: Pretrained language models often rely on superficial features that appear predictive during training yet fail to generalize at test time, a phenomenon k

MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models

SafetyDGX agent

arXiv:2604.12537v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have achieved remarkable progress in multimodal understanding, yet their positional encoding mechanisms remain suboptima

MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization

SafetyDGX agent

arXiv:2604.12237v1 Announce Type: cross Abstract: In drug discovery, molecular optimization aims to iteratively refine a lead compound to improve molecular properties while preserving structural simil

Mutual Information Surprise: Rethinking Unexpectedness in Autonomous Systems

SafetyDGX agent

arXiv:2508.17403v3 Announce Type: replace Abstract: A community of researchers appears to think that a machine can be surprised and have introduced various surprise measures, principally the Shannon S

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning

SafetyDGX agent

arXiv:2601.06794v2 Announce Type: replace Abstract: Critique-guided reinforcement learning (RL) has emerged as a powerful paradigm for training LLM agents by augmenting sparse outcome rewards with nat

Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning

SafetyDGX agent

arXiv:2604.05164v2 Announce Type: replace-cross Abstract: As LLM reasoning performance plateau, improving inference-time compute efficiency is crucial to mitigate overthinking and long thinking traces

Offline-Online Reinforcement Learning for Linear Mixture MDPs

SafetyDGX agent

arXiv:2604.11994v1 Announce Type: new Abstract: We study offline-online reinforcement learning in linear mixture Markov decision processes (MDPs) under environment shift. In the offline phase, data ar

PAINT: Partner-Agnostic Intent-Aware Cooperative Transport with Legged Robots

SafetyDGX agent

arXiv:2604.12852v1 Announce Type: new Abstract: Collaborative transport requires robots to infer partner intent through physical interaction while maintaining stable loco-manipulation. This becomes pa

Parallax: Why AI Agents That Think Must Never Act

SafetyDGX agent

arXiv:2604.12986v1 Announce Type: cross Abstract: Autonomous AI agents are rapidly transitioning from experimental tools to operational infrastructure, with projections that 80% of enterprise applicat

Perception-Aware Policy Optimization for Multimodal Reasoning

SafetyDGX agent

arXiv:2507.06448v5 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has proven to be a highly effective strategy for endowing Large Language Models (LLMs) with ro

Physics-Grounded Monocular Vehicle Distance Estimation Using Standardized License Plate Typography

SafetyDGX agent

arXiv:2604.12239v1 Announce Type: new Abstract: Accurate inter-vehicle distance estimation is a cornerstone of Advanced Driver Assistance Systems (ADAS) and autonomous driving. While LiDAR and radar p

🇧🇪 Positive news for FSD Supervised in Belgium! I just received an official response from the cabinet of @MDiependaele , Minister-Presiden…

SafetyDGX agent

🇧🇪 Positive news for FSD Supervised in Belgium! I just received an official response from the cabinet of @MDiependaele , Minister-President of the Flemish Government. A few days ago, the Dutch vehicle

PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation

SafetyDGX agent

arXiv:2604.12113v1 Announce Type: cross Abstract: Visual Foundation Models (VFMs) such as the Segment Anything Model (SAM) have significantly advanced broad use of image segmentation. However, SAM and

Prediction from a year ago about AI backlash that looks to be on track:

SafetyDGX agent

Prediction from a year ago about AI backlash that looks to be on track: Public backlash against AI will so be strong by 2028 that anti-AI sentiment will probably be a major factor in the 2028 US Presi

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints

SafetyDGX agent

arXiv:2604.12384v1 Announce Type: new Abstract: Safety alignment in Large Language Models (LLMs) remains highly fragile during fine-tuning, where even benign adaptation can degrade pre-trained refusal

Probabilistic Feature Imputation and Uncertainty-Aware Multimodal Federated Aggregation

SafetyDGX agent

arXiv:2604.12970v1 Announce Type: cross Abstract: Multimodal federated learning enables privacy-preserving collaborative model training across healthcare institutions. However, a fundamental challenge

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation

SafetyDGX agent

arXiv:2511.17097v2 Announce Type: replace Abstract: Vision-Language Navigation requires agents to act coherently over long horizons by understanding not only local visual context but also how far they

PubSwap: Public-Data Off-Policy Coordination for Federated RLVR

SafetyDGX agent

arXiv:2604.12160v1 Announce Type: new Abstract: Reasoning post-training with reinforcement learning from verifiable rewards (RLVR) is typically studied in centralized settings, yet many realistic appl

RACF: A Resilient Autonomous Car Framework with Object Distance Correction

SafetyDGX agent

arXiv:2604.12418v1 Announce Type: cross Abstract: Autonomous vehicles are increasingly deployed in safety-critical applications, where sensing failures or cyberphysical attacks can lead to unsafe oper

Reasoning about Intent for Ambiguous Requests

SafetyDGX agent

arXiv:2511.10453v3 Announce Type: replace-cross Abstract: Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating safet

Redefining Quality Criteria and Distance-Aware Score Modeling for Image Editing Assessment

SafetyDGX agent

arXiv:2604.12175v1 Announce Type: new Abstract: Recent advances in image editing have heightened the need for reliable Image Editing Quality Assessment (IEQA). Unlike traditional methods, IEQA require

Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models

SafetyDGX agent

arXiv:2604.12582v1 Announce Type: new Abstract: Recent Video Large Language Models (Video-LLMs) have demonstrated strong capability in video understanding, yet they still suffer from hallucinations. E

Reliability-Guided Depth Fusion for Glare-Resilient Navigation Costmaps

SafetyDGX agent

arXiv:2604.12753v1 Announce Type: new Abstract: Specular glare on reflective floors and glass surfaces frequently corrupts RGB-D depth measurements, producing holes and spikes that accumulate as persi

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

SafetyDGX agent

arXiv:2604.13016v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a core technique in the post-training of large language models, yet its training dynamics remain poorly unders

Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG

SafetyDGX agent

arXiv:2511.09803v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) improves factuality but retrieving for every query often hurts quality while inflating tokens and latency. We p

Risk-Calibrated Learning: Minimizing Fatal Errors in Medical AI

SafetyDGX agent

arXiv:2604.12693v1 Announce Type: new Abstract: Deep learning models often achieve expert-level accuracy in medical image classification but suffer from a critical flaw: semantic incoherence. These hi

Robust Optimization for Mitigating Reward Hacking with Correlated Proxies

SafetyDGX agent

arXiv:2604.12086v1 Announce Type: new Abstract: Designing robust reinforcement learning (RL) agents in the presence of imperfect reward signals remains a core challenge. In practice, agents are often

Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design

SafetyDGX agent

arXiv:2604.12500v1 Announce Type: new Abstract: Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the condi

SAM3-I: Segment Anything with Instructions

SafetyDGX agent

arXiv:2512.04585v3 Announce Type: replace Abstract: Segment Anything Model 3 (SAM3) advances open-vocabulary segmentation through promptable concept segmentation, enabling users to segment all instanc

Scaffold-Conditioned Preference Triplets for Controllable Molecular Optimization with Large Language Models

SafetyDGX agent

arXiv:2604.12350v1 Announce Type: cross Abstract: Molecular property optimization is central to drug discovery, yet many deep learning methods rely on black-box scoring and offer limited control over

Scalable and General Whole-Body Control for Cross-Humanoid Locomotion

SafetyDGX agent

arXiv:2602.05791v2 Announce Type: replace Abstract: Learning-based whole-body controllers have become a key driver for humanoid robots, yet most existing approaches require robot-specific training. In

Scalable Verification of Neural Control Barrier Functions Using Linear Bound Propagation

SafetyDGX agent

arXiv:2511.06341v2 Announce Type: replace Abstract: Control barrier functions (CBFs) are a popular tool for safety certification of nonlinear dynamical control systems. Recently, CBFs represented as n

Schema-Adaptive Tabular Representation Learning with LLMs for Generalizable Multimodal Clinical Reasoning

SafetyDGX agent

arXiv:2604.11835v1 Announce Type: cross Abstract: Machine learning for tabular data remains constrained by poor schema generalization, a challenge rooted in the lack of semantic understanding of struc

Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

SafetyDGX agent

arXiv:2604.12002v1 Announce Type: new Abstract: Current post-training methods in verifiable settings fall into two categories. Reinforcement learning (RLVR) relies on binary rewards, which are broadly

Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery

SafetyDGX agent

arXiv:2507.08977v4 Announce Type: replace-cross Abstract: Scientific modeling faces a tradeoff between the interpretability of mechanistic theory and the predictive power of machine learning. While ex

Skill-informed Data-driven Haptic Nudges for High-dimensional Human Motor Learning

SafetyDGX agent

arXiv:2603.12583v2 Announce Type: replace Abstract: In this work, we propose a data-driven framework to design optimal haptic nudge feedback leveraging the learner's estimated skill to address the cha

So true. “Being right too soon is socially unacceptable”

SafetyDGX agent

So true. “Being right too soon is socially unacceptable” I remember reading this from Heinlein as a young man and it took experience for me to fully understand it. People and institutions will persist

SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models

SafetyDGX agent

arXiv:2604.12617v1 Announce Type: cross Abstract: The post-training pipeline for diffusion models currently has two stages: supervised fine-tuning (SFT) on curated data and reinforcement learning (RL)

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback

SafetyDGX agent

arXiv:2510.20093v2 Announce Type: replace-cross Abstract: Although recent advancements in diffusion models have significantly enriched the quality of generated images, challenges remain in synthesizin

Task Alignment: A simple and effective proxy for model merging in computer vision

SafetyDGX agent

arXiv:2604.12935v1 Announce Type: new Abstract: Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Des

Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement Learning

SafetyDGX agent

arXiv:2604.12770v1 Announce Type: new Abstract: Editing human-written text has become a standard use case of large language models (LLMs), for example, to make one's arguments more appropriate for a d

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs

SafetyDGX agent

arXiv:2604.12232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed across diverse domains, yet their vulnerability to jailbreak attacks, where adversarial inputs

Ternary Logic Encodings of Temporal Behavior Trees with Application to Control Synthesis

SafetyDGX agent

arXiv:2604.12092v1 Announce Type: new Abstract: Behavior Trees (BTs) provide designers an intuitive graphical interface to construct long-horizon plans for autonomous systems. To ensure their correctn

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment

SafetyDGX agent

arXiv:2604.12116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents capable of executing system-level operations. While existing benchmarks

The front door to the internet hasn't changed -- but the person going through that door has changed from a human browsing 5 blue links, to a…

SafetyDGX agent

The front door to the internet hasn't changed -- but the person going through that door has changed from a human browsing 5 blue links, to an agent browsing a similar index in vastly different ways. A

The role of System 1 and System 2 semantic memory structure in human and LLM biases

SafetyDGX agent

arXiv:2604.12816v1 Announce Type: new Abstract: Implicit biases in both humans and large language models (LLMs) pose significant societal risks. Dual process theories propose that biases arise primari

The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games

SafetyDGX agent

arXiv:2510.09087v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown remarkable progress in social deduction games (SDGs). However, existing approaches primarily focus on i

← Previous
1…196197198199200…210
Next →