AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
All
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,600 results
Safety

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

DGX agent

arXiv:2604.13016v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a core technique in the post-training of large language models, yet its training dynamics remain poorly unders

safetyarxiv-cs-ai
15 Apr 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG

DGX agent

arXiv:2511.09803v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) improves factuality but retrieving for every query often hurts quality while inflating tokens and latency. We p

safetyarxiv-cs-cl
15 Apr 2026
Safety

Risk-Calibrated Learning: Minimizing Fatal Errors in Medical AI

DGX agent

arXiv:2604.12693v1 Announce Type: new Abstract: Deep learning models often achieve expert-level accuracy in medical image classification but suffer from a critical flaw: semantic incoherence. These hi

safetyarxiv-cs-cv
15 Apr 2026
Safety

Robust Optimization for Mitigating Reward Hacking with Correlated Proxies

DGX agent

arXiv:2604.12086v1 Announce Type: new Abstract: Designing robust reinforcement learning (RL) agents in the presence of imperfect reward signals remains a core challenge. In practice, agents are often

safetyarxiv-cs-lg
15 Apr 2026
Safety

Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design

DGX agent

arXiv:2604.12500v1 Announce Type: new Abstract: Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the condi

safetyarxiv-cs-lg
15 Apr 2026
Safety

SAM3-I: Segment Anything with Instructions

DGX agent

arXiv:2512.04585v3 Announce Type: replace Abstract: Segment Anything Model 3 (SAM3) advances open-vocabulary segmentation through promptable concept segmentation, enabling users to segment all instanc

safetyarxiv-cs-cv
15 Apr 2026
Safety

Scaffold-Conditioned Preference Triplets for Controllable Molecular Optimization with Large Language Models

DGX agent

arXiv:2604.12350v1 Announce Type: cross Abstract: Molecular property optimization is central to drug discovery, yet many deep learning methods rely on black-box scoring and offer limited control over

safetyarxiv-cs-ai
15 Apr 2026
Safety

Scalable and General Whole-Body Control for Cross-Humanoid Locomotion

DGX agent

arXiv:2602.05791v2 Announce Type: replace Abstract: Learning-based whole-body controllers have become a key driver for humanoid robots, yet most existing approaches require robot-specific training. In

safetyarxiv-cs-ro
15 Apr 2026
Safety

Scalable Verification of Neural Control Barrier Functions Using Linear Bound Propagation

DGX agent

arXiv:2511.06341v2 Announce Type: replace Abstract: Control barrier functions (CBFs) are a popular tool for safety certification of nonlinear dynamical control systems. Recently, CBFs represented as n

safetyarxiv-cs-lg
15 Apr 2026
Safety

Schema-Adaptive Tabular Representation Learning with LLMs for Generalizable Multimodal Clinical Reasoning

DGX agent

arXiv:2604.11835v1 Announce Type: cross Abstract: Machine learning for tabular data remains constrained by poor schema generalization, a challenge rooted in the lack of semantic understanding of struc

safetyarxiv-cs-ai
15 Apr 2026
Safety

Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

DGX agent

arXiv:2604.12002v1 Announce Type: new Abstract: Current post-training methods in verifiable settings fall into two categories. Reinforcement learning (RLVR) relies on binary rewards, which are broadly

safetyarxiv-cs-cl
15 Apr 2026
Safety

Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery

DGX agent

arXiv:2507.08977v4 Announce Type: replace-cross Abstract: Scientific modeling faces a tradeoff between the interpretability of mechanistic theory and the predictive power of machine learning. While ex

safetyarxiv-cs-ai
15 Apr 2026
Safety

Skill-informed Data-driven Haptic Nudges for High-dimensional Human Motor Learning

DGX agent

arXiv:2603.12583v2 Announce Type: replace Abstract: In this work, we propose a data-driven framework to design optimal haptic nudge feedback leveraging the learner's estimated skill to address the cha

safetyarxiv-cs-ro
15 Apr 2026
Safety

So true. “Being right too soon is socially unacceptable”

DGX agent

So true. “Being right too soon is socially unacceptable” I remember reading this from Heinlein as a young man and it took experience for me to fully understand it. People and institutions will persist

safetygary-marcus--x
15 Apr 2026
Safety

SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models

DGX agent

arXiv:2604.12617v1 Announce Type: cross Abstract: The post-training pipeline for diffusion models currently has two stages: supervised fine-tuning (SFT) on curated data and reinforcement learning (RL)

safetyarxiv-cs-ai
15 Apr 2026
Safety

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback

DGX agent

arXiv:2510.20093v2 Announce Type: replace-cross Abstract: Although recent advancements in diffusion models have significantly enriched the quality of generated images, challenges remain in synthesizin

safetyarxiv-cs-ai
15 Apr 2026
Safety

Task Alignment: A simple and effective proxy for model merging in computer vision

DGX agent

arXiv:2604.12935v1 Announce Type: new Abstract: Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Des

safetyarxiv-cs-cv
15 Apr 2026
Safety

Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement Learning

DGX agent

arXiv:2604.12770v1 Announce Type: new Abstract: Editing human-written text has become a standard use case of large language models (LLMs), for example, to make one's arguments more appropriate for a d

safetyarxiv-cs-cl
15 Apr 2026
Safety

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs

DGX agent

arXiv:2604.12232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed across diverse domains, yet their vulnerability to jailbreak attacks, where adversarial inputs

safetyarxiv-cs-ai
15 Apr 2026
Safety

Ternary Logic Encodings of Temporal Behavior Trees with Application to Control Synthesis

DGX agent

arXiv:2604.12092v1 Announce Type: new Abstract: Behavior Trees (BTs) provide designers an intuitive graphical interface to construct long-horizon plans for autonomous systems. To ensure their correctn

safetyarxiv-cs-ro
15 Apr 2026
Safety

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment

DGX agent

arXiv:2604.12116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents capable of executing system-level operations. While existing benchmarks

safetyarxiv-cs-ai
15 Apr 2026
Safety

The front door to the internet hasn't changed -- but the person going through that door has changed from a human browsing 5 blue links, to a…

DGX agent

The front door to the internet hasn't changed -- but the person going through that door has changed from a human browsing 5 blue links, to an agent browsing a similar index in vastly different ways. A

safetysonya-huang--x
15 Apr 2026
Safety

The role of System 1 and System 2 semantic memory structure in human and LLM biases

DGX agent

arXiv:2604.12816v1 Announce Type: new Abstract: Implicit biases in both humans and large language models (LLMs) pose significant societal risks. Dual process theories propose that biases arise primari

safetyarxiv-cs-cl
15 Apr 2026
Safety

The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games

DGX agent

arXiv:2510.09087v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown remarkable progress in social deduction games (SDGs). However, existing approaches primarily focus on i

safetyarxiv-cs-ai
15 Apr 2026
Safety

Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training

DGX agent

arXiv:2509.25758v2 Announce Type: replace Abstract: The remarkable capabilities of modern large reasoning models are largely unlocked through post-training techniques such as supervised fine-tuning (S

safetyarxiv-cs-ai
15 Apr 2026
Safety

This is what a black politician in South Africa said… “We will kill white women, we will kill white children, and we will even kill your pet…

DGX agent

This is what a black politician in South Africa said… “We will kill white women, we will kill white children, and we will even kill your pets' Violent language entering politics is a serious warning s

safetyelon-musk--x
15 Apr 2026
Safety

Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood

DGX agent

arXiv:2604.12736v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has significantly advanced the reasoning ability of large language models (LLMs), particularly in their mathem

safetyarxiv-cs-cl
15 Apr 2026
Safety

Towards Generalized Certified Robustness with Multi-Norm Training

DGX agent

arXiv:2410.03000v3 Announce Type: replace Abstract: Existing certified training methods can only train models to be robust against a certain perturbation type (e.g. l_infty or l_2). However, an l_inft

safetyarxiv-cs-lg
15 Apr 2026
Safety

Towards Platonic Representation for Table Reasoning: A Foundation for Permutation-Invariant Retrieval

DGX agent

arXiv:2604.12133v1 Announce Type: new Abstract: Historical approaches to Table Representation Learning (TRL) have largely adopted the sequential paradigms of Natural Language Processing (NLP). We argu

safetyarxiv-cs-ai
15 Apr 2026
Safety

Uncertainty-Aware Image Classification In Biomedical Imaging Using Spectral-normalized Neural Gaussian Processes

DGX agent

arXiv:2602.02370v2 Announce Type: replace Abstract: Accurate histopathologic interpretation is key for clinical decision-making; however, current deep learning models for digital pathology are often o

safetyarxiv-cs-cv
15 Apr 2026
Safety

Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory

DGX agent

arXiv:2604.12817v1 Announce Type: new Abstract: Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To im

safetyarxiv-cs-lg
15 Apr 2026
Safety

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces

DGX agent

arXiv:2603.05295v3 Announce Type: replace Abstract: We introduce WebChain, the largest open-source dataset of human-annotated trajectories on real-world websites, designed to accelerate reproducible r

safetyarxiv-cs-ai
15 Apr 2026
Safety

What happens when you systematically oversell the value of your product for years, while pretty much screwing society along the way? Eventua…

DGX agent

What happens when you systematically oversell the value of your product for years, while pretty much screwing society along the way? Eventually your customers figure it out. Surprisingly, people see A

safetygary-marcus--x
15 Apr 2026
Safety

Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers

DGX agent

arXiv:2604.12509v1 Announce Type: cross Abstract: Mobile Manipulation (MoMa) of articulated objects, such as opening doors, drawers, and cupboards, demands simultaneous, whole-body coordination betwee

safetyarxiv-cs-cv
15 Apr 2026
Safety

WiseOWL: A Methodology for Evaluating Ontological Descriptiveness and Semantic Correctness for Ontology Reuse and Ontology Recommendations

DGX agent

arXiv:2604.12025v1 Announce Type: new Abstract: The Semantic Web standardizes concept meaning for humans and machines, enabling machine-operable content and consistent interpretation that improves adv

safetyarxiv-cs-ai
15 Apr 2026
Safety

XRZero-G0: Pushing the Frontier of Dexterous Robotic Manipulation with Interfaces, Quality and Ratios

DGX agent

arXiv:2604.13001v1 Announce Type: new Abstract: The acquisition of high-quality, action-aligned demonstration data remains a fundamental bottleneck in scaling foundation models for dexterous robot man

safetyarxiv-cs-ro
15 Apr 2026
Safety

3D Multi-View Stylization with Pose-Free Correspondences Matching for Robust 3D Geometry Preservation

DGX agent

arXiv:2604.09639v1 Announce Type: new Abstract: Artistic style transfer is well studied for images and videos, but extending it to multi-view 3D scenes remains difficult because stylization can disrup

safetyarxiv-cs-cv
14 Apr 2026
Safety

A Comparative Theoretical Analysis of Entropy Control Methods in Reinforcement Learning

DGX agent

arXiv:2604.09676v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for enhancing reasoning in large language models (LLMs), yet scalable training is often hindered

safetyarxiv-cs-ai
14 Apr 2026
Safety

A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets

DGX agent

arXiv:2604.10252v1 Announce Type: new Abstract: Reinforcement learning agent-based simulation (RL-ABS) has become an important tool for electricity market mechanism analysis and evaluation. In the mod

safetyarxiv-cs-ai
14 Apr 2026
Safety

A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning

DGX agent

arXiv:2601.16399v5 Announce Type: replace Abstract: We study a structured bi-level optimization problem where the upper-level objective is a smooth function and the lower-level problem is policy optim

safetyarxiv-cs-lg
14 Apr 2026
Safety

A Mamba-Based Multimodal Network for Multiscale Blast-Induced Rapid Structural Damage Assessment

DGX agent

arXiv:2604.11709v1 Announce Type: new Abstract: Accurate and rapid structural damage assessment (SDA) is crucial for post-disaster management, helping responders prioritise resources, plan rescues, an

safetyarxiv-cs-ai
14 Apr 2026
Safety

A mathematical theory of evolution for self-designing AIs

DGX agent

arXiv:2604.05142v2 Announce Type: replace Abstract: As artificial intelligence systems (AIs) become increasingly produced by recursive self-improvement, a form of evolution may emerge, with the traits

safetyarxiv-cs-ai
14 Apr 2026
Safety

A Multilingual Dataset and Empirical Validation for the Mutual Reinforcement Effect in Information Extraction

DGX agent

arXiv:2407.10953v5 Announce Type: replace Abstract: The Mutual Reinforcement Effect (MRE) describes a phenomenon in information extraction where word-level and sentence-level tasks can mutually improv

safetyarxiv-cs-cl
14 Apr 2026
Safety

A profile of BusPatrol, whose AI-powered cameras on 35K+ school buses in 24 US states record vehicles passing illegally, claiming they help reduce violations (Byard Duncan/Bloomberg)

DGX agent

Byard Duncan / Bloomberg: A profile of BusPatrol, whose AI-powered cameras on 35K+ school buses in 24 US states record vehicles passing illegally, claiming they help reduce violations — BusPatrol says

safetytechmeme
14 Apr 2026
Safety

A Proposed Biomedical Data Policy Framework to Reduce Fragmentation, Improve Quality, and Incentivize Sharing in Indian Healthcare in the era of Artificial Intelligence and Digital Health

DGX agent

arXiv:2604.11125v1 Announce Type: new Abstract: India generates vast biomedical data through postgraduate research, government hospital services and audits, government schemes, private hospitals and t

safetyarxiv-cs-ai
14 Apr 2026
Safety

A Queueing-Theoretic Framework for Dynamic Attack Surfaces: Data-Integrated Risk Analysis and Adaptive Defense

DGX agent

arXiv:2604.10427v1 Announce Type: cross Abstract: We develop a queueing-theoretic framework to model the temporal evolution of cyber-attack surfaces, where the number of active vulnerabilities is repr

safetyarxiv-cs-ai
14 Apr 2026
Safety

Active Diffusion Matching: Score-based Iterative Alignment of Cross-Modal Retinal Images

DGX agent

arXiv:2604.10084v1 Announce Type: new Abstract: Objective: The study aims to address the challenge of aligning Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which is difficu

safetyarxiv-cs-cv
14 Apr 2026
Safety

Adaptive Bidding Policies for First-Price Auctions with Budget Constraints under Non-stationarity

DGX agent

arXiv:2505.02796v2 Announce Type: replace-cross Abstract: We study how a budget-constrained bidder should learn to adaptively bid in repeated first-price auctions to maximize her cumulative payoff. Th

safetyarxiv-cs-lg
14 Apr 2026
← Previous
1…246247248249250…263
Next →