AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,704 results
16 Jul 2026

Pretraining in Actor-Critic Reinforcement Learning for Locomotion

SafetyDGX agent

arXiv:2510.12363v4 Announce Type: replace-cross Abstract: The pretraining-finetuning paradigm has facilitated numerous transformative advancements in artificial intelligence research in recent years.

Price of Fairness in Bandits: A Tight Minimax Characterization

SafetyDGX agent

arXiv:2607.13402v1 Announce Type: cross Abstract: In bandit problems, standard regret-minimizing algorithms treat exploration as an amortized cost, which can expose early participants to unfair ex-ant

Privacy Preserving Recommender Systems Balancing Personalization with Privacy

SafetyDGX agent

arXiv:2607.13328v1 Announce Type: cross Abstract: Personalized recommendation systems are central to modern e-commerce and retail platforms, but they typically rely on centralized storage of detailed


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities

SafetyDGX agent

arXiv:2607.13596v1 Announce Type: cross Abstract: When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledgi

PUe: Biased Positive-Unlabeled Learning Enhancement by Causal Inference

SafetyDGX agent

arXiv:2607.13428v1 Announce Type: new Abstract: Positive-Unlabeled (PU) learning aims to achieve high-accuracy binary classification with limited labeled positive examples and numerous unlabeled ones.

RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset

SafetyDGX agent

arXiv:2603.11811v2 Announce Type: replace-cross Abstract: The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by the p

Reverse to Advance: Teleoperation-Cost Effective Hard Policy Learning from Reversed Easy Tasks

SafetyDGX agent

arXiv:2607.13455v1 Announce Type: new Abstract: High-quality teleoperation datasets are costly to collect, particularly for hard tasks. We observe that many tasks exhibit directional asymmetry: comple

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing

SafetyDGX agent

arXiv:2607.13594v1 Announce Type: new Abstract: LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a gua

SARFA: Segment Anything with Radiomic Feature Alignment

SafetyDGX agent

arXiv:2607.13323v1 Announce Type: new Abstract: The Segment Anything Model (SAM) has demonstrated strong generalizability across a variety of segmentation tasks. However, SAM often struggles in situat

SCOPE-RL: Optimizing Reasoning Paths Before and After Success

SafetyDGX agent

arXiv:2607.11506v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) optimizes LLMs using sparse verifiable final-answer rewards. This sparse anchor reliably verif

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

SafetyDGX agent

arXiv:2607.13124v1 Announce Type: cross Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compre

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

SafetyDGX agent

arXiv:2607.13931v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-languag

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

SafetyDGX agent

arXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrop

STITCHER: Constrained Trajectory Planning in Complex Environments with Real-Time Motion Primitive Search

SafetyDGX agent

arXiv:2510.14893v4 Announce Type: replace Abstract: Autonomous high-speed navigation through large, complex environments requires real-time generation of agile trajectories that are dynamically feasib

Structured Reinforcement Learning for Bayesian Persuasion : Application to Intelligent Interactive Driving

SafetyDGX agent

arXiv:2607.13576v1 Announce Type: new Abstract: Interactive driving, wherein an intelligent lead vehicle equipped with real-time traffic data coordinates route choices of connected vehicles, offers a

Temperature Scaling Is Not Enough: Calibration Gaps Under Human Label Distributions

SafetyDGX agent

arXiv:2607.13423v1 Announce Type: new Abstract: Temperature scaling is the dominant post-hoc calibration method in modern deep learning. Its theoretical justification rests on an assumption that is ra

The Cafe in Amsterdam: When the Incumbent Becomes the Oracle

SafetyDGX agent

arXiv:2607.13393v1 Announce Type: cross Abstract: A field can reformulate its computations freely exactly where its demand is stated independently of any incumbent implementation, and finds itself una

The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models

SafetyDGX agent

arXiv:2607.13612v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) are the dominant design for latent world models, yet they are usually justified by empirical performa

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

SafetyDGX agent

arXiv:2607.13539v1 Announce Type: new Abstract: While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based genera

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation

SafetyDGX agent

arXiv:2512.10607v2 Announce Type: replace Abstract: We present TCAM (Track and Caption Any Motion), a generative framework that watches a video and with no text query and no region prompt decides what

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

SafetyDGX agent

arXiv:2607.13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at

Vision-Based Obstacle Separation for Strawberry Harvesting in Clusters Using Hierarchical Reinforcement Learning

SafetyDGX agent

arXiv:2607.13799v1 Announce Type: new Abstract: Selective harvesting in clustered strawberry environments is challenging because ripe fruits are often occluded by surrounding unripe fruits, making dir

Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback

SafetyDGX agent

arXiv:2607.13389v1 Announce Type: new Abstract: Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pi

15 Jul 2026

A Neurosymbolic Approach to Natural Language Formalization and Verification

SafetyDGX agent

arXiv:2511.09008v2 Announce Type: replace-cross Abstract: Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits th

AAAI-26 Dual Submissions: Novel Challenges

SafetyDGX agent

arXiv:2607.11918v1 Announce Type: cross Abstract: Dual submissions, in which identical or substantially similar papers are simultaneously submitted to one or more archival venues, without cross-citati

Accuracy and Normalized Accuracy under Length Bias: Analysis, Guidelines, and a Bayesian Alternative

SafetyDGX agent

arXiv:2607.12767v1 Announce Type: new Abstract: Multiple-choice benchmarks that rank candidate completions by conditional log-probability suffer from a length bias: because log-probabilities sum over

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to …

SafetyDGX agent

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models c

Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity

SafetyDGX agent

arXiv:2607.11957v1 Announce Type: new Abstract: Maintenance of critical infrastructures, such as railways and power plants, is essential for ensuring operational safety and reliability. However, the d

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment

SafetyDGX agent

arXiv:2510.13698v4 Announce Type: replace Abstract: Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety

Auditable Context-Aware HFMD Forecasting with Structured LLM Agents

SafetyDGX agent

arXiv:2511.23276v2 Announce Type: replace Abstract: Effective HFMD surveillance requires forecasts capturing both time-series patterns and contextual drivers such as school calendars, weather, and pol

Beyond Coordinate Gauge: An Audited Protocol for Detecting Donor-Specific Functional Fingerprints after Neural Collapse

SafetyDGX agent

arXiv:2607.11967v1 Announce Type: cross Abstract: Independently trained neural networks have no shared neuron-index reference frame, so comparing them requires accounting for coordinate freedom. Neura

Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings

SafetyDGX agent

arXiv:2607.12071v1 Announce Type: new Abstract: Continuous semantic reconstruction from non-invasive neural recordings remains limited by the representational mismatch between semantic feature spaces

Calculating Mutual Information between a Reward Maximizer and its Environment

SafetyDGX agent

arXiv:2602.12963v2 Announce Type: replace Abstract: An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world. In this work,

Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation

SafetyDGX agent

arXiv:2607.12075v1 Announce Type: cross Abstract: Background: Deep learning models can classify thyroid nodules on ultrasound, but reliable clinical decision support also requires calibrated probabili

Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses

SafetyDGX agent

arXiv:2607.11959v1 Announce Type: new Abstract: Greenhouse reinforcement learning can test climate-control ideas at a speed and scale that is difficult to achieve with crop experiments alone. For smar

Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?

SafetyDGX agent

arXiv:2607.12631v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents in high-stakes domains, understanding contextual factors that may modul

Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

SafetyDGX agent

arXiv:2607.12835v1 Announce Type: new Abstract: Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, whe

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation

SafetyDGX agent

arXiv:2601.08010v3 Announce Type: replace Abstract: Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasonin

CGRL: Concept-Guided Pruning and Representation Learning for Whole-Slide Image Classification

SafetyDGX agent

arXiv:2607.12556v1 Announce Type: new Abstract: Weakly supervised whole-slide image (WSI) classification is widely used in computational pathology because slide-level labels are easier to obtain than

ChunkFlow: Towards Continuity-Consistent Chunked Policy Learning

SafetyDGX agent

arXiv:2607.12992v1 Announce Type: new Abstract: Vision-language action (VLA) models increasingly adopt chunked action heads to satisfy real-time constraints; however, this introduces boundary jitter:

CityBehavEx: A Scalable and Empirically Validated LLM-Assisted Urban Simulation Platform

SafetyDGX agent

arXiv:2607.12086v1 Announce Type: new Abstract: Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validat

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

SafetyDGX agent

arXiv:2607.12273v1 Announce Type: cross Abstract: As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, w

Compos3D: Interactive Part-Based Composition for Creative Control in Generative 3D Models

SafetyDGX agent

arXiv:2607.12193v1 Announce Type: cross Abstract: While generative AI has unlocked new opportunities for 3D content creation, current workflows often rely on multiple regenerations, which provides lim

Continuous Policy and Value Iteration for Stochastic Control Problems and Its Convergence

SafetyDGX agent

arXiv:2506.08121v2 Announce Type: replace-cross Abstract: We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and

Data Safety: Synthetic Data Quality Analysis Using CIFAKE Dataset

SafetyDGX agent

arXiv:2607.12165v1 Announce Type: new Abstract: Recently, the societal implementation of high-performance image classification models has expanded rapidly. While these models require vast amounts of t

DenseReward: Dense Reward Learning via Failure Synthesis for Robotic Manipulation

SafetyDGX agent

arXiv:2607.13033v1 Announce Type: new Abstract: Reinforcement learning holds great promise for improving robot policies beyond the limits of imitation learning. However, its practical adoption remains

Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences

SafetyDGX agent

arXiv:2607.12466v1 Announce Type: new Abstract: Aligning robot policies with human preferences is essential for deployment to diverse end users. In per-user alignment approach, preference feedback is

Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

SafetyDGX agent

arXiv:2607.12784v1 Announce Type: cross Abstract: Reinforcement Learning has revolutionized the landscape of robotic research, allowing robust learning of complex robotic skills in simulation. However

Expert Knowledge-driven Reinforcement Learning for Autonomous Racing via Trajectory Guidance and Dynamics Constraints

SafetyDGX agent

arXiv:2603.05842v2 Announce Type: replace Abstract: Reinforcement learning has demonstrated significant potential in the field of autonomous driving. However, it suffers from defects such as training

ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning

SafetyDGX agent

arXiv:2607.12931v1 Announce Type: new Abstract: Reinforcement Learning (RL) has demonstrated significant potential for improving Vision-Language-Action (VLA) models on complex manipulation tasks. Howe

ExtraGS: Enhancing Endoscopic View Extrapolation via Diffusion-Guided 3D Gaussian Splatting

SafetyDGX agent

arXiv:2607.12785v1 Announce Type: new Abstract: Robot-assisted minimally invasive surgery (MIS) critically depends on reliable endoscopic perception for navigation and safety. However, conventional en

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

SafetyDGX agent

arXiv:2607.13017v1 Announce Type: cross Abstract: World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveragin

From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness

SafetyDGX agent

arXiv:2607.12166v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are the standard for decomposing superposed neural representations into interpretable features, and evaluation relies predomi

From Sentiment to Actionable Insights: Public Sentiment Analysis of Advanced Air Mobility

SafetyDGX agent

arXiv:2606.20751v2 Announce Type: replace Abstract: Advanced Air Mobility (AAM) is an emerging low-altitude transportation system whose successful deployment depends on both technological progress and

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

SafetyDGX agent

arXiv:2607.12463v1 Announce Type: new Abstract: Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in

G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis

SafetyDGX agent

arXiv:2607.11892v1 Announce Type: cross Abstract: Human-factor event diagnosis is essential for learning from operational events in nuclear power plants, yet its quality depends strongly on expert int

GaitSpan: Growing Humanoid Locomotion from Walking to Running

SafetyDGX agent

arXiv:2607.12114v1 Announce Type: cross Abstract: A humanoid that can walk should not relearn locomotion from scratch to jog or run. Yet current approaches often obtain gait diversity by prescribing g

Git-Assistant: Planning-Based Support for Updating Git Repositories

SafetyDGX agent

arXiv:2607.09224v2 Announce Type: replace-cross Abstract: Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners. Re

Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization

SafetyDGX agent

arXiv:2607.12332v1 Announce Type: new Abstract: We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme

Gradient-free learning of a closed-loop wall controller for turbulent drag reduction

SafetyDGX agent

arXiv:2607.12626v1 Announce Type: cross Abstract: Closed-loop wall control learnt by multi-agent reinforcement learning can lower skin-friction drag in turbulent channels, but these gradient-based pol

← Previous
1…3435363738…212
Next →