AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,708 results
12 May 2026

Positive Alignment: Artificial Intelligence for Human Flourishing

SafetyDGX agent

arXiv:2605.10310v1 Announce Type: new Abstract: Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of ali

Primal-Dual Guided Decoding for Constrained Discrete Diffusion

SafetyDGX agent

arXiv:2605.09749v1 Announce Type: new Abstract: Discrete diffusion models generate structured sequences by progressively unmasking tokens, but enforcing global property constraints during generation r

Princeton faculty votes to require proctoring in all in-person exams starting this summer, reversing an 1893 policy amid concerns about AI-fueled cheating (Douglas Belkin/Wall Street Journal)

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Douglas Belkin / Wall Street Journal: Princeton faculty votes to require proctoring in all in-person exams starting this summer, reversing an 1893 policy amid concerns about AI-fueled cheating — The c

Privacy-Aware Video Anomaly Detection through Orthogonal Subspace Projection

SafetyDGX agent

arXiv:2605.08651v1 Announce Type: cross Abstract: Video anomaly detection (VAD) systems often prioritize accuracy while overlooking privacy concerns, limiting their suitability for real-world deployme

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation

SafetyDGX agent

arXiv:2605.08774v1 Announce Type: cross Abstract: Long-horizon robotic manipulation requires dense feedback that reflects how a task advances through its procedural stages, not merely whether the fina

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

SafetyDGX agent

arXiv:2501.03544v5 Announce Type: replace-cross Abstract: Recent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, the

ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design

SafetyDGX agent

arXiv:2605.10189v1 Announce Type: cross Abstract: Designing proteins with desired functions or properties represents a core goal in synthetic biology and drug discovery. Recent advances in protein lan

Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions

SafetyDGX agent

arXiv:2605.09893v1 Announce Type: cross Abstract: Large language models (LLMs) are often evaluated based on their stated values, yet these do not reliably translate into their actions, a discrepancy t

Q-learning with Adjoint Matching

SafetyDGX agent

arXiv:2601.14234v3 Announce Type: replace-cross Abstract: We propose Q-learning with Adjoint Matching (QAM), a novel TD-based reinforcement learning (RL) algorithm that tackles a long-standing challen

Quantile-Coupled Flow Matching for Distributional Reinforcement Learning

SafetyDGX agent

arXiv:2605.08515v1 Announce Type: new Abstract: Unlike standard expected-return Reinforcement Learning (RL), Distributional RL (DRL) models the full return distribution, making it better-suited for un

Re-Triggering Safeguards within LLMs for Jailbreak Detection

SafetyDGX agent

arXiv:2605.10611v1 Announce Type: cross Abstract: This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs

Reasoning Compression with Mixed-Policy Distillation

SafetyDGX agent

arXiv:2605.08776v1 Announce Type: new Abstract: Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive t

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

SafetyDGX agent

arXiv:2605.10805v1 Announce Type: new Abstract: Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings r

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

SafetyDGX agent

arXiv:2605.09614v1 Announce Type: new Abstract: Long chain-of-thought (CoT) reasoning improves large vision--language models, but visual information often fades during generation, limiting long-horizo

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias

SafetyDGX agent

arXiv:2605.08315v1 Announce Type: new Abstract: Existing LLM-based policy optimizers see only scalar rewards: that a policy scored 0.45, but not whether the agent got stuck in a loop, fell into a hole

Reinforcement learning for inverse structural design and rapid laser cutting of kirigami prototypes

SafetyDGX agent

arXiv:2605.08098v1 Announce Type: new Abstract: Kirigami is an increasingly useful fabrication method to produce shape-programmable metamaterial structures. However, inverse design remains difficult b

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems

SafetyDGX agent

arXiv:2605.08378v1 Announce Type: cross Abstract: Reinforcement learning has become a powerful paradigm for improving the capability of intelligent systems, but its practical deployment faces two cent

Reinforcement Learning with Action Chunking

SafetyDGX agent

arXiv:2507.07969v4 Announce Type: replace-cross Abstract: We present Q-chunking, a simple yet effective recipe for improving reinforcement learning (RL) algorithms for long-horizon, sparse-reward task

Reinforcing Multimodal Reasoning Against Visual Degradation

SafetyDGX agent

arXiv:2605.09262v1 Announce Type: cross Abstract: Reinforcement Learning has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet the resulting policies r

Relational reasoning and inductive bias in transformers and large language models

SafetyDGX agent

arXiv:2506.04289v3 Announce Type: replace Abstract: Transformer-based models have demonstrated remarkable reasoning abilities, but the mechanisms underlying relational reasoning remain poorly understo

Relational Retrieval: Leveraging Known-Novel Interactions for Generalized Category Discovery

SafetyDGX agent

arXiv:2605.09420v1 Announce Type: cross Abstract: In this study, we tackle Generalized Category Discovery (GCD) via a Relational Retrieval perspective, explicitly coupling labeled and unlabeled data t

Relations Are Channels: Knowledge Graph Embedding via Kraus Decompositions

SafetyDGX agent

arXiv:2605.10317v1 Announce Type: cross Abstract: Knowledge graph embedding (KGE) models typically represent each relation as an operator on entity embeddings. In this work, we identify three structur

Relative Score Policy Optimization for Diffusion Language Models

SafetyDGX agent

arXiv:2605.10218v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability require

Remember to Forget: Gated Adaptive Positional Encoding

SafetyDGX agent

arXiv:2605.10414v1 Announce Type: new Abstract: Rotary Positional Encoding (RoPE) is widely used in modern large language models. However, when sequences are extended beyond the range seen during trai

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models

SafetyDGX agent

arXiv:2605.09410v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervisi

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

SafetyDGX agent

arXiv:2605.08257v1 Announce Type: cross Abstract: Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develo

Responsible Benchmarking of Fairness for Automatic Speech Recognition

SafetyDGX agent

arXiv:2605.10615v1 Announce Type: new Abstract: Many studies have shown automatic speech processing (ASR) systems have unequal performance across speakergroups (SG's). However, the manner in which suc

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models

SafetyDGX agent

arXiv:2605.08186v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) via entropy minimization (EM) has proven effective for classification tasks, yet its application to generative autoregressi

Rethinking Loss Reweighting for Imbalance Learning as an Inverse Problem: A Neural Collapse Point of View

SafetyDGX agent

arXiv:2605.10047v1 Announce Type: cross Abstract: Loss reweighting is a widely used strategy for long-tailed classification, but existing reweighting strategies often rely on heuristics and rarely def

Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.09212v1 Announce Type: new Abstract: Centralized training with decentralized execution (CTDE) is a standard framework for cooperative multi-agent policy-gradient reinforcement learning, all

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

SafetyDGX agent

arXiv:2605.06241v2 Announce Type: replace Abstract: Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not

Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with k-step Policy Gradients

SafetyDGX agent

arXiv:2605.10909v1 Announce Type: new Abstract: This work revisits standard policy gradient methods used on restricted policy classes, which are known to get stuck in suboptimal critical points. We id

Revitalizing the Beginning: Avoiding Storage Dependency for Model Merging in Continual Learning

SafetyDGX agent

arXiv:2605.08311v1 Announce Type: cross Abstract: Model merging provides a compelling paradigm for integrating specialized expertise into a unified multi-task model, a goal that aligns naturally with

Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios

SafetyDGX agent

arXiv:2512.00920v4 Announce Type: replace Abstract: Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods fo

Reward-Conditioned Reinforcement Learning

SafetyDGX agent

arXiv:2603.05066v2 Announce Type: replace Abstract: Single-task RL agents are typically trained under a fixed reward function, which limits their robustness to reward misspecification and their abilit

RigidFormer: Learning Rigid Dynamics using Transformers

SafetyDGX agent

arXiv:2605.09196v1 Announce Type: cross Abstract: Learning-based simulation of multi-object rigid-body dynamics remains difficult because contact is discontinuous and errors compound over long horizon

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning

SafetyDGX agent

arXiv:2605.10293v1 Announce Type: cross Abstract: In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide gua

Route by State, Recover from Trace: STAR with Failure-Aware Markov Routing for Multi-Agent Spatiotemporal Reasoning

SafetyDGX agent

arXiv:2605.10057v1 Announce Type: new Abstract: Compositional spatiotemporal reasoning often requires a system to invoke multiple heterogeneous specialists, such as geometric, temporal, topological, a

RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems

SafetyDGX agent

arXiv:2605.10862v1 Announce Type: new Abstract: This paper demonstrates RUBEN, an interactive tool for discovering minimal rules to explain the outputs of retrieval-augmented large language models (LL

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

SafetyDGX agent

arXiv:2605.10899v1 Announce Type: new Abstract: Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyo

RuPLaR : Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors From Multi-Step to One-Step

SafetyDGX agent

arXiv:2605.09346v1 Announce Type: cross Abstract: The Chain-of-Thought (CoT) paradigm, while enhancing the interpretability of Large Language Models (LLMs), is constrained by the inefficiencies and ex

Safe and Real-Time Consistent Planning for Autonomous Vehicles in Partially Observed Environments via Parallel Consensus Optimization

SafetyDGX agent

arXiv:2409.10310v3 Announce Type: replace Abstract: Ensuring safety and driving consistency is a significant challenge for autonomous vehicles operating in partially observed environments. This work i

Safe Exploration for Nonlinear Processes Using Online Gaussian Process Learning

SafetyDGX agent

arXiv:2605.09772v1 Announce Type: cross Abstract: This paper proposes a safe data-driven control framework for nonlinear systems with partially known dynamics. The method ensures stability and constra

Safety-Critical LiDAR-Inertial Odometry with On-Manifold Deterministic Protection Level

SafetyDGX agent

arXiv:2605.09383v1 Announce Type: new Abstract: In safety-critical scenarios, the protection level of the autonomous navigation system is crucial for enabling mobile robots to perform safe tasks. Howe

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

SafetyDGX agent

arXiv:2510.20129v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, where adversarially crafted prompts induce policy-violating responses des

SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

SafetyDGX agent

arXiv:2605.08334v1 Announce Type: new Abstract: We present SalesSim, a framework and testbed for evaluating the ability of Multimodal Large Language Models (MLLMs) to simulate realistic, persona-drive

Sam Altman swearing to tell the whole truth, and then failing to do so. May 2023.

SafetyDGX agent

Sam Altman made statements under oath in May 2023 regarding AI safety and OpenAI's practices, but Gary Marcus critiqued these statements as incomplete or misleading, suggesting Altman failed to fully

Sample-Mean Anchored Thompson Sampling for Offline-to-Online Learning with Distribution Shift

SafetyDGX agent

arXiv:2605.10289v1 Announce Type: new Abstract: Offline-to-online learning aims to improve online decision-making by leveraging offline logged data. A central challenge in this setting is the distribu

SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology

SafetyDGX agent

arXiv:2603.27977v2 Announce Type: replace Abstract: Reinforcement learning is critical to improving large reasoning models, but its success relies heavily on verifiable rewards (RLVR), making it hard

SceneFactory: GPU-Accelerated Multi-Agent Driving Simulation with Physics-Based Vehicle Dynamics

SafetyDGX agent

arXiv:2605.08528v1 Announce Type: cross Abstract: Autonomous-driving simulators typically trade physical fidelity for scalable parallelism. Physics-based platforms such as CARLA and MetaDrive provide

SCOT: Multi-Source Cross-City Transfer with Optimal-Transport Soft-Correspondence Objective

SafetyDGX agent

arXiv:2604.07383v2 Announce Type: replace Abstract: Cross-city transfer improves prediction in label-scarce cities by leveraging labeled data from other cities, but it becomes challenging when cities

SDFlow: Similarity-Driven Flow Matching for Time Series Generation

SafetyDGX agent

arXiv:2605.05736v2 Announce Type: replace Abstract: Vector quantization (VQ) with autoregressive (AR) token modeling is a widely adopted and highly competitive paradigm for time-series generation. How

Segment Anything with Robust Uncertainty-Accuracy Correlation

SafetyDGX agent

arXiv:2605.10603v1 Announce Type: new Abstract: Despite strong zero-shot performance, SAM is unreliable under domain shift due to Mask-level Confidence Confusion (MCC), where a single IoU-based mask s

Selection of the Best Policy under Fairness Constraints for Subpopulations

SafetyDGX agent

arXiv:2605.09945v1 Announce Type: new Abstract: Many high-stakes decisions in health care, public policy, and clinical development require committing to a single policy that will be applied uniformly

Selection Plateau and a Sparsity-Dependent Hierarchy of Pruning Features

SafetyDGX agent

arXiv:2605.09345v1 Announce Type: new Abstract: We identify a Selection Plateau phenomenon in one-shot neural network pruning: all rank-monotone weight scorers converge to identical accuracy at fixed

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories

SafetyDGX agent

arXiv:2605.08936v1 Announce Type: new Abstract: Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe r

Semantic Alignment in Hyperbolic Space for Open-Vocabulary Semantic Segmentation

SafetyDGX agent

arXiv:2605.08874v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation requires adapting image-level vision-language models such as CLIP to dense pixel-level prediction, which is challe

SGC-RML: A reliable and interpretable longitudinal assessment for PD in real-world DNS

SafetyDGX agent

arXiv:2605.08302v1 Announce Type: cross Abstract: Real-world digital Parkinson's disease assessment faces challenges such as heterogeneous modalities, cross-device bias, and incomplete labeling. Exist

SHIELD: Scalable Optimal Control with Certification using Duality and Convexity

SafetyDGX agent

arXiv:2605.09171v1 Announce Type: new Abstract: We present SHIELD, a hierarchical algorithm that reduces both the decision-variable dimension and the constraint set in ell_1-regularized convex program

Shields to Guarantee Probabilistic Safety in MDPs

SafetyDGX agent

arXiv:2605.10888v1 Announce Type: cross Abstract: Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens

← Previous
1…151152153154155…212
Next →