AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
All
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,806 results
Safety

Positional Encoding via Token-Aware Phase Attention

DGX agent

arXiv:2509.12635v3 Announce Type: replace-cross Abstract: We prove under practical assumptions that Rotary Positional Embedding (RoPE) introduces an intrinsic distance-dependent bias in attention scor

safetyarxiv-cs-ai
12 May 2026
Safety

Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases

Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2605.09472v1 Announce Type: new Abstract: Positional encoding in transformers is commonly implemented through positional embeddings, attention masks, or bias terms, but formal connections betwee

safetyarxiv-cs-lg
12 May 2026
Safety

Positive Alignment: Artificial Intelligence for Human Flourishing

DGX agent

arXiv:2605.10310v1 Announce Type: new Abstract: Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of ali

safetyarxiv-cs-ai
12 May 2026
Safety

Primal-Dual Guided Decoding for Constrained Discrete Diffusion

DGX agent

arXiv:2605.09749v1 Announce Type: new Abstract: Discrete diffusion models generate structured sequences by progressively unmasking tokens, but enforcing global property constraints during generation r

safetyarxiv-cs-ai
12 May 2026
Safety

Princeton faculty votes to require proctoring in all in-person exams starting this summer, reversing an 1893 policy amid concerns about AI-fueled cheating (Douglas Belkin/Wall Street Journal)

DGX agent

Douglas Belkin / Wall Street Journal: Princeton faculty votes to require proctoring in all in-person exams starting this summer, reversing an 1893 policy amid concerns about AI-fueled cheating — The c

safetytechmeme
12 May 2026
Safety

Privacy-Aware Video Anomaly Detection through Orthogonal Subspace Projection

DGX agent

arXiv:2605.08651v1 Announce Type: cross Abstract: Video anomaly detection (VAD) systems often prioritize accuracy while overlooking privacy concerns, limiting their suitability for real-world deployme

safetyarxiv-cs-ai
12 May 2026
Safety

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation

DGX agent

arXiv:2605.08774v1 Announce Type: cross Abstract: Long-horizon robotic manipulation requires dense feedback that reflects how a task advances through its procedural stages, not merely whether the fina

safetyarxiv-cs-lg
12 May 2026
Safety

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

DGX agent

arXiv:2501.03544v5 Announce Type: replace-cross Abstract: Recent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, the

safetyarxiv-cs-ai
12 May 2026
Safety

ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design

DGX agent

arXiv:2605.10189v1 Announce Type: cross Abstract: Designing proteins with desired functions or properties represents a core goal in synthetic biology and drug discovery. Recent advances in protein lan

safetyarxiv-cs-ai
12 May 2026
Safety

Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions

DGX agent

arXiv:2605.09893v1 Announce Type: cross Abstract: Large language models (LLMs) are often evaluated based on their stated values, yet these do not reliably translate into their actions, a discrepancy t

safetyarxiv-cs-ai
12 May 2026
Safety

Q-learning with Adjoint Matching

DGX agent

arXiv:2601.14234v3 Announce Type: replace-cross Abstract: We propose Q-learning with Adjoint Matching (QAM), a novel TD-based reinforcement learning (RL) algorithm that tackles a long-standing challen

safetyarxiv-cs-ai
12 May 2026
Safety

Quantile-Coupled Flow Matching for Distributional Reinforcement Learning

DGX agent

arXiv:2605.08515v1 Announce Type: new Abstract: Unlike standard expected-return Reinforcement Learning (RL), Distributional RL (DRL) models the full return distribution, making it better-suited for un

safetyarxiv-cs-lg
12 May 2026
Safety

Re-Triggering Safeguards within LLMs for Jailbreak Detection

DGX agent

arXiv:2605.10611v1 Announce Type: cross Abstract: This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs

safetyarxiv-cs-ai
12 May 2026
Safety

Reasoning Compression with Mixed-Policy Distillation

DGX agent

arXiv:2605.08776v1 Announce Type: new Abstract: Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive t

safetyarxiv-cs-ai
12 May 2026
Safety

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

DGX agent

arXiv:2605.10805v1 Announce Type: new Abstract: Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings r

safetyarxiv-cs-ai
12 May 2026
Safety

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

DGX agent

arXiv:2605.09614v1 Announce Type: new Abstract: Long chain-of-thought (CoT) reasoning improves large vision--language models, but visual information often fades during generation, limiting long-horizo

safetyarxiv-cs-cv
12 May 2026
Safety

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias

DGX agent

arXiv:2605.08315v1 Announce Type: new Abstract: Existing LLM-based policy optimizers see only scalar rewards: that a policy scored 0.45, but not whether the agent got stuck in a loop, fell into a hole

safetyarxiv-cs-lg
12 May 2026
Safety

Reinforcement learning for inverse structural design and rapid laser cutting of kirigami prototypes

DGX agent

arXiv:2605.08098v1 Announce Type: new Abstract: Kirigami is an increasingly useful fabrication method to produce shape-programmable metamaterial structures. However, inverse design remains difficult b

safetyarxiv-cs-lg
12 May 2026
Safety

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems

DGX agent

arXiv:2605.08378v1 Announce Type: cross Abstract: Reinforcement learning has become a powerful paradigm for improving the capability of intelligent systems, but its practical deployment faces two cent

safetyarxiv-cs-ai
12 May 2026
Safety

Reinforcement Learning with Action Chunking

DGX agent

arXiv:2507.07969v4 Announce Type: replace-cross Abstract: We present Q-chunking, a simple yet effective recipe for improving reinforcement learning (RL) algorithms for long-horizon, sparse-reward task

safetyarxiv-cs-ai
12 May 2026
Safety

Reinforcing Multimodal Reasoning Against Visual Degradation

DGX agent

arXiv:2605.09262v1 Announce Type: cross Abstract: Reinforcement Learning has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet the resulting policies r

safetyarxiv-cs-cl
12 May 2026
Safety

Relational reasoning and inductive bias in transformers and large language models

DGX agent

arXiv:2506.04289v3 Announce Type: replace Abstract: Transformer-based models have demonstrated remarkable reasoning abilities, but the mechanisms underlying relational reasoning remain poorly understo

safetyarxiv-cs-lg
12 May 2026
Safety

Relational Retrieval: Leveraging Known-Novel Interactions for Generalized Category Discovery

DGX agent

arXiv:2605.09420v1 Announce Type: cross Abstract: In this study, we tackle Generalized Category Discovery (GCD) via a Relational Retrieval perspective, explicitly coupling labeled and unlabeled data t

safetyarxiv-cs-ai
12 May 2026
Safety

Relations Are Channels: Knowledge Graph Embedding via Kraus Decompositions

DGX agent

arXiv:2605.10317v1 Announce Type: cross Abstract: Knowledge graph embedding (KGE) models typically represent each relation as an operator on entity embeddings. In this work, we identify three structur

safetyarxiv-cs-ai
12 May 2026
Safety

Relative Score Policy Optimization for Diffusion Language Models

DGX agent

arXiv:2605.10218v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability require

safetyarxiv-cs-cl
12 May 2026
Safety

Remember to Forget: Gated Adaptive Positional Encoding

DGX agent

arXiv:2605.10414v1 Announce Type: new Abstract: Rotary Positional Encoding (RoPE) is widely used in modern large language models. However, when sequences are extended beyond the range seen during trai

safetyarxiv-cs-lg
12 May 2026
Safety

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models

DGX agent

arXiv:2605.09410v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervisi

safetyarxiv-cs-ai
12 May 2026
Safety

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

DGX agent

arXiv:2605.08257v1 Announce Type: cross Abstract: Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develo

safetyarxiv-cs-ai
12 May 2026
Safety

Responsible Benchmarking of Fairness for Automatic Speech Recognition

DGX agent

arXiv:2605.10615v1 Announce Type: new Abstract: Many studies have shown automatic speech processing (ASR) systems have unequal performance across speakergroups (SG's). However, the manner in which suc

safetyarxiv-cs-cl
12 May 2026
Safety

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models

DGX agent

arXiv:2605.08186v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) via entropy minimization (EM) has proven effective for classification tasks, yet its application to generative autoregressi

safetyarxiv-cs-ai
12 May 2026
Safety

Rethinking Loss Reweighting for Imbalance Learning as an Inverse Problem: A Neural Collapse Point of View

DGX agent

arXiv:2605.10047v1 Announce Type: cross Abstract: Loss reweighting is a widely used strategy for long-tailed classification, but existing reweighting strategies often rely on heuristics and rarely def

safetyarxiv-cs-ai
12 May 2026
Safety

Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.09212v1 Announce Type: new Abstract: Centralized training with decentralized execution (CTDE) is a standard framework for cooperative multi-agent policy-gradient reinforcement learning, all

safetyarxiv-cs-lg
12 May 2026
Safety

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

DGX agent

arXiv:2605.06241v2 Announce Type: replace Abstract: Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not

safetyarxiv-cs-cl
12 May 2026
Safety

Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with k-step Policy Gradients

DGX agent

arXiv:2605.10909v1 Announce Type: new Abstract: This work revisits standard policy gradient methods used on restricted policy classes, which are known to get stuck in suboptimal critical points. We id

safetyarxiv-cs-lg
12 May 2026
Safety

Revitalizing the Beginning: Avoiding Storage Dependency for Model Merging in Continual Learning

DGX agent

arXiv:2605.08311v1 Announce Type: cross Abstract: Model merging provides a compelling paradigm for integrating specialized expertise into a unified multi-task model, a goal that aligns naturally with

safetyarxiv-cs-cv
12 May 2026
Safety

Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios

DGX agent

arXiv:2512.00920v4 Announce Type: replace Abstract: Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods fo

safetyarxiv-cs-cl
12 May 2026
Safety

Reward-Conditioned Reinforcement Learning

DGX agent

arXiv:2603.05066v2 Announce Type: replace Abstract: Single-task RL agents are typically trained under a fixed reward function, which limits their robustness to reward misspecification and their abilit

safetyarxiv-cs-lg
12 May 2026
Safety

RigidFormer: Learning Rigid Dynamics using Transformers

DGX agent

arXiv:2605.09196v1 Announce Type: cross Abstract: Learning-based simulation of multi-object rigid-body dynamics remains difficult because contact is discontinuous and errors compound over long horizon

safetyarxiv-cs-ai
12 May 2026
Safety

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning

DGX agent

arXiv:2605.10293v1 Announce Type: cross Abstract: In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide gua

safetyarxiv-cs-ai
12 May 2026
Safety

Route by State, Recover from Trace: STAR with Failure-Aware Markov Routing for Multi-Agent Spatiotemporal Reasoning

DGX agent

arXiv:2605.10057v1 Announce Type: new Abstract: Compositional spatiotemporal reasoning often requires a system to invoke multiple heterogeneous specialists, such as geometric, temporal, topological, a

safetyarxiv-cs-ai
12 May 2026
Safety

RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems

DGX agent

arXiv:2605.10862v1 Announce Type: new Abstract: This paper demonstrates RUBEN, an interactive tool for discovering minimal rules to explain the outputs of retrieval-augmented large language models (LL

safetyarxiv-cs-cl
12 May 2026
Safety

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

DGX agent

arXiv:2605.10899v1 Announce Type: new Abstract: Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyo

safetyarxiv-cs-cl
12 May 2026
Safety

RuPLaR : Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors From Multi-Step to One-Step

DGX agent

arXiv:2605.09346v1 Announce Type: cross Abstract: The Chain-of-Thought (CoT) paradigm, while enhancing the interpretability of Large Language Models (LLMs), is constrained by the inefficiencies and ex

safetyarxiv-cs-ai
12 May 2026
Safety

Safe and Real-Time Consistent Planning for Autonomous Vehicles in Partially Observed Environments via Parallel Consensus Optimization

DGX agent

arXiv:2409.10310v3 Announce Type: replace Abstract: Ensuring safety and driving consistency is a significant challenge for autonomous vehicles operating in partially observed environments. This work i

safetyarxiv-cs-ro
12 May 2026
Safety

Safe Exploration for Nonlinear Processes Using Online Gaussian Process Learning

DGX agent

arXiv:2605.09772v1 Announce Type: cross Abstract: This paper proposes a safe data-driven control framework for nonlinear systems with partially known dynamics. The method ensures stability and constra

safetyarxiv-cs-ro
12 May 2026
Safety

Safety-Critical LiDAR-Inertial Odometry with On-Manifold Deterministic Protection Level

DGX agent

arXiv:2605.09383v1 Announce Type: new Abstract: In safety-critical scenarios, the protection level of the autonomous navigation system is crucial for enabling mobile robots to perform safe tasks. Howe

safetyarxiv-cs-ro
12 May 2026
Safety

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

DGX agent

arXiv:2510.20129v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, where adversarially crafted prompts induce policy-violating responses des

safetyarxiv-cs-ai
12 May 2026
Safety

SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

DGX agent

arXiv:2605.08334v1 Announce Type: new Abstract: We present SalesSim, a framework and testbed for evaluating the ability of Multimodal Large Language Models (MLLMs) to simulate realistic, persona-drive

safetyarxiv-cs-cl
12 May 2026
← Previous
1…191192193194195…267
Next →