AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,481 results
Safety

OpenAI’s Sam Altman’s personal investments are coming under intensifying scrutiny from Republicans following an April article in The Wall St…

DGX agent

Sam Altman's personal investments faced increased scrutiny from Republican lawmakers following reporting by The Wall Street Journal in April regarding potential conflicts of interest. The controversy

safetygary-marcus--x
12 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

OpenClaw-RL: Train Any Agent Simply by Talking

DGX agent

arXiv:2603.10165v2 Announce Type: replace-cross Abstract: Every agent interaction generates a next-state signal, namely the user reply, tool output, terminal or GUI state change that follows each acti

safetyarxiv-cs-ai
12 May 2026
Safety

Opinion: It is rare for the US public to agree on anything these days. Fear of AI is as close to a national consensus as it gets. A clear ma…

DGX agent

Opinion: It is rare for the US public to agree on anything these days. Fear of AI is as close to a national consensus as it gets. A clear majority says that AI will do more harm than good. https://ft.

safetygary-marcus--x
12 May 2026
Safety

Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning

DGX agent

arXiv:2605.09640v1 Announce Type: new Abstract: Recent studies suggest that Reinforcement Fine-Tuning (RFT) is inherently more resilient to catastrophic forgetting than Supervised Fine-Tuning (SFT). H

safetyarxiv-cs-cv
12 May 2026
Safety

Path-Coupled Bellman Flows for Distributional Reinforcement Learning

DGX agent

arXiv:2605.08253v1 Announce Type: cross Abstract: Distributional reinforcement learning (DRL) models the full return distribution, but existing finite-support or quantile-based methods rely on project

safetyarxiv-cs-ai
12 May 2026
Safety

PATRA: Pattern-Aware Alignment and Balanced Reasoning for Time Series Question Answering

DGX agent

arXiv:2602.23161v2 Announce Type: replace Abstract: Time series reasoning demands both the perception of complex dynamics and logical depth. However, existing LLM-based approaches exhibit two limitati

safetyarxiv-cs-ai
12 May 2026
Safety

Pay attention to this one if you build research or knowledge-work agents. Most research-agent systems produce uniform outputs regardless of …

DGX agent

Pay attention to this one if you build research or knowledge-work agents. Most research-agent systems produce uniform outputs regardless of who is driving them. This new work, NanoResearch, argues tha

safetydair-ai--x
12 May 2026
Safety

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs

DGX agent

arXiv:2605.09422v1 Announce Type: new Abstract: Although Large Multimodal Models (LMMs) have achieved strong performance on general video understanding, their susceptibility to textual prior shortcuts

safetyarxiv-cs-cl
12 May 2026
Safety

Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework

DGX agent

arXiv:2605.10043v1 Announce Type: cross Abstract: Large Language Model (LLM) personalization aims to align model behaviors with individual user preferences. Existing methods often focus on isolated us

safetyarxiv-cs-ai
12 May 2026
Safety

PFN-TS: Thompson Sampling for Contextual Bandits via Prior-Data Fitted Networks

DGX agent

arXiv:2605.10137v1 Announce Type: cross Abstract: Thompson sampling is a widely used strategy for contextual bandits: at each round, it samples a reward function from a Bayesian posterior and acts gre

safetyarxiv-cs-lg
12 May 2026
Safety

PHAGE: Patent Heterogeneous Attention-Guided Graph Encoder for Representation Learning

DGX agent

arXiv:2605.10073v1 Announce Type: new Abstract: Patent claims form a directed dependency structure in which dependent claims inherit and refine the scope of earlier claims; however, existing patent en

safetyarxiv-cs-cl
12 May 2026
Safety

PhysEDA: Physics-Aware Learning Framework for Efficient EDA With Manhattan Distance Decay

DGX agent

arXiv:2605.10547v1 Announce Type: new Abstract: Electronic design automation (EDA) addresses placement, routing, timing analysis, and power-integrity verification for integrated circuits. Learning met

safetyarxiv-cs-lg
12 May 2026
Safety

Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation

DGX agent

arXiv:2605.10118v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated exceptional general reasoning capabilities. However, their performance in embodied navigation remains hi

safetyarxiv-cs-ro
12 May 2026
Safety

Plan2Cleanse: Test-Time Backdoor Defense via Monte-Carlo Planning in Deep Reinforcement Learning

DGX agent

arXiv:2605.09638v1 Announce Type: new Abstract: Ensuring the security of reinforcement learning (RL) models is critical, particularly when they are trained by third parties and deployed in real-world

safetyarxiv-cs-lg
12 May 2026
Safety

PMCTS: Particle Monte Carlo Tree Search for Principled Parallelized Inference Time Scaling

DGX agent

arXiv:2605.08982v1 Announce Type: new Abstract: Monte Carlo Tree Search (MCTS) is a widely used approach for policy improvement through search with increasing popularity for real world applications. D

safetyarxiv-cs-lg
12 May 2026
Safety

Policy Gradient Methods for Non-Markovian Reinforcement Learning

DGX agent

arXiv:2605.10816v1 Announce Type: cross Abstract: We study policy gradient methods for reinforcement learning in non-Markovian decision processes (NMDPs), where observations and rewards depend on the

safetyarxiv-cs-ai
12 May 2026
Safety

Political Plasticity: An Analysis of Ideological Adaptability in Large Language Models

DGX agent

arXiv:2605.08415v1 Announce Type: new Abstract: Since the advent of Large Language Models (LLMs), a significant area of research has focused on their intrinsic biases, particularly in political discou

safetyarxiv-cs-ai
12 May 2026
Safety

Position: Academic Conferences are Potentially Facing Denominator Gaming Caused by Fully Automated Scientific Agents

DGX agent

arXiv:2605.09915v1 Announce Type: cross Abstract: The implicit policy of maintaining relatively stable acceptance rates at top AI conferences, despite exponentially growing submissions, introduces a c

safetyarxiv-cs-ai
12 May 2026
Safety

Positional Encoding via Token-Aware Phase Attention

DGX agent

arXiv:2509.12635v3 Announce Type: replace-cross Abstract: We prove under practical assumptions that Rotary Positional Embedding (RoPE) introduces an intrinsic distance-dependent bias in attention scor

safetyarxiv-cs-ai
12 May 2026
Safety

Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases

DGX agent

arXiv:2605.09472v1 Announce Type: new Abstract: Positional encoding in transformers is commonly implemented through positional embeddings, attention masks, or bias terms, but formal connections betwee

safetyarxiv-cs-lg
12 May 2026
Safety

Primal-Dual Guided Decoding for Constrained Discrete Diffusion

DGX agent

arXiv:2605.09749v1 Announce Type: new Abstract: Discrete diffusion models generate structured sequences by progressively unmasking tokens, but enforcing global property constraints during generation r

safetyarxiv-cs-ai
12 May 2026
Safety

Princeton faculty votes to require proctoring in all in-person exams starting this summer, reversing an 1893 policy amid concerns about AI-fueled cheating (Douglas Belkin/Wall Street Journal)

DGX agent

Douglas Belkin / Wall Street Journal: Princeton faculty votes to require proctoring in all in-person exams starting this summer, reversing an 1893 policy amid concerns about AI-fueled cheating — The c

safetytechmeme
12 May 2026
Safety

Privacy-Aware Video Anomaly Detection through Orthogonal Subspace Projection

DGX agent

arXiv:2605.08651v1 Announce Type: cross Abstract: Video anomaly detection (VAD) systems often prioritize accuracy while overlooking privacy concerns, limiting their suitability for real-world deployme

safetyarxiv-cs-ai
12 May 2026
Safety

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation

DGX agent

arXiv:2605.08774v1 Announce Type: cross Abstract: Long-horizon robotic manipulation requires dense feedback that reflects how a task advances through its procedural stages, not merely whether the fina

safetyarxiv-cs-lg
12 May 2026
Safety

ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design

DGX agent

arXiv:2605.10189v1 Announce Type: cross Abstract: Designing proteins with desired functions or properties represents a core goal in synthetic biology and drug discovery. Recent advances in protein lan

safetyarxiv-cs-ai
12 May 2026
Safety

Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions

DGX agent

arXiv:2605.09893v1 Announce Type: cross Abstract: Large language models (LLMs) are often evaluated based on their stated values, yet these do not reliably translate into their actions, a discrepancy t

safetyarxiv-cs-ai
12 May 2026
Safety

Q-learning with Adjoint Matching

DGX agent

arXiv:2601.14234v3 Announce Type: replace-cross Abstract: We propose Q-learning with Adjoint Matching (QAM), a novel TD-based reinforcement learning (RL) algorithm that tackles a long-standing challen

safetyarxiv-cs-ai
12 May 2026
Safety

Quantile-Coupled Flow Matching for Distributional Reinforcement Learning

DGX agent

arXiv:2605.08515v1 Announce Type: new Abstract: Unlike standard expected-return Reinforcement Learning (RL), Distributional RL (DRL) models the full return distribution, making it better-suited for un

safetyarxiv-cs-lg
12 May 2026
Safety

Re-Triggering Safeguards within LLMs for Jailbreak Detection

DGX agent

arXiv:2605.10611v1 Announce Type: cross Abstract: This paper proposes a jailbreaking prompt detection method for large language models (LLMs) to defend against jailbreak attacks. Although recent LLMs

safetyarxiv-cs-ai
12 May 2026
Safety

Reasoning Compression with Mixed-Policy Distillation

DGX agent

arXiv:2605.08776v1 Announce Type: new Abstract: Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive t

safetyarxiv-cs-ai
12 May 2026
Safety

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

DGX agent

arXiv:2605.10805v1 Announce Type: new Abstract: Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings r

safetyarxiv-cs-ai
12 May 2026
Safety

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

DGX agent

arXiv:2605.09614v1 Announce Type: new Abstract: Long chain-of-thought (CoT) reasoning improves large vision--language models, but visual information often fades during generation, limiting long-horizo

safetyarxiv-cs-cv
12 May 2026
Safety

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias

DGX agent

arXiv:2605.08315v1 Announce Type: new Abstract: Existing LLM-based policy optimizers see only scalar rewards: that a policy scored 0.45, but not whether the agent got stuck in a loop, fell into a hole

safetyarxiv-cs-lg
12 May 2026
Safety

Reinforcement learning for inverse structural design and rapid laser cutting of kirigami prototypes

DGX agent

arXiv:2605.08098v1 Announce Type: new Abstract: Kirigami is an increasingly useful fabrication method to produce shape-programmable metamaterial structures. However, inverse design remains difficult b

safetyarxiv-cs-lg
12 May 2026
Safety

Reinforcement Learning with Action Chunking

DGX agent

arXiv:2507.07969v4 Announce Type: replace-cross Abstract: We present Q-chunking, a simple yet effective recipe for improving reinforcement learning (RL) algorithms for long-horizon, sparse-reward task

safetyarxiv-cs-ai
12 May 2026
Safety

Reinforcing Multimodal Reasoning Against Visual Degradation

DGX agent

arXiv:2605.09262v1 Announce Type: cross Abstract: Reinforcement Learning has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet the resulting policies r

safetyarxiv-cs-cl
12 May 2026
Safety

Relational reasoning and inductive bias in transformers and large language models

DGX agent

arXiv:2506.04289v3 Announce Type: replace Abstract: Transformer-based models have demonstrated remarkable reasoning abilities, but the mechanisms underlying relational reasoning remain poorly understo

safetyarxiv-cs-lg
12 May 2026
Safety

Relational Retrieval: Leveraging Known-Novel Interactions for Generalized Category Discovery

DGX agent

arXiv:2605.09420v1 Announce Type: cross Abstract: In this study, we tackle Generalized Category Discovery (GCD) via a Relational Retrieval perspective, explicitly coupling labeled and unlabeled data t

safetyarxiv-cs-ai
12 May 2026
Safety

Relations Are Channels: Knowledge Graph Embedding via Kraus Decompositions

DGX agent

arXiv:2605.10317v1 Announce Type: cross Abstract: Knowledge graph embedding (KGE) models typically represent each relation as an operator on entity embeddings. In this work, we identify three structur

safetyarxiv-cs-ai
12 May 2026
Safety

Relative Score Policy Optimization for Diffusion Language Models

DGX agent

arXiv:2605.10218v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability require

safetyarxiv-cs-cl
12 May 2026
Safety

Remember to Forget: Gated Adaptive Positional Encoding

DGX agent

arXiv:2605.10414v1 Announce Type: new Abstract: Rotary Positional Encoding (RoPE) is widely used in modern large language models. However, when sequences are extended beyond the range seen during trai

safetyarxiv-cs-lg
12 May 2026
Safety

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models

DGX agent

arXiv:2605.09410v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervisi

safetyarxiv-cs-ai
12 May 2026
Safety

Responsible Benchmarking of Fairness for Automatic Speech Recognition

DGX agent

arXiv:2605.10615v1 Announce Type: new Abstract: Many studies have shown automatic speech processing (ASR) systems have unequal performance across speakergroups (SG's). However, the manner in which suc

safetyarxiv-cs-cl
12 May 2026
Safety

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models

DGX agent

arXiv:2605.08186v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) via entropy minimization (EM) has proven effective for classification tasks, yet its application to generative autoregressi

safetyarxiv-cs-ai
12 May 2026
Safety

Rethinking Loss Reweighting for Imbalance Learning as an Inverse Problem: A Neural Collapse Point of View

DGX agent

arXiv:2605.10047v1 Announce Type: cross Abstract: Loss reweighting is a widely used strategy for long-tailed classification, but existing reweighting strategies often rely on heuristics and rarely def

safetyarxiv-cs-ai
12 May 2026
Safety

Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.09212v1 Announce Type: new Abstract: Centralized training with decentralized execution (CTDE) is a standard framework for cooperative multi-agent policy-gradient reinforcement learning, all

safetyarxiv-cs-lg
12 May 2026
Safety

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

DGX agent

arXiv:2605.06241v2 Announce Type: replace Abstract: Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not

safetyarxiv-cs-cl
12 May 2026
Safety

Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with k-step Policy Gradients

DGX agent

arXiv:2605.10909v1 Announce Type: new Abstract: This work revisits standard policy gradient methods used on restricted policy classes, which are known to get stuck in suboptimal critical points. We id

safetyarxiv-cs-lg
12 May 2026
← Previous
1…220221222223224…302
Next →