AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,814 results
Safety

Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation

DGX agent

arXiv:2605.25170v1 Announce Type: cross Abstract: Training data for olfaction is scattered through disparate, non-standardized datasets that limit the ability to build representative world models. Olf

safetyarxiv-cs-ai
26 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models

DGX agent

arXiv:2605.25443v1 Announce Type: new Abstract: Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Gro

safetyarxiv-cs-cl
26 May 2026
Safety

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

DGX agent

arXiv:2605.25967v1 Announce Type: new Abstract: As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autore

safetyarxiv-cs-lg
26 May 2026
Safety

Hidden-State Privacy Has an Empty Middle

DGX agent

arXiv:2605.24042v1 Announce Type: cross Abstract: Of 1{,}536 Gaussian release covariances we tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy again

safetyarxiv-cs-ai
26 May 2026
Safety

Hide-and-Shill: A Reinforcement Learning Framework for Market Manipulation Detection in Symphony-a Decentralized Multi-Agent System

DGX agent

arXiv:2507.09179v3 Announce Type: replace Abstract: Decentralized finance (DeFi) has introduced a new era of permissionless financial innovation but also led to unprecedented market manipulation. With

safetyarxiv-cs-ai
26 May 2026
Safety

Hide to Guide: Learning via Semantic Masking

DGX agent

arXiv:2605.25198v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but i

safetyarxiv-cs-ai
26 May 2026
Safety

HoLoArm: Deformable Arms for Collision-Tolerant Quadrotor Flight

DGX agent

arXiv:2605.25790v1 Announce Type: new Abstract: The increasing use of drones in human-centric applications highlights the need for designs that can survive collisions and recover rapidly, minimizing r

safetyarxiv-cs-ro
26 May 2026
Safety

How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description

DGX agent

arXiv:2605.24351v1 Announce Type: new Abstract: Large language models (LLMs) can support scientific literature synthesis, but remain prone to hallucinated references, uneven coverage, and weakly groun

safetyarxiv-cs-cl
26 May 2026
Safety

How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis

DGX agent

arXiv:2605.24749v1 Announce Type: cross Abstract: Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed po

safetyarxiv-cs-lg
26 May 2026
Safety

How to Mitigate the Distribution Shift Problem in Robotics Control: A Robust and Adaptive Approach Based on Offline to Online Imitation Learning

DGX agent

arXiv:2605.25414v1 Announce Type: new Abstract: Distribution shift in imitation learning refers to the problem that the agent cannot plan proper actions for a state that has not been visited during th

safetyarxiv-cs-ro
26 May 2026
Safety

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

DGX agent

arXiv:2605.24934v1 Announce Type: cross Abstract: Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challengi

safetyarxiv-cs-ai
26 May 2026
Safety

HumanFlow -- Diffusion-Driven MAV Navigation Among Humans via Tightly-Coupled Motion Tracking, Forecasting, and Control

DGX agent

arXiv:2605.25685v1 Announce Type: new Abstract: Robust and accurate perception of humans in their 3D scene context is essential for integrating robots into everyday environments. Existing approaches,

safetyarxiv-cs-ro
26 May 2026
Safety

I feel like I'm eating crazy pills when I read the countless bad takes around how the Vatican would have virtually anointed Anthropic. When …

DGX agent

I feel like I'm eating crazy pills when I read the countless bad takes around how the Vatican would have virtually anointed Anthropic. When if you read the Pope's encyclical it's actually a COMPLETE r

safetygary-marcus--x
26 May 2026
Safety

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks

DGX agent

arXiv:2605.24217v1 Announce Type: new Abstract: As Large Language Models (LLMs) transition from research environments to production deployments, evaluating their performance against strict Service Lev

safetyarxiv-cs-ai
26 May 2026
Safety

⚠️⚠️⚠️if you believe this you and think you are safe, you don’t understand how S&P is about to change the rules and what that means. run don…

DGX agent

⚠️⚠️⚠️if you believe this you and think you are safe, you don’t understand how S&P is about to change the rules and what that means. run don’t walk to read my essay “This one weird trick might cost yo

safetygary-marcus--x
26 May 2026
Safety

if you don’t follow or think you know better i urge you to read my essay “This one weird trick might cost your retirement fund billions”

DGX agent

Gary Marcus discusses a potentially overlooked financial risk that could have massive implications for retirement savings, presenting his analysis in an essay format on X. The post appears to emphasiz

safetygary-marcus--x
26 May 2026
Safety

Import AI 458: Reckoning with the future; and a singularity story

DGX agent

Import AI 458 discusses perspectives on AI's future trajectory and potential long-term scenarios, likely including analysis of singularity concepts and their implications. The newsletter entry examine

safetyimport-ai
26 May 2026
Safety

Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression

DGX agent

arXiv:2603.05691v2 Announce Type: replace Abstract: It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenom

safetyarxiv-cs-lg
26 May 2026
Safety

Improving Ensemble CAPE Forecasts with a Diffusion Model Incorporating Aerosol Information

DGX agent

arXiv:2605.24009v1 Announce Type: cross Abstract: Convective available potential energy (CAPE) is an important variable for forecasting severe weather and understanding deep convection and precipitati

safetyarxiv-cs-lg
26 May 2026
Safety

Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation

DGX agent

arXiv:2605.24247v1 Announce Type: cross Abstract: Many automated labeling pipelines classify inputs into categories defined by a written specification, content moderation being a prominent use case. S

safetyarxiv-cs-ai
26 May 2026
Safety

Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach

DGX agent

arXiv:2605.23924v1 Announce Type: new Abstract: Segment-level disclosures are a central component of financial reporting, providing insight into firms' internal organization and the allocation of econ

safetyarxiv-cs-cl
26 May 2026
Safety

In our new paper, EDGE-OPD (EviDence GuidEd On-Policy Distillation), we introduce two key ideas for improving On-Policy Distillation (OPD). …

DGX agent

In our new paper, EDGE-OPD (EviDence GuidEd On-Policy Distillation), we introduce two key ideas for improving On-Policy Distillation (OPD). First, we use guided rollouts that inject privileged context

safetyemad-mostaque--x
26 May 2026
Safety

Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

DGX agent

arXiv:2605.25123v1 Announce Type: cross Abstract: We study inference-time alignment for diffusion-based generative models, aiming to steer a base model toward high-reward outputs without updating its

safetyarxiv-cs-ai
26 May 2026
Safety

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning

DGX agent

arXiv:2605.05226v2 Announce Type: replace-cross Abstract: The central challenge of reinforcement learning for reasoning lies not only in the sparsity of outcome-level supervision, but more fundamental

safetyarxiv-cs-ai
26 May 2026
Safety

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

DGX agent

arXiv:2605.24883v1 Announce Type: new Abstract: The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on con

safetyarxiv-cs-ai
26 May 2026
Safety

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

DGX agent

arXiv:2605.24960v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated und

safetyarxiv-cs-ai
26 May 2026
Safety

Is Decentralized AI Governable? From Regulative Policy to Constitutive Protocol

DGX agent

arXiv:2605.24538v1 Announce Type: cross Abstract: Every major framework for governing artificial intelligence presupposes an identifiable entity -- a developer, deployer, or operator -- who can be hel

safetyarxiv-cs-ai
26 May 2026
Safety

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection

DGX agent

arXiv:2509.13608v2 Announce Type: replace Abstract: As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Align

safetyarxiv-cs-lg
26 May 2026
Safety

IsaacIPC: Coupling High-Fidelity Simulation and Realistic Rendering for Contact-Rich Robotic Systems

DGX agent

arXiv:2605.24339v1 Announce Type: new Abstract: We present IsaacIPC, a robotic simulation framework that couples GPU accelerated incremental potential contact (IPC) with IsaacSim/Lab. IsaacIPC maps si

safetyarxiv-cs-ro
26 May 2026
Safety

Iterative Feature Space Optimization through Incremental Adaptive Evaluation

DGX agent

arXiv:2501.14889v2 Announce Type: replace Abstract: Iterative feature space optimization involves systematically evaluating and adjusting the feature space to improve downstream task performance. Howe

safetyarxiv-cs-lg
26 May 2026
Safety

Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation

DGX agent

arXiv:2605.24041v1 Announce Type: cross Abstract: Neural operators serve as fast, data-driven surrogates for scientific modeling but typically rely on a monolithic, single-pass inference procedure tha

safetyarxiv-cs-ai
26 May 2026
Safety

It's a shame the Pope didn't ask Chris Olah what happened to his plan to give up to 10% of Anthropic to the authors of the work they train o…

DGX agent

It's a shame the Pope didn't ask Chris Olah what happened to his plan to give up to 10% of Anthropic to the authors of the work they train on. (Spoiler: it never happened, and the authors on whose wor

safetygary-marcus--x
26 May 2026
Safety

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

DGX agent

arXiv:2605.23997v1 Announce Type: cross Abstract: Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they

safetyarxiv-cs-ai
26 May 2026
Safety

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models

DGX agent

arXiv:2605.24550v1 Announce Type: new Abstract: Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs), but it can weaken safety-alignment under harmful fine-tuning at

safetyarxiv-cs-ai
26 May 2026
Safety

Joint Optimization of Training and Inference in Federated Edge Learning via Constrained Multi-Objective Deep Reinforcement Learning

DGX agent

arXiv:2605.25916v1 Announce Type: new Abstract: Federated edge learning (FEEL) has recently emerged as a promising paradigm for achieving edge intelligence (EI) via enabling collaborative model traini

safetyarxiv-cs-lg
26 May 2026
Safety

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

DGX agent

arXiv:2605.24414v1 Announce Type: new Abstract: We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe

safetyarxiv-cs-ai
26 May 2026
Safety

KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable Provenance and Hierarchical Policy Composition

DGX agent

arXiv:2605.25376v1 Announce Type: cross Abstract: Observability tells operators when an agent is slow. KYA tells operators when an agent is wrong, drifting, leaking, or quietly going rogue. We present

safetyarxiv-cs-ai
26 May 2026
Safety

Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime

DGX agent

arXiv:2605.25275v1 Announce Type: new Abstract: The Neural Tangent Kernel (NTK) framework explains optimization in over-parameterized neural networks via approximately linearized dynamics, yielding ex

safetyarxiv-cs-lg
26 May 2026
Safety

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation

DGX agent

arXiv:2605.25036v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) extend large language models with visual understanding, but remain vulnerable to hallucination, where outputs are

safetyarxiv-cs-ai
26 May 2026
Safety

LAPLEX: The FFT of Learnable Laplace Kernels

DGX agent

arXiv:2605.24584v1 Announce Type: cross Abstract: Fast linear algebra in deep learning usually comes with a choice: fixed geometry and exact computation, as in the Fourier transform, or adaptive geome

safetyarxiv-cs-ai
26 May 2026
Safety

Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

DGX agent

arXiv:2605.25740v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However,

safetyarxiv-cs-lg
26 May 2026
Safety

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition

DGX agent

arXiv:2605.24005v1 Announce Type: new Abstract: The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous

safetyarxiv-cs-ai
26 May 2026
Safety

Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models

DGX agent

arXiv:2603.29123v2 Announce Type: replace Abstract: The next-token prediction (NTP) objective trains language models to predict a single token at each step, even though many continuations can express

safetyarxiv-cs-cl
26 May 2026
Safety

Learning High-Frequency Continuous Action Chunks in Latent Space

DGX agent

arXiv:2605.24931v1 Announce Type: new Abstract: Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal con

safetyarxiv-cs-ro
26 May 2026
Safety

Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning

DGX agent

arXiv:2605.26012v1 Announce Type: cross Abstract: Deep reinforcement learning (RL) agents commonly rely on high-dimensional neural representations, despite growing evidence that task-relevant value an

safetyarxiv-cs-ai
26 May 2026
Safety

Learning to Route Languages for Multilingual Policy Optimization

DGX agent

arXiv:2605.25360v1 Announce Type: new Abstract: Large language models~(LLMs) are trained on heterogeneous multilingual corpora, yet existing policy optimization methods often implicitly restrict each

safetyarxiv-cs-cl
26 May 2026
Safety

LipoAgent: Coordinating Fine-Tuned LLM Agents for Safer Lipid Design

DGX agent

arXiv:2605.25250v1 Announce Type: new Abstract: Lipid nanoparticles (LNPs) are among the most clinically mature platforms for nucleic acid delivery, yet designing lipids that are both effective and bi

safetyarxiv-cs-ai
26 May 2026
Safety

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models

DGX agent

arXiv:2605.25179v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cos

safetyarxiv-cs-cl
26 May 2026
← Previous
1…146147148149150…267
Next →