AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,707 results
3 Jul 2026

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning

SafetyDGX agent

arXiv:2602.02762v2 Announce Type: replace Abstract: Semi-supervised imitation learning (SSIL) consists in learning a policy from a small dataset of action-labeled trajectories and a much larger datase

One Demonstration Is Enough for Real-World Robotic Reinforcement Learning

SafetyDGX agent

arXiv:2607.01651v1 Announce Type: new Abstract: Learning effective robot control policies on physical hardware is challenging due to costly data collection and the difficulty of reward specification.

Online Resource Allocation with Continuous Random Consumption: Regret under Degeneracy

SafetyDGX agent

arXiv:2607.02196v1 Announce Type: new Abstract: We study online resource allocation when both rewards and consumption sizes may be continuously distributed. Requests arrive sequentially and must be ac


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Online Safety Monitoring for LLMs

SafetyDGX agent

arXiv:2607.02510v1 Announce Type: new Abstract: Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safet

OpenAI offers feds a stake, Anthropic gets out of AI model jail and Meta wants to be a neocloud

SafetyDGX agent

OpenAI reportedly has floated giving the U.S. government a 5% stake in the company, perhaps the start of a series of such stakes in other AI companies as well. This no doubt has traditional anti-indus

Optimizing Visual Generative Models via Distribution-wise Rewards

SafetyDGX agent

arXiv:2607.02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems

SafetyDGX agent

arXiv:2607.01518v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have been increasingly integrated into robotic systems. However, these models may exhibit overthinking behaviors,

Playing 20 Question Game with Policy-Based Reinforcement Learning

SafetyDGX agent

arXiv:1808.07645v5 Announce Type: replace-cross Abstract: The 20 Questions (Q20) game is a well known game which encourages deductive reasoning and creativity. In the game, the answerer first thinks o

Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander

SafetyDGX agent

arXiv:2607.01736v1 Announce Type: cross Abstract: We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone. Choosing the ri

Prediction Sets for Counterfactual Decisions: Coverage, Optimality, and Conformal Prediction

SafetyDGX agent

arXiv:2607.02206v1 Announce Type: cross Abstract: Predictions are increasingly used to guide high-stakes decisions, from treatment selection to policy making. To ensure reliability with imperfect pred

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

SafetyDGX agent

arXiv:2607.02234v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference

Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score

SafetyDGX agent

arXiv:2607.01527v1 Announce Type: cross Abstract: Room embeddings derived from reverberant speech are often unreliable: speech content and recording degradation can alter the representation even when

Quantum-Inspired Vision: Leveraging Wave-Particle Duality for Low-Illumination Enhancement

SafetyDGX agent

arXiv:2607.01731v1 Announce Type: cross Abstract: This study provides a theoretical expansion of the recent Data Relativistic Uncertainty (DRU) framework by formalizing a physics-to-AI paradigm for im

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

SafetyDGX agent

arXiv:2607.01897v1 Announce Type: cross Abstract: We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards. RTA trains a

RedCoder: Automated Multi-Turn Red Teaming for Code LLMs

SafetyDGX agent

arXiv:2507.22063v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) for code generation (i.e., Code LLMs) have demonstrated impressive capabilities in AI-assisted software developme

Rethinking Post-Hoc Calibration in Semantic Segmentation

SafetyDGX agent

arXiv:2607.01902v1 Announce Type: cross Abstract: Reliable confidence estimates are essential in semantic segmentation, especially in safety-critical settings where overconfident errors can mislead do

Risk Architecture for AI-Native Engineering Teams: An Organizational Framework for Agentic System Governance

SafetyDGX agent

arXiv:2607.01421v1 Announce Type: cross Abstract: Engineering management research has produced mature frameworks for software risk: ownership by feature, escalation by severity, and assurance by test

SABER: A Semantic-Aligned Brain Network Analysis Framework via Multi-scale Hypergraphs

SafetyDGX agent

arXiv:2607.01901v1 Announce Type: cross Abstract: Effective brain disease diagnosis requires the synergy of brain connectivity patterns and high-level semantic knowledge. Existing methods, however, la

Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model

SafetyDGX agent

arXiv:2607.01595v1 Announce Type: new Abstract: As the scale and complexity of cloud-based AI systems continue to escalate, ensuring service reliability through rapid fault detection and adaptive reco

Safeguarding LLM Agents from Misalignment through Provenance Analysis

SafetyDGX agent

arXiv:2607.01236v1 Announce Type: cross Abstract: As LLM agents gain increasing access to powerful tools, ensuring that their actions are aligned with the user's intent becomes critical. When an agent

Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving

SafetyDGX agent

arXiv:2604.03497v2 Announce Type: replace-cross Abstract: Vision-language-model (VLM)-guided reinforcement learning (RL) has recently attracted significant attention for it, replacing brittle hand-cra

SPLC: Social Preference Learning for Crowd Robot Navigation

SafetyDGX agent

arXiv:2607.01925v1 Announce Type: new Abstract: Offline reinforcement learning (RL) holds significant potential for crowd robot navigation in human-robot coexistence applications. However, the inheren

Structuring the Space of Sociotechnical Alignment

SafetyDGX agent

arXiv:2607.01250v1 Announce Type: cross Abstract: Sociotechnical alignment concerns the social desirability of AI behavior and is thus inherently normative, not merely technical. While NLP research in

The Rising Unsustainability of AI Graphics Cards Production

SafetyDGX agent

arXiv:2607.01258v1 Announce Type: cross Abstract: The rapid advancement of Artificial Intelligence (AI) has been accompanied by significant increases in computational and environmental costs, driven b

Tight Lower Bounds for the Multi-Secretary Problem via Bellman Certificates

SafetyDGX agent

arXiv:2607.02150v1 Announce Type: cross Abstract: This paper studies additive regret in the multi-secretary problem, defined as the gap between the expected offline prophet reward and the reward of th

Towards Learning Representations of Policies in Two-Player Zero-Sum Imperfect-Information Games

SafetyDGX agent

arXiv:2607.01498v1 Announce Type: new Abstract: We investigate the problem of learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games. We make three cont

Transformer Geometry Observatory TGO-II: Representational Similarity Observatory

SafetyDGX agent

arXiv:2607.02386v1 Announce Type: cross Abstract: While Vision Transformers have achieved remarkable success across computer vision and language applications, the geometric evolution of their internal

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models

SafetyDGX agent

arXiv:2512.01715v2 Announce Type: replace Abstract: Vision-language-action (VLA) models that generate continuous action chunks via flow matching lack an internal signal for judging whether a given pre

VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment

SafetyDGX agent

arXiv:2607.01586v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remai

WaveLander: A Generalizable Hierarchical Control Framework for UAV Landing on Wave-Disturbed Platforms via Reinforcement Learning

SafetyDGX agent

arXiv:2607.01281v1 Announce Type: new Abstract: Autonomous landing of unmanned aerial vehicles (UAVs) on wave-disturbed marine platforms remains challenging due to stochastic platform motion, time-var

WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field Convolution

SafetyDGX agent

arXiv:2607.02097v1 Announce Type: cross Abstract: Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory ac

What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates

SafetyDGX agent

arXiv:2607.02507v1 Announce Type: new Abstract: LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly t

When Sample Selection Bias Precipitates Model Collapse

SafetyDGX agent

arXiv:2606.13732v2 Announce Type: replace Abstract: The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training erodes distr

When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations

SafetyDGX agent

arXiv:2607.01426v1 Announce Type: new Abstract: Autonomous customer-service agents are shifting from conversational interfaces toward operational execution roles: they retrieve firm records, apply ser

Wind-Aware Reinforcement Learning Control of a Small Quadrotor Using Learned Onboard Wind Estimation in Simulated Atmospheric Turbulence

SafetyDGX agent

arXiv:2607.01528v1 Announce Type: new Abstract: Small multirotor aircraft are increasingly tasked with operations in the atmospheric boundary layer, where turbulent winds comparable to the vehicle's a

WorldSample: Closed-loop Real-robot RL with World Modelling

SafetyDGX agent

arXiv:2607.02431v1 Announce Type: cross Abstract: Reinforcement learning (RL) can overcome the demonstration-coverage limitation of imitation learning (IL) by allowing robots to improve through trial-

Wow, even I was surprised how high AI ranked! Very good to see!

SafetyDGX agent

Wow, even I was surprised how high AI ranked! Very good to see! Interesting poll of Hill staffers from @PunchbowlNews. 250 years is a long time! But interesting to see that 'losing control of AI' is t

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models

SafetyDGX agent

arXiv:2601.15588v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained fil

2 Jul 2026

A Category Theory Account of AI Identity

SafetyDGX agent

arXiv:2607.00220v1 Announce Type: cross Abstract: Artificial intelligence (AI) systems are routinely modified after deployment through retraining and changes in their environments. These transformatio

A Filtered Mixture-of-Generators for Fully Synthetic Survival Training

SafetyDGX agent

arXiv:2607.00127v1 Announce Type: new Abstract: Survival analysis models time-to-event data, but in clinical settings training data are costly and scarce: events accrue over years of follow-up, cohort

A Mechanism-Driven Theory of Phase Transitions in Active Learning

SafetyDGX agent

arXiv:2607.00144v1 Announce Type: cross Abstract: Active learning (AL) performance is known to be budget-dependent, yet regimes are typically defined by heuristic label counts that fail to generalize

A Multi-Resolution Finite-Volume Inspired Deep Learning Framework for Spatiotemporal Dynamics Prediction

SafetyDGX agent

arXiv:2607.00460v1 Announce Type: cross Abstract: Predicting complex spatiotemporal dynamics in physical processes often demands computationally expensive numerical methods or data-driven neural netwo

A small tax on every token produced could be transformative, without putting the government into bed with a specific company. And because ev…

SafetyDGX agent

A small tax on every token produced could be transformative, without putting the government into bed with a specific company. And because every token draws on uncompensated contributions from multiple

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization

SafetyDGX agent

arXiv:2607.00531v1 Announce Type: cross Abstract: Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of training such re

Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers

SafetyDGX agent

arXiv:2607.00580v1 Announce Type: new Abstract: Vision Transformers (ViTs) commonly rely on injected positional mechanisms to address self-attention's permutation invariance. Motivated by the spatial

AI Native Games: A Survey and Roadmap

SafetyDGX agent

arXiv:2607.00527v1 Announce Type: new Abstract: Generative AI now enables games to produce dialogue, quests, characters, images, and worlds at runtime. Yet generation alone does not make a game AI-nat

Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders

SafetyDGX agent

arXiv:2607.00023v1 Announce Type: cross Abstract: Dense sentence embeddings are fundamental to modern Retrieval-Augmented Generation (RAG) systems but suffer from a lack of interpretability due to fea

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire

SafetyDGX agent

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire Altman’s AI safety proposal: let us win, or everybody loses https://ft.trib.al/UrDI

ASPIRE: Agentic /Skills Discovery for Robotics

SafetyDGX agent

arXiv:2607.00272v1 Announce Type: cross Abstract: Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling divers

Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval

SafetyDGX agent

arXiv:2607.00379v1 Announce Type: cross Abstract: Unsupervised cross-modal hashing enables efficient retrieval of semantically related instances across different modalities without requiring manual se

AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation

SafetyDGX agent

arXiv:2607.01051v1 Announce Type: new Abstract: Different stages of manipulation tasks exhibit varying levels of difficulty, suggesting stage-dependent motion speeds and temporal prediction horizons.

Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces

SafetyDGX agent

arXiv:2607.00481v1 Announce Type: cross Abstract: Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs). While prior work has primarily studied attacks and

Bounded Morality: Defining the Space of Moral Computation

SafetyDGX agent

arXiv:2607.00002v1 Announce Type: new Abstract: Moral cognition has traditionally been modeled as adherence to fixed ethical theories--deontology, consequentialism, virtue ethics--implemented as stati

BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure

SafetyDGX agent

arXiv:2607.00573v1 Announce Type: new Abstract: Diffusion MRI probes brain microstructure with particular sensitivity to early cerebrovascular and neurodegenerative changes. Neurite Orientation Disper

Caption Bottleneck Models

SafetyDGX agent

arXiv:2607.00578v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) provide interpretability by routing predictions through a layer of human-understandable concepts. However, defining an

ClinRAG-GRAPH: Clinical-prior Retrieval-Augmented Graph Model with Domain Adversarial Learning for Breast pCR Prediction

SafetyDGX agent

arXiv:2607.00798v1 Announce Type: new Abstract: Neoadjuvant chemotherapy (NAC) response prediction is clinically important for treatment stratification in breast cancer. However, robust pre-treatment

congrats @yudapearl!

SafetyDGX agent

congrats @yudapearl! Judea Pearl Named AI Pioneer by Boston Global Forum in Honor of America’s 250th Anniversary https://samueli.ucla.edu/judea-pearl-named-ai-pioneer-by-boston-global-forum-in-honor-o

'consensus' can't just built by big tech! let's all be careful of regulatory capture and extreme concentration of power!

SafetyDGX agent

'consensus' can't just built by big tech! let's all be careful of regulatory capture and extreme concentration of power! NEW: Anthropic announces it is drafting a consensus framework with Amazon, Micr

Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction

SafetyDGX agent

arXiv:2607.00001v1 Announce Type: new Abstract: Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized. This assumption conflicts with extensive empirica

Continuous Speculative Decoding for Autoregressive Image Generation

SafetyDGX agent

arXiv:2411.11925v3 Announce Type: replace Abstract: Continuous visual autoregressive (AR) models have demonstrated promising performance in image generation, but their inherently sequential nature res

← Previous
1…4748495051…212
Next →