AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,702 results
31 Jul 2026

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

SafetyDGX agent

arXiv:2607.28457v1 Announce Type: cross Abstract: Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refine

TAPO: Transition-Aware Policy Optimization for LLM Agents

SafetyDGX agent

arXiv:2607.27973v1 Announce Type: new Abstract: Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents. However, existing me

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

SafetyDGX agent

arXiv:2607.28058v1 Announce Type: new Abstract: Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

SafetyDGX agent

arXiv:2602.08159v2 Announce Type: replace-cross Abstract: When a language model asserts that 'the capital of Australia is Sydney,' does it know this is wrong? Models assert misconceptions with the sam

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

SafetyDGX agent

arXiv:2607.26067v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

SafetyDGX agent

arXiv:2607.26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models

SafetyDGX agent

arXiv:2607.27281v1 Announce Type: new Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth noth

this is the whole problem in a nutshell. LLM-centered systems just can’t be trusted to follow hard constraints, we absolutely must find alte…

SafetyDGX agent

this is the whole problem in a nutshell. LLM-centered systems just can’t be trusted to follow hard constraints, we absolutely must find alternatives that can, or we are screwed. Anybody remember this

Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles

SafetyDGX agent

arXiv:2607.28483v1 Announce Type: new Abstract: Real-time anomaly segmentation is essential for the safety of autonomous systems. Although recent approaches offer high accuracy, their computational co

Uncertainty quantification for trustworthy deep learning: Methods and measures

SafetyDGX agent

arXiv:2607.28248v1 Announce Type: cross Abstract: The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures

UniCross: Unified Cross-Skill Dexterous Manipulation Synthesis

SafetyDGX agent

arXiv:2607.28198v1 Announce Type: cross Abstract: Many dexterous manipulation tasks require the object to remain securely held throughout the interaction. From the perspective of hand-object relationa

Unifying Adversarially Robust Model Experts in Vision-Language Models

SafetyDGX agent

arXiv:2607.27897v1 Announce Type: new Abstract: Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

SafetyDGX agent

arXiv:2607.28590v1 Announce Type: cross Abstract: Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view t

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards

SafetyDGX agent

arXiv:2511.23310v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective paradigm for post-training large language models, yet the de

When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment

SafetyDGX agent

arXiv:2607.27530v1 Announce Type: new Abstract: Graph-text retrieval typically maps a graph and its description to a single embedding, even when a query concerns only one semantic aspect, such as a cl

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models

SafetyDGX agent

arXiv:2607.27599v1 Announce Type: cross Abstract: Building generalizable agents for diverse applications remains a fundamental challenge. While imitation learning-based policies succeed in specific tr

30 Jul 2026

A fundamental flaw leaves LLMs strikingly vulnerable to attack

SafetyDGX agent

It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conferen

A Persona-based Rate Action Index

SafetyDGX agent

arXiv:2607.26545v1 Announce Type: cross Abstract: We propose an index for predicting the U.S. Federal Open Market Committee (FOMC) decision to hike/hold/cut the current federal funds target rate based

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

SafetyDGX agent

arXiv:2607.26170v1 Announce Type: new Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting imag

AgentGFM: A Graph Foundation Model with Node-Agent Information-Flow Control

SafetyDGX agent

arXiv:2607.26533v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) aim to learn transferable knowledge from multi-domain graphs and adapt to unseen scenarios. As a fundamental source of re

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

SafetyDGX agent

arXiv:2607.26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by

AI Alignment in Medical Imaging: Unveiling Hidden Biases Through Counterfactual Analysis

SafetyDGX agent

arXiv:2504.19621v2 Announce Type: replace Abstract: Machine learning (ML) systems for medical imaging have demonstrated remarkable diagnostic capabilities, but their susceptibility to biases poses sig

AlloyDB adds group authentication to secure enterprise scale and AI agents

SafetyDGX agent

Database security traditionally relies on a fragile balance between the granular control developers need and the administrative overhead of managing thousands of individual database passwords. Between

Anatomy Contextualized Adaption of CT Foundation Models

SafetyDGX agent

arXiv:2607.27154v1 Announce Type: new Abstract: CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume repres

Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time

SafetyDGX agent

arXiv:2607.26647v1 Announce Type: new Abstract: While text-to-image diffusion models achieve impressive visual quality, they frequently struggle to maintain precise alignment with complex compositiona

Atomic Chat signed the Open Weights letter! We believe everyone should be able to run AI on their own device. When a model is open, thousand…

SafetyDGX agent

Atomic Chat signed the Open Weights letter! We believe everyone should be able to run AI on their own device. When a model is open, thousands of teams fine-tune it, quantize it and build new tools on

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

SafetyDGX agent

arXiv:2607.26914v1 Announce Type: new Abstract: Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed

CASIAL: Geometric Distortion Robust Image Watermarking

SafetyDGX agent

arXiv:2607.26729v1 Announce Type: new Abstract: Deep learning-based watermarking has shown strong robustness against non-geometric distortions, yet its performance under geometric transformations rema

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

SafetyDGX agent

arXiv:2607.26452v1 Announce Type: cross Abstract: World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

SafetyDGX agent

arXiv:2607.26789v1 Announce Type: new Abstract: Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions withou

CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents

SafetyDGX agent

arXiv:2607.26910v1 Announce Type: new Abstract: Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

SafetyDGX agent

arXiv:2607.26529v1 Announce Type: new Abstract: Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained contr

Cohere has joined @NVIDIA alongside industry leaders in founding the Open Secure AI Alliance. Everybody should have the capability to keep t…

SafetyDGX agent

Cohere has joined @NVIDIA alongside industry leaders in founding the Open Secure AI Alliance. Everybody should have the capability to keep their infrastructure secure. Everybody deserves access to mod

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

SafetyDGX agent

arXiv:2607.26509v1 Announce Type: new Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improveme

Constitutional Midtraining: Content Presence Drives Alignment Gains

SafetyDGX agent

arXiv:2607.26654v1 Announce Type: new Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce

Controlled Experiments on Lane Changing by Transitional Autonomous Vehicle: Dataset and Behavioral Insights

SafetyDGX agent

arXiv:2607.27085v1 Announce Type: new Abstract: This paper presents the North Carolina Transitional Autonomous Vehicle Lane-Changing (NC-tALC) dataset and uses it to characterize mandatory lane-changi

Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation

SafetyDGX agent

arXiv:2607.26164v1 Announce Type: new Abstract: Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicab

DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models

SafetyDGX agent

arXiv:2607.26891v1 Announce Type: new Abstract: Sequence labeling is a fine-grained information extraction task, yet existing large language model-based approaches suffer from insufficient domain alig

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

SafetyDGX agent

arXiv:2607.26811v1 Announce Type: new Abstract: Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they t

Do Methods Support the Claims? Intra-Paper Verification for Peer Review

SafetyDGX agent

arXiv:2607.26066v1 Announce Type: new Abstract: The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review. Existing automated nove

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

SafetyDGX agent

arXiv:2607.26411v1 Announce Type: new Abstract: Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabi

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

SafetyDGX agent

arXiv:2607.27203v1 Announce Type: new Abstract: Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) thi

Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

SafetyDGX agent

arXiv:2607.26735v1 Announce Type: new Abstract: Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target images without

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

SafetyDGX agent

arXiv:2607.26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all resp

Eddeep: a deep-learning framework for fast eddy-current distortion correction in diffusion MRI

SafetyDGX agent

arXiv:2607.26292v1 Announce Type: new Abstract: Diffusion MRI (dMRI) relies on diffusion-weighted echo-planar imaging, which is highly susceptible to eddy-current-induced geometric distortions. These

Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making

SafetyDGX agent

arXiv:2607.27022v1 Announce Type: new Abstract: Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

SafetyDGX agent

arXiv:2607.27145v1 Announce Type: new Abstract: As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitorin

FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets

SafetyDGX agent

arXiv:2607.26423v1 Announce Type: cross Abstract: As autonomous drone deployments scale from individual units to coordinated swarms, the human operator's role shifts from direct piloting to high-level

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

SafetyDGX agent

arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk

FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows

SafetyDGX agent

arXiv:2607.26645v1 Announce Type: new Abstract: Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point clouds are cons

From Found to Designed: Concepts as a Design Axis for Large Language Models

SafetyDGX agent

arXiv:2607.26825v1 Announce Type: new Abstract: Large language models (LLMs) encode rich concept-like information, but represent it implicitly through distributed statistical associations rather than

From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data

SafetyDGX agent

arXiv:2607.26521v1 Announce Type: new Abstract: Conventional subgroup analyses can yield unstable and difficult-to-interpret conclusions, especially in observational biomedical data where each individ

Graph Signal Diffusion Models for Wireless Resource Allocation

SafetyDGX agent

arXiv:2604.05175v2 Announce Type: replace-cross Abstract: We consider constrained ergodic resource optimization in wireless networks with graph-structured interference. We train a diffusion model poli

Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning

SafetyDGX agent

arXiv:2607.14117v2 Announce Type: replace Abstract: Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

SafetyDGX agent

arXiv:2607.26515v1 Announce Type: new Abstract: We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and ba

Journey Operators for Structured Multi-Axis Composition

SafetyDGX agent

arXiv:2607.26775v1 Announce Type: new Abstract: Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D vo

Lag-aware cross-hand alignment for dual-hand action segmentation

SafetyDGX agent

arXiv:2607.26215v1 Announce Type: new Abstract: Dual-hand action segmentation commonly fuses left- and right-hand representations at identical temporal indices, although coordinated hand transitions m

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

SafetyDGX agent

arXiv:2607.26060v1 Announce Type: new Abstract: LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical

Latent-IM: Latent Interaction Management for Speech LLMs

SafetyDGX agent

arXiv:2607.26928v1 Announce Type: new Abstract: Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a gener

Learning Implicit Causal World Models from Multi-Agent Demonstrations

SafetyDGX agent

arXiv:2607.26336v1 Announce Type: new Abstract: In model-based reinforcement learning, world models exist as internal simulators, but their training often conflates statistical correlations with causa

← Previous
1…2122232425…212
Next →