AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,488 results
2 Jun 2026

V-LynX: Token Interface Alignment for Video+X LLMs

SafetyDGX agent

arXiv:2606.00508v1 Announce Type: cross Abstract: This study introduces an intriguing phenomenon in Video LLMs: rather than merely translating frames into textual embeddings, Video LLMs establish a co

Value-Free Policy Optimization via Reward Partitioning

SafetyDGX agent

arXiv:2506.13702v4 Announce Type: replace-cross Abstract: Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative

VERA: Variational Inference Framework for Jailbreaking Large Language Models

SafetyDGX agent

arXiv:2506.22666v3 Announce Type: replace-cross Abstract: The rise of API-only access to state-of-the-art LLMs highlights the need for effective black-box jailbreak methods to identify model vulnerabi

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Visualizing definitional divergence in high-dimensional data by manifold alignment: Application to 3D right ventricular strain computations

SafetyDGX agent

arXiv:2501.12178v2 Announce Type: replace Abstract: Medical imaging studies often rely on a single sample per subject, assuming it is representative of their physiological traits. However, variations

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

SafetyDGX agent

arXiv:2601.03309v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models, which integrate pretrained large Vision-Language Models (VLM) into their policy backbone, are gaining sig

Wavelet-Fusion Diffusion Model for Multimodal Brain MRI Synthesis with Modality and Metadata Conditioning

SafetyDGX agent

arXiv:2606.00689v1 Announce Type: new Abstract: Multimodal MRI provides complementary information for neuroimaging analysis, where different imaging modalities capture distinct anatomical, tissue, and

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight

SafetyDGX agent

arXiv:2606.00424v1 Announce Type: new Abstract: As large language models become stronger, weak supervisors may fail to provide reliable labels, preferences, or final judgments for complex outputs, lim

When Does Predictive Inverse Dynamics Outperform Behavior Cloning?

SafetyDGX agent

arXiv:2601.21718v2 Announce Type: replace-cross Abstract: Behavior cloning (BC) is a practical offline imitation learning method, but it often fails when expert demonstrations are limited. Recent work

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models

SafetyDGX agent

arXiv:2606.01671v1 Announce Type: new Abstract: In the contemporary epoch of multilingual education, learning idioms provides a fascinating gateway towards creativity, cultural values, historical cont

Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations

SafetyDGX agent

arXiv:2511.05613v2 Announce Type: replace-cross Abstract: Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risk

Why things will eventually fall apart: 1. Everybody, even Google, seems to be treating AI as if it were some kind of winner take all competi…

SafetyDGX agent

Why things will eventually fall apart: 1. Everybody, even Google, seems to be treating AI as if it were some kind of winner take all competition like web search was, in which Google taking over 95% 2.

World Models for Robotic Manipulation: A Survey

SafetyDGX agent

arXiv:2606.00113v1 Announce Type: new Abstract: Robotic manipulation depends on the ability to anticipate how actions reshape objects, contacts, and scene geometry before execution. Learned world mode

World-Task Factorization for Robot Learning

SafetyDGX agent

arXiv:2606.02027v1 Announce Type: cross Abstract: Robot learning must produce policies that generalize to new combinations of constraints, teammates, and environments. To achieve this, we must structu

1 Jun 2026

A hitchhiker's guide to Poisson gradient estimation

SafetyDGX agent

arXiv:2602.03896v2 Announce Type: replace-cross Abstract: Poisson-distributed latent variable models are widely used in computational neuroscience, but differentiating through discrete stochastic samp

A Lecture Note on Offline RL and IRL, Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models

SafetyDGX agent

arXiv:2605.30843v1 Announce Type: new Abstract: In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn

A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI

SafetyDGX agent

arXiv:2605.31021v1 Announce Type: new Abstract: Current alignment paradigms for generative artificial intelligence rely predominantly on monolithic benchmarking frameworks that reduce the plurality of

A Unified Framework for Gradient Aggregation in Multi-Objective Optimization

SafetyDGX agent

arXiv:2605.30452v1 Announce Type: cross Abstract: Many machine learning problems involve multiple inherent trade-offs that are best addressed by gradient-based multi-objective optimization (MOO) algor

Active Timepoint Selection for Learning Measure-Valued Trajectories

SafetyDGX agent

arXiv:2605.30625v1 Announce Type: cross Abstract: Inferring continuous probability paths from sparse snapshots is a fundamental challenge in domains like single-cell biology, where high-fidelity data

AI Loss of Control Incident Management: Response & Resilience

SafetyDGX agent

arXiv:2605.30406v1 Announce Type: cross Abstract: Recent research demonstrating AI systems exhibiting deception and shutdown resistance suggests that AI loss of control (LOC) is an urgent policy conce

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

SafetyDGX agent

arXiv:2605.31034v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling

Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding

SafetyDGX agent

arXiv:2605.30742v1 Announce Type: new Abstract: This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topi

Are Full Rollouts Necessary for On-Policy Distillation?

SafetyDGX agent

arXiv:2605.31490v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense teacher feedback along rollouts generated by the student and has emerged as a promising post-training paradi

BIAS-ID: A Framework for Analyzing Transformation Biases in AI-Generated Image Detectors

SafetyDGX agent

arXiv:2605.31153v1 Announce Type: new Abstract: Given the surge of harmful AI-generated imagery online, reliably distinguishing authentic images from generated ones has become an urgent research topic

Biases in the Blind Spot: Detecting What LLMs Fail to Mention

SafetyDGX agent

arXiv:2602.10117v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We cal

BiSegMamba: Efficient Bidirectional Tri-Oriented Mamba for 3D Medical Image Segmentation

SafetyDGX agent

arXiv:2605.30972v1 Announce Type: new Abstract: Accurate 3D medical image segmentation requires both long-range volumetric context and fine boundary preservation. CNN-based methods have limited global

Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models

SafetyDGX agent

arXiv:2510.11683v3 Announce Type: replace-cross Abstract: A key challenge in applying reinforcement learning (RL) to diffusion large language models (dLLMs) is the intractability of their likelihood f

Breaking Information Cocoons: A Hyperbolic Framework for Balancing Exploration and Exploitation in Recommender Systems

SafetyDGX agent

arXiv:2411.13865v4 Announce Type: replace-cross Abstract: Modern recommender systems often create information cocoons, restricting users' exposure to diverse content. The central challenge is to balan

Building Generalization Into Behavior Generation Via Adaptive Compositions of Regularities

SafetyDGX agent

arXiv:2605.31110v1 Announce Type: new Abstract: Generalization in robotics requires prior knowledge about how the world is structured, yet this structure changes from one situation to the next. This p

Calibrated Uncertainty for Trustworthy Clinical Gait Analysis Using Probabilistic Multiview Markerless Motion Capture

SafetyDGX agent

arXiv:2601.22412v2 Announce Type: replace Abstract: Video-based human movement analysis holds potential for movement assessment in clinical practice and research. However, the clinical implementation

Can Aerial VLA Models Cooperate? Evaluating Closed-Loop Air-Ground Coordination with CARLA-Air

SafetyDGX agent

arXiv:2605.31066v1 Announce Type: new Abstract: Recent aerial vision-language-action (VLA) models show promising single-UAV capabilities, such as tracking moving objects and navigating to language-spe

Causal Evaluation of Membership Inference Attacks

SafetyDGX agent

arXiv:2602.02819v3 Announce Type: replace Abstract: Membership Inference Attacks (MIAs) aim to distinguish training points (members) from unseen data (non-members), and are widely used to quantify mem

CellBRIDGE: Learning Cellular Trajectories via Interaction-Aware Alignment

SafetyDGX agent

arXiv:2605.30635v1 Announce Type: new Abstract: Inferring dynamics from population snapshots is a fundamental challenge in machine learning and biology. In scRNA-sequencing (scRNA-seq), destructive me

COFT: Counterfactual-Conformal Decoding for Fair Chain-of-Thought Reasoning in Large Language Models

SafetyDGX agent

arXiv:2605.30641v1 Announce Type: cross Abstract: Large language models (LLMs) can reveal and amplify societal biases during chain-of-thought (CoT) generation. We present COFT (Chain of Fair Thought),

ConSensus: Multi-Agent Collaboration for Multimodal Sensing

SafetyDGX agent

arXiv:2601.06453v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However,

Constrained Multi-Objective Reinforcement Learning with Max-Min Criterion

SafetyDGX agent

arXiv:2605.31388v1 Announce Type: new Abstract: Multi-Objective Reinforcement Learning (MORL) extends standard RL by optimizing policies with respect to multiple, often conflicting, objectives. While

Contextual Scalarisation Thompson Sampling for multi-objective decisions in public media

SafetyDGX agent

arXiv:2605.31291v1 Announce Type: cross Abstract: Recommender systems may operate under multiple, competing objectives. For example, audience reach, cultural values, public service mandate, and operat

Cost-aware Stopping for Bayesian Optimization

SafetyDGX agent

arXiv:2507.12453v5 Announce Type: replace Abstract: In automated machine learning, scientific discovery, and other applications of Bayesian optimization, deciding when to stop evaluating expensive bla

Cross-Modal Attention Calibration for LVLM Hallucination Mitigation

SafetyDGX agent

arXiv:2501.01926v3 Announce Type: replace-cross Abstract: Large vision-language models (LVLMs) have shown remarkable capabilities in visual-language understanding. Despite their success, LVLMs still s

Current AI is bandaids all the way down; LLMs can’t play nicely with basic tools like databases and knowledge graphs, and you never know wha…

SafetyDGX agent

Current AI is bandaids all the way down; LLMs can’t play nicely with basic tools like databases and knowledge graphs, and you never know what you are going to get. It’s long time to face facts: LLMs h

Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning

SafetyDGX agent

arXiv:2605.31174v1 Announce Type: new Abstract: Object detection in real-world scenarios remains challenging due to diverse image degradations and heterogeneous object distributions, which significant

Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents

SafetyDGX agent

arXiv:2605.31354v1 Announce Type: new Abstract: Modular visual reasoning systems increasingly rely on shared working memory for multi-step collaboration, yet the failure dynamics of intermediate state

Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

SafetyDGX agent

arXiv:2602.00521v2 Announce Type: replace Abstract: While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offer

Differentially Private Preference Data Synthesis for Large Language Model Alignment

SafetyDGX agent

arXiv:2605.30808v1 Announce Type: cross Abstract: Preference alignment is a crucial post-training step for large language models (LLMs) to ensure their outputs align with human values. However, post-t

DISCO: Mitigating Bias in Deep Learning with Conditional Distance Correlation

SafetyDGX agent

arXiv:2506.11653v3 Announce Type: replace-cross Abstract: Dataset bias often leads deep learning models to exploit spurious correlations instead of task-relevant signals. We introduce the Standard Ant

Distilling LLM Feedback for Lean Theorem Proving

SafetyDGX agent

arXiv:2605.30861v1 Announce Type: new Abstract: Post-training for reasoning models typically combines supervised fine-tuning with reinforcement learning from verifiable rewards, most commonly with GRP

DiTTo: Scalable Order-aware All-in-One Image Restoration Agent

SafetyDGX agent

arXiv:2605.30915v1 Announce Type: new Abstract: Real-world images rarely suffer from a single degradation, and the order in which degradations are removed substantially affects the final restoration q

DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs

SafetyDGX agent

arXiv:2605.31432v1 Announce Type: cross Abstract: Simultaneous speech-to-text translation (SimulST) generates translations while speech is still unfolding, requiring a streaming policy that decides wh

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization

SafetyDGX agent

arXiv:2605.31455v1 Announce Type: cross Abstract: Large language models are increasingly deployed in multi-turn interactive settings where users or environments can iteratively provide lightweight fee

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models

SafetyDGX agent

arXiv:2509.24319v4 Announce Type: replace-cross Abstract: Large language models can express values in two main ways: (1) intrinsic expression, reflecting the model's inherent values learned during tra

EchoRL: Reinforcement Learning via Rollout Echoing

SafetyDGX agent

arXiv:2605.31228v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards is an effective route for post-training to strengthen the reasoning capability of large language models

Efficient and Uncertainty-Aware Diffusion Framework for Offline-to-Online Reinforcement Learning

SafetyDGX agent

arXiv:2605.30776v1 Announce Type: new Abstract: Offline-to-Online Reinforcement Learning (O2O-RL) leverages an offline, pre-trained policy to minimize costly online interactions. Although data-efficie

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation

SafetyDGX agent

arXiv:2605.30484v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing ac

Elon was right that what OpenAI did in its for profit turn was shameful. But what he is doing with the SpaceX, at the likely expense of many…

SafetyDGX agent

Elon was right that what OpenAI did in its for profit turn was shameful. But what he is doing with the SpaceX, at the likely expense of many people’s retirement funds, is just as shameful. SpaceX bein

Enhancing Regime Shift Detection Using Unstructured Data: A Study on the Treasury Market

SafetyDGX agent

arXiv:2605.30363v1 Announce Type: cross Abstract: Regime shifts in financial markets reorganise the joint dynamics of asset prices and macro variables, breaking any single-regime calibration. They are

Entropic Projection Alignment: Estimating, Explaining, and Improving Model Performance Under Distribution Shift

SafetyDGX agent

arXiv:2605.31250v1 Announce Type: cross Abstract: We propose a unified framework for addressing three key challenges of distribution shift: (1) estimating a model's performance on an unlabeled target

Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation

SafetyDGX agent

arXiv:2605.31266v1 Announce Type: cross Abstract: The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I met

Equivariant Latent Alignment via Flow Matching under Group Symmetries

SafetyDGX agent

arXiv:2605.30705v1 Announce Type: new Abstract: Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency. In parallel, equiva

Extending the UXR Point of View Pyramid: A Generative AI-Augmented Methodology for Human-Centred AI Systems

SafetyDGX agent

arXiv:2605.31143v1 Announce Type: cross Abstract: Rising household debt and cost-of-living pressures in the United Kingdom have intensified the role of AI-driven financial technologies in mediating cr

Fair Decisions from Calibrated Scores: Achieving Optimal Classification While Satisfying Sufficiency

SafetyDGX agent

arXiv:2602.07285v2 Announce Type: replace Abstract: Binary classification based on predicted probabilities (scores) is a fundamental task in supervised machine learning. While thresholding scores is B

Feat2Go: Visual Feature-Grounded Value Estimation for Embodied Reinforcement Learning

SafetyDGX agent

arXiv:2605.30795v1 Announce Type: new Abstract: Reinforcement learning is a promising approach for improving the capabilities of vision-language-action (VLA) models while avoiding the heavy data requi

← Previous
1…133134135136137…242
Next →