AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,809 results
19 May 2026

White-Box Sensitivity Auditing with Steering Vectors

SafetyDGX agent

arXiv:2601.16398v2 Announce Type: replace-cross Abstract: Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators. Current audits of

Why Do Safety Guardrails Degrade Across Languages?

SafetyDGX agent

arXiv:2605.17173v1 Announce Type: cross Abstract: Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds

World Model-Enabled Causal Digital Twins for Semantic Communications in Physical AI Systems

SafetyDGX agent

arXiv:2605.16547v1 Announce Type: new Abstract: Semantic communication has emerged as a promising paradigm for enabling goal-oriented networking. However, most existing semantic communication solution


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Zero-Shot Textual Explanations via Translating Decision-Critical Features

SafetyDGX agent

arXiv:2512.07245v2 Announce Type: replace Abstract: Textual explanations make image classifier decisions transparent by describing the prediction rationale in natural language. Large vision-language m

ZeroSiam: An Efficient Asymmetry for Test-Time Entropy Optimization without Collapse

SafetyDGX agent

arXiv:2509.23183v3 Announce Type: replace Abstract: Test-time entropy minimization helps adapt a model to novel environments and incentivize its reasoning capability, unleashing the model's potential

18 May 2026

A Differentiable Measure of Algebraic Complexity: Provably Exact Discovery of Group Structures

SafetyDGX agent

arXiv:2511.23152v3 Announce Type: replace Abstract: Discovering discrete algebraic rules from data is a fundamental challenge in machine learning. We formalize this problem through Cayley-table comple

A Generative AI Framework for Intelligent Utility Billing CO 2 Analytics and Sustainable Resource Optimisation

SafetyDGX agent

arXiv:2605.16250v1 Announce Type: cross Abstract: Distribution utilities are now expected to deliver bills that customers can actually read attach a defensible carbon number to every kWh sold and sche

A Split-Client Approach to Second-Order Optimization

SafetyDGX agent

arXiv:2510.15714v3 Announce Type: replace-cross Abstract: Second-order optimization methods offer superior convergence rates but are often bottlenecked by the wall-clock cost of Hessian computation an

a very procedural end. we will never know what the world might be like had OpenAI been forced to fully follow its original mission.

SafetyDGX agent

a very procedural end. we will never know what the world might be like had OpenAI been forced to fully follow its original mission. Breaking News: A jury rejected Elon Musk’s lawsuit accusing OpenAI o

Accelerated Gradient Descent for Faster Convergence with Minimal Overhead

SafetyDGX agent

arXiv:2605.16017v1 Announce Type: new Abstract: In this paper, we present CT-AGD (Curvature-Tuned Accelerated Gradient Descent), an optimization method for non-convex optimization problems in deep lea

ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment

SafetyDGX agent

arXiv:2505.19241v2 Announce Type: replace-cross Abstract: The recent success in using human preferences to align large language models (LLMs) has significantly improved their performance in various do

Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-Making

SafetyDGX agent

arXiv:2605.16054v1 Announce Type: cross Abstract: Recent work has framed decision-making as a sequence modeling problem using generative models such as diffusion models. Although promising, these appr

Adaptive Outer-Loop Control of Quadrotors via Reinforcement Learning

SafetyDGX agent

arXiv:2605.16015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (DRL) for quadrotor flight control typically relies on Domain Randomization (DR) for sim-to-real transfer, resulting in ov

AI Consciousness and Existential Risk

SafetyDGX agent

arXiv:2511.19115v2 Announce Type: replace Abstract: In AI, the existential risk denotes the hypothetical threat posed by an artificial system that would possess both the capability and the objective,

AI-Mediated Communication Can Steer Collective Opinion

SafetyDGX agent

arXiv:2605.16245v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) is increasingly integrated into the online platforms where humans exchange opinions; large language models (LL

Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time

SafetyDGX agent

arXiv:2605.15220v1 Announce Type: cross Abstract: Data mixing decides how to combine different sources or types of data and is a consequential problem throughout language model training. In pretrainin

“Americans are now more comfortable living near a nuclear power plant than an AI data center” -@RachelBitecofer The AI oligarchs took a winn…

SafetyDGX agent

“Americans are now more comfortable living near a nuclear power plant than an AI data center” -@RachelBitecofer The AI oligarchs took a winning hand, and with a mixture cigarette-industry level greed

An Algebraic Exposition of the Theory of Dyadic Morality

SafetyDGX agent

arXiv:2605.16153v1 Announce Type: new Abstract: This paper provides an algebraic exposition of the theory of dyadic morality (TDM), a psychological model of moral judgment grounded in a simple two-nod

An Introduction to Deep Reinforcement and Imitation Learning

SafetyDGX agent

arXiv:2512.08052v3 Announce Type: replace-cross Abstract: Embodied agents, such as robots and virtual characters, must continuously select actions to execute tasks effectively, solving complex sequent

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models

SafetyDGX agent

arXiv:2605.15687v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) may memorize sensitive cross-modal information during pretraining, making machine unlearning (MU) crucial. Ex

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs

SafetyDGX agent

arXiv:2605.15565v1 Announce Type: cross Abstract: Reinforcement learning (RL) is increasingly used to improve the reasoning, coding, and tool-use capabilities of large language models, but agentic RL

Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes

SafetyDGX agent

arXiv:2602.01295v3 Announce Type: replace Abstract: We investigate episodic Markov Decision Processes with heavy-tailed losses (HTMDPs). Existing approaches for HTMDPs are conservative in stochastic e

Beyond Objective-Based Improvement: Stationarity-Aware Expected Improvement for Bayesian Optimization

SafetyDGX agent

arXiv:2601.21357v2 Announce Type: replace Abstract: Bayesian Optimization (BO) is a principled framework for optimizing expensive black-box functions, with Expected Improvement (EI) among its most wid

Beyond Performance Disparities: A Three-Level Audit of Representational Harm in CelebA

SafetyDGX agent

arXiv:2605.15312v1 Announce Type: cross Abstract: Large-scale facial datasets like CelebA are widely used in computer vision, yet the cultural biases embedded in their labels remain underexplored. Fai

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling

SafetyDGX agent

arXiv:2507.01679v3 Announce Type: replace-cross Abstract: Existing LLMs-post-training techniques are broadly categorized into supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT). Each par

Can we all agree that Dario played the “Ooh! AI scary!” card one time too many?

SafetyDGX agent

Can we all agree that Dario played the “Ooh! AI scary!” card one time too many? “Americans are now more comfortable living near a nuclear power plant than an AI data center” -@RachelBitecofer The AI o

Constrained MPC-Based Motion Planning for Morphing Quadrotors in Ultra-Narrow Passages under Limited Perception

SafetyDGX agent

arXiv:2605.15999v1 Announce Type: new Abstract: This paper introduces a motion planning framework to plan morphology and trajectory for morphing quadrotors under extremely constrained environments. We

Controllable Molecular Generative Foundation Models

SafetyDGX agent

arXiv:2605.15354v1 Announce Type: new Abstract: Despite the success of foundation models in language and vision, molecular graph generation still lacks a unified framework for heterogeneous design tas

CTF4Nuclear: Common Task Framework for Nuclear Fission and Fusion Models

SafetyDGX agent

arXiv:2605.15549v1 Announce Type: cross Abstract: The demand for clean energy is ever increasing, with new nuclear technologies presenting a complementary solution to renewable energies. However, desi

DebiasRAG: A Tuning-Free Path to Fair Generation in Large Language Models through Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2605.16113v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved unprecedented success due to their exceptional generative capabilities. However, because they depend on kno

Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation

SafetyDGX agent

arXiv:2605.15942v1 Announce Type: cross Abstract: Open-vocabulary segmentation models often struggle to generalize to unseen combinations of object categories and attributes, because fine-grained desc

Deep Double Q-learning

SafetyDGX agent

arXiv:2507.00275v2 Announce Type: replace-cross Abstract: Double Q-learning is a classical control algorithm that mitigates the maximization bias of Q-learning. To do so, it explicitly trains two inde

DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation

SafetyDGX agent

arXiv:2605.15532v1 Announce Type: cross Abstract: Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically

Designing Datacenter Power Delivery Hierarchies for the AI Era

SafetyDGX agent

arXiv:2605.16255v1 Announce Type: cross Abstract: Demand for AI accelerators is rapidly increasing rack power density, with projections approaching 1MW per deployment by 2027. This poses a major chall

Detecting Heel Strike and toe off Events Using Kinematic Methods and LSTM Models

SafetyDGX agent

arXiv:2503.00794v2 Announce Type: replace Abstract: Accurate gait event detection is crucial for gait analysis, rehabilitation, and assistive technology, particularly in exoskeleton control, where pre

Differentially Private Motif-Preserving Multi-modal Hashing

SafetyDGX agent

arXiv:2605.15460v1 Announce Type: cross Abstract: Cross-modal hashing enables efficient retrieval by encoding images and text into compact binary codes. State-of-the-art methods rely on semantic simil

Diffusion Policy for Coordinated Control of a Nonholonomic Mobile Base and Dual Arms in Door Opening and Passing

SafetyDGX agent

arXiv:2605.15352v1 Announce Type: new Abstract: Opening heavy, self closing doors, especially those that require pulling remains a long standing challenge in robotics. Humans naturally employ both arm

DiffVAS: Diffusion-Guided Visual Active Search in Partially Observable Environments

SafetyDGX agent

arXiv:2605.15519v1 Announce Type: cross Abstract: Visual active search (VAS) has been introduced as a modeling framework that leverages visual cues to direct aerial (e.g., UAV-based) exploration and p

Discretizing Group-Convolutional Neural Networks for 3D Geometry in Feature Space

SafetyDGX agent

arXiv:2605.15368v1 Announce Type: new Abstract: Group-convolutional neural networks (GCNNs) are among the most important methods for introducing symmetry as an inductive bias in deep learning: In each

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

SafetyDGX agent

arXiv:2605.15855v1 Announce Type: new Abstract: Despite strong image-generation performance, diffusion models' reconstruction objectives limit alignment with human preferences. RL enables such alignme

DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

SafetyDGX agent

arXiv:2511.19399v3 Announce Type: replace-cross Abstract: Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are tr

Drawback of Enforcing Equivariance and its Compensation via the Lens of Expressive Power

SafetyDGX agent

arXiv:2512.09673v3 Announce Type: replace-cross Abstract: Equivariant neural networks encode the intrinsic symmetry of data as an inductive bias, which has achieved impressive performance in wide doma

Driving Through the Network: Performance and Workload Under Latency and Video Impairment

SafetyDGX agent

arXiv:2605.15952v1 Announce Type: cross Abstract: Teleoperation promises to extend the operational envelope of automated vehicles, yet it critically depends on network latency and video quality. We re

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts

SafetyDGX agent

arXiv:2605.15422v1 Announce Type: new Abstract: Modern RL post-training methods such as GRPO and DAPO train on N response sequences of R tokens sampled from a shared prompt of P tokens, but standard F

DualReg: Dual-Space Filtering and Reinforcement for Rigid Registration

SafetyDGX agent

arXiv:2508.17034v2 Announce Type: cross Abstract: Noisy, partially overlapping data and the need for real-time processing pose major challenges for rigid registration. Considering that feature-based m

Dynamic Plasma Shape Control with Arbitrary Sensor Subsets

SafetyDGX agent

arXiv:2605.15935v1 Announce Type: new Abstract: Plasma shape control in tokamaks requires a real-time controller that tracks dynamically changing shape targets while tolerating diagnostic failures. Cl

Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling

SafetyDGX agent

arXiv:2509.23352v3 Announce Type: replace-cross Abstract: The integration of Reinforcement Learning (RL) into flow matching models for text-to-image (T2I) generation has driven substantial advances in

Efficiently Solving Mixed-Hierarchy Games with Quasi-Policy Approximations

SafetyDGX agent

arXiv:2602.01568v2 Announce Type: replace-cross Abstract: Multi-robot coordination often exhibits hierarchical structure, with some robots' decisions depending on the planned behaviors of others. Whil

EgoExo-WM: Unlocking Exo Video for Ego World Models

SafetyDGX agent

arXiv:2605.15477v1 Announce Type: new Abstract: Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited avail

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices

SafetyDGX agent

arXiv:2605.15684v1 Announce Type: new Abstract: The Diffusion Transformer (DiT) architecture is the state-of-the-art paradigm for high-fidelity image generation, underpinning models like Stable Diffus

Embedding-perturbed Exploration Preference Optimization for Flow Models

SafetyDGX agent

arXiv:2605.15803v1 Announce Type: new Abstract: Recent advancements have established Reinforcement Learning (RL) as a pivotal paradigm for aligning generative models with human intent. However, group-

Embracing Biased Transition Matrices for Complementary-Label Learning with Many Classes

SafetyDGX agent

arXiv:2605.15586v1 Announce Type: cross Abstract: Complementary-label learning (CLL) is a weakly supervised paradigm where instances are labeled with classes they do not belong to. Despite a decade of

EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy

SafetyDGX agent

arXiv:2605.15711v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to backdoor attacks. Exi

Eskwai for Students: Generative AI Assistant for Legal Education in Ghana

SafetyDGX agent

arXiv:2605.15380v1 Announce Type: new Abstract: Recent advances in generative AI have shown their potential to be leveraged for legal education. Yet, work on the development and deployment of such sys

Explainable AI Isn't Enough! Rethinking Algorithmic Contestability

SafetyDGX agent

arXiv:2605.16041v1 Announce Type: cross Abstract: Machine learning systems increasingly make life-changing decisions about individuals, such as loan approvals, hiring, and cheating detection, raising

f-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data

SafetyDGX agent

arXiv:2605.15417v1 Announce Type: cross Abstract: In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities is an effective, low v

FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures

SafetyDGX agent

arXiv:2604.05966v2 Announce Type: replace Abstract: Financial reporting systems increasingly leverage Large Language Models (LLMs) to extract and summarize corporate disclosures. However, most existin

FLASH: Efficient Visuomotor Policy via Sparse Sampling

SafetyDGX agent

arXiv:2605.15492v1 Announce Type: cross Abstract: Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative d

FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control

SafetyDGX agent

arXiv:2604.04539v2 Announce Type: replace Abstract: Reinforcement learning (RL) is a core approach for robot control when expert demonstrations are unavailable. On-policy methods such as Proximal Poli

FlipAttack: Jailbreak LLMs via Flipping

SafetyDGX agent

arXiv:2410.02832v2 Announce Type: replace-cross Abstract: This paper proposes a simple yet effective jailbreak attack named FlipAttack against black-box LLMs. First, from the autoregressive nature, we

← Previous
1…136137138139140…214
Next →