AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
26 May 2026

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

SafetyDGX agent

arXiv:2605.24960v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated und

Is Decentralized AI Governable? From Regulative Policy to Constitutive Protocol

SafetyDGX agent

arXiv:2605.24538v1 Announce Type: cross Abstract: Every major framework for governing artificial intelligence presupposes an identifiable entity -- a developer, deployer, or operator -- who can be hel

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2509.13608v2 Announce Type: replace Abstract: As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Align

IsaacIPC: Coupling High-Fidelity Simulation and Realistic Rendering for Contact-Rich Robotic Systems

SafetyDGX agent

arXiv:2605.24339v1 Announce Type: new Abstract: We present IsaacIPC, a robotic simulation framework that couples GPU accelerated incremental potential contact (IPC) with IsaacSim/Lab. IsaacIPC maps si

Iterative Feature Space Optimization through Incremental Adaptive Evaluation

SafetyDGX agent

arXiv:2501.14889v2 Announce Type: replace Abstract: Iterative feature space optimization involves systematically evaluating and adjusting the feature space to improve downstream task performance. Howe

Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation

SafetyDGX agent

arXiv:2605.24041v1 Announce Type: cross Abstract: Neural operators serve as fast, data-driven surrogates for scientific modeling but typically rely on a monolithic, single-pass inference procedure tha

It's a shame the Pope didn't ask Chris Olah what happened to his plan to give up to 10% of Anthropic to the authors of the work they train o…

SafetyDGX agent

It's a shame the Pope didn't ask Chris Olah what happened to his plan to give up to 10% of Anthropic to the authors of the work they train on. (Spoiler: it never happened, and the authors on whose wor

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

SafetyDGX agent

arXiv:2605.23997v1 Announce Type: cross Abstract: Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models

SafetyDGX agent

arXiv:2605.24550v1 Announce Type: new Abstract: Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs), but it can weaken safety-alignment under harmful fine-tuning at

Joint Optimization of Training and Inference in Federated Edge Learning via Constrained Multi-Objective Deep Reinforcement Learning

SafetyDGX agent

arXiv:2605.25916v1 Announce Type: new Abstract: Federated edge learning (FEEL) has recently emerged as a promising paradigm for achieving edge intelligence (EI) via enabling collaborative model traini

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

SafetyDGX agent

arXiv:2605.24414v1 Announce Type: new Abstract: We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe

KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable Provenance and Hierarchical Policy Composition

SafetyDGX agent

arXiv:2605.25376v1 Announce Type: cross Abstract: Observability tells operators when an agent is slow. KYA tells operators when an agent is wrong, drifting, leaking, or quietly going rogue. We present

Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime

SafetyDGX agent

arXiv:2605.25275v1 Announce Type: new Abstract: The Neural Tangent Kernel (NTK) framework explains optimization in over-parameterized neural networks via approximately linearized dynamics, yielding ex

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation

SafetyDGX agent

arXiv:2605.25036v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) extend large language models with visual understanding, but remain vulnerable to hallucination, where outputs are

LAPLEX: The FFT of Learnable Laplace Kernels

SafetyDGX agent

arXiv:2605.24584v1 Announce Type: cross Abstract: Fast linear algebra in deep learning usually comes with a choice: fixed geometry and exact computation, as in the Fourier transform, or adaptive geome

Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

SafetyDGX agent

arXiv:2605.25740v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However,

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition

SafetyDGX agent

arXiv:2605.24005v1 Announce Type: new Abstract: The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous

Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models

SafetyDGX agent

arXiv:2603.29123v2 Announce Type: replace Abstract: The next-token prediction (NTP) objective trains language models to predict a single token at each step, even though many continuations can express

Learning High-Frequency Continuous Action Chunks in Latent Space

SafetyDGX agent

arXiv:2605.24931v1 Announce Type: new Abstract: Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal con

Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning

SafetyDGX agent

arXiv:2605.26012v1 Announce Type: cross Abstract: Deep reinforcement learning (RL) agents commonly rely on high-dimensional neural representations, despite growing evidence that task-relevant value an

Learning to Route Languages for Multilingual Policy Optimization

SafetyDGX agent

arXiv:2605.25360v1 Announce Type: new Abstract: Large language models~(LLMs) are trained on heterogeneous multilingual corpora, yet existing policy optimization methods often implicitly restrict each

LipoAgent: Coordinating Fine-Tuned LLM Agents for Safer Lipid Design

SafetyDGX agent

arXiv:2605.25250v1 Announce Type: new Abstract: Lipid nanoparticles (LNPs) are among the most clinically mature platforms for nucleic acid delivery, yet designing lipids that are both effective and bi

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models

SafetyDGX agent

arXiv:2605.25179v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cos

lol. OpenAI as the WeWork of AI. Literally called it, in those exact words, @CNBC w @carlquintanilla, 2024. Now even SoftBank is worried.

SafetyDGX agent

lol. OpenAI as the WeWork of AI. Literally called it, in those exact words, @CNBC w @carlquintanilla, 2024. Now even SoftBank is worried. SoftBank's own executives think Sam Altman is scamming their C

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_!

SafetyDGX agent

Looking forward to speaking at BloombergTech next week with @shiringhaffary to discuss AI safety and our research progress at @LawZero_! Is AI development progressing too quickly? @business' @shiringh

Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence

SafetyDGX agent

arXiv:2605.23952v1 Announce Type: new Abstract: Artificial agents now generate behavior rich enough to invite trust, surprise, and concern, yet our evaluation tools still privilege capability scores o

MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models

SafetyDGX agent

arXiv:2605.26004v1 Announce Type: cross Abstract: Instruction tuning of large vision-language models (LVLMs) increasingly depends on massive multimodal corpora, yet these datasets contain samples with

MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games

SafetyDGX agent

arXiv:2605.24139v1 Announce Type: new Abstract: Imperfect-information games (IIGs) are challenging, as players must make decisions without fully observing the true game state. While AlphaZero has achi

MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling

SafetyDGX agent

arXiv:2602.17658v2 Announce Type: replace-cross Abstract: Reward modeling is central to alignment pipelines such as RLHF, RLAIF, and PPO-based policy optimization, yet its reliability is constrained b

MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models

SafetyDGX agent

arXiv:2605.25342v1 Announce Type: new Abstract: Aligning large language models (LLMs) with diverse and multifaceted user preferences is a fundamental challenge in personalized AI systems. Existing mul

Measuring the Depth of LLM Unlearning via Activation Patching

SafetyDGX agent

arXiv:2605.24614v1 Announce Type: cross Abstract: Large language model (LLM) unlearning has emerged as a crucial post-hoc mechanism for privacy protection and AI safety, yet auditing whether target kn

MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

SafetyDGX agent

arXiv:2602.02474v2 Announce Type: replace-cross Abstract: Most Large Language Model (LLM) agent memory systems rely on a small set of static, hand-designed operations for extracting memory. These fixe

Metacognition Should Be the Scientific Framework for Bounded and Effective Self-Governance in Generative AI

SafetyDGX agent

arXiv:2605.23981v1 Announce Type: cross Abstract: Generative AI research increasingly confronts a shared problem: systems must sustain yet govern their own generative activity when uncertainty is high

Micro-Swarm Locomotion Optimization in Dynamic Flow using Multi-Objective Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.25025v1 Announce Type: new Abstract: Coordinating micro-robotic swarms in physiologically realistic, time-dependent fluid environments remains an unsolved challenge for biomedical and envir

Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation

SafetyDGX agent

arXiv:2512.16189v3 Announce Type: replace Abstract: In healthcare, it is essential for any LLM-generated output to be reliable and accurate, particularly in cases involving decision-making and patient

MMUEChange: A Generalized LLM Agent Framework for Intelligent Multi-Modal Urban Environment Change Analysis

SafetyDGX agent

arXiv:2601.05483v2 Announce Type: replace Abstract: Understanding urban environment change is essential for sustainable development. However, current approaches, particularly remote sensing change det

Motion-Compensated Weight Compression

SafetyDGX agent

arXiv:2605.24754v1 Announce Type: cross Abstract: Neural network weights are increasingly a bottleneck for deployment, yet most compression pipelines treat layers independently and overlook cross-laye

MuGen: Multi-Skill Generative Locomotion Controller for Humanoid Robots

SafetyDGX agent

arXiv:2605.24592v1 Announce Type: new Abstract: This paper presents MuGen, a data-driven framework for learning and deploying multi-skill locomotion on humanoid robots. MuGen enables a robot to perfor

Multi-Agent Coordination Adaptation via Structure-Guided Orchestration

SafetyDGX agent

arXiv:2605.25746v1 Announce Type: cross Abstract: As large language model (LLM)-based multi-agent systems scale to handle increasingly complex tasks, balancing structural stability and dynamic adaptab

Multi-Agent Systems are Mixtures of Experts: Who Becomes an Influencer?

SafetyDGX agent

arXiv:2605.25929v1 Announce Type: cross Abstract: The effectiveness of multi-agent LLM deliberation depends not only on the agents' individual predictions, but also on how they communicate and collabo

Multi-Alignment Contrastive Learning for Enzyme--Reaction Retrieval

SafetyDGX agent

arXiv:2512.08508v2 Announce Type: replace-cross Abstract: Identifying enzymes that catalyze target biochemical reactions is a key step in computational enzyme discovery and biocatalyst design. Recent

Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval

SafetyDGX agent

arXiv:2209.11572v3 Announce Type: replace-cross Abstract: As an increasingly popular task in multimedia information retrieval, video moment retrieval (VMR) aims to localize the target moment from an u

Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning

SafetyDGX agent

arXiv:2605.25210v1 Announce Type: cross Abstract: Diffusion models are increasingly used as powerful conditional generators, yet real deployments often involve multiple target distributions arising fr

Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation

SafetyDGX agent

arXiv:2605.25220v1 Announce Type: cross Abstract: High-fidelity 3D Gaussian head avatar generation is critical for applications such as AR/VR, telepresence, and digital humans. Existing methods depend

Multicalibration Boosting: Theory, Convergence, and Transferability

SafetyDGX agent

arXiv:2605.24364v1 Announce Type: cross Abstract: Multicalibration extends classical calibration by requiring predictions to be unbiased over a rich collection of functions, encompassing both predicti

Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation

SafetyDGX agent

arXiv:2605.23961v1 Announce Type: cross Abstract: The design of RNA molecules that interact with specific proteins is a critical challenge in experimental and computational biology. Despite recent pro

Multimodal Functional Maximum Correlation for Emotion Recognition

SafetyDGX agent

arXiv:2512.23076v2 Announce Type: replace-cross Abstract: Emotional states manifest as coordinated yet heterogeneous physiological responses across central and autonomic systems, posing a fundamental

MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection

SafetyDGX agent

arXiv:2505.23803v2 Announce Type: replace-cross Abstract: Phishing email detection faces significant challenges due to evolving adversarial tactics and heterogeneous attack patterns. Traditional appro

Music Transcription with (Almost) No Supervision

SafetyDGX agent

arXiv:2605.24193v1 Announce Type: cross Abstract: Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty

NeuralTouch: Neural Descriptors for Precise Sim-to-Real Tactile Robot Control

SafetyDGX agent

arXiv:2510.20390v2 Announce Type: replace Abstract: Grasping accuracy is a critical prerequisite for precise object manipulation, often requiring careful alignment between the robot hand and object. N

Neuro-Inspired Inverse Learning for Planning and Control

SafetyDGX agent

arXiv:2605.24152v1 Announce Type: new Abstract: We present a neuro-inspired framework for embodied planning and control. Building on three principles that enable fast and highly effective goal-directe

Not All Transitions Matter: Evidence from PPO

SafetyDGX agent

arXiv:2605.24071v1 Announce Type: cross Abstract: Training a reinforcement learning agent on-policy means collecting fresh experience at every update, and that experience comes with a hidden problem.

Not only where, But when: Temporal Scheduling for RLVR

SafetyDGX agent

arXiv:2605.25381v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a core technique for post-training of Large Language Models (LLMs). While policy optimi

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation

SafetyDGX agent

arXiv:2605.25829v1 Announce Type: cross Abstract: Recent vision-language-action (VLA) models and world action models (WAMs) advance robotic manipulation by enriching intermediate representations with

oh. my. god. could this word cloud diagram be … conscious?

SafetyDGX agent

Gary Marcus, a prominent AI researcher and critic, questions whether a word cloud diagram could possess consciousness, likely engaging in ironic commentary on overclaimed AI capabilities or consciousn

OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization

SafetyDGX agent

arXiv:2602.10635v2 Announce Type: replace Abstract: Socially intelligent AI systems must entail reasoning across diverse human behavioral tasks, and generalization to new contexts. However, AI has yet

On Reliability of Efficient Membership Inference Vulnerability Evaluation

SafetyDGX agent

arXiv:2605.25819v1 Announce Type: new Abstract: Membership inference attacks (MIAs) are popular methods for empirically assessing the leakage of sensitive information in the training data through mode

On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits

SafetyDGX agent

arXiv:2605.25789v1 Announce Type: cross Abstract: We study a stochastic multi-armed bandit problem where an agent is granted a free exploration budget before regret accumulates, a setting not captured

On the Stability and Realizability of Recurrent Polynomial Surrogate Ternary Logic Gate Networks

SafetyDGX agent

arXiv:2605.24649v1 Announce Type: cross Abstract: Recurrent Neural Networks (RNNs) can learn to predict Signal Temporal Logic (STL) verdicts online from partial trajectories, but deploying them as run

One-Step Bellman Alignment Enables Provably Efficient Transfer in Online RL

SafetyDGX agent

arXiv:2601.21924v2 Announce Type: replace Abstract: We study online transfer reinforcement learning (RL) in episodic Markov decision processes, where experience from related source tasks is available

← Previous
1…117118119120121…214
Next →