AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,708 results
12 May 2026

Users as Annotators: LLM Preference Learning from Comparison Mode

SafetyDGX agent

arXiv:2510.13830v2 Announce Type: replace-cross Abstract: Pairwise preference data have played an important role in the alignment of large language models (LLMs). Each sample of such data consists of

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning

SafetyDGX agent

arXiv:2605.10172v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a per

Value-Decomposed Reinforcement Learning Framework for Taxiway Routing with Hierarchical Conflict-Aware Observations

SafetyDGX agent

arXiv:2605.08754v1 Announce Type: new Abstract: Taxiway routing and on-surface conflict avoidance are coupled safety-critical decision problems in airport surface operations. Existing planning and opt


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Variational Inference for Levy Process-Driven SDEs via Neural Tilting

SafetyDGX agent

arXiv:2605.10934v1 Announce Type: cross Abstract: Modelling extreme events and heavy-tailed phenomena is central to building reliable predictive systems in domains such as finance, climate science, an

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA

SafetyDGX agent

arXiv:2605.10850v1 Announce Type: new Abstract: Self-verification, re-invoking the same vision language model (VLM) in a fresh context to check its own generated answer, is increasingly used as a defa

Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward

SafetyDGX agent

arXiv:2605.09920v1 Announce Type: cross Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a promising post-training paradigm for Large Language Models (LLMs

VISTA: A Generative Egocentric Video Framework for Daily Assistance

SafetyDGX agent

arXiv:2605.10579v1 Announce Type: new Abstract: Training AI agents to proactively assist humans in daily activities, from routine household tasks to urgent safety situations, requires large-scale visu

Wavelet Policy: Imitation Learning in the Scale Domain with World Prior Memory

SafetyDGX agent

arXiv:2504.04991v4 Announce Type: replace Abstract: Conventional visuomotor imitation learning usually predicts future robot actions directly in the time domain. Such formulations often have limited p

'We're going to reach a point of diminishing returns from scaling. That's what actually happened. The industry knows this. But doesn't want …

SafetyDGX agent

'We're going to reach a point of diminishing returns from scaling. That's what actually happened. The industry knows this. But doesn't want to tell you.' @GaryMarcus, AI Expert, Scientist & Author at

What should post-training optimize? A test-time scaling law perspective

SafetyDGX agent

arXiv:2605.10716v1 Announce Type: new Abstract: Large language models are increasingly deployed with test-time strategies: sample N responses, score them with a reward model or verifier, and return th

What Structural Inductive Bias Helps Transformers Reason Over Knowledge Graphs? A Study with Tabula RASA

SafetyDGX agent

arXiv:2602.02834v3 Announce Type: replace-cross Abstract: What structural inductive bias helps transformers reason over knowledge graphs? Through controlled ablations of a minimal transformer modifica

When a Robot is More Capable than a Human: Learning from Constrained Demonstrators

SafetyDGX agent

arXiv:2510.09096v3 Announce Type: replace-cross Abstract: Learning from demonstrations enables experts to teach robots complex tasks using interfaces such as kinesthetic teaching, joystick control, an

When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents

SafetyDGX agent

arXiv:2605.08828v1 Announce Type: new Abstract: Large language model agents increasingly operate through environment-facing scaffolds that expose files, web pages, APIs, and logs. These observations i

When Can Digital Personas Reliably Approximate Human Survey Findings?

SafetyDGX agent

arXiv:2605.10659v1 Announce Type: cross Abstract: Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

SafetyDGX agent

arXiv:2605.08245v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate,

When More Parameters Hurt: Foundation Model Priors Amplify Worst-Client Disparity Under Extreme Federated Heterogeneity

SafetyDGX agent

arXiv:2605.08992v1 Announce Type: new Abstract: Federated learning (FL) is increasingly used to fine-tune foundation models (FMs) on distributed private data. The community largely assumes that large-

Where Do Flow Semantics Reside? A Protocol-Native Tabular Pretraining Paradigm for Encrypted Traffic Classification

SafetyDGX agent

arXiv:2603.10051v2 Announce Type: replace-cross Abstract: Self-supervised masked modeling shows promise for encrypted traffic classification by masking and reconstructing raw bytes. Yet recent work re

White Circle raises $11M to help companies secure and monitor AI model behavior

SafetyDGX agent

Artificial intelligence guardrail and monitoring startup Pumpkin Intelligence Inc., which operates as White Circle, announced today it raised 11 million in seed funding from a who’s who of AI leadersh

Why Adam Works Better with eta_1 = eta_2: The Missing Gradient Scale Invariance Principle

SafetyDGX agent

arXiv:2601.21739v2 Announce Type: replace-cross Abstract: Adam has been at the core of large-scale training for almost a decade, yet a simple empirical fact remains unaccounted for: both validation sc

Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space

SafetyDGX agent

arXiv:2605.08250v1 Announce Type: cross Abstract: Recent advances in diffusion transformers (DiTs) have enabled promising single-turn image editing capabilities. However, multi-turn editing often lead

WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records

SafetyDGX agent

arXiv:2605.09765v1 Announce Type: cross Abstract: Representation learning in electronic health records (EHR) has largely followed paradigms inherited from natural language processing, relying on seque

wow! “It is the [House] Committee’s understanding that the new board of directors at OpenAI tried to address these problems upon your return…

SafetyDGX agent

wow! “It is the [House] Committee’s understanding that the new board of directors at OpenAI tried to address these problems upon your return by creating an “audit committee to review potential conflic

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

SafetyDGX agent

arXiv:2605.05611v2 Announce Type: replace-cross Abstract: In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to spea

XQCfD: Accelerating Fast Actor-Critic Algorithms with Prior Data and Prior Policies

SafetyDGX agent

arXiv:2605.10734v1 Announce Type: new Abstract: For reinforcement learning in the real world online exploration is expensive A common practice in robotic reinforcement learning is to incorporate addit

Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers

SafetyDGX agent

arXiv:2603.25074v2 Announce Type: replace Abstract: Concept erasure serves as a vital safety mechanism for removing unwanted concepts from text-to-image (T2I) models. While extensively studied in U-Ne

11 May 2026

A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning

SafetyDGX agent

arXiv:2605.06866v1 Announce Type: new Abstract: Recent non-asymptotic analyses have substantially advanced the theory of distributional policy evaluation, but they largely concern synchronous full-sta

A Generalized Singular Value Theory for Neural Networks

SafetyDGX agent

arXiv:2605.06938v1 Announce Type: cross Abstract: Building on the abstract Generalized Singular Value Decomposition (GSVD) theory of Brown et al. [2025], we prove that most modern neural architectures

A Geometric Taxonomy of Hallucinations in LLMs

SafetyDGX agent

arXiv:2602.13224v3 Announce Type: replace Abstract: Hallucinations in deployed language models can have real consequences for downstream decisions in domains such as healthcare, legal, and financial s

A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method

SafetyDGX agent

arXiv:2602.02320v3 Announce Type: replace-cross Abstract: Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential fo

A Statistical Framework for Algorithmic Collective Action with Multiple Collectives

SafetyDGX agent

arXiv:2605.06749v1 Announce Type: cross Abstract: As learning systems increasingly shape everyday decisions, Algorithmic Collective Action (ACA), i.e., users coordinating changes to shared data to ste

A Systematic Investigation of The RL-Jailbreaker in LLMs

SafetyDGX agent

arXiv:2605.07032v1 Announce Type: cross Abstract: The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardening. Adversa

Accurate and Efficient Statistical Testing for Word Semantic Breadth

SafetyDGX agent

arXiv:2605.08048v1 Announce Type: new Abstract: Measuring the breadth of a word's meaning, or its spread across contexts, has become feasible with contextualized token embeddings. A word type can be r

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures

SafetyDGX agent

arXiv:2605.07324v1 Announce Type: cross Abstract: Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior w

Actor-Critic Algorithm for Dynamic Expectile and CVaR

SafetyDGX agent

arXiv:2605.07857v1 Announce Type: new Abstract: Optimizing dynamic risk with stochastic policies is challenging in both policy updates and value learning. The former typically requires transition pert

Actor-Critic with Active Importance Sampling

SafetyDGX agent

arXiv:2605.07094v1 Announce Type: new Abstract: This paper introduces the Active-Importance-Sampling Actor-Critic (AISAC) algorithm, an extension of the Actor-Critic framework for reducing variance in

Adaptive Subspace Projection for Generative Personalization

SafetyDGX agent

arXiv:2605.07257v1 Announce Type: new Abstract: Generative personalization often suffers from the semantic collapsing problem (SCP), where a learned personalized concept overpowers the rest of the tex

Agentic Coding Needs Proactivity, Not Just Autonomy

SafetyDGX agent

arXiv:2605.06717v1 Announce Type: cross Abstract: Coding agents are rapidly changing the landscape of software development, moving from inline completion to autonomous systems that edit repositories,

Ah yes, chatGPT, tell me more about the 'colony mind', 'foraging patrol' and 'nost building' behaviors of the human liver. For anyone predic…

SafetyDGX agent

Ah yes, chatGPT, tell me more about the 'colony mind', 'foraging patrol' and 'nost building' behaviors of the human liver. For anyone predicting near-term physician replacement by LLM-based AI, please

AI existential crisis for software engineers is to go live in the woods and read poetry. AI existential crisis for creatives is to make thin…

SafetyDGX agent

AI existential crisis for software engineers is to go live in the woods and read poetry. AI existential crisis for creatives is to make things until 4am every day because you can't stop now. is this w

Am old enough to remember when @GeoffreyHinton told me I was stupid for saying that LLMs regurgitate training data. He was wrong. LLM regurg…

SafetyDGX agent

Am old enough to remember when @GeoffreyHinton told me I was stupid for saying that LLMs regurgitate training data. He was wrong. LLM regurgitation is now one of the best-established findings in the f

Anisotropic Modality Align

SafetyDGX agent

arXiv:2605.07825v1 Announce Type: cross Abstract: Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show that the s

APEX: Assumption-free Projection-based Embedding eXamination Metric for Image Quality Assessment

SafetyDGX agent

arXiv:2605.07786v1 Announce Type: cross Abstract: As generative models achieve unprecedented visual quality, the gold standard for image evaluation remains traditional feature-distribution metrics (e.

Approximation-Free Differentiable Oblique Decision Trees

SafetyDGX agent

arXiv:2605.07837v1 Announce Type: cross Abstract: Decision Trees (DTs) are widely used in safety-critical domains such as medical diagnosis, valued for their interpretability and effectiveness on tabu

ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule

SafetyDGX agent

arXiv:2601.18681v2 Announce Type: replace-cross Abstract: We consider time discretization for score-based diffusion models to generate samples from a learned reverse-time dynamic on a finite grid. Uni

ASPECT: Node-Level Adaptive Spectral Fusion for Graph Contrastive Learning

SafetyDGX agent

arXiv:2604.01878v2 Announce Type: replace-cross Abstract: Spectral graph contrastive learning often constructs low- and high-frequency views to capture complementary graph signals, but these views are

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level

SafetyDGX agent

arXiv:2605.06387v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher feedback and often outperforms off-policy disti

BEAVER: An Efficient Deterministic LLM Verifier

SafetyDGX agent

arXiv:2512.05439v2 Announce Type: replace Abstract: As large language models (LLMs) transition from research prototypes to production systems, practitioners often need reliable methods to verify model

Bellman Calibration for V-Learning in Offline Reinforcement Learning

SafetyDGX agent

arXiv:2512.23694v2 Announce Type: replace-cross Abstract: Reliable long-horizon value prediction is difficult in offline reinforcement learning because fitted value methods combine bootstrapping, func

Better Protein Function Prediction by Modeling Survivorship Bias

SafetyDGX agent

arXiv:2605.06879v1 Announce Type: new Abstract: Protein sequence data from nature exhibits survivorship bias: we only observe data from those organisms that survive and reproduce, while non-functional

Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

SafetyDGX agent

arXiv:2605.07806v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved fro

Beyond 'I cannot fulfill this request': Alleviating Rigid Rejection in LLMs via Label Enhancement

SafetyDGX agent

arXiv:2605.07883v1 Announce Type: new Abstract: Large Language Models (LLMs) rely on safety alignment to obey safe requests while refusing harmful ones. However, traditional refusal mechanisms often l

Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph

SafetyDGX agent

arXiv:2605.08037v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) aligns language models using pairwise preference comparisons, offering a simple and effective alternative to Rein

Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies

SafetyDGX agent

arXiv:2602.23811v4 Announce Type: replace-cross Abstract: We investigate the theoretical aspects of offline reinforcement learning (RL) under general function approximation. While prior works (e.g., X

Bias and Uncertainty in LLM-as-a-Judge Estimation

SafetyDGX agent

arXiv:2605.06939v1 Announce Type: new Abstract: LLM-as-a-Judge evaluation has become a standard tool for assessing base model performance. However, characterizing performance via the naive estimator,

🚨BREAKING: Ilya just confirmed under oath what he saw: Sam lying. And he confirmed that he thought it was appropriate to fire Altman for it…

SafetyDGX agent

🚨BREAKING: Ilya just confirmed under oath what he saw: Sam lying. And he confirmed that he thought it was appropriate to fire Altman for it. And that he had been concerned for about Sam for a long tim

CalexNet: Soft Cascade-Aligned Training and Calibration for Lightweight Early-Exit Branches

SafetyDGX agent

arXiv:2509.08318v2 Announce Type: replace Abstract: Early-exit cascades over a frozen convolutional backbone enable adaptive inference but suffer from three sources of train-inference mismatch: branch

Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents

SafetyDGX agent

arXiv:2601.21699v2 Announce Type: replace Abstract: Multi-turn reasoning agents solve complex questions by decomposing them into intermediate retrieval or tool-use steps, for accumulating supporting e

Causal EpiNets: Precision-corrected Bounds on Individual Treatment Effects using Epistemic Neural Networks

SafetyDGX agent

arXiv:2605.07065v1 Announce Type: cross Abstract: Individual treatment effects are not point-identified from data. The Probability of Necessity and Sufficiency (PNS) circumvents this limitation by cha

Checkmate to everyone on X who doubted my coverage of Sam’s firing. Checkmate.

SafetyDGX agent

Checkmate to everyone on X who doubted my coverage of Sam’s firing. Checkmate. 🚨BREAKING: Ilya just confirmed under oath what he saw: Sam lying. And he confirmed that he thought it was appropriate to

Cognitive Agent Compilation for Explicit Problem Solver Modeling

SafetyDGX agent

arXiv:2605.07040v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used for tutoring, feedback generation, and content creation, but their broad pretraining makes them hard to c

← Previous
1…153154155156157…212
Next →