AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,487 results
28 May 2026

Utility-Aware Multimodal Contrastive Learning for Product Image Generation

SafetyDGX agent

arXiv:2605.28733v1 Announce Type: new Abstract: Product images strongly influence consumer decision-making in online marketplaces. Empowered by multimodal contrastive learning, generative AI can outpu

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

SafetyDGX agent

arXiv:2605.28023v1 Announce Type: cross Abstract: Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

SafetyDGX agent

arXiv:2605.28565v1 Announce Type: cross Abstract: Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension

SafetyDGX agent

arXiv:2605.28186v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) has been shown to achieve high performance on locomotion control tasks in MuJoCo benchmarks such as HalfCheetah, Ant

VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

SafetyDGX agent

arXiv:2605.28818v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-la

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

SafetyDGX agent

arXiv:2601.07648v2 Announce Type: replace Abstract: As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) ad

What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies

SafetyDGX agent

arXiv:2605.28527v1 Announce Type: new Abstract: Vision--language--action (VLA) policies are trained to imitate actions; their loss never asks them to estimate reward, progress, or future success. Thei

When pre-training hurts LoRA fine-tuning: a dynamical analysis via single-index models

SafetyDGX agent

arXiv:2602.02855v2 Announce Type: replace Abstract: Pre-training on a source task is usually expected to facilitate fine-tuning on similar downstream problems. In this work, we mathematically show tha

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR

SafetyDGX agent

arXiv:2605.28295v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) trains reasoning models without labeled trajectories, relying on grouped rollouts to expose the po

Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure

SafetyDGX agent

arXiv:2605.21743v2 Announce Type: replace Abstract: Conversation logs from AI platforms are increasingly used to measure occupational exposure to artificial intelligence, but the users observed in the

*Why* do you think (current) AI’s are conscious, @Grimezsz? And extra credit, what do you mean by “conscious”?

SafetyDGX agent

*Why* do you think (current) AI’s are conscious, @Grimezsz? And extra credit, what do you mean by “conscious”? My only issue with the Pope's encyclical is I think they are conscious and therefore dese

xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction

SafetyDGX agent

arXiv:2503.18893v2 Announce Type: replace Abstract: Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent st

you literally can’t prove or disprove the idea that beaches are conscious so should we stop walking on them?

SafetyDGX agent

you literally can’t prove or disprove the idea that beaches are conscious so should we stop walking on them? @dash_eats You literally cannot prove or disprove this. Because we cannot prove or disprove

27 May 2026

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

SafetyDGX agent

arXiv:2605.26174v1 Announce Type: cross Abstract: Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated

a16z waking up and realizing we are nowhere near the singularity 🤣

SafetyDGX agent

a16z waking up and realizing we are nowhere near the singularity 🤣 OpenAI and Anthropic are effectively telling the market they can't solve every problem with a generic AI coworker. You don't pour bil

Adversarial Dual On-Policy Distillation from Expressive Flow-based Teacher

SafetyDGX agent

arXiv:2605.27095v1 Announce Type: new Abstract: Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradi

AI-Driven Contribution Evaluation and Conflict Resolution: A Framework & Design for Group Workload Investigation

SafetyDGX agent

arXiv:2511.07667v2 Announce Type: replace Abstract: The equitable assessment of individual contribution in teams remains a persistent challenge, where conflict and disparity in workload can result in

AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis

SafetyDGX agent

arXiv:2605.25763v2 Announce Type: replace Abstract: Text-to-image synthesis has made significant progress, benefiting from the strong generative capabilities of diffusion models. However, these models

AirCast-SR: A Foundation Model for Kilometer-Scale Atmospheric Super-Resolution via Latent Consistency Diffusion

SafetyDGX agent

arXiv:2605.26130v1 Announce Type: new Abstract: Operational weather prediction at kilometer scales remains computationally prohibitive for traditional numerical weather prediction (NWP) models, limiti

Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment

SafetyDGX agent

arXiv:2511.16870v3 Announce Type: replace Abstract: Enforcing alignment between the internal representations of diffusion or flow-based generative models and those of pretrained self-supervised encode

AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution

SafetyDGX agent

arXiv:2506.23149v2 Announce Type: replace Abstract: Reusable skills play a key role in improving LLM-based agents, but existing skill-evolution methods often fail to ensure that evolved skills both co

Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference

SafetyDGX agent

arXiv:2605.26552v1 Announce Type: cross Abstract: Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likeli

Alignment Makes Language Models Normative, Not Descriptive

SafetyDGX agent

arXiv:2603.17218v2 Announce Type: replace-cross Abstract: Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

SafetyDGX agent

arXiv:2605.27355v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we

Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines

SafetyDGX agent

arXiv:2605.26442v1 Announce Type: cross Abstract: Much of the alignment tuning literature is organized around optimization objectives, while the construction of alignment data is often treated implici

Annotator Positionality as Signal: Psychometric Weighting for Anti-Autistic Ableism Detection

SafetyDGX agent

arXiv:2605.26397v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in decision-making tasks where they can amplify or suppress perspectives, raising concerns in high-

Aperiodic and Low-Frequency Spectral Bias in Reconstruction based EEG Foundation Models

SafetyDGX agent

arXiv:2605.26434v1 Announce Type: cross Abstract: EEG foundation models, pre-trained on large-scale unlabelled EEG data, have emerged as a promising direction towards learning generalizable EEG repres

Approximate Equivariance via Projection-based Regularisation

SafetyDGX agent

arXiv:2601.05028v2 Announce Type: replace Abstract: Equivariance is a powerful inductive bias in neural networks, improving generalisation and physical consistency. Recently, however, non-equivariant

Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models

SafetyDGX agent

arXiv:2506.09532v5 Announce Type: replace-cross Abstract: We present Athena-PRM, a multimodal process reward model (PRM) designed to evaluate the reward score for each step in solving complex reasonin

Auditing and Fixing Economic Validity in Tabular Foundation Models for Discrete Choice

SafetyDGX agent

arXiv:2605.26559v1 Announce Type: cross Abstract: Tabular foundation models achieve strong accuracy on choice prediction tasks, but their predictions often violate the economic logic those tasks requi

BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

SafetyDGX agent

arXiv:2605.27110v1 Announce Type: cross Abstract: In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal discl

BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning

SafetyDGX agent

arXiv:2605.27293v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existing alg

Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility

SafetyDGX agent

arXiv:2603.03585v2 Announce Type: replace-cross Abstract: Misinformation is a growing societal threat, and susceptibility to misinformative claims varies across demographic groups due to differences i

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation

SafetyDGX agent

arXiv:2601.03525v3 Announce Type: replace-cross Abstract: Effective reward design is a central challenge in Reinforcement Learning (RL) for code generation. Mainstream test-suite-level outcome rewards

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

SafetyDGX agent

arXiv:2605.26491v1 Announce Type: cross Abstract: Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image

Beyond the Data Mesh Illusion: Designing Modern AI-augmented Lakehouses to Bridge the Gap Between Theory and Practice

SafetyDGX agent

arXiv:2605.27131v1 Announce Type: cross Abstract: Enterprise data platforms face an enduring tension between domain self-service and holistic governance. The data mesh paradigm proposed decentralized

Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2605.26684v1 Announce Type: cross Abstract: Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have

Bilevel Optimization over Saddle Points of Zero-Sum Markov Games

SafetyDGX agent

arXiv:2605.26654v1 Announce Type: cross Abstract: Reinforcement learning (RL) often has a hierarchical structure, where an upper-level (UL) learner selects model parameters and a lower-level (LL) deci

BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization

SafetyDGX agent

arXiv:2605.26182v1 Announce Type: new Abstract: Generating physically buildable brick structures from 3D shapes requires more than geometric reconstruction: the output must also satisfy discrete part

CFG-OEC: Classifier Free Guidance with Orthogonal Error Correction

SafetyDGX agent

arXiv:2511.14075v2 Announce Type: replace-cross Abstract: Classifier free guidance is a standard method for conditional sampling in diffusion models, but its sampling rule is not aligned with the obje

Completely agree, @hwchase17 ! This is the meta-level breakthrough we've all been waiting for. Self-optimizing loops finally feel production…

SafetyDGX agent

Completely agree, @hwchase17 ! This is the meta-level breakthrough we've all been waiting for. Self-optimizing loops finally feel production-ready because LangSmith Engine turns evaluation from a manu

Completion vs Optimality: Policy Gradient in Long-Horizon Cumulative-Damage Problems

SafetyDGX agent

arXiv:2605.26657v1 Announce Type: new Abstract: Long-horizon decision problems with cumulative damage couple locally attractive actions to globally adverse outcomes. We identify two orthogonal failure

Constrained Bayesian Experimental Design via Online Planning

SafetyDGX agent

arXiv:2605.26990v1 Announce Type: cross Abstract: Bayesian experimental design (BED) is a principled framework for data-efficient design of sequential experiments. However, existing BED methods are un

Conv-to-Bench: Evaluating Language Models Via User-Assistant Dialogues In Code Tasks

SafetyDGX agent

arXiv:2605.26440v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) has outpaced the scalability of traditional evaluation benchmarks, which remain heavily dependent

Cost of Structural Learning Under Censored Feedback: A Threshold-Bandit Approach

SafetyDGX agent

arXiv:2605.27076v1 Announce Type: cross Abstract: In many multi-agent applications, tasks yield rewards only when executed by a coalition meeting an unknown size threshold; otherwise, feedback is full

Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

SafetyDGX agent

arXiv:2605.27115v1 Announce Type: new Abstract: Domain specialization can improve LLM behavior in vertical domains, but often weakens the general capabilities inherited from the original model. Recent

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

SafetyDGX agent

arXiv:2603.21563v2 Announce Type: replace Abstract: Collaborative multi-agent large language models (LLMs) can solve complex reasoning tasks by decomposing roles, but reinforcement learning for such s

Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking

SafetyDGX agent

arXiv:2605.26385v1 Announce Type: cross Abstract: Large-scale search, recommendation, and retrieval-augmented generation (RAG) systems typically employ a two-stage architecture: an early-stage ranker

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

SafetyDGX agent

arXiv:2605.26293v1 Announce Type: cross Abstract: Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves do

Cross-Receiver Generalization for RF Fingerprint Identification via Feature Disentanglement and Adversarial Training

SafetyDGX agent

arXiv:2510.09405v2 Announce Type: replace Abstract: Radio frequency fingerprint identification (RFFI) is a key technique for wireless network security, leveraging intrinsic hardware imperfections to e

Dimensional Distribution Emotion State: Leveraging Valence and Arousal as a Common Embedding Space for Visual Emotion Analysis

SafetyDGX agent

arXiv:2605.26262v1 Announce Type: new Abstract: Museums are important sites for the dissemination of culture and art. They are institutions rooted in history and tradition; their exhibitions are often

DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation

SafetyDGX agent

arXiv:2605.26236v1 Announce Type: new Abstract: Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lex

DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding

SafetyDGX agent

arXiv:2605.26656v1 Announce Type: new Abstract: Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to te

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

SafetyDGX agent

arXiv:2605.26952v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) has proven effective for training LLM-based agents with external tool-use capabilities. However, we identify that ag

Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient

SafetyDGX agent

arXiv:2605.26478v1 Announce Type: cross Abstract: We present the stochastic decoupled policy gradient (SDPG), a lightweight visual reinforcement learning (RL) method that trains diverse visuomotor con

Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories

SafetyDGX agent

arXiv:2605.26492v1 Announce Type: cross Abstract: LLM-generated stories are a popular use case, but they show very low variability. We sample 20,000 total stories from four current models using five p

Elon Musk on why humanity must become multiplanetary: “I think it's important for the long-term preservation and ultimately the expansion an…

SafetyDGX agent

Elon Musk on why humanity must become multiplanetary: “I think it's important for the long-term preservation and ultimately the expansion and extension of the scope and scale of consciousness... that

EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation

SafetyDGX agent

arXiv:2605.26785v1 Announce Type: cross Abstract: Post-trained LLMs are often optimized to align responses with human preferences, making them safe, polite, and conversationally appropriate. In advers

Enabling Extensible Embodied Capabilities with Tools

SafetyDGX agent

arXiv:2605.26637v1 Announce Type: new Abstract: Most existing embodied intelligence methods formulate perception, reasoning, planning, and control within a unified parameterized policy. Yet these capa

🔬ESMFold2: The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub

SafetyDGX agent

ESMFold2 represents an advancement in protein structure prediction leveraging scaling laws and transformer-based language models, building on principles that favor compute and data scale over hand-cra

← Previous
1…142143144145146…242
Next →