AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
27 May 2026

Auditing and Fixing Economic Validity in Tabular Foundation Models for Discrete Choice

SafetyDGX agent

arXiv:2605.26559v1 Announce Type: cross Abstract: Tabular foundation models achieve strong accuracy on choice prediction tasks, but their predictions often violate the economic logic those tasks requi

BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

SafetyDGX agent

arXiv:2605.27110v1 Announce Type: cross Abstract: In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal discl

BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning

SafetyDGX agent

arXiv:2605.27293v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existing alg


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility

SafetyDGX agent

arXiv:2603.03585v2 Announce Type: replace-cross Abstract: Misinformation is a growing societal threat, and susceptibility to misinformative claims varies across demographic groups due to differences i

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation

SafetyDGX agent

arXiv:2601.03525v3 Announce Type: replace-cross Abstract: Effective reward design is a central challenge in Reinforcement Learning (RL) for code generation. Mainstream test-suite-level outcome rewards

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

SafetyDGX agent

arXiv:2605.06213v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) today rests on fixed benchmarks that apply the same set of items to any model, producing ceiling and floor e

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

SafetyDGX agent

arXiv:2605.26491v1 Announce Type: cross Abstract: Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image

Beyond the Data Mesh Illusion: Designing Modern AI-augmented Lakehouses to Bridge the Gap Between Theory and Practice

SafetyDGX agent

arXiv:2605.27131v1 Announce Type: cross Abstract: Enterprise data platforms face an enduring tension between domain self-service and holistic governance. The data mesh paradigm proposed decentralized

Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2605.26684v1 Announce Type: cross Abstract: Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have

Bilevel Optimization over Saddle Points of Zero-Sum Markov Games

SafetyDGX agent

arXiv:2605.26654v1 Announce Type: cross Abstract: Reinforcement learning (RL) often has a hierarchical structure, where an upper-level (UL) learner selects model parameters and a lower-level (LL) deci

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty

SafetyDGX agent

arXiv:2605.26627v1 Announce Type: cross Abstract: Deploying reinforcement learning in safety critical domains, from autonomous vehicles to medical decision support, is constrained by failures arising

BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization

SafetyDGX agent

arXiv:2605.26182v1 Announce Type: new Abstract: Generating physically buildable brick structures from 3D shapes requires more than geometric reconstruction: the output must also satisfy discrete part

Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial

SafetyDGX agent

arXiv:2605.26577v1 Announce Type: cross Abstract: Learning-based methods for synthesizing controllers have gained popularity due to their high expressiveness and strong empirical performance. However,

CFG-OEC: Classifier Free Guidance with Orthogonal Error Correction

SafetyDGX agent

arXiv:2511.14075v2 Announce Type: replace-cross Abstract: Classifier free guidance is a standard method for conditional sampling in diffusion models, but its sampling rule is not aligned with the obje

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

SafetyDGX agent

arXiv:2605.26542v1 Announce Type: cross Abstract: Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterp

CmIVTP: Cross-modal Interaction-based Vessel Trajectory Prediction for Maritime Intelligence

SafetyDGX agent

arXiv:2605.26524v1 Announce Type: cross Abstract: Maritime intelligent transportation systems (MITS) are essential for ensuring navigation safety and efficiency in busy waterways. However, accurate ve

CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment

SafetyDGX agent

arXiv:2603.07211v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes tra

Completely agree, @hwchase17 ! This is the meta-level breakthrough we've all been waiting for. Self-optimizing loops finally feel production…

SafetyDGX agent

Completely agree, @hwchase17 ! This is the meta-level breakthrough we've all been waiting for. Self-optimizing loops finally feel production-ready because LangSmith Engine turns evaluation from a manu

Completion vs Optimality: Policy Gradient in Long-Horizon Cumulative-Damage Problems

SafetyDGX agent

arXiv:2605.26657v1 Announce Type: new Abstract: Long-horizon decision problems with cumulative damage couple locally attractive actions to globally adverse outcomes. We identify two orthogonal failure

Constrained Bayesian Experimental Design via Online Planning

SafetyDGX agent

arXiv:2605.26990v1 Announce Type: cross Abstract: Bayesian experimental design (BED) is a principled framework for data-efficient design of sequential experiments. However, existing BED methods are un

Constrained Meta Reinforcement Learning with Provable Test-Time Safety

SafetyDGX agent

arXiv:2601.21845v2 Announce Type: replace Abstract: Meta reinforcement learning (RL) allows agents to leverage experience across a distribution of tasks on which the agent can train at will, enabling

Conv-to-Bench: Evaluating Language Models Via User-Assistant Dialogues In Code Tasks

SafetyDGX agent

arXiv:2605.26440v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) has outpaced the scalability of traditional evaluation benchmarks, which remain heavily dependent

Cost of Structural Learning Under Censored Feedback: A Threshold-Bandit Approach

SafetyDGX agent

arXiv:2605.27076v1 Announce Type: cross Abstract: In many multi-agent applications, tasks yield rewards only when executed by a coalition meeting an unknown size threshold; otherwise, feedback is full

Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

SafetyDGX agent

arXiv:2605.27115v1 Announce Type: new Abstract: Domain specialization can improve LLM behavior in vertical domains, but often weakens the general capabilities inherited from the original model. Recent

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

SafetyDGX agent

arXiv:2603.21563v2 Announce Type: replace Abstract: Collaborative multi-agent large language models (LLMs) can solve complex reasoning tasks by decomposing roles, but reinforcement learning for such s

Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking

SafetyDGX agent

arXiv:2605.26385v1 Announce Type: cross Abstract: Large-scale search, recommendation, and retrieval-augmented generation (RAG) systems typically employ a two-stage architecture: an early-stage ranker

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

SafetyDGX agent

arXiv:2605.26293v1 Announce Type: cross Abstract: Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves do

Cross-Receiver Generalization for RF Fingerprint Identification via Feature Disentanglement and Adversarial Training

SafetyDGX agent

arXiv:2510.09405v2 Announce Type: replace Abstract: Radio frequency fingerprint identification (RFFI) is a key technique for wireless network security, leveraging intrinsic hardware imperfections to e

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

SafetyDGX agent

arXiv:2605.26365v1 Announce Type: new Abstract: Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping h

Curriculum Learning for Safety Alignment

SafetyDGX agent

arXiv:2605.26315v1 Announce Type: cross Abstract: Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and exhibit

Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

SafetyDGX agent

arXiv:2605.27157v1 Announce Type: new Abstract: Retrieval-augmented LLMs are deployed for tasks where evidence quality determines action safety, yet evaluation protocols assume that single-turn robust

Dimensional Distribution Emotion State: Leveraging Valence and Arousal as a Common Embedding Space for Visual Emotion Analysis

SafetyDGX agent

arXiv:2605.26262v1 Announce Type: new Abstract: Museums are important sites for the dissemination of culture and art. They are institutions rooted in history and tradition; their exhibitions are often

DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation

SafetyDGX agent

arXiv:2605.26236v1 Announce Type: new Abstract: Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lex

DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding

SafetyDGX agent

arXiv:2605.26656v1 Announce Type: new Abstract: Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to te

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

SafetyDGX agent

arXiv:2605.26952v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) has proven effective for training LLM-based agents with external tool-use capabilities. However, we identify that ag

Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient

SafetyDGX agent

arXiv:2605.26478v1 Announce Type: cross Abstract: We present the stochastic decoupled policy gradient (SDPG), a lightweight visual reinforcement learning (RL) method that trains diverse visuomotor con

Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories

SafetyDGX agent

arXiv:2605.26492v1 Announce Type: cross Abstract: LLM-generated stories are a popular use case, but they show very low variability. We sample 20,000 total stories from four current models using five p

Elon Musk on why humanity must become multiplanetary: “I think it's important for the long-term preservation and ultimately the expansion an…

SafetyDGX agent

Elon Musk on why humanity must become multiplanetary: “I think it's important for the long-term preservation and ultimately the expansion and extension of the scope and scale of consciousness... that

EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation

SafetyDGX agent

arXiv:2605.26785v1 Announce Type: cross Abstract: Post-trained LLMs are often optimized to align responses with human preferences, making them safe, polite, and conversationally appropriate. In advers

Enabling Extensible Embodied Capabilities with Tools

SafetyDGX agent

arXiv:2605.26637v1 Announce Type: new Abstract: Most existing embodied intelligence methods formulate perception, reasoning, planning, and control within a unified parameterized policy. Yet these capa

Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models

SafetyDGX agent

arXiv:2605.26332v1 Announce Type: cross Abstract: Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been

🔬ESMFold2: The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub

SafetyDGX agent

ESMFold2 represents an advancement in protein structure prediction leveraging scaling laws and transformer-based language models, building on principles that favor compute and data scale over hand-cra

Ethical Fairness without Demographics in Human-Centered AI

SafetyDGX agent

arXiv:2603.13373v3 Announce Type: replace-cross Abstract: In ubiquitous and mobile health systems, computational models infer human states from wearable, behavioral, and physiological sensing data. In

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights

SafetyDGX agent

arXiv:2501.06708v5 Announce Type: replace-cross Abstract: Large-scale web-crawled datasets contain noise, bias, and irrelevant information, necessitating data selection techniques. Existing methods de

FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions

SafetyDGX agent

arXiv:2605.27062v1 Announce Type: new Abstract: State-of-the-art performance for Automatic Speech Recognition (ASR) largely depends on the availability of large-scale labeled corpora. This creates a d

Few-shot Cross-country Generalization of Tabular Machine Learning and Foundation Models for Childhood Anemia Prediction under Distribution Shift

SafetyDGX agent

arXiv:2605.26589v1 Announce Type: cross Abstract: Childhood anemia affects around 40% of children aged 6-59 months globally and arises from heterogeneous factors, limiting model generalizability. We e

FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

SafetyDGX agent

arXiv:2605.27333v1 Announce Type: new Abstract: Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary

Flow Matching Policy Optimization with Mirror Descent and Entropy Constraints

SafetyDGX agent

arXiv:2603.17685v3 Announce Type: replace Abstract: Balancing policy expressiveness with the exploration-exploitation trade-off is a core challenge in online Reinforcement Learning (RL). While Stochas

FM-fMRI: Event Conditioned Flow Matching for Rest-to-Task fMRI Time-Series Synthesis

SafetyDGX agent

arXiv:2605.26423v1 Announce Type: new Abstract: Task-based fMRI provides a direct readout of task-evoked neural dynamics, but it is expensive and difficult to acquire at scale, motivating rest-to-task

Foundations of a Time-Consistent Counterfactual Actuarial Runtime for Autonomous AI Agents

SafetyDGX agent

arXiv:2605.26508v1 Announce Type: cross Abstract: We propose a foundational runtime actuarial layer for autonomous AI agents in which every side-effect-bearing action carries a time-consistent, counte

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

SafetyDGX agent

arXiv:2605.26926v1 Announce Type: new Abstract: Computing legal indicators from normative texts is a key task in legal monitoring and policy evaluation, but presents significant challenges due to the

From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator

SafetyDGX agent

arXiv:2605.26403v1 Announce Type: new Abstract: A long-standing goal of the research community is to develop highly interactive LLM-based dialogue agents. Recent research focuses on optimizing policie

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

SafetyDGX agent

arXiv:2605.26601v1 Announce Type: new Abstract: Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible trainin

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

SafetyDGX agent

arXiv:2605.26158v1 Announce Type: cross Abstract: Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshol

Gary's no way as bad as some people make him out to be.

SafetyDGX agent

Gary's no way as bad as some people make him out to be. some data i shared yesterday on anthropic revenue possibly slowing down aren’t as a compelling as i thought; i have deleted my posts and await b

Generalist Graph Anomaly Detection via Prototype-Based Distillation

SafetyDGX agent

arXiv:2605.26857v1 Announce Type: new Abstract: Driven by the pressing demand for graph anomaly detection (GAD) in high-stakes domains, the generalist GAD paradigm, which trains a single detector tran

GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation

SafetyDGX agent

arXiv:2602.16449v2 Announce Type: replace-cross Abstract: Generative model evaluation commonly relies on high-dimensional embedding spaces to compute distances between samples. We show that dataset re

Grounding Text Embeddings in Stakeholder Associations

SafetyDGX agent

arXiv:2605.27168v1 Announce Type: cross Abstract: Text embeddings are widely used to analyse large corpora of complex texts. However, it is unclear whether the embeddings capture the same semantic dis

Heterogeneous AAV Logistics Task Allocation: A Reinforcement Learning Enhanced Overlapping Coalition Formation Game Approach

SafetyDGX agent

arXiv:2605.26471v1 Announce Type: new Abstract: In dynamic urban logistics, the stochastic emergence of time-sensitive tasks poses a significant optimality challenge for heterogeneous AAVs logistics t

Hi-SAM: A Hierarchical Structure-Aware Multi-modal Framework for Large-Scale Recommendation

SafetyDGX agent

arXiv:2602.11799v2 Announce Type: replace Abstract: Multi-modal recommendation has gained traction as items possess rich attributes like text and images. Semantic ID-based approaches effectively discr

← Previous
1…112113114115116…214
Next →