AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,435 results
Safety

Securing SIM-Assisted Wireless Networks via Quantum Reinforcement Learning

DGX agent

arXiv:2602.13238v2 Announce Type: replace-cross Abstract: Stacked intelligent metasurfaces (SIMs) have recently emerged as a powerful wave-domain technology that enables multi-stage manipulation of el

safetyarxiv-cs-lg
29 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Self-Play Reinforcement Learning under Imperfect Information in Big 2

DGX agent

arXiv:2605.28863v1 Announce Type: cross Abstract: Imperfect-information multiplayer games test whether agents can act under hidden information, sparse rewards, and non-stationary opponents. We study t

safetyarxiv-cs-ai
29 May 2026
Safety

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

DGX agent

arXiv:2605.30116v1 Announce Type: new Abstract: Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models. However, DMD-style vid

safetyarxiv-cs-cv
29 May 2026
Safety

Statistical Embeddings for Similarity, Retrieval, and Interpretable Alignment of Numeric Tabular Datasets

DGX agent

arXiv:2605.30289v1 Announce Type: new Abstract: Numeric tabular datasets are the dominant data format in scientific practice, yet large language models lack native mechanisms for representing numeric

safetyarxiv-cs-lg
29 May 2026
Safety

Teaching Values to Machines: Simulating Human-Like Behavior in LLMs

DGX agent

arXiv:2605.30036v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manif

safetyarxiv-cs-ai
29 May 2026
Safety

The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation

DGX agent

arXiv:2512.10388v2 Announce Type: replace-cross Abstract: Conventional Sequential Recommender Systems (SRS) typically assign unique hash IDs (HID) to construct item embeddings, which mainly capture co

safetyarxiv-cs-ai
29 May 2026
Safety

The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models

DGX agent

arXiv:2605.29123v1 Announce Type: new Abstract: Masked diffusion language models (MDMs) uniquely support any-order generation, with confidence-based decoding currently serving as the de facto standard

safetyarxiv-cs-ai
29 May 2026
Safety

The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane

DGX agent

arXiv:2605.29082v1 Announce Type: new Abstract: AI agents are increasingly expected to operate as digital employees: accessing enterprise data, making decisions, and taking actions autonomously. But a

safetyarxiv-cs-ai
29 May 2026
Safety

The Sample Complexity of Multiclass and Sparse Contextual Bandits

DGX agent

arXiv:2605.29645v1 Announce Type: cross Abstract: We study contextual bandits in the stochastic i.i.d. setting, where a learner observes contexts drawn from an unknown distribution, selects actions fr

safetyarxiv-cs-ai
29 May 2026
Safety

Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

DGX agent

arXiv:2605.29032v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably

safetyarxiv-cs-lg
29 May 2026
Safety

Think Fast, Talk Smart: Partitioning Deterministic and Neural Computation for Structured Health Text Generation

DGX agent

arXiv:2605.29652v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being used to generate health text from structured records such as wearable time series, biomarkers, vital

safetyarxiv-cs-ai
29 May 2026
Safety

Toward AI Systems That Understand Self and Others: A Multi-Phase Inference Framework for Human Cognitive Diversity and World-Model Alignment

DGX agent

arXiv:2605.29930v1 Announce Type: new Abstract: Mutual misunderstanding in contemporary society does not arise merely because people hold different opinions or values. Even under the same observations

safetyarxiv-cs-ai
29 May 2026
Safety

Toward User Preference Alignment in LLM Recommendation via Explicit Context Feedback

DGX agent

arXiv:2605.29141v1 Announce Type: cross Abstract: Traditional recommender systems (RecSys) primarily infer user preferences from implicit signals (such as clicks, watches, and purchases), often neglec

safetyarxiv-cs-ai
29 May 2026
Safety

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation

DGX agent

arXiv:2605.29430v1 Announce Type: new Abstract: Automatic speech recognition (ASR) is a core component of human--computer interaction and an increasingly important front-end for LLM-based assistants a

safetyarxiv-cs-ai
29 May 2026
Safety

TraceCodec: A Compiler-Backed Neural Codec for Stateful Multi-Flow Network Traffic Traces

DGX agent

arXiv:2605.29941v1 Announce Type: cross Abstract: Critical networking workflows require high-fidelity packet captures (PCAPs) for testing, security analysis, and protocol validation, not just statisti

safetyarxiv-cs-lg
29 May 2026
Safety

TRACER: Persistent Regularization for Robust Multimodal Finetuning

DGX agent

arXiv:2605.29380v1 Announce Type: cross Abstract: Mainstream strategies for finetuning pretrained multimodal models often degrade out-of-distribution (OOD) robustness, a phenomenon known as catastroph

safetyarxiv-cs-ai
29 May 2026
Safety

Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

DGX agent

arXiv:2605.29894v1 Announce Type: new Abstract: Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual task

safetyarxiv-cs-cv
29 May 2026
Safety

TriSearch: Learning to Optimize Triangulations via Bistellar Flips

DGX agent

arXiv:2605.30220v1 Announce Type: new Abstract: We introduce TriSearch, a reinforcement learning framework for optimizing objectives over triangulations of a polytope via bistellar flips. The key idea

safetyarxiv-cs-lg
29 May 2026
Safety

Uncertainty Estimation via Hyperspherical Confidence Mapping

DGX agent

arXiv:2605.05964v2 Announce Type: replace Abstract: Quantifying uncertainty in neural network predictions is essential for high-stakes domains such as autonomous driving, healthcare, and manufacturing

safetyarxiv-cs-lg
29 May 2026
Safety

User-Aware Active Knowledge Acquisition for Emotional Support Dialogue

DGX agent

arXiv:2605.29715v1 Announce Type: new Abstract: Emotional support plays an important role in dialogue systems, and its success depends on adapting to a user's evolving and implicit needs across multi-

safetyarxiv-cs-cl
29 May 2026
Safety

ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems

DGX agent

arXiv:2602.08567v2 Announce Type: replace-cross Abstract: Multi-agent large language model (LLM) systems increasingly consist of agents that observe and respond to one another's outputs. While value a

safetyarxiv-cs-cl
29 May 2026
Safety

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

DGX agent

arXiv:2605.30117v1 Announce Type: new Abstract: Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Tra

safetyarxiv-cs-ai
29 May 2026
Safety

When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop

DGX agent

arXiv:2605.29267v1 Announce Type: new Abstract: Foundation models are increasingly trained on synthetic data generated by prior model iterations rather than exclusively on real data. This self-consumi

safetyarxiv-cs-ai
29 May 2026
Safety

xModel-KD: Cross-modal Knowledge Distillation for 3D Scene Perception using LiDAR

DGX agent

arXiv:2605.30111v1 Announce Type: cross Abstract: Point cloud segmentation is a fundamental task in 3D scene understanding. Its progress is constrained by the high cost and time required for dense 3D

safetyarxiv-cs-ai
29 May 2026
Safety

A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons

DGX agent

arXiv:2605.27461v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies have shown promising manipulation capabilities, yet their practical impact is often limited by the reliability dem

safetyarxiv-cs-ro
28 May 2026
Safety

A Structural Theory of Position Bias in Transformers

DGX agent

arXiv:2602.16837v2 Announce Type: replace Abstract: Transformer models systematically favor certain token positions, yet the architectural origins of this position bias remain poorly understood. This

safetyarxiv-cs-lg
28 May 2026
Safety

AdaMemento: Adaptive Memory-Assisted Policy Optimization for Reinforcement Learning

DGX agent

arXiv:2410.04498v2 Announce Type: replace Abstract: In sparse reward scenarios of reinforcement learning (RL), the memory mechanism provides promising shortcuts to policy optimization by reflecting on

safetyarxiv-cs-lg
28 May 2026
Safety

ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation

DGX agent

arXiv:2605.28396v1 Announce Type: cross Abstract: On-policy distillation (OPD) transfers reasoning behavior by training a student on teacher feedback along student-generated trajectories, but standard

safetyarxiv-cs-ai
28 May 2026
Safety

Affective Music Recommendation: A Rollout-Based World Model for Offline Preference Optimization

DGX agent

arXiv:2605.28810v1 Announce Type: new Abstract: Functional music applications, from consumer focus and sleep aids to clinical interventions, share a distinctive recommendation problem: success is defi

safetyarxiv-cs-lg
28 May 2026
Safety

AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems

DGX agent

arXiv:2605.27466v1 Announce Type: cross Abstract: Multi-agent systems built on large language models (LLMs) require many coordination choices that are difficult to fix a priori: which skill protocol t

safetyarxiv-cs-ai
28 May 2026
Safety

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning

DGX agent

arXiv:2605.28774v1 Announce Type: new Abstract: Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning

safetyarxiv-cs-cl
28 May 2026
Safety

Agents that Matter: Optimizing Multi-Agent LLMs via Removal-Based Attribution

DGX agent

arXiv:2605.27621v1 Announce Type: cross Abstract: As multi-agent systems (MAS) become increasingly complex, identifying the contributions of individual agents is critical for system optimization. Howe

safetyarxiv-cs-cl
28 May 2026
Safety

AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

DGX agent

arXiv:2605.28255v1 Announce Type: new Abstract: AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration re

safetyarxiv-cs-ai
28 May 2026
Safety

AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning

DGX agent

arXiv:2605.28809v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) is important in building real-world learning systems. In CLIP-based CIL, the model performs classification by comparing

safetyarxiv-cs-cv
28 May 2026
Safety

Artemis: Structured Visual Reasoning for Perception Policy Learning

DGX agent

arXiv:2512.01988v2 Announce Type: replace Abstract: Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural languag

safetyarxiv-cs-cv
28 May 2026
Safety

Auditable Decision Models with Learned Abstention and Real-Time Steering

DGX agent

arXiv:2605.27768v1 Announce Type: new Abstract: Production AI systems often operate with incomplete, conflicting, or insufficient evidence. Forced classifiers collapse such cases into action labels, w

safetyarxiv-cs-ai
28 May 2026
Safety

Auditing Stance Asymmetry in Generative Explanations

DGX agent

arXiv:2605.27988v1 Announce Type: new Abstract: Bias evaluation for language models has made substantial progress on bounded comparisons, such as overt derogation, stereotype association, or label-sen

safetyarxiv-cs-cl
28 May 2026
Safety

Automated Estimation of Impact Time, Impact Location, and Shuttlecock Speed in Badminton Smashes Using Event Cameras

DGX agent

arXiv:2605.28011v1 Announce Type: new Abstract: Quantifying impact phenomena in badminton smashes is important for evaluating both athletic performance and equipment; however, conventional measurement

safetyarxiv-cs-cv
28 May 2026
Safety

Behavioural Analysis of Alignment Faking

DGX agent

arXiv:2605.27681v1 Announce Type: new Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deploym

safetyarxiv-cs-ai
28 May 2026
Safety

Beyond Binary: Sim-to-Real Dexterous Manipulation with Physics-Grounded Contact Representation

DGX agent

arXiv:2605.28812v1 Announce Type: cross Abstract: A primary bottleneck in contact-rich manipulation is the difficulty of collecting real-world data. Sim-to-real reinforcement learning offers a scalabl

safetyarxiv-cs-ai
28 May 2026
Safety

BiasEdit: A Training-Free Bias-Detect-and-Edit Framework for Learning Fair Visual Classifiers

DGX agent

arXiv:2605.28450v1 Announce Type: cross Abstract: Visual data from the Web power image classifiers, which often underpin many web services, such as recommendation and content moderation. However, the

safetyarxiv-cs-ai
28 May 2026
Safety

Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking

DGX agent

arXiv:2605.28632v1 Announce Type: cross Abstract: Cryptographic watermarking is a leading defense for attributing text generated by large language models (LLMs). Existing schemes, including KGW, Unigr

safetyarxiv-cs-ai
28 May 2026
Safety

Boundary Suppression Asymmetry in Post-trained Assistants: Over-expansion as a Controllability Cost

DGX agent

arXiv:2605.27969v1 Announce Type: new Abstract: Post-trained language-model assistants are often optimized to avoid under-answering, encouraging complete, helpful, cautious, and proactive responses. W

safetyarxiv-cs-cl
28 May 2026
Safety

BPPO: Binary Prefix Policy Optimization for Efficient GRPO-Style Reasoning RL with Concise Responses

DGX agent

arXiv:2605.28028v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is widely used for training reasoning models, but updating all sampled completions in each group incurs substa

safetyarxiv-cs-lg
28 May 2026
Safety

Breaking the Script Barrier: Enabling Automatic Alignment for PoS-based ASR Error Analysis in Non-Latin Scripts

DGX agent

arXiv:2605.28438v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) systems are commonly evaluated using aggregate metrics such as Word Error Rate (WER), which do not capture the lingui

safetyarxiv-cs-cl
28 May 2026
Safety

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information

DGX agent

arXiv:2605.28070v1 Announce Type: new Abstract: We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified

safetyarxiv-cs-ai
28 May 2026
Safety

Calibrated Inference for the Conditional Average Treatment Effect in the Few-Placebo Regime via Gaussian Processes

DGX agent

arXiv:2605.27473v1 Announce Type: cross Abstract: Estimating how much an intervention helps a given individual the conditional average treatment effect (CATE) is increasingly central to decision-makin

safetyarxiv-cs-lg
28 May 2026
Safety

CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation

DGX agent

arXiv:2412.08052v2 Announce Type: replace Abstract: Off-policy evaluation (OPE) is critical for applying contextual bandit algorithms to high-stakes decision-making settings such as healthcare, where

safetyarxiv-cs-lg
28 May 2026
← Previous
1…152153154155156…260
Next →