AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,816 results
Safety

Boundary Suppression Asymmetry in Post-trained Assistants: Over-expansion as a Controllability Cost

DGX agent

arXiv:2605.27969v1 Announce Type: new Abstract: Post-trained language-model assistants are often optimized to avoid under-answering, encouraging complete, helpful, cautious, and proactive responses. W

safetyarxiv-cs-cl
28 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

BPPO: Binary Prefix Policy Optimization for Efficient GRPO-Style Reasoning RL with Concise Responses

DGX agent

arXiv:2605.28028v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is widely used for training reasoning models, but updating all sampled completions in each group incurs substa

safetyarxiv-cs-lg
28 May 2026
Safety

Breaking the Script Barrier: Enabling Automatic Alignment for PoS-based ASR Error Analysis in Non-Latin Scripts

DGX agent

arXiv:2605.28438v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) systems are commonly evaluated using aggregate metrics such as Word Error Rate (WER), which do not capture the lingui

safetyarxiv-cs-cl
28 May 2026
Safety

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information

DGX agent

arXiv:2605.28070v1 Announce Type: new Abstract: We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified

safetyarxiv-cs-ai
28 May 2026
Safety

Calibrated Inference for the Conditional Average Treatment Effect in the Few-Placebo Regime via Gaussian Processes

DGX agent

arXiv:2605.27473v1 Announce Type: cross Abstract: Estimating how much an intervention helps a given individual the conditional average treatment effect (CATE) is increasingly central to decision-makin

safetyarxiv-cs-lg
28 May 2026
Safety

CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation

DGX agent

arXiv:2412.08052v2 Announce Type: replace Abstract: Off-policy evaluation (OPE) is critical for applying contextual bandit algorithms to high-stakes decision-making settings such as healthcare, where

safetyarxiv-cs-lg
28 May 2026
Safety

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

DGX agent

arXiv:2603.22335v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior dis

safetyarxiv-cs-ai
28 May 2026
Safety

Causal Machine Learning: A Survey and Open Problems

DGX agent

arXiv:2206.15475v3 Announce Type: replace Abstract: Causal Machine Learning (CausalML) is an umbrella term for machine learning methods that formalize the data-generation process as a structural causa

safetyarxiv-cs-lg
28 May 2026
Safety

Chance-Constrained MPPI under State and Dynamic Object Prediction Uncertainty and the Evaluation of Collision Risk Calibration

DGX agent

arXiv:2605.28330v1 Announce Type: new Abstract: Chance-constrained Model Predictive Path Integral (MPPI) control is increasingly adopted for navigation in dynamic environments to explicitly bound coll

safetyarxiv-cs-ro
28 May 2026
Safety

CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models

DGX agent

arXiv:2605.28292v1 Announce Type: new Abstract: Implicit Chain-of-Thought (CoT) reduces the inference cost of large language models by internalizing the explicit rationales. However, existing approach

safetyarxiv-cs-cl
28 May 2026
Safety

CodeGENCAT: Generative Computerized Adaptive Testing for Open-ended Coding Problems

DGX agent

arXiv:2602.20020v2 Announce Type: replace Abstract: Existing Computerized Adaptive Testing (CAT) frameworks typically select questions based on the predicted likelihood that the student will answer co

safetyarxiv-cs-cl
28 May 2026
Safety

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

DGX agent

arXiv:2602.15198v2 Announce Type: replace-cross Abstract: Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperativ

safetyarxiv-cs-ai
28 May 2026
Safety

Commit to the Bit: Reactive Reinforcement Learning Done Right

DGX agent

arXiv:2605.28276v1 Announce Type: new Abstract: Reinforcement learning algorithms are commonly analyzed (and designed) under the Markov assumption. This is unrealistic, as most environments encountere

safetyarxiv-cs-lg
28 May 2026
Safety

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization

DGX agent

arXiv:2605.28615v1 Announce Type: new Abstract: Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bi

safetyarxiv-cs-cv
28 May 2026
Safety

Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry

DGX agent

arXiv:2605.27952v1 Announce Type: new Abstract: Visual odometry (VO) is a fundamental component in robotics and augmented reality. RGB-D direct VO benefits from metric depth measurements, but it can d

safetyarxiv-cs-cv
28 May 2026
Safety

COTTA: Context-Aware Transfer Adaptation for Trajectory Prediction in Autonomous Driving

DGX agent

arXiv:2604.00402v2 Announce Type: replace-cross Abstract: Developing robust models to accurately predict the trajectories of surrounding agents is fundamental to autonomous driving safety. However, mo

safetyarxiv-cs-ai
28 May 2026
Safety

Counterfactually Fair Regression via Optimal Transport

DGX agent

arXiv:2605.28251v1 Announce Type: cross Abstract: We consider the problem of learning a counterfactually fair regressor. We adopt a causal uncertainty view in which counterfactual fairness is defined

safetyarxiv-cs-lg
28 May 2026
Safety

CPPO: Contrastive Perception Policy Optimization for VLM Agents

DGX agent

arXiv:2601.00501v2 Announce Type: replace Abstract: We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core

safetyarxiv-cs-cv
28 May 2026
Safety

CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders

DGX agent

arXiv:2604.01604v2 Announce Type: replace Abstract: While modern LLMs are aligned to refuse harmful requests, it is essential to understand the underlying mechanistic basis of this refusal behavior fo

safetyarxiv-cs-ai
28 May 2026
Safety

crazy that this was announced on the day tokenmaxxing died.

DGX agent

crazy that this was announced on the day tokenmaxxing died. Anthropic raised 65 billion in a funding round that valued the artificial intelligence company at 965 billion including the new investment,

safetygary-marcus--x
28 May 2026
Safety

Cross-Entropy Games and Frost Training

DGX agent

arXiv:2605.27701v1 Announce Type: new Abstract: We present Frost Training, a method for improving Monte Carlo-based policy optimization for a large family of LLM-as-a-judge tasks called Cross-Entropy

safetyarxiv-cs-ai
28 May 2026
Safety

Cyberbullying Governance on Social Media: A Unified Framework from Content Identification to Intervention

DGX agent

arXiv:2605.27584v1 Announce Type: new Abstract: The proliferation of social media platforms and online communities has inadvertently catalyzed the spread of cyberbullying, hate speech, and other forms

safetyarxiv-cs-ai
28 May 2026
Safety

DebFilter: Eradicating Biases Stashed in Value

DGX agent

arXiv:2605.28167v1 Announce Type: new Abstract: Text-to-image diffusion models, which are theoretically equivalent to score-based generative models, generate images through a multi-step denoising proc

safetyarxiv-cs-cv
28 May 2026
Safety

Deconstructing Spatial Complexity: Hierarchical Decomposition for LLM Spatial Reasoning

DGX agent

arXiv:2605.28144v1 Announce Type: new Abstract: LLMs have shown remarkable proficiency in general language understanding and reasoning. However, they consistently underperform in spatial reasoning tha

safetyarxiv-cs-ai
28 May 2026
Safety

Delay-Aware Reinforcement Learning for Highway On-Ramp Merging under Stochastic Communication Latency

DGX agent

arXiv:2403.11852v5 Announce Type: replace-cross Abstract: Delayed and partially observable state information poses significant challenges for reinforcement learning (RL)-based control in real-world au

safetyarxiv-cs-ai
28 May 2026
Safety

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes

DGX agent

arXiv:2605.28421v1 Announce Type: new Abstract: Reinforcement learning has become a central paradigm for advancing reasoning in large language models, yet most existing methods still depend on stronge

safetyarxiv-cs-ai
28 May 2026
Safety

Diagnosing Live Within-Policy Instruction Conflicts in LLM Agents with Witnessed Resolution Profiles

DGX agent

arXiv:2605.27784v1 Announce Type: new Abstract: LLM agents are governed by long-lived natural-language prompt policies, but individually reasonable standing rules can interact in uninspected ways. We

safetyarxiv-cs-ai
28 May 2026
Safety

Did AI just solve math? Cal Newport podcast: https://youtu.be/fhZRWZ6J4k4?si=Gzut_5x856n5DW6o

DGX agent

Cal Newport discusses whether recent AI advances represent a breakthrough in mathematical problem-solving capabilities, examining the implications of AI systems' improved performance on complex mathem

safetygary-marcus--x
28 May 2026
Safety

Differentiable Model Predictive Safety for Heterogeneous Mobility at Urban Intersections

DGX agent

arXiv:2605.27418v1 Announce Type: cross Abstract: The imminent integration of autonomous vehicles and mobile robots in urban settings presents a critical safety challenge for future intelligent transp

safetyarxiv-cs-ro
28 May 2026
Safety

Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning

DGX agent

arXiv:2512.02019v3 Announce Type: replace-cross Abstract: Diffusion models excel at sampling from complex, unnormalized distributions. In this work, we extend Maximum Entropy Reinforcement Learning (M

safetyarxiv-cs-ai
28 May 2026
Safety

Diffusion Large Language Models for Visual Speech Recognition

DGX agent

arXiv:2605.28456v1 Announce Type: new Abstract: Existing Visual Speech Recognition (VSR) systems commonly rely on left-to-right autoregressive decoding, which can force premature decisions on visually

safetyarxiv-cs-ai
28 May 2026
Safety

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

DGX agent

arXiv:2605.28491v1 Announce Type: new Abstract: We study real-time audio-responsive character control as a deployment-faithful problem: strictly causal, bounded-latency streaming that must generate co

safetyarxiv-cs-cv
28 May 2026
Safety

Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security

DGX agent

arXiv:2605.27823v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulti

safetyarxiv-cs-ai
28 May 2026
Safety

DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution

DGX agent

arXiv:2605.28678v1 Announce Type: new Abstract: Speculative reasoning has recently been proposed as a means to accelerate reasoning-intensive generation in large multimodal models, but its effectivene

safetyarxiv-cs-ai
28 May 2026
Safety

EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA

DGX agent

arXiv:2605.27846v1 Announce Type: new Abstract: Large Reasoning Models are typically trained via reinforcement learning from verifiable rewards (RLVR). However, existing approaches adopt fixed weights

safetyarxiv-cs-ai
28 May 2026
Safety

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning

DGX agent

arXiv:2602.02150v2 Announce Type: replace-cross Abstract: Test-time reinforcement learning generates multiple candidate answers via repeated rollouts and performs online updates using pseudo-labels co

safetyarxiv-cs-ai
28 May 2026
Safety

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning

DGX agent

arXiv:2603.09882v2 Announce Type: replace-cross Abstract: Extrinsic dexterity leverages environmental contact to overcome the limitations of prehensile manipulation. However, achieving such dexterity

safetyarxiv-cs-ai
28 May 2026
Safety

EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection

DGX agent

arXiv:2605.28630v1 Announce Type: new Abstract: Zero-Shot Anomaly Detection (ZSAD) aims to detect anomalies in unseen domains without target-domain adaptation. Recent CLIP-based methods have shown pro

safetyarxiv-cs-cv
28 May 2026
Safety

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization

DGX agent

arXiv:2605.27741v1 Announce Type: new Abstract: Audio and omni-modal large language models exhibit impressive cross-modal reasoning capabilities. However, applying standard reinforcement learning post

safetyarxiv-cs-cl
28 May 2026
Safety

Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News

DGX agent

arXiv:2605.28598v1 Announce Type: cross Abstract: LLM-powered social agents are increasingly used to simulate online social behavior, yet their realism remains difficult to validate. Existing work has

safetyarxiv-cs-ai
28 May 2026
Safety

Evolving and Detecting Multi-Turn Deception using Geometric Signatures

DGX agent

arXiv:2605.27671v1 Announce Type: cross Abstract: Safety defenses for large language models (LLMs) are typically trained and evaluated on single-turn prompts, yet real attacks often unfold as indirect

safetyarxiv-cs-lg
28 May 2026
Safety

Examining Agents' Bias Amplification versus Suppression in Multi-Agent Systems

DGX agent

arXiv:2605.28098v1 Announce Type: new Abstract: Multi-agent systems are increasingly deployed to support various tasks where agents interact to achieve individual and collective objectives. Although t

safetyarxiv-cs-ai
28 May 2026
Safety

F Scott Fitzgerald, writing about someone who may as well have been Peter Thiel: “They were careless people, Tom and Daisy- they smashed up …

DGX agent

F Scott Fitzgerald, writing about someone who may as well have been Peter Thiel: “They were careless people, Tom and Daisy- they smashed up things and creatures and then retreated back into their mone

safetygary-marcus--x
28 May 2026
Safety

FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning

DGX agent

arXiv:2605.28389v1 Announce Type: new Abstract: While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own sol

safetyarxiv-cs-cl
28 May 2026
Safety

FedEHR-Gen: Federated Synthetic Time-Series EHR Generation via Latent Space Alignment and Distribution-Aware Aggregation

DGX agent

arXiv:2605.27892v1 Announce Type: new Abstract: Synthetic Electronic Health Record (EHR) generation provides a promising avenue for data augmentation and cross-hospital modeling in privacy-constrained

safetyarxiv-cs-lg
28 May 2026
Safety

From Affect to Complex Behavior: Advancing Multimodal Human-Centered AI at the 10th ABAW Workshop & Competition

DGX agent

arXiv:2605.27451v1 Announce Type: new Abstract: The 10th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop and Competition, held at CVPR 2026, continues to advance research on modelling, analy

safetyarxiv-cs-cv
28 May 2026
Safety

From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons

DGX agent

arXiv:2605.27387v1 Announce Type: cross Abstract: Diffusion models promise efficient parallel text generation but rely on bidirectional attention, creating a structural mismatch with pre-trained Autor

safetyarxiv-cs-ai
28 May 2026
Safety

From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constraints

DGX agent

arXiv:2605.28483v1 Announce Type: new Abstract: Linking learning resources to a structured competency framework is key to enabling competency-based search and curriculum analytics in Learning Manageme

safetyarxiv-cs-ai
28 May 2026
← Previous
1…136137138139140…267
Next →