AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,490 results
Safety

Human-like in-group bias in instruction-tuned language model agents

DGX agent

arXiv:2605.28114v1 Announce Type: new Abstract: As autonomous AI agents are deployed in persistent, interacting networks -- coordinating tasks, routing resources, and accumulating reputational histori

safetyarxiv-cs-ai
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

I read this garbage (in a big UK newspaper) when Google can’t even count reliably and wonder why people don’t spend more time learning about…

DGX agent

I read this garbage (in a big UK newspaper) when Google can’t even count reliably and wonder why people don’t spend more time learning about AI’s actual strengths and weakness before running their mou

safetygary-marcus--x
28 May 2026
Safety

ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment

DGX agent

arXiv:2605.27374v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) and diffusion models (DMs) have opened new possibilities for AI-generated content. Yet, pers

safetyarxiv-cs-cl
28 May 2026
Safety

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

DGX agent

arXiv:2601.04716v3 Announce Type: replace Abstract: While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing q

safetyarxiv-cs-cl
28 May 2026
Safety

Imitating and Finetuning Model Predictive Control for Robust and Symmetric Quadrupedal Locomotion

DGX agent

arXiv:2311.02304v3 Announce Type: replace Abstract: Control of legged robots is a challenging problem that has been investigated by different approaches, such as model-based control and learning algor

safetyarxiv-cs-ro
28 May 2026
Safety

Important context for latest OpenAI announcement. Especially (5:30): 'The model spit out a long transcript of an answer. Then a team of expe…

DGX agent

Important context for latest OpenAI announcement. Especially (5:30): 'The model spit out a long transcript of an answer. Then a team of expert mathematicians poured over this [transcript] and identifi

safetygary-marcus--x
28 May 2026
Safety

IMU Propagation as Preintegration

DGX agent

arXiv:2605.28279v1 Announce Type: new Abstract: IMU preintegration is widely used in factor-graph-based visual--inertial, lidar--inertial, and radar--inertial state estimation, yet it is often treated

safetyarxiv-cs-ro
28 May 2026
Safety

Information-theoretic Multimodal Representation Learning for Electrocardiogram Signals

DGX agent

arXiv:2605.27583v1 Announce Type: new Abstract: Electrocardiograms (ECGs) are widely used non-invasive measurements of cardiac activity and play a central role in clinical diagnosis. Recent multimodal

safetyarxiv-cs-lg
28 May 2026
Safety

Informing AI Policy Assessment using Large-Scale Simulation of Interventions

DGX agent

arXiv:2605.27395v1 Announce Type: cross Abstract: As the rapid proliferation of AI systems and harms spurs efforts in AI governance around the world, prioritizing among competing policy options has be

safetyarxiv-cs-ai
28 May 2026
Safety

Insurance Pricing Optimization via Off-Policy Evaluation

DGX agent

arXiv:2605.28327v1 Announce Type: cross Abstract: Traditional insurance pricing relies on risk-based principles that ensure actuarial fairness and solvency but do not explicitly account for policyhold

safetyarxiv-cs-lg
28 May 2026
Safety

JECA^2: Judgment-Explanation Consistent Adversarial Attack against Forensic Vision-Language Models

DGX agent

arXiv:2605.28609v1 Announce Type: new Abstract: Forensic vision-language models (VLMs) have recently been developed to detect image tampering and provide natural-language explanations. However, their

safetyarxiv-cs-cv
28 May 2026
Safety

Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration

DGX agent

arXiv:2605.28184v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as the standard paradigm for improving reasoning capability of large language models,

safetyarxiv-cs-lg
28 May 2026
Safety

just imagine what will happen to the economy and people’s retirement funds if these projections from FT are correct. brace for bailouts.

DGX agent

just imagine what will happen to the economy and people’s retirement funds if these projections from FT are correct. brace for bailouts. The AI numbers are starting to look very ugly. Even under 'best

safetygary-marcus--x
28 May 2026
Safety

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs

DGX agent

arXiv:2507.06999v2 Announce Type: replace-cross Abstract: Reasoning is essential for large language models (LLMs), especially in complex tasks such as mathematical problem solving. However, multimodal

safetyarxiv-cs-cl
28 May 2026
Safety

Learning High-Dimensional Parity Functions with Product Networks using Gradient Descent

DGX agent

arXiv:2605.28612v1 Announce Type: new Abstract: Parity functions are fundamental Boolean operations with critical applications across machine learning, cryptography, and error correction. Yet, learnin

safetyarxiv-cs-lg
28 May 2026
Safety

Learning to Assign Prediction Tasks to Agents with Capacity Constraints

DGX agent

arXiv:2605.27999v1 Announce Type: cross Abstract: We address the problem of learning to assign prediction tasks to one agent from a set of available human or AI agents. In particular, we focus on the

safetyarxiv-cs-ai
28 May 2026
Safety

Learning to Bid in Repeated Second-Price Auctions with Dynamic Values and Aggregated Feedback

DGX agent

arXiv:2605.28133v1 Announce Type: new Abstract: We study the problem of learning to bid when the bidder's value is dynamic, i.e., when the current value depends on past outcomes. Specifically, we cons

safetyarxiv-cs-lg
28 May 2026
Safety

Learning with Importance Weighted Variational Inference

DGX agent

arXiv:2410.12035v2 Announce Type: replace-cross Abstract: Several variational bounds involving importance weighting ideas generalize the Evidence Lower BOund (ELBO) for marginal likelihood optimizatio

safetyarxiv-cs-lg
28 May 2026
Safety

Let Relations Speak: An End-to-End LLM-GNN Soft Prompt Framework for Fraud Detection

DGX agent

arXiv:2605.28524v1 Announce Type: new Abstract: In recent years, Large Language Models (LLMs) have shown great capability in processing graph tasks such as fraud detection. However, most existing meth

safetyarxiv-cs-ai
28 May 2026
Safety

LLM Watermark Evasion via Bias Inversion

DGX agent

arXiv:2509.23019v5 Announce Type: replace-cross Abstract: Watermarking offers a promising solution for detecting LLM-generated content, yet its robustness under realistic query-free (black-box) evasio

safetyarxiv-cs-ai
28 May 2026
Safety

Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization

DGX agent

arXiv:2605.28109v1 Announce Type: new Abstract: Recent advances in online reinforcement learning (RL) for large language models (LLMs) have demonstrated promising performance in complex reasoning task

safetyarxiv-cs-lg
28 May 2026
Safety

Mag-VLA: Vision-Language-Action Model for Bimanual Magnetically Actuated Microrobot Manipulation

DGX agent

arXiv:2605.28486v1 Announce Type: new Abstract: Magnetically actuated microrobots have been used as wireless, non-contact manipulation tools at microscales, making them promising for minimally invasiv

safetyarxiv-cs-ro
28 May 2026
Safety

Mathematical Modelling of Ethical AI Use in Higher Education: A Coordination Game Framework for Future-Facing Learning

DGX agent

arXiv:2605.27400v1 Announce Type: cross Abstract: The rapid uptake of generative artificial intelligence (AI) in higher education is reshaping assessment practices and intensifying concerns around aca

safetyarxiv-cs-ai
28 May 2026
Safety

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification

DGX agent

arXiv:2605.28604v1 Announce Type: cross Abstract: Identifying key individuals in video scenes is essential for applications such as automated video editing and intelligent surveillance. Current method

safetyarxiv-cs-ai
28 May 2026
Safety

Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation

DGX agent

arXiv:2605.12515v2 Announce Type: replace Abstract: Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's langua

safetyarxiv-cs-cl
28 May 2026
Safety

Mobile-Aptus: Confidence-Driven Proactive and Robust Interaction in MLLM-based Mobile-Using Agents

DGX agent

arXiv:2605.28629v1 Announce Type: new Abstract: Recent advancements in multimodal large language models (MLLMs) have shown exceptional potential in enabling mobile-using agents to autonomously execute

safetyarxiv-cs-cl
28 May 2026
Safety

Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities

DGX agent

arXiv:2605.27388v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly utilized as proxies for computational social analysis; yet, their ability to faithfully represent the 't

safetyarxiv-cs-ai
28 May 2026
Safety

Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction

DGX agent

arXiv:2605.27878v1 Announce Type: new Abstract: Large language models produce fluent fiction, yet their creative output is widely seen as flat. We ask where this quality originates in the training and

safetyarxiv-cs-cl
28 May 2026
Safety

NCSAM Noise-Compensated Sharpness-Aware Minimization for Noisy Label Learning

DGX agent

arXiv:2601.19947v2 Announce Type: replace-cross Abstract: Learning from Noisy Labels (LNL) remains a fundamental challenge in deep learning because real-world datasets often contain corrupted annotati

safetyarxiv-cs-ai
28 May 2026
Safety

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

DGX agent

arXiv:2603.01766v2 Announce Type: replace Abstract: Despite the rapid progress of vision-language-action (VLA) models, the prevailing practice of predicting action chunks as discrete waypoints remains

safetyarxiv-cs-ro
28 May 2026
Safety

Off-Policy Learning to Reason Works Because It Is More Pessimistic Than You Think

DGX agent

arXiv:2605.28150v1 Announce Type: new Abstract: Large scale reinforcement learning has become a central tool for improving reasoning in large language models. At this scale, generation is often lagged

safetyarxiv-cs-lg
28 May 2026
Safety

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning

DGX agent

arXiv:2604.18530v2 Announce Type: replace Abstract: Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet m

safetyarxiv-cs-ai
28 May 2026
Safety

@OpenAI Foundation just launched with a 25B commitment and equity in OpenAI valued at ~130B. Sounds like a philanthropy giant. But look cl…

DGX agent

@OpenAI Foundation just launched with a 25B commitment and equity in OpenAI valued at ~130B. Sounds like a philanthropy giant. But look closer: the big money goes to programs they control internally (

safetygary-marcus--x
28 May 2026
Safety

Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensitive Deployment Conditions -- A Governance Framework for High-Stakes AI Systems

DGX agent

arXiv:2605.27827v1 Announce Type: new Abstract: AI governance frameworks increasingly emphasize fairness, transparency, accountability, and lifecycle risk management in high-stakes domains. However, m

safetyarxiv-cs-ai
28 May 2026
Safety

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

DGX agent

arXiv:2605.28675v1 Announce Type: new Abstract: Data acquisition efficiency is a central challenge in deploying reinforcement learning in business and healthcare operations, where interactions are cos

safetyarxiv-cs-lg
28 May 2026
Safety

Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains

DGX agent

arXiv:2605.28345v1 Announce Type: new Abstract: Progress in Prognostics and Health Management (PHM) is hindered by the lack of standardized and reusable evaluation practices across tasks, datasets, an

safetyarxiv-cs-ai
28 May 2026
Model Releases

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

DGX agent

arXiv:2605.28201v1 Announce Type: new Abstract: Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into ext

model-releasesarxiv-cs-ai
28 May 2026
Safety

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

DGX agent

arXiv:2602.01745v2 Announce Type: replace-cross Abstract: Token-level reweighting is a simple yet effective mechanism for controlling supervised fine-tuning, but common indicators are largely one-dime

safetyarxiv-cs-ai
28 May 2026
Safety

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

DGX agent

arXiv:2510.06974v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases

safetyarxiv-cs-cl
28 May 2026
Safety

ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation

DGX agent

arXiv:2605.28293v1 Announce Type: cross Abstract: Proactive Recommender Systems (PRSs) aim to guide user preference shift toward target items by generating paths of intermediate recommendations. Reinf

safetyarxiv-cs-ai
28 May 2026
Safety

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

DGX agent

arXiv:2602.15894v2 Announce Type: replace Abstract: In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, exist

safetyarxiv-cs-cl
28 May 2026
Safety

RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?

DGX agent

arXiv:2605.27436v1 Announce Type: cross Abstract: Multimodal alignment is critical for bridging the semantic gap in information retrieval. However, traditional pairwise strategies introduce a geometri

safetyarxiv-cs-ai
28 May 2026
Safety

Reasoning Matters: Mitigate Hallucination in Multimodal Large Reasoning Models via Reasoning-Conditioned Preference Optimization

DGX agent

arXiv:2605.27906v1 Announce Type: new Abstract: Multimodal Large Reasoning Models introduce the reasoning paradigm, demonstrating strong capabilities on complex vision-language tasks. However, they st

safetyarxiv-cs-ai
28 May 2026
Safety

Refining Multidimensional Video Reward Models via Disentangled Influence Functions

DGX agent

arXiv:2605.28203v1 Announce Type: new Abstract: As Text-to-Video (T2V) generation models continue to evolve, the complexity of video evaluation necessitates a fine-grained assessment across various ax

safetyarxiv-cs-lg
28 May 2026
Safety

Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning

DGX agent

arXiv:2605.27765v1 Announce Type: cross Abstract: Self-Distillation Policy Optimization (SDPO) provides dense token-level credit assignment for reinforcement learning with large language models by lev

safetyarxiv-cs-ai
28 May 2026
Safety

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

DGX agent

arXiv:2605.27881v1 Announce Type: new Abstract: Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reaso

safetyarxiv-cs-cl
28 May 2026
Safety

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure

DGX agent

arXiv:2605.27996v1 Announce Type: new Abstract: Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto cor

safetyarxiv-cs-ai
28 May 2026
Safety

Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach

DGX agent

arXiv:2605.27834v1 Announce Type: new Abstract: We study the transfer of rewards learned using inverse reinforcement learning from expert demonstrations in one environment to reinforcement learning in

safetyarxiv-cs-lg
28 May 2026
← Previous
1…176177178179180…302
Next →