AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
29 May 2026

yes, absolutely, many companies are experimenting. but also: most of those experiments are failing to yield significant RoI. (weird for an e…

SafetyDGX agent

yes, absolutely, many companies are experimenting. but also: most of those experiments are failing to yield significant RoI. (weird for an economist to not even ask or address that question.) Looks li

you break it, you buy it peter thiel has broken the united states, and now he is abandoning it

SafetyDGX agent

you break it, you buy it peter thiel has broken the united states, and now he is abandoning it Peter Thiel has temporarily relocated his family to Argentina, enrolled his children in school there, and

28 May 2026

1. Agreed that OpenAI is in deep trouble; that’s why I have long suggested that it might be the WeWork of AI but 2. Anthropic is not out of …


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

1. Agreed that OpenAI is in deep trouble; that’s why I have long suggested that it might be the WeWork of AI but 2. Anthropic is not out of the woods; their best quarter was exactly when tokenmaxxing

A Comparative Study of Rule-Based and Data-Driven Approaches in Industrial Monitoring

SafetyDGX agent

arXiv:2509.15848v2 Announce Type: replace Abstract: Industrial monitoring systems, especially when deployed in Industry 4.0 environments, are experiencing a shift in paradigm from traditional rule-bas

A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons

SafetyDGX agent

arXiv:2605.27461v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies have shown promising manipulation capabilities, yet their practical impact is often limited by the reliability dem

A Policy-Driven Runtime Layer for Agentic LLM Serving

SafetyDGX agent

arXiv:2605.27744v1 Announce Type: new Abstract: Multi-agent LLM systems have become the dominant production workload, but the serving stack was not built for them. The agent framework above knows agen

A Structural Theory of Position Bias in Transformers

SafetyDGX agent

arXiv:2602.16837v2 Announce Type: replace Abstract: Transformer models systematically favor certain token positions, yet the architectural origins of this position bias remain poorly understood. This

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection

SafetyDGX agent

arXiv:2605.28664v1 Announce Type: cross Abstract: Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are sca

AdaMemento: Adaptive Memory-Assisted Policy Optimization for Reinforcement Learning

SafetyDGX agent

arXiv:2410.04498v2 Announce Type: replace Abstract: In sparse reward scenarios of reinforcement learning (RL), the memory mechanism provides promising shortcuts to policy optimization by reflecting on

ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation

SafetyDGX agent

arXiv:2605.28396v1 Announce Type: cross Abstract: On-policy distillation (OPD) transfers reasoning behavior by training a student on teacher feedback along student-generated trajectories, but standard

Affective Music Recommendation: A Rollout-Based World Model for Offline Preference Optimization

SafetyDGX agent

arXiv:2605.28810v1 Announce Type: new Abstract: Functional music applications, from consumer focus and sleep aids to clinical interventions, share a distinctive recommendation problem: success is defi

AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems

SafetyDGX agent

arXiv:2605.27466v1 Announce Type: cross Abstract: Multi-agent systems built on large language models (LLMs) require many coordination choices that are difficult to fix a priori: which skill protocol t

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning

SafetyDGX agent

arXiv:2605.28774v1 Announce Type: new Abstract: Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning

Agents that Matter: Optimizing Multi-Agent LLMs via Removal-Based Attribution

SafetyDGX agent

arXiv:2605.27621v1 Announce Type: cross Abstract: As multi-agent systems (MAS) become increasingly complex, identifying the contributions of individual agents is critical for system optimization. Howe

AI systems may soon help run economies, infrastructure, and military operations. But these systems are not reliably loyal or secure. An adve…

SafetyDGX agent

AI systems may soon help run economies, infrastructure, and military operations. But these systems are not reliably loyal or secure. An adversary can make an AI work against its own operator. In our n

AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

SafetyDGX agent

arXiv:2605.28255v1 Announce Type: new Abstract: AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration re

all these decades i have been using calculators, and it never occurred to me that they might be … conscious!

SafetyDGX agent

Gary Marcus raises a philosophical question about whether calculators, despite their computational capabilities, might possess consciousness. The post likely explores the disconnect between functional

Almost everyone is building agent harness systems the wrong way. The default move: pick LangChain or LangGraph or the OpenAI Agents SDK, acc…

SafetyDGX agent

Almost everyone is building agent harness systems the wrong way. The default move: pick LangChain or LangGraph or the OpenAI Agents SDK, accept the loop, the tools, the memory, the orchestration, the

ANTHROPIC VALUATION: 965B WALMART VALUATION: 940B ANTHROPIC REVENUE: 20B WALMART REVENUE: 725B BUT AI IS NOT A BUBBLE, RIGHT?

SafetyDGX agent

ANTHROPIC VALUATION: 965B WALMART VALUATION: 940B ANTHROPIC REVENUE: 20B WALMART REVENUE: 725B BUT AI IS NOT A BUBBLE, RIGHT? Media JUST IN: Anthropic raises 65B at 965B valuation 70% chance of IPO th

AOE: Exhaustive Out-of-Distribution Detection via Recalibrating Outlier Labels

SafetyDGX agent

arXiv:2605.28021v1 Announce Type: new Abstract: Out-of-distribution (OOD) detection is essential for deploying machine learning models in open-world and safety-critical scenarios, where test inputs ma

AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning

SafetyDGX agent

arXiv:2605.28809v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) is important in building real-world learning systems. In CLIP-based CIL, the model performs classification by comparing

Artemis: Structured Visual Reasoning for Perception Policy Learning

SafetyDGX agent

arXiv:2512.01988v2 Announce Type: replace Abstract: Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural languag

Auditable Decision Models with Learned Abstention and Real-Time Steering

SafetyDGX agent

arXiv:2605.27768v1 Announce Type: new Abstract: Production AI systems often operate with incomplete, conflicting, or insufficient evidence. Forced classifiers collapse such cases into action labels, w

Auditing Stance Asymmetry in Generative Explanations

SafetyDGX agent

arXiv:2605.27988v1 Announce Type: new Abstract: Bias evaluation for language models has made substantial progress on bounded comparisons, such as overt derogation, stereotype association, or label-sen

Automated Estimation of Impact Time, Impact Location, and Shuttlecock Speed in Badminton Smashes Using Event Cameras

SafetyDGX agent

arXiv:2605.28011v1 Announce Type: new Abstract: Quantifying impact phenomena in badminton smashes is important for evaluating both athletic performance and equipment; however, conventional measurement

AWS CEO Matt Garman: The idea that AI will replace junior developers is “the dumbest thing I have ever heard.”

SafetyDGX agent

AWS CEO Matt Garman stated that the notion AI will replace junior developers is fundamentally misguided, suggesting instead that AI tools will augment and enhance developer productivity rather than el

Bayesian Deployment Approval for Learned Landing Controllers under Finite Rollout Validation

SafetyDGX agent

arXiv:2605.27720v1 Announce Type: new Abstract: Reinforcement learning and data-driven autonomous controllers are commonly evaluated using cumulative reward and empirical success frequency under finit

Bayesian Gated Non-Negative Contrastive Learning

SafetyDGX agent

arXiv:2605.28441v1 Announce Type: cross Abstract: While Contrastive Learning (CL) has revolutionized self-supervised representation learning, its latent representations remain highly entangled and opa

Becoming an AI-native company is existential for every business, says Dell CMO

SafetyDGX agent

Becoming an AI-native company is no longer the competitive advantage it was five minutes ago. Now, it’s an empirical obligation, and Dell Technologies Inc. made that case to customers at its annual fl

Behavioural Analysis of Alignment Faking

SafetyDGX agent

arXiv:2605.27681v1 Announce Type: new Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deploym

Beyond Binary: Sim-to-Real Dexterous Manipulation with Physics-Grounded Contact Representation

SafetyDGX agent

arXiv:2605.28812v1 Announce Type: cross Abstract: A primary bottleneck in contact-rich manipulation is the difficulty of collecting real-world data. Sim-to-real reinforcement learning offers a scalabl

BiasEdit: A Training-Free Bias-Detect-and-Edit Framework for Learning Fair Visual Classifiers

SafetyDGX agent

arXiv:2605.28450v1 Announce Type: cross Abstract: Visual data from the Web power image classifiers, which often underpin many web services, such as recommendation and content moderation. However, the

Blind PRNG Hijacking: An Undetectable Integrity-Preserving Attack Against LLM Watermarking

SafetyDGX agent

arXiv:2605.28632v1 Announce Type: cross Abstract: Cryptographic watermarking is a leading defense for attributing text generated by large language models (LLMs). Existing schemes, including KGW, Unigr

Boundary Suppression Asymmetry in Post-trained Assistants: Over-expansion as a Controllability Cost

SafetyDGX agent

arXiv:2605.27969v1 Announce Type: new Abstract: Post-trained language-model assistants are often optimized to avoid under-answering, encouraging complete, helpful, cautious, and proactive responses. W

BPPO: Binary Prefix Policy Optimization for Efficient GRPO-Style Reasoning RL with Concise Responses

SafetyDGX agent

arXiv:2605.28028v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is widely used for training reasoning models, but updating all sampled completions in each group incurs substa

Breaking the Script Barrier: Enabling Automatic Alignment for PoS-based ASR Error Analysis in Non-Latin Scripts

SafetyDGX agent

arXiv:2605.28438v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) systems are commonly evaluated using aggregate metrics such as Word Error Rate (WER), which do not capture the lingui

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information

SafetyDGX agent

arXiv:2605.28070v1 Announce Type: new Abstract: We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified

Calibrated Inference for the Conditional Average Treatment Effect in the Few-Placebo Regime via Gaussian Processes

SafetyDGX agent

arXiv:2605.27473v1 Announce Type: cross Abstract: Estimating how much an intervention helps a given individual the conditional average treatment effect (CATE) is increasingly central to decision-makin

CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation

SafetyDGX agent

arXiv:2412.08052v2 Announce Type: replace Abstract: Off-policy evaluation (OPE) is critical for applying contextual bandit algorithms to high-stakes decision-making settings such as healthcare, where

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

SafetyDGX agent

arXiv:2603.22335v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior dis

Causal Machine Learning: A Survey and Open Problems

SafetyDGX agent

arXiv:2206.15475v3 Announce Type: replace Abstract: Causal Machine Learning (CausalML) is an umbrella term for machine learning methods that formalize the data-generation process as a structural causa

Chance-Constrained MPPI under State and Dynamic Object Prediction Uncertainty and the Evaluation of Collision Risk Calibration

SafetyDGX agent

arXiv:2605.28330v1 Announce Type: new Abstract: Chance-constrained Model Predictive Path Integral (MPPI) control is increasingly adopted for navigation in dynamic environments to explicitly bound coll

CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models

SafetyDGX agent

arXiv:2605.28292v1 Announce Type: new Abstract: Implicit Chain-of-Thought (CoT) reduces the inference cost of large language models by internalizing the explicit rationales. However, existing approach

CodeGENCAT: Generative Computerized Adaptive Testing for Open-ended Coding Problems

SafetyDGX agent

arXiv:2602.20020v2 Announce Type: replace Abstract: Existing Computerized Adaptive Testing (CAT) frameworks typically select questions based on the predicted likelihood that the student will answer co

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

SafetyDGX agent

arXiv:2602.15198v2 Announce Type: replace-cross Abstract: Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperativ

Commit to the Bit: Reactive Reinforcement Learning Done Right

SafetyDGX agent

arXiv:2605.28276v1 Announce Type: new Abstract: Reinforcement learning algorithms are commonly analyzed (and designed) under the Markov assumption. This is unrealistic, as most environments encountere

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization

SafetyDGX agent

arXiv:2605.28615v1 Announce Type: new Abstract: Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bi

Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry

SafetyDGX agent

arXiv:2605.27952v1 Announce Type: new Abstract: Visual odometry (VO) is a fundamental component in robotics and augmented reality. RGB-D direct VO benefits from metric depth measurements, but it can d

COTTA: Context-Aware Transfer Adaptation for Trajectory Prediction in Autonomous Driving

SafetyDGX agent

arXiv:2604.00402v2 Announce Type: replace-cross Abstract: Developing robust models to accurately predict the trajectories of surrounding agents is fundamental to autonomous driving safety. However, mo

Counterfactually Fair Regression via Optimal Transport

SafetyDGX agent

arXiv:2605.28251v1 Announce Type: cross Abstract: We consider the problem of learning a counterfactually fair regressor. We adopt a causal uncertainty view in which counterfactual fairness is defined

CPPO: Contrastive Perception Policy Optimization for VLM Agents

SafetyDGX agent

arXiv:2601.00501v2 Announce Type: replace Abstract: We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core

CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders

SafetyDGX agent

arXiv:2604.01604v2 Announce Type: replace Abstract: While modern LLMs are aligned to refuse harmful requests, it is essential to understand the underlying mechanistic basis of this refusal behavior fo

crazy that this was announced on the day tokenmaxxing died.

SafetyDGX agent

crazy that this was announced on the day tokenmaxxing died. Anthropic raised 65 billion in a funding round that valued the artificial intelligence company at 965 billion including the new investment,

Cross-Entropy Games and Frost Training

SafetyDGX agent

arXiv:2605.27701v1 Announce Type: new Abstract: We present Frost Training, a method for improving Monte Carlo-based policy optimization for a large family of LLM-as-a-judge tasks called Cross-Entropy

Cyberbullying Governance on Social Media: A Unified Framework from Content Identification to Intervention

SafetyDGX agent

arXiv:2605.27584v1 Announce Type: new Abstract: The proliferation of social media platforms and online communities has inadvertently catalyzed the spread of cyberbullying, hate speech, and other forms

DebFilter: Eradicating Biases Stashed in Value

SafetyDGX agent

arXiv:2605.28167v1 Announce Type: new Abstract: Text-to-image diffusion models, which are theoretically equivalent to score-based generative models, generate images through a multi-step denoising proc

Deconstructing Spatial Complexity: Hierarchical Decomposition for LLM Spatial Reasoning

SafetyDGX agent

arXiv:2605.28144v1 Announce Type: new Abstract: LLMs have shown remarkable proficiency in general language understanding and reasoning. However, they consistently underperform in spatial reasoning tha

Delay-Aware Reinforcement Learning for Highway On-Ramp Merging under Stochastic Communication Latency

SafetyDGX agent

arXiv:2403.11852v5 Announce Type: replace-cross Abstract: Delayed and partially observable state information poses significant challenges for reinforcement learning (RL)-based control in real-world au

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes

SafetyDGX agent

arXiv:2605.28421v1 Announce Type: new Abstract: Reinforcement learning has become a central paradigm for advancing reasoning in large language models, yet most existing methods still depend on stronge

Diagnosing Live Within-Policy Instruction Conflicts in LLM Agents with Witnessed Resolution Profiles

SafetyDGX agent

arXiv:2605.27784v1 Announce Type: new Abstract: LLM agents are governed by long-lived natural-language prompt policies, but individually reasonable standing rules can interact in uninspected ways. We

← Previous
1…108109110111112…214
Next →