AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,704 results
30 Apr 2026

Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents

SafetyDGX agent

arXiv:2505.02077v2 Announce Type: replace-cross Abstract: AI agents are beginning to interact with each other directly and across internet platforms and physical environments, creating security challe

Open Problems in Frontier AI Risk Management

SafetyDGX agent

arXiv:2604.25982v1 Announce Type: cross Abstract: Frontier AI both amplifies existing risks and introduces qualitatively novel challenges. Not only is there a notable lack of stable scientific consens

OpenAI’s lawyer plays dirty. Why I am not surprised?

SafetyDGX agent

OpenAI’s lawyer plays dirty. Why I am not surprised? 🚨 OpenAI's lawyer read a 2016 email from Sutskever to Musk to argue Musk knew OpenAI would go closed-source Sutskever: 'As we get closer to buildin


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Operating-Layer Controls for Onchain Language-Model Agents Under Real Capital

SafetyDGX agent

arXiv:2604.26091v1 Announce Type: new Abstract: We study reliability in autonomous language-model agents that translate user mandates into validated tool actions under real capital. The setting is DX

Oracle has dropped 50% since I declared that the OpenAI deal they made was “peak bubble”. And, as @edzitron lays out below, they may fall a …

SafetyDGX agent

Oracle has dropped 50% since I declared that the OpenAI deal they made was “peak bubble”. And, as @edzitron lays out below, they may fall a lot further yet. At which point the entire house of cards co

OT Score: An OT based Confidence Score for Prototype-Assisted Source Free Unsupervised Domain Adaptation

SafetyDGX agent

arXiv:2505.11669v3 Announce Type: replace-cross Abstract: We address the computational and theoretical limitations of current distributional alignment methods for source-free unsupervised domain adapt

Ouch! Current AI assistants often corrupt documents. Sounds like an intern you can’t trust — once again. A trillion dollar investment in sca…

SafetyDGX agent

Ouch! Current AI assistants often corrupt documents. Sounds like an intern you can’t trust — once again. A trillion dollar investment in scaling hasn’t solved this. New Microsoft paper shows that curr

PAINT: Partial-Solution Adaptive Interpolated Training for Self-Distilled Reasoners

SafetyDGX agent

arXiv:2604.26573v1 Announce Type: new Abstract: Improving large language model (LLM) reasoning requires supervision that is both aligned with the model's own test-time states and informative at the to

PBiLoss: Popularity-Aware Regularization to Improve Fairness in Graph-Based Recommender Systems

SafetyDGX agent

arXiv:2507.19067v2 Announce Type: replace-cross Abstract: Recommender systems based on graph neural networks (GNNs) have been proved to perform well on user-item interactions. However, they commonly s

Probe-then-Plan: Environment-Aware Planning for Industrial E-commerce Search

SafetyDGX agent

arXiv:2603.15262v2 Announce Type: replace Abstract: Modern e-commerce search is evolving to resolve complex user intents. While Large Language Models (LLMs) offer strong reasoning, existing LLM-based

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation

SafetyDGX agent

arXiv:2510.08547v2 Announce Type: replace-cross Abstract: Towards the aim of generalized robotic manipulation, spatial generalization is the most fundamental capability that requires the policy to wor

Recipes for Calibration Checks in Safety-Critical Applications

SafetyDGX agent

arXiv:2604.26479v1 Announce Type: cross Abstract: Safety-critical prediction systems, such as autonomous vehicles, weather forecasters, and medical monitors, commonly rely on probabilistic forecasters

Risk Reporting for Developers' Internal AI Model Use

SafetyDGX agent

arXiv:2604.24966v1 Announce Type: cross Abstract: Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a p

Roblox reports Q1 bookings up 43% YoY to 1.7B, vs. 1.73B est., and DAUs up 35% to 132M, below analysts' estimates of 143.8M; RBLX drops 16%+ after hours (Cecilia D'Anastasio/Bloomberg)

SafetyDGX agent

Cecilia D'Anastasio / Bloomberg: Roblox reports Q1 bookings up 43% YoY to 1.7B, vs. 1.73B est., and DAUs up 35% to 132M, below analysts' estimates of 143.8M; RBLX drops 16%+ after hours — Roblox Corp.

Robust Alignment: Harmonizing Clean Accuracy and Adversarial Robustness in Adversarial Training

SafetyDGX agent

arXiv:2604.26496v1 Announce Type: new Abstract: Adversarial Training (AT) is one of the most effective methods for developing robust deep neural networks (DNNs). However, AT faces a trade-off problem

Rule-based High-Level Coaching for Goal-Conditioned Reinforcement Learning in Search-and-Rescue UAV Missions Under Limited-Simulation Training

SafetyDGX agent

arXiv:2604.26833v1 Announce Type: cross Abstract: This paper presents a hierarchical decision-making framework for unmanned aerial vehicle (UAV) missions motivated by search-and-rescue (SAR) scenarios

SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling

SafetyDGX agent

arXiv:2604.26630v1 Announce Type: new Abstract: Effective mental health counseling is a complex, theory-driven process requiring the simultaneous integration of psychological frameworks, real-time dis

SD2AIL: Adversarial Imitation Learning from Synthetic Demonstrations via Diffusion Models

SafetyDGX agent

arXiv:2512.18583v2 Announce Type: replace Abstract: Adversarial Imitation Learning (AIL) is a dominant framework in imitation learning that infers rewards from expert demonstrations to guide policy op

Seeking Consensus: Geometric-Semantic On-the-Fly Recalibration for Open-Vocabulary Remote Sensing Semantic Segmentation

SafetyDGX agent

arXiv:2604.26221v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) in remote sensing images is a promising task that employs textual descriptions for identifying undefined

Sergey Brin-backed Building a Better California says two ballot countermeasures to the proposed wealth tax are on track to qualify for the November ballot (Laura J. Nelson/Wall Street Journal)

SafetyDGX agent

Laura J. Nelson / Wall Street Journal: Sergey Brin-backed Building a Better California says two ballot countermeasures to the proposed wealth tax are on track to qualify for the November ballot — Two

“shitting away money at scale”. @mcuban nails what I have been trying to say:

SafetyDGX agent

“shitting away money at scale”. @mcuban nails what I have been trying to say: .@mcuban is not impressed with OpenAI's business: 'They're shitting away money at scale,' he told me. Full discussion live

Sociodemographic Biases in Educational Counselling by Large Language Models

SafetyDGX agent

arXiv:2604.25932v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly integrated into educational settings, understanding their potential biases is critical. This study ex

Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection

SafetyDGX agent

arXiv:2604.26409v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have demonstrated significant success in interpreting Large Language Models (LLMs) by decomposing dense representations into

STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation

SafetyDGX agent

arXiv:2604.26848v1 Announce Type: new Abstract: Robotic manipulation critically requires reasoning about future spatial-temporal interactions, yet existing VLA policies and world-model-enhanced polici

Student Guides Teacher: Weak-to-Strong Inference via Spectral Orthogonal Exploration

SafetyDGX agent

arXiv:2601.06160v2 Announce Type: replace Abstract: Large Language Models (LLMs) often suffer from ''Reasoning Collapse'' on challenging mathematical reasoning tasks, where stochastic sampling produce

Talent or Luck? Evaluating Attribution Bias in Large Language Models

SafetyDGX agent

arXiv:2505.22910v2 Announce Type: replace Abstract: When a student fails an exam, do we tend to blame their effort or the test's difficulty? Attribution, defined as how reasons are assigned to event o

Tatemae: Detecting Alignment Faking via Tool Selection in LLMs

SafetyDGX agent

arXiv:2604.26511v1 Announce Type: cross Abstract: Alignment faking (AF) occurs when an LLM strategically complies with training objectives to avoid value modification, reverting to prior preferences o

Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards

SafetyDGX agent

arXiv:2510.04214v3 Announce Type: replace Abstract: We deploy large language models (LLMs) as business development (BD) agents for persuasive price negotiation in online travel agencies (OTAs). The ag

Test-Time Safety Alignment

SafetyDGX agent

arXiv:2604.26167v1 Announce Type: cross Abstract: Recent work has shown that a model's input word embeddings can serve as effective control variables for steering its behavior toward outputs that sati

Text Style Transfer with Machine Translation for Graphic Designs

SafetyDGX agent

arXiv:2604.26361v1 Announce Type: cross Abstract: Globalization of graphic designs such as those used in marketing materials and magazines is increasingly important for communication to broad audience

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

SafetyDGX agent

arXiv:2603.02259v2 Announce Type: replace-cross Abstract: Multi-agent systems provide mature methodologies for role decomposition, coordination, and normative governance, capabilities that remain esse

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

SafetyDGX agent

arXiv:2604.26347v1 Announce Type: cross Abstract: Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring em

The Zig project's rationale for their firm anti-AI contribution policy

SafetyDGX agent

Zig has one of the most stringent anti-LLM policies of any major open source project: No LLMs for issues. No LLMs for pull requests. No LLMs for comments on the bug tracker, including translation. Eng

Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning

SafetyDGX agent

arXiv:2604.26370v1 Announce Type: new Abstract: Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language lea

Translating Under Pressure: Domain-Aware LLMs for Crisis Communication

SafetyDGX agent

arXiv:2604.26597v1 Announce Type: cross Abstract: Timely and reliable multilingual communication is critical during natural and human-induced disasters, but developing effective solutions for crisis c

Uncertainty-Aware Information Pursuit for Interpretable and Reliable Medical Image Analysis

SafetyDGX agent

arXiv:2506.16742v3 Announce Type: replace Abstract: To be adopted in safety-critical domains like medical image analysis, AI systems must provide human-interpretable decisions. Variational Information

Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics

SafetyDGX agent

arXiv:2604.26836v1 Announce Type: new Abstract: Predictive safety filters (PSFs) leverage model predictive control to enforce constraint satisfaction during deep reinforcement learning (RL) exploratio

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking

SafetyDGX agent

arXiv:2604.26360v1 Announce Type: cross Abstract: Reinforcement learning (RL) systems typically optimize scalar reward functions that assume precise and reliable evaluation of outcomes. However, real-

Unifying Runtime Monitoring Approaches for Safety-Critical Machine Learning: Application to Vision-Based Landing

SafetyDGX agent

arXiv:2604.26411v1 Announce Type: new Abstract: Runtime monitoring is essential to ensure the safety of ML applications in safety-critical domains. However, current research is fragmented, with indepe

Verified Critical Step Optimization for LLM Agents

SafetyDGX agent

arXiv:2602.03412v2 Announce Type: replace Abstract: As large language model agents tackle increasingly complex long-horizon tasks, effective post-training becomes critical. Prior work faces fundamenta

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks

SafetyDGX agent

arXiv:2509.09870v2 Announce Type: replace-cross Abstract: Large language models (LLMs) enable conversational agents (CAs) to express distinctive personalities, raising new questions about how such des

ViBE: Visual-to-M/EEG Brain Encoding via Spatio-Temporal VAE and Distribution-Aligned Projection

SafetyDGX agent

arXiv:2604.26218v1 Announce Type: new Abstract: Brain encoding models not only serve to decipher how visual stimuli are transformed into neural responses, but also represent a critical step toward vis

ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers

SafetyDGX agent

arXiv:2505.20032v3 Announce Type: replace Abstract: Tactile sensing provides local essential information that is complementary to visual perception, such as texture, compliance, and force. Despite rec

Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance

SafetyDGX agent

arXiv:2604.26839v1 Announce Type: new Abstract: Assisting humans in open-world outdoor environments requires robots to translate high-level natural-language intentions into safe, long-horizon, and soc

What Kind of Language is Easy to Language-Model Under Curriculum Learning?

SafetyDGX agent

arXiv:2604.26844v1 Announce Type: new Abstract: Many of the thousands of attested languages share common configurations of features, creating a spectrum from typologically very rare (e.g., object-verb

When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity

SafetyDGX agent

arXiv:2510.17548v2 Announce Type: replace Abstract: Language models are often evaluated with scalar metrics like accuracy, but such measures fail to capture how models internally represent ambiguity,

// When to Retrieve During Reasoning // Pay attention to this one, AI devs. (bookmark it) Most RAG systems retrieve once, before the model s…

SafetyDGX agent

// When to Retrieve During Reasoning // Pay attention to this one, AI devs. (bookmark it) Most RAG systems retrieve once, before the model starts reasoning. Large reasoning models like o1 and R1 don't

Who is accountable? the latest from @marketoonist

SafetyDGX agent

This post likely discusses accountability in AI systems, featuring commentary or a cartoon from Marketoonist (a popular cartoonist who creates comics about business and technology). Gary Marcus, an AI

wow. incredible.

SafetyDGX agent

wow. incredible. A Chinese court has ruled it illegal to replace human workers with AI purely for the sake of cost-cutting. The court decided that companies hold a social responsibility to treat worke

Zuck is zucking up Meta

SafetyDGX agent

Zuck is zucking up Meta Mark Zuckerberg lacks the cloud-computing business to truly justify Meta's AI spending, @DaveLeeBBG says (via @opinion) https://www.bloomberg.com/opinion/articles/2026-04-30/me

29 Apr 2026

A Blueprint for AI-Driven Software Quality: Integrating LLMs with Established Standards

SafetyDGX agent

arXiv:2505.13766v5 Announce Type: replace-cross Abstract: Software Quality Assurance (SQA) is critical for delivering reliable, secure, and efficient software products. The Software Quality Assurance

A Deep Reinforcement Learning Approach to Automated Stock Trading, using xLSTM Networks

SafetyDGX agent

arXiv:2503.09655v2 Announce Type: replace-cross Abstract: Traditional Long Short-Term Memory (LSTM) networks are effective for handling sequential data but have limitations such as gradient vanishing

A Systematic Post-Train Framework for Video Generation

SafetyDGX agent

arXiv:2604.25427v1 Announce Type: new Abstract: While large-scale video diffusion models have demonstrated impressive capabilities in generating high-resolution and semantically rich content, a signif

agents are going and are already changing every industry & vertical, awesome blog from the @MadrigalPharma team on using Deep Agents, Skills…

SafetyDGX agent

agents are going and are already changing every industry & vertical, awesome blog from the @MadrigalPharma team on using Deep Agents, Skills, & LangSmith at the frontier of biopharma it’s an awesome t

AI-Residual Economies (Data → Income) Users begin to expect compensation for their contribution to AI systems, not just free usage. These pl…

SafetyDGX agent

AI-Residual Economies (Data → Income) Users begin to expect compensation for their contribution to AI systems, not just free usage. These platforms track, attribute, and pay micro-royalties when user

ANCHOR: A Physically Grounded Closed-Loop Framework for Robust Home-Service Mobile Manipulation

SafetyDGX agent

arXiv:2604.25323v1 Announce Type: new Abstract: Recent advances in open-vocabulary mobile manipulation have brought robots into real domestic environments. In such settings, reliable long-horizon exec

BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate

SafetyDGX agent

arXiv:2604.25203v1 Announce Type: new Abstract: Deploying guardrails for custom policies remains challenging, as generic safety models fail to capture task-specific requirements, while prompting LLMs

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

SafetyDGX agent

arXiv:2604.25072v1 Announce Type: new Abstract: Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evalua

Carbon-Taxed Transformers: A Green Compression Pipeline for Overgrown Language Models

SafetyDGX agent

arXiv:2604.25903v1 Announce Type: cross Abstract: The accelerating adoption of Large Language Models (LLMs) in software engineering (SE) has brought with it a silent crisis: unsustainable computationa

CHUCKLE -- When Humans Teach AI To Learn Emotions The Easy Way

SafetyDGX agent

arXiv:2510.09382v2 Announce Type: replace Abstract: Curriculum learning (CL) structures training from simple to complex samples, facilitating progressive learning. However, existing CL approaches for

← Previous
1…172173174175176…212
Next →