AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,814 results
1 Jun 2026

What Am I Missing? Question-Answering as Hidden State Probing

SafetyDGX agent

arXiv:2605.31561v1 Announce Type: new Abstract: Test-time reasoning has become a significant field of study since the introduction of chain-of-thought reasoning in large language models (LLMs). Howeve

What if the skeptics are right and superintelligence is impossible? Great! Then a proactive ban costs us nothing, prevents massive compute w…

SafetyDGX agent

What if the skeptics are right and superintelligence is impossible? Great! Then a proactive ban costs us nothing, prevents massive compute waste, and hurts no one. But if they are wrong? We face an un

When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?

SafetyDGX agent

arXiv:2605.30719v1 Announce Type: cross Abstract: We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.e., when can we


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

which is NOT new; see this quote from 5 years ago. the fact that all this is still true says a lot.

SafetyDGX agent

which is NOT new; see this quote from 5 years ago. the fact that all this is still true says a lot. This was right five years ago, and still is: “Large scale pretrained models are certainly likely to

Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems

SafetyDGX agent

arXiv:2506.00175v5 Announce Type: replace-cross Abstract: Modern AI systems are typically developed through multiple stages-pretraining, fine-tuning rounds, and subsequent adaptation or alignment, whe

Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning

SafetyDGX agent

arXiv:2605.31261v1 Announce Type: cross Abstract: The family of linear recurrent neural networks has shown strong performance as recurrent memory units in partially observable reinforcement learning.

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

SafetyDGX agent

arXiv:2604.01985v2 Announce Type: replace-cross Abstract: General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness re

World2Act: Latent Action Post-Training from World Model Dynamics

SafetyDGX agent

arXiv:2603.10422v2 Announce Type: replace Abstract: World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve gen

Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

SafetyDGX agent

arXiv:2605.30833v1 Announce Type: cross Abstract: On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from

ZAPS-DA: Zero-Phase Action Policy Smoothing with Decoupled Actor for Continuous Control in Reinforcement Learning

SafetyDGX agent

arXiv:2605.30612v1 Announce Type: cross Abstract: Continuous control policies trained with off-policy reinforcement learning frequently exhibit high-frequency action jitter, rendering direct deploymen

Zero Collapse: A Failure Mode of Policy Gradient Methods in Discontinuous Reward Environments

SafetyDGX agent

arXiv:2605.30896v1 Announce Type: new Abstract: Bidding in repeated auctions is a central challenge for reinforcement learning (RL), combining continuous control with the strategic complexities of dig

31 May 2026

Anyone remember how I said in January 2025 that AI was going to be stumble and be dubbed “too big to fail”, along with cries for bailouts? I…

SafetyDGX agent

Anyone remember how I said in January 2025 that AI was going to be stumble and be dubbed “too big to fail”, along with cries for bailouts? If that call was correct – which increasingly seems likely, a

further discussion here: https://open.substack.com/pub/garymarcus/p/the-pope-appears-to-understand-ai?r=8tdk6&utm_campaign=post-expanded-sha…

SafetyDGX agent

Gary Marcus discusses the Pope's understanding and perspective on artificial intelligence, examining statements or positions the religious leader has taken regarding AI technology and its implications

I am absolutely with @GaryMarcus and the Pope on this. (Not something you expect to say everyday). We are not creating beings. Systems don’t…

SafetyDGX agent

I am absolutely with @GaryMarcus and the Pope on this. (Not something you expect to say everyday). We are not creating beings. Systems don’t exp. grief nor hope, hold a value construct, or the ability

Listen, I’m a big “the index is the index” guy, but they are openly looting the coffers. This is 100% fraud.

SafetyDGX agent

Listen, I’m a big “the index is the index” guy, but they are openly looting the coffers. This is 100% fraud. Rule changes for the SpaceX SPCX IPO: Index providers waived the profitability requirement

@ParValue26 @Hedgeye Let’s be clear what this actually is: the administration and the world’s richest man working together to screw over ord…

SafetyDGX agent

@ParValue26 @Hedgeye Let’s be clear what this actually is: the administration and the world’s richest man working together to screw over ordinary investors to goose returns on the IPO for the wealthie

serious accusation. does this fit with people’s experience?

SafetyDGX agent

serious accusation. does this fit with people’s experience? Is Anthropic altering model performance to force costly upgrades? Chapter Co-Founder and CEO @CobiBGantz outlines a shift his team recently

SpaceX being rammed into indices with no profit requirements, seasoning, and generally looser constraints is economic terrorism. Index track…

SafetyDGX agent

SpaceX being rammed into indices with no profit requirements, seasoning, and generally looser constraints is economic terrorism. Index trackers will eat the loss when reality catches up and retail inv

The backlash against AI - generated content is so visceral - expect the 'human-authored' certification gain momentum, esp for fiction. https…

SafetyDGX agent

Gary Marcus argues that consumer backlash against AI-generated content will be increasingly visceral, particularly in creative fields like fiction, leading to growing demand for 'human-authored' certi

Three companies are about to IPO at a higher combined value than all 2,600 dot-com IPOs from 1995 to 2000 combined. – SpaceX, OpenAI, Anthro…

SafetyDGX agent

Three companies are about to IPO at a higher combined value than all 2,600 dot-com IPOs from 1995 to 2000 combined. – SpaceX, OpenAI, Anthropic (2026): ~3.75 trillion – Every dot-com IPO from 1995–200

Weird how the Pope seems to understand AI better than @geoffreyhinton, but I am 100% with the Pope on this. We are NOT creating beings. The …

SafetyDGX agent

Weird how the Pope seems to understand AI better than @geoffreyhinton, but I am 100% with the Pope on this. We are NOT creating beings. The Pope is right. We are creating interactive fiction that is t

weird the way this tweet was getting a ton of traffic and then just stopped. 🤷‍♂️

SafetyDGX agent

weird the way this tweet was getting a ton of traffic and then just stopped. 🤷‍♂️ I honestly think Elon’s best days are behind him: BYD is crushing Tesla in EVs. Waymo is crushing Tesla in AVs. Anthro

30 May 2026

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the o…

SafetyDGX agent

AI safety can't happen behind closed doors! Super cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the open on @huggingface, so researchers everywhere can scrutiniz

calling someone a retard when you don’t how apostrophes work 🙄

SafetyDGX agent

This post likely critiques the irony of someone insulting another person's intelligence while making a basic grammatical error (omitting an apostrophe in 'don't'). Gary Marcus, a cognitive scientist a

genius reply, full of reasoned intellectual argument, from the kind of retail investor who is probably going to get burned on the SpaceX IPO…

SafetyDGX agent

genius reply, full of reasoned intellectual argument, from the kind of retail investor who is probably going to get burned on the SpaceX IPO. Going to MIT and NYU may have meant something in the past,

i am only blocking tesla supporters that come at me with insult rather than argument. but that’s a lot of them. and a not great sign for the…

SafetyDGX agent

Gary Marcus discusses his moderation approach on social media, noting that he blocks Tesla supporters primarily when they resort to insults rather than substantive arguments, and observes that this oc

I honestly think Elon’s best days are behind him: BYD is crushing Tesla in EVs. Waymo is crushing Tesla in AVs. Anthropic, Openai, and Googl…

SafetyDGX agent

I honestly think Elon’s best days are behind him: BYD is crushing Tesla in EVs. Waymo is crushing Tesla in AVs. Anthropic, Openai, and Google are crushing Xai on AI. The SpaceX S-1 is so ridiculous th

“I'm more concerned about the lack of intellectual diversity within the frontier AI commentariat/research world. This improved a lot over th…

SafetyDGX agent

“I'm more concerned about the lack of intellectual diversity within the frontier AI commentariat/research world. This improved a lot over the last two years, but we're still far from a healthy ecosyst

is having a four month lead a sustainable multitrillion dollar business model?

SafetyDGX agent

is having a four month lead a sustainable multitrillion dollar business model? We took another look at the capability gap between open-weight and proprietary models. Since the start of the year, open-

it is the single most clever move Elon ever pulled off

SafetyDGX agent

it is the single most clever move Elon ever pulled off Rule changes for the SpaceX SPCX IPO: Index providers waived the profitability requirement and cut the seasoning window from 90 days to 5. This f

More of my thoughts on the topic in @Corriere: https://www.corriere.it/esteri/26_maggio_29/yoshua-bengio-ai-pope-6aa95de2-33f7-4f73-8905-34b…

SafetyDGX agent

More of my thoughts on the topic in @Corriere: https://www.corriere.it/esteri/26_maggio_29/yoshua-bengio-ai-pope-6aa95de2-33f7-4f73-8905-34b1abedbxlk.shtml “Like nuclear energy, AI must be at the serv

My feed is suddenly filled w Elon supporters who can’t understand the difference between an insult and a genuine counterargument. As always,…

SafetyDGX agent

Gary Marcus expresses frustration about an influx of Elon Musk supporters on his social media feed who conflate personal insults with substantive counterarguments. The post appears to critique a lack

so much for multitrillion dollar candy companies

SafetyDGX agent

so much for multitrillion dollar candy companies Note to all staff: Turns out that super tasty AI candy we've been putting out in large bowls isn't free and it isn't necessarily leading to better prod

this won’t end well. it will end with a bailout.

SafetyDGX agent

this won’t end well. it will end with a bailout. Cash flow no longer covers the AI capex bill, so hyperscalers are funding it with record debt: Hyperscaler bond issuance has soared to $150 billion YTD

what a time to be alive

SafetyDGX agent

what a time to be alive wow, Opus 4.8 is very... argument-happy? it picked a fight with me about my usage of the word 'ontology', and when we eventually got back on the same page philosophically, told

Where is the power/value of money physically located? It's clearly not in the actual physical bills. Why are we more scared of an elderly ma…

SafetyDGX agent

Where is the power/value of money physically located? It's clearly not in the actual physical bills. Why are we more scared of an elderly mafia boss than their much more physically dangerous underling

Why LLMs rarely payoff—and what I have been saying literally for 7 years—confirmed yet again: LLMs can’t handle the truth. (Nor apparently c…

SafetyDGX agent

Why LLMs rarely payoff—and what I have been saying literally for 7 years—confirmed yet again: LLMs can’t handle the truth. (Nor apparently can my critics, who keep saying I am “always wrong”, when I h

29 May 2026

A 25B fund just refused SpaceX at any price. It says the company can't be worth more than 1T, half the $1.8T IPO target, and Musk's 85% co…

SafetyDGX agent

A 25B fund just refused SpaceX at any price. It says the company can't be worth more than 1T, half the $1.8T IPO target, and Musk's 85% control makes it impossible to fix from inside. https://thenextw

A Fully Convolutional Approach to Denoising Structural Dynamics Data from X-Ray Photon Correlation Spectroscopy

SafetyDGX agent

arXiv:2605.29975v1 Announce Type: new Abstract: We present a fully convolutional denoising autoencoder (FC-DAE) for denoising two-time intensity-intensity correlation functions (C_2) in X-ray photon c

A Geometric View of SRC: Learning Representations for Stable Residual Inference

SafetyDGX agent

arXiv:2605.29673v1 Announce Type: cross Abstract: Reconstruction-based inference assigns a class by comparing class-wise reconstruction residuals; Sparse Representation Classification (SRC) is a canon

A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms

SafetyDGX agent

arXiv:2605.30313v1 Announce Type: new Abstract: Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning a

A Modular Architecture for Typologically Controlled Lexicon Generation

SafetyDGX agent

arXiv:2605.28824v1 Announce Type: new Abstract: Constructing artificial lexicons that are pronounceable, typologically plausible, and semantically structured remains an open challenge in computational

A Predictive Law for On-Policy Self-Distillation From World Feedback

SafetyDGX agent

arXiv:2605.30070v1 Announce Type: cross Abstract: Moving beyond simple scalar rewards toward richer world feedback is a natural path to more scalable RL post-training. On-policy self-distillation (OPS

A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach

SafetyDGX agent

arXiv:2512.11944v2 Announce Type: replace-cross Abstract: Motion planning for autonomous driving (AD) faces a critical trade-off. While traditional rule-based pipelines offer verifiable safety and int

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

SafetyDGX agent

arXiv:2605.29340v1 Announce Type: new Abstract: In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual an

ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation

SafetyDGX agent

arXiv:2605.29791v1 Announce Type: new Abstract: While Large Language Models (LLMs) can convincingly simulate personas in explicit self-reports, they often deviate in implicit behavioral decisions, rev

Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment

SafetyDGX agent

arXiv:2605.29458v1 Announce Type: cross Abstract: Accurately simulating the decisions of a specific individual remains challenging for large language models (LLMs), partly because persona information

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

SafetyDGX agent

arXiv:2603.01006v2 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher

Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents

SafetyDGX agent

arXiv:2605.29910v1 Announce Type: cross Abstract: Consensus protocols form the backbone of distributed systems and blockchains, where implementation bugs can cause data corruption and financial losses

AIRGuard: Guarding Agent Actions with Runtime Authority Control

SafetyDGX agent

arXiv:2605.28914v1 Announce Type: cross Abstract: Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model C

AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing

SafetyDGX agent

arXiv:2605.29434v1 Announce Type: cross Abstract: Existing sentence-level watermarking methods enhance robustness to paraphrasing by anchoring watermarks in sentence semantics. However, their prefix-b

Anytime-Valid Federated Conformal RAG for LLM Swarms

SafetyDGX agent

arXiv:2605.29139v1 Announce Type: cross Abstract: Federated Conformal RAG (FC-RAG) provides distribution-free coverage for a bandwidth-limited swarm of weak language models, but only at a fixed horizo

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

SafetyDGX agent

arXiv:2605.30031v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsaf

Auditing Training Data in Generative Music Models via Black-Box Membership Inference

SafetyDGX agent

arXiv:2605.29202v1 Announce Type: new Abstract: Recent advances in text-to-music generation enable high-fidelity synthesis of structured musical audio, raising growing concerns about data provenance,

Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency

SafetyDGX agent

arXiv:2605.30208v1 Announce Type: cross Abstract: AI-assisted coding tools have altered software production. At Meta, significant lines of code per human-landed diff grew by 105.9% year over year and

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

SafetyDGX agent

arXiv:2605.28994v1 Announce Type: new Abstract: AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable.

Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction

SafetyDGX agent

arXiv:2605.28849v1 Announce Type: new Abstract: Gradient temporal-difference methods provide stable off-policy prediction with linear function approximation, but their practical performance is strongl

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures

SafetyDGX agent

arXiv:2605.29629v1 Announce Type: new Abstract: Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not ho

Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning

SafetyDGX agent

arXiv:2605.29414v1 Announce Type: cross Abstract: Recent studies have shown that code-switching data (CSD), in which multiple languages are mixed within the same context, can improve cross-lingual tra

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

SafetyDGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

← Previous
1…104105106107108…214
Next →