AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,598 results
13 Apr 2026

Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence

SafetyDGX agent

arXiv:2604.09104v1 Announce Type: cross Abstract: Scheming, the covert pursuit of misaligned goals by AI systems, represents a potentially catastrophic risk, yet scheming research suffers from signifi

SCoRe: Clean Image Generation from Diffusion Models Trained on Noisy Images

SafetyDGX agent

arXiv:2604.09436v1 Announce Type: new Abstract: Diffusion models trained on noisy datasets often reproduce high-frequency training artifacts, significantly degrading generation quality. To address thi

Score-Driven Rating System for Sports

SafetyDGX agent

arXiv:2604.09143v1 Announce Type: new Abstract: This paper introduces a score-driven rating system, a generalization of the classical Elo rating system that employs the score, i.e. the gradient of the


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines

SafetyDGX agent

arXiv:2604.08608v1 Announce Type: cross Abstract: We introduce Semantic Intent Fragmentation (SIF), an attack class against LLM orchestration systems where a single, legitimately phrased request cause

SHIFT: Steering Hidden Intermediates in Flow Transformers

SafetyDGX agent

arXiv:2604.09213v1 Announce Type: new Abstract: Diffusion models have become leading approaches for high-fidelity image generation. Recent DiT-based diffusion models, in particular, achieve strong pro

Sim-to-Real Transfer for Muscle-Actuated Robots via Generalized Actuator Networks

SafetyDGX agent

arXiv:2604.09487v1 Announce Type: cross Abstract: Tendon drives paired with soft muscle actuation enable faster and safer robots while potentially accelerating skill acquisition. Still, these systems

SPEAR: An Engineering Case Study of Multi-Agent Coordination for Smart Contract Auditing

SafetyDGX agent

arXiv:2602.04418v3 Announce Type: replace-cross Abstract: We present SPEAR, a multi-agent coordination framework for smart contract auditing that applies established MAS patterns in a realistic securi

SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks

SafetyDGX agent

arXiv:2604.08865v1 Announce Type: new Abstract: Proximal Policy Optimization (PPO) is central to aligning Large Language Models (LLMs) in reasoning tasks with verifiable rewards. However, standard tok

SSPO: Subsentence-level Policy Optimization

SafetyDGX agent

arXiv:2511.04256v2 Announce Type: replace Abstract: As a key component of large language model (LLM) post-training, Reinforcement Learning from Verifiable Rewards (RLVR) has substantially improved rea

StaRPO: Stability-Augmented Reinforcement Policy Optimization

SafetyDGX agent

arXiv:2604.08905v1 Announce Type: new Abstract: Reinforcement learning (RL) is effective in enhancing the accuracy of large language models in complex reasoning tasks. Existing RL policy optimization

STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting

SafetyDGX agent

arXiv:2509.25210v3 Announce Type: replace-cross Abstract: To gain finer regional forecasts, many works have explored the regional integration from the global atmosphere, e.g., by solving boundary equa

StructRL: Recovering Dynamic Programming Structure from Learning Dynamics in Distributional Reinforcement Learning

SafetyDGX agent

arXiv:2604.08620v1 Announce Type: cross Abstract: Reinforcement learning is typically treated as a uniform, data-driven optimization process, where updates are guided by rewards and temporal-differenc

SubQuad: Near-Quadratic-Free Structure Inference with Distribution-Balanced Objectives in Adaptive Receptor framework

SafetyDGX agent

arXiv:2602.17330v3 Announce Type: replace-cross Abstract: Comparative analysis of adaptive immune repertoires at population scale is hampered by two practical bottlenecks: the near-quadratic cost of p

Summary: AI Governance to Avoid Extinction

SafetyDGX agent

With AI capabilities rapidly increasing, humans appear close to developing AI systems that are better than human experts across all domains. This raises a series of questions about how the world will—

The causal relation between off-street parking and electric vehicle adoption in Scotland

SafetyDGX agent

arXiv:2604.09271v1 Announce Type: new Abstract: The transition to electric mobility hinges on maximising aggregate adoption while also facilitating equitable access. This study examines whether the 'c

The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?

SafetyDGX agent

arXiv:2601.23045v2 Announce Type: replace Abstract: As AI becomes more capable, we entrust it with more general and consequential tasks. The risks from failure grow more severe with increasing task sc

The trend of treating all of AI as One Big Thing that always includes data centers & job changes & education changes & power & accelerating …

SafetyDGX agent

The trend of treating all of AI as One Big Thing that always includes data centers & job changes & education changes & power & accelerating science & misinformation & national security & corporate con

The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs

SafetyDGX agent

arXiv:2601.01580v2 Announce Type: replace-cross Abstract: Self-reflection capabilities emerge in Large Language Models after RL post-training, with multi-turn RL achieving substantial gains over SFT c

'There is no advantage from being first. If you get first to a superintelligence you don't control, the superintelligence wins. Not the USA.…

SafetyDGX agent

'There is no advantage from being first. If you get first to a superintelligence you don't control, the superintelligence wins. Not the USA. Not China. Not the UK.' @andreamiotti of @ControlAI & TBC o

Think Less, Know More: State-Aware Reasoning Compression with Knowledge Guidance for Efficient Reasoning

SafetyDGX agent

arXiv:2604.09150v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks by leveraging long Chain-of-Thought (CoT), but often suffer from overthinking,

Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation

SafetyDGX agent

arXiv:2604.09368v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as scalable user simulators for recommender system evaluation. Yet existing simulators per

TME-PSR: Time-aware, Multi-interest, and Explanation Personalization for Sequential Recommendation

SafetyDGX agent

arXiv:2604.09439v1 Announce Type: cross Abstract: In this paper, we propose a sequential recommendation model that integrates Time-aware personalization, Multi-interest personalization, and Explanatio

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence

SafetyDGX agent

arXiv:2604.09057v1 Announce Type: new Abstract: Audio-video (AV) generation has recently made strong progress in perceptual quality and multimodal coherence, yet generating content with plausible moti

Toward World Models for Epidemiology

SafetyDGX agent

arXiv:2604.09519v1 Announce Type: new Abstract: World models have emerged as a unifying paradigm for learning latent dynamics, simulating counterfactual futures, and supporting planning under uncertai

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

SafetyDGX agent

arXiv:2604.08815v1 Announce Type: new Abstract: Medical vision-language models (VLMs) show strong performance on radiology tasks but often produce fluent yet weakly grounded conclusions due to over-re

Training event-based neural networks with exact gradients via Differentiable ODE Solving in JAX

SafetyDGX agent

arXiv:2603.08146v3 Announce Type: replace Abstract: Existing frameworks for gradient-based training of spiking neural networks face a trade-off: discrete-time methods using surrogate gradients support

Traj2Action: A Co-Denoising Framework for Trajectory-Guided Human-to-Robot Skill Transfer

SafetyDGX agent

arXiv:2510.00491v3 Announce Type: replace-cross Abstract: Learning diverse manipulation skills for real-world robots is severely bottlenecked by the reliance on costly and hard-to-scale teleoperated d

Truncated Rectified Flow Policy for Reinforcement Learning with One-Step Sampling

SafetyDGX agent

arXiv:2604.09159v1 Announce Type: new Abstract: Maximum entropy reinforcement learning (MaxEnt RL) has become a standard framework for sequential decision making, yet its standard Gaussian policy para

Unbiased Rectification for Sequential Recommender Systems Under Fake Orders

SafetyDGX agent

arXiv:2604.08550v1 Announce Type: cross Abstract: Fake orders pose increasing threats to sequential recommender systems by misleading recommendation results through artificially manipulated interactio

UniSemAlign: Text-Prototype Alignment with a Foundation Encoder for Semi-Supervised Histopathology Segmentation

SafetyDGX agent

arXiv:2604.09169v1 Announce Type: new Abstract: Semi-supervised semantic segmentation in computational pathology remains challenging due to scarce pixel-level annotations and unreliable pseudo-label s

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

SafetyDGX agent

arXiv:2604.09330v1 Announce Type: cross Abstract: Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-wo

Verbalizing LLMs' assumptions to explain and control sycophancy

SafetyDGX agent

arXiv:2604.03058v2 Announce Type: replace-cross Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like 'am I in the wrong?' rather than providing genuine assessment.

Violence is not the answer. But maybe boycotts are?

SafetyDGX agent

Violence is not the answer. But maybe boycotts are? 🚨 NOW: The FBI is RAIDING the home of a 20-year-old man who threw a molotov cocktail at the home of OpenAI CEO Sam Altman Over a DOZEN federal agent

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning

SafetyDGX agent

arXiv:2604.09508v1 Announce Type: cross Abstract: Visual Retrieval-Augmented Generation (VRAG) empowers Vision-Language Models to retrieve and reason over visually rich documents. To tackle complex qu

Visually-Guided Policy Optimization for Multimodal Reasoning

SafetyDGX agent

arXiv:2604.09349v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly advanced the reasoning ability of vision-language models (VLMs). However, the

When & How to Write for Personalized Demand-aware Query Rewriting in Video Search

SafetyDGX agent

arXiv:2602.17667v2 Announce Type: replace-cross Abstract: In video search systems, user historical behaviors provide rich context for identifying search intent and resolving ambiguity. However, tradit

Wireless Communication Enhanced Value Decomposition for Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2604.08728v1 Announce Type: new Abstract: Cooperation in multi-agent reinforcement learning (MARL) benefits from inter-agent communication, yet most approaches assume idealized channels and exis

⚡️ with @staysaasy, our second anonymous pod ever: https://www.youtube.com/watch?v=5KnCKadxSPY A conversation with the Stay Sassy duo on how…

SafetyDGX agent

⚡️ with @staysaasy, our second anonymous pod ever: https://www.youtube.com/watch?v=5KnCKadxSPY A conversation with the Stay Sassy duo on how AI is changing software teams, management, internal tooling

You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector

SafetyDGX agent

arXiv:2603.15757v2 Announce Type: replace-cross Abstract: What happens when a pretrained generative robot policy is provided a constant initial noise as input, rather than repeatedly sampling it from

12 Apr 2026

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to se…

SafetyDGX agent

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to see a debunking of their data, then we safety advocates should

Each workflow in Thoth is a full LangGraph agent which can call subagents which themselves are LangGraph agents. Just describe what you want…

SafetyDGX agent

Each workflow in Thoth is a full LangGraph agent which can call subagents which themselves are LangGraph agents. Just describe what you want in plain English and it builds a full multi-step pipeline.

🇧🇪Good news for Belgian Tesla owners! The Netherlands’ RDW just issued the first European type approval for Tesla FSD Supervised. Attached…

SafetyDGX agent

🇧🇪Good news for Belgian Tesla owners! The Netherlands’ RDW just issued the first European type approval for Tesla FSD Supervised. Attached letter from the Belgian federal administration (Minister Jean

I just love the language of this study...it speaks of the shifting 'community language'...and that is so true...have you noticed the new 'co…

SafetyDGX agent

I just love the language of this study...it speaks of the shifting 'community language'...and that is so true...have you noticed the new 'community language' is 'harness', it was 'contextual prompting

Is Elon Musk right about Sam Altman? just read the new Ronan Farrow and Andrew Marantz piece on Sam Altman in The New Yorker, and it is damn…

SafetyDGX agent

Is Elon Musk right about Sam Altman? just read the new Ronan Farrow and Andrew Marantz piece on Sam Altman in The New Yorker, and it is damning. And the answer is yes. It’s a deep dive based on never-

Restore Britain would allow free public speech without imprisonment

SafetyDGX agent

Restore Britain would allow free public speech without imprisonment Platforms hosting lawful content must be shielded from government pressure to censor. We would require transparency in content moder

Sources: the US' AI chip export push risks being undermined by licensing bottlenecks, staff attrition, and unclear policy at the Bureau of Industry and Security (Maggie Eastland/Bloomberg)

SafetyDGX agent

Maggie Eastland / Bloomberg: Sources: the US' AI chip export push risks being undermined by licensing bottlenecks, staff attrition, and unclear policy at the Bureau of Industry and Security — Presiden

The Biden administration actively flew illegals into America with no vetting of their violent criminal past into America and paid for their …

SafetyDGX agent

The Biden administration actively flew illegals into America with no vetting of their violent criminal past into America and paid for their flights via NGOs. This is a war crime. Mayorkas and his budd

Under South Africa's Employment Equity policy, every employer with more than 50 staff must comply with RACIAL quota targets set by the gover…

SafetyDGX agent

Under South Africa's Employment Equity policy, every employer with more than 50 staff must comply with RACIAL quota targets set by the government. Under these rules, in roles such as 'skilled technici

Wow, time for a social media detox. Just today - a call for violence against me - insults (that’s every day) - flagrant lies about my creden…

SafetyDGX agent

Wow, time for a social media detox. Just today - a call for violence against me - insults (that’s every day) - flagrant lies about my credentials - claims that I advocated violence when i repeatedly c

11 Apr 2026

After @aiDotEngineer, which was full of useful criticism, I remembered that the most confident takes on AI often came from the least exposur…

SafetyDGX agent

After @aiDotEngineer, which was full of useful criticism, I remembered that the most confident takes on AI often came from the least exposure. Rejection is easy, trial and error is expensive. I wrote

Angel investor calls for violence against AI skeptic.* *AI skeptic himself repeatedly decried violence, in favor of boycotts, despite multip…

SafetyDGX agent

An angel investor reportedly called for violence against AI skeptic Gary Marcus, who has consistently and publicly advocated against violent responses, instead promoting boycotts as a form of protest.

At this point how can anybody take seriously @sama’s claim that “Working towards prosperity for everyone, empowering all people, and advanci…

SafetyDGX agent

At this point how can anybody take seriously @sama’s claim that “Working towards prosperity for everyone, empowering all people, and advancing science and technology are moral obligations for me”, whe

Don’t fuck him. Don’t bomb him. Boycott him.

SafetyDGX agent

Gary Marcus, an AI researcher and cognitive scientist, posted a tweet advocating for a boycott as a nonviolent, non-confrontational form of protest or opposition against an unnamed individual. The pos

@GaryMarcus @sama the gap between altmans public statements and openais actual behavior has been widening steadily for 2 years now. at some …

SafetyDGX agent

@GaryMarcus @sama the gap between altmans public statements and openais actual behavior has been widening steadily for 2 years now. at some point 'we want to benefit humanity' and 'we want zero liabil

Given the extremely high rate of recidivism, this is important for community safety

SafetyDGX agent

Given the extremely high rate of recidivism, this is important for community safety Murder registry is a good idea from Elon. People should know if they are close to a murderer so proper precautions c

Home safe!!!

SafetyDGX agent

I was unable to retrieve the specific content of the tweet at the URL provided (https://x.com/GaryMarcus/status/2042756801201082637). X (formerly Twitter) content is generally not accessible via we...

Homicides per 100,000 in El Salvador: 2015: 103 2016: 81.0 2017: 60.2 2018: 50.4 2019: 35.8 2020: 21.2 2021: 18.1 2022: 7.8 2023: 2.4 2024: …

SafetyDGX agent

El Salvador's homicide rate per 100,000 people declined dramatically from 103 in 2015 to 2.4 in 2023, representing a reduction of over 97% during that period. This sharp decline is widely attributed t

👇 “I still maintain that LLMs are really dumb and of limited use. it’s my opinion as a practitioner and professional that the hype being pe…

SafetyDGX agent

👇 “I still maintain that LLMs are really dumb and of limited use. it’s my opinion as a practitioner and professional that the hype being peddled by the AI corporations and self promoters (from X grift

memory isn't a retrieval widget. it's write policy. the harness decides what survives compaction, what gets promoted, and what becomes reusa…

SafetyDGX agent

memory isn't a retrieval widget. it's write policy. the harness decides what survives compaction, what gets promoted, and what becomes reusable state. https://x.com/hwchase17/status/204297850056760973

One investor today called for violence against me. Another lied about me, in a pretty deep and fundamental way. They are feeling the heat.

SafetyDGX agent

Gary Marcus, a prominent AI researcher and critic, posted on X (formerly Twitter) describing hostile reactions from investors, including one allegedly calling for violence against him and another maki

← Previous
1…204205206207208…210
Next →