AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
29 Apr 2026

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

Model ReleasesDGX agent

arXiv:2509.09708v3 Announce Type: replace Abstract: Refusal on harmful prompts is a key safety behaviour in instruction-tuned large language models (LLMs), yet the internal causes of this behaviour re

Carbon-Taxed Transformers: A Green Compression Pipeline for Overgrown Language Models

SafetyDGX agent

arXiv:2604.25903v1 Announce Type: cross Abstract: The accelerating adoption of Large Language Models (LLMs) in software engineering (SE) has brought with it a silent crisis: unsustainable computationa

CHUCKLE -- When Humans Teach AI To Learn Emotions The Easy Way

SafetyDGX agent

arXiv:2510.09382v2 Announce Type: replace Abstract: Curriculum learning (CL) structures training from simple to complex samples, facilitating progressive learning. However, existing CL approaches for

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

- Comey indicted for tweeting a number. - Trump FCC threatens ABC's broadcast license. - Trump defacing more govt institutions with his name…

SafetyDGX agent

- Comey indicted for tweeting a number. - Trump FCC threatens ABC's broadcast license. - Trump defacing more govt institutions with his name and picture. - Trump's kids cashing in on huge govt contrac

Compute Aligned Training: Optimizing for Test Time Inference

SafetyDGX agent

arXiv:2604.24957v1 Announce Type: new Abstract: Scaling test-time compute has emerged as a powerful mechanism for enhancing Large Language Model (LLM) performance. However, standard post-training para

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

SafetyDGX agent

arXiv:2604.25891v1 Announce Type: new Abstract: Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavio

CORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual RAG

SafetyDGX agent

arXiv:2604.25676v1 Announce Type: new Abstract: Multilingual retrieval-augmented generation (mRAG) is often implemented within a fixed retrieval space, typically via query or document translation or m

Crazy how many people expect some company or country to achieve a “insurmountable lead” in AI, when in reality no company seems to hold a le…

SafetyDGX agent

Crazy how many people expect some company or country to achieve a “insurmountable lead” in AI, when in reality no company seems to hold a lead of more that a few weeks or at most a few months before s

CroSearch-R1: Better Leveraging Cross-lingual Knowledge for Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2604.25182v1 Announce Type: new Abstract: A multilingual collection may contain useful knowledge in other languages to supplement and correct the facts in the original language for Retrieval-Aug

Cross-Lingual Jailbreak Detection via Semantic Codebooks

Model ReleasesDGX agent

arXiv:2604.25716v1 Announce Type: new Abstract: Safety mechanisms for large language models (LLMs) remain predominantly English-centric, creating systematic vulnerabilities in multilingual deployment.

Damning

SafetyDGX agent

Damning 🚨 Altman texts Musk: 'we offered you equity when we established the capped profit. you didn't want it at the time.' (Altman drafted this with Shivon Zilis, then sent it to Musk that night.) Op

DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework

SafetyDGX agent

arXiv:2506.05199v3 Announce Type: replace Abstract: A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pa

DGLight: DQN-Guided GRPO Fine-Tuning of Large Language Models for Traffic Signal Control

SafetyDGX agent

arXiv:2604.25259v1 Announce Type: new Abstract: Traffic signal control (TSC) plays a central role in reducing congestion and maintaining urban mobility. This dissertation introduces DGLight, a critic-

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors

SafetyDGX agent

arXiv:2604.25050v1 Announce Type: new Abstract: Unlike chatbots, physical AI must act while the world keeps evolving. Therefore, the inter-chunk pause of synchronous executors are fatal for dynamic ta

DSO: Direct Steering Optimization for Bias Mitigation

SafetyDGX agent

Generative models are often deployed to make decisions on behalf of users, such as vision-language models (VLMs) identifying which person in a room is a doctor to help visually impaired individuals. Y

Elon’s lawsuit might have merit even his motivations are questionable, as I just explained on @cnni https://video.snapstream.net/Play/7vHDLC…

SafetyDGX agent

Gary Marcus discusses on CNN the legal merits of Elon Musk's lawsuit, arguing that despite questionable motivations behind the case, there may be legitimate legal grounds supporting it. The post refer

Enterprises turn to runtime security to close the agentic AI trust gap

SafetyDGX agent

As enterprises push agentic AI out of the proof-of-concept phase and into production, AI runtime security — the ability to enforce policy at the exact moment an agent acts — is proving to be the bedro

Frictive Policy Optimization for LLMs: Epistemic Intervention, Risk-Sensitive Control, and Reflective Alignment

SafetyDGX agent

arXiv:2604.25136v1 Announce Type: new Abstract: We propose Frictive Policy Optimization (FPO), a framework for learning language model policies that regulate not only what to say, but when and how to

@GaryMarcus @TheParisSpleen Gary, that was a great article. Wonderful work as always sir. You have been proven accurate once again.

SafetyDGX agent

This is a complimentary response to an article by Gary Marcus, praising his accuracy and work quality. The tweet does not provide specific details about the article's content, only that it was well-re

How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuum

SafetyDGX agent

arXiv:2604.25907v1 Announce Type: new Abstract: Adapting reasoning models to new tasks during post-training with only output-level supervision stalls under reinforcement learning from verifiable rewar

How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning

SafetyDGX agent

arXiv:2603.01070v2 Announce Type: replace Abstract: Solving complex geometric problems inherently requires interleaved reasoning: a tight alternation between constructing diagrams and performing logic

I-INR: Iterative Implicit Neural Representations

SafetyDGX agent

arXiv:2504.17364v4 Announce Type: replace Abstract: Implicit Neural Representations (INRs) have revolutionized signal processing and computer vision by modeling signals as continuous, differentiable f

👇 I was there for this, sitting right next to Altman. Realizing a few months later he lied about it (by omission) was what turned me agains…

SafetyDGX agent

👇 I was there for this, sitting right next to Altman. Realizing a few months later he lied about it (by omission) was what turned me against him. We had sworn to tell the whole truth and nothing but t

Interactive Episodic Memory with User Feedback

SafetyDGX agent

arXiv:2604.24893v1 Announce Type: new Abstract: In episodic memory with natural language queries (EM-NLQ), a user may ask a question (e.g., 'Where did I place the mug?') that requires searching a long

La industria de las IA vive exclusivamente de nuestra creencias y supuesto de que podría hacer lo que no hace.

SafetyDGX agent

La industria de las IA vive exclusivamente de nuestra creencias y supuesto de que podría hacer lo que no hace. Sheer insanity. Amazon, Google, Microsoft, and Meta collectively are spending more money

Learning-Based Dynamics Modeling and Robust Control for Tendon-Driven Continuum Robots

SafetyDGX agent

arXiv:2604.25691v1 Announce Type: new Abstract: Tendon-Driven Continuum Robots (TDCRs) pose significant modeling and control challenges due to complex nonlinearities, such as frictional hysteresis and

Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization

SafetyDGX agent

arXiv:2604.24952v1 Announce Type: new Abstract: Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets

Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System

SafetyDGX agent

arXiv:2604.24921v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are a promising paradigm for generalist robotic manipulation by grounding high-level semantic instructions into ex

LOL. Just drop the stupid supply chain risk designation, and admit your mistake.

SafetyDGX agent

LOL. Just drop the stupid supply chain risk designation, and admit your mistake. SCOOP: The White House is developing guidance that would allow agencies to get around Anthropic's supply chain risk des

MAIC-UI: Making Interactive Courseware with Generative UI

SafetyDGX agent

arXiv:2604.25806v1 Announce Type: new Abstract: Creating interactive STEM courseware traditionally requires HTML/CSS/JavaScript expertise, leaving barriers for educators. While generative AI can produ

Misleading indeed.

SafetyDGX agent

Misleading indeed. 🚨 OpenAI's lawyer William Savitt just attacked Musk on cross-examination: 'you pledged 1 billion, only delivered 38 million.' That's misleading. The $1B was a COLLECTIVE pledge from

MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts

SafetyDGX agent

arXiv:2411.14721v2 Announce Type: replace Abstract: Molecule discovery is a pivotal research field, impacting everything from medicine to materials. Recently, Large Language Models (LLMs) have been wi

Navigating Global AI Regulation: A Multi-Jurisdictional Retrieval-Augmented Generation System

SafetyDGX agent

arXiv:2604.25448v1 Announce Type: new Abstract: Navigating AI regulation across jurisdictions is increasingly difficult for policymakers, legal professionals, and researchers. To address this, we pres

NimbleReg: A light-weight deep-learning framework for diffeomorphic image registration

SafetyDGX agent

arXiv:2503.07768v2 Announce Type: replace Abstract: This paper presents NimbleReg, a light-weight deep-learning (DL) framework for diffeomorphic image registration leveraging surface representation of

One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query Refinement

SafetyDGX agent

arXiv:2604.25444v1 Announce Type: new Abstract: Large Language Models (LLMs) often fail to utilize their latent reasoning capabilities due to a distributional mismatch between ambiguous human inquirie

OpenAI may no longer be a nonprofit but it continues to be—and may always be—a “never made a profit”

SafetyDGX agent

OpenAI may no longer be a nonprofit but it continues to be—and may always be—a “never made a profit” @GaryMarcus The irony may be that OpenAI won't make the incomes it is projecting and will be a non

Policy Improvement Reinforcement Learning

SafetyDGX agent

arXiv:2604.00860v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a central post-training paradigm for improving the reasoning capabilities of large

Progressing beyond Art Masterpieces or Touristic Cliches: how to assess your LLMs for cultural alignment?

SafetyDGX agent

arXiv:2604.25654v1 Announce Type: new Abstract: Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention -- often framed in terms of cultural bias -- unt

Records Sam Altman might set: • Most money burned • Biggest lead squandered • Most promises broken • Most nonprofits seized (tie, probably n…

SafetyDGX agent

Records Sam Altman might set: • Most money burned • Biggest lead squandered • Most promises broken • Most nonprofits seized (tie, probably nobody has two) • Most times fired from one job (anyone know

Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots

SafetyDGX agent

arXiv:2604.25698v1 Announce Type: new Abstract: Tendon-Driven Continuum Robots (TDCRs) pose significant control challenges due to their highly nonlinear, path-dependent dynamics and non-Markovian char

Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models

SafetyDGX agent

arXiv:2604.25636v1 Announce Type: new Abstract: Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified ca

ResetEdit: Precise Text-guided Editing of Generated Image via Resettable Starting Latent

SafetyDGX agent

arXiv:2604.25128v1 Announce Type: new Abstract: Recent advances in diffusion models have enabled high-quality image generation, leading to increasing demand for post-generation editing that modifies l

ReSim: Reliable World Simulation for Autonomous Driving

SafetyDGX agent

arXiv:2506.09981v2 Announce Type: replace Abstract: How can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusivel

Resource Orchestration & Redistribution Systems As production becomes more efficient (energy, goods, services), the constraint shifts to all…

SafetyDGX agent

Resource Orchestration & Redistribution Systems As production becomes more efficient (energy, goods, services), the constraint shifts to allocation. These systems dynamically route resources based on

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

SafetyDGX agent

arXiv:2510.10150v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) serves as a cornerstone technique for enhancing the reasoning capabilities of Large Language M

“roughly $1.35 trillion of pure narrative” is the best euphemism for pure bullshit that I have ever seen, @gnoble79. 👏

SafetyDGX agent

“roughly 1.35 trillion of pure narrative” is the best euphemism for pure bullshit that I have ever seen, @gnoble79. 👏 This is the most OUTRAGEOUS deal I've seen in my 45 years on Wall Street. SpaceX j

Sam’s not the right CEO for OpenAI anymore. He squandered their lead, failed to deliver the AGI he promised, flailed productwise, and torche…

SafetyDGX agent

Sam’s not the right CEO for OpenAI anymore. He squandered their lead, failed to deliver the AGI he promised, flailed productwise, and torched his own reputation; meanwhile he is burning money at a unp

Sheer insanity. Amazon, Google, Microsoft, and Meta collectively are spending more money than the Manhattan Project *every single month*. Mo…

SafetyDGX agent

Sheer insanity. Amazon, Google, Microsoft, and Meta collectively are spending more money than the Manhattan Project *every single month*. More than 12x the Manhattan Project every year. And what they

Spark Policy Toolkit: Semantic Contracts and Scalable Execution for Policy Learning in Spark

SafetyDGX agent

arXiv:2604.25061v1 Announce Type: cross Abstract: Custom policy-learning pipelines in Spark fail for two coupled systems reasons: rowwise Python execution makes inference impractical, and driver-side

Spectral bandits

SafetyDGX agent

arXiv:2604.25272v1 Announce Type: cross Abstract: Smooth functions on graphs have wide applications in manifold and semi-supervised learning. In this work, we study a bandit problem where the payoffs

Subliminal Steering: Stronger Encoding of Hidden Signals

SafetyDGX agent

arXiv:2604.25783v1 Announce Type: new Abstract: Subliminal learning describes a student language model inheriting a behavioral bias by fine-tuning on seemingly innocuous data generated by a biased tea

the ability of this man to sell things he doesn’t believe in is truly extraordinary

SafetyDGX agent

the ability of this man to sell things he doesn’t believe in is truly extraordinary SAM ALTMAN “OpenAI is structured as a nonprofit because we don’t ever want to be making decisions to benefit shareho

The Musk-Altman trial is clearly a battle of egos—and it’s hard to root for either side—𝗯𝘂𝘁 𝗶𝘁’𝘀 𝗮𝗹𝘀𝗼 𝗮 𝘁𝗿𝗶𝗮𝗹 𝗮𝗯𝗼𝘂𝘁 𝘄…

SafetyDGX agent

The Musk-Altman trial is clearly a battle of egos—and it’s hard to root for either side—𝗯𝘂𝘁 𝗶𝘁’𝘀 𝗮𝗹𝘀𝗼 𝗮 𝘁𝗿𝗶𝗮𝗹 𝗮𝗯𝗼𝘂𝘁 𝘄𝗵𝗲𝘁𝗵𝗲𝗿 𝗢𝗽𝗲𝗻𝗔𝗜 𝘀𝗵𝗼𝘂𝗹𝗱 𝗯𝗲 𝗵𝗲𝗹𝗱 𝘁𝗼 𝗶𝘁𝘀 𝗽𝗿𝗼𝗺𝗶𝘀𝗲𝘀 𝘁𝗼 𝗯𝗲 𝗮 𝗻𝗼𝗻𝗽𝗿𝗼𝗳𝗶𝘁 𝘄𝗼𝗿𝗸𝗶𝗻𝗴 𝗳𝗼𝗿 𝘁𝗵𝗲 𝗯𝗲𝗻𝗲

The President of the United States trying to take $10 billion dollars from US taxpayers for his own personal gain—and he is probably going t…

SafetyDGX agent

This post appears to discuss allegations that a U.S. President sought to obtain $10 billion in taxpayer funds for personal use. The content is incomplete in the provided title, making it difficult to

The Rashomon Effect for Visualizing High-Dimensional Data

SafetyDGX agent

arXiv:2604.00485v2 Announce Type: replace Abstract: Dimension reduction (DR) is inherently non-unique: multiple embeddings can preserve the structure of high-dimensional data equally well while differ

The Russian Legislative Corpus

SafetyDGX agent

arXiv:2406.04855v3 Announce Type: replace Abstract: We present a comprehensive corpus of Russian primary and secondary legislation adopted between 1991 and 2025, comprising 304,382 texts (194,425,905

Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models

SafetyDGX agent

arXiv:2510.16340v2 Announce Type: replace Abstract: Recent advances in post-training techniques have endowed Large Language Models (LLMs) with enhanced capabilities for tackling complex, logic-intensi

this kind of armchair nonsense is especially hilarious on day when I defended Elon’s lawsuit on CNN and in my newsletter. 🙄

SafetyDGX agent

this kind of armchair nonsense is especially hilarious on day when I defended Elon’s lawsuit on CNN and in my newsletter. 🙄 @GaryMarcus @gnoble79 I'm not sure history is going to be kind to this take.

Three Models of RLHF Annotation: Extension, Evidence, and Authority

SafetyDGX agent

arXiv:2604.25895v1 Announce Type: cross Abstract: Preference-based alignment methods, most prominently Reinforcement Learning with Human Feedback (RLHF), use the judgments of human annotators to shape

To be a bit more clear for people who did not read the full article: A domestic data-center ban does not directly address the risks of extin…

SafetyDGX agent

To be a bit more clear for people who did not read the full article: A domestic data-center ban does not directly address the risks of extinction from ASI, and weakens the US. People may or may not al

← Previous
1…192193194195196…240
Next →