AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,488 results
11 Jun 2026

Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning

SafetyDGX agent

arXiv:2603.14867v4 Announce Type: replace-cross Abstract: Many strategic decision-making problems, such as environment design for warehouse robots, can be naturally formulated as bi-level reinforcemen

Signed Compression Progress on a Sealed Audit is Goodhart-Resistant

SafetyDGX agent

arXiv:2606.11417v1 Announce Type: cross Abstract: Compression progress is a long-standing proposal for intrinsic motivation: reward an agent when its world model becomes better at predicting or compre

SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation

SafetyDGX agent

arXiv:2511.05203v3 Announce Type: replace Abstract: Today's autonomous agents, largely driven by foundation models (FMs), can understand natural language instructions and solve long-horizon tasks with

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Sovereign Assurance Boundary: Certificate-Bound Admission for Agentic Infrastructure

SafetyDGX agent

arXiv:2606.11632v1 Announce Type: cross Abstract: Agentic infrastructure introduces a critical control-plane authorization problem: non-deterministic reasoning systems can propose high-stakes mutation

Spectrally Regularized Latent Flow Matching for Turbulence Generation

SafetyDGX agent

arXiv:2606.11691v1 Announce Type: new Abstract: Latent diffusion and flow matching have emerged as leading approaches for synthetic turbulence generation, yet they systematically under-represent dissi

Steering Multirobot Behavior via Closed-Loop Affine Activation Editing

SafetyDGX agent

arXiv:2606.11489v1 Announce Type: new Abstract: Real-world robots need to adapt their behavior beyond the envelope of their pre-trained policy. Policy finetuning or retraining are options, but they ri

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning

SafetyDGX agent

arXiv:2606.11770v1 Announce Type: new Abstract: Spatial reasoning remains a challenge for Multimodal Large Language Models (MLLMs), as it requires reliable multi-hop inference over both intermediate s

TacCoRL: Integrating Tactile Feedback into VLA via Simulation

SafetyDGX agent

arXiv:2606.11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often mis

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning

SafetyDGX agent

arXiv:2606.11853v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) depend on in-context learning (ICL) for rapid task adaptation, but their scalability is severely limited by

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

SafetyDGX agent

arXiv:2606.11918v1 Announce Type: new Abstract: Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approa

The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning

SafetyDGX agent

arXiv:2505.03296v2 Announce Type: replace-cross Abstract: We present Mixture of Discrete-time Gaussian Processes (MiDiGap), a novel approach for flexible policy representation and imitation learning i

This was perhaps the most controversial aspect of the guardrails around Fable, now being rolled back.

SafetyDGX agent

This was perhaps the most controversial aspect of the guardrails around Fable, now being rolled back. Very pleased to hear Anthropic have walked back this policy https://simonwillison.net/2026/Jun/11/

To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending

SafetyDGX agent

arXiv:2606.11201v1 Announce Type: cross Abstract: The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions. Among

Toward Preference-aligned Large Language Models via Residual-based Model Steering

SafetyDGX agent

arXiv:2509.23982v2 Announce Type: replace-cross Abstract: Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences. Existing approache

Towards a Bridge Layer Between Bibliographic and Formalized Mathematical Knowledge

SafetyDGX agent

arXiv:2606.11430v1 Announce Type: cross Abstract: Mathematical knowledge is split between bibliographic databases (e.g., MathSciNet, zbMATH Open) and formal proof libraries (e.g., Lean mathlib), preve

Towards Conditional Feature Alignment for Cross-Domain Counting

SafetyDGX agent

arXiv:2506.17137v3 Announce Type: replace Abstract: Object counting models often degrade under cross-domain deployment because density composition varies across domains and is itself task-relevant. St

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Personality Assessment

SafetyDGX agent

arXiv:2606.11269v1 Announce Type: new Abstract: Personality assessment aims to infer stable personality traits from dynamic behaviors across language, voice, and facial cues. Since different personali

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning

SafetyDGX agent

arXiv:2606.12372v1 Announce Type: cross Abstract: Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy imp

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA

SafetyDGX agent

arXiv:2606.11740v1 Announce Type: cross Abstract: We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

SafetyDGX agent

arXiv:2606.11681v1 Announce Type: new Abstract: We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phonem

Urban Heat MiniCubes: An AI-Ready dataset for urban heat research

SafetyDGX agent

arXiv:2606.11534v1 Announce Type: cross Abstract: Urban heat is amplified by impermeable surfaces and heterogeneous built environments, yet street-level variability remains difficult to quantify becau

Very pleased to hear Anthropic have walked back this policy https://simonwillison.net/2026/Jun/11/anthropic-walks-back-policy/

SafetyDGX agent

Very pleased to hear Anthropic have walked back this policy https://simonwillison.net/2026/Jun/11/anthropic-walks-back-policy/ BREAKING NEWS: Anthropic's latest model will NOT help you if it thinks yo

ViT-FREE: Efficient Face Recognition via Early Exiting and Synthetic Adaptation

SafetyDGX agent

arXiv:2606.12023v1 Announce Type: new Abstract: Vision Transformers (ViTs) have gained significant attention in computer vision and shown strong potential for face recognition (FR). However, their hig

VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving

SafetyDGX agent

arXiv:2606.12396v1 Announce Type: new Abstract: Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D wo

What are smart people are saying about OpenAI's IPO filing? 🤔

SafetyDGX agent

Gary Marcus discusses expert commentary and analysis regarding OpenAI's initial public offering filing, likely examining implications for the AI industry, company valuation, and competitive landscape.

When Context Returns: Toward Robust Internalization in On-Policy Distillation

SafetyDGX agent

arXiv:2606.11627v1 Announce Type: cross Abstract: Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so th

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?

SafetyDGX agent

arXiv:2510.02660v2 Announce Type: replace-cross Abstract: When researchers claim AI systems possess ToM or mental models, they are fundamentally discussing behavioral predictions and bias corrections

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

SafetyDGX agent

arXiv:2606.12199v1 Announce Type: cross Abstract: Spoken dialogue models typically start from text LLM backbones, yet reasoning often degrades when conditioning on speech instead of text. We attribute

Wrong horse and too early. Gary nailed it. Time will show this.

SafetyDGX agent

Wrong horse and too early. Gary nailed it. Time will show this. Masa can hype AI all he likes, but apparently he can’t even get a loan on his OpenAI shares. 🤔 But I actually half agree—and it might su

10 Jun 2026

100%, hallucinations are baked in (as I have been saying since 2001) and that is why basically no LLM company afford to operate in Germany n…

SafetyDGX agent

100%, hallucinations are baked in (as I have been saying since 2001) and that is why basically no LLM company afford to operate in Germany now. We need a better technology. @GaryMarcus Lol, and Google

3SPO: State-Score-Supervised Policy Optimization for LLM Agents

SafetyDGX agent

arXiv:2606.09961v1 Announce Type: cross Abstract: Training large language models (LLMs) as autonomous agents via reinforcement learning (RL) has enabled frontier models to achieve superhuman performan

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

SafetyDGX agent

arXiv:2410.15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct P

A fine-grained attention and geometric correspondence model for musculoskeletal risk classification in athletes using multimodal visual and skeletal features

SafetyDGX agent

arXiv:2509.05913v3 Announce Type: replace Abstract: Musculoskeletal disorders pose significant risks to athletes, and early risk assessment is essential for prevention. However, most existing methods

A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation

SafetyDGX agent

arXiv:2606.10366v1 Announce Type: cross Abstract: Simulation has become an essential tool for evaluating and improving vision-language-action (VLA) policies, offering scalable, reproducible, and contr

A Source Domain is All You Need: Source-Only Cross-OS Transfer Learning for APT Anomaly Detection via Semantic Alignment and Optimal Transport

SafetyDGX agent

arXiv:2606.10216v1 Announce Type: cross Abstract: Advanced Persistent Threats (APTs) are stealthy, multi-stage cyberattacks whose detection is difficult due to scarce labeled traces, severe class imba

A Unified Multi-Modal Framework for Intelligent Financial Systems: Integrating Reinforcement Learning, High-Frequency Trading, and Game-Theoretic Approaches with Cross-Modal Sentiment Analysis

SafetyDGX agent

arXiv:2606.10412v1 Announce Type: new Abstract: The rapid evolution of financial technology demands sophisticated artificial intelligence systems capable of handling diverse challenges across multiple

Adoption of Generative Artificial Intelligence in the German Software Engineering Industry: An Empirical Study

SafetyDGX agent

arXiv:2601.16700v2 Announce Type: replace-cross Abstract: Generative artificial intelligence (GenAI) tools have seen rapid adoption among software developers. While adoption rates in the industry are

AI oligarchs are trying to buy elections (again)

SafetyDGX agent

AI oligarchs are trying to buy elections (again) AI titans are spending 1.3M to shape Utah primaries. Defending Our Values, run by Chris Stewart, spent 880K to support Rep. Celeste Maloy. DOV operates

Alignment Defends LLMs from Property Inference Attacks

SafetyDGX agent

arXiv:2606.10217v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly fine-tuned on domain-specific datasets that may contain sensitive, dataset-level properties. Recent work h

Alongside it, Anthropic is releasing a proposal for how governments can address the risks posed by frontier AI and a policy framework for jo…

SafetyDGX agent

Alongside it, Anthropic is releasing a proposal for how governments can address the risks posed by frontier AI and a policy framework for job displacement, for which we intend to provide substantial f

An LLM-Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models

SafetyDGX agent

arXiv:2606.09843v1 Announce Type: cross Abstract: Large language models (LLMs) produce stable self-reports on personality inventories, but these self-reports do not predict observed behavior. Whether

and see for deeper analysis https://open.substack.com/pub/garymarcus/p/breaking-news-and-how-the-end-might?r=8tdk6&utm_medium=ios

SafetyDGX agent

Gary Marcus discusses breaking news and potential existential risks or significant developments related to AI systems, likely examining how current AI capabilities and limitations might lead to critic

AnimaSpark: A Feed-Forward Method for Animating Arbitrary 3D Objects

SafetyDGX agent

arXiv:2606.10988v1 Announce Type: new Abstract: While recent advancements in generative AI have substantially accelerated static 3D model creation workflows, the synthesis of category-agnostic 3D anim

Anthropic releases two policy proposals on how governments should address catastrophic risks and manage labor market disruption from advanced AI systems (Anthropic)

SafetyDGX agent

Anthropic: Anthropic releases two policy proposals on how governments should address catastrophic risks and manage labor market disruption from advanced AI systems — AI is advancing at exponential spe

Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans

SafetyDGX agent

arXiv:2606.10953v1 Announce Type: new Abstract: Furnished floor plans are fundamental to real estate visualization, interior design, and architectural workflows. However, progress in automatic furnitu

are you kidding me

SafetyDGX agent

are you kidding me Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was built to handle. The essay lays out where I

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

SafetyDGX agent

arXiv:2606.11188v1 Announce Type: new Abstract: This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-

As vertically integrated platforms start to dominate they lock out third party access to the most valuable portions of the platform. Of cour…

SafetyDGX agent

As vertically integrated platforms start to dominate they lock out third party access to the most valuable portions of the platform. Of course, Anthropic is has the right to implement whatever policy

AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents

SafetyDGX agent

arXiv:2606.05597v2 Announce Type: replace Abstract: Training vision-language web agents with multi-step RL is compute-intensive, with two dominant forms of inefficiency: idle GPUs in synchronous RL, a

Automated Alignment between Elicitation Interviews and Requirements

SafetyDGX agent

arXiv:2510.08622v2 Announce Type: replace Abstract: Software requirements are derived from a variety of elicitation techniques, many of which have a conversational nature, like interviews. However, ev

Automated Scoring of Arabic Text Using Large Language Models: A Literature Review

SafetyDGX agent

arXiv:2606.09830v1 Announce Type: new Abstract: In modern educational systems, Automatic Text Scoring (ATS) plays a central role by enabling scalable and consistent evaluation of learner responses wit

Banger post from @testdrivenzen There is no good outcome for the world unless countries coordinate to stop ASI development.

SafetyDGX agent

Banger post from @testdrivenzen There is no good outcome for the world unless countries coordinate to stop ASI development. Any plan for surviving superintelligent AI that doesn't go through strong in

Baseline-Free Policy Optimization for Neural Combinatorial Optimization

SafetyDGX agent

arXiv:2606.10321v1 Announce Type: cross Abstract: Neural combinatorial optimization (NCO) trains autoregressive policies to solve routing problems. The standard training algorithm, REINFORCE with a ro

Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets

SafetyDGX agent

arXiv:2606.10979v1 Announce Type: new Abstract: Many Markov decision processes (MDPs) in operations research have feasible actions that are state dependent and defined implicitly by various operationa

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

SafetyDGX agent

arXiv:2606.10968v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing PPO-style trust-region mechan

🚨Breaking news that could be huge, and enormously bad for GenAI, if other countries make similar decisions. https://the-decoder.com/landmar…

SafetyDGX agent

🚨Breaking news that could be huge, and enormously bad for GenAI, if other countries make similar decisions. https://the-decoder.com/landmark-german-ruling-declares-googles-ai-overviews-are-googles-own

BREAKING! This may well be the beginning of the end.

SafetyDGX agent

BREAKING! This may well be the beginning of the end. 🚨BREAKING: SoftBank tried to borrow 6 billion against its 13% OpenAI stock… to keep funding OpenAI Banks said NO They don’t believe OpenAI is worth

Bypassing Copyright Protection in Diffusion-based Customization via Two-Stage Latent Feature Optimization

SafetyDGX agent

arXiv:2606.09909v1 Announce Type: cross Abstract: With the growing concerns over copyright infringement in diffusion-based customization, adversarial attacks have emerged as a prominent defense strate

Causal Ensemble Agent: Hierarchical Causal Discovery with LLM-guided Expert Reweighting

SafetyDGX agent

arXiv:2606.10607v1 Announce Type: cross Abstract: Causal discovery aims to uncover causal structures from observational data, which is crucial for real-world decision-making. However, different causal

Closing the Modality Gap in Zero-Shot HAR: Contrastive Training and Separability-Optimized Prototypes on IMU Data

SafetyDGX agent

arXiv:2606.10789v1 Announce Type: new Abstract: Zero-shot learning (ZSL) for inertial measurement unit (IMU)-based human activity recognition (HAR) faces a central challenge: bridging the gap between

← Previous
1…111112113114115…242
Next →