AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,490 results
Safety

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

DGX agent

arXiv:2606.11709v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distri

safetyarxiv-cs-cl
11 Jun 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment

DGX agent

arXiv:2606.11512v1 Announce Type: new Abstract: Large language models increasingly express uncertainty through natural-language statements, yet these expressions often fail to reflect the model's samp

safetyarxiv-cs-cl
11 Jun 2026
Safety

Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning

DGX agent

arXiv:2603.14867v4 Announce Type: replace-cross Abstract: Many strategic decision-making problems, such as environment design for warehouse robots, can be naturally formulated as bi-level reinforcemen

safetyarxiv-cs-ai
11 Jun 2026
Safety

Signed Compression Progress on a Sealed Audit is Goodhart-Resistant

DGX agent

arXiv:2606.11417v1 Announce Type: cross Abstract: Compression progress is a long-standing proposal for intrinsic motivation: reward an agent when its world model becomes better at predicting or compre

safetyarxiv-cs-ai
11 Jun 2026
Safety

SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation

DGX agent

arXiv:2511.05203v3 Announce Type: replace Abstract: Today's autonomous agents, largely driven by foundation models (FMs), can understand natural language instructions and solve long-horizon tasks with

safetyarxiv-cs-ro
11 Jun 2026
Safety

Sovereign Assurance Boundary: Certificate-Bound Admission for Agentic Infrastructure

DGX agent

arXiv:2606.11632v1 Announce Type: cross Abstract: Agentic infrastructure introduces a critical control-plane authorization problem: non-deterministic reasoning systems can propose high-stakes mutation

safetyarxiv-cs-ai
11 Jun 2026
Safety

Spectrally Regularized Latent Flow Matching for Turbulence Generation

DGX agent

arXiv:2606.11691v1 Announce Type: new Abstract: Latent diffusion and flow matching have emerged as leading approaches for synthetic turbulence generation, yet they systematically under-represent dissi

safetyarxiv-cs-lg
11 Jun 2026
Safety

Steering Multirobot Behavior via Closed-Loop Affine Activation Editing

DGX agent

arXiv:2606.11489v1 Announce Type: new Abstract: Real-world robots need to adapt their behavior beyond the envelope of their pre-trained policy. Policy finetuning or retraining are options, but they ri

safetyarxiv-cs-ro
11 Jun 2026
Safety

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning

DGX agent

arXiv:2606.11770v1 Announce Type: new Abstract: Spatial reasoning remains a challenge for Multimodal Large Language Models (MLLMs), as it requires reliable multi-hop inference over both intermediate s

safetyarxiv-cs-ai
11 Jun 2026
Safety

TacCoRL: Integrating Tactile Feedback into VLA via Simulation

DGX agent

arXiv:2606.11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often mis

safetyarxiv-cs-lg
11 Jun 2026
Safety

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning

DGX agent

arXiv:2606.11853v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) depend on in-context learning (ICL) for rapid task adaptation, but their scalability is severely limited by

safetyarxiv-cs-ai
11 Jun 2026
Safety

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

DGX agent

arXiv:2606.11918v1 Announce Type: new Abstract: Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approa

safetyarxiv-cs-ai
11 Jun 2026
Safety

The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning

DGX agent

arXiv:2505.03296v2 Announce Type: replace-cross Abstract: We present Mixture of Discrete-time Gaussian Processes (MiDiGap), a novel approach for flexible policy representation and imitation learning i

safetyarxiv-cs-ai
11 Jun 2026
Safety

This was perhaps the most controversial aspect of the guardrails around Fable, now being rolled back.

DGX agent

This was perhaps the most controversial aspect of the guardrails around Fable, now being rolled back. Very pleased to hear Anthropic have walked back this policy https://simonwillison.net/2026/Jun/11/

safetyethan-mollick--x
11 Jun 2026
Safety

To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending

DGX agent

arXiv:2606.11201v1 Announce Type: cross Abstract: The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions. Among

safetyarxiv-cs-ai
11 Jun 2026
Safety

Toward Preference-aligned Large Language Models via Residual-based Model Steering

DGX agent

arXiv:2509.23982v2 Announce Type: replace-cross Abstract: Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences. Existing approache

safetyarxiv-cs-ai
11 Jun 2026
Safety

Towards a Bridge Layer Between Bibliographic and Formalized Mathematical Knowledge

DGX agent

arXiv:2606.11430v1 Announce Type: cross Abstract: Mathematical knowledge is split between bibliographic databases (e.g., MathSciNet, zbMATH Open) and formal proof libraries (e.g., Lean mathlib), preve

safetyarxiv-cs-ai
11 Jun 2026
Safety

Towards Conditional Feature Alignment for Cross-Domain Counting

DGX agent

arXiv:2506.17137v3 Announce Type: replace Abstract: Object counting models often degrade under cross-domain deployment because density composition varies across domains and is itself task-relevant. St

safetyarxiv-cs-cv
11 Jun 2026
Safety

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Personality Assessment

DGX agent

arXiv:2606.11269v1 Announce Type: new Abstract: Personality assessment aims to infer stable personality traits from dynamic behaviors across language, voice, and facial cues. Since different personali

safetyarxiv-cs-cv
11 Jun 2026
Safety

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning

DGX agent

arXiv:2606.12372v1 Announce Type: cross Abstract: Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy imp

safetyarxiv-cs-lg
11 Jun 2026
Safety

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA

DGX agent

arXiv:2606.11740v1 Announce Type: cross Abstract: We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a

safetyarxiv-cs-cl
11 Jun 2026
Safety

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

DGX agent

arXiv:2606.11681v1 Announce Type: new Abstract: We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phonem

safetyarxiv-cs-cl
11 Jun 2026
Safety

Urban Heat MiniCubes: An AI-Ready dataset for urban heat research

DGX agent

arXiv:2606.11534v1 Announce Type: cross Abstract: Urban heat is amplified by impermeable surfaces and heterogeneous built environments, yet street-level variability remains difficult to quantify becau

safetyarxiv-cs-lg
11 Jun 2026
Safety

Very pleased to hear Anthropic have walked back this policy https://simonwillison.net/2026/Jun/11/anthropic-walks-back-policy/

DGX agent

Very pleased to hear Anthropic have walked back this policy https://simonwillison.net/2026/Jun/11/anthropic-walks-back-policy/ BREAKING NEWS: Anthropic's latest model will NOT help you if it thinks yo

safetyjeremy-howard--x
11 Jun 2026
Safety

ViT-FREE: Efficient Face Recognition via Early Exiting and Synthetic Adaptation

DGX agent

arXiv:2606.12023v1 Announce Type: new Abstract: Vision Transformers (ViTs) have gained significant attention in computer vision and shown strong potential for face recognition (FR). However, their hig

safetyarxiv-cs-cv
11 Jun 2026
Safety

VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving

DGX agent

arXiv:2606.12396v1 Announce Type: new Abstract: Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D wo

safetyarxiv-cs-cv
11 Jun 2026
Safety

What are smart people are saying about OpenAI's IPO filing? 🤔

DGX agent

Gary Marcus discusses expert commentary and analysis regarding OpenAI's initial public offering filing, likely examining implications for the AI industry, company valuation, and competitive landscape.

safetygary-marcus--x
11 Jun 2026
Safety

When Context Returns: Toward Robust Internalization in On-Policy Distillation

DGX agent

arXiv:2606.11627v1 Announce Type: cross Abstract: Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so th

safetyarxiv-cs-ai
11 Jun 2026
Safety

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?

DGX agent

arXiv:2510.02660v2 Announce Type: replace-cross Abstract: When researchers claim AI systems possess ToM or mental models, they are fundamentally discussing behavioral predictions and bias corrections

safetyarxiv-cs-ai
11 Jun 2026
Safety

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

DGX agent

arXiv:2606.12199v1 Announce Type: cross Abstract: Spoken dialogue models typically start from text LLM backbones, yet reasoning often degrades when conditioning on speech instead of text. We attribute

safetyarxiv-cs-cl
11 Jun 2026
Safety

Wrong horse and too early. Gary nailed it. Time will show this.

DGX agent

Wrong horse and too early. Gary nailed it. Time will show this. Masa can hype AI all he likes, but apparently he can’t even get a loan on his OpenAI shares. 🤔 But I actually half agree—and it might su

safetygary-marcus--x
11 Jun 2026
Safety

100%, hallucinations are baked in (as I have been saying since 2001) and that is why basically no LLM company afford to operate in Germany n…

DGX agent

100%, hallucinations are baked in (as I have been saying since 2001) and that is why basically no LLM company afford to operate in Germany now. We need a better technology. @GaryMarcus Lol, and Google

safetygary-marcus--x
10 Jun 2026
Safety

3SPO: State-Score-Supervised Policy Optimization for LLM Agents

DGX agent

arXiv:2606.09961v1 Announce Type: cross Abstract: Training large language models (LLMs) as autonomous agents via reinforcement learning (RL) has enabled frontier models to achieve superhuman performan

safetyarxiv-cs-ai
10 Jun 2026
Safety

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

DGX agent

arXiv:2410.15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct P

safetyarxiv-cs-ai
10 Jun 2026
Safety

A fine-grained attention and geometric correspondence model for musculoskeletal risk classification in athletes using multimodal visual and skeletal features

DGX agent

arXiv:2509.05913v3 Announce Type: replace Abstract: Musculoskeletal disorders pose significant risks to athletes, and early risk assessment is essential for prevention. However, most existing methods

safetyarxiv-cs-cv
10 Jun 2026
Safety

A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation

DGX agent

arXiv:2606.10366v1 Announce Type: cross Abstract: Simulation has become an essential tool for evaluating and improving vision-language-action (VLA) policies, offering scalable, reproducible, and contr

safetyarxiv-cs-ai
10 Jun 2026
Safety

A Source Domain is All You Need: Source-Only Cross-OS Transfer Learning for APT Anomaly Detection via Semantic Alignment and Optimal Transport

DGX agent

arXiv:2606.10216v1 Announce Type: cross Abstract: Advanced Persistent Threats (APTs) are stealthy, multi-stage cyberattacks whose detection is difficult due to scarce labeled traces, severe class imba

safetyarxiv-cs-ai
10 Jun 2026
Safety

A Unified Multi-Modal Framework for Intelligent Financial Systems: Integrating Reinforcement Learning, High-Frequency Trading, and Game-Theoretic Approaches with Cross-Modal Sentiment Analysis

DGX agent

arXiv:2606.10412v1 Announce Type: new Abstract: The rapid evolution of financial technology demands sophisticated artificial intelligence systems capable of handling diverse challenges across multiple

safetyarxiv-cs-ai
10 Jun 2026
Safety

Adoption of Generative Artificial Intelligence in the German Software Engineering Industry: An Empirical Study

DGX agent

arXiv:2601.16700v2 Announce Type: replace-cross Abstract: Generative artificial intelligence (GenAI) tools have seen rapid adoption among software developers. While adoption rates in the industry are

safetyarxiv-cs-ai
10 Jun 2026
Safety

AI oligarchs are trying to buy elections (again)

DGX agent

AI oligarchs are trying to buy elections (again) AI titans are spending 1.3M to shape Utah primaries. Defending Our Values, run by Chris Stewart, spent 880K to support Rep. Celeste Maloy. DOV operates

safetygary-marcus--x
10 Jun 2026
Safety

Alignment Defends LLMs from Property Inference Attacks

DGX agent

arXiv:2606.10217v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly fine-tuned on domain-specific datasets that may contain sensitive, dataset-level properties. Recent work h

safetyarxiv-cs-lg
10 Jun 2026
Safety

Alongside it, Anthropic is releasing a proposal for how governments can address the risks posed by frontier AI and a policy framework for jo…

DGX agent

Alongside it, Anthropic is releasing a proposal for how governments can address the risks posed by frontier AI and a policy framework for job displacement, for which we intend to provide substantial f

safetydario-amodei--x
10 Jun 2026
Safety

An LLM-Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models

DGX agent

arXiv:2606.09843v1 Announce Type: cross Abstract: Large language models (LLMs) produce stable self-reports on personality inventories, but these self-reports do not predict observed behavior. Whether

safetyarxiv-cs-ai
10 Jun 2026
Safety

and see for deeper analysis https://open.substack.com/pub/garymarcus/p/breaking-news-and-how-the-end-might?r=8tdk6&utm_medium=ios

DGX agent

Gary Marcus discusses breaking news and potential existential risks or significant developments related to AI systems, likely examining how current AI capabilities and limitations might lead to critic

safetygary-marcus--x
10 Jun 2026
Safety

AnimaSpark: A Feed-Forward Method for Animating Arbitrary 3D Objects

DGX agent

arXiv:2606.10988v1 Announce Type: new Abstract: While recent advancements in generative AI have substantially accelerated static 3D model creation workflows, the synthesis of category-agnostic 3D anim

safetyarxiv-cs-cv
10 Jun 2026
Safety

Anthropic releases two policy proposals on how governments should address catastrophic risks and manage labor market disruption from advanced AI systems (Anthropic)

DGX agent

Anthropic: Anthropic releases two policy proposals on how governments should address catastrophic risks and manage labor market disruption from advanced AI systems — AI is advancing at exponential spe

safetytechmeme
10 Jun 2026
Safety

Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans

DGX agent

arXiv:2606.10953v1 Announce Type: new Abstract: Furnished floor plans are fundamental to real estate visualization, interior design, and architectural workflows. However, progress in automatic furnitu

safetyarxiv-cs-ai
10 Jun 2026
Safety

are you kidding me

DGX agent

are you kidding me Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was built to handle. The essay lays out where I

safetyjeremy-howard--x
10 Jun 2026
← Previous
1…139140141142143…302
Next →