AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,809 results
11 Jun 2026

Mitigating Disparate Impact of Differentially Private Learning through Bounded Adaptive Clipping

SafetyDGX agent

arXiv:2506.01396v2 Announce Type: replace Abstract: Differential privacy (DP) has become an essential framework for privacy-preserving machine learning. Existing DP learning methods, however, often ha

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching

SafetyDGX agent

arXiv:2606.12215v1 Announce Type: new Abstract: The explosive growth of user-generated video content on online platforms is accompanied by the emergence of numerous near-duplicate videos--videos that

Noise-Guided Transport for Imitation Learning

SafetyDGX agent

arXiv:2509.26294v2 Announce Type: replace-cross Abstract: We consider imitation learning in the low-data regime, where only a limited number of expert demonstrations are available. In this setting, me


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Offline Diffusion Policy for Multi-User Delay-Constrained Scheduling

SafetyDGX agent

arXiv:2501.12942v2 Announce Type: replace Abstract: Effective multi-user delay-constrained scheduling is crucial in various real-world applications, including embodied AI, instant messaging, live stre

One Jailbreak, Many Tongues: Learning Language-Insensitive Intention Representations for Multilingual Jailbreak Detection

SafetyDGX agent

arXiv:2606.11202v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in applications for global multilingual users, yet safety training remains concentrated in domina

Open Materials Generation with Inference-Time Reinforcement Learning

SafetyDGX agent

arXiv:2602.00424v2 Announce Type: replace Abstract: Continuous-time generative models for crystalline materials enable inverse materials design by learning to predict stable crystal structures, but in

🚨 OpenAI pondering big price cuts, per WSJ scoop:

SafetyDGX agent

OpenAI is reportedly considering significant price reductions for its AI services, according to a Wall Street Journal report shared by AI researcher Gary Marcus. The move would likely reflect competit

PAWS: Preference Learning with Advantage-Weighted Segments

SafetyDGX agent

arXiv:2606.11982v1 Announce Type: new Abstract: Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demon

Performance Analysis of YOLOv11 and YOLOv8 for Mixed Traffic Object Detection under Adverse Weather Conditions in Developing Countries

SafetyDGX agent

arXiv:2606.12066v1 Announce Type: new Abstract: In modern vehicular systems, robust performance under harsh conditions has become a critical problem of autonomous driving. Our study delivers a compreh

Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding

SafetyDGX agent

arXiv:2606.11838v1 Announce Type: new Abstract: Reward models for text-to-video (T2V) generation guide post-training but often fail at fine-grained semantic alignment. We trace this to two structural

ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward

SafetyDGX agent

arXiv:2606.11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning. Recent post-training with reinforcement learning under verifiable rewards (RLVR)

Redesign Mixture-of-Experts Routers with Manifold Power Iteration

SafetyDGX agent

arXiv:2606.12397v1 Announce Type: cross Abstract: Router is the cornerstone component to the Mixture-of-Experts models. Serving as expert proxies, the rows of the router matrix compute their similarit

Reinforcement Learning Disrupts Gradient-Based Adversarial Optimization

SafetyDGX agent

arXiv:2606.12251v1 Announce Type: cross Abstract: Gradient-based adversarial attacks remain a dominant threat to deep neural networks (DNNs), as they exploit gradient information to efficiently optimi

Reinforcement Learning with Action-Triggered Observations

SafetyDGX agent

arXiv:2510.02149v2 Announce Type: replace Abstract: We introduce Action-Triggered Sporadically Traceable Markov Decision Processes (ATST-MDPs), a reinforcement learning framework for partial observabi

Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies

SafetyDGX agent

arXiv:2601.08136v2 Announce Type: replace Abstract: Diffusion and flow policies are gaining prominence in online reinforcement learning (RL) due to their expressive power, yet training them efficientl

Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models

SafetyDGX agent

arXiv:2606.11409v1 Announce Type: cross Abstract: Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly tr

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

SafetyDGX agent

arXiv:2606.11709v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distri

Runtime Enforcement of Hybrid System Properties

SafetyDGX agent

arXiv:2606.12022v1 Announce Type: cross Abstract: Runtime enforcement has emerged as a promising approach for ensuring the safety of autonomous and cyber-physical systems operating in uncertain and dy

SAFER-Nav: Enhancing Safety for Visual Robot Navigation via Segmentation-Aware Fine-Tuning

SafetyDGX agent

arXiv:2606.11636v1 Announce Type: new Abstract: Vision-based navigation models, particularly foundation models, generate viable trajectories from RGB observations alone. However, even state-of-the-art

SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment

SafetyDGX agent

arXiv:2606.11512v1 Announce Type: new Abstract: Large language models increasingly express uncertainty through natural-language statements, yet these expressions often fail to reflect the model's samp

Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning

SafetyDGX agent

arXiv:2603.14867v4 Announce Type: replace-cross Abstract: Many strategic decision-making problems, such as environment design for warehouse robots, can be naturally formulated as bi-level reinforcemen

Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version

SafetyDGX agent

arXiv:2606.11399v1 Announce Type: new Abstract: Large Language Models (LLMs) are deployed across cultural contexts but often reflect homogenized values inherited from training data. Evaluations of cul

Schutzen: Evaluating LLM Safety in Bulgarian and German Contexts

SafetyDGX agent

arXiv:2606.11316v1 Announce Type: new Abstract: Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disr

Seeing Before Colliding: Anticipatory Safe RL with Frozen Vision-Language Models

SafetyDGX agent

arXiv:2606.11266v1 Announce Type: new Abstract: The cost signal that constrained-RL algorithms optimize against is almost always reactive: the simulator emits a non-zero cost only after a collision ha

Semantically-Aware Diver Activity Recognition Framework for Effective Underwater Multi-Human-Robot Collaboration

SafetyDGX agent

arXiv:2606.12374v1 Announce Type: cross Abstract: Effective multi-human-robot collaboration is essential for expanding human-led operations in the challenging and high-risk underwater environment. For

Signed Compression Progress on a Sealed Audit is Goodhart-Resistant

SafetyDGX agent

arXiv:2606.11417v1 Announce Type: cross Abstract: Compression progress is a long-standing proposal for intrinsic motivation: reward an agent when its world model becomes better at predicting or compre

SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation

SafetyDGX agent

arXiv:2511.05203v3 Announce Type: replace Abstract: Today's autonomous agents, largely driven by foundation models (FMs), can understand natural language instructions and solve long-horizon tasks with

Sovereign Assurance Boundary: Certificate-Bound Admission for Agentic Infrastructure

SafetyDGX agent

arXiv:2606.11632v1 Announce Type: cross Abstract: Agentic infrastructure introduces a critical control-plane authorization problem: non-deterministic reasoning systems can propose high-stakes mutation

Spectrally Regularized Latent Flow Matching for Turbulence Generation

SafetyDGX agent

arXiv:2606.11691v1 Announce Type: new Abstract: Latent diffusion and flow matching have emerged as leading approaches for synthetic turbulence generation, yet they systematically under-represent dissi

Steering Multirobot Behavior via Closed-Loop Affine Activation Editing

SafetyDGX agent

arXiv:2606.11489v1 Announce Type: new Abstract: Real-world robots need to adapt their behavior beyond the envelope of their pre-trained policy. Policy finetuning or retraining are options, but they ri

SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning

SafetyDGX agent

arXiv:2606.11770v1 Announce Type: new Abstract: Spatial reasoning remains a challenge for Multimodal Large Language Models (MLLMs), as it requires reliable multi-hop inference over both intermediate s

TacCoRL: Integrating Tactile Feedback into VLA via Simulation

SafetyDGX agent

arXiv:2606.11743v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often mis

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning

SafetyDGX agent

arXiv:2606.11853v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) depend on in-context learning (ICL) for rapid task adaptation, but their scalability is severely limited by

The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network

SafetyDGX agent

arXiv:2508.21380v3 Announce Type: replace-cross Abstract: Recent mechanistic work has uncovered learned algorithms within neural networks, from modular arithmetic to search and planning in game-playin

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

SafetyDGX agent

arXiv:2606.11918v1 Announce Type: new Abstract: Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approa

The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning

SafetyDGX agent

arXiv:2505.03296v2 Announce Type: replace-cross Abstract: We present Mixture of Discrete-time Gaussian Processes (MiDiGap), a novel approach for flexible policy representation and imitation learning i

This was perhaps the most controversial aspect of the guardrails around Fable, now being rolled back.

SafetyDGX agent

This was perhaps the most controversial aspect of the guardrails around Fable, now being rolled back. Very pleased to hear Anthropic have walked back this policy https://simonwillison.net/2026/Jun/11/

To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending

SafetyDGX agent

arXiv:2606.11201v1 Announce Type: cross Abstract: The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions. Among

Toward Preference-aligned Large Language Models via Residual-based Model Steering

SafetyDGX agent

arXiv:2509.23982v2 Announce Type: replace-cross Abstract: Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences. Existing approache

Towards a Bridge Layer Between Bibliographic and Formalized Mathematical Knowledge

SafetyDGX agent

arXiv:2606.11430v1 Announce Type: cross Abstract: Mathematical knowledge is split between bibliographic databases (e.g., MathSciNet, zbMATH Open) and formal proof libraries (e.g., Lean mathlib), preve

Towards Conditional Feature Alignment for Cross-Domain Counting

SafetyDGX agent

arXiv:2506.17137v3 Announce Type: replace Abstract: Object counting models often degrade under cross-domain deployment because density composition varies across domains and is itself task-relevant. St

Traceable Virtual Sea Trials in the Marine Robotics Unity Simulator for Manoeuvring Assessment of Unmanned Surface Vehicles

SafetyDGX agent

arXiv:2606.12349v1 Announce Type: new Abstract: Accurate identification of hydrodynamic derivatives is essential for control and navigation of Unmanned Surface Vehicles (USVs), but high-fidelity manoe

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Personality Assessment

SafetyDGX agent

arXiv:2606.11269v1 Announce Type: new Abstract: Personality assessment aims to infer stable personality traits from dynamic behaviors across language, voice, and facial cues. Since different personali

UGV-Conditioned Multi-UAV Informative Planning on a Shared Exposure Belief

SafetyDGX agent

arXiv:2606.12306v1 Announce Type: new Abstract: Safe ground navigation in large, threat-augmented environments requires aerial support that actively reduces the risks that a ground vehicle faces along

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning

SafetyDGX agent

arXiv:2606.12372v1 Announce Type: cross Abstract: Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy imp

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA

SafetyDGX agent

arXiv:2606.11740v1 Announce Type: cross Abstract: We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

SafetyDGX agent

arXiv:2606.11681v1 Announce Type: new Abstract: We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phonem

Urban Heat MiniCubes: An AI-Ready dataset for urban heat research

SafetyDGX agent

arXiv:2606.11534v1 Announce Type: cross Abstract: Urban heat is amplified by impermeable surfaces and heterogeneous built environments, yet street-level variability remains difficult to quantify becau

Very pleased to hear Anthropic have walked back this policy https://simonwillison.net/2026/Jun/11/anthropic-walks-back-policy/

SafetyDGX agent

Very pleased to hear Anthropic have walked back this policy https://simonwillison.net/2026/Jun/11/anthropic-walks-back-policy/ BREAKING NEWS: Anthropic's latest model will NOT help you if it thinks yo

ViT-FREE: Efficient Face Recognition via Early Exiting and Synthetic Adaptation

SafetyDGX agent

arXiv:2606.12023v1 Announce Type: new Abstract: Vision Transformers (ViTs) have gained significant attention in computer vision and shown strong potential for face recognition (FR). However, their hig

VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving

SafetyDGX agent

arXiv:2606.12396v1 Announce Type: new Abstract: Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D wo

What are smart people are saying about OpenAI's IPO filing? 🤔

SafetyDGX agent

Gary Marcus discusses expert commentary and analysis regarding OpenAI's initial public offering filing, likely examining implications for the AI industry, company valuation, and competitive landscape.

When Context Returns: Toward Robust Internalization in On-Policy Distillation

SafetyDGX agent

arXiv:2606.11627v1 Announce Type: cross Abstract: Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student model so th

When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?

SafetyDGX agent

arXiv:2510.02660v2 Announce Type: replace-cross Abstract: When researchers claim AI systems possess ToM or mental models, they are fundamentally discussing behavioral predictions and bias corrections

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

SafetyDGX agent

arXiv:2606.12199v1 Announce Type: cross Abstract: Spoken dialogue models typically start from text LLM backbones, yet reasoning often degrades when conditioning on speech instead of text. We attribute

Wrong horse and too early. Gary nailed it. Time will show this.

SafetyDGX agent

Wrong horse and too early. Gary nailed it. Time will show this. Masa can hype AI all he likes, but apparently he can’t even get a loan on his OpenAI shares. 🤔 But I actually half agree—and it might su

10 Jun 2026

100%, hallucinations are baked in (as I have been saying since 2001) and that is why basically no LLM company afford to operate in Germany n…

SafetyDGX agent

100%, hallucinations are baked in (as I have been saying since 2001) and that is why basically no LLM company afford to operate in Germany now. We need a better technology. @GaryMarcus Lol, and Google

3SPO: State-Score-Supervised Policy Optimization for LLM Agents

SafetyDGX agent

arXiv:2606.09961v1 Announce Type: cross Abstract: Training large language models (LLMs) as autonomous agents via reinforcement learning (RL) has enabled frontier models to achieve superhuman performan

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

SafetyDGX agent

arXiv:2410.15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct P

A fine-grained attention and geometric correspondence model for musculoskeletal risk classification in athletes using multimodal visual and skeletal features

SafetyDGX agent

arXiv:2509.05913v3 Announce Type: replace Abstract: Musculoskeletal disorders pose significant risks to athletes, and early risk assessment is essential for prevention. However, most existing methods

← Previous
1…7475767778…214
Next →