AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,490 results
10 Jun 2026

Using Probabilistic Programs to Train Inductive Reasoning in Large Language Models

SafetyDGX agent

arXiv:2606.09856v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is veri

Visual-TCAV: Concept-based Attribution and Saliency Maps for Post-hoc Explainability in Image Classification

SafetyDGX agent

arXiv:2411.05698v3 Announce Type: replace-cross Abstract: Convolutional Neural Networks (CNNs) have shown remarkable performance in image classification. However, interpreting their predictions is cha

Warren to SEC, lightly paraphrased: “Do your f’ing job, and don’t let retail investors get screwed”

SafetyDGX agent

Senator Elizabeth Warren criticized the SEC for insufficient enforcement and investor protection, urging the agency to strengthen oversight and prevent harm to retail investors. The post, shared by AI

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

What if, Germany locked itself out of the LLM race and • Had students who actually learned things in high school, instead of turning in prom…

SafetyDGX agent

What if, Germany locked itself out of the LLM race and • Had students who actually learned things in high school, instead of turning in prompt outputs they barely read • Emerged from the sea of slop •

What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents

SafetyDGX agent

arXiv:2606.09421v2 Announce Type: replace Abstract: Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, valida

When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models

SafetyDGX agent

arXiv:2512.06343v3 Announce Type: replace-cross Abstract: Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

SafetyDGX agent

arXiv:2606.11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic

Wow. This is worth watching, might have huge impact. @SenWarren makes some valid points, and the SEC owes her public answers

SafetyDGX agent

Wow. This is worth watching, might have huge impact. @SenWarren makes some valid points, and the SEC owes her public answers Sen. Warren calls on SEC to delay SpaceX IPO @CNBC https://www.cnbc.com/202

YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale

SafetyDGX agent

arXiv:2606.10244v1 Announce Type: cross Abstract: We introduce Yielding Universal Bidigital Interface (YUBI), a finger-aligned gripper designed to enable intuitive, ergonomic, and scalable data collec

9 Jun 2026

A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales

SafetyDGX agent

arXiv:2606.09470v1 Announce Type: cross Abstract: Automated L2 speech assessment can assign proficiency labels, but often lacks interpretability. We propose a rubric-guided SpeechLLM for multi-aspect,

A Geometric Unification of Concept Learning with Concept Cones

SafetyDGX agent

arXiv:2512.07355v2 Announce Type: replace Abstract: Two traditions of interpretability have evolved side by side but seldom spoken to each other: Concept Bottleneck Models (CBMs), which prescribe what

A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control

SafetyDGX agent

arXiv:2606.08517v1 Announce Type: new Abstract: Selective predictors answer on confident inputs and abstain elsewhere; deploying one safely needs a single finite-sample certificate that simultaneously

A Mixed Diet Makes DINO An Omnivorous Vision Encoder

SafetyDGX agent

arXiv:2602.24181v2 Announce Type: replace-cross Abstract: Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their features a

A Unifying Lens on Reward Uncertainty in RLHF

SafetyDGX agent

arXiv:2606.09073v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) is bottlenecked by reward hacking, where the policy exploits errors in a proxy reward model (RM) and

Ablation-Reversible Heads Don't Transfer: A Stress Test for Mechanistic Role Claims in Transformers

SafetyDGX agent

arXiv:2606.08292v1 Announce Type: new Abstract: In mechanistic interpretability, attention heads are commonly elevated to role claims (e.g., 'this head represents addition') when they are necessary fo

absolutely correct

SafetyDGX agent

absolutely correct @GaryMarcus I've recently learned not to believe anything that anyone says anymore when it comes to AI companies. They're all competing to win against each other while exploiting, s

ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies

SafetyDGX agent

arXiv:2606.08508v1 Announce Type: cross Abstract: Generative robot policies fail unpredictably at deployment: they hesitate at critical moments, drift off-task, or commit to unrecoverable actions. Exi

Adaptive Loss Balancing for Noise-Robust GRPO in Generative Recommendation

SafetyDGX agent

arXiv:2606.08480v1 Announce Type: cross Abstract: Reinforcement learning (RL) presents a promising avenue for enhancing generative recommendation beyond supervised imitation, leveraging reward signals

Agent Economics: An Entropy-Controlled Pluralistic Alignment Framework for Preventing Artificial Hivemind in Autonomous Agents

SafetyDGX agent

arXiv:2606.09039v1 Announce Type: new Abstract: This study proposes the Behavioral Protocol Framework (BPF), an entropy-controlled pluralistic alignment framework designed to address two critical chal

AgriGov: A Structured Multilingual Dataset Curation for Indian Government Schemes for Farmers

SafetyDGX agent

arXiv:2606.08272v1 Announce Type: cross Abstract: AgriGov is a curated, trilingual (English-Hindi-Marathi) dataset designed to address the scarcity of domain-grounded multilingual resources for agricu

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

SafetyDGX agent

arXiv:2606.09811v1 Announce Type: cross Abstract: World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical

AI Code Sandboxes: A Comparative Security Study. Part 1 of 2 -- Engine-Level Properties (Attack Surface, Leakage, Stackability, CVE History, Patch Cadence, Fuzzing)

SafetyDGX agent

arXiv:2606.08433v1 Announce Type: cross Abstract: This paper reads six engine-level measurements together -- 1.1 host attack surface, 1.2 information leakage, 1.3 defense-in-depth stackability, 1.4 pu

AI-Integrated Learning Management System for Middle School: A Longitudinal Study of Learning Outcomes Through High School and Beyond

SafetyDGX agent

arXiv:2606.07544v1 Announce Type: cross Abstract: Middle school is a key window for building core academic skills and the learning routines students carry into later grades, yet many students still fa

Aligned but Not Partner-Specific: Distinguishing How Multimodal LLM Agents Succeed in Reference Games Without Human-Like Conventions

SafetyDGX agent

arXiv:2606.08081v1 Announce Type: cross Abstract: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in s

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

SafetyDGX agent

arXiv:2507.08920v4 Announce Type: replace-cross Abstract: We introduce AMix-1, a powerful protein foundation model built on Bayesian Flow Networks and empowered by a systematic training methodology, e

An Agency-Transferring Model-Free Policy Enhancement Technique

SafetyDGX agent

arXiv:2606.09825v1 Announce Type: cross Abstract: Training reinforcement learning (RL) policies from scratch is costly: it requires careful reward and environment design, extensive tuning, and substan

Anchor-Conditioned Compositional Control for Landscape Image Generation

SafetyDGX agent

arXiv:2606.07638v1 Announce Type: cross Abstract: Image generative models, though widely used as creative tools, offer limited support for the kind of compositional control that photographers and visu

Anthropic didn’t just add guardrails to make Mythos safer; they added guardrails to protect their own IP. *Their own IP*. They are still as …

SafetyDGX agent

Anthropic didn’t just add guardrails to make Mythos safer; they added guardrails to protect their own IP. *Their own IP*. They are still as happy as fuck to build their AI on other people’s IP. Intere

Anthropic played the media like a fiddle. From “untold catastrophe” to “check our latest model”, in two months and a day 🙄 (cc @tomfriedman…

SafetyDGX agent

Anthropic played the media like a fiddle. From “untold catastrophe” to “check our latest model”, in two months and a day 🙄 (cc @tomfriedman) This is the scary phase of AI — a model deemed so powerful

Autonomous Aerial Manipulation via Contextual Contrastive Meta Reinforcement Learning

SafetyDGX agent

arXiv:2606.08533v1 Announce Type: new Abstract: Unmanned aerial vehicles (UAVs) are increasingly being deployed in logistics, service robotics, and other real-world applications, creating a growing de

Autonomous FPV Flight with Translational Optical Flow and Uncertainty Mask

SafetyDGX agent

arXiv:2606.09088v1 Announce Type: new Abstract: Autonomous FPV quadrotor flight in complex environments using a monocular RGB camera as the sole exteroceptive sensor remains a fundamental challenge. R

Autonomous Obstacle Removal for Excavators through Policy Learning with Particle Simulation

SafetyDGX agent

arXiv:2606.09183v1 Announce Type: new Abstract: Autonomous obstacle removal from the ground is an important earthwork task, but this is difficult to automate because an excavator must adapt its excava

Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection

SafetyDGX agent

arXiv:2606.09258v1 Announce Type: new Abstract: Vision-language-action (VLA) policies can deviate from nominal trajectories during manipulation, even when tasks remain physically feasible. Recovering

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

SafetyDGX agent

arXiv:2606.09802v1 Announce Type: cross Abstract: We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, eac

BareWave: Waveform-Native Flow-Matching Text-to-Speech

SafetyDGX agent

arXiv:2606.09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling. In text-to-spee

Beware of GeeksBearing Gifts: Building True EU Frontier AI Sovereignty

SafetyDGX agent

arXiv:2606.07536v1 Announce Type: cross Abstract: Frontier artificial intelligence is reshaping all aspects of society, from economic output or military capability to democratic institutions. The EU i

Beyond Homophily: Towards Generalized Graph Reconstruction Attack and Defense

SafetyDGX agent

arXiv:2606.08067v1 Announce Type: new Abstract: Graph neural networks (GNNs) are widely deployed on relational data, yet they can leak sensitive or proprietary information about the training graph adj

Beyond Neural Collapse: Task-Intrinsic Geometry Governs Neural Representations in Modular Arithmetic

SafetyDGX agent

arXiv:2606.08985v1 Announce Type: new Abstract: While neural collapse (NC) predicts that a K-class-balanced classifier should organize terminal representations as a (K-1)-dimensional simplex equiangul

Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

SafetyDGX agent

arXiv:2606.09076v1 Announce Type: new Abstract: Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric score

Boundary Variance Inflation Causes Acquisition Bias in Gaussian Processes

SafetyDGX agent

arXiv:2606.07561v1 Announce Type: new Abstract: Gaussian processes with stationary kernels on bounded domains exhibit inflated posterior variance near the boundary. Despite being a long-recognized art

Breaking the Tokenizer Barrier: On-Policy Distillation across Model Families

SafetyDGX agent

arXiv:2606.09456v1 Announce Type: new Abstract: On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain exp

Bridging Expert Knowledge and Automated Feature Engineering via Self-Evolution

SafetyDGX agent

arXiv:2606.08800v1 Announce Type: new Abstract: In high-stakes settings such as brand compliance, clinical care, and content moderation, machine learning cannot be deployed as opaque oracles: practiti

Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis

SafetyDGX agent

arXiv:2606.07533v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) effectively integrate text and audio to interpret context in complex interactive dialogues. However, the inte

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs

SafetyDGX agent

arXiv:2606.09371v1 Announce Type: new Abstract: Tool learning enables LLMs to invoke external tools to accomplish tasks. Prior studies have demonstrated the effectiveness of a hierarchical structure:

Causal Semantic Alignment for LLM-based Time Series Forecasting

SafetyDGX agent

arXiv:2606.08262v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have opened new possibilities for time series forecasting by enabling alignment between temporal pattern

Causal Transfer in Medical Image Analysis

SafetyDGX agent

arXiv:2603.24388v2 Announce Type: replace Abstract: Medical imaging models frequently fail when deployed across hospitals, scanners, populations, or imaging protocols due to domain shift, limiting the

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

SafetyDGX agent

arXiv:2606.08420v1 Announce Type: new Abstract: Vision-language models (VLMs) pretrained on large-scale image-text pairs demonstrate strong image-level understanding, but are primarily optimized for g

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

SafetyDGX agent

arXiv:2606.09639v1 Announce Type: new Abstract: The fidelity and structural diversity of training datasets fundamentally determine the capabilities of video generation models. While commercial systems

Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.09138v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) has become an important post-training paradigm for turning LLMs from static chatbots into interactive agents, giving

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning

SafetyDGX agent

arXiv:2509.25004v2 Announce Type: replace Abstract: Online reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning abilities of large languag

Cold feet about coating the entire surface of the earth in data centers?

SafetyDGX agent

Gary Marcus likely discusses concerns about the environmental and practical implications of exponentially expanding data center infrastructure across the globe, questioning whether covering Earth's su

Comparative evaluation of training strategies using partially labelled datasets for segmentation of white matter hyperintensities and stroke lesions in FLAIR MRI

SafetyDGX agent

arXiv:2601.20503v2 Announce Type: replace-cross Abstract: White matter hyperintensities (WMH) and ischaemic stroke lesions (ISL) are key imaging biomarkers of cerebral small vessel disease (SVD) detec

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning

SafetyDGX agent

arXiv:2606.08088v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has recently become a key paradigm for improving the reasoning abilities of Large Language Models

Constrained Paraphrase Consistency for LLM Hallucination Detection

SafetyDGX agent

arXiv:2606.08158v1 Announce Type: cross Abstract: Large language models (LLMs) can generate factually inconsistent claims, motivating accurate and scalable hallucination detectors. Prior work largely

Constrained user-item allocation for e-commerce marketing campaigns

SafetyDGX agent

arXiv:2606.09623v1 Announce Type: new Abstract: When running marketing campaigns, retailers must decide which products to promote and which users to target. These decisions are inherently coupled: eff

Constraint-Aware Optimization for Robust Protein Stability Prediction

SafetyDGX agent

arXiv:2606.08100v1 Announce Type: new Abstract: Multimodal DeltaDelta G predictors integrating protein language models with inverse-folding representations achieve strong in-distribution accuracy on t

Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps

SafetyDGX agent

arXiv:2606.09084v1 Announce Type: cross Abstract: Tool-using LLM agents interact with the world through actions that persist state in artifacts (e.g., workspace files or logs). Consequently, jailbreak

Context Over Compute Human-in-the-Loop Outperforms Iterative Chain-of-Thought Prompting in Interview Answer Quality

SafetyDGX agent

arXiv:2603.09995v2 Announce Type: replace-cross Abstract: Behavioral interview evaluation using large language models presents unique challenges that require structured assessment, realistic interview

Contrast encodes inductive bias: separating slow noise from dynamics in predictive representation learning

SafetyDGX agent

arXiv:2606.07770v1 Announce Type: new Abstract: Self-supervised methods that learn representations and predict dynamics fully in the latent space, such as JEPA, have been shown to confuse slowly varyi

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

SafetyDGX agent

arXiv:2606.07604v1 Announce Type: cross Abstract: Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs). However, this approa

← Previous
1…114115116117118…242
Next →