AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
Safety

A Unifying Lens on Reward Uncertainty in RLHF

DGX agent

arXiv:2606.09073v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) is bottlenecked by reward hacking, where the policy exploits errors in a proxy reward model (RM) and

safetyarxiv-cs-ai
9 Jun 2026
Safety

A VideoMAE-v2 Approach to Zero-Shot Traffic Accident Anticipation

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.09542v1 Announce Type: new Abstract: Traffic accident anticipation -- predicting the likelihood of an imminent collision at every frame of a dashcam video -- is safety-critical yet difficul

safetyarxiv-cs-cv
9 Jun 2026
Safety

Ablation-Reversible Heads Don't Transfer: A Stress Test for Mechanistic Role Claims in Transformers

DGX agent

arXiv:2606.08292v1 Announce Type: new Abstract: In mechanistic interpretability, attention heads are commonly elevated to role claims (e.g., 'this head represents addition') when they are necessary fo

safetyarxiv-cs-ai
9 Jun 2026
Safety

absolutely correct

DGX agent

absolutely correct @GaryMarcus I've recently learned not to believe anything that anyone says anymore when it comes to AI companies. They're all competing to win against each other while exploiting, s

safetygary-marcus--x
9 Jun 2026
Safety

ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies

DGX agent

arXiv:2606.08508v1 Announce Type: cross Abstract: Generative robot policies fail unpredictably at deployment: they hesitate at critical moments, drift off-task, or commit to unrecoverable actions. Exi

safetyarxiv-cs-ai
9 Jun 2026
Safety

Adaptive Loss Balancing for Noise-Robust GRPO in Generative Recommendation

DGX agent

arXiv:2606.08480v1 Announce Type: cross Abstract: Reinforcement learning (RL) presents a promising avenue for enhancing generative recommendation beyond supervised imitation, leveraging reward signals

safetyarxiv-cs-ai
9 Jun 2026
Safety

AeroSpectra Sentinel: An Auditable LLM Prompt-Chaining Decision-Support Workflow for Acute Asthma Risk Assessment from Respiratory Sounds and Clinical Signals

DGX agent

arXiv:2606.08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing,

safetyarxiv-cs-ai
9 Jun 2026
Safety

Agent Economics: An Entropy-Controlled Pluralistic Alignment Framework for Preventing Artificial Hivemind in Autonomous Agents

DGX agent

arXiv:2606.09039v1 Announce Type: new Abstract: This study proposes the Behavioral Protocol Framework (BPF), an entropy-controlled pluralistic alignment framework designed to address two critical chal

safetyarxiv-cs-ai
9 Jun 2026
Safety

AgriGov: A Structured Multilingual Dataset Curation for Indian Government Schemes for Farmers

DGX agent

arXiv:2606.08272v1 Announce Type: cross Abstract: AgriGov is a curated, trilingual (English-Hindi-Marathi) dataset designed to address the scarcity of domain-grounded multilingual resources for agricu

safetyarxiv-cs-ai
9 Jun 2026
Safety

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

DGX agent

arXiv:2606.09811v1 Announce Type: cross Abstract: World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical

safetyarxiv-cs-ai
9 Jun 2026
Safety

AI Assurance in UK Defence: Challenges in Operationalising JSP 936

DGX agent

arXiv:2606.09414v1 Announce Type: cross Abstract: This report examines practical challenges in operationalising JSP 936 Part 1 for AI assurance in UK Defence. Using a structured interpretive review of

safetyarxiv-cs-ai
9 Jun 2026
Safety

AI Code Sandboxes: A Comparative Security Study. Part 1 of 2 -- Engine-Level Properties (Attack Surface, Leakage, Stackability, CVE History, Patch Cadence, Fuzzing)

DGX agent

arXiv:2606.08433v1 Announce Type: cross Abstract: This paper reads six engine-level measurements together -- 1.1 host attack surface, 1.2 information leakage, 1.3 defense-in-depth stackability, 1.4 pu

safetyarxiv-cs-ai
9 Jun 2026
Safety

AI-Integrated Learning Management System for Middle School: A Longitudinal Study of Learning Outcomes Through High School and Beyond

DGX agent

arXiv:2606.07544v1 Announce Type: cross Abstract: Middle school is a key window for building core academic skills and the learning routines students carry into later grades, yet many students still fa

safetyarxiv-cs-ai
9 Jun 2026
Safety

alienating your customers before you IPO is maybe not the best idea, @AnthropicAI

DGX agent

alienating your customers before you IPO is maybe not the best idea, @AnthropicAI Brilliant idea! Next up: Apple randomly reboots your Mac if you're building competing tech, Gmail silently edits your

safetygary-marcus--x
9 Jun 2026
Safety

Aligned but Not Partner-Specific: Distinguishing How Multimodal LLM Agents Succeed in Reference Games Without Human-Like Conventions

DGX agent

arXiv:2606.08081v1 Announce Type: cross Abstract: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in s

safetyarxiv-cs-ai
9 Jun 2026
Safety

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

DGX agent

arXiv:2507.08920v4 Announce Type: replace-cross Abstract: We introduce AMix-1, a powerful protein foundation model built on Bayesian Flow Networks and empowered by a systematic training methodology, e

safetyarxiv-cs-ai
9 Jun 2026
Safety

An Agency-Transferring Model-Free Policy Enhancement Technique

DGX agent

arXiv:2606.09825v1 Announce Type: cross Abstract: Training reinforcement learning (RL) policies from scratch is costly: it requires careful reward and environment design, extensive tuning, and substan

safetyarxiv-cs-ai
9 Jun 2026
Safety

Anchor-Conditioned Compositional Control for Landscape Image Generation

DGX agent

arXiv:2606.07638v1 Announce Type: cross Abstract: Image generative models, though widely used as creative tools, offer limited support for the kind of compositional control that photographers and visu

safetyarxiv-cs-ai
9 Jun 2026
Safety

Anthropic didn’t just add guardrails to make Mythos safer; they added guardrails to protect their own IP. *Their own IP*. They are still as …

DGX agent

Anthropic didn’t just add guardrails to make Mythos safer; they added guardrails to protect their own IP. *Their own IP*. They are still as happy as fuck to build their AI on other people’s IP. Intere

safetygary-marcus--x
9 Jun 2026
Safety

Anthropic played the media like a fiddle. From “untold catastrophe” to “check our latest model”, in two months and a day 🙄 (cc @tomfriedman…

DGX agent

Anthropic played the media like a fiddle. From “untold catastrophe” to “check our latest model”, in two months and a day 🙄 (cc @tomfriedman) This is the scary phase of AI — a model deemed so powerful

safetygary-marcus--x
9 Jun 2026
Safety

Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks

DGX agent

arXiv:2604.01039v2 Announce Type: replace-cross Abstract: System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive

safetyarxiv-cs-ai
9 Jun 2026
Safety

Autonomous Aerial Manipulation via Contextual Contrastive Meta Reinforcement Learning

DGX agent

arXiv:2606.08533v1 Announce Type: new Abstract: Unmanned aerial vehicles (UAVs) are increasingly being deployed in logistics, service robotics, and other real-world applications, creating a growing de

safetyarxiv-cs-lg
9 Jun 2026
Safety

Autonomous FPV Flight with Translational Optical Flow and Uncertainty Mask

DGX agent

arXiv:2606.09088v1 Announce Type: new Abstract: Autonomous FPV quadrotor flight in complex environments using a monocular RGB camera as the sole exteroceptive sensor remains a fundamental challenge. R

safetyarxiv-cs-ro
9 Jun 2026
Safety

Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations

DGX agent

arXiv:2606.09122v1 Announce Type: cross Abstract: Cloud network infrastructure at hyperscale presents unique operational challenges where traditional human-driven incident response cannot keep pace wi

safetyarxiv-cs-ai
9 Jun 2026
Safety

Autonomous Obstacle Removal for Excavators through Policy Learning with Particle Simulation

DGX agent

arXiv:2606.09183v1 Announce Type: new Abstract: Autonomous obstacle removal from the ground is an important earthwork task, but this is difficult to automate because an excavator must adapt its excava

safetyarxiv-cs-ro
9 Jun 2026
Safety

Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection

DGX agent

arXiv:2606.09258v1 Announce Type: new Abstract: Vision-language-action (VLA) policies can deviate from nominal trajectories during manipulation, even when tasks remain physically feasible. Recovering

safetyarxiv-cs-ro
9 Jun 2026
Safety

Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care

DGX agent

arXiv:2606.08982v1 Announce Type: new Abstract: Baichuan-M4 is Baichuan Intelligence's clinical-grade medical large model, designed for continuous care rather than single-turn medical question answeri

safetyarxiv-cs-ai
9 Jun 2026
Safety

Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Context Drifts

DGX agent

arXiv:2606.09802v1 Announce Type: cross Abstract: We consider a variant of the linear contextual stochastic multi-armed bandits, where the learner must provide recommendations to a group of users, eac

safetyarxiv-cs-ai
9 Jun 2026
Safety

BareWave: Waveform-Native Flow-Matching Text-to-Speech

DGX agent

arXiv:2606.09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling. In text-to-spee

safetyarxiv-cs-ai
9 Jun 2026
Safety

Beware of GeeksBearing Gifts: Building True EU Frontier AI Sovereignty

DGX agent

arXiv:2606.07536v1 Announce Type: cross Abstract: Frontier artificial intelligence is reshaping all aspects of society, from economic output or military capability to democratic institutions. The EU i

safetyarxiv-cs-ai
9 Jun 2026
Safety

Beyond Accuracy: Interpreting Topic Representation in Suicide Ideation Detection Models

DGX agent

arXiv:2606.07714v1 Announce Type: cross Abstract: Suicide ideation detection models are typically evaluated using aggregate performance metrics, yet little is known about how they internally represent

safetyarxiv-cs-ai
9 Jun 2026
Safety

Beyond Homophily: Towards Generalized Graph Reconstruction Attack and Defense

DGX agent

arXiv:2606.08067v1 Announce Type: new Abstract: Graph neural networks (GNNs) are widely deployed on relational data, yet they can leak sensitive or proprietary information about the training graph adj

safetyarxiv-cs-lg
9 Jun 2026
Safety

Beyond Linear Activation Steering: Invertible Latent Transformations for Controlling LLM Behavior

DGX agent

arXiv:2606.08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation

safetyarxiv-cs-lg
9 Jun 2026
Safety

Beyond Neural Collapse: Task-Intrinsic Geometry Governs Neural Representations in Modular Arithmetic

DGX agent

arXiv:2606.08985v1 Announce Type: new Abstract: While neural collapse (NC) predicts that a K-class-balanced classifier should organize terminal representations as a (K-1)-dimensional simplex equiangul

safetyarxiv-cs-lg
9 Jun 2026
Safety

Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

DGX agent

arXiv:2606.09076v1 Announce Type: new Abstract: Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric score

safetyarxiv-cs-cv
9 Jun 2026
Safety

Boundary Variance Inflation Causes Acquisition Bias in Gaussian Processes

DGX agent

arXiv:2606.07561v1 Announce Type: new Abstract: Gaussian processes with stationary kernels on bounded domains exhibit inflated posterior variance near the boundary. Despite being a long-recognized art

safetyarxiv-cs-lg
9 Jun 2026
Safety

Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents

DGX agent

arXiv:2606.09315v1 Announce Type: cross Abstract: BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agents, exposing a new attack surface we call brain-pro

safetyarxiv-cs-ai
9 Jun 2026
Safety

Breaking the Tokenizer Barrier: On-Policy Distillation across Model Families

DGX agent

arXiv:2606.09456v1 Announce Type: new Abstract: On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain exp

safetyarxiv-cs-lg
9 Jun 2026
Safety

Bridging Expert Knowledge and Automated Feature Engineering via Self-Evolution

DGX agent

arXiv:2606.08800v1 Announce Type: new Abstract: In high-stakes settings such as brand compliance, clinical care, and content moderation, machine learning cannot be deployed as opaque oracles: practiti

safetyarxiv-cs-ai
9 Jun 2026
Safety

Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis

DGX agent

arXiv:2606.07533v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) effectively integrate text and audio to interpret context in complex interactive dialogues. However, the inte

safetyarxiv-cs-ai
9 Jun 2026
Safety

Brilliant idea! Next up: Apple randomly reboots your Mac if you're building competing tech, Gmail silently edits your email if you mention r…

DGX agent

Brilliant idea! Next up: Apple randomly reboots your Mac if you're building competing tech, Gmail silently edits your email if you mention rival platforms, and Tesla Autopilot swerves if it detects yo

safetyjeremy-howard--x
9 Jun 2026
Safety

Can Data Work be Reparative?

DGX agent

arXiv:2606.09408v1 Announce Type: cross Abstract: We present an ethnographic study of an alternative approach to data work, developed by a civic-tech initiative that builds datasets for training and b

safetyarxiv-cs-ai
9 Jun 2026
Safety

Can the Environment Speak for Itself? T^{2}-GRPO: A Turn-Trajectory Group Relative Policy Optimization for Caregiver Agents

DGX agent

arXiv:2606.08875v1 Announce Type: new Abstract: Optimizing large language models (LLMs) for long-horizon caregiver agents requires balancing delayed task objectives with immediate environment dynamics

safetyarxiv-cs-ai
9 Jun 2026
Safety

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs

DGX agent

arXiv:2606.09371v1 Announce Type: new Abstract: Tool learning enables LLMs to invoke external tools to accomplish tasks. Prior studies have demonstrated the effectiveness of a hierarchical structure:

safetyarxiv-cs-ai
9 Jun 2026
Safety

CARE: A Conformal Safety Layer for Medical Summarization

DGX agent

arXiv:2606.08969v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical summarization, but their outputs can omit medically important information and introduce

safetyarxiv-cs-ai
9 Jun 2026
Safety

Causal Semantic Alignment for LLM-based Time Series Forecasting

DGX agent

arXiv:2606.08262v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have opened new possibilities for time series forecasting by enabling alignment between temporal pattern

safetyarxiv-cs-lg
9 Jun 2026
Safety

Causal Transfer in Medical Image Analysis

DGX agent

arXiv:2603.24388v2 Announce Type: replace Abstract: Medical imaging models frequently fail when deployed across hospitals, scanners, populations, or imaging protocols due to domain shift, limiting the

safetyarxiv-cs-cv
9 Jun 2026
Safety

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

DGX agent

arXiv:2606.08420v1 Announce Type: new Abstract: Vision-language models (VLMs) pretrained on large-scale image-text pairs demonstrate strong image-level understanding, but are primarily optimized for g

safetyarxiv-cs-cv
9 Jun 2026
← Previous
1…9899100101102…267
Next →