AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
3 Jun 2026

When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation

Model ReleasesDGX agent

arXiv:2606.03532v1 Announce Type: cross Abstract: Self on-policy distillation trains a student policy against a teacher derived from its own parameter history, yet the teacher's update schedule -- whi

2 Jun 2026

Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities

Model ReleasesDGX agent

arXiv:2505.24621v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have transformed natural language understanding and generation, leading to extensive benchmarkin

ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2606.01494v1 Announce Type: cross Abstract: Agent skills extend AI agents with reusable instructions, tools, scripts, references, and workflows, establishing a security boundary distinct from bo

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

Model ReleasesDGX agent

arXiv:2606.01850v1 Announce Type: new Abstract: Model compression techniques such as quantization and pruning are widely used to reduce the deployment cost of large language models (LLMs), with existi

Dynamic Proxy-Mixing: Transferring Replay Controllers from Small to Large Models for Continual Instruction Tuning

Model ReleasesDGX agent

arXiv:2606.00400v1 Announce Type: new Abstract: Continual instruction tuning updates a language model through a sequence of new domains, yet each update can progressively erode previously learned capa

From Segments to Scenes: Temporal Understanding in Autonomous Driving via Vision-Language Model

Model ReleasesDGX agent

arXiv:2512.05277v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning

Model ReleasesDGX agent

arXiv:2509.12263v3 Announce Type: replace Abstract: Large multimodal models (LMMs) encode physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs

Investigating and Alleviating Harm Amplification in LLM Interactions

Model ReleasesDGX agent

arXiv:2606.02423v1 Announce Type: new Abstract: Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve ha

Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture

Model ReleasesDGX agent

arXiv:2606.00288v1 Announce Type: new Abstract: Large language models are undergoing a transition from model technology to system technology. As developers use Codex, Claude Code, AutoGPT, and related

Product-Aware Deep Autoencoders for Robust Process Monitoring in Multi-Product Cyber-Physical Systems

Model ReleasesDGX agent

arXiv:2606.00052v1 Announce Type: new Abstract: As Industry 4.0 accelerates the integration of Cyber-Physical Systems (CPS) in manufacturing, robust anomaly detection has become critical for ensuring

SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents

Model ReleasesDGX agent

arXiv:2606.02302v1 Announce Type: cross Abstract: Autonomous LLM agents increasingly operate in stateful environments where they access tools, files, memory, and external services. While such capabili

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence

Model ReleasesDGX agent

arXiv:2606.02380v1 Announce Type: cross Abstract: As LLM-based agents expand their operational scope, reliability becomes a prerequisite for real-world deployment. However, in practical applications,

TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models

Model ReleasesDGX agent

arXiv:2606.00023v1 Announce Type: cross Abstract: The rapid development of Language Diffusion Models (LDMs) challenges the dominant position of auto-regressive competitors in language processing. Howe

Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning

Model ReleasesDGX agent

arXiv:2606.00105v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on vision-language tasks, but they may also memorize and expose sensitive o

1 Jun 2026

AbstainGNN: Teaching Graph Neural Networks to Abstain for Graph Classification

Model ReleasesDGX agent

arXiv:2605.30786v1 Announce Type: new Abstract: Graph classification is a core task in graph data mining with widespread real-world applications. Recent advances in graph neural networks (GNNs) have l

BOKBO (Best of K Bad Options): Calibrated Abstention for VLA Policies

Model ReleasesDGX agent

arXiv:2605.30660v1 Announce Type: new Abstract: Test-time scaling for vision-language-action (VLA) policies, methods such as RoboMonkey, SEAL, MG-Select, and V-GPS, samples K candidate action chunks a

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Model ReleasesDGX agent

arXiv:2503.08679v5 Announce Type: replace Abstract: Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) o

EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.30924v1 Announce Type: new Abstract: MLLM-powered embodied agents deployed in real-world environments encounter physical hazards. However, existing approaches lack explicit mechanisms for i

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

Model ReleasesDGX agent

arXiv:2605.30654v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but t

Mellum2 Technical Report

Model ReleasesDGX agent

arXiv:2605.31268v1 Announce Type: new Abstract: We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-p

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs

Model ReleasesDGX agent

arXiv:2605.30646v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in clinical applications. However, their behavior remains highly sensitive to subtle linguistic var

TAGA: A Tangent-Based Reactive Approach for Socially Compliant Robot Navigation Around Human Groups

Model ReleasesDGX agent

arXiv:2503.21168v3 Announce Type: replace Abstract: Robots navigating human-populated environments must avoid collisions while respecting the social structure of crowds, particularly the implicit boun

Target-Agnostic Calibration under Distribution Shift with Frequency-Aware Gradient Rectification

Model ReleasesDGX agent

arXiv:2508.19830v2 Announce Type: replace-cross Abstract: Real-world model deployments inevitably encounter distribution shifts, rendering the confidence estimates of deep neural networks highly unrel

The state of AI right now

IndustryDGX agent

Elon Musk shared observations about the current state of artificial intelligence development and capabilities on X (formerly Twitter). The post likely discusses recent AI advances, challenges, or Musk

When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception

Model ReleasesDGX agent

arXiv:2605.30381v1 Announce Type: cross Abstract: Deceptive alignment, in which models maintain accurate internal representations while deliberately producing false outputs, remains a central challeng

29 May 2026

A shared playbook for trustworthy third party evaluations

TutorialsDGX agent

This document outlines OpenAI's framework and recommendations for conducting independent third-party evaluations of AI systems to ensure trustworthiness and accountability. It establishes shared stand

Auto-review mode is now available in Cursor. It allows agents to run tool calls with fewer approval prompts and safer execution.

ToolsDGX agent

Cursor has introduced an auto-review mode feature that enables AI agents to execute tool calls with reduced approval requirements while maintaining safer execution practices. This feature streamlines

ChatGPT diagnosed 40 million people with a disease that was invented as a joke. Not a real disease. Not a misunderstood disease. A completel…

Model ReleasesDGX agent

ChatGPT diagnosed 40 million people with a disease that was invented as a joke. Not a real disease. Not a misunderstood disease. A completely fictional condition with a fake name, fake papers, and fak

DOJ sues states that rejected ICE requests for undercover license plates

IndustryDGX agent

The Department of Justice filed lawsuits against Maine, Massachusetts, Oregon, and Washington state alleging their refusal to issue undercover license plates to federal agents imposes unconstitutional

Feature Geometry of LoRA Adapters: A Sparse Autoencoder Analysis of Representational Divergence in Fine-Tuned Language Models

Model ReleasesDGX agent

arXiv:2605.28896v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has emerged as a widely adopted approach for adapting large language models, yet the internal representational changes induce

FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions

Model ReleasesDGX agent

arXiv:2605.29427v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in financial services, a single non-compliant interaction can expose institutions to regulator

GPIC: A Giant Permissive Image Corpus for Visual Generation

Model ReleasesDGX agent

arXiv:2605.30341v1 Announce Type: cross Abstract: Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

Model ReleasesDGX agent

arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials

Latent Performance Profiling of Large Language Models

Model ReleasesDGX agent

arXiv:2605.30018v1 Announce Type: new Abstract: Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabili

MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models

Model ReleasesDGX agent

arXiv:2605.28825v1 Announce Type: new Abstract: Large language models (LLMs) frequently encode factual and reasoning knowledge in their internal representations that is not faithfully reflected in the

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

Model ReleasesDGX agent

arXiv:2605.29685v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social

Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs

Model ReleasesDGX agent

arXiv:2605.30021v1 Announce Type: new Abstract: Many open-ended instructions have multiple valid answers that users can benefit from seeing, but post-training often narrows an LLM's output space towar

self-modification

AgentsDGX agent

Self-modification in AI systems refers to the capability of an artificial intelligence to alter its own code, parameters, or behavior patterns without external intervention. This concept, discussed by

Steering at the Source: Style Modulation Heads for Robust Persona Control

Local AiDGX agent

arXiv:2603.13249v2 Announce Type: replace-cross Abstract: Activation steering offers a computationally efficient mechanism for controlling Large Language Models (LLMs) without fine-tuning. While effec

SURGENT: A Surgical Multi-Agent Assistance System Across the Perioperative Workflow

Model ReleasesDGX agent

arXiv:2605.29368v1 Announce Type: cross Abstract: The intricate nature of modern surgical care necessitates intelligent systems that can synthesize extensive patient records, support collaborative dec

TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech

Model ReleasesDGX agent

arXiv:2601.11178v2 Announce Type: replace Abstract: Social media platforms are increasingly dominated by long-form multimodal content, where harmful narratives are constructed through a complex interp

28 May 2026

AI in SRE: Where and how Google is deploying agentic AI to improve operations

Model ReleasesDGX agent

Since its inception over 20 years ago, Google has used Site Reliability Engineering (SRE) to keep services like Search, Gmail, Maps, YouTube and Google Cloud reliable and highly available, adhering to

Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2512.00349v2 Announce Type: replace Abstract: Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their per

Localizing Input Uncertainty Quantification for Large Language Models via Shapley Values

Local AiDGX agent

arXiv:2605.28170v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly integrated into high-stakes decision-making, the ability to reliably quantify uncertainty has become a

Multi-Adapter Representation Interventions via Energy Calibration

Model ReleasesDGX agent

arXiv:2605.28722v1 Announce Type: new Abstract: Representation intervention has emerged as a promising paradigm for aligning large language models toward desired behaviors without modifying model weig

PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI

Model ReleasesDGX agent

arXiv:2605.27545v1 Announce Type: new Abstract: Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text

Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models

Model ReleasesDGX agent

arXiv:2503.01829v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) demonstrate persuasive capabilities that rival human-level persuasion. While these capabilities can be used for s

27 May 2026

Edge AI Deployment Beyond Models: A BSP-Aware Systems Framework for Industrial Embedded Platforms

Local AiDGX agent

arXiv:2605.26119v1 Announce Type: cross Abstract: Industrial Edge AI programs often begin with the model and only later confront the platform. That sequencing is attractive because it allows early dem

On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs

Model ReleasesDGX agent

arXiv:2510.05864v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly operate on long inputs, yet their behavior when harmful sentences are sparsely embedded within such inputs

Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study

Local AiDGX agent

arXiv:2605.26870v1 Announce Type: cross Abstract: Background: Large language models are typically evaluated as models, benchmarks, or short conversational episodes. Less is known about what happens wh

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

Model ReleasesDGX agent

arXiv:2605.26548v1 Announce Type: cross Abstract: Large language models (LLMs) now support automated software security tasks, including vulnerability discovery and proof-of-concept (PoC) generation. E

Sentinel: Embodied Cooperative Spatial Reasoning and Planning

Model ReleasesDGX agent

arXiv:2605.26239v1 Announce Type: new Abstract: In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmen

What Molecular Structure Cannot Tell Us: A Taxonomy of Explainability Gaps in GNN-Based Drug Toxicity Prediction

Model ReleasesDGX agent

arXiv:2605.26183v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have emerged as a structurally natural approach for molecular toxicity prediction, operating directly on atomic connectiv

26 May 2026

AI Content Moderation in Therapy Conversations

Model ReleasesDGX agent

arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL

Benchmarking and Learning Real-World Customer Service Dialogue

Model ReleasesDGX agent

arXiv:2510.22143v3 Announce Type: replace Abstract: Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction

Model ReleasesDGX agent

arXiv:2605.24686v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly integrated into emotionally sensitive domains, the structural integrity of their emotional intelligence

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

Model ReleasesDGX agent

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can

PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction

Model ReleasesDGX agent

arXiv:2605.24562v1 Announce Type: cross Abstract: Pedestrian intention and trajectory prediction are critical for the safe deployment of autonomous driving systems, directly influencing navigation dec

Reward-free Alignment for Conflicting Objectives

Model ReleasesDGX agent

arXiv:2602.02495v3 Announce Type: replace-cross Abstract: Direct alignment methods are increasingly used to align large language models (LLMs) with human preferences. However, many real-world alignmen

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions

Model ReleasesDGX agent

arXiv:2605.25073v1 Announce Type: cross Abstract: Background: Fine-tuning is central to adapting pre-trained Large Language Models (LLMs) to downstream tasks, but its reliance on training data, parame

← Previous
1…230231232233234…238
Next →