AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
28 May 2026

LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks

Model ReleasesDGX agent

arXiv:2605.27375v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context

The Alignment Floor: When Persona Customization Is Safe

Model ReleasesDGX agent

arXiv:2605.27382v1 Announce Type: cross Abstract: A key promise of pluralistic AI is behavioral adaptation: persona prompts like 'be creative' or 'be thorough' let systems respect diverse user values

Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models

Local AiDGX agent

arXiv:2605.27997v1 Announce Type: cross Abstract: Large language models frequently generate toxic, hateful, or harmful content, yet existing mitigation methods rely on costly retraining or output-leve

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
27 May 2026

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation

Model ReleasesDGX agent

arXiv:2605.26918v1 Announce Type: new Abstract: Video generation models (VGMs) are rapidly entering classrooms, yet existing benchmarks evaluate only perceptual quality, intrinsic faithfulness, generi

26 May 2026

Chain-of-Thought Hijacking

Model ReleasesDGX agent

arXiv:2510.26418v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reas

SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation

Model ReleasesDGX agent

arXiv:2605.25984v1 Announce Type: cross Abstract: Ensuring safe and contextually appropriate behaviour in Large Language Models (LLMs) remains a critical challenge for real-world deployment. We presen

25 May 2026

Learning Safely Without Knowing the World:COMPASS-Hedge

Model ReleasesDGX agent

arXiv:2603.22348v3 Announce Type: replace Abstract: Online learning algorithms often face a fundamental trilemma: balancing regret guarantees between adversarial and stochastic settings and providing

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

Model ReleasesDGX agent

arXiv:2605.23157v1 Announce Type: new Abstract: The attack surface of a multimodal large language model (MLLM) is language-dependent in ways that reveal the mechanistic structure of alignment failures

22 May 2026

Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.22001v1 Announce Type: cross Abstract: Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads that announce themselves as override directives.

21 May 2026

HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation

Model ReleasesDGX agent

arXiv:2605.20469v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used for medical image interpretation, yet they frequently hallucinate, generating clinically plausible b

Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining

Model ReleasesDGX agent

arXiv:2605.20296v1 Announce Type: new Abstract: Fine-tuning a language model for a target task routinely degrades capabilities the training data never explicitly threatened. We study this phenomenon,

20 May 2026

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

Model ReleasesDGX agent

arXiv:2605.18915v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are vulnerable to jailbreak attacks, which can elicit harmful responses from MLLMs. Many MLLMs support multi-

19 May 2026

Alignment Dynamics in LLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.18309v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alig

On-Device Interpretable Tsetlin Machine-Based Intrusion Detection for Secure IoMT

Local AiDGX agent

arXiv:2605.16707v1 Announce Type: cross Abstract: The rapid evolution of digital health technologies is redefining healthcare services worldwide. The integration of wireless communication and Internet

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts

Model ReleasesDGX agent

arXiv:2510.07239v2 Announce Type: replace Abstract: Automated red-teaming has emerged as a scalable approach for auditing Large Language Models (LLMs) prior to deployment, yet existing approaches lack

The agentic era: Architecting the blueprint for mission impact across the public sector

Model ReleasesDGX agent

This is a new era — the agentic era – and the question is no longer, “what’s possible?” but rather, “what creates impact?” Today, organizations across industries around the world are swiftly moving fr

18 May 2026

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute

Model ReleasesDGX agent

arXiv:2605.15377v1 Announce Type: new Abstract: As AI systems are increasingly deployed in autonomous agentic settings at scale, it is important to ensure the actions they take are safe and aligned wi

To protect passengers or cargo, the powered rear seats & trunk in Model Y will automatically pop back up if detecting an obstruction while f…

IndustryDGX agent

Tesla Model Y's powered rear seats and trunk are equipped with automatic obstruction detection that causes them to automatically reverse and pop back up if an obstruction is detected during operation,

15 May 2026

CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering

Model ReleasesDGX agent

arXiv:2506.08584v4 Announce Type: replace Abstract: Medical question answering (QA) benchmarks often focus on multiple-choice or fact-based tasks, leaving open-ended answers to real patient questions

14 May 2026

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

Model ReleasesDGX agent

arXiv:2605.13542v1 Announce Type: new Abstract: Intensive care units (ICU) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states u

13 May 2026

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

Model ReleasesDGX agent

arXiv:2605.11533v1 Announce Type: new Abstract: Clinical check-up reports are multimodal documents that combine page layouts, tables, numerical biomarkers, abnormality flags, imaging findings, and dom

Fine-Tuning Large Language Models for Cooperative Tactical Deconfliction of Small Unmanned Aerial Systems

Model ReleasesDGX agent

arXiv:2603.28561v2 Announce Type: replace Abstract: The growing deployment of small Unmanned Aerial Systems (sUASs) in low-altitude airspaces has increased the need for reliable tactical deconfliction

12 May 2026

AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2601.01762v2 Announce Type: replace-cross Abstract: Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. W

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

Model ReleasesDGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

9 May 2026

The new Wild West of AI kids’ toys

IndustryDGX agent

AI-powered toys are rapidly emerging in the consumer market, including pocket pets, autonomous robots, and AI emotional companions that use advanced language models and adaptive learning capabilities.

6 May 2026

CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification

Model ReleasesDGX agent

arXiv:2605.03476v1 Announce Type: new Abstract: Discharge summaries require extracting critical information from lengthy electronic health records (EHRs), a process that is labor-intensive when perfor

5 May 2026

The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure

Model ReleasesDGX agent

arXiv:2605.02398v1 Announce Type: cross Abstract: As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability -- knowing what they do not kn

4 May 2026

Cloud Is Closer Than It Appears: Revisiting the Tradeoffs of Distributed Real-Time Inference

Local AiDGX agent

arXiv:2605.00005v1 Announce Type: new Abstract: The increasing deployment of deep neural networks (DNNs) in cyber-physical systems (CPS) enhances perception fidelity, but imposes substantial computati

3 May 2026

Wiki Lint Report — 2026-05-03

SynthesesDGX agent

Automated lint: 45 errors, 11 warnings, 3 info

1 May 2026

MAEO: Multiobjective Animorphic Ensemble Optimization for Scalable Large-scale Engineering Applications

Model ReleasesDGX agent

arXiv:2604.26973v1 Announce Type: cross Abstract: Multiobjective optimization remains challenging for many scientific and engineering problems due to the need to balance convergence, diversity, and co

30 Apr 2026

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Model ReleasesDGX agent

arXiv:2508.04325v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However,

28 Apr 2026

50+ fully managed MCP servers now available for Google Cloud services

Model ReleasesDGX agent

At Google Cloud Next ‘26, we announced that more than 50 Google-managed Model Context Protocol (MCP) servers are generally available or in preview, with more on the way. Why it matters: To move beyond

AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation

Model ReleasesDGX agent

arXiv:2604.24086v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have been demonstrated possessing strong zero-shot generalization for robot control, their massive parameter

Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features

Local AiDGX agent

arXiv:2604.23829v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) extract millions of interpretable features from a language model, but flat feature inventories aren't very useful on their ow

Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models

Model ReleasesDGX agent

arXiv:2604.24542v1 Announce Type: cross Abstract: Large language models deployed at runtime can misbehave in ways that clean-data validation cannot anticipate: training-time backdoors lie dormant unti

Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models

Model ReleasesDGX agent

arXiv:2604.23460v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has emerged as a key technique for eliciting complex reasoning in Large Language Models (LLMs). Although interpretable,

27 Apr 2026

Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems

Model ReleasesDGX agent

arXiv:2604.22136v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly issue API calls that mutate real systems, yet many current architectures pass stochastic model outputs

TreeCoder: Systematic Exploration and Optimisation of Decoding and Constraints for LLM Code Generation

Model ReleasesDGX agent

arXiv:2511.22277v2 Announce Type: replace Abstract: Large language models (LLMs) have shown remarkable ability to generate code, yet their outputs often violate syntactic or semantic constraints when

26 Apr 2026

Wiki Lint Report — 2026-04-26

SynthesesDGX agent

Automated lint: 44 errors, 10 warnings, 3 info

24 Apr 2026

RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation

Model ReleasesDGX agent

arXiv:2603.27112v2 Announce Type: replace Abstract: As Automatic Train Operation (ATO) advances toward GoA4 and beyond, it increasingly depends on efficient, reliable cab-view visual perception and de

23 Apr 2026

Mythos and the Unverified Cage: Z3-Based Pre-Deployment Verification for Frontier-Model Sandbox Infrastructure

Model ReleasesDGX agent

arXiv:2604.20496v1 Announce Type: cross Abstract: The April 2026 Claude Mythos sandbox escape exposed a critical weakness in frontier AI containment: the infrastructure surrounding advanced models rem

SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models

Model ReleasesDGX agent

arXiv:2509.15174v3 Announce Type: replace-cross Abstract: WARNING: This paper contains examples of offensive materials. To address the proliferation of toxic content on social media, we introduce SMAR

Spatio-temporal modelling of electric vehicle charging demand

Model ReleasesDGX agent

arXiv:2604.19841v1 Announce Type: cross Abstract: Accurate forecasting of electric vehicle (EV) charging demand is critical for grid management and infrastructure planning. Yet the field continues to

22 Apr 2026

Cross-Model Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Across Three Large Language Models

Model ReleasesDGX agent

arXiv:2604.19598v1 Announce Type: cross Abstract: This study compared repeated generation consistency of exercise prescription outputs across three large language models (LLMs), specifically GPT-4.1,

Deep sprite-based image models: An analysis

Model ReleasesDGX agent

arXiv:2604.19480v1 Announce Type: new Abstract: While foundation models drive steady progress in image segmentation and diffusion algorithms compose always more realistic images, the seemingly simple

Detoxification for LLM: From Dataset Itself

Local AiDGX agent

arXiv:2604.19124v1 Announce Type: new Abstract: Existing detoxification methods for large language models mainly focus on post-training stage or inference time, while few tackle the source of toxicity

HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing

Model ReleasesDGX agent

arXiv:2604.19274v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as co-authors in collaborative writing, where users begin with rough drafts and rely on LLMs to compl

IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text

Model ReleasesDGX agent

arXiv:2604.19298v1 Announce Type: cross Abstract: We introduce IndiaFinBench, to our knowledge the first publicly available evaluation benchmark for assessing large language model (LLM) performance on

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators

Model ReleasesDGX agent

arXiv:2512.09427v5 Announce Type: replace-cross Abstract: Existing memory management techniques severely hinder efficient Large Language Model serving on accelerators constrained by poor random-access

21 Apr 2026

Bridging the Culture Gap: A Framework for LLM-Driven Socio-Cultural Localization of Math Word Problems in Low-Resource Languages

Local AiDGX agent

arXiv:2508.14913v4 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant capabilities in solving mathematical problems expressed in natural language. However, mul

Camo-M3FD: A New Benchmark Dataset for Cross-Spectral Camouflaged Pedestrian Detection

Model ReleasesDGX agent

arXiv:2604.16582v1 Announce Type: new Abstract: Pedestrian detection is fundamental to autonomous driving, robotics, and surveillance. Despite progress in deep learning, reliable identification remain

Chain Of Interaction Benchmark (COIN): When Reasoning meets Embodied Interaction

Model ReleasesDGX agent

arXiv:2604.16886v1 Announce Type: new Abstract: Generalist embodied agents must perform interactive, causally-dependent reasoning, continually interacting with the environment, acquiring information,

Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs

Model ReleasesDGX agent

arXiv:2510.11288v4 Announce Type: replace Abstract: Recent work has shown that narrow finetuning can produce broadly misaligned LLMs, a phenomenon termed emergent misalignment (EM). While concerning,

ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models

Model ReleasesDGX agent

arXiv:2604.16405v1 Announce Type: cross Abstract: Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physi

IncreFA: Breaking the Static Wall of Generative Model Attribution

Model ReleasesDGX agent

arXiv:2604.17736v1 Announce Type: new Abstract: As AI generative models evolve at unprecedented speed, image attribution has become a moving target. New diffusion, adversarial and autoregressive gener

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

Model ReleasesDGX agent

arXiv:2604.17282v1 Announce Type: new Abstract: Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general

Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning

Local AiDGX agent

arXiv:2510.16054v2 Announce Type: replace-cross Abstract: When users submit queries to Large Language Models (LLMs), their prompts can often contain sensitive data, forcing a difficult choice: Send th

Trustworthy Endoscopic Super-Resolution

Local AiDGX agent

arXiv:2604.18001v1 Announce Type: new Abstract: Super-resolution (SR) models are attracting growing interest for enhancing minimally invasive surgery and diagnostic videos under hardware constraints.

20 Apr 2026

Beyond Single-Model Optimization: Preserving Plasticity in Continual Reinforcement Learning

Local AiDGX agent

arXiv:2604.15414v1 Announce Type: cross Abstract: Continual reinforcement learning must balance retention with adaptation, yet many methods still rely on single-model preservation, committing to one e

Power to the Clients: Federated Learning in a Dictatorship Setting

Local AiDGX agent

arXiv:2510.22149v3 Announce Type: replace-cross Abstract: Federated learning (FL) has emerged as a promising paradigm for decentralized model training, enabling multiple clients to collaboratively lea

← Previous
1…223224225226227…238
Next →