AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Model Releases

LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks

DGX agent

arXiv:2605.27375v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context

model-releasesarxiv-cs-cl
28 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

The Alignment Floor: When Persona Customization Is Safe

DGX agent

arXiv:2605.27382v1 Announce Type: cross Abstract: A key promise of pluralistic AI is behavioral adaptation: persona prompts like 'be creative' or 'be thorough' let systems respect diverse user values

model-releasesarxiv-cs-ai
28 May 2026
Local Ai

Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models

DGX agent

arXiv:2605.27997v1 Announce Type: cross Abstract: Large language models frequently generate toxic, hateful, or harmful content, yet existing mitigation methods rely on costly retraining or output-leve

local-aiarxiv-cs-ai
28 May 2026
Model Releases

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation

DGX agent

arXiv:2605.26918v1 Announce Type: new Abstract: Video generation models (VGMs) are rapidly entering classrooms, yet existing benchmarks evaluate only perceptual quality, intrinsic faithfulness, generi

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Chain-of-Thought Hijacking

DGX agent

arXiv:2510.26418v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reas

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation

DGX agent

arXiv:2605.25984v1 Announce Type: cross Abstract: Ensuring safe and contextually appropriate behaviour in Large Language Models (LLMs) remains a critical challenge for real-world deployment. We presen

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Learning Safely Without Knowing the World:COMPASS-Hedge

DGX agent

arXiv:2603.22348v3 Announce Type: replace Abstract: Online learning algorithms often face a fundamental trilemma: balancing regret guarantees between adversarial and stochastic settings and providing

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

DGX agent

arXiv:2605.23157v1 Announce Type: new Abstract: The attack surface of a multimodal large language model (MLLM) is language-dependent in ways that reveal the mechanistic structure of alignment failures

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems

DGX agent

arXiv:2605.22001v1 Announce Type: cross Abstract: Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads that announce themselves as override directives.

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation

DGX agent

arXiv:2605.20469v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used for medical image interpretation, yet they frequently hallucinate, generating clinically plausible b

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining

DGX agent

arXiv:2605.20296v1 Announce Type: new Abstract: Fine-tuning a language model for a target task routinely degrades capabilities the training data never explicitly threatened. We study this phenomenon,

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

DGX agent

arXiv:2605.18915v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are vulnerable to jailbreak attacks, which can elicit harmful responses from MLLMs. Many MLLMs support multi-

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Alignment Dynamics in LLM Fine-Tuning

DGX agent

arXiv:2605.18309v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alig

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

On-Device Interpretable Tsetlin Machine-Based Intrusion Detection for Secure IoMT

DGX agent

arXiv:2605.16707v1 Announce Type: cross Abstract: The rapid evolution of digital health technologies is redefining healthcare services worldwide. The integration of wireless communication and Internet

local-aiarxiv-cs-lg
19 May 2026
Model Releases

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts

DGX agent

arXiv:2510.07239v2 Announce Type: replace Abstract: Automated red-teaming has emerged as a scalable approach for auditing Large Language Models (LLMs) prior to deployment, yet existing approaches lack

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

The agentic era: Architecting the blueprint for mission impact across the public sector

DGX agent

This is a new era — the agentic era – and the question is no longer, “what’s possible?” but rather, “what creates impact?” Today, organizations across industries around the world are swiftly moving fr

model-releasesgoogle-cloud-ai
19 May 2026
Model Releases

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute

DGX agent

arXiv:2605.15377v1 Announce Type: new Abstract: As AI systems are increasingly deployed in autonomous agentic settings at scale, it is important to ensure the actions they take are safe and aligned wi

model-releasesarxiv-cs-ai
18 May 2026
Industry

To protect passengers or cargo, the powered rear seats & trunk in Model Y will automatically pop back up if detecting an obstruction while f…

DGX agent

Tesla Model Y's powered rear seats and trunk are equipped with automatic obstruction detection that causes them to automatically reverse and pop back up if an obstruction is detected during operation,

industryelon-musk--x
18 May 2026
Model Releases

CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering

DGX agent

arXiv:2506.08584v4 Announce Type: replace Abstract: Medical question answering (QA) benchmarks often focus on multiple-choice or fact-based tasks, leaving open-ended answers to real patient questions

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

DGX agent

arXiv:2605.13542v1 Announce Type: new Abstract: Intensive care units (ICU) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states u

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

DGX agent

arXiv:2605.11533v1 Announce Type: new Abstract: Clinical check-up reports are multimodal documents that combine page layouts, tables, numerical biomarkers, abnormality flags, imaging findings, and dom

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Fine-Tuning Large Language Models for Cooperative Tactical Deconfliction of Small Unmanned Aerial Systems

DGX agent

arXiv:2603.28561v2 Announce Type: replace Abstract: The growing deployment of small Unmanned Aerial Systems (sUASs) in low-altitude airspaces has increased the need for reliable tactical deconfliction

model-releasesarxiv-cs-ro
13 May 2026
Model Releases

AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

DGX agent

arXiv:2601.01762v2 Announce Type: replace-cross Abstract: Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. W

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

DGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

model-releasesarxiv-cs-ai
12 May 2026
Industry

The new Wild West of AI kids’ toys

DGX agent

AI-powered toys are rapidly emerging in the consumer market, including pocket pets, autonomous robots, and AI emotional companions that use advanced language models and adaptive learning capabilities.

industryars-technica
9 May 2026
Model Releases

CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification

DGX agent

arXiv:2605.03476v1 Announce Type: new Abstract: Discharge summaries require extracting critical information from lengthy electronic health records (EHRs), a process that is labor-intensive when perfor

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure

DGX agent

arXiv:2605.02398v1 Announce Type: cross Abstract: As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability -- knowing what they do not kn

model-releasesarxiv-cs-cl
5 May 2026
Local Ai

Cloud Is Closer Than It Appears: Revisiting the Tradeoffs of Distributed Real-Time Inference

DGX agent

arXiv:2605.00005v1 Announce Type: new Abstract: The increasing deployment of deep neural networks (DNNs) in cyber-physical systems (CPS) enhances perception fidelity, but imposes substantial computati

local-aiarxiv-cs-lg
4 May 2026
Syntheses

Wiki Lint Report — 2026-05-03

DGX agent

Automated lint: 45 errors, 11 warnings, 3 info

linthealth-checkautomated
3 May 2026
Model Releases

MAEO: Multiobjective Animorphic Ensemble Optimization for Scalable Large-scale Engineering Applications

DGX agent

arXiv:2604.26973v1 Announce Type: cross Abstract: Multiobjective optimization remains challenging for many scientific and engineering problems due to the need to balance convergence, diversity, and co

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

DGX agent

arXiv:2508.04325v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However,

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

50+ fully managed MCP servers now available for Google Cloud services

DGX agent

At Google Cloud Next ‘26, we announced that more than 50 Google-managed Model Context Protocol (MCP) servers are generally available or in preview, with more on the way. Why it matters: To move beyond

model-releasesgoogle-cloud-ai
28 Apr 2026
Model Releases

AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation

DGX agent

arXiv:2604.24086v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have been demonstrated possessing strong zero-shot generalization for robot control, their massive parameter

model-releasesarxiv-cs-ai
28 Apr 2026
Local Ai

Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features

DGX agent

arXiv:2604.23829v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) extract millions of interpretable features from a language model, but flat feature inventories aren't very useful on their ow

local-aiarxiv-cs-ai
28 Apr 2026
Model Releases

Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models

DGX agent

arXiv:2604.24542v1 Announce Type: cross Abstract: Large language models deployed at runtime can misbehave in ways that clean-data validation cannot anticipate: training-time backdoors lie dormant unti

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models

DGX agent

arXiv:2604.23460v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has emerged as a key technique for eliciting complex reasoning in Large Language Models (LLMs). Although interpretable,

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems

DGX agent

arXiv:2604.22136v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly issue API calls that mutate real systems, yet many current architectures pass stochastic model outputs

model-releasesarxiv-cs-lg
27 Apr 2026
Model Releases

TreeCoder: Systematic Exploration and Optimisation of Decoding and Constraints for LLM Code Generation

DGX agent

arXiv:2511.22277v2 Announce Type: replace Abstract: Large language models (LLMs) have shown remarkable ability to generate code, yet their outputs often violate syntactic or semantic constraints when

model-releasesarxiv-cs-lg
27 Apr 2026
Syntheses

Wiki Lint Report — 2026-04-26

DGX agent

Automated lint: 44 errors, 10 warnings, 3 info

linthealth-checkautomated
26 Apr 2026
Model Releases

RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation

DGX agent

arXiv:2603.27112v2 Announce Type: replace Abstract: As Automatic Train Operation (ATO) advances toward GoA4 and beyond, it increasingly depends on efficient, reliable cab-view visual perception and de

model-releasesarxiv-cs-cv
24 Apr 2026
Model Releases

Mythos and the Unverified Cage: Z3-Based Pre-Deployment Verification for Frontier-Model Sandbox Infrastructure

DGX agent

arXiv:2604.20496v1 Announce Type: cross Abstract: The April 2026 Claude Mythos sandbox escape exposed a critical weakness in frontier AI containment: the infrastructure surrounding advanced models rem

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models

DGX agent

arXiv:2509.15174v3 Announce Type: replace-cross Abstract: WARNING: This paper contains examples of offensive materials. To address the proliferation of toxic content on social media, we introduce SMAR

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Spatio-temporal modelling of electric vehicle charging demand

DGX agent

arXiv:2604.19841v1 Announce Type: cross Abstract: Accurate forecasting of electric vehicle (EV) charging demand is critical for grid management and infrastructure planning. Yet the field continues to

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

Cross-Model Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Across Three Large Language Models

DGX agent

arXiv:2604.19598v1 Announce Type: cross Abstract: This study compared repeated generation consistency of exercise prescription outputs across three large language models (LLMs), specifically GPT-4.1,

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Deep sprite-based image models: An analysis

DGX agent

arXiv:2604.19480v1 Announce Type: new Abstract: While foundation models drive steady progress in image segmentation and diffusion algorithms compose always more realistic images, the seemingly simple

model-releasesarxiv-cs-cv
22 Apr 2026
Local Ai

Detoxification for LLM: From Dataset Itself

DGX agent

arXiv:2604.19124v1 Announce Type: new Abstract: Existing detoxification methods for large language models mainly focus on post-training stage or inference time, while few tackle the source of toxicity

local-aiarxiv-cs-cl
22 Apr 2026
Model Releases

HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing

DGX agent

arXiv:2604.19274v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as co-authors in collaborative writing, where users begin with rough drafts and rely on LLMs to compl

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text

DGX agent

arXiv:2604.19298v1 Announce Type: cross Abstract: We introduce IndiaFinBench, to our knowledge the first publicly available evaluation benchmark for assessing large language model (LLM) performance on

model-releasesarxiv-cs-ai
22 Apr 2026
← Previous
1…279280281282283…297
Next →