AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation

DGX agent

arXiv:2605.26918v1 Announce Type: new Abstract: Video generation models (VGMs) are rapidly entering classrooms, yet existing benchmarks evaluate only perceptual quality, intrinsic faithfulness, generi

model-releasesarxiv-cs-cl
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Chain-of-Thought Hijacking

DGX agent

arXiv:2510.26418v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reas

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation

DGX agent

arXiv:2605.25984v1 Announce Type: cross Abstract: Ensuring safe and contextually appropriate behaviour in Large Language Models (LLMs) remains a critical challenge for real-world deployment. We presen

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Learning Safely Without Knowing the World:COMPASS-Hedge

DGX agent

arXiv:2603.22348v3 Announce Type: replace Abstract: Online learning algorithms often face a fundamental trilemma: balancing regret guarantees between adversarial and stochastic settings and providing

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

DGX agent

arXiv:2605.23157v1 Announce Type: new Abstract: The attack surface of a multimodal large language model (MLLM) is language-dependent in ways that reveal the mechanistic structure of alignment failures

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems

DGX agent

arXiv:2605.22001v1 Announce Type: cross Abstract: Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads that announce themselves as override directives.

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation

DGX agent

arXiv:2605.20469v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used for medical image interpretation, yet they frequently hallucinate, generating clinically plausible b

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining

DGX agent

arXiv:2605.20296v1 Announce Type: new Abstract: Fine-tuning a language model for a target task routinely degrades capabilities the training data never explicitly threatened. We study this phenomenon,

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

DGX agent

arXiv:2605.18915v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are vulnerable to jailbreak attacks, which can elicit harmful responses from MLLMs. Many MLLMs support multi-

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Alignment Dynamics in LLM Fine-Tuning

DGX agent

arXiv:2605.18309v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alig

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

On-Device Interpretable Tsetlin Machine-Based Intrusion Detection for Secure IoMT

DGX agent

arXiv:2605.16707v1 Announce Type: cross Abstract: The rapid evolution of digital health technologies is redefining healthcare services worldwide. The integration of wireless communication and Internet

local-aiarxiv-cs-lg
19 May 2026
Model Releases

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts

DGX agent

arXiv:2510.07239v2 Announce Type: replace Abstract: Automated red-teaming has emerged as a scalable approach for auditing Large Language Models (LLMs) prior to deployment, yet existing approaches lack

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute

DGX agent

arXiv:2605.15377v1 Announce Type: new Abstract: As AI systems are increasingly deployed in autonomous agentic settings at scale, it is important to ensure the actions they take are safe and aligned wi

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering

DGX agent

arXiv:2506.08584v4 Announce Type: replace Abstract: Medical question answering (QA) benchmarks often focus on multiple-choice or fact-based tasks, leaving open-ended answers to real patient questions

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

DGX agent

arXiv:2605.13542v1 Announce Type: new Abstract: Intensive care units (ICU) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states u

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation

DGX agent

arXiv:2605.11533v1 Announce Type: new Abstract: Clinical check-up reports are multimodal documents that combine page layouts, tables, numerical biomarkers, abnormality flags, imaging findings, and dom

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Fine-Tuning Large Language Models for Cooperative Tactical Deconfliction of Small Unmanned Aerial Systems

DGX agent

arXiv:2603.28561v2 Announce Type: replace Abstract: The growing deployment of small Unmanned Aerial Systems (sUASs) in low-altitude airspaces has increased the need for reliable tactical deconfliction

model-releasesarxiv-cs-ro
13 May 2026
Model Releases

AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

DGX agent

arXiv:2601.01762v2 Announce Type: replace-cross Abstract: Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. W

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

DGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification

DGX agent

arXiv:2605.03476v1 Announce Type: new Abstract: Discharge summaries require extracting critical information from lengthy electronic health records (EHRs), a process that is labor-intensive when perfor

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure

DGX agent

arXiv:2605.02398v1 Announce Type: cross Abstract: As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability -- knowing what they do not kn

model-releasesarxiv-cs-cl
5 May 2026
Local Ai

Cloud Is Closer Than It Appears: Revisiting the Tradeoffs of Distributed Real-Time Inference

DGX agent

arXiv:2605.00005v1 Announce Type: new Abstract: The increasing deployment of deep neural networks (DNNs) in cyber-physical systems (CPS) enhances perception fidelity, but imposes substantial computati

local-aiarxiv-cs-lg
4 May 2026
Model Releases

MAEO: Multiobjective Animorphic Ensemble Optimization for Scalable Large-scale Engineering Applications

DGX agent

arXiv:2604.26973v1 Announce Type: cross Abstract: Multiobjective optimization remains challenging for many scientific and engineering problems due to the need to balance convergence, diversity, and co

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

DGX agent

arXiv:2508.04325v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However,

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation

DGX agent

arXiv:2604.24086v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have been demonstrated possessing strong zero-shot generalization for robot control, their massive parameter

model-releasesarxiv-cs-ai
28 Apr 2026
Local Ai

Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features

DGX agent

arXiv:2604.23829v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) extract millions of interpretable features from a language model, but flat feature inventories aren't very useful on their ow

local-aiarxiv-cs-ai
28 Apr 2026
Model Releases

Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models

DGX agent

arXiv:2604.24542v1 Announce Type: cross Abstract: Large language models deployed at runtime can misbehave in ways that clean-data validation cannot anticipate: training-time backdoors lie dormant unti

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models

DGX agent

arXiv:2604.23460v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has emerged as a key technique for eliciting complex reasoning in Large Language Models (LLMs). Although interpretable,

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems

DGX agent

arXiv:2604.22136v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly issue API calls that mutate real systems, yet many current architectures pass stochastic model outputs

model-releasesarxiv-cs-lg
27 Apr 2026
Model Releases

TreeCoder: Systematic Exploration and Optimisation of Decoding and Constraints for LLM Code Generation

DGX agent

arXiv:2511.22277v2 Announce Type: replace Abstract: Large language models (LLMs) have shown remarkable ability to generate code, yet their outputs often violate syntactic or semantic constraints when

model-releasesarxiv-cs-lg
27 Apr 2026
Model Releases

RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation

DGX agent

arXiv:2603.27112v2 Announce Type: replace Abstract: As Automatic Train Operation (ATO) advances toward GoA4 and beyond, it increasingly depends on efficient, reliable cab-view visual perception and de

model-releasesarxiv-cs-cv
24 Apr 2026
Model Releases

Mythos and the Unverified Cage: Z3-Based Pre-Deployment Verification for Frontier-Model Sandbox Infrastructure

DGX agent

arXiv:2604.20496v1 Announce Type: cross Abstract: The April 2026 Claude Mythos sandbox escape exposed a critical weakness in frontier AI containment: the infrastructure surrounding advanced models rem

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models

DGX agent

arXiv:2509.15174v3 Announce Type: replace-cross Abstract: WARNING: This paper contains examples of offensive materials. To address the proliferation of toxic content on social media, we introduce SMAR

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Spatio-temporal modelling of electric vehicle charging demand

DGX agent

arXiv:2604.19841v1 Announce Type: cross Abstract: Accurate forecasting of electric vehicle (EV) charging demand is critical for grid management and infrastructure planning. Yet the field continues to

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

Cross-Model Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Across Three Large Language Models

DGX agent

arXiv:2604.19598v1 Announce Type: cross Abstract: This study compared repeated generation consistency of exercise prescription outputs across three large language models (LLMs), specifically GPT-4.1,

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Deep sprite-based image models: An analysis

DGX agent

arXiv:2604.19480v1 Announce Type: new Abstract: While foundation models drive steady progress in image segmentation and diffusion algorithms compose always more realistic images, the seemingly simple

model-releasesarxiv-cs-cv
22 Apr 2026
Local Ai

Detoxification for LLM: From Dataset Itself

DGX agent

arXiv:2604.19124v1 Announce Type: new Abstract: Existing detoxification methods for large language models mainly focus on post-training stage or inference time, while few tackle the source of toxicity

local-aiarxiv-cs-cl
22 Apr 2026
Model Releases

HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing

DGX agent

arXiv:2604.19274v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as co-authors in collaborative writing, where users begin with rough drafts and rely on LLMs to compl

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text

DGX agent

arXiv:2604.19298v1 Announce Type: cross Abstract: We introduce IndiaFinBench, to our knowledge the first publicly available evaluation benchmark for assessing large language model (LLM) performance on

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators

DGX agent

arXiv:2512.09427v5 Announce Type: replace-cross Abstract: Existing memory management techniques severely hinder efficient Large Language Model serving on accelerators constrained by poor random-access

model-releasesarxiv-cs-ai
22 Apr 2026
Local Ai

Bridging the Culture Gap: A Framework for LLM-Driven Socio-Cultural Localization of Math Word Problems in Low-Resource Languages

DGX agent

arXiv:2508.14913v4 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant capabilities in solving mathematical problems expressed in natural language. However, mul

local-aiarxiv-cs-cl
21 Apr 2026
Model Releases

Camo-M3FD: A New Benchmark Dataset for Cross-Spectral Camouflaged Pedestrian Detection

DGX agent

arXiv:2604.16582v1 Announce Type: new Abstract: Pedestrian detection is fundamental to autonomous driving, robotics, and surveillance. Despite progress in deep learning, reliable identification remain

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Chain Of Interaction Benchmark (COIN): When Reasoning meets Embodied Interaction

DGX agent

arXiv:2604.16886v1 Announce Type: new Abstract: Generalist embodied agents must perform interactive, causally-dependent reasoning, continually interacting with the environment, acquiring information,

model-releasesarxiv-cs-ro
21 Apr 2026
Model Releases

Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs

DGX agent

arXiv:2510.11288v4 Announce Type: replace Abstract: Recent work has shown that narrow finetuning can produce broadly misaligned LLMs, a phenomenon termed emergent misalignment (EM). While concerning,

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models

DGX agent

arXiv:2604.16405v1 Announce Type: cross Abstract: Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physi

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

IncreFA: Breaking the Static Wall of Generative Model Attribution

DGX agent

arXiv:2604.17736v1 Announce Type: new Abstract: As AI generative models evolve at unprecedented speed, image attribution has become a moving target. New diffusion, adversarial and autoregressive gener

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

DGX agent

arXiv:2604.17282v1 Announce Type: new Abstract: Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general

model-releasesarxiv-cs-cl
21 Apr 2026
Local Ai

Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning

DGX agent

arXiv:2510.16054v2 Announce Type: replace-cross Abstract: When users submit queries to Large Language Models (LLMs), their prompts can often contain sensitive data, forcing a difficult choice: Send th

local-aiarxiv-cs-cl
21 Apr 2026
← Previous
1…240241242243244…255
Next →