AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

AgentOpt v0.1 Technical Report: Client-Side Optimization for LLM-Based Agent

DGX agent

arXiv:2604.06296v1 Announce Type: cross Abstract: AI agents are increasingly deployed in real-world applications, including systems such as Manus, OpenClaw, and coding agents. Existing research has pr

model-releasesarxiv-cs-ai
10 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

CHiQPM: Calibrated Hierarchical Interpretable Image Classification

DGX agent

arXiv:2511.20779v2 Announce Type: replace Abstract: Globally interpretable models are a promising approach for trustworthy AI in safety-critical domains. Alongside global explanations, detailed local

local-aiarxiv-cs-lg
10 Apr 2026
Model Releases

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

DGX agent

arXiv:2604.08008v1 Announce Type: new Abstract: Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dat

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

DGX agent

arXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server

model-releasesarxiv-cs-lg
12 Aug 2026
Local Ai

Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

DGX agent

arXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar

local-aiarxiv-cs-cl
12 Aug 2026
Model Releases

Diffract: Spectral View of LLM Domain Adaptation

DGX agent

arXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

DGX agent

arXiv:2608.10300v1 Announce Type: new Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human re

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages

DGX agent

arXiv:2606.15821v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, fo

model-releasesarxiv-cs-ai
12 Aug 2026
Local Ai

Toward the Cognitive--Physical Limits of Embodied Intelligence through a World-Model-Centric Autonomous Racing Agent

DGX agent

arXiv:2608.10618v1 Announce Type: new Abstract: Embodied artificial intelligence aims to develop agents that perceive, reason, and act through continuous interaction with the physical world. However,

local-aiarxiv-cs-ro
12 Aug 2026
Model Releases

Automating Deception: Scalable Multi-Turn LLM Jailbreaks

DGX agent

arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

DGX agent

arXiv:2608.09900v1 Announce Type: new Abstract: Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably wa

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

DGX agent

arXiv:2608.07533v1 Announce Type: new Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodied agents pri

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design

DGX agent

arXiv:2608.07978v1 Announce Type: cross Abstract: Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their output,leaving

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition

DGX agent

arXiv:2608.09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predi

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

DGX agent

arXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Quantization Damage Is Multiplicative, Not Additive

DGX agent

arXiv:2608.06564v1 Announce Type: cross Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. What nobody can say is which of the model's

model-releasesarxiv-cs-cl
10 Aug 2026
Model Releases

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

DGX agent

arXiv:2607.18056v2 Announce Type: replace Abstract: Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace c

model-releasesarxiv-cs-cl
7 Aug 2026
Model Releases

Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation

DGX agent

arXiv:2608.05210v1 Announce Type: cross Abstract: Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by

model-releasesarxiv-cs-ai
7 Aug 2026
Local Ai

Unified Planning-Learning Framework for Robust UUV Navigation Under Partial Observability

DGX agent

arXiv:2608.05365v1 Announce Type: new Abstract: This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that in

local-aiarxiv-cs-ro
7 Aug 2026
Model Releases

Dynamic Jailbreaking Attack

DGX agent

arXiv:2510.02422v4 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suffix toward a predefined target response with a stat

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

K-EXAONE 2.0 Technical Report

DGX agent

arXiv:2608.04505v1 Announce Type: new Abstract: This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward glo

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

DGX agent

arXiv:2608.04514v1 Announce Type: new Abstract: Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-co

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Robust and Personalized Federated Learning for Aircraft-Engine Prognostics under Benign and Adversarial Client Heterogeneity

DGX agent

arXiv:2608.04045v1 Announce Type: cross Abstract: Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without shar

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control

DGX agent

arXiv:2608.04732v1 Announce Type: cross Abstract: Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

DGX agent

arXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and educatio

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests

DGX agent

arXiv:2608.02685v1 Announce Type: cross Abstract: Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixe

model-releasesarxiv-cs-ai
5 Aug 2026
Local Ai

Risk Occupancy: A New and Efficient Paradigm through Vehicle-Road-Cloud Collaboration

DGX agent

arXiv:2408.07367v3 Announce Type: replace Abstract: This paper proposes a novel 4D risk occupancy (RiskOcc) perception paradigm under the Vehicle-Road-Cloud integrated architecture, which unifies obje

local-aiarxiv-cs-ro
5 Aug 2026
Model Releases

TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology

DGX agent

arXiv:2608.03190v1 Announce Type: new Abstract: Neuro-oncology decisions require coordinated interpretation of serial MRI, pathology, molecular markers, treatment history, performance status, and evol

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

DGX agent

arXiv:2607.09842v2 Announce Type: replace-cross Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state tra

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Gaokerena: A Small Persian Medical Language Model Family

DGX agent

arXiv:2608.00932v1 Announce Type: new Abstract: The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

SAFE-Merge: Data-Free Continual Model Merging with General Knowledge Preservation

DGX agent

arXiv:2608.01184v1 Announce Type: new Abstract: Data-free continual model merging must incorporate a stream of specialized models while retaining both pretrained general knowledge and previously acqui

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

DGX agent

arXiv:2607.28661v1 Announce Type: new Abstract: Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding

model-releasesarxiv-cs-cl
3 Aug 2026
Model Releases

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

DGX agent

arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monit

model-releasesarxiv-cs-cl
3 Aug 2026
Model Releases

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

DGX agent

arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of do

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy

DGX agent

arXiv:2607.27782v1 Announce Type: new Abstract: Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused

model-releasesarxiv-cs-ro
31 Jul 2026
Model Releases

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

DGX agent

arXiv:2607.28128v1 Announce Type: new Abstract: LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Scaling medical imaging report generation with multimodal reinforcement learning

DGX agent

arXiv:2601.17151v2 Announce Type: replace-cross Abstract: Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit ma

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

GPT-Red: Automated Red Teaming via Self-Play at Scale

DGX agent

arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce extbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

DGX agent

arXiv:2607.26121v1 Announce Type: new Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures

model-releasesarxiv-cs-ro
30 Jul 2026
Model Releases

Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning

DGX agent

arXiv:2607.25754v1 Announce Type: new Abstract: Cooperative navigation of multi-agent UAVs in complex environments faces key challenges including local optima traps, sparse rewards, learning imbalance

model-releasesarxiv-cs-ro
29 Jul 2026
Model Releases

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

DGX agent

arXiv:2607.22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuni

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs

DGX agent

arXiv:2607.22555v1 Announce Type: new Abstract: Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with expla

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

DGX agent

arXiv:2607.24665v1 Announce Type: new Abstract: Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capa

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window

DGX agent

arXiv:2607.22695v1 Announce Type: new Abstract: Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment.

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

DGX agent

arXiv:2607.23458v1 Announce Type: new Abstract: Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-b

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

DGX agent

arXiv:2607.20476v1 Announce Type: new Abstract: We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three c

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

DGX agent

arXiv:2607.21063v1 Announce Type: new Abstract: Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is as

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

ETPDesigner: Multi-Agent Orchestration for Interactive Multimodal Electronic Theater Program

DGX agent

arXiv:2607.19947v1 Announce Type: new Abstract: Electronic Theater Programs (ETPs) serve as critical promotional media in the performing arts, comprising a multi-page collection of heterogeneous visua

model-releasesarxiv-cs-cv
23 Jul 2026
← Previous
1…238239240241242…255
Next →