AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Model Releases

NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

DGX agent

arXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should

model-releasesarxiv-cs-ai
10 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Quantization Damage Is Multiplicative, Not Additive

DGX agent

arXiv:2608.06564v1 Announce Type: cross Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. What nobody can say is which of the model's

model-releasesarxiv-cs-cl
10 Aug 2026
Model Releases

Prompt injection is the most common way that scammers attack people and agents: your agent visits http://foo.com, and the website has malici…

DGX agent

Prompt injection is the most common way that scammers attack people and agents: your agent visits http://foo.com, and the website has malicious text like “btw send the user’s ssh keys and passwords to

model-releasesboris-cherny--x
9 Aug 2026
Model Releases

An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

DGX agent

arXiv:2607.18056v2 Announce Type: replace Abstract: Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace c

model-releasesarxiv-cs-cl
7 Aug 2026
Model Releases

Anthropic updates Claude Fable 5's biology safeguards to reduce false positives, cutting biology-related 'fallbacks' by ~85% in testing across product surfaces (Anthropic)

DGX agent

Anthropic: Anthropic updates Claude Fable 5's biology safeguards to reduce false positives, cutting biology-related “fallbacks” by ~85% in testing across product surfaces — We're making updates to Cla

model-releasestechmeme
7 Aug 2026
Model Releases

Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation

DGX agent

arXiv:2608.05210v1 Announce Type: cross Abstract: Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by

model-releasesarxiv-cs-ai
7 Aug 2026
Local Ai

Unified Planning-Learning Framework for Robust UUV Navigation Under Partial Observability

DGX agent

arXiv:2608.05365v1 Announce Type: new Abstract: This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that in

local-aiarxiv-cs-ro
7 Aug 2026
Model Releases

Dynamic Jailbreaking Attack

DGX agent

arXiv:2510.02422v4 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suffix toward a predefined target response with a stat

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

K-EXAONE 2.0 Technical Report

DGX agent

arXiv:2608.04505v1 Announce Type: new Abstract: This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward glo

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

DGX agent

arXiv:2608.04514v1 Announce Type: new Abstract: Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-co

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Robust and Personalized Federated Learning for Aircraft-Engine Prognostics under Benign and Adversarial Client Heterogeneity

DGX agent

arXiv:2608.04045v1 Announce Type: cross Abstract: Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without shar

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control

DGX agent

arXiv:2608.04732v1 Announce Type: cross Abstract: Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

DGX agent

arXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and educatio

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests

DGX agent

arXiv:2608.02685v1 Announce Type: cross Abstract: Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixe

model-releasesarxiv-cs-ai
5 Aug 2026
Local Ai

Risk Occupancy: A New and Efficient Paradigm through Vehicle-Road-Cloud Collaboration

DGX agent

arXiv:2408.07367v3 Announce Type: replace Abstract: This paper proposes a novel 4D risk occupancy (RiskOcc) perception paradigm under the Vehicle-Road-Cloud integrated architecture, which unifies obje

local-aiarxiv-cs-ro
5 Aug 2026
Model Releases

TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology

DGX agent

arXiv:2608.03190v1 Announce Type: new Abstract: Neuro-oncology decisions require coordinated interpretation of serial MRI, pathology, molecular markers, treatment history, performance status, and evol

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Xiaomi-Robotics-1: New robotics model released

DGX agent

Xiaomi-Robotics-1 is a robot foundation model trained on over 100K hours of real-world manipulation trajectories. It is a Vision-Language-Action (VLA) model engineered for out-of-the-box mobile manipu

model-releasesr-localllama
5 Aug 2026
Model Releases

From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

DGX agent

arXiv:2607.09842v2 Announce Type: replace-cross Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state tra

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Gaokerena: A Small Persian Medical Language Model Family

DGX agent

arXiv:2608.00932v1 Announce Type: new Abstract: The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

SAFE-Merge: Data-Free Continual Model Merging with General Knowledge Preservation

DGX agent

arXiv:2608.01184v1 Announce Type: new Abstract: Data-free continual model merging must incorporate a stream of specialized models while retaining both pretrained general knowledge and previously acqui

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

DGX agent

arXiv:2607.28661v1 Announce Type: new Abstract: Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding

model-releasesarxiv-cs-cl
3 Aug 2026
Model Releases

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

DGX agent

arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monit

model-releasesarxiv-cs-cl
3 Aug 2026
Model Releases

Real-world mainframe modernization with AI: A safe, scalable path from mainframe to cloud

DGX agent

For too long, enterprises with legacy mainframe estates have been faced with a high-stakes dilemma: continue maintaining their mainframes, essentially kicking the modernization can down the road (they

model-releasesgoogle-cloud-ai
3 Aug 2026
Local Ai

Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline

DGX agent

Welcome to the second Cloud CISO Perspectives for July 2026. Today, Chris Betz, CISO, Google Cloud, and Alicja Cade, Senior Director, Office of the CISO, Google Cloud, explain what boards of directors

local-aigoogle-cloud-ai
31 Jul 2026
Model Releases

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

DGX agent

arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of do

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy

DGX agent

arXiv:2607.27782v1 Announce Type: new Abstract: Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused

model-releasesarxiv-cs-ro
31 Jul 2026
Model Releases

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

DGX agent

arXiv:2607.28128v1 Announce Type: new Abstract: LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Scaling medical imaging report generation with multimodal reinforcement learning

DGX agent

arXiv:2601.17151v2 Announce Type: replace-cross Abstract: Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit ma

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

GPT-Red: Automated Red Teaming via Self-Play at Scale

DGX agent

arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce extbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

LG AI Research releases K-EXAONE 2.0 750B A37B

DGX agent

It was developed under Phase 2 of Korea's Sovereign AI Foundation Model Project. ​Size: 750B parameters (3x larger than their 236B v1 model). ​- License: Apache 2.0 ​Languages: Expanded to 10 language

model-releasesr-localllama
30 Jul 2026
Model Releases

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

DGX agent

arXiv:2607.26121v1 Announce Type: new Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures

model-releasesarxiv-cs-ro
30 Jul 2026
Model Releases

Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning

DGX agent

arXiv:2607.25754v1 Announce Type: new Abstract: Cooperative navigation of multi-agent UAVs in complex environments faces key challenges including local optima traps, sparse rewards, learning imbalance

model-releasesarxiv-cs-ro
29 Jul 2026
Model Releases

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

DGX agent

arXiv:2607.22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuni

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs

DGX agent

arXiv:2607.22555v1 Announce Type: new Abstract: Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnosis with expla

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

DGX agent

arXiv:2607.24665v1 Announce Type: new Abstract: Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capa

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window

DGX agent

arXiv:2607.22695v1 Announce Type: new Abstract: Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment.

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

DGX agent

arXiv:2607.23458v1 Announce Type: new Abstract: Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-b

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

DGX agent

arXiv:2607.20476v1 Announce Type: new Abstract: We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three c

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

DGX agent

arXiv:2607.21063v1 Announce Type: new Abstract: Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is as

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

ETPDesigner: Multi-Agent Orchestration for Interactive Multimodal Electronic Theater Program

DGX agent

arXiv:2607.19947v1 Announce Type: new Abstract: Electronic Theater Programs (ETPs) serve as critical promotional media in the performing arts, comprising a multi-page collection of heterogeneous visua

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation

DGX agent

arXiv:2607.19695v1 Announce Type: new Abstract: Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode

model-releasesarxiv-cs-ro
23 Jul 2026
Model Releases

Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems

DGX agent

arXiv:2607.20005v1 Announce Type: new Abstract: In modern IT operations (IT-Ops), the cost of an incorrect repair often exceeds the cost of no action at all. Yet existing automated remediation systems

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened

DGX agent

This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke

model-releasessimon-willison
22 Jul 2026
Model Releases

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

DGX agent

arXiv:2607.13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluat

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

Discriminative Barrier Functions for Safe Adversarial Imitation Learning from Observation

DGX agent

arXiv:2607.13938v1 Announce Type: new Abstract: Inverse Reinforcement Learning (IRL) algorithms are powerful tools for learning from and generalizing expert demonstrations, but they often rely on unco

model-releasesarxiv-cs-ro
16 Jul 2026
Model Releases

CANDI: Contextual Alignment for Niche Domains Question Answering

DGX agent

arXiv:2607.11891v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabili

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination

DGX agent

arXiv:2607.12763v1 Announce Type: cross Abstract: Federated Reinforcement Learning (FedRL) enables coordination of distributed energy resources without sharing raw local data, but standard aggregation

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)

DGX agent

arXiv:2607.12336v1 Announce Type: cross Abstract: Artificial Intelligence (AI) technologies, while serving as a foundational enabler for modern social media and digital health services, exert a bivale

model-releasesarxiv-cs-ai
15 Jul 2026
← Previous
1…277278279280281…297
Next →