AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models

DGX agent

arXiv:2604.15967v1 Announce Type: cross Abstract: Despite the remarkable synthesis capabilities of text-to-image (T2I) models, safeguarding them against content violations remains a persistent challen

model-releasesarxiv-cs-cv
20 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ECM Contracts: Contract-Aware, Versioned, and Governable Capability Interfaces for Embodied Agents

DGX agent

arXiv:2604.13097v1 Announce Type: cross Abstract: Embodied agents increasingly rely on modular capabilities that can be installed, upgraded, composed, and governed at runtime. Prior work has introduce

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

Mechanistic Decoding of Cognitive Constructs in LLMs

DGX agent

arXiv:2604.14593v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate increasingly sophisticated affective capabilities, the internal mechanisms by which they process complex

model-releasesarxiv-cs-cl
17 Apr 2026
Local Ai

Rethinking AI Hardware: A Three-Layer Cognitive Architecture for Autonomous Agents

DGX agent

arXiv:2604.13757v1 Announce Type: new Abstract: The next generation of autonomous AI systems will be constrained not only by model capability, but by how intelligence is structured across heterogeneou

local-aiarxiv-cs-ai
17 Apr 2026
Model Releases

A Proactive EMR Assistant for Doctor-Patient Dialogue: Streaming ASR, Belief Stabilization, and Preliminary Controlled Evaluation

DGX agent

arXiv:2604.13059v1 Announce Type: new Abstract: Most dialogue-based electronic medical record (EMR) systems still behave as passive pipelines: transcribe speech, extract information, and generate the

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Document-tuning for robust alignment to animals

DGX agent

arXiv:2604.13076v1 Announce Type: new Abstract: We investigate the robustness of value alignment via finetuning with synthetic documents, using animal compassion as a value that is both important in i

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

DGX agent

arXiv:2505.19662v3 Announce Type: replace-cross Abstract: This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agent

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

DGX agent

arXiv:2604.13803v1 Announce Type: new Abstract: Vision-language models are increasingly deployed in high-stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

Mosaic: An Extensible Framework for Composing Rule-Based and Learned Motion Planners

DGX agent

arXiv:2604.13853v1 Announce Type: new Abstract: Safe and explainable motion planning remains a central challenge in autonomous driving. While rule-based planners offer predictable and explainable beha

model-releasesarxiv-cs-ro
16 Apr 2026
Model Releases

Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub

DGX agent

arXiv:2604.13064v1 Announce Type: new Abstract: Skill ecosystems have emerged as an increasingly important layer in Large Language Model (LLM) agent systems, enabling reusable task packaging, public d

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Robust Reward Modeling for Large Language Models via Causal Decomposition

DGX agent

arXiv:2604.13833v1 Announce Type: new Abstract: Reward models are central to aligning large language models, yet they often overfit to spurious cues such as response length and overly agreeable tone.

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious

DGX agent

arXiv:2604.13051v1 Announce Type: new Abstract: There is debate about whether LLMs can be conscious. We investigate a distinct question: if a model claims to be conscious, how does this affect its dow

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Benchmarking Deflection and Hallucination in Large Vision-Language Models

DGX agent

arXiv:2604.12033v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) increasingly rely on retrieval to answer knowledge-intensive multimodal questions. Existing benchmarks overlook c

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Latent Chain-of-Thought World Modeling for End-to-End Driving

DGX agent

arXiv:2512.10226v2 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models for autonomous driving explore inference-time reasoning as a way to improve driving performance and safet

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

Memory as Metabolism: A Design for Companion Knowledge Systems

DGX agent

arXiv:2604.12034v1 Announce Type: new Abstract: Retrieval-Augmented Generation remains the dominant pattern for giving LLMs persistent memory, but a visible cluster of personal wiki-style memory archi

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

DGX agent

arXiv:2604.12102v1 Announce Type: new Abstract: We introduce compute-grounded reasoning (CGR), a design paradigm for spatial-aware research agents in which every answerable sub-problem is resolved by

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Beyond the Beep: Scalable Collision Anticipation and Real-Time Explainability with BADAS-2.0

DGX agent

arXiv:2604.05767v2 Announce Type: replace-cross Abstract: We present BADAS-2.0, the second generation of our collision anticipation system, building on BADAS-1.0, which showed that fine-tuning V-JEPA2

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

BLUEmed: Retrieval-Augmented Multi-Agent Debate for Clinical Error Detection

DGX agent

arXiv:2604.10389v1 Announce Type: new Abstract: Terminology substitution errors in clinical notes, where one medical term is replaced by a linguistically valid but clinically different term, pose a pe

model-releasesarxiv-cs-cl
14 Apr 2026
Local Ai

ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection

DGX agent

arXiv:2604.11790v1 Announce Type: cross Abstract: Tool-augmented Large Language Model (LLM) agents have demonstrated impressive capabilities in automating complex, multi-step real-world tasks, yet rem

local-aiarxiv-cs-ai
14 Apr 2026
Model Releases

Comparative Analysis of Large Language Models in Healthcare

DGX agent

arXiv:2604.10316v1 Announce Type: new Abstract: Background: Large Language Models (LLMs) are transforming artificial intelligence applications in healthcare due to their ability to understand, generat

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

Detecting Corporate AI-Washing via Cross-Modal Semantic Inconsistency Learning

DGX agent

arXiv:2604.09644v1 Announce Type: cross Abstract: Corporate AI-washing-the strategic misrepresentation of AI capabilities via exaggerated or fabricated cross-channel disclosures-has emerged as a syste

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Edu-MMBias: A Three-Tier Multimodal Benchmark for Auditing Social Bias in Vision-Language Models under Educational Contexts

DGX agent

arXiv:2604.10200v1 Announce Type: new Abstract: As Vision-Language Models (VLMs) become integral to educational decision-making, ensuring their fairness is paramount. However, current text-centric eva

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Evaluating Small Open LLMs for Medical Question Answering: A Practical Framework

DGX agent

arXiv:2604.10535v1 Announce Type: cross Abstract: Incorporating large language models (LLMs) in medical question answering demands more than high average accuracy: a model that returns substantively d

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

From GPT-3 to GPT-5: Mapping their capabilities, scope, limitations, and consequences

DGX agent

arXiv:2604.10332v1 Announce Type: new Abstract: We present the progress of the GPT family from GPT-3 through GPT-3.5, GPT-4, GPT-4 Turbo, GPT-4o, GPT-4.1, and the GPT-5 family. Our work is comparative

model-releasesarxiv-cs-ai
14 Apr 2026
Local Ai

Gait Recognition with Temporal Kolmogorov-Arnold Networks

DGX agent

arXiv:2604.09990v1 Announce Type: new Abstract: Gait recognition is a biometric modality that identifies individuals from their characteristic walking patterns. Unlike conventional biometric traits, g

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

Generation-Augmented Generation: A Plug-and-Play Framework for Private Knowledge Injection in Large Language Models

DGX agent

arXiv:2601.08209v3 Announce Type: replace Abstract: In domains such as materials science, biomedicine, and finance, high-stakes deployment of large language models (LLMs) requires injecting private, d

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

HG-Lane: High-Fidelity Generation of Lane Scenes under Adverse Weather and Lighting Conditions without Re-annotation

DGX agent

arXiv:2603.10128v2 Announce Type: replace Abstract: Lane detection is a crucial task in autonomous driving, as it helps ensure the safe operation of vehicles. However, existing datasets such as CULane

model-releasesarxiv-cs-cv
14 Apr 2026
Local Ai

Intelligent bear deterrence system based on computer vision: Reducing human bear conflicts in remote areas

DGX agent

arXiv:2503.23178v2 Announce Type: replace Abstract: Conflicts between humans and bears on the Tibetan Plateau present substantial threats to local communities and hinder wildlife preservation initiati

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models

DGX agent

arXiv:2604.11609v1 Announce Type: new Abstract: Large language models exhibit sycophantic tendencies--validating incorrect user beliefs to appear agreeable. We investigate whether this behavior varies

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

PSF-Med: Measuring and Explaining Paraphrase Sensitivity in Medical Vision Language Models

DGX agent

arXiv:2602.21428v2 Announce Type: replace Abstract: Medical Vision Language Models (VLMs) can change their answers when clinicians rephrase the same question, a failure mode that threatens deployment

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

RCBSF: A Multi-Agent Framework for Automated Contract Revision via Stackelberg Game

DGX agent

arXiv:2604.10740v1 Announce Type: new Abstract: Despite the widespread adoption of Large Language Models (LLMs) in Legal AI, their utility for automated contract revision remains impeded by hallucinat

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

RiTeK: A Dataset for Large Language Models Complex Reasoning over Textual Knowledge Graphs in Medicine

DGX agent

arXiv:2410.13987v3 Announce Type: replace Abstract: Answering complex real-world questions in the medical domain often requires accurate retrieval from medical Textual Knowledge Graphs (medical TKGs),

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

SignReasoner: Compositional Reasoning for Complex Traffic Sign Understanding via Functional Structure Units

DGX agent

arXiv:2604.10436v1 Announce Type: new Abstract: Accurate semantic understanding of complex traffic signs-including those with intricate layouts, multi-lingual text, and composite symbols-is critical f

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models

DGX agent

arXiv:2604.10733v1 Announce Type: cross Abstract: Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean

DGX agent

arXiv:2604.08595v1 Announce Type: cross Abstract: Existing evaluation methods for LLM-based AI systems, such as LLM-as-a-Judge, verdict systems, and NLI, do not always align well with human assessment

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Agentic Jackal: Live Execution and Semantic Value Grounding for Text-to-JQL

DGX agent

arXiv:2604.09470v1 Announce Type: new Abstract: Translating natural language into Jira Query Language (JQL) requires resolving ambiguous field references, instance-specific categorical values, and com

model-releasesarxiv-cs-cl
13 Apr 2026
Model Releases

Hidden in Plain Sight: Visual-to-Symbolic Analytical Solution Inference from Field Visualizations

DGX agent

arXiv:2604.08863v1 Announce Type: new Abstract: Recovering analytical solutions of physical fields from visual observations is a fundamental yet underexplored capability for AI-assisted scientific rea

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Medical Reasoning with Large Language Models: A Survey and MR-Bench

DGX agent

arXiv:2604.08559v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-wor

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Retrieval Augmented Classification for Confidential Documents

DGX agent

arXiv:2604.08628v1 Announce Type: cross Abstract: Unauthorized disclosure of confidential documents demands robust, low-leakage classification. In real work environments, there is a lot of inflow and

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SAGE: A Service Agent Graph-guided Evaluation Benchmark

DGX agent

arXiv:2604.09285v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has catalyzed automation in customer service, yet benchmarking their performance remains challenging. Ex

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SenBen: Sensitive Scene Graphs for Explainable Content Moderation

DGX agent

arXiv:2604.08819v1 Announce Type: cross Abstract: Content moderation systems classify images as safe or unsafe but lack spatial grounding and interpretability: they cannot explain what sensitive behav

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning

DGX agent

arXiv:2604.08780v1 Announce Type: cross Abstract: World models promise a paradigm shift in robotics, where an agent learns the underlying physics of its environment once to enable efficient planning a

model-releasesarxiv-cs-lg
13 Apr 2026
Model Releases

VAGNet: Vision-based accident anticipation with global features

DGX agent

arXiv:2604.09305v1 Announce Type: new Abstract: Traffic accidents are a leading cause of fatalities and injuries across the globe. Therefore, the ability to anticipate hazardous situations in advance

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

AgentGate: A Lightweight Structured Routing Engine for the Internet of Agents

DGX agent

arXiv:2604.06696v1 Announce Type: new Abstract: The rapid development of AI agent systems is leading to an emerging Internet of Agents, where specialized agents operate across local devices, edge node

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

ATANT: An Evaluation Framework for AI Continuity

DGX agent

arXiv:2604.06710v1 Announce Type: new Abstract: We present ATANT (Automated Test for Acceptance of Narrative Truth), an open evaluation framework for measuring continuity in AI systems: the ability to

model-releasesarxiv-cs-ai
10 Apr 2026
Local Ai

Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test

DGX agent

arXiv:2506.06975v5 Announce Type: replace-cross Abstract: As API access becomes a primary interface to large language models (LLMs), users often interact with black-box systems that offer little trans

local-aiarxiv-cs-cl
10 Apr 2026
Model Releases

Benchmarking LLM Tool-Use in the Wild

DGX agent

arXiv:2604.06185v1 Announce Type: cross Abstract: Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inh

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses

DGX agent

arXiv:2604.06216v1 Announce Type: cross Abstract: As LLM-powered chatbots are increasingly deployed in mental health services, detecting hallucinations and omissions has become critical for user safet

model-releasesarxiv-cs-ai
10 Apr 2026
← Previous
1…252253254255
Next →