AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Model Releases

SignReasoner: Compositional Reasoning for Complex Traffic Sign Understanding via Functional Structure Units

DGX agent

arXiv:2604.10436v1 Announce Type: new Abstract: Accurate semantic understanding of complex traffic signs-including those with intricate layouts, multi-lingual text, and composite symbols-is critical f

model-releasesarxiv-cs-cv
14 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models

DGX agent

arXiv:2604.10733v1 Announce Type: cross Abstract: Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean

DGX agent

arXiv:2604.08595v1 Announce Type: cross Abstract: Existing evaluation methods for LLM-based AI systems, such as LLM-as-a-Judge, verdict systems, and NLI, do not always align well with human assessment

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Agentic Jackal: Live Execution and Semantic Value Grounding for Text-to-JQL

DGX agent

arXiv:2604.09470v1 Announce Type: new Abstract: Translating natural language into Jira Query Language (JQL) requires resolving ambiguous field references, instance-specific categorical values, and com

model-releasesarxiv-cs-cl
13 Apr 2026
Model Releases

Hidden in Plain Sight: Visual-to-Symbolic Analytical Solution Inference from Field Visualizations

DGX agent

arXiv:2604.08863v1 Announce Type: new Abstract: Recovering analytical solutions of physical fields from visual observations is a fundamental yet underexplored capability for AI-assisted scientific rea

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Medical Reasoning with Large Language Models: A Survey and MR-Bench

DGX agent

arXiv:2604.08559v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-wor

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Retrieval Augmented Classification for Confidential Documents

DGX agent

arXiv:2604.08628v1 Announce Type: cross Abstract: Unauthorized disclosure of confidential documents demands robust, low-leakage classification. In real work environments, there is a lot of inflow and

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SAGE: A Service Agent Graph-guided Evaluation Benchmark

DGX agent

arXiv:2604.09285v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has catalyzed automation in customer service, yet benchmarking their performance remains challenging. Ex

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SenBen: Sensitive Scene Graphs for Explainable Content Moderation

DGX agent

arXiv:2604.08819v1 Announce Type: cross Abstract: Content moderation systems classify images as safe or unsafe but lack spatial grounding and interpretability: they cannot explain what sensitive behav

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning

DGX agent

arXiv:2604.08780v1 Announce Type: cross Abstract: World models promise a paradigm shift in robotics, where an agent learns the underlying physics of its environment once to enable efficient planning a

model-releasesarxiv-cs-lg
13 Apr 2026
Model Releases

VAGNet: Vision-based accident anticipation with global features

DGX agent

arXiv:2604.09305v1 Announce Type: new Abstract: Traffic accidents are a leading cause of fatalities and injuries across the globe. Therefore, the ability to anticipate hazardous situations in advance

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

AgentGate: A Lightweight Structured Routing Engine for the Internet of Agents

DGX agent

arXiv:2604.06696v1 Announce Type: new Abstract: The rapid development of AI agent systems is leading to an emerging Internet of Agents, where specialized agents operate across local devices, edge node

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

ATANT: An Evaluation Framework for AI Continuity

DGX agent

arXiv:2604.06710v1 Announce Type: new Abstract: We present ATANT (Automated Test for Acceptance of Narrative Truth), an open evaluation framework for measuring continuity in AI systems: the ability to

model-releasesarxiv-cs-ai
10 Apr 2026
Local Ai

Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test

DGX agent

arXiv:2506.06975v5 Announce Type: replace-cross Abstract: As API access becomes a primary interface to large language models (LLMs), users often interact with black-box systems that offer little trans

local-aiarxiv-cs-cl
10 Apr 2026
Model Releases

Benchmarking LLM Tool-Use in the Wild

DGX agent

arXiv:2604.06185v1 Announce Type: cross Abstract: Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inh

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses

DGX agent

arXiv:2604.06216v1 Announce Type: cross Abstract: As LLM-powered chatbots are increasingly deployed in mental health services, detecting hallucinations and omissions has become critical for user safet

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Distributed Multi-Layer Editing for Rule-Level Knowledge in Large Language Models

DGX agent

arXiv:2604.08284v1 Announce Type: new Abstract: Large language models store not only isolated facts but also rules that support reasoning across symbolic expressions, natural language explanations, an

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction

DGX agent

arXiv:2604.07659v1 Announce Type: new Abstract: Large language models (LLMs) hold significant promise for healthcare, yet their reliability in high-stakes clinical settings is often compromised by hal

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Knowledge Graphs Generation from Cultural Heritage Texts: Combining LLMs and Ontological Engineering for Scholarly Debates

DGX agent

arXiv:2511.10354v1 Announce Type: cross Abstract: Cultural Heritage texts contain rich knowledge that is difficult to query systematically due to the challenges of converting unstructured discourse in

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces

DGX agent

arXiv:2604.06188v1 Announce Type: cross Abstract: People increasingly hold sustained, open-ended conversations with large language models (LLMs). Public reports and early studies suggest that, in such

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

DGX agent

arXiv:2604.08031v1 Announce Type: cross Abstract: Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an int

model-releasesarxiv-cs-cv
10 Apr 2026
Research

ReCellTy: Domain-Specific Knowledge Graph Retrieval-Augmented LLMs Reasoning Workflow for Single-Cell Annotation

DGX agent

arXiv:2505.00017v2 Announce Type: replace Abstract: With the rapid development of large language models (LLMs), their application to cell type annotation has drawn increasing attention. However, gener

researcharxiv-cs-cl
10 Apr 2026
Model Releases

Robustness Risk of Conversational Retrieval: Identifying and Mitigating Noise Sensitivity in Qwen3-Embedding Model

DGX agent

arXiv:2604.06176v1 Announce Type: cross Abstract: We present an empirical study of embedding-based retrieval under realistic conversational settings, where queries are short, dialogue-like, and weakly

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems

DGX agent

arXiv:2604.06811v1 Announce Type: cross Abstract: Skill-based agent systems tackle complex tasks by composing reusable skills, improving modularity and scalability while introducing a largely unexamin

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Validated Intent Compilation for Constrained Routing in LEO Mega-Constellations

DGX agent

arXiv:2604.07264v1 Announce Type: cross Abstract: Operating LEO mega-constellations requires translating high-level operator intents ('reroute financial traffic away from polar links under 80 ms') int

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

DGX agent

arXiv:2604.06182v1 Announce Type: cross Abstract: Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of

model-releasesarxiv-cs-ai
10 Apr 2026
← Previous
1…255256257
Next →