AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “automated”

GridTimelineEvolution
3,914 results
Model Releases

From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering

DGX agent

arXiv:2604.04948v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) systems depend critically on the quality of document preprocessing, yet no prior study has evaluated PDF

model-releasesarxiv-cs-ai
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering

DGX agent

arXiv:2605.27249v1 Announce Type: new Abstract: An effective method of teaching across disciplines is to provide examples of high-quality work. However, an example may be significantly different from

researcharxiv-cs-ai
27 May 2026
Local Ai

HRVConformer: Neonatal Hypoxic-Ischemic Encephalopathy Classification from the Heart Rate signals

DGX agent

arXiv:2605.26190v1 Announce Type: cross Abstract: This paper presents the HRVConformer, a novel deep learning architecture for the classification of hypoxic-ischemic encephalopathy (HIE) using the ins

local-aiarxiv-cs-ai
27 May 2026
Safety

LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback

DGX agent

arXiv:2509.18384v2 Announce Type: replace Abstract: Large language models (LLMs) can translate natural language instructions into executable action plans for robotics, autonomous driving, and other do

safetyarxiv-cs-ro
27 May 2026
Model Releases

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

DGX agent

arXiv:2605.26781v1 Announce Type: new Abstract: Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

DGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

model-releasesarxiv-cs-ai
27 May 2026
Tutorials

Optimising Factual Consistency in Summarisation via Preference Learning from Multiple Imperfect Metrics

DGX agent

arXiv:2605.26840v1 Announce Type: new Abstract: Reinforcement learning with evaluation metrics as rewards is widely used to enhance specific capabilities of language models. However, for tasks such as

tutorialsarxiv-cs-cl
27 May 2026
Local Ai

Prototyping an End-to-End Multi-Modal Tiny-CNN for Cardiovascular Sensor Patches

DGX agent

arXiv:2510.18668v2 Announce Type: replace-cross Abstract: The vast majority of cardiovascular diseases may be preventable if early signs and risk factors are detected. Cardiovascular monitoring with b

local-aiarxiv-cs-cv
27 May 2026
Local Ai

ReVEL: Multi-Turn Reflective LLM-Guided Heuristic Evolution via Structured Performance Feedback

DGX agent

arXiv:2604.04940v2 Announce Type: replace Abstract: Designing effective heuristics for NP-hard combinatorial optimization problems remains challenging and often requires substantial domain expertise.

local-aiarxiv-cs-ai
27 May 2026
Model Releases

RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction

DGX agent

arXiv:2605.26862v1 Announce Type: new Abstract: Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scen

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

DGX agent

arXiv:2510.09606v2 Announce Type: replace Abstract: With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still strug

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

SteelDS: A High-Resolution Video Dataset of E40 Steel Scrap for Object Detection and Instance Segmentation

DGX agent

arXiv:2605.26682v1 Announce Type: cross Abstract: This dataset provides high-resolution, annotated video sequences of shredded E40-grade steel and copper scrap on a conveyor belt. Captured in a contro

model-releasesarxiv-cs-cv
27 May 2026
Research

Structure-Adaptive Conformal Inference for Large-Scale Out-of-Distribution Testing

DGX agent

arXiv:2605.26429v1 Announce Type: cross Abstract: This paper addresses structured out-of-distribution (OOD) testing in high-stakes machine learning applications. Traditional conformal methods rely on

researcharxiv-cs-ai
27 May 2026
Applications

Towards Real-World Identification of Fatigued Muscle Groups via Musculoskeletal Simulation

DGX agent

arXiv:2605.26151v1 Announce Type: cross Abstract: Contactless diagnosis of musculoskeletal disorders can potentially improve population health as well as robot behaviours in collaborative settings. Ho

applicationsarxiv-cs-ro
27 May 2026
Applications

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action

DGX agent

arXiv:2510.17790v3 Announce Type: replace-cross Abstract: Computer-use agents face a fundamental limitation. They rely exclusively on primitive GUI actions (click, type, scroll), creating brittle exec

applicationsarxiv-cs-cl
27 May 2026
Research

Understanding the Challenges in Iterative Generative Optimization with LLMs

DGX agent

arXiv:2603.23994v2 Announce Type: replace-cross Abstract: Generative optimization uses large language models (LLMs) to iteratively improve artifacts (such as code, workflows or prompts) using executio

researcharxiv-cs-ai
27 May 2026
Hardware

Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU

DGX agent

arXiv:2605.26118v1 Announce Type: cross Abstract: Porting deep learning algorithms to new hardware accelerators requires developers to repeatedly apply the same low-level optimizations -- quantization

hardwarearxiv-cs-ai
27 May 2026
Model Releases

Zero-Shot MARL Benchmark in the Cyber-Physical Mobility Lab

DGX agent

arXiv:2601.16578v2 Announce Type: replace Abstract: We present a reproducible benchmark for evaluating sim-to-real transfer of Multi-Agent Reinforcement Learning (MARL) policies for Connected and Auto

model-releasesarxiv-cs-ro
27 May 2026
Agents

A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback

DGX agent

arXiv:2605.25440v1 Announce Type: cross Abstract: Verbal feedback delivered by attending surgeons in the operating room plays a critical formative role in resident trainee skill acquisition. Yet, asse

agentsarxiv-cs-ai
26 May 2026
Agents

Architecting Agentic Communities using Design Patterns

DGX agent

arXiv:2601.03624v3 Announce Type: replace Abstract: The rapid evolution of Large Language Models (LLM) and subsequent Agentic AI technologies requires systematic architectural guidance for building so

agentsarxiv-cs-ai
26 May 2026
Safety

Auditing medical multi-agent AI reveals risks of false consensus

DGX agent

arXiv:2510.10185v2 Announce Type: replace-cross Abstract: Large language models are increasingly being assembled into medical multi-agent systems that emulate multidisciplinary consultation through sp

safetyarxiv-cs-ai
26 May 2026
Research

Beyond Control-Flow: Integrating the Resource Perspective into Multi-Collaborative Process Modeling from Text

DGX agent

arXiv:2605.24546v1 Announce Type: new Abstract: Process modeling is a sub-domain of Business Process Management (BPM) focused on the translation of process artifacts into formal models. This task trad

researcharxiv-cs-ai
26 May 2026
Model Releases

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

DGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

BODHI: Precise OS Kernel Specification Inference

DGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

DGX agent

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Y

model-releasesarxiv-cs-ai
26 May 2026
Agents

CRISP -- Clustering-Based Redundancy-Reduced Instance Sampling for Pathology Case Representation and Retrieval

DGX agent

arXiv:2605.24253v1 Announce Type: cross Abstract: Digital pathology archives increasingly contain multiple whole-slide images (WSIs) per case, capturing spatially distinct tumour regions and reflectin

agentsarxiv-cs-ai
26 May 2026
Model Releases

DRInQ: Evaluating Conversational Implicature with Controlled Context Variation

DGX agent

arXiv:2605.24267v1 Announce Type: new Abstract: Human conversation relies heavily on conversational implicature, in which speakers convey meanings that are suggested rather than explicitly stated. Alt

model-releasesarxiv-cs-cl
26 May 2026
Safety

Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters

DGX agent

arXiv:2505.18979v2 Announce Type: replace Abstract: Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters.

safetyarxiv-cs-lg
26 May 2026
Model Releases

Empirical Analysis and Detection of Hallucinations in LLM-Generated Bug Report Summaries

DGX agent

arXiv:2605.24137v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to generate summaries of software bug reports, including sections such as Steps-to-Reproduce (S2R),

model-releasesarxiv-cs-ai
26 May 2026
Safety

EPPC-OASIS: Ontology-Aware Adaptation and Structured Inference Refinement for Electronic Patient-Provider Communication Mining in Secure Messages

DGX agent

arXiv:2605.24172v1 Announce Type: new Abstract: Secure patient-provider messages contain clinically important communication behaviors that are difficult to characterize manually at scale. The Electron

safetyarxiv-cs-ai
26 May 2026
Safety

FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models

DGX agent

arXiv:2510.22827v3 Announce Type: replace-cross Abstract: Evaluating text-to-image (T2I) systems requires judging not only whether an image matches a prompt, but also whether socially salient attribut

safetyarxiv-cs-lg
26 May 2026
Model Releases

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

DGX agent

arXiv:2605.25052v1 Announce Type: new Abstract: Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these t

model-releasesarxiv-cs-cl
26 May 2026
Safety

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

DGX agent

arXiv:2605.25447v1 Announce Type: new Abstract: Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-pu

safetyarxiv-cs-cl
26 May 2026
Model Releases

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

DGX agent

arXiv:2605.24636v1 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios re

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

DGX agent

arXiv:2507.19219v2 Announce Type: replace Abstract: Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalan

model-releasesarxiv-cs-cl
26 May 2026
Safety

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

DGX agent

arXiv:2605.24883v1 Announce Type: new Abstract: The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on con

safetyarxiv-cs-ai
26 May 2026
Agents

Iterate Until Retrieved: Factual Nugget Optimization for Discoverable Continual Corrections in Agentic RAG

DGX agent

arXiv:2605.25641v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) systems in complex B2B (business-to-business) settings may often receive free-form response feedback. Rathe

agentsarxiv-cs-cl
26 May 2026
Safety

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

DGX agent

arXiv:2605.23997v1 Announce Type: cross Abstract: Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they

safetyarxiv-cs-ai
26 May 2026
Applications

Lake Detection and Water Quality Estimation in Sentinel-2 Data

DGX agent

arXiv:2605.24515v1 Announce Type: new Abstract: With climate change and increasing human pressure on natural landscapes, inland water resources are becoming progressively scarcer, more vulnerable, and

applicationsarxiv-cs-lg
26 May 2026
Tutorials

Leveraging Spreading Activation for Improved Document Retrieval in Knowledge-Graph-Based RAG Systems

DGX agent

arXiv:2512.15922v3 Announce Type: replace Abstract: Despite initial successes and a variety of architectures, retrieval-augmented generation systems still struggle to reliably retrieve and connect the

tutorialsarxiv-cs-ai
26 May 2026
Model Releases

Memory-Induced Tool-Drift in LLM Agents

DGX agent

arXiv:2605.24941v1 Announce Type: cross Abstract: Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinn

model-releasesarxiv-cs-lg
26 May 2026
Agents

Meta-Engineering Harnesses for AI-Native Software Production: A Contract-Driven Adversarial Verification Architecture with Early Deployment Report

DGX agent

arXiv:2605.25665v1 Announce Type: cross Abstract: AI-native software development is often evaluated at the level of individual models, prompts, or generated artifacts. This framing is insufficient for

agentsarxiv-cs-ai
26 May 2026
Research

Methodology for Creating a Clinically Verified Dermoscopic Image Dataset

DGX agent

arXiv:2605.25168v1 Announce Type: cross Abstract: This study presents a methodology for constructing a clinically verified dataset of dermatoscopic images for medical informatics research. The relevan

researcharxiv-cs-ai
26 May 2026
Model Releases

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

DGX agent

arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, h

model-releasesarxiv-cs-cl
26 May 2026
Safety

Multi-Agent Coordination Adaptation via Structure-Guided Orchestration

DGX agent

arXiv:2605.25746v1 Announce Type: cross Abstract: As large language model (LLM)-based multi-agent systems scale to handle increasingly complex tasks, balancing structural stability and dynamic adaptab

safetyarxiv-cs-ai
26 May 2026
Safety

PrivFusion: A Privacy-preserving Multi-Agent Framework for Harmonizing Distributed Datasets

DGX agent

arXiv:2605.24249v1 Announce Type: new Abstract: The growing availability of clinical data has increased the use of machine learning, yet centralized data aggregation is often infeasible for sensitive

safetyarxiv-cs-lg
26 May 2026
Safety

ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents

DGX agent

arXiv:2605.24900v1 Announce Type: new Abstract: Proactive task-oriented agents must autonomously anticipate user needs, identify actionable opportunities, and trigger software actions at appropriate m

safetyarxiv-cs-ai
26 May 2026
Tutorials

QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis

DGX agent

arXiv:2601.06870v2 Announce Type: replace-cross Abstract: Multimodal large language models have demonstrated strong ability in capturing semantic representations for multimodal sentiment analysis. The

tutorialsarxiv-cs-ai
26 May 2026
← Previous
1…6263646566…82
Next →