AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “automated”

GridTimelineEvolution
4,980 results
27 May 2026

Automatic Layer Selection for Hallucination Detection

TutorialsDGX agent

arXiv:2605.26366v1 Announce Type: new Abstract: Recent studies on hallucination detection have shown that hallucination-related signals are more strongly encoded in intermediate layers than in the fin

Been using Grok Build these past few days, and the thing that really got me hooked is Imagine and Imagine Video. I built a full dinosaur enc…

ApplicationsDGX agent

Been using Grok Build these past few days, and the thing that really got me hooked is Imagine and Imagine Video. I built a full dinosaur encyclopedia site — every image, every video clip on it, all ge

Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

ResearchDGX agent

arXiv:2605.27189v1 Announce Type: new Abstract: This study examines the relationship between speech representations and the hierarchical structure of cognitive assessment in mild cognitive impairment.


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Beyond Holistic Models: Systematic Component-level Benchmarking of Deep Multivariate Time-Series Forecasting

Model ReleasesDGX agent

arXiv:2605.26562v1 Announce Type: new Abstract: While previous research in multivariate time series forecasting has focused on developing complex holistic models, this work advocates for a shift towar

Building self-improving tax agents with Codex

AgentsDGX agent

This article describes how OpenAI's Codex model can be used to build autonomous tax agents capable of self-improvement through code generation and execution. The work demonstrates using large language

CNNs, Transformers, Hybrid, and Vision Language Models for Skin Cancer Detection

Model ReleasesDGX agent

arXiv:2605.26294v1 Announce Type: new Abstract: Skin cancer is a common and fast rising malignancy worldwide. Early detection is critical for improving outcomes. Deep learning models trained on dermos

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving

Model ReleasesDGX agent

arXiv:2601.14702v2 Announce Type: replace Abstract: Autonomous driving requires reliable perception and safe decision-making in complex scenarios. Recent vision-language models (VLMs) demonstrate reas

Experiments in Agentic AI for Science

Local AiDGX agent

arXiv:2605.26305v1 Announce Type: new Abstract: This paper details two novel frameworks for developing autonomous, agentic AI in scientific workflows. Both systems leverage a hybrid Local Body, Remote

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

SafetyDGX agent

arXiv:2605.26926v1 Announce Type: new Abstract: Computing legal indicators from normative texts is a key task in legal monitoring and policy evaluation, but presents significant challenges due to the

From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering

Model ReleasesDGX agent

arXiv:2604.04948v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) systems depend critically on the quality of document preprocessing, yet no prior study has evaluated PDF

Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering

ResearchDGX agent

arXiv:2605.27249v1 Announce Type: new Abstract: An effective method of teaching across disciplines is to provide examples of high-quality work. However, an example may be significantly different from

HRVConformer: Neonatal Hypoxic-Ischemic Encephalopathy Classification from the Heart Rate signals

Local AiDGX agent

arXiv:2605.26190v1 Announce Type: cross Abstract: This paper presents the HRVConformer, a novel deep learning architecture for the classification of hypoxic-ischemic encephalopathy (HIE) using the ins

I really appreciate the lessons and technical ideas @samaysham & team were able to share about their tax agent system, which learns from pro…

AgentsDGX agent

I really appreciate the lessons and technical ideas @samaysham & team were able to share about their tax agent system, which learns from production traces to self-improve via detailed tracing tightly

I think Anthropic and OpenAI have found product-market fit

Model ReleasesDGX agent

Anthropic are strongly rumored to be about to have their first profitable quarter. Stories are circulating of companies surprised at how expensive their LLM bills are becoming from usage by their staf

LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback

SafetyDGX agent

arXiv:2509.18384v2 Announce Type: replace Abstract: Large language models (LLMs) can translate natural language instructions into executable action plans for robotics, autonomous driving, and other do

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

Model ReleasesDGX agent

arXiv:2605.26781v1 Announce Type: new Abstract: Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

Model ReleasesDGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

Optimising Factual Consistency in Summarisation via Preference Learning from Multiple Imperfect Metrics

TutorialsDGX agent

arXiv:2605.26840v1 Announce Type: new Abstract: Reinforcement learning with evaluation metrics as rewards is widely used to enhance specific capabilities of language models. However, for tasks such as

Prototyping an End-to-End Multi-Modal Tiny-CNN for Cardiovascular Sensor Patches

Local AiDGX agent

arXiv:2510.18668v2 Announce Type: replace-cross Abstract: The vast majority of cardiovascular diseases may be preventable if early signs and risk factors are detected. Cardiovascular monitoring with b

ReVEL: Multi-Turn Reflective LLM-Guided Heuristic Evolution via Structured Performance Feedback

Local AiDGX agent

arXiv:2604.04940v2 Announce Type: replace Abstract: Designing effective heuristics for NP-hard combinatorial optimization problems remains challenging and often requires substantial domain expertise.

RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction

Model ReleasesDGX agent

arXiv:2605.26862v1 Announce Type: new Abstract: Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scen

Role-Based Access Control for Humans and Agents

ApplicationsDGX agent

This article discusses implementing role-based access control (RBAC) systems that work for both human users and AI agents, likely addressing how to manage permissions and authentication in environment

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

Model ReleasesDGX agent

arXiv:2510.09606v2 Announce Type: replace Abstract: With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still strug

SteelDS: A High-Resolution Video Dataset of E40 Steel Scrap for Object Detection and Instance Segmentation

Model ReleasesDGX agent

arXiv:2605.26682v1 Announce Type: cross Abstract: This dataset provides high-resolution, annotated video sequences of shredded E40-grade steel and copper scrap on a conveyor belt. Captured in a contro

Structure-Adaptive Conformal Inference for Large-Scale Out-of-Distribution Testing

ResearchDGX agent

arXiv:2605.26429v1 Announce Type: cross Abstract: This paper addresses structured out-of-distribution (OOD) testing in high-stakes machine learning applications. Traditional conformal methods rely on

Towards Real-World Identification of Fatigued Muscle Groups via Musculoskeletal Simulation

ApplicationsDGX agent

arXiv:2605.26151v1 Announce Type: cross Abstract: Contactless diagnosis of musculoskeletal disorders can potentially improve population health as well as robot behaviours in collaborative settings. Ho

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action

ApplicationsDGX agent

arXiv:2510.17790v3 Announce Type: replace-cross Abstract: Computer-use agents face a fundamental limitation. They rely exclusively on primitive GUI actions (click, type, scroll), creating brittle exec

Understanding the Challenges in Iterative Generative Optimization with LLMs

ResearchDGX agent

arXiv:2603.23994v2 Announce Type: replace-cross Abstract: Generative optimization uses large language models (LLMs) to iteratively improve artifacts (such as code, workflows or prompts) using executio

Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU

HardwareDGX agent

arXiv:2605.26118v1 Announce Type: cross Abstract: Porting deep learning algorithms to new hardware accelerators requires developers to repeatedly apply the same low-level optimizations -- quantization

Zero-Shot MARL Benchmark in the Cyber-Physical Mobility Lab

Model ReleasesDGX agent

arXiv:2601.16578v2 Announce Type: replace Abstract: We present a reproducible benchmark for evaluating sim-to-real transfer of Multi-Agent Reinforcement Learning (MARL) policies for Connected and Auto

26 May 2026

A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback

AgentsDGX agent

arXiv:2605.25440v1 Announce Type: cross Abstract: Verbal feedback delivered by attending surgeons in the operating room plays a critical formative role in resident trainee skill acquisition. Yet, asse

AgentWatch: Proactive AWS monitoring with ambient agents

AgentsDGX agent

In this post, we demonstrate the capabilities of AgentWatch through practical implementation. You will see how the solution performs infrastructure checks every 15 minutes, summarizing CloudWatch metr

Architecting Agentic Communities using Design Patterns

AgentsDGX agent

arXiv:2601.03624v3 Announce Type: replace Abstract: The rapid evolution of Large Language Models (LLM) and subsequent Agentic AI technologies requires systematic architectural guidance for building so

Auditing medical multi-agent AI reveals risks of false consensus

SafetyDGX agent

arXiv:2510.10185v2 Announce Type: replace-cross Abstract: Large language models are increasingly being assembled into medical multi-agent systems that emulate multidisciplinary consultation through sp

Beyond Control-Flow: Integrating the Resource Perspective into Multi-Collaborative Process Modeling from Text

ResearchDGX agent

arXiv:2605.24546v1 Announce Type: new Abstract: Process modeling is a sub-domain of Business Process Management (BPM) focused on the translation of process artifacts into formal models. This task trad

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

Model ReleasesDGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

BODHI: Precise OS Kernel Specification Inference

Model ReleasesDGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

Build high-performance generative AI systems with Strands Agents, NVIDIA NIM, and Amazon Bedrock AgentCore

HardwareDGX agent

In this post you'll learn how to build a multi-agent campaign review system that demonstrates parallel reasoning, context persistence, and traceable execution paths using an integrated architecture th

Choosing to Stay Human

ApplicationsDGX agent

This essay by Ethan Mollick explores how individuals can maintain their humanity and agency in an increasingly AI-driven world, likely addressing practical strategies for preserving human skills, crea

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

Model ReleasesDGX agent

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Y

CRISP -- Clustering-Based Redundancy-Reduced Instance Sampling for Pathology Case Representation and Retrieval

AgentsDGX agent

arXiv:2605.24253v1 Announce Type: cross Abstract: Digital pathology archives increasingly contain multiple whole-slide images (WSIs) per case, capturing spatially distinct tumour regions and reflectin

DRInQ: Evaluating Conversational Implicature with Controlled Context Variation

Model ReleasesDGX agent

arXiv:2605.24267v1 Announce Type: new Abstract: Human conversation relies heavily on conversational implicature, in which speakers convey meanings that are suggested rather than explicitly stated. Alt

Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters

SafetyDGX agent

arXiv:2505.18979v2 Announce Type: replace Abstract: Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters.

Empirical Analysis and Detection of Hallucinations in LLM-Generated Bug Report Summaries

Model ReleasesDGX agent

arXiv:2605.24137v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to generate summaries of software bug reports, including sections such as Steps-to-Reproduce (S2R),

EPPC-OASIS: Ontology-Aware Adaptation and Structured Inference Refinement for Electronic Patient-Provider Communication Mining in Secure Messages

SafetyDGX agent

arXiv:2605.24172v1 Announce Type: new Abstract: Secure patient-provider messages contain clinically important communication behaviors that are difficult to characterize manually at scale. The Electron

FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models

SafetyDGX agent

arXiv:2510.22827v3 Announce Type: replace-cross Abstract: Evaluating text-to-image (T2I) systems requires judging not only whether an image matches a prompt, but also whether socially salient attribut

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

Model ReleasesDGX agent

arXiv:2605.25052v1 Announce Type: new Abstract: Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these t

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

SafetyDGX agent

arXiv:2605.25447v1 Announce Type: new Abstract: Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-pu

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

Model ReleasesDGX agent

arXiv:2605.24636v1 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios re

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

Model ReleasesDGX agent

arXiv:2507.19219v2 Announce Type: replace Abstract: Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalan

How we evolved Google’s global and data center networks for the AI era

Model ReleasesDGX agent

Over the last 25 years of building Google’s global network, we’ve navigated major architectural eras — from the Internet, to streaming, and the cloud. Today, we are squarely in the midst of a fourth:

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

SafetyDGX agent

arXiv:2605.24883v1 Announce Type: new Abstract: The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on con

Iterate Until Retrieved: Factual Nugget Optimization for Discoverable Continual Corrections in Agentic RAG

AgentsDGX agent

arXiv:2605.25641v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) systems in complex B2B (business-to-business) settings may often receive free-form response feedback. Rathe

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

SafetyDGX agent

arXiv:2605.23997v1 Announce Type: cross Abstract: Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they

Lake Detection and Water Quality Estimation in Sentinel-2 Data

ApplicationsDGX agent

arXiv:2605.24515v1 Announce Type: new Abstract: With climate change and increasing human pressure on natural landscapes, inland water resources are becoming progressively scarcer, more vulnerable, and

Leveraging Spreading Activation for Improved Document Retrieval in Knowledge-Graph-Based RAG Systems

TutorialsDGX agent

arXiv:2512.15922v3 Announce Type: replace Abstract: Despite initial successes and a variety of architectures, retrieval-augmented generation systems still struggle to reliably retrieve and connect the

Memory-Induced Tool-Drift in LLM Agents

Model ReleasesDGX agent

arXiv:2605.24941v1 Announce Type: cross Abstract: Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinn

Meta-Engineering Harnesses for AI-Native Software Production: A Contract-Driven Adversarial Verification Architecture with Early Deployment Report

AgentsDGX agent

arXiv:2605.25665v1 Announce Type: cross Abstract: AI-native software development is often evaluated at the level of individual models, prompts, or generated artifacts. This framing is insufficient for

Methodology for Creating a Clinically Verified Dermoscopic Image Dataset

ResearchDGX agent

arXiv:2605.25168v1 Announce Type: cross Abstract: This study presents a methodology for constructing a clinically verified dataset of dermatoscopic images for medical informatics research. The relevan

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Model ReleasesDGX agent

arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, h

← Previous
1…6364656667…83
Next →