AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “automated”

GridTimelineEvolution
4,980 results
9 Jun 2026

A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline

AgentsDGX agent

arXiv:2606.07718v1 Announce Type: new Abstract: Agentic AI tools offer a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that ta

LATTEArena: An Evaluation Framework for LLM-powered Tabular Feature Engineering (Extended Version)

ResearchDGX agent

arXiv:2606.09004v1 Announce Type: new Abstract: Feature engineering remains essential for tabular data analysis, and Large Language Models (LLMs) have emerged as a promising paradigm for automating th

Real-Time Industrial Defect Detection on Edge Hardware Using Fine-Tuned YOLOv8: A Systematic Benchmark on the NEU Surface Defect Database and MVTec AD with Automotive & Battery Manufacturing Extensions


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2606.07659v1 Announce Type: new Abstract: Automated surface defect detection is critical for ensuring rigorous quality control in high-speed manufacturing environments. While deep learning model

When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents

Model ReleasesDGX agent

arXiv:2602.08235v2 Announce Type: replace-cross Abstract: Although computer-use agents (CUAs) hold significant potential to automate increasingly complex OS workflows, they can demonstrate unsafe unin

8 Jun 2026

ReclAIm: A Multi-Agent Framework for Monitoring and Correcting Performance Decline in Medical Imaging AI

Model ReleasesDGX agent

arXiv:2510.17004v2 Announce Type: replace-cross Abstract: Purpose: To develop and evaluate a multi-agent framework (ReclAIm) for automated monitoring, detection, and correction of performance decline

4 Jun 2026

What's new for Managed Service for Apache Spark clusters

Model ReleasesDGX agent

At Google Cloud, our goal is to let you run large-scale analytical and data science workloads with maximum efficiency so you can process big data pipelines, machine learning, and ETL tasks. We recentl

3 Jun 2026

SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model

ResearchDGX agent

arXiv:2603.26738v3 Announce Type: replace-cross Abstract: While automated sleep staging has achieved expert-level accuracy, its clinical adoption is hindered by a lack of auditable reasoning. We intro

2 Jun 2026

ANDES: Agent Native Data Evolving Synthesis Tool for Autonomous Instruction Alignment

SafetyDGX agent

arXiv:2606.01279v1 Announce Type: new Abstract: AI agents are increasingly being tasked with automating AI research itself, particularly the critical post-training phase that transforms base LLMs into

MiCU: End-to-End Smart Home Command Understanding with Large Language Model

ApplicationsDGX agent

arXiv:2606.01099v1 Announce Type: cross Abstract: Command understanding systems in smart home ecosystems can automate device control and substantially improve user experience. However, while they perf

VESTA: Visual Exploration with Statistical Tool Agents

Model ReleasesDGX agent

arXiv:2606.00384v1 Announce Type: new Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems lev

1 Jun 2026

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

SafetyDGX agent

arXiv:2605.31381v1 Announce Type: new Abstract: We evaluate the consistency of automated judges in conducting a multi-dimensional safety evaluation in a reference-free setup. Our results indicate that

28 May 2026

An LLM-Based Assistance System for Intuitive and Flexible Capability-Based Planning

AgentsDGX agent

arXiv:2605.28666v1 Announce Type: new Abstract: In modern industry, dynamic environments and the complexity of modular and reconfigurable resources require automated planning of process sequences. Cap

NVIDIA Research Advances Robotics From Simulation to the Real World

HardwareDGX agent

Robotics is entering a new phase: moving from controlled demos and scripted automation toward generalizable, reliable embodied autonomy in the real world. At the International Conference on Robotics a

27 May 2026

Benchmarking Convolutional, Transformer, Hybrid, and Vision Language Models for Multi Disease Retinal Screening

Model ReleasesDGX agent

arXiv:2605.26283v1 Announce Type: new Abstract: Modern deep learning offers powerful tools for automated retinal screening, but it remains unclear how different visual model families compare in realis

Improving your agent has been a manual process of: ✅ Reading traces ✅ Looking for patterns ✅ Writing evals ✅ Creating fixes Now, LangSmith E…

AgentsDGX agent

LangSmith has introduced automated tools to streamline agent improvement, eliminating the manual workflow of reading execution traces, identifying patterns, writing evaluations, and implementing fixes

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

Model ReleasesDGX agent

arXiv:2605.26730v1 Announce Type: new Abstract: The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automate

26 May 2026

AutoSG: LLM-Driven Solver Generation Solely from Task Prompts for Expensive Optimization

Local AiDGX agent

arXiv:2605.25658v1 Announce Type: cross Abstract: Expensive optimization tasks are ubiquitous in real-world applications, demanding highly specialized solvers. While LLM-driven automated solver genera

Catching MRI outliers: unsupervised detection and localization of MRI artefacts and clinical anomalies using deep learning

Local AiDGX agent

arXiv:2605.24609v1 Announce Type: cross Abstract: Artificial intelligence is increasingly integrated into radiotherapy workflows, yet such pipelines remain vulnerable to out-of-distribution image data

Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence

ResearchDGX agent

arXiv:2509.23573v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to help security analysts manage the surge of cyber threats, automating tasks from vulnerab

22 May 2026

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

Model ReleasesDGX agent

arXiv:2605.22200v1 Announce Type: new Abstract: Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment ho

20 May 2026

CircuitHub raises $28M to scale electronics production in days rather than months

ApplicationsDGX agent

CircuitHub Inc., a company building automated manufacturing and assembly systems for electronics hardware at scale that supports industries such as self-driving cars and satellites, today announced it

Geo-Data-Driven HD Map Generation Workflow with Integrated Reference-Free Constraint-Based Verification

ApplicationsDGX agent

arXiv:2605.18921v1 Announce Type: new Abstract: High-definition (HD) maps are core artifacts for automated driving systems, but their generation commonly relies on sensor-intensive mobile mapping camp

19 May 2026

AI for Auto-Research: Roadmap & User Guide

Model ReleasesDGX agent

arXiv:2605.18661v1 Announce Type: new Abstract: AI-assisted research is crossing a threshold: fully automated systems can now generate research papers for as little as $15, while long-horizon agents c

Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech

ApplicationsDGX agent

arXiv:2605.16280v1 Announce Type: cross Abstract: Automating legal reasoning forces a choice between imperfect alternatives: symbolic systems offer transparency but struggle with ambiguity, whereas ne

Lean Meets Theoretical Computer Science: Scalable Synthesis of Theorem Proving Challenges in Formal-Informal Pairs

ResearchDGX agent

arXiv:2508.15878v2 Announce Type: replace-cross Abstract: Formal theorem proving (FTP) has emerged as a critical foundation for evaluating the reasoning capabilities of large language models, enabling

MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair

AgentsDGX agent

arXiv:2605.17444v1 Announce Type: cross Abstract: Modern software ecosystems face a rapidly growing number of disclosed vulnerabilities, increasing the need for automated repair techniques that can op

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility

Model ReleasesDGX agent

arXiv:2605.16616v1 Announce Type: new Abstract: Autonomous research systems capable of generating complete scientific manuscripts have advanced rapidly, yet robust and realistic evaluation frameworks

TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT

Model ReleasesDGX agent

arXiv:2605.16572v1 Announce Type: new Abstract: Automated segmentation of liver lesions on non-contrast computed tomography (NCCT) is clinically important but fundamentally challenging, particularly i

18 May 2026

Surrogate Neural Architecture Codesign Package (SNAC-Pack)

Local AiDGX agent

arXiv:2605.16138v1 Announce Type: cross Abstract: Neural architecture search (NAS) is a powerful approach for automating model design, but existing methods often optimize for accuracy alone or rely on

17 May 2026

WebWright - Agentic Extension for your Browser is now live on Chrome 😁

AgentsDGX agent

WebWright is a Chrome browser extension that enables AI agents to automate web tasks directly within the browser, leveraging local models like those available through Ollama. The announcement on r/oll

15 May 2026

CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation

Model ReleasesDGX agent

arXiv:2602.20571v2 Announce Type: replace Abstract: Many benchmarks for automated causal inference evaluate a system's performance based on a single numerical output, such as an Average Treatment Effe

How business operations teams use Codex

Model ReleasesDGX agent

This document from OpenAI describes how business operations teams leverage Codex, OpenAI's code generation model, to automate and streamline their workflows. It likely covers practical applications su

SkillFlow: Flow-Driven Recursive Skill Evolution for Agentic Orchestration

SafetyDGX agent

arXiv:2605.14089v1 Announce Type: new Abstract: In recent years, a variety of powerful LLM-based agentic systems have been applied to automate complex tasks through task orchestration. However, existi

14 May 2026

Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics

Model ReleasesDGX agent

arXiv:2605.13171v1 Announce Type: new Abstract: As automated reasoning systems advance rapidly, there is a growing need for research-level formal mathematical problems to accurately evaluate their cap

VERA-MH Concept Paper

Model ReleasesDGX agent

arXiv:2510.15297v4 Announce Type: replace-cross Abstract: We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in

13 May 2026

Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation

SafetyDGX agent

arXiv:2605.11730v1 Announce Type: new Abstract: Automated red-teaming for LLMs often discovers narrow attack slices, missing diverse real-world threats, and yielding insufficient data for safety fine-

Read, Extract, Classify: A Tool for Smarter Requirements Engineering

ResearchDGX agent

arXiv:2605.11045v1 Announce Type: cross Abstract: This paper presents the ReXCL tool, which automates the extraction and classification processes in requirements engineering, enhancing the software de

12 May 2026

CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs

Model ReleasesDGX agent

arXiv:2605.08467v1 Announce Type: new Abstract: Large language models show promise for automated CUDA programming, however even the strongest coding models (e.g., Claude-Opus-4.6) may still fall short

How finance teams use Codex

ApplicationsDGX agent

This OpenAI Academy article explores how finance teams leverage Codex, OpenAI's code generation model, to automate financial analysis, reporting, and data processing tasks. It likely covers practical

LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges

HardwareDGX agent

arXiv:2605.10807v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) and hardware security is rapidly reshaping the semiconductor i

Reliable LLM-Based Edge-Cloud-Expert Cascades for Telecom Knowledge Systems

Model ReleasesDGX agent

arXiv:2512.20012v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are emerging as key enablers of automation in domains such as telecommunications, assisting with tasks including

11 May 2026

LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning

Local AiDGX agent

arXiv:2605.07505v1 Announce Type: new Abstract: Developing lightweight, on-device vision-language GUI agents is essential for efficient cross-platform automated interaction. However, current on-device

@trycua https://hermes-agent.nousresearch.com/docs/user-guide/features/computer-use

AgentsDGX agent

This documentation page covers the computer-use feature in Hermes Agent, detailing how the AI system can interact with computer interfaces and perform automated tasks through screen interaction and in

6 May 2026

A Neuro-Symbolic Framework for Accountability in Public-Sector AI

ApplicationsDGX agent

arXiv:2512.12109v3 Announce Type: replace-cross Abstract: Automated eligibility systems increasingly determine access to essential public benefits, but the explanations they generate often fail to ref

An Empirical Study of Agent Skills for Healthcare: Practice, Gaps, and Governance

Local AiDGX agent

arXiv:2605.02709v1 Announce Type: new Abstract: Healthcare automation is shaped by local procedures and organizational constraints, so agent capabilities rarely transfer unchanged across settings. Age

arXiv Papers → LLM Artifacts This is how I keep up with AI research now. It's like having access to the most personalized arXiv feed. Automa…

AgentsDGX agent

arXiv Papers → LLM Artifacts This is how I keep up with AI research now. It's like having access to the most personalized arXiv feed. Automations run everyday to curate papers based a set of rules and

Conventional Commit Classification using Large Language Models and Prompt Engineering

Model ReleasesDGX agent

arXiv:2605.02033v1 Announce Type: cross Abstract: Conventional commits provide a structured format for writing commit messages, which improves readability, software maintenance, and enables automation

Fully Automatic Trace Gas Plume Detection

ResearchDGX agent

arXiv:2605.03372v1 Announce Type: new Abstract: Future imaging spectrometers will increase data volumes by orders of magnitude, requiring automated detection of trace gas point sources. We present a f

LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation

Model ReleasesDGX agent

arXiv:2605.01394v1 Announce Type: cross Abstract: Formal specification is essential for rigorous program verification, yet writing correct specifications remains costly and difficult to automate. Alth

MILD: Mediator Agent System with Bidirectional Perception and Multi-Layered Alignment for Human-Vehicle Collaboration

SafetyDGX agent

arXiv:2605.01507v1 Announce Type: new Abstract: Prior studies report that partial driving automation can increase the cognitive demands on human drivers. This effect largely arises from human drivers'

PERSA: Reinforcement Learning for Professor-Style Personalized Feedback with LLMs

Model ReleasesDGX agent

arXiv:2605.01123v1 Announce Type: new Abstract: Large language models (LLMs) can provide automated feedback in educational settings, but aligning an LLMs style with a specific instructors tone while m

4 May 2026

Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments

ResearchDGX agent

arXiv:2505.09901v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate or automate human behavior in complex sequential decision-making settings. A na

Import AI 455: AI systems are about to start building themselves.

SafetyDGX agent

This newsletter entry discusses advancements in automating AI research and development processes, exploring how artificial intelligence systems are becoming capable of autonomously designing and impro

Scaling data and AI with Managed Service for Apache Airflow

Model ReleasesDGX agent

Orchestration is no longer just about moving data; it is about governing enterprise intelligence. To reflect our deep commitment to and embrace of open-source software, we shared earlier that Cloud Co

3 May 2026

The number of jobs in the future is endless because the problems to solve are endless. Jobs multiply as we get more complex. No AI or human …

HardwareDGX agent

The number of jobs in the future is endless because the problems to solve are endless. Jobs multiply as we get more complex. No AI or human can solve all problems and all the work to do in the Univers

30 Apr 2026

Cuando sucede un problema en un puesto automatizado por la IA... ¿Quién es el responsable?

SafetyDGX agent

This post discusses the liability and accountability questions that arise when problems occur in AI-automated workplaces, exploring who bears responsibility—whether the AI developer, the employer, the

29 Apr 2026

From six months to four days: Inside how AI slashed wait times for essential autism care

IndustryDGX agent

Organizations are shifting away from broad digital transformation efforts and instead targeting high-impact bottlenecks where AI and automation can deliver immediate results. These targeted changes ar

SciDER: Scientific Data-centric End-to-end Researcher

AgentsDGX agent

arXiv:2603.01421v2 Announce Type: replace-cross Abstract: Automated scientific discovery with large language models is transforming the research lifecycle from ideation to experimentation, yet existin

28 Apr 2026

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics

Model ReleasesDGX agent

arXiv:2509.20374v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong performance across general NLP tasks, but their utility in automating numerical experime

Quantifying and Mitigating Self-Preference Bias of LLM Judges

SafetyDGX agent

arXiv:2604.22891v1 Announce Type: cross Abstract: LLM-as-a-Judge has become a dominant approach in automated evaluation systems, playing critical roles in model alignment, leaderboard construction, qu

← Previous
1…1617181920…83
Next →