AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “automated”

GridTimelineEvolution
4,980 results
1 Jul 2026

Stealthy Multi-Task Adversarial Attacks

SafetyDGX agent

arXiv:2411.17936v2 Announce Type: replace-cross Abstract: Deep neural networks are highly vulnerable to adversarial perturbations, raising serious safety concerns in the real-world systems. While prio

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability

ResearchDGX agent

arXiv:2601.18778v3 Announce Type: replace-cross Abstract: RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal. We investigat

Test-Time Verification for Text-to-SQL via Outcome Reward Models

SafetyDGX agent

arXiv:2606.30851v1 Announce Type: cross Abstract: Improving the reliability of large language models (LLMs) at inference time is a central challenge in structured reasoning tasks such as Text-to-SQL.


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed Claims

Model ReleasesDGX agent

arXiv:2606.31273v1 Announce Type: new Abstract: AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce

TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data

ResearchDGX agent

arXiv:2601.00260v2 Announce Type: replace Abstract: While foundation models in radiology are expected to be applied to various clinical tasks, computational cost constraints remain a major challenge w

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

Model ReleasesDGX agent

arXiv:2606.30755v1 Announce Type: cross Abstract: Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on

30 Jun 2026

Accelerating scientific discovery with Co-Scientist

Model ReleasesDGX agent

arXiv:2502.18864v2 Announce Type: replace Abstract: Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation. To augm

An expanded Vercel Agent: chat, investigations, and approved actions, now in public beta

AgentsDGX agent

Vercel has released an expanded version of its Vercel Agent in public beta, introducing new capabilities including chat functionality, investigation tools, and an approved actions feature. The update

Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems

ApplicationsDGX agent

arXiv:2510.20640v2 Announce Type: replace Abstract: In this paper, we present DiRecGNN, an attention-enhanced entity recommendation framework for monitoring cloud services at Microsoft. We provide ins

AutoB2G: Agentic Simulation and Reinforcement Learning for Spatio-Temporal Grid-Interactive Building Control

AgentsDGX agent

arXiv:2603.26005v2 Announce Type: replace Abstract: Grid-interactive building control has emerged as a promising approach for improving demand-side flexibility in modern power systems. Realistic studi

Bricker to BRACE: A Bracket Exposure RAW Dataset and Restoration Model for Flicker-Banding

ResearchDGX agent

arXiv:2606.29845v1 Announce Type: new Abstract: Flicker-banding (FB), arises from temporal aliasing between a camera's rolling shutter and a display's brightness modulation, degrading screen-captured

Build agents even faster with Gemini Enterprise Agent Platform’s fully-managed, remote MCP server

Model ReleasesDGX agent

A couple of months ago, we announced that over 50 Google-managed MCP servers are available. Today, we’ll dive into how to use the Gemini Enterprise Agent Platform remote MCP server to securely connect

Carolina Guide: A Multi-Agent RAG System with Institutional Guardrails for Academic Policy Assistance

SafetyDGX agent

arXiv:2606.28360v1 Announce Type: cross Abstract: University students often struggle to navigate complex academic policies, leading to advising bottlenecks and delayed access to critical information.

CaveAgent: Transforming LLMs into Stateful Runtime Operators

AgentsDGX agent

arXiv:2601.01569v4 Announce Type: replace Abstract: LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradigms that s

CellDETR: A Detection-Guided Framework for Scalable Cell Representation Learning from Histopathology Images

ResearchDGX agent

arXiv:2606.29463v1 Announce Type: new Abstract: Recent advances in pathology foundation models have substantially improved patch and slide level representation learning from whole-slide images (WSIs).

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation

AgentsDGX agent

arXiv:2602.12089v3 Announce Type: replace-cross Abstract: As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove bot

Claude Sonnet 5 is now available in Devin Desktop and Devin CLI. Sonnet 5 pairs frontier-level coding performance with a more affordable pri…

Model ReleasesDGX agent

Claude Sonnet 5 has been integrated into Devin Desktop and Devin CLI, offering advanced coding capabilities at a more competitive price point than previous models. This release represents an update to

COHORT: Collaborative Orchestration for Hardening via Offensive Replay on Emulated Topologies

Model ReleasesDGX agent

arXiv:2606.30479v1 Announce Type: cross Abstract: Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to that adver

CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph

Model ReleasesDGX agent

arXiv:2606.30175v1 Announce Type: new Abstract: The continuous evolution of large language models drives escalating demands on data scale and quality, and as different training stages impose increasin

Cybersecurity is the True Frontier for Generative AI Success or Failure

ResearchDGX agent

arXiv:2606.28929v1 Announce Type: cross Abstract: Cybersecurity is a real-life test-bed for many machine learning problems at once, especially when considering modern strides in using Large Language M

CytoCLIP: Learning Cytoarchitectural Characteristics in Developing Human Brain Using Contrastive Language Image Pre-Training

TutorialsDGX agent

arXiv:2601.12282v2 Announce Type: replace-cross Abstract: The functions of different regions of the human brain are closely linked to their distinct cytoarchitecture, which is defined by the spatial a

Digitizing Coaching Intelligence: An Agentic Framework for Holistic Athlete Profiling using VLM and RAG

Model ReleasesDGX agent

arXiv:2606.28570v1 Announce Type: cross Abstract: Athlete assessment is a critical process for tracking physical progress and identifying elite talent. However, during mass recruitment drives, traditi

Evolutionary Hyperparameter Optimization to Find Lightweight CNN Models for Autonomous Steering

AgentsDGX agent

arXiv:2606.29684v1 Announce Type: cross Abstract: This research investigates the optimization of Convolutional and Dense Neural Networks (CNNs and DNNs) for autonomous steering using the (N+M) Evoluti

Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving

TutorialsDGX agent

arXiv:2606.29493v1 Announce Type: new Abstract: Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-check

Few-class Fidelity: Evaluating Explanations of Real-conditions CNN classifiers with Optimized Perturbations

Local AiDGX agent

arXiv:2606.28391v1 Announce Type: cross Abstract: The wide use of Convolutional Neural Networks (CNN) in numerous domains and real-world classification applications is justified by their high precisio

FLAME 3 Dataset: Unleashing the Power of Radiometric Thermal UAV Imagery for Wildfire Management

ResearchDGX agent

arXiv:2412.02831v2 Announce Type: replace-cross Abstract: The increasing accessibility of radiometric thermal imaging sensors for unmanned aerial vehicles (UAVs) offers significant potential for advan

Implementation of Hyperelastic Physics-Augmented Neural Networks in the Explicit Finite Element Codes Simcenter Radioss and OpenRadioss with Applications to Impact Events

ResearchDGX agent

arXiv:2606.29874v1 Announce Type: cross Abstract: Data-driven material modeling techniques have gained significant attention due to their ability to capture complex constitutive behaviors beyond the l

It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents

Model ReleasesDGX agent

arXiv:2606.27944v1 Announce Type: cross Abstract: Phone-use Agents can execute complex tasks end to end across real mobile applications. By operating a real device on the user's behalf, they reach far

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

SafetyDGX agent

arXiv:2606.30642v1 Announce Type: cross Abstract: Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts.

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue

ResearchDGX agent

arXiv:2509.02292v2 Announce Type: replace Abstract: What if large language models could not only infer human mindsets but also expose every blind spot in team dialogue such as discrepancies in the tea

MIThinker: A Plug-and-Play Policy-Optimized Thinker For Motivational Interviewing Counseling

SafetyDGX agent

arXiv:2606.29265v1 Announce Type: new Abstract: Reasoning large language models (LLMs) have recently made much progress in complex problem-solving, leveraging internal reasoning (or thought) to guide

ML-Powered LDAP Reconnaissance Detection using Weak Supervision

ResearchDGX agent

arXiv:2606.28917v1 Announce Type: new Abstract: Lightweight Directory Access Protocol (LDAP) is a protocol that allows users to query and modify Active Directory (AD) data. By default, all users have

Multi-Agentic System Leveraging Open-Source LLMs to Mitigate Disinformation Threats

Model ReleasesDGX agent

arXiv:2606.30259v1 Announce Type: new Abstract: In contemporary societies, the threat of disinformation has reached alarming levels, exacerbated by the proliferation of electronic communication, socia

PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science

Model ReleasesDGX agent

arXiv:2508.17117v3 Announce Type: replace-cross Abstract: Existing plant-disease datasets target classification and detection, leaving vision-language models unable to support interactive, reasoning-b

Project gallery and whitepaper: https://research.nvidia.com/labs/gear/aspire/ ASPIRE is a great collaboration between NVIDIA GEAR lab, UMich…

HardwareDGX agent

Project gallery and whitepaper: https://research.nvidia.com/labs/gear/aspire/ ASPIRE is a great collaboration between NVIDIA GEAR lab, UMich, Berkeley, and CMU. Kudos to all the coauthors who pour the

Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

Model ReleasesDGX agent

arXiv:2606.30294v1 Announce Type: new Abstract: Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch the correspo

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought

SafetyDGX agent

arXiv:2606.15753v2 Announce Type: replace Abstract: Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding th

SAT-RTS: A systematic framework for tactical knowledge extraction and visualization-based analysis in real-time strategy games

ResearchDGX agent

arXiv:2606.30090v1 Announce Type: new Abstract: Efficient tactical knowledge extraction and analysis in real-time strategy (RTS) games micromanagement are constrained by the high-dimensional coupled s

Self-Supervised Calibration of Scientific Instruments Using Physical Consistency Constraints

AgentsDGX agent

arXiv:2606.29466v1 Announce Type: new Abstract: Calibration remains one of the principal obstacles to the deployment of machine learning in scientific instrumentation because it typically relies on ex

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

Model ReleasesDGX agent

arXiv:2606.29955v1 Announce Type: cross Abstract: Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks

Supporting Workflow Reproducibility by Linking Bioinformatics Tools across Papers and Executable Code

ResearchDGX agent

arXiv:2603.08195v2 Announce Type: replace Abstract: Motivation: The rapid growth of biological data has intensified the need for transparent, reproducible, and well-documented computational workflows.

The gap in autonomous agentic loops that gets ignored: agents can plan and call APIs but can't acquire tools they don't have access to. x402…

AgentsDGX agent

The gap in autonomous agentic loops that gets ignored: agents can plan and call APIs but can't acquire tools they don't have access to. x402 + Apify's 20,000+ Actors is a concrete fix for that. Worth

Towards Improved Anomaly Detection for Cloud Cybersecurity via Graph Neural Networks

ApplicationsDGX agent

arXiv:2606.28923v1 Announce Type: new Abstract: Detecting security threats in an organization's cloud computing environment has become necessary due to the increased reliance on cloud infrastructure.

TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation

SafetyDGX agent

arXiv:2606.29097v1 Announce Type: new Abstract: Recent research has investigated the use of large language models (LLMs) to generate traffic scenarios for autonomous driving. However, pretrained LLMs

ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models

Model ReleasesDGX agent

arXiv:2606.28804v1 Announce Type: new Abstract: Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluat

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs

Model ReleasesDGX agent

arXiv:2606.28438v1 Announce Type: cross Abstract: Recursive self-training can degrade neural generative models when generated data is reused without fresh human data or external quality control. We st

29 Jun 2026

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

SafetyDGX agent

arXiv:2606.28270v1 Announce Type: new Abstract: The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has funda

Boris sat down with Spotify VP of Engineering Niklas Gustavsson. Spotify ships 4,500 production deploys a day, and 73% of PRs are now AI-ass…

ApplicationsDGX agent

Spotify's VP of Engineering Niklas Gustavsson discussed the company's deployment practices, revealing they execute 4,500 production deploys daily and that 73% of pull requests now utilize AI assistanc

Building a Scalable, Reproducible, Evaluatable, and Closed-Loop Simulation Environment Foundation for Embodied Intelligence Cloud-Native Simulation Infrastructure for Embodied Intelligence Training, Evaluation, and Data Collection

Model ReleasesDGX agent

arXiv:2606.27962v1 Announce Type: new Abstract: This paper presents a cloud-native simulation infrastructure framework for embodied intelligence that supports large-scale training, standardized evalua

COOPA: A Modular LLM Agent Architecture for Operations Research Problems

AgentsDGX agent

arXiv:2606.27611v1 Announce Type: new Abstract: Operations Research (OR) provides a rigorous framework for high-stakes decision-making, but effective OR modeling requires substantial domain knowledge,

Drifting in the Future: Stabilizing Path Following Drifting on High-Latency Vehicle Systems

SafetyDGX agent

arXiv:2606.27914v1 Announce Type: new Abstract: Autonomously controlling and handling a vehicle at and beyond its stability limit is a mathematically and computationally demanding task. Prior demonstr

Forecasting Technological Directions in Wireless Networks and Mobile Computing via AutoML Framework

ResearchDGX agent

arXiv:2606.27394v1 Announce Type: cross Abstract: The exponential increase in scientific publications has driven the emergence of new trends. Accurate forecasting of these developments is essential fo

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

Model ReleasesDGX agent

arXiv:2601.17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal o

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications

ApplicationsDGX agent

arXiv:2606.28070v1 Announce Type: new Abstract: JD.com, one of the world's largest e-commerce platforms, serves over 700 million active users and millions of merchants, with a catalog of tens of billi

Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents

Model ReleasesDGX agent

arXiv:2606.27595v1 Announce Type: new Abstract: Web-agent benchmarks overwhelmingly measure depth -- pinning one obscure answer behind a chain of constraints -- while breadth, exhaustively enumerating

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

Model ReleasesDGX agent

arXiv:2606.27537v1 Announce Type: new Abstract: Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most ass

Physics-Guided Robotic Radiation Source Localization along Arbitrary Measurement Paths in Unstructured Environments

Local AiDGX agent

arXiv:2606.27624v1 Announce Type: cross Abstract: Using robots to estimate the location of the radiation source is an effective way to improve efficiency and safety. Existing methods focus on planning

Scaling Network Analysis for Fraud Prevention with BigQuery Graph

IndustryDGX agent

Based in the UK, Curve are building a financial super-app, a smart wallet that consolidates all your debit and credit cards into a single app and card, simplifying how millions of users spend, send an

ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation

SafetyDGX agent

arXiv:2606.27736v1 Announce Type: new Abstract: The rapid spread of fake news poses increasing threats to information ecosystems, especially as AI-generated misinformation under Generative Engine Opti

Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors

ResearchDGX agent

arXiv:2606.28237v1 Announce Type: new Abstract: Quadruped robots have achieved remarkable locomotion, yet their behavioral repertoire remains confined to a few gaits--far from the expressive, companio

← Previous
1…5455565758…83
Next →