AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,088 results
19 May 2026

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

Model ReleasesDGX agent

arXiv:2605.16679v1 Announce Type: cross Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions

DeepArrhythmia: Segment-Contextualized ECG Arrhythmia Classification via Selective Evidence Acquisition

AgentsDGX agent

arXiv:2605.16441v1 Announce Type: cross Abstract: Beat-level Electrocardiography (ECG) arrhythmia detection aims to assign an arrhythmia class to each beat in a recording, yet many existing systems tr

Detecting Verbatim LLM Copy-Paste in Homework

Model ReleasesDGX agent

arXiv:2605.16336v1 Announce Type: cross Abstract: Large language models (LLMs) have made fluent essay writing, code drafting, and quiz answering instantly available to students at every level, from se

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Evidence of a Cognitive Shift in AI Education: How Students Are Rethinking Human Intelligence?

TutorialsDGX agent

arXiv:2605.16292v1 Announce Type: cross Abstract: Perceptions of intelligence shape how learners evaluate and rely on artificial intelligence (AI) systems. Despite rapid advances in AI capabilities, t

llm-gemini 0.32a0

Model ReleasesDGX agent

I don't have current information about this specific entry, so I'll describe what it likely covers based on the available details. This entry documents the release or update of llm-gemini version 0.32

Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework

AgentsDGX agent

arXiv:2605.16821v1 Announce Type: new Abstract: The rapid evolution of Large Language Model (LLM) agents has produced diverse interaction paradigms, yet few production systems integrate multiple parad

OpenJarvis: Personal AI, On Personal Devices

Model ReleasesDGX agent

arXiv:2605.17172v1 Announce Type: cross Abstract: Personal AI stacks, like OpenClaw and Hermes Agent, are becoming central to daily work, yet they route nearly every query (often over sensitive local

Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review

SafetyDGX agent

arXiv:2605.17548v1 Announce Type: cross Abstract: Code review has evolved for decades, from informal peer checking to today's pull request (PR) workflows, yet it remains a largely manual, uneven, and

SEDD: Scalable and Efficient Dataset Deduplication with GPUs

Model ReleasesDGX agent

arXiv:2501.01046v4 Announce Type: replace Abstract: Dataset deduplication is widely recognized as a crucial preprocessing step that enhances data quality and improves the performance of large language

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain

Model ReleasesDGX agent

arXiv:2605.17946v1 Announce Type: new Abstract: Multimodal large language models are increasingly used as agent backbones that understand multimodal inputs, plan retrieval actions, invoke external too

TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks

HardwareDGX agent

arXiv:2605.17170v1 Announce Type: new Abstract: Agentic workloads have emerged as a major workload for LLM inference. They differ significantly from chat-only workloads, requiring long-context process

Your SaaS Is an Insurance Product: A Modeling Framework

Model ReleasesDGX agent

arXiv:2605.16699v1 Announce Type: new Abstract: Capped-usage SaaS products -- LLM subscriptions such as Claude Code and ChatGPT, cloud platforms such as Vercel and Cloudflare Workers, corporate benefi

18 May 2026

Context, Reasoning, and Hierarchy: A Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP

AgentsDGX agent

arXiv:2605.16205v1 Announce Type: new Abstract: Deploying compound LLM agents in adversarial, partially observable sequential environments requires navigating several design dimensions: (1) what the a

How Data Augmentation Shapes Neural Representations

ResearchDGX agent

arXiv:2605.15306v1 Announce Type: new Abstract: Data augmentation is widely recognized for improving generalization in deep networks, yet its impact on the geometry of learned representations remains

How Google Does It: Fleet-wide, large-scale A/B experimentation

SafetyDGX agent

When most people think of A/B experimentation, they think of button colors, landing page layouts, or checkout flows. At Google, many fundamental infrastructure improvements also need the rigor of A/B

Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data

AgentsDGX agent

arXiv:2509.21465v3 Announce Type: replace Abstract: Tabular foundation models are becoming increasingly popular for low-resource tabular problems. These models make up for small training datasets by p

15 May 2026

From User Preferences to Base Score Extraction Functions in Gradual Argumentation (with Appendix)

ResearchDGX agent

arXiv:2602.14674v4 Announce Type: replace Abstract: Gradual argumentation is a field of symbolic AI which is attracting attention for its ability to support transparent and contestable AI systems. It

Gemini Live Agent Challenge: Announcing the winners and highlights

Model ReleasesDGX agent

The Gemini Live Agent Challenge is officially in the books! We challenged developers worldwide to break out of the traditional 'text box' paradigm by building next-generation AI agents. From our initi

MicroscopyMatching: Towards a Ready-to-use Framework for Microscopy Image Analysis in Diverse Conditions

ResearchDGX agent

arXiv:2605.14980v1 Announce Type: cross Abstract: Analyzing microscopy images to extract biological object properties (e.g., their morphological organization, temporal dynamics, and population density

Near-Miss: Latent Policy Failure Detection in Agentic Workflows

Model ReleasesDGX agent

arXiv:2603.29665v2 Announce Type: replace Abstract: Agentic systems for business process automation often require compliance with policies governing conditional updates to the system state. Evaluation

OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling

Model ReleasesDGX agent

arXiv:2601.19924v2 Announce Type: replace-cross Abstract: We investigate the capabilities and scalability of Large Language Models (LLMs) in optimization modeling, a domain requiring structured reason

PALMS: A Computational Implementation for Pavlovian Associative Learning Models' Simulation

ResearchDGX agent

arXiv:2602.07519v3 Announce Type: replace Abstract: In contrast to static formalisms, computational definitions describe the operational mechanisms of a model. Simulations are an essential part of the

Probabilistic Verification of Recurrent Neural Networks for Single and Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.14758v1 Announce Type: new Abstract: History-dependent policies induced by recurrent neural networks (RNNs) rely on latent hidden state dynamics, making verification in partially observable

SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks

Model ReleasesDGX agent

arXiv:2605.14051v1 Announce Type: new Abstract: Industrial LLM agent systems often separate planning from execution, yet LLM planners frequently produce structurally invalid or unnecessarily long work

Web Agents Should Adopt the Plan-Then-Execute Paradigm

Model ReleasesDGX agent

arXiv:2605.14290v1 Announce Type: cross Abstract: ReAct has become the default architecture across LLM agents, and many existing web agents follow this paradigm. We argue that it is the wrong default

Welcome to BlackFile: Inside a Vishing Extortion Operation

Model ReleasesDGX agent

Written by: Austin Larsen, Tyler McLellan, Genevieve Stark, Dan Ebreo Introduction Google Threat Intelligence Group (GTIG) has continued to track an expansive extortion campaign by UNC6671, a threat a

14 May 2026

Cloud CISO Perspectives: How Google + Wiz changes multicloud strategy for CISOs

Model ReleasesDGX agent

Welcome to the first Cloud CISO Perspectives for May 2026. Today, Vinod D’Souza, director, Office of the CISO, shares highlights from his RSA Conference fireside chat with Anthony Belfiore, chief stra

Context Training with Active Information Seeking

ResearchDGX agent

arXiv:2605.13050v1 Announce Type: cross Abstract: Most existing large language models (LLMs) are expensive to adapt after deployment, especially when a task requires newly produced information or nich

Integration of an Agent Model into an Open Simulation Architecture for Scenario-Based Testing of Automated Vehicles

SafetyDGX agent

arXiv:2605.13539v1 Announce Type: new Abstract: Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are r

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

SafetyDGX agent

arXiv:2605.12729v1 Announce Type: cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), includ

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

Model ReleasesDGX agent

arXiv:2603.24649v2 Announce Type: replace Abstract: Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition.

xAI just released Grok Build CLI and it’s a game changer for developers Grok Build is a powerful AI coding agent and CLI built for professio…

Model ReleasesDGX agent

xAI just released Grok Build CLI and it’s a game changer for developers Grok Build is a powerful AI coding agent and CLI built for professional software engineering and complex coding workflows runnin

13 May 2026

ABRA: Agent Benchmark for Radiology Applications

Model ReleasesDGX agent

arXiv:2605.11224v1 Announce Type: new Abstract: Existing medical-agent benchmarks deliver imaging as pre-selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiolo

How Glance turns hours of video into mobile-ready clips with AI

Model ReleasesDGX agent

Every day, thousands of hours of new video content sits waiting to be discovered. Most of it lives in long-form, horizontal formats, while audiences are scrolling through vertical feeds on their phone

HTML Artifacts are a big part of how I work with agents now. Artifacts can be more than just static files. When combined with agents, they c…

Model ReleasesDGX agent

HTML Artifacts are a big part of how I work with agents now. Artifacts can be more than just static files. When combined with agents, they can take action or help you take action. This unlocks all kin

MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces

HardwareDGX agent

arXiv:2605.11333v1 Announce Type: cross Abstract: The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed ma

Posterior Contraction Rates for Sparse Kolmogorov-Arnold Networks in Anisotropic Besov Spaces

Model ReleasesDGX agent

arXiv:2605.11652v1 Announce Type: cross Abstract: We study posterior contraction rates for sparse Bayesian Kolmogorov-Arnold networks (KANs) over anisotropic Besov spaces, providing a statistical foun

Unpacking the Eye of the Beholder: Social Location, Identity, and the Moving Target of Political Perspectives

ResearchDGX agent

arXiv:2605.11166v1 Announce Type: new Abstract: Political and social identities structure how people evaluate political information, a finding decades deep in political science and routinely discarded

12 May 2026

A probabilistic framework for crystal structure denoising, phase classification, and order parameters

Local AiDGX agent

arXiv:2512.11077v3 Announce Type: replace-cross Abstract: Atomistic simulations generate large volumes of noisy structural data, yet extracting phase labels and continuous order parameters (OPs) in a

AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design

Model ReleasesDGX agent

arXiv:2605.08756v1 Announce Type: new Abstract: Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent works show t

AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care

SafetyDGX agent

arXiv:2605.08480v1 Announce Type: new Abstract: Individuals with Alzheimer's disease (AD) and Alzheimer's disease-related dementia (ADRD) experience memory and thinking changes that impact their abili

AI security startup Grego AI debuts, claims record $250,000 bounty for AI-found exploit

IndustryDGX agent

Artificial intelligence cybersecurity startup Grego AI formally launched today with a claimed method of using existing AI models to find critical software vulnerabilities that human auditors and other

Applying Graph Analysis for Unsupervised Fast Malware Fingerprinting

ResearchDGX agent

arXiv:2510.12811v2 Announce Type: replace-cross Abstract: Malware proliferation is increasing at a tremendous rate, with hundreds of thousands of new samples identified daily. Manual investigation of

ChladniSonify: A Visual-Acoustic Mapping Method for Chladni Patterns in New Media Art Creation

ResearchDGX agent

arXiv:2605.09846v1 Announce Type: cross Abstract: In new media art creation, the mapping between vision and hearing is often subjective. As a classic carrier of sound visualization, Chladni patterns h

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents

Model ReleasesDGX agent

arXiv:2605.09675v1 Announce Type: new Abstract: Clinical reasoning agents based on large language models (LLMs) aim to automate tasks such as intensive care unit (ICU) monitoring and patient state tra

CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs

Model ReleasesDGX agent

arXiv:2605.08467v1 Announce Type: new Abstract: Large language models show promise for automated CUDA programming, however even the strongest coding models (e.g., Claude-Opus-4.6) may still fall short

EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving

Model ReleasesDGX agent

arXiv:2605.10556v1 Announce Type: new Abstract: As large language models span dense, mixture-of-experts, and state-space architectures and are deployed on heterogeneous accelerators under increasingly

FairHealth: An Open-Source Python Library for Trustworthy Healthcare AI in Low-Resource Settings

SafetyDGX agent

arXiv:2605.08198v1 Announce Type: cross Abstract: We present FairHealth, an open-source Python library that provides a unified, modular framework for trustworthy machine learning in healthcare applica

FORTIS: Benchmarking Over-Privilege in Agent Skills

Model ReleasesDGX agent

arXiv:2605.09163v1 Announce Type: new Abstract: Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This

Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation

Model ReleasesDGX agent

arXiv:2601.11258v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) face the 'knowledge cutoff' challenge, where their frozen parametric memory prevents direct internalization of ne

MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents

ResearchDGX agent

arXiv:2603.13131v3 Announce Type: replace Abstract: Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A centra

MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring

Model ReleasesDGX agent

arXiv:2605.09684v1 Announce Type: cross Abstract: We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-eli

$NBIS announced a partnership with LangChain to integrate Nebius Token Factory with LangChain’s Deep Agents. The goal is to make it easier f…

Model ReleasesDGX agent

$NBIS announced a partnership with LangChain to integrate Nebius Token Factory with LangChain’s Deep Agents. The goal is to make it easier for teams building AI agents on LangChain to run those worklo

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

Model ReleasesDGX agent

arXiv:2605.08762v1 Announce Type: cross Abstract: Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start

Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark

Model ReleasesDGX agent

arXiv:2509.26574v4 Announce Type: replace Abstract: While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason

Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

Model ReleasesDGX agent

arXiv:2605.10848v1 Announce Type: cross Abstract: Does a lexical retriever suffice as large language models (LLMs) become more capable in an agentic loop? This question naturally arises when building

Strategic commitments shape collective cybersecurity under AI inequality

Model ReleasesDGX agent

arXiv:2605.09415v1 Announce Type: new Abstract: The growing integration of AI into cybersecurity is reshaping the balance between attackers and defenders. When access to advanced AI-enabled defence to

TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit

AgentsDGX agent

arXiv:2507.09788v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLM) have led to a new class of autonomous agents, renewing and expanding interest in the area. LLM-

What Cohort INRs Encode and Where to Freeze Them

TutorialsDGX agent

arXiv:2605.08298v1 Announce Type: cross Abstract: Reusing the early layers of cohort-trained INRs as initialization for new signals has been shown to accelerate and improve signal fitting, yet it rema

xAI is rolling out Skills on http://grok.com. This new feature lets you create custom skills that Grok can reuse across conversations. Skill…

IndustryDGX agent

xAI is rolling out Skills on http://grok.com. This new feature lets you create custom skills that Grok can reuse across conversations. Skills run inside Grok’s sandbox environment, so they can edit fi

← Previous
1…9495969798…169
Next →