AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,030 results
29 Jul 2026

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Model ReleasesDGX agent

arXiv:2607.25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context,

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

Model ReleasesDGX agent

arXiv:2607.24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as mat

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering

AgentsDGX agent

arXiv:2607.25090v1 Announce Type: new Abstract: Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expensive and feed

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image Segmentation

SafetyDGX agent

arXiv:2602.10508v2 Announce Type: replace Abstract: Modern segmentation models achieve strong predictive performance but remain largely opaque, limiting our ability to diagnose failures, understand da

Modular Robotic Catheters for Endovascular Aneurysm Repair

ResearchDGX agent

arXiv:2607.25807v1 Announce Type: new Abstract: Fenestrated/Branched endovascular aneurysm repair (FEVAR/BEVAR) require surgeons to navigate catheters and guidewires into various branches of the abdom

Multiclass Classification without Labels via Posterior Simplex Geometry

ResearchDGX agent

arXiv:2607.24943v1 Announce Type: cross Abstract: In many classification problems, reliable instance-level labels are unavailable. However, it is often possible to construct weakly enriched unlabeled

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures …

Model ReleasesDGX agent

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures whether the agent reached it the permitted way, which is the

On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?

Model ReleasesDGX agent

arXiv:2607.24784v1 Announce Type: new Abstract: Specialised translation relies on the use of documentary and terminological resources, including corpora. These resources are particularly useful for te

OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation

Model ReleasesDGX agent

arXiv:2607.25656v1 Announce Type: new Abstract: Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (

SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

SafetyDGX agent

arXiv:2607.24850v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning h

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play

SafetyDGX agent

arXiv:2607.25425v1 Announce Type: new Abstract: Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web explo

Towards an Agent Operating System - Lessons from Classical and Cloud OS

AgentsDGX agent

arXiv:2607.25076v1 Announce Type: new Abstract: Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc implementations, f

Towards Reliable Stain Transfer: An Iterative Data-Model Co-Optimization Framework Based on Multimodal Expert-Guided Assessment

ResearchDGX agent

arXiv:2607.25393v1 Announce Type: new Abstract: Histopathological examination primarily relies on hematoxylin and eosin (H&E) and immunohistochemistry (IHC) staining. Although IHC provides critical mo

Tripody: An Overconstrained 3-SPR-like Parallel Robot for High-Reach Construction Tasks

SafetyDGX agent

arXiv:2607.25781v1 Announce Type: new Abstract: Many ceiling construction tasks still rely on heavy serial manipulators that are difficult to deploy in cluttered interiors, motivating lightweight, fie

// Unfolding Sub-Agents for Long-Horizon ML Engineering // Watch a single agent work a machine learning engineering task for six hours you s…

AgentsDGX agent

// Unfolding Sub-Agents for Long-Horizon ML Engineering // Watch a single agent work a machine learning engineering task for six hours you see issues like context fills with stack traces, dead experim

We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use …

Model ReleasesDGX agent

We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use it to scan repositories, track findings across runs, verify

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https…

SafetyDGX agent

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https://youtu.be/MW8-kqd2SD8?si=jSKDogcUWbJ8N2k7 Our main takeawa

'We'll have to see how it works': An interview study to understand collaborative practices in interdisciplinary artificial intelligence and healthcare research

ApplicationsDGX agent

arXiv:2311.18424v3 Announce Type: replace-cross Abstract: Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients an

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

SafetyDGX agent

arXiv:2607.25152v1 Announce Type: new Abstract: Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation

When Thinking Before Retrieval Hurts: TraceBound Diagnostics for Adaptive Knowledge-Graph Retrieval

SafetyDGX agent

arXiv:2607.24800v1 Announce Type: cross Abstract: Adaptive retrieval promises to make knowledge-graph question answering more robust by letting a controller search, inspect neighborhoods, revise actio

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing

SafetyDGX agent

arXiv:2607.25648v1 Announce Type: cross Abstract: Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressur

28 Jul 2026

A Diagnostic Gap Framework for Evaluating Reconstruction Fidelity in Weakly Supervised Mammography

ResearchDGX agent

arXiv:2607.22740v1 Announce Type: new Abstract: Weakly supervised pipelines for medical imaging have become increasingly popular over the years. These systems often include multiple stages and compone

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

Model ReleasesDGX agent

arXiv:2607.23784v1 Announce Type: cross Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies tha

A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection

ResearchDGX agent

arXiv:2607.22722v1 Announce Type: cross Abstract: Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation

ACM: Agentic Context Management for Long Horizon Tasks

AgentsDGX agent

arXiv:2607.23809v1 Announce Type: new Abstract: Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existing context co

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models

ResearchDGX agent

arXiv:2603.05147v2 Announce Type: replace Abstract: Current research on Vision-Language-Action (VLA) models predominantly focuses on enhancing generalization through reasoning techniques. While effect

Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents

Model ReleasesDGX agent

arXiv:2607.24625v1 Announce Type: cross Abstract: Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynam

Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction

AgentsDGX agent

arXiv:2602.00575v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) provides a promising pathway for continuously advancing GUI agents, yet existing reward modeli

AI Strategy: How to Choose What AI Product to Implement

TutorialsDGX agent

arXiv:2607.23733v1 Announce Type: cross Abstract: Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve opposite d

An adaptive multi-fuzzy logic model for diagnosing transformer faults using dynamic weight optimization

ResearchDGX agent

arXiv:2607.23486v1 Announce Type: new Abstract: Dissolved gas analysis (DGA) is crucial for diagnosing early power transformer failures. Traditional DGA interpretation methods like Duval Triangle, IEC

An Unofficial FastLAS Tutorial: A Programmer's Guide

SafetyDGX agent

arXiv:2607.23557v1 Announce Type: cross Abstract: FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of examples, and

ARdena: Scenario-driven control of real-time LLM agents

SafetyDGX agent

arXiv:2607.22651v1 Announce Type: new Abstract: Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time interactive e

At the core of our mission is working through how to ensure increasingly powerful AI benefits everyone. We believe that, at some point in th…

TutorialsDGX agent

At the core of our mission is working through how to ensure increasingly powerful AI benefits everyone. We believe that, at some point in the future, AI acceleration for frontier model development may

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

Model ReleasesDGX agent

arXiv:2607.22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuni

CausAdv: A Causal-based Framework for Detecting Adversarial Examples

ResearchDGX agent

arXiv:2411.00839v4 Announce Type: replace-cross Abstract: Deep learning has led to tremendous success in computer vision, largely due to Convolutional Neural Networks (CNNs). However, CNNs have been s

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

Local AiDGX agent

arXiv:2607.22688v1 Announce Type: new Abstract: Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research traj

Coding agents are helping scientists spend more time advancing research, taking on everything from routine maintenance and targeted optimiza…

Model ReleasesDGX agent

Coding agents are helping scientists spend more time advancing research, taking on everything from routine maintenance and targeted optimization to complete redesigns and new systems. While agents can

Concept-based Visual Counterfactual Explanations with Diffusion Models

SafetyDGX agent

arXiv:2607.22544v1 Announce Type: new Abstract: Visual counterfactual explanations aim to answer 'what minimal change to this image would flip the model's prediction?', and are increasingly important

Continual-RL for Generalization in Autonomous Racing on the RoboRacer Platform

SafetyDGX agent

arXiv:2607.24320v1 Announce Type: new Abstract: A key challenge in modern robotics is to adapt to changing environments, a challenge that is exacerbated when simulations cannot encompass every possibl

Continuous surrogates versus threshold Boolean networks for modeling Arabidopsis ISR gene regulation

ResearchDGX agent

arXiv:2607.23289v1 Announce Type: cross Abstract: Gene regulatory network modeling often requires balancing predictive accuracy and mechanistic interpretability. In this work, we compare continuous su

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

ResearchDGX agent

arXiv:2607.24717v1 Announce Type: cross Abstract: Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixe

Directional Influence Function: Estimating Training Data Influence in Constrained Learning

SafetyDGX agent

arXiv:2607.23388v1 Announce Type: cross Abstract: As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustnes

DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video

AgentsDGX agent

arXiv:2607.23132v1 Announce Type: new Abstract: Assessing the severity of a traffic accident scenario is important to decide which emergency service to dispatch. Missing an ambulance dispatch on a ped

DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

Model ReleasesDGX agent

arXiv:2607.22644v1 Announce Type: new Abstract: Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or typ

Evaluating LLMs as Interpretable Controllers for Dynamical Systems

Model ReleasesDGX agent

arXiv:2607.22609v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for decision-making and reasoning tasks, yet their potential as controllers for physical systems rema

Extending Desbordante with Probabilistic Functional Dependency Discovery Support

TutorialsDGX agent

arXiv:2607.23636v1 Announce Type: cross Abstract: Data profiling aims to extract complex patterns from data for further analysis and use that data in domains such as data cleaning, data deduplication,

False Prophets: On the Security of World Models in Agentic Systems

Model ReleasesDGX agent

arXiv:2607.23147v1 Announce Type: cross Abstract: Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of t

From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages

Model ReleasesDGX agent

arXiv:2607.24542v1 Announce Type: new Abstract: Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient lang

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

AgentsDGX agent

arXiv:2607.24117v1 Announce Type: new Abstract: Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existi

ID-V2V: Identity-Preserving Video Restylization

ApplicationsDGX agent

arXiv:2607.22830v1 Announce Type: new Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance whil

Invariant Discovery for Networked Systems

TutorialsDGX agent

arXiv:2607.22944v1 Announce Type: cross Abstract: Invariants, the relations expected to hold among measured signals of a network, underpin applications from verification to traffic generation, telemet

LanteRn: Latent Visual Structured Reasoning

ResearchDGX agent

arXiv:2603.25629v2 Announce Type: replace Abstract: While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs). As a result, m

Let AI Agents Translate Networks, Not Reason About Them

Local AiDGX agent

arXiv:2607.22947v1 Announce Type: new Abstract: A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no production network

LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models

Model ReleasesDGX agent

arXiv:2411.00918v5 Announce Type: replace-cross Abstract: Mixture of experts (MoE) architectures have become a cornerstone for scaling up and are a key component in most large language models such as

LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports

Model ReleasesDGX agent

arXiv:2607.24573v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes

Making Mathematical Knowledge Explainable, Accessible and Interoperable Through Large Language Model Integration

SafetyDGX agent

arXiv:2607.24512v1 Announce Type: new Abstract: Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles. Knowledge bases such as

MANGO: A Global Single-Date Paired Dataset for Mangrove Segmentation

Model ReleasesDGX agent

arXiv:2601.17039v2 Announce Type: replace-cross Abstract: Mangroves are critical for climate-change mitigation, requiring reliable monitoring for effective conservation. While deep learning has emerge

ML-based Predictive Models for Power Consumption in Virtualised O-RANs

ResearchDGX agent

arXiv:2607.24256v1 Announce Type: cross Abstract: As communication networks adopt virtualized and disaggregated architectures, achieving energy efficiency has become increasingly important for both ec

Mwando: Leveraging AI to Preserve and Teach shiKomori

AgentsDGX agent

arXiv:2607.23481v1 Announce Type: new Abstract: This paper presents Mwando, a virtual educational assistant designed to support the teaching and preservation of shiKomori, the language of the Comoros

Neptuna: A Comprehensive Machine Learning Framework for Benchmarking Complex Multiphase Flows

Model ReleasesDGX agent

arXiv:2607.22280v2 Announce Type: replace-cross Abstract: Compressible multiphase flows involving shocks and material interfaces arise in applications such as bubble collapse and droplet breakup, wher

← Previous
1…109110111112113…168
Next →