AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,005 results
6 May 2026

The Design and Composition of Structural Causal Decision Processes

SafetyDGX agent

arXiv:2605.02681v1 Announce Type: cross Abstract: We present two new classes of causal models of decision-making agents. Our approach is motivated by the needs of modeling the economics of computing s

Using LLMs in Software Design: An Empirical Study of GitHub and A Practitioner Survey

ResearchDGX agent

arXiv:2605.01392v1 Announce Type: cross Abstract: Recent advancements in Large Language Models (LLMs) have demonstrated significant potential across a wide range of software engineering tasks, includi

Valley3: Scaling Omni Foundation Models for E-commerce

Model ReleasesDGX agent

arXiv:2605.01278v1 Announce Type: new Abstract: In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understandi

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

very fun to collab with @harvey on their Long Horizon Legal Agent Benchmark. We need more industry specific benchmarks, and Harvey is paving…

Model ReleasesDGX agent

Harrison Chase expresses enthusiasm about collaborating with Harvey on their Long Horizon Legal Agent Benchmark, highlighting the value of developing industry-specific benchmarks for AI evaluation. Th

Will the Carbon Border Adjustment Mechanism Impact European Electricity Prices? A GNN-Based Network Analysis

SafetyDGX agent

arXiv:2605.03304v1 Announce Type: new Abstract: The European Union's Carbon Border Adjustment Mechanism (CBAM) creates a complex challenge for the interconnected European electricity market. Tradition

5 May 2026

10 trillion samples a day: Scaling beyond traditional monitoring infra at Databricks

IndustryDGX agent

Databricks describes their approach to handling massive-scale observability and monitoring infrastructure that processes 10 trillion samples daily, moving beyond conventional monitoring solutions. The

2D-ThermAl: Physics-Informed Framework for Thermal Analysis of Circuits using Generative AI

ResearchDGX agent

arXiv:2512.01163v2 Announce Type: replace Abstract: Thermal analysis is increasingly critical in modern integrated circuits, where non-uniform power dissipation and high transistor densities can cause

A Category-Theoretic Analysis of Conformal Prediction

ResearchDGX agent

arXiv:2507.04441v4 Announce Type: replace-cross Abstract: Conformal prediction (CP) produces prediction regions with finite-sample, distribution free coverage guarantees, but its interpretation as a q

A Language for Describing Agentic LLM Contexts

AgentsDGX agent

arXiv:2605.01920v1 Announce Type: cross Abstract: Large language models are increasingly used within larger systems ('LLM agents'). These make a sequence of LLM calls, each call providing the LLM with

Active Reasoning Vision-Language Models via Sequential Experimental Design

ResearchDGX agent

arXiv:2605.01345v1 Announce Type: new Abstract: Visual perception in modern Vision-Language Models (VLMs) is constrained by a fundamental perceptual bandwidth bottleneck: a broad field of view inevita

Adaptive Interpolation-Synthesis for Motion In-Betweening on Keyframe-Based Animation

SafetyDGX agent

arXiv:2605.02742v1 Announce Type: cross Abstract: Motion in-betweening is one of the most artistically demanding and time consuming stages of 3D animation, where the expressivity and rhythm of motion

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction

AgentsDGX agent

arXiv:2602.05353v3 Announce Type: replace-cross Abstract: Large Language Models have shown strong capabilities in complex problem solving, yet many agentic systems remain difficult to interpret and co

Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure

SafetyDGX agent

arXiv:2605.00055v1 Announce Type: cross Abstract: We report a safety incident in a deployed multi-agent research system in which a primary AI agent installed 107 unauthorized software components, over

Analyzing Adversarial Inputs in Deep Reinforcement Learning

SafetyDGX agent

arXiv:2402.05284v2 Announce Type: replace Abstract: In recent years, Deep Reinforcement Learning (DRL) has become a popular paradigm in machine learning due to its successful applications to real-worl

Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction

ResearchDGX agent

arXiv:2605.02234v1 Announce Type: cross Abstract: We present a method for diagnosing interpretation in neural networks by identifying an input subspace where a proposed interpretation is highly faithf

Codex is gaining steam

IndustryDGX agent

Codex, likely referring to OpenAI's code generation model, is experiencing increased adoption and usage. The article from Ben's Bites discusses the growing momentum and applications of this AI coding

Constructing Interpretable Features from Compositional Neuron Groups

Model ReleasesDGX agent

arXiv:2506.10920v2 Announce Type: replace Abstract: A central goal for mechanistic interpretability has been to identify the right units of analysis in large language models (LLMs) that causally expla

Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features

Model ReleasesDGX agent

arXiv:2602.10437v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) decompose language model activations into interpretable features, but existing methods reveal only which features a

Creating and Evaluating Figurative Language Dataset for Sindhi

Model ReleasesDGX agent

arXiv:2605.01323v1 Announce Type: new Abstract: In this article, we introduce SiNFluD, a novel benchmark dataset for Sindhi figurative language classification. We first collect raw text from various b

Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages

ResearchDGX agent

arXiv:2605.02608v1 Announce Type: new Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-

DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA

ResearchDGX agent

arXiv:2605.00905v1 Announce Type: new Abstract: Diagram question answering (Diagram QA) requires reasoning-level attribution that links each question-answer pair to all visual regions needed to derive

FeedbackLLM: Metadata driven Multi-Agentic Language Agnostic Test Case Generator with Evolving prompt and Coverage Feedback

Model ReleasesDGX agent

arXiv:2605.01264v1 Announce Type: cross Abstract: Traditional approaches to test case generation often involve manual effort and incur significant computational overhead. Additionally, these approache

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2605.02130v1 Announce Type: new Abstract: Human-level agentic intelligence extends beyond low-level geometric perception, evolving from recognizing where things are to understanding what they ar

GA-VisAgent: A Multi-Agent application for code generation and visualization in interactive learning

AgentsDGX agent

arXiv:2605.01299v1 Announce Type: new Abstract: Geometric Algebra (GA) presents challenges to learners due to its highly abstract mathematical structure and complex operational rules, as translating a

Hallucinations Undermine Trust; Metacognition is a Way Forward

AgentsDGX agent

arXiv:2605.01428v1 Announce Type: new Abstract: Despite significant strides in factual reliability, errors -- often termed hallucinations -- remain a major concern for generative AI, especially as LLM

How to Build In-Vehicle AI Agents with NVIDIA: From Cloud to Car

HardwareDGX agent

The automotive cockpit is undergoing a fundamental shift from rule-based interfaces to agentic, multimodal AI systems capable of reasoning, planning, and acting. This is enabled by an agentic AI pipel

IMPACT-Scribe: Interactive Temporal Action Segmentation with Boundary Scribbles and Query Planning

ResearchDGX agent

arXiv:2605.01668v1 Announce Type: new Abstract: Dense temporal annotation of procedural activity videos is vital for action understanding and embodied intelligence but remains labor-intensive due to r

KANs need curvature: penalties for compositional smoothness

ResearchDGX agent

arXiv:2605.02190v1 Announce Type: new Abstract: Kolmogorov-Arnold networks (KANs) offer a potent combination of accuracy and interpretability, thanks to their compositions of learnable univariate acti

Last Week in AI #340 - OpenAI vs Musk + Microsoft, DeepSeek v4, Vision Banana

Model ReleasesDGX agent

This newsletter episode covers recent AI industry developments including a legal dispute between OpenAI and Elon Musk, Microsoft's involvement in AI developments, the release of DeepSeek's v4 model, a

Learning Koopman operators for coupled systems via information on governing equations of subsystems

TutorialsDGX agent

arXiv:2605.01835v1 Announce Type: new Abstract: Nonlinear coupled systems are ubiquitous in science and engineering. The analysis and modeling of such systems is challenging due to their high dimensio

Linear-Readout Floors and Threshold Recovery in Computation in Superposition

ResearchDGX agent

arXiv:2605.01192v1 Announce Type: new Abstract: Two recent approaches to computation in superposition reach different recursive capacity regimes: Hanni et al. certify ilde{O}(d^{3/2}) computable featu

LITcoder: A General-Purpose Library for Building and Comparing Encoding Models

ResearchDGX agent

arXiv:2509.09152v2 Announce Type: replace Abstract: We introduce LITcoder, an open-source library for building and benchmarking neural encoding models. Designed as a flexible backend, LITcoder provide

MedScribe: Clinically Grounded CT Reporting through Agentic Workflows

AgentsDGX agent

arXiv:2605.01779v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown potential for automated radiology report generation, yet existing approaches rely on global embedding compressi

MolViBench: Evaluating LLMs on Molecular Vibe Coding

Model ReleasesDGX agent

arXiv:2605.02351v1 Announce Type: new Abstract: Molecular Vibe Coding, a paradigm where chemists interact with LLMs to generate executable programs for molecular tasks, has emerged as a flexible alter

Open-access model for detecting openly dumped dispersed municipal solid waste from crowdsourced UAV imagery in Sub-Saharan Africa

Local AiDGX agent

arXiv:2605.02316v1 Announce Type: new Abstract: Managing municipal solid waste in rapidly urbanizing Sub-Saharan Africa remains challenging due to dispersed informal dumping and limited high-resolutio

OpenAI GPT-5 System Card

Model ReleasesDGX agent

arXiv:2601.03267v2 Announce Type: replace Abstract: This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers

Parllama -- a terminal UI for Ollama model management and multi-provider LLM chat

Local AiDGX agent

Parllama is a TUI (Text UI) application designed for easy management and use of Ollama-based LLMs that also works with major cloud-provided LLMs. It provides core model management features including f

Planner Matters! An Efficient and Unbalanced Multi-agent Collaboration Framework for Long-horizon Planning

Model ReleasesDGX agent

arXiv:2605.02168v1 Announce Type: cross Abstract: Language model (LM)-based agents have demonstrated promising capabilities in automating complex tasks from natural language instructions, yet they con

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

ResearchDGX agent

arXiv:2605.02050v1 Announce Type: cross Abstract: This work establishes a foundational framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established ex

Random-Effects Algorithm for Random Objects in Metric Spaces

ResearchDGX agent

arXiv:2605.02693v1 Announce Type: cross Abstract: Across many scientific disciplines, multiple observations are collected from the same experimental units, and in modern datasets these observations of

Reconstructing conformal field theoretical compositions with Transformers

ResearchDGX agent

arXiv:2605.01072v1 Announce Type: cross Abstract: We study the use of transformers to reconstruct the compositions of tensor products of two-dimensional rational conformal field theories (RCFTs) based

Reinforcement Learning from Compiler and Language Server Feedback

SafetyDGX agent

arXiv:2510.22907v2 Announce Type: replace Abstract: Coding agents fail when text-level guesses outrun program facts: they hallucinate APIs, drift to the wrong symbol, and apply edits without evidence

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

Model ReleasesDGX agent

arXiv:2605.01489v1 Announce Type: cross Abstract: Frontier scientific reasoning is rapidly emerging as a key foundation for advancing AI agents in automated scientific discovery. Deep research agents

Semia: Auditing Agent Skills via Constraint-Guided Representation Synthesis

AgentsDGX agent

arXiv:2605.00314v1 Announce Type: cross Abstract: An agent skill is a configuration package that equips an LLM-driven agent with a concrete capability, such as reading email, executing shell commands,

so well explained ! this is like applying Richard Sutton’s Bitter Lesson to agent systems. leverage comes from feedback loops, not handcraft…

AgentsDGX agent

so well explained ! this is like applying Richard Sutton’s Bitter Lesson to agent systems. leverage comes from feedback loops, not handcrafted prompts. attaching feedback to traces is the missing piec

Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes

AgentsDGX agent

arXiv:2510.24332v3 Announce Type: replace-cross Abstract: Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly re

Spectral Model eXplainer: a chemically-grounded explainability framework for spectral-based machine learning models

Model ReleasesDGX agent

arXiv:2605.02684v1 Announce Type: new Abstract: Spectral-based machine learning models have been increasingly deployed in chemometrics and spectroscopy, where predictive accuracy is as important as ex

StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer

Model ReleasesDGX agent

arXiv:2605.00924v1 Announce Type: new Abstract: AI-generated content (AIGC) detectors are increasingly deployed in high-stakes settings such as academic integrity screening, yet their reliability rest

SUDP: Secret-Use Delegation Protocol for Agentic Systems

AgentsDGX agent

arXiv:2604.24920v2 Announce Type: replace-cross Abstract: Agentic systems increasingly act with user secrets for APIs, messaging platforms, and cloud services. Today's bearer-secret interfaces impleme

SwiftPie: Lightning-fast Subject-driven Image Personalization via One step Diffusion

SafetyDGX agent

arXiv:2605.01510v1 Announce Type: new Abstract: Diffusion models have achieved remarkable success in high-quality image synthesis, sparking interest in image-guided generation tasks such as subject-dr

TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs

Model ReleasesDGX agent

arXiv:2510.15545v4 Announce Type: replace Abstract: Accelerating the inference of large language models (LLMs) has been a critical challenge in generative AI. Speculative decoding (SD) substantially i

UnGAP: Uncertainty-Guided Affine Prompting for Real-Time Crack Segmentation

Local AiDGX agent

arXiv:2605.02380v1 Announce Type: new Abstract: Real-time crack segmentation is vital for structural health monitoring but is plagued by aleatoric uncertainties arising from varying lighting, blur, an

Universality in Deep Neural Networks: An approach via the Lindeberg exchange principle

ResearchDGX agent

arXiv:2605.02771v1 Announce Type: cross Abstract: We consider the infinite-width limit of a fully connected deep neural network with general weights, and we prove quantitative general bounds on the 2-

Virtual Scanning for NSCLC Histology: Investigating the Discriminatory Power of Synthetic PET

ResearchDGX agent

arXiv:2605.02746v1 Announce Type: new Abstract: Accurate histological differentiation between adenocarcinoma (ADC) and squamous cell carcinoma (SCC) is critical for personalized treatment in non-small

4 May 2026

A $1 Billion one person company may look like a video game. That’s the view Andrew Pignanelli and the team at Intelligence Co are taking. We…

AgentsDGX agent

A $1 Billion one person company may look like a video game. That’s the view Andrew Pignanelli and the team at Intelligence Co are taking. We sat down with him on @11AMdotclub ahead of today’s launch.

A Novel Patch-Based TDA Approach for Computed Tomography Imaging

ResearchDGX agent

arXiv:2512.12108v5 Announce Type: replace Abstract: The development of machine learning models based on computed tomography (CT) imaging has been a major focus due to the promise that imaging holds fo

Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines

SafetyDGX agent

arXiv:2605.00410v1 Announce Type: new Abstract: A multi-agent pipeline with N agents typically issues N LLM calls per run. Merging agents into fewer calls (compound execution) promises token savings,

Can Coding Agents Reproduce Findings in Computational Materials Science?

Model ReleasesDGX agent

arXiv:2605.00803v1 Announce Type: cross Abstract: Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering be

ComfyStudio Update Since announcing the program, we've made the decision to part ways with the founding studio leads and take a step back to…

ApplicationsDGX agent

ComfyStudio Update Since announcing the program, we've made the decision to part ways with the founding studio leads and take a step back to re-scope the program. ComfyStudio remains a priority for Co

Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments

ResearchDGX agent

arXiv:2505.09901v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate or automate human behavior in complex sequential decision-making settings. A na

← Previous
1…147148149150151…167
Next →