AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,030 results
4 Aug 2026

MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflow Routing

Model ReleasesDGX agent

arXiv:2608.00107v1 Announce Type: new Abstract: Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an in

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study

AgentsDGX agent

arXiv:2608.00902v1 Announce Type: new Abstract: LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck. KV c

PyDPF: A Python Package for Differentiable Particle Filtering

ApplicationsDGX agent

arXiv:2510.25693v3 Announce Type: replace-cross Abstract: State-space models (SSMs) are a widely used tool in time series analysis. In the complex systems that arise from real-world data, it is common

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

Model ReleasesDGX agent

arXiv:2608.02358v1 Announce Type: new Abstract: To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone,

Thermalizing Stochastic Programs

ResearchDGX agent

arXiv:2608.01615v1 Announce Type: cross Abstract: We present a set of tools for mapping general stochastic programs to thermodynamic hardware designed for energy-efficient stochastic sampling. Given a

this post from 3 years ago and commercial LLMs *still* can’t play chess anywhere near as well serious players (except by calling external to…

Model ReleasesDGX agent

this post from 3 years ago and commercial LLMs *still* can’t play chess anywhere near as well serious players (except by calling external tools) When @GaryMarcus and others point out that GPT-4 is bad

Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety

Model ReleasesDGX agent

arXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequenc

3 Aug 2026

An analysis of machine learning approaches for enhancing decision-making in complex discrete choice tasks

SafetyDGX agent

arXiv:2607.28854v1 Announce Type: new Abstract: Discrete choice modeling is a common tool used for preference elicitation during policy-making, but this is typically done through parametric models. Ma

CrowdStrike finds AI systems under direct attack as exploit windows shrink

Model ReleasesDGX agent

Artificial intelligence has become a target for attackers rather than only a tool they use, according to CrowdStrike Holdings Inc.’s “2026 Threat Hunting Report,” released today. The annual report dra

Fracture Risk Prediction in Adults Over 50 Years Old Using DXA and EHR: Comparison of Traditional and Machine Learning Models in Two Large Cohorts

ResearchDGX agent

arXiv:2607.28671v1 Announce Type: cross Abstract: Accurate fracture risk prediction is important for osteoporosis management, but commonly used clinical tools may not fully use information available i

From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale

AgentsDGX agent

arXiv:2607.29516v1 Announce Type: cross Abstract: AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code review tools o

future generations will have no idea what was real and what was not.

SafetyDGX agent

Gary Marcus warns that, with increasingly sophisticated generative‑AI tools, future generations may struggle to tell what is real from what is fabricated. Matt Stoller adds that generative AI will tra

Improving scDiffusion with Sparsity-Biased Classifier-Free Guidance

ResearchDGX agent

arXiv:2607.29043v1 Announce Type: cross Abstract: Single-cell RNA sequencing (scRNA-seq) has become an essential tool in modern cellular biology, and generating accurate synthetic scRNA-seq data is be

MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

Model ReleasesDGX agent

arXiv:2607.29002v1 Announce Type: new Abstract: Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult

Sakana Namazu: An LLM API with Japanese-vibes! 🎏 Built for Japanese enterprises, featuring frontier-level reasoning and built-in agentic to…

AgentsDGX agent

Sakana Namazu: An LLM API with Japanese-vibes! 🎏 Built for Japanese enterprises, featuring frontier-level reasoning and built-in agentic tools. 開発者の皆様、大変お待たせしました!Sakana Chatのモデルが遂にAPIとして公開です。ぜひお試しください

The Checking Problem: What must be true before AI ships in a regulated firm

ApplicationsDGX agent

arXiv:2607.28666v1 Announce Type: new Abstract: Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of

2 Aug 2026

I built a self-hosted studio that turns one reference photo into a curated, captioned, trained and tested LoRA — one browser tab, open source, MIT

Local AiDGX agent

I shared this tool here a week ago and the feedback shaped a big new version, so here's the full tour of what it does today. Screenshots of every screen: github.com/perfectgf/lora-dataset-studio — plu

Real-world reality check on Qwen for autonomous coding agents

Model ReleasesDGX agent

TLDR below 👇🏼 I’ve seen a lot of hype around Qwen 3.6 35B and 3.5 120B lately, especially regarding coding and tool-use capabilities. On this subreddit it is the defacto recommended model for everyone

Try handling complex tasks to your local models with GraphARC, graph engineering yes !

Local AiDGX agent

🚀 We just built our first real-time implementation of Graph Engineering, inspired by our experience building graph tooling used by 4,000+ developers. 🔗 Repo: https://github.com/CodeGraphContext/grapha

1 Aug 2026

b10217

Model ReleasesDGX agent

chat : enable tool call in thinking for DS4 (#26269) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFr

great write up, evals are hard! here are 2 broad buckets we use to evaluate agents: 1. Measure the State of the World 2. Agent as a Judge on…

AgentsDGX agent

great write up, evals are hard! here are 2 broad buckets we use to evaluate agents: 1. Measure the State of the World 2. Agent as a Judge on the Trajectory 1. Measure the state of the environment befo

31 Jul 2026

AI as Friction for Reflection Support in Ideation

AgentsDGX agent

arXiv:2607.26827v1 Announce Type: cross Abstract: Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration and faster o

AI-native software development requires a new engineering model

IndustryDGX agent

Artificial intelligence has quickly become a standard part of modern software development. Coding assistants, code completion tools and AI-powered integrated development environments are now widely av

An analysis of binary isotonic regression: degrees of freedom and implications for calibration

ResearchDGX agent

arXiv:2607.27301v1 Announce Type: cross Abstract: Isotonic regression is a canonical tool for estimating monotone functions and calibrating probabilistic predictors. We provide a fully sharp finite-sa

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

AgentsDGX agent

arXiv:2607.15715v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components

Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories

Model ReleasesDGX agent

arXiv:2607.27595v1 Announce Type: new Abstract: Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-pa

Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments

AgentsDGX agent

arXiv:2607.28591v1 Announce Type: cross Abstract: Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a r

Could we all crowdsource a dataset/model/finetune?

Local AiDGX agent

I know it’s been discussed to try to make our own model through crowdsourcing, but finetuning seems like it would be even easier. We could edit and proofread and write our own datasets at a large scal

Expanding Data-Agnostic Pivotal Instances Selection Models with Proximity Trees and Ensemble Learning

ResearchDGX agent

arXiv:2607.27522v1 Announce Type: new Abstract: As decision-making processes grow more complex, machine learning tools have become essential for tackling business and societal challenges. However, man

FADEx: Feature Attribution and Distortion-based Explanation of Dimensionality Reduction

ResearchDGX agent

arXiv:2607.27463v1 Announce Type: new Abstract: Dimensionality Reduction (DR) is a fundamental tool for high-dimensional data exploration, reducing the complexity of latent spaces of machine learning

How does downsampling affect needle electromyography signals? A generalisable workflow for understanding downsampling effects on high-frequency time series

ResearchDGX agent

arXiv:2601.10191v2 Announce Type: replace Abstract: Automated analysis of needle electromyography (nEMG) signals is emerging as a tool to support the detection of neuromuscular diseases (NMDs), yet th

Language Diversity: Evaluating Language Usage and AI Performance on African Languages in Digital Spaces

Model ReleasesDGX agent

arXiv:2512.01557v3 Announce Type: replace Abstract: This study examines the digital representation of African languages and the challenges this presents for current language detection tools. We evalua

On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems

Model ReleasesDGX agent

arXiv:2607.28080v1 Announce Type: cross Abstract: We extend a recently introduced Entropy-Optimal Manifold Clustering (EOMC) to allow for a joint simultaneous identification of subsets and subspaces o

Optimal Realistic Local AI for Most

Model ReleasesDGX agent

So you’ve got a 3090 or maybe even a 5090? Or more likely a 4060 8GB Ti. You wanna try local AI, you don’t know what it can/can’t do. 1) Install the best model you can. If you have a 3090 or a 5090, t

Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation

ApplicationsDGX agent

arXiv:2607.27210v1 Announce Type: new Abstract: The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting me

SciDataSailor: Deep Scientific Data Exploring

AgentsDGX agent

arXiv:2607.28098v1 Announce Type: cross Abstract: Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, in

smevals - a small eval suite for evaluating models, prompts, and harnesses

Model ReleasesDGX agent

smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer

Sparsity Induced Identifiability in Matrix Tri-Factorisation

ResearchDGX agent

arXiv:2607.27507v1 Announce Type: new Abstract: Matrix factorisation is a fundamental tool for exploiting low-dimensional structure in high-dimensional data, with applications such as data compression

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

TutorialsDGX agent

arXiv:2607.27179v1 Announce Type: cross Abstract: Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among

Theatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography

ResearchDGX agent

arXiv:2607.27266v1 Announce Type: new Abstract: We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books. This provides a tool that accelerates

What's your local AI coding setup on a MacBook Pro M4?

Model ReleasesDGX agent

I've spent the last couple of days trying different setups (Ollama, Continue, Claude Code, Gemini CLI, OpenRouter...) and at this point I feel like I've spent more time configuring tools than actually

30 Jul 2026

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

SafetyDGX agent

arXiv:2607.26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by

Atomic Chat signed the Open Weights letter! We believe everyone should be able to run AI on their own device. When a model is open, thousand…

SafetyDGX agent

Atomic Chat signed the Open Weights letter! We believe everyone should be able to run AI on their own device. When a model is open, thousands of teams fine-tune it, quantize it and build new tools on

Do more with less: How GKE can reduce your cost per agent by 75%

AgentsDGX agent

In today’s agentic era, modern cloud applications are evolving from a set of passive tools to fleets of autonomous digital workers that reason, plan, and take action across a wide range of tasks. For

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Model ReleasesDGX agent

Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic app

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

ApplicationsDGX agent

arXiv:2607.26769v1 Announce Type: new Abstract: Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether

29 Jul 2026

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

SafetyDGX agent

arXiv:2607.25489v1 Announce Type: new Abstract: Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

AgentsDGX agent

arXiv:2607.25379v1 Announce Type: new Abstract: Cyber-capable AI agents combine language models with tools, memory, and execution en- vironments to perform multi-step offensive-security tasks. Existin

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

SafetyDGX agent

arXiv:2607.24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observa

Endpoint Replay: Compressing the Recency Buffer in Deep Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.25123v1 Announce Type: new Abstract: Experience replay remains one of the most practical and useful algorithmic tools in the deep reinforcement learning (DRL) toolbox. Aside from the limite

Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

SafetyDGX agent

arXiv:2507.11548v3 Announce Type: replace-cross Abstract: The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias re

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

Model ReleasesDGX agent

arXiv:2607.25583v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute budgets, yet the

Install the open-source Codex Security CLI,: npm install @OpenAI/codex-security Or start with: npx @OpenAI/codex-security@latest --help NPM:…

Model ReleasesDGX agent

OpenAI has released the open‑source Codex Security CLI, which can be installed with `npm install @OpenAI/codex-security` or run directly via `npx @OpenAI/codex-security@latest --help`. The tool scans

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Model ReleasesDGX agent

arXiv:2607.25904v1 Announce Type: new Abstract: Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluat

Jensen Huang was asked about the people technology left behind. He didn't offer a plan. He said the gap already closed. Huang: 'All of a sud…

TutorialsDGX agent

Jensen Huang was asked about the people technology left behind. He didn't offer a plan. He said the gap already closed. Huang: 'All of a sudden artificial intelligence closed that technology divide.'

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning

SafetyDGX agent

arXiv:2607.25369v1 Announce Type: new Abstract: Agentic systems have rapidly advanced in their ability to interact with real-world environments, leverage external tools, and provide services for users

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Model ReleasesDGX agent

arXiv:2607.25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Pri

Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course

TutorialsDGX agent

arXiv:2607.24755v1 Announce Type: cross Abstract: This full research paper examines how different forms of learner-AI interaction relate to learning outcomes in object-oriented programming (OOP) cours

Probabilistic Symbolic Regression for Equation Discovery via Operator-induced and Regularized Symbolic Forests

Model ReleasesDGX agent

arXiv:2509.19710v2 Announce Type: cross Abstract: Symbolic regression has emerged as a powerful tool for artificial intelligence-driven scientific discovery by learning interpretable analytical expres

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

Model ReleasesDGX agent

arXiv:2607.25565v1 Announce Type: new Abstract: Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since edita

← Previous
1…6162636465…168
Next →