AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,005 results
Tools

Pelicans for Meta's new Muse Spark models - plus I did a bit of a deep dive into the Code Interpreter and fascinating 'container.visual_grou…

DGX agent

Pelicans for Meta's new Muse Spark models - plus I did a bit of a deep dive into the Code Interpreter and fascinating 'container.visual_grounding' tools in their http://meta.ai chat UI https://simonwi

toolssimon-willison--x
8 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Tools

Quoting Giles Turnbull

DGX agent

I have a feeling that everyone likes using AI tools to try doing someone else’s profession. They’re much less keen when someone else uses it for their profession. — Giles Turnbull, AI and the human vo

toolssimon-willison
8 Apr 2026
Model Releases

Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs

DGX agent

arXiv:2608.11232v1 Announce Type: cross Abstract: Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest outputs requ

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

What unique, custom QOL upgrades have you given your local agents?

DGX agent

Warning: Kinda long post. If you don't like reading, please skip for your own sanity. Also, I've got nothing to sell, just a tinkerer, so I just want to share ideas and learn from you guys too. When I

model-releasesr-localllama
12 Aug 2026
Model Releases

Learning to Triage Vulnerability Reports from Program Analysis: An Empirical Study in Node.js

DGX agent

arXiv:2510.20739v2 Announce Type: replace-cross Abstract: Program analysis tools often produce large volumes of candidate vulnerability reports that require costly manual review, creating a practical

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

One Adapter Pair per Model: A Universal Activation Interface for Language Models

DGX agent

arXiv:2608.09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

You Don't Need To Stay in The Loop: An Agentic Robotics Loop for Robot-Policy Improvement

DGX agent

arXiv:2608.07555v1 Announce Type: new Abstract: Coding agents such as Claude Code and Codex close the software loop: a main agent manages the loop, subagents analyze and execute, tools do the work. We

model-releasesarxiv-cs-ro
11 Aug 2026
Agents

Agentic Planning for Symbolic Execution

DGX agent

arXiv:2608.06397v1 Announce Type: cross Abstract: Symbolic execution seeks to explore feasible program paths, yet a practical run may exhaust its resources while much program behaviour remains unreach

agentsarxiv-cs-ai
10 Aug 2026
Model Releases

b10291

DGX agent

vulkan: fix submission batching size, add debug tools for diagnosing causes of DeviceLost drivers errors (#26371) vulkan: add debug tooling to get more information about a DeviceLost error fix submiss

model-releasesllama-cpp-releases
6 Aug 2026
Model Releases

Uber burned through its 2026 AI coding budget in four months. Microsoft canceled most of its Claude Code licenses six months after rolling t…

DGX agent

Uber burned through its 2026 AI coding budget in four months. Microsoft canceled most of its Claude Code licenses six months after rolling them out. The mechanics are simple: per-token cost keeps fall

model-releasesitamar-friedman--x
6 Aug 2026
Industry

Reddit is introducing a new moderator: AI

DGX agent

Reddit is enlisting AI to help moderate new subreddits - and eventually the rest of site. The company is introducing automated moderation tools that rely on LLMs to help mods manage their communities,

industrythe-verge-ai
5 Aug 2026
Model Releases

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

DGX agent

arXiv:2608.03499v1 Announce Type: new Abstract: Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served b

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

b10249

DGX agent

server: add get_info tool (#26522) server: add get_info tool fix --rpc in docs server: harden get_info probe result handling Report the OS as unknown when the probe process fails to spawn or times out

model-releasesllama-cpp-releases
4 Aug 2026
Model Releases

SERL-SQL: Selective Hindsight Distillation for Text-to-SQL Reinforcement Agentic Learning

DGX agent

arXiv:2608.00485v1 Announce Type: new Abstract: Recent Text-to-SQL systems increasingly rely on multi-turn interaction, execution feedback, and reinforcement learning. However, most existing methods u

model-releasesarxiv-cs-cl
4 Aug 2026
Agents

SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection

DGX agent

arXiv:2512.06716v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabi

agentsarxiv-cs-cl
4 Aug 2026
Model Releases

b10227

DGX agent

chat : add qwen3 specialized parser (#26252) Add tagged thinking tool parser chat : refactor and add permute helper cont : add support for <tool_call> omission cont : update tool delimiters cont : add

model-releasesllama-cpp-releases
2 Aug 2026
Model Releases

Mechanistic interpretability streamlined for everyday users like us😎 🧠

DGX agent

Context: I want to give the community an Open Research (well open under Apache 2.0 clause) - tool that allows everyday users like us to look deeper into the local models we use consistently. Mechanist

model-releasesr-localllama
30 Jul 2026
Model Releases

Building a training dataset: pulling and restoring stills from video sources

DGX agent

I was looking for a tool to help me train a character lora from an old movie (think 1980's low-budget movie). The digital transfer was low-quality; modern upscales exist and they are horrible. So I wa

model-releasesr-stablediffusion
28 Jul 2026
Model Releases

Detect early and enforce firmly with Google Cloud's enhanced cost controls for AI spend

DGX agent

Generative AI can make cloud costs difficult to predict. A single five-word prompt can run complex operations and generate significant costs. Traditional metrics like requests per second no longer hel

model-releasesgoogle-cloud-ai
28 Jul 2026
Safety

MemTX: Transactional Belief Commit for Stateful Agent Memory

DGX agent

arXiv:2607.23929v1 Announce Type: new Abstract: LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a tool call with

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

DGX agent

arXiv:2607.23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced

model-releasesarxiv-cs-ai
28 Jul 2026
Tools

You can now fine-tune Kimi K3 on Fireworks. Conduct supervised fine-tuning, preference tuning, and reinforcement learning via Training API. …

DGX agent

You can now fine-tune Kimi K3 on Fireworks. Conduct supervised fine-tuning, preference tuning, and reinforcement learning via Training API. Run across dedicated, and serverless training. The first ope

toolsfireworks-ai--x
28 Jul 2026
Model Releases

90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B

DGX agent

Last week someone here said ThinkingCap and Fable Fusion 'really do beat the OG' for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the

model-releasesr-localllama
26 Jul 2026
Agents

Causal-AgentIR: Self-Evolving Causal Memory for Adaptive Image Restoration Agents

DGX agent

arXiv:2607.21125v1 Announce Type: new Abstract: Image restoration agents have recently emerged as a flexible paradigm for handling diverse and unpredictable degradations in real-world scenarios. Exist

agentsarxiv-cs-cv
24 Jul 2026
Model Releases

GuardianAgentBench: Where Agents Fail and How to Guard Them

DGX agent

arXiv:2607.20982v1 Announce Type: new Abstract: As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavi

model-releasesarxiv-cs-ai
24 Jul 2026
Agents

ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems

DGX agent

arXiv:2607.19432v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) is an open-source standard that allows AI agents to connect to external tools, databases, and services. While this co

agentsarxiv-cs-ai
23 Jul 2026
Model Releases

One encoder, seven heads: what we learned training a unified security classifier with masked losses [P]

DGX agent

We spent the last months consolidating seven separate sequence classifiers into one multi-head model, our apex model, so to speak, and since the weights are now public, I wanted to share what worked a

model-releasesr-machinelearning
22 Jul 2026
Model Releases

v0.32.1

DGX agent

What's Changed Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations Fixed a recurrent MLX model cache leak that could increase memory use across

model-releasesollama-releases
16 Jul 2026
Model Releases

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

DGX agent

arXiv:2512.01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

How Inference Compute Shapes Frontier LLM Evaluation

DGX agent

arXiv:2606.17930v2 Announce Type: replace Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result,

model-releasesarxiv-cs-ai
15 Jul 2026
Local Ai

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

DGX agent

arXiv:2607.13027v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and it

local-aiarxiv-cs-ai
15 Jul 2026
Model Releases

Securing the AI supply chain on GKE: Introducing k8s-aibom for automated AI BOMs

DGX agent

How should your security team manage shadow AI? Workloads deployed by developers without formal registration can often evade traditional security scanners, because organizations are reluctant to slow

model-releasesgoogle-cloud-ai
13 Jul 2026
Model Releases

GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads

DGX agent

arXiv:2607.02518v1 Announce Type: cross Abstract: OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the c

model-releasesarxiv-cs-ai
7 Jul 2026
Agents

Rethinking Scientific Discovery in an Agentic Era

DGX agent

arXiv:2607.03863v1 Announce Type: new Abstract: Artificial intelligence has advanced scientific discovery, but most AI4Science systems remain fragmented tools that rely on humans to coordinate problem

agentsarxiv-cs-cl
7 Jul 2026
Hardware

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference

DGX agent

arXiv:2607.03333v1 Announce Type: cross Abstract: LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often serial: th

hardwarearxiv-cs-ai
7 Jul 2026
Safety

ElephantAgent: Contextual State Continuity in Agentic Systems

DGX agent

arXiv:2607.01919v1 Announce Type: new Abstract: Agentic systems enhance their capabilities by invoking external tools and maintaining persistent memory. However, these external dependencies introduce

safetyarxiv-cs-ai
3 Jul 2026
Safety

Safeguarding LLM Agents from Misalignment through Provenance Analysis

DGX agent

arXiv:2607.01236v1 Announce Type: cross Abstract: As LLM agents gain increasing access to powerful tools, ensuring that their actions are aligned with the user's intent becomes critical. When an agent

safetyarxiv-cs-ai
3 Jul 2026
Agents

Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents

DGX agent

arXiv:2606.31229v1 Announce Type: new Abstract: Ideation plays a pivotal role in scientific discovery. Recent LLM, especially AI Scientist systems, show promising potential for automated ideation. How

agentsarxiv-cs-ai
1 Jul 2026
Model Releases

Bringing speed and strong cost performance to the market with Gemini Omni Flash and Nano Banana 2 Lite

DGX agent

Great creative happens when your tools move at the speed of your ideas. To help you create rich, reliable experiences while reducing regeneration time and costs, we’re adding two new models to Gemini

model-releasesgoogle-cloud-ai
30 Jun 2026
Model Releases

Inside Genebench-Pro

DGX agent

Genebench-Pro appears to be a case study or tool from OpenAI focused on benchmarking or evaluating genetic/genomic analysis capabilities, likely demonstrating how OpenAI's models or tools can be appli

model-releasesopenai
30 Jun 2026
Model Releases

MCP Server Architecture Patterns for LLM-Integrated Applications

DGX agent

arXiv:2606.30317v1 Announce Type: cross Abstract: The Model Context Protocol (MCP), introduced by Anthropic in November 2024, defines a standardized interface for connecting large language models (LLM

model-releasesarxiv-cs-ai
30 Jun 2026
Agents

wanted a quick way to visualize side events for @aidotengineer world fair next week https://yoheinakajima.github.io/aie-side-events/ filter …

DGX agent

Yohei Nakajima created an interactive visualization tool for side events at the AI Dojo Engineer World Fair, accessible via a GitHub Pages site. The tool appears to include filtering capabilities to h

agentsyohei-nakajima--x
26 Jun 2026
Local Ai

b9786

DGX agent

b9786 is a release of llama.cpp, a tool for LLM inference in C/C++ . Llama.cpp is a free and open-source tool that allows users to run AI models locally on Windows, Linux, and macOS . The b9786 releas

local-aillama-cpp-releases
25 Jun 2026
Agents

In the meantime, for parsing human-native documents, check out LlamaParse: https://cloud.llamaindex.ai/

DGX agent

LlamaParse is a document parsing tool from LlamaIndex designed to extract and process data from human-readable documents with high accuracy. The tool is available as a cloud service and represents Lla

agentsjerry-liu--x
20 Jun 2026
Safety

APPO: Agentic Procedural Policy Optimization

DGX agent

arXiv:2606.12384v1 Announce Type: cross Abstract: Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents

safetyarxiv-cs-ai
11 Jun 2026
Local Ai

Demo: Turn Research Into a Client-Ready Report with Row-Bot

DGX agent

Row-Bot is a local-first desktop AI assistant that orchestrates tools and models to handle reasoning and workflows while keeping data local. The demo likely showcases how Row-Bot's integrated tools, k

local-air-ollama
10 Jun 2026
Model Releases

Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures

DGX agent

arXiv:2606.08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observabili

model-releasesarxiv-cs-ai
9 Jun 2026
Tutorials

How to use Row-Bot to turn unread emails into a daily action plan

DGX agent

Row-Bot is a local-first desktop AI assistant that can check emails and integrate multiple tools in a single conversation turn . The application includes integrated tools, a personal knowledge graph,

tutorialsr-ollama
9 Jun 2026
← Previous
1…2627282930…209
Next →