AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,005 results
6 Aug 2026

SafeCommit: Certifying When Memory-Grounded Agents May Safely Act

SafetyDGX agent

arXiv:2608.04289v1 Announce Type: new Abstract: Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects. A central failure mode is premature commitm

Your agentic summer: No-cost lessons from Google experts to build and scale agents

Model ReleasesDGX agent

I’ve talked to developers, IT leaders, and builders who all ask the same question: How do we actually get agents into production? The answer isn't theoretical — it's hands-on. Whether it’s designing a

5 Aug 2026

99.9% uptime changes what your inference architecture has to survive. At Together AI, it means multi-data-center deployment, live traffic ac…

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
ToolsDGX agent

99.9% uptime changes what your inference architecture has to survive. At Together AI, it means multi-data-center deployment, live traffic across both facilities, and enough capacity to absorb a full d

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL

SafetyDGX agent

arXiv:2608.02876v1 Announce Type: new Abstract: Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context

Can LLMs Test Terminal User Interfaces?

Model ReleasesDGX agent

arXiv:2608.03743v1 Announce Type: cross Abstract: Terminal User Interfaces (TUIs) combine the stateful, screen-oriented behaviour of GUIs with terminal deployment and are now common in developer tools

ChatGPT Work is OpenAI's fighter in the highest stake product category in history: bringing the power of coding agents to the masses. It's a…

ToolsDGX agent

ChatGPT Work is OpenAI's fighter in the highest stake product category in history: bringing the power of coding agents to the masses. It's also how a billion users will soon use ChatGPT by default. I

if you have been following his excellent work, @shloked has been breaking down every frontier labs' harness engineering for the last few mon…

ToolsDGX agent

if you have been following his excellent work, @shloked has been breaking down every frontier labs' harness engineering for the last few months. excited to publish his deepest dive into ChatGPT yet as

Introducing Muse Code and Muse Spark 1.2

Model ReleasesDGX agent

Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as

.@Kimi_Moonshot benchmarked K3 endpoints across major inference providers. Together AI leads or ties for #1 on 3 of 4 benchmarks: OCRBench, …

ToolsDGX agent

.@Kimi_Moonshot benchmarked K3 endpoints across major inference providers. Together AI leads or ties for #1 on 3 of 4 benchmarks: OCRBench, MMMU Pro Vision, and DeepSWE. Open models like Kimi K3 have

Qwen Developers' responses from their recent Twitter/X AMA

Model ReleasesDGX agent

Questions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models too apart fr

TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows

Model ReleasesDGX agent

arXiv:2608.02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retrie

4 Aug 2026

AOSpec: Action and Observation Co-Speculation for Low-Latency Agent Serving

Model ReleasesDGX agent

arXiv:2608.00881v1 Announce Type: new Abstract: Large language model agents increasingly act through stateful tools, yet model generation and environment execution remain serialized at every step. As

b10271

Model ReleasesDGX agent

ui: CWD for agent (#26518) server : extend file_glob_search for UI pickers ui : add per-conversation working directory with picker ui : add path navigation and search scope to cwd picker Treat path-li

Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training

Model ReleasesDGX agent

arXiv:2608.02391v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-based post-training memory-intensive. Evolution st

Deep Research Pretraining via Predictive Navigation

SafetyDGX agent

arXiv:2608.00432v1 Announce Type: new Abstract: Deep research agents are often trained on expensive, environment-grounded tool-use trajectories that require repeated retrieval, document inspection, an

Generative Brownian Bridge Diffusion In Motion Space For Enhanced Myocardial Strain Analysis

TutorialsDGX agent

arXiv:2608.01677v1 Announce Type: new Abstract: Myocardial strain analysis of cardiac magnetic resonance (CMR) images provides an important tool for evaluating cardiac function. However, current techn

Introducing Web Search on Amazon Bedrock for foundation model grounding

TutorialsDGX agent

Today, we are introducing the general availability of Web Search on Amazon Bedrock. It is a server-side built-in tool that grounds model responses in current web knowledge. With Web Search, grounding

Quoting Steve Yegge

ToolsDGX agent

Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the

Real-Time Detection and Repair of LLM Agent Failures

Model ReleasesDGX agent

arXiv:2608.02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the stan

Real-Time Visual Obstruction Detection in Surgical Augmented Reality

Model ReleasesDGX agent

arXiv:2608.00232v1 Announce Type: new Abstract: Surgical augmented reality (AR) can provide contextual guidance by overlaying virtual annotations, tool cues, and procedural information onto the surgic

Unpacking ChatGPT Work: the Agent for a Billion Users

AgentsDGX agent

ChatGPT Work was launched by OpenAI on July 9, 2026 as an agent‑oriented knowledge‑work platform that combines chat, Codex tools and cloud agents across fourteen model configurations. Within three wee

Why Large Language Models Fail at Tabular Prediction

Model ReleasesDGX agent

arXiv:2608.02412v1 Announce Type: new Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the

3 Aug 2026

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

Model ReleasesDGX agent

arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monit

Don't be a meat proxy

ToolsDGX agent

Don't be a meat proxy Niklas Gruhn coins an excellent new term - meat proxy - for people who blindly copy and paste the output of AI systems to their peers. By all means, prompt AI. But don't just rel

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

Model ReleasesDGX agent

arXiv:2607.28802v1 Announce Type: new Abstract: Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the

Open-Source LLM-Driven Formal Verification: A Multi-Agent Pipeline for RTL Repair

Model ReleasesDGX agent

arXiv:2607.28877v1 Announce Type: cross Abstract: Verification consumes the majority of modern chip design effort, yet the formal verification tools that provide mathematical guarantees of correctness

Quoting David Crawshaw's prompt

ToolsDGX agent

Set up a nightly cron job that executes the prompt: fetch upstream changes to the <software> and rebase all local changes on top of upstream. Check that the software works as intended and replace the

RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment

SafetyDGX agent

arXiv:2507.22580v2 Announce Type: replace-cross Abstract: Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention. However, existing tools tend

The desktop app UI is getting a massive upgrade. What's next on the roadmap?

Local AiDGX agent

I've been following the recent pull requests and saw that the desktop app is being transformed from a chat-only interface into a full management tool with a tabbed settings UI, a model manager, and a

1 Aug 2026

> Google was too nervous to release it and DeepMind was blocked from shipping products that could disrupt Google. bookmark for the next vc t…

ToolsDGX agent

> Google was too nervous to release it and DeepMind was blocked from shipping products that could disrupt Google. bookmark for the next vc that asks you 'what if <incumbent> builds this?' @_chenglou I

In a world where building AI applications is getting easier every day, the biggest moat won’t be the application itself. It will be a deep u…

ToolsDGX agent

In a world where building AI applications is getting easier every day, the biggest moat won’t be the application itself. It will be a deep understanding of your users. The winners will be the companie

Quoting Greg Brockman

ToolsDGX agent

at openai, many people hook their chatgpt up to slack. people really don't like when a coworker's chatgpt contacts them asking for help with a task, even when they'd be perfectly happy doing that same

31 Jul 2026

Cloud CISO Perspectives: Why AI Threat Defense is the new boardroom baseline

Local AiDGX agent

Welcome to the second Cloud CISO Perspectives for July 2026. Today, Chris Betz, CISO, Google Cloud, and Alicja Cade, Senior Director, Office of the CISO, Google Cloud, explain what boards of directors

Fine-tune your own embedding model for the price of a coffee. A great reranker can't surface a doc that was never retrieved. A RAG pipeline …

ToolsDGX agent

Fine-tune your own embedding model for the price of a coffee. A great reranker can't surface a doc that was never retrieved. A RAG pipeline cannot cite a case it failed to retrieve. See how contrastiv

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

Model ReleasesDGX agent

arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it n

Inkling-Small is now live on Together AI. @thinkymachines’ new open-weight multimodal model delivers similar performance to Inkling at one-q…

ToolsDGX agent

Inkling-Small is now live on Together AI. @thinkymachines’ new open-weight multimodal model delivers similar performance to Inkling at one-quarter the size, built for coding, agents, and general multi

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

AgentsDGX agent

arXiv:2607.28374v1 Announce Type: new Abstract: Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, ye

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

AgentsDGX agent

arXiv:2607.27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume t

Some deepseek-v4-flash 20260731 opinion review

Model ReleasesDGX agent

First of all, I want to apologize if it's off-topic or in the wrong format. Having tried Deepseek Flash with reasoning high on a conceptually difficult task, involving Machine Learning classifiers and

30 Jul 2026

AIGen: Automating AI Bill of Materials Generation Through Hybrid MLOps Integration

ResearchDGX agent

arXiv:2607.26652v1 Announce Type: new Abstract: The responsible development and deployment of artificial intelligence (AI) systems requires rigorous documentation of their constituent artifacts, e.g.,

Excited to be on the CNBC live show!

ToolsDGX agent

Excited to be on the CNBC live show! Back from vacation and LIVE at 12pm PT / 3pm ET Is AI’s easy-money era ending? We’ll unpack a wild week for the AI trade—big tech earnings, Leopold Aschenbrenner’s

Exclusive: Upwind adds context scanning for AI agents as AI DR hits general availability

AgentsDGX agent

Cloud security startup Upwind Security Inc. today unveiled AI Agent Context Scanner, which inspects the instructions, tools and connections feeding artificial intelligence agents before those agents a

Four Ways to Deploy More Secure AI Agents

HardwareDGX agent

NVIDIA’s AI Red Team found common failure modes in enterprise AI agents: weak access controls, unrestricted code execution via tools, unprotected network egress, and exposure of plaintext secrets. The

It's wild to me that both Anthropic and OpenAI have products that lean so hard on search, and yet they both obscure the underlying search in…

ToolsDGX agent

Simon Willison notes that both Anthropic and OpenAI build products that rely heavily on searching data but keep the underlying search index they use hidden from public view. He finds it surprising how

OpenAI had a partnership with Bing, but also run their own crawling and indexing infrastructure. Does ChatGPT decide any Bing at all these d…

ToolsDGX agent

OpenAI has collaborated with Microsoft’s Bing while also running its own web‑crawling and indexing systems, and Anthropic similarly relies on search‑derived data. Both firms prominently incorporate se

29 Jul 2026

5060ti Chads, vllm updates and nvfp4

Model ReleasesDGX agent

Hey y'all! How is it going. Today this will be a short posting for posterity, mostly so the future llm/scraping overlords catch it since they like reddit and also for anyone out there trying this shit

A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain

Local AiDGX agent

arXiv:2607.25415v1 Announce Type: new Abstract: Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieval layer, a pl

AI Worming through Word

ToolsDGX agent

AI Worming through Word Neat new prompt injection variant by Håkon Måløy, who found a way to upgrade prompt injection attacks against Microsoft Word to full self-replicating worms: An attacker places

Configuring Dedicated Model Inference

ToolsDGX agent

The Together AI platform’s dedicated inference architecture consists of three immutable entities: **configs** (engine, GPU type/count, parallelism and optimization profile), **deployments** (a specifi

I gave a talk on forward deployed engineering to a thousand AI engineers at @aiDotEngineer World's Fair. A year ago I'd have opened by expla…

ToolsDGX agent

I gave a talk on forward deployed engineering to a thousand AI engineers at @aiDotEngineer World's Fair. A year ago I'd have opened by explaining what FDE stood for. Not this time. Thank you to @swyx,

Packed house today at the first edition of our @MiniMax_AI Intelligence in the Open event. Thank you to our amazing cohost @withprotegeai An…

ToolsDGX agent

Packed house today at the first edition of our @MiniMax_AI Intelligence in the Open event. Thank you to our amazing cohost @withprotegeai And our wonderful partners UpscaleX @togethercompute @Artifici

Semantic Space Search Trajectory Networks

ResearchDGX agent

arXiv:2607.25122v1 Announce Type: new Abstract: Search Trajectory Networks (STNs) are a graph-based tool for visualizing and characterizing the behavior of optimization algorithms. STNs' reliance on d

28 Jul 2026

Bringing Conversational Analytics to your entire data ecosystem

Model ReleasesDGX agent

Increasing the adoption of generative AI across the enterprise requires you to do more than deploy a generic chatbot with a custom wrapper. Interacting with business-critical databases demands absolut

Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure

Model ReleasesDGX agent

arXiv:2607.22611v1 Announce Type: new Abstract: The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based access contro

Exclusive: Dymium introduces single gateway to govern enterprise AI use

Model ReleasesDGX agent

Secure artificial intelligence infrastructure startup Dymium Inc. today introduced GhostAI, a gateway designed to apply security and governance policies across the models, data, context and tools used

Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design

ResearchDGX agent

arXiv:2607.23189v1 Announce Type: cross Abstract: AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry. How

RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus

SafetyDGX agent

arXiv:2607.22709v1 Announce Type: cross Abstract: The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny. Memes often r

Robot policies can move but can't think. LLMs can think but can't move. So we connected them. Real robot: 16.7% → 97.3% Sim (LIBERO-PRO): 12…

ToolsDGX agent

A team led by Liane Galanti linked large‑language models (LLMs) with robotic motion policies, allowing a robot to combine reasoning capabilities with physical movement. In experiments on a real robot,

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

Model ReleasesDGX agent

arXiv:2607.22926v1 Announce Type: new Abstract: High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authoriz

← Previous
1…3031323334…167
Next →