AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,564 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
12 Aug 2026We have her dash camera which shows they are lying. The agents are wearing body cameras and should have dash cameras of their own. If what t…

We have her dash camera which shows they are lying. The agents are wearing body cameras and should have dash cameras of their own. If what they say happened was true they wouldn’t be issuing statement

→12 Aug 2026Situation Graph Prediction for User Perspective Modeling

arXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data

3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→12 Aug 2026Grok Bot

Grok Bot Here's my Grok Bot team: - Webby: Web designer - Shotry: Short-form content creator - Writey: Article/Newsletter writer - Claude Code: Grok agent that specializes in CC - Codex: Same as the a

→12 Aug 2026GitSkills: A Dataset of Agent Skills on GitHub

arXiv:2608.10906v1 Announce Type: cross Abstract: An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference fi

→12 Aug 2026DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

arXiv:2608.10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coor

→12 Aug 2026DHS has lied about these incidents over and over again and been caught in the lies over and over again. They are lying now, and nobody shoul…

DHS has lied about these incidents over and over again and been caught in the lies over and over again. They are lying now, and nobody should believe them. NOW: DHS response to Virginia woman's video

→12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses

arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these

→12 Aug 2026AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting

arXiv:2608.09959v1 Announce Type: cross Abstract: AI weather models are in the process of revolutionising weather forecasting. While these models have been shown to achieve superior performance to phy

CompanyOpenAI8 recent entries
11 Aug 2026The small open weight models are scarier in AI development

Imagine if your everyday laptop could run an AI model smart enough to take care of 90% of your work—totally private, lightning fast, and completely free of monthly fees. That is the exact tipping poin

→11 Aug 2026Stealing Reasoning Traces from Proprietary LLM APIs

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

→11 Aug 2026Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to give enterprise AI capability options

Artificial intelligence silicon and software giant Nvidia Corp. today announced two new services: a highly customizable Nemotron model and an agentic AI model router named NeMo Switchyard. As enterpri

→11 Aug 2026Multi-modal Interactive Control of Robotic Arm based on Offline Large Language Models

arXiv:2608.08183v1 Announce Type: new Abstract: Large Language Models (LLMs) have significantly revolutionized the modern society with numerous advanced interactions between humans and AI agents, wher

→11 Aug 2026I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

→11 Aug 2026Epic talk: the cheat code for how to build your own in-house lab, featuring @gabepereyra of @harvey

Epic talk: the cheat code for how to build your own in-house lab, featuring @gabepereyra of @harvey Want world class research capabilities, but don’t have the resources of a big lab? At our recent Sov

→11 Aug 2026Curriculum Generation under Structured Parametric Environments for Robust Navigation Policies

arXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f

→12 Aug 2026Grok is now an AI ‘teammate’ you can assign work

SpaceXAI has introduced Grok Bot, an always-on AI agent service designed to behave like independent 'AI teammates' that can do your work for you. The bots share their own cloud-based computer environm

CompanyGoogle8 recent entries
11 Aug 2026Google’s Gemini AI app passes 1 billion monthly active users

Google LLC’s Gemini artificial intelligence app has passed 1 billion monthly active users, making it the 14th product in the company’s history to reach that mark. The company announced the milestone t

→11 Aug 2026Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on

→11 Aug 2026BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

→11 Aug 2026An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.

→11 Aug 2026360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

arXiv:2608.08814v1 Announce Type: cross Abstract: We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment construc

→12 Aug 2026Situation Graph Prediction for User Perspective Modeling

arXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data

→12 Aug 2026Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents

arXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure be

→12 Aug 2026Google unveils the $399 Pixel Watch 5 with a satin pyrite case finish, offline Gemini, proactive AI suggestions, better GPS maps, and insulin resistance trends (Victoria Song/The Verge)

Victoria Song / The Verge: Google unveils the 399 Pixel Watch 5 with a satin pyrite case finish, offline Gemini, proactive AI suggestions, better GPS maps, and insulin resistance trends — The 399 Goog

CompanyMeta8 recent entries
12 Aug 2026SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents

arXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over t

→12 Aug 2026Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory

arXiv:2608.11066v1 Announce Type: cross Abstract: We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a future-acce

→12 Aug 2026New Muse-Glimmer-30B SoTA Quants - hopefully a new lineup :)

Hey Folks, I've been making quants for a while - recently I took a short break to get into hardcore research (submitted my first EMNLP paper during it!). Along the way, I built up a little arsenal of

→12 Aug 2026Muse Glimmer is live on Fireworks. The new open-weight model from Meta Superintelligence Labs is a 30B dense model built for always-on agent…

Muse Glimmer is live on Fireworks. The new open-weight model from Meta Superintelligence Labs is a 30B dense model built for always-on agents that reason across many sequential tool calls and can reco

→12 Aug 2026MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph

arXiv:2608.10504v1 Announce Type: new Abstract: As coding agents increasingly handle implementation, the central challenge shifts from building individual agents to building an infrastructure that sys

→12 Aug 2026How OneAdvanced deployed over 50 AI agents on UK-sovereign AWS

Learn how OneAdvanced, a UK enterprise software provider, built a UK-sovereign AI platform by self-hosting Llama 4 Maverick and Llama Guard 4 on Amazon SageMaker AI, with a RAG pipeline on pgvector an

→12 Aug 2026Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief Systems

arXiv:2509.05624v3 Announce Type: replace-cross Abstract: How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach

→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique

arXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions

CompanyMistral8 recent entries
4 Aug 2026GPT-OSS has turned one year old today!

It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that

→5 Aug 2026Prime Agent - a new coding harness surpassing Codex/CC/PI

Prime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficien

→5 Aug 2026Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models

French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperform

→5 Aug 2026Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all…

Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all document processing solutions - just screenshot the page an

→6 Aug 2026The death of SLMs?

I love to see these impressive models coming out that compete with the giants from companies like Z.ai, Moonshot, Alibaba, etc. A win for the open source/weight community is always welcome. While I am

→6 Aug 2026nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging Face

NVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blu

→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

→11 Aug 2026💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to c…

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to choose the right model for the right task. As part of this, we

CompanyxAI8 recent entries
10 Aug 2026Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration

arXiv:2608.06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting gener

→11 Aug 2026The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games

arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move fro

→11 Aug 2026SpaceXAI rolls out Grok Bot AI agent app in beta on Mac, iOS, Windows, and Linux, initially for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium users (Zac Hall/9to5Mac)

Zac Hall / 9to5Mac: SpaceXAI rolls out Grok Bot AI agent app in beta on Mac, iOS, Windows, and Linux, initially for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium users — SpaceXAI and Cursor

→11 Aug 2026Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models

arXiv:2512.11614v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuristic

→11 Aug 2026Anyone Using (Koreas) 'Solar Open 2' (250B, 15B) Model?

I just heard of this model. Seems to be a competitor to DeepSeek V4 Flash. About the same size and active parameters. Anyone tested it compared to V4 Flash? Link: https://huggingface.co/upstage/Solar-

→11 Aug 20261 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-cases

A few things right off the bat: it reasons very efficiently. Like Grok 4.5 levels of efficient thinking it quantizes very well. My first few tests with iq3_xxs were better than Qwen/Gemma behaved at t

→12 Aug 2026Grok is now an AI ‘teammate’ you can assign work

SpaceXAI has introduced Grok Bot, an always-on AI agent service designed to behave like independent 'AI teammates' that can do your work for you. The bots share their own cloud-based computer environm

→12 Aug 2026Grok Bot

Grok Bot Here's my Grok Bot team: - Webby: Web designer - Shotry: Short-form content creator - Writey: Article/Newsletter writer - Claude Code: Grok agent that specializes in CC - Codex: Same as the a

CompanyDeepSeek8 recent entries
11 Aug 2026I ran Muse Glimmer @ 1M context - All tests passed.

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

→11 Aug 2026I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon

→11 Aug 2026Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses

arXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harne

→11 Aug 2026DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO…

DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO, then deploy the fine-tuned model on Together AI for produc

→11 Aug 2026DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integration

Spent two nights getting deepseek-ai/DeepSeek-V4-Flash-0731 (284B MoE, 13B active, native FP4/FP8, 1M context) running production-grade on two DGX Sparks connected by one QSFP DAC cable. Everything —

→11 Aug 2026Anyone Using (Koreas) 'Solar Open 2' (250B, 15B) Model?

I just heard of this model. Seems to be a competitor to DeepSeek V4 Flash. About the same size and active parameters. Anyone tested it compared to V4 Flash? Link: https://huggingface.co/upstage/Solar-

→12 Aug 2026Persistent Recursive Worlds Enable Autonomous Software Evolution

arXiv:2608.10450v1 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve conti

→12 Aug 2026Idea for a deepseek-v4-flash-0731 backed automated research workflow to be leveraged via qwen3.6/3.8 27b for difficult tasks that require highly technical, not easy to find information.

Sometimes you have tasks that are outside of your expertise and the idea is this workflow automation could be leveraged to manage to have local AI figure it out using research from his workflow gather

CompanyNVIDIA8 recent entries
11 Aug 2026Multi-tier storage rewrites the economics of AI inference

As inference becomes the dominant workload in AI infrastructure, multi-tier storage architectures are emerging as a key method for cost control and enhanced performance. These architectures combine fl

→11 Aug 2026Looking for a faster specialized model for your Agent Work? @nvidia Nemotron 3.5 Lightning (30B MoE, 3B active params) is now live on Firewo…

Looking for a faster specialized model for your Agent Work? @nvidia Nemotron 3.5 Lightning (30B MoE, 3B active params) is now live on Fireworks. It’s distilled from NVIDIA Nemotron 3 Ultra to be your

→11 Aug 2026It's been exciting for Ollama to partner with @JensenHuang and the @NVIDIAAI team on launching open models. Open models have no boundaries, …

It's been exciting for Ollama to partner with @JensenHuang and the @NVIDIAAI team on launching open models. Open models have no boundaries, and let's continue to work together to make this ecosystem b

→11 Aug 2026I ran Muse Glimmer @ 1M context - All tests passed.

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

→11 Aug 2026I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Everything started with the sudden death of my old ASRock J1900. While looking for the perfect ITX replacement, I stumbled upon the Chinese CW-NAS-ADLN-K motherboard, which looked perfect on paper: In

→11 Aug 2026FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

arXiv:2607.18171v2 Announce Type: replace Abstract: Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose effici

→12 Aug 2026Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work

Ran the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quali

→12 Aug 2026Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4

arXiv:2608.10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matrix loads wit