AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
Model Releases

LaCy: What Small Language Models Can and Should Learn is Not Just a Question of Loss

DGX agent

This paper was accepted at the Workshop on Memory for LLM-Based Agentic Systems at ICLR. Language models have consistently grown to compress more world knowledge into their parameters, but the knowled

model-releasesapple-ml-research
9 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Tools

live now: @steipete on the State of the Claw link in qt

DGX agent

live now: @steipete on the State of the Claw link in qt 🎥🔴 LIVESTREAM ANNOUNCEMENT: Keynotes + @openclaw / Personal Agents track 🔔 Hit the bell on our YT channel 👉 https://www.youtube.com/watch?v=O_IM

toolsswyx--x
9 Apr 2026
Model Releases

you'll need to explicitly prompt Claude Code to use it, but the Monitor Tool is super powerful e.g. 'start my dev server and use the Monitor…

DGX agent

you'll need to explicitly prompt Claude Code to use it, but the Monitor Tool is super powerful e.g. 'start my dev server and use the MonitorTool to observe for errors' Thrilled to announce the Monitor

model-releasesthariq--x
9 Apr 2026
Model Releases

Google's Gemma 4 is pretty wild. You can now run it locally with OpenClaw in 3 steps. 1. Install Ollama 2. Pull Gemma 4 model 3. Launch Open…

DGX agent

Google's Gemma 4 is pretty wild. You can now run it locally with OpenClaw in 3 steps. 1. Install Ollama 2. Pull Gemma 4 model 3. Launch OpenClaw with Gemma as the backend Private local AI agents in mi

model-releasesollama--x
8 Apr 2026
Research

Happening in 15 minutes! Come meet our CTO, @theemozilla

DGX agent

Happening in 15 minutes! Come meet our CTO, @theemozilla You have heard of @openclaw competitor from @NousResearch called “Hermes.” Tomorrow at 4 pm we will get nerdy with @theemozilla. Live. I will g

researchnous-research--x
8 Apr 2026
Model Releases

The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything …

DGX agent

The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlate

model-releasesfrancois-chollet--x
8 Apr 2026
Tools

The top 10 teams will present their working business live on June 9. Register: https://www.perplexity.ai/computer/a/the-billion-dollar-build…

DGX agent

Perplexity AI has launched the 'Billion Dollar Build,' an eight-week competition inviting founders to use its agentic AI system, Perplexity Computer, to build a company with a path to a $1 billion...

toolsperplexity--x
8 Apr 2026
Model Releases

Trying to DIY your own document parser by screenshotting into a frontier VLM (Opus, 5.4, Gemini) carries when you try to scale it up into pr…

DGX agent

Trying to DIY your own document parser by screenshotting into a frontier VLM (Opus, 5.4, Gemini) carries when you try to scale it up into production workflows. Here are two edge cases we've observed:

model-releasesjerry-liu--x
8 Apr 2026
Model Releases

Curious about vibe coding? Or are you already shipping apps and just want an easier way to explain your new favorite hobby to your friends, …

DGX agent

Curious about vibe coding? Or are you already shipping apps and just want an easier way to explain your new favorite hobby to your friends, parents, grandparents, etc.? Either way, this video is for y

model-releasesgoogle-ai--x
7 Apr 2026
Model Releases

got opencode working in taugentic, so now i can use glm-5.1 with my coding plan.

DGX agent

User Kevin Kern reported successfully integrating OpenCode into Taugentic (an AI-powered browser/agent environment), enabling use of GLM-5.1 via Z.ai's GLM Coding Plan. GLM-5.1 is Z.ai's next-gener...

model-releaseszhipu-ai--x
7 Apr 2026
Model Releases

Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework

DGX agent

arXiv:2608.11891v1 Announce Type: cross Abstract: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Asse

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

DGX agent

arXiv:2608.12262v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a fre

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

DGX agent

arXiv:2608.11210v1 Announce Type: new Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Grok 4.6 ranks #1 on the GPQA Diamond leaderboard 🧠 Grok 4.6 (high) scores 95% - the highest score on the chart for graduate-level scientif…

DGX agent

Grok 4.6 ranks #1 on the GPQA Diamond leaderboard 🧠 Grok 4.6 (high) scores 95% - the highest score on the chart for graduate-level scientific reasoning It outperforms Claude Fable 5, Opus 5, GPT-5.6 S

model-releaseselon-musk--x
13 Aug 2026
Safety

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

DGX agent

arXiv:2608.12063v1 Announce Type: cross Abstract: Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely

safetyarxiv-cs-ai
13 Aug 2026
Model Releases

OEIS Open: How many conjectures can language models turn into theorems?

DGX agent

arXiv:2608.11941v1 Announce Type: new Abstract: We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conje

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

out of curiosity, i tasked fable to figure this thing out and it looks like the 'saas killer' is done. ▲ it maxed out Fable 5 + used most of…

DGX agent

out of curiosity, i tasked fable to figure this thing out and it looks like the 'saas killer' is done. ▲ it maxed out Fable 5 + used most of Codex limits ▲ most of the work was done within 45h session

model-releasesswyx--x
13 Aug 2026
Applications

To use Computer History, opt in under Settings → Integrations in the ChatGPT desktop app on Mac. Rolling out globally now to Pro, Business, …

DGX agent

To use Computer History, opt in under Settings → Integrations in the ChatGPT desktop app on Mac. Rolling out globally now to Pro, Business, and Enterprise users, with access in the EEA, UK, and Switze

applicationsopenai--x
13 Aug 2026
Model Releases

Try Grok 4.6

DGX agent

Try Grok 4.6 Grok 4.6 wins again. 👑 Grok 4.6 takes the #1 spot on GPQA Diamond with a score of 94.9%, beating GPT-5.6, Gemini 3.1 Pro, Claude Opus 5, and every other model tested by Artificial Analysi

model-releaseselon-musk--x
13 Aug 2026
Safety

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

DGX agent

arXiv:2608.11410v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Err

safetyarxiv-cs-lg
13 Aug 2026
Model Releases

A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

DGX agent

arXiv:2608.10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

An adaptive and evolvable deep reinforcement learning framework for weather prediction

DGX agent

arXiv:2608.09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times. Rather than building yet another architecture, we reframe the for

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

DGX agent

arXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

DGX agent

arXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri

model-releasesarxiv-cs-cl
12 Aug 2026
Safety

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

DGX agent

arXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world mod

safetyarxiv-cs-lg
12 Aug 2026
Safety

ELMER: Evolutionary Language Model that Explores and Refines

DGX agent

arXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

DGX agent

arXiv:2608.10679v1 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are wor

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition

DGX agent

arXiv:2608.10396v1 Announce Type: new Abstract: Transforming table-form documents into machine-processable records requires recovering not only their visible content but also the multilevel structure

model-releasesarxiv-cs-cv
12 Aug 2026
Hardware

How to Choose Full-Stack Observability for NVIDIA AI Factories

DGX agent

A full‑stack observability framework for NVIDIA AI factories links telemetry from compute, networking, storage, orchestration and application layers using specialized tools (DCGM, NVSM, UFM, NetQ, NMX

hardwarenvidia-developer
12 Aug 2026
Safety

IO Factory: Simulating AI-Enabled Influence Campaigns at Scale

DGX agent

arXiv:2608.10920v1 Announce Type: new Abstract: We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

MindTopo reveals VLMs’ spatial reasoning abilities

DGX agent

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post

model-releasesmicrosoft-research
12 Aug 2026
Model Releases

Quoting Florian Herrengt

DGX agent

But then users start to report a weird bug. It's the 4th time your team has been trying to fix it. I mean... asking AI to fix it. Unfortunately, it seems like not even Fable can figure it out. You go

model-releasessimon-willison
12 Aug 2026
Model Releases

Qwen 3.8 2.4T is out , no 27b today RIP.

DGX agent

i was hoping they would release both today but the big one just dropped now , and it says its been listed since 5 hours ago on huggingface. I guess the real date for 27b is https://modelscope.cn/model

model-releasesr-localllama
12 Aug 2026
Model Releases

Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator

DGX agent

arXiv:2608.09971v1 Announce Type: cross Abstract: Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models whose medium-

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Sources detail moves behind Google's AI reshuffle; Sergey Brin urged key staff to go all in on Gemini, and some teams shifted from DeepMind to corporate Google (Kenrick Cai/Reuters)

DGX agent

Kenrick Cai / Reuters: Sources detail moves behind Google's AI reshuffle; Sergey Brin urged key staff to go all in on Gemini, and some teams shifted from DeepMind to corporate Google — Google co-found

model-releasestechmeme
12 Aug 2026
Model Releases

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

DGX agent

arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached o

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

V-FiLLM: Verified Financial LLM Reasoning Benchmark

DGX agent

arXiv:2608.11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains compar

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑‍🔬 , our effort to create the most comprehensive, schema-guided, real-world document …

DGX agent

We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑‍🔬 , our effort to create the most comprehensive, schema-guided, real-world document extraction benchmark. It’s extremely detailed and covers every

model-releasesjerry-liu--x
12 Aug 2026
Model Releases

Why does CLAUDE.MD keep growing? If you maintain a CLAUDE.md or an AGENTS.md, this one is worth your time. (bookmark it) This work traces wh…

DGX agent

Why does CLAUDE.MD keep growing? If you maintain a CLAUDE.md or an AGENTS.md, this one is worth your time. (bookmark it) This work traces why these files grow without bound. Appending an instruction i

model-releasesdair-ai--x
12 Aug 2026
Safety

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

DGX agent

arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream dec

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

DGX agent

arXiv:2607.10180v2 Announce Type: replace-cross Abstract: We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. T

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

DGX agent

arXiv:2608.07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist en

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Anyone Using (Koreas) 'Solar Open 2' (250B, 15B) Model?

DGX agent

I just heard of this model. Seems to be a competitor to DeepSeek V4 Flash. About the same size and active parameters. Anyone tested it compared to V4 Flash? Link: https://huggingface.co/upstage/Solar-

model-releasesr-localllama
11 Aug 2026
Model Releases

Back to the Future: A workbook time machine for spread sheet creation benchmarks

DGX agent

arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived obj

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Blumira launches Hearth, an AI command center that spans rival security tools

DGX agent

Security operations platform startup Blumira Inc. today launched Hearth, a vendor-agnostic artificial intelligence command center that lets security teams investigate and act across their existing too

model-releasessiliconangle
11 Aug 2026
Model Releases

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

DGX agent

arXiv:2608.08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the 'right' answer, but how they decide what matters most when moral principle

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integration

DGX agent

Spent two nights getting deepseek-ai/DeepSeek-V4-Flash-0731 (284B MoE, 13B active, native FP4/FP8, 1M context) running production-grade on two DGX Sparks connected by one QSFP DAC cable. Everything —

model-releasesr-localllama
11 Aug 2026
Model Releases

EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility

DGX agent

arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity

model-releasesarxiv-cs-ai
11 Aug 2026
← Previous
1…300301302303304…371
Next →