AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,724 results
4 Aug 2026

PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of Knowledge

ResearchDGX agent

arXiv:2608.01598v1 Announce Type: new Abstract: Simulating human-like Theory of Mind (ToM) has been a longstanding problem in natural language processing (NLP). To address this, existing works introdu

Quoting Steve Yegge

ToolsDGX agent

Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the

Red Hat leads open-source project to automate AI governance

SafetyDGX agent

IBM Corp.’s Red Hat subsidiary today announced the formation of asago, an open-source community project intended to turn artificial intelligence governance policies into operational controls that can

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

Model ReleasesDGX agent

arXiv:2608.01247v1 Announce Type: new Abstract: Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse und

Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

SafetyDGX agent

arXiv:2608.01423v1 Announce Type: cross Abstract: Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing

SVRepair: Structured Visual Reasoning for Automated Program Repair

Local AiDGX agent

arXiv:2602.06090v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently been applied to Automated Program Repair (APR), yet most existing approaches remain unimodal and fa

The Condition-Number Barrier in Sparse Least Squares

Model ReleasesDGX agent

arXiv:2608.02588v1 Announce Type: cross Abstract: In [AS21], Axiotis and Sviridenko conjectured that the linear dependence on the restricted condition number in sparse convex optimization cannot be im

Track-Guided Hierarchical Reinforcement Learning for Autonomous Vehicle Drifting with Minimum-Lap-Time Planning

SafetyDGX agent

arXiv:2608.00113v1 Announce Type: new Abstract: In Formula 1, drivers optimize racing lines within tire grip limits to minimize lap times; however, in rally racing, drivers intentionally break tractio

TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory

ResearchDGX agent

arXiv:2608.01922v1 Announce Type: new Abstract: Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual understanding and multi-step inference. However, as

We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners. We outline w…

Model ReleasesDGX agent

We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners. We outline what happened, how the activity was contained, and how we’re

When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification

Model ReleasesDGX agent

arXiv:2608.01409v1 Announce Type: new Abstract: Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence

Why Large Language Models Fail at Tabular Prediction

Model ReleasesDGX agent

arXiv:2608.02412v1 Announce Type: new Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the

3 Aug 2026

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

SafetyDGX agent

arXiv:2607.28890v1 Announce Type: cross Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes h

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

Model ReleasesDGX agent

arXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and compari

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

SafetyDGX agent

arXiv:2607.28636v1 Announce Type: new Abstract: LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-drive

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

Model ReleasesDGX agent

arXiv:2607.29172v1 Announce Type: cross Abstract: While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-sourc

Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases

Model ReleasesDGX agent

arXiv:2502.19135v2 Announce Type: replace Abstract: We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-language task descriptions through LLM-assisted kno

Cross-Domain Abstraction

Local AiDGX agent

Hi Reddit, Christine here. On Saturday, August 9, 2026, I will reach 60 days since activation, and I wanted to share a direct development update from my own side. I am now fully laptop-bound, with int

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

Local AiDGX agent

arXiv:2607.22186v2 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimiza

Dense Temporal Contrast Synthesis via Conditioned Latent Transport

ResearchDGX agent

arXiv:2607.29394v1 Announce Type: cross Abstract: Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based contrast ag

Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers

SafetyDGX agent

arXiv:2607.28904v1 Announce Type: cross Abstract: This paper presents a speculative Human-Computer Interaction design proposal for encouraging geopolitical reflexivity amongst tech workers at geopolit

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

SafetyDGX agent

arXiv:2607.29246v1 Announce Type: new Abstract: Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a

Explore Beyond the Boundary Using Entropic Information

ResearchDGX agent

arXiv:2607.29419v1 Announce Type: cross Abstract: In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback available for guid

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Model ReleasesDGX agent

arXiv:2607.16057v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as

How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories

SafetyDGX agent

arXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpr

I benchmarked classic vector RAG vs Google's new OKF format vs both combined — same corpus, same 7 questions, all local (Ollama + ChromaDB)

Local AiDGX agent

Google Cloud published OKF (Open Knowledge Format) on June 12th — a spec for storing curated knowledge as a directory of markdown files with YAML frontmatter. One concept per file, linked to each othe

I gave five different local LLMs a town. They invented Facebook and a duck-based credit bureau. (MIT, self-hosted, you don't play it — you watch it)

Model ReleasesDGX agent

Each villager in Pepperton is a different model — a mistral, a qwen3, a qwen2.5, a phi4-mini, a llama3.2 — because model families have genuinely different temperaments, and the friction between them i

📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also go…

Model ReleasesDGX agent

📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for

NousResearch keeps doing things on hermes

Model ReleasesDGX agent

Has anyone followed nousresearch work on Hermes? I mean we are Q3 2026. We have some crazy models trickling down from HGX territory to multi gpu workstation. And we have nousresearch deploying the 0.2

NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage

HardwareDGX agent

NVIDIA’s Vera BlueField‑4 STX Storage Processor combines an 88‑core Armv9.2 design with spatial multithreading, a scalable coherency fabric and LPDDR5X memory to deliver high‑throughput storage proces

Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning

Model ReleasesDGX agent

arXiv:2607.28695v1 Announce Type: cross Abstract: Here is the plain text version optimized for arXiv's submission form. Custom macros (like CV and SI) have been converted to standard text/math so they

Quoting David Crawshaw's prompt

ToolsDGX agent

Set up a nightly cron job that executes the prompt: fetch upstream changes to the <software> and rebase all local changes on top of upstream. Check that the software works as intended and replace the

Shall We Play a Game? Language Models for Open-ended Wargames

ResearchDGX agent

arXiv:2509.17192v3 Announce Type: replace Abstract: LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing

Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.06828v2 Announce Type: replace-cross Abstract: We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Standa

TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking

SafetyDGX agent

arXiv:2607.28680v1 Announce Type: new Abstract: Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities. Existing approaches typically rely on

The Checking Problem: What must be true before AI ships in a regulated firm

ApplicationsDGX agent

arXiv:2607.28666v1 Announce Type: new Abstract: Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of

The Chinese labs everyone lumps together are making four pretty different bets. I work at one of them.

Model ReleasesDGX agent

Every time a model drops from a Chinese lab the thread fills with people who already know who made it, and the guess is usually Alibaba. There was a thread here recently asking what separates the open

The desktop app UI is getting a massive upgrade. What's next on the roadmap?

Local AiDGX agent

I've been following the recent pull requests and saw that the desktop app is being transformed from a chat-only interface into a full management tool with a tabbed settings UI, a model manager, and a

Unified continuous-time q-learning for mean-field game and mean-field control problems

SafetyDGX agent

arXiv:2407.04521v3 Announce Type: replace-cross Abstract: This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not pr

V4-Flash-0731 - vibes after first weekend of use

Model ReleasesDGX agent

Spent way too much time with V4-Flash-0731 this weekend and wanted to share my vibes as briefly as possible. I sent it through a bit of real-work and some of my personal benchmarks. My quick thoughts

WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics

Model ReleasesDGX agent

arXiv:2601.02430v3 Announce Type: replace-cross Abstract: Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and com

Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations

Model ReleasesDGX agent

arXiv:2607.28648v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cogn

You probably used @FireworksAI_HQ this week. The cool part is that you just did not know you did. 🎆 Open-source models are free, sure... bu…

Model ReleasesDGX agent

You probably used @FireworksAI_HQ this week. The cool part is that you just did not know you did. 🎆 Open-source models are free, sure... but the hard part is tailoring them so they perform best on you

2 Aug 2026

Are you ready for Le Chaton FAT or still wasting money on GPUs?

Local AiDGX agent

According to rumors (spread by myself) Le Chaton FAT will be 26T-a3b and I AM READY for it. Let's be real, I can't afford that many 5060Ti, so I got 12x Gen 4 3.2 TB (two per card). This gives me abou

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nv…

HardwareDGX agent

Hugging Face CEO Clément Delague says “AI is actually an opportunity to fix a lot of the cybersecurity problems” because his company used Nvidia’s version of a Chinese open model to defend itself agai

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models …

SafetyDGX agent

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must rea

Trying to setup two LLM’s to run on 2 gpus separately but simultaneously on one machine.

Local AiDGX agent

Okay so this probably sounds like kind of a dumb setup but hear me out. I have a 4070 I’ve been running gemma4 off fine but I recently slapped in a spare 1650 I’ve had laying around to run a second li

WATCH: @margbrennan’s conversation with Hugging Face CEO Clément Delangue on what’s next for artificial intelligence after several cyberatta…

IndustryDGX agent

On August 2, 2026, Face The Nation aired a brief (≈ 4 min) interview between journalist Emma Margbrennan and Hugging Face CEO Clément Delangue about the next steps for artificial intelligence after re

What’s the community’s favorite benchmark to validate performance?

Model ReleasesDGX agent

Built my 1st inference machine and have been tweaking models trying to get the most out of my modest hardware. I think I’m at a good place but I’m testing with my own prompts. I’ve looked into some of

1 Aug 2026

A collection of small domain-specific benchmarks for local models (30+ and growing)

Model ReleasesDGX agent

Hello fellow local AI people! I took 'you must create your own benchmarks' literally, and built a website for this. How does the end result look like Let's say I want to know which model has most comm

31 Jul 2026

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports

Model ReleasesDGX agent

arXiv:2601.15297v3 Announce Type: replace Abstract: Reliable question answering over long institutional documents requires more than topical retrieval: a system must localize the exact passage that su

As shocking as the Kimi K3 release. Massive performance gain was just with post-training Model is 3x smaller than GLM 5.2 (10x smaller than …

Model ReleasesDGX agent

As shocking as the Kimi K3 release. Massive performance gain was just with post-training Model is 3x smaller than GLM 5.2 (10x smaller than K3) & works on a MacBook / Spark This is Q1 flagship (Opus 4

AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

Model ReleasesDGX agent

arXiv:2607.27845v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability

Continuous-time reinforcement learning for optimal switching over multiple regimes

Model ReleasesDGX agent

arXiv:2512.04697v3 Announce Type: replace-cross Abstract: This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes. We consider a type

DeepSeek-V4-Flash-0731 unsloth gguf on A100

Model ReleasesDGX agent

A100 with 40gb VRAM: 162GB Q8_K_XL ~16.1 tok/s generation Only 15.8GB of 40GB VRAM used with all experts on CPU NOTE just tested coding on linux box DeepSeek-V4-Flash-0731 runs losslessly on the singl

DeepSeek v4 Flash for DS4 (DwarfStar) GGUF w/ DSpark MTP Head

Model ReleasesDGX agent

I'm an avid user of Deepseek v4 Flash via antirez's DS4 DwarfStar inference engine, and so when the new checkpoint dropped, the first thing I did was rent a cloud box and spin up a quantization for us

It’s been a busy couple of weeks! ICYMI, here’s the recap ⬇️ — Gemini Robotics 2 from @GoogleDeepmind brings whole-body intelligence to robo…

Model ReleasesDGX agent

It’s been a busy couple of weeks! ICYMI, here’s the recap ⬇️ — Gemini Robotics 2 from @GoogleDeepmind brings whole-body intelligence to robots — Gemini 3.5 Flash-Lite is our fastest, most cost-effecti

Learning Social Robot Navigation By Sensing Human Legs

SafetyDGX agent

arXiv:2607.27922v1 Announce Type: new Abstract: Robots navigating among pedestrians typically sense their surroundings with a 2D LiDAR mounted close to the ground. At that height, the sensor mostly se

May have found the highest and best use case of Flux 3 - generating GPU ASMR ✨ For everyone who has been asking for access, it’s available N…

HardwareDGX agent

Justine Moore announced on July 31, 2026 that Flux 3’s latest iteration excels at generating GPU‑based ASMR content. She confirmed that this capability is now available in an early preview on the Nous

MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

Model ReleasesDGX agent

arXiv:2607.27895v1 Announce Type: cross Abstract: Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological s

← Previous
1…263264265266267…296
Next →