AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,030 results
4 Aug 2026

SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach

Model ReleasesDGX agent

arXiv:2608.00030v1 Announce Type: new Abstract: Specialised retrieval agents typically surface higher quality results than general-purpose search, but selecting the optimal agent for a given query rem

SPIRIT: Spatio-temporal Pairwise Relational Modeling of Instrument-Tissue Interactions for Surgical Action Triplet Recognition

SafetyDGX agent

arXiv:2608.02188v1 Announce Type: new Abstract: Fine-grained understanding of surgical activity is essential for context-aware assistance in the operating room, including safety monitoring, adverse ev

Start Classifying: Categorical Critics for LLM Reinforcement Learning

SafetyDGX agent

arXiv:2608.02181v1 Announce Type: new Abstract: Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets.

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Label Defines the Timescale: Trait-State Limits of Temporal-Aggregate Learning

ResearchDGX agent

arXiv:2608.01587v1 Announce Type: cross Abstract: Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows. Their a

TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

Model ReleasesDGX agent

arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high

UAV-Based Environmental Monitoring of Rip-Current Indicators Using Wavelet-Derived Texture Features

SafetyDGX agent

arXiv:2608.02448v1 Announce Type: new Abstract: Rip currents are recurrent coastal natural hazards that threaten beachgoers and create operational challenges for lifeguards and coastal managers. Relia

v0.32.6

Model ReleasesDGX agent

What's Changed Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically /v1/chat/completions streaming now matches OpenAI's wire format: rol

Visualising Information Flow in Word Embeddings with Diffusion Tensor Imaging

ResearchDGX agent

arXiv:2601.05713v2 Announce Type: replace Abstract: Understanding how large language models (LLMs) represent natural language is a central challenge in natural language processing (NLP) research. Many

Why are Gamers so incredibly hostile to AI? Is it just a tiny vocal minority that spreads such toxic vitriol online?

Model ReleasesDGX agent

It's more accurate to say that many highly engaged online gamers are hostile to AI, not that 'gamers' as a whole are. Gaming is a huge community with hundreds of millions of people, and opinions vary

Zero-Cost Virtual RNA: Approximating Immunotherapy Signatures via Cross-Modal WSI Retrieval

ResearchDGX agent

arXiv:2608.00544v1 Announce Type: new Abstract: Identifying the ``Inflamed'' immunophenotype in Gastric Adenocarcinoma predicts immunotherapy response but requires an expensive 10-gene RNA signature.

3 Aug 2026

A Model-Driven Approach for Developing Families of Reinforcement Learning Environments

Local AiDGX agent

arXiv:2606.20324v2 Announce Type: replace-cross Abstract: Virtual training environments are software-intensive systems in which reinforcement learning (RL) agents learn, adapt, and demonstrate meaning

Agentic Harness for Real-World Compilers

Model ReleasesDGX agent

arXiv:2603.20075v2 Announce Type: replace-cross Abstract: Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements enable

AI9Stars released G9v3-39A5B

Model ReleasesDGX agent

AI9Stars has released G9v3-39A5B an open weights language model designed to deliver even stronger reasoning capabilities than ai9stars/G9v3-3B with its 39B and 5 active experts. It is released under t

Application of machine learning to monster level prediction in tabletop RPG game design

ResearchDGX agent

arXiv:2607.09196v2 Announce Type: replace Abstract: Designing balanced adversaries is a central but labor-intensive task in tabletop role-playing game (TTRPG) development. In systems such as Pathfinde

Beyond Retrieval: Analytic Memory for Multimodal Agents

ResearchDGX agent

arXiv:2607.29440v1 Announce Type: new Abstract: Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions.

Code Is the Body: Agent-Owned Software Bodies for Recursive Evolution and Descent

AgentsDGX agent

arXiv:2607.28691v1 Announce Type: cross Abstract: Personalized AI agents are often configurable without giving users control over the artifacts that determine their future behavior. We present OurArk,

Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment

SafetyDGX agent

arXiv:2607.29593v1 Announce Type: new Abstract: This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equ

Devtools must be open source (exe.dev)

Model ReleasesDGX agent

My comment on Devtools must be open source (exe.dev) — Hacker News.One of the arguments for open source software for end-users has always been the freedom to examine and modify how that software works

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

SafetyDGX agent

arXiv:2607.29246v1 Announce Type: new Abstract: Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a

DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing

ApplicationsDGX agent

arXiv:2607.28750v1 Announce Type: cross Abstract: As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhead, and cros

Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml

ResearchDGX agent

arXiv:2602.15751v2 Announce Type: replace-cross Abstract: This paper presents an end-to-end demonstration of a viable, ultra-fast, radiation-hard machine learning (ML) application on FPGAs, which coul

Fisher Information, Training and Bias in Fourier Regression Models

SafetyDGX agent

arXiv:2510.06945v2 Announce Type: replace Abstract: Motivated by the growing interest in quantum machine learning, in particular quantum neural networks (QNNs), we study how recently introduced evalua

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Model ReleasesDGX agent

arXiv:2607.16057v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as

GPT-Live can listen while it speaks. To make that feel natural at ChatGPT scale, we rebuilt the voice stack from client to model. This new a…

AgentsDGX agent

GPT-Live can listen while it speaks. To make that feel natural at ChatGPT scale, we rebuilt the voice stack from client to model. This new architecture keeps audio flowing continuously, so deeper reas

HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution

AgentsDGX agent

arXiv:2607.13683v2 Announce Type: replace Abstract: Large Language Models (LLMs) have enabled capable agents across diverse applications. Beyond the foundation model, the performance of an agent is go

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

Model ReleasesDGX agent

arXiv:2607.29196v1 Announce Type: new Abstract: Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, track

I got tired of ad-filled mobile wrappers for Ollama, so I built PocketLLM Lite an open-source, offline Android client (Local GGUF, SKILL.md plugins, local RAG)

Local AiDGX agent

Hey, Like a lot of people here, I use local models via Ollama on my desktop/server and wanted a mobile client that actually felt responsive, worked offline, and respected privacy. Most apps on the Pla

// Model or Harness // Great paper if you are building with agents in production. (bookmark it) It organizes 41 agent failure modes by the i…

AgentsDGX agent

// Model or Harness // Great paper if you are building with agents in production. (bookmark it) It organizes 41 agent failure modes by the interaction they originate in. Each mode gets assigned to an

Nice benchmark to measure agentic e-commerce capabilities. They ran an agent for one simulated year of e-commerce operations and it ends up …

Model ReleasesDGX agent

Nice benchmark to measure agentic e-commerce capabilities. They ran an agent for one simulated year of e-commerce operations and it ends up with 27.3% of the money a human makes. MerchantBench is a 36

OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation

HardwareDGX agent

arXiv:2607.29266v1 Announce Type: cross Abstract: Artificial Intelligence (AI) and Deep Learning (DL) have notably advanced medical image analysis, yet many health- care organizations struggle to adop

SATViz: Real-Time Visualization of Clausal Proofs

ResearchDGX agent

arXiv:2209.05838v2 Announce Type: replace Abstract: Visual layouts of graphs representing SAT instances can highlight the community structure of SAT instances. The community structure of SAT instances

SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery

Model ReleasesDGX agent

arXiv:2607.29347v1 Announce Type: cross Abstract: Modern neuroscience relies on integrating multi-scale, multimodal datasets to uncover the neural principles underlying intelligence. However, analytic

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

SafetyDGX agent

arXiv:2607.29468v1 Announce Type: new Abstract: Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failures affect gra

StaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent

SafetyDGX agent

arXiv:2506.13862v2 Announce Type: replace-cross Abstract: In Reinforcement Learning (RL), regularization with a Kullback-Leibler divergence that penalizes large deviations between successive policies

Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions

ResearchDGX agent

arXiv:2607.28687v1 Announce Type: cross Abstract: As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades, yet routi

The result is a faster, more natural conversation with ChatGPT Voice from the moment a session starts. How we built it: https://openai.com/i…

Model ReleasesDGX agent

OpenAI has redesigned the ChatGPT Voice stack—from client to model—to enable continuous audio streaming, allowing GPT‑Live to listen while speaking without interruption. The new architecture supports

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

HardwareDGX agent

arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, w

two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apo…

Model ReleasesDGX agent

two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apology. i'm sorry that i was right about every single thing. a

Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?

Model ReleasesDGX agent

arXiv:2607.28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that treatment is

What Is Missing in Surgical Risk Stratification and Outcome Prediction: A Scoping Review of End-to-End Machine Learning Approaches

Model ReleasesDGX agent

arXiv:2607.29090v1 Announce Type: new Abstract: Postoperative adverse events, including mortality and morbidity, remain a major global burden, many of which are preventable through early identificatio

You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structured data directly from your PDFs: form fiel…

Model ReleasesDGX agent

You shouldn't need a vision model to know your PDF has checkboxes. LiteParse can now pull structured data directly from your PDFs: form field values, checkbox states, annotations, embedded images, vec

YouTuber Hank Green faces online criticism after using ChatGPT to help research a script, and says his LLM usage 'is not healthy for me or good for the world' (Anthony Ha/TechCrunch)

IndustryDGX agent

Anthony Ha / TechCrunch: YouTuber Hank Green faces online criticism after using ChatGPT to help research a script, and says his LLM usage “is not healthy for me or good for the world” — Hank Green, a

Zero-Mem: Zero-Token Memory Operations for LLM Agents

Local AiDGX agent

arXiv:2607.29377v1 Announce Type: new Abstract: LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermed

2 Aug 2026

Comfyui VRAM tracker

Model ReleasesDGX agent

Hello! VRAM tracker is a node that track the full memory lifecycle of a comfyui run: when each weight is reserved, paged into VRAM, computed on, evicted, and freed. It renders it as an interactive HTM

Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines…

Model ReleasesDGX agent

Cybersecurity isn’t a fortress problem, it’s an immunity problem. Think vaccines. Eliminating pathogen is not practically possible. Vaccines don’t eliminate pathogens. They teach the immune system to

Nine iterations of BabyAGI in three years, and yet the bit that @yoheinakajima kept coming back to was graphs. @aiDotEngineer published wher…

AgentsDGX agent

Nine iterations of BabyAGI in three years, and yet the bit that @yoheinakajima kept coming back to was graphs. @aiDotEngineer published where that landed, 'Active Graph Agent Runtime (BabyAGI 4)', on

Open letters about AI development

Model ReleasesDGX agent

Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and Ame

The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were …

HardwareDGX agent

The funny thing about Anthropic and OpenAI people saying they want to slow down AI progress is that this is what Anti Trust mechanisms were built for. It's illegal to collude and slow down AI progress

1 Aug 2026

A collection of small domain-specific benchmarks for local models (30+ and growing)

Model ReleasesDGX agent

Hello fellow local AI people! I took 'you must create your own benchmarks' literally, and built a website for this. How does the end result look like Let's say I want to know which model has most comm

A look at the deluge of AI computing power set to come online in the coming years; Epoch AI expects the number of AI chips in use to double every nine months (New York Times)

IndustryDGX agent

New York Times: A look at the deluge of AI computing power set to come online in the coming years; Epoch AI expects the number of AI chips in use to double every nine months — The milestones for artif

Hot take on OpenAI’s Astra: - Obviously impressive - But math is different from most other problems in that it is more amenable to to formal…

ApplicationsDGX agent

Hot take on OpenAI’s Astra: - Obviously impressive - But math is different from most other problems in that it is more amenable to to formal verification and synthetic data. How well it works in open-

Is there a point where models just cannot get any smaller without losing intelligence?

Model ReleasesDGX agent

DeepSeek V4 Flash got me thinking... We keep seeing smaller models get way better. A model at a certain parameter count today can be much smarter than a model of the same size from a year or two ago.

“Stochastic parrots” is not my term (it’s @emilymbender’s). But a lot of people today commenting on it are confused. To some extent (though …

Model ReleasesDGX agent

“Stochastic parrots” is not my term (it’s @emilymbender’s). But a lot of people today commenting on it are confused. To some extent (though I don’t think it’s a perfect metaphor, and have said that be

What's currently the 'smartest' LLM to use on 8GB vram and 16 RAM and same thing for 8 VRAM and 64 RAM?

Model ReleasesDGX agent

Been trying to find something that actually handles my workload well instead of just being 'fine.' Started on Qwen 2.5 7B, moved to Qwen 3 8B, and right now I'm using Nemotron 3 Ultra (the big 550B on

31 Jul 2026

A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response

SafetyDGX agent

arXiv:2607.27597v1 Announce Type: new Abstract: Recent advances in Vision Language Models (VLMs) have created new opportunities for disaster response, where responders must interpret large volumes of

AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages

SafetyDGX agent

arXiv:2505.10300v2 Announce Type: replace-cross Abstract: Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle through

Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer

Local AiDGX agent

arXiv:2607.17100v2 Announce Type: replace-cross Abstract: Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judged

AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

Model ReleasesDGX agent

arXiv:2607.27845v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability

Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance

AgentsDGX agent

arXiv:2607.27283v1 Announce Type: new Abstract: Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself ex

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

Model ReleasesDGX agent

arXiv:2607.27747v1 Announce Type: new Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either

← Previous
1…107108109110111…168
Next →