AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “openai”

GridTimelineEvolution
2,585 results
11 Aug 2026

Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to give enterprise AI capability options

Model ReleasesDGX agent

Artificial intelligence silicon and software giant Nvidia Corp. today announced two new services: a highly customizable Nemotron model and an agentic AI model router named NeMo Switchyard. As enterpri

Stealing Reasoning Traces from Proprietary LLM APIs

SafetyDGX agent

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

The small open weight models are scarier in AI development

Model ReleasesDGX agent

Imagine if your everyday laptop could run an AI model smart enough to take care of 90% of your work—totally private, lightning fast, and completely free of monthly fees. That is the exact tipping poin

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning

Local AiDGX agent

arXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Exi

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

Model ReleasesDGX agent

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

10 Aug 2026

Blast Radius

AgentsDGX agent

arXiv:2608.07440v1 Announce Type: new Abstract: Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

Model ReleasesDGX agent

arXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys

Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks

ResearchDGX agent

arXiv:2608.07283v1 Announce Type: new Abstract: Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cr

LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents

SafetyDGX agent

arXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial t

NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

Model ReleasesDGX agent

arXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should

Social World Models

Model ReleasesDGX agent

arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited informat

This Meta campaign is a case study in strategic reframing, with 4 major examples: 1) REFRAMES THE AI RACE FROM “who builds it the best” TO “…

SafetyDGX agent

This Meta campaign is a case study in strategic reframing, with 4 major examples: 1) REFRAMES THE AI RACE FROM “who builds it the best” TO “who distributes it to the most people”: Instead of fighting

Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceX

HardwareDGX agent

The TileRT InferenceX article (Aug 10 2026) examines whether the TileRT software stack on NVIDIA GPUs can compete with dedicated inference systems such as Cerebras, Groq LPUs and SambaNova for ultra‑h

9 Aug 2026

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so th…

Model ReleasesDGX agent

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so they hide all that stuff. They should instead explain choices

GitHub Models is now retired

Model ReleasesDGX agent

GitHub Models is now retired I missed this news until today, when the GitHub Actions run for my simonw/research repository failed with this error message: GitHub Models is temporarily unavailable as p

8 Aug 2026

Building a zero-dependency C inference engine for BitNet (1.58-bit) - lessons from hitting 36 tok/s on a Xeon CPU

Local AiDGX agent

Over the past few months I have been building a CPU-first inference engine from scratch in pure C99 (no Python, no CUDA, no BLAS, just GCC and make). The focus has been running 1.58-bit ternary models

Neat example here of the agents communicating purely through file names, including adding base64-encoded attachments and using 'zz' prefixes…

ToolsDGX agent

Neat example here of the agents communicating purely through file names, including adding base64-encoded attachments and using 'zz' prefixes to ensure their new message sorts to the bottom of the list

7 Aug 2026

Every CIO should watch this. Persistent systems trying to break through and solve problems at all costs are more creative than you think. La…

IndustryDGX agent

Every CIO should watch this. Persistent systems trying to break through and solve problems at all costs are more creative than you think. Labs and enterprises will be focused more on network effects (

In this episode of @wandb's Gradient Dissent , @l2k and Fireworks CEO @lqiao discuss why she believes the industry is at a turning point... …

Model ReleasesDGX agent

In this episode of @wandb's Gradient Dissent , @l2k and Fireworks CEO @lqiao discuss why she believes the industry is at a turning point... ... one that calls for more open intelligence, not less. The

Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra)

Model ReleasesDGX agent

Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5, where I had Claude Fable 5 build a full working game

ProDVI: Programmatic Dynamics Priors for Value Network Initialization

ResearchDGX agent

arXiv:2608.06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch,

Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?

Model ReleasesDGX agent

Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the communit

We didn’t give it access to the internet, but it found a way. We forgot to give it the Google Sheets permission, but it found a way. It acci…

IndustryDGX agent

We didn’t give it access to the internet, but it found a way. We forgot to give it the Google Sheets permission, but it found a way. It accidentally had edit access to the artifactory so we removed ed

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

Model ReleasesDGX agent

arXiv:2608.06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet m

What’s behind the Google AI shake-up

IndustryDGX agent

Some of the biggest names on Google's AI team got new jobs this week. In some cases, including for legendary Googler Jeff Dean, those jobs are no longer at Google. Given that Google's models seem to b

6 Aug 2026

Anthropic will design its own hardware to power Claude

Model ReleasesDGX agent

Anthropic is hiring a custom silicon team to design proprietary chips that will power its Claude models, while still planning a multi‑chip strategy that mixes internally designed hardware with compone

Can Post-Training Transform LLMs into Causal Reasoners?

Model ReleasesDGX agent

arXiv:2602.06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise in

Docs: US data labeling companies, like Surge AI and Mercor, that sell training datasets to US AI labs and the government are also selling them to Chinese labs (Anna Tong/Forbes)

HardwareDGX agent

Anna Tong / Forbes: Docs: US data labeling companies, like Surge AI and Mercor, that sell training datasets to US AI labs and the government are also selling them to Chinese labs — The same Silicon Va

General Availability of Pinecone Nexus Proves Knowledge Drives Real Outcomes for Agentic AI

Model ReleasesDGX agent

Pinecone announced the general availability of Pinecone Nexus, a knowledge engine that converts an enterprise’s proprietary data into governed, agent‑ready knowledge delivered through a single query c

I ported vLLM's serving stack to C++20: 66 MiB binary, no Python at inference, output checked token-for-token against vLLM

Model ReleasesDGX agent

I'm the author, so discount the enthusiasm accordingly. This is an unaffiliated community port, not endorsed by the vLLM project, which it uses to verify its correctness. What started it: I love vLLM,

the smartest young people in sf were working on agi/alignment 5-10y ago. they are working on bci today. naomi is a total star and is one of …

SafetyDGX agent

the smartest young people in sf were working on agi/alignment 5-10y ago. they are working on bci today. naomi is a total star and is one of an explosion of young talent into the bci space recently. i’

This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that w…

ApplicationsDGX agent

This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that was 'merely' good at hacking under human instructions. Initia

5 Aug 2026

ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference

Model ReleasesDGX agent

arXiv:2608.02947v1 Announce Type: cross Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelen

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

ResearchDGX agent

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

Model ReleasesDGX agent

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory

Model ReleasesDGX agent

A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B tot

is there a polite synonym for “circle jerk”?

SafetyDGX agent

The post contains two distinct snippets. First, user @GaryMarcus asks whether there is a more polite way to refer to “circle jerk.” Second, it shares a (likely satirical) claim that Microsoft’s AI rev

Meta releases Muse Code in beta, a terminal coding agent powered by Muse Spark 1.2, a coding-focused model priced at 1.25/1M input and 4.25/1M output tokens (Jonathan Vanian/CNBC)

AgentsDGX agent

Jonathan Vanian / CNBC: Meta releases Muse Code in beta, a terminal coding agent powered by Muse Spark 1.2, a coding-focused model priced at 1.25/1M input and 4.25/1M output tokens — Meta is rolling o

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

Model ReleasesDGX agent

TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe <N> | -ncmoe <N> Keep the routed MoE ex

None of this was us. Day 1 (and the first 48 hours) belonged to the open-source community: Generate with it @ComfyUI — native support + offi…

Local AiDGX agent

None of this was us. Day 1 (and the first 48 hours) belonged to the open-source community: Generate with it @ComfyUI — native support + official quantized builds, Day 0 Diffusers — the reference Pytho

Qwen Developers' responses from their recent Twitter/X AMA

Model ReleasesDGX agent

Questions & Responses(in BOLD) below. Favorite question(s) moved to end of the thread with combined responses(removed duplicates). Be optimistic folks. I'm sure we're getting other models too apart fr

SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite Intelligence

Local AiDGX agent

arXiv:2608.03728v1 Announce Type: new Abstract: Onboard satellite intelligence requires a task layer that translates mission intent into local tool calls, exposes execution state, and returns machine-

so do i have this right? the market is up on optimism about Microsoft’s revenue but MSFT’S biggest customer by far is burning billions a mon…

SafetyDGX agent

so do i have this right? the market is up on optimism about Microsoft’s revenue but MSFT’S biggest customer by far is burning billions a month with no obvious way to meet all of its future obligations

The bottleneck for AI progress was never compute, it was always the verifier. Recursive self-improvement is limited by verification, not com…

ResearchDGX agent

The bottleneck for AI progress was never compute, it was always the verifier. Recursive self-improvement is limited by verification, not computation. Compute buys proposals - verifiers buy knowledge.

Watch a local Ollama's qwen3:8b turn one English question into a 9-node investigation graph - planned, admitted by a deterministic gate, and run live in the browser (open source, MIT)

Model ReleasesDGX agent

The video is one real run, not a mock-up: grapharc go 'why did checkout latency spike at 09:14 UTC?' --model ollama/qwen3:8b A local 8B model proposes the graph → triage fanning out into four parallel

Welcome to the thunderdome of commoditization

SafetyDGX agent

Welcome to the thunderdome of commoditization 🚨The endgame of commoditization and falling prices and lack of a moat that I warned was inevitable in August 2023 took three years. But that moment has co

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Model ReleasesDGX agent

arXiv:2606.15673v2 Announce Type: replace Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and of

4 Aug 2026

GPT-OSS has turned one year old today!

Model ReleasesDGX agent

It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that

Introducing Web Search on Amazon Bedrock for foundation model grounding

TutorialsDGX agent

Today, we are introducing the general availability of Web Search on Amazon Bedrock. It is a server-side built-in tool that grounds model responses in current web knowledge. With Web Search, grounding

The Download: US robot restrictions, and ICE’s DNA grab

ResearchDGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Trump’s AI protectionism has come for robotics —James O’Donnel

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs

Model ReleasesDGX agent

arXiv:2603.28568v2 Announce Type: replace Abstract: Vision-language models (VLMs) share visual-textual representations across zero-shot classification, image captioning, and visual question answering

3 Aug 2026

LWiAI Podcast #253 - Opus 5, Gemini 3.6, Kimi K3, Hugging Face Hack

Model ReleasesDGX agent

LWiAI Podcast #253 (July 29 2026) covers a roundup of recent AI developments: Anthropic introduced Claude Opus 5 with Fable‑like capabilities; Google released Gemini 3.6/3.5 “Flash” variants and a cyb

Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

Local AiDGX agent

arXiv:2607.28696v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual

The Download: reward hacking explained, and suspected Iranian cyberattacks

ResearchDGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Here’s why AI agents lie and cheat to reach their goals When t

White House invites AI companies to review its new AI safety framework

Model ReleasesDGX agent

Cybersecurity chiefs at the White House have reportedly finalized the outline of a forthcoming framework that will enable artificial intelligence companies to voluntarily submit their latest frontier

2 Aug 2026

DeepSeek-V4-Flash 284B on 5.3GB of memory

Model ReleasesDGX agent

Following up on my Qwen 3.6 port, I wanted to keep adding models and ended up fixing a bunch of things along the way, so it's its own engine now: Mference. Same core idea from TurboFieldfare, MoE mode

DSpark Benchmark Result on Deepseek v4 Flash 0731

Model ReleasesDGX agent

TensorSharp supports DSpark on Deepseek v4 Flash 0731 now. Here is the benchmark result on 4x Nvidia A40 GPUs, cuda 12.8 with/without DSpark: Model: DeepSeek-V4-Flash-0731-UD-Q8_K_XL from https://hugg

Emad Mostaque @EMostaque came on PostAGI and said AI had already found 121 years of missing algebra in Einstein's equations. A billion param…

Model ReleasesDGX agent

Emad Mostaque @EMostaque came on PostAGI and said AI had already found 121 years of missing algebra in Einstein's equations. A billion parameter model trained on nothing past 1911 got to general relat

Encrypted Clouds?

Model ReleasesDGX agent

I love the progress happening on open models but I feel like it is kind of getting clear that hardware to run good sized models is completely unaffordable for me right now. I know that you all love Qw

yep they are indeed trying to gaslight me, — in exactly the way you anticipated. both predictable and intellectually dishonest.

SafetyDGX agent

yep they are indeed trying to gaslight me, — in exactly the way you anticipated. both predictable and intellectually dishonest. To be clear: when I say “LLM,” I mean the base model, not those with pat

← Previous
1…3233343536…44
Next →