AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “concepts”

GridTimelineEvolution
2,525 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
19 Jul 2026Wiki Lint Report — 2026-07-19

Automated lint: 20 errors, 8743 warnings, 3 info

→28 Jul 2026Do LLMs Know Their Vulnerable Scenarios?

arXiv:2607.23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their sa

→4 Aug 2026New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider t

→5 Aug 2026One-shotting a Raccoon Heist game using Claude Fable 5

Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept 'art' created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5

→5 Aug 2026... four years later, I got the new Claude Fable 5 to actually build the game https://x.com/simonw/status/2085089518223602058

... four years later, I got the new Claude Fable 5 to actually build the game https://x.com/simonw/status/2085089518223602058 Four years ago today I tweeted about having GPT-3 and DALL-E come up with

→6 Aug 2026One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIP

arXiv:2505.19840v3 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) have achieved widespread success yet remain prone to adversarial attacks. Typically, such attacks either involve f

→7 Aug 2026You can play it here: https://simonw.github.io/raccoon-heist-codex/ For comparison, here's Fable 5 + Claude Code's game, built from the exac…

You can play it here: https://simonw.github.io/raccoon-heist-codex/ For comparison, here's Fable 5 + Claude Code's game, built from the exact same prompt https://x.com/simonw/status/208508951822360205

→7 Aug 2026M^3R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

arXiv:2608.05817v1 Announce Type: new Abstract: Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visu

CompanyOpenAI8 recent entries
29 Jul 2026From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new

→30 Jul 2026Under the Hood: Serving Kimi K3

DigitalOcean launched Kimi K3 on day 0. It’s already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a

→2 Aug 2026Top eight misconceptions about OpenAI’s amazing new Astra math results. 1. Expertise in one domain does not at all guarantee expertise in al…

Top eight misconceptions about OpenAI’s amazing new Astra math results. 1. Expertise in one domain does not at all guarantee expertise in all or even most domains. There is an important, principled re

→4 Aug 2026New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider t

→5 Aug 2026One-shotting a Raccoon Heist game using Claude Fable 5

Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept 'art' created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5

→7 Aug 2026Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

arXiv:2608.05411v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learni

→9 Aug 2026GitHub Models is now retired

GitHub Models is now retired I missed this news until today, when the GitHub Actions run for my simonw/research repository failed with this error message: GitHub Models is temporarily unavailable as p

→10 Aug 2026Social World Models

arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited informat

CompanyGoogle8 recent entries
31 Jul 2026TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning

arXiv:2607.26307v1 Announce Type: new Abstract: Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benc

→4 Aug 2026New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider t

→5 Aug 2026Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

arXiv:2608.03569v1 Announce Type: new Abstract: Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in eval

→6 Aug 2026NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts

arXiv:2608.04030v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domains remains

→6 Aug 2026Agentic Future Ready With BigQuery: Continually Improving Price-Performance, Zero Effort

In the modern data landscape, query performance tuning and managing system price-performance is challenging, especially as the number of agentic workloads increase. Even for experienced developers and

→7 Aug 2026Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index

arXiv:2608.05411v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learni

→10 Aug 2026Need real world ML problems to evaluate my educational ML tools

I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept

→11 Aug 2026LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that

CompanyMeta8 recent entries
4 Aug 2026DS@GT ARC at MEDIQA-CORE-Task-1 2026: Trimodal Model Fusion with Task-Specific Gates for Brain Tumor Subtype Classification

arXiv:2608.00086v1 Announce Type: new Abstract: Brain tumor diagnosis is a time-sensitive process in which patients may wait weeks for a finalized pathology report. This problem motivates automated sy

→5 Aug 2026Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICS

arXiv:2608.03524v1 Announce Type: new Abstract: AGENTONOMICS is a framework that treats AI agents as economic entities that can be designed, managed, and governed through an integrated management arch

→5 Aug 2026Disentangling MLP Neuron Weights in Vocabulary Space

arXiv:2604.06005v2 Announce Type: replace Abstract: Interpreting the information encoded in language model weights remains a fundamental challenge in mechanistic interpretability. In this work, we int

→7 Aug 2026PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden

→7 Aug 2026Got job as Director of AI and Systems development self-taught

Hey everyone, I just wanted to share my journey here for some motivation. Three years ago, I saw the sudden spike in AI and realized it was the future of tech. My goal at the time was to be an indie g

→8 Aug 2026Showoff Saturday: Local 4x 6000 Pro (multi-year progression)

Not the biggest or shiniest, but it's mine From gaming machine inference on the original llama models, to a 4x RTX 6000 Pro Max Q + 4x 3090s local AI cluster. Pictures are in reverse chronological ord

→10 Aug 2026MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering

arXiv:2608.06638v1 Announce Type: cross Abstract: Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public t

→10 Aug 2026Latent Fact-Checking: Detecting Misinformation through Activation Engineering

arXiv:2608.06417v1 Announce Type: cross Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level ling

CompanyMistral8 recent entries
25 Jun 2026Steering Vision-Language Models with Joint Sparse Autoencoders

arXiv:2606.25657v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representat

→15 Jul 2026RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories

arXiv:2512.04144v3 Announce Type: replace Abstract: Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often propagat

→15 Jul 2026Wiki Lint Report — 2026-07-15

Automated lint: 26 errors, 6728 warnings, 3 info

→19 Jul 2026Wiki Lint Report — 2026-07-19

Automated lint: 20 errors, 8743 warnings, 3 info

→28 Jul 2026Might need math+code benchmark for frontier model(LLMs Silently Replace Math)[D]

Hello guys. I found some problems in current frontier models. And want to share. # math_code_hallucination > Record of a failure caused by combining mathematics and code in a single prompt. --- ## Cas

→31 Jul 2026STEREODISCO: Discovering Stereotypicality in LLMs

arXiv:2607.27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psycholog

→4 Aug 2026New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider t

→7 Aug 2026PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden

CompanyxAI8 recent entries
19 Jul 2026Wiki Lint Report — 2026-07-19

Automated lint: 20 errors, 8743 warnings, 3 info

→24 Jul 2026Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit

arXiv:2512.10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases

→25 Jul 2026Old Coder Needs help with New AI Development and wants to get up to speed to understand it all.

Hi Guys, I'm an old coder and DBA that has been in the field for almost 40 years. More and more the jobs I was doing for work are being taken over by AI and the need for my type of work is diminishing

→28 Jul 2026Context-Aware Concept Distillation for Trustworthy Flood Prediction

arXiv:2607.23237v1 Announce Type: cross Abstract: Effective flood risk management relies on accurate forecasting, yet the 'black box' nature of stateof-the-art Deep Learning models creates a barrier t

→7 Aug 2026Challenges in Evaluating Explanation Methods for Static and Evolving Data

arXiv:2608.06351v1 Announce Type: new Abstract: This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through

→10 Aug 2026Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided Design

arXiv:2608.07091v1 Announce Type: cross Abstract: Edge Artificial Intelligence (Edge AI) enables the deployment of AI models directly on local edge devices, while such deployments are subject to stric

→12 Aug 2026Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information

arXiv:2608.10766v1 Announce Type: new Abstract: Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rul

→12 Aug 2026Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding

arXiv:2603.25251v2 Announce Type: replace-cross Abstract: Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which est

CompanyDeepSeek8 recent entries
15 Jul 2026Wiki Lint Report — 2026-07-15

Automated lint: 26 errors, 6728 warnings, 3 info

→19 Jul 2026Wiki Lint Report — 2026-07-19

Automated lint: 20 errors, 8743 warnings, 3 info

→24 Jul 2026What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

arXiv:2607.21491v1 Announce Type: new Abstract: Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-c

→31 Jul 2026Some deepseek-v4-flash 20260731 opinion review

First of all, I want to apologize if it's off-topic or in the wrong format. Having tried Deepseek Flash with reasoning high on a conceptually difficult task, involving Machine Learning classifiers and

→3 Aug 2026Döner Bench DeepSeek-V4-Flash IQ2_XS running on a single RTX 3090

https://preview.redd.it/3zcvpbds14hh1.png?width=1911&format=png&auto=webp&s=a79aafb71eeca97638da93d2591902631e897fd5 I tried a test similar to the recent model-quant comparisons, but this time I focus

→8 Aug 2026Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks?

(I am not a native speaker, written by myself, so please bear with me) I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence bench

→10 Aug 2026Need real world ML problems to evaluate my educational ML tools

I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept

→11 Aug 2026I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples

I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon

CompanyNVIDIA8 recent entries
19 Jul 2026Wiki Lint Report — 2026-07-19

Automated lint: 20 errors, 8743 warnings, 3 info

→29 Jul 2026From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new

→30 Jul 2026Under the Hood: Serving Kimi K3

DigitalOcean launched Kimi K3 on day 0. It’s already one of the most popular models on the platform and across the market: second most likes on Hugging Face, sixth most traffic on OpenCode. Getting a

→1 Aug 2026Github repo to learn the OPD/OPSD and how they perform compared to GRPO, on a consumer grade GPU [P]

I am trying to learn concepts like On Policy Distillation (OPD), On Policy Self Distillation (OPSD) and how do they compare to RL algorithms like GRPO. There are a lot of papers on this, but because o

→4 Aug 2026Llama.cpp PR 8% speed boost

Llama.cpp currently uses cpu based sampling for user with mtp enabled. The PR moves sampling to the gpu, which on a 5090 boasts an 8% increase in tok/s for qwen3.6:35b. I tested it on my P40 and obser

→10 Aug 2026Need real world ML problems to evaluate my educational ML tools

I'm a retired platform engineer, coding mainly in Rust, and involved with a ML study group. I developed a ML programming language (alternative to Python, Colab) to help me learn (and teach) ML concept

→10 Aug 2026Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-Tuning

arXiv:2608.07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class

→11 Aug 2026LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs

arXiv:2608.07733v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems