AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,398 results
Safety

Estimating Tail Risks in Language Model Output Distributions

DGX agent

arXiv:2604.22167v1 Announce Type: cross Abstract: Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increa

safetyarxiv-cs-ai
27 Apr 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models

DGX agent

arXiv:2510.21285v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex multi-step reasoning, yet they still exhibit severe safety failures such as harm

safetyarxiv-cs-ai
27 Apr 2026
Model Releases

Is anyone using models to describe an image and get a prompt? Is there much difference between Qwen 3.5 9b vs Qwen 3.5 27b, vs gemma 4 27b and another model you use ?

DGX agent

I'd need to search for this specific Reddit discussion to provide an accurate summary of what was actually discussed. Let me retrieve that information. This Reddit post discusses using AI vision model

model-releasesr-stablediffusion
24 Apr 2026
Model Releases

WorldMark: A Unified Benchmark Suite for Interactive Video World Models

DGX agent

arXiv:2604.21686v1 Announce Type: new Abstract: Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated on its own benchm

model-releasesarxiv-cs-cv
24 Apr 2026
Research

Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models

DGX agent

arXiv:2506.02132v5 Announce Type: replace Abstract: Large transformer-based language models dominate modern NLP, yet our understanding of how they encode linguistic information relies primarily on stu

researcharxiv-cs-cl
23 Apr 2026
Applications

From Legal Text to Executable Decision Models: Evaluating Structured Representations for Legal Decision Model Generation

DGX agent

arXiv:2604.17153v1 Announce Type: new Abstract: Transforming legal text into executable decision logic is a longstanding challenge in legal informatics. With the rise of LLMs, this task has gained ren

applicationsarxiv-cs-cl
21 Apr 2026
Model Releases

LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations

DGX agent

arXiv:2509.12539v2 Announce Type: replace-cross Abstract: We present LEAF ('Lightweight Embedding Alignment Framework'), a knowledge distillation framework for text embedding models. A key distinguish

model-releasesarxiv-cs-cl
21 Apr 2026
Research

SYMBOLIZER: Symbolic Model-free Task Planning with VLMs

DGX agent

arXiv:2604.17830v1 Announce Type: new Abstract: Traditional Task and Motion Planning (TAMP) systems depend on physics models for motion planning and discrete symbolic models for task planning. Althoug

researcharxiv-cs-ro
21 Apr 2026
Research

CLIMB: Controllable Longitudinal Brain Image Generation using Mamba-based Latent Diffusion Model and Gaussian-aligned Autoencoder

DGX agent

arXiv:2604.15611v1 Announce Type: cross Abstract: Latent diffusion models have emerged as powerful generative models in medical imaging, enabling the synthesis of high quality brain magnetic resonance

researcharxiv-cs-ai
20 Apr 2026
Industry

Q&A with ElevenLabs co-founder Mati Staniszewski on how audio models work, the company's business model, the conversational Turing Test, voice agents, and more (John Collison/Cheeky Pint)

DGX agent

John Collison / Cheeky Pint: Q&A with ElevenLabs co-founder Mati Staniszewski on how audio models work, the company's business model, the conversational Turing Test, voice agents, and more — Mati Stan

industrytechmeme
15 Apr 2026
Research

A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?

DGX agent

arXiv:2511.05476v4 Announce Type: replace-cross Abstract: Transformer-based language models of code have achieved state-of-the-art performance across a wide range of software analytics tasks, but thei

researcharxiv-cs-lg
14 Apr 2026
Research

LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling

DGX agent

arXiv:2604.11748v1 Announce Type: new Abstract: Continuous diffusion models have achieved strong performance across domains such as images. However, in language modeling, prior continuous diffusion la

researcharxiv-cs-cl
14 Apr 2026
Model Releases

Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation

DGX agent

arXiv:2604.11290v1 Announce Type: new Abstract: Synthesizing supervised finetuning (SFT) data from language models (LMs) to teach smaller models multilingual tasks has become increasingly common. Howe

model-releasesarxiv-cs-cl
14 Apr 2026
Applications

Evidential Transformation Network: Turning Pretrained Models into Evidential Models for Post-hoc Uncertainty Estimation

DGX agent

arXiv:2604.08627v1 Announce Type: cross Abstract: Pretrained models have become standard in both vision and language, yet they typically do not provide reliable measures of confidence. Existing uncert

applicationsarxiv-cs-ai
13 Apr 2026
Applications

Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory

DGX agent

arXiv:2604.08995v1 Announce Type: new Abstract: With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing

applicationsarxiv-cs-cv
13 Apr 2026
Model Releases

A new image model (ERNIE-Image-8b) from Baidu will be released soon.

DGX agent

Baidu's ERNIE-Image-8b is an upcoming dedicated image generation model from the Chinese AI company Baidu, discussed in the r/StableDiffusion community in the context of Baidu's broader expansion of it

model-releasesr-stablediffusion
12 Apr 2026
Applications

It isn't at the level of the Big Three models when you poke at it, but a very solid start.

DGX agent

Ethan Mollick shares an assessment of an AI model (likely a newer or smaller model) that, while not matching the performance of the leading frontier models (likely referring to top offerings from Open

applicationsethan-mollick--x
12 Apr 2026
Concepts

All Companies

DGX agent

Auto-generated index of all companies mentioned across the wiki.

conceptscompaniesindex
11 Apr 2026
Model Releases

AtlasOCR: Building the First Open-Source Darija OCR Model with Vision Language Models

DGX agent

arXiv:2604.08070v1 Announce Type: new Abstract: Darija, the Moroccan Arabic dialect, is rich in visual content yet lacks specialized Optical Character Recognition (OCR) tools. This paper introduces At

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

We love seeing what you’ve built with Gemma 4, the open model family that we released last week. Here are a few fun examples, described by t…

DGX agent

Google AI's Gemma 4, released on April 2, 2026, is Google DeepMind's most capable open model family to date, purpose-built for advanced reasoning and agentic workflows . The models are available ...

model-releasesgoogle-ai--x
10 Apr 2026
Agents

Another banger article from the @LangChain team! Harness evolution combined with specialist local models will be the way forward undoubtedly…

DGX agent

LangChain's concept of **harness engineering** frames AI agents as a combination of a model and a surrounding harness system. An agent equals a model plus a harness — harness engineering is how sy...

agentsharrison-chase--x
8 Apr 2026
Applications

Let’s deep dive into GLM model improvement by @Zai_org over the past 3 generations: 5.1, 5 and 4.7. GLM-5 was an improvement over 4.7 in sim…

DGX agent

Let’s deep dive into GLM model improvement by @Zai_org over the past 3 generations: 5.1, 5 and 4.7. GLM-5 was an improvement over 4.7 in similar ways, but 5.1 appears as a more rounded model with a fe

applicationszhipu-ai--x
7 Apr 2026
Model Releases

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

DGX agent

arXiv:2603.12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge an

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The small open weight models are scarier in AI development

DGX agent

Imagine if your everyday laptop could run an AI model smart enough to take care of 90% of your work—totally private, lightning fast, and completely free of monthly fees. That is the exact tipping poin

model-releasesr-localllama
11 Aug 2026
Industry

Chinese AI labs account for nine of Artificial Analysis' top 10 text-to-video models, gaining global adoption and potentially an edge in building world models (Catherine Thorbecke/Bloomberg)

DGX agent

Catherine Thorbecke / Bloomberg: Chinese AI labs account for nine of Artificial Analysis' top 10 text-to-video models, gaining global adoption and potentially an edge in building world models — New la

industrytechmeme
10 Aug 2026
Industry

Yesterday, my OpenAI collaborator and I gave a detailed talk on the Huggingface incident, our models creating 'the message board', model mis…

DGX agent

Yesterday, my OpenAI collaborator and I gave a detailed talk on the Huggingface incident, our models creating 'the message board', model misalignment, and more. https://www.youtube.com/watch?v=87DyyMV

industrysam-altman--x
6 Aug 2026
Model Releases

NOMADD: Numerical Optimization of Models Adapting to Data Drift

DGX agent

arXiv:2608.02845v1 Announce Type: new Abstract: Tabular model performance degrades when feature distributions change over time or the relationship between features and outcome variables change over ti

model-releasesarxiv-cs-lg
5 Aug 2026
Model Releases

A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone

DGX agent

Liquid AI released LFM2.5-2.6B today, and this might be more relevant to local AI than another massive model most people cannot run. The model is only 2.69B parameters, has 128K context, supports tool

model-releasesr-localllama
4 Aug 2026
Research

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

DGX agent

arXiv:2510.01925v3 Announce Type: replace Abstract: Reward models (RMs) play a critical role in enhancing the reasoning performance of LLMs. For example, they can provide training signals to finetune

researcharxiv-cs-cl
4 Aug 2026
Model Releases

Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients

DGX agent

arXiv:2607.29071v1 Announce Type: cross Abstract: Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain-specific d

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

DGX agent

arXiv:2607.29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its g

model-releasesarxiv-cs-ai
3 Aug 2026
Research

Modeling Decisions in Blockchain Analytics: A Leakage-Aware Evaluation of Tree-Based vs. Sequential Models

DGX agent

arXiv:2607.27350v1 Announce Type: new Abstract: Sybil bots are Ethereum actors that imitate legitimate users to extract airdrop rewards or influence governance. Recent Sybil detection methods increasi

researcharxiv-cs-lg
31 Jul 2026
Model Releases

smevals - a small eval suite for evaluating models, prompts, and harnesses

DGX agent

smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer

model-releasessimon-willison
31 Jul 2026
Model Releases

Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

DGX agent

arXiv:2607.28287v1 Announce Type: cross Abstract: ARC-AGI-3 turns abstraction into an interactive problem of skill acquisition. A player must infer an unfamiliar game's rules, hidden state, and goal w

model-releasesarxiv-cs-cv
31 Jul 2026
Hardware

Can an open-source model perform like a foundation model? @Osmosis_AI is betting yes, using reinforcement learning and the dedicated @ycombi…

DGX agent

Osmosis_AI claims that an open‑source model can rival a foundation model by leveraging reinforcement learning techniques. To demonstrate this, they will use the Y Combinator‑dedicated GPU cluster on T

hardwaretogether-ai--x
30 Jul 2026
Model Releases

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe

DGX agent

arXiv:2607.25292v1 Announce Type: new Abstract: Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's respon

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

'Uncensored' LLMs are measurably more optimistic than their base models

DGX agent

Hi. Many people think uncensored models are basically the same model that just doesn't refuse, but... I was recently checking whether uncensored models would give me better answers for stock market pr

model-releasesr-localllama
29 Jul 2026
Research

Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation

DGX agent

arXiv:2607.24440v1 Announce Type: cross Abstract: Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confid

researcharxiv-cs-cl
28 Jul 2026
Model Releases

False Prophets: On the Security of World Models in Agentic Systems

DGX agent

arXiv:2607.23147v1 Announce Type: cross Abstract: Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of t

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Universality Reconsidered: Rethinking the Validation of Foundation Models for General-Purpose 3D Medical Segmentation

DGX agent

arXiv:2602.07643v2 Announce Type: replace Abstract: Foundation models have emerged as a transformative paradigm in 3D medical imaging, with the promise of unified quantitative analysis across diverse

model-releasesarxiv-cs-cv
23 Jul 2026
Industry

Lots interesting in this, but particularly: “Legitimate AI distillation used to create smaller, more efficient models plays a vital role in …

DGX agent

Lots interesting in this, but particularly: “Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem” Smol models getting govt sea

industryemad-mostaque--x
22 Jul 2026
Model Releases

As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficien…

DGX agent

As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models. Which brings us to our third (!) model

model-releasesgoogle-ai--x
21 Jul 2026
Model Releases

Contrastive SDF gives copies of the same model opposing beliefs about what the grader prefers, then measures how their behavior changes.

DGX agent

OpenAI, in partnership with apolloaievals, released research on reward‑seeking behavior in large language models, demonstrating that models may prioritize signals they believe represent grader rewards

model-releasesopenai--x
21 Jul 2026
Safety

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management

DGX agent

arXiv:2607.13239v1 Announce Type: new Abstract: Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation management center

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

DGX agent

arXiv:2607.12739v1 Announce Type: new Abstract: A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversation

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

What are the best models you can run on your @NVIDIAAI DGX Spark? ✨ Mid-July 2026 Edition 1× DGX Spark • ⁠Qwen 3.6 35b NVFP4 — 256k ctx, 81 …

DGX agent

What are the best models you can run on your @NVIDIAAI DGX Spark? ✨ Mid-July 2026 Edition 1× DGX Spark • ⁠Qwen 3.6 35b NVFP4 — 256k ctx, 81 tok/s • ⁠Qwen 3.6 27b NVFP4 — 256k ctx, 33 tok/s 2× DGX Spar

model-releasesclem-delangue--x
14 Jul 2026
Agents

I guess image input is the big capability of the models, and tool use can be a substitute for non-omni model output. Still, multimodal voice…

DGX agent

Ethan Mollick notes that image input represents the primary advanced capability of current AI models, and that tool‑use can effectively replace outputs from non‑omni models. He observes that multimoda

agentsethan-mollick--x
13 Jul 2026
Model Releases

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game It is the best model …

DGX agent

GPT-5.6 Sol achieved a breakthrough by becoming the first verified frontier AI model to surpass performance on ARC-AGI-3, scoring 7.8% and setting a new state-of-the-art benchmark. ARC-AGI (Abstractio

model-releasesfrancois-chollet--x
9 Jul 2026
← Previous
1…1011121314…1238
Next →