AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,428 results
12 Apr 2026

MiniMax M2.7 is available on Ollama's cloud, and is licensed for commercial usage. Use it with OpenClaw: ollama launch openclaw --model mini…

Model ReleasesDGX agent

MiniMax M2.7 is available on Ollama's cloud, and is licensed for commercial usage. Use it with OpenClaw: ollama launch openclaw --model minimax-m2.7:cloud Coding agents, such as Claude: ollama launch

OstrisAI-Toolkit Lora --> Anima model.

Local AiDGX agent

This Reddit post discusses using the Ostris AI-Toolkit to train a LoRA (Low-Rank Adaptation) model targeting the 'Anima' Stable Diffusion model. The Ostris AI Toolkit is a powerful framework that allo

11 Apr 2026

Hermes Agent is good. What makes it impressive is how performant it is with even small or locally hosted models. This used to be stuff that …

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
AgentsDGX agent

Nous Research highlights Hermes Agent as a notably impressive AI agent framework, distinguished by its strong performance even when running small or locally hosted models. This capability is significa

10 Apr 2026

Best part? The spontaneous booth conversations with builders focused on turning AI ambition into production systems that scale. Open models …

ApplicationsDGX agent

Best part? The spontaneous booth conversations with builders focused on turning AI ambition into production systems that scale. Open models are moving fast. The infrastructure around them is too. Grea

Beyond Mamba: Enhancing State-space Models with Deformable Dilated Convolutions for Multi-scale Traffic Object Detection

Model ReleasesDGX agent

arXiv:2604.08038v1 Announce Type: new Abstract: In a real-world traffic scenario, varying-scale objects are usually distributed in a cluttered background, which poses great challenges to accurate dete

Code Sharing In Prediction Model Research: A Scoping Review

ResearchDGX agent

arXiv:2604.06212v1 Announce Type: cross Abstract: Analytical code is essential for reproducing diagnostic and prognostic prediction model research, yet code availability in the published literature re

Completely agreed. Everyone in SF knows how good these models are at coding / work automation. When it comes to writing, there’s still a ins…

AgentsDGX agent

Completely agreed. Everyone in SF knows how good these models are at coding / work automation. When it comes to writing, there’s still a insane amount of model-generated AI slop because there’s no eas

Controller Design for Structured State-space Models via Contraction Theory

ResearchDGX agent

arXiv:2604.07069v1 Announce Type: cross Abstract: This paper presents an indirect data-driven output feedback controller synthesis for nonlinear systems, leveraging Structured State-space Models (SSMs

DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models

TutorialsDGX agent

arXiv:2306.14685v5 Announce Type: replace-cross Abstract: We demonstrate that pre-trained text-to-image diffusion models, despite being trained on raster images, possess a remarkable capacity to guide

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing

Model ReleasesDGX agent

arXiv:2509.01986v4 Announce Type: replace-cross Abstract: In recent years, integrating multimodal understanding and generation into a single unified model has emerged as a promising paradigm. While th

E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task

Model ReleasesDGX agent

arXiv:2510.14509v3 Announce Type: replace-cross Abstract: The rapid advancement in large language models (LLMs) has demonstrated significant potential in End-to-End Software Development (E2ESD). Howev

From Synthetic Data to Real Restorations: Diffusion Model for Patient-specific Dental Crown Completion

Local AiDGX agent

arXiv:2603.26588v2 Announce Type: replace-cross Abstract: We present ToothCraft, a diffusion-based model for the contextual generation of tooth crowns, trained on artificially created incomplete teeth

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs

SafetyDGX agent

arXiv:2604.07655v1 Announce Type: cross Abstract: Hard-gated safety checkers often over-refuse and misalign with a vendor's model spec; prevailing taxonomies also neglect robustness and honesty, yield

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

HardwareDGX agent

arXiv:2604.07812v1 Announce Type: new Abstract: In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making th

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles

SafetyDGX agent

arXiv:2604.07650v1 Announce Type: cross Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Shared pretra

HST-HGN: Heterogeneous Spatial-Temporal Hypergraph Networks with Bidirectional State Space Models for Global Fatigue Assessment

ResearchDGX agent

arXiv:2604.08435v1 Announce Type: new Abstract: It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of modeling long-ra

Interventional Time Series Priors for Causal Foundation Models

ResearchDGX agent

arXiv:2603.11090v2 Announce Type: replace Abstract: Prior-data fitted networks (PFNs) have emerged as powerful foundation models for tabular causal inference, yet their extension to time series remain

Models randomly becoming corrupted?

Local AiDGX agent

I was unable to retrieve the specific Reddit thread at the provided URL, and the search results did not return content from that exact post. The search results returned related but distinct issues ...

Negative Binomial Variational Autoencoders for Overdispersed Latent Modeling

Model ReleasesDGX agent

arXiv:2508.05423v2 Announce Type: replace Abstract: Although artificial neural networks are often described as brain-inspired, their representations typically rely on continuous activations, such as t

NEWS: The Tesla Model Y was the top-selling passenger vehicle in China in March 2026, with 39,827 retail registrations, beating all EVs and …

IndustryDGX agent

NEWS: The Tesla Model Y was the top-selling passenger vehicle in China in March 2026, with 39,827 retail registrations, beating all EVs and ICE models. It outpaced competitors across sedans, SUVs, and

Not All Tokens See Equally: Perception-Grounded Policy Optimization for Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.01840v2 Announce Type: replace Abstract: While Reinforcement Learning from Verifiable Rewards (RLVR) has advanced reasoning in Large Vision-Language Models (LVLMs), prevailing frameworks su

Phantasia: Context-Adaptive Backdoors in Vision Language Models

ResearchDGX agent

arXiv:2604.08395v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have greatly enhanced the integration of visual perception and linguistic reasoning, driving rapid prog

Recommended Model for a 4060ti 8gb and 16gb ram

Local AiDGX agent

For users running Ollama on an NVIDIA RTX 4060 Ti with 8GB VRAM and 16GB system RAM, the community consensus recommends 7B–8B parameter models (such as Llama 3.1 8B, Mistral 7B, or Qwen 8B) using Q...

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

SafetyDGX agent

arXiv:2604.06628v1 Announce Type: new Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit t

$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models

ResearchDGX agent

arXiv:2604.06260v1 Announce Type: cross Abstract: Test-time scaling investigates whether a fixed diffusion language model (DLM) can generate better outputs when given more inference compute, without a

The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models

ResearchDGX agent

arXiv:2604.06374v1 Announce Type: cross Abstract: Latent reasoning via continuous chain-of-thoughts (Latent CoT) has emerged as a promising alternative to discrete CoT reasoning. Operating in continuo

TREASURE: The Visa Payment Foundation Model for High-Volume Transaction Understanding

ApplicationsDGX agent

arXiv:2511.19693v3 Announce Type: replace-cross Abstract: Payment networks form the backbone of modern commerce, generating high volumes of transaction records from daily activities. Properly modeling

Variational Feature Compression for Model-Specific Representations

Model ReleasesDGX agent

arXiv:2604.06644v1 Announce Type: cross Abstract: As deep learning inference is increasingly deployed in shared and cloud-based settings, a growing concern is input repurposing, in which data submitte

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment

ApplicationsDGX agent

arXiv:2604.08212v1 Announce Type: new Abstract: General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring preci

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

Model ReleasesDGX agent

arXiv:2604.07705v1 Announce Type: new Abstract: Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonom

What if a model became the computer itself?

ResearchDGX agent

What if a model became the computer itself? NEW paper from Meta. (bookmark this one) What if the model wasn't just using the computer, but became the computer? New research from Meta AI and KAUST make

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

SafetyDGX agent

arXiv:2604.08546v1 Announce Type: new Abstract: Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a

9 Apr 2026

>8 out of 8 [cheap oss] models detected Mythos's flagship FreeBSD exploit Completely disingenuous They gave it just ~20 lines of code to rea…

ResearchDGX agent

>8 out of 8 [cheap oss] models detected Mythos's flagship FreeBSD exploit Completely disingenuous They gave it just ~20 lines of code to read. They baked in custom, relevant context pertinent to the e

expect everyone to be speaking about world models over the next 6-8 months.

IndustryDGX agent

expect everyone to be speaking about world models over the next 6-8 months. expect everyone to be speaking about video models over the next 6-8 months. this might be the most important time for image/

LaCy: What Small Language Models Can and Should Learn is Not Just a Question of Loss

Model ReleasesDGX agent

This paper was accepted at the Workshop on Memory for LLM-Based Agentic Systems at ICLR. Language models have consistently grown to compress more world knowledge into their parameters, but the knowled

Science needs a way to process models that are only 'mostly correct' in terms of their predictions, but are very compressive (high ratio bet…

ResearchDGX agent

Science needs a way to process models that are only 'mostly correct' in terms of their predictions, but are very compressive (high ratio between predictive power and model complexity). They are likely

8 Apr 2026

AI Gateway now supports team-wide Zero Data Retention (ZDR). Building safely with multiple AI models means wrestling with fragmented data po…

Model ReleasesDGX agent

AI Gateway now supports team-wide Zero Data Retention (ZDR). Building safely with multiple AI models means wrestling with fragmented data policies, per-provider negotiations, and the hope that develop

Experimenting with GPUs: GKE managed DRANET and Inference Gateway AI Deployment

Model ReleasesDGX agent

Building and serving models on infrastructure is a strong use case for businesses. In Google Cloud, you have the ability to design your AI infrastructure to suit your workloads. Recently, I experiment

ok i read the cyber part of the mythos model card. some thoughts. 250 'trials' across 50 crash categories but almost every full exploit is a…

ResearchDGX agent

ok i read the cyber part of the mythos model card. some thoughts. 250 'trials' across 50 crash categories but almost every full exploit is a permutation of the same 2 bugs, rediscovered from different

Seems like a good model from Meta that is still trailing the current series of releases. The most important thing to note is that it is not …

ApplicationsDGX agent

Seems like a good model from Meta that is still trailing the current series of releases. The most important thing to note is that it is not open weights. That was the main reason that Meta's models we

7 Apr 2026

Anthropic's Project Glasswing - restricting Claude Mythos to security researchers - sounds necessary to me

Model ReleasesDGX agent

Anthropic didn't release their latest model, Claude Mythos (system card PDF), today. They have instead made it available to a very restricted set of preview partners under their newly announced Projec

http://Z.ai releases GLM-5.1, a 754B-parameter model that it says outperforms GPT-5.4 and Claude Opus 4.6 on SWE-bench Pro, available under …

Model ReleasesDGX agent

http://Z.ai releases GLM-5.1, a 754B-parameter model that it says outperforms GPT-5.4 and Claude Opus 4.6 on SWE-bench Pro, available under an MIT license (@carlfranzen / VentureBeat) https://ventureb

13 Aug 2026

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Model ReleasesDGX agent

arXiv:2608.12307v1 Announce Type: cross Abstract: Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forc

12 Aug 2026

CohereLabs/North-Micro-Vision-Instruct · Hugging Face

Model ReleasesDGX agent

North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation fo

10 Aug 2026

Quantization Damage Is Multiplicative, Not Additive

Model ReleasesDGX agent

arXiv:2608.06564v1 Announce Type: cross Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt. What nobody can say is which of the model's

6 Aug 2026

Advancing brain tumor research with privacy-first AI

Model ReleasesDGX agent

The intersection of medicine and AI has led to remarkable innovations. However, developers now face the thorny challenge of building robust medical AI tools that have been tested and evaluated on dive

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

Model ReleasesDGX agent

arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Mo

4 Aug 2026

New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

Model ReleasesDGX agent

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider t

SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials

Model ReleasesDGX agent

arXiv:2509.21079v2 Announce Type: replace Abstract: Foundation models have shown remarkable capabilities in various domains, but their performance on complex, multimodal engineering problems remains l

2 Aug 2026

Open letters about AI development

Model ReleasesDGX agent

Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and Ame

31 Jul 2026

FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing

Model ReleasesDGX agent

arXiv:2508.02092v3 Announce Type: replace-cross Abstract: Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuable

29 Jul 2026

Automate your agent development lifecycle using any coding agent

Model ReleasesDGX agent

Welcome to our latest Gemini Enterprise Agent Platform deep dive, a practical walkthrough where we’ll teach you how to build real-world, production-ready agents starting from step 1. If you haven’t al

24 Jul 2026

swiss-ai/Apertus-v1.5 70B/8B

Model ReleasesDGX agent

https://huggingface.co/swiss-ai/Apertus-v1.5-70B https://huggingface.co/swiss-ai/Apertus-v1.5-8B Apertus 1.5 is a family of 8B and 70B parameter language models designed to advance the state of multil

23 Jul 2026

[Paper] SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Model ReleasesDGX agent

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlappe

15 Jul 2026

PRISM Edit: One Vector for All Temporal Answers

Model ReleasesDGX agent

arXiv:2607.11327v2 Announce Type: replace-cross Abstract: Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing locat

26 Jun 2026

Quoting OpenAI

Model ReleasesDGX agent

We're beginning a limited preview of the GPT‑5.6 series: Sol, our flagship model; Terra, a balanced model for everyday work; and Luna, a fast and affordable model. Terra has competitive performance to

10 Jun 2026

Everyone Operating At The Frontier Satya Nadella, Chairman & CEO, Microsoft, interviewed by @saranormous & @eladgil (No Priors) and @swyx (L…

Model ReleasesDGX agent

Everyone Operating At The Frontier Satya Nadella, Chairman & CEO, Microsoft, interviewed by @saranormous & @eladgil (No Priors) and @swyx (Latent Space) Crossover special at Microsoft Build 2026. Summ

2 Jun 2026

Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2606.01682v1 Announce Type: cross Abstract: Selecting the best response from multiple small-model samples using a stronger scorer is a simple inference-time strategy, but fails when the small mo

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation

Model ReleasesDGX agent

arXiv:2606.01481v1 Announce Type: new Abstract: With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos fro

12 May 2026

jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition

Model ReleasesDGX agent

arXiv:2605.08384v1 Announce Type: new Abstract: In this work, we introduce frozen-encoder model composition, a novel approach to multimodal embedding models. We build on the VLM-style architecture, in

← Previous
1…7576777879…1008
Next →