AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
85,202Total entries
1Added by human
85,201Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
61,046 results
Model Releases

Qwen3.8-2.4T-A95B is live on Together AI. The Qwen team’s new open-weight MoE packs 2.4T total parameters with 95B active, a 1M context wind…

DGX agent

Qwen3.8-2.4T-A95B is live on Together AI. The Qwen team’s new open-weight MoE packs 2.4T total parameters with 95B active, a 1M context window, and strong coding & agent capabilities. Start building:

model-releasestogether-ai--x
14 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Steering the Language Axis: From Linear Decodability to Causal Control

DGX agent

arXiv:2608.12334v1 Announce Type: cross Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood.

model-releasesarxiv-cs-ai
14 Aug 2026
Model Releases

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

DGX agent

arXiv:2608.13495v1 Announce Type: new Abstract: Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured a

model-releasesarxiv-cs-cv
14 Aug 2026
Safety

Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

DGX agent

arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in r

safetyarxiv-cs-cl
13 Aug 2026
Model Releases

Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release

DGX agent

Qwen just released their first 3.8 model. The main addition in 3.8 is prompt-steered reasoning effort. You can tell the model how deeply to think by setting reasoning_effort to xhigh, medium, or low.

model-releasesr-localllama
13 Aug 2026
Model Releases

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

DGX agent

arXiv:2608.12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain p

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

DGX agent

arXiv:2608.11755v1 Announce Type: cross Abstract: Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards inc

model-releasesarxiv-cs-cl
13 Aug 2026
Model Releases

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

DGX agent

arXiv:2608.12127v1 Announce Type: new Abstract: Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is

model-releasesarxiv-cs-cv
13 Aug 2026
Model Releases

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

DGX agent

arXiv:2608.11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same pr

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates

DGX agent

arXiv:2608.11435v1 Announce Type: new Abstract: Forward and inverse modeling of parametric dynamical systems requires surrogate models that are not only accurate for state prediction, but also informa

model-releasesarxiv-cs-lg
13 Aug 2026
Agents

Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning

DGX agent

arXiv:2608.10438v1 Announce Type: new Abstract: Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. Fo

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

DGX agent

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

DGX agent

arXiv:2608.10764v1 Announce Type: new Abstract: Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

Hidden Reasoning from Claude and GPT are Decoded, and it is interesting

DGX agent

Yesteday a paper showed a gap that allows to see 100% of the reasoning tokens form ALL Claude and GPT models Stealing Reasoning Traces from Proprietary LLM APIs. check it out, they have published lots

model-releasesr-localllama
12 Aug 2026
Model Releases

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

DGX agent

arXiv:2608.10300v1 Announce Type: new Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human re

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator

DGX agent

arXiv:2608.09971v1 Announce Type: cross Abstract: Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models whose medium-

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

DGX agent

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

DGX agent

arXiv:2608.10258v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

DGX agent

arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached o

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

DGX agent

arXiv:2608.09164v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms.

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

FEAST: Federated Shared-Space Training for Resource-Heterogeneous Clients

DGX agent

arXiv:2608.09250v1 Announce Type: new Abstract: Federated learning (FL) must serve devices with varying computational capabilities. A fixed model cannot suit all devices, while training one model per

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

I put Gemma 4 E4B and E2B into an e-reader so I can ask my weird questions and share my thoughts in private directly in app.

DGX agent

Here's how it works in the app: Framework: Runs on LiteRT-LM (like Google's AI Edge). Models: Downloads either the E2B (~2.5 GB) or E4B (~3.6 GB) INT4 quantized models directly from ungated litert-com

model-releasesr-localllama
11 Aug 2026
Model Releases

Staying True to the Origin: Continuous Image Stylization with Smooth Transitions

DGX agent

arXiv:2608.08125v1 Announce Type: new Abstract: Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a g

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

DGX agent

arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

DGX agent

arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to c…

DGX agent

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to choose the right model for the right task. As part of this, we

model-releasesmistral-ai--x
11 Aug 2026
Model Releases

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

DGX agent

arXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

DGX agent

arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Explanation Stability of Test-Time Adaptation in Computational Pathology: A Large-Scale Benchmark

DGX agent

arXiv:2608.07062v1 Announce Type: new Abstract: Test-time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in comput

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?

DGX agent

arXiv:2608.06972v1 Announce Type: new Abstract: Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess rep

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.

DGX agent

Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots a

model-releasesr-localllama
10 Aug 2026
Model Releases

24 GB of VRAM is not really 24 GB for a local LLM. Here is the worksheet I use

DGX agent

I kept seeing model file size compared directly with the number printed on the GPU box. That misses several memory buckets. A simple planning model is: usable capacity = advertised VRAM x 0.90 total t

model-releasesr-ollama
9 Aug 2026
Model Releases

https://x.com/sama/status/2085862292311396515

DGX agent

https://x.com/sama/status/2085862292311396515 astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few

model-releasesopenai--x
7 Aug 2026
Model Releases

The Bitter Lesson of Tool Calling

DGX agent

arXiv:2608.06370v1 Announce Type: new Abstract: Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by

model-releasesarxiv-cs-cl
7 Aug 2026
Model Releases

Item Response Theory for AI Safety

DGX agent

arXiv:2608.05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to tr

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

They almost catched up on Frontier performance, so now catching up on prices

DGX agent

Users also report that the free version was significantly downgraded after the release of the new models this is very important for us when considering local hosting. A lot of people decided not to bu

model-releasesr-localllama
6 Aug 2026
Model Releases

Unsloth's Gemma 4 mmproj silently broke vision & audio on newer llama.cpp builds — anyone else hit this?

DGX agent

So I had been building ScreenMind, kinda like local ai desktop assistant that uses Gemma 4 for screen analysis, voice memo transcription, and meeting transcription — all through llama-server. Everythi

model-releasesr-localllama
6 Aug 2026
Model Releases

Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

DGX agent

arXiv:2608.03892v1 Announce Type: new Abstract: We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, r

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

SocietyBench: Forecasting Counterfactual Social-World Evolution

DGX agent

arXiv:2608.04009v1 Announce Type: new Abstract: Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a b

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Sulphur 3 is looking for funding

DGX agent

Hello, I'm the guy who made Sulphur 2. With the recent release of a certain video model, we are looking to mobilize and train Sulphur 3 on this new model. We are targeting $10,000 USD. This certain ne

model-releasesr-stablediffusion
5 Aug 2026
Model Releases

Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

DGX agent

arXiv:2512.08163v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are increasingly used as functional models of human vision, yet standard monocular depth estimation (MDE) benchmarks lar

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

b10270

DGX agent

mtmd: support Qwen3-TTS (note: breaking change to llama-tts binary) (#26254) convert text model main model load ok convert encoder ok speaker encoder loading ok speaker enc graph adapt vocab for backb

model-releasesllama-cpp-releases
4 Aug 2026
Model Releases

Evaluating VLMs on Multimodal Aristotelian Persuasion Tasks

DGX agent

arXiv:2608.01238v1 Announce Type: new Abstract: Vision Language Models (VLMs) have demonstrated exceptional performance across various tasks. However, they have not yet been thoroughly evaluated on mo

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Image-Space Rule Discovery

DGX agent

arXiv:2608.00490v1 Announce Type: new Abstract: Can image-editing models discover visual rules in image space and complete problem-solving end-to-end? We tackle this question in the spirit of a human

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

DGX agent

arXiv:2608.01522v1 Announce Type: cross Abstract: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning t

researcharxiv-cs-cl
4 Aug 2026
Model Releases

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

DGX agent

arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challen

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

DGX agent

arXiv:2607.28658v1 Announce Type: cross Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, e

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

Anthropic says Claude accidentally hacked real companies too

DGX agent

Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation co

model-releasesthe-verge-ai
31 Jul 2026
← Previous
1…245246247248249…1272
Next →