AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,597 results
9 Apr 2026

Refiant raises $5M to refine AI models with ‘nature-inspired’ energy efficiency

IndustryDGX agent

Artificial intelligence model compression startup Refiant AI said today it has raised $5 million in seed funding from VoLo Earth Ventures to try to put an end to the “arms race” that has ignited a mul

8 Apr 2026

Pelicans for Meta's new Muse Spark models - plus I did a bit of a deep dive into the Code Interpreter and fascinating 'container.visual_grou…

ToolsDGX agent

Pelicans for Meta's new Muse Spark models - plus I did a bit of a deep dive into the Code Interpreter and fascinating 'container.visual_grounding' tools in their http://meta.ai chat UI https://simonwi

Want more proof that Anthropic's PR has no idea what it's talking about? The talk of Mythos being 'their most aligned model ever'. They coul…

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

Want more proof that Anthropic's PR has no idea what it's talking about? The talk of Mythos being 'their most aligned model ever'. They could perhaps truthfully speak about 'new high scores on our ali

7 Apr 2026

Wrote up some thoughts on Anthropic's Project Glassing, where their latest Opus-beating model is available to partnered security research or…

ToolsDGX agent

Wrote up some thoughts on Anthropic's Project Glassing, where their latest Opus-beating model is available to partnered security research organizations only Given recent alarm bells raised by credible

15 Aug 2026

club-5060ti refresh: tested RTX 5060 Ti presets, a proper high-context harness, and Qwen3.8 27B

Model ReleasesDGX agent

Quick update on the RTX 5060 Ti local LLM repo. It has changed quite a bit since my previous posts. The project started as a collection of practical notes and benchmark results. That was useful, but a

14 Aug 2026

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

Model ReleasesDGX agent

arXiv:2608.12345v1 Announce Type: new Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasur

Qwen3.8-2.4T-A95B is live on Together AI. The Qwen team’s new open-weight MoE packs 2.4T total parameters with 95B active, a 1M context wind…

Model ReleasesDGX agent

Qwen3.8-2.4T-A95B is live on Together AI. The Qwen team’s new open-weight MoE packs 2.4T total parameters with 95B active, a 1M context window, and strong coding & agent capabilities. Start building:

Steering the Language Axis: From Linear Decodability to Causal Control

Model ReleasesDGX agent

arXiv:2608.12334v1 Announce Type: cross Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood.

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

Model ReleasesDGX agent

arXiv:2608.13495v1 Announce Type: new Abstract: Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured a

13 Aug 2026

Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

SafetyDGX agent

arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in r

Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release

Model ReleasesDGX agent

Qwen just released their first 3.8 model. The main addition in 3.8 is prompt-steered reasoning effort. You can tell the model how deeply to think by setting reasoning_effort to xhigh, medium, or low.

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Model ReleasesDGX agent

arXiv:2608.12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain p

MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

Model ReleasesDGX agent

arXiv:2608.11755v1 Announce Type: cross Abstract: Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards inc

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

Model ReleasesDGX agent

arXiv:2608.12127v1 Announce Type: new Abstract: Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

Model ReleasesDGX agent

arXiv:2608.11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same pr

Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates

Model ReleasesDGX agent

arXiv:2608.11435v1 Announce Type: new Abstract: Forward and inverse modeling of parametric dynamical systems requires surrogate models that are not only accurate for state prediction, but also informa

12 Aug 2026

Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning

AgentsDGX agent

arXiv:2608.10438v1 Announce Type: new Abstract: Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. Fo

DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

Model ReleasesDGX agent

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

Model ReleasesDGX agent

arXiv:2608.10764v1 Announce Type: new Abstract: Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ

Hidden Reasoning from Claude and GPT are Decoded, and it is interesting

Model ReleasesDGX agent

Yesteday a paper showed a gap that allows to see 100% of the reasoning tokens form ALL Claude and GPT models Stealing Reasoning Traces from Proprietary LLM APIs. check it out, they have published lots

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

Model ReleasesDGX agent

arXiv:2608.10300v1 Announce Type: new Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human re

Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator

Model ReleasesDGX agent

arXiv:2608.09971v1 Announce Type: cross Abstract: Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models whose medium-

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

Model ReleasesDGX agent

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

Model ReleasesDGX agent

arXiv:2608.10258v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do

The Evaluation Protocol Determines the Result: An Independent Reproduction of LeWorldModel on TwoRoom

Model ReleasesDGX agent

arXiv:2608.10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached o

11 Aug 2026

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

Model ReleasesDGX agent

arXiv:2608.09164v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms.

FEAST: Federated Shared-Space Training for Resource-Heterogeneous Clients

Model ReleasesDGX agent

arXiv:2608.09250v1 Announce Type: new Abstract: Federated learning (FL) must serve devices with varying computational capabilities. A fixed model cannot suit all devices, while training one model per

I put Gemma 4 E4B and E2B into an e-reader so I can ask my weird questions and share my thoughts in private directly in app.

Model ReleasesDGX agent

Here's how it works in the app: Framework: Runs on LiteRT-LM (like Google's AI Edge). Models: Downloads either the E2B (~2.5 GB) or E4B (~3.6 GB) INT4 quantized models directly from ungated litert-com

Staying True to the Origin: Continuous Image Stylization with Smooth Transitions

Model ReleasesDGX agent

arXiv:2608.08125v1 Announce Type: new Abstract: Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a g

Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

Model ReleasesDGX agent

arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

Model ReleasesDGX agent

arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to c…

Model ReleasesDGX agent

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to choose the right model for the right task. As part of this, we

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

Model ReleasesDGX agent

arXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

Model ReleasesDGX agent

arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c

10 Aug 2026

Explanation Stability of Test-Time Adaptation in Computational Pathology: A Large-Scale Benchmark

Model ReleasesDGX agent

arXiv:2608.07062v1 Announce Type: new Abstract: Test-time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in comput

Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?

Model ReleasesDGX agent

arXiv:2608.06972v1 Announce Type: new Abstract: Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess rep

Needle 2: 14MB agentic LLM for phones, wearables, smart home and robots.

Model ReleasesDGX agent

Hey LocalLlaMa, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots a

9 Aug 2026

24 GB of VRAM is not really 24 GB for a local LLM. Here is the worksheet I use

Model ReleasesDGX agent

I kept seeing model file size compared directly with the number printed on the GPU box. That misses several memory buckets. A simple planning model is: usable capacity = advertised VRAM x 0.90 total t

7 Aug 2026

https://x.com/sama/status/2085862292311396515

Model ReleasesDGX agent

https://x.com/sama/status/2085862292311396515 astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few

The Bitter Lesson of Tool Calling

Model ReleasesDGX agent

arXiv:2608.06370v1 Announce Type: new Abstract: Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by

6 Aug 2026

Item Response Theory for AI Safety

Model ReleasesDGX agent

arXiv:2608.05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to tr

They almost catched up on Frontier performance, so now catching up on prices

Model ReleasesDGX agent

Users also report that the free version was significantly downgraded after the release of the new models this is very important for us when considering local hosting. A lot of people decided not to bu

Unsloth's Gemma 4 mmproj silently broke vision & audio on newer llama.cpp builds — anyone else hit this?

Model ReleasesDGX agent

So I had been building ScreenMind, kinda like local ai desktop assistant that uses Gemma 4 for screen analysis, voice memo transcription, and meeting transcription — all through llama-server. Everythi

5 Aug 2026

Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

Model ReleasesDGX agent

arXiv:2608.03892v1 Announce Type: new Abstract: We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, r

SocietyBench: Forecasting Counterfactual Social-World Evolution

Model ReleasesDGX agent

arXiv:2608.04009v1 Announce Type: new Abstract: Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a b

Sulphur 3 is looking for funding

Model ReleasesDGX agent

Hello, I'm the guy who made Sulphur 2. With the recent release of a certain video model, we are looking to mobilize and train Sulphur 3 on this new model. We are targeting $10,000 USD. This certain ne

4 Aug 2026

Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2512.08163v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are increasingly used as functional models of human vision, yet standard monocular depth estimation (MDE) benchmarks lar

b10270

Model ReleasesDGX agent

mtmd: support Qwen3-TTS (note: breaking change to llama-tts binary) (#26254) convert text model main model load ok convert encoder ok speaker encoder loading ok speaker enc graph adapt vocab for backb

Evaluating VLMs on Multimodal Aristotelian Persuasion Tasks

Model ReleasesDGX agent

arXiv:2608.01238v1 Announce Type: new Abstract: Vision Language Models (VLMs) have demonstrated exceptional performance across various tasks. However, they have not yet been thoroughly evaluated on mo

Image-Space Rule Discovery

Model ReleasesDGX agent

arXiv:2608.00490v1 Announce Type: new Abstract: Can image-editing models discover visual rules in image space and complete problem-solving end-to-end? We tackle this question in the spirit of a human

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

ResearchDGX agent

arXiv:2608.01522v1 Announce Type: cross Abstract: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning t

3 Aug 2026

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

Model ReleasesDGX agent

arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challen

Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation

Model ReleasesDGX agent

arXiv:2607.28658v1 Announce Type: cross Abstract: Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets. However, e

31 Jul 2026

Anthropic says Claude accidentally hacked real companies too

Model ReleasesDGX agent

Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation co

Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction

Model ReleasesDGX agent

arXiv:2607.28079v1 Announce Type: new Abstract: Chemical property prediction plays a critical role in accelerating scientific discovery in chemistry, materials science, and drug development. However,

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

Model ReleasesDGX agent

Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Model ReleasesDGX agent

arXiv:2607.27919v1 Announce Type: new Abstract: Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independent

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

Model ReleasesDGX agent

arXiv:2607.27378v1 Announce Type: new Abstract: Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-g

30 Jul 2026

APEX-Accounting

Model ReleasesDGX agent

arXiv:2607.27189v1 Announce Type: new Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountant

GPT-Red: Automated Red Teaming via Self-Play at Scale

Model ReleasesDGX agent

arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce extbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal

← Previous
1…194195196197198…1010
Next →