AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,512 results
25 Jul 2026

Quoting Boris Cherny

Model ReleasesDGX agent

More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red t

Ruff v0.16.0

Model ReleasesDGX agent

Ruff v0.16.0 Astral shipped a significant new version of their Ruff Python linting tool a few days ago on July 23rd. I noticed today because my various CI jobs all started failing thanks to new defaul

Sources: DeepSeek told investors it is suspending its second funding round after remarks attributed to Liang Wenfeng on US-China AI competition went viral (Pei Li/Bloomberg)

Model ReleasesDGX agent

Pei Li / Bloomberg: Sources: DeepSeek told investors it is suspending its second funding round after remarks attributed to Liang Wenfeng on US-China AI competition went viral — DeepSeek has told prosp


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Gemma series is an amazing set of highly performant open-weight models. They have proven extremely effective in industrial settings wher…

Model ReleasesDGX agent

The Gemma series is an amazing set of highly performant open-weight models. They have proven extremely effective in industrial settings where site-deployed agents need exactly this as a base for domai

The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts 📊 LlamaParse does a surpr…

Model ReleasesDGX agent

The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts 📊 LlamaParse does a surprisingly good job on agentic (1.25c per page) and agentic plus

Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact hav…

Model ReleasesDGX agent

Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact have consistently made open weights models with Gemma available

24 Jul 2026

A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset

Model ReleasesDGX agent

arXiv:2607.21274v1 Announce Type: cross Abstract: We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance judgments. W

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction

Model ReleasesDGX agent

arXiv:2607.20453v1 Announce Type: cross Abstract: Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge,

A new model launch is not a product update ‼ It's one of three things, and you don't know which until you test it. Sometimes it's nothing: t…

Model ReleasesDGX agent

A new model launch is not a product update ‼ It's one of three things, and you don't know which until you test it. Sometimes it's nothing: the model improved inside the same distribution, your harness

A reminder this is 20% off via the Nous Portal <3

Model ReleasesDGX agent

mr‑r0b0t announced that Claude Opus 5 is available at a 20 % discount through the Nous Portal. The model can be accessed via the Hermes Agent on the Nous Portal, as well as through OpenRouter and Anth

A Sovereign, Open-Source Foundation Model for German and English

Model ReleasesDGX agent

arXiv:2607.09424v3 Announce Type: replace-cross Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English

Achieving Text-based Person Retrieval with Any Granularity

Model ReleasesDGX agent

arXiv:2607.21057v1 Announce Type: new Abstract: Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This p

ADABORD: a novel AdaBoost approach for ordinal classification

Model ReleasesDGX agent

arXiv:2607.21003v1 Announce Type: new Abstract: Ordinal Classification (OC) deals with classification tasks where the classes follow a natural order. Despite the progress in OC, many existing approach

Adaptive Multi-Horizon Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.20656v1 Announce Type: cross Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL

Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

Model ReleasesDGX agent

arXiv:2607.21482v1 Announce Type: new Abstract: Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their

Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation

Model ReleasesDGX agent

arXiv:2607.20866v1 Announce Type: new Abstract: Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamenta

AI Assistants Overassist

Model ReleasesDGX agent

arXiv:2607.21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assis

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs

Model ReleasesDGX agent

arXiv:2607.20498v1 Announce Type: new Abstract: Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-h

Am I supposed to interpret that it's better than Fable 5 from the benchmarks? Or it's ~close but cheaper?

Model ReleasesDGX agent

Am I supposed to interpret that it's better than Fable 5 from the benchmarks? Or it's ~close but cheaper? Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the front

An Analytically Trained Variational Surrogate for Quantum Phase Estimation on NISQ Hardware

Model ReleasesDGX agent

arXiv:2607.20943v1 Announce Type: cross Abstract: Quantum Phase Estimation (QPE) is a foundational algorithm for molecular ground-state energy estimation, but its deep circuit requirements make direct

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus …

Model ReleasesDGX agent

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus 5 max thinking leads to a degradation in performance compare

An LLM-Driven Workflow for Automated Process Control Strategy Generation and Tuning from Dynamic Process Models

Model ReleasesDGX agent

arXiv:2607.21292v1 Announce Type: new Abstract: We present a structured large-language-model-driven workflow for automated multi-variable control design from dynamic process models. The workflow decom

Announcing Fugu-Ultra v1.1 🐡 We’ve been thrilled by the reception to the Fugu model family. Thanks to everyone who tried it, shared feedbac…

Model ReleasesDGX agent

Announcing Fugu-Ultra v1.1 🐡 We’ve been thrilled by the reception to the Fugu model family. Thanks to everyone who tried it, shared feedback, and trusted Fugu with real work. Today, we’re releasing Fu

Anthropic launches Claude Opus 5, which it says comes close to Fable 5 performance at half the price; it is the new default model on Claude Max (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic launches Claude Opus 5, which it says comes close to Fable 5 performance at half the price; it is the new default model on Claude Max — Claude Opus 5 is available today. It's a th

Anthropic releases Opus 5 with ‘close’ to Fable 5’s capabilities

Model ReleasesDGX agent

Weeks after Anthropic's latest toe-to-toe with the US government, and days after an OpenAI security incident that dominated tech industry discussions, Anthropic on Thursday released its newest model,

Anthropic says Opus 5 is the company's 'most aligned model to date'; it is Anthropic's fourth model release in less than two months (Madison Mills/Axios)

Model ReleasesDGX agent

Madison Mills / Axios: Anthropic says Opus 5 is the company's “most aligned model to date”; it is Anthropic's fourth model release in less than two months — Anthropic on Thursday is releasing Claude O

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

Model ReleasesDGX agent

arXiv:2607.20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.

ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management

Model ReleasesDGX agent

arXiv:2607.20764v1 Announce Type: new Abstract: We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevan

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

Model ReleasesDGX agent

arXiv:2607.20596v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE fami

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Model ReleasesDGX agent

arXiv:2607.21461v1 Announce Type: new Abstract: Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candida

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchb…

Model ReleasesDGX agent

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a

Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models

Model ReleasesDGX agent

arXiv:2607.20524v1 Announce Type: new Abstract: Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval rema

Attribution Markets: A Fisher-Market Formulation for Fractional Credit Assignment Between Planned Tasks and Performed Actions

Model ReleasesDGX agent

arXiv:2607.20694v1 Announce Type: new Abstract: Personal and organizational planning systems maintain two records that drift apart: what was planned (a task's effort budget) and what was done (a logge

[audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains

Model ReleasesDGX agent

audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new: Added Higgs Audio v3 TTS 4B, Fish Audio

Automated Synthesis and Adversarial Validation of Executable Causal Research Pipelines

Model ReleasesDGX agent

arXiv:2607.21173v1 Announce Type: new Abstract: While automated research systems promise to accelerate empirical analysis, they are prone to silent failures: instances in which analysis code executes

Autonomous disproofs of the sum-product conjecture over mathbb R with GPT-5.5 Pro

Model ReleasesDGX agent

arXiv:2607.20525v1 Announce Type: new Abstract: OpenAI's recent disproof of the Erdos unit distance conjecture marked a milestone for AI in mathematics. It also inspired another breakthrough: a human

Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants

Model ReleasesDGX agent

arXiv:2607.20488v1 Announce Type: new Abstract: Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for example by mixing

AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

Model ReleasesDGX agent

arXiv:2607.21588v1 Announce Type: new Abstract: Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale b

b10103

Model ReleasesDGX agent

metal : add f16 type support to leaky relu (#25981) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFra

b10105

Model ReleasesDGX agent

args: refactor mlock/mmap/directio into load-mode (#20834) args: overhaul mmap/mlock/dio into single arg Signed-off-by: Aaron Teo aaron.teo1@ibm.com docs: update docs with llama-gen-docs Signed-off-by

b10106

Model ReleasesDGX agent

CUDA: fix external compilation of q1_0 MMQ (#25778) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFra

b10107

Model ReleasesDGX agent

hexagon: fix Windows crash when op_poll is enabled (#26029) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) i

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

Model ReleasesDGX agent

arXiv:2607.20476v1 Announce Type: new Abstract: We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three c

Benchmarking the Personalization Capabilities of Large Language Models

Model ReleasesDGX agent

arXiv:2607.20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradit

Benchmarking Unlearning for Vision Transformers

Model ReleasesDGX agent

arXiv:2602.20114v2 Announce Type: replace-cross Abstract: Machine unlearning (MU) refers to the post-training capability to remove (the influence of) training examples that are incorrect, biased, or l

Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

Model ReleasesDGX agent

arXiv:2607.20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail

Break Through the Compression Bottleneck: From Theory to Practice

Model ReleasesDGX agent

arXiv:2607.20434v1 Announce Type: cross Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead.

CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful

Model ReleasesDGX agent

I’m not affiliated with this project, but I’ve been running it recently and I’m surprised it hasn’t received more attention here: https://github.com/fewtarius/CachyLLama CachyLLama is a fork of llama.

CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents

Model ReleasesDGX agent

arXiv:2607.20458v1 Announce Type: cross Abstract: Large language model (LLM) agents operating over extended dialogues accumulate vast amounts of information, yet existing memory systems either retain

Can LLMs solve mazes?

Model ReleasesDGX agent

https://reddit.com/link/1v5rvuq/video/bgmwc754i9fh1/player My goal was to create a benchmark to measure the spatial awareness and memory of models. Eventually, I came up with the simple idea of a maze

CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits

Model ReleasesDGX agent

arXiv:2607.20518v1 Announce Type: new Abstract: AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchma

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

Model ReleasesDGX agent

arXiv:2607.21340v1 Announce Type: new Abstract: In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defens

Cardinality-Decomposed Loss: Matching Training Objectives to Relation Structure in Heterogeneous Recommendation Graphs

Model ReleasesDGX agent

arXiv:2607.20737v1 Announce Type: new Abstract: Graph Neural Networks trained on heterogenous bipartite graphs form a common basis in recommendation systems. These graphs often express relations that

Case study: solving P-99 with LPTP and an LLM

Model ReleasesDGX agent

arXiv:2607.21196v1 Announce Type: cross Abstract: Ninety-Nine Prolog Problems (P-99) is a famous set of Prolog exercises. We solved the first thirty three just by prompting an LLM (Large Language Mode

Chronofy: A Temporal-Logical Decay Architecture for Information Validity in Time-Aware Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.20560v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems retrieve and integrate external knowledge to ground large language model (LLM) outputs. However, current RA

Classifier-pruned Bayesian optimization for particle accelerator tuning: Exploring temporally structured manifold of 6D beam phase space

Model ReleasesDGX agent

arXiv:2412.01748v2 Announce Type: replace Abstract: Complex dynamical systems, such as particle accelerators, often require intricate and time-consuming tuning procedures to achieve optimal performanc

ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

Model ReleasesDGX agent

arXiv:2607.20526v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone i

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

Model ReleasesDGX agent

arXiv:2607.21155v1 Announce Type: cross Abstract: Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiri

CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging

Model ReleasesDGX agent

arXiv:2607.20561v1 Announce Type: cross Abstract: LoRA adapters provide an efficient way to specialize a pretrained model for many downstream tasks, but deploying one adapter per task requires adapter

CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages

Model ReleasesDGX agent

arXiv:2607.21016v1 Announce Type: new Abstract: Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping aw

← Previous
1…6869707172…376
Next →