AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
Model Releases

v0.32.1

DGX agent

What's Changed Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations Fixed a recurrent MLX model cache leak that could increase memory use across

model-releasesollama-releases
16 Jul 2026
Model Releases

Value Drifts: Tracing Value Alignment During LLM Post-Training

DGX agent

arXiv:2510.26707v2 Announce Type: replace-cross Abstract: As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
model-releasesarxiv-cs-lg
16 Jul 2026
Model Releases

VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation

DGX agent

arXiv:2607.13527v1 Announce Type: new Abstract: Recent video generation models (VGMs) have made substantial progress in visual fidelity, yet their ability to follow long, compositional instructions re

model-releasesarxiv-cs-cv
16 Jul 2026
Model Releases

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

DGX agent

arXiv:2607.14088v1 Announce Type: new Abstract: Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimi

model-releasesarxiv-cs-cv
16 Jul 2026
Model Releases

WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency

DGX agent

arXiv:2607.13099v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable success but raise growing concerns about content provenance and misuse, motivating the need for

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

Weight Feedback Computes the Jacobian Transpose Locally in Modern Deep Networks

DGX agent

arXiv:2607.13380v1 Announce Type: new Abstract: Predictive Coding (PC) offers a biologically motivated alternative to backpropagation via local weight updates, yet routing error between layers still r

model-releasesarxiv-cs-lg
16 Jul 2026
Model Releases

We’re excited to collaborate with NVIDIA to build the next generation of Fugu orchestration models together, by incorporating leading open-w…

DGX agent

We’re excited to collaborate with NVIDIA to build the next generation of Fugu orchestration models together, by incorporating leading open-weights models. Sakana AI Teams With NVIDIA to Advance Open M

model-releasesdavid-ha--x
16 Jul 2026
Model Releases

When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents

DGX agent

arXiv:2602.11619v2 Announce Type: replace Abstract: Running the same LLM agent on identical inputs yields 2.3-4.2 distinct action sequences per 10 runs; this behavioral variance constitutes a training

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

DGX agent

arXiv:2602.17659v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow langu

model-releasesarxiv-cs-cv
16 Jul 2026
Model Releases

1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World

DGX agent

arXiv:2602.18548v2 Announce Type: replace-cross Abstract: Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to inco

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

🥉 3rd place: CashFromChaos, by David Diaz (@davddiazm) CashFromChaos starts from a single seller input and automates everything up until a …

DGX agent

🥉 3rd place: CashFromChaos, by David Diaz (@davddiazm) CashFromChaos starts from a single seller input and automates everything up until a completed sale. You send a photo and a one-line clue, and Her

model-releasesnous-research--x
15 Jul 2026
Model Releases

A Calibrated Multimodal Ensemble for Ambivalence/Hesitancy Recognition: System Description and Private-Test Submission Strategy

DGX agent

arXiv:2607.12176v1 Announce Type: new Abstract: Ambivalence and hesitancy (A/H) undermine digital behaviour-change interventions, and recognizing them automatically from video is the goal of the ABAW

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

A closer look at improved intelligence in GPT-Live: the model can keep a conversation going while helping with multiple tasks at once, like …

DGX agent

A closer look at improved intelligence in GPT-Live: the model can keep a conversation going while helping with multiple tasks at once, like checking flights, pulling up local weather, and shaping an i

model-releasesopenai--x
15 Jul 2026
Model Releases

A Comparative Analysis of Institutional and Course Generative AI Policies within Higher Education: Implications for Instruction in Computing Education

DGX agent

arXiv:2607.12296v1 Announce Type: cross Abstract: With the increased use of generative AI (GenAI) applications such as ChatGPT, higher education institutions (HEIs) have released a range of guidelines

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs

DGX agent

arXiv:2607.12550v1 Announce Type: cross Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference. It grows with batch size, context length, and depth, and at lon

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

A Shared Subcircuit Lets LLMs Count Down Across Tasks

DGX agent

arXiv:2607.12279v1 Announce Type: new Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table. These are all tasks that language model

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

DGX agent

arXiv:2607.10350v1 Announce Type: cross Abstract: Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime la

model-releasesarxiv-cs-ro
15 Jul 2026
Model Releases

ABot-N1: Toward a General Visual Language Navigation Foundation Model

DGX agent

arXiv:2607.10383v2 Announce Type: replace-cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse emb

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Accepted Prefixes Are Not All You Need: A Negative Result on PEFT-Based Block-Diffusion Drafting

DGX agent

arXiv:2607.12422v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive language model inference by using a cheap drafter to propose multiple future tokens and a target model t

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction

DGX agent

arXiv:2607.12293v1 Announce Type: new Abstract: Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs or

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

DGX agent

arXiv:2511.04689v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

DGX agent

arXiv:2607.12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Agentic orchestration: Enterprise AI organizations have a deployment problem, not a platform problem — and most are calling chatbots agents

DGX agent

Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on re

model-releasesventurebeat-ai
15 Jul 2026
Model Releases

Agentic systems for breast cancer treatment recommendations

DGX agent

arXiv:2607.12051v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

Agents-A1-4B (Qwen3.7-4B ???) : Scaling the Horizon, Not the Parameters

DGX agent

MODEL + GGUF : https://huggingface.co/InternScience/models?search=a1-4b Technical Report Benchmark Qwen3.5-4B Agents-A1-4B Qwen3.5 Qwen3.6 Nex-N2-mini Agents-A1 🧠 Dense Models (~4B) 🔀 MoE Models (35B-

model-releasesr-localllama
15 Jul 2026
Model Releases

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents

DGX agent

arXiv:2607.10526v2 Announce Type: replace Abstract: Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversationa

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

AI DevOps startup MyDecisive launches with $12M and open-source SmartHub

DGX agent

Artificial intelligence DevOps startup MyDecisive formally launched today and announced 12 million in new funding to bring to market an open-source foundation for managing observability data and a com

model-releasessiliconangle
15 Jul 2026
Model Releases

An Agentic AI Scientific Community for Automated Neural Operator Discovery

DGX agent

arXiv:2607.12122v1 Announce Type: new Abstract: We present an agentic approach to autonomous neural operator discovery based on an AI scientific community, which consists of a swarm of virtual laborat

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

An Empirical Study for Android-to-OpenHarmony GUI Test Migration

DGX agent

arXiv:2607.11245v2 Announce Type: replace-cross Abstract: To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existing G

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Anthropic, Blackstone, and Hellman & Friedman's $1.5B AI implementation company, announced in May, launches with the name Ode with Anthropic and 100 engineers (Rebecca Bellan/TechCrunch)

DGX agent

Rebecca Bellan / TechCrunch: Anthropic, Blackstone, and Hellman & Friedman's $1.5B AI implementation company, announced in May, launches with the name Ode with Anthropic and 100 engineers — AI models

model-releasestechmeme
15 Jul 2026
Model Releases

APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model

DGX agent

arXiv:2603.08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots. Classical navigation approaches offer safety a

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

DGX agent

arXiv:2607.11997v1 Announce Type: cross Abstract: Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice m

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Audio perception layer for LLM agents, with a memory that grows through use

DGX agent

LLMs handle speech well once you run speech-to-text. They don't hear the rest: a bird outside, a glass breaking two rooms away, a smoke alarm two floors down. I've been working on an experimental open

model-releasesr-localllama
15 Jul 2026
Model Releases

Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks

DGX agent

arXiv:2607.12278v1 Announce Type: new Abstract: Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answer

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration

DGX agent

arXiv:2607.12058v1 Announce Type: cross Abstract: Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe oper

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

DGX agent

arXiv:2607.12820v1 Announce Type: new Abstract: Omni-modal video captioning is not merely combining visual captioning with audio transcription: a useful caption must describe how visual actions, speec

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

b10016

DGX agent

[SYCL] Flash Attention with XMX engine via oneDNN (#25222) [SYCL] F16 (default) Flash Attention with XMX engine via oneDNN graph API; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at p=80

model-releasesllama-cpp-releases
15 Jul 2026
Model Releases

b10017

DGX agent

sycl: Increase minimum buffer size for USM system allocations (#25525) Raise the threshold for minimum buffer size from 1 GiB to 4 GiB, based on real-world experiments of overcommitting device memory

model-releasesllama-cpp-releases
15 Jul 2026
Model Releases

b10021

DGX agent

DeepseekV4: reduce graph splits (#25702) macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu

model-releasesllama-cpp-releases
15 Jul 2026
Model Releases

Belief-reality separation lives in routing over a shared value slot in language models

DGX agent

arXiv:2607.11945v1 Announce Type: new Abstract: Capable language models hold what a character believes apart from what is true: told 'Anna believes the cup is blue; in reality it is red,' they answer

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark

DGX agent

arXiv:2607.11915v1 Announce Type: cross Abstract: Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. In practice, fusion diagnostics f

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

Beyond Perfect Priors: Adaptive Gaussian Graph for 4D Driving Reconstruction in the Wild

DGX agent

arXiv:2607.12214v1 Announce Type: new Abstract: Reconstructing 4D driving scenes in the wild (e.g., internet and AI-generated videos) is critical for diverse autonomous driving simulation. While recen

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)

DGX agent

Below Upstream Status sections are from https://github.com/PrismML-Eng/Bonsai-demo Upstream Status for Binary Q1_0 is supported out of the box in upstream llama.cpp across many backends: CPU (generic,

model-releasesr-localllama
15 Jul 2026
Model Releases

Breaking Deja Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

DGX agent

arXiv:2607.12818v1 Announce Type: new Abstract: Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop clos

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

BREAKING: Grok 4.5 has climbed to #2 on the FrontierSWE benchmark. The result places Grok 4.5 among the world's top-performing AI models for…

DGX agent

BREAKING: Grok 4.5 has climbed to #2 on the FrontierSWE benchmark. The result places Grok 4.5 among the world's top-performing AI models for software engineering tasks, highlighting its growing streng

model-releaseselon-musk--x
15 Jul 2026
Model Releases

BREAKING: Inkling by @thinkymachines is 9th overall on Agentic Web App Arena by Design Arena with an Elo of 1257 It's an open-weight model i…

DGX agent

BREAKING: Inkling by @thinkymachines is 9th overall on Agentic Web App Arena by Design Arena with an Elo of 1257 It's an open-weight model in the same performance band as Claude Opus 4.6 by @Anthropic

model-releasessoumith-chintala--x
15 Jul 2026
Model Releases

Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans

DGX agent

arXiv:2607.11263v2 Announce Type: replace Abstract: Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity o

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Calibratable Disambiguation Loss for Multi-Instance Partial-Label Learning

DGX agent

arXiv:2512.17788v2 Announce Type: replace Abstract: Multi-instance partial-label learning (MIPL) is a weakly supervised framework that extends the principles of multi-instance learning (MIL) and parti

model-releasesarxiv-cs-lg
15 Jul 2026
← Previous
1…102103104105106…471
Next →