AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,323 results
31 Jul 2026

Epistemic diversity across language models mitigates knowledge collapse

Model ReleasesDGX agent

arXiv:2512.15011v3 Announce Type: replace Abstract: Artificial intelligence (AI) increasingly generates the very content used to train future AI systems. This feedback loop can degrade model quality,

eta-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

Model ReleasesDGX agent

arXiv:2607.28582v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reli

Evidence-Ledger Adjudication for Claim-Evidence Traceability

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.26512v1 Announce Type: new Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a

EvoCause: LLM-Guided Evolution of Causal Graphs for Root Cause Analysis

Model ReleasesDGX agent

arXiv:2607.27290v1 Announce Type: new Abstract: Modern telecommunication, cloud, and microservice systems emit correlated alarm cascades when components fail. Root cause analysis (RCA) aims to identif

Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control

Model ReleasesDGX agent

arXiv:2607.27914v1 Announce Type: new Abstract: Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators. Model

Expected Survival-Time Bounds for Robust Optimization Over Time under Isotropic Gaussian Dynamics

Model ReleasesDGX agent

arXiv:2607.27280v1 Announce Type: cross Abstract: Robust Optimization Over Time (ROOT) is a recent branch of evolutionary dynamic optimization that seeks solutions capable of remaining effective acros

Experience sharing: How do you use your local models and for what kind of tasks?

Model ReleasesDGX agent

Here is my experience, which I would like to share with you and I also would like to hear your thoughts and valuable tips&tricks. Hardware: Mac Mini M4 (32GB Unified Memory) Model Server: Ollama Orche

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

Model ReleasesDGX agent

arXiv:2607.27372v1 Announce Type: cross Abstract: The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generat

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

Model ReleasesDGX agent

arXiv:2607.26367v1 Announce Type: new Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent t

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

Model ReleasesDGX agent

arXiv:2607.09306v3 Announce Type: replace Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

Model ReleasesDGX agent

arXiv:2607.28319v1 Announce Type: new Abstract: This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias

Fantastic Adaptive Taxonomies and How to Use Them

Model ReleasesDGX agent

arXiv:2607.16387v2 Announce Type: replace-cross Abstract: An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajectory s

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

Model ReleasesDGX agent

arXiv:2607.26611v1 Announce Type: new Abstract: AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in us

FinanceHarness: Autonomous Financial Deep Research Framework

Model ReleasesDGX agent

arXiv:2607.27853v1 Announce Type: new Abstract: Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

Model ReleasesDGX agent

arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it n

FlexiGrad: Adaptive Gradient Modulation for Hierarchical Fine-Grained Classification

Model ReleasesDGX agent

arXiv:2607.17563v2 Announce Type: replace Abstract: Many fine-grained recognition tasks contain hierarchical labels such as order, family and species. Although this supervision should be beneficial, j

FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing

Model ReleasesDGX agent

arXiv:2508.02092v3 Announce Type: replace-cross Abstract: Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuable

From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference

Model ReleasesDGX agent

arXiv:2607.28097v1 Announce Type: new Abstract: Mathematically equivalent expert-reduction orders can produce observably different sparse-MoE executions. We isolate this effect in native DeepSeek-V4-F

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Model ReleasesDGX agent

arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of do

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Model ReleasesDGX agent

arXiv:2607.28568v1 Announce Type: new Abstract: Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a

Generalization and Trade-off in Adversarial Training: An RKHS Perspective via Kernel Integral Operators

Model ReleasesDGX agent

arXiv:2607.27995v1 Announce Type: cross Abstract: Adversarial training has emerged as a powerful approach for protecting models against adversarial attacks in a broad range of real-world applications.

Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

Model ReleasesDGX agent

arXiv:2607.27422v1 Announce Type: new Abstract: Good action rankings do not make a contrastive critic safe to maximize. These critics increasingly act as value-like objectives for best-of-K selection,

Google starts rolling out access to Gemini Spark for Google AI Pro subscribers to over 160 countries and adds a Chrome auto browse integration on desktop (Abner Li/9to5Google)

Model ReleasesDGX agent

Abner Li / 9to5Google: Google starts rolling out access to Gemini Spark for Google AI Pro subscribers to over 160 countries and adds a Chrome auto browse integration on desktop — Gemini Spark is getti

Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning

Model ReleasesDGX agent

arXiv:2607.27766v1 Announce Type: new Abstract: On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. Thi

Graph Neural Multilevel Preconditioners for Iterative Solvers

Model ReleasesDGX agent

arXiv:2607.28456v1 Announce Type: cross Abstract: Solving large, sparse linear systems is a core task in scientific computing, and efficient iterative solvers rely critically on effective and robust p

Graph Neural Network Force Fields for Spin Dynamics in Metallic Magnets

Model ReleasesDGX agent

arXiv:2607.28537v1 Announce Type: cross Abstract: Metallic magnets exhibit complex spin dynamics governed by electronically generated interactions. Predictive simulations of such dynamics typically re

GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

Model ReleasesDGX agent

arXiv:2607.26160v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather t

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference

Model ReleasesDGX agent

arXiv:2607.27694v1 Announce Type: cross Abstract: Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. Ho

Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA

Model ReleasesDGX agent

arXiv:2502.10497v2 Announce Type: replace Abstract: Recent advancements in Generative AI have significantly improved the efficiency and adaptability of natural language processing (NLP) systems, parti

HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks

Model ReleasesDGX agent

arXiv:2607.28301v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race d

Has anyone actually benchmarked where the 'big-model orchestrator + local-model worker' split breaks down?

Model ReleasesDGX agent

I keep seeing the 'use a big model via API as the architect, run local small/mid models as workers' pattern recommended for people with modest local hardware. I've been running it myself (orchestrator

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent…

Model ReleasesDGX agent

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent - 0.47 Codex - 0.51 OpenCode - 0.54 Kimi Code - 1.47 Claude C

I have trained a model to predict my blood sugar [P]

Model ReleasesDGX agent

It's an encoder-only transformer that consumes past(blood glucose + carbs + insulin) and future(carbs + insulin) and predicts future blood glucose for the next 2 hours. Announced meals and boluses/bas

I predict DeepSeek V4 Flash 0731's Artificial Analysis score to be 57 ± 1 point (Kimi K3 Level)

Model ReleasesDGX agent

Deepseek's new model V4 Flash 0731 is much better, I (Claude lol) did a bit of linear regression with a leave one out style verification to predict its AA Score, and that puts it at Kimi K3 level, whi

i see your moore's law and i raise you 20x

Model ReleasesDGX agent

i see your moore's law and i raise you 20x GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs 2.50/15; Luna now costs 0.20/1.20. In other words, roughly four months late

I switched my https://agent.datasette.io instance to Luna (it was previously on Gemini 3.1 Flash-Lite - Luna is cheaper now) - you can sign …

Model ReleasesDGX agent

Simon Willison switched his Datasette Agent instance from Gemini 3.1 Flash‑Lite to GPT‑5.6 “Luna” after a recent 80% price drop. He reports the new model is significantly faster and automatically gene

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

Model ReleasesDGX agent

arXiv:2607.26075v1 Announce Type: cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tun

If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, o…

Model ReleasesDGX agent

If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, optimization, rank) to see if we could close the gap between

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

Model ReleasesDGX agent

arXiv:2607.26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or pers

IFHierBench: Hierarchical Instruction Following for Large Language Models

Model ReleasesDGX agent

arXiv:2607.27912v1 Announce Type: cross Abstract: Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the ou

'Intelligence too cheap to meter' battle is on! Given that DeepSeek-V4-Flash-Preview is already great for agentic tasks, there is no doubt t…

Model ReleasesDGX agent

'Intelligence too cheap to meter' battle is on! Given that DeepSeek-V4-Flash-Preview is already great for agentic tasks, there is no doubt this new checkpoint must be an absolute beast. 20+ point jump

Introducing Qwen-Audio-3.0-ASR-Flash: More context-aware. Stronger domain-term recognition. 🚀Our latest ASR model upgrades: • Context consi…

Model ReleasesDGX agent

Introducing Qwen-Audio-3.0-ASR-Flash: More context-aware. Stronger domain-term recognition. 🚀Our latest ASR model upgrades: • Context consistency • Domain-term recognition • Custom hotwords • Speech p

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

Model ReleasesDGX agent

Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without

It’s been a busy couple of weeks! ICYMI, here’s the recap ⬇️ — Gemini Robotics 2 from @GoogleDeepmind brings whole-body intelligence to robo…

Model ReleasesDGX agent

It’s been a busy couple of weeks! ICYMI, here’s the recap ⬇️ — Gemini Robotics 2 from @GoogleDeepmind brings whole-body intelligence to robots — Gemini 3.5 Flash-Lite is our fastest, most cost-effecti

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Model ReleasesDGX agent

arXiv:2607.27670v1 Announce Type: new Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that creat

K-EXAONE 2.0 released

Model ReleasesDGX agent

https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B-FP8 https://huggingface.co/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B-NVFP4 https://huggingf

KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

Model ReleasesDGX agent

arXiv:2607.28608v1 Announce Type: new Abstract: Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pip

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

Model ReleasesDGX agent

arXiv:2607.27610v1 Announce Type: new Abstract: Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critical

KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation

Model ReleasesDGX agent

arXiv:2607.27231v1 Announce Type: cross Abstract: Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specia

Language Diversity: Evaluating Language Usage and AI Performance on African Languages in Digital Spaces

Model ReleasesDGX agent

arXiv:2512.01557v3 Announce Type: replace Abstract: This study examines the digital representation of African languages and the challenges this presents for current language detection tools. We evalua

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

Learning-Augmented and Randomized Algorithms for Line Aggregation with Delays

Model ReleasesDGX agent

arXiv:2607.27807v1 Announce Type: new Abstract: This paper studies learning-augmented and randomized online aggregation with delays on a line metric. We consider advice given as online suggested servi

Learning Color Grading, No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement

Model ReleasesDGX agent

arXiv:2607.27659v1 Announce Type: new Abstract: Personalized image enhancement should reflect individual aesthetic taste, yet learning such preferences commonly depends on private photos and ratings t

Learning features from Newton's algorithm: a way to accelerate nonlinear parametrized PDE solvers

Model ReleasesDGX agent

arXiv:2607.28036v1 Announce Type: new Abstract: It is well known that Newton's method converges faster when the initial guess is closer to a root of a system of nonlinear equations. In this paper, a t

Learning path to fully understand the Kimi K3 technical report?

Model ReleasesDGX agent

Hi everyone, Can anyone suggest a learning path to fully understand the technical report for Kimi K3? My background: • I've taken a graduate-level deep learning course. • I understand the Transformer

Learning to Trace Seiberg Dualities

Model ReleasesDGX agent

arXiv:2607.28628v1 Announce Type: cross Abstract: Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical systems. In practice, though, it

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models

Model ReleasesDGX agent

arXiv:2607.28449v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning tha

LLM-Guided Initialization for Accelerated Hybrid Quantum-Classical Medical Image Classification

Model ReleasesDGX agent

arXiv:2607.27262v1 Announce Type: cross Abstract: Variational quantum algorithms often encounter barren plateaus, where cost gradients decay rapidly with increasing circuit depth, undermining the trai

LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints

Model ReleasesDGX agent

arXiv:2410.06458v2 Announce Type: replace Abstract: Instruction following is a key capability for LLMs. However, recent studies have shown that LLMs often struggle with instructions containing multipl

LLM2Vec-Gen: Generative Embeddings from Large Language Models

Model ReleasesDGX agent

arXiv:2603.10913v3 Announce Type: replace Abstract: Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output

← Previous
1…4647484950…373
Next →