AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
3 Jul 2026

SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses

Model ReleasesDGX agent

arXiv:2607.02049v1 Announce Type: cross Abstract: Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-lingual abiliti

SPOT: Spatio-Temporal Obstacle-free Trajectory Planning for UAVs in Unknown Dynamic Environments

Model ReleasesDGX agent

arXiv:2602.01189v3 Announce Type: replace Abstract: We address the problem of reactive motion planning for quadrotors operating in unknown environments with dynamic obstacles. Our approach leverages a

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.09517v2 Announce Type: replace Abstract: Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do no

Steerability via constraints: a substrate for scalable oversight of coding agents

Model ReleasesDGX agent

arXiv:2607.02389v1 Announce Type: new Abstract: Coding agents are capable; human oversight is the bottleneck. Unconstrained agents introduce security risks, erode codebase scalability, and make human

Structured Gaussian Processes for Uncertainty-Aware Classification of High-Dimensional, Small-Sampled Omics Data

Model ReleasesDGX agent

arXiv:2607.02103v1 Announce Type: cross Abstract: Classifying heterogeneous omics data remains a fundamental challenge in computational biology, particularly in high-dimensional, small-sample settings

TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution

Model ReleasesDGX agent

arXiv:2607.02469v1 Announce Type: cross Abstract: Software tests and code evolve together: a code change should be followed by new or updated tests that record the new software behavior. Yet existing

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

Model ReleasesDGX agent

arXiv:2607.02407v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods ofte

The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits

Model ReleasesDGX agent

arXiv:2607.02201v1 Announce Type: cross Abstract: The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains fragmented

the scroll feature in the claude code CLI is really nice

Model ReleasesDGX agent

Jerry Liu highlights the scroll feature in the Claude Code CLI as a beneficial functionality. The post suggests this feature improves the user experience when working with the Claude Code command-line

The team at @vercel recently released the Eve agent framework, so we built a template that integrates LiteParse with it🦙 The template provi…

Model ReleasesDGX agent

The team at @vercel recently released the Eve agent framework, so we built a template that integrates LiteParse with it🦙 The template provides a set of read-only filesystem tools that let Eve resolve

The Wiola Architecture for Efficient Small Language Models

Model ReleasesDGX agent

arXiv:2607.01394v1 Announce Type: new Abstract: We present Wiola, a fully original Small Language Model (SLM) architecture built from first principles, sharing no structural lineage with any existing

This is actually useful. LangChain just released OpenWiki. It's an open-source agent that creates a wiki for your codebase, connects it to y…

Model ReleasesDGX agent

This is actually useful. LangChain just released OpenWiki. It's an open-source agent that creates a wiki for your codebase, connects it to your coding agent, and keeps it updated as your repo changes.

This is true… but maybe less important than the fact that people don’t try ambitious things with these systems. Many models are excellent as…

Model ReleasesDGX agent

This is true… but maybe less important than the fact that people don’t try ambitious things with these systems. Many models are excellent as a Google replacement, for homework “help,” etc. It is someo

Token Geometry

Model ReleasesDGX agent

arXiv:2607.01455v1 Announce Type: cross Abstract: Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface between them.

Towards a Phonology-Informed Evaluation of Multilingual TTS

Model ReleasesDGX agent

arXiv:2607.01965v1 Announce Type: new Abstract: Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words fro

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

Model ReleasesDGX agent

arXiv:2607.02043v1 Announce Type: cross Abstract: Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates a new asymm

Towards Robustness against Typographic Attack with Training-free Concept Localization

Model ReleasesDGX agent

arXiv:2607.02494v1 Announce Type: cross Abstract: Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Model

TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B

Model ReleasesDGX agent

arXiv:2607.01927v1 Announce Type: cross Abstract: This paper presents TUDUM (Turkce Dusunen Uretken Model), a project pipeline for adapting a Qwen-family 27B thinking model toward Turkish reasoning. T

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

Model ReleasesDGX agent

arXiv:2607.01345v1 Announce Type: cross Abstract: Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluations often re

UA-ChatDev: Uncertainty-Aware Multi-Agent Collaboration for Reliable Software Development

Model ReleasesDGX agent

arXiv:2607.02186v1 Announce Type: new Abstract: Software development is a complex task that demands cooperation among agents with diverse roles. Large language models (LLMs) have enabled autonomous mu

Understanding Agent-Based Patching of Compiler Missed Optimizations

Model ReleasesDGX agent

arXiv:2607.02370v1 Announce Type: cross Abstract: Compiler missed optimizations refer to cases in which compilers failed to optimize certain code. It takes many compiler developers' efforts to impleme

Unpopular opinion: While everyone is so hyped about Fable, GPT5.6 and other huge and expensive models, I think the real hero of the last few…

Model ReleasesDGX agent

Unpopular opinion: While everyone is so hyped about Fable, GPT5.6 and other huge and expensive models, I think the real hero of the last few months is *Qwen 27b*. Our ML/AI engineering teams are have

VisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval

Model ReleasesDGX agent

arXiv:2607.02371v1 Announce Type: cross Abstract: Over 285 million people worldwide live with a visual impairment, for whom everyday tasks such as avoiding obstacles, locating personal belongings, rec

VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer

Model ReleasesDGX agent

arXiv:2512.11891v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalizing across diverse robotic manipulation tasks. However, de

WARP: Weight-Space Analysis for Recovering Training Data Portfolios

Model ReleasesDGX agent

arXiv:2607.01686v1 Announce Type: new Abstract: Foundation models are routinely released to the public, yet the data recipes used to train them -- such as domain mixture weights that determine how dif

We analyzed GLM 5.2 vs Sonnet 5 for software engineering tasks using DeepSWE. GLM 5.2 gets you ~80% of Sonnet 5's capability at ~20% of the …

Model ReleasesDGX agent

We analyzed GLM 5.2 vs Sonnet 5 for software engineering tasks using DeepSWE. GLM 5.2 gets you ~80% of Sonnet 5's capability at ~20% of the price. More insights in the thread! Deepdive: Sonnet 5 and G

we distilled 2.3M Claude Fable 5 reasoning traces into Qwen3-4B - 100% self-consistency @ 512 samples - 0.00 bits output entropy - zero hall…

Model ReleasesDGX agent

we distilled 2.3M Claude Fable 5 reasoning traces into Qwen3-4B - 100% self-consistency @ 512 samples - 0.00 bits output entropy - zero hallucination variance turns out the student is not bounded by t

Yes they can move and dance and stuff You can assume they will talk and sing and more as good as anyone https://x.com/ubtechrobotics/status/…

Model ReleasesDGX agent

Yes they can move and dance and stuff You can assume they will talk and sing and more as good as anyone https://x.com/ubtechrobotics/status/2072651508710285419?s=46 UBTECH Launches UWORLD U1 — The Wor

2 Jul 2026

3 years ago I gave a talk at the first @aiDotEngineer conference on 'Advanced RAG' techniques in order to work around the limitations of nai…

Model ReleasesDGX agent

3 years ago I gave a talk at the first @aiDotEngineer conference on 'Advanced RAG' techniques in order to work around the limitations of naive RAG. It's insane how much the world has changed since the

A Lightweight Self-Supervised Learning Framework for Multivariate Time Series using Hierarchical-JEPA on ECG Data

Model ReleasesDGX agent

arXiv:2607.01145v1 Announce Type: new Abstract: Data analysis in the medical domain often encounters scenarios involving a limited target dataset and a large, unannotated dataset with a general distri

A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models

Model ReleasesDGX agent

arXiv:2607.00309v1 Announce Type: cross Abstract: We present a real-time musical interface that converts natural-language scene descriptions into evolving procedural soundscapes. A performer types a p

A Unified Benchmark for RCM-Constrained Visual Servoing: Modeling-Controller Interaction and Robustness Analysis in Laparoscopic Robots

Model ReleasesDGX agent

arXiv:2607.00030v1 Announce Type: new Abstract: In robot-assisted laparoscopic minimally invasive surgery (MIS), accurate enforcement of the remote center of motion (RCM) constraint is critical for sa

ActivityNarrated: An Open-Ended Narrative Paradigm for Wearable Human Activity Understanding

Model ReleasesDGX agent

arXiv:2604.00767v2 Announce Type: replace Abstract: Wearable human activity recognition (HAR) has made steady progress, yet much of this progress remains grounded in fixed-window, closed-set classific

AD-MPCC: Adaptive Differentiable Model Predictive Contouring Control for Autonomous Racing

Model ReleasesDGX agent

arXiv:2607.00141v1 Announce Type: new Abstract: This paper presents Adaptive Differentiable Model Predictive Contouring Control (AD-MPCC), a framework for autonomous racing that integrates differentia

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Model ReleasesDGX agent

arXiv:2607.01153v1 Announce Type: cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an in

AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware

Model ReleasesDGX agent

arXiv:2602.16249v2 Announce Type: replace Abstract: Self-supervised pretraining has transformed computer vision by enabling data-efficient fine-tuning, yet high-resolution pretraining typically requir

AGC-Bench: Measuring Artificial General Creativity

Model ReleasesDGX agent

arXiv:2607.01152v1 Announce Type: new Abstract: Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from gen

AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.00052v1 Announce Type: cross Abstract: GraphRAG is an extension of retrieval-augmented generation (RAG) that supports large language models (LLMs) by referring to graph-structured data as e

AGI Maze as a Benchmark Framework for World-Modeling Agents

Model ReleasesDGX agent

arXiv:2607.00627v1 Announce Type: new Abstract: Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context

AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation

Model ReleasesDGX agent

arXiv:2607.00062v1 Announce Type: cross Abstract: High pass rates on established programming benchmarks such as HumanEval and LiveCodeBench do not always show whether a model can reason about algorith

Amortized Maximum Inner Product Search with Learned Support Functions

Model ReleasesDGX agent

arXiv:2603.08001v3 Announce Type: replace Abstract: Maximum inner product search (MIPS) is a crucial subroutine in machine learning, requiring the identification of a vector taken within a database (t

An LLM-Based Framework for Intent-Driven Network Topology Design

Model ReleasesDGX agent

arXiv:2607.00292v1 Announce Type: cross Abstract: Designing deployable and resilient network topologies from natural language requirements remains a challenging problem in network automation. This wor

AnF-DiffPET: Anatomy- and Frequency-Guided Diffusion for PET/CT Denoising

Model ReleasesDGX agent

arXiv:2607.00509v1 Announce Type: new Abstract: Positron emission tomography (PET) provides essential functional information for disease assessment, however reducing injected activity or acquisition t

Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their…

Model ReleasesDGX agent

Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their scores is a trap. 'Overall, we establish that robust aggreg

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

Model ReleasesDGX agent

arXiv:2607.01211v1 Announce Type: cross Abstract: Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real reposi

Artifacts in Claude Code have been life changing. Excited to expand to Pro and Max!

Model ReleasesDGX agent

Artifacts in Claude Code have been life changing. Excited to expand to Pro and Max! Artifacts in Claude Code are now also available on Pro and Max plans. Ask for an artifact, Claude writes the code, p

ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis

Model ReleasesDGX agent

arXiv:2607.00041v1 Announce Type: cross Abstract: Multi-agent LLM systems can decompose software-engineering work into planning, generation, validation, and repair, but a narrower systems problem rema

Auditing Forgetting in Limited Memory Language Models

Model ReleasesDGX agent

arXiv:2607.00605v1 Announce Type: cross Abstract: Limited Memory Language Models (LMLMs) externalize factual knowledge to a database to enable deletion-based unlearning without retraining. Existing ev

AutoMem: Automated Learning of Memory as a Cognitive Skill

Model ReleasesDGX agent

arXiv:2607.01224v1 Announce Type: new Abstract: Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capacity known in cognitive science as m

// AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory management as a trai…

Model ReleasesDGX agent

// AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory management as a trainable skill instead of a fixed module. The model decides wha

Autonomous Scientific Discovery via Iterative Meta-Reflection

Model ReleasesDGX agent

arXiv:2607.01131v1 Announce Type: cross Abstract: Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation.

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

Model ReleasesDGX agent

arXiv:2607.00726v1 Announce Type: new Abstract: Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for

BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal

Model ReleasesDGX agent

arXiv:2607.00501v1 Announce Type: cross Abstract: We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference throughput on

Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering

Model ReleasesDGX agent

arXiv:2607.00972v1 Announce Type: new Abstract: Trustworthy deployment of Agentic Retrieval-Augmented Generation (RAG) systems requires mechanisms for estimating when multi-stage reasoning pipelines m

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

Model ReleasesDGX agent

arXiv:2607.00139v1 Announce Type: new Abstract: The cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, high-stakes domains. This is particularly acu

Beyond Activation Alignment:The Alignment-Diversity Tradeoff in Task-Aware LLM Quantization

Model ReleasesDGX agent

arXiv:2607.00908v1 Announce Type: new Abstract: Mixed-precision quantization (MPQ) has become a key technique for deploying large language models under stringent memory and compute constraints. We fir

Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents

Model ReleasesDGX agent

arXiv:2607.00895v1 Announce Type: new Abstract: Hallucination detection for retrieval-augmented generation (RAG) is usually evaluated on natural-language document evidence. However, grounded generatio

big week at langchain, with a lot of launches: 1/ OpenWiki - auto generate a wiki of a github repo 2/ two different voice agent tutorials 3/…

Model ReleasesDGX agent

big week at langchain, with a lot of launches: 1/ OpenWiki - auto generate a wiki of a github repo 2/ two different voice agent tutorials 3/ Harbor integration and tutorial for long running, stateful

Bridgewater just published numbers that should make every frontier lab nervous. The world's largest hedge fund tested Gemini, Claude, and GP…

Model ReleasesDGX agent

Bridgewater just published numbers that should make every frontier lab nervous. The world's largest hedge fund tested Gemini, Claude, and GPT on six document filtering tasks its investors do every day

But yikes does Fable write text that sounds like a parody of a Claude model on overdrive.

Model ReleasesDGX agent

This post critiques Fable AI's text generation style, suggesting it produces overly verbose or exaggerated outputs that parody Claude's characteristic writing patterns taken to an extreme. The comment

← Previous
1…108109110111112…377
Next →