AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,620 results
3 Jun 2026

TeX-1500: A Paired Real-World LWIR Hyperspectral Dataset and Benchmark for Temperature-Emissivity-Texture Decomposition

Model ReleasesDGX agent

arXiv:2606.03806v1 Announce Type: new Abstract: Temperature-emissivity-texture (TeX) decomposition seeks to recover object heat state, material spectral response, and visible-like geometric texture fr

The DeepSpeak-Agentic Dataset

Model ReleasesDGX agent

arXiv:2606.03686v1 Announce Type: new Abstract: We present DeepSpeak-Agentic, a dataset of videos comprising over 37 hours of semi-structured conversations between a human and an embodied AI agent. We

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

Model ReleasesDGX agent

arXiv:2606.03043v1 Announce Type: new Abstract: LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions: A Study Protocol

Model ReleasesDGX agent

arXiv:2606.03907v1 Announce Type: cross Abstract: Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from s

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

Model ReleasesDGX agent

arXiv:2606.03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for de

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

Model ReleasesDGX agent

arXiv:2606.02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence. We derive a two-paramete

The Word and the Way: Strategies for Domain-Specific BERT Pre-Training in German Medical NLP

Model ReleasesDGX agent

arXiv:2606.03250v1 Announce Type: new Abstract: Digital healthcare generates vast amounts of clinical text that can support AI-assisted applications, yet German biomedical language models remain limit

This is really intellectually dishonest. I have not been arguing that LLM token prices are increasing (though the all you can eat buffet is …

Model ReleasesDGX agent

This is really intellectually dishonest. I have not been arguing that LLM token prices are increasing (though the all you can eat buffet is over), I have been arguing the *opposite*, viz that they wil

This story was so implausible that the only way it even (kind of) made sense if it is some sort of internal accounting placeholder at a clou…

Model ReleasesDGX agent

This story was so implausible that the only way it even (kind of) made sense if it is some sort of internal accounting placeholder at a cloud provider using their own compute. And even then it seems u

thought I was supposed to be the shit-posting account here

Model ReleasesDGX agent

thought I was supposed to be the shit-posting account here shower thought If: 1. AI is smarter than humans at law, therapy, etc. 2. Humans still like talking to other humans. Then: Humans are just an

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

Model ReleasesDGX agent

arXiv:2606.03503v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (Co

Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech mod…

Model ReleasesDGX agent

Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a

Topology-Aware Gaussian Graph Repair for Robust Graph Neural Networks

Model ReleasesDGX agent

arXiv:2606.03462v1 Announce Type: new Abstract: Graph neural networks have achieved strong performance on graph-structured data, but their effectiveness depends heavily on the quality of the observed

Towards Characterizing Scientific Image Utility and Upgradability

Model ReleasesDGX agent

arXiv:2606.03401v1 Announce Type: new Abstract: Scientific images function as critical evidence in research communication, yet their integrity faces unprecedented threats from AI-generated content tha

Towards Fair Graph Prompting: A Dual-Prompt Mechanism for Mitigating Attribute and Structural Bias

Model ReleasesDGX agent

arXiv:2510.23469v2 Announce Type: replace Abstract: Self-supervised pre-training on unlabeled graph data has become a common paradigm for Graph Neural Networks (GNNs). However, an objective gap often

Trading Human Curation for Synthetic Augmentation in RLVR

Model ReleasesDGX agent

arXiv:2606.03800v1 Announce Type: cross Abstract: The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic language models

TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment

Model ReleasesDGX agent

arXiv:2606.03036v1 Announce Type: new Abstract: LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-w

TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

Model ReleasesDGX agent

arXiv:2606.03629v1 Announce Type: new Abstract: Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently,

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics

Model ReleasesDGX agent

arXiv:2606.03626v1 Announce Type: cross Abstract: Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focu

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs

Model ReleasesDGX agent

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs I wrote the other day about Uber blowing its 2026 AI budget in four months, and how that wasn't particularly surprising given they would ha

Using open models and inference clouds (which serve open models) is a leading indicator of what is to come. The advantage of open weights is…

Model ReleasesDGX agent

Using open models and inference clouds (which serve open models) is a leading indicator of what is to come. The advantage of open weights is that you can train, serve, and continually improve your own

v0.30.4: llama-server: fix gemma4 patch wiring (#16477)

Model ReleasesDGX agent

Ollama v0.30.4 is a patch release addressing a bug in the llama-server component related to incorrect parameter wiring in the Gemma 4 model implementation. This fix ensures Gemma 4 models operate corr

v0.30.4-rc0: Kill llama-server during Windows cleanup (#16458)

Model ReleasesDGX agent

This release candidate fixes a Windows-specific issue where the llama-server process wasn't being properly terminated during cleanup operations. The fix addresses GitHub issue #16458 and improves the

VidMsg: A Benchmark for Implicit Message Inference in Short Videos

Model ReleasesDGX agent

arXiv:2606.03635v1 Announce Type: cross Abstract: Understanding short online videos involves more than identifying visible objects and actions; video makers often include an underlying message or purp

VistaHop: Benchmarking Multi-hop Visual Reasoning for Visual DeepSearch

Model ReleasesDGX agent

arXiv:2606.03273v1 Announce Type: cross Abstract: Visual DeepSearch requires multimodal large reasoning model (MLRM) agents to answer complex visual queries by repeatedly inspecting image regions, gro

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2512.22539v2 Announce Type: replace-cross Abstract: While Vision-Language-Action models (VLAs) are rapidly advancing towards generalist robot policies, it remains difficult to quantitatively und

VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

Model ReleasesDGX agent

arXiv:2606.03954v1 Announce Type: new Abstract: As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible conse

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

Model ReleasesDGX agent

arXiv:2603.04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- se

WaterSIC: Information-Theoretically (Near) Optimal Linear Layer Quantization

Model ReleasesDGX agent

arXiv:2603.04956v2 Announce Type: replace Abstract: This paper considers the problem of converting a given dense linear layer to low precision. The tradeoff between compressed length and output discre

Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning

Model ReleasesDGX agent

arXiv:2509.19305v2 Announce Type: replace-cross Abstract: Diffusion probability models have shown significant promise in offline reinforcement learning by directly modeling trajectory sequences. Howev

We partnered with @FireworksAI_HQ to train open-source models for legal. Here's what we found: 1) Hybrid legal agents can beat frontier mode…

Model ReleasesDGX agent

We partnered with @FireworksAI_HQ to train open-source models for legal. Here's what we found: 1) Hybrid legal agents can beat frontier models on quality and cost by routing selectively to a frontier

We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-…

Model ReleasesDGX agent

We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-5.5’s agentic coding and tool use together with stronger int

What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

Model ReleasesDGX agent

arXiv:2606.02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceed

What Makes Interaction Trajectories Effective for Training Terminal Agents?

Model ReleasesDGX agent

arXiv:2606.03461v1 Announce Type: new Abstract: Stronger code agents are commonly assumed to be superior teachers for post-training, yet this assumption remains poorly disentangled from task difficult

When Model Merging Breaks Routing: Training-Free Calibration for MoE

Model ReleasesDGX agent

arXiv:2606.03391v1 Announce Type: cross Abstract: Model merging has emerged as a cost-effective approach for consolidating the capabilities of multiple LLMs without retraining. However, existing mergi

When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation

Model ReleasesDGX agent

arXiv:2606.03532v1 Announce Type: cross Abstract: Self on-policy distillation trains a student policy against a teacher derived from its own parameter history, yet the teacher's update schedule -- whi

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?

Model ReleasesDGX agent

arXiv:2606.03837v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it a

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

Model ReleasesDGX agent

arXiv:2606.02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-regis

Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

Model ReleasesDGX agent

arXiv:2602.08873v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now used for academic expert recommendation. Existing audits typically evaluate such recommendations in isola

Will Accurate Fields Mislead Photonic Design? FromGlobal Accuracy to Port Readout

Model ReleasesDGX agent

arXiv:2606.03038v1 Announce Type: new Abstract: Neural field surrogates can accelerate photonic design loops, but a surrogate that looks accurate in global field error can still mis-rank candidate dev

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2503.07265v4 Announce Type: replace-cross Abstract: Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content. However, existing research and evalua

Wordle 1,809 5/6 🟨⬛🟨⬛⬛ ⬛🟨⬛⬛⬛ ⬛🟩🟨🟨⬛ 🟩🟩🟩🟩⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a completed Wordle puzzle (puzzle #1,809) solved in five attempts, showing the progression of letter guesses through color-coded feedback (yellow for correct letters in wrong posit

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

Model ReleasesDGX agent

arXiv:2606.02908v1 Announce Type: cross Abstract: Multi-turn user-facing agents must infer user intent from incomplete requests, collect missing information through dialogue and tools, and execute val

2 Jun 2026

3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum

Model ReleasesDGX agent

arXiv:2606.00489v1 Announce Type: new Abstract: Placenta Accreta Spectrum (PAS) is a rare but highly dangerous obstetric disease. Early and accurate PAS diagnosis is critical for maternal health. Trad

3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

Model ReleasesDGX agent

arXiv:2606.01057v1 Announce Type: cross Abstract: Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neur

3rd Place at CVPR 2026 CASTLE Challenge: Agentic Multi-View Long-Context Video Understanding via Hierarchical Knowledge Graph Retrieval

Model ReleasesDGX agent

arXiv:2606.01933v1 Announce Type: new Abstract: This paper presents our winning methodology for the CASTLE 2026 Challenge at the CVPR 2026 EgoVis Workshop, where our team secured third place globally.

A Closer Look at In-Distribution vs. Out-of-Distribution Accuracy for Open-Set Test-time Adaptation

Model ReleasesDGX agent

arXiv:2606.01973v1 Announce Type: cross Abstract: Open-set test-time adaptation (TTA) updates models on new data in the presence of input shifts and unknown output classes. While recent methods have m

A Comparative Analysis of Machine Learning Algorithms for Multi-Task Prediction of the Parameters of the Pectin Hydrolysis--Extraction Process

Model ReleasesDGX agent

arXiv:2606.00821v1 Announce Type: new Abstract: This study addresses the challenge of controlling a complex, multi-parameter technological process -- pectin hydrolysis--extraction -- using machine lea

A Finite-Calibration Regime Map for LLM Judge Panels

Model ReleasesDGX agent

arXiv:2606.01034v1 Announce Type: new Abstract: We study when LLM judge panels should be calibrated with low-dimensional stackers versus joint output tables under finite human-label budgets. Low-dimen

A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL

Model ReleasesDGX agent

arXiv:2606.02398v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves large language models (LLMs) on individual domains such as mathematical reasoning, code generation,

A Machine-to-Machine Knowledge-Guided LLM Agent for Generalizable Radiotherapy Treatment Planning

Model ReleasesDGX agent

arXiv:2606.00922v1 Announce Type: cross Abstract: In this work, we propose a prototype machine-to-machine (M2M) knowledge-guided Large Language Model (LLM) framework for automated radiotherapy treatme

A Methodological Framework for Explicit Control of the Speed-Accuracy Trade-off in Brain-Computer Interfaces

Model ReleasesDGX agent

arXiv:2606.00106v1 Announce Type: cross Abstract: Brain-computer interfaces (BCIs) are limited by low signal-to-noise ratio in modalities such as electroencephalography, which requires multiple trials

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

Model ReleasesDGX agent

arXiv:2606.00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or

A Novel Data Augmentation Strategy for Robust Deep Learning Classification of Biomedical Time-Series Data: Application to ECG and EEG Analysis

Model ReleasesDGX agent

arXiv:2507.12645v1 Announce Type: cross Abstract: The increasing need for accurate and unified analysis of diverse biological signals, such as ECG and EEG, is paramount for comprehensive patient asses

A Per-Component Diagnostic Protocol for Neural HJB-PIDE Solvers under Control-Dependent Levy Jumps

Model ReleasesDGX agent

arXiv:2606.01122v1 Announce Type: new Abstract: We propose a five-step diagnostic protocol for residual-trained neural HJB-PIDE solvers with control-dependent Levy jumps, targeting a general failure m

A Structured Benchmark for Text-Guided Anomaly Detection: When Language Stops Conditioning the Decision

Model ReleasesDGX agent

arXiv:2606.01992v1 Announce Type: cross Abstract: Industrial anomaly detection has historically been a unimodal task. Recent multimodal vision-language models have produced systems that admit textual

A Systematic Benchmark of Intraoperative Ultrasound-to-MR Synthesis for Brain Tumour Surgery

Model ReleasesDGX agent

arXiv:2606.00630v1 Announce Type: new Abstract: Intraoperative ultrasound (ioUS) is a versatile, cost-effective modality in brain tumour surgery, but its interpretation is difficult: acquisition plane

AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research

Model ReleasesDGX agent

arXiv:2507.08038v3 Announce Type: replace-cross Abstract: Language model agents are increasingly used to automate scientific research, yet evaluating their scientific contributions remains a challenge

Absorbing Complexity: An Interaction-Native Knowledge Harness for Financial LLM Agents

Model ReleasesDGX agent

arXiv:2606.01886v1 Announce Type: new Abstract: Financial AI agents often fail for a simple reason: they make users carry the complexity. A user must repeatedly restate goals, risk preferences, portfo

Accelerating data lakes: Optimizing Apache Iceberg and Spark with gcs-analytics-core

Model ReleasesDGX agent

Many data engineers spend significant time managing compatibility and getting best performance across multiple analytics engines. To help solve this pain point, we are excited to announce gcs-analytic

← Previous
1…177178179180181…377
Next →