AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,603 results
26 May 2026

UtilityMax Prompting: A Formal Framework for Multi-Objective Large Language Model Tasks

Model ReleasesDGX agent

arXiv:2603.11583v4 Announce Type: replace-cross Abstract: The success of a Large Language Model (LLM) task depends heavily on its prompt. Most use-cases specify prompts using natural language, which i

UWM-JEPA: Predictive World Models That Imagine in Belief Space

Model ReleasesDGX agent

arXiv:2605.25313v1 Announce Type: cross Abstract: World models for partially observed environments must imagine multiple compatible hidden futures and steer between them under counterfactual actions.

v0.13.0 of Exo just dropped, and it's one I've been looking forward to for a while. tl;dr swapping out anthropic for @ollama cloud dropped m…

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

v0.13.0 of Exo just dropped, and it's one I've been looking forward to for a while. tl;dr swapping out anthropic for @ollama cloud dropped my usage costs from 30/day to just 20 per MONTH with no notic

v0.30.0-rc26: Merge remote-tracking branch 'upstream/main' into llama-runner-phase-0

Model ReleasesDGX agent

This release candidate merges updates from the upstream main branch into the llama-runner-phase-0 branch, likely incorporating recent improvements and bug fixes into the development version. Version 0

V3H: View Variation and View Heredity for Incomplete Multi-view Clustering

Model ReleasesDGX agent

arXiv:2011.11194v4 Announce Type: replace Abstract: Real data often appear in the form of multiple incomplete views. Incomplete multi-view clustering is an effective method to integrate these incomple

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation

Model ReleasesDGX agent

arXiv:2605.24675v1 Announce Type: cross Abstract: Translating text embedded in Web images is crucial for improving content accessibility and cross-lingual information retrieval, particularly within so

vAttention: Verified Sparse Attention

Model ReleasesDGX agent

arXiv:2510.05688v2 Announce Type: replace-cross Abstract: State-of-the-art sparse attention methods for reducing decoding latency fall into two main categories: approximate top-k (and its extension, t

VeriTrace: Evolving Mental Models for Deep Research Agents

Model ReleasesDGX agent

arXiv:2605.26081v1 Announce Type: new Abstract: Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representatio

ViroBench: Benchmarking Nucleotide Foundation Models on Viral Genomics Tasks

Model ReleasesDGX agent

arXiv:2605.25388v1 Announce Type: new Abstract: Nucleotide sequences constitute the fundamental genetic basis of biological systems, rendering viral genomic analysis critical for biomedical advancemen

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.25820v1 Announce Type: new Abstract: Diffusion-based multimodal large language models (dMLLMs) decode by iteratively predicting tokens at multiple masked positions in parallel. This turns e

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

Model ReleasesDGX agent

arXiv:2509.25339v3 Announce Type: replace-cross Abstract: Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answerin

We have, as far as I can tell, no good tests of the productivity impact of the autonomous coding tools that appeared starting in December 20…

Model ReleasesDGX agent

We have, as far as I can tell, no good tests of the productivity impact of the autonomous coding tools that appeared starting in December 2025. Every paper out there is from prior to the Claude Code/C

We’ve shipped a security-guidance plugin for Claude Code that helps identify and fix vulnerabilities as you’re writing code. Available for a…

Model ReleasesDGX agent

We’ve shipped a security-guidance plugin for Claude Code that helps identify and fix vulnerabilities as you’re writing code. Available for all Claude Code users. Install from the plugin marketplace (/

What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA

Model ReleasesDGX agent

arXiv:2605.25988v1 Announce Type: new Abstract: Medical RAG needs evidence-grounded claims, so plugging a claim-level NLI checker into retrieval-augmented RL is intuitive. extbf{We find that the check

When Can We Trust Early Warnings? Leakage-Excluded Early Outcome Prediction from LMS Interaction Logs

Model ReleasesDGX agent

arXiv:2605.25794v1 Announce Type: new Abstract: Early-warning models built from Learning Management System (LMS) logs aim to predict end-of-course outcomes early enough to enable timely learner suppor

When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure

Model ReleasesDGX agent

arXiv:2605.23932v1 Announce Type: new Abstract: Despite strong medical benchmark accuracy, LLMs can exhibit severe multi-turn sycophancy in clinical dialogue, abandoning initial correct diagnosis unde

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

Model ReleasesDGX agent

arXiv:2605.25981v1 Announce Type: new Abstract: We document an empirical phenomenon in chain-of-thought and ReAct agents driven by ten large language models from seven architecture families: meaning-b

When Mean CE Fails: Median CE Can Better Track Language Model Quality

Model ReleasesDGX agent

arXiv:2605.24667v1 Announce Type: new Abstract: Mean cross-entropy is the standard validation metric for language models, but it can fail to track model quality during training. We examine this in two

When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation

Model ReleasesDGX agent

arXiv:2605.24902v1 Announce Type: cross Abstract: Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical do

When Search Becomes Memory: Turning Robot Design Trials into Transferable Skills

Model ReleasesDGX agent

arXiv:2605.25832v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as proposal generators for evolutionary robot design, yet most loops remain memoryless: simulator r

When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents

Model ReleasesDGX agent

arXiv:2605.24069v1 Announce Type: cross Abstract: The rise of tool-using Large Language Model (LLM) agents, standardized by protocols like the Model Context Protocol (MCP), has unlocked unprecedented

WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems

Model ReleasesDGX agent

arXiv:2605.24579v1 Announce Type: new Abstract: Long-context memory systems often fail under fixed budgets, but end-to-end evaluation does not reveal whether evidence was discarded during compression

Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring

Model ReleasesDGX agent

arXiv:2605.24737v1 Announce Type: cross Abstract: Current approaches to AI compliance treat conformity as a binary, audit-time verdict rather than a continuous, measurable property of production syste

WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification

Model ReleasesDGX agent

arXiv:2605.26070v1 Announce Type: new Abstract: Annotating speaker attributes from text is inherently ambiguous, particularly in multilingual settings where demographic and social cues are implicit an

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

Model ReleasesDGX agent

arXiv:2605.25256v1 Announce Type: new Abstract: Aligning AI systems with organizational decision-making is typically framed as a single-target problem: make the model behave like the organization. We

WideDepth: Millimeter-Accurate Benchmark for Fisheye Depth Estimation

Model ReleasesDGX agent

arXiv:2605.24074v1 Announce Type: cross Abstract: Fisheye cameras are increasingly adopted in robotics for near-field manipulation, navigation, and immersive perception, yet indoor depth benchmarks wi

WLNO: Wavelet-Laplace Neural Operator for Solving Partial Differential Equations

Model ReleasesDGX agent

arXiv:2605.24658v1 Announce Type: new Abstract: This work introduces the Wavelet-Laplace Neural Operator (WLNO), a novel neural operator that fuses Haar wavelet multi-scale spatial decomposition with

World-State Transformations for Neuro-symbolic Interactive Storytelling

Model ReleasesDGX agent

arXiv:2605.24719v1 Announce Type: cross Abstract: Large Language Models (LLMs) have changed the possibilities of Interactive Storytelling systems that process free-text user input. However, as more of

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point

Model ReleasesDGX agent

arXiv:2502.08047v5 Announce Type: replace Abstract: Recent progress in GUI agents has substantially improved visual grounding, yet robust planning remains challenging, particularly when the environmen

25 May 2026

A Comparative Evaluation of Structural Topic Models and BERTopic for Short, Open-Ended Survey Responses

Model ReleasesDGX agent

arXiv:2605.23093v1 Announce Type: new Abstract: Topic modeling in applied psychology increasingly spans two methodological traditions: probabilistic bag-of-words models and newer embedding-based appro

A European Multi-Center Breast Cancer MRI Dataset

Model ReleasesDGX agent

arXiv:2506.00474v3 Announce Type: replace-cross Abstract: Early detection of breast cancer is critical for improving patient outcomes. While mammography remains the primary screening modality, magneti

A look at DeepSeek's model optimization to reduce HBM use, potentially enabling domestic memory, ASIC, and CPU makers to create a Chinese AI hardware ecosystem (@bookwormengr)

Model ReleasesDGX agent

@bookwormengr: A look at DeepSeek's model optimization to reduce HBM use, potentially enabling domestic memory, ASIC, and CPU makers to create a Chinese AI hardware ecosystem — Have you ever wondered,

A measurement substrate for agentic Kubernetes operations: Methodology and a case study in retrieval-compounding falsification

Model ReleasesDGX agent

arXiv:2605.23058v1 Announce Type: cross Abstract: Empirical claims about autonomous Kubernetes operations agents are largely unfalsifiable. Published work reports observational results without control

A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek Parliamentary Text

Model ReleasesDGX agent

arXiv:2605.22978v1 Announce Type: new Abstract: Katharevousa Greek remains poorly served by contemporary NLP pipelines despite its importance for legal, administrative, and parliamentary archives. We

A Systematic Evaluation of Co-folding Model Representations for Small-Molecule Learning

Model ReleasesDGX agent

arXiv:2602.13249v2 Announce Type: replace-cross Abstract: Small-molecule foundation models are typically pretrained on standalone molecular data, unlike vision and language models that often benefit f

Agentic Proving for Program Verification

Model ReleasesDGX agent

arXiv:2605.23772v1 Announce Type: new Abstract: Agentic systems have recently emerged as state-of-the-art approaches for automated theorem proving in formal mathematics. To assess how far these capabi

Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.22896v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by leveraging pre-trained vision-language representa

AI Evaluation Should Require Standardized Item-Level Data Releases

Model ReleasesDGX agent

arXiv:2604.03244v2 Announce Type: replace Abstract: This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluatio

An open source model has returned to #1 on the 3D Design leaderboard by Design Arena. Kimi K2.6 has reached the top of the leaderboard for 3…

Model ReleasesDGX agent

An open source model has returned to #1 on the 3D Design leaderboard by Design Arena. Kimi K2.6 has reached the top of the leaderboard for 3D Design, ahead of models 10X more expensive like Opus 4.7 b

Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking

Model ReleasesDGX agent

arXiv:2605.23733v1 Announce Type: cross Abstract: Whole-body tracking (WBT) models have become a key foundation for humanoid robots, enabling them to imitate diverse motions with high fidelity. Traini

Anytime Training with Schedule-Free Spectral Optimization

Model ReleasesDGX agent

arXiv:2605.23061v1 Announce Type: cross Abstract: Standard neural network training relies on learning-rate schedules tied to a fixed horizon, leading to strong path dependence and costly re-tuning as

Approaching I/O-optimality for Approximate Attention

Model ReleasesDGX agent

arXiv:2605.23751v1 Announce Type: new Abstract: We revisit the I/O complexity of attention in large language models. Given query-key-value matrices Q,K,VinR^{nimes d}, and a machine with fast memory s

AraHopeCorpus: Annotation Guidelines and Dataset for Hope Speech in Arabic Social Media Crisis Discourse

Model ReleasesDGX agent

arXiv:2605.23325v1 Announce Type: new Abstract: Social media has become a crucial arena for shaping public narratives during armed conflicts, providing space for both harmful and constructive communic

Archimedean Copula Inference via Taylor-Mode AD

Model ReleasesDGX agent

arXiv:2605.23134v1 Announce Type: new Abstract: No existing nested Archimedean copula tool handles all three of (a) arbitrary per-variable (right-)censoring in survival analysis, (b) arbitrary nesting

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

Model ReleasesDGX agent

arXiv:2605.23243v1 Announce Type: cross Abstract: We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM

As X, Do Y: How Persona and Task Combine in Instruction-Tuned LLMs

Model ReleasesDGX agent

arXiv:2605.23147v1 Announce Type: cross Abstract: Role prompts of the form As X, do Y admit a clean linear decomposition at one specific site in the residual stream: the prompt-to-answer transition --

Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering

Model ReleasesDGX agent

arXiv:2605.23497v1 Announce Type: new Abstract: Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds

Atom-level Protein Representation Learning Improves Protein Structure Prediction

Model ReleasesDGX agent

arXiv:2605.22133v2 Announce Type: replace-cross Abstract: Recent advances in generative modeling show that pretrained representations can improve generation as conditioning features or alignment targe

Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics

Model ReleasesDGX agent

arXiv:2510.12787v4 Announce Type: replace Abstract: We present Ax-Prover, a multi-agent system for automated theorem proving in Lean that can solve problems across diverse scientific domains and opera

Benchmarking and Enhancing VLM for Compressed Image Understanding

Model ReleasesDGX agent

arXiv:2512.20901v2 Announce Type: replace Abstract: With the rapid development of Vision-Language Models (VLMs) and the growing demand for their applications, efficient compression of the image inputs

Benchmarking Google Embeddings 2 against Open-Source Models for Multilingual Dense Retrieval and RAG Systems

Model ReleasesDGX agent

arXiv:2605.23618v1 Announce Type: new Abstract: We benchmark Google Embeddings (GE2), a Vertex-AI-hosted bi-encoder with 2,048-token context and explicit task-type conditioning, against five open-sour

BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems

Model ReleasesDGX agent

arXiv:2605.22866v1 Announce Type: new Abstract: Compound AI systems route tasks through hierarchies of specialised components. Attribution is dominated by Shapley-based methods (SHAP), which decompose

Brain-LLM Alignment Tracks Training Data, Not Typology

Model ReleasesDGX agent

arXiv:2605.23032v1 Announce Type: cross Abstract: Brain-LLM alignment is well established in English, yet the brain's language network is neuroanatomically universal across languages. Does alignment a

BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models

Model ReleasesDGX agent

arXiv:2602.18788v3 Announce Type: replace Abstract: We introduce BURMESE-SAN, the first holistic benchmark that systematically evaluates large language models (LLMs) for Burmese across three core NLP

Can AI Guess What You Know? Performance Comparison of Large Language Models for Human Domain Knowledge Estimation From Communication Logs

Model ReleasesDGX agent

arXiv:2605.22971v1 Announce Type: new Abstract: Employees often struggle to identify ``who knows what,'' leading to organizational productivity losses. We investigate whether Large Language Models (LL

Canadians be like come to our Toronto tech week party

Model ReleasesDGX agent

Toronto Tech Week is an annual event in Toronto, Canada that brings together technology professionals, entrepreneurs, and innovators for networking, learning, and celebration of the local tech ecosyst

CARE: Class-Adaptive Expert Consensus for Reliable Learning with Long-Tailed Noisy Labels

Model ReleasesDGX agent

arXiv:2605.23254v1 Announce Type: new Abstract: Learning from real-world data is frequently hindered by the compound challenge of long-tailed class distributions and noisy annotations. Existing method

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

Model ReleasesDGX agent

arXiv:2605.23216v1 Announce Type: new Abstract: Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception t

ChartFI: Benchmarking Faithfulness and Insightfulness of Chart Descriptions from Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.23694v1 Announce Type: new Abstract: Chart descriptions are essential for accessibility, cross-modal retrieval, and assisting readers in extracting insights from complex visualizations. As

CHRONOS: Temporally-Aware Multi-Agent Coordination for Evolving Data Marketplaces

Model ReleasesDGX agent

arXiv:2605.23887v1 Announce Type: cross Abstract: Temporal knowledge-graph data marketplaces face three coupled failures in static designs: stale hybrid index shortcuts reduce recall as edges evolve,

← Previous
1…215216217218219…377
Next →