AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,611 results
27 May 2026

Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes

Model ReleasesDGX agent

arXiv:2601.07737v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in mainstream visual understanding tasks, but their ability

Self-Ensembling Vision-Language Models for Chart Data Extraction

Model ReleasesDGX agent

arXiv:2605.27298v1 Announce Type: new Abstract: Charts effectively convey quantitative information, but the underlying data are often locked in image form, hindering reuse and analysis. Manually digit

Self-Verified Distillation: Your Language Model Is Secretly Its Own Synthetic Data Pipeline

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.26132v1 Announce Type: new Abstract: Can post-trained large language models (LLMs) further improve themselves using only unlabeled prompts, without external teachers or feedback from tools?

Sentinel: Embodied Cooperative Spatial Reasoning and Planning

Model ReleasesDGX agent

arXiv:2605.26239v1 Announce Type: new Abstract: In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmen

Separating Semantic Competition from Context Length in RAG Reading

Model ReleasesDGX agent

arXiv:2605.27294v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems can respond incorrectly even when the correct passage was retrieved. The model must still read the retrieve

Shedding Light on Dark Matter at the LHC with Machine Learning

Model ReleasesDGX agent

arXiv:2509.15121v2 Announce Type: replace-cross Abstract: We investigate a WIMP dark matter (DM) candidate in the form of a singlino-dominated lightest supersymmetric particle (LSP) within the Z_3-sym

Shopping Companion: A Memory-Augmented LLM Agent for Real-World E-Commerce Tasks

Model ReleasesDGX agent

arXiv:2603.14864v2 Announce Type: replace Abstract: In e-commerce, LLM agents show promise for shopping tasks such as recommendations, budget management, and bundle deals, where accurately capturing u

SilIF: Silhouette-Augmented Isolation Forest for Unsupervised Transaction Fraud Detection

Model ReleasesDGX agent

arXiv:2605.26135v1 Announce Type: new Abstract: Unsupervised anomaly detection is widely used in transaction fraud detection where labels are scarce. Isolation Forest (IF) is among the most popular cl

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning

Model ReleasesDGX agent

arXiv:2603.28730v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown impressive capabilities across diverse tasks, motivating efforts to leverage these models to supervis

Someone just asked me “What I don’t get is why you see it as some kind of failure that AI labs are now using harnesses and neurosymbolic too…

Model ReleasesDGX agent

Someone just asked me “What I don’t get is why you see it as some kind of failure that AI labs are now using harnesses and neurosymbolic tools…?” I don’t! As I argued in my newsletter on Claude Code (

SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens

Model ReleasesDGX agent

arXiv:2508.05305v2 Announce Type: replace Abstract: The recently proposed Large Concept Model (LCM) generates text by predicting a sequence of sentence-level embeddings and training with either mean-s

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

Model ReleasesDGX agent

arXiv:2510.09606v2 Announce Type: replace Abstract: With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still strug

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

Model ReleasesDGX agent

arXiv:2605.27367v1 Announce Type: new Abstract: While spatial foundation models have demonstrated impressive performance on standard datasets, a critical question remains: are they truly all-round pla

Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context Prefilling

Model ReleasesDGX agent

arXiv:2604.18103v2 Announce Type: replace Abstract: Prefilling computational costs pose a significant bottleneck for Large Language Models (LLMs) and Large Multimodal Models (LMMs) in long-context set

SteelDS: A High-Resolution Video Dataset of E40 Steel Scrap for Object Detection and Instance Segmentation

Model ReleasesDGX agent

arXiv:2605.26682v1 Announce Type: cross Abstract: This dataset provides high-resolution, annotated video sequences of shredded E40-grade steel and copper scrap on a conveyor belt. Captured in a contro

Stochastic global optimization of continuous functions via random walks on Grassmannians

Model ReleasesDGX agent

arXiv:2605.14151v1 Announce Type: cross Abstract: We introduce a stochastic global optimization method based on random walks on Grassmannian manifolds. To minimize a continuous objective ell:R^drighta

Strategies for Guiding LLMs to Use Software Design Patterns: A Case of Singleton

Model ReleasesDGX agent

arXiv:2605.26898v1 Announce Type: cross Abstract: Large Language Models (LLMs) can generate functional source code from natural-language prompts, but often fail to consistently follow higher-level arc

Stronger models do not always need lighter harnesses. Everyone believes more structured harnesses universally improve reliability, and that …

Model ReleasesDGX agent

Stronger models do not always need lighter harnesses. Everyone believes more structured harnesses universally improve reliability, and that higher-capability models need proportionally less structural

Structured Relational Reasoning for Group Activity Assessment

Model ReleasesDGX agent

arXiv:2508.07996v2 Announce Type: replace Abstract: Group Activity Detection (GAD) involves recognizing social groups and their collective behaviors in videos. Vision Foundation Models (VFMs), like DI

SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution

Model ReleasesDGX agent

arXiv:2603.01327v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong performance on self-contained programming tasks. However, they still struggle with repository-leve

TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews

Model ReleasesDGX agent

arXiv:2605.26911v1 Announce Type: new Abstract: LLM-generated peer reviews are increasingly common at major venues, yet their deficiencies are hard to detect because they are uniformly fluent and well

Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora

Model ReleasesDGX agent

arXiv:2605.27239v1 Announce Type: new Abstract: Annotation quality is difficult to sustain when campaigns span weeks or months with small annotator pools. We present a Setswana sentiment dataset of 3,

The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works

Model ReleasesDGX agent

arXiv:2605.26246v1 Announce Type: new Abstract: Knowledge distillation (KD) transfers knowledge from a large teacher model to a smaller student. In language modeling, the student is trained either on

The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?

Model ReleasesDGX agent

arXiv:2605.27176v1 Announce Type: new Abstract: Knowledge graphs (KGs) can provide structured scientific context to language models, but it remains unclear which graph facts actually shape the generat

The MiniMax M2 series was one of the most widely used open-weight LLM series earlier this year. Now, we got a technical report with some int…

Model ReleasesDGX agent

The MiniMax M2 series was one of the most widely used open-weight LLM series earlier this year. Now, we got a technical report with some interesting tidbits. I summarized some of them below: 1. Full a

The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training

Model ReleasesDGX agent

arXiv:2605.26489v1 Announce Type: new Abstract: Large language model pre-training typically exhibits a two-phase trajectory: a fast initial loss drop followed by a prolonged slow improvement. We ident

Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V

Model ReleasesDGX agent

arXiv:2605.27003v1 Announce Type: cross Abstract: W4A4 quantization of large video diffusion Transformers offers substantial memory savings but is hindered by two main challenges: sparse large-magnitu

Today we're announcing ESMFold2, an open scientific engine to power prediction, design, and discovery across protein biology. The new model …

Model ReleasesDGX agent

Today we're announcing ESMFold2, an open scientific engine to power prediction, design, and discovery across protein biology. The new model delivers state of the art performance on protein interaction

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets

Model ReleasesDGX agent

arXiv:2605.26165v1 Announce Type: cross Abstract: Agentic RAG systems that equip language models with dozens to hundreds of tool definitions face a critical resource conflict: tool schemas consume the

Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records

Model ReleasesDGX agent

arXiv:2605.26463v1 Announce Type: cross Abstract: Data consistency between unstructured clinical notes and structured tables in Electronic Health Records (EHRs) is essential for patient safety and cli

Towards Generalization-Oriented Models for Vehicle Routing Problems with Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2605.26776v1 Announce Type: cross Abstract: In recent years, Deep Reinforcement Learning (DRL) has achieved substantial progress on Vehicle Routing Problems (VRPs). However, existing DRL-based m

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

Model ReleasesDGX agent

arXiv:2601.05899v2 Announce Type: replace Abstract: Recent breakthroughs in Large Language Models (LLMs) have positioned them as a promising paradigm for agents, with long-term planning and decision-m

Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry

Model ReleasesDGX agent

arXiv:2605.27071v1 Announce Type: new Abstract: Key knowledge for steel-industry volatile organic compounds (VOCs) governance is scattered across unstructured scientific literature, making it difficul

Trust Region Q Adjoint Matching

Model ReleasesDGX agent

arXiv:2605.27079v1 Announce Type: cross Abstract: Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step s

Two-Parameter Flows for Learning Population Dynamics of Physical Systems

Model ReleasesDGX agent

arXiv:2605.26285v1 Announce Type: new Abstract: This work addresses the problem of learning the dynamics of high-dimensional probability densities over time using unlabeled samples, without assuming a

Underwater360: Reconstructing Underwater Scenes from Panoramic Images with Omnidirectional Gaussian Splatting

Model ReleasesDGX agent

arXiv:2605.26447v1 Announce Type: new Abstract: Underwater scene reconstruction is essential for immersive exploration of aquatic environments, yet remains challenging due to complex participating-med

UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.26646v1 Announce Type: new Abstract: LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules

Variational Inference for Evidential Deep Learning

Model ReleasesDGX agent

arXiv:2605.26477v1 Announce Type: new Abstract: While Deep Neural Networks (DNNs) achieve remarkable performance, their tendency to produce overconfident predictions. Evidential Deep Learning (EDL) mi

Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization

Model ReleasesDGX agent

arXiv:2605.26433v1 Announce Type: new Abstract: Large language model (LLM) summarization systems may pass compact vector representations of private inputs to downstream retrieval, monitoring, audit, o

Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

Model ReleasesDGX agent

arXiv:2605.26457v1 Announce Type: cross Abstract: AI coding agents are increasingly used to write real-world software, but ensuring that their outputs are correct remains a fundamental challenge. Form

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

Model ReleasesDGX agent

arXiv:2605.26144v1 Announce Type: cross Abstract: We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

Model ReleasesDGX agent

arXiv:2605.26380v1 Announce Type: cross Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

Model ReleasesDGX agent

arXiv:2605.27141v1 Announce Type: new Abstract: Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks. Effective collaboration in such setti

Warp’s big bet on building open source with GPT-5.5

Model ReleasesDGX agent

Warp is making a significant investment in developing open source tools and integrations built on GPT-5.5, OpenAI's advanced language model. The initiative aims to leverage GPT-5.5's capabilities to c

We’ve been putting a lot of effort into making Claude Code more responsive & reliable. Here’s an update on everything we’ve done:

Model ReleasesDGX agent

Anthropic has made improvements to Claude Code's responsiveness and reliability, with Thariq providing details on the enhancements implemented. The update likely covers performance optimizations, bug

We've created the world's fastest PDF parser ⚡️ And it's more accurate than any other open-source, model-free PDF parser out there (pymupdf,…

Model ReleasesDGX agent

We've created the world's fastest PDF parser ⚡️ And it's more accurate than any other open-source, model-free PDF parser out there (pymupdf, pypdf, markitdown, pdftotext, opendataloader, pymupdf4llm)

What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies

Model ReleasesDGX agent

arXiv:2507.06513v3 Announce Type: replace Abstract: Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To

What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

Model ReleasesDGX agent

arXiv:2605.26795v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting reliably improves language-model accuracy, but which properties of a rationale text drive the improvement is poorly und

What Molecular Structure Cannot Tell Us: A Taxonomy of Explainability Gaps in GNN-Based Drug Toxicity Prediction

Model ReleasesDGX agent

arXiv:2605.26183v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have emerged as a structurally natural approach for molecular toxicity prediction, operating directly on atomic connectiv

When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control

Model ReleasesDGX agent

arXiv:2605.26418v1 Announce Type: cross Abstract: A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every wor

When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation

Model ReleasesDGX agent

arXiv:2509.26600v2 Announce Type: replace-cross Abstract: As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inp

Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis

Model ReleasesDGX agent

arXiv:2605.26655v1 Announce Type: new Abstract: Automated prompt optimization methods (e.g., DSpy, TextGrad) can substantially improve the performance of large language model (LLM), however, their gen

WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization

Model ReleasesDGX agent

arXiv:2605.26660v1 Announce Type: new Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in

Workflow Closure Is Not Scientific Closure in Auto-Research Systems

Model ReleasesDGX agent

arXiv:2605.26200v1 Announce Type: cross Abstract: This paper argues that workflow closure is not scientific closure in auto-research systems. Current systems can increasingly complete research-like lo

Xreal launches a new sub-brand, X by Xreal, and the $299 a01 display glasses with micro OLED displays, 50° FOV, and 62g weight, set for July release in the US (Scott Stein/CNET)

Model ReleasesDGX agent

Scott Stein / CNET: Xreal launches a new sub-brand, X by Xreal, and the 299 a01 display glasses with micro OLED displays, 50° FOV, and 62g weight, set for July release in the US — X by Xreal is arrivi

Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems

Model ReleasesDGX agent

arXiv:2605.26302v1 Announce Type: new Abstract: Long-lived AI agents are increasingly deployed as persistent operational systems, yet they are still evaluated like freshly initialized models. Day-one

// Your Agents are Aging Too // Huh!? They need 'sleep,' and now they are aging? Joke aside, great write-up on reliable agentic engineering.…

Model ReleasesDGX agent

// Your Agents are Aging Too // Huh!? They need 'sleep,' and now they are aging? Joke aside, great write-up on reliable agentic engineering. This new research introduces AgingBench, a longitudinal rel

Zero-Shot MARL Benchmark in the Cyber-Physical Mobility Lab

Model ReleasesDGX agent

arXiv:2601.16578v2 Announce Type: replace Abstract: We present a reproducible benchmark for evaluating sim-to-real transfer of Multi-Agent Reinforcement Learning (MARL) policies for Connected and Auto

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

Model ReleasesDGX agent

arXiv:2605.26383v1 Announce Type: new Abstract: Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and l

26 May 2026

7AI launches PLAID ELITE fully managed agentic security operations service

Model ReleasesDGX agent

Agentic artificial intelligence security startup 7AI Inc. today announced the launch of PLAID ELITE, a fully managed AI-native security operations service. The new service combines autonomous investig

← Previous
1…208209210211212…377
Next →