AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
All
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,628 results
Model Releases

SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution

DGX agent

arXiv:2603.01327v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong performance on self-contained programming tasks. However, they still struggle with repository-leve

model-releasesarxiv-cs-cl
27 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews

DGX agent

arXiv:2605.26911v1 Announce Type: new Abstract: LLM-generated peer reviews are increasingly common at major venues, yet their deficiencies are hard to detect because they are uniformly fluent and well

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora

DGX agent

arXiv:2605.27239v1 Announce Type: new Abstract: Annotation quality is difficult to sustain when campaigns span weeks or months with small annotator pools. We present a Setswana sentiment dataset of 3,

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

The Bridge-Garden Dilemma in LLM Distillation: Why Mixing Hard and Soft Labels Works

DGX agent

arXiv:2605.26246v1 Announce Type: new Abstract: Knowledge distillation (KD) transfers knowledge from a large teacher model to a smaller student. In language modeling, the student is trained either on

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?

DGX agent

arXiv:2605.27176v1 Announce Type: new Abstract: Knowledge graphs (KGs) can provide structured scientific context to language models, but it remains unclear which graph facts actually shape the generat

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

The MiniMax M2 series was one of the most widely used open-weight LLM series earlier this year. Now, we got a technical report with some int…

DGX agent

The MiniMax M2 series was one of the most widely used open-weight LLM series earlier this year. Now, we got a technical report with some interesting tidbits. I summarized some of them below: 1. Full a

model-releasessebastian-raschka--x
27 May 2026
Model Releases

The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training

DGX agent

arXiv:2605.26489v1 Announce Type: new Abstract: Large language model pre-training typically exhibits a two-phase trajectory: a fast initial loss drop followed by a prolonged slow improvement. We ident

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V

DGX agent

arXiv:2605.27003v1 Announce Type: cross Abstract: W4A4 quantization of large video diffusion Transformers offers substantial memory savings but is hindered by two main challenges: sparse large-magnitu

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Today we're announcing ESMFold2, an open scientific engine to power prediction, design, and discovery across protein biology. The new model …

DGX agent

Today we're announcing ESMFold2, an open scientific engine to power prediction, design, and discovery across protein biology. The new model delivers state of the art performance on protein interaction

model-releasesyann-lecun--x
27 May 2026
Model Releases

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets

DGX agent

arXiv:2605.26165v1 Announce Type: cross Abstract: Agentic RAG systems that equip language models with dozens to hundreds of tool definitions face a critical resource conflict: tool schemas consume the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records

DGX agent

arXiv:2605.26463v1 Announce Type: cross Abstract: Data consistency between unstructured clinical notes and structured tables in Electronic Health Records (EHRs) is essential for patient safety and cli

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Towards Generalization-Oriented Models for Vehicle Routing Problems with Mixture-of-Experts

DGX agent

arXiv:2605.26776v1 Announce Type: cross Abstract: In recent years, Deep Reinforcement Learning (DRL) has achieved substantial progress on Vehicle Routing Problems (VRPs). However, existing DRL-based m

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

DGX agent

arXiv:2601.05899v2 Announce Type: replace Abstract: Recent breakthroughs in Large Language Models (LLMs) have positioned them as a promising paradigm for agents, with long-term planning and decision-m

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry

DGX agent

arXiv:2605.27071v1 Announce Type: new Abstract: Key knowledge for steel-industry volatile organic compounds (VOCs) governance is scattered across unstructured scientific literature, making it difficul

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Trust Region Q Adjoint Matching

DGX agent

arXiv:2605.27079v1 Announce Type: cross Abstract: Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step s

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Two-Parameter Flows for Learning Population Dynamics of Physical Systems

DGX agent

arXiv:2605.26285v1 Announce Type: new Abstract: This work addresses the problem of learning the dynamics of high-dimensional probability densities over time using unlabeled samples, without assuming a

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Underwater360: Reconstructing Underwater Scenes from Panoramic Images with Omnidirectional Gaussian Splatting

DGX agent

arXiv:2605.26447v1 Announce Type: new Abstract: Underwater scene reconstruction is essential for immersive exploration of aquatic environments, yet remains challenging due to complex participating-med

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.26646v1 Announce Type: new Abstract: LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Variational Inference for Evidential Deep Learning

DGX agent

arXiv:2605.26477v1 Announce Type: new Abstract: While Deep Neural Networks (DNNs) achieve remarkable performance, their tendency to produce overconfident predictions. Evidential Deep Learning (EDL) mi

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization

DGX agent

arXiv:2605.26433v1 Announce Type: new Abstract: Large language model (LLM) summarization systems may pass compact vector representations of private inputs to downstream retrieval, monitoring, audit, o

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

DGX agent

arXiv:2605.26457v1 Announce Type: cross Abstract: AI coding agents are increasingly used to write real-world software, but ensuring that their outputs are correct remains a fundamental challenge. Form

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

DGX agent

arXiv:2605.26144v1 Announce Type: cross Abstract: We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

DGX agent

arXiv:2605.26380v1 Announce Type: cross Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

DGX agent

arXiv:2605.27141v1 Announce Type: new Abstract: Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks. Effective collaboration in such setti

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Warp’s big bet on building open source with GPT-5.5

DGX agent

Warp is making a significant investment in developing open source tools and integrations built on GPT-5.5, OpenAI's advanced language model. The initiative aims to leverage GPT-5.5's capabilities to c

model-releasesopenai
27 May 2026
Model Releases

We’ve been putting a lot of effort into making Claude Code more responsive & reliable. Here’s an update on everything we’ve done:

DGX agent

Anthropic has made improvements to Claude Code's responsiveness and reliability, with Thariq providing details on the enhancements implemented. The update likely covers performance optimizations, bug

model-releasesthariq--x
27 May 2026
Model Releases

We've created the world's fastest PDF parser ⚡️ And it's more accurate than any other open-source, model-free PDF parser out there (pymupdf,…

DGX agent

We've created the world's fastest PDF parser ⚡️ And it's more accurate than any other open-source, model-free PDF parser out there (pymupdf, pypdf, markitdown, pdftotext, opendataloader, pymupdf4llm)

model-releasesjerry-liu--x
27 May 2026
Model Releases

What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies

DGX agent

arXiv:2507.06513v3 Announce Type: replace Abstract: Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

DGX agent

arXiv:2605.26795v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting reliably improves language-model accuracy, but which properties of a rationale text drive the improvement is poorly und

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

What Molecular Structure Cannot Tell Us: A Taxonomy of Explainability Gaps in GNN-Based Drug Toxicity Prediction

DGX agent

arXiv:2605.26183v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have emerged as a structurally natural approach for molecular toxicity prediction, operating directly on atomic connectiv

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control

DGX agent

arXiv:2605.26418v1 Announce Type: cross Abstract: A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every wor

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation

DGX agent

arXiv:2509.26600v2 Announce Type: replace-cross Abstract: As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inp

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis

DGX agent

arXiv:2605.26655v1 Announce Type: new Abstract: Automated prompt optimization methods (e.g., DSpy, TextGrad) can substantially improve the performance of large language model (LLM), however, their gen

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization

DGX agent

arXiv:2605.26660v1 Announce Type: new Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Workflow Closure Is Not Scientific Closure in Auto-Research Systems

DGX agent

arXiv:2605.26200v1 Announce Type: cross Abstract: This paper argues that workflow closure is not scientific closure in auto-research systems. Current systems can increasingly complete research-like lo

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Xreal launches a new sub-brand, X by Xreal, and the $299 a01 display glasses with micro OLED displays, 50° FOV, and 62g weight, set for July release in the US (Scott Stein/CNET)

DGX agent

Scott Stein / CNET: Xreal launches a new sub-brand, X by Xreal, and the 299 a01 display glasses with micro OLED displays, 50° FOV, and 62g weight, set for July release in the US — X by Xreal is arrivi

model-releasestechmeme
27 May 2026
Model Releases

Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems

DGX agent

arXiv:2605.26302v1 Announce Type: new Abstract: Long-lived AI agents are increasingly deployed as persistent operational systems, yet they are still evaluated like freshly initialized models. Day-one

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

// Your Agents are Aging Too // Huh!? They need 'sleep,' and now they are aging? Joke aside, great write-up on reliable agentic engineering.…

DGX agent

// Your Agents are Aging Too // Huh!? They need 'sleep,' and now they are aging? Joke aside, great write-up on reliable agentic engineering. This new research introduces AgingBench, a longitudinal rel

model-releasesdair-ai--x
27 May 2026
Model Releases

Zero-Shot MARL Benchmark in the Cyber-Physical Mobility Lab

DGX agent

arXiv:2601.16578v2 Announce Type: replace Abstract: We present a reproducible benchmark for evaluating sim-to-real transfer of Multi-Agent Reinforcement Learning (MARL) policies for Connected and Auto

model-releasesarxiv-cs-ro
27 May 2026
Model Releases

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

DGX agent

arXiv:2605.26383v1 Announce Type: new Abstract: Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and l

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

7AI launches PLAID ELITE fully managed agentic security operations service

DGX agent

Agentic artificial intelligence security startup 7AI Inc. today announced the launch of PLAID ELITE, a fully managed AI-native security operations service. The new service combines autonomous investig

model-releasessiliconangle
26 May 2026
Model Releases

A Blended Likelihood Approach for Achieving Fairness Using Naive Bayes

DGX agent

arXiv:2605.25228v1 Announce Type: new Abstract: Concerns about algorithmic bias and fairness have increased as artificial intelligence has been incorporated into high-stakes decision-making. Tradition

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

A Comprehensive Dataset for Human vs. AI Generated Text Detection

DGX agent

arXiv:2510.22874v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authentic

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

A computational phase transition for learning-to-sample from Ising models

DGX agent

arXiv:2605.24752v1 Announce Type: new Abstract: We study learning-to-sample -- a basic algorithmic task underlying generative modeling -- for Ising models, a standard testbed for algorithmic ideas in

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis

DGX agent

arXiv:2605.25502v1 Announce Type: cross Abstract: Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because e

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?

DGX agent

arXiv:2605.24045v1 Announce Type: cross Abstract: Protein-ligand modeling underpins computational drug discovery and molecular design. Existing protein-ligand benchmarks typically evaluate whether a p

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

A Learning Stability Profile for Finite-Dimensional Learning Dynamics

DGX agent

arXiv:2512.21208v3 Announce Type: replace Abstract: We develop a finite-dimensional sensitivity framework for studying stability in learning systems whose states include representations, parameters, a

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

A lift for input-convex neural network training

DGX agent

arXiv:2605.24274v1 Announce Type: new Abstract: Input-convex neural networks (ICNNs) are widely used for log-concave density estimation, convex-potential normalizing flows, optimal transport, and tran

model-releasesarxiv-cs-lg
26 May 2026
← Previous
1…261262263264265…472
Next →