AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,593 results
Model Releases

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks

DGX agent

arXiv:2605.10286v1 Announce Type: new Abstract: Building effective clinical decision support systems requires the synthesis of complex heterogeneous multimodal data. Such modalities include temporal e

model-releasesarxiv-cs-ai
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design

DGX agent

arXiv:2605.08756v1 Announce Type: new Abstract: Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent works show t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AI has a secondary intent problem. I was negotiating a rent renewal (I live in NYC and they raised rent 10% because I guess they were bored)…

DGX agent

AI has a secondary intent problem. I was negotiating a rent renewal (I live in NYC and they raised rent 10% because I guess they were bored). The property manager said she would 'do everything she cou

model-releasesallie-k--miller--x
12 May 2026
Model Releases

Alice v1: Distillation-Enhanced Video Generation Surpassing Closed-Source Models

DGX agent

arXiv:2605.08115v1 Announce Type: cross Abstract: Wepresent Alice v1, a 14-billion parameter open-source video generation model that achieves state-of-the-art quality through consistency distillation

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

DGX agent

arXiv:2601.01762v2 Announce Type: replace-cross Abstract: Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. W

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling

DGX agent

arXiv:2604.08178v2 Announce Type: replace Abstract: In classical Reinforcement Learning from Human Feedback (RLHF), Reward Models (RMs) serve as the fundamental signal provider for model alignment. As

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play

DGX agent

arXiv:2605.09150v1 Announce Type: new Abstract: Poker is an imperfect information game that has served as a long-standing benchmark for decision-making under uncertainty. To maximize utility beyond th

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment

DGX agent

arXiv:2603.26680v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) evolve into lifelong AI assistants, LLM personalization has become a critical frontier. However, progress is c

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents

DGX agent

arXiv:2605.09698v1 Announce Type: new Abstract: As data-science agents shift from co-pilots to auto-pilots, silent misframing becomes a critical failure mode. Agents quietly commit to plausible but un

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

An Annotation Scheme and Classifier for Personal Facts in Dialogue

DGX agent

arXiv:2605.10339v1 Announce Type: new Abstract: The advancement of Large Language Models (LLMs) has enabled their application in personalized dialogue systems. We present an extended annotation scheme

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

An Empirical Study of Multi-Agent Collaboration for Automated Research

DGX agent

arXiv:2603.29632v2 Announce Type: replace-cross Abstract: As AI agents evolve, the community is rapidly shifting from single Large Language Models (LLMs) to Multi-Agent Systems (MAS) to overcome cogni

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation

DGX agent

arXiv:2605.10397v1 Announce Type: cross Abstract: Visual anomaly detection (VAD) is crucial in many real-world fields, such as industrial inspection, medical imaging, infrastructure monitoring, and re

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Anthropic announces 12 Claude plugins for the legal sector, including a 'commercial counsel' tool for reviewing vendor agreements and a bar exam study tool (Rachel Metz/Bloomberg)

DGX agent

Rachel Metz / Bloomberg: Anthropic announces 12 Claude plugins for the legal sector, including a “commercial counsel” tool for reviewing vendor agreements and a bar exam study tool — Anthropic PBC is

model-releasestechmeme
12 May 2026
Model Releases

AnyDepth-DETR/-YOLO: Any-depth object detection with a single network

DGX agent

arXiv:2605.09407v1 Announce Type: new Abstract: Modern object detectors are static, fixed-depth networks optimized for a single operating point, requiring separate models for different deployment scen

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Arcane: An Assertion Reduction Framework through Semantic Clustering and MCTS-Guided Rule Exploring

DGX agent

arXiv:2605.10107v1 Announce Type: new Abstract: Assertion-based Verification (ABV) is essential for ensuring that hardware designs conform to their intended specifications. However, existing automated

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Architecture, Not Scale: Circuit Localization in Large Language Models

DGX agent

arXiv:2605.08853v1 Announce Type: new Abstract: Mechanistic interpretability assumes that circuit analysis becomes harder as models scale. We challenge this assumption by showing that the attention ar

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Are vision-language models ready to zero-shot replace supervised classification models in agriculture?

DGX agent

arXiv:2512.15977v3 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly proposed as general-purpose solutions for visual recognition tasks, yet their reliability for agricul

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Artificial Intelligence in Number Theory: LLMs for Algorithm Generation and Ensemble Methods for Conjecture Verification

DGX agent

arXiv:2504.19451v3 Announce Type: cross Abstract: This paper presents two concrete applications of Artificial Intelligence to algorithmic and analytic number theory. Recent benchmarks of large languag

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

DGX agent

arXiv:2605.10876v1 Announce Type: cross Abstract: Recent advances in machine learning and large-scale biological data collections have revived the prospect of building a virtual cell, a computational

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AssemPlanner: A Multi-Agent Based Task Planning Framework for Flexible Assembly System

DGX agent

arXiv:2605.08831v1 Announce Type: new Abstract: In flexible assembly systems, existing task planning methods require a time-consuming configuration process by multiple experts to establish a productio

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

ASTRA-QA: A Benchmark for Abstract Question Answering over Documents

DGX agent

arXiv:2605.10168v1 Announce Type: new Abstract: Document-based question answering (QA) increasingly includes abstract questions that require synthesizing scattered information from long documents or a

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Attention Grounded Enhancement for Visual Document Retrieval

DGX agent

arXiv:2511.13415v2 Announce Type: replace-cross Abstract: Visual document retrieval requires understanding heterogeneous and multi-modal content to satisfy implicit information needs. Recent advances

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

AUHead: Realistic Emotional Talking Head Generation via Action Units Control

DGX agent

arXiv:2602.09534v2 Announce Type: replace Abstract: Realistic talking-head video generation is critical for virtual avatars, film production, and interactive systems. Current methods struggle with nua

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Automated Approach for Solving Infinite-state Polynomial Reachability Games

DGX agent

arXiv:2605.10169v1 Announce Type: new Abstract: Reachability games are two-player games played on a graph, where the objective of exttt{REACH} player is to reach the target set whereas the objective o

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation

DGX agent

arXiv:2605.10845v1 Announce Type: cross Abstract: As global cross-lingual communication intensifies, language barriers in visually rich documents such as PDFs remain a practical bottleneck. Existing d

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

BCJR-QAT: A Differentiable Relaxation of Trellis-Coded Weight Quantization

DGX agent

arXiv:2605.10655v1 Announce Type: new Abstract: Trellis-coded quantization sets the current 2-bit post-training frontier for LLMs (QTIP), but pushing below the PTQ ceiling requires quantization-aware

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

BEACON: A Multimodal Dataset for Learning Behavioral Fingerprints from Gameplay Data

DGX agent

arXiv:2605.10867v1 Announce Type: cross Abstract: Continuous authentication in high-stakes digital environments requires datasets with fine-grained behavioral signals under realistic cognitive and mot

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD

DGX agent

arXiv:2605.10865v1 Announce Type: new Abstract: Industrial Computer-Aided Design (CAD) code generation requires models to produce executable parametric programs from visual or textual inputs. Beyond r

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

BenchHAR: Benchmarking Self-Supervised Learning for Generalizable Sensor-based Activity Recognition

DGX agent

arXiv:2605.08296v1 Announce Type: new Abstract: Human Activity Recognition (HAR) from wearable sensors supports broad healthcare and behavior science applications. However, data heterogeneity and the

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Benchmarking Compositional Generalisation for Machine Learning Interatomic Potentials

DGX agent

arXiv:2605.08988v1 Announce Type: cross Abstract: Machine Learning Interatomic Potentials play a fundamental role in computational chemistry and materials science, enabling applications from molecular

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing

DGX agent

arXiv:2605.10146v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces criti

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Benchmarking Sensor-Fault Robustness in Forecasting

DGX agent

arXiv:2605.10822v1 Announce Type: new Abstract: Cyber-physical system (CPS) forecasting models depend on sensor streams with noisy, biased, missing, or temporally misaligned readings, yet standard for

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Benchmarking Transformer and xLSTM for Time-Series Forecasting of Heat Consumption

DGX agent

arXiv:2605.09722v1 Announce Type: new Abstract: Obtaining an accurate short-term forecasting for heat demand is an essential part of operating district heating networks cost-efficient and reliable. He

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

@BereznevKi20669 @ggerganov Yes I believe the real llama.cpp revolution is yet to happen at its full scale. As computers will have more RAM …

DGX agent

@BereznevKi20669 @ggerganov Yes I believe the real llama.cpp revolution is yet to happen at its full scale. As computers will have more RAM and models will improve, and *if* China will continue shippi

model-releasesgeorgi-gerganov--x
12 May 2026
Model Releases

Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning

DGX agent

arXiv:2605.09292v1 Announce Type: new Abstract: Large language models now achieve high final-answer accuracy on mathematical reasoning benchmarks, but accuracy alone does not capture reasoning flexibi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation

DGX agent

arXiv:2605.09441v1 Announce Type: new Abstract: The pursuit of general-purpose embodied agents is hindered by fragmented evaluation protocols that isolate navigation skills and fixate on specific robo

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models

DGX agent

arXiv:2605.09496v1 Announce Type: new Abstract: Large language models represent the same reasoning in vastly different surface forms -- English prose, Python code, mathematical notation -- yet whether

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Beyond Local Edits: Embedding-Virtualized Knowledge for Broader Evaluation and Preservation of Model Editing

DGX agent

arXiv:2602.01977v2 Announce Type: replace Abstract: Knowledge editing methods for large language models are commonly evaluated using predefined benchmarks that assess edited facts together with a limi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers

DGX agent

arXiv:2605.10901v1 Announce Type: new Abstract: Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving

DGX agent

arXiv:2605.10034v1 Announce Type: new Abstract: Recent Autonomous Driving (AD) works such as GigaFlow and PufferDrive have unlocked Reinforcement Learning (RL) at scale as a training strategy for driv

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

Beyond source code: The files AI coding agents trust — and attackers exploit

DGX agent

As AI coding agents become deeply embedded in developer workflows, defenders must evolve their definition of malicious files and rethink how to protect against them. Autonomous AI agents operate acros

model-releasesgoogle-cloud-ai
12 May 2026
Model Releases

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

DGX agent

arXiv:2605.08761v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles,

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Beyond the False Trade-off: Adaptive EWC for Stealthy and Generalizable T2I Backdoors

DGX agent

arXiv:2605.08280v1 Announce Type: cross Abstract: Preserving model fidelity is essential for stealthy text-to-image (T2I) backdoor attacks. Existing methods such as Learning without Forgetting (LwF) r

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation

DGX agent

arXiv:2502.08943v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated significant utility in real-world applications, exhibiting impressive capabilities in natural l

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Beyond Toy Benchmarks: A Systematic Evaluation of OOD Detection Methods For Plant Pathology Classification

DGX agent

arXiv:2605.08618v1 Announce Type: new Abstract: Out-of-distribution (OOD) detection is essential for reliable deployment of deep learning systems, yet the majority of existing methods are evaluated on

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization

DGX agent

arXiv:2605.10345v1 Announce Type: new Abstract: Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence

DGX agent

arXiv:2605.09041v1 Announce Type: new Abstract: Bias audits of large language models now operate within governance frameworks such as the EU AI Act, making benchmark reliability a security concern in

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Bilinear autoencoders find interpretable manifolds

DGX agent

arXiv:2605.08891v1 Announce Type: new Abstract: Sparse autoencoders have become a standard tool for uncovering interpretable latent representations in neural networks. Yet salient concepts often span

model-releasesarxiv-cs-lg
12 May 2026
← Previous
1…325326327328329…471
Next →