AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,577 results
Model Releases

Are vision-language models ready to zero-shot replace supervised classification models in agriculture?

DGX agent

arXiv:2512.15977v3 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly proposed as general-purpose solutions for visual recognition tasks, yet their reliability for agricul

model-releasesarxiv-cs-cv
12 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Artificial Intelligence in Number Theory: LLMs for Algorithm Generation and Ensemble Methods for Conjecture Verification

DGX agent

arXiv:2504.19451v3 Announce Type: cross Abstract: This paper presents two concrete applications of Artificial Intelligence to algorithmic and analytic number theory. Recent benchmarks of large languag

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

DGX agent

arXiv:2605.10876v1 Announce Type: cross Abstract: Recent advances in machine learning and large-scale biological data collections have revived the prospect of building a virtual cell, a computational

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

AssemPlanner: A Multi-Agent Based Task Planning Framework for Flexible Assembly System

DGX agent

arXiv:2605.08831v1 Announce Type: new Abstract: In flexible assembly systems, existing task planning methods require a time-consuming configuration process by multiple experts to establish a productio

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

ASTRA-QA: A Benchmark for Abstract Question Answering over Documents

DGX agent

arXiv:2605.10168v1 Announce Type: new Abstract: Document-based question answering (QA) increasingly includes abstract questions that require synthesizing scattered information from long documents or a

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Attention Grounded Enhancement for Visual Document Retrieval

DGX agent

arXiv:2511.13415v2 Announce Type: replace-cross Abstract: Visual document retrieval requires understanding heterogeneous and multi-modal content to satisfy implicit information needs. Recent advances

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

AUHead: Realistic Emotional Talking Head Generation via Action Units Control

DGX agent

arXiv:2602.09534v2 Announce Type: replace Abstract: Realistic talking-head video generation is critical for virtual avatars, film production, and interactive systems. Current methods struggle with nua

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Automated Approach for Solving Infinite-state Polynomial Reachability Games

DGX agent

arXiv:2605.10169v1 Announce Type: new Abstract: Reachability games are two-player games played on a graph, where the objective of exttt{REACH} player is to reach the target set whereas the objective o

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation

DGX agent

arXiv:2605.10845v1 Announce Type: cross Abstract: As global cross-lingual communication intensifies, language barriers in visually rich documents such as PDFs remain a practical bottleneck. Existing d

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

BCJR-QAT: A Differentiable Relaxation of Trellis-Coded Weight Quantization

DGX agent

arXiv:2605.10655v1 Announce Type: new Abstract: Trellis-coded quantization sets the current 2-bit post-training frontier for LLMs (QTIP), but pushing below the PTQ ceiling requires quantization-aware

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

BEACON: A Multimodal Dataset for Learning Behavioral Fingerprints from Gameplay Data

DGX agent

arXiv:2605.10867v1 Announce Type: cross Abstract: Continuous authentication in high-stakes digital environments requires datasets with fine-grained behavioral signals under realistic cognitive and mot

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD

DGX agent

arXiv:2605.10865v1 Announce Type: new Abstract: Industrial Computer-Aided Design (CAD) code generation requires models to produce executable parametric programs from visual or textual inputs. Beyond r

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

BenchHAR: Benchmarking Self-Supervised Learning for Generalizable Sensor-based Activity Recognition

DGX agent

arXiv:2605.08296v1 Announce Type: new Abstract: Human Activity Recognition (HAR) from wearable sensors supports broad healthcare and behavior science applications. However, data heterogeneity and the

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Benchmarking Compositional Generalisation for Machine Learning Interatomic Potentials

DGX agent

arXiv:2605.08988v1 Announce Type: cross Abstract: Machine Learning Interatomic Potentials play a fundamental role in computational chemistry and materials science, enabling applications from molecular

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing

DGX agent

arXiv:2605.10146v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces criti

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Benchmarking Sensor-Fault Robustness in Forecasting

DGX agent

arXiv:2605.10822v1 Announce Type: new Abstract: Cyber-physical system (CPS) forecasting models depend on sensor streams with noisy, biased, missing, or temporally misaligned readings, yet standard for

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Benchmarking Transformer and xLSTM for Time-Series Forecasting of Heat Consumption

DGX agent

arXiv:2605.09722v1 Announce Type: new Abstract: Obtaining an accurate short-term forecasting for heat demand is an essential part of operating district heating networks cost-efficient and reliable. He

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

@BereznevKi20669 @ggerganov Yes I believe the real llama.cpp revolution is yet to happen at its full scale. As computers will have more RAM …

DGX agent

@BereznevKi20669 @ggerganov Yes I believe the real llama.cpp revolution is yet to happen at its full scale. As computers will have more RAM and models will improve, and *if* China will continue shippi

model-releasesgeorgi-gerganov--x
12 May 2026
Model Releases

Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning

DGX agent

arXiv:2605.09292v1 Announce Type: new Abstract: Large language models now achieve high final-answer accuracy on mathematical reasoning benchmarks, but accuracy alone does not capture reasoning flexibi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation

DGX agent

arXiv:2605.09441v1 Announce Type: new Abstract: The pursuit of general-purpose embodied agents is hindered by fragmented evaluation protocols that isolate navigation skills and fixate on specific robo

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models

DGX agent

arXiv:2605.09496v1 Announce Type: new Abstract: Large language models represent the same reasoning in vastly different surface forms -- English prose, Python code, mathematical notation -- yet whether

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Beyond Local Edits: Embedding-Virtualized Knowledge for Broader Evaluation and Preservation of Model Editing

DGX agent

arXiv:2602.01977v2 Announce Type: replace Abstract: Knowledge editing methods for large language models are commonly evaluated using predefined benchmarks that assess edited facts together with a limi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers

DGX agent

arXiv:2605.10901v1 Announce Type: new Abstract: Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving

DGX agent

arXiv:2605.10034v1 Announce Type: new Abstract: Recent Autonomous Driving (AD) works such as GigaFlow and PufferDrive have unlocked Reinforcement Learning (RL) at scale as a training strategy for driv

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

Beyond source code: The files AI coding agents trust — and attackers exploit

DGX agent

As AI coding agents become deeply embedded in developer workflows, defenders must evolve their definition of malicious files and rethink how to protect against them. Autonomous AI agents operate acros

model-releasesgoogle-cloud-ai
12 May 2026
Model Releases

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

DGX agent

arXiv:2605.08761v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles,

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Beyond the False Trade-off: Adaptive EWC for Stealthy and Generalizable T2I Backdoors

DGX agent

arXiv:2605.08280v1 Announce Type: cross Abstract: Preserving model fidelity is essential for stealthy text-to-image (T2I) backdoor attacks. Existing methods such as Learning without Forgetting (LwF) r

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation

DGX agent

arXiv:2502.08943v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated significant utility in real-world applications, exhibiting impressive capabilities in natural l

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Beyond Toy Benchmarks: A Systematic Evaluation of OOD Detection Methods For Plant Pathology Classification

DGX agent

arXiv:2605.08618v1 Announce Type: new Abstract: Out-of-distribution (OOD) detection is essential for reliable deployment of deep learning systems, yet the majority of existing methods are evaluated on

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization

DGX agent

arXiv:2605.10345v1 Announce Type: new Abstract: Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence

DGX agent

arXiv:2605.09041v1 Announce Type: new Abstract: Bias audits of large language models now operate within governance frameworks such as the EU AI Act, making benchmark reliability a security concern in

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Bilinear autoencoders find interpretable manifolds

DGX agent

arXiv:2605.08891v1 Announce Type: new Abstract: Sparse autoencoders have become a standard tool for uncovering interpretable latent representations in neural networks. Yet salient concepts often span

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

BoostLLM: Boosting-inspired LLM Fine-tuning for Few-shot Tabular Classification

DGX agent

arXiv:2605.06117v2 Announce Type: replace Abstract: Large language models (LLMs) have recently been adapted to tabular prediction by serializing structured features into natural language, but their pe

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

DGX agent

arXiv:2605.08271v1 Announce Type: cross Abstract: Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Bridging Spectral Operator Learning and U-Net Hierarchies: SpectraNet for Stable Autoregressive PDE Surrogates

DGX agent

arXiv:2605.09096v1 Announce Type: new Abstract: Neural operators for time-dependent PDEs face a structural tension: spectral architectures (FNO and descendants) inherit exponential rollout-error growt

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Built Environment Reasoning from Remote Sensing Imagery Using Large Vision--Language Models

DGX agent

arXiv:2605.08404v1 Announce Type: cross Abstract: This work investigates the use of large language models (LLMs) for tasks in smart cities. The core idea is to leverage remote sensing imagery to chara

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

bViT: Investigating Single-Block Recurrence in Vision Transformers for Image Recognition

DGX agent

arXiv:2605.10661v1 Announce Type: cross Abstract: Vision Transformers (ViTs) are built by stacking independently parameterized blocks, but it remains unclear how much of this depth requires layer spec

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks

DGX agent

arXiv:2605.09611v1 Announce Type: new Abstract: This preprint presents an empirical analysis of byte-exact chunk-level deduplication in Retrieval-Augmented Generation (RAG) pipelines. We measure conte

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving

DGX agent

arXiv:2605.10744v1 Announce Type: new Abstract: Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

DGX agent

arXiv:2605.10873v1 Announce Type: cross Abstract: Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure because existin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

CALYREX: Cross-Attention LaYeR EXtended Transformers for System Prompt Anchoring

DGX agent

arXiv:2605.09737v1 Announce Type: new Abstract: Modern large language models (LLMs) rely on system prompts to establish behavioral constraints and safety rules. Standard causal self-attention treats p

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation

DGX agent

arXiv:2605.10448v1 Announce Type: new Abstract: Interactive agent benchmarks map an agent run to a binary outcome through outcome checks. When these checks rely on surface level signals or fail to cap

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

DGX agent

arXiv:2601.12369v3 Announce Type: replace Abstract: Deep Research Agents increasingly automate survey generation, yet whether they match human experts at retrieving essential papers and organizing the

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Can LLMs Predict Polymer Physics Just by Reading Synthesis and Processing Prose?

DGX agent

arXiv:2605.08255v1 Announce Type: cross Abstract: Can large language models predict physical and mechanical polymer properties simply by reading unstructured scientific prose? Polymer performance is r

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Can We Go Beyond Visual Features? Neural Tissue Relation Modeling for Relational Graph Analysis in Non-Melanoma Skin Histology

DGX agent

arXiv:2512.06949v3 Announce Type: replace Abstract: Histopathology image segmentation is essential for delineating tissue structures in skin cancer diagnostics, but modeling spatial context and inter-

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness

DGX agent

arXiv:2605.09634v1 Announce Type: new Abstract: LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability ac

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts

DGX agent

arXiv:2503.05066v5 Announce Type: replace-cross Abstract: The Mixture of Experts (MoE) is an effective architecture for scaling large language models by leveraging sparse expert activation to balance

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models

DGX agent

arXiv:2605.10903v1 Announce Type: new Abstract: This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adapta

model-releasesarxiv-cs-cv
12 May 2026
← Previous
1…324325326327328…471
Next →