AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,797 results
Model Releases

NestedKV: Nested Memory Routing for Long-Context KV Cache Compression

DGX agent

arXiv:2605.26678v1 Announce Type: new Abstract: Long-context language models are limited by the memory footprint of the key-value (KV) cache. Existing training-free KV compression methods usually rank

model-releasesarxiv-cs-cl
27 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Neural Autoregressive Control Variates for the Quantum Monte Carlo Sign Problem

DGX agent

arXiv:2605.26814v1 Announce Type: cross Abstract: We train a pair of autoregressive models to construct zero-mean control variates to mitigate the sign problem in quantum Monte Carlo simulations. The

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation

DGX agent

arXiv:2605.26844v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uni

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

DGX agent

arXiv:2605.26584v1 Announce Type: new Abstract: Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

DGX agent

arXiv:2505.17163v2 Announce Type: replace-cross Abstract: Recent advancements in multimodal slow-thinking systems have demonstrated remarkable performance across various visual reasoning tasks. Howeve

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection

DGX agent

arXiv:2508.01253v2 Announce Type: replace Abstract: Existing studies typically investigate domain shift and category shift as independent problems, however, in real-world scenarios, the two types of s

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

DGX agent

arXiv:2603.16654v2 Announce Type: replace-cross Abstract: Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, espec

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

OMD-GraphRAG: Enhancing GraphRAG with Ontology-Guided Extraction, Multi-Dimensional Clustering and Dual-Channel Fusion

DGX agent

arXiv:2603.25152v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems face significant challenges in complex reasoning, multi-hop queries, and domain-specific QA. While exis

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

DGX agent

arXiv:2605.26485v1 Announce Type: cross Abstract: We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-vi

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation

DGX agent

arXiv:2605.26641v1 Announce Type: new Abstract: Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) e

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

DGX agent

arXiv:2605.26322v1 Announce Type: new Abstract: Theory of Mind (ToM), the ability to infer others' knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning

DGX agent

arXiv:2605.27083v1 Announce Type: new Abstract: Counterfactual tuning (CFT) has emerged as a promising paradigm for Large Language Model (LLM) unlearning by training models to generate alternative fic

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs

DGX agent

arXiv:2510.05864v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly operate on long inputs, yet their behavior when harmful sentences are sparsely embedded within such inputs

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research

DGX agent

arXiv:2601.21008v3 Announce Type: replace-cross Abstract: Operations Research practitioners debug infeasible models through an iterative process: inspecting Irreducible Infeasible Subsystems ( IIS), i

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

OSMa-Bench++: Toward Open-Ended Benchmarking of Semantic Mapping for Manipulation with Prompt-Generated Synthetic Scenes

DGX agent

arXiv:2605.26831v1 Announce Type: new Abstract: Semantic mapping methods are increasingly used as intermediate scene representations for downstream robotic reasoning and manipulation, yet their evalua

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech

DGX agent

arXiv:2605.26978v1 Announce Type: new Abstract: Text-to-speech (TTS) evaluation for low-resource non-Latin-script languages can fail when it relies on a single ASR round-trip word error rate (WER). A

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

DGX agent

arXiv:2511.20586v4 Announce Type: replace Abstract: Trustworthiness has become a key requirement for the deployment of artificial intelligence systems in safety-critical applications. Conventional eva

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Periodic Topological Deep Learning for Polymer Design and Discovery

DGX agent

arXiv:2605.26833v1 Announce Type: cross Abstract: Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging.

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark

DGX agent

arXiv:2506.00250v4 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions

DGX agent

arXiv:2605.27015v1 Announce Type: new Abstract: Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We in

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

DGX agent

arXiv:2602.17003v2 Announce Type: replace-cross Abstract: Large language models have advanced web agents, yet current agents lack personalization capabilities. Since users rarely specify every detail

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

'PhyWorldBench': A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

DGX agent

arXiv:2507.13428v3 Announce Type: replace-cross Abstract: Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurate

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

PIDM-DP: Physics-Informed Diffusion with Dormand-Prince Integration for Chaotic System Identification and State Reconstruction across Multiple Dynamical Regimes

DGX agent

arXiv:2605.26619v1 Announce Type: new Abstract: Reconstructing continuous state trajectories of chaotic dynamical systems from sparse, noisy observations remains a fundamental open problem in nonlinea

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

DGX agent

arXiv:2605.27258v1 Announce Type: cross Abstract: Building state-of-the-art text-to-speech (TTS) systems typically demands millions of hours of proprietary data and complex multi-stage architectures,

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

PitchBench: Measuring Pitch Hearing in Audio-Language Models

DGX agent

arXiv:2605.26176v1 Announce Type: cross Abstract: Audio-language models (ALMs) are increasingly used in real-world applications that require understanding music, from music tutoring and transcription

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Position: AI Safety Requires Effective Controllability

DGX agent

arXiv:2605.27117v1 Announce Type: new Abstract: AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing ha

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

PRBench: A Standardized Probabilistic Robustness Benchmark

DGX agent

arXiv:2511.01724v3 Announce Type: replace Abstract: Deep learning models are notoriously vulnerable to imperceptible perturbations. Most existing research centers on adversarial robustness (AR), which

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Pretrained Approximators for Low-Thrust Trajectory Cost and Reachability

DGX agent

arXiv:2605.26790v1 Announce Type: new Abstract: Low-thrust trajectory design relies heavily on repeated evaluations of fuel consumption and transfer feasibility, which require expensive optimal contro

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

DGX agent

arXiv:2605.26730v1 Announce Type: new Abstract: The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automate

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

PRISM: Position-encoded Regressive Inverse Spectral Model for Multilayer Thin-Film Design

DGX agent

arXiv:2605.26502v1 Announce Type: new Abstract: The inverse problem of multilayer thin-film optical coatings design represents a complex combinatorial-continuous optimization challenge. We present PRI

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Probing Cultural Awareness in LLMs: A Case Study of Cross-Culture Aesthetic Stylistics

DGX agent

arXiv:2605.27296v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in diverse cultural contexts, yet their ability to master aesthetic stylistics, i.e., the strateg

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals

DGX agent

arXiv:2605.26999v1 Announce Type: new Abstract: Prompt injection poses a critical threat to the safe deployment of large language models, yet existing detection approaches are typically evaluated unde

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Provably Communication-Efficient and Privacy-Preserving Federated Graph Neural Networks

DGX agent

arXiv:2605.26243v1 Announce Type: new Abstract: Graph neural networks (GNNs) achieve strong performance on relational data, but real-world graphs are often distributed across organizations that cannot

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Qiskit QuantumKatas: Adapting Microsoft's Quantum Computing exercises for LLM evaluation

DGX agent

arXiv:2605.27210v1 Announce Type: cross Abstract: We adapt Microsoft's QuantumKatas -- a well-established quantum computing curriculum -- from Q# to Qiskit, the most widely-adopted quantum computing f

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering

DGX agent

arXiv:2605.27164v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems for question answering typically retrieve evidence by semantic similarity between the query and document ch

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

🚀🚀 Qwen3.7-Max just hit #4 on Code Arena, on par with Claude Opus 4.6 ,top-ranked Chinese lab on the board! @arena More to ship. Stay tune…

DGX agent

🚀🚀 Qwen3.7-Max just hit #4 on Code Arena, on par with Claude Opus 4.6 ,top-ranked Chinese lab on the board! @arena More to ship. Stay tuned. 🕶️ Qwen3.7 Max (20250517) debuts at #4 in Code Arena: Front

model-releasesqwen--x
27 May 2026
Model Releases

RCSP: Risk-Sensitive Conjectural Scenario Planning for Safe Dynamic Robot Navigation

DGX agent

arXiv:2605.26348v1 Announce Type: new Abstract: Mobile robots can fail before they collide: a velocity that is safe now may commit the robot to a passage that moving obstacles will soon close. We stud

model-releasesarxiv-cs-ro
27 May 2026
Model Releases

Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions

DGX agent

arXiv:2605.26414v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve impressive accuracy on mathematical reasoning benchmarks, yet their performance drops when problems are modified wi

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Receipt Replay OOD: A Small Benchmark for Screen Replay Detection Under Domain Shift

DGX agent

arXiv:2605.26855v1 Announce Type: new Abstract: Public datasets such as DLC-2021, SynID, and KID34K have significantly contributed to research on presentation attack detection for identity documents,

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Reliable Extraction of Clinical Follow-Up Instructions: A Hybrid Neural-Symbolic Pipeline

DGX agent

arXiv:2605.26560v1 Announce Type: cross Abstract: Objective. Outpatient notes carry follow-up instructions pairing actions with future times ('MRI brain in two weeks'). Extracting (action, date) pairs

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference

DGX agent

arXiv:2605.27081v1 Announce Type: cross Abstract: Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining hi

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure

DGX agent

arXiv:2601.10566v5 Announce Type: replace Abstract: Entity-level unlearning is usually evaluated by what a model says: whether it stops naming the target, refuses a query, or shifts a Truth Ratio dist

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Resolving Ambiguity in Composed Image Retrieval via Calibrated Interaction

DGX agent

arXiv:2605.24634v2 Announce Type: replace Abstract: Composed image retrieval (CIR) searches a corpus with a reference image and a text describing how to modify it. Despite rapid progress from triplet-

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Rethinking the Multilingual Reasoning Gap with Layer Swap

DGX agent

arXiv:2605.26735v1 Announce Type: new Abstract: Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior wor

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Risk Averse Alert Prioritization for IDS Using Subnormal Gaussian Fuzzy Models

DGX agent

arXiv:2605.27299v1 Announce Type: cross Abstract: Modern intrusion detection systems generate thousands of alerts daily, but alert fatigue severely limits security operations effectiveness due to too

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data

DGX agent

arXiv:2605.26971v1 Announce Type: new Abstract: The proliferation of Reinforcement Learning from Verifiable Rewards (RLVR) datasets has exacerbated provenance collapse due to unclear lineage among exi

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction

DGX agent

arXiv:2605.26862v1 Announce Type: new Abstract: Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scen

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

DGX agent

arXiv:2605.27134v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking,

model-releasesarxiv-cs-ai
27 May 2026
← Previous
1…262263264265266…475
Next →