AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
9 Jun 2026

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

Model ReleasesDGX agent

arXiv:2606.08959v1 Announce Type: new Abstract: We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO Wo

End-to-End Training for Discrete Token LLM based TTS System

Model ReleasesDGX agent

arXiv:2606.09234v1 Announce Type: cross Abstract: Recent state-of-the-art (SOTA) text-to-speech (TTS) systems typically adopt a cascaded pipeline consisting of a speech tokenizer, an autoregressive la

ERBench: A Benchmark and Testsuite for Equation Discovery Algorithms

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.09276v1 Announce Type: new Abstract: Equation discovery aims to automate the discovery of scientific models in the form of mathematical equations from data. Technically, equation discovery

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

Model ReleasesDGX agent

arXiv:2606.09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost

GD-MIL: Grade-Disentangled Multiple Instance Learning for Multimodal Biochemical Recurrence Prediction in Prostate Cancer

Model ReleasesDGX agent

arXiv:2606.09453v1 Announce Type: new Abstract: Biochemical recurrence (BCR) after radical prostatectomy is a critical endpoint in prostate cancer, yet risk stratification relies almost entirely on va

Generalization in Nonlinear Least Squares via Learned Feature Geometry

Model ReleasesDGX agent

arXiv:2606.08799v1 Announce Type: cross Abstract: We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

Model ReleasesDGX agent

arXiv:2606.08036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in academic research workflows, but scholarly tasks require high factual precision and therefore ex

GRPO Does Not Close the Multi-Agent Coordination Gap

Model ReleasesDGX agent

arXiv:2606.07845v1 Announce Type: cross Abstract: We measure how well current large language models coordinate as multiple agents sharing a common resource, using the dining philosophers problem as a

Guided Discovery of New Behaviors using Diffusion Policies

SafetyDGX agent

arXiv:2606.08743v1 Announce Type: new Abstract: Diffusion models have become a powerful tool for generative modeling in robotics, with diffusion policies excelling at modeling multimodal action-trajec

Hybrid Robustness Verification for Spatio-Temporal Neural Networks

Model ReleasesDGX agent

arXiv:2606.09746v1 Announce Type: cross Abstract: With AI increasingly deployed in safety-critical systems, providing formal robustness guarantees for the underlying models is essential. Existing veri

Language-based Trial and Error Falls Behind in the Era of Experience

Model ReleasesDGX agent

arXiv:2601.21754v3 Announce Type: replace Abstract: While Large Language Models (LLMs) excel in language-based agentic tasks, their applicability to unseen, nonlinguistic environments (e.g., symbolic

Learning from Human Driving: A Human-in-the-Loop Online Behavior Cloning Framework for Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.08170v1 Announce Type: new Abstract: With the evolution of large foundation models (LFMs), data-driven autonomous driving has made significant strides. However, existing paradigms still fac

Mean Teacher based SSL Framework for Indoor Localization Using Wi-Fi RSSI Fingerprinting

Local AiDGX agent

arXiv:2407.13303v2 Announce Type: replace Abstract: Conventional large-scale indoor localization based on Wi-Fi RSSI fingerprinting faces issues of time-consuming and labor-intensive labeled data coll

MedicalRec: Medical recommender system for image classification without retraining

ApplicationsDGX agent

arXiv:2606.07553v1 Announce Type: cross Abstract: The emergence of machine learning and deep learning has revolutionized the efficiency of diagnostic, therapeutic, and administrative systems in health

MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting

Model ReleasesDGX agent

arXiv:2601.09085v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) has become a standard approach for training mathematical reasoning models; however, its reliance on

Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs

Model ReleasesDGX agent

arXiv:2606.09411v1 Announce Type: cross Abstract: Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs. This creates a steganographic exfiltrati

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

Model ReleasesDGX agent

arXiv:2606.08572v1 Announce Type: new Abstract: While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability t

Report: GKE Inference Gateway delivers up to 92% faster AI responses

Model ReleasesDGX agent

As generative AI moves from experimental pilots to massive production environments, the efficiency of your infrastructure becomes the ultimate differentiator. One way to get the most out of it and min

Rewrite to Translate, Translate to Reward: Reinforcement Learning for Source Rewriting in Machine Translation

ResearchDGX agent

arXiv:2606.08011v1 Announce Type: cross Abstract: Although directly prompting off-the-shelf Large Language Models (LLMs) to generate meaning-preserving source rewrites can effectively enhance Machine

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding

Model ReleasesDGX agent

arXiv:2606.09064v1 Announce Type: cross Abstract: Recent advances in Video Large Language Models (Video-LLMs) have enabled performance on long-video understanding tasks. However, existing methods stil

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation

Model ReleasesDGX agent

arXiv:2606.08278v1 Announce Type: new Abstract: Humanoid foundation models are advancing faster than we can evaluate them. While real-world testing is expensive and difficult to reproduce, existing si

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

Model ReleasesDGX agent

arXiv:2606.09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However,

Still: Amortized KV Cache Compaction in a Single Forward Pass

Model ReleasesDGX agent

arXiv:2606.07878v1 Announce Type: new Abstract: The KV cache is the memory bottleneck of long-horizon language model deployment. Practically, a deployable compactor must be lightweight enough to call

Trajectory-Refined Distillation

Model ReleasesDGX agent

arXiv:2606.08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision alo

Vision-Language Asymmetry in Bistable Image Captioning

SafetyDGX agent

arXiv:2606.08031v1 Announce Type: new Abstract: Wittgenstein's duck-rabbit poses a question for vision-language models: when a model captions an ambiguous image, where in the model is the commitment t

What's the Point? Spatial Grammar & Index Resolution for Sign Language Processing

ResearchDGX agent

arXiv:2606.08056v1 Announce Type: cross Abstract: Sign language models are predominantly trained with gloss-sequence or text supervision, thereby under-modeling non-lexical and productive construction

XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad

Model ReleasesDGX agent

arXiv:2601.14063v2 Announce Type: replace-cross Abstract: Cross-cultural competence in large language models (LLMs) requires understanding and adapting Culture-Specific Items (CSIs) across varying cul

8 Jun 2026

A Comprehensive Anatomy of Human and DeepSeek-R1 LLM Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2606.07410v1 Announce Type: cross Abstract: The emergence of 'Aha moments' in large language models, particularly DeepSeek-R1-0120, has raised the question of whether these systems genuinely rea

A Held-Out Transition-Pair Falsifier for Long-Horizon Non-Abelian State Tracking

Model ReleasesDGX agent

arXiv:2606.07254v1 Announce Type: new Abstract: State tracking exposes a sharp limitation of sequence models: the relevant signal is often not a summary of observed tokens, but an ordered latent state

ADAGE: Active Defenses Against GNN Extraction

Model ReleasesDGX agent

arXiv:2503.00065v4 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) achieve high performance in various real-world applications, such as drug discovery, traffic states prediction, a

ARAPDiffusion: ARAP Regularization for Diffusion-Based Deformable Shape Space Learning

TutorialsDGX agent

arXiv:2606.06887v1 Announce Type: new Abstract: This paper introduces ARAPDiffusion, a latent diffusion model to learn the underlying continuous shape space of a deformation shape collection. The key

MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights

Model ReleasesDGX agent

arXiv:2606.07020v1 Announce Type: new Abstract: Multilingual and multicultural benchmarks now cover dozens of languages and model families, but the resulting score landscapes remain metric-rich and in

Nex-N2-Pro running locally https://huggingface.co/nex-agi/Nex-N2-Pro

IndustryDGX agent

Nex-N2-Pro is a model available on Hugging Face that can be run locally, enabling users to execute the model on their own infrastructure rather than relying on cloud services. The model is distributed

NTILC: Neural Tool Invocation via Learned Compression

Model ReleasesDGX agent

arXiv:2606.06566v1 Announce Type: cross Abstract: Agentic tool-calling language models depend on large registries of callable APIs, functions, and local actions. Placing full tool specifications direc

OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios

Model ReleasesDGX agent

arXiv:2606.06959v1 Announce Type: cross Abstract: Hallucination detection is essential for the reliable deployment of large language models (LLMs). However, existing evaluations face two core challeng

Product units in gated recurrent units improve nuclear-mass prediction

Model ReleasesDGX agent

arXiv:2606.06866v1 Announce Type: new Abstract: The prediction of masses of atomic nuclei using machine learning can complement theoretical models and advance the exploration of poorly known domains o

ReclAIm: A Multi-Agent Framework for Monitoring and Correcting Performance Decline in Medical Imaging AI

Model ReleasesDGX agent

arXiv:2510.17004v2 Announce Type: replace-cross Abstract: Purpose: To develop and evaluate a multi-agent framework (ReclAIm) for automated monitoring, detection, and correction of performance decline

Reversible Foundations: Training a 120B Sparse MoE through State-Preserving Scaling

Model ReleasesDGX agent

arXiv:2606.07404v1 Announce Type: new Abstract: This paper reports on training a hundred-billion-parameter sparse mixture of experts on a single eight-GPU node, end to end. LightningLM 0.1V is a recur

RhinoVLA Technical Report

Model ReleasesDGX agent

arXiv:2606.07383v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but real-time deployment on edge hardware remains challengin

Robustly estimating heterogeneity in factorial data using Rashomon Partitions

ResearchDGX agent

arXiv:2404.02141v5 Announce Type: replace-cross Abstract: In both observational data and randomized control trials, researchers select statistical models to articulate how the outcome of interest vari

Sparsely gated tiny linear experts

ApplicationsDGX agent

arXiv:2606.07414v1 Announce Type: new Abstract: Sparsity allows scaling model parameters without proportionally increasing computational cost. While mixture of experts (MoE) models are made increasing

The Post-GCN Decade Revisited: Curvature-Stratified Evaluation of Relational Learning

Model ReleasesDGX agent

arXiv:2606.06397v2 Announce Type: replace Abstract: Current evaluation practices in relational learning rely heavily on flat leaderboards that average performance across heterogeneous datasets, implic

Trading Engagement for Sustainability: Carbon-Aware Re-ranking for E-commerce Recommendations

Model ReleasesDGX agent

arXiv:2606.04550v1 Announce Type: cross Abstract: E-commerce recommender systems strongly influence which products users consider and purchase, yet sustainability signals such as Product Carbon Footpr

6 Jun 2026

Can LLMs Write Correct TLA+ Specifications? Evaluating Natural-Language-to-TLA+ Generation

Model ReleasesDGX agent

arXiv:2606.05792v1 Announce Type: new Abstract: TLA+ has supported industrial verification at companies such as Amazon and Microsoft, yet writing correct TLA+ specifications from natural language stil

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs

Model ReleasesDGX agent

arXiv:2606.05966v1 Announce Type: cross Abstract: Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) sti

GITCO: Gated Inference-Time Context Optimization in TSFMs

Model ReleasesDGX agent

arXiv:2606.05332v1 Announce Type: new Abstract: Patch-based Time Series Foundation Models (TSFMs) suffer from context poisoning: structurally anomalous patches capture disproportionate attention and s

Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing

Model ReleasesDGX agent

arXiv:2605.04733v2 Announce Type: replace Abstract: Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for

5 Jun 2026

Before the week ends, let's acknowledge one of the most INSANE week ever for open AI, with 25+ notable open-weight drops across every modali…

Model ReleasesDGX agent

Before the week ends, let's acknowledge one of the most INSANE week ever for open AI, with 25+ notable open-weight drops across every modality: 🧠 LLMs → NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-MoE,

Can LLMs Be Constrained to the Past? Improving Knowledge Cutoff through Recall-Based Prompting

Model ReleasesDGX agent

arXiv:2606.05804v1 Announce Type: new Abstract: Prompted knowledge cutoff instructs a large language model (LLM) to act as if information beyond a specified cutoff date were unavailable. However, prio

Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation

Model ReleasesDGX agent

arXiv:2606.05988v1 Announce Type: cross Abstract: Reasoning models produce long chain-of-thought traces that are costly to distill and encourage verbose student outputs. We study post-hoc compression

From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment

ResearchDGX agent

arXiv:2606.05180v1 Announce Type: new Abstract: Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts,

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

Model ReleasesDGX agent

arXiv:2511.20158v2 Announce Type: replace Abstract: While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominan

Improving Answer Extraction in Context-based Question Answering Systems Using LLMs

Model ReleasesDGX agent

arXiv:2606.06197v1 Announce Type: new Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face challenges in a

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

ResearchDGX agent

arXiv:2606.06286v1 Announce Type: new Abstract: Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather th

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

Model ReleasesDGX agent

arXiv:2606.05677v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon ta

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

Model ReleasesDGX agent

arXiv:2606.05176v1 Announce Type: new Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-s

ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces

Model ReleasesDGX agent

arXiv:2606.05402v1 Announce Type: new Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluat

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework

Model ReleasesDGX agent

arXiv:2606.05570v1 Announce Type: new Abstract: Repository-level coding benchmarks face a trade-off between task difficulty and evaluation reliability: tasks that challenge frontier models often invol

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

Model ReleasesDGX agent

arXiv:2606.06476v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning

Model ReleasesDGX agent

arXiv:2606.05576v1 Announce Type: new Abstract: Vision-language models (VLMs) excel on visual question answering and multimodal reasoning benchmarks. Yet their capability on ultra-resolution images -

← Previous
1…305306307308309…1036
Next →