AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

DGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

model-releasesarxiv-cs-cl
22 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

DGX agent

arXiv:2602.01851v2 Announce Type: replace Abstract: Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

DGX agent

arXiv:2605.22064v1 Announce Type: new Abstract: Hy-MT2 is a family of fast-thinking multilingual translation models designed for complex real-world scenarios. It includes three model sizes: 1.8B, 7B,

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering

DGX agent

arXiv:2605.22035v1 Announce Type: cross Abstract: Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions

DGX agent

arXiv:2605.22247v1 Announce Type: new Abstract: Idioms pose a fundamental challenge for language models, as their meaning cannot be inferred from surface form alone. Understanding such expressions, th

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

InfVSR: Breaking Length Limits of Generic Video Super-Resolution

DGX agent

arXiv:2510.00948v2 Announce Type: replace Abstract: Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent c

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

InnerQ: Hardware-Aware Tuning-Free Quantization of KV Cache for Large Language Models

DGX agent

arXiv:2602.23200v2 Announce Type: replace-cross Abstract: When transformer-based language models are deployed for text generation, most of the inference time is spent in the decoding stage, where outp

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation

DGX agent

arXiv:2510.09724v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly capable of generating complete applications from natural language instructions, creating new opp

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements

DGX agent

arXiv:2605.22079v1 Announce Type: new Abstract: Large language models (LLMs) are widely used to generate structured outputs such as JSON, SQL, and code, yet public resources remain limited for evaluat

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

DGX agent

arXiv:2605.22080v1 Announce Type: new Abstract: We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF material

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models

DGX agent

arXiv:2605.21861v1 Announce Type: new Abstract: Multi-modality medical vision (MV) foundation models (FM) are fundamentally challenged by pronounced Non-IID feature statistics across heterogeneous ima

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

DGX agent

arXiv:2605.21988v1 Announce Type: new Abstract: Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

DGX agent

arXiv:2605.21573v1 Announce Type: new Abstract: We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Linear Dynamics in the RLVR Training of Large Language Models

DGX agent

arXiv:2601.04537v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven significant performance gains in reasoning-oriented large language models (LL

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

LLM Readiness Harness: Evaluation, Observability, and CI Gates for LLM/RAG Applications

DGX agent

arXiv:2603.27355v2 Announce Type: replace-cross Abstract: We present a readiness harness for LLM and RAG applications that turns evaluation into a deployment decision workflow. The system combines aut

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

LongVT: Incentivizing 'Thinking with Long Videos' via Native Tool Calling

DGX agent

arXiv:2511.20785v3 Announce Type: replace Abstract: Large multimodal models (LMMs) have shown great potential for video reasoning with textual Chain-of-Thought. However, they remain vulnerable to hall

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

DGX agent

arXiv:2605.22089v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sp

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

M3: Conversational LLMs Simplify Secure Clinical Data Access, Understanding, and Analysis

DGX agent

arXiv:2507.01053v4 Announce Type: replace-cross Abstract: Large-scale clinical databases offer opportunities for medical research, but their complexity creates barriers to effective use. The Medical I

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

DGX agent

arXiv:2605.22177v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and modular skills has endowed autonomous agents with increasingly powerful capabilities. Existing f

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MAP4TS: A Multi-Aspect Prompting Framework for Time-Series Forecasting with Large Language Models

DGX agent

arXiv:2510.23090v2 Announce Type: replace Abstract: Recent advances have investigated the use of pretrained large language models (LLMs) for time-series forecasting by aligning numerical inputs with L

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation

DGX agent

arXiv:2605.22469v1 Announce Type: new Abstract: Evaluating single-concept personalization in text-to-image diffusion requires measuring both concept preservation, which captures identity fidelity to a

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Matching with Deliberation: Test-Time Evolutionary Hierarchical Multi-Agents for Zero-Shot Compositional Image Retrieval

DGX agent

arXiv:2605.22478v1 Announce Type: new Abstract: Zero-Shot Compositional Image Retrieval (ZS-CIR) requires both preserving the visual continuity of the reference image and faithfully executing the sema

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

DGX agent

arXiv:2605.21917v1 Announce Type: new Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

DGX agent

arXiv:2605.21954v1 Announce Type: new Abstract: Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal la

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

DGX agent

arXiv:2605.21796v1 Announce Type: cross Abstract: Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

DGX agent

arXiv:2605.22818v1 Announce Type: new Abstract: Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally inco

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MOTOR: A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding

DGX agent

arXiv:2605.22550v1 Announce Type: new Abstract: Two-wheelers account for a disproportionately high share of road fatalities in the Global South. Research on two-wheeler rider behavior, however, lags f

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

DGX agent

arXiv:2505.17123v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Not All Starting Points Are Equal: Pre-trained Priors and Their Outsized Impact on Person Identification

DGX agent

arXiv:2507.17640v3 Announce Type: replace Abstract: Recent years have seen an explosion of diverse general purpose pre-training methodologies for computer vision. However, the impact that these pre-tr

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems

DGX agent

arXiv:2605.22144v1 Announce Type: new Abstract: Existing approaches for digital short-drama production typically rely on one-shot LLM generated scripts and loosely coupled pipelines, which fail to sat

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Open-World Evaluations for Measuring Frontier AI Capabilities

DGX agent

arXiv:2605.20520v1 Announce Type: new Abstract: Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind

DGX agent

arXiv:2605.20423v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on many language tasks, but their Theory of Mind (ToM) reasoning is still uneven in complex social settings. E

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

DGX agent

arXiv:2605.22200v1 Announce Type: new Abstract: Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment ho

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

PartCo: Part-Level Correspondence Priors Enhance Category Discovery

DGX agent

arXiv:2509.22769v2 Announce Type: replace Abstract: Generalized Category Discovery (GCD) aims to identify both known and novel categories within unlabeled data by leveraging a set of labeled examples

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

DGX agent

arXiv:2605.22109v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchm

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Physiology and Anatomy Aware Inverse Inference of Myocardial Infarction for Cardiac Digital Twin

DGX agent

arXiv:2605.22044v1 Announce Type: new Abstract: Accurate localization of myocardial infarction is essential for risk stratification. While LGE-MRI remains the gold standard, it is resource-intensive.

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Polite on the Surface, Wrong in Practice: A Curated Dataset for Fixing Honorific Failures in Multilingual Bangla Generation

DGX agent

arXiv:2605.22487v1 Announce Type: new Abstract: Recent advances in Multilingual Large Language Models (MLLMs) have significantly enhanced cross-lingual conversational capabilities, yet modeling cultur

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

DGX agent

arXiv:2605.22446v1 Announce Type: new Abstract: While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deplo

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

DGX agent

arXiv:2605.20251v2 Announce Type: cross Abstract: Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limi

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts

DGX agent

arXiv:2605.21776v1 Announce Type: new Abstract: Estimating mutual information from text usually requires training a task-specific critic, which limits its use in low-data settings. We ask whether larg

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP

DGX agent

arXiv:2603.20405v2 Announce Type: replace-cross Abstract: We report on an experiment in which Claude Opus~4.6, equipped with a suite of Model Context Protocol (MCP) tools for the Rocq proof assistant,

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems

DGX agent

arXiv:2510.13910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallu

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

DGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

DGX agent

arXiv:2605.22148v1 Announce Type: cross Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

DGX agent

arXiv:2605.20204v1 Announce Type: cross Abstract: LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstraine

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Reflective Prompt Tuning through Language Model Function-Calling

DGX agent

arXiv:2605.21781v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable of following instructions and complex reasoning, making prompting a flexible interface for

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Residual Skill Optimization for Text-to-SQL Ensembles

DGX agent

arXiv:2605.21792v1 Announce Type: new Abstract: Text-to-SQL ensembles improve over single-candidate generation by drawing multiple SQL candidates and selecting one, but their effectiveness is bounded

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Rethinking Noise-Robust Training for Frozen Vision Foundation Models: A Cross-Dataset Benchmark with a Case Study of Small-Loss Failure

DGX agent

arXiv:2605.22591v1 Announce Type: new Abstract: Frozen Vision Foundation Models (VFMs) with lightweight classification heads are increasingly used in medical imaging because they offer efficient and r

model-releasesarxiv-cs-cv
22 May 2026
← Previous
1…210211212213214…361
Next →