AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,585 results
Model Releases

Towards Compact Sign Language Translation: Frame Rate and Model Size Trade-offs

DGX agent

arXiv:2605.09554v1 Announce Type: new Abstract: Sign Language Translation (SLT) converts sign language videos into spoken-language text, bridging communication between Deaf and hearing communities. Cu

model-releasesarxiv-cs-cl
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Towards Conversational Medical AI with Eyes, Ears and a Voice

DGX agent

arXiv:2605.09272v1 Announce Type: new Abstract: The practice of medicine relies not only upon skillful dialogue but also on the nuanced exchange and interpretation of rich auditory and visual cues bet

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Towards Cross-lingual Values Judgment: A Consensus-Pluralism Perspective

DGX agent

arXiv:2602.17283v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are employed worldwide, existing evaluation paradigms for their multilingual capabilities primarily focus on f

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse

DGX agent

arXiv:2605.09965v1 Announce Type: new Abstract: The real world unfolds along a single set of physics laws, yet human intelligence demonstrates a remarkable capacity to generalize experiences from this

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models

DGX agent

arXiv:2605.09670v1 Announce Type: cross Abstract: Teleoperation systems are fundamentally limited by communication latency, which degrades situational awareness and control performance. Predictive dis

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

DGX agent

arXiv:2605.10832v1 Announce Type: new Abstract: Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visua

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs

DGX agent

arXiv:2604.17502v2 Announce Type: replace Abstract: Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length trajectori

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias

DGX agent

arXiv:2605.09087v1 Announce Type: cross Abstract: Audio deepfake detection systems are increasingly deployed in high-stakes security applications, yet their fairness across demographic groups remains

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Towards Universal Gene Regulatory Network Inference: Unlocking Generalizable Regulatory Knowledge in Single-cell Foundation Models

DGX agent

arXiv:2605.08128v1 Announce Type: cross Abstract: Gene Regulatory Network (GRN) inference is essential for understanding complex cellular mechanisms, rendered tractable through single-cell transcripto

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

DGX agent

arXiv:2605.09934v1 Announce Type: new Abstract: Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, a

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Tracing Moral Foundations in Large Language Models

DGX agent

arXiv:2601.05437v2 Announce Type: replace-cross Abstract: Large language models often produce human-like moral judgments, but it is unclear whether this reflects an internal conceptual structure or su

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models

DGX agent

arXiv:2605.08974v1 Announce Type: cross Abstract: While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We arg

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Trajectory Supervision for Continual Tool-Use Learning in LLMs

DGX agent

arXiv:2605.09734v1 Announce Type: cross Abstract: Most language-model training data shows final artifacts, not the process that produced them. We study a tractable version of this question in tool use

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

TrajPrism: A Multi-Task Benchmark for Language-Grounded Urban Trajectory Understanding

DGX agent

arXiv:2605.10782v1 Announce Type: new Abstract: Urban mobility is naturally expressed both as trajectories in space and as natural-language descriptions of travel intent, constraints, and preferences.

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training

DGX agent

arXiv:2605.10835v1 Announce Type: new Abstract: Optical Music Recognition (OMR), the task of transcribing sheet music into a structured textual representation, is currently bottlenecked by a lack of l

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Transfer Learning of Multiobjective Indirect Low-Thrust Trajectories Using Diffusion Models and Markov Chain Monte Carlo

DGX agent

arXiv:2605.09125v1 Announce Type: cross Abstract: Preliminary low-thrust spacecraft mission design is a global search problem characterized by a complex solution landscape, multiple objectives, and nu

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models

DGX agent

arXiv:2601.22478v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Transformers Provably Learn Sparse XOR with Polylogarithmic Parameters

DGX agent

arXiv:2502.07553v2 Announce Type: replace Abstract: Learning sparse parity functions has become a theoretical testbed for studying feature learning in neural networks. However, existing analyses prima

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Try Grok Voice

DGX agent

Try Grok Voice Grok Voice Think Fast 1.0 ranks #1 on the Artificial Analysis τ-Voice benchmark for real-world agentic customer service resolution Absolutely outperforming GPT-Realtime-2 (High) and Gem

model-releaseselon-musk--x
12 May 2026
Model Releases

TTCD:Transformer Integrated Temporal Causal Discovery from Non-Stationary Time Series Data

DGX agent

arXiv:2605.08111v1 Announce Type: cross Abstract: The widespread availability of complex time series data in various domains such as environmental science, epidemiology, and economics demands robust c

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Two Ways to De-Bias an LLM-as-a-Judge: A Continuous-Score Comparison of Hierarchical Bayesian Calibration and Neural-ODE Score Transport

DGX agent

arXiv:2605.09227v1 Announce Type: new Abstract: [Abridged] Using a Large Language Model (LLM) as an automatic rater (LLM-as-a-judge) is cheap but potentially biased: some judges run lenient, others st

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

UFO: A Unified Flow-Oriented Framework for Robust Continual Graph Learning

DGX agent

arXiv:2605.09862v1 Announce Type: cross Abstract: Graph learning research has increasingly shifted toward continual graph learning (CGL), which better reflects real-world scenarios where graphs evolve

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

UMEDA: Unified Multi-modal Efficient Data Fusion for Privacy-Preserving Graph Federated Learning via Spectral-Gated Attention and Diffusion-Based Operator Alignment

DGX agent

arXiv:2605.08288v1 Announce Type: cross Abstract: Device-free localization trains models from heterogeneous wireless and visual sensors (e.g., Wi-Fi, LiDAR) distributed across edge devices. Federated

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Understanding Asynchronous Inference Methods for Vision-Language-Action Models

DGX agent

arXiv:2605.08168v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models offer a promising path to generalist robot control, but their inference latency causes observation staleness when

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning

DGX agent

arXiv:2605.10445v1 Announce Type: new Abstract: Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works la

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Unified Modeling of Lane and Lane Topology for Driving Scene Reasoning

DGX agent

arXiv:2605.08911v1 Announce Type: new Abstract: Autonomous vehicles need to perceive not only physical elements in the driving scene, such as lane lines and traffic lights, but also logical elements l

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Unifying Scientific Communication: Fine-Grained Correspondence Across Scientific Media

DGX agent

arXiv:2605.05831v2 Announce Type: replace Abstract: The communication of scientific knowledge has become increasingly multimodal, spanning text, visuals, and speech through materials such as research

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

UniShield: Unified Face Attack Detection via KG-Informed Multimodal Reasoning

DGX agent

arXiv:2605.08709v1 Announce Type: new Abstract: Unified face attack detection (UAD) requires recognizing physical spoofing and digital forgery within a shared decision space, yet existing discriminati

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

DGX agent

arXiv:2605.10889v1 Announce Type: cross Abstract: On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this sign

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Unpredictability dissociates from structured control in language agents

DGX agent

arXiv:2605.09692v1 Announce Type: new Abstract: Unpredictable behavior is often taken as evidence of control, yet stochastic dispersion and structured action control need not coincide. This paper test

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Unveiling High-Probability Generalization in Decentralized SGD

DGX agent

arXiv:2605.10205v1 Announce Type: new Abstract: Decentralized stochastic gradient descent (D-SGD) is an efficient method for large-scale distributed learning. Existing generalization studies mainly ad

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception

DGX agent

arXiv:2605.09936v1 Announce Type: new Abstract: We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imager

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

UserGPT Technical Report

DGX agent

arXiv:2605.08766v1 Announce Type: cross Abstract: Personalized user understanding from large-scale digital traces remains a fundamental challenge. Traditional user profiling methods rely on discrimina

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

UTS at PsyDefDetect: Multi-Agent Councils and Absence-Based Reasoning for Defense Mechanism Classification

DGX agent

arXiv:2605.09769v1 Announce Type: new Abstract: This paper describes our system for classifying psychological defense mechanisms in emotional support dialogues using the Defense Mechanism Rating Scale

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

V4FinBench: Benchmarking Tabular Foundation Models, LLMs, and Standard Methods on Corporate Bankruptcy Prediction

DGX agent

arXiv:2605.10896v1 Announce Type: new Abstract: Corporate bankruptcy prediction is a high-stakes financial task characterized by severe class imbalance and multi-horizon forecasting demands. Public da

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization

DGX agent

arXiv:2605.10405v1 Announce Type: new Abstract: Selecting the best large language model (LLM) for a fixed benchmark is often expensive, since exhaustive evaluation requires running every model on ever

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Vapi nabs $50M to make voice AI more human

DGX agent

Voice artificial intelligence startup Vapi Inc. said today it has raised 50 million in new funding to change the way people talk to computers, experience phone calls and interact with customer support

model-releasessiliconangle
12 May 2026
Model Releases

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models

DGX agent

arXiv:2603.18113v2 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly shape content generation, interaction, and decision-making across the Web, aligning them with hum

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

DGX agent

arXiv:2605.10485v1 Announce Type: new Abstract: Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominan

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation

DGX agent

arXiv:2605.08553v1 Announce Type: cross Abstract: Large language models can generate useful code from natural language, but their outputs come without correctness guarantees. Verifiable code generatio

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

VFM-SDM: A vision foundation model-based framework for training-free, marker-free, and calibration-free structural displacement measurement

DGX agent

arXiv:2605.09677v1 Announce Type: new Abstract: Reliable displacement measurement is fundamental for structural health monitoring and digital engineering workflows, as it provides direct structural re

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning

DGX agent

arXiv:2604.03701v2 Announce Type: replace Abstract: Video-based numerical reasoning provides a premier arena for testing whether Vision-Language Models (VLMs) truly 'understand' real-world dynamics, a

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

VISOR: A Vision-Language Model-based Test Oracle for Testing Robot

DGX agent

arXiv:2605.10408v1 Announce Type: cross Abstract: Testing robots requires assessing whether they perform their intended tasks correctly, dependably, and with high quality, a challenge known as the tes

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

VISTA: A Benchmark for Real-Time Video Streaming under Network Impairments in Surgical Teleoperation

DGX agent

arXiv:2605.08886v1 Announce Type: cross Abstract: Real-time video streaming is crucial in surgical teleoperation, yet reproducible evaluation under realistic network impairments remains limited. This

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

Visual-ERM: Reward Modeling for Visual Equivalence

DGX agent

arXiv:2603.13224v2 Announce Type: replace-cross Abstract: Vision-to-code tasks require models to reconstruct structured visual inputs, such as charts, tables, and SVGs, into executable or structured r

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving

DGX agent

arXiv:2605.08133v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

VORT: Adaptive Power-Law Memory for NLP Transformers

DGX agent

arXiv:2605.08966v1 Announce Type: new Abstract: Standard Transformers impose near-exponential decay on the influence of distant tokens, conflicting with the power-law structure of long-range dependenc

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning

DGX agent

arXiv:2605.08146v1 Announce Type: cross Abstract: Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domai

model-releasesarxiv-cs-ai
12 May 2026
← Previous
1…337338339340341…471
Next →