AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization

DGX agent

arXiv:2605.02223v1 Announce Type: cross Abstract: Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within a

model-releasesarxiv-cs-cv
5 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Towards High Fidelity Face Swapping: A Comprehensive Survey and New Benchmark

DGX agent

arXiv:2605.00883v1 Announce Type: new Abstract: Face swapping has witnessed significant progress in recent years, largely driven by advances in deep generative models such as GANs and diffusion models

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Towards Lightest Low-Light Image Enhancement Architecture for Mobile Devices

DGX agent

arXiv:2507.04277v2 Announce Type: replace Abstract: Real-time low-light image enhancement on mobile and embedded devices requires models that balance visual quality and computational efficiency. Exist

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Towards Visual Query Localization in the 3D World

DGX agent

arXiv:2605.01498v1 Announce Type: new Abstract: Visual query localization (VQL) aims to predict the spatio-temporal response of the most recent occurrence in a sequence given a query. Currently, most

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

TRACED: In vivo imaging of extracellular intrinsic diffusivity, tortuosity, cell size distribution and cell density in human glioma patients

DGX agent

arXiv:2605.02615v1 Announce Type: cross Abstract: The lack of analytical models describing diffusion time dependence at intermediate time scales in complex tissue microstructure limits the accurate qu

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Training-Free Time Series Classification via In-Context Reasoning with LLM Agents

DGX agent

arXiv:2510.05950v2 Announce Type: replace Abstract: Time series classification (TSC) spans diverse application scenarios, yet labeled data are often scarce, making task-specific training costly and in

model-releasesarxiv-cs-ai
5 May 2026
Model Releases

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation

DGX agent

arXiv:2605.00907v1 Announce Type: new Abstract: Large language models (LLMs) and multimodal large models (MLLMs) are increasingly used for transportation tasks such as regulation question answering, t

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Triple Spectral Fusion for Sensor-based Human Activity Recognition

DGX agent

arXiv:2605.02743v1 Announce Type: cross Abstract: The field of sensor-based human activity recognition (HAR) mainly uses posture, motion and context data of Inertial Measurement Units (IMUs) to identi

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Environments

DGX agent

arXiv:2510.04142v2 Announce Type: replace Abstract: This paper identifies a critical yet underexplored challenge in reasoning alignment from multiple multi-modal large language models (MLLMs): In non-

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

TwistNet-2D: Learning Second-Order Channel Interactions via Spiral Twisting for Texture Recognition

DGX agent

arXiv:2602.07262v3 Announce Type: replace Abstract: Second-order feature statistics are central to texture recognition, yet existing mechanisms exhibit a structural tension: bilinear pooling and Gram

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Two-Pass Zero-Shot Temporal-Spatial Grounding of Rare Traffic Events in Surveillance Video

DGX agent

arXiv:2605.01512v1 Announce Type: new Abstract: Grounding traffic accidents in real CCTV footage is a rare-event problem where training on labeled accident video is often prohibited, yet accurate join

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Understanding Emergent Misalignment via Feature Superposition Geometry

DGX agent

arXiv:2605.00842v1 Announce Type: cross Abstract: Emergent misalignment, where fine-tuning on narrow, non-harmful tasks induces harmful behaviors, poses a key challenge for AI safety in LLMs. Despite

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis

DGX agent

arXiv:2605.00826v1 Announce Type: cross Abstract: Text-to-video retrieval enables users to find relevant video content using natural language queries, a task that has grown increasingly important with

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Unsupervised full-field Bayesian inference of orthotropic hyperelasticity from a single biaxial test: a myocardial case study

DGX agent

arXiv:2510.09498v3 Announce Type: replace-cross Abstract: Cardiac muscle tissue exhibits highly non-linear hyperelastic and orthotropic material behavior during passive deformation. Traditional consti

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation

DGX agent

arXiv:2605.01517v1 Announce Type: new Abstract: Scalable Vector Graphics (SVG) animation generation is pivotal for professional design due to their structural editability and resolution independence.

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Variational Matrix-Learning Fourier Networks for Parametric Multiphysics Surrogates

DGX agent

arXiv:2605.02280v1 Announce Type: new Abstract: Multiphysics simulation is critical for system-technology co-optimization (STCO) in chiplet-based design, but repeated finite-element solutions of PDE-g

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

VeRO: An Evaluation Harness for Agents to Optimize Agents

DGX agent

arXiv:2602.22480v2 Announce Type: replace-cross Abstract: An important emerging application of coding agents is agent optimization: the iterative improvement of a target agent through edit-execute-eva

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models

DGX agent

arXiv:2605.01662v1 Announce Type: new Abstract: Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition

DGX agent

arXiv:2605.02834v1 Announce Type: new Abstract: Videos are unique in their ability to capture actions which transcend multiple frames. Accordingly, for many years action recognition was the quintessen

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

VILAS: A VLA-Integrated Low-cost Architecture with Soft Grasping for Robotic Manipulation

DGX agent

arXiv:2605.02037v1 Announce Type: new Abstract: We present VILAS, a fully low-cost, modular robotic manipulation platform designed to support end-to-end vision-language-action (VLA) policy learning an

model-releasesarxiv-cs-ro
5 May 2026
Model Releases

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

DGX agent

arXiv:2605.01391v1 Announce Type: new Abstract: Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Visual Implicit Autoregressive Modeling

DGX agent

arXiv:2605.01220v1 Announce Type: new Abstract: Visual Autoregressive Modeling (VAR) based on next-scale prediction achieves strong generation quality, but their explicit deep stacks fix the amount of

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

DGX agent

arXiv:2605.02735v1 Announce Type: new Abstract: Continuous latent-space reasoning offers a compact alternative to textual chain-of-thought for multimodal models, enabling high-dimensional visual evide

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Visualizing Critic Match Loss Landscapes for Interpretation of Online Reinforcement Learning Control Algorithms

DGX agent

arXiv:2603.14535v2 Announce Type: replace Abstract: Reinforcement learning has proven its power on various occasions. However, its performance is not always guaranteed when system dynamics change. Ins

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model

DGX agent

arXiv:2605.01194v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities and generalization in embodied manipulation. However, their decision-makin

model-releasesarxiv-cs-ro
5 May 2026
Model Releases

Watermarking LLM Agent Trajectories

DGX agent

arXiv:2602.18700v2 Announce Type: replace-cross Abstract: LLM agents rely heavily on high-quality trajectory data to guide their problem-solving behaviors, yet producing such data requires substantial

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models

DGX agent

arXiv:2605.02038v1 Announce Type: new Abstract: Single-prompt accuracy is the dominant way to benchmark language models, but it can miss reliability failures that matter. We evaluate a 15-model open-w

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition

DGX agent

arXiv:2605.02782v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models

DGX agent

arXiv:2605.02363v1 Announce Type: new Abstract: Deployed language models must produce outputs that are both correct and format-compliant. We study this structured-output reliability gap using two math

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation

DGX agent

arXiv:2605.00911v1 Announce Type: new Abstract: Industrial Retrieval-Augmented Generation (RAG) systems depend on optical character recognition (OCR) to transform visual documents into text. Existing

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

DGX agent

arXiv:2601.19827v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative retrieval-re

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

When Less Is More: Simplicity Beats Complexity for Physics-Constrained InSAR Phase Unwrapping

DGX agent

arXiv:2605.00896v1 Announce Type: new Abstract: Operational phase unwrapping is the primary computational bottleneck in InSAR-based volcanic and seismic monitoring. We challenge the industry trend of

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System

DGX agent

arXiv:2602.06932v3 Announce Type: replace Abstract: Speculative decoding can significantly accelerate LLM serving, yet most deployments today disentangle speculator training from serving, treating spe

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

WILD SAM: A Simulated-and-Real Data Augmentation for Autonomous Driving Perception under Challenging Weather

DGX agent

arXiv:2605.01081v1 Announce Type: new Abstract: The performance of state-of-the-art object detectors degrades significantly under adverse weather, causing a safety-critical domain shift problem for au

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

DGX agent

arXiv:2605.01018v1 Announce Type: new Abstract: Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

X2SAM: Any Segmentation in Images and Videos

DGX agent

arXiv:2605.00891v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception acros

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Zero-Shot Interpretable Image Steganalysis for Invertible Image Hiding

DGX agent

arXiv:2605.01331v1 Announce Type: new Abstract: Image steganalysis, which aims at detecting secret information concealed within images, has become a critical countermeasure for assessing the security

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

ZNO: Stable Rational Neural Operators in the Z-Domain for Discrete-Time Dynamic

DGX agent

arXiv:2605.02356v1 Announce Type: new Abstract: We introduce the Z-Domain Neural Operator (ZNO), a causal neural operator whose layers are stable low-rank multiple-input multiple-output (MIMO) rationa

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

A Dirac-Frenkel-Onsager principle: Instantaneous residual minimization with gauge momentum for nonlinear parametrizations of PDE solutions

DGX agent

arXiv:2605.00284v1 Announce Type: new Abstract: Dirac-Frenkel instantaneous residual minimization evolves nonlinear parametrizations of PDE solutions in time, but ill-conditioning can render the param

model-releasesarxiv-cs-lg
4 May 2026
Model Releases

A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction

DGX agent

arXiv:2605.00551v1 Announce Type: new Abstract: AI agents that interact with graphical user interfaces (GUIs) require effective observation representations for reliable grounding. The accessibility tr

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Adaptive Dual-Teacher Distillation with Subnetwork Rectification for Bridging Semantic Gaps in Black-Box Domain Adaptation

DGX agent

arXiv:2603.22908v3 Announce Type: replace Abstract: Assuming that neither source data nor source model parameters are accessible, black-box domain adaptation (BBDA) represents a highly practical yet c

model-releasesarxiv-cs-cv
4 May 2026
Model Releases

Agent Factories for High Level Synthesis: How Far Can General-Purpose Coding Agents Go in Hardware Optimization?

DGX agent

arXiv:2603.25719v2 Announce Type: replace-cross Abstract: We present an empirical study of how far general-purpose coding agents -- without hardware-specific training -- can optimize hardware designs

model-releasesarxiv-cs-lg
4 May 2026
Model Releases

AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?

DGX agent

arXiv:2605.00334v1 Announce Type: cross Abstract: Production agentic systems make many model calls per user request, and most of those calls are short, structured, and routine. This raises a practical

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs

DGX agent

arXiv:2605.00539v1 Announce Type: new Abstract: Quantization is a key method for reducing the GPU memory requirement of training large language models (LLMs). Yet, current approaches are ineffective f

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Alethia: A Foundational Encoder for Voice Deepfakes

DGX agent

arXiv:2605.00251v1 Announce Type: cross Abstract: Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, dow

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

AlphaInventory: Evolving White-Box Inventory Policies via Large Language Models with Deployment Guarantees

DGX agent

arXiv:2605.00369v1 Announce Type: new Abstract: We study how large language models can be used to evolve inventory policies in online, non-stationary environments. Our work is motivated by recent adva

model-releasesarxiv-cs-lg
4 May 2026
Model Releases

BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social Interaction

DGX agent

arXiv:2603.15949v3 Announce Type: replace Abstract: Large Language Models have demonstrated strong multilingual fluency, yet fluency alone does not guarantee socially appropriate language use. In high

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

DGX agent

arXiv:2605.00674v1 Announce Type: new Abstract: Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating

model-releasesarxiv-cs-cl
4 May 2026
← Previous
1…282283284285286…361
Next →