AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
60,292 results
Safety

VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation

DGX agent

arXiv:2601.23286v2 Announce Type: replace Abstract: While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, o

safetyarxiv-cs-cv
5 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition

DGX agent

arXiv:2605.02834v1 Announce Type: new Abstract: Videos are unique in their ability to capture actions which transcend multiple frames. Accordingly, for many years action recognition was the quintessen

model-releasesarxiv-cs-cv
5 May 2026
Research

ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking

DGX agent

arXiv:2605.02638v1 Announce Type: new Abstract: Cross-view Referring Multi-Object Tracking (CRMOT) aims to track multiple objects specified by natural language across multiple camera views, with globa

researcharxiv-cs-cv
5 May 2026
Model Releases

VILAS: A VLA-Integrated Low-cost Architecture with Soft Grasping for Robotic Manipulation

DGX agent

arXiv:2605.02037v1 Announce Type: new Abstract: We present VILAS, a fully low-cost, modular robotic manipulation platform designed to support end-to-end vision-language-action (VLA) policy learning an

model-releasesarxiv-cs-ro
5 May 2026
Hardware

ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA

DGX agent

arXiv:2605.01935v1 Announce Type: cross Abstract: Vision Mamba (ViM) models offer a compelling efficiency advantage over Transformers by leveraging the linear complexity of State Space Models (SSMs),

hardwarearxiv-cs-cv
5 May 2026
Research

Virtual Scanning for NSCLC Histology: Investigating the Discriminatory Power of Synthetic PET

DGX agent

arXiv:2605.02746v1 Announce Type: new Abstract: Accurate histological differentiation between adenocarcinoma (ADC) and squamous cell carcinoma (SCC) is critical for personalized treatment in non-small

researcharxiv-cs-cv
5 May 2026
Safety

Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy

DGX agent

arXiv:2605.01101v1 Announce Type: cross Abstract: This paper develops Virtual Speech Therapist (VST), an intelligent agent-based platform that streamlines stuttering assessment and delivers customized

safetyarxiv-cs-cl
5 May 2026
Safety

Visibility-Aware Mobile Grasping in Dynamic Environments

DGX agent

arXiv:2605.02487v1 Announce Type: new Abstract: This paper addresses the problem of mobile grasping in dynamic, unknown environments where a robot must operate under a limited field-of-view. The funda

safetyarxiv-cs-ro
5 May 2026
Model Releases

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

DGX agent

arXiv:2605.01391v1 Announce Type: new Abstract: Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute

model-releasesarxiv-cs-cv
5 May 2026
Research

Visual Chart Representations for Cryptocurrency Regime Prediction: A Systematic Deep Learning Study

DGX agent

arXiv:2605.00875v1 Announce Type: new Abstract: Technical traders have long relied on visual analysis of candlestick charts to identify market patterns and predict price movements. While deep learning

researcharxiv-cs-cv
5 May 2026
Model Releases

Visual Implicit Autoregressive Modeling

DGX agent

arXiv:2605.01220v1 Announce Type: new Abstract: Visual Autoregressive Modeling (VAR) based on next-scale prediction achieves strong generation quality, but their explicit deep stacks fix the amount of

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

DGX agent

arXiv:2605.02735v1 Announce Type: new Abstract: Continuous latent-space reasoning offers a compact alternative to textual chain-of-thought for multimodal models, enabling high-dimensional visual evide

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Visualizing Critic Match Loss Landscapes for Interpretation of Online Reinforcement Learning Control Algorithms

DGX agent

arXiv:2603.14535v2 Announce Type: replace Abstract: Reinforcement learning has proven its power on various occasions. However, its performance is not always guaranteed when system dynamics change. Ins

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model

DGX agent

arXiv:2605.01194v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities and generalization in embodied manipulation. However, their decision-makin

model-releasesarxiv-cs-ro
5 May 2026
Safety

VOFA: Visual Object Goal Pushing with Force-Adaptive Control for Humanoids

DGX agent

arXiv:2605.01518v1 Announce Type: new Abstract: The ability to push large objects in a goal-directed manner using onboard egocentric perception is an essential skill for humanoid robots to perform com

safetyarxiv-cs-ro
5 May 2026
Research

VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection

DGX agent

arXiv:2605.01365v1 Announce Type: new Abstract: Open-vocabulary 3D affordance detection requires localizing interaction regions on point clouds given novel affordance descriptions. Recent methods exte

researcharxiv-cs-cv
5 May 2026
Research

VRGaussianAvatar: Integrating 3D Gaussian Avatars into VR

DGX agent

arXiv:2602.01674v2 Announce Type: replace Abstract: We present VRGaussianAvatar, an integrated system that enables real-time full-body 3D Gaussian Splatting (3DGS) avatars in virtual reality using onl

researcharxiv-cs-cv
5 May 2026
Research

Watch Your Step: Information Injection in Diffusion Models via Shadow Timestep Embedding

DGX agent

arXiv:2605.00935v1 Announce Type: cross Abstract: Diffusion models have become the foundation of modern generative systems, with most research focusing primarily on improving generation efficiency and

researcharxiv-cs-cv
5 May 2026
Model Releases

Watermarking LLM Agent Trajectories

DGX agent

arXiv:2602.18700v2 Announce Type: replace-cross Abstract: LLM agents rely heavily on high-quality trajectory data to guide their problem-solving behaviors, yet producing such data requires substantial

model-releasesarxiv-cs-cl
5 May 2026
Applications

Weight Clipping for Robust Conformal Inference under Unbounded Covariate Shifts

DGX agent

arXiv:2605.02072v1 Announce Type: new Abstract: Conformal prediction (CP) provides powerful, distribution-free prediction sets, but its guarantees rely on the exchangeability of training and test data

applicationsarxiv-cs-lg
5 May 2026
Research

What Does a Meow Mean? In Search of Intuitively Understandable Communication by a Nonverbal Companion Robot

DGX agent

arXiv:2605.01251v1 Announce Type: cross Abstract: Older adults living alone have a number of challenges, and robots can help with some of them--by providing reminders, initiating activity, or offering

researcharxiv-cs-ro
5 May 2026
Applications

What price to pay? Auto-tuning a building MPC controller for optimal economic cost

DGX agent

arXiv:2501.10859v2 Announce Type: replace-cross Abstract: Demand-side management (DSM) programs introduce complex pricing, requiring advanced control for cost minimization. Model Predictive Control (M

applicationsarxiv-cs-lg
5 May 2026
Model Releases

What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models

DGX agent

arXiv:2605.02038v1 Announce Type: new Abstract: Single-prompt accuracy is the dominant way to benchmark language models, but it can miss reliability failures that matter. We evaluate a 15-model open-w

model-releasesarxiv-cs-cl
5 May 2026
Safety

When Attention Collapses: Residual Evidence Modeling for Compositional Inference

DGX agent

arXiv:2605.02323v1 Announce Type: new Abstract: Compositional inference - the decomposition of observations into an unknown number of latent components - is central to perception and scientific data a

safetyarxiv-cs-lg
5 May 2026
Model Releases

When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition

DGX agent

arXiv:2605.02782v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models

DGX agent

arXiv:2605.02363v1 Announce Type: new Abstract: Deployed language models must produce outputs that are both correct and format-compliant. We study this structured-output reliability gap using two math

model-releasesarxiv-cs-cl
5 May 2026
Safety

When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.01133v1 Announce Type: cross Abstract: Large language model (LLM)-powered multi-agent systems (MAS) enable agents to communicate and share information, achieving strong performance on compl

safetyarxiv-cs-lg
5 May 2026
Model Releases

When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation

DGX agent

arXiv:2605.00911v1 Announce Type: new Abstract: Industrial Retrieval-Augmented Generation (RAG) systems depend on optical character recognition (OCR) to transform visual documents into text. Existing

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

DGX agent

arXiv:2601.19827v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative retrieval-re

model-releasesarxiv-cs-cl
5 May 2026
Research

When Less is Enough: Efficient Inference via Collaborative Reasoning

DGX agent

arXiv:2605.01111v1 Announce Type: cross Abstract: In this work, we introduce DUET (Dual-model Efficient Two-stage inference), a collaborative inference framework in which a capable model and a lightwe

researcharxiv-cs-cl
5 May 2026
Model Releases

When Less Is More: Simplicity Beats Complexity for Physics-Constrained InSAR Phase Unwrapping

DGX agent

arXiv:2605.00896v1 Announce Type: new Abstract: Operational phase unwrapping is the primary computational bottleneck in InSAR-based volcanic and seismic monitoring. We challenge the industry trend of

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System

DGX agent

arXiv:2602.06932v3 Announce Type: replace Abstract: Speculative decoding can significantly accelerate LLM serving, yet most deployments today disentangle speculator training from serving, treating spe

model-releasesarxiv-cs-lg
5 May 2026
Research

When To Adapt? Adapting the Model or Data in Federated Medical Imaging

DGX agent

arXiv:2605.00892v1 Announce Type: new Abstract: Federated learning enables collaborative model training across medical institutions without sharing raw data, but its performance is often limited by do

researcharxiv-cs-cv
5 May 2026
Research

Where Do Prompt Perturbations Break Generation? A Segment-Level View of Robustness in LoRA-Tuned Language Models

DGX agent

arXiv:2605.01605v1 Announce Type: new Abstract: Large language models are sensitive to minor prompt perturbations, yet existing robustness methods usually enforce consistency at the whole-sequence lev

researcharxiv-cs-cl
5 May 2026
Safety

Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework

DGX agent

arXiv:2605.01416v1 Announce Type: cross Abstract: The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy

safetyarxiv-cs-cl
5 May 2026
Model Releases

WILD SAM: A Simulated-and-Real Data Augmentation for Autonomous Driving Perception under Challenging Weather

DGX agent

arXiv:2605.01081v1 Announce Type: new Abstract: The performance of state-of-the-art object detectors degrades significantly under adverse weather, causing a safety-critical domain shift problem for au

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

DGX agent

arXiv:2605.01018v1 Announce Type: new Abstract: Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its

model-releasesarxiv-cs-cv
5 May 2026
Hardware

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization

DGX agent

arXiv:2605.02262v1 Announce Type: cross Abstract: Recently, video language models (VLMs) have been applied in various fields. However, the visual token sequence of the VLM is too long, which may cause

hardwarearxiv-cs-cl
5 May 2026
Model Releases

X2SAM: Any Segmentation in Images and Videos

DGX agent

arXiv:2605.00891v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception acros

model-releasesarxiv-cs-cv
5 May 2026
Tutorials

XekRung Technical Report

DGX agent

arXiv:2605.00072v1 Announce Type: cross Abstract: We present XekRung, a frontier large language model for cybersecurity, designed to provide comprehensive security capabilities. To achieve this, we de

tutorialsarxiv-cs-ai
5 May 2026
Safety

Zero-Shot Adaptation of Behavioral Foundation Models to Unseen Dynamics

DGX agent

arXiv:2505.13150v2 Announce Type: replace Abstract: Behavioral Foundation Models (BFMs) proved successful in producing policies for arbitrary tasks in a zero-shot manner, requiring no test-time traini

safetyarxiv-cs-lg
5 May 2026
Local Ai

Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training

DGX agent

arXiv:2605.02241v1 Announce Type: cross Abstract: How reliably can a small language model estimate its own correctness? The answer determines whether local-to-cloud routing-escalating queries a cheap

local-aiarxiv-cs-cl
5 May 2026
Model Releases

Zero-Shot Interpretable Image Steganalysis for Invertible Image Hiding

DGX agent

arXiv:2605.01331v1 Announce Type: new Abstract: Image steganalysis, which aims at detecting secret information concealed within images, has become a critical countermeasure for assessing the security

model-releasesarxiv-cs-cv
5 May 2026
Safety

Zero-Shot, Safe and Time-Efficient UAV Navigation via Potential-Based Reward Shaping, Control Lyapunov and Barrier Functions

DGX agent

arXiv:2605.01787v1 Announce Type: cross Abstract: Autonomous navigation and obstacle avoidance remain a core challenge of modern Unmanned Aerial Vehicles (UAVs). While traditional control methods stru

safetyarxiv-cs-lg
5 May 2026
Model Releases

ZNO: Stable Rational Neural Operators in the Z-Domain for Discrete-Time Dynamic

DGX agent

arXiv:2605.02356v1 Announce Type: new Abstract: We introduce the Z-Domain Neural Operator (ZNO), a causal neural operator whose layers are stable low-rank multiple-input multiple-output (MIMO) rationa

model-releasesarxiv-cs-lg
5 May 2026
Research

2D-SuGaR: Surface-Aware Gaussian Splatting for Geometrically Accurate Mesh Reconstruction

DGX agent

arXiv:2605.00569v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful technique for generating photorealistic renderings of a scene in real-time. However, the volumetr

researcharxiv-cs-cv
4 May 2026
Local Ai

A Comparative Analysis of Machine Learning Models for Intrusion Detection in Intelligent Transport Systems

DGX agent

arXiv:2605.00279v1 Announce Type: cross Abstract: AI-powered edge computing security is moving Intelligent Transportation Systems (ITS) from passive, rule-based protections to proactive, smart, zero-t

local-aiarxiv-cs-lg
4 May 2026
Research

A Comparative Study of QSPR Methods on a Unique Multitask PAMPA dataset

DGX agent

arXiv:2605.00508v1 Announce Type: new Abstract: We present a unique, multitask dataset comprising 143 drug and drug candidate molecules, each evaluated on in vitro, parallel artificial-membrane permea

researcharxiv-cs-lg
4 May 2026
← Previous
1…994995996997998…1257
Next →