AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models

DGX agent

arXiv:2605.28051v1 Announce Type: new Abstract: Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. Existing methods typically rel

researcharxiv-cs-cv
28 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Bias Leaves a Gradient Trail: Label-Free Bias Identification via Gradient Probes on Concept Decompositions

DGX agent

arXiv:2605.28780v1 Announce Type: new Abstract: Vision classifiers can exploit spurious correlations, achieving high in-distribution accuracy yet failing under distribution shift. Existing approaches

model-releasesarxiv-cs-cv
28 May 2026
Research

Bound-Constrained Sparse Representation for Electrical Impedance Tomography

DGX agent

arXiv:2605.28392v1 Announce Type: new Abstract: This study proposes a bound-constrained sparse representation (BC-SR) framework for electrical impedance tomography (EIT), aimed at improving conductivi

researcharxiv-cs-cv
28 May 2026
Model Releases

Bounded-Compute Multimodal Regression for Product-Rating Prediction

DGX agent

arXiv:2605.27737v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generatio

model-releasesarxiv-cs-cv
28 May 2026
Research

Bridging the Generalization Gap in Adverse Weather Segmentation: A Training Recipe Perspective

DGX agent

arXiv:2605.27962v1 Announce Type: new Abstract: This paper describes our approach for the 8th UG2+ Workshop (CVPR 2026) Track~2, which targets semantic segmentation of outdoor scenes degraded by five

researcharxiv-cs-cv
28 May 2026
Research

Bridging the Sampling Distribution Shift in Radio Map Estimation: A Trajectory-Aware Paradigm

DGX agent

arXiv:2605.28234v1 Announce Type: new Abstract: Learning-based radio map estimation (RME) plays a critical role in UAV-assisted wireless sensing, enabling tasks such as coverage prediction and network

researcharxiv-cs-cv
28 May 2026
Model Releases

Category-Level 3D Correspondence in Camera Space via Morphable Object Priors

DGX agent

arXiv:2605.28257v1 Announce Type: new Abstract: Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estim

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Chirpy3D: Part-Aware Multi-View Diffusion for Creative Fine-Grained Object Generation

DGX agent

arXiv:2501.04144v3 Announce Type: replace Abstract: Understanding and generating the fine-grained structure of objects -- such as birds with species-specific beaks, wings, and tails -- is a long-stand

model-releasesarxiv-cs-cv
28 May 2026
Tutorials

CLEAR-NeRF: Collinearity and Local-region Enhanced Accurate 3D Reconstruction in Unbounded Scenes

DGX agent

arXiv:2605.28125v1 Announce Type: new Abstract: Many real-world 3D reconstruction applications demand photorealism and metric accuracy across unbounded, complex scenes with challenging lighting and im

tutorialsarxiv-cs-cv
28 May 2026
Research

ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation

DGX agent

arXiv:2605.27852v1 Announce Type: cross Abstract: Unified and scalable Transformers have recently achieved remarkable success in modeling diverse phenomena traditionally associated with computer graph

researcharxiv-cs-cv
28 May 2026
Model Releases

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

DGX agent

arXiv:2605.28056v1 Announce Type: new Abstract: Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

DGX agent

arXiv:2508.21046v3 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high com

model-releasesarxiv-cs-cv
28 May 2026
Safety

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization

DGX agent

arXiv:2605.28615v1 Announce Type: new Abstract: Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bi

safetyarxiv-cs-cv
28 May 2026
Safety

Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry

DGX agent

arXiv:2605.27952v1 Announce Type: new Abstract: Visual odometry (VO) is a fundamental component in robotics and augmented reality. RGB-D direct VO benefits from metric depth measurements, but it can d

safetyarxiv-cs-cv
28 May 2026
Safety

CPPO: Contrastive Perception Policy Optimization for VLM Agents

DGX agent

arXiv:2601.00501v2 Announce Type: replace Abstract: We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core

safetyarxiv-cs-cv
28 May 2026
Research

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies

DGX agent

arXiv:2605.24302v2 Announce Type: replace Abstract: Egocentric action recognition is a challenging task due to erratic camera motion, frequent hand occlusion, and the difficulty of maintaining consist

researcharxiv-cs-cv
28 May 2026
Research

CuriosAI Submission to the CASTLE Challenge at EgoVis 2026

DGX agent

arXiv:2605.27800v1 Announce Type: new Abstract: CASTLE 2026 asks 185 multiple-choice questions over 600+ hours of synchronized multi-view egocentric video. We explore two approaches on top of a shared

researcharxiv-cs-cv
28 May 2026
Tutorials

D^2Turb: Depth-Aware Simulation and Decoupled Learning for Single-Frame Atmospheric Turbulence Mitigation

DGX agent

arXiv:2605.27460v1 Announce Type: new Abstract: Single-frame atmospheric turbulence mitigation is inherently ill-posed due to spatially varying blur coupled with non-rigid geometric distortion. Existi

tutorialsarxiv-cs-cv
28 May 2026
Safety

DebFilter: Eradicating Biases Stashed in Value

DGX agent

arXiv:2605.28167v1 Announce Type: new Abstract: Text-to-image diffusion models, which are theoretically equivalent to score-based generative models, generate images through a multi-step denoising proc

safetyarxiv-cs-cv
28 May 2026
Model Releases

Decoupled Training with Local Reinforcement Fine-Tuning in Federated Learning

DGX agent

arXiv:2605.27900v1 Announce Type: new Abstract: Federated Learning (FL) with pre-trained Vision-Language Models (VLMs) has emerged as a promising paradigm for various downstream tasks. By leveraging i

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation

DGX agent

arXiv:2605.28587v1 Announce Type: new Abstract: Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. Howeve

model-releasesarxiv-cs-cv
28 May 2026
Safety

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

DGX agent

arXiv:2605.28491v1 Announce Type: new Abstract: We study real-time audio-responsive character control as a deployment-faithful problem: strictly causal, bounded-latency streaming that must generate co

safetyarxiv-cs-cv
28 May 2026
Research

DODO: Discrete OCR Diffusion Models

DGX agent

arXiv:2602.16872v2 Announce Type: replace Abstract: Optical Character Recognition (OCR) is a fundamental task for digitizing information, serving as a critical bridge between visual data and textual u

researcharxiv-cs-cv
28 May 2026
Model Releases

DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving

DGX agent

arXiv:2605.28544v1 Announce Type: new Abstract: Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primaril

model-releasesarxiv-cs-cv
28 May 2026
Local Ai

Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking

DGX agent

arXiv:2605.28018v1 Announce Type: new Abstract: Given the real-time demands of UAV tracking, many methods simplify the backbone to reduce computation, but this often weakens feature representation and

local-aiarxiv-cs-cv
28 May 2026
Research

EchoAvatar: Real-time Generative Avatar Animation from Audio Streams

DGX agent

arXiv:2605.28272v1 Announce Type: new Abstract: Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistant

researcharxiv-cs-cv
28 May 2026
Research

EgoRelight: Egocentric Human Capture and Illumination Recovery for Relightable and Photoreal Avatar Rendering

DGX agent

arXiv:2605.28401v1 Announce Type: new Abstract: Mixed Reality (MR) headsets promise a future of immersive telepresence where virtual humans blend indistinguishably into real or virtual surroundings. A

researcharxiv-cs-cv
28 May 2026
Model Releases

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning

DGX agent

arXiv:2510.27266v2 Announce Type: replace Abstract: Autonomous graphical user interface (GUI) agents rely on accurate GUI grounding, which maps language instructions to on-screen coordinates, to execu

model-releasesarxiv-cs-cv
28 May 2026
Research

Enhancing Ultra-low-field MRI with Segmentation-guided Adversarial Learning

DGX agent

arXiv:2605.28016v1 Announce Type: new Abstract: Ultra-low-field (ULF) MRI offers portable and low-cost imaging but suffers from poor image quality. To address this, we present our submission to the 20

researcharxiv-cs-cv
28 May 2026
Safety

EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection

DGX agent

arXiv:2605.28630v1 Announce Type: new Abstract: Zero-Shot Anomaly Detection (ZSAD) aims to detect anomalies in unseen domains without target-domain adaptation. Recent CLIP-based methods have shown pro

safetyarxiv-cs-cv
28 May 2026
Research

Evaluating the Feasibility of Inferring Dietary Behavior Change Receptivity from Egocentric Images of Eating Environment

DGX agent

arXiv:2605.27950v1 Announce Type: new Abstract: Accurately assessing dietary behavior change receptivity is essential for designing effective just-in-time adaptive interventions (JITAIs) that promote

researcharxiv-cs-cv
28 May 2026
Research

Event-based Motion & Appearance Fusion for 6D Object Pose Tracking

DGX agent

arXiv:2603.08264v2 Announce Type: replace Abstract: Object pose tracking is a fundamental and essential task for robotics to perform tasks in the home and industrial settings. The most commonly used s

researcharxiv-cs-cv
28 May 2026
Model Releases

EventShiftFlow: Towards Hardware-efficient FPGA-based Flow Estimation

DGX agent

arXiv:2605.28312v1 Announce Type: cross Abstract: Event-based vision sensors offer asynchronous, high-temporal-resolution measurements that are attractive for low-latency robotic perception, but many

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Every9D-21M: Large-Scale Real-World 9D Canonicalization of Everyday Objects

DGX agent

arXiv:2605.28270v1 Announce Type: new Abstract: Estimating the 9D pose of everyday objects from a single real-world image remains challenging. This is largely due to the lack of large-scale supervisio

model-releasesarxiv-cs-cv
28 May 2026
Research

Explaining Digital Pathology Models via Clustering Activations

DGX agent

arXiv:2511.14558v2 Announce Type: replace Abstract: We present a clustering-based explainability technique for digital pathology models based on convolutional neural networks. Unlike commonly used met

researcharxiv-cs-cv
28 May 2026
Research

Explicit Critic Guidance for Aligning Diffusion Models

DGX agent

arXiv:2605.27736v1 Announce Type: cross Abstract: Online reinforcement learning is becoming increasingly important for aligning diffusion models with non-differentiable objectives. However, existing m

researcharxiv-cs-cv
28 May 2026
Research

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models

DGX agent

arXiv:2512.16483v2 Announce Type: replace Abstract: Visual Autoregressive (VAR) modeling departs from the next-token prediction paradigm of traditional Autoregressive (AR) models through next-scale pr

researcharxiv-cs-cv
28 May 2026
Agents

Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent

DGX agent

arXiv:2605.27452v1 Announce Type: new Abstract: Bridge inspection in Japan requires mandatory visual assessments every five years, yet qualitative damage ratings (levels a-e) assigned by different eng

agentsarxiv-cs-cv
28 May 2026
Model Releases

ForestHG-Trace: Traceable Long-Horizon Ecological Reasoning over Large-Scale Forest Scenes

DGX agent

arXiv:2605.27590v1 Announce Type: new Abstract: Remote sensing question answering (RS-QA) often requires more than direct semantic prediction, especially in large-scale forest scenes where ecological

model-releasesarxiv-cs-cv
28 May 2026
Safety

From Affect to Complex Behavior: Advancing Multimodal Human-Centered AI at the 10th ABAW Workshop & Competition

DGX agent

arXiv:2605.27451v1 Announce Type: new Abstract: The 10th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop and Competition, held at CVPR 2026, continues to advance research on modelling, analy

safetyarxiv-cs-cv
28 May 2026
Research

From Kellgren-Lawrence to Calcium Pyrophosphate Crystal Deposition: A Soft-Labelling Framework for Knee Osteoarthritis Assessmen

DGX agent

arXiv:2605.28176v1 Announce Type: new Abstract: Background and objective. Conventional Deep Learning (DL) approaches for Knee Osteoarthritis (KOA) grading rely on one-hot labels, which fail to capture

researcharxiv-cs-cv
28 May 2026
Safety

From Pixels to Words -- Towards Native One-Vision Models at Scale

DGX agent

arXiv:2605.28820v1 Announce Type: new Abstract: Current vision-language models (VLMs) typically stitch together separate image encoders and language decoders via multi-stage alignment, a modular frame

safetyarxiv-cs-cv
28 May 2026
Model Releases

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

DGX agent

arXiv:2605.28816v1 Announce Type: new Abstract: World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single contr

model-releasesarxiv-cs-cv
28 May 2026
Applications

GEM: Generative Supervision Helps Embodied Intelligence

DGX agent

arXiv:2605.28548v1 Announce Type: new Abstract: Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Acti

applicationsarxiv-cs-cv
28 May 2026
Local Ai

HarmoVid: Relightful Video Portrait Harmonization

DGX agent

arXiv:2605.28811v1 Announce Type: new Abstract: We present a method for harmonizing the lighting of a foreground video to match a target background scene, adjusting shadows, color tone, and illuminati

local-aiarxiv-cs-cv
28 May 2026
Tutorials

Hierarchical Relation-augmented Representation Generalization for Few-shot Action Recognition

DGX agent

arXiv:2504.10079v4 Announce Type: replace Abstract: Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level repres

tutorialsarxiv-cs-cv
28 May 2026
Safety

HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment

DGX agent

arXiv:2508.15130v2 Announce Type: replace Abstract: Despite significant progress in no-reference image quality assessment (NR-IQA), dataset biases and reliance on subjective labels continue to hinder

safetyarxiv-cs-cv
28 May 2026
Model Releases

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction

DGX agent

arXiv:2510.06928v2 Announce Type: replace Abstract: Autoregressive models have emerged as a powerful paradigm for visual content creation, but often overlook the intrinsic structural properties of vis

model-releasesarxiv-cs-cv
28 May 2026
← Previous
1…138139140141142…263
Next →