AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

DGX agent

arXiv:2605.21605v1 Announce Type: new Abstract: Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal

safetyarxiv-cs-cv
22 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Applications

GenHAR: Generalizing Cross-domain Human Activity Recognition for Last-mile Delivery

DGX agent

arXiv:2605.22086v1 Announce Type: new Abstract: Human Activity Recognition (HAR) has shown remarkable effectiveness in various applications, such as smart healthcare and intelligent manufacturing. How

applicationsarxiv-cs-cv
22 May 2026
Model Releases

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

DGX agent

arXiv:2605.22558v1 Announce Type: new Abstract: Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearan

model-releasesarxiv-cs-cv
22 May 2026
Applications

GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations

DGX agent

arXiv:2605.22812v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, exi

applicationsarxiv-cs-cv
22 May 2026
Safety

GLeVE: Graph-Guided Lesion Grounding with Proposal Verification in 3D CT

DGX agent

arXiv:2605.22619v1 Announce Type: new Abstract: Grounding radiology report descriptions to 3D CT volumes is essential for verifiable clinical interpretation, yet remains challenging due to the semanti

safetyarxiv-cs-cv
22 May 2026
Research

Guided Trajectory Optimization with Sparse Scaling for Test-Time Diffusion

DGX agent

arXiv:2605.21907v1 Announce Type: new Abstract: The efficient Test-Time Scaling (TTS) paradigm offers a promising perspective for enhancing the generation performance of diffusion models. However, cur

researcharxiv-cs-cv
22 May 2026
Model Releases

H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning

DGX agent

arXiv:2605.22629v1 Announce Type: new Abstract: Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimate

model-releasesarxiv-cs-cv
22 May 2026
Safety

Hierarchical Variational Policies for Reward-Guided Diffusion

DGX agent

arXiv:2605.21661v1 Announce Type: cross Abstract: Adapting pretrained diffusion models to downstream objectives such as inverse problems often requires expensive test-time guidance or optimization. We

safetyarxiv-cs-cv
22 May 2026
Model Releases

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

DGX agent

arXiv:2602.01851v2 Announce Type: replace Abstract: Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In

model-releasesarxiv-cs-cv
22 May 2026
Applications

HyperBench: Standardizing and Scaling Synthetic Evaluation for Hyperspectral Super-Resolution

DGX agent

arXiv:2605.21671v1 Announce Type: cross Abstract: Hyperspectral super-resolution (HSR) reconstructs a high-spatial-resolution hyperspectral image by fusing a low-resolution hyperspectral image (LR-HSI

applicationsarxiv-cs-cv
22 May 2026
Research

Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors

DGX agent

arXiv:2605.22272v1 Announce Type: cross Abstract: Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising

researcharxiv-cs-cv
22 May 2026
Applications

Impact of Atmospheric Turbulence and Pointing Error on Earth Observation

DGX agent

arXiv:2605.22268v1 Announce Type: cross Abstract: Earth Observation (EO) imagery is often degraded by atmospheric turbulence and pointing jitter; yet, these effects are rarely considered in datasets u

applicationsarxiv-cs-cv
22 May 2026
Research

Improved DDIM Sampling with Moment Matching Gaussian Mixtures

DGX agent

arXiv:2311.04938v5 Announce Type: replace Abstract: We propose using a Gaussian Mixture Model (GMM) as reverse transition operator (kernel) within the Denoising Diffusion Implicit Models (DDIM) framew

researcharxiv-cs-cv
22 May 2026
Research

Improving 3D Labeling in Self-Driving by Inferring Vehicle Information using Vision Language Models

DGX agent

arXiv:2605.21747v1 Announce Type: new Abstract: We present an approach to improve 3D vehicle labeling in self-driving applications through zero-shot inference of vehicle information, leveraging Vehicl

researcharxiv-cs-cv
22 May 2026
Research

Improving Viewpoint-Invariance and Temporal Consistency for Action Detection

DGX agent

arXiv:2605.22695v1 Announce Type: new Abstract: Viewpoint change invariance and action temporal consistency are critical aspects for the effective deployment of human action detection of untrimmed vid

researcharxiv-cs-cv
22 May 2026
Model Releases

InfVSR: Breaking Length Limits of Generic Video Super-Resolution

DGX agent

arXiv:2510.00948v2 Announce Type: replace Abstract: Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent c

model-releasesarxiv-cs-cv
22 May 2026
Research

Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow

DGX agent

arXiv:2605.21980v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) represent a significant leap towards empathetic agents, demonstrating remarkable capabilities in emotion understand

researcharxiv-cs-cv
22 May 2026
Model Releases

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

DGX agent

arXiv:2605.22080v1 Announce Type: new Abstract: We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF material

model-releasesarxiv-cs-cv
22 May 2026
Research

Label tree semantic losses for rich multi-class medical image segmentation

DGX agent

arXiv:2507.15777v3 Announce Type: replace Abstract: Rich and accurate medical image segmentation is poised to underpin the next generation of AI-defined clinical practice by delineating critical anato

researcharxiv-cs-cv
22 May 2026
Safety

LACO: Adaptive Latent Communication for Collaborative Driving

DGX agent

arXiv:2605.22504v1 Announce Type: cross Abstract: Collaborative driving aims to improve safety and efficiency by enabling connected vehicles to coordinate under partial observability. Recent approache

safetyarxiv-cs-cv
22 May 2026
Model Releases

Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models

DGX agent

arXiv:2605.21861v1 Announce Type: new Abstract: Multi-modality medical vision (MV) foundation models (FM) are fundamentally challenged by pronounced Non-IID feature statistics across heterogeneous ima

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

DGX agent

arXiv:2605.21988v1 Announce Type: new Abstract: Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

DGX agent

arXiv:2605.21573v1 Announce Type: new Abstract: We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with

model-releasesarxiv-cs-cv
22 May 2026
Research

LFX: Towards Unified Light Field Dense Semantic Segmentation and Salient Object Detection

DGX agent

arXiv:2503.00747v2 Announce Type: replace Abstract: Light field cameras capture multi-view observations within a single exposure. However, existing studies are typically tailored to specific LF repres

researcharxiv-cs-cv
22 May 2026
Model Releases

LongVT: Incentivizing 'Thinking with Long Videos' via Native Tool Calling

DGX agent

arXiv:2511.20785v3 Announce Type: replace Abstract: Large multimodal models (LMMs) have shown great potential for video reasoning with textual Chain-of-Thought. However, they remain vulnerable to hall

model-releasesarxiv-cs-cv
22 May 2026
Local Ai

Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming

DGX agent

arXiv:2605.21652v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have significantly advanced medical visual question answering, yet their performance in ultrasound remains suboptimal. In

local-aiarxiv-cs-cv
22 May 2026
Model Releases

LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

DGX agent

arXiv:2605.22089v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sp

model-releasesarxiv-cs-cv
22 May 2026
Tutorials

MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement

DGX agent

arXiv:2602.01760v2 Announce Type: replace Abstract: This paper focuses on a highly practical scenario: how to continue benefiting from the advantages of multi-modal image fusion under harsh conditions

tutorialsarxiv-cs-cv
22 May 2026
Safety

Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low Light

DGX agent

arXiv:2605.22455v1 Announce Type: new Abstract: Real-world deployment of AI vision models is both fueled and limited by the data available for training and testing. Real datasets are sparse and uneven

safetyarxiv-cs-cv
22 May 2026
Safety

Mapping Tomato Cropping Systems in California Using AlphaEarth Geospatial Embeddings and Deep Learning Analysis

DGX agent

arXiv:2605.21804v1 Announce Type: cross Abstract: Field-scale crop maps support supply-chain forecasting and policy, yet statewide crop identification still often depends on retrospective surveys or r

safetyarxiv-cs-cv
22 May 2026
Model Releases

MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation

DGX agent

arXiv:2605.22469v1 Announce Type: new Abstract: Evaluating single-concept personalization in text-to-image diffusion requires measuring both concept preservation, which captures identity fidelity to a

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Matching with Deliberation: Test-Time Evolutionary Hierarchical Multi-Agents for Zero-Shot Compositional Image Retrieval

DGX agent

arXiv:2605.22478v1 Announce Type: new Abstract: Zero-Shot Compositional Image Retrieval (ZS-CIR) requires both preserving the visual continuity of the reference image and faithfully executing the sema

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

DGX agent

arXiv:2605.21917v1 Announce Type: new Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

DGX agent

arXiv:2605.21954v1 Announce Type: new Abstract: Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal la

model-releasesarxiv-cs-cv
22 May 2026
Applications

Moment-Reenacting: Inverse Motion Degradation with Cross-shutter Guidance

DGX agent

arXiv:2605.22423v1 Announce Type: new Abstract: Motion degradation, manifested as blur in global shutter (GS) images or rolling shutter (RS) distortion in RS counterparts, remains a fundamental challe

applicationsarxiv-cs-cv
22 May 2026
Model Releases

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

DGX agent

arXiv:2605.22818v1 Announce Type: new Abstract: Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally inco

model-releasesarxiv-cs-cv
22 May 2026
Research

MotionDPS: Motion-Compensated 3D Brain MRI Reconstruction

DGX agent

arXiv:2605.22121v1 Announce Type: new Abstract: Magnetic resonance imaging (MRI) is highly susceptible to patient motion due to its relatively long acquisition times and the fact that data are acquire

researcharxiv-cs-cv
22 May 2026
Model Releases

MOTOR: A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding

DGX agent

arXiv:2605.22550v1 Announce Type: new Abstract: Two-wheelers account for a disproportionately high share of road fatalities in the Global South. Research on two-wheeler rider behavior, however, lags f

model-releasesarxiv-cs-cv
22 May 2026
Research

MRecover: A Conditional Generative Model for Recovering Motion-Corrupted MR images Using AI Generated Contrast

DGX agent

arXiv:2605.21669v1 Announce Type: new Abstract: Hippocampal subfield segmentation requires high-resolution T2w turbo spin echo (TSE) MRI, yet this sequence is susceptible to motion artifacts, leading

researcharxiv-cs-cv
22 May 2026
Research

MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering

DGX agent

arXiv:2605.22269v1 Announce Type: new Abstract: Long streaming video QA remains challenging due to growing visual tokens and limited reasoning length of large language models (LLMs). KV-caching stores

researcharxiv-cs-cv
22 May 2026
Research

Multi-scale interaction network for stereo image super-resolution

DGX agent

arXiv:2605.21913v1 Announce Type: new Abstract: Stereo image super-resolution aims to generate high-resolution images by leveraging complementary information from binocular systems. Although previous

researcharxiv-cs-cv
22 May 2026
Research

No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos

DGX agent

arXiv:2605.22190v1 Announce Type: new Abstract: Recent feed-forward 3D gaussian splatting methods have made dramatic progress on individual aspects of 3D scene reconstruction, but no existing method j

researcharxiv-cs-cv
22 May 2026
Model Releases

Not All Starting Points Are Equal: Pre-trained Priors and Their Outsized Impact on Person Identification

DGX agent

arXiv:2507.17640v3 Announce Type: replace Abstract: Recent years have seen an explosion of diverse general purpose pre-training methodologies for computer vision. However, the impact that these pre-tr

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems

DGX agent

arXiv:2605.22144v1 Announce Type: new Abstract: Existing approaches for digital short-drama production typically rely on one-shot LLM generated scripts and loosely coupled pipelines, which fail to sat

model-releasesarxiv-cs-cv
22 May 2026
Agents

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization

DGX agent

arXiv:2605.22104v1 Announce Type: new Abstract: Real-world image restoration is challenging due to complex and interacting mixed degradations. Recent agent-based approaches address this problem by com

agentsarxiv-cs-cv
22 May 2026
Hardware

ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration

DGX agent

arXiv:2605.22015v1 Announce Type: new Abstract: Diffusion Transformer (DiT) has emerged as a powerful model architecture for generating high-quality images and videos. In the case of video DiT, 3D Spa

hardwarearxiv-cs-cv
22 May 2026
Model Releases

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

DGX agent

arXiv:2605.22200v1 Announce Type: new Abstract: Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment ho

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

PartCo: Part-Level Correspondence Priors Enhance Category Discovery

DGX agent

arXiv:2509.22769v2 Announce Type: replace Abstract: Generalized Category Discovery (GCD) aims to identify both known and novel categories within unlabeled data by leveraging a set of labeled examples

model-releasesarxiv-cs-cv
22 May 2026
← Previous
1…147148149150151…263
Next →