AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Masked Diffusion Vision-Language Models for Temporal Action Localization

DGX agent

arXiv:2605.29858v1 Announce Type: new Abstract: Temporal action localization (TAL) requires recognizing the target event and localizing its start and end times precisely in untrimmed videos. Recent vi

researcharxiv-cs-cv
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species

DGX agent

arXiv:2601.03729v2 Announce Type: replace Abstract: Fine-grained recognition of marine organisms is important for ecological research, biodiversity monitoring, habitat conservation, and evidence-based

safetyarxiv-cs-cv
29 May 2026
Research

Mesh-Aware Epipolar Matching for Multi-View Multi-Person 3D Pose Estimation in Basketball

DGX agent

arXiv:2605.29953v1 Announce Type: new Abstract: Multi-view multi-person 3D pose estimation in team sports scenarios remains challenging due to player occlusions, appearance similarity caused by team u

researcharxiv-cs-cv
29 May 2026
Safety

MetaRanker: Human-in-the-loop Active Ranking for Metalens Image Quality

DGX agent

arXiv:2605.29212v1 Announce Type: new Abstract: Image quality in modern imaging systems emerges from the coupled effects of the sensor, optics, and computational reconstruction. Ultra-thin metalenses

safetyarxiv-cs-cv
29 May 2026
Research

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

DGX agent

arXiv:2605.30263v1 Announce Type: new Abstract: Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive

researcharxiv-cs-cv
29 May 2026
Safety

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning

DGX agent

arXiv:2605.29577v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a promising framework that unifies perception, reasoning, and control for robot manipulation by adap

safetyarxiv-cs-cv
29 May 2026
Safety

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds

DGX agent

arXiv:2510.27391v2 Announce Type: replace Abstract: Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods e

safetyarxiv-cs-cv
29 May 2026
Safety

MonoPhysics: Estimating Geometry, Appearance, and Physical Parameters from Monocular Videos

DGX agent

arXiv:2605.30320v1 Announce Type: new Abstract: Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D struc

safetyarxiv-cs-cv
29 May 2026
Applications

Motion-guided sparse correction enables expert-quality point tracking across diverse microscopy regimes

DGX agent

arXiv:2605.29220v1 Announce Type: new Abstract: Tracking the dynamics of non-canonical biological systems in microscopy videos remains a persistent challenge. Both classical and learning-based tracker

applicationsarxiv-cs-cv
29 May 2026
Tutorials

Multi-level Collaborative Distillation Meets Global Workspace Model: A Unified Framework for OCIL

DGX agent

arXiv:2508.08677v2 Announce Type: replace-cross Abstract: Online Class-Incremental Learning (OCIL) enables models to learn continuously from non-i.i.d. data streams. Since samples of the data streams

tutorialsarxiv-cs-cv
29 May 2026
Local Ai

Multi-Scale Local Speculative Decoding for Image Generation

DGX agent

arXiv:2601.05149v2 Announce Type: replace Abstract: Autoregressive (AR) models have achieved remarkable success in image synthesis, yet their sequential nature imposes significant latency constraints.

local-aiarxiv-cs-cv
29 May 2026
Research

Multi-Stage VLM Pipeline for Zero-Shot Traffic Accident Understanding

DGX agent

arXiv:2605.29325v1 Announce Type: new Abstract: We present the 1st-place solution to the ACCIDENT challenge at the CVPR 2026 AUTOPILOT Workshop, which asks for zero-shot prediction of accident timing,

researcharxiv-cs-cv
29 May 2026
Model Releases

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models

DGX agent

arXiv:2602.14399v2 Announce Type: replace Abstract: Multi-turn jailbreak attacks have proven effective against text-only large language models (LLMs), where malicious content is gradually introduced t

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Multimodal LLMs See Sentiment

DGX agent

arXiv:2508.16873v3 Announce Type: replace Abstract: Understanding how visual content conveys sentiment is increasingly important in a digital landscape dominated by imagery. However, sentiment percept

model-releasesarxiv-cs-cv
29 May 2026
Safety

Native Audio-Visual Alignment for Generation

DGX agent

arXiv:2605.30073v1 Announce Type: new Abstract: Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual-acoustic content. However, existing open-source

safetyarxiv-cs-cv
29 May 2026
Tutorials

NeuROK: Generative 4D Neural Object Kinematics

DGX agent

arXiv:2605.30347v1 Announce Type: new Abstract: Data-driven approaches have revolutionized 3D vision, enabling transformers to effectively reconstruct and generate static 3D objects. However, generati

tutorialsarxiv-cs-cv
29 May 2026
Research

Non-Forgetting Knowledge Allocation with Bi-level Competition for Class-Incremental Learning

DGX agent

arXiv:2605.29592v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) with pre-trained models (PTMs) aims to sequentially adapt PTMs to new categories without forgetting old knowledge. Buil

researcharxiv-cs-cv
29 May 2026
Tutorials

Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval Using Language

DGX agent

arXiv:2605.29812v1 Announce Type: new Abstract: Video Moment Retrieval (VMR) targets to retrieve the specific moment corresponding to a sentence query from an untrimmed video. Although recent works ha

tutorialsarxiv-cs-cv
29 May 2026
Tutorials

OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the Wild

DGX agent

arXiv:2511.08423v3 Announce Type: replace Abstract: A truly universal AI-Generated Image (AIGI) detector must simultaneously generalize across diverse generative models and varied semantic content. Cu

tutorialsarxiv-cs-cv
29 May 2026
Research

OmniCD: A Foundational Framework for Remote Sensing Image Change Detection Guided by Multimodal Semantics

DGX agent

arXiv:2605.30168v1 Announce Type: new Abstract: Change detection (CD) in remote sensing is vital for applications such as urban monitoring and disaster assessment, yet traditional methods struggle wit

researcharxiv-cs-cv
29 May 2026
Research

One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation

DGX agent

arXiv:2605.29429v1 Announce Type: new Abstract: Cell instance segmentation models trained on cell-specific datasets suffer severe performance drops on out-of-distribution cell types, while interactive

researcharxiv-cs-cv
29 May 2026
Model Releases

Optimizing Latent Representations for Robust Building Damage Assessment Onboard Earth Observation Satellites

DGX agent

arXiv:2605.29575v1 Announce Type: new Abstract: Rapid identification of damaged buildings after natural disasters or on war areas is crucial to support emergency response and prioritize interventions.

model-releasesarxiv-cs-cv
29 May 2026
Safety

Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation

DGX agent

arXiv:2605.29390v1 Announce Type: new Abstract: Text-to-image (T2I) models have become increasingly capable of generating high-quality images. Yet, enforcing the explicit absence of a specified object

safetyarxiv-cs-cv
29 May 2026
Model Releases

Parameter-Efficient Subspace Decoupling ViT for Mitigating Multi-Task Negative Transfer in Histological Scoring

DGX agent

arXiv:2605.29852v1 Announce Type: new Abstract: Histological scoring is essential for diagnosing Non-Alcoholic Fatty Liver Disease (NAFLD), yet its automation remains challenging due to the high annot

model-releasesarxiv-cs-cv
29 May 2026
Research

ParCo-SDF: Learning Prior-Free Partial-to-Complete Signed Distance Fields of Deformable Objects

DGX agent

arXiv:2605.29417v1 Announce Type: new Abstract: This study addresses the partial-to-complete geometry reconstruction of deformable objects (DOs) from point-cloud observations toward precise DO manipul

researcharxiv-cs-cv
29 May 2026
Local Ai

RadioFormer3D: Weakly Supervised 3D Radio Map Estimation in Low-Altitude Airspace via Generative Modeling

DGX agent

arXiv:2605.29538v1 Announce Type: new Abstract: With the emergence of wireless applications in three-dimensional environments, such as the low-altitude airspace and 3D heterogeneous networks, radio ma

local-aiarxiv-cs-cv
29 May 2026
Model Releases

ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation

DGX agent

arXiv:2605.29579v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinat

model-releasesarxiv-cs-cv
29 May 2026
Tutorials

Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations

DGX agent

arXiv:2602.01456v2 Announce Type: replace-cross Abstract: Joint-Embedding Predictive Architectures (JEPA) learn view-invariant representations and admit projection-based distribution matching for coll

tutorialsarxiv-cs-cv
29 May 2026
Research

Reducing Experimental Testing in Space Propulsion Film Cooling Analyses by Pixelwise Generative Image Interpolation

DGX agent

arXiv:2605.29911v1 Announce Type: cross Abstract: We propose a machine learning approach for image regression from sparse experimental measurements. We show the application of the proposed method on f

researcharxiv-cs-cv
29 May 2026
Safety

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching

DGX agent

arXiv:2605.26108v2 Announce Type: replace Abstract: Recent advances in few-step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains

safetyarxiv-cs-cv
29 May 2026
Safety

Resolution as a Direction: Vector-Panning Feature Alignment for Cross-Resolution Re-Identification

DGX agent

arXiv:2510.00936v2 Announce Type: replace Abstract: Cross-resolution person re-identification (CR-ReID) remains challenging in practical surveillance, where camera quality and capture distance lead to

safetyarxiv-cs-cv
29 May 2026
Safety

Resolving Endpoint Underfitting in Diffusion Bridges via Noise Alignment

DGX agent

arXiv:2605.28962v1 Announce Type: new Abstract: Diffusion bridge models offer a powerful framework for connecting two data distributions, such as in image restoration and translation. Many existing me

safetyarxiv-cs-cv
29 May 2026
Safety

REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image

DGX agent

arXiv:2605.30338v1 Announce Type: new Abstract: Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applic

safetyarxiv-cs-cv
29 May 2026
Model Releases

RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization

DGX agent

arXiv:2603.27758v2 Announce Type: replace Abstract: Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In

model-releasesarxiv-cs-cv
29 May 2026
Tutorials

Robust Cross-Domain Generalization Using Unlabeled Target Data with Source-Domain Supervision

DGX agent

arXiv:2605.29122v1 Announce Type: new Abstract: It is often desirable to generalize medical imaging AI models trained with dense annotations to data acquired from different ultrasound scanners or clin

tutorialsarxiv-cs-cv
29 May 2026
Research

S2MDF: A Plug-And-Play Layer for Intersection-Free Multi-Object Signed Distance Fields

DGX agent

arXiv:2605.29761v1 Announce Type: new Abstract: Compositional implicit surface representations model scenes as collections of objects, each encoded by a Signed Distance Field (SDF). A fundamental limi

researcharxiv-cs-cv
29 May 2026
Applications

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation

DGX agent

arXiv:2605.29662v1 Announce Type: new Abstract: Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for a

applicationsarxiv-cs-cv
29 May 2026
Model Releases

SAGE: Segment-Aware Gloss-Free Encoding for Token-Efficient Sign Language Translation

DGX agent

arXiv:2507.09266v2 Announce Type: replace Abstract: Gloss-free Sign Language Translation (SLT) has advanced rapidly, achieving strong performances without relying on gloss annotations. However, these

model-releasesarxiv-cs-cv
29 May 2026
Research

SalsaAgent: A multimodal embodied language model for interactive dance generation

DGX agent

arXiv:2605.29219v1 Announce Type: new Abstract: Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive

researcharxiv-cs-cv
29 May 2026
Applications

SAM3D-Phys: Towards Multi-Object Interactive Simulation in Real World

DGX agent

arXiv:2605.30239v1 Announce Type: new Abstract: This work addresses the problem of recovering complete, simulatable object geometry from reconstructed real-world scenes, enabling physics-based interac

applicationsarxiv-cs-cv
29 May 2026
Research

SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification

DGX agent

arXiv:2602.13600v2 Announce Type: replace Abstract: A line of recent training-free methods for mitigating hallucinations in large vision-language models (LVLMs) operates by amplifying attention to vis

researcharxiv-cs-cv
29 May 2026
Model Releases

SDF-Net: Structure-Aware Disentangled Feature Learning for Opticall-SAR Ship Re-identification

DGX agent

arXiv:2603.12588v2 Announce Type: replace Abstract: Cross-modal ship re-identification (ReID) between optical and synthetic aperture radar (SAR) imagery is fundamentally challenged by the severe radio

model-releasesarxiv-cs-cv
29 May 2026
Tutorials

Seeing through boxes: Non-Line-of-Sight 3D Reconstruction from Radar Signals

DGX agent

arXiv:2605.29098v1 Announce Type: new Abstract: Reconstructing object geometry from radio frequency (RF) signals is fundamentally challenging due to the lensless imaging nature of RF sensing, which le

tutorialsarxiv-cs-cv
29 May 2026
Safety

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

DGX agent

arXiv:2605.30116v1 Announce Type: new Abstract: Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models. However, DMD-style vid

safetyarxiv-cs-cv
29 May 2026
Model Releases

SLAD : Shared LoRA Adapters for Task Specific Distillation

DGX agent

arXiv:2605.29726v1 Announce Type: new Abstract: In the context of resource-constrained environments such as embedded systems, adapting reduced-size foundation models to downstream tasks has become inc

model-releasesarxiv-cs-cv
29 May 2026
Research

Soften the Mask: Adaptive Temporal Soft Mask for Efficient Dynamic Facial Expression Recognition

DGX agent

arXiv:2502.21004v2 Announce Type: replace Abstract: Dynamic Facial Expression Recognition (DFER) facilitates the understanding of psychological intentions through non-verbal communication. Existing me

researcharxiv-cs-cv
29 May 2026
Tutorials

SRUG: Shadow-Guided Relightable Urban Scene with Generation Model

DGX agent

arXiv:2605.24700v2 Announce Type: replace Abstract: Creating relightable urban scenes from images or videos is widely useful but highly ill-posed. Urban environments are typically unbounded and extend

tutorialsarxiv-cs-cv
29 May 2026
Model Releases

Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning

DGX agent

arXiv:2605.30257v1 Announce Type: new Abstract: We present Stable-Layers, a reinforcement learning framework that eliminates the need for paired supervision by fine-tuning a pretrained layer decomposi

model-releasesarxiv-cs-cv
29 May 2026
← Previous
1…136137138139140…263
Next →