AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips

DGX agent

arXiv:2502.07408v2 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) can be catastrophically disrupted by flipping only a handful of parameter bits. We introduce Deep Neural Lesion (D

model-releasesarxiv-cs-cv
17 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models

DGX agent

arXiv:2505.20122v2 Announce Type: replace Abstract: This paper introduces MEBench, a novel benchmark for evaluating mutual exclusivity (ME) bias, a cognitive phenomenon observed in children during wor

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry

DGX agent

arXiv:2604.14866v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remains

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation

DGX agent

arXiv:2604.07021v2 Announce Type: replace Abstract: Weakly supervised semantic segmentation aims to achieve pixel-level predictions using image-level labels. Existing methods typically entangle semant

model-releasesarxiv-cs-cv
17 Apr 2026
Local Ai

MS-SSE-Net: A Multi-Scale Spatial Squeeze-and-Excitation Network for Structural Damage Detection in Civil and Geotechnical Engineering

DGX agent

arXiv:2604.14711v1 Announce Type: new Abstract: Structural damage detection is essential for maintaining the safety and reliability of civil infrastructure. However, accurately identifying different t

local-aiarxiv-cs-cv
17 Apr 2026
Model Releases

Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening

DGX agent

arXiv:2604.14622v1 Announce Type: new Abstract: In this work, we propose a Multigrain-aware Semantic Prototype Scanning paradigm for pan-sharpening, built upon a high-order RWKV architecture and a tri

model-releasesarxiv-cs-cv
17 Apr 2026
Safety

NG-GS: NeRF-Guided 3D Gaussian Splatting Segmentation

DGX agent

arXiv:2604.14706v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled highly efficient and photorealistic novel view synthesis. However, segmenting objects accur

safetyarxiv-cs-cv
17 Apr 2026
Research

NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results

DGX agent

arXiv:2604.14816v1 Announce Type: new Abstract: This paper presents an overview of the NTIRE 2026 Challenge on Video Saliency Prediction. The goal of the challenge participants was to develop automati

researcharxiv-cs-cv
17 Apr 2026
Model Releases

OmniGCD: Abstracting Generalized Category Discovery for Modality Agnosticism

DGX agent

arXiv:2604.14762v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) challenges methods to identify known and novel classes using partially labeled data, mirroring human category learn

model-releasesarxiv-cs-cv
17 Apr 2026
Applications

OmniLight: One Model to Rule All Lighting Conditions

DGX agent

arXiv:2604.15170v1 Announce Type: new Abstract: Adverse lighting conditions, such as cast shadows and irregular illumination, pose significant challenges to computer vision systems by degrading visibi

applicationsarxiv-cs-cv
17 Apr 2026
Research

One-shot Compositional 3D Head Avatars with Deformable Hair

DGX agent

arXiv:2604.14782v1 Announce Type: new Abstract: We propose a compositional method for constructing a complete 3D head avatar from a single image. Prior one-shot holistic approaches frequently fail to

researcharxiv-cs-cv
17 Apr 2026
Model Releases

Open-Set Vein Biometric Recognition with Deep Metric Learning

DGX agent

arXiv:2604.14874v1 Announce Type: new Abstract: Most state-of-the-art vein recognition methods rely on closed-set classification, which inherently limits their scalability and prevents the adaptive en

model-releasesarxiv-cs-cv
17 Apr 2026
Applications

PAGE-4D: Disentangled Pose and Geometry Estimation for VGGT-4D Perception

DGX agent

arXiv:2510.17568v5 Announce Type: replace Abstract: Recent 3D feed-forward models, such as the Visual Geometry Grounded Transformer (VGGT), have shown strong capability in inferring 3D attributes of s

applicationsarxiv-cs-cv
17 Apr 2026
Model Releases

Physically-Induced Atmospheric Adversarial Perturbations: Enhancing Transferability and Robustness in Remote Sensing Image Classification

DGX agent

arXiv:2604.14643v1 Announce Type: new Abstract: Adversarial attacks pose a severe threat to the reliability of deep learning models in remote sensing (RS) image classification. Most existing methods r

model-releasesarxiv-cs-cv
17 Apr 2026
Research

PixelDiT: Pixel Diffusion Transformers for Image Generation

DGX agent

arXiv:2511.20645v2 Announce Type: replace Abstract: Latent-space modeling has been the standard for Diffusion Transformers (DiTs). However, it relies on a two-stage pipeline where the pretrained autoe

researcharxiv-cs-cv
17 Apr 2026
Model Releases

PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation

DGX agent

arXiv:2604.03611v2 Announce Type: replace Abstract: Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coar

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models

DGX agent

arXiv:2604.14591v1 Announce Type: new Abstract: We address the problem of prompt-guided image editing in visual autoregressive models. Given a source image and a target text prompt, we aim to modify t

model-releasesarxiv-cs-cv
17 Apr 2026
Research

Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation

DGX agent

arXiv:2604.14953v1 Announce Type: new Abstract: Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or im

researcharxiv-cs-cv
17 Apr 2026
Research

Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration

DGX agent

arXiv:2503.21970v3 Announce Type: replace Abstract: State-Space Models (SSMs) have attracted considerable attention in Image Restoration (IR) due to their ability to scale linearly sequence length whi

researcharxiv-cs-cv
17 Apr 2026
Research

QualiaNet: An Experience-Before-Inference Network

DGX agent

arXiv:2604.14193v1 Announce Type: new Abstract: Human 3D vision involves two distinct stages: an Experience Module, where stereo depth is extracted relative to fixation, and an Inference Module, where

researcharxiv-cs-cv
17 Apr 2026
Applications

Quality-Aware Calibration for AI-Generated Image Detection in the Wild

DGX agent

arXiv:2604.15027v1 Announce Type: new Abstract: Significant progress has been made in detecting synthetic images, however most existing approaches operate on a single image instance and overlook a key

applicationsarxiv-cs-cv
17 Apr 2026
Safety

R3D: Revisiting 3D Policy Learning

DGX agent

arXiv:2604.15281v1 Announce Type: new Abstract: 3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe o

safetyarxiv-cs-cv
17 Apr 2026
Safety

RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework

DGX agent

arXiv:2604.15308v1 Announce Type: new Abstract: High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interac

safetyarxiv-cs-cv
17 Apr 2026
Research

Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors

DGX agent

arXiv:2604.14563v1 Announce Type: new Abstract: Vision Transformer (ViT)-based sparse multi-view 3D object detectors have achieved remarkable accuracy but still suffer from high inference latency due

researcharxiv-cs-cv
17 Apr 2026
Safety

Reward-Aware Trajectory Shaping for Few-step Visual Generation

DGX agent

arXiv:2604.14910v1 Announce Type: new Abstract: Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely

safetyarxiv-cs-cv
17 Apr 2026
Research

Robustness of Vision Foundation Models to Common Perturbations

DGX agent

arXiv:2604.14973v1 Announce Type: cross Abstract: A vision foundation model outputs an embedding vector for an image, which can be affected by common editing operations (e.g., JPEG compression, bright

researcharxiv-cs-cv
17 Apr 2026
Research

S2AM3D: Scale-controllable Part Segmentation of 3D Point Cloud

DGX agent

arXiv:2512.00995v3 Announce Type: replace Abstract: Part-level point cloud segmentation has recently attracted significant attention in 3D computer vision. Nevertheless, existing research is constrain

researcharxiv-cs-cv
17 Apr 2026
Safety

SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning

DGX agent

arXiv:2604.14373v1 Announce Type: new Abstract: Rural environmental risks are shaped by place-based conditions (e.g., housing quality, road access, land-surface patterns), yet standard vulnerability i

safetyarxiv-cs-cv
17 Apr 2026
Research

Seen-to-Scene: Keep the Seen, Generate the Unseen for Video Outpainting

DGX agent

arXiv:2604.14648v1 Announce Type: new Abstract: Video outpainting aims to expand the visible content of a video beyond the original frame boundaries while preserving spatial fidelity and temporal cohe

researcharxiv-cs-cv
17 Apr 2026
Research

SegviGen: Repurposing 3D Generative Model for Part Segmentation

DGX agent

arXiv:2603.16869v2 Announce Type: replace Abstract: We introduce SegviGen, a framework that repurposes native 3D generative models for 3D part segmentation. Existing pipelines either lift strong 2D pr

researcharxiv-cs-cv
17 Apr 2026
Research

SegWithU: Uncertainty as Perturbation Energy for Single-Forward-Pass Risk-Aware Medical Image Segmentation

DGX agent

arXiv:2604.15271v1 Announce Type: new Abstract: Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and clinical decisio

researcharxiv-cs-cv
17 Apr 2026
Model Releases

Speak, Segment, Track, Navigate: An Interactive System for Video-Guided Skull-Base Surgery

DGX agent

arXiv:2603.16024v2 Announce Type: replace Abstract: We introduce a speech-guided embodied agent framework for video-guided skull base surgery that dynamically executes perception and image-guidance ta

model-releasesarxiv-cs-cv
17 Apr 2026
Safety

Step-level Denoising-time Diffusion Alignment with Multiple Objectives

DGX agent

arXiv:2604.14379v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a powerful tool for aligning diffusion models with human preferences, typically by optimizing a single rewa

safetyarxiv-cs-cv
17 Apr 2026
Research

STEP-Parts: Geometric Partitioning of Boundary Representations for Large-Scale CAD Processing

DGX agent

arXiv:2604.14927v1 Announce Type: cross Abstract: Many CAD learning pipelines discretize Boundary Representations (B-Reps) into triangle meshes, discarding analytic surface structure and topological a

researcharxiv-cs-cv
17 Apr 2026
Research

StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression

DGX agent

arXiv:2604.15237v1 Announce Type: new Abstract: Reconstructing dense 3D geometry from continuous video streams requires stable inference under a constant memory budget. Existing O(1) frameworks primar

researcharxiv-cs-cv
17 Apr 2026
Safety

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models

DGX agent

arXiv:2604.14629v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown remarkable capabilities in joint vision-language understanding, but their large scale poses significant challen

safetyarxiv-cs-cv
17 Apr 2026
Model Releases

TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?

DGX agent

arXiv:2509.15602v5 Announce Type: replace Abstract: Multimodal large language models (MLLMs) excel at general video understanding but struggle with fast, high-frequency sports like tennis, where rally

model-releasesarxiv-cs-cv
17 Apr 2026
Local Ai

The Courtroom Trial of Pixels: Robust Image Manipulation Localization via Adversarial Evidence and Reinforcement Learning Judgment

DGX agent

arXiv:2604.14703v1 Announce Type: new Abstract: Although some existing image manipulation localization (IML) methods incorporate authenticity-related supervision, this information is typically utilize

local-aiarxiv-cs-cv
17 Apr 2026
Model Releases

The Fourth Challenge on Image Super-Resolution (imes4) at NTIRE 2026: Benchmark Results and Method Overview

DGX agent

arXiv:2604.14558v1 Announce Type: new Abstract: This paper presents the NTIRE 2026 image super-resolution (imes4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026.

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Think in Latent Thoughts: A New Paradigm for Gloss-Free Sign Language Translation

DGX agent

arXiv:2604.15301v1 Announce Type: new Abstract: Many SLT systems quietly assume that brief chunks of signing map directly to spoken-language words. That assumption breaks down because signers often cr

model-releasesarxiv-cs-cv
17 Apr 2026
Safety

To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs

DGX agent

arXiv:2603.18373v2 Announce Type: replace Abstract: When VLMs answer correctly, do they genuinely rely on visual information or exploit language shortcuts? We introduce the Tri-Layer Diagnostic Framew

safetyarxiv-cs-cv
17 Apr 2026
Research

TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens

DGX agent

arXiv:2604.15239v1 Announce Type: new Abstract: In this work, we revisit several key design choices of modern Transformer-based approaches for feed-forward 3D Gaussian Splatting (3DGS) prediction. We

researcharxiv-cs-cv
17 Apr 2026
Research

TokenLight: Precise Lighting Control in Images using Attribute Tokens

DGX agent

arXiv:2604.15310v1 Announce Type: new Abstract: This paper presents a method for image relighting that enables precise and continuous control over multiple illumination attributes in a photograph. We

researcharxiv-cs-cv
17 Apr 2026
Research

Towards Design Compositing

DGX agent

arXiv:2604.14605v1 Announce Type: new Abstract: Graphic design creation involves harmoniously assembling multimodal components such as images, text, logos, and other visual assets collected from diver

researcharxiv-cs-cv
17 Apr 2026
Applications

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation

DGX agent

arXiv:2604.14580v1 Announce Type: new Abstract: Existing audio-driven video digital human generation models rely on multi-step denoising, resulting in substantial computational overhead that severely

applicationsarxiv-cs-cv
17 Apr 2026
Safety

TwinOR: Photorealistic Digital Twins of Dynamic Operating Rooms for Embodied AI Research

DGX agent

arXiv:2511.07412v2 Announce Type: replace Abstract: Developing embodied AI for intelligent surgical systems requires safe, controllable environments for continual learning and evaluation. However, saf

safetyarxiv-cs-cv
17 Apr 2026
Safety

UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

DGX agent

arXiv:2604.14967v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) extends Large Vision-Language Models (LVLMs) with external visual knowledge. However, existing visual RAG systems t

safetyarxiv-cs-cv
17 Apr 2026
Safety

Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization

DGX agent

arXiv:2604.15196v1 Announce Type: new Abstract: We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first intr

safetyarxiv-cs-cv
17 Apr 2026
← Previous
1…237238239240241…261
Next →