AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Applications

MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models

DGX agent

arXiv:2509.22151v3 Announce Type: replace Abstract: Material node graphs are programs that generate the 2D channels of procedural materials, including geometry such as roughness and displacement maps,

applicationsarxiv-cs-cv
15 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation

DGX agent

arXiv:2508.05008v2 Announce Type: replace Abstract: Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their ap

researcharxiv-cs-cv
15 May 2026
Model Releases

Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models

DGX agent

arXiv:2605.14938v1 Announce Type: cross Abstract: Continual learning in multimodal large language models (MLLMs) aims to sequentially acquire knowledge while mitigating catastrophic forgetting, yet ex

model-releasesarxiv-cs-cv
15 May 2026
Research

PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models

DGX agent

arXiv:2505.22394v2 Announce Type: replace Abstract: We present PacTure, a novel framework for generating physically-based rendering (PBR) material textures for an untextured 3D mesh from a text descri

researcharxiv-cs-cv
15 May 2026
Research

PanoPlane: Plane-Aware Panoramic Completion for Sparse-View Indoor 3D Gaussian Splatting

DGX agent

arXiv:2605.14135v1 Announce Type: new Abstract: We present PanoPlane, an approach for high-fidelity sparse-view indoor novel view synthesis that reconstructs closed room geometry via panoramic scene c

researcharxiv-cs-cv
15 May 2026
Model Releases

Physics-Grounded Adversarial Stain Augmentation with Calibrated Coverage Guarantees

DGX agent

arXiv:2605.13889v1 Announce Type: cross Abstract: Stain variation across hospitals degrades histopathology models at deployment. Existing augmentation methods perturb color spaces with arbitrary hyper

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Probing into Camera Control of Video Models

DGX agent

arXiv:2605.14815v1 Announce Type: new Abstract: Video is a rich and scalable source of 3D/4D visual observations, and camera control is a key capability for video generation models to produce geometri

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

PVRF: All-in-one Adverse Weather Removal via Prior-modulated and Velocity-constrained Rectified Flow

DGX agent

arXiv:2605.14045v1 Announce Type: new Abstract: Adverse weather removal (AWR) in real-world images remains challenging due to heterogeneous and unseen degradations, while distortion-driven training of

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

RAM-W600: A Multi-Task Wrist Dataset and Benchmark for Rheumatoid Arthritis

DGX agent

arXiv:2507.05193v4 Announce Type: replace-cross Abstract: Rheumatoid arthritis (RA) is a common autoimmune disease that has been the focus of research in computer-aided diagnosis (CAD) and disease mon

model-releasesarxiv-cs-cv
15 May 2026
Safety

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling

DGX agent

arXiv:2510.20206v2 Announce Type: replace Abstract: Prompt design plays a crucial role in text-to-video (T2V) generation, yet user-provided prompts are often short, unstructured, and misaligned with t

safetyarxiv-cs-cv
15 May 2026
Safety

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO

DGX agent

arXiv:2605.15190v1 Announce Type: new Abstract: Causal autoregressive video diffusion models support real-time streaming generation by extrapolating future chunks from previously generated content. Di

safetyarxiv-cs-cv
15 May 2026
Model Releases

Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

DGX agent

arXiv:2605.14462v1 Announce Type: new Abstract: Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based l

model-releasesarxiv-cs-cv
15 May 2026
Tutorials

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning

DGX agent

arXiv:2605.13852v1 Announce Type: cross Abstract: We often aim to generate images that are both photorealistic and 3D-consistent, adhering to precise geometry, material, and viewpoint controls. Typica

tutorialsarxiv-cs-cv
15 May 2026
Model Releases

Reduce the Artifacts Bias for More Generalizable AI-Generated Image Detection

DGX agent

arXiv:2605.14486v1 Announce Type: new Abstract: As the misuse of AI-generated images grows, generalizable image detection techniques are urgently needed. Recent state-of-the-art (SOTA) methods adopt a

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

RefDecoder: Enhancing Visual Generation with Conditional Video Decoding

DGX agent

arXiv:2605.15196v1 Announce Type: new Abstract: Video generation powers a vast array of downstream applications. However, while the de facto standard, i.e., latent diffusion models, typically employ h

model-releasesarxiv-cs-cv
15 May 2026
Research

RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model

DGX agent

arXiv:2512.12083v3 Announce Type: replace Abstract: Semantic-rich features from Vision Foundation Models (VFMs) have been leveraged to enhance Latent Diffusion Models (LDMs). However, raw VFM features

researcharxiv-cs-cv
15 May 2026
Research

Representative Attention For Vision Transformers

DGX agent

arXiv:2605.14913v1 Announce Type: new Abstract: Linear attention has emerged as a promising direction for scaling Vision Transformers beyond the quadratic cost of dense self-attention. A prevalent str

researcharxiv-cs-cv
15 May 2026
Research

Rethinking the Good Enough Embedding for Easy Few-Shot Learning

DGX agent

arXiv:2605.14145v1 Announce Type: new Abstract: The field of deep visual recognition is undergoing a paradigm shift toward universal representations. The Platonic Representation Hypothesis suggests th

researcharxiv-cs-cv
15 May 2026
Safety

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

DGX agent

arXiv:2511.13026v3 Announce Type: replace Abstract: Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied

safetyarxiv-cs-cv
15 May 2026
Safety

Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse

DGX agent

arXiv:2605.14925v1 Announce Type: new Abstract: Drone-view geo-localization aims to match a query drone image, often captured under adverse weather conditions (e.g., rain, snow, fog), against a galler

safetyarxiv-cs-cv
15 May 2026
Research

SAGE3D: Soft-guided attention and graph excitation for 3D point cloud corner detection

DGX agent

arXiv:2605.15088v1 Announce Type: new Abstract: We present SAGE3D, a hybrid Transformer-based model for corner detection in airborne LiDAR point clouds. We propose a multi-stage solution built on a hi

researcharxiv-cs-cv
15 May 2026
Model Releases

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

DGX agent

arXiv:2605.15178v1 Announce Type: new Abstract: We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p,

model-releasesarxiv-cs-cv
15 May 2026
Research

SceneForge: Structured World Supervision from 3D Interventions

DGX agent

arXiv:2605.14399v1 Announce Type: new Abstract: Many multimodal learning tasks require supervision that remains consistent across edits, viewpoints, and scene-level interventions. However, such superv

researcharxiv-cs-cv
15 May 2026
Model Releases

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding

DGX agent

arXiv:2605.14923v1 Announce Type: new Abstract: General scene perception has progressed from object recognition toward open-vocabulary grounding, part localization, and affordance prediction. Yet thes

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples

DGX agent

arXiv:2507.07776v3 Announce Type: replace Abstract: Unrestricted adversarial attacks aim to fool computer vision models without being constrained by ell_p-norm bounds to remain imperceptible to humans

model-releasesarxiv-cs-cv
15 May 2026
Applications

SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation

DGX agent

arXiv:2605.14926v1 Announce Type: new Abstract: Achieving pixel-level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face signific

applicationsarxiv-cs-cv
15 May 2026
Research

SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification

DGX agent

arXiv:2504.09549v3 Announce Type: replace Abstract: Aerial-Ground Person Re-IDentification (AG-ReID) aims to retrieve specific persons across cameras with different viewpoints. Previous works focus on

researcharxiv-cs-cv
15 May 2026
Local Ai

SEDiT: Mask-Free Video Subtitle Erasure via One-step Diffusion Transformer

DGX agent

arXiv:2605.14894v1 Announce Type: new Abstract: Recent breakthroughs in video diffusion models have significantly accelerated the development of video editing techniques. However, existing methods oft

local-aiarxiv-cs-cv
15 May 2026
Local Ai

Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation

DGX agent

arXiv:2605.13862v1 Announce Type: cross Abstract: We present Seed3D 2.0, an advanced 3D content generation system built on Seed3D 1.0, with substantial improvements across generation fidelity, simulat

local-aiarxiv-cs-cv
15 May 2026
Safety

SpectraFlow: Unifying Structural Pretraining and Frequency Adaptation for Medical Image Segmentation

DGX agent

arXiv:2605.14566v1 Announce Type: new Abstract: Medical image segmentation remains challenging in low-data regimes, where scarce annotations often yield poor generalization and ambiguous boundaries wi

safetyarxiv-cs-cv
15 May 2026
Model Releases

SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation

DGX agent

arXiv:2605.14847v1 Announce Type: new Abstract: Modern image super-resolution methods generate detailed, visually appealing results, but they often introduce visual artifacts: unnatural patterns and t

model-releasesarxiv-cs-cv
15 May 2026
Local Ai

SteerSeg: Attention Steering for Reasoning Video Segmentation

DGX agent

arXiv:2605.14908v1 Announce Type: new Abstract: Video reasoning segmentation requires localizing objects across video frames from natural language expressions, often involving spatial reasoning and im

local-aiarxiv-cs-cv
15 May 2026
Model Releases

SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection

DGX agent

arXiv:2605.14110v1 Announce Type: new Abstract: Vision Transformers (ViTs) enable strong multi-view 3D detection but are limited by high inference latency from dense token and query processing across

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

StyleTextGen: Style-Conditioned Multilingual Scene Text Generation

DGX agent

arXiv:2605.14708v1 Announce Type: new Abstract: Style-conditioned scene text generation faces unique challenges in extracting precise text styles from complex backgrounds and maintaining fine-grained

model-releasesarxiv-cs-cv
15 May 2026
Applications

SuperADD: Training-free Class-agnostic Anomaly Segmentation -- CVPR 2026 VAND 4.0 Workshop Challenge Industrial Track

DGX agent

arXiv:2605.14808v1 Announce Type: new Abstract: Visual anomaly detection (AD) for industrial inspection is a highly relevant task in modern production environments. The problem becomes particularly ch

applicationsarxiv-cs-cv
15 May 2026
Safety

SuperF: Neural Implicit Fields for Multi-Image Super-Resolution

DGX agent

arXiv:2512.09115v2 Announce Type: replace Abstract: High-resolution imagery is often hindered by limitations in sensor technology, atmospheric conditions, and costs. Such challenges occur in satellite

safetyarxiv-cs-cv
15 May 2026
Model Releases

SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding

DGX agent

arXiv:2510.13016v3 Announce Type: replace Abstract: A truly capable AI system must do more than detect objects or recognize activities in isolation. It must form unified, grounded representations of w

model-releasesarxiv-cs-cv
15 May 2026
Applications

SyncLight: Single-Edit Multi-View Relighting

DGX agent

arXiv:2601.16981v2 Announce Type: replace Abstract: We present SyncLight, a method to enable consistent, parametric control over light sources across multiple uncalibrated views of a static scene cond

applicationsarxiv-cs-cv
15 May 2026
Safety

Systematic Discovery of Semantic Attacks in Online Map Construction through Conditional Diffusion

DGX agent

arXiv:2605.14396v1 Announce Type: new Abstract: Autonomous vehicles depend on online HD map construction to perceive lane boundaries, dividers, and pedestrian crossings -- safety-critical road element

safetyarxiv-cs-cv
15 May 2026
Research

TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion

DGX agent

arXiv:2605.14136v1 Announce Type: new Abstract: Recent text-to-video diffusion transformers generate visually compelling frames, yet still struggle with temporal coherence, often producing flickering,

researcharxiv-cs-cv
15 May 2026
Model Releases

TERRA-CD: Multi-Temporal Framework for Multi-class and Semantic Change Detection

DGX agent

arXiv:2605.14651v1 Announce Type: new Abstract: Urban vegetation monitoring plays a vital role in understanding environmental changes, yet comprehensive datasets for this purpose remain limited. To ad

model-releasesarxiv-cs-cv
15 May 2026
Applications

The Potential of Convolutional Neural Networks for Cancer Detection

DGX agent

arXiv:2412.17155v4 Announce Type: replace Abstract: Early detection is crucial for successful cancer treatment and increasing survivability rates, particularly in the most common forms. Ten different

applicationsarxiv-cs-cv
15 May 2026
Research

The Velocity Deficit: Initial Energy Injection for Flow Matching

DGX agent

arXiv:2605.14819v1 Announce Type: new Abstract: While Flow Matching theoretically guarantees constant-velocity trajectories, we identify a critical breakdown in high-dimensional practice: the Velocity

researcharxiv-cs-cv
15 May 2026
Applications

TOPOS: High-Fidelity and Efficient Industry-Grade 3D Head Generation

DGX agent

arXiv:2605.14594v1 Announce Type: new Abstract: High-fidelity 3D head generation plays a crucial role in the film, animation and video game industries. In industrial pipelines, studios typically enfor

applicationsarxiv-cs-cv
15 May 2026
Research

Towards Accurate Single Panoramic 3D Detection: A Semantic Gaussian Centric Approach

DGX agent

arXiv:2605.14601v1 Announce Type: new Abstract: Three-dimensional object detection in panoramic imagery is crucial for comprehensive scene understanding, yet accurately mapping 2D features to 3D remai

researcharxiv-cs-cv
15 May 2026
Safety

Towards Continuous Sign Language Conversation from Isolated Signs

DGX agent

arXiv:2605.14705v1 Announce Type: new Abstract: Sign language is the primary language for many Deaf and Hard-of-Hearing (DHH) signers, yet most conversational AI systems still mediate interaction thro

safetyarxiv-cs-cv
15 May 2026
Local Ai

Towards Real-Time Autonomous Navigation: Transformer-Based Catheter Tip Tracking in Fluoroscopy

DGX agent

arXiv:2605.14253v1 Announce Type: new Abstract: Purpose: Mechanical thrombectomy (MT) improves stroke outcomes, but is limited by a lack of local treatment access. Widespread distribution of reinforce

local-aiarxiv-cs-cv
15 May 2026
Local Ai

Training-Free Inference for High-Resolution Sinogram Completion

DGX agent

arXiv:2506.08809v5 Announce Type: replace Abstract: High-resolution sinogram completion is critical for computed tomography reconstruction, as missing projections can introduce severe artifacts. While

local-aiarxiv-cs-cv
15 May 2026
← Previous
1…170171172173174…263
Next →