AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Cyclic Denoising Reveals Ultrastable Memories in Diffusion Models

DGX agent

arXiv:2606.24000v1 Announce Type: cross Abstract: We introduce cyclic denoising -- repeated forward and reverse diffusion at controlled noise amplitudes -- as an extraction attack for image diffusion

researcharxiv-cs-cv
24 Jun 2026
Tutorials
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation

DGX agent

arXiv:2606.18478v2 Announce Type: replace Abstract: Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matchin

tutorialsarxiv-cs-cv
24 Jun 2026
Safety

DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection

DGX agent

arXiv:2606.24805v1 Announce Type: new Abstract: Stereo-based 3D object detection still faces two critical safety challenges: real-time performance and open-set generalization. Existing stereo 3D metho

safetyarxiv-cs-cv
24 Jun 2026
Hardware

Differentiable Packing of Irregular 3D Objects with Adaptive Container Estimation

DGX agent

arXiv:2606.16333v2 Announce Type: replace Abstract: Most existing approaches either fix the container in advance or optimize only a single container dimension through an outer search loop, leaving the

hardwarearxiv-cs-cv
24 Jun 2026
Model Releases

Differential Unfolding: Efficient Unfolding Reconstruction for Video Snapshot Compressive Imaging

DGX agent

arXiv:2606.24153v1 Announce Type: new Abstract: While Deep Unfolding Networks (DUNs) dominate video Snapshot Compressive Imaging (SCI), they remain constrained by a uniform design philosophy. Existing

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

DiffusionBench: On Holistic Evaluation of Diffusion Transformers

DGX agent

arXiv:2606.24888v1 Announce Type: new Abstract: Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While met

model-releasesarxiv-cs-cv
24 Jun 2026
Research

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation

DGX agent

arXiv:2606.23950v1 Announce Type: new Abstract: Subject-driven image generation faces an 'Identity-Diversity Paradox', where strong identity preservation often leads to rigid and low-diversity outputs

researcharxiv-cs-cv
24 Jun 2026
Model Releases

DLTPose: 6DoF Pose Estimation From Accurate Dense Surface Point Estimates

DGX agent

arXiv:2504.07335v3 Announce Type: replace Abstract: We propose DLTPose, a novel method for 6DoF object pose estimation from RGBD images that combines the accuracy of sparse keypoint methods with the r

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Does it matter which Gaussians you pick in 4D Gaussian streaming?

DGX agent

arXiv:2603.17227v3 Announce Type: replace Abstract: Anchor-driven 4D Gaussian streaming methods such as Instant Gaussian Stream (IGS) update a dynamic scene each frame from a compact set of Gaussian a

model-releasesarxiv-cs-cv
24 Jun 2026
Safety

DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model

DGX agent

arXiv:2606.24051v1 Announce Type: new Abstract: Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follow

safetyarxiv-cs-cv
24 Jun 2026
Model Releases

Dual-Branch Cross-Projection Debiasing through Diffusion-based Disentanglement

DGX agent

arXiv:2606.24161v1 Announce Type: new Abstract: Foundation models trained on biased datasets often rely on spurious correlations between target labels and non-causal attributes, resulting in poor gene

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

DGX agent

arXiv:2512.24731v2 Announce Type: replace Abstract: Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite r

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

EERLoss: A Novel Loss Function for Training Deep Biometric Models. A Case Study in Keystroke Dynamics

DGX agent

arXiv:2606.24586v1 Announce Type: new Abstract: Deep learning approaches to biometric verification are commonly trained by optimizing indirect objectives, creating a misalignment between the optimizat

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

DGX agent

arXiv:2606.24422v1 Announce Type: new Abstract: We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of mo

model-releasesarxiv-cs-cv
24 Jun 2026
Tutorials

Emotion Diffusion Classifier with Adaptive Margin Discrepancy Training for Facial Expression Recognition

DGX agent

arXiv:2603.29578v2 Announce Type: replace Abstract: Facial Expression Recognition (FER) is essential for human-machine interaction, as it enables machines to interpret human emotions and internal stat

tutorialsarxiv-cs-cv
24 Jun 2026
Research

EPEdit: Redefining Image Editing with Generative AI and User-Centric Design

DGX agent

arXiv:2606.24057v1 Announce Type: new Abstract: The demand for image manipulation has seen a significant increase recently. Traditional tools like Photoshop and Capture One, while powerful, require co

researcharxiv-cs-cv
24 Jun 2026
Model Releases

EPMF: Efficient Perception-aware Multi-sensor Fusion for 3D Semantic Segmentation

DGX agent

arXiv:2106.15277v4 Announce Type: replace Abstract: We study multi-sensor fusion for 3D semantic segmentation that is important to scene understanding for many applications, such as autonomous driving

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Fabric Image Demoireing Benchmark from Synthesis to Restoration

DGX agent

arXiv:2606.24072v1 Announce Type: new Abstract: Fabric moire is a sampling-induced aliasing artifact caused by the interaction between fine textile patterns and camera sensor grids, producing structur

model-releasesarxiv-cs-cv
24 Jun 2026
Research

FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait Image

DGX agent

arXiv:2606.24232v1 Announce Type: new Abstract: We introduce FiCA, a Feed-forward, instant Gaussian Codec Avatar generation pipeline that creates lifelike avatars from a single portrait image. Generat

researcharxiv-cs-cv
24 Jun 2026
Model Releases

Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark

DGX agent

arXiv:2503.14862v3 Announce Type: replace Abstract: Open-vocabulary detectors are proposed to locate and recognize objects in novel classes. However, variations in vision-aware language vocabulary dat

model-releasesarxiv-cs-cv
24 Jun 2026
Research

FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation

DGX agent

arXiv:2606.24876v1 Announce Type: new Abstract: Generating explorable 3D scenes from a single image requires strong generative priors and accurate geometric representations suitable for downstream use

researcharxiv-cs-cv
24 Jun 2026
Research

Flood Mapping from RGB imagery using a Vision Foundation Model

DGX agent

arXiv:2606.24120v1 Announce Type: new Abstract: Timely, high-resolution maps of flood extent around settlements are essential for emergency response and damage assessment. We consider airborne RGB ima

researcharxiv-cs-cv
24 Jun 2026
Model Releases

FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation

DGX agent

arXiv:2511.21029v3 Announce Type: replace Abstract: Music-to-dance generation aims to translate auditory signals into expressive human motion, with broad applications in virtual reality, choreography,

model-releasesarxiv-cs-cv
24 Jun 2026
Local Ai

ForensicsTok: Forensics-Guided Tokenized Modeling for Image Tampering Localization

DGX agent

arXiv:2606.24538v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) offer powerful reasoning for forensic tasks, yet existing approaches utilizing exogenous segmentation decoders

local-aiarxiv-cs-cv
24 Jun 2026
Agents

From Open Waters to Enclosed Cabins: ProteusVPR for Cross-Scene Visual Place Recognition in Maritime Perception and Cabin Inspection

DGX agent

arXiv:2606.24234v1 Announce Type: new Abstract: Autonomous robotic inspection in maritime environments presents unique challenges for Visual Place Recognition (VPR) due to cross-scene perceptual shift

agentsarxiv-cs-cv
24 Jun 2026
Research

Full-resolution MLPs Empower Medical Dense Prediction

DGX agent

arXiv:2311.16707v2 Announce Type: replace-cross Abstract: Dense prediction is a fundamental requirement for many medical vision tasks such as medical image restoration, registration, and segmentation.

researcharxiv-cs-cv
24 Jun 2026
Safety

GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence

DGX agent

arXiv:2511.21945v3 Announce Type: replace Abstract: Generating complete 3D objects under partial occlusions (i.e., amodal scenarios) is a practically important yet challenging problem, as large portio

safetyarxiv-cs-cv
24 Jun 2026
Safety

Generative Manifold Distillation: Aligning Restoration Trajectories with Natural Image Prior

DGX agent

arXiv:2512.11121v2 Announce Type: replace Abstract: Pre-trained image restoration models often fail on out-of-distribution (OOD) real-world degradations. Adapting to these domains is challenging as re

safetyarxiv-cs-cv
24 Jun 2026
Local Ai

GeoIMO: Geometry-Driven Independent Motion Classification for Event Cameras

DGX agent

arXiv:2606.24499v1 Announce Type: new Abstract: Existing automotive event datasets rely on appearance-based annotations from frame pipelines, making them poorly suited for motion-aware event perceptio

local-aiarxiv-cs-cv
24 Jun 2026
Safety

Geometric Action Model for Robot Policy Learning

DGX agent

arXiv:2606.17046v2 Announce Type: replace-cross Abstract: Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physi

safetyarxiv-cs-cv
24 Jun 2026
Research

Geometry-Aware Style Transfer in 3D Gaussian Splatting

DGX agent

arXiv:2606.24144v1 Announce Type: new Abstract: In this paper, we present a novel geometry-aware style transfer framework for 3D Gaussian splatting (3DGS) that simultaneously transfers appearance attr

researcharxiv-cs-cv
24 Jun 2026
Research

Geometry-Instructed Video Editing

DGX agent

arXiv:2606.24225v1 Announce Type: new Abstract: Object-level geometric edits, including translating, rotating, scaling, duplicating, or removing an object, are routine operations in digital content cr

researcharxiv-cs-cv
24 Jun 2026
Model Releases

GeoT2V-Bench: Benchmarking 3D Consistency in Text-to-Video Models via 3D Reconstruction

DGX agent

arXiv:2606.24829v1 Announce Type: new Abstract: Camera-prompted text-to-video (T2V) models are increasingly used to synthesize virtual camera captures, such as orbiting objects or moving through stati

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models

DGX agent

arXiv:2606.23843v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are typically pre-trained on large-scale image-text datasets to capture semantic correspondences between visual content an

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Heterogeneous Knowledge Distillation via Geometry Decoupling and Momentum-Aware Gradient Regulation

DGX agent

arXiv:2606.24557v1 Announce Type: new Abstract: Heterogeneous Knowledge Distillation (HKD) aims to transfer knowledge across varying architectures (e.g., from Transformer to CNN) but inherently suffer

model-releasesarxiv-cs-cv
24 Jun 2026
Safety

Hierarchical Spatial and Channel Aggregation for Cross-domain Few-shot Segmentation

DGX agent

arXiv:2606.24296v1 Announce Type: new Abstract: Cross-domain Few-shot Segmentation (CD-FSS) aims to learn generalizable segmentation capability from abundant annotated samples in the source domain, en

safetyarxiv-cs-cv
24 Jun 2026
Research

High-Fidelity Synthetic Transmission Electron Microscopy Image Generation Using Diffusion Probabilistic Models for Data-Limited Semiconductor Metrology

DGX agent

arXiv:2606.24817v1 Announce Type: new Abstract: Advanced semiconductor nodes drastically increased demand for Transmission Electron Microscopy (TEM), yet destructive sample preparation, slow imaging a

researcharxiv-cs-cv
24 Jun 2026
Research

Hybrid Event Frame Sensors: Modeling, Calibration, and Simulation

DGX agent

arXiv:2511.18037v2 Announce Type: replace Abstract: Hybrid event-frame sensors integrate an Event Vision Sensor (EVS) and an Active Pixel Sensor (APS) within a single chip, combining the high dynamic

researcharxiv-cs-cv
24 Jun 2026
Local Ai

Ill-Posed by Design: Probing Evidence Use in VLMs

DGX agent

arXiv:2606.24335v1 Announce Type: new Abstract: Counterfactual analysis is widely used to study evidence use in vision-language models, but its diagnostic value is limited on well-posed tasks: when se

local-aiarxiv-cs-cv
24 Jun 2026
Research

Ingredient-Level Food Image Segmentation for Nutrition Awareness

DGX agent

arXiv:2606.24059v1 Announce Type: new Abstract: Food images often contain several visible ingredients, so assigning one dish label to an entire image hides important visual structure. This work studie

researcharxiv-cs-cv
24 Jun 2026
Model Releases

Jolia: Concept-Level Vision-Language Alignment for 3D CT Contrastive Learning

DGX agent

arXiv:2606.24570v1 Announce Type: new Abstract: Vision-language contrastive pretraining has become the dominant recipe for 3D medical foundation models, leveraging the large volumes of paired scans an

model-releasesarxiv-cs-cv
24 Jun 2026
Safety

Latent Visual States for Efficient Multimodal Reasoning

DGX agent

arXiv:2606.24233v1 Announce Type: new Abstract: The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relie

safetyarxiv-cs-cv
24 Jun 2026
Applications

Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching

DGX agent

arXiv:2606.24457v1 Announce Type: new Abstract: Recent advances in stereo matching have achieved remarkable accuracy, but often rely on large models, heavy computation, or additional foundation-model

applicationsarxiv-cs-cv
24 Jun 2026
Model Releases

LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation

DGX agent

arXiv:2509.17773v2 Announce Type: replace Abstract: The rapid progress of image-guided video generation (I2V) has raised concerns about its potential misuse in misinformation and fraud, underscoring t

model-releasesarxiv-cs-cv
24 Jun 2026
Research

M^2C-EvDet: Multi-Domain Multi-Order Cross-Modal Knowledge Distillation for Event-based Object Detection

DGX agent

arXiv:2606.24248v1 Announce Type: new Abstract: Event-based object Detection (EvDet), as a biologically inspired visual perception paradigm, demonstrates superior performance in scenarios demanding hi

researcharxiv-cs-cv
24 Jun 2026
Model Releases

M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for optical-SAR Object Detection

DGX agent

arXiv:2505.10931v4 Announce Type: replace Abstract: Single-source remote sensing object detection using optical or SAR images struggles in complex environments. Optical images offer rich textural deta

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Machine Learning Modeling for Real-Time Melt Pool Monitoring in Laser Powder Bed Fusion Additive Manufacturing: A Hybrid Approach

DGX agent

arXiv:2606.23851v1 Announce Type: cross Abstract: This work investigates the implementation of artificial intelligence and machine learning (AI/ML) for real-time monitoring in laser powder bed fusion

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Mamba-FSCIL: Dynamic Adaptation with Selective State Space Model for Few-Shot Class-Incremental Learning

DGX agent

arXiv:2407.06136v4 Announce Type: replace Abstract: Few-shot class-incremental learning (FSCIL) aims to incrementally learn novel classes from limited examples while preserving knowledge of previously

model-releasesarxiv-cs-cv
24 Jun 2026
← Previous
1…9293949596…263
Next →