AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Tutorials

Sketch2Motion: Text-driven 2D Sketch to 3D Animation via Diffusion-guided Skeleton Optimization

DGX agent

arXiv:2605.28394v1 Announce Type: new Abstract: Animation of 2D hand-drawn sketches provides an effective medium for visual communication. However, these sketches pose challenges, particularly in hand

tutorialsarxiv-cs-cv
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

ST-ColoNet: Spatio-Temporal Colon Segment Recognition via Hybrid Attention and Edge-Guided Feature Learning

DGX agent

arXiv:2605.28119v1 Announce Type: new Abstract: Colo-segment recognition in colonoscopy videos is a key requirement for many downstream tasks, but existing automatic recognition methods only use colon

researcharxiv-cs-cv
28 May 2026
Model Releases

Stay Fair! Ensuring Group Fairness in Diffusion Models Across Guidance Scales

DGX agent

arXiv:2605.28036v1 Announce Type: new Abstract: Diffusion models steer conditional generation with a tunable guidance scale to trade off prompt alignment and diversity. However, existing debiasing tec

model-releasesarxiv-cs-cv
28 May 2026
Safety

Structure-Guided Visual Perturbation Neutralization for LVLMs

DGX agent

arXiv:2605.27927v1 Announce Type: new Abstract: Image inputs enable Large Vision Language Models (LVLMs) to perceive fine-grained visual information, but also introduce a pixel-level attack surface th

safetyarxiv-cs-cv
28 May 2026
Research

Structure over Pixels: Learning Variable-Length Visual Programs

DGX agent

arXiv:2605.27696v1 Announce Type: new Abstract: Discrete visual tokenizers translate images into ordered sequences of codes, providing a natural representation for structural description of scenes. Ye

researcharxiv-cs-cv
28 May 2026
Research

Super-Resolved Canopy Height Mapping from Sentinel-2 Time Series Using Airborne LiDAR HD Reference Data across Metropolitan France

DGX agent

arXiv:2512.11524v3 Announce Type: replace Abstract: Fine-scale forest monitoring is essential for understanding canopy structure and its dynamics, which are key indicators of carbon stocks, biodiversi

researcharxiv-cs-cv
28 May 2026
Applications

Toward Semantic-Agnostic and Shape-Aware Vision-Language Segmentation Models

DGX agent

arXiv:2605.28348v1 Announce Type: new Abstract: Vision-language segmentation models have recently achieved strong performance by leveraging high-level semantic object categories expressed in natural l

applicationsarxiv-cs-cv
28 May 2026
Safety

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

DGX agent

arXiv:2605.27894v1 Announce Type: new Abstract: Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these

safetyarxiv-cs-cv
28 May 2026
Research

Transfer learning RGB models to hyperspectral images with trainable tensor decompositions

DGX agent

arXiv:2605.28331v1 Announce Type: new Abstract: Transfer learning makes it possible to use large vision networks on a variety of domains, by specializing their models' general filters to new tasks. Ho

researcharxiv-cs-cv
28 May 2026
Hardware

Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation

DGX agent

arXiv:2605.27582v1 Announce Type: cross Abstract: Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environme

hardwarearxiv-cs-cv
28 May 2026
Agents

VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding

DGX agent

arXiv:2605.25059v2 Announce Type: replace Abstract: Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally constructs dense spatial representations on the fly. Ho

agentsarxiv-cs-cv
28 May 2026
Model Releases

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

DGX agent

arXiv:2510.08555v2 Announce Type: replace Abstract: Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpaint

model-releasesarxiv-cs-cv
28 May 2026
Safety

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking

DGX agent

arXiv:2605.28083v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly

safetyarxiv-cs-cv
28 May 2026
Model Releases

WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation

DGX agent

arXiv:2602.22096v2 Announce Type: replace Abstract: Editable high-fidelity 4D scenes are crucial for autonomous driving, as they can be applied to end-to-end training and closed-loop simulation. Howev

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios

DGX agent

arXiv:2605.27589v1 Announce Type: new Abstract: Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not

model-releasesarxiv-cs-cv
28 May 2026
Tutorials

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

DGX agent

arXiv:2605.28132v1 Announce Type: new Abstract: Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this,

tutorialsarxiv-cs-cv
28 May 2026
Model Releases

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

DGX agent

arXiv:2506.22726v4 Announce Type: replace Abstract: Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hin

model-releasesarxiv-cs-cv
28 May 2026
Agents

3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation

DGX agent

arXiv:2605.26500v1 Announce Type: new Abstract: Vision-language navigation (VLN) requires an agent to traverse complex 3D environments based on natural language instructions, necessitating a thorough

agentsarxiv-cs-cv
27 May 2026
Research

A Dynamic Programming Framework for Discovering Count and Values of Multilevel Image Thresholding

DGX agent

arXiv:2605.27287v1 Announce Type: new Abstract: Multilevel Image thresholding is an important preprocessing algorithm in computer vision applications nowadays. Since most common thresholding methods t

researcharxiv-cs-cv
27 May 2026
Research

A multifractal-based masked auto-encoder: an application to medical images

DGX agent

arXiv:2605.26287v1 Announce Type: new Abstract: Masked autoencoders (MAE) have shown great promise in medical image classification. However, the random masking strategy employed by traditional MAEs ma

researcharxiv-cs-cv
27 May 2026
Research

A Unified Framework for Diffusion Model Unlearning with f-Divergence

DGX agent

arXiv:2509.21167v2 Announce Type: replace-cross Abstract: Most existing methods for concept unlearning in text-to-image diffusion models minimize a mean squared error (MSE) loss between the denoiser o

researcharxiv-cs-cv
27 May 2026
Agents

AD-H: Language-guided Autonomous Driving with Hierarchical Agents

DGX agent

arXiv:2406.03474v2 Announce Type: replace Abstract: Language-guided autonomous driving requires bridging a large abstraction gap between high-level natural-language instructions and low-level vehicle

agentsarxiv-cs-cv
27 May 2026
Agents

Adaptation-Free Heterogeneous Collaborative Perception with Unseen Agent Configurations

DGX agent

arXiv:2605.26642v1 Announce Type: new Abstract: Collaborative perception improves 3D object detection by enabling agents to share complementary observations, but most existing methods assume fixed or

agentsarxiv-cs-cv
27 May 2026
Research

Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset

DGX agent

arXiv:2509.18919v2 Announce Type: replace Abstract: The pretraining-finetuning paradigm is a crucial strategy in metallic surface defect detection for mitigating the challenges posed by data scarcity.

researcharxiv-cs-cv
27 May 2026
Safety

AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis

DGX agent

arXiv:2605.25763v2 Announce Type: replace Abstract: Text-to-image synthesis has made significant progress, benefiting from the strong generative capabilities of diffusion models. However, these models

safetyarxiv-cs-cv
27 May 2026
Safety

Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment

DGX agent

arXiv:2511.16870v3 Announce Type: replace Abstract: Enforcing alignment between the internal representations of diffusion or flow-based generative models and those of pretrained self-supervised encode

safetyarxiv-cs-cv
27 May 2026
Model Releases

An uncertainty-aware Bayesian framework for machine learning classification models: A case study in land cover classification

DGX agent

arXiv:2503.21510v3 Announce Type: replace-cross Abstract: Ensuring that predictions of machine learning (ML) classification models are accompanied by uncertainty estimates is one of the main pillars o

model-releasesarxiv-cs-cv
27 May 2026
Research

AnySurf: Any Surface Generation with Directed Edge

DGX agent

arXiv:2605.26149v1 Announce Type: cross Abstract: Open surface components prevail in real industrial 3D content and support rendering, physical simulation and geometric editing. Garments serve as a ty

researcharxiv-cs-cv
27 May 2026
Research

Attenuation-Resilient Alternating Optimization for Laparoscopic Liver Landmark Detection

DGX agent

arXiv:2605.26630v1 Announce Type: new Abstract: Liver surface landmark detection is a fundamental prerequisite for anatomical guidance in laparoscopic liver surgery. However, it remains unreliable in

researcharxiv-cs-cv
27 May 2026
Model Releases

Axial-Centric Cross-Plane Attention for 3D Medical Image Classification

DGX agent

arXiv:2602.21636v2 Announce Type: replace Abstract: Abridged: Clinicians commonly interpret 3D medical images by examining multiple anatomical planes rather than relying on volumetric views. In clinic

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation

DGX agent

arXiv:2605.27067v1 Announce Type: new Abstract: Automatic movie trailer generation must select shots from a full-length film and synchronize them with background music. Existing methods either relegat

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Benchmarking Convolutional, Transformer, Hybrid, and Vision Language Models for Multi Disease Retinal Screening

DGX agent

arXiv:2605.26283v1 Announce Type: new Abstract: Modern deep learning offers powerful tools for automated retinal screening, but it remains unclear how different visual model families compare in realis

model-releasesarxiv-cs-cv
27 May 2026
Safety

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

DGX agent

arXiv:2605.26491v1 Announce Type: cross Abstract: Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image

safetyarxiv-cs-cv
27 May 2026
Agents

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

DGX agent

arXiv:2605.27243v1 Announce Type: new Abstract: Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories

agentsarxiv-cs-cv
27 May 2026
Model Releases

Cesarean Scar Defect Segmentation in Transvaginal Ultrasound Images: a Dataset and Benchmark

DGX agent

arXiv:2605.26774v1 Announce Type: new Abstract: Cesarean Scar Defect (CSD) is one of the most prevalent complications following cesarean delivery. Transvaginal ultrasonography is widely used for prima

model-releasesarxiv-cs-cv
27 May 2026
Tutorials

Chaos-SSL: An Attention-Based Self-Supervised Learning Framework with Chaotic Transformation for Medical Image Classification

DGX agent

arXiv:2605.27146v1 Announce Type: new Abstract: Self-Supervised Learning (SSL) has emerged as a powerful paradigm to mitigate the reliance on large, annotated datasets, a common bottleneck in medical

tutorialsarxiv-cs-cv
27 May 2026
Model Releases

ChartAct: A Benchmark for Dynamic Chart Understanding

DGX agent

arXiv:2605.26994v1 Announce Type: new Abstract: Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts,

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains

DGX agent

arXiv:2605.26734v1 Announce Type: new Abstract: Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-history consistency and are restricted to the fashion domain. To address the

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis

DGX agent

arXiv:2605.26483v1 Announce Type: new Abstract: Medical video diagnosis involves inferring clinical decisions from dynamic tissue responses throughout examination processes. Existing methods rely on a

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

CNNs, Transformers, Hybrid, and Vision Language Models for Skin Cancer Detection

DGX agent

arXiv:2605.26294v1 Announce Type: new Abstract: Skin cancer is a common and fast rising malignancy worldwide. Early detection is critical for improving outcomes. Deep learning models trained on dermos

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning

DGX agent

arXiv:2605.26967v1 Announce Type: new Abstract: Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, wher

model-releasesarxiv-cs-cv
27 May 2026
Applications

ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement

DGX agent

arXiv:2605.25569v2 Announce Type: replace Abstract: Existing deep learning-based low-light enhancement methods are typically trained on limited datasets with single enhancement targets, which restrict

applicationsarxiv-cs-cv
27 May 2026
Model Releases

COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection

DGX agent

arXiv:2605.27116v1 Announce Type: new Abstract: Open-vocabulary object detection (OVD) has made significant progress, enabling detectors to generalize from seen to unseen categories. However, real-wor

model-releasesarxiv-cs-cv
27 May 2026
Tutorials

CRoFT: Robust Fine-Tuning with Concurrent Optimization for OOD Generalization and Open-Set OOD Detection

DGX agent

arXiv:2405.16417v2 Announce Type: replace Abstract: Recent vision-language pre-trained models (VL-PTMs) have shown remarkable success in open-vocabulary tasks. However, downstream use cases often invo

tutorialsarxiv-cs-cv
27 May 2026
Agents

Datasets for Lane Detection in Autonomous Driving: A Comprehensive Review

DGX agent

arXiv:2504.08540v2 Announce Type: replace Abstract: Accurate lane detection is essential for automated driving, enabling safe and reliable vehicle navigation across a variety of road scenarios. Numero

agentsarxiv-cs-cv
27 May 2026
Research

DeepInterestGR: Mining Deep Multi-Interest Using Multi-Modal LLMs for Generative Recommendation

DGX agent

arXiv:2602.18907v2 Announce Type: replace-cross Abstract: We introduce DeepInterestGR, a novel framework that integrates deep interest mining into the generative recommendation pipeline. This addresse

researcharxiv-cs-cv
27 May 2026
Model Releases

DelowlightSplat: Feed-Forward Gaussian Splatting for Lowlight 3D Scene Reconstruction

DGX agent

arXiv:2605.26629v1 Announce Type: new Abstract: Novel-view synthesis and 3D reconstruction from sparse posed images are central to robotics and AR/VR. Yet, feed-forward 3D Gaussian reconstruction fail

model-releasesarxiv-cs-cv
27 May 2026
Agents

Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation

DGX agent

arXiv:2605.26451v1 Announce Type: cross Abstract: Producing presentation slides automatically entails coordinating narrative structure with page-level graphic design under strict spatial constraints.

agentsarxiv-cs-cv
27 May 2026
← Previous
1…140141142143144…263
Next →