AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA

DGX agent

arXiv:2512.10521v2 Announce Type: replace Abstract: Few-shot semantic segmentation (FSS) aims to segment novel classes in query images using only a small annotated support set. While prior research ha

researcharxiv-cs-cv
4 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Tutorials

TGSD: Topology-Guided State-Space Diffusion for EEG Spatial Super-Resolution

DGX agent

arXiv:2606.03998v1 Announce Type: cross Abstract: Low-density EEG is more suitable for wearable and IoT-based brain sensing, but sparse electrode sampling often lacks sufficient spatial information to

tutorialsarxiv-cs-cv
4 Jun 2026
Safety

Toward Multi-Domain and Long-Tailed Quantization via Feature Alignment and Scaling

DGX agent

arXiv:2606.04920v1 Announce Type: cross Abstract: Quantizing deep neural networks is essential for efficient inference on resource-constrained devices. However, most existing methods are designed for

safetyarxiv-cs-cv
4 Jun 2026
Model Releases

Toward Trustworthy Portrait Editing: Evaluation of Demographic Misrepresentation in I2I Models

DGX agent

arXiv:2602.16149v2 Announce Type: replace Abstract: Instruction-guided image-to-image (I2I) editors are increasingly used in consumer and professional visual workflows, where trustworthiness depends n

model-releasesarxiv-cs-cv
4 Jun 2026
Applications

Towards Evaluating the Robustness of Visual State Space Models

DGX agent

arXiv:2406.09407v3 Announce Type: replace Abstract: Vision State Space Models (VSSMs), a novel architecture that combines the strengths of recurrent neural networks and latent variable models, have de

applicationsarxiv-cs-cv
4 Jun 2026
Safety

Transferable Multi-Bit Watermarking Across Frozen Diffusion Models via Latent Consistency Bridges

DGX agent

arXiv:2603.20304v2 Announce Type: replace Abstract: As generative AI advances, global governance frameworks increasingly mandate verifiable content provenance. However, existing watermarking technique

safetyarxiv-cs-cv
4 Jun 2026
Applications

Ultra-Fast Neural Video Compression

DGX agent

arXiv:2606.04410v1 Announce Type: new Abstract: While neural video codecs (NVCs) have demonstrated superior compression ratio, their prohibitive computational complexity remains a critical barrier to

applicationsarxiv-cs-cv
4 Jun 2026
Research

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation

DGX agent

arXiv:2606.04264v1 Announce Type: new Abstract: Recent years have seen remarkable progress in unified vision-language models handling both multimodal understanding and generation within a single archi

researcharxiv-cs-cv
4 Jun 2026
Research

ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models

DGX agent

arXiv:2512.14099v3 Announce Type: replace Abstract: Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by conti

researcharxiv-cs-cv
4 Jun 2026
Research

Vision Transformer Finetuning Benefits from Non-Smooth Components

DGX agent

arXiv:2602.06883v3 Announce Type: replace-cross Abstract: The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversar

researcharxiv-cs-cv
4 Jun 2026
Safety

VT-3DAD: Cross-Category 3D Anomaly Detection via Visual-Text Normal Space Alignment

DGX agent

arXiv:2606.04369v1 Announce Type: new Abstract: Few-shot cross-category 3D anomaly detection aims to determine whether an unknown point cloud belongs to a target normal category using only a few norma

safetyarxiv-cs-cv
4 Jun 2026
Research

Weakly Supervised Incremental Segmentation via Semantic Anchors and Spatial Arbitration

DGX agent

arXiv:2606.04060v1 Announce Type: new Abstract: Weakly Incremental Learning for Semantic Segmentation (WILSS) suffers from the continuous introduction of noisy supervision, which progressively corrupt

researcharxiv-cs-cv
4 Jun 2026
Research

When Detectors Forget Forensics: Blocking Semantic Shortcuts for Generalizable AI-Generated Image Detection

DGX agent

arXiv:2603.09242v2 Announce Type: replace Abstract: The growing realism of generative models has blurred the boundary between real and synthetic content, posing significant challenges to reliable AI-g

researcharxiv-cs-cv
4 Jun 2026
Model Releases

When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection

DGX agent

arXiv:2606.04098v1 Announce Type: new Abstract: Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spli

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

XSSR: Cross-Domain Self-Supervised Representative Selection for Efficient Annotation in Medical Image Segmentation

DGX agent

arXiv:2606.04301v1 Announce Type: new Abstract: Acquiring labeled medical image data is resource-intensive and a challenge further exacerbated in cross-domain scenarios where source and target dataset

model-releasesarxiv-cs-cv
4 Jun 2026
Applications

Z-FLoc: Zero-Shot Floorplan Localization via Geometric Primitives

DGX agent

arXiv:2606.04788v1 Announce Type: new Abstract: Visual localization -- estimating a camera pose within a pre-existing map -- is a fundamental problem in computer vision. Floorplans are an attractive m

applicationsarxiv-cs-cv
4 Jun 2026
Research

ZipSplat: Fewer Gaussians, Better Splats

DGX agent

arXiv:2606.05102v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict o

researcharxiv-cs-cv
4 Jun 2026
Model Releases

A Benchmark for Semi-supervised Multi-modal Crowd Counting

DGX agent

arXiv:2606.03646v1 Announce Type: new Abstract: This paper constructs the first benchmark on semi-supervised multi-modal crowd counting. To lay the foundation for this unexplored task, we first formul

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

A Fast Methane Detection Pipeline on Board Satellites Based on Mag1c-SAS and LinkNet

DGX agent

arXiv:2606.03675v1 Announce Type: new Abstract: Methane is a potent greenhouse gas, and detecting leaks early via hyperspectral satellite imagery can help climate change mitigation efforts. Meanwhile,

model-releasesarxiv-cs-cv
3 Jun 2026
Local Ai

A unified multi-task framework enables interpretable chest radiograph analysis

DGX agent

arXiv:2606.03417v1 Announce Type: new Abstract: While multimodal deep learning has advanced medical imaging analysis, existing black-box systems extcolor{black}{may remain confined to isolated tasks,

local-aiarxiv-cs-cv
3 Jun 2026
Local Ai

A^2: Smaller Self-Supervised ViTs Localize Better than Larger Ones

DGX agent

arXiv:2606.03148v1 Announce Type: new Abstract: Robust visual classification often depends on localizing the main foreground objects in an image while ignoring contextual distractors. Surprisingly, we

local-aiarxiv-cs-cv
3 Jun 2026
Research

AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation

DGX agent

arXiv:2606.03972v1 Announce Type: new Abstract: We present AAD-1, an Asymmetric Adversarial Distillation framework for One-step autoregressive image-to-video generation. State-of-the-art methods adopt

researcharxiv-cs-cv
3 Jun 2026
Research

Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning

DGX agent

arXiv:2603.00667v3 Announce Type: replace Abstract: Computational pathology has advanced rapidly in recent years, driven by domain-specific image encoders and growing interest in using vision-language

researcharxiv-cs-cv
3 Jun 2026
Safety

Adaptive Causal Alignment for High-Confidence Adversarial Training

DGX agent

arXiv:2606.03925v1 Announce Type: new Abstract: Inverse adversarial training leverages high-confidence predictions to stabilize robust learning, yet we uncover a critical paradox: high confidence ofte

safetyarxiv-cs-cv
3 Jun 2026
Model Releases

AmbientEye: A Dataset for Pupil Segmentation under Natural Ambient Infrared Illumination

DGX agent

arXiv:2606.03774v1 Announce Type: new Abstract: Eye tracking is essential for smart glasses, as it provides insight into user attention for ambient intelligence applications. However, most existing ey

model-releasesarxiv-cs-cv
3 Jun 2026
Research

An Attention-Based Denoising Model for Diffusion Weighted Imaging

DGX agent

arXiv:2606.03903v1 Announce Type: new Abstract: Diffusion-weighted imaging (DWI) is used for whole-body cancer screening, but it typically requires a long acquisition time. When the scan time is reduc

researcharxiv-cs-cv
3 Jun 2026
Research

An Improved Method for Personalizing Diffusion Models

DGX agent

arXiv:2407.05312v2 Announce Type: replace Abstract: Diffusion models have demonstrated impressive image generation capabilities. Personalized approaches, such as textual inversion and Dreambooth, enha

researcharxiv-cs-cv
3 Jun 2026
Model Releases

Any2Poster: Any-Source Poster Generation Across Modalities and Domains

DGX agent

arXiv:2606.02915v1 Announce Type: new Abstract: Visual posters are a compact medium for communicating dense information, yet progress on automatic poster generation remains difficult to measure becaus

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation

DGX agent

arXiv:2606.03175v1 Announce Type: new Abstract: Instance Goal Navigation (IGN) requires an embodied agent to find a specific object instance among distractors from an underspecified natural-language d

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

ATLAS: A Large-Scale Evaluation Benchmark for Adversarial LiDAR Perception

DGX agent

arXiv:2606.02924v1 Announce Type: new Abstract: Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and pot

model-releasesarxiv-cs-cv
3 Jun 2026
Applications

Attend to Anything: Foundation Model for Unified Human Attention Modeling

DGX agent

arXiv:2606.03540v1 Announce Type: new Abstract: Existing human attention (saliency) modeling methods persist as highly fragmented across modalities, scenes, and task formulations. Consequently, even w

applicationsarxiv-cs-cv
3 Jun 2026
Local Ai

Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models

DGX agent

arXiv:2604.06052v2 Announce Type: replace Abstract: Text-to-image diffusion models exhibit remarkable generative capabilities, yet their internal operations remain opaque, particularly when handling p

local-aiarxiv-cs-cv
3 Jun 2026
Model Releases

Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging

DGX agent

arXiv:2606.02809v1 Announce Type: new Abstract: Evaluating vision-language models (VLMs) on medical images requires benchmarks that are clinically grounded, scalable, and controlled for evaluation con

model-releasesarxiv-cs-cv
3 Jun 2026
Research

AvatarMix: Identity-Preserving Cross-Avatar Composition for Outfit Personalization

DGX agent

arXiv:2606.03506v1 Announce Type: new Abstract: Existing 3D avatar outfit transfer methods face distinct challenges: approaches that lift 2D edits to 3D often suffer from outfit or identity quality de

researcharxiv-cs-cv
3 Jun 2026
Research

BA-T: An Iterative Transformer for Two-View Bundle Adjustment

DGX agent

arXiv:2606.03287v1 Announce Type: new Abstract: Feed-forward models for 3D reconstruction have achieved strong performance using deep cross-view attention to exchange information across images. Howeve

researcharxiv-cs-cv
3 Jun 2026
Research

BEAST3D: Animal behavioral analysis and neural encoding from multi-view video via Gaussian splatting

DGX agent

arXiv:2606.02937v1 Announce Type: cross Abstract: Multi-view video recordings are increasingly used to capture the 3D movements of animals in experimental settings, yet extracting rich 3D representati

researcharxiv-cs-cv
3 Jun 2026
Model Releases

Benchmarking Visual State Tracking in Multimodal Video Understanding

DGX agent

arXiv:2606.03920v1 Announce Type: new Abstract: Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacit

model-releasesarxiv-cs-cv
3 Jun 2026
Safety

Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

DGX agent

arXiv:2602.12221v2 Announce Type: replace Abstract: We propose UniDFlow, a unified discrete flow-matching framework for multimodal understanding, generation, and editing. It decouples understanding an

safetyarxiv-cs-cv
3 Jun 2026
Research

Beyond Compression: Quantifying Spectral Accessibility in Vision Representations

DGX agent

arXiv:2606.03795v1 Announce Type: new Abstract: Vision-language models map visual features into a shared embedding space through learned projection layers, yet it remains unclear how these transformat

researcharxiv-cs-cv
3 Jun 2026
Research

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models

DGX agent

arXiv:2606.03730v1 Announce Type: new Abstract: Vision-language models (VLMs) such as CLIP show strong zero-shot generalization but remain highly vulnerable to adversarial attacks. Adversarial trainin

researcharxiv-cs-cv
3 Jun 2026
Research

Beyond Single Solution: Multi-Hypothesis Collaborative Deep Unfolding Network for Image Compressive Sensing

DGX agent

arXiv:2606.03666v1 Announce Type: new Abstract: Recent deep unfolding networks (DUNs) have advanced Compressive Sensing (CS) by effectively integrating iterative optimization with deep learning archit

researcharxiv-cs-cv
3 Jun 2026
Research

Bootstrap Your Generator: Unpaired Visual Editing with Flow Matching

DGX agent

arXiv:2606.03911v1 Announce Type: new Abstract: Modern generative models possess a deep understanding of visual content, yet training them for image editing typically requires massive datasets of pair

researcharxiv-cs-cv
3 Jun 2026
Research

BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks

DGX agent

arXiv:2606.02947v1 Announce Type: cross Abstract: Supervised fine-tuning is the predominant approach for adapting autoregressive vision-language models to downstream tasks. Recent work has shown that

researcharxiv-cs-cv
3 Jun 2026
Research

CAD-to-CT Registration of Cylindrical Objects via Ellipse-Based Axis Estimation

DGX agent

arXiv:2606.02935v1 Announce Type: new Abstract: Accurate registration of CAD models to CT scans is essential for establishing ground truth geometry in volumetric imaging. Obtaining reliable object mas

researcharxiv-cs-cv
3 Jun 2026
Model Releases

Characterizing Detectability in 3DGS Poisoning: A Stage-wise Benchmark

DGX agent

arXiv:2606.03499v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has rapidly emerged as a leading representation for real-time novel view synthesis, but recent work shows it is vulnerable

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

COD10K-C: Benchmarking Robustness of Camouflaged Object Detection Under Natural Image Corruptions

DGX agent

arXiv:2606.02603v1 Announce Type: new Abstract: Camouflaged object detection has improved substantially, but most standard benchmarks evaluate models only on clean images. This is not realistic becaus

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Consistent Yet Wrong: Evidence Insensitivity in Spatial Vision-Language Models

DGX agent

arXiv:2606.02742v1 Announce Type: new Abstract: Spatial reasoning is fundamental to robotics, autonomy, and embodied AI, yet modern vision-language models (VLMs) remain unreliable on metric distance q

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

CoralBay: A Self-Supervised CT Foundation Model

DGX agent

arXiv:2606.03888v1 Announce Type: new Abstract: Self-supervised learning has enabled large-scale pre-training on 2D natural images, producing general-purpose visual representations that transfer effec

model-releasesarxiv-cs-cv
3 Jun 2026
← Previous
1…121122123124125…263
Next →