AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Model Releases

Addressable Memory for Video World Models

DGX agent

arXiv:2608.07408v1 Announce Type: new Abstract: We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward p

model-releasesarxiv-cs-cv
10 Aug 2026
Applications

AdvTiles: Physical Adversarial Camouflage Clothing against Person Detectors via Learnable Tiles

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2608.06801v1 Announce Type: new Abstract: Physical adversarial attacks against person detectors have evolved from localized patches to full-body textures. However, achieving both visual naturaln

applicationsarxiv-cs-cv
10 Aug 2026
Safety

An AI4AI Framework for Visual Token Pruning

DGX agent

arXiv:2608.07193v1 Announce Type: cross Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fix

safetyarxiv-cs-cv
10 Aug 2026
Research

AnyTrack: Unifying Visual Object Tracking with Any Modalities

DGX agent

arXiv:2608.06773v1 Announce Type: new Abstract: Visual object tracking aims to continuously locate specific targets within sequential frames, evolving from single-modal methods to multi-modal ones. Ho

researcharxiv-cs-cv
10 Aug 2026
Research

Are Visual Place Recognition Models Recognizing Places or Conditions? Distractor-Augmented Evaluation and Condition Suppression

DGX agent

arXiv:2608.06847v1 Announce Type: cross Abstract: Long-term Visual Place Recognition (VPR) is typically evaluated by matching queries from one condition against a database from another. Crowdsourced m

researcharxiv-cs-cv
10 Aug 2026
Applications

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

DGX agent

arXiv:2608.06729v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially ob

applicationsarxiv-cs-cv
10 Aug 2026
Model Releases

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

DGX agent

arXiv:2608.06930v1 Announce Type: new Abstract: Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main

model-releasesarxiv-cs-cv
10 Aug 2026
Research

Bend the Basics: Degradation-Aware Deformable Tokenization for All-in-One Image Restoration

DGX agent

arXiv:2608.06832v1 Announce Type: new Abstract: All-in-one image restoration seeks a single model that can recover images degraded by diverse and spatially non-uniform corruptions. However, many unifi

researcharxiv-cs-cv
10 Aug 2026
Model Releases

Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation

DGX agent

arXiv:2608.07117v1 Announce Type: new Abstract: Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for

model-releasesarxiv-cs-cv
10 Aug 2026
Tutorials

C2Dex: Contact-Consistent Reconstruction and Retargeting for Dexterous Manipulation from Monocular Video

DGX agent

arXiv:2608.07045v1 Announce Type: cross Abstract: High-quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocular human videos provide a scalable so

tutorialsarxiv-cs-cv
10 Aug 2026
Model Releases

CADSpotting: Robust Panoptic Symbol Spotting on Large-Scale CAD Drawings

DGX agent

arXiv:2412.07377v5 Announce Type: replace Abstract: We introduce CADSpotting, an effective method for panoptic symbol spotting in large-scale architectural CAD drawings. Existing approaches often stru

model-releasesarxiv-cs-cv
10 Aug 2026
Applications

CANIS: Generation-Assisted 3D Canonicalization via an Image-Semantic Bridge

DGX agent

arXiv:2608.07256v1 Announce Type: new Abstract: Canonicalizing 3D object orientation is fundamental to 3D understanding and analysis. Existing approaches often rely on geometric cues, although 3D cano

applicationsarxiv-cs-cv
10 Aug 2026
Research

CAS2UML: A Handwritten Sketch-to-PlantUML Dataset for Class and Activity Diagrams

DGX agent

arXiv:2608.07036v1 Announce Type: cross Abstract: Automated UML generation from sketches and images is gaining renewed attention with the rise of large language models and multimodal AI. However, repr

researcharxiv-cs-cv
10 Aug 2026
Applications

Casting the Net! Revisiting MasterFace Impersonation Attacks

DGX agent

arXiv:2608.06952v1 Announce Type: cross Abstract: Impersonation is a fundamental security threat in face recognition systems (FRSs). While the security of FRSs has been challenged by various attack ve

applicationsarxiv-cs-cv
10 Aug 2026
Agents

CloudDiffusion: Diffusion-Based Scene Completion in the Point Cloud Domain

DGX agent

arXiv:2606.16048v3 Announce Type: replace Abstract: Reconstructing dense 3D scenes from sparse LiDAR point clouds (LiDAR scene completion) is a fundamental challenge in autonomous driving, where diffu

agentsarxiv-cs-cv
10 Aug 2026
Model Releases

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

DGX agent

arXiv:2608.06691v1 Announce Type: new Abstract: Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency

model-releasesarxiv-cs-cv
10 Aug 2026
Local Ai

Conformal Coverage Guarantees for Any Video Temporal Grounder

DGX agent

arXiv:2608.07434v1 Announce Type: new Abstract: Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less t

local-aiarxiv-cs-cv
10 Aug 2026
Local Ai

ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE

DGX agent

arXiv:2608.06878v1 Announce Type: new Abstract: Layout-guided multi-instance generation is essential for controllable image synthesis in Multi-Modal Diffusion Transformers (MM-DiTs). However, integrat

local-aiarxiv-cs-cv
10 Aug 2026
Safety

Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers

DGX agent

arXiv:2608.06674v1 Announce Type: new Abstract: Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in

safetyarxiv-cs-cv
10 Aug 2026
Safety

DA-Cal: Towards Cross-Domain Calibration in Semantic Segmentation

DGX agent

arXiv:2602.20860v2 Announce Type: replace Abstract: While existing unsupervised domain adaptation (UDA) methods greatly enhance target domain performance in semantic segmentation, they often neglect n

safetyarxiv-cs-cv
10 Aug 2026
Model Releases

Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery

DGX agent

arXiv:2608.06406v1 Announce Type: new Abstract: Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosys

model-releasesarxiv-cs-cv
10 Aug 2026
Safety

Degradation-Aware Prompt Learning with Cross-Modal Compensation for Adverse Weather Removal

DGX agent

arXiv:2608.06939v1 Announce Type: new Abstract: Adverse weather causes diverse and complex image degradations, severely compromising the reliability of computer vision systems. Existing all-in-one res

safetyarxiv-cs-cv
10 Aug 2026
Research

Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model

DGX agent

arXiv:2608.07361v1 Announce Type: cross Abstract: Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself

researcharxiv-cs-cv
10 Aug 2026
Applications

Direct Visual Grounding by Directing Attention of Visual Tokens

DGX agent

arXiv:2511.12738v2 Announce Type: replace Abstract: Vision Language Models (VLMs) mix visual tokens and text tokens. A puzzling issue is the fact that visual tokens most related to the query receive l

applicationsarxiv-cs-cv
10 Aug 2026
Research

DREAMS: Diverse Reactions of Engagement and Attention Mind States Dataset

DGX agent

arXiv:2608.06382v1 Announce Type: cross Abstract: Active attention and engagement are important in improving users' learning experiences. Engagement refers to the level of involvement and interest ind

researcharxiv-cs-cv
10 Aug 2026
Safety

Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification

DGX agent

arXiv:2608.06943v1 Announce Type: new Abstract: Cross-modal Re-Identification (ReID) aims to retrieve the same identity across heterogeneous imaging modalities and has been widely studied in visible-i

safetyarxiv-cs-cv
10 Aug 2026
Model Releases

ECAD: Expanding Class-Agnostic Detection Beyond Thing-Centric Objectness

DGX agent

arXiv:2608.06841v1 Announce Type: new Abstract: Object detection is a fundamental task in visual perception, providing structured region representations for recognition, grounding, reasoning, and inte

model-releasesarxiv-cs-cv
10 Aug 2026
Research

ELMZip: Onboard Satellite Image Compression via Extreme Learning Machines for Efficient Downlink

DGX agent

arXiv:2608.06942v1 Announce Type: cross Abstract: The acquisition of multispectral imagery via small satellites (e.g., CubeSats) presents significant data downlink challenges due to high data volumes

researcharxiv-cs-cv
10 Aug 2026
Model Releases

Explanation Stability of Test-Time Adaptation in Computational Pathology: A Large-Scale Benchmark

DGX agent

arXiv:2608.07062v1 Announce Type: new Abstract: Test-time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in comput

model-releasesarxiv-cs-cv
10 Aug 2026
Safety

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models

DGX agent

arXiv:2608.06768v1 Announce Type: new Abstract: Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distr

safetyarxiv-cs-cv
10 Aug 2026
Research

Flow-Corrected Shape Optimization: Taming Manifold Drift in High-Dimensional 3D Models

DGX agent

arXiv:2608.07199v1 Announce Type: new Abstract: Optimizing 3D shapes within the latent spaces of deep generative models is fundamental to computer assisted engineering, yet remains prone to a critical

researcharxiv-cs-cv
10 Aug 2026
Research

Foundation Models Adaptation for Multi-View Multi-modal Cardiac MRI Segmentation and Direct Ejection Fraction Estimation

DGX agent

arXiv:2608.07291v1 Announce Type: new Abstract: Foundation models have shown strong transferability in cardiac MRI (CMR), but their effectiveness for heterogeneous multi-view and multi-sequence CMR an

researcharxiv-cs-cv
10 Aug 2026
Model Releases

Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?

DGX agent

arXiv:2608.06972v1 Announce Type: new Abstract: Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess rep

model-releasesarxiv-cs-cv
10 Aug 2026
Safety

GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes

DGX agent

arXiv:2608.06836v1 Announce Type: new Abstract: We study the problem of inserting new furniture into indoor scene images. Under masked single-view 2D image-plane conditioning, however, the physical sc

safetyarxiv-cs-cv
10 Aug 2026
Safety

Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP

DGX agent

arXiv:2502.18816v3 Announce Type: replace Abstract: Significant progress has been achieved on the improvement and downstream usages of the Contrastive Language-Image Pre-training (CLIP) vision-languag

safetyarxiv-cs-cv
10 Aug 2026
Model Releases

GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models

DGX agent

arXiv:2608.06769v1 Announce Type: new Abstract: Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision-language tasks, creating an urgent need for more

model-releasesarxiv-cs-cv
10 Aug 2026
Local Ai

HazeSpikeMamba: Coupling Spiking-Inspired and State-Space Features for Self-Supervised Real-World Dehazing

DGX agent

arXiv:2608.06886v1 Announce Type: new Abstract: Dehazing networks are commonly trained on synthetic hazy-clear pairs, but their performance often drops on real photographs. Synthetic haze generated us

local-aiarxiv-cs-cv
10 Aug 2026
Research

HRDiT: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models

DGX agent

arXiv:2608.07003v1 Announce Type: new Abstract: Training-free text-to-high-resolution image generation has recently attracted growing research attention. However, existing studies on this task primari

researcharxiv-cs-cv
10 Aug 2026
Local Ai

Human-AI Perceptual Alignment by Playing Hues and Cues

DGX agent

arXiv:2608.07141v1 Announce Type: new Abstract: Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks tha

local-aiarxiv-cs-cv
10 Aug 2026
Research

IceHorizon: A Dataset for Horizon Detection in Ice-Covered Maritime Environments and Comparative Evaluation of Detection Methods

DGX agent

arXiv:2608.07018v1 Announce Type: cross Abstract: Horizon detection in images of ice-covered waters is a challenging problem for maritime navigation due to low contrast between water and sky, cluttere

researcharxiv-cs-cv
10 Aug 2026
Safety

Identity as Presence: Towards Appearance and Voice Personalized Joint Audio-Video Generation

DGX agent

arXiv:2603.17889v4 Announce Type: replace Abstract: Recent advances in video synthesis have enabled realistic integration of real individuals, driving demand for identity-aware generation. While emerg

safetyarxiv-cs-cv
10 Aug 2026
Research

Implicit Neural Speckle Denoising

DGX agent

arXiv:2608.06574v1 Announce Type: cross Abstract: Speckle fundamentally limits coherent imaging by introducing multiplicative, spatially correlated noise that obscures scene structure. Removing speckl

researcharxiv-cs-cv
10 Aug 2026
Applications

Improving Low-Resolution Face Recognition under Limited Data: How Synthetic Data Generation Can Close the Domain Gap

DGX agent

arXiv:2608.06580v1 Announce Type: new Abstract: Face Recognition (FR) systems in surveillance settings often encounter Low Resolution (LR) faces, those whose face region falls below the standard 112 i

applicationsarxiv-cs-cv
10 Aug 2026
Model Releases

InsertFuse: A Unified Framework for Multi-Category Reference-Guided Image Insertion

DGX agent

arXiv:2608.06490v1 Announce Type: new Abstract: We present InsertFuse, a unified framework for multi-category reference-guided image insertion. Its key idea is to decouple category-specific expertise

model-releasesarxiv-cs-cv
10 Aug 2026
Tutorials

InstanceSplat: Instance-Aware Feed-Forward 3D Gaussian Splatting for Scene Understanding

DGX agent

arXiv:2608.07144v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) enables efficient and generalizable 3D reconstruction, but current feed-forward 3DGS methods for scene underst

tutorialsarxiv-cs-cv
10 Aug 2026
Safety

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

DGX agent

arXiv:2608.06799v1 Announce Type: cross Abstract: Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-

safetyarxiv-cs-cv
10 Aug 2026
Safety

KnifeHunter: Structured Local Representation Learning for Fine-Grained Knife Image Retrieval in Law Enforcement

DGX agent

arXiv:2608.07057v1 Announce Type: new Abstract: Knife-enabled violence presents a major public safety challenge, and law enforcement agencies require scalable tools for catalogue-level knife identific

safetyarxiv-cs-cv
10 Aug 2026
Applications

Learning Ordinal Degradation Representations with Textual Priors for Diffusion-Based Blind Image Super-Resolution

DGX agent

arXiv:2512.10340v2 Announce Type: replace Abstract: Blind image super-resolution (Blind SR) has achieved remarkable perceptual quality via generative priors. However, lacking clear degradation represe

applicationsarxiv-cs-cv
10 Aug 2026
← Previous
1…7891011…259
Next →