AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Accelerating Merge with Motion Vector Difference via Filter Difference Analysis for VVenC

DGX agent

arXiv:2606.31084v1 Announce Type: cross Abstract: Merge with Motion Vector Difference (MMVD) is a key coding tool in Versatile Video Coding for improving motion prediction accuracy. However, its exhau

researcharxiv-cs-cv
1 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

AeroVerse-SatAgent: UAV-Satellite Collaborative Spatial Reasoning Inspired by the Dual Visual Pathway Theory of Cognitive Neuroscience

DGX agent

arXiv:2606.31467v1 Announce Type: new Abstract: With the rapid advancement of aerospace embodied intelligence, enabling Unmanned Aerial Vehicles (UAVs) to autonomously understand and reason about comp

safetyarxiv-cs-cv
1 Jul 2026
Safety

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

DGX agent

arXiv:2606.19100v2 Announce Type: replace Abstract: Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open-

safetyarxiv-cs-cv
1 Jul 2026
Safety

Anchoring on Reality: Breaking the Pseudo-Target Ceiling in Makeup Transfer

DGX agent

arXiv:2606.31089v1 Announce Type: new Abstract: Makeup transfer applies a reference cosmetic style to a source face while preserving its identity and geometry. However, this task is severely hindered

safetyarxiv-cs-cv
1 Jul 2026
Applications

AnyBokeh: Physics-Guided Any-to-Any Bokeh Editing with Optical Fingerprint Transfer

DGX agent

arXiv:2606.31959v1 Announce Type: new Abstract: Depth-of-field control is a fundamental tool in photography, yet post-capture bokeh editing from a single image remains challenging. A practical editor

applicationsarxiv-cs-cv
1 Jul 2026
Applications

AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images

DGX agent

arXiv:2606.31077v1 Announce Type: new Abstract: Multi-modal image matching is essential for visual localization and multi-sensor fusion, but it is hindered by the scarcity of large-scale training data

applicationsarxiv-cs-cv
1 Jul 2026
Safety

AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization

DGX agent

arXiv:2603.17461v2 Announce Type: replace Abstract: Streaming autoregressive (AR) video generators combined with few-step distillation achieve low-latency, high-quality synthesis, yet remain difficult

safetyarxiv-cs-cv
1 Jul 2026
Research

Auditing Generalization in AI-Generated Video Detection: A Six-Control Protocol and the VidAudit Toolkit

DGX agent

arXiv:2606.31004v1 Announce Type: new Abstract: AI-generated video detection benchmarks such as GenVidBench and AIGVDBench are the de facto leaderboards, yet most evaluation protocols leave uncontroll

researcharxiv-cs-cv
1 Jul 2026
Research

AugSplat: Radiance Field-Informed Gaussian Splatting for Sparse-View Settings

DGX agent

arXiv:2606.31556v1 Announce Type: new Abstract: Generating high-quality novel views at real-time frame rates remains a central challenge in 3D vision, particularly in sparse-view scenarios. Neural rad

researcharxiv-cs-cv
1 Jul 2026
Research

Automated Background Swapping for Robustness against Spurious Backgrounds

DGX agent

arXiv:2606.32018v1 Announce Type: new Abstract: Classifiers based on Deep Neural Networks exhibit strong performance across domains, yet can fail catastrophically if they rely on spurious correlations

researcharxiv-cs-cv
1 Jul 2026
Safety

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

DGX agent

arXiv:2606.30811v1 Announce Type: new Abstract: Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality sounding video content with fine-grained

safetyarxiv-cs-cv
1 Jul 2026
Model Releases

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

DGX agent

arXiv:2606.31169v1 Announce Type: new Abstract: Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form

model-releasesarxiv-cs-cv
1 Jul 2026
Safety

Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration

DGX agent

arXiv:2601.19506v4 Announce Type: replace Abstract: Blind face restoration remains a persistent challenge due to the inherent ill-posedness of reconstructing holistic structures from severely constrai

safetyarxiv-cs-cv
1 Jul 2026
Research

Bridging Video Understanding and Generation in a Unified Framework

DGX agent

arXiv:2606.31326v1 Announce Type: new Abstract: Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video dom

researcharxiv-cs-cv
1 Jul 2026
Applications

Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning

DGX agent

arXiv:2510.14819v3 Announce Type: replace Abstract: Trajectory representation learning (TRL) aims to encode raw trajectory data into low-dimensional embeddings for downstream tasks such as travel time

applicationsarxiv-cs-cv
1 Jul 2026
Safety

CasaMaestro: Multi-View Panoramas for House-Scale 3D Reconstruction

DGX agent

arXiv:2606.31086v1 Announce Type: new Abstract: The rise of home-deployed embodied AI systems is driving a growing need for fast, metric 3D reconstruction of residential spaces to support navigation,

safetyarxiv-cs-cv
1 Jul 2026
Model Releases

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts

DGX agent

arXiv:2606.31986v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning has enabled multi-modal large language models (MLLMs) to tackle complex visual reasoning tasks by generating explicit i

model-releasesarxiv-cs-cv
1 Jul 2026
Research

CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation

DGX agent

arXiv:2606.32012v1 Announce Type: cross Abstract: Uncertainty estimation has been a long-standing challenge in AI models; it amounts to 'knowing what you don't know,' and metacognition is notoriously

researcharxiv-cs-cv
1 Jul 2026
Tutorials

CoMNet: A MedNeXt-CorrDiff Framework for Multi-Site Brain Tumor Segmentation

DGX agent

arXiv:2606.15305v2 Announce Type: replace Abstract: Accurate brain tumor segmentation from multiparametric magnetic resonance imaging (MRI) is critical for treatment planning, response assessment, and

tutorialsarxiv-cs-cv
1 Jul 2026
Model Releases

CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization

DGX agent

arXiv:2606.31219v1 Announce Type: new Abstract: Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However,

model-releasesarxiv-cs-cv
1 Jul 2026
Research

Cross-Resolution Distribution Matching for Diffusion Distillation

DGX agent

arXiv:2603.06136v2 Announce Type: replace Abstract: Diffusion distillation is central to accelerating image and video generation, yet existing methods are fundamentally limited by the denoising proces

researcharxiv-cs-cv
1 Jul 2026
Safety

Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers

DGX agent

arXiv:2606.32020v1 Announce Type: new Abstract: Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, they rely on a critical assumption: T

safetyarxiv-cs-cv
1 Jul 2026
Model Releases

DANTE-W: Diffuse Albedo Neural Texturing in the Wild

DGX agent

arXiv:2606.30677v1 Announce Type: cross Abstract: Classical mesh texturing techniques blend captured multi-view images directly, which inevitably suffer from baked-in shading and casted shadows that c

model-releasesarxiv-cs-cv
1 Jul 2026
Safety

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation

DGX agent

arXiv:2606.31537v1 Announce Type: new Abstract: Text-rich image generation is one of the most challenging settings in image generation, since models must simultaneously produce visually realistic imag

safetyarxiv-cs-cv
1 Jul 2026
Applications

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning

DGX agent

arXiv:2606.31257v1 Announce Type: new Abstract: The standard way to read latent knowledge out of a model, a linear probe confirmed by a steering recovery, can systematically overstate what a vision-la

applicationsarxiv-cs-cv
1 Jul 2026
Applications

Deep Spectral Models for Robust Dental Shape Generation

DGX agent

arXiv:2606.31293v1 Announce Type: new Abstract: Accurate modeling of dental crown morphology is fundamental for diagnosis, orthodontic planning, and computer-aided restoration design. However, dataset

applicationsarxiv-cs-cv
1 Jul 2026
Research

DEMUN: Fast and accurate discovery of music notation in very large collections

DGX agent

arXiv:2606.31956v1 Announce Type: cross Abstract: Much of written musical heritage is preserved and digitised at memory institutions: libraries, museums, and archives. Owing to their collection struct

researcharxiv-cs-cv
1 Jul 2026
Local Ai

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

DGX agent

arXiv:2606.31007v1 Announce Type: new Abstract: Vision foundation models such as SAM 3 can provide transferable object-level structure across diverse surgical video conditions, but segmentation output

local-aiarxiv-cs-cv
1 Jul 2026
Research

DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection

DGX agent

arXiv:2603.23455v2 Announce Type: replace Abstract: Multi-Modal LLMs (MLLMs) demonstrate strong visual grounding capabilities on popular object detection benchmarks like OdinW-13 and RefCOCO. However,

researcharxiv-cs-cv
1 Jul 2026
Research

Diffusion-Based Material Regularization for Physics-Based Inverse Rendering

DGX agent

arXiv:2606.31065v1 Announce Type: new Abstract: Reconstructing physics-based 3D assets -- geometry, materials, and illumination -- from multi-view images is a core problem in computer graphics and vis

researcharxiv-cs-cv
1 Jul 2026
Model Releases

Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation

DGX agent

arXiv:2606.20196v2 Announce Type: replace Abstract: Continual Test-Time Adaptation (CTTA) aims to maintain model performance under evolving target domains by adapting online without labeled data. Howe

model-releasesarxiv-cs-cv
1 Jul 2026
Research

Distortion-Corrected Diffusion MRI Using Rotated-View EPI and Joint Field-Map/Image Estimation with Gaussian Primitives

DGX agent

arXiv:2606.31521v1 Announce Type: cross Abstract: Echo Planar Imaging (EPI) is the standard acquisition technique for diffusion and functional neuroimaging, enabling rapid imaging but suffering from g

researcharxiv-cs-cv
1 Jul 2026
Research

Do Not Break the Vessels: Structure-Preserving Mean Flow for Vascular Image Translation

DGX agent

arXiv:2606.31095v1 Announce Type: new Abstract: Reconstructing anatomically faithful vascular structures from clinically accessible imaging modalities is of substantial clinical significance. However,

researcharxiv-cs-cv
1 Jul 2026
Research

Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization

DGX agent

arXiv:2606.31373v1 Announce Type: new Abstract: Cycle self-training (CST) breaks the shared classifier assumption of the standard self-training framework, which is effective for unsupervised domain ad

researcharxiv-cs-cv
1 Jul 2026
Agents

DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation

DGX agent

arXiv:2606.31918v1 Announce Type: new Abstract: A pivotal step in autonomous driving simulation involves inserting foreground vehicles with predefined trajectories into simulated scenes. This process

agentsarxiv-cs-cv
1 Jul 2026
Agents

DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation

DGX agent

arXiv:2606.31488v1 Announce Type: new Abstract: Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry witho

agentsarxiv-cs-cv
1 Jul 2026
Research

Drop-In Perceptual Optimization for 3D Gaussian Splatting

DGX agent

arXiv:2603.23297v2 Announce Type: replace Abstract: Despite their output being ultimately consumed by human viewers, 3D Gaussian Splatting (3DGS) methods often rely on ad-hoc combinations of pixel-lev

researcharxiv-cs-cv
1 Jul 2026
Model Releases

Dual Sparse Aggregation Transformer for Multispectral Object Detection

DGX agent

arXiv:2606.31015v1 Announce Type: new Abstract: Transformer-based approaches have obtained excellent performance in multispectral object detection tasks due to their ability to model long-range depend

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments

DGX agent

arXiv:2606.31654v1 Announce Type: cross Abstract: Recent advances in multimodal large models have significantly improved UAV vision-language navigation (UAV-VLN) by enhancing high-level perception and

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes

DGX agent

arXiv:2604.04834v2 Announce Type: replace Abstract: Robotic Vision-Language-Action (VLA) models generalize well for open-ended manipulation, but their perception is fragile under sensing-stage degrada

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

Editing Everything Everywhere All at Once

DGX agent

arXiv:2606.31278v1 Announce Type: new Abstract: Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency

model-releasesarxiv-cs-cv
1 Jul 2026
Applications

EgoCogNav: Cognition-aware Human Egocentric Navigation

DGX agent

arXiv:2511.17581v3 Announce Type: replace-cross Abstract: Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction

applicationsarxiv-cs-cv
1 Jul 2026
Safety

EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning

DGX agent

arXiv:2511.18242v3 Announce Type: replace Abstract: Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal la

safetyarxiv-cs-cv
1 Jul 2026
Research

EpiMask: Leveraging Epipolar Distance Based Masks in Cross-Attention for Satellite Image Matching

DGX agent

arXiv:2603.21463v2 Announce Type: replace Abstract: The deep-learning based image matching networks can now handle significantly larger variations in viewpoints and illuminations while providing match

researcharxiv-cs-cv
1 Jul 2026
Safety

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs

DGX agent

arXiv:2606.31982v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) incur prohibitive inference costs due to long visual token sequences. Training-free visual token reduction prov

safetyarxiv-cs-cv
1 Jul 2026
Research

Estimating Velocity of Spheres from Rolling-Shutter Image(s)

DGX agent

arXiv:2606.31760v1 Announce Type: new Abstract: Rolling-shutter cameras introduce characteristic distortions when imaging fast moving objects, and these effects are typically treated as artifacts to b

researcharxiv-cs-cv
1 Jul 2026
Local Ai

Event-Driven Video Generation

DGX agent

arXiv:2603.13402v3 Announce Type: replace Abstract: Current text-to-video models can make individual frames look convincing while still getting simple interactions wrong: objects move before contact,

local-aiarxiv-cs-cv
1 Jul 2026
Model Releases

Evidence Triangulation for Multimodal Fact-Checking in the Wild

DGX agent

arXiv:2606.31367v1 Announce Type: cross Abstract: The proliferation of multimedia content on social platforms has fueled multimodal misinformation, where images are used to reinforce false claims. Con

model-releasesarxiv-cs-cv
1 Jul 2026
← Previous
1…7374757677…263
Next →