AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

CouCE: A Unified Causal Framework for Debiased Deep Metric Learning

DGX agent

arXiv:2606.30365v1 Announce Type: new Abstract: Deep Metric Learning (DML) often struggles with zero-shot generalization because standard objectives inherently capture what co-occurs rather than what

researcharxiv-cs-cv
30 Jun 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Cross-Modal Iteration Distillation for Robust IHD Screening: The IDNet Framework and A New Benchmark

DGX agent

arXiv:2606.30027v1 Announce Type: new Abstract: Color Fundus Photography (CFP) offers a low-cost and non-invasive route for ischemic heart disease (IHD) screening, but current studies are limited by s

model-releasesarxiv-cs-cv
30 Jun 2026
Applications

Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance

DGX agent

arXiv:2606.30458v1 Announce Type: new Abstract: Text-to-image person re-identification (TIPR) retrieves target persons using natural language descriptions. However, existing methods largely overlook r

applicationsarxiv-cs-cv
30 Jun 2026
Research

CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation

DGX agent

arXiv:2604.02948v2 Announce Type: replace Abstract: Multimodal semantic segmentation has shown great potential in leveraging complementary information across diverse sensing modalities. However, exist

researcharxiv-cs-cv
30 Jun 2026
Model Releases

CylindTrack: Depth-Aware Cylindrical Motion Modeling for Panoramic Multi-Object Tracking

DGX agent

arXiv:2606.30097v1 Announce Type: new Abstract: Multi-Object Tracking (MOT) is a core capability for embodied perception, and panoramic cameras are attractive for embodied systems because their 360{eg

model-releasesarxiv-cs-cv
30 Jun 2026
Applications

D^{2}R^{2}OSR: Degradation-Disentangled Representation for Real-World Omnidirectional Image Super-Resolution

DGX agent

arXiv:2606.29314v1 Announce Type: new Abstract: With the growing demand for immersive visual experiences, high-quality omnidirectional images (ODIs) have become increasingly important. However, limita

applicationsarxiv-cs-cv
30 Jun 2026
Research

DCGrasp: Distance-aware Controllable Grasp Generation

DGX agent

arXiv:2606.29924v1 Announce Type: new Abstract: Generating 3D hand-object interactions is essential for applications in robotics, XR, and synthetic data generation, where flexible controllability and

researcharxiv-cs-cv
30 Jun 2026
Research

DCSNet: Multiscale Feature Aggregation for Small Medical Object Segmentation with Detection-guided Hierarchical Cropping

DGX agent

arXiv:2606.28402v1 Announce Type: new Abstract: Small object segmentation in medical imaging is primarily hindered by class imbalance and inherent boundary complexity. Consequently, conventional globa

researcharxiv-cs-cv
30 Jun 2026
Research

DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering

DGX agent

arXiv:2602.19323v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful paradigm for real-time and high-fidelity 3D reconstruction from posed images. However, recent

researcharxiv-cs-cv
30 Jun 2026
Safety

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation

DGX agent

arXiv:2512.20117v2 Announce Type: replace Abstract: Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by integrating auditory and visual cues. However, existi

safetyarxiv-cs-cv
30 Jun 2026
Research

DeVAR: Low-Dose CT Denoising via Visual Autoregressive Modeling

DGX agent

arXiv:2606.28453v1 Announce Type: cross Abstract: Computed tomography (CT) plays a crucial role in medical diagnosis, but minimizing radiation exposure while maintaining image quality remains a critic

researcharxiv-cs-cv
30 Jun 2026
Research

DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent

DGX agent

arXiv:2606.28417v1 Announce Type: new Abstract: Recently, diffusion models have been widely adopted in generative modeling and have served as foundational models for many image generation tasks. To co

researcharxiv-cs-cv
30 Jun 2026
Safety

Distribution Matching Variational AutoEncoder

DGX agent

arXiv:2512.07778v2 Announce Type: replace Abstract: Most visual generative models compress images into a latent space before applying diffusion or autoregressive modelling. Yet, existing approaches su

safetyarxiv-cs-cv
30 Jun 2026
Research

DivAS: Interactive 3D Segmentation by Depth-Weighted Voxel Aggregation

DGX agent

arXiv:2601.04860v2 Announce Type: replace Abstract: Interactive 3D segmentation of a reconstructed scene should not require a representation-specific optimization loop. We observe that the recipe for

researcharxiv-cs-cv
30 Jun 2026
Research

DLGStream: Dynamic Language-embedded Guassian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming

DGX agent

arXiv:2606.28840v1 Announce Type: new Abstract: 3D Gaussian Splatting~(3DGS) has emerged as a promising paradigm for reconstructing streamable free-viewpoint video~(FVV) from multi-view videos. Howeve

researcharxiv-cs-cv
30 Jun 2026
Hardware

DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model

DGX agent

arXiv:2606.30292v1 Announce Type: cross Abstract: We present DreamForge-World 0.1 Preview, a preview foundational world model for real-time interactive world simulation. The system adapts the LongLive

hardwarearxiv-cs-cv
30 Jun 2026
Research

DRESS: Disentangled Representation-based Self-Supervised Meta-Learning for Diverse Tasks

DGX agent

arXiv:2503.09679v2 Announce Type: replace-cross Abstract: Meta-learning represents a strong class of approaches for solving few-shot learning tasks. Nonetheless, recent research suggests that simply p

researcharxiv-cs-cv
30 Jun 2026
Research

Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting

DGX agent

arXiv:2603.28049v3 Announce Type: replace Abstract: Autoregressive (AR)-Diffusion hybrid paradigms combine AR's structured semantic modeling with diffusion's high-fidelity synthesis, yet suffer from a

researcharxiv-cs-cv
30 Jun 2026
Safety

DrivenMorph: Bridging Attention Mechanism and Variational Image Registration via Difference Modeling

DGX agent

arXiv:2606.30183v1 Announce Type: new Abstract: Medical image registration benefits significantly from deep learning, yet existing approaches often lack physical explainability and fine-grained deform

safetyarxiv-cs-cv
30 Jun 2026
Safety

DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution

DGX agent

arXiv:2606.29198v1 Announce Type: new Abstract: As the most perceptually powerful Face Video Super-Resolution (FVSR) method, existing works in Generative FVSR (GFVSR) mainly exploit the generative pri

safetyarxiv-cs-cv
30 Jun 2026
Research

Dynamic High-frequency Convolution for Infrared Small Target Detection

DGX agent

arXiv:2602.02969v2 Announce Type: replace Abstract: Infrared small targets are typically tiny and locally salient, which belong to high-frequency components (HFCs) in images. Single-frame infrared sma

researcharxiv-cs-cv
30 Jun 2026
Model Releases

Early Estimation of Language to Latent Alignment in Diffusion Models

DGX agent

arXiv:2512.08505v2 Announce Type: replace Abstract: Conditional diffusion models frequently suffer from language-image misalignments. Due to the ambiguity of intermediate noise corrupted latents, asse

model-releasesarxiv-cs-cv
30 Jun 2026
Research

EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics

DGX agent

arXiv:2606.30557v1 Announce Type: new Abstract: DiT video generation is latency-intensive due to iterative full-frame denoising, while prior cloud-edge methods largely rely on static inter-step decoup

researcharxiv-cs-cv
30 Jun 2026
Agents

Efficient 3D Gaussian Splatting with Axis-Shared Rasterization and Order-independent Transmittance

DGX agent

arXiv:2506.07069v2 Announce Type: replace-cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, combining high-quality reconstruction with efficien

agentsarxiv-cs-cv
30 Jun 2026
Model Releases

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction

DGX agent

arXiv:2606.29850v1 Announce Type: new Abstract: Visual pointing maps a language instruction to pixel co ordinates, a core skill for embodied AI. We describe our PointArena 2026 solution, which achieve

model-releasesarxiv-cs-cv
30 Jun 2026
Agents

Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation

DGX agent

arXiv:2512.10310v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated significant promise in Vision-Language Navigation (VLN), existing agents remain hea

agentsarxiv-cs-cv
30 Jun 2026
Research

Emergence of a Shared Canonical Object Frame from In-the-Wild Videos

DGX agent

arXiv:2606.30058v1 Announce Type: new Abstract: Comparing object orientations and positions across different instances requires their poses to be expressed in a shared canonical frame. Establishing su

researcharxiv-cs-cv
30 Jun 2026
Research

Empirical Evaluation of Multi-Modal Touch Detection in Over-the-Shoulder Video Surveillance

DGX agent

arXiv:2606.29504v1 Announce Type: new Abstract: Video Intelligence Surveillance (VIDINT) on over-the-shoulder footage is a proposed vector for monitoring human-computer interaction patterns without di

researcharxiv-cs-cv
30 Jun 2026
Applications

End-to-End Facial Expression Detection in Long Videos

DGX agent

arXiv:2504.07660v2 Announce Type: replace Abstract: Facial expression detection requires spotting when expressions occur and recognizing which emotional category they belong to. Despite their close re

applicationsarxiv-cs-cv
30 Jun 2026
Research

Enhancing Layer Interaction Using Key-Correlated Layer Attention

DGX agent

arXiv:2606.28405v1 Announce Type: new Abstract: Recent advances in network architecture design have introduced layer attention to enhance inter-layer interactions. In such frameworks, each layer queri

researcharxiv-cs-cv
30 Jun 2026
Research

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

DGX agent

arXiv:2606.29267v1 Announce Type: new Abstract: Visual grounding aims to associate free-form textual queries with specific regions in an image. While recent Multimodal Large Language Models (MLLMs) ha

researcharxiv-cs-cv
30 Jun 2026
Model Releases

Envisage: Diffusion-Based Rhinoplasty Goal Visualization with Mask-Decomposed Evaluation

DGX agent

arXiv:2606.28628v1 Announce Type: cross Abstract: Localized generative editing needs localized evaluation: full-image identity metrics are structurally confounded under hard-composited edits. We prese

model-releasesarxiv-cs-cv
30 Jun 2026
Research

EpiSAM: Character Segmentation in Challenging Stone Inscriptions

DGX agent

arXiv:2606.28859v1 Announce Type: new Abstract: Stone inscriptions are invaluable sources of historical and linguistic knowledge, yet their automated analysis remains a major challenge due to surface

researcharxiv-cs-cv
30 Jun 2026
Safety

EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal

DGX agent

arXiv:2512.21545v2 Announce Type: replace Abstract: Object removal must prevent the masked target from reappearing and reconstruct the occluded background with structural and contextual fidelity, rath

safetyarxiv-cs-cv
30 Jun 2026
Applications

Establishing the Minimal Clinically Important Difference (MCID) for Smartphone-Derived Gait Measures in Multiple Sclerosis

DGX agent

arXiv:2606.28449v1 Announce Type: cross Abstract: Background: Digital health technologies allow for frequent, remote gait monitoring in people with multiple sclerosis (MS). However, to differentiate d

applicationsarxiv-cs-cv
30 Jun 2026
Agents

Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems

DGX agent

arXiv:2504.02918v3 Announce Type: replace Abstract: Recent advances in image and video generation raise hopes that these models possess world modeling capabilities-the ability to generate realistic, p

agentsarxiv-cs-cv
30 Jun 2026
Applications

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model

DGX agent

arXiv:2606.29384v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable

applicationsarxiv-cs-cv
30 Jun 2026
Model Releases

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

DGX agent

arXiv:2606.20092v2 Announce Type: replace Abstract: Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-r

model-releasesarxiv-cs-cv
30 Jun 2026
Research

EvLIR: Learning Illumination Residuals from Ordered Events for Low-Light Image Enhancement

DGX agent

arXiv:2606.29430v1 Announce Type: new Abstract: Low-light image enhancement is severely ill-posed when the input frame contains missing structure, saturated noise, and weak local contrast. Event camer

researcharxiv-cs-cv
30 Jun 2026
Agents

Evolutionary Hyperparameter Optimization to Find Lightweight CNN Models for Autonomous Steering

DGX agent

arXiv:2606.29684v1 Announce Type: cross Abstract: This research investigates the optimization of Convolutional and Dense Neural Networks (CNNs and DNNs) for autonomous steering using the (N+M) Evoluti

agentsarxiv-cs-cv
30 Jun 2026
Local Ai

ExACT: Exemplar-Driven Calibrated Refinement for Training-Free Visual Grounding in Remote Sensing Images

DGX agent

arXiv:2606.28920v1 Announce Type: new Abstract: Remote sensing visual grounding (RSVG) aims to locate specific objects in high-resolution RS imagery using free-form natural language descriptions. Whil

local-aiarxiv-cs-cv
30 Jun 2026
Model Releases

Explainability-Aware Frustum Attack: Exposing Structural Vulnerabilities in LiDAR-Based 3D Object Detectors

DGX agent

arXiv:2606.29963v1 Announce Type: new Abstract: The structural vulnerabilities of point cloud-based 3D object detectors remain poorly understood. Prior work has studied adversarial robustness primaril

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving

DGX agent

arXiv:2604.02714v2 Announce Type: replace Abstract: End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies t

safetyarxiv-cs-cv
30 Jun 2026
Safety

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation

DGX agent

arXiv:2602.12155v2 Announce Type: replace Abstract: Post-training of flow matching models-aligning the output distribution with a high-quality target-is mathematically equivalent to imitation learning

safetyarxiv-cs-cv
30 Jun 2026
Research

Fast Equivariant Imaging: Accelerating Unsupervised Learning and Model Adaptation via Inexact Splitting

DGX agent

arXiv:2507.06764v5 Announce Type: replace-cross Abstract: In this work, we propose Fast Equivariant Imaging (FEI), a novel unsupervised learning framework to rapidly and efficiently train deep imaging

researcharxiv-cs-cv
30 Jun 2026
Model Releases

FastPano3D: Feed-Forward Indoor Panoramic 3D Reconstruction from a Single Image

DGX agent

arXiv:2606.30352v1 Announce Type: new Abstract: Recent advances in 3D scene reconstruction have highlighted the intricate trade-offs among rendering quality, inference efficiency, and data dependency.

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On

DGX agent

arXiv:2606.29319v1 Announce Type: new Abstract: Image-based Virtual Try-On (IVTON) has greatly advanced through diffusion models, yet existing methods require many sampling steps and depend on masks w

safetyarxiv-cs-cv
30 Jun 2026
Model Releases

FiRe: Frequency Reparameterization as a Preconditioner for Periodic Implicit Neural Representations

DGX agent

arXiv:2606.29414v1 Announce Type: new Abstract: Periodic Implicit Neural Representations (INRs) such as SIREN and FINER assign every neuron, the same global frequency, spending the representational bu

model-releasesarxiv-cs-cv
30 Jun 2026
← Previous
1…7879808182…263
Next →