AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

Bayesian adaptively-weighted ensembles for few-shot abdominal segmentation

DGX agent

arXiv:2608.05815v1 Announce Type: new Abstract: Few-shot learning has emerged as a promising approach for anatomical segmentation when labelled data are scarce. However, different few-shot learning al

researcharxiv-cs-cv
7 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

BendTwin: Robust Dense-to-Sparse Physical Reconstruction with Bending-Aware Differentiable Spring-Mass Models

DGX agent

arXiv:2608.06164v1 Announce Type: new Abstract: Reconstructing objects with mechanical properties from video observations enables physically consistent dynamic prediction, benefiting robotics planning

researcharxiv-cs-cv
7 Aug 2026
Research

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

DGX agent

arXiv:2608.05592v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation m

researcharxiv-cs-cv
7 Aug 2026
Agents

Beyond Relevance: Bayesian Evidence Acquisition for Agentic Whole-Slide Image Reasoning

DGX agent

arXiv:2608.05757v1 Announce Type: new Abstract: Whole-slide image (WSI) reasoning requires an agent to sequentially acquire visual evidence before answering a diagnostic question. Existing training-fr

agentsarxiv-cs-cv
7 Aug 2026
Research

CDSeg: A Renderable Gaussian Carrier for Image-to-3D Label Transfer

DGX agent

arXiv:2608.05482v1 Announce Type: new Abstract: Modern image models provide strong cues about what should be segmented in each view, but their masks do not by themselves determine where those labels s

researcharxiv-cs-cv
7 Aug 2026
Research

CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection

DGX agent

arXiv:2608.06205v1 Announce Type: new Abstract: RGB--T object detection exploits the complementary strengths of visible and infrared imagery, supporting robust perception in low-light, adverse-weather

researcharxiv-cs-cv
7 Aug 2026
Model Releases

ChronoVision: Temporal Reasoning via Latent State Reconstruction

DGX agent

arXiv:2608.05631v1 Announce Type: new Abstract: Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. T

model-releasesarxiv-cs-cv
7 Aug 2026
Local Ai

ConceptADapt: Concept-guided Adaptive Feature Reconstruction with Dynamic Attention for Few-Shot Industrial Anomaly Detection

DGX agent

arXiv:2608.05743v1 Announce Type: new Abstract: Few-shot industrial anomaly detection (FS-IAD) focuses on detecting and localizing visual defects in industrial inspection during the cold-start phase,

local-aiarxiv-cs-cv
7 Aug 2026
Research

Confidence matters: Leveraging Multi-view Geometric Priors for GS-based Reconstruction

DGX agent

arXiv:2608.06117v1 Announce Type: new Abstract: 3D Gaussian splatting (3DGS) has emerged as a widely-used tool for novel view synthesis, offering real-time rendering in a sparse representation. Howeve

researcharxiv-cs-cv
7 Aug 2026
Research

Context Matters: Support Set Selection and Failure Detection for In-Context Medical Image Segmentation

DGX agent

arXiv:2608.05333v1 Announce Type: new Abstract: In-context learning (ICL) adapts medical image segmentation models to unseen structures and modalities without retraining by conditioning on a task-spec

researcharxiv-cs-cv
7 Aug 2026
Research

Controllable Clothing: Precise Labels and Generation for Virtual Try-On with Latent Diffusion Models

DGX agent

arXiv:2608.05834v1 Announce Type: new Abstract: In this technical report, I present a new method for guiding image generation in the context of Virtual- Try-On (VITON). The proposed method leverages n

researcharxiv-cs-cv
7 Aug 2026
Safety

CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images

DGX agent

arXiv:2608.05569v1 Announce Type: new Abstract: Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localizatio

safetyarxiv-cs-cv
7 Aug 2026
Local Ai

Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation

DGX agent

arXiv:2608.05844v1 Announce Type: new Abstract: Radiology foundation models learn transferable representations that can be adapted to new tasks by training only small layers on top of a frozen encoder

local-aiarxiv-cs-cv
7 Aug 2026
Safety

DARAD: Dual Adapters and Ranking-Aware Distillation for Continual Remote Sensing Image-Text Retrieval

DGX agent

arXiv:2608.06059v1 Announce Type: new Abstract: With the rapid growth of Earth observation technologies, remote sensing archives are rapidly expanding, making remote sensing image-text retrieval (RS-I

safetyarxiv-cs-cv
7 Aug 2026
Research

Dense-Cast: A lightweight ensemble of deep learning architectures for precipitation nowcasting

DGX agent

arXiv:2608.06082v1 Announce Type: new Abstract: Proper short-term forecasting of precipitation is crucial in disaster management and preparedness. Nonetheless, the variability and nonlinearity of prec

researcharxiv-cs-cv
7 Aug 2026
Safety

Deterministic World Models for Closed-loop Reachability Analysis of End-to-End Vision-based Control

DGX agent

arXiv:2512.08991v3 Announce Type: replace Abstract: End-to-end image controllers that map raw camera frames directly to control actions are increasingly deployed in safety-critical systems. However, f

safetyarxiv-cs-cv
7 Aug 2026
Applications

DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching

DGX agent

arXiv:2603.26320v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models that encode actions using a discrete tokenization scheme have been widely adopted for robotic manipulation

applicationsarxiv-cs-cv
7 Aug 2026
Research

Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model

DGX agent

arXiv:2608.05976v1 Announce Type: new Abstract: Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, an

researcharxiv-cs-cv
7 Aug 2026
Agents

Disentangling 3D Modeling from Spatial Reasoning

DGX agent

arXiv:2608.05242v1 Announce Type: cross Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly a

agentsarxiv-cs-cv
7 Aug 2026
Applications

DTRNet: Dual Text-Radical Decoding for Handwritten Chinese Text Recognition with Faked Character Detection

DGX agent

arXiv:2608.05848v1 Announce Type: new Abstract: In K-12 educational scenarios, handwritten Chinese text recognition should not only transcribe student writing, but also detect faked characters. Howeve

applicationsarxiv-cs-cv
7 Aug 2026
Applications

Dual-Attention and Adversarial Transfer Networks for Sim-to-Real Cross-Orientation Wireless Sensing

DGX agent

arXiv:2608.05664v1 Announce Type: new Abstract: Millimeter-wave human activity recognition suffers significant performance degradation when the user's orientation changes relative to the sensing syste

applicationsarxiv-cs-cv
7 Aug 2026
Research

Dual-Output Multi-Exposure HDR Reconstruction via SDR Fusion and Gain Map Inverse Tone Mapping

DGX agent

arXiv:2608.05626v1 Announce Type: new Abstract: We propose DOME-HDR, a dual-output multi-exposure HDR reconstruction framework that jointly produces a perceptually balanced SDR image and a consistent

researcharxiv-cs-cv
7 Aug 2026
Applications

Dynamic Object Masks as Goal Representations for Visual Goal-Conditioned Reinforcement Learning

DGX agent

arXiv:2510.06277v2 Announce Type: replace Abstract: Goal-conditioned reinforcement learning (GCRL) offers a unified way to pursue diverse tasks, yet most existing methods rely on state- or position-ba

applicationsarxiv-cs-cv
7 Aug 2026
Model Releases

DynaPix: Can Vision-Language Models Identify the Exact Future?

DGX agent

arXiv:2608.05505v1 Announce Type: new Abstract: Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking ima

model-releasesarxiv-cs-cv
7 Aug 2026
Tutorials

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

DGX agent

arXiv:2608.05565v1 Announce Type: new Abstract: Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coheren

tutorialsarxiv-cs-cv
7 Aug 2026
Safety

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

DGX agent

arXiv:2608.06231v1 Announce Type: new Abstract: Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progr

safetyarxiv-cs-cv
7 Aug 2026
Model Releases

Energy-Guided Flow Matching

DGX agent

arXiv:2608.05811v1 Announce Type: new Abstract: Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dim

model-releasesarxiv-cs-cv
7 Aug 2026
Research

Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams

DGX agent

arXiv:2608.05728v1 Announce Type: new Abstract: Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-

researcharxiv-cs-cv
7 Aug 2026
Research

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

DGX agent

arXiv:2608.05265v1 Announce Type: cross Abstract: Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in r

researcharxiv-cs-cv
7 Aug 2026
Safety

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

DGX agent

arXiv:2608.05780v1 Announce Type: new Abstract: Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selectin

safetyarxiv-cs-cv
7 Aug 2026
Model Releases

EvReflection: Event-Driven Micro-Dynamics for Reflection Removal

DGX agent

arXiv:2608.06184v1 Announce Type: new Abstract: Despite remarkable progress in reflection removal, current methods primarily exploit static image priors from a single frame and still suffer from sever

model-releasesarxiv-cs-cv
7 Aug 2026
Safety

FashionPose: Unified Text-Driven Fashion Synthesis with Joint Geometric and Photometric Control

DGX agent

arXiv:2507.13311v2 Announce Type: replace Abstract: Realistic and controllable garment synthesis is essential for fashion e-commerce, yet it demands precise coordination between human pose geometry an

safetyarxiv-cs-cv
7 Aug 2026
Safety

Faster and Better Alignment for Flow Matching Models via Step-aware Advantages

DGX agent

arXiv:2602.01591v2 Announce Type: replace Abstract: Recent advances in flow matching models, particularly with reinforcement learning (RL), have significantly enhanced human preference alignment in fe

safetyarxiv-cs-cv
7 Aug 2026
Local Ai

Floating Radiance Networks

DGX agent

arXiv:2608.05920v1 Announce Type: new Abstract: Recent advances in neural scene representations enable photorealistic novel-view synthesis, yet most methods remain tightly coupled to a single renderin

local-aiarxiv-cs-cv
7 Aug 2026
Research

Flow-Map Distillation on Relation Manifolds for Image Restoration

DGX agent

arXiv:2608.05769v1 Announce Type: new Abstract: Knowledge distillation for image restoration typically aligns intermediate features or relation matrices between teacher and student networks as static

researcharxiv-cs-cv
7 Aug 2026
Local Ai

G^2ARD-GS: Geometry-Guided Anchor-Regularized Gaussian Splatting Distillation

DGX agent

arXiv:2608.05704v1 Announce Type: new Abstract: Dense colored LiDAR maps provide accurate city-scale geometry, but lifting them into 3D Gaussian Splatting (3DGS) retains millions of primitives, making

local-aiarxiv-cs-cv
7 Aug 2026
Research

Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI

DGX agent

arXiv:2608.05258v1 Announce Type: new Abstract: Gradient-weighted Class Activation Mapping (Grad-CAM) is widely used to visualize model decisions, but it was originally formulated for convolutional ne

researcharxiv-cs-cv
7 Aug 2026
Model Releases

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

DGX agent

arXiv:2608.05747v1 Announce Type: new Abstract: Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overloo

model-releasesarxiv-cs-cv
7 Aug 2026
Research

HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models

DGX agent

arXiv:2608.05523v1 Announce Type: new Abstract: Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challe

researcharxiv-cs-cv
7 Aug 2026
Research

Hierarchical Flow Matching for 3D Point Cloud Generation

DGX agent

arXiv:2608.05557v1 Announce Type: new Abstract: Generating high-quality 3D point clouds requires capturing both global shape topology and local geometric details. Existing flow-based methods rely on c

researcharxiv-cs-cv
7 Aug 2026
Research

HOPE: Hand-Object Pressure Estimation from Monocular Videos

DGX agent

arXiv:2608.06192v1 Announce Type: new Abstract: Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estim

researcharxiv-cs-cv
7 Aug 2026
Applications

IDperturb: Enhancing Variation in Synthetic Face Generation via Angular Perturbation

DGX agent

arXiv:2602.18831v2 Announce Type: replace Abstract: Synthetic data has emerged as a practical alternative to authentic face datasets for training face recognition (FR) systems, especially as privacy a

applicationsarxiv-cs-cv
7 Aug 2026
Safety

Inspecting Training Dynamics of Similarity Development in Supervised Vision Networks

DGX agent

arXiv:2505.21338v2 Announce Type: replace Abstract: For trustworthy and human-aware artificial intelligence, models should be evaluated beyond accuracy, among others through error predictability and s

safetyarxiv-cs-cv
7 Aug 2026
Research

Invisible Shortcuts: Why Vision Encoders Know Your Camera

DGX agent

arXiv:2608.05424v1 Announce Type: new Abstract: Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-

researcharxiv-cs-cv
7 Aug 2026
Model Releases

IPV-Bench: Benchmarking Image Protection Methods under Diverse Image-to-Video Generation Scenarios

DGX agent

arXiv:2603.26154v2 Announce Type: replace Abstract: Image-to-video (I2V) generation models can be misused to animate a single image into a convincing fake video, motivating perturbation-based image pr

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

Iterate or Widen? When Test-Time Refinement Helps LiDAR Scene Completion: A Controlled Study of Evidence Geometry, Training Coverage, and Compute

DGX agent

arXiv:2608.06014v1 Announce Type: new Abstract: Should a completion model spend extra test-time compute by iterating, or spend a similar parameter budget on a wider one-shot predictor? The answer is e

model-releasesarxiv-cs-cv
7 Aug 2026
Local Ai

Iterative Hybrid Discrete-Continuous Viewpoint Planning for UAV Photogrammetry

DGX agent

arXiv:2608.05718v1 Announce Type: new Abstract: Unmanned aerial vehicle (UAV) photogrammetry requires camera networks that provide sufficient surface coverage, image overlap, parallax, and resolution,

local-aiarxiv-cs-cv
7 Aug 2026
Research

KVAE: Family of Tokenizers for Multimodal Generative Models

DGX agent

arXiv:2608.05798v1 Announce Type: new Abstract: Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions t

researcharxiv-cs-cv
7 Aug 2026
← Previous
1…910111213…259
Next →