AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
7 Aug 2026

Bayesian adaptively-weighted ensembles for few-shot abdominal segmentation

ResearchDGX agent

arXiv:2608.05815v1 Announce Type: new Abstract: Few-shot learning has emerged as a promising approach for anatomical segmentation when labelled data are scarce. However, different few-shot learning al

BendTwin: Robust Dense-to-Sparse Physical Reconstruction with Bending-Aware Differentiable Spring-Mass Models

ResearchDGX agent

arXiv:2608.06164v1 Announce Type: new Abstract: Reconstructing objects with mechanical properties from video observations enables physically consistent dynamic prediction, benefiting robotics planning

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

ResearchDGX agent

arXiv:2608.05592v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation m


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Beyond Relevance: Bayesian Evidence Acquisition for Agentic Whole-Slide Image Reasoning

AgentsDGX agent

arXiv:2608.05757v1 Announce Type: new Abstract: Whole-slide image (WSI) reasoning requires an agent to sequentially acquire visual evidence before answering a diagnostic question. Existing training-fr

CDSeg: A Renderable Gaussian Carrier for Image-to-3D Label Transfer

ResearchDGX agent

arXiv:2608.05482v1 Announce Type: new Abstract: Modern image models provide strong cues about what should be segmented in each view, but their masks do not by themselves determine where those labels s

CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection

ResearchDGX agent

arXiv:2608.06205v1 Announce Type: new Abstract: RGB--T object detection exploits the complementary strengths of visible and infrared imagery, supporting robust perception in low-light, adverse-weather

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Model ReleasesDGX agent

arXiv:2608.05631v1 Announce Type: new Abstract: Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. T

ConceptADapt: Concept-guided Adaptive Feature Reconstruction with Dynamic Attention for Few-Shot Industrial Anomaly Detection

Local AiDGX agent

arXiv:2608.05743v1 Announce Type: new Abstract: Few-shot industrial anomaly detection (FS-IAD) focuses on detecting and localizing visual defects in industrial inspection during the cold-start phase,

Confidence matters: Leveraging Multi-view Geometric Priors for GS-based Reconstruction

ResearchDGX agent

arXiv:2608.06117v1 Announce Type: new Abstract: 3D Gaussian splatting (3DGS) has emerged as a widely-used tool for novel view synthesis, offering real-time rendering in a sparse representation. Howeve

Context Matters: Support Set Selection and Failure Detection for In-Context Medical Image Segmentation

ResearchDGX agent

arXiv:2608.05333v1 Announce Type: new Abstract: In-context learning (ICL) adapts medical image segmentation models to unseen structures and modalities without retraining by conditioning on a task-spec

Controllable Clothing: Precise Labels and Generation for Virtual Try-On with Latent Diffusion Models

ResearchDGX agent

arXiv:2608.05834v1 Announce Type: new Abstract: In this technical report, I present a new method for guiding image generation in the context of Virtual- Try-On (VITON). The proposed method leverages n

CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images

SafetyDGX agent

arXiv:2608.05569v1 Announce Type: new Abstract: Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localizatio

Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation

Local AiDGX agent

arXiv:2608.05844v1 Announce Type: new Abstract: Radiology foundation models learn transferable representations that can be adapted to new tasks by training only small layers on top of a frozen encoder

DARAD: Dual Adapters and Ranking-Aware Distillation for Continual Remote Sensing Image-Text Retrieval

SafetyDGX agent

arXiv:2608.06059v1 Announce Type: new Abstract: With the rapid growth of Earth observation technologies, remote sensing archives are rapidly expanding, making remote sensing image-text retrieval (RS-I

Dense-Cast: A lightweight ensemble of deep learning architectures for precipitation nowcasting

ResearchDGX agent

arXiv:2608.06082v1 Announce Type: new Abstract: Proper short-term forecasting of precipitation is crucial in disaster management and preparedness. Nonetheless, the variability and nonlinearity of prec

Deterministic World Models for Closed-loop Reachability Analysis of End-to-End Vision-based Control

SafetyDGX agent

arXiv:2512.08991v3 Announce Type: replace Abstract: End-to-end image controllers that map raw camera frames directly to control actions are increasingly deployed in safety-critical systems. However, f

DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching

ApplicationsDGX agent

arXiv:2603.26320v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models that encode actions using a discrete tokenization scheme have been widely adopted for robotic manipulation

Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model

ResearchDGX agent

arXiv:2608.05976v1 Announce Type: new Abstract: Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, an

Disentangling 3D Modeling from Spatial Reasoning

AgentsDGX agent

arXiv:2608.05242v1 Announce Type: cross Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly a

DTRNet: Dual Text-Radical Decoding for Handwritten Chinese Text Recognition with Faked Character Detection

ApplicationsDGX agent

arXiv:2608.05848v1 Announce Type: new Abstract: In K-12 educational scenarios, handwritten Chinese text recognition should not only transcribe student writing, but also detect faked characters. Howeve

Dual-Attention and Adversarial Transfer Networks for Sim-to-Real Cross-Orientation Wireless Sensing

ApplicationsDGX agent

arXiv:2608.05664v1 Announce Type: new Abstract: Millimeter-wave human activity recognition suffers significant performance degradation when the user's orientation changes relative to the sensing syste

Dual-Output Multi-Exposure HDR Reconstruction via SDR Fusion and Gain Map Inverse Tone Mapping

ResearchDGX agent

arXiv:2608.05626v1 Announce Type: new Abstract: We propose DOME-HDR, a dual-output multi-exposure HDR reconstruction framework that jointly produces a perceptually balanced SDR image and a consistent

Dynamic Object Masks as Goal Representations for Visual Goal-Conditioned Reinforcement Learning

ApplicationsDGX agent

arXiv:2510.06277v2 Announce Type: replace Abstract: Goal-conditioned reinforcement learning (GCRL) offers a unified way to pursue diverse tasks, yet most existing methods rely on state- or position-ba

DynaPix: Can Vision-Language Models Identify the Exact Future?

Model ReleasesDGX agent

arXiv:2608.05505v1 Announce Type: new Abstract: Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking ima

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

TutorialsDGX agent

arXiv:2608.05565v1 Announce Type: new Abstract: Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coheren

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

SafetyDGX agent

arXiv:2608.06231v1 Announce Type: new Abstract: Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progr

Energy-Guided Flow Matching

Model ReleasesDGX agent

arXiv:2608.05811v1 Announce Type: new Abstract: Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dim

Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams

ResearchDGX agent

arXiv:2608.05728v1 Announce Type: new Abstract: Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

ResearchDGX agent

arXiv:2608.05265v1 Announce Type: cross Abstract: Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in r

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

SafetyDGX agent

arXiv:2608.05780v1 Announce Type: new Abstract: Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selectin

EvReflection: Event-Driven Micro-Dynamics for Reflection Removal

Model ReleasesDGX agent

arXiv:2608.06184v1 Announce Type: new Abstract: Despite remarkable progress in reflection removal, current methods primarily exploit static image priors from a single frame and still suffer from sever

FashionPose: Unified Text-Driven Fashion Synthesis with Joint Geometric and Photometric Control

SafetyDGX agent

arXiv:2507.13311v2 Announce Type: replace Abstract: Realistic and controllable garment synthesis is essential for fashion e-commerce, yet it demands precise coordination between human pose geometry an

Faster and Better Alignment for Flow Matching Models via Step-aware Advantages

SafetyDGX agent

arXiv:2602.01591v2 Announce Type: replace Abstract: Recent advances in flow matching models, particularly with reinforcement learning (RL), have significantly enhanced human preference alignment in fe

Floating Radiance Networks

Local AiDGX agent

arXiv:2608.05920v1 Announce Type: new Abstract: Recent advances in neural scene representations enable photorealistic novel-view synthesis, yet most methods remain tightly coupled to a single renderin

Flow-Map Distillation on Relation Manifolds for Image Restoration

ResearchDGX agent

arXiv:2608.05769v1 Announce Type: new Abstract: Knowledge distillation for image restoration typically aligns intermediate features or relation matrices between teacher and student networks as static

G^2ARD-GS: Geometry-Guided Anchor-Regularized Gaussian Splatting Distillation

Local AiDGX agent

arXiv:2608.05704v1 Announce Type: new Abstract: Dense colored LiDAR maps provide accurate city-scale geometry, but lifting them into 3D Gaussian Splatting (3DGS) retains millions of primitives, making

Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI

ResearchDGX agent

arXiv:2608.05258v1 Announce Type: new Abstract: Gradient-weighted Class Activation Mapping (Grad-CAM) is widely used to visualize model decisions, but it was originally formulated for convolutional ne

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Model ReleasesDGX agent

arXiv:2608.05747v1 Announce Type: new Abstract: Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overloo

HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models

ResearchDGX agent

arXiv:2608.05523v1 Announce Type: new Abstract: Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challe

Hierarchical Flow Matching for 3D Point Cloud Generation

ResearchDGX agent

arXiv:2608.05557v1 Announce Type: new Abstract: Generating high-quality 3D point clouds requires capturing both global shape topology and local geometric details. Existing flow-based methods rely on c

HOPE: Hand-Object Pressure Estimation from Monocular Videos

ResearchDGX agent

arXiv:2608.06192v1 Announce Type: new Abstract: Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estim

IDperturb: Enhancing Variation in Synthetic Face Generation via Angular Perturbation

ApplicationsDGX agent

arXiv:2602.18831v2 Announce Type: replace Abstract: Synthetic data has emerged as a practical alternative to authentic face datasets for training face recognition (FR) systems, especially as privacy a

Inspecting Training Dynamics of Similarity Development in Supervised Vision Networks

SafetyDGX agent

arXiv:2505.21338v2 Announce Type: replace Abstract: For trustworthy and human-aware artificial intelligence, models should be evaluated beyond accuracy, among others through error predictability and s

Invisible Shortcuts: Why Vision Encoders Know Your Camera

ResearchDGX agent

arXiv:2608.05424v1 Announce Type: new Abstract: Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-

IPV-Bench: Benchmarking Image Protection Methods under Diverse Image-to-Video Generation Scenarios

Model ReleasesDGX agent

arXiv:2603.26154v2 Announce Type: replace Abstract: Image-to-video (I2V) generation models can be misused to animate a single image into a convincing fake video, motivating perturbation-based image pr

Iterate or Widen? When Test-Time Refinement Helps LiDAR Scene Completion: A Controlled Study of Evidence Geometry, Training Coverage, and Compute

Model ReleasesDGX agent

arXiv:2608.06014v1 Announce Type: new Abstract: Should a completion model spend extra test-time compute by iterating, or spend a similar parameter budget on a wider one-shot predictor? The answer is e

Iterative Hybrid Discrete-Continuous Viewpoint Planning for UAV Photogrammetry

Local AiDGX agent

arXiv:2608.05718v1 Announce Type: new Abstract: Unmanned aerial vehicle (UAV) photogrammetry requires camera networks that provide sufficient surface coverage, image overlap, parallax, and resolution,

KVAE: Family of Tokenizers for Multimodal Generative Models

ResearchDGX agent

arXiv:2608.05798v1 Announce Type: new Abstract: Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions t

LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

SafetyDGX agent

arXiv:2608.05706v1 Announce Type: new Abstract: World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied int

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

ResearchDGX agent

arXiv:2608.06060v1 Announce Type: new Abstract: Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-

Learning visual representations for compositional analysis of artworks and photographs

SafetyDGX agent

arXiv:2608.06142v1 Announce Type: new Abstract: Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it re

LiteKD-Net: Lightweight Knowledge-Distilled Network for Mobile Image Denoising

ApplicationsDGX agent

arXiv:2608.05739v1 Announce Type: new Abstract: Mobile image denoising requires both good restoration quality and low computational cost. In addition, it's annoying to collect large-scale LQ-GT clean

LoDA: A Level of Detection Aware Method and a Multimodal Sensing Benchmark for Object Level Change Detection

Model ReleasesDGX agent

arXiv:2608.05356v1 Announce Type: new Abstract: High-definition 3D LiDAR maps are important for autonomous driving and smart-city services, which require reliable detection of object-level changes in

MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding

Model ReleasesDGX agent

arXiv:2607.16284v2 Announce Type: replace Abstract: Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communicat

Mapping Armenian Paris: Extracting and Geocoding Commercial Advertisements from the 20th-Century Diaspora Press

ResearchDGX agent

arXiv:2608.05911v1 Announce Type: new Abstract: This paper presents an end-to-end, IIIF-based pipeline that turns the digitised Armenian press of France into an interactive map of the 20th-century Par

MapTCL: Temporal Consistency Learning via Bidirectional Alignment for Vectorized HD Map Construction

SafetyDGX agent

arXiv:2608.05209v1 Announce Type: new Abstract: Constructing reliable online HD maps remains challenging in dynamic urban environments due to moving objects and occlusions. While recent works employ f

MASS: Multiplayer World Models with Authoritative Shared State

Model ReleasesDGX agent

arXiv:2608.06257v1 Announce Type: new Abstract: Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redunda

MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers

Model ReleasesDGX agent

arXiv:2608.05878v1 Announce Type: new Abstract: Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-s

MirrorNet: Can Medical Image Anonymization Really Protect Patient Identity?

ResearchDGX agent

arXiv:2608.05938v1 Announce Type: new Abstract: Medical images are routinely de-identified---names, dates, and other metadata removed---and then shared for research, teaching, and public benchmarks un

MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation

ResearchDGX agent

arXiv:2608.05450v1 Announce Type: new Abstract: Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substanti

← Previous
1…7891011…207
Next →