AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
11 May 2026

GraphFusion3D: Dynamic Graph Attention Convolution with Adaptive Cross-Modal Transformer for 3D Object Detection

ResearchDGX agent

arXiv:2512.02991v2 Announce Type: replace Abstract: Despite significant progress in 3D object detection, point clouds remain challenging due to sparse data, incomplete structures, and limited semantic

Head Similarity: Modeling Structured Whole-Head Appearance Beyond Face Recognition

Model ReleasesDGX agent

arXiv:2605.07766v1 Announce Type: new Abstract: Many vision applications require identity consistency beyond strict biometric recognition, especially under non-frontal views or when facial cues are mi

HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.07973v1 Announce Type: new Abstract: Text-to-image diffusion models can generate visually stunning images, yet, controlling what appears and how it appears, remains surprisingly difficult,

Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.07512v1 Announce Type: new Abstract: Class-incremental learning aims to continuously acquire new knowledge while preserving previously learned information, thereby mitigating catastrophic f

Hierarchical Perfusion Graphs for Tumor Heterogeneity Modeling in Glioma Molecular Subtyping

ResearchDGX agent

arXiv:2605.07156v1 Announce Type: new Abstract: Precise molecular subtyping of gliomas, including isocitrate dehydrogenase (IDH) mutation and 1p/19q codeletion, directly guides surgical and therapeuti

High-Fidelity Surface Splatting-Based 3D Reconstruction from Multi-View Images

ResearchDGX agent

arXiv:2605.07254v1 Announce Type: new Abstract: Multi-view mesh reconstruction remains a core challenge in computer graphics and vision, especially for recovering high-frequency geometry from sparse o

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings

Model ReleasesDGX agent

arXiv:2605.07492v1 Announce Type: new Abstract: The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanual

HumanNet: Scaling Human-centric Video Learning to One Million Hours

Model ReleasesDGX agent

arXiv:2605.06747v1 Announce Type: new Abstract: Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, lea

ICDAR 2026 Competition on Writer Identification and Pen Classification from Hand-Drawn Circles

ResearchDGX agent

arXiv:2605.07816v1 Announce Type: new Abstract: This paper presents CircleID, a large-scale ICDAR 2026 competition on writer identification and pen classification from scanned hand-drawn circles. The

ImplantMamba: Long-range Sequential Modeling Mamba For Dental Implant Position Prediction

ResearchDGX agent

arXiv:2605.07082v1 Announce Type: new Abstract: In the design of surgical guides for implant placement, determining the precise implant position is a critical step. However, the implant region itself

Implicit Multi-Camera System Calibration Using Gaussian Processes

TutorialsDGX agent

arXiv:2605.07491v1 Announce Type: new Abstract: This paper proposes a novel framework for implicit multi-camera system calibration utilizing Gaussian Process (GP) regression. Conventional explicit cal

InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-Localization

SafetyDGX agent

arXiv:2605.07099v1 Announce Type: new Abstract: Cross-view geo-localization (CVGL) is fundamental for precise localization and navigation in GPS-denied environments, aiming to match ground or UAV imag

InsHuman: Towards Natural and Identity-Preserving Human Insertion

ResearchDGX agent

arXiv:2605.07402v1 Announce Type: new Abstract: Human insertion aims to naturally place specific individuals into a target background. Although existing image editing models may have such ability, the

InterCoG: Towards Spatially Precise Image Editing with Interleaved Chain-of-Grounding Reasoning

SafetyDGX agent

arXiv:2603.01586v3 Announce Type: replace Abstract: Emerging unified editing models have demonstrated strong capabilities in general object editing tasks. However, it remains a significant challenge t

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models

SafetyDGX agent

arXiv:2605.07514v1 Announce Type: cross Abstract: World Action Models (WAMs) enable decision-making through imagined rollouts by predicting future observations and actions. However, the reliability of

LAMES: A Large-Scale and Artisanal Mining Environmental Segmentation Dataset

ApplicationsDGX agent

arXiv:2605.07740v1 Announce Type: new Abstract: Mining operations are of utmost importance to the economy of some nations. However, such operations result in land-use change, very high energy consumpt

Large Video Planner Enables Generalizable Robot Control

ApplicationsDGX agent

arXiv:2512.15840v2 Announce Type: replace-cross Abstract: General-purpose robots require decision-making models that generalize across diverse tasks and environments. Recent works build robot foundati

Learning Image-Adaptive Scale Fields for Metric Depth Recovery

ResearchDGX agent

arXiv:2605.07418v1 Announce Type: new Abstract: Monocular depth estimation (MDE) typically produces depth estimations that are defined up to an unknown scale or shift. When only sparse metric anchors

Learning to Track Instance from Single Nature Language Description

SafetyDGX agent

arXiv:2605.07064v1 Announce Type: new Abstract: How to achieve vision-language (VL) tracking using natural language descriptions from a video sequence extbf{without relying on any bounding-box ground

LENS: Low-Frequency Eigen Noise Shaping for Efficient Diffusion Sampling

ResearchDGX agent

arXiv:2605.07253v1 Announce Type: new Abstract: Distilled diffusion models accelerate image generation by reducing the number of denoising steps, but often suffer from degraded image quality. To mitig

Lightweight Unpaired Smartphone ISP Transfer with Semantic Pseudo-Pairing

SafetyDGX agent

arXiv:2605.07495v1 Announce Type: new Abstract: Unpaired smartphone ISP is a challenging problem due to the lack of scene and color alignment between RAW and target RGB images. Many existing methods e

LoHGNet: Infrared Small Target Detection through Lorentz Geometric Encoding with High-Order Relation Learning

Local AiDGX agent

arXiv:2605.07213v1 Announce Type: new Abstract: Infrared small target detection (IRSTD) remains challenging due to the scarcity of useful target cues and the presence of severe background clutter. Mos

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute

TutorialsDGX agent

arXiv:2605.06809v1 Announce Type: new Abstract: Transformers dominate video recognition. They split videos into tokens, and processing them has expensive superlinear computational cost. Yet videos are

Lossy Common Information in a Learnable Gray-Wyner Network

ResearchDGX agent

arXiv:2601.21424v3 Announce Type: replace-cross Abstract: Many computer vision tasks share substantial overlapping information, yet conventional codecs tend to ignore this, leading to redundant and in

Masks Can Talk: Extracting Structured Text Information from Single-Modal Images for Remote Sensing Change Detection

SafetyDGX agent

arXiv:2605.07178v1 Announce Type: new Abstract: Remote sensing change detection is pivotal for urban monitoring, disaster assessment, and environmental resource management. Yet, unimodal deep learning

Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers

ResearchDGX agent

arXiv:2605.06169v1 Announce Type: cross Abstract: Scaling Diffusion Transformers (DiTs) to hundreds of layers introduces a structural vulnerability: networks can enter a silent, mean-dominated collaps

MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence

Model ReleasesDGX agent

arXiv:2605.07919v1 Announce Type: new Abstract: Medical vision--language models (VLMs) are usually evaluated on intact image--question pairs, but trustworthy clinical use requires a stronger property:

MicroBi-ConvLSTM: An Ultra-Lightweight Efficient Model for Human Activity Recognition on Resource Constrained Devices

Model ReleasesDGX agent

arXiv:2602.06523v2 Announce Type: replace Abstract: Human Activity Recognition (HAR) on resource constrained wearables requires models that balance accuracy against strict memory and computational bud

Mind the Gap: Geometrically Accurate Generative Reconstruction from Disjoint Views

SafetyDGX agent

arXiv:2605.07550v1 Announce Type: new Abstract: 3D vision systems are fundamentally constrained by their reliance on visual overlap: reconstruction methods require it for geometric alignment, while ge

MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation

SafetyDGX agent

arXiv:2605.08050v1 Announce Type: new Abstract: Talking-head generation requires joint modeling of identity, head pose, facial expression, and mouth dynamics. Existing methods typically address only a

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation

ResearchDGX agent

arXiv:2411.16748v5 Announce Type: replace Abstract: Long-duration talking video synthesis faces enduring challenges in achieving high video quality, portrait consistency, temporal coherence, and compu

Multimodal Latent Reasoning via Hierarchical Visual Cues Injection

ResearchDGX agent

arXiv:2602.05359v2 Announce Type: replace Abstract: The advancement of multimodal large language models (MLLMs) has enabled impressive perception capabilities. However, their reasoning process often r

Multimodal Stepwise Clinically-Guided Attention Learning for Pathological Complete Response Prediction in Breast Cancer

ResearchDGX agent

arXiv:2605.07561v1 Announce Type: new Abstract: Pathological complete response (pCR) is a key prognostic factor in breast cancer patients undergoing neoadjuvant therapy, strongly associated with long-

Multispectral Indices for Wildfire Management

ResearchDGX agent

arXiv:2309.01751v3 Announce Type: replace-cross Abstract: The increasing frequency and severity of wildfires necessitates advanced methods for effective surveillance and management, as traditional gro

Normalizing Trajectory Models

ResearchDGX agent

arXiv:2605.08078v1 Announce Type: new Abstract: Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a

Not All Tokens Need 40 Steps: Heterogeneous Step Allocation in Diffusion Transformers for Efficient Video Generation

ResearchDGX agent

arXiv:2605.06892v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have achieved state-of-the-art video generation quality, but they incur immense computational cost because standard infere

NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2508.01248v4 Announce Type: replace Abstract: The rapid progress of generative models, such as GANs and diffusion models, has facilitated the creation of highly realistic images, raising growing

Object Hallucination-Free Reinforcement Unlearning for Vision-Language Models

SafetyDGX agent

arXiv:2605.08031v1 Announce Type: new Abstract: Vision-language models (VLMs) raise growing concerns about privacy, copyright, and bias, motivating machine unlearning to remove sensitive knowledge. Ho

On the Role of Strain and Vorticity in Numerical Integration Error for Flow Matching

ResearchDGX agent

arXiv:2605.06680v1 Announce Type: cross Abstract: Flow matching generates data by integrating a learned velocity field, where the number of integration steps (NFE) directly determines inference cost.

One World, Dual Timeline: Decoupled Spatio-Temporal Gaussian Scene Graph for 4D Cooperative Driving Reconstruction

AgentsDGX agent

arXiv:2605.07910v1 Announce Type: new Abstract: Reconstructing dynamic scenes from Vehicle-to-Infrastructure Cooperative Autonomous Driving (VICAD) data is fundamentally complicated by temporal asynch

OneViewAll: Semantic Prior Guided One-View 6D Pose Estimation for Novel Objects

ApplicationsDGX agent

arXiv:2605.07023v1 Announce Type: new Abstract: In many practical 6D object pose estimation scenarios, we often have access to only a single real-world RGB-D reference view per object, typically witho

OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical Videos

ResearchDGX agent

arXiv:2605.07695v1 Announce Type: new Abstract: High-fidelity surgical video generation can greatly improve medical training and the development of AI, adapting these generative models for precise vid

PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation

ResearchDGX agent

arXiv:2605.07252v1 Announce Type: cross Abstract: Co-speech gesture generation aims to synthesize realistic body movements that are semantically coherent with speech and faithful to a user-specified g

PET-Adapter: Test-Time Domain Adaptation for Full and Limited-Angle PET Image Reconstruction

ResearchDGX agent

arXiv:2605.08030v1 Announce Type: new Abstract: Positron Emission Tomography (PET) image reconstruction is inherently challenged by Poisson noise and physical degradation factors, which are further ex

PicoEyes: Unified Gaze Estimation Framework for Mixed Reality with a Large-Scale Multi-View Dataset

ResearchDGX agent

arXiv:2605.07188v1 Announce Type: new Abstract: We present PicoEyes, a unified gaze estimation framework that directly predicts all key attributes of gaze, including 3D eye parameters, eye-region segm

PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.07574v1 Announce Type: new Abstract: Mainstream vision-language models (VLMs) fundamentally struggle with severe optical ambiguities, such as reflections and transparent objects, due to the

Pre-training Enables Extraordinary All-optical Image Denoising

ResearchDGX agent

arXiv:2605.07810v1 Announce Type: cross Abstract: Optical neural networks are emerging as powerful machine learning and information processing tools because of their potential advantages in speed and

Pretty Good Measurement for Radiomics: A Quantum-Inspired Multi-Class Classifier for Lung Cancer Subtyping and Prostate Cancer Risk Stratification

ResearchDGX agent

arXiv:2603.00223v2 Announce Type: replace Abstract: We investigate a quantum-inspired approach to supervised multi-class classification based on the Pretty Good Measurement (PGM), viewed as an operato

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition

Model ReleasesDGX agent

arXiv:2605.07154v1 Announce Type: new Abstract: Referring Audio-Visual Segmentation (Ref-AVS) seeks to localize and segment target objects in video frames based on visual, auditory, and textual referr

Probabilistic Object Detection with Conformal Prediction

SafetyDGX agent

arXiv:2605.07549v1 Announce Type: new Abstract: Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a su

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos

Model ReleasesDGX agent

arXiv:2512.03479v2 Announce Type: replace Abstract: Procedural activities are fundamentally driven by object state transitions, yet existing instructional video benchmarks remain action-centric and ca

Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment

SafetyDGX agent

arXiv:2605.08064v1 Announce Type: new Abstract: Spatial intelligence in vision-language models (VLMs) attracts research interest with the practical demand to reason in the 3D world.Despite promising r

PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection

SafetyDGX agent

arXiv:2509.26272v3 Announce Type: replace Abstract: The rapid rise of synthetic media has made deepfake detection a critical challenge for online safety and trust. Progress remains constrained by the

Radiologist-Guided Causal Concept Bottleneck Models for Chest X-Ray Interpretation

SafetyDGX agent

arXiv:2605.07785v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) in medical imaging aim to improve model interpretability by predicting intermediate clinical concepts before final diag

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation

Model ReleasesDGX agent

arXiv:2605.07334v1 Announce Type: new Abstract: Video Reasoning Segmentation (VRS) aims to segment target objects in videos based on implicit instructions that convey human intent and temporal logic.

Real-IAD MVN: A Multi-View Normal Vector Dataset and Benchmark for High-Fidelity Industrial Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.07149v1 Announce Type: new Abstract: Industrial Anomaly Detection (IAD) is critical for quality control, but existing methods struggle with subtle, geometric defects. Standard 2D (RGB) imag

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning

SafetyDGX agent

arXiv:2605.07477v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have achieved remarkable progress, however, many edited results still suffer from artifacts, unintended mo

Rebalancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictions

ResearchDGX agent

arXiv:2605.07945v1 Announce Type: new Abstract: We present CoopNet, an approach that improves the cooperation of co-trained networks by dynamically adapting the apportionment of gradient, to ensure eq

ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation

SafetyDGX agent

arXiv:2408.06747v4 Announce Type: replace Abstract: Recent works utilize CLIP to perform the challenging unsupervised semantic segmentation task where only images without annotations are available. Ho

RECON: Robust symmetry discovery via Explicit Canonical Orientation Normalization

ApplicationsDGX agent

arXiv:2505.13289v5 Announce Type: replace-cross Abstract: Real world data often exhibits unknown, instance-specific symmetries that rarely exactly match a transformation group G fixed a priori. Class-

← Previous
1…148149150151152…209
Next →