AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
29 Jul 2026

Adversarial Deepfake Generation and an Investigation of Purification-Based Adversarial Detection

ResearchDGX agent

arXiv:2607.25842v1 Announce Type: new Abstract: This paper describes the participation of team 'Go To Germany' in the ImageCLEF 2026 Deepfake Detection and Generation Task. For the image generation ta

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

SafetyDGX agent

arXiv:2607.25489v1 Announce Type: new Abstract: Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and

ANFI: Rethinking Neighbor Feature Interaction in Person Re-ID

ResearchDGX agent

arXiv:2607.25407v1 Announce Type: new Abstract: In person re-identification, neighbor-based methods have achieved significant success by interacting with neighbor samples to obtain more robust represe


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

Model ReleasesDGX agent

arXiv:2607.24821v1 Announce Type: cross Abstract: While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modali

Beyond Background Bias: Saliency-Driven Prototype Alignment for Dataset Distillation

Model ReleasesDGX agent

arXiv:2607.25318v1 Announce Type: new Abstract: Dataset distillation aims to synthesize compact datasets that can approximate the performance of full-data training while significantly reducing computa

Beyond Facial Consistency: Personalized Person Image Generation with Holistic Identity Preservation

Model ReleasesDGX agent

arXiv:2607.25622v1 Announce Type: new Abstract: Personalized person image generation requires preserving subject identity across both local facial details and broader appearance cues. Existing methods

Beyond Static Costs: Learning-Dynamics Aware Loss Functions for Long-Tailed Classification

Model ReleasesDGX agent

arXiv:2607.25830v1 Announce Type: new Abstract: Deep learning models in computer vision face significant challenges when trained on long-tailed datasets, where a few majority classes dominate while ma

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

Local AiDGX agent

arXiv:2607.25993v1 Announce Type: new Abstract: Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental

Bi-Level Collaborative Learning for Few-Shot Scribble-Supervised Medical Image Segmentation

TutorialsDGX agent

arXiv:2607.25432v1 Announce Type: new Abstract: Scribble annotations offer an efficient alternative to costly pixel-wise labeling for medical image segmentation, yet in real clinical scenarios, scribb

CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2607.25239v1 Announce Type: new Abstract: Referring multi-object tracking (RMOT) extends tracking from category-driven perception to language-guided understanding by grounding object trajectorie

CFR-Net:Collaborative Feature Refinement Network for Medical Image Anomaly Detection

ResearchDGX agent

arXiv:2607.11509v2 Announce Type: replace Abstract: Medical image anomaly detection remains challenging because networks pretrained on natural images often exhibit limited adaptability to medical imag

Cinematic Compositing Using Character-Environment-Harmonized Video Generation Models

ResearchDGX agent

arXiv:2606.20233v2 Announce Type: replace Abstract: Cinematic compositing aims to integrate green-screen characters into novel environments while maintaining physical and photometric realism. Previous

COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding

Model ReleasesDGX agent

arXiv:2409.12760v3 Announce Type: replace Abstract: To help address the occlusion problem in panoptic segmentation and image understanding, this paper proposes a new large-scale dataset named COCO-OLA

DensFiLM: Density-Conditioned Video Saliency for Crowd Scenes

SafetyDGX agent

arXiv:2607.25465v1 Announce Type: new Abstract: Video saliency models typically apply a single fixation strategy across crowd scenes, despite systematic changes in attention with crowd density. Sparse

Depth to Anatomy: Organ Localization from Depth Images for Automated Patient Table Positioning in Radiology Workflow

ApplicationsDGX agent

arXiv:2601.18260v3 Announce Type: replace Abstract: In clinical radiology, accurate patient table positioning is essential to align specific internal organs of interest with the scanner imaging isocen

Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models

ResearchDGX agent

arXiv:2607.25078v1 Announce Type: new Abstract: Generative diffusion models have revolutionized facial image synthesis, yet robust identity preservation in high resolution outputs remains a critical c

Diff2DGS: Reliable Reconstruction of Occluded Surgical Scenes via 2D Gaussian Splatting

ResearchDGX agent

arXiv:2602.18314v2 Announce Type: replace Abstract: Real-time reconstruction of deformable surgical scenes is vital for advancing robotic surgery, improving intraoperative guidance, and enabling autom

DopQ-ViT: Towards Distribution-Friendly and Outlier-Aware Post-Training Quantization for Vision Transformers

ResearchDGX agent

arXiv:2408.03291v4 Announce Type: replace Abstract: Vision Transformers (ViTs) have gained significant attention, but their high computing cost limits the practical applications. While post-training q

Enabling Fully Integer-Only Inference for Lightweight Detection Transformers

ApplicationsDGX agent

arXiv:2607.24981v1 Announce Type: new Abstract: Vision Transformer detectors now approach the accuracy of CNNs but remain difficult to deploy on NPUs and microcontrollers because key components, inclu

End-to-End 3-D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration

AgentsDGX agent

arXiv:2512.21831v2 Announce Type: replace Abstract: Multiview cooperative perception and multimodal fusion are essential for reliable 3-D spatiotemporal understanding in autonomous driving, especially

Explicit Layer Modeling for Video Object Insertion and Layer Decomposition

TutorialsDGX agent

arXiv:2607.25802v1 Announce Type: new Abstract: Most video editing systems still lack explicit layered video representations, limiting their ability to perform realistic compositing, object reuse, and

Faces of Fairness: Examining Bias in Facial Expression Recognition Datasets and Models

SafetyDGX agent

arXiv:2502.11049v3 Announce Type: replace Abstract: Automated Facial Expression Recognition (FER), involves two critical aspects: data and model design. Both significantly influence bias and fairness

Few-Shot Open-Vocabulary Remote Sensing Segmentation via Textual Inversion

Model ReleasesDGX agent

arXiv:2607.25563v1 Announce Type: new Abstract: Open-vocabulary segmentation labels arbitrary categories from a text query without per-class training, yet on remote sensing imagery it underperforms on

FIDAC: An Easy-to-use Pipeline to Extract and Interpret Interpersonal Distance From Video

ResearchDGX agent

arXiv:2607.25146v1 Announce Type: new Abstract: The distance between persons reveals significant information about their perception of each other. However, such information is not easily extractable a

Fine-Grained Food Image Understanding via Target-Aware Data Alignment

SafetyDGX agent

arXiv:2607.25794v1 Announce Type: new Abstract: Fine-grained food visual--semantic understanding requires models to capture subtle distinctions across ingredients, cooking methods, doneness, color, te

FLASH: Efficient Impact Fall Detection with Unified Hypergraph State-Space Model

ResearchDGX agent

arXiv:2607.25791v1 Announce Type: new Abstract: Falls represent a critical public health challenge, and accurate detection of the impact moment when an individual hits the ground is crucial for timely

Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion

Model ReleasesDGX agent

arXiv:2607.25820v1 Announce Type: new Abstract: Food image segmentation plays a vital role in health-related applications such as nutrition tracking and personalized health monitoring. However, existi

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

ResearchDGX agent

arXiv:2607.25266v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. However, the density of

FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation

ResearchDGX agent

arXiv:2605.06421v2 Announce Type: replace Abstract: Pixel-space diffusion has re-emerged as a promising alternative to latent-space generation because it avoids the representation bottleneck introduce

Freq-RemoteVAR: Next-Frequency Autoregressive Modeling for Remote Sensing Change Detection

ResearchDGX agent

arXiv:2607.25815v1 Announce Type: new Abstract: Remote sensing change detection aims to identify land-cover changes from bi-temporal images. Most existing methods follow a one-shot dense prediction pa

Fundamental Recovery Bounds for SPAD Signals under Stationary Flux

ResearchDGX agent

arXiv:2601.07599v4 Announce Type: replace Abstract: Single-photon avalanche diodes (SPADs) record light as a discrete stream of individual detections. The signal is stochastic. Its statistical structu

FunnelAL: Retrieve-then-Rank Active Learning for Single-Class Discovery

ResearchDGX agent

arXiv:2607.25276v1 Announce Type: new Abstract: We present FunnelAL, a retrieve-then-rank active learning system for single-class discovery, which adapts the multi-stage funnel architecture of industr

Gaussian Volumetric Representation for Efficient Shear-Warp Visualization

TutorialsDGX agent

arXiv:2607.25377v1 Announce Type: new Abstract: Medical image visualization requires volumetric rendering algorithms that preserve anatomical fidelity while maintaining high rendering speeds. To addre

GeoMFD: Continual Drone-View Geo-Localization with Geometry-Aware Adapter and Margin-Field Distillation

Local AiDGX agent

arXiv:2607.25788v1 Announce Type: new Abstract: Existing drone-view geo-localization (DVGL) methods are mainly developed under a static training paradigm, where models are optimized for fixed environm

Gradient-Based Latent Decomposition Reveals Mechanisms of Feature Degradation in Weakly Supervised Mammography

ResearchDGX agent

arXiv:2607.24835v1 Announce Type: new Abstract: Weakly supervised hierarchical models exhibit a persistent asymmetry: coarse lesion-type features are preserved under reconstruction while fine-grained

Group Equivariant Diffusion for Anomaly Detection in Computational Cytology

SafetyDGX agent

arXiv:2607.25503v1 Announce Type: new Abstract: Computational cytology on whole-slide images is challenging because malignant cells are rare, heterogeneous, and annotated slides are scarce. Anomaly de

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Local AiDGX agent

arXiv:2607.25895v1 Announce Type: cross Abstract: Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is

HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing

ResearchDGX agent

arXiv:2508.05899v3 Announce Type: replace Abstract: 3D scene generation plays a crucial role in gaming, artistic creation, virtual reality, and many other domains. However, current 3D scene design sti

HOME: Robust Hough-space Matching Method for Structured and Textureless Videos

ResearchDGX agent

arXiv:2607.25389v1 Announce Type: new Abstract: Visual front-ends for robotic localization typically rely on point-based features such as Oriented FAST and Rotated BRIEF (ORB), which frequently fail i

Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection

ResearchDGX agent

arXiv:2607.25310v1 Announce Type: new Abstract: Hyperspectral imaging (HSI) is useful for material discrimination, but operational mine screening also depends on how many false alarms must be inspecte

Hyperspectral Intrinsic Decomposition: Joint Recovery of Reflectance and Photometric Components for Non-Lambertian Scenes

ApplicationsDGX agent

arXiv:2607.25371v1 Announce Type: new Abstract: Hyperspectral intrinsic decomposition (HID) aims to disentangle material-related spectral properties and photometric effects in hyperspectral images (HS

Impact Detection in Fall Events: Leveraging Spatio-Temporal Graph Convolutional Networks and Recurrent Neural Networks Using 3D Skeletons Data

ApplicationsDGX agent

arXiv:2607.25710v1 Announce Type: new Abstract: Fall represents a significant risk of accidental death among individuals aged over 65, presenting a global health concern. A fall is defined as any even

IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation

Model ReleasesDGX agent

arXiv:2607.25106v1 Announce Type: new Abstract: Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (Obj

Intrinsic and Triangulation-Agnostic Attention: A Simple and Powerful Approach for Learning on Meshes

ResearchDGX agent

arXiv:2607.24954v1 Announce Type: cross Abstract: This work proposes an adaptation of the attention mechanism for triangle meshes. The core observation is that endowing the attention mechanism with cr

Keypoint-Guided Optimal Transport: Models, Algorithms, and Applications

SafetyDGX agent

arXiv:2303.13102v2 Announce Type: replace Abstract: Existing Optimal Transport (OT) methods mainly derive the optimal transport plan/matching under the criterion of transport cost/distance minimizatio

LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

Model ReleasesDGX agent

arXiv:2607.25962v1 Announce Type: new Abstract: Recent generative models can produce images with few obvious visual artifacts, weakening detectors and explanations that rely only on surface appearance

Leak-Free Cross-Validated Stacking with Per-Architecture Calibration for Sand-Boil Segmentation in Earthen Levees

ResearchDGX agent

arXiv:2607.25367v1 Announce Type: new Abstract: Sand boils, points where water seeping beneath an earthen levee re-emerges at the surface, are early warnings of internal erosion, and deep segmentation

LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos

ResearchDGX agent

arXiv:2607.25125v1 Announce Type: new Abstract: Despite rapid progress in Multi-modal Large Language Models (MLLMs), understanding long-form videos is still bottlenecked by limited context windows. Wh

Leveraging Semantic Maps for City-Scale Cross-View Localization

ResearchDGX agent

arXiv:2607.25215v1 Announce Type: cross Abstract: We want robots to localize in previously untraversed environments against commonly available prior data. Rich semantic data available from OpenStreetM

LGFNet: A CTC-Guided Local-Global Fusion Framework for Single-Channel Sleep Staging

SafetyDGX agent

arXiv:2607.25197v1 Announce Type: new Abstract: Sleep staging remains challenging due to long-range temporal dependencies, ambiguous stage transitions-particularly in N1-and substantial distribution s

Med-SegLens: Latent-Level Model Diffing for Interpretable Medical Image Segmentation

SafetyDGX agent

arXiv:2602.10508v2 Announce Type: replace Abstract: Modern segmentation models achieve strong predictive performance but remain largely opaque, limiting our ability to diagnose failures, understand da

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence

ApplicationsDGX agent

arXiv:2603.27176v2 Announce Type: replace Abstract: Lesion detection, symptom tracking, and visual explainability are central to real-world medical image analysis, yet current medical Vision-Language

Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation

SafetyDGX agent

arXiv:2607.25242v1 Announce Type: new Abstract: Medical world models offer a framework for extending medical artificial intelligence beyond static prediction by representing evolving patient states an

MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

Model ReleasesDGX agent

arXiv:2607.25300v1 Announce Type: new Abstract: Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes

Mondrian: On-Device High-Performance Video Analytics with Compressive Packed Inference

Local AiDGX agent

arXiv:2403.07598v2 Announce Type: replace Abstract: In this paper, we present Mondrian, an edge system that enables high-performance object detection on high-resolution video streams. Many lightweight

MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks

Model ReleasesDGX agent

arXiv:2607.25092v1 Announce Type: new Abstract: Face morphing attacks create synthetic images verifiable against multiple identities, threatening border control and identity verification systems. We i

NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model

SafetyDGX agent

arXiv:2607.24789v1 Announce Type: cross Abstract: We present NEXT (Next-interest EXploration Transformer), a reasoning-driven video recommendation framework that reasons over the video a user has just

Noise-Free One-Step LoRA for Task-Driven Image Restoration with Diffusion Priors

ApplicationsDGX agent

arXiv:2607.25390v1 Announce Type: new Abstract: Degraded images not only reduce visual quality but also impair downstream high-level vision tasks. Task-driven image restoration (TDIR) addresses this i

ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction

Model ReleasesDGX agent

arXiv:2607.25210v1 Announce Type: new Abstract: Oblique-view urban remote sensing imagery inevitably exhibits geometric projection displacements between building roofs and footprints, leading to signi

On the Use of Synthetic Data for Threshold Calibration in Face Recognition: Performance and Security Implications for Border Control Systems

SafetyDGX agent

arXiv:2607.25990v1 Announce Type: new Abstract: The recently deployed Entry/Exit System (EES) introduces large-scale biometric verification into European border control, requiring face recognition sys

← Previous
1…2728293031…209
Next →