AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
27 Jul 2026

Projection Pursuit CPCANet for Domain Generalization

SafetyDGX agent

arXiv:2607.22117v1 Announce Type: new Abstract: Domain Generalization (DG) aims to learn representations robust to distribution shifts. Recent geometric alignment methods, such as CPCANet, extract dom

Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models

TutorialsDGX agent

arXiv:2507.16257v2 Announce Type: replace Abstract: Defending pre-trained vision-language models (VLMs), such as CLIP, against adversarial attacks is crucial, as these models are widely used in divers

RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

Model ReleasesDGX agent

arXiv:2607.22293v1 Announce Type: new Abstract: Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ReCowGnition: A Realistic Biometric Benchmark for Cow Face Recognition

Model ReleasesDGX agent

arXiv:2607.22071v1 Announce Type: new Abstract: With the development of precision livestock farming and the advances in computer vision, visual animal biometrics has gained attention. Using biometric

Rethinking Layer-Wise Information Allocation for Vision Foundation Model Adaptation

Model ReleasesDGX agent

arXiv:2607.21973v1 Announce Type: new Abstract: Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central

Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification in the Foundation-Model Era

ResearchDGX agent

arXiv:2607.22068v1 Announce Type: new Abstract: Multi-branch architectures and CNN-Transformer fusion have long been regarded as effective ways to improve vehicle re-identification (Re-ID) by combinin

Risk-Routed Implicit Boundary Refinement for Robust Ultrasound Image Segmentation

Model ReleasesDGX agent

arXiv:2607.21787v1 Announce Type: new Abstract: Medical ultrasound (US) image segmentation faces significant challenges due to speckle noise, low-contrast boundaries, acoustic shadowing, and acquisiti

Robot-Factored World Models via Robot Rendering

TutorialsDGX agent

arXiv:2607.22535v1 Announce Type: cross Abstract: Action-conditioned video world models predict future observations from an initial observation and an action signal. In robotics, actions influence fut

SCALE: Self-Supervised Constraint-Aware Layout GEneration for Local P&R DRV Fixing at Advanced Nodes

Local AiDGX agent

arXiv:2607.21850v1 Announce Type: new Abstract: As semiconductor manufacturing advances toward sub-2nm nodes, local place-and-route (P&R) design-rule violation (DRV) fixing is increasingly limited by

SceneActBench: Can Agents Act on the 3D Scenes They See?

Model ReleasesDGX agent

arXiv:2607.22393v1 Announce Type: cross Abstract: Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual res

Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free Calibration

ResearchDGX agent

arXiv:2607.21673v1 Announce Type: cross Abstract: Test-time adaptive out-of-distribution (OOD) detectors update a memory bank from the unlabelled stream. We show this adaptation obeys a provable dynam

SiPhy: Single-Image Physical Property Reasoning

ResearchDGX agent

arXiv:2607.22355v1 Announce Type: new Abstract: Inferring physical properties such as mass, stiffness, and elasticity from a single image is essential for simulation and embodied AI, yet most existing

SLIP: Segmentation with Low-latency Interactive Prompting for 3D Medical Images

ResearchDGX agent

arXiv:2607.22332v1 Announce Type: new Abstract: Interactive deep image segmentation enables efficient medical image annotation by iteratively refining predictions from user prompts, such as positive a

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

ApplicationsDGX agent

arXiv:2607.22534v1 Announce Type: new Abstract: Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding rem

Spectral Prior for Reducing Exposure Bias in Diffusion Models

SafetyDGX agent

arXiv:2607.22091v1 Announce Type: new Abstract: Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequen

TDiR: Transformer based Diffusion for Image Restoration Tasks

ResearchDGX agent

arXiv:2506.20302v2 Announce Type: replace Abstract: Images captured in challenging environments often experience various types of degradation, such as noise, color cast, blur, and light scattering. Th

TextSLIP: Text Self-Supervised CLIP for Medical Report Generation

SafetyDGX agent

arXiv:2607.21970v1 Announce Type: new Abstract: Automating radiology report generation is important for improving reporting consistency and clinical workflows . While Contrastive Language--Image Pretr

The 3D Mirage: Probing and Taming 3D Hallucinations

Model ReleasesDGX agent

arXiv:2512.15423v2 Announce Type: replace Abstract: Monocular depth foundation models achieve remarkable generalization by learning large-scale semantic priors, but this creates a critical vulnerabili

The Lift Spectrum: How Measurement-to-Space Adaptivity Shapes Robustness in Image-Free Single-Pixel Sensing

Model ReleasesDGX agent

arXiv:2607.22077v1 Announce Type: cross Abstract: Single-pixel sensing encodes a scene as a short sequence of coded measurements, and image-free methods infer the task directly from that sequence. Rem

The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis

Local AiDGX agent

arXiv:2601.19618v2 Announce Type: replace Abstract: Differential privacy protects the patients whose images train medical imaging models, but it lowers diagnostic accuracy, and the initialization is t

Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions

Model ReleasesDGX agent

arXiv:2607.22352v1 Announce Type: new Abstract: We introduce time-reversed imaging, a new paradigm that infers what just happened in a scene from fading multimodal traces. Instead of extrapolating or

Toward High-Fidelity 3D Point-Cloud Learning for Brain Folding Morphology Prediction Using Trans-Unet

Local AiDGX agent

arXiv:2607.21840v1 Announce Type: new Abstract: Learning high-fidelity point-cloud features in the 3D space poses significant challenges, including permutation invariance, lack of local context, diffi

TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution

Local AiDGX agent

arXiv:2607.22231v1 Announce Type: new Abstract: Video super-resolution (VSR) using large-scale Diffusion Transformer (DiT) priors achieves exceptional perceptual quality but is often impractical due t

Twins: Learn to Predict Unified Representations with Focal Loss

SafetyDGX agent

arXiv:2607.22531v1 Announce Type: new Abstract: Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the

Visual Relocalization from Sparse Views in Aliased and Low-Texture Environments via Novel View Synthesis

Local AiDGX agent

arXiv:2607.22147v1 Announce Type: new Abstract: Visual localization becomes extremely challenging in planetary-like terrains characterized by low texture, perceptual aliasing, harsh illumination, and

Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning

TutorialsDGX agent

arXiv:2607.22013v1 Announce Type: new Abstract: Multimodal chain-of-thought (CoT) reasoning integrates visual and textual cues through step-by-step inference. In small models with limited token budget

VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory

AgentsDGX agent

arXiv:2607.14514v2 Announce Type: replace Abstract: Training-free ObjectNav agents increasingly use vision-language models (VLMs), yet typically discard acquired scene knowledge after each request. We

What Happens to Accuracy When Photo Lineups Contain Non-Mated Rank-One Images From Large Galleries?

ResearchDGX agent

arXiv:2607.21792v1 Announce Type: new Abstract: One-to-many facial identification is commonly used to match a probe image from surveillance video against a gallery of driver's licenses and/or booking

24 Jul 2026

3D-GIMP: When 3D Gaussian Inpainting Meets PatchMatch

ResearchDGX agent

arXiv:2607.20789v1 Announce Type: new Abstract: Recent advances in 3D scene editing have leveraged iterative diffusion models to update input views. However, this process is computationally expensive

A real-time RGB-D perception pipeline for autonomous impact hammers in mining: self-filtering, rock segmentation and rock-breaking poses generation

AgentsDGX agent

arXiv:2607.20748v1 Announce Type: cross Abstract: Impact hammers, also known as rock-breakers, are essential machines in mining operations, where they perform secondary reduction. In underground minin

Achieving Text-based Person Retrieval with Any Granularity

Model ReleasesDGX agent

arXiv:2607.21057v1 Announce Type: new Abstract: Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This p

Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation

Model ReleasesDGX agent

arXiv:2607.20866v1 Announce Type: new Abstract: Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamenta

ASTRA-Net: Anatomy-Specific Transfer and Representation Alignment for Drug-Induced Sleep Endoscopy Segmentation

SafetyDGX agent

arXiv:2607.21370v1 Announce Type: new Abstract: Quantitative drug-induced sleep endoscopy (DISE) requires reliable airway boundaries at specific anatomical levels. Pixel-level DISE annotations are sca

AUCH-Net: Action Unit-Based Consistency-Aware Hypergraph Network for Cross-Domain Few-Shot Facial Expression Recognition

ResearchDGX agent

arXiv:2607.21004v1 Announce Type: new Abstract: Recently, cross-domain few-shot facial expression recognition (CF-FER) has received considerable attention. However, the performance of existing CF-FER

Axolotl3D: a Unified Framework for Faithful 3D Shape Completion

SafetyDGX agent

arXiv:2607.20660v1 Announce Type: new Abstract: Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and diffusion architectures. However, they assume

Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving

AgentsDGX agent

arXiv:2607.21526v1 Announce Type: new Abstract: Self-supervised depth estimation is challenging for safe autonomous driving under various adverse weather conditions due to sensor perception degradatio

C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs

TutorialsDGX agent

arXiv:2607.21076v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) require huge memory and computational costs, which limits their practical deployment. Post-training quantizatio

Causal-AgentIR: Self-Evolving Causal Memory for Adaptive Image Restoration Agents

AgentsDGX agent

arXiv:2607.21125v1 Announce Type: new Abstract: Image restoration agents have recently emerged as a flexible paradigm for handling diverse and unpredictable degradations in real-world scenarios. Exist

CLUIE: Clustering-Aware Recurrent Propagation with Local Structural Compensation for Underwater Image Enhancement

ResearchDGX agent

arXiv:2607.21467v1 Announce Type: new Abstract: Underwater image enhancement remains challenging due to wavelength-dependent light absorption, scattering, and backscattering, which jointly cause color

Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification

Local AiDGX agent

arXiv:2607.21068v1 Announce Type: cross Abstract: Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reduci

CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging

Model ReleasesDGX agent

arXiv:2607.20561v1 Announce Type: cross Abstract: LoRA adapters provide an efficient way to specialize a pretrained model for many downstream tasks, but deploying one adapter per task requires adapter

DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV

AgentsDGX agent

arXiv:2607.21438v1 Announce Type: new Abstract: Monocular depth estimation is a fundamental prerequisite for 3D reconstruction and autonomous navigation in Unmanned Aerial Vehicles (UAVs). In practica

DART: A Degradation-Aware Recurrent Transformer for Archival Film Restoration

TutorialsDGX agent

arXiv:2607.21219v1 Announce Type: new Abstract: Archival film restoration is a challenging problem because historical footage contains compound degradations such as scratches, dust, blur, noise, flick

DCVC-MV: Deep Contextual Multiview Video Compression with Efficient Inter-View Prediction

TutorialsDGX agent

arXiv:2509.03922v2 Announce Type: replace Abstract: Multiview video is a key format for 3D applications such as free-viewpoint broadcasting and virtual reality, yet its large data volume poses signifi

Decoupling Cross-Modality Manifold Discrepancy: Leveraging Visible Diffusion Priors for Infrared Super-Resolution

Local AiDGX agent

arXiv:2607.21174v1 Announce Type: new Abstract: Infrared image super-resolution (IISR) mitigates the limitations imposed by low spatial resolution. Existing methods have recognized that IISR should pr

Detecting Neural Network Failures through Spectral Analysis of Internal Activations

Model ReleasesDGX agent

arXiv:2607.20590v1 Announce Type: cross Abstract: Neural network misclassifications exhibit characteristic spectral instability in internal activations that is invisible at the output layer. This phen

Detectors Learn the Wrong Thing: Shortcut-Resistant Adversarial Training Against Physically Realizable Attacks

SafetyDGX agent

arXiv:2607.21243v1 Announce Type: new Abstract: AI-enabled visual perception systems are increasingly deployed in intelligent transportation infrastructure and autonomous vehicle related applications.

DINO-VPT: Hierarchical Visual Prompt Tuning for Joint Physical-Digital Face Anti-Spoofing

Model ReleasesDGX agent

arXiv:2607.20900v1 Announce Type: new Abstract: With the increasing diversity of spoofing attacks, there is a growing demand for unified Face Anti-Spoofing (FAS) models capable of detecting both physi

Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

SafetyDGX agent

arXiv:2607.20984v1 Announce Type: new Abstract: This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task

Do Pathology Vision-Language Models Truly See Pathology?

Model ReleasesDGX agent

arXiv:2607.21065v1 Announce Type: new Abstract: Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. Howe

DTIF: Robust Loop Closure Detection via Delaunay Triangle Topology in Complex Forests

SafetyDGX agent

arXiv:2607.21138v1 Announce Type: new Abstract: Accurate forest inventory and large-scale mapping are essential for ecosystem monitoring and sustainable forest management. Multiple low-cost edge platf

EmoSpace: Immersive Affective Image Generation Guided by Fine-Grained Emotion Prototypes

SafetyDGX agent

arXiv:2602.11658v2 Announce Type: replace Abstract: Immersive affective content generation aims to create visually compelling VR imagery with controllable emotional nuance, yet existing methods typica

Engine-Native Editable 3D World Reconstruction with Objects and Lighting

Model ReleasesDGX agent

arXiv:2607.20889v1 Announce Type: new Abstract: Editable 3D scene creation requires object instances and lights that can be inspected, moved, and imported into standard engines, yet existing single-im

Evaluation and Prognostic Validation of Deep Regression Models for WSI-Based Gene-Expression Prediction

ResearchDGX agent

arXiv:2410.00945v2 Announce Type: replace-cross Abstract: Gene-expression profiling is widely used in research and central to many areas of precision oncology, but remains costly and not universally a

Explainable Deepfake Detection Challenge

Model ReleasesDGX agent

arXiv:2607.21007v1 Announce Type: new Abstract: Deepfake detection is moving beyond binary classification decisions toward systems that can also explain the visual evidence supporting those decisions.

Explainable graph attention network for stress recognition (StressGAT) via differential action units

SafetyDGX agent

arXiv:2607.20819v1 Announce Type: new Abstract: Stress is a dynamic process characterized by significant individual variability in facial expression. Traditional architectures, such as Recurrent Neura

FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

ResearchDGX agent

arXiv:2607.20922v1 Announce Type: new Abstract: We propose FA-LAM, a Focus-Aware Large Avatar Model for one-shot animatable Gaussian head creation, while simultaneously enabling static 3D and dynamic

Fitting Generalized Power Diagrams to 3D Image Data: A Prerequisite for Virtual Materials Testing

ResearchDGX agent

arXiv:2507.14268v2 Announce Type: replace Abstract: This paper reviews algorithmic and modeling approaches for fitting generalized power diagrams to three-dimensional image data, a key step in virtual

Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform

Model ReleasesDGX agent

arXiv:2607.21271v1 Announce Type: new Abstract: Equivariant networks embed geometric symmetries as structural priors through weight sharing, achieving remarkable parameter efficiency across vision tas

Focus on What Matters: Constraining Spatial-Temporal Attention via Action-Units for Noise-Resilient AQA

TutorialsDGX agent

arXiv:2511.05611v2 Announce Type: replace Abstract: The core challenge in Action Quality Assessment (AQA) lies in extracting fine-grained motion features from redundant and complex video backgrounds.

← Previous
1…3334353637…209
Next →