AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
23 Jun 2026

ZeroGVC: Zero-Shot Generative Video Compression with Autoregressive Diffusion Priors

ResearchDGX agent

arXiv:2606.22371v1 Announce Type: cross Abstract: Recent generative video compression methods leverage powerful generative priors to achieve perceptually pleasing reconstructions. However, most existi

11 Jun 2026

3D-CBM: A Framework for Concept-Based Interpretability in Generative 3D Modeling

SafetyDGX agent

arXiv:2606.11446v1 Announce Type: new Abstract: This research introduces a framework for incorporating Concept Bottleneck Models (CBMs) into 3D generative architectures to address the inherent 'semant

4DP-QA: Scalable QA for 4D Perception in Vision Language Models


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2606.11568v1 Announce Type: new Abstract: Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D

A Comprehensive Ecosystem for Open-Domain Customized Video Generation

Model ReleasesDGX agent

arXiv:2606.11783v1 Announce Type: new Abstract: Recent progress in video generation has shown impressive visual synthesis capabilities. However, open-domain customized video generation remains limited

A Scalable PyTorch Abstraction for Multi-GPU Gaussian Splatting

HardwareDGX agent

arXiv:2606.11390v1 Announce Type: new Abstract: Gaussian splatting methods have become increasingly popular for neural reconstruction of the real world. However, they are often limited in scale and re

A Turbo-Inference Strategy for Object Detection and Instance Segmentation

Local AiDGX agent

arXiv:2606.12371v1 Announce Type: new Abstract: Object detection and instance segmentation tasks are closely related. Existing top-down instance segmentation methods usually follow a detect-then-segme

A2SG:Adaptive and Asymmetric Surrogate Gradients for Training Deep Spiking Neural Networks

ResearchDGX agent

arXiv:2606.11236v1 Announce Type: cross Abstract: Training deep spiking neural networks (SNNs) remains challenging due to sharp loss landscapes and temporal inconsistency caused by surrogate gradients

Adapting Vision-Language Models from Iconic to Inclusive for Multi-Label Recognition Without Labels

SafetyDGX agent

arXiv:2606.11626v1 Announce Type: new Abstract: Understanding multi-label images remains a challenging task in computer vision. With the rapid progress of vision-language multimodal learning, vision-l

Adv-TGD: Adversarial Text-Guided Diffusion for Face Recognition Impersonation Attacks

SafetyDGX agent

arXiv:2606.11615v1 Announce Type: new Abstract: The widespread adoption of face recognition (FR) technologies raises serious privacy concerns, as facial data can be exploited without consent. To addre

AerialClaw: An Open-Source Framework for LLM-Driven Autonomous Aerial Agents

SafetyDGX agent

arXiv:2606.12142v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs) are increasingly used in inspection, search and rescue, environmental monitoring, and emergency response. However, mos

AGE-MIL: Anchor-Guided Evidence Learning for Patient-Level Prediction

TutorialsDGX agent

arXiv:2606.12126v1 Announce Type: new Abstract: Existing computational pathology methods predominantly operate within whole-slide image (WSI)-level multiple instance learning (MIL) paradigms, while pa

An Electric Potential-Augmented Benchmark Dataset for Physics-Guided Image Reconstruction of Electrical Capacitance Tomography

Model ReleasesDGX agent

arXiv:2606.12226v1 Announce Type: new Abstract: While deep learning has significantly advanced image reconstruction of Electrical Capacitance Tomography (ECT), most data-driven methods map directly be

Anatomically Conditioned Recurrent Refinement for Topology-Aware Circle of Willis Segmentation

ResearchDGX agent

arXiv:2606.12319v1 Announce Type: new Abstract: Segmenting the Circle of Willis (CoW) from Magnetic Resonance Angiography (MRA) is challenging due to complex topology and thin vascular structures that

Auditing Demographic Bias in Facial Landmark Detection for Fair Human-Robot Interaction

SafetyDGX agent

arXiv:2604.06961v2 Announce Type: replace Abstract: Fairness in human-robot interaction critically depends on the reliability of the perceptual models that enable robots to interpret human behavior. W

Battery detection of XRay images using transfer learning

ResearchDGX agent

arXiv:2606.11779v1 Announce Type: new Abstract: The need for detecting and sorting batteries is drastically increasing for many applications. This study proves the potential of transfer learning in pr

Benchmarking Cross-Domain Audio-Visual Deception Detection

Model ReleasesDGX agent

arXiv:2405.06995v4 Announce Type: replace-cross Abstract: Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Convent

Beyond Dark Knowledge: Mixup-Based Distillation for Reliable Predictions

ResearchDGX agent

arXiv:2606.12171v1 Announce Type: new Abstract: Knowledge Distillation (KD) and mixup have proven effective at inducing smoothness in class boundaries; KD captures inherent class relationships in prob

Bridging Day and Night: Unsupervised Cross-Domain Re-Identification with Synergistic Prompt and Prototype Learning

SafetyDGX agent

arXiv:2606.12258v1 Announce Type: new Abstract: Cross-domain day-night re-identification (ReID) is fundamentally challenged by the substantial visual appearance discrepancies between daytime and night

Bridging the Modality Gap in Forensic Image Retrieval

ApplicationsDGX agent

arXiv:2606.12294v1 Announce Type: new Abstract: Automated image retrieval plays an increasingly critical role in modern forensic analysis, supporting investigative workflows that rely on efficient com

Causal Clothes-Invariant Feature Learning for Cloth-Changing Person Re-ID

TutorialsDGX agent

arXiv:2305.06145v2 Announce Type: replace Abstract: In cloth-changing person re-identification (CCReID), it is critical to learn clothes-invariant feature, which can provide discriminative ID features

CellNet -- Localizing Cells using Sparse and Noisy Point Annotations

ResearchDGX agent

arXiv:2606.12286v1 Announce Type: new Abstract: Counting living cells is an important step in many biological research workflows. Our collaborators at the Wellcome Sanger Institute study vital genes i

CFCamo: A Counterfactual Detect-or-Abstain Framework for Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2606.11231v1 Announce Type: new Abstract: Vision-language reinforcement learning has recently shown strong target-present localization for camouflaged object detection (COD). Yet localization is

Contactless 3D Human Body Measurement Using Depth Cameras for Smart Health Monitoring

ApplicationsDGX agent

arXiv:2606.11578v1 Announce Type: new Abstract: Contactless body measurement technologies are becoming increasingly significant for smart health monitoring, digital health applications, and remote pat

Continual Learning with Support Boundary Experience Blending

ResearchDGX agent

arXiv:2507.23534v3 Announce Type: replace-cross Abstract: Continual learning (CL) seeks to mitigate catastrophic forgetting when models are trained with sequential tasks. A common approach, experience

Corpus Augmentation for Sign Language Translation via LLM-Guided Video Stitching

SafetyDGX agent

arXiv:2606.11925v1 Announce Type: new Abstract: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and ena

CountZES: Counting via Zero-Shot Exemplar Selection

ResearchDGX agent

arXiv:2512.16415v3 Announce Type: replace Abstract: Object counting in complex scenes is particularly challenging in the zero-shot (ZS) setting, where instances of unseen categories are counted using

CoVR-R:Reason-Aware Composed Video Retrieval

Model ReleasesDGX agent

arXiv:2603.20190v2 Announce Type: replace Abstract: Composed Video Retrieval (CoVR) aims to find a target video given a reference video and a textual modification. Prior work assumes the modification

Cross-Domain Multi-Person Human Activity Recognition via Near-Field Wi-Fi Sensing

ResearchDGX agent

arXiv:2510.17816v2 Announce Type: replace-cross Abstract: Wi-Fi-based human activity recognition (HAR) provides substantial convenience and has emerged as a thriving research field, yet the coarse spa

Cross-Modal Benchmarking for Robotic Perception in Natural Environments

Model ReleasesDGX agent

arXiv:2606.11563v1 Announce Type: new Abstract: Natural environments present a complex challenge to robotics perception systems. Current models, particularly vision foundation models, are largely trai

DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model

ApplicationsDGX agent

arXiv:2606.12105v1 Announce Type: cross Abstract: Vision-language-action (VLA) models inherit a shared synchronous clock from vision-language pretraining, processing every input at one rate. This is m

Damage-TriageFormer: A Foundation-Model Framework for Typology-Based Building Damage Assessment from Mono-Temporal Imagery

Model ReleasesDGX agent

arXiv:2606.12248v1 Announce Type: new Abstract: Decision-relevant building damage assessment is critical for prioritizing resources and recovery after a disaster, yet most automated methods either fla

DarkVGGT: Seeing Through Darkness Using Thermal Geometry without Daylight Tax

ResearchDGX agent

arXiv:2606.11326v1 Announce Type: new Abstract: Recent feed-forward 3D reconstruction methods have demonstrated strong performance and flexibility in efficient end-to-end scene geometry estimation fro

DeceptionX: Explainable Deception Detection with Multimodal Large Language Models

ApplicationsDGX agent

arXiv:2606.11385v1 Announce Type: new Abstract: Deception detection is a critical and highly challenging task within affective computing and behavioral analysis. Existing deep learning methods typical

DepthMaster: Unified Monocular Depth Estimation for Perspective and Panoramic Images

TutorialsDGX agent

arXiv:2606.12368v1 Announce Type: new Abstract: While monocular depth estimation has achieved significant progress, achieving generalized metric depth estimation for both narrow field-of-view (FoV) pe

DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems

AgentsDGX agent

arXiv:2606.12236v1 Announce Type: cross Abstract: Many autonomous driving systems are increasingly incorporating foundation models to improve generalization and handle long-tail scenarios. However, th

DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace

Model ReleasesDGX agent

arXiv:2606.11687v1 Announce Type: new Abstract: Unmanned Aerial Vehicle (UAV) threats have emerged as a defining security challenge of the 21st century. This paper presents DroneShield-AI, a unified o

DynaTok: Token-Based 4D Reconstruction from Partial Point Clouds

ResearchDGX agent

arXiv:2606.12189v1 Announce Type: new Abstract: We address 4D reconstruction from partial point cloud sequences, where depth-sensor observations are incomplete, unordered, and lack explicit temporal c

Echoes of the Prior: A Computational Phenomenology of Forgetting

HardwareDGX agent

arXiv:2606.12340v1 Announce Type: new Abstract: Memory is not merely the storage of data; it is the scaffolding of reality. When biological memory fades, the world does not simply turn black; it regre

ERN-Net : Evolving Reason Node-Net for Document Binarization

ResearchDGX agent

arXiv:2606.11710v1 Announce Type: new Abstract: This paper presents ERN-Net, an Evolving Reason Node-Net for efficient document image binarization. ERN-Net enhances degradation-sensitive regions, such

EventRadar: Long-Range Visual UAV Discovery through Spatiotemporal Event Sensing

ResearchDGX agent

arXiv:2606.11285v1 Announce Type: new Abstract: Unauthorized unmanned aerial vehicle (UAV) activity around airports, public venues, and other sensitive sites has made protected-airspace monitoring inc

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

ResearchDGX agent

arXiv:2511.16672v4 Announce Type: replace Abstract: Recent advances in large multimodal models (LMMs) have enabled impressive reasoning and perception abilities, yet most existing training pipelines s

Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition

ResearchDGX agent

arXiv:2606.11450v1 Announce Type: new Abstract: Recently, masked skeleton reconstruction models have emerged as strong action representation learners, driving significant progress in self-supervised s

Feature extraction for plant growth estimation

AgentsDGX agent

arXiv:2606.11966v1 Announce Type: new Abstract: Precision agriculture requires the estimation of plant growth stages in real-time. When the plant growth stage is known, the wastage of resources in cul

Finding Sparse Subnetworks in One Training Cycle via Progressive Magnitude-Based Pruning

ResearchDGX agent

arXiv:2606.12278v1 Announce Type: new Abstract: Neural network pruning reduces model size by removing less important parameters while aiming to preserve predictive performance. Although the Lottery Ti

FitVTON: Fit-aware Virtual Try-On via Body-Garment Size Control

TutorialsDGX agent

arXiv:2606.12012v1 Announce Type: new Abstract: While diffusion-based virtual try-on has achieved impressive visual realism, most methods treat the task as 2D inpainting, prioritizing texture preserva

FreqKD: Frequency-Decoupled Cross-Modal Knowledge Distillation for Infrared Object Detection

SafetyDGX agent

arXiv:2606.11572v1 Announce Type: new Abstract: Transfer learning from large-scale RGB foundation models to infrared (IR) imagery through knowledge distillation (KD) remains challenging due to fundame

From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion

ResearchDGX agent

arXiv:2606.12303v1 Announce Type: new Abstract: Multimodal image fusion aims to integrate complementary information from different modalities into a fused image that preserves rich local details while

From Content to Knowledge: Lightning Fast Long-Video Understanding with Neural Knowledge Representations

Model ReleasesDGX agent

arXiv:2606.11913v1 Announce Type: new Abstract: We propose a new paradigm for long video understanding by treating a long video as a Neural Knowledge Representation (NKR). NKR represents video content

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models

SafetyDGX agent

arXiv:2602.08735v3 Announce Type: replace Abstract: While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, whic

From Nominal Intensity to Equivalent Rainfall: A Path-Based Credibility Evaluation Framework for Simulated Rainfall in Autonomous-Driving Perception Tests

AgentsDGX agent

arXiv:2606.11989v1 Announce Type: new Abstract: Credible simulated-rainfall conditions are essential for identifying perception-system boundaries and supporting SOTIF-oriented risk assessment in autom

From Simulation to Real-World: An In-Field 6D Pose Dataset and Baseline for Robotic Strawberry Harvesting

HardwareDGX agent

arXiv:2606.11381v1 Announce Type: new Abstract: Robotic strawberry harvesting requires precise 6D pose estimation; however, collecting 6D pose ground truth in real agricultural fields is inherently ch

Frozen Foundation-Model Embeddings Discard Small-Lesion Signal in Chest Radiography: Implications for Pre-Deployment Evaluation

Local AiDGX agent

arXiv:2606.11606v1 Announce Type: new Abstract: Frozen vision-transformer (ViT) foundation-model embeddings increasingly serve as the substrate for downstream chest-radiography (CXR) pipelines, yet wh

Higher order PCA-like rotation-invariant features for detailed shape descriptors modulo rotation

ResearchDGX agent

arXiv:2601.03326v2 Announce Type: replace Abstract: PCA can be used for rotation invariant features, describing a shape with its p_{ab}=E[(x_i-E[x_a])(x_b-E[x_b])] covariance matrix approximating shap

How Auxiliary Reasoning Unleashes GUI Grounding in VLMs

Model ReleasesDGX agent

arXiv:2509.11548v2 Announce Type: replace Abstract: Graphical user interface (GUI) grounding is a fundamental task for building GUI agents. However, general vision-language models (VLMs) struggle with

How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology

Model ReleasesDGX agent

arXiv:2606.12407v1 Announce Type: new Abstract: General-purpose large language models (LLMs) are routinely used as baselines when evaluating specialized pathology models on whole-slide images (WSIs).

i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models

Model ReleasesDGX agent

arXiv:2606.11289v1 Announce Type: new Abstract: Diffusion models have consistently driven progress in text-to-image generation. However, it is challenging to attribute recent progress to specific mode

Image Quality Assessment of Identity Cards Using Measures from Open Face Image Quality

ResearchDGX agent

arXiv:2606.11884v1 Announce Type: new Abstract: This paper addresses the challenge of assessing image quality in ID cards in remote verification systems by applying capture-related quality measures fr

Intelligent Skin Cancer Detection Using a Multispectral Metasurface and a Hybrid

Local AiDGX agent

arXiv:2606.11287v1 Announce Type: cross Abstract: Skin cancer is among the most prevalent malignancies worldwiAdbe satnradcitts early detection is essential for improving patient survival and reducing

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

SafetyDGX agent

arXiv:2606.12195v1 Announce Type: new Abstract: Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts large

ISAP-3D: Identity-Slot Aligned Part-Aware 3D Generation

SafetyDGX agent

arXiv:2606.12099v1 Announce Type: new Abstract: Part-aware 3D generation aims to synthesize structured objects with semantically meaningful components, yet often suffers from structural ambiguity due

← Previous
1…8384858687…211
Next →