AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
24 Jun 2026

A Geometry-Informed Computer Vision Method for Detecting and Examining Overtaking Vehicles From A Bicycle

SafetyDGX agent

arXiv:2606.23699v1 Announce Type: new Abstract: Instrumented bicycle studies have produced direct field evidence on vehicle passing behavior, but extracting overtaking events from continuous rear-faci

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP

ResearchDGX agent

arXiv:2603.05962v2 Announce Type: replace Abstract: To address the limitations of existing open-vocabulary object recognition methods, including high system complexity, substantial training costs, and

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.23835v1 Announce Type: new Abstract: ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generati

Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction

ResearchDGX agent

arXiv:2606.24156v1 Announce Type: new Abstract: Visual token reduction has emerged as an effective strategy for accelerating Multimodal Large Language Models (MLLMs). Many existing methods prune token

ActiveScope: Actively Seeking and Correcting Perception for MLLMs

Local AiDGX agent

arXiv:2606.24292v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated impressive vision-language understanding, yet still struggle with fine-grained perception in

Adaptive Hebbian Memory Routing in Vision Transformers for Few-Shot Learning

ResearchDGX agent

arXiv:2606.24756v1 Announce Type: new Abstract: Few-shot image recognition requires models to adapt to new classes from a small labeled support set. Hebbian fast-weight memory can provide temporary as

Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods

ResearchDGX agent

arXiv:2606.24484v1 Announce Type: new Abstract: WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially mo

AerialFusionMapNet: Online HD Map Construction with Aerial-Onboard BEV Fusion

ResearchDGX agent

arXiv:2606.24784v1 Announce Type: new Abstract: High-resolution aerial imagery has recently emerged as a complementary modality for automated driving perception and has shown potential to improve bird

Agentic Collaborative Cognition for Zero-Shot 3D Understanding

AgentsDGX agent

arXiv:2606.24649v1 Announce Type: new Abstract: Recent advancements have explored agentic zero-shot 3D understanding by reformulating it as video keyframe understanding with Multimodal Large Language

An LMM for Precisely Grounding Elements in Documents

Local AiDGX agent

arXiv:2606.24118v1 Announce Type: new Abstract: Visual grounding in documents is a crucial ability for Large Multimodal Models (LMMs) in areas such as document understanding, deep research and documen

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

Model ReleasesDGX agent

arXiv:2606.24548v1 Announce Type: new Abstract: Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it rem

ArtiTwinSplat: Interactable Digital Twin Reconstruction via Gaussian Splatting from RGB-D videos

ApplicationsDGX agent

arXiv:2606.24628v1 Announce Type: cross Abstract: Deploying robots in unstructured real-world environments needs accurate, interactive models of the objects. Constructing these models at scale remains

Automated Residual Plot Assessment With the R Package autovi and the Shiny Application autovi.web

ResearchDGX agent

arXiv:2606.24236v1 Announce Type: cross Abstract: Visual assessment of residual plots is a common approach for diagnosing linear models, but it relies on manual evaluation, which does not scale well a

Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models

AgentsDGX agent

arXiv:2606.24152v1 Announce Type: new Abstract: Existing literature claims that video generation essentially is world modelling. On the one hand, the claim is productive because it pushes generative A

BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases

Model ReleasesDGX agent

arXiv:2606.24883v1 Announce Type: new Abstract: Artificial intelligence (AI) has achieved remarkable success in medical imaging, but it is widely recognized that these models often perform inconsisten

Bengal-HP_RU: A Dataset of Bengal People For Head Pose Estimation

ResearchDGX agent

arXiv:2606.24122v1 Announce Type: new Abstract: Existing head pose datasets predominantly feature subjects of Western or East Asian origin, leaving South Asian populations, particularly Bengali indivi

BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming

Model ReleasesDGX agent

arXiv:2606.24740v1 Announce Type: new Abstract: Recent advances in vision-language models (VLMs) such as CLIP have demonstrated strong generalization across natural-image domains. However, adapting th

Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation

SafetyDGX agent

arXiv:2606.24464v1 Announce Type: new Abstract: Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing mod

Bridging the Manifold Gap: Riemannian Residual Line Search for One-Step Image Editing

SafetyDGX agent

arXiv:2606.24844v1 Announce Type: new Abstract: One-step diffusion editors are fast because they avoid inversion and iterative optimization, but a single transport update must be aggressive enough to

CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities

Model ReleasesDGX agent

arXiv:2506.08690v3 Announce Type: replace Abstract: Canada experienced in 2023 one of the most severe wildfire seasons in recent history, causing damage across ecosystems, destroying communities, and

Compact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual Relocalization

Model ReleasesDGX agent

arXiv:2606.24767v1 Announce Type: new Abstract: Indoor visual relocalization plays a critical role in emerging spatial and embodied AI applications. However, prior research was predominantly devoted t

Configurable Holography: Towards Display and Scene Adaptation

ResearchDGX agent

arXiv:2405.01558v4 Announce Type: replace Abstract: Rendering holograms for holographic displays is often an iterative and computationally costly process. Emerging learned holography methods have alle

Counting Trees from Satellite Imagery with Noisy Supervision

Model ReleasesDGX agent

arXiv:2606.24786v1 Announce Type: new Abstract: Counting individual trees is a fundamental task for environmental monitoring, yet remains largely unexplored with satellite imagery. At these resolution

CrossFusion: A Multi-Scale Cross-Attention Convolutional Fusion Model for Cancer Survival Prediction

ResearchDGX agent

arXiv:2503.02064v2 Announce Type: replace-cross Abstract: Cancer survival prediction from whole slide images (WSIs) is a challenging task in computational pathology due to the large size, irregular sh

Cyclic Denoising Reveals Ultrastable Memories in Diffusion Models

ResearchDGX agent

arXiv:2606.24000v1 Announce Type: cross Abstract: We introduce cyclic denoising -- repeated forward and reverse diffusion at controlled noise amplitudes -- as an extraction attack for image diffusion

Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation

TutorialsDGX agent

arXiv:2606.18478v2 Announce Type: replace Abstract: Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matchin

DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection

SafetyDGX agent

arXiv:2606.24805v1 Announce Type: new Abstract: Stereo-based 3D object detection still faces two critical safety challenges: real-time performance and open-set generalization. Existing stereo 3D metho

Differentiable Packing of Irregular 3D Objects with Adaptive Container Estimation

HardwareDGX agent

arXiv:2606.16333v2 Announce Type: replace Abstract: Most existing approaches either fix the container in advance or optimize only a single container dimension through an outer search loop, leaving the

Differential Unfolding: Efficient Unfolding Reconstruction for Video Snapshot Compressive Imaging

Model ReleasesDGX agent

arXiv:2606.24153v1 Announce Type: new Abstract: While Deep Unfolding Networks (DUNs) dominate video Snapshot Compressive Imaging (SCI), they remain constrained by a uniform design philosophy. Existing

DiffusionBench: On Holistic Evaluation of Diffusion Transformers

Model ReleasesDGX agent

arXiv:2606.24888v1 Announce Type: new Abstract: Diffusion transformer (DiT) research on image generation has converged to a single evaluation setup: class-conditional generation on ImageNet. While met

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation

ResearchDGX agent

arXiv:2606.23950v1 Announce Type: new Abstract: Subject-driven image generation faces an 'Identity-Diversity Paradox', where strong identity preservation often leads to rigid and low-diversity outputs

DLTPose: 6DoF Pose Estimation From Accurate Dense Surface Point Estimates

Model ReleasesDGX agent

arXiv:2504.07335v3 Announce Type: replace Abstract: We propose DLTPose, a novel method for 6DoF object pose estimation from RGBD images that combines the accuracy of sparse keypoint methods with the r

Does it matter which Gaussians you pick in 4D Gaussian streaming?

Model ReleasesDGX agent

arXiv:2603.17227v3 Announce Type: replace Abstract: Anchor-driven 4D Gaussian streaming methods such as Instant Gaussian Stream (IGS) update a dynamic scene each frame from a compact set of Gaussian a

DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model

SafetyDGX agent

arXiv:2606.24051v1 Announce Type: new Abstract: Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follow

Dual-Branch Cross-Projection Debiasing through Diffusion-based Disentanglement

Model ReleasesDGX agent

arXiv:2606.24161v1 Announce Type: new Abstract: Foundation models trained on biased datasets often rely on spurious correlations between target labels and non-causal attributes, resulting in poor gene

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

Model ReleasesDGX agent

arXiv:2512.24731v2 Announce Type: replace Abstract: Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite r

EERLoss: A Novel Loss Function for Training Deep Biometric Models. A Case Study in Keystroke Dynamics

Model ReleasesDGX agent

arXiv:2606.24586v1 Announce Type: new Abstract: Deep learning approaches to biometric verification are commonly trained by optimizing indirect objectives, creating a misalignment between the optimizat

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

Model ReleasesDGX agent

arXiv:2606.24422v1 Announce Type: new Abstract: We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of mo

Emotion Diffusion Classifier with Adaptive Margin Discrepancy Training for Facial Expression Recognition

TutorialsDGX agent

arXiv:2603.29578v2 Announce Type: replace Abstract: Facial Expression Recognition (FER) is essential for human-machine interaction, as it enables machines to interpret human emotions and internal stat

EPEdit: Redefining Image Editing with Generative AI and User-Centric Design

ResearchDGX agent

arXiv:2606.24057v1 Announce Type: new Abstract: The demand for image manipulation has seen a significant increase recently. Traditional tools like Photoshop and Capture One, while powerful, require co

EPMF: Efficient Perception-aware Multi-sensor Fusion for 3D Semantic Segmentation

Model ReleasesDGX agent

arXiv:2106.15277v4 Announce Type: replace Abstract: We study multi-sensor fusion for 3D semantic segmentation that is important to scene understanding for many applications, such as autonomous driving

Fabric Image Demoireing Benchmark from Synthesis to Restoration

Model ReleasesDGX agent

arXiv:2606.24072v1 Announce Type: new Abstract: Fabric moire is a sampling-induced aliasing artifact caused by the interaction between fine textile patterns and camera sensor grids, producing structur

FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait Image

ResearchDGX agent

arXiv:2606.24232v1 Announce Type: new Abstract: We introduce FiCA, a Feed-forward, instant Gaussian Codec Avatar generation pipeline that creates lifelike avatars from a single portrait image. Generat

Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark

Model ReleasesDGX agent

arXiv:2503.14862v3 Announce Type: replace Abstract: Open-vocabulary detectors are proposed to locate and recognize objects in novel classes. However, variations in vision-aware language vocabulary dat

FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation

ResearchDGX agent

arXiv:2606.24876v1 Announce Type: new Abstract: Generating explorable 3D scenes from a single image requires strong generative priors and accurate geometric representations suitable for downstream use

Flood Mapping from RGB imagery using a Vision Foundation Model

ResearchDGX agent

arXiv:2606.24120v1 Announce Type: new Abstract: Timely, high-resolution maps of flood extent around settlements are essential for emergency response and damage assessment. We consider airborne RGB ima

FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation

Model ReleasesDGX agent

arXiv:2511.21029v3 Announce Type: replace Abstract: Music-to-dance generation aims to translate auditory signals into expressive human motion, with broad applications in virtual reality, choreography,

ForensicsTok: Forensics-Guided Tokenized Modeling for Image Tampering Localization

Local AiDGX agent

arXiv:2606.24538v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) offer powerful reasoning for forensic tasks, yet existing approaches utilizing exogenous segmentation decoders

From Open Waters to Enclosed Cabins: ProteusVPR for Cross-Scene Visual Place Recognition in Maritime Perception and Cabin Inspection

AgentsDGX agent

arXiv:2606.24234v1 Announce Type: new Abstract: Autonomous robotic inspection in maritime environments presents unique challenges for Visual Place Recognition (VPR) due to cross-scene perceptual shift

Full-resolution MLPs Empower Medical Dense Prediction

ResearchDGX agent

arXiv:2311.16707v2 Announce Type: replace-cross Abstract: Dense prediction is a fundamental requirement for many medical vision tasks such as medical image restoration, registration, and segmentation.

GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence

SafetyDGX agent

arXiv:2511.21945v3 Announce Type: replace Abstract: Generating complete 3D objects under partial occlusions (i.e., amodal scenarios) is a practically important yet challenging problem, as large portio

Generative Manifold Distillation: Aligning Restoration Trajectories with Natural Image Prior

SafetyDGX agent

arXiv:2512.11121v2 Announce Type: replace Abstract: Pre-trained image restoration models often fail on out-of-distribution (OOD) real-world degradations. Adapting to these domains is challenging as re

GeoIMO: Geometry-Driven Independent Motion Classification for Event Cameras

Local AiDGX agent

arXiv:2606.24499v1 Announce Type: new Abstract: Existing automotive event datasets rely on appearance-based annotations from frame pipelines, making them poorly suited for motion-aware event perceptio

Geometric Action Model for Robot Policy Learning

SafetyDGX agent

arXiv:2606.17046v2 Announce Type: replace-cross Abstract: Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physi

Geometry-Aware Style Transfer in 3D Gaussian Splatting

ResearchDGX agent

arXiv:2606.24144v1 Announce Type: new Abstract: In this paper, we present a novel geometry-aware style transfer framework for 3D Gaussian splatting (3DGS) that simultaneously transfers appearance attr

Geometry-Instructed Video Editing

ResearchDGX agent

arXiv:2606.24225v1 Announce Type: new Abstract: Object-level geometric edits, including translating, rotating, scaling, duplicating, or removing an object, are routine operations in digital content cr

GeoT2V-Bench: Benchmarking 3D Consistency in Text-to-Video Models via 3D Reconstruction

Model ReleasesDGX agent

arXiv:2606.24829v1 Announce Type: new Abstract: Camera-prompted text-to-video (T2V) models are increasingly used to synthesize virtual camera captures, such as orbiting objects or moving through stati

HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models

Model ReleasesDGX agent

arXiv:2606.23843v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are typically pre-trained on large-scale image-text datasets to capture semantic correspondences between visual content an

Heterogeneous Knowledge Distillation via Geometry Decoupling and Momentum-Aware Gradient Regulation

Model ReleasesDGX agent

arXiv:2606.24557v1 Announce Type: new Abstract: Heterogeneous Knowledge Distillation (HKD) aims to transfer knowledge across varying architectures (e.g., from Transformer to CNN) but inherently suffer

Hierarchical Spatial and Channel Aggregation for Cross-domain Few-shot Segmentation

SafetyDGX agent

arXiv:2606.24296v1 Announce Type: new Abstract: Cross-domain Few-shot Segmentation (CD-FSS) aims to learn generalizable segmentation capability from abundant annotated samples in the source domain, en

← Previous
1…7374757677…211
Next →