AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
8 Jun 2026

ActionMap: Robot Policy Learning via Voxel Action Heatmap

SafetyDGX agent

arXiv:2606.06904v1 Announce Type: cross Abstract: Vision-language-action (VLA) models have advanced rapidly across backbones, training recipes, and data scale, yet the action decoder, which converts t

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO

SafetyDGX agent

arXiv:2606.06828v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. Howeve

Adaptive Band Selection for Hyperspectral Classification with Spatially Disjoint Evaluation

ResearchDGX agent

arXiv:2606.06684v1 Announce Type: new Abstract: Hyperspectral band selection methods based on differentiable selectors can be sensitive to initialization and to extracting a final discrete subset, whi


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

AdaTok: Self-Budgeting Image Tokenization with Quality-Preserving Dynamic Tokens

SafetyDGX agent

arXiv:2606.07185v1 Announce Type: new Abstract: Image tokenizers, from 2D grids to recent 1D sequences, typically encode every image with the same fixed number of tokens. Yet visual complexity is high

Advanced Flood Prediction with Physics-Guided Deep Learning: Combining UNet, FNO, and SAR/Optical Imagery

ResearchDGX agent

arXiv:2606.06524v1 Announce Type: cross Abstract: Accurate and scalable flood mapping remains challenging due to limited ground observations, heterogeneous terrain conditions, and the difficulty of en

AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired

ResearchDGX agent

arXiv:2511.06080v4 Announce Type: replace Abstract: This paper presents AIDEN, an artificial intelligence-based assistant designed to enhance the autonomy and daily quality of life of visually impaire

An Adaptive Data cleaning Framework for Noisy Label Detection

ApplicationsDGX agent

arXiv:2606.07086v1 Announce Type: new Abstract: Deep neural networks (DNNs) excel in computer vision tasks given large annotated datasets. In real-world applications, however, labels are often corrupt

An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?

Model ReleasesDGX agent

arXiv:2605.25806v2 Announce Type: replace Abstract: Women's safety and security are paramount for a modern society. Crimes against women occur in daylight as well as in low-light conditions. Often, su

An Integrated Roadside Sensing and Communication Framework for Vulnerable Road User Safety at Signalized Intersections

Model ReleasesDGX agent

arXiv:2606.07016v1 Announce Type: cross Abstract: Vulnerable road users (VRUs) account for approximately half of urban traffic deaths globally, with intersections concentrating a disproportionate shar

Anchored, Not Graded: Vision-Language Models Fail at Slant-from-Texture Perception

ResearchDGX agent

arXiv:2606.06714v1 Announce Type: new Abstract: Human perception of surface slant from texture exhibits systematic, graded biases that emerge reliably in psychophysical experiments. Prior work showed

AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization

AgentsDGX agent

arXiv:2606.07326v1 Announce Type: new Abstract: Despite being a pivotal frontier, interactive world modeling remains underexplored in terms of the versatile controllability required by practical scena

Applying Deep Learning for cockpit segmentation in the context of mixed reality

ResearchDGX agent

arXiv:2606.06520v1 Announce Type: new Abstract: Computer vision is an area that has been growing continuously. With the advance of technologies with a first-person view, new development opportunities

ARAPDiffusion: ARAP Regularization for Diffusion-Based Deformable Shape Space Learning

TutorialsDGX agent

arXiv:2606.06887v1 Announce Type: new Abstract: This paper introduces ARAPDiffusion, a latent diffusion model to learn the underlying continuous shape space of a deformation shape collection. The key

Architecture-Adaptive Uncertainty Fusion for Deepfake Detection

ResearchDGX agent

arXiv:2606.06666v1 Announce Type: new Abstract: Deepfake detection systems achieve near-perfect accuracy on benchmarks, yet forensic deployment demands reliable prediction uncertainty. Existing uncert

AsyncPatch Diffusion: spatially-flexible image generation

TutorialsDGX agent

arXiv:2606.07079v1 Announce Type: new Abstract: Standard diffusion models corrupt an entire sample with a single shared noise level, forcing all spatial regions to follow the same denoising trajectory

Audio-Visual World Models: Grounding Multisensory Imagination for Embodied Agents

Model ReleasesDGX agent

arXiv:2512.00883v3 Announce Type: replace-cross Abstract: World models simulate environmental dynamics to enable agents to plan and reason about future states. While existing approaches have primarily

Beyond Backscatter: InSAR coherence from detected SAR images

TutorialsDGX agent

arXiv:2606.07374v1 Announce Type: cross Abstract: In this work, we propose a deep learning framework for coherence regression directly from detected SAR images, without the need for accurate coregistr

Beyond Universality: The GCC-FER Dataset and Culture-Aware Adaptation for Dynamic Facial Expression Recognition

Model ReleasesDGX agent

arXiv:2606.07063v1 Announce Type: cross Abstract: Dynamic Facial Expression Recognition (DFER) is a key enabling technology in affective computing, human-computer interaction, and intelligent multimed

Broadband Hyperspectral 3D Imaging using Dispersed Structured Light

HardwareDGX agent

arXiv:2605.25757v2 Announce Type: replace Abstract: Hyperspectral 3D imaging enables the capture of dense spectral information and scene geometry but has traditionally been confined to narrow spectral

Certified Robustness to Data Poisoning in Gradient-Based Training

Model ReleasesDGX agent

arXiv:2406.05670v3 Announce Type: replace-cross Abstract: Modern machine learning pipelines leverage large amounts of public data, making it infeasible to guarantee data quality and leaving models ope

CFRNet: Cycle-Consistent Fixed-Point Training for Real-Time Blind Face Restoration on Consumer Embedded NPUs

Model ReleasesDGX agent

arXiv:2606.06850v1 Announce Type: new Abstract: Blind face restoration on consumer devices has to balance image quality against speed and memory. Strong methods such as GFPGAN and CodeFormer give good

CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection

ApplicationsDGX agent

arXiv:2606.06978v1 Announce Type: new Abstract: Continual Object Detection (COD) requires a detector to acquire new categories over time while preserving previously learned ones. This goal is closely

Closed-Form Spectral Regularization for Multi-Task Model Merging

Model ReleasesDGX agent

arXiv:2606.07289v1 Announce Type: cross Abstract: Model merging combines several independently fine-tuned experts into a single multi-task model without any training data, reducing the storage, servin

COMPOSE: Hypergraph Cover Optimization for Multi-view 3D Human Pose Estimation

ResearchDGX agent

arXiv:2601.09698v2 Announce Type: replace Abstract: 3D human pose estimation from sparse multi-view camera rigs is an essential task for numerous applications, including action recognition, sports ana

Compute-Optimal Network Design for Echocardiography Myocardial Segmentation and Perfusion Quantification using Neural Scaling Laws

Model ReleasesDGX agent

arXiv:2606.06725v1 Announce Type: cross Abstract: Myocardial perfusion quantification using contrast-enhanced ultrasound offers a bedside non-ionizing alternative to nuclear imaging modalities. Howeve

Consistency-Preserving Diverse Video Generation

ResearchDGX agent

arXiv:2602.15287v2 Announce Type: replace Abstract: Text-to-video generation is expensive, so only a few samples are typically produced per prompt. In this low-sample regime, maximizing the value of e

Consistent-Inversion: Reverse Consistency Guidance for Structure-Preserving Visual Editing

SafetyDGX agent

arXiv:2606.07145v1 Announce Type: new Abstract: Text-guided diffusion models have become effective tools for real-image visual editing, where the edited image must follow a target instruction while pr

Constructing VAE Latent Spaces with Prescribed Topology

TutorialsDGX agent

arXiv:2606.07058v1 Announce Type: cross Abstract: Variational autoencoders (VAEs) learn low-dimensional latent representations of high-dimensional data. When the data lies on a manifold with non-Eucli

Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos

Model ReleasesDGX agent

arXiv:2606.07366v1 Announce Type: new Abstract: Self-driving simulations typically rely on data collected in a small number of cities or on hand-authored synthetic scenarios. Dashcam videos cover a fa

Deep Learning Pose Estimation for Multi-Label Recognition of Combined Hyperkinetic Movement Disorders

ResearchDGX agent

arXiv:2602.00163v2 Announce Type: replace Abstract: Hyperkinetic movement disorders (HMDs) such as dystonia, tremor, chorea, myoclonus, and tics are disabling motor manifestations across childhood and

Detecting Temporally Localized Manipulations in Authentic Video Streams

Model ReleasesDGX agent

arXiv:2606.07090v1 Announce Type: new Abstract: The rapid advancement of video editing and generative artificial intelligence technologies has made realistic video manipulation increasingly accessible

Diagnosing Visual Ignorance in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.06890v1 Announce Type: new Abstract: Vision-Language Models (VLMs) frequently rely on language priors, producing confident answers that are weakly grounded in visual evidence. While this be

Differences in Detection: Explainability Where it Matters

TutorialsDGX agent

arXiv:2606.07503v1 Announce Type: new Abstract: We propose Differences in Detection (DnD), an intuitive method to compare two object detection models. Based on the same matching algorithm, it compleme

DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose Estimation

Model ReleasesDGX agent

arXiv:2606.07419v1 Announce Type: new Abstract: Recovering 3D human poses for multiple individuals from different camera views is a fundamental bottleneck for analyzing interacting behaviors. Existing

Does Appearance Help? A Systematic Study of Image-Based Re-Identification in Online 3D Multi-Pedestrian Tracking

ResearchDGX agent

arXiv:2606.07233v1 Announce Type: new Abstract: LiDAR-based 3D Multi-Object Tracking (MOT) typically relies solely on geometric information, which is often insufficient to distinguish between targets

DRIFT: From Robustness Gaps to Invariance Manifolds for AI-Generated Image Detection

ResearchDGX agent

arXiv:2606.06918v1 Announce Type: new Abstract: The rapid evolution of generative image models challenges existing AI-generated image detectors, particularly in open-world settings with unseen generat

DSU-Net: An Attention-Enhanced Dense Skip U-Net for Breast Lesion Segmentation in Mammographic Images

ResearchDGX agent

arXiv:2606.06537v1 Announce Type: cross Abstract: Breast cancer remains one of the leading causes of cancer-related mortality among women worldwide, making early detection essential for effective trea

ErA: Error-Aware Deep Unrolling Network for Single Image Defocus Deblurring

ResearchDGX agent

arXiv:2606.06540v1 Announce Type: cross Abstract: We introduce ErA (Error-Aware Deep Unrolling Network), an end-to-end frame work for single-image defocus deblurring. ErA jointly learns a compact kern

EvoGS: Constructing Continuous-Layered Gaussian Splatting with Evolution Tree for Scalable 3D Streaming

Local AiDGX agent

arXiv:2606.07179v1 Announce Type: new Abstract: Streaming 3D Gaussian Splatting requires highly scalable, progressive representations. Existing progressive methods rely on extit{discrete layering}, ac

ExMesh: EXplicit Mesh Reconstruction with Topology Adaptation

TutorialsDGX agent

arXiv:2606.07288v1 Announce Type: new Abstract: Reconstructing surface meshes from multi-view images has remained a core challenge in recent years. Most existing methods, whether implicit or explicit,

ForensicConcept: Transferable Forensic Concepts for AIGI Detection

SafetyDGX agent

arXiv:2606.07034v1 Announce Type: new Abstract: AI-generated image detectors achieve high accuracy on in-distribution data but often fail on unseen generators. A key obstacle to understanding this fai

From Pixels to Newtons: Predicting In Vivo Joint Contact Forces from Monocular Video

ResearchDGX agent

arXiv:2606.06631v1 Announce Type: new Abstract: Joint contact forces govern implant longevity, cartilage health, and rehabilitation outcomes, shaping who develops osteoarthritis, who recovers well fro

From Vision to Text: A Compact Multimodal Approach for Robust, Cross-Domain Presentation Attack Detection on ID Cards

Model ReleasesDGX agent

arXiv:2606.06966v1 Announce Type: new Abstract: Cross-domain shifts challenge Presentation Attack Detection (PAD) on ID Cards, given the restricted data available due to privacy concerns. This work pr

FS-DVS: A Frequency-Selective Dynamic Visual Sensing Paradigm for Enhancing Information Completeness

ResearchDGX agent

arXiv:2606.06856v1 Announce Type: new Abstract: Dynamic vision sensors (DVS) offer exceptional temporal resolution and dynamic range by asynchronously reporting pixel-level intensity changes. However,

Generalization of Diffusion Models Arises with a Balanced Representation Space

Local AiDGX agent

arXiv:2512.20963v3 Announce Type: replace-cross Abstract: Diffusion models excel at generating high-quality, diverse samples, yet they risk memorizing training data when overfit to the training object

Geometric-Aware Hypergraph Reasoning for Novel Class Discovery in Point Cloud Segmentation

ResearchDGX agent

arXiv:2606.07280v1 Announce Type: new Abstract: Novel class discovery in point cloud segmentation aims to transfer knowledge from known classes to automatically identify and segment unlabeled novel cl

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning

AgentsDGX agent

arXiv:2606.06532v1 Announce Type: new Abstract: Despite significant progress in agentic long video understanding, existing methods still lack detailed motion comprehension coupled with an efficient me

GuideCAD: A Lightweight Multimodal Framework for 3D CAD Model Generation via Prefix Embedding

Model ReleasesDGX agent

arXiv:2606.07024v1 Announce Type: new Abstract: Multi-modal approaches used for 3D CAD generation require substantial computational resources, necessitating efficient training. To address this, we pro

Image class translation: visual inspection of class-specific hypotheticals and classification based on translation distance

ResearchDGX agent

arXiv:2408.08973v3 Announce Type: replace Abstract: Purpose: A major barrier to the implementation of artificial intelligence for medical applications is automated CNNs' lack of explainability and hig

Implicit Data Synthesis for Contrastive Unsupervised Data Augmentation

ResearchDGX agent

arXiv:2606.07498v1 Announce Type: new Abstract: Scientific observations generate large quantities of unlabeled data which is laborious to hand-label, making unsupervised learning techniques valuable f

JA-SIREN: Deterministic Initialization for Sinusoidal Networks via Spectral Matching

ResearchDGX agent

arXiv:2606.06671v1 Announce Type: new Abstract: Existing implicit neural representation (INR) approaches suffer from stochastic initialization that does not guarantee consistent or high-quality perfor

LARA: Latent Action Representation Alignment for Vision-Language-Action Models

SafetyDGX agent

arXiv:2606.07100v1 Announce Type: new Abstract: Visual-language action (VLA) models enable robots to predict actions directly from observations and language instructions, but their performance depends

Lighting-Aware Representation Learning under Controllable Lighting Variation

TutorialsDGX agent

arXiv:2606.06899v1 Announce Type: new Abstract: Variations in illumination remain a major challenge for visual representation learning, as they induce substantial appearance changes both across and wi

LRMIL: Efficient Low-Resolution Multiple Instance Learning via High-Resolution Knowledge Distillation for Whole Slide Image Classification

ApplicationsDGX agent

arXiv:2606.06864v1 Announce Type: new Abstract: Multiple instance learning (MIL) has become a standard paradigm for whole slide image (WSI) analysis in digital pathology, as it enables slide-level pre

LUCID: Learning Unified Control for Image Deflaring and Exposure Mastery in Nighttime Photography

ApplicationsDGX agent

arXiv:2606.06901v1 Announce Type: new Abstract: Photography is the art of painting with light, yet nighttime scenes are shaped by competing degradations: intense flares obscure scene structure, while

MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language Models

ResearchDGX agent

arXiv:2606.06760v1 Announce Type: new Abstract: Medical large vision-language models (Med-LVLMs) have recently achieved remarkable progress in vision-language comprehension and medical image segmentat

Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation

ResearchDGX agent

arXiv:2606.07394v1 Announce Type: new Abstract: In Video Instance Segmentation (VIS), classification, segmentation, and tracking objectives are jointly evaluated, but their individual contributions to

Multi-FRuGaL: Multimodal Flexible Redundancy-aware Decomposed Gated Learning for Cancer Diagnosis and Prognosis

ApplicationsDGX agent

arXiv:2606.06867v1 Announce Type: new Abstract: Modern medicine relies on heterogeneous data sources spanning radiology, pathology, text reports, and structured clinical information. However, real-wor

MVSegNet: A Lightweight Boundary-Aware Network for Fetal Lateral Ventricle Segmentation and Atrial Width Estimation in Prenatal Ultrasound

HardwareDGX agent

arXiv:2606.06958v1 Announce Type: new Abstract: Fetal ventriculomegaly is assessed by measuring the atrial width of the lateral ventricle in prenatal ultrasound. Accurate segmentation is essential for

OpenGlass: Open-Source Smart Glasses for On-Device Event-Based Gesture Recognition

Model ReleasesDGX agent

arXiv:2606.07431v1 Announce Type: new Abstract: Smart eyewear enables unobtrusive, context-aware interaction through multimodal sensors and on-device intelligence, but is severely limited by power, me

← Previous
1…9192939495…211
Next →