AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
1 Jul 2026

Distortion-Corrected Diffusion MRI Using Rotated-View EPI and Joint Field-Map/Image Estimation with Gaussian Primitives

ResearchDGX agent

arXiv:2606.31521v1 Announce Type: cross Abstract: Echo Planar Imaging (EPI) is the standard acquisition technique for diffusion and functional neuroimaging, enabling rapid imaging but suffering from g

Do Not Break the Vessels: Structure-Preserving Mean Flow for Vascular Image Translation

ResearchDGX agent

arXiv:2606.31095v1 Announce Type: new Abstract: Reconstructing anatomically faithful vascular structures from clinically accessible imaging modalities is of substantial clinical significance. However,

Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.31373v1 Announce Type: new Abstract: Cycle self-training (CST) breaks the shared classifier assumption of the standard self-training framework, which is effective for unsupervised domain ad

DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation

AgentsDGX agent

arXiv:2606.31918v1 Announce Type: new Abstract: A pivotal step in autonomous driving simulation involves inserting foreground vehicles with predefined trajectories into simulated scenes. This process

DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation

AgentsDGX agent

arXiv:2606.31488v1 Announce Type: new Abstract: Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry witho

Drop-In Perceptual Optimization for 3D Gaussian Splatting

ResearchDGX agent

arXiv:2603.23297v2 Announce Type: replace Abstract: Despite their output being ultimately consumed by human viewers, 3D Gaussian Splatting (3DGS) methods often rely on ad-hoc combinations of pixel-lev

Dual Sparse Aggregation Transformer for Multispectral Object Detection

Model ReleasesDGX agent

arXiv:2606.31015v1 Announce Type: new Abstract: Transformer-based approaches have obtained excellent performance in multispectral object detection tasks due to their ability to model long-range depend

DynFly: Dynamic-Aware Continuous Trajectory Generation for UAV Vision-Language Navigation in Urban Environments

Model ReleasesDGX agent

arXiv:2606.31654v1 Announce Type: cross Abstract: Recent advances in multimodal large models have significantly improved UAV vision-language navigation (UAV-VLN) by enhancing high-level perception and

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes

Model ReleasesDGX agent

arXiv:2604.04834v2 Announce Type: replace Abstract: Robotic Vision-Language-Action (VLA) models generalize well for open-ended manipulation, but their perception is fragile under sensing-stage degrada

Editing Everything Everywhere All at Once

Model ReleasesDGX agent

arXiv:2606.31278v1 Announce Type: new Abstract: Editing multiple elements of an image in a single forward pass is a practical alternative to multi-turn image manipulation, offering improved efficiency

EgoCogNav: Cognition-aware Human Egocentric Navigation

ApplicationsDGX agent

arXiv:2511.17581v3 Announce Type: replace-cross Abstract: Modeling the cognitive and experiential factors of human navigation is central to deepening our understanding of human-environment interaction

EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning

SafetyDGX agent

arXiv:2511.18242v3 Announce Type: replace Abstract: Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal la

EpiMask: Leveraging Epipolar Distance Based Masks in Cross-Attention for Satellite Image Matching

ResearchDGX agent

arXiv:2603.21463v2 Announce Type: replace Abstract: The deep-learning based image matching networks can now handle significantly larger variations in viewpoints and illuminations while providing match

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs

SafetyDGX agent

arXiv:2606.31982v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) incur prohibitive inference costs due to long visual token sequences. Training-free visual token reduction prov

Estimating Velocity of Spheres from Rolling-Shutter Image(s)

ResearchDGX agent

arXiv:2606.31760v1 Announce Type: new Abstract: Rolling-shutter cameras introduce characteristic distortions when imaging fast moving objects, and these effects are typically treated as artifacts to b

Event-Driven Video Generation

Local AiDGX agent

arXiv:2603.13402v3 Announce Type: replace Abstract: Current text-to-video models can make individual frames look convincing while still getting simple interactions wrong: objects move before contact,

Evidence Triangulation for Multimodal Fact-Checking in the Wild

Model ReleasesDGX agent

arXiv:2606.31367v1 Announce Type: cross Abstract: The proliferation of multimedia content on social platforms has fueled multimodal misinformation, where images are used to reinforce false claims. Con

ExPLoRe: Expert Patch-Level Loss Routing for Multi-Objective Masked Image Modeling

SafetyDGX agent

arXiv:2606.31201v1 Announce Type: new Abstract: Multi-objective masked image modeling (MIM) combines complementary learning signals (token distillation, CLS alignment, and pixel reconstruction) but ex

FaceMoE: Mixture of Experts for Low-Resolution Face Recognition

ResearchDGX agent

arXiv:2606.32040v1 Announce Type: new Abstract: Low-resolution face recognition (LR-FR) remains a challenging task due to poor feature extraction and aggregation, as probe images often contain limited

FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning

Model ReleasesDGX agent

arXiv:2511.17979v2 Announce Type: replace Abstract: Diffusion models have achieved remarkable success in generative modeling, yet how to effectively adapt large pretrained models to new tasks remains

Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples

ApplicationsDGX agent

arXiv:2509.25682v2 Announce Type: replace Abstract: AI-generated image (AIGI) attribution presents a pressing challenge that goes beyond mere AIGI detection, aiming to identify the source model or tec

Few to Big: Prototype Expansion Network via Diffusion Learner for Point Cloud Few-shot Semantic Segmentation

ResearchDGX agent

arXiv:2509.12878v2 Announce Type: replace Abstract: Few-shot 3D point cloud semantic segmentation aims to segment novel categories using a minimal number of annotated support samples. However, prototy

Filterless Snapshot Hyperspectral Imaging using Guided Patch Diffusion

Local AiDGX agent

arXiv:2412.02798v3 Announce Type: replace Abstract: We consider the problem of reconstructing a HxWx31 hyperspectral image from a Himes W grayscale snapshot measurement that is captured using only a s

Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction

SafetyDGX agent

arXiv:2603.09930v2 Announce Type: replace Abstract: Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences

Fleet: Few Shots Lead Effective AI-generated Image Detection

Model ReleasesDGX agent

arXiv:2606.31082v1 Announce Type: new Abstract: AI-generated image (AIGI) detection is undergoing a critical transition from laboratory benchmarks to open-world adversarial defense. The prevalent para

FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers

ResearchDGX agent

arXiv:2606.31938v1 Announce Type: cross Abstract: Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heteroge

ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving

AgentsDGX agent

arXiv:2606.31226v1 Announce Type: new Abstract: World-model-based autonomous driving endows the model with the ability to understand scene evolution. Yet this promise is undermined by the prevailing i

FROST: Training-Free Few-Shot Segmentation with Frozen Features and Nonparametric Statistics

ResearchDGX agent

arXiv:2606.31136v1 Announce Type: new Abstract: Few-shot segmentation asks a model to delineate a target class in a query image from only a handful of annotated examples, a setting most acute in remot

Fully Automated High-Precision Segmentation of Retinal Atrophy and Ellipsoid Zone Thickness in OCT: A Reliable Tool for Real-World GA Monitoring

ApplicationsDGX agent

arXiv:2606.31502v1 Announce Type: new Abstract: Geographic atrophy (GA) secondary to age-related macular degeneration (AMD) requires precise monitoring of relevant structural biomarkers to assess dise

G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation

SafetyDGX agent

arXiv:2601.03510v3 Announce Type: replace Abstract: Point cloud segmentation is critical for 3D scene understanding. However, sparse and irregular point distributions provide limited appearance eviden

Gaussian Belief Propagation Network for Depth Completion

ResearchDGX agent

arXiv:2601.21291v2 Announce Type: replace Abstract: Depth completion aims to predict a dense depth map from a color image with sparse depth measurements. Although deep learning methods have achieved s

GaussianMap: Learning Gaussian Representation for Multi-Sensor Online HD Map Construction

Local AiDGX agent

arXiv:2606.31177v1 Announce Type: new Abstract: Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online construction o

GaussLite: Online Task-Conditioned 3D Gaussian Splatting for Real-Time Robotic Mapping

ResearchDGX agent

arXiv:2606.30809v1 Announce Type: new Abstract: Existing 3D Gaussian Splatting (3DGS) systems distribute representation capacity uniformly across a scene, ignoring the fact that many downstream roboti

GEAR: Guided End-to-End AutoRegression for Image Synthesis

SafetyDGX agent

arXiv:2606.32039v1 Announce Type: new Abstract: Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and then frozen, after which a generator i

Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior

Model ReleasesDGX agent

arXiv:2606.31814v1 Announce Type: new Abstract: Lane topology reasoning aims to construct a lane graph from onboard sensor observations. Existing methods follow a detection and association paradigm th

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

ResearchDGX agent

arXiv:2603.14965v2 Announce Type: replace Abstract: Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While

GRAPE: Graph-Augmented Prototype Explanations for Interactive Medical Image Diagnosis

SafetyDGX agent

arXiv:2606.30901v1 Announce Type: new Abstract: Prototype-based medical image classifiers present three clinical limitations: they treat findings as independent, silently amplify unsafe physician feed

Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation

ResearchDGX agent

arXiv:2606.31071v1 Announce Type: new Abstract: Semantic navigation is a fundamental task for embodied agents operating in unseen environments, requiring both semantic understanding and long-term deci

Horizon3D: Sparse Radar-Camera Fusion for Long-Range 3D Perception in Autonomous Driving

AgentsDGX agent

arXiv:2606.31096v1 Announce Type: new Abstract: Long-range 3D object detection is critical for safe autonomous driving at highway speeds, yet existing radar-camera fusion methods remain limited at ext

HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection

Model ReleasesDGX agent

arXiv:2606.31172v1 Announce Type: new Abstract: Monocular 3D lane detection plays a critical role in autonomous driving, yet recovering reliable 3D geometry from a single image remains challenging due

HVPNet: A Bio-Inspired Network for General Salient and Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2606.31496v1 Announce Type: new Abstract: In recent years, most research on multimodal salient object detection (SOD) and camouflaged object detection (COD) typically aims to improve performance

Hybrid Unet-Transformer Model for Generating Stress and Strain Fields from Composite Geometrics

ResearchDGX agent

arXiv:2606.31068v1 Announce Type: new Abstract: Accurate prediction of stress and strain fields in hierarchical composite microstructures is critical for physics-informed material design, yet conventi

HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space

ResearchDGX agent

arXiv:2606.31245v1 Announce Type: new Abstract: Surgical vision-language foundation models typically adopt educational materials, such as surgical lecture videos, to transfer surgical knowledge encode

Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility

TutorialsDGX agent

arXiv:2505.18521v2 Announce Type: replace Abstract: The substantial training cost of diffusion models hinders their deployment. Immiscible Diffusion recently showed that reducing diffusion trajectory

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

SafetyDGX agent

arXiv:2606.31109v1 Announce Type: new Abstract: Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving communit

InstanceControl: Controllable Complex Image Generation without Instance Labeling

TutorialsDGX agent

arXiv:2606.31924v1 Announce Type: new Abstract: Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to g

Intrinsically Stable Spiking Neural Networks: Overcoming the Performance Barrier in the Absence of Batch Normalization

ResearchDGX agent

arXiv:2606.31695v1 Announce Type: new Abstract: The performance of deep spiking neural networks (SNNs) often relies on batch normalization (BN). However, the advanced dynamic BN variants used in state

JacobianAvatar: Temporally Consistent Semi-rigid Avatar Reconstruction from a Monocular Video

Model ReleasesDGX agent

arXiv:2606.31115v1 Announce Type: new Abstract: Generating realistic human avatars in complex motions--such as clothing dynamics--requires modeling of global and local deformations which remains chall

Joint Optimization for 4D Human-Scene Reconstruction in the Wild

ResearchDGX agent

arXiv:2501.02158v3 Announce Type: replace Abstract: Reconstructing human motion and its surrounding environment is crucial for understanding human-scene interaction and predicting human movements in t

Knowledge-Driven Dimension Estimation from a Single Image -3D Asset Generation Technology for Digital Twin Construction

AgentsDGX agent

arXiv:2606.30896v1 Announce Type: new Abstract: In the verification of in-vehicle cameras, simulation technology using virtual spaces has advanced, enabling pre-evaluation of false detections and miss

LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior

SafetyDGX agent

arXiv:2603.25399v2 Announce Type: replace Abstract: We introduce extbf{LaMP}, a dual-expert Vision-Language-Action framework that embeds dense 3D scene flow as a latent motion prior for robotic manipu

Language-Assisted Super-Resolution from Real-World Low-Resolution Patches

SafetyDGX agent

arXiv:2606.31363v1 Announce Type: new Abstract: Single image super-resolution aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs. Training SR models typically requires pai

Learning a Sampling-Free Variational DNN Plugin from Tiny Training Sets to Refine OOD Segmentation With Uncertainty Estimation

ResearchDGX agent

arXiv:2606.15837v2 Announce Type: replace Abstract: Deep neural networks (DNNs) frequently fail to generalize to out-of-distribution (OOD) medical images because of variations in scanners and acquisit

Learning to Deny: Action Denial in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2606.31187v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero-shot and few-shot recognition across standard

LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR

Model ReleasesDGX agent

arXiv:2601.14251v2 Announce Type: replace Abstract: We present LightOnOCR-2-1B, a 1B-parameter end-to-end multilingual vision--language model that converts document images (e.g., PDFs) into clean, nat

LiteMatch: Lightweight Zero-Shot Stereo Matching via Cost Volume Stabilization

ResearchDGX agent

arXiv:2606.31636v1 Announce Type: new Abstract: Despite rapid progress in learning-based stereo matching, high accuracy is often achieved at the cost of heavy backbones and computationally intensive 3

Localized Conformal Prediction for Image Classification with Vision-Language Models

ResearchDGX agent

arXiv:2606.31577v1 Announce Type: new Abstract: Conformal predictions have attracted significant attention in the field of uncertainty quantification, mainly because of their strong marginal coverage

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

Model ReleasesDGX agent

arXiv:2601.08758v4 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) reasoning has proven effective in enhancing large language models by encouraging step-by-step intermediate reasoning, a

MAPE: Defending Against Transferable Adversarial Attacks Using Multi-Source Adversarial Perturbations Elimination

ResearchDGX agent

arXiv:2606.31378v1 Announce Type: new Abstract: Neural networks are vulnerable to meticulously crafted adversarial examples, leading to high-confidence misclassifications in image classification tasks

Medical Image Spatial Grounding with Semantic Sampling

Model ReleasesDGX agent

arXiv:2603.14579v3 Announce Type: replace Abstract: Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs rep

← Previous
1…5758596061…209
Next →