AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

ExPLoRe: Expert Patch-Level Loss Routing for Multi-Objective Masked Image Modeling

DGX agent

arXiv:2606.31201v1 Announce Type: new Abstract: Multi-objective masked image modeling (MIM) combines complementary learning signals (token distillation, CLS alignment, and pixel reconstruction) but ex

safetyarxiv-cs-cv
1 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

FaceMoE: Mixture of Experts for Low-Resolution Face Recognition

DGX agent

arXiv:2606.32040v1 Announce Type: new Abstract: Low-resolution face recognition (LR-FR) remains a challenging task due to poor feature extraction and aggregation, as probe images often contain limited

researcharxiv-cs-cv
1 Jul 2026
Model Releases

FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning

DGX agent

arXiv:2511.17979v2 Announce Type: replace Abstract: Diffusion models have achieved remarkable success in generative modeling, yet how to effectively adapt large pretrained models to new tasks remains

model-releasesarxiv-cs-cv
1 Jul 2026
Applications

Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples

DGX agent

arXiv:2509.25682v2 Announce Type: replace Abstract: AI-generated image (AIGI) attribution presents a pressing challenge that goes beyond mere AIGI detection, aiming to identify the source model or tec

applicationsarxiv-cs-cv
1 Jul 2026
Research

Few to Big: Prototype Expansion Network via Diffusion Learner for Point Cloud Few-shot Semantic Segmentation

DGX agent

arXiv:2509.12878v2 Announce Type: replace Abstract: Few-shot 3D point cloud semantic segmentation aims to segment novel categories using a minimal number of annotated support samples. However, prototy

researcharxiv-cs-cv
1 Jul 2026
Local Ai

Filterless Snapshot Hyperspectral Imaging using Guided Patch Diffusion

DGX agent

arXiv:2412.02798v3 Announce Type: replace Abstract: We consider the problem of reconstructing a HxWx31 hyperspectral image from a Himes W grayscale snapshot measurement that is captured using only a s

local-aiarxiv-cs-cv
1 Jul 2026
Safety

Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction

DGX agent

arXiv:2603.09930v2 Announce Type: replace Abstract: Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences

safetyarxiv-cs-cv
1 Jul 2026
Model Releases

Fleet: Few Shots Lead Effective AI-generated Image Detection

DGX agent

arXiv:2606.31082v1 Announce Type: new Abstract: AI-generated image (AIGI) detection is undergoing a critical transition from laboratory benchmarks to open-world adversarial defense. The prevalent para

model-releasesarxiv-cs-cv
1 Jul 2026
Research

FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers

DGX agent

arXiv:2606.31938v1 Announce Type: cross Abstract: Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heteroge

researcharxiv-cs-cv
1 Jul 2026
Agents

ForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving

DGX agent

arXiv:2606.31226v1 Announce Type: new Abstract: World-model-based autonomous driving endows the model with the ability to understand scene evolution. Yet this promise is undermined by the prevailing i

agentsarxiv-cs-cv
1 Jul 2026
Research

FROST: Training-Free Few-Shot Segmentation with Frozen Features and Nonparametric Statistics

DGX agent

arXiv:2606.31136v1 Announce Type: new Abstract: Few-shot segmentation asks a model to delineate a target class in a query image from only a handful of annotated examples, a setting most acute in remot

researcharxiv-cs-cv
1 Jul 2026
Applications

Fully Automated High-Precision Segmentation of Retinal Atrophy and Ellipsoid Zone Thickness in OCT: A Reliable Tool for Real-World GA Monitoring

DGX agent

arXiv:2606.31502v1 Announce Type: new Abstract: Geographic atrophy (GA) secondary to age-related macular degeneration (AMD) requires precise monitoring of relevant structural biomarkers to assess dise

applicationsarxiv-cs-cv
1 Jul 2026
Safety

G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation

DGX agent

arXiv:2601.03510v3 Announce Type: replace Abstract: Point cloud segmentation is critical for 3D scene understanding. However, sparse and irregular point distributions provide limited appearance eviden

safetyarxiv-cs-cv
1 Jul 2026
Research

Gaussian Belief Propagation Network for Depth Completion

DGX agent

arXiv:2601.21291v2 Announce Type: replace Abstract: Depth completion aims to predict a dense depth map from a color image with sparse depth measurements. Although deep learning methods have achieved s

researcharxiv-cs-cv
1 Jul 2026
Local Ai

GaussianMap: Learning Gaussian Representation for Multi-Sensor Online HD Map Construction

DGX agent

arXiv:2606.31177v1 Announce Type: new Abstract: Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online construction o

local-aiarxiv-cs-cv
1 Jul 2026
Research

GaussLite: Online Task-Conditioned 3D Gaussian Splatting for Real-Time Robotic Mapping

DGX agent

arXiv:2606.30809v1 Announce Type: new Abstract: Existing 3D Gaussian Splatting (3DGS) systems distribute representation capacity uniformly across a scene, ignoring the fact that many downstream roboti

researcharxiv-cs-cv
1 Jul 2026
Safety

GEAR: Guided End-to-End AutoRegression for Image Synthesis

DGX agent

arXiv:2606.32039v1 Announce Type: new Abstract: Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and then frozen, after which a generator i

safetyarxiv-cs-cv
1 Jul 2026
Model Releases

Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior

DGX agent

arXiv:2606.31814v1 Announce Type: new Abstract: Lane topology reasoning aims to construct a lane graph from onboard sensor observations. Existing methods follow a detection and association paradigm th

model-releasesarxiv-cs-cv
1 Jul 2026
Research

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

DGX agent

arXiv:2603.14965v2 Announce Type: replace Abstract: Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While

researcharxiv-cs-cv
1 Jul 2026
Safety

GRAPE: Graph-Augmented Prototype Explanations for Interactive Medical Image Diagnosis

DGX agent

arXiv:2606.30901v1 Announce Type: new Abstract: Prototype-based medical image classifiers present three clinical limitations: they treat findings as independent, silently amplify unsafe physician feed

safetyarxiv-cs-cv
1 Jul 2026
Research

Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation

DGX agent

arXiv:2606.31071v1 Announce Type: new Abstract: Semantic navigation is a fundamental task for embodied agents operating in unseen environments, requiring both semantic understanding and long-term deci

researcharxiv-cs-cv
1 Jul 2026
Agents

Horizon3D: Sparse Radar-Camera Fusion for Long-Range 3D Perception in Autonomous Driving

DGX agent

arXiv:2606.31096v1 Announce Type: new Abstract: Long-range 3D object detection is critical for safe autonomous driving at highway speeds, yet existing radar-camera fusion methods remain limited at ext

agentsarxiv-cs-cv
1 Jul 2026
Model Releases

HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection

DGX agent

arXiv:2606.31172v1 Announce Type: new Abstract: Monocular 3D lane detection plays a critical role in autonomous driving, yet recovering reliable 3D geometry from a single image remains challenging due

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

HVPNet: A Bio-Inspired Network for General Salient and Camouflaged Object Detection

DGX agent

arXiv:2606.31496v1 Announce Type: new Abstract: In recent years, most research on multimodal salient object detection (SOD) and camouflaged object detection (COD) typically aims to improve performance

model-releasesarxiv-cs-cv
1 Jul 2026
Research

Hybrid Unet-Transformer Model for Generating Stress and Strain Fields from Composite Geometrics

DGX agent

arXiv:2606.31068v1 Announce Type: new Abstract: Accurate prediction of stress and strain fields in hierarchical composite microstructures is critical for physics-informed material design, yet conventi

researcharxiv-cs-cv
1 Jul 2026
Research

HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space

DGX agent

arXiv:2606.31245v1 Announce Type: new Abstract: Surgical vision-language foundation models typically adopt educational materials, such as surgical lecture videos, to transfer surgical knowledge encode

researcharxiv-cs-cv
1 Jul 2026
Tutorials

Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility

DGX agent

arXiv:2505.18521v2 Announce Type: replace Abstract: The substantial training cost of diffusion models hinders their deployment. Immiscible Diffusion recently showed that reducing diffusion trajectory

tutorialsarxiv-cs-cv
1 Jul 2026
Safety

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

DGX agent

arXiv:2606.31109v1 Announce Type: new Abstract: Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving communit

safetyarxiv-cs-cv
1 Jul 2026
Tutorials

InstanceControl: Controllable Complex Image Generation without Instance Labeling

DGX agent

arXiv:2606.31924v1 Announce Type: new Abstract: Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to g

tutorialsarxiv-cs-cv
1 Jul 2026
Research

Intrinsically Stable Spiking Neural Networks: Overcoming the Performance Barrier in the Absence of Batch Normalization

DGX agent

arXiv:2606.31695v1 Announce Type: new Abstract: The performance of deep spiking neural networks (SNNs) often relies on batch normalization (BN). However, the advanced dynamic BN variants used in state

researcharxiv-cs-cv
1 Jul 2026
Model Releases

JacobianAvatar: Temporally Consistent Semi-rigid Avatar Reconstruction from a Monocular Video

DGX agent

arXiv:2606.31115v1 Announce Type: new Abstract: Generating realistic human avatars in complex motions--such as clothing dynamics--requires modeling of global and local deformations which remains chall

model-releasesarxiv-cs-cv
1 Jul 2026
Research

Joint Optimization for 4D Human-Scene Reconstruction in the Wild

DGX agent

arXiv:2501.02158v3 Announce Type: replace Abstract: Reconstructing human motion and its surrounding environment is crucial for understanding human-scene interaction and predicting human movements in t

researcharxiv-cs-cv
1 Jul 2026
Agents

Knowledge-Driven Dimension Estimation from a Single Image -3D Asset Generation Technology for Digital Twin Construction

DGX agent

arXiv:2606.30896v1 Announce Type: new Abstract: In the verification of in-vehicle cameras, simulation technology using virtual spaces has advanced, enabling pre-evaluation of false detections and miss

agentsarxiv-cs-cv
1 Jul 2026
Safety

LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior

DGX agent

arXiv:2603.25399v2 Announce Type: replace Abstract: We introduce extbf{LaMP}, a dual-expert Vision-Language-Action framework that embeds dense 3D scene flow as a latent motion prior for robotic manipu

safetyarxiv-cs-cv
1 Jul 2026
Safety

Language-Assisted Super-Resolution from Real-World Low-Resolution Patches

DGX agent

arXiv:2606.31363v1 Announce Type: new Abstract: Single image super-resolution aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs. Training SR models typically requires pai

safetyarxiv-cs-cv
1 Jul 2026
Research

Learning a Sampling-Free Variational DNN Plugin from Tiny Training Sets to Refine OOD Segmentation With Uncertainty Estimation

DGX agent

arXiv:2606.15837v2 Announce Type: replace Abstract: Deep neural networks (DNNs) frequently fail to generalize to out-of-distribution (OOD) medical images because of variations in scanners and acquisit

researcharxiv-cs-cv
1 Jul 2026
Model Releases

Learning to Deny: Action Denial in Multimodal Large Language Models

DGX agent

arXiv:2606.31187v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero-shot and few-shot recognition across standard

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR

DGX agent

arXiv:2601.14251v2 Announce Type: replace Abstract: We present LightOnOCR-2-1B, a 1B-parameter end-to-end multilingual vision--language model that converts document images (e.g., PDFs) into clean, nat

model-releasesarxiv-cs-cv
1 Jul 2026
Research

LiteMatch: Lightweight Zero-Shot Stereo Matching via Cost Volume Stabilization

DGX agent

arXiv:2606.31636v1 Announce Type: new Abstract: Despite rapid progress in learning-based stereo matching, high accuracy is often achieved at the cost of heavy backbones and computationally intensive 3

researcharxiv-cs-cv
1 Jul 2026
Research

Localized Conformal Prediction for Image Classification with Vision-Language Models

DGX agent

arXiv:2606.31577v1 Announce Type: new Abstract: Conformal predictions have attracted significant attention in the field of uncertainty quantification, mainly because of their strong marginal coverage

researcharxiv-cs-cv
1 Jul 2026
Model Releases

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

DGX agent

arXiv:2601.08758v4 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) reasoning has proven effective in enhancing large language models by encouraging step-by-step intermediate reasoning, a

model-releasesarxiv-cs-cv
1 Jul 2026
Research

MAPE: Defending Against Transferable Adversarial Attacks Using Multi-Source Adversarial Perturbations Elimination

DGX agent

arXiv:2606.31378v1 Announce Type: new Abstract: Neural networks are vulnerable to meticulously crafted adversarial examples, leading to high-confidence misclassifications in image classification tasks

researcharxiv-cs-cv
1 Jul 2026
Model Releases

Medical Image Spatial Grounding with Semantic Sampling

DGX agent

arXiv:2603.14579v3 Announce Type: replace Abstract: Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs rep

model-releasesarxiv-cs-cv
1 Jul 2026
Applications

MemLearner: Learning to Query Context memory for Video World Models

DGX agent

arXiv:2606.31734v1 Announce Type: new Abstract: Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical c

applicationsarxiv-cs-cv
1 Jul 2026
Research

Mesh BDF: Barycentric Dominance Field for 3D Native Mesh Generation

DGX agent

arXiv:2606.31777v1 Announce Type: new Abstract: Autoregressive (AR) modeling has recently achieved remarkable progress in native 3D mesh generation, largely due to its natural ability to handle variab

researcharxiv-cs-cv
1 Jul 2026
Local Ai

MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images

DGX agent

arXiv:2506.09919v4 Announce Type: replace Abstract: We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle

local-aiarxiv-cs-cv
1 Jul 2026
Research

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs

DGX agent

arXiv:2606.31383v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) typically employ resampling-based projectors to transform dense visual features into a compact token sequence f

researcharxiv-cs-cv
1 Jul 2026
Research

MSNN-LINet: Cross-Modal Learning via Continuous Linear Integration

DGX agent

arXiv:2606.31135v1 Announce Type: new Abstract: We present LINet (Linear Integration Network), a Multi-Stream Neural Network (MSNN) for RGB-D scene classification. Current multi-modal architectures tr

researcharxiv-cs-cv
1 Jul 2026
← Previous
1…7475767778…263
Next →