AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
30 Jul 2026

Neural Radiance Fields for the Real World: A Survey

ApplicationsDGX agent

arXiv:2501.13104v3 Announce Type: replace Abstract: Neural Radiance Fields (NeRFs) have remodeled 3D scene representation since release. NeRFs can effectively reconstruct complex 3D scenes from 2D ima

Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation

Local AiDGX agent

arXiv:2607.27058v1 Announce Type: new Abstract: Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2505.22039v2 Announce Type: replace Abstract: While anomaly detection has made significant progress, generating detailed analyses that incorporate industrial knowledge remains a challenge. To ad

One-Frame Calibration with Siamese Network in Facial Action Unit Recognition

SafetyDGX agent

arXiv:2409.00240v2 Announce Type: replace Abstract: Automatic facial action unit (AU) recognition is used widely in facial expression analysis. Most existing AU recognition systems aim for cross-parti

Online Handwriting Trajectory Reconstruction from Kinematic Sensors using Temporal Convolutional Network

Model ReleasesDGX agent

arXiv:2607.26733v1 Announce Type: new Abstract: Handwriting with digital pens is a common way to facilitate human-computer interaction through the use of Online Handwriting (OH) trajectory reconstruct

Particle-Filtering-based Latent Diffusion for Inverse Problems

ResearchDGX agent

arXiv:2408.13868v2 Announce Type: replace Abstract: Current strategies for solving image-based inverse problems apply latent diffusion models to perform posterior sampling.However, almost all approach

PatchDenoiser: Parameter-efficient multi-scale patch learning and fusion denoiser for Low-dose CT imaging

Model ReleasesDGX agent

arXiv:2602.21987v3 Announce Type: replace Abstract: Low-dose CT images are essential for reducing radiation exposure in cancer screening, pediatric imaging, and longitudinal monitoring protocols, but

Physically Real-time Infrared Attack against Optical Flow Estimation Networks

SafetyDGX agent

arXiv:2607.26651v1 Announce Type: new Abstract: With the promising performance of deep neural networks on image-based tasks, different real-world applications such as autonomous driving and motion det

Prior Directions: Why GUI Grounding Gets Locked in the Past

ResearchDGX agent

arXiv:2607.26913v1 Announce Type: new Abstract: Vision-language models often use descriptions of earlier visual states to make decisions about the current scene. When the scene changes, stale language

PRISM-Net: Patient-specific reference-guided inter-breast symmetry matching for three-class breast DCE-MRI classification

ResearchDGX agent

arXiv:2607.26799v1 Announce Type: new Abstract: Breast DCE-MRI AI is increasingly being explored for breast-level classification of no-lesion, benign, and malignant findings, beyond conventional lesio

Progressive Multimodal Alignment for Continual Instruction Tuning

Model ReleasesDGX agent

arXiv:2607.26947v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cro

R-SLPR: Region-based Small-to-Large Point-cloud Registration with Contrastive Learning

SafetyDGX agent

arXiv:2607.26583v1 Announce Type: new Abstract: Point-cloud (PC) registration is fundamental to three-dimensional (3D) perception in robotic systems. However, classic registration algorithms falter wh

RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry

ResearchDGX agent

arXiv:2603.01544v2 Announce Type: replace Abstract: Recent image generators produce photo-realistic content that undermines the reliability of downstream recognition systems. As visual appearance cues

Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography

Model ReleasesDGX agent

arXiv:2607.26196v1 Announce Type: new Abstract: Self-supervised pretraining is central to 3D medical image analysis, where unlabeled CT volumes are abundant but expert annotations are scarce. Yet exis

Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation

ResearchDGX agent

arXiv:2607.26395v1 Announce Type: new Abstract: White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spati

Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification

ResearchDGX agent

arXiv:2607.26565v1 Announce Type: new Abstract: Vision models do not form a representation at once; each block revises it. We ask whether the resulting computation path contains evidence that the fina

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

SafetyDGX agent

arXiv:2607.26333v1 Announce Type: cross Abstract: Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. However

Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory

TutorialsDGX agent

arXiv:2607.26818v1 Announce Type: new Abstract: Audio-video generative models achieve impressive quality but suffer from high latency, making them unsuitable for real-time applications. Although sever

Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences

ResearchDGX agent

arXiv:2607.26973v1 Announce Type: new Abstract: Accurate refinement of Rational Polynomial Camera (RPC) models is essential for high-quality satellite image geolocation. In ground control point (GCP)-

ScalablePromptus: Scalable and High-Fidelity Prompt-Based Video Streaming

ApplicationsDGX agent

arXiv:2607.26106v1 Announce Type: cross Abstract: Prompt-based video streaming transmits compact semantic prompts instead of pixel-level content for generative reconstruction, enabling ultra-low-bitra

SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation

SafetyDGX agent

arXiv:2607.26885v1 Announce Type: new Abstract: Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are

SceneExpander: Text-Guided 3D Scene Expansion via Free-Form View Insertion

ResearchDGX agent

arXiv:2603.27084v3 Announce Type: replace Abstract: World building with 3D scene representations is increasingly important for content creation, simulation, and interactive experiences, yet real workf

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

SafetyDGX agent

arXiv:2607.27066v1 Announce Type: new Abstract: Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully s

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Model ReleasesDGX agent

arXiv:2607.27084v1 Announce Type: new Abstract: Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in

ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection

Local AiDGX agent

arXiv:2607.27065v1 Announce Type: new Abstract: While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annota

Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification

SafetyDGX agent

arXiv:2607.26765v1 Announce Type: new Abstract: Background/Objectives: Dermoscopic skin lesion classifiers often lose accuracy under domain shift across imaging devices, illumination, and capture arti

SeasonStereo: Robust Dense Stereo Matching for Multi-Date Satellite Imagery via Generative AI

ResearchDGX agent

arXiv:2607.27139v1 Announce Type: new Abstract: Accurate 3D reconstruction from satellite imagery typically relies on near-simultaneous stereo pairs, limiting its applicability to diachronic settings

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

ApplicationsDGX agent

arXiv:2607.26769v1 Announce Type: new Abstract: Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.26326v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. How

Semantic-Aware Temporal Adaptation for UAV Anti-UAV Tracking

Model ReleasesDGX agent

arXiv:2607.26511v1 Announce Type: new Abstract: UAV Anti-UAV tracking is an emerging low-altitude security task for localizing an adversarial UAV using the onboard camera of a moving observer UAV. It

Sequence-SOD: Bio-inspired Sequence-aware Spiking ObjectDetection for Event Cameras

Local AiDGX agent

arXiv:2607.26703v1 Announce Type: new Abstract: Event cameras follow a retina-inspired sensing principle, reporting local intensity changes asynchronously with hightemporal resolution and a wide dynam

Shape-Based Inductive Bias for Glioma Grading from Tumor Contours

SafetyDGX agent

arXiv:2607.26090v1 Announce Type: cross Abstract: Glioma grading from tumor contours is often treated as a pixel problem even when the signal of interest is shape. We align closed contours with a func

SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions

SafetyDGX agent

arXiv:2603.23118v2 Announce Type: replace Abstract: Recent works have shown that multimodal large language models (MLLMs) are highly vulnerable to hidden-pattern visual illusions, where the hidden con

SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM

Model ReleasesDGX agent

arXiv:2607.26595v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing dem

Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots

ApplicationsDGX agent

arXiv:2607.26567v1 Announce Type: cross Abstract: Humanoid robots increasingly require multi-modal understanding for natural interaction with humans. Despite the prominence of vision-language models,

Spline-Based Boundary Representations for Sparse View Reconstruction and Simulation Using Isogeometric Analysis

ResearchDGX agent

arXiv:2607.26234v1 Announce Type: new Abstract: Image-based reconstruction aims to recover three-dimensional geometry from images. Recent advances have enabled the recovery of visually detailed models

SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision

ResearchDGX agent

arXiv:2603.27519v2 Announce Type: replace Abstract: Image-based plant phenotyping depends on dense structural understanding of crops, yet pixel-level annotation remains expensive across species, organ

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

TutorialsDGX agent

arXiv:2607.26754v1 Announce Type: new Abstract: Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by p

Step-Attention Refinement of DINOv3 Features for Efficient Anterior Eye Segmentation

ResearchDGX agent

arXiv:2607.27087v1 Announce Type: new Abstract: Anterior eye segment (AES) segmentation is a key component of both ocular biometrics and emerging clinical image analysis applications. However, heterog

StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction

ResearchDGX agent

arXiv:2607.26889v1 Announce Type: cross Abstract: Reconstructing articulated objects with multiple movable parts is essential for understanding object structure and enabling physical interaction. Howe

Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs

SafetyDGX agent

arXiv:2607.27122v1 Announce Type: new Abstract: Gastrointestinal (GI) endoscopic image analysis has shifted from single-label classification toward visual question answering (VQA), where a model must

TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.26536v1 Announce Type: new Abstract: High-pressure prompts can push vision-language models (VLMs) into unsupported commitments, such as reading illegible text, reporting indeterminate times

TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models

ResearchDGX agent

arXiv:2607.26706v1 Announce Type: new Abstract: Text-to-video diffusion models generate temporally coherent content from natural language, yet when a prompt describes an early scene that persists whil

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

SafetyDGX agent

arXiv:2607.26107v1 Announce Type: new Abstract: Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Local AiDGX agent

arXiv:2607.27205v1 Announce Type: new Abstract: Vision-language-action (VLA) models commonly adopt an LLM-centric V o L o A pathway, where visual observations are projected into the representation spa

Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach

Model ReleasesDGX agent

arXiv:2607.26608v1 Announce Type: new Abstract: Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvem

Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

SafetyDGX agent

arXiv:2607.27113v1 Announce Type: new Abstract: The growing capability of image generation models has made synthetic images a routine presence in open media, making robust and generalizable AI-Generat

VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion

ResearchDGX agent

arXiv:2607.27194v1 Announce Type: new Abstract: Accurately recovering the camera's calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and sc

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

HardwareDGX agent

arXiv:2607.26694v1 Announce Type: new Abstract: We present Visko Orbis 1.0, a Live Model for real-time, interactive long-video generation. Users can change the prompt at any moment during generation,

Visual Credit Audit for Multimodal Spatial Reasoning

Model ReleasesDGX agent

arXiv:2607.27069v1 Announce Type: new Abstract: Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choi

Walk through Paintings: Egocentric World Models from Internet Priors

ResearchDGX agent

arXiv:2601.15284v2 Announce Type: replace Abstract: What if a video generation model could not only imagine a plausible future, but the correct one -- accurately reflecting how the world changes with

Weight and Height Estimation from a Single Human Image Captured in the Wild

ResearchDGX agent

arXiv:2607.26104v1 Announce Type: new Abstract: A person's physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routines and financ

When Fish Look Alike: Tracking Identities with Dual-branch Elasticity

Model ReleasesDGX agent

arXiv:2607.26412v1 Announce Type: new Abstract: Tracking dense, homogeneous targets like schooling fish remains a major challenge for multiple object tracking due to extreme inter-individual homogenei

Where Physics Meets Privacy: Federated PINNs for Privacy-Preserving Brain Tumor Biomechanical Modeling

Local AiDGX agent

arXiv:2607.26207v1 Announce Type: new Abstract: Brain tumors such as glioma, meningioma, and pituitary adenoma alter the mechanical behavior of soft brain tissue, yet common diagnostic methods rely on

WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

Model ReleasesDGX agent

arXiv:2607.26203v1 Announce Type: new Abstract: Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its imp

Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment

Model ReleasesDGX agent

arXiv:2607.26381v1 Announce Type: new Abstract: Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require labeled Wi-Fi

29 Jul 2026

A Functional Approach to Curve Alignment and Shape Analysis

SafetyDGX agent

arXiv:2503.05632v2 Announce Type: cross Abstract: In many image analysis problems, the contours of objects carry important statistical information about shape. Such contours are typically affected by

A systematic evaluation of machine learning classifiers for event-by-event background rejection in LAFOV PET scanners

ResearchDGX agent

arXiv:2607.25732v1 Announce Type: new Abstract: The introduction of LAFOV PET scanners brings significant sensitivity gains but also a substantial increase in the background rate from accidental coinc

A Unified Benchmark and Modality-Adaptive Network for Day-and-Night Drone-View Geo-Localization

Model ReleasesDGX agent

arXiv:2607.25778v1 Announce Type: new Abstract: Most existing drone-view geo-localization (DVGL) benchmarks contain drone imagery captured under a single illumination condition and lack geographically

A2D2: Audi Autonomous Driving Dataset

SafetyDGX agent

arXiv:2004.06320v2 Announce Type: replace Abstract: Research in machine learning, mobile robotics, and autonomous driving is accelerated by the availability of high quality annotated data. To this end

← Previous
1…2627282930…209
Next →