AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
11 Aug 2026

Progressive Learned Image Compression for Machine Perception

ApplicationsDGX agent

arXiv:2512.20070v2 Announce Type: replace Abstract: Recent advances in learned image codecs have extended from human perception toward machine perception However, progressive image compression with fi

RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing

Model ReleasesDGX agent

arXiv:2608.09186v1 Announce Type: new Abstract: Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geom

RayLift: Lifting Complementary Ray-Wise Evidence with 3D Geometry Priors for Semantic Scene Completion

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.08476v1 Announce Type: new Abstract: Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing method

Real Data Closes Synthetic-to-Real Gap in Optical Chemical Structure Recognition

Model ReleasesDGX agent

arXiv:2608.09100v1 Announce Type: cross Abstract: Millions of chemical structures appear in patents and papers only as drawings, and using that information at scale requires reading the drawings. OCSR

Real-time physics inversion for retrieval of sub-pixel wildfire temperatures from VSWIR imaging spectroscopy

HardwareDGX agent

arXiv:2608.07580v1 Announce Type: new Abstract: In this work, we present a wildfire temperature retrieval framework for VSWIR imaging spectroscopy data, employed on data from NASA's Airborne Visible I

RealDenseFace: Real-time Monocular 3D Face Reconstruction from Dense UV-space Priors

Model ReleasesDGX agent

arXiv:2608.09238v1 Announce Type: new Abstract: Recent monocular 3D face reconstruction methods achieve high fidelity by fitting a 3D Morphable Model (3DMM) to dense priors predicted by networks, but

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection

Local AiDGX agent

arXiv:2608.09147v1 Announce Type: new Abstract: Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that l

Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable Cells

SafetyDGX agent

arXiv:2608.09658v1 Announce Type: cross Abstract: Human-Robot Collaboration (HRC) plays a vital role in dynamic, high mix, low volume industrial scenarios such as remanufacturing, which frequently fac

RenderMatte: Exact-Alpha Rendering and Group-Relative Alignment for Image Matting

Model ReleasesDGX agent

arXiv:2608.08487v1 Announce Type: new Abstract: Image matting is an essential enabling technology for modern visual content production, where foreground extraction determines the realism and editabili

ResemBrick: Brick Reconstruction from Photographs with Perceptual Fidelity and Buildability

ResearchDGX agent

arXiv:2608.09597v1 Announce Type: new Abstract: Producing a hand-buildable, colored brick model of a 3D object from a few casual photographs is a clean testbed for a broader challenge: generating 3D c

Rethinking 3D Segmentation from Individual LiDAR Scans: Incidence-Aware Sampling on the SIP Benchmark

Model ReleasesDGX agent

arXiv:2608.07757v1 Announce Type: new Abstract: 3D scene understanding is increasingly important in construction, yet most methods are developed on curated datasets that do not fully reflect real site

Rethinking Attention Locality in Spiking Transformers

Model ReleasesDGX agent

arXiv:2608.08541v1 Announce Type: new Abstract: Spiking Transformers provide a promising paradigm for efficient visual processing with spike-driven computation, yet their Softmax-free Spiking Self-Att

Retrieval-Augmented Generation-Based Color Restoration for Low-Light Image Enhancement

SafetyDGX agent

arXiv:2608.08211v1 Announce Type: cross Abstract: Recent low-light image enhancement (LLIE) methods have driven brightness and structural fidelity close to that of normally-exposed images, yet their o

Revisiting the Current Frame: Physical-Trace-Guided Network Output Correction for Video Restoration

ResearchDGX agent

arXiv:2608.09342v1 Announce Type: new Abstract: Video restoration methods exploit temporal information to recover information missing from degraded observations. However, reference frames within the s

Right Answer, Wrong Heat: Explanation-Aware Evaluation and Thermal-Grounded Feedback for MLLMs on Infrared Images

ResearchDGX agent

arXiv:2608.09145v1 Announce Type: new Abstract: General-purpose multimodal large language models (MLLMs) are increasingly applied to infrared images, where they are commonly scored by answer accuracy

RMR-Net: Degradation-Evidence-Guided Road-Image Restoration for Defect Detection

ResearchDGX agent

arXiv:2608.08957v1 Announce Type: new Abstract: Vehicle-mounted road cameras are vulnerable to motion blur, defocus, poor illumination, and noise, which can erase thin cracks and pothole boundaries ne

RobustDefect-LLM: Explainable and Robustness-Aware Industrial Surface Defect Classification with Decision Support and AI-Assisted Reporting

SafetyDGX agent

arXiv:2608.08589v1 Announce Type: new Abstract: This paper presents RobustDefect-LLM, an industrial surface-defect inspection framework integrating deep-learning classification, operator-facing visual

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

SafetyDGX agent

arXiv:2608.09853v1 Announce Type: cross Abstract: General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from

SC-Diff: Semantically Calibrated Diffusion for Visible-to-Infrared Image Translation

SafetyDGX agent

arXiv:2608.08555v1 Announce Type: new Abstract: Visible-to-infrared image translation provides a practical way to expand infrared training data using abundant visible images. Diffusion models are prom

SC^{2}-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments

ResearchDGX agent

arXiv:2608.07548v1 Announce Type: cross Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observabili

SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange

ResearchDGX agent

arXiv:2608.07923v1 Announce Type: new Abstract: Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training

SCTD 3.0: Sonar Common Target Detection in the Wild - A Large-Scale, Multi-Scene Dataset from Real Marine Surveys

Model ReleasesDGX agent

arXiv:2608.08106v1 Announce Type: new Abstract: Synthetic Aperture Sonar (SAS) is core for wide-area detection of small underwater targets. However, large-scale, high-quality SAS datasets are scarce,

Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence

ApplicationsDGX agent

arXiv:2608.08075v1 Announce Type: cross Abstract: Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archi

SegDem: Segmentation helps Demosaicing

ResearchDGX agent

arXiv:2608.07916v1 Announce Type: new Abstract: Image demosaicing reconstructs a full-color image from incomplete color measurements produced by a sensor covered with a color filter array (CFA). Most

Sekai2: From World Exploration to Interactive World Modeling

Model ReleasesDGX agent

arXiv:2608.09449v1 Announce Type: new Abstract: Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefor

Semi-Dense Matching Uncertainty Is Not Just Local Confidence

ResearchDGX agent

arXiv:2608.08685v1 Announce Type: new Abstract: Reliable semi-dense matching is essential for modern geometric vision systems. Designed under a coarse-to-fine paradigm, it achieves an optimal balance

SeqLoc: Beyond the Single Frame for Cross-View Geo-Localization in Feature-Sparse Scenes

Model ReleasesDGX agent

arXiv:2608.07835v1 Announce Type: new Abstract: Cross-View Geo-Localization (CVGL) with OpenStreetMap (OSM) performs well in structure-rich urban environments but collapses in feature-sparse scenes su

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models

ApplicationsDGX agent

arXiv:2608.08839v1 Announce Type: cross Abstract: World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actio

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

Model ReleasesDGX agent

arXiv:2608.09097v1 Announce Type: new Abstract: Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly

SIP: Site in Pieces- A Dataset of Disaggregated Construction-Phase 3D Scans for Semantic Segmentation and Scene Understanding

SafetyDGX agent

arXiv:2512.09062v2 Announce Type: replace Abstract: Accurate 3D scene interpretation in active construction sites is essential for progress monitoring, safety assessment, and digital twin development.

SLAP: Selective Local Vision-Language Alignment for Fish Re-Identification via Partial Optimal Transport

Local AiDGX agent

arXiv:2608.08840v1 Announce Type: new Abstract: Individual fish re-identification (ReID) is a fine-grained recognition problem in which identity-discriminative cues are often localized to specific bod

Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football

ResearchDGX agent

arXiv:2608.09887v1 Announce Type: new Abstract: Ball possession is the most-cited and most-misleading number in football: 60% recycled in one's own half is not 60% spent pinning the opponent back. Exi

Sparse Attention to Emotion: Efficient Facial Emotion Recognition via Token Reduction

ApplicationsDGX agent

arXiv:2608.08873v1 Announce Type: new Abstract: Facial Emotion Recognition (FER) is an important task that has significant implications across various fields such as biometrics, health, and human-comp

SplitGaussian: Reconstructing Dynamic Scenes via Visual Geometry Decomposition

ResearchDGX agent

arXiv:2508.04224v2 Announce Type: replace Abstract: Reconstructing dynamic 3D scenes from monocular video remains fundamentally challenging due to the need to jointly infer motion, structure, and appe

SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning

Local AiDGX agent

arXiv:2608.07932v1 Announce Type: new Abstract: Sports video analysis is crucial for athletic analytics and broadcasting enhancement. Dense sports video reasoning, however, demands a fine-grained unde

SRE-FER: Regional residual evidence learning for mitigating local evidence dilution in fine-grained facial expression recognition

Local AiDGX agent

arXiv:2608.08702v1 Announce Type: new Abstract: Fine-grained facial expression recognition (FER) hinges on capturing subtle muscular cues that distinguish adjacent emotions. Yet capturing these cues p

Staying True to the Origin: Continuous Image Stylization with Smooth Transitions

Model ReleasesDGX agent

arXiv:2608.08125v1 Announce Type: new Abstract: Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a g

SUMI: Scalable Unified Model for 3D Point Cloud Inference

ResearchDGX agent

arXiv:2608.08115v1 Announce Type: new Abstract: Point cloud completion commonly follows a coarse-to-fine paradigm, where a low-density coarse shape is first predicted and then upsampled to the target

SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping

Model ReleasesDGX agent

arXiv:2608.09497v1 Announce Type: new Abstract: Operational crop mapping requires models that generalise across years, resolve fine-grained crop taxonomies, and distinguish cropland from surrounding l

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model

SafetyDGX agent

arXiv:2608.07948v1 Announce Type: new Abstract: VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling comp

Task-Adaptive 3D Cross-Field MRI Translation via Field-Conditioned Content-Style Pretraining

ResearchDGX agent

arXiv:2608.09264v1 Announce Type: new Abstract: Magnetic field strength is a major source of domain shift in magnetic resonance imaging (MRI), affecting signal-to-noise ratio, tissue contrast, spatial

TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching

ResearchDGX agent

arXiv:2608.09590v1 Announce Type: new Abstract: Learning reliable correspondences between images and point clouds is fundamental for 2D-3D matching. Despite recent progress in detection-free methods,

Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation

Model ReleasesDGX agent

arXiv:2608.08290v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning

AgentsDGX agent

arXiv:2608.09682v1 Announce Type: new Abstract: Tool-augmented vision-language models increasingly 'think with images': they call crop, zoom, or code tools and reason over the returned pixels. However

Tokenizer Generator Coupling in Medical Image Generation

ResearchDGX agent

arXiv:2608.07713v1 Announce Type: new Abstract: Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation is valid in a controlled ChestMNIST

Topology-Aware Global-Local Mamba Networks for Palm Vein Biometrics

Model ReleasesDGX agent

arXiv:2608.08951v1 Announce Type: new Abstract: Palm-vein recognition is a fine-grained biometric task in which both local vascular texture and the global layout of the vessel tree carry discriminativ

Toward Mask Annotation-Free Surgical Instrument Segmentation from Endoscopic Images Using Text-Prompted Segment Anything Model 3 (SAM3)

Model ReleasesDGX agent

arXiv:2608.08844v1 Announce Type: new Abstract: Surgical instrument segmentation is a fundamental task for computer-assisted interventions, yet most existing methods rely on pixel-level annotations or

Towards Adaptive Super-Resolution and Quality Assessment via Test-Time Adaptation

ApplicationsDGX agent

arXiv:2608.08508v1 Announce Type: new Abstract: This paper presents doctoral research on adaptive video super-resolution and perceptual quality modeling under real-world conditions. Existing video sup

Towards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives

AgentsDGX agent

arXiv:2608.09541v1 Announce Type: new Abstract: Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enablin

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

SafetyDGX agent

arXiv:2608.09529v1 Announce Type: new Abstract: As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) gen

TriView-YOLO: Early Multi-View Fusion for Ground Penetrating Radar Cavity Detection in Soft, High-Water-Content Soils

ResearchDGX agent

arXiv:2608.09522v1 Announce Type: new Abstract: Automated detection of subsurface cavities from Ground Penetrating Radar (GPR) is most difficult in soft, high-water-content ground, where conductive, w

Tropical Cyclone Forecasting via Latent Rectified Flow using Satellite Imagery and Atmospheric Fields

ResearchDGX agent

arXiv:2608.08354v1 Announce Type: new Abstract: Tropical cyclones are growing more destructive in a changing climate, and efficient forecasting of their structure and track has become a necessity. Dee

Uncertainty-Aware 4D Gaussian Splatting for Monocular Occluded Human Rendering

ResearchDGX agent

arXiv:2602.06343v3 Announce Type: replace Abstract: High-fidelity rendering of dynamic humans from monocular videos typically degrades catastrophically under occlusions. Existing solutions incorporate

UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

Local AiDGX agent

arXiv:2608.09143v1 Announce Type: new Abstract: Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodifi

UniScale: Arbitrary-Scale Industrial Anomaly Generation

TutorialsDGX agent

arXiv:2608.07864v1 Announce Type: new Abstract: Industrial anomaly inspection faces a major challenge due to the lack of real-world anomaly samples. While generative models are used to create anomaly

Unsupervised Domain Adaptation for Multitask Image Analysis in Realistic Context with Extreme Label Shift; Application to the CTAO first Large Sized Telescope

ResearchDGX agent

arXiv:2608.09630v1 Announce Type: cross Abstract: Unsupervised domain adaptation is a widespread set of methods that leverages the knowledge of a labeled source domain to train a model to perform well

Unsupervised Point Cloud Registration with Self-Distillation

Model ReleasesDGX agent

arXiv:2409.07558v2 Announce Type: replace Abstract: Rigid point cloud registration is a fundamental problem and highly relevant in robotics and autonomous driving. Nowadays deep learning methods can b

Unveiling the Secret of AdaLN-Zero in Diffusion Transformer

ResearchDGX agent

arXiv:2608.09438v1 Announce Type: new Abstract: Diffusion transformer (DiT), a rapidly emerging architecture for image generation, has gained much attention. However, despite ongoing efforts to improv

UPolarSQ: Polar Representation Learning for Optic Disc and Peripapillary Atrophy Segmentation and Quantification in Fundus Photographs

ResearchDGX agent

arXiv:2608.08771v1 Announce Type: new Abstract: Myopia-induced posterior-pole remodeling is frequently accompanied by Optic Disc (OD) deformation and Peripapillary Atrophy (PPA), both of which provide

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models

SafetyDGX agent

arXiv:2608.08622v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses u

← Previous
1…45678…207
Next →