AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
7 Jul 2026

Diverse Normal Prototypes-Guided Contrastive Reconstruction for Medical Anomaly Detection

SafetyDGX agent

arXiv:2508.19573v2 Announce Type: replace Abstract: Anomaly detection in medical images is challenging due to limited annotations and the domain gap. Existing reconstruction-based methods often rely o

Diversity-aware View Partitioning for Scalable VGGT

ResearchDGX agent

arXiv:2607.01885v2 Announce Type: replace Abstract: Geometry transformers such as VGGT achieve strong performance by jointly reasoning over multiple views with global attention. However, scaling them

Do Diabetic Foot Ulcer Segmentation Models Generalize? A Cross-Dataset Benchmark of CNN and Transformer Architectures

Model ReleasesDGX agent

arXiv:2607.02555v1 Announce Type: new Abstract: Deep learning models for diabetic foot ulcer (DFU) segmentation routinely report high accuracy, but they are almost always trained and tested on the sam


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Do Flat Minima Improve Sparse Novel View Synthesis?

ResearchDGX agent

arXiv:2511.17918v2 Announce Type: replace Abstract: Despite the success of recent novel view synthesis methods, they tend to struggle in sparse-view settings. This poor generalization to unseen viewpo

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs

Model ReleasesDGX agent

arXiv:2607.03647v1 Announce Type: new Abstract: Large vision language models (VLMs) report strong accuracy on medical question-answering, yet it remains unclear whether they reason from visual evidenc

DriftST: One-Step Generative Inference of Spatial Transcriptomics from H&E Histology

ResearchDGX agent

arXiv:2607.04740v1 Announce Type: new Abstract: Spatial Transcriptomics (ST) measures gene expression while preserving spatial context, but its high cost and low throughput leave public datasets small

DS-SAC: Density Search for Sample Consensus

ApplicationsDGX agent

arXiv:2607.03972v1 Announce Type: new Abstract: Robust geometric model estimation is a fundamental problem in computer vision. RANSAC and its variants remain widely used for this task; however, they r

Dual-Adaptive SAM3: Hierarchical Routing over Low-Rank Expert Layers for Parameter-Efficient Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.02571v1 Announce Type: new Abstract: The Segment Anything Model with Concepts (SAM3) heralds a new paradigm for open-vocabulary segmentation through natural language interaction, offering s

DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation

Model ReleasesDGX agent

arXiv:2607.02604v1 Announce Type: new Abstract: Although vision-language-action (VLA) models have received widespread attention, many challenges remain in manipulating dynamic moving objects. In most

E-TraMamba: A New Paradigm for Efficient Long-Term 3D Feature Tracking with Event Cameras

ResearchDGX agent

arXiv:2607.02866v1 Announce Type: new Abstract: Event-based 3D tracking enables low-latency and high-speed perception, while existing CNN- and Transformer-based trackers struggle to capture long-range

EgoInertia-MI: A Multimodal Egocentric Vision and IMU Benchmark for Motor Impairment Assessment

Model ReleasesDGX agent

arXiv:2607.03934v1 Announce Type: new Abstract: Motor impairments, including tremor, bradykinesia, gait abnormalities, and postural instability, are common across many neurological and movement-relate

EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation

Model ReleasesDGX agent

arXiv:2508.16239v2 Announce Type: replace Abstract: Quantitative microstructural characterization is fundamental to materials science, and electron micrographs (EMs) provide indispensable high-resolut

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions

Model ReleasesDGX agent

arXiv:2607.02674v1 Announce Type: new Abstract: Precise control of 3D facial expressions from text is crucial for virtual avatars, animation, and human-computer interaction, yet existing text-to-3D me

EMPURPLE: A Free Lunch for Diffusion Distillation based on the Information Bottleneck

TutorialsDGX agent

arXiv:2607.04276v1 Announce Type: new Abstract: Diffusion models achieve impressive image-generation quality but remain expensive at inference time. Diffusion distillation reduces sampling steps, yet

Enhancing Facial Expression Recognition in Head-Mounted Displays with Synthetic Data

SafetyDGX agent

arXiv:2607.04490v1 Announce Type: new Abstract: Facial expression recognition (FER) is crucial for social interaction in mixed reality environments that employ head-mounted displays (HMD). However, co

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

Local AiDGX agent

arXiv:2607.04636v1 Announce Type: new Abstract: Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance

Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors

SafetyDGX agent

arXiv:2508.09629v2 Announce Type: replace Abstract: We revisit the role of texture in monocular 3D hand reconstruction, not as an afterthought for photorealism, but as a dense, spatially grounded cue

Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising

Model ReleasesDGX agent

arXiv:2607.04653v1 Announce Type: new Abstract: While modern video diffusion models excel in visual fidelity, maintaining long-range physical consistency remains a formidable challenge. Conventional p

Entropy-Coded MS-VQ-VAE with Learned Priors for Ultra-Low Bitrate Video Compression

ResearchDGX agent

arXiv:2607.02562v1 Announce Type: new Abstract: Learned video codecs based on continuous latent representations struggle to operate reliably below 0.1 bits per pixel~(bpp): without a differentiable ra

Erasing Without Collateral Damage: Precise Concept Removal in Diffusion Models

Model ReleasesDGX agent

arXiv:2607.05274v1 Announce Type: new Abstract: Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of

EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots

SafetyDGX agent

arXiv:2607.02646v1 Announce Type: cross Abstract: We present EVA-Client, an open-source framework for deployment, data collection, and evaluation of trained manipulation policies on real robots. Sitti

Evaluating Agentic Harness Systems for Autonomous Computational Pathology

Model ReleasesDGX agent

arXiv:2607.02598v1 Announce Type: new Abstract: Autonomous computational pathology (ACP) converts high-level pathology analysis goals into executable, traceable and clinically bounded workflows. Reali

Evaluating Intellectual Property Guardrails of Generative Image Models: A Technical Report

Model ReleasesDGX agent

arXiv:2607.02582v1 Announce Type: new Abstract: Generative image models are capable of producing images that bear a strong resemblance to, or replicate, recognizable intellectual property (IP). In thi

EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration

Model ReleasesDGX agent

arXiv:2607.04472v1 Announce Type: new Abstract: The rapid proliferation of artificial intelligence-generated content necessitates reliable multimodal forensics. Beyond video-level binary classificatio

Event Detection in Videos: A Framework for the Development of New Methods

ResearchDGX agent

arXiv:2607.04372v1 Announce Type: new Abstract: Event detection tasks in videos, the most important aspect of video surveillance, aim to detect events either at the pixel-level, frame-level, or clip-l

Explainable Flood Segmentation on Sentinel-1 SAR1 Imagery Using CNN and Transformer Architectures

Model ReleasesDGX agent

arXiv:2606.16302v2 Announce Type: replace Abstract: Rapid and accurate flood prediction is essential for disaster response and mitigation planning. Synthetic Aperture Radar (SAR) sensors in satellites

Exploring SAM Supervision for Fine-Grained UAV Target Segmentation under Data Scarcity

ApplicationsDGX agent

arXiv:2607.03754v1 Announce Type: new Abstract: Unmanned aerial vehicle (UAV) target segmentation remains challenging due to the small size of objects, appearance variations, cluttered backgrounds, an

ExpoMotion: A Large-Scale Benchmark and A Householder Projection Network for Multi-Exposure Fusion

Model ReleasesDGX agent

arXiv:2607.03110v1 Announce Type: new Abstract: Multi-Exposure Fusion (MEF) effectively extends dynamic range, but practical deployment is hindered by motion-induced ghosting and the scarcity of high-

FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers

Model ReleasesDGX agent

arXiv:2607.03180v1 Announce Type: new Abstract: Multimodal diffusion transformers (MM-DiTs) have emerged as the prevalent backbone for modern text-to-image generation systems. However, they exhibit cr

Fast 3D Foundation Model Initialized Gaussian Splatting

AgentsDGX agent

arXiv:2607.03209v1 Announce Type: new Abstract: This paper introduces a fast method for high-quality 3D Gaussian Splatting (3DGS) reconstruction without traditional Structure-from-Motion (SfM). The pr

FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction

Local AiDGX agent

arXiv:2607.03822v1 Announce Type: new Abstract: Vision-based 3D occupancy prediction fundamentally relies on the 2D-to-3D view transformation. Current paradigms predominantly utilize explicit physical

FedProIn: Mitigating Client Drift for Learnable Prototypes in Federated Medical Imaging

Model ReleasesDGX agent

arXiv:2607.04158v1 Announce Type: cross Abstract: Federated learning (FL) is severely hindered by statistical heterogeneity due to variations in scanners, acquisition protocols, and patient population

Fields of the Planet: Field Boundary Mapping Beyond 10m

ResearchDGX agent

arXiv:2607.04449v1 Announce Type: new Abstract: Field-boundary maps support crop monitoring, irrigation planning, and yield estimation, but many smallholder parcels span only a few 10 m Sentinel-2 pix

Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

ResearchDGX agent

arXiv:2607.04461v1 Announce Type: new Abstract: Inference-time scaling for text-to-image generation has progressed from simple Best-of-N (BoN) sampling to guided search methods that verify and steer c

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

SafetyDGX agent

arXiv:2607.03509v1 Announce Type: new Abstract: Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inferen

FlowMark: Mask-Guided Video Watermarking

ApplicationsDGX agent

arXiv:2607.05261v1 Announce Type: new Abstract: We present FlowMark, a video watermarking framework guided by automatically predicted object masks. In contrast to prior region-based approaches that re

Fortifying Fully Convolutional Generative Adversarial Networks for Image Super-Resolution Using Divergence Measures

Model ReleasesDGX agent

arXiv:2404.06294v2 Announce Type: replace-cross Abstract: Super-Resolution (SR) is a time-hallowed image processing problem that aims to improve the quality of a Low-Resolution (LR) sample up to the s

Fourier Splatting: Generalized Fourier encoded primitives for scalable radiance fields

ResearchDGX agent

arXiv:2603.19834v3 Announce Type: replace Abstract: Novel view synthesis has recently been revolutionized by 3D Gaussian Splatting (3DGS), which enables real-time rendering through explicit primitive

Framework and Multi-modal Dataset for Roadwork Zone Detection and Geo-localization

Model ReleasesDGX agent

arXiv:2607.04330v1 Announce Type: new Abstract: Autonomous vehicles often rely on high-definition (HD) maps for navigation; however, these maps are not frequently updated and often lack semi-static in

FRFDet: Efficient UAV Small Object Detection with Symmetric Sampling and Scalable Fusion

SafetyDGX agent

arXiv:2607.04125v1 Announce Type: new Abstract: Small object detection in Unmanned Aerial Vehicle (UAV) imagery remains challenging under adverse conditions, including complex weather, low illuminatio

From General Actions to Domain-Specific Monitoring: Prior-Adaptive Transfer for Skeleton-Based Action Recognition

Model ReleasesDGX agent

arXiv:2607.03327v1 Announce Type: new Abstract: Skeleton-based action recognition models have recently shown strong performance on large-scale benchmarks with general actions. However, directly transf

From Geometric Labels to Semantic Understanding of Indoor Building Components Using Multimodal Large Language Models

ApplicationsDGX agent

arXiv:2607.03661v1 Announce Type: new Abstract: Point cloud-based understanding has become an important enabler for facility operation and maintenance involving indoor building components. However, ex

From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation

SafetyDGX agent

arXiv:2607.04691v1 Announce Type: new Abstract: While controllable image generation has made significant strides by incorporating visual reference conditions, existing methods predominantly operate as

From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation

SafetyDGX agent

arXiv:2607.03792v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding

FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small Object Detection

ApplicationsDGX agent

arXiv:2607.05176v1 Announce Type: new Abstract: Small object detection (SOD) remains a challenging task in real-world applications. Despite recent advances, existing detectors remain limited by rigid

Fully Rotation-Equivariant Spectral-Spatial Learning for Multispectral Object Detection

Model ReleasesDGX agent

arXiv:2607.05148v1 Announce Type: new Abstract: Existing multispectral detectors are limited by discrete spectral processing, a scale-dependent shift in the relative reliability of spectral and spatia

FunPhase: A Periodic Functional Autoencoder for Motion Generation via Phase Manifolds

Local AiDGX agent

arXiv:2512.09423v2 Announce Type: replace Abstract: Learning natural body motion remains challenging due to the strong coupling between spatial geometry and temporal dynamics. Embedding motion in phas

G^2TAM: Geometry Grounded Track Anything Model

Local AiDGX agent

arXiv:2607.03789v1 Announce Type: new Abstract: Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across vie

G3Splat: Geometrically Consistent Generalizable Gaussian Splatting

Model ReleasesDGX agent

arXiv:2512.17547v2 Announce Type: replace Abstract: 3D Gaussians have become a powerful scene representation for real-time splatting and high-quality novel-view synthesis. This has motivated generaliz

GALOSH: Blind, Training-Free Denoising of Raw Bayer and sRGB Images by Parallel-Friendly Local Shrinkage

Model ReleasesDGX agent

arXiv:2607.03768v1 Announce Type: cross Abstract: Classical training-free denoisers such as BM3D and non-local means owe much of their strength to search: content-dependent block matching whose memory

GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects

Model ReleasesDGX agent

arXiv:2508.14891v3 Announce Type: replace Abstract: Reconstructing articulated objects is essential for building digital twins of interactive environments. However, prior methods typically decouple ge

Geographic Diversity Beats Data Volume for Cross-Domain Generalization in Zero-Label JEPA Driving World Models

ResearchDGX agent

arXiv:2607.04500v1 Announce Type: new Abstract: Self-supervised latent world models can assign a surprise score to driving scenarios without any human labels. A natural follow-up question is whether s

Geometric Observability Index: An Operator-Theoretic Framework for Per-Feature Sensitivity, Weak Observability, and Dynamic Effects in SE(3) Pose Estimation

Model ReleasesDGX agent

arXiv:2602.05582v2 Announce Type: replace Abstract: We introduce the Geometric Observability Index (GOI), a per-feature sensitivity measure for pose estimation on SE(3). For a Gauss-Newton curvature m

Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation

ResearchDGX agent

arXiv:2607.05354v1 Announce Type: new Abstract: Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR

Geometry-aware Depth-guided Representation Learning for Structure-preserving Low-light Image Enhancement

Model ReleasesDGX agent

arXiv:2607.05005v1 Announce Type: new Abstract: Low-light degradation reduces image visibility and weakens structural cues that are important for visual representation and scene understanding. Existin

GeoSAM-Lite: A Lightweight Foundation Model for Onboard Remote Sensing Segmentation

ResearchDGX agent

arXiv:2607.03760v1 Announce Type: new Abstract: The deployment of large-scale foundation models like Segment Anything Model (SAM) on resource-constrained Earth observation platforms is hindered by pro

GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation

ApplicationsDGX agent

arXiv:2511.23191v2 Announce Type: replace Abstract: Previous works that leverage video models for image-to-3D scene generation often suffer from geometric distortions and blurry content. Using video g

GestaltMML: Enhancing Rare Genetic Disease Diagnosis through Multimodal Machine Learning Combining Facial Images and Clinical Text

ResearchDGX agent

arXiv:2312.15320v3 Announce Type: replace-cross Abstract: Individuals with suspected rare genetic disorders often undergo multiple clinical evaluations, imaging studies, laboratory tests, and genetic

Ghosts Beneath Textures: Texture-Relation Cues for Cross-Paradigm AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2607.03862v1 Announce Type: new Abstract: AI-generated images have proliferated rapidly, motivating extensive research. Most existing AI-generated image detectors are developed and evaluated und

GlaBoost: A Multimodal Structured Framework for Glaucoma Risk Stratification

ApplicationsDGX agent

arXiv:2508.03750v2 Announce Type: replace-cross Abstract: Early and accurate glaucoma detection is critical to prevent irreversible vision loss, yet existing AI methods often rely on unimodal inputs a

← Previous
1…4950515253…209
Next →