AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
9 Jun 2026

Echo-DM: Ultrasound Marker Removal via Conditional Latent Diffusion and Region-Aware Fusion

Model ReleasesDGX agent

arXiv:2606.09378v1 Announce Type: new Abstract: Clinical ultrasound images often contain artificial markers, such as measurement calipers and text, to assist diagnostic interpretation and comparison.

Echo-Memory: A Controlled Study of Memory in Action World Models

ResearchDGX agent

arXiv:2606.09803v1 Announce Type: new Abstract: We present extbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models. These models generate multi-segment videos fr

Edge-Constrained UAV Small-Object Detection with P2 Enhancement and Quantum-Inspired Lightweight Structure Search

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.09081v1 Announce Type: new Abstract: Unmanned aerial vehicle (UAV) object detection requires compact detectors that retain small-object details under onboard computation and memory constrai

EditSSC: Toward Editable Semantic Occupancy Scenes with Unconditional Diffusion Models

AgentsDGX agent

arXiv:2606.09273v1 Announce Type: new Abstract: 3D semantic scene generation is crucial for autonomous driving applications, yet most methods rely on complex 3D-specific architectures such as triplane

Efficient Minimal Solvers for Relative Pose Estimation in Autonomous Driving Applications

Model ReleasesDGX agent

arXiv:2606.09569v1 Announce Type: cross Abstract: With the advancement of visual sensing systems, computer vision is playing an increasingly important role in autonomous driving and robot navigation.

Efficient Minimal Solvers for Visual-Inertial Relative Pose Estimation in Multi-Camera Systems

Model ReleasesDGX agent

arXiv:2606.09477v1 Announce Type: new Abstract: Estimating the relative poses of multi-camera systems is a fundamental problem in computer vision, with critical applications in autonomous vehicles, mo

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

ResearchDGX agent

arXiv:2606.08495v1 Announce Type: cross Abstract: Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified traje

Embedded Graph Convolutional Networks for Real-Time Event Data Processing on SoC FPGAs

ResearchDGX agent

arXiv:2406.07318v3 Announce Type: replace Abstract: The utilisation of event cameras represents an important and swiftly evolving trend aimed at addressing the constraints of traditional video systems

Empowering Feed-Forward Reconstruction Models with Metric Scale via Satellite Images

Local AiDGX agent

arXiv:2606.08205v1 Announce Type: new Abstract: Feed-forward 3D reconstruction models have recently shown strong generalization across diverse scenes, yet most of them recover geometry only up to an u

End-to-End Optimization of Incoherent Imaging for Classification Under Detector-Limited Readout

ResearchDGX agent

arXiv:2606.09792v1 Announce Type: new Abstract: End-to-end co-optimization of optical front-ends (e.g. metasurfaces) and neural network back-ends has been widely applied to imaging tasks, yet a formal

Enhanced Detection of Tiny Objects in Aerial Images

Model ReleasesDGX agent

arXiv:2509.17078v3 Announce Type: replace Abstract: While one-stage detectors like YOLOv8 offer fast training speed, they often under-perform on detecting small objects as a trade-off. This becomes ev

Enhancing Adversarial Robustness with Signed Distance Fields for Harmonizing Geometric Invariance and Texture

Local AiDGX agent

arXiv:2602.05175v2 Announce Type: replace Abstract: Deep neural networks demonstrate impressive performance in visual recognition but remain highly vulnerable to imperceptible adversarial attacks. Exi

EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation

ResearchDGX agent

arXiv:2606.08980v1 Announce Type: new Abstract: This paper introduces EPS3D, a new end-to-end feed-forward framework for open-vocabulary 3D panoptic segmentation. Unlike existing methods relying on ad

Evaluating the Representation Space of Diffusion Models via Self-Supervised Principles

ResearchDGX agent

arXiv:2606.09718v1 Announce Type: cross Abstract: Diffusion models have demonstrated remarkable generative capabilities and have also emerged as powerful self-supervised representation learners, yet t

Event-driven dynamic trajectories reconstruction and measurement of mechanical parameters for fragments

Local AiDGX agent

arXiv:2606.09208v1 Announce Type: new Abstract: During warhead detonation, high-density, high-speed, and mutually occluded fragments are generated. Their mechanical parameters (position, velocity, kin

ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification

ResearchDGX agent

arXiv:2606.09360v1 Announce Type: new Abstract: Open-domain open-vocabulary detection (ODOVD) requires detectors to generalize to both novel categories and unseen domains, making it more challenging t

Facial Expression Recognition in the Deep Learning Era: A Systematic Multi-Criteria Review of Methods, Models, Datasets, Performance, Challenges, and Future Research Directions

ResearchDGX agent

arXiv:2606.08612v1 Announce Type: new Abstract: Facial Expression Recognition (FER) has advanced rapidly over the last decade, driven by the shift from handcrafted descriptors and shallow classifiers

FADRW: A Feature-Aware Modulated and Dynamically Reweighted Loss for Few-Shot Linguistic Steganalysis

SafetyDGX agent

arXiv:2606.07655v1 Announce Type: cross Abstract: The ubiquity of social media platforms facilitates malicious linguistic steganography, posing significant security risks. However, detection is severe

Feasibility to detect rapid change and disappearance of seagrass: Lessons from nearly 80 years of vegetation change in the Ako, Seto Inland Sea, Japan

ResearchDGX agent

arXiv:2606.07949v1 Announce Type: cross Abstract: This study analyses the Ako tidal flat in the Seto Inland Sea, Japan, where nearly all Zostera marina disappeared within a single year in 2025. Using

FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching

SafetyDGX agent

arXiv:2601.05212v2 Announce Type: replace Abstract: Brain Magnetic Resonance Imaging (MRI) plays a central role in studying neurological development, aging, and diseases. One key application is Brain

FMRFusion: Frequency-Aware Multi-View Representation Learning for Heterogeneous Image Fusion

Model ReleasesDGX agent

arXiv:2606.07985v1 Announce Type: new Abstract: Infrared and visible image fusion aims to generate a composite image that retains significant target information and preserves detailed textures, integr

Frankenstein in the Pipeline: Computational Epistemicide in Facial Recognition

SafetyDGX agent

arXiv:2606.07628v1 Announce Type: cross Abstract: While the eugenic roots of computer vision are well-documented in critical technology studies, less attention has been paid to the operational mechani

Frequency Decoupled Framework for Screen Content Image Super-Resolution

ResearchDGX agent

arXiv:2606.09029v1 Announce Type: new Abstract: Methods based on implicit neural representations have demonstrated superior performance in Screen Content Image Super-Resolution (SCISR) . However, they

Frequency-Scale Saliency for Spectral Descriptor Analysis in 3D Shape Retrieval

ResearchDGX agent

arXiv:2606.07791v1 Announce Type: cross Abstract: Classical spectral descriptors such as the Heat Kernel Signature and Wave Kernel Signature are widely used for non-rigid 3D shape retrieval, yet their

Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking

ResearchDGX agent

arXiv:2603.27493v2 Announce Type: replace Abstract: Spiking Neural Networks (SNNs), characterized by their event-driven computation and low power consumption, have shown great potential for energy-eff

G2G: Exploiting Intra-Group Geometry for Inter-Group Pose Estimation

ApplicationsDGX agent

arXiv:2606.08284v1 Announce Type: new Abstract: Recovering the relative 6-DoF pose between two image groups underlies cross-sequence relocalization and multi-camera rig odometry. Each group carries kn

GD-MIL: Grade-Disentangled Multiple Instance Learning for Multimodal Biochemical Recurrence Prediction in Prostate Cancer

Model ReleasesDGX agent

arXiv:2606.09453v1 Announce Type: new Abstract: Biochemical recurrence (BCR) after radical prostatectomy is a critical endpoint in prostate cancer, yet risk stratification relies almost entirely on va

Generalizing Geometry-Guided Mamba as a Plug-and-Play Context Module for CNN-based Semantic Segmentation

ResearchDGX agent

arXiv:2606.08866v1 Announce Type: new Abstract: CNN-based semantic segmentation networks usually rely on context heads such as ASPP, PPM, or attention modules to enlarge the receptive field. These hea

GenEyePose: Patient-Free, Knowledge-Based Saccadic Eye Movement Modeling for Digital Neurophysiologic Biomarker Development

Local AiDGX agent

arXiv:2606.09681v1 Announce Type: new Abstract: Eye movements, including saccades, are widely regarded as highly sensitive and objective biomarkers of neurophysiologic states. Detecting saccadic signa

Geometric Analysis of Magnetic Labyrinthine Stripe Evolution via Deep Learning Segmentation

ResearchDGX agent

arXiv:2509.11485v3 Announce Type: replace-cross Abstract: Labyrinthine stripe patterns are common in many physical systems, yet their lack of long-range order makes quantitative characterization chall

Geometry-Aware Fisheye-LiDAR Fusion for Robust 3D Object Detection in Low-Overlap Setups

AgentsDGX agent

arXiv:2606.08844v1 Announce Type: new Abstract: As autonomous systems expand from capital-intensive robotaxis to cost-sensitive logistics, sensor configurations are increasingly optimized for coverage

Geometry-Driven Flow Analysis of Brain Sulcal Pattern

ResearchDGX agent

arXiv:2606.08404v1 Announce Type: new Abstract: Cortical folding reflects coordinated neurodevelopmental processes and is increasingly recognized as a sensitive marker of neurological disease. However

GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization

ApplicationsDGX agent

arXiv:2601.18585v2 Announce Type: replace Abstract: Fine-tuning-based adaptation is widely used to customize diffusion-based image generation, leading to large collections of community-created adapter

GraspFoM: Towards Reconstruction-Driven Robotic Grasping with 3D Foundation Priors

ResearchDGX agent

arXiv:2606.08440v1 Announce Type: cross Abstract: Robotic grasping is a fundamental capability in robotic manipulation. Yet grasping remains challenging under partial observations. Reliable grasping d

Gravity-guided Contact Dynamics Estimation from 3D Human Motions

ResearchDGX agent

arXiv:2606.08133v1 Announce Type: new Abstract: Ground contact forces acting on the human body, are crucial for biomechanics studies or sport performance analysis. Prior methods rely on force plates o

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling

ResearchDGX agent

arXiv:2606.08302v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality generation with substantially fewer decoding steps. How

Harnessing Streaming Video in the Wild

Model ReleasesDGX agent

arXiv:2606.08615v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentar

HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging

SafetyDGX agent

arXiv:2606.09110v1 Announce Type: new Abstract: Most existing multi-exposure HDR methods follow a fixed feed-forward reconstruction paradigm, making them prone to ghosting artifacts in complex dynamic

HDSL: A Hierarchical Domain-Specific Language for Structured 3D Indoor Scene Generation and Localized Editing with LLM Agents

Model ReleasesDGX agent

arXiv:2606.09738v1 Announce Type: new Abstract: Text-driven indoor scene generation and editing require an intermediate representation that language models can both produce and revise. Existing LLM-ba

HiMat: DiT-based Ultra-High Resolution SVBRDF Generation

SafetyDGX agent

arXiv:2508.07011v5 Announce Type: replace Abstract: Creating ultra-high-resolution spatially varying bidirectional reflectance functions (SVBRDFs) is critical for photorealistic 3D content creation, t

How Much MRI Preprocessing Is Enough? A Cost-Utility Study for Brain MRI Foundation Models

ResearchDGX agent

arXiv:2606.08164v1 Announce Type: new Abstract: MRI preprocessing defines the input distribution seen by brain MRI foundation models, yet it is usually treated as routine data cleaning rather than a m

Hummus: A Dataset of Humorous Multimodal Metaphor Use

ResearchDGX agent

arXiv:2504.02983v3 Announce Type: replace-cross Abstract: Metaphor and humor share a lot of common ground, and metaphor is one of the most common humorous mechanisms. This study focuses on the humorou

Hyperspectral Smoke Segmentation via Mixture of Prototypes

SafetyDGX agent

arXiv:2602.10858v2 Announce Type: replace Abstract: Smoke segmentation is critical for wildfire management and industrial safety applications. Traditional visible-light-based methods face limitations

IB-HFN: Information Bottleneck-Driven SAR-Optical Fusion Network for High-Fidelity Cloud Removal

ResearchDGX agent

arXiv:2606.09347v1 Announce Type: new Abstract: Synthetic aperture radar (SAR)-assisted optical cloud removal aims to recover surface information obscured by clouds in optical remote sensing images by

IDDM: Identity-Decoupled Personalized Diffusion Models with a Tunable Privacy-Utility Trade-off

Model ReleasesDGX agent

arXiv:2604.00903v2 Announce Type: replace Abstract: Personalized text-to-image diffusion models (e.g., DreamBooth, LoRA) enable users to synthesize high-fidelity avatars from a few reference photos fo

IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation

Model ReleasesDGX agent

arXiv:2601.04498v2 Announce Type: replace-cross Abstract: Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information

Illumination-Invariant Anomaly Detection for Sub-Canopy UAV Multispectral Point Clouds

ResearchDGX agent

arXiv:2606.09111v1 Announce Type: new Abstract: Unmanned Aerial Vehicle (UAV) multispectral point clouds (MPC) provide high-dimensional spatial-spectral data for sub-canopy target detection; however,

iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

ApplicationsDGX agent

arXiv:2606.09813v1 Announce Type: cross Abstract: Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environment simulation. However, conventio

IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval

ResearchDGX agent

arXiv:2606.08144v1 Announce Type: new Abstract: Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing met

KITE: A Tri-Modal Transformer Integrating Text, Images, and Knowledge Graphs for Fake News Detection

Model ReleasesDGX agent

arXiv:2606.07651v1 Announce Type: cross Abstract: Traditional fake news detection methods are falling behind as multimodal misinformation grows more advanced, seamlessly blending deceptive text, manip

Latent Spatial Memory for Video World Models

ResearchDGX agent

arXiv:2606.09828v1 Announce Type: new Abstract: Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space.

Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning

TutorialsDGX agent

arXiv:2606.08641v1 Announce Type: new Abstract: The processing of gigapixel whole slide images within vision language models faces a major difficulty due to an excessive number of visual tokens. Exist

Learning a Semantic Calibration Network for Open-Vocabulary Semantic Segmentation

ResearchDGX agent

arXiv:2606.08001v1 Announce Type: new Abstract: Semantic image segmentation assigns a predefined category label to each pixel, has achieved significant progress lately. Open-Vocabulary Segmentation (O

Learning to Solve Generative ODEs Beyond the Linear Span

ResearchDGX agent

arXiv:2606.08672v1 Announce Type: new Abstract: Diffusion and flow generative models sample by integrating a learned ODE, but high quality still requires many sequential model evaluations. Solver lear

LEGS: Laplacian-Enhanced Gaussian Splatting with a Nonlinear Weighted Loss

ResearchDGX agent

arXiv:2606.07932v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has become an efficient explicit representation for radiance field reconstruction and real-time novel view synthesis. Howev

Less Is More: Training-Free Acceleration Framework of 3D Diffusion Models for Low-Count PET Denoising via Global-Local Trajectory Reduction

Local AiDGX agent

arXiv:2606.08751v1 Announce Type: new Abstract: Accurate quantification and uptake measurement in PET are critical for assessing disease progression and supporting clinical decision-making. While high

Leveraging Morphology for Historical Script Metrological Analysis

TutorialsDGX agent

arXiv:2606.09446v1 Announce Type: new Abstract: Advances in handwritten text recognition have enabled large-scale transcription of historical documents, but still provide limited access to interpretab

Leveraging NeRF-Rendered Images for 3D Gaussian Splatting

ResearchDGX agent

arXiv:2606.09034v1 Announce Type: new Abstract: Neural radiance field (NeRF) and 3D Gaussian splatting (3DGS) are two mainstream approaches for novel view synthesis. They often show complementary perf

Light-WAM: Efficient World Action Models with State-Fusion Action Decoding

SafetyDGX agent

arXiv:2606.08242v1 Announce Type: new Abstract: World Action Models (WAMs) extend robot policy learning by incorporating future prediction as an additional training objective, encouraging the policy t

LiteVSR: Lightweight Adaptation of Frozen Diffusion Transformers for Video Super-Resolution

SafetyDGX agent

arXiv:2606.09250v1 Announce Type: new Abstract: Adapting large-scale pre-trained video generators for Video Super-Resolution (VSR) in novel domains remains computationally prohibitive. Methods that re

← Previous
1…8889909192…211
Next →