AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
9 Jul 2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

SafetyDGX agent

arXiv:2607.07675v1 Announce Type: new Abstract: Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For e

Scaling Quantum Machine Learning without Tricks: Full-Resolution and Diverse Image Generation

ResearchDGX agent

arXiv:2603.00233v2 Announce Type: replace-cross Abstract: Quantum generative modeling is a rapidly evolving discipline at the intersection of quantum computing and machine learning. Contemporary quant

Seeing What Matters: Lesion-Aware High-Resolution Patch Discovery and Fusion for Chest X-ray Report Generation

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.06909v1 Announce Type: new Abstract: Despite rapid advances in chest X-ray (CXR) foundation models, most radiology report generation (RRG) systems still rely on heavily downsampled inputs (

Segmenting Low-Contrast XCTs of Concrete: An Unsupervised Approach

Local AiDGX agent

arXiv:2603.00127v2 Announce Type: replace Abstract: X-Ray Computed Tomography (XCT) is a compelling tool in experimental mechanics, capable of non-destructively extracting information pertaining to th

SHTA: Semantic Hard Token Correction and Center Alignment for Semi-Supervised Medical Image Segmentation

SafetyDGX agent

arXiv:2607.07019v1 Announce Type: new Abstract: Recent advances in semi-supervised medical image segmentation have achieved remarkable performance through prediction consistency, pseudo-label supervis

Smart Scissor: Coupling Spatial Redundancy Reduction and CNN Compression for Embedded Hardware

Local AiDGX agent

arXiv:2607.06915v1 Announce Type: new Abstract: Scaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising fo

SoccerNet 2026 Challenges Results

ApplicationsDGX agent

arXiv:2607.07320v1 Announce Type: new Abstract: The SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision rese

SonoRank: Towards Calibration-Free Real-Time Finger Flexion Detection from Forearm Ultrasound Sequences

ResearchDGX agent

arXiv:2607.07542v1 Announce Type: cross Abstract: Powered prosthetic hands are frequently abandoned, largely due to the limited functionality of current devices that rely on surface electromyography (

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP

Local AiDGX agent

arXiv:2607.07135v1 Announce Type: new Abstract: Contrastive Language-Image Pre-training (CLIP) relies on softmax-based self-attention, a strictly positive distribution that assigns probability mass to

SpiS-GAN: Spiral-Modulated Handwriting Synthesis with Star Operation

ResearchDGX agent

arXiv:2607.06949v1 Announce Type: new Abstract: Training robust handwriting recognition (HTR) systems requires massive amounts of annotated data, which is often difficult to acquire. While synthetic h

Stage-Aware Adaptation and Distribution Calibration for Subject-Driven Personalized Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2607.07173v1 Announce Type: new Abstract: Subject-driven personalized text-to-image generation requires a pretrained diffusion model to acquire a specific subject from a few reference images whi

T^{3}S: Think in Thermal Time for Generalizable Crop Mapping from Satellite Image Time Series

Model ReleasesDGX agent

arXiv:2506.12885v4 Announce Type: replace Abstract: Crop type classification from optical satellite time series remains limited in its ability to generalize across growing seasons, particularly when c

TACoS: Weakly Supervised Learning of Two-Dimensional Materials from Scribble Annotations to Precise Segmentation

SafetyDGX agent

arXiv:2607.07169v1 Announce Type: new Abstract: The precise pixel-level localization of 2D material flakes is crucial for high-throughput screening. However, traditional fully supervised methods rely

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration

SafetyDGX agent

arXiv:2603.27742v2 Announce Type: replace Abstract: Vision-language agents that orchestrate specialized tools for image restoration (IR) have emerged as a promising method, yet most existing framework

Towards Accurate and Fast Clinical Body Composition: A Resource-Efficient Hierarchical Segmentation Framework for Multi-Source CT

HardwareDGX agent

arXiv:2607.07177v1 Announce Type: cross Abstract: Background: Automated 3D segmentation of muscles and adipose tissue from CT is vital for body composition analysis, but multi-source data heterogeneit

TRACE-Seg3D: Counterfactual Context Auditing For Robust 3D Glioma Segmentation Under Institutional Shift

Model ReleasesDGX agent

arXiv:2607.07038v1 Announce Type: new Abstract: Medical image segmentation models can achieve strong benchmark performance while remaining sensitive to scanner, protocol, and institutional variation.

Trexplorer Super: Topologically Correct Centerline Tree Tracking of Tubular Objects in CT Volumes

ResearchDGX agent

arXiv:2507.10881v2 Announce Type: replace Abstract: Tubular tree structures, such as blood vessels and airways, are essential in human anatomy and accurately tracking them while preserving their topol

Two-Stage Multi-Modal Fusion with Adaptive Alignment for Action Quality Assessment

Model ReleasesDGX agent

arXiv:2607.07438v1 Announce Type: new Abstract: Action Quality Assessment (AQA) aims to evaluate how well a person performs a movement, which is essential in applications such as sports scoring, skill

Unraveling Machine Behavior by Multi-Level Bias Analysis and Detection: Methodology and Application to Computer Vision

Model ReleasesDGX agent

arXiv:2607.07236v1 Announce Type: new Abstract: This study investigates the presence and propagation of bias within Neural Networks through a comprehensive multi-level analysis spanning the learned la

URS-Stereo: Uncertainty-Guided Residual Search for Real-Time Stereo Matching

Local AiDGX agent

arXiv:2607.06779v1 Announce Type: new Abstract: Real-time stereo matching is crucial for robotics, autonomous systems, and embedded vision applications, where both computational efficiency and dispari

VCDP: Variation-Conditioned Distributional Proxy Learning for Semi-Supervised Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.07416v1 Announce Type: new Abstract: Semi-supervised 3D medical image segmentation reduces the need for dense voxel-level annotations by exploiting unlabeled volumes. Although existing meth

VFM-Loc: Training-Free Cross-View Geo-Localization via Aligning Discriminative Visual Hierarchies

SafetyDGX agent

arXiv:2603.13855v2 Announce Type: replace Abstract: Cross-View Geo-Localization (CVGL) in remote sensing aims to locate a drone-view query by matching it to geo-tagged satellite images. Although super

Video-Based Detection of squint and cataract for accessibility-aware adaptive web interface rendering

ResearchDGX agent

arXiv:2607.07099v1 Announce Type: new Abstract: Squint and cataract are major ocular disorders that majorly affect visual perception and interaction capability. This paper proposes a real-time video-b

Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild

Model ReleasesDGX agent

arXiv:2607.06875v1 Announce Type: new Abstract: Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis

When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators

ResearchDGX agent

arXiv:2506.19348v2 Announce Type: replace Abstract: Training-free motion customization imposes motion patterns from reference videos onto video generators through test-time computation. Most existing

Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors

Local AiDGX agent

arXiv:2607.07216v1 Announce Type: new Abstract: Deepfake (DF) technology poses a significant threat to information integrity, driving the need for robust detection methods. Most DF detectors only cons

Widest-Path Reachability Fields for Connectivity-Preserving Slender Structure Segmentation

ResearchDGX agent

arXiv:2607.07123v1 Announce Type: new Abstract: Segmenting slender curvilinear structures such as retinal vessels, cracks, and roads demands topological correctness, as even a single-pixel discontinui

WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence

AgentsDGX agent

arXiv:2607.06838v1 Announce Type: new Abstract: Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilometers. Can AI build spatial represe

Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer

SafetyDGX agent

arXiv:2510.24108v2 Announce Type: replace-cross Abstract: Human demonstrations are widely considered the cornerstone of end-to-end (E2E) autonomous driving despite human demonstration's scarcity for l

8 Jul 2026

A Task-Driven Evaluation of UAV Detection and Tracking under Synthetic Fog

ResearchDGX agent

arXiv:2607.05467v1 Announce Type: new Abstract: Fog severely degrades the visibility of small unmanned aerial vehicles (UAVs) in skydominant, long-range imagery, reducing the reliability of downstream

A VLM-Enhanced Framework for Comprehensive Traffic Sign Condition Assessment Integrating Daytime Visual Performance and Nighttime Retroreflectivity Evaluation

Model ReleasesDGX agent

arXiv:2607.06478v1 Announce Type: new Abstract: Traffic signs are crucial components of road safety, serving as visual tools under all lighting conditions. The Manual on Uniform Traffic Control Device

Abductive Corroboration of Probabilistic AI Models for Forensic Synthetic Media Detection

ApplicationsDGX agent

arXiv:2607.05434v1 Announce Type: cross Abstract: Artificial Intelligence (AI) models, at their core, apply general learnings from broad datasets to individual circumstances using probabilistic behavi

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

Local AiDGX agent

arXiv:2607.06120v1 Announce Type: new Abstract: Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploi

AlayaWorld: Long-Horizon and Playable Video World Generation

ApplicationsDGX agent

arXiv:2607.06291v1 Announce Type: new Abstract: Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and e

Andha-Dhun: A First Look at Audio Descriptions in Hindi

ResearchDGX agent

arXiv:2607.06457v1 Announce Type: new Abstract: Audio Descriptions (ADs) narrate visual content for Blind and Low Vision (BLV) audiences during gaps in audiovisual media. There is growing momentum aro

Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation

ResearchDGX agent

arXiv:2607.01499v2 Announce Type: replace Abstract: Recent advances in Image-to-Video generation allow a single image to be animated into a convincing video under text guidance, raising serious copyri

ARMS: Anchor-Relational Motion Streaming for Seamless Solo-Social Motion Transitions

SafetyDGX agent

arXiv:2607.05733v1 Announce Type: new Abstract: Generating temporally continuous and socially coherent human motion from text remains a fundamental challenge, particularly in realistic streams where p

Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles

AgentsDGX agent

arXiv:2607.06484v1 Announce Type: cross Abstract: Poisoning attacks against public datasets lead to major concerns, such as (i) misclassification of perceived objects when the poisoned data is used fo

Association Restoration Test: Revealing Restorable Shortcuts after Unlearning

ResearchDGX agent

arXiv:2607.05726v1 Announce Type: new Abstract: Association unlearning aims to disable learned label-attribute shortcuts while preserving task performance. Existing evaluations mainly measure output-l

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring

Local AiDGX agent

arXiv:2607.05859v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are promising for construction-site monitoring, and recent construction-tailored VLMs have primarily adapted pretrained VL

Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective

Model ReleasesDGX agent

arXiv:2607.05783v1 Announce Type: new Abstract: Environmental illusions (eg., shadows, reflections, and tire marks) are naturally existing yet overlooked phenomena in real-world driving environments.

BitFair: A 12nm Bit-Serial CNN Accelerator with Learnable Early Termination and Adaptive Bit Ordering for Ultra-Low-Power XR Vision

ResearchDGX agent

arXiv:2607.05445v1 Announce Type: cross Abstract: Extended Reality (XR) wearables require always-on perception within tight power envelopes of a few watts and motion-to-photon latency budgets below 20

Blind Quality Enhancement of Compressed Video via Fine-Grained Degradation-Guided Sequential Inference

TutorialsDGX agent

arXiv:2511.16137v2 Announce Type: replace Abstract: Existing studies on quality enhancement for compressed video (QECV) predominantly rely on known quantization parameters (QPs), training separate enh

Breaking Spurious Correlations via Generative Randomization and Cross-Variant Self-Supervised Learning

TutorialsDGX agent

arXiv:2607.05850v1 Announce Type: new Abstract: Deep neural networks trained with Empirical Risk Minimization (ERM) often fail under distribution shifts because they exploit spurious correlations betw

Bridging Diffusion Pruning and Step Distillation with Teacher-Aligned Repair

SafetyDGX agent

arXiv:2607.06335v1 Announce Type: new Abstract: Diffusion models generate high-quality images, but their inference cost comes from two sources: large denoising networks and repeated denoising steps. E

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

Model ReleasesDGX agent

arXiv:2607.06534v1 Announce Type: new Abstract: Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the abilit

Clustered Codebook Quantization for 2D Gaussian-based Image Compression

Model ReleasesDGX agent

arXiv:2607.05667v1 Announce Type: new Abstract: Gaussian-based image representations effectively model image content using compact parametric primitives while preserving high visual fidelity, yet stor

Conformal Prediction Sets for Instance Segmentation

Model ReleasesDGX agent

arXiv:2602.10045v2 Announce Type: replace Abstract: Current instance segmentation models achieve high performance on average predictions, but lack principled uncertainty quantification: their outputs

Cross-Contextual Vision-Language Adaptation with LoRA for Personalized Severe Adverse Event Detection in Clinical Wound Monitoring

SafetyDGX agent

arXiv:2607.05625v1 Announce Type: new Abstract: Wound monitoring is a critical yet underserved clinical challenge, where timely identification of severe adverse events (SAEs) such as infection, tissue

DeSeG: Decoupling Semantic Intent and Geometric Constraints for Physically Plausible Human-Scene Interaction

SafetyDGX agent

arXiv:2607.05787v1 Announce Type: new Abstract: Synthesizing physically plausible human-scene interactions (HSI) remains a critical challenge in computer vision and the development of human avatars. A

EeveeDark: A Binary Neural Framework for Low-Light Video Enhancement via Event-Guided Sensor-Level Fusion

ApplicationsDGX agent

arXiv:2607.06217v1 Announce Type: new Abstract: Enhancing videos under extreme low-light conditions remains challenging due to the difficulty of balancing restoration quality and computational efficie

EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage

Model ReleasesDGX agent

arXiv:2607.06468v1 Announce Type: new Abstract: We introduce EgoPolice, a carefully curated dataset of real, egocentric police-civilian interactions, sourced from publicly available body-worn camera v

Enhanced Seam Segmentation for Automated Welding Robot in Construction Through Transfer Learning: Addressing Limitations of Bilateral Segmentation Network

Model ReleasesDGX agent

arXiv:2607.06150v1 Announce Type: new Abstract: Reliable seam segmentation is essential for autonomous robotic welding in construction, where harsh illumination, specular reflections, and thin weld ge

FADRA: Frequency-Aware Diffusion with Residual Adaptation for Video Face Restoration

SafetyDGX agent

arXiv:2607.06389v1 Announce Type: new Abstract: Video face restoration (VFR) aims to recover high-quality and temporally consistent facial details from severely degraded video sequences; however, exis

FedDAF: Federated Domain Adaptation Using Model Functional Distance

ApplicationsDGX agent

arXiv:2509.11819v2 Announce Type: replace-cross Abstract: Federated Domain Adaptation (FDA) is a federated learning (FL) approach that improves model performance at the target client by collaborating

FGAA-FPN: Foreground-Guided Angle-Aware Feature Pyramid Network for Oriented Object Detection

TutorialsDGX agent

arXiv:2602.10710v2 Announce Type: replace Abstract: With the increasing availability of high-resolution remote sensing and aerial imagery, oriented object detection has become a key capability for geo

FIELDS: Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision

ResearchDGX agent

arXiv:2511.21245v3 Announce Type: replace Abstract: Monocular 3D face reconstruction estimates a 3D morphable model (3DMM) representation from a single image, providing geometry-aware expression codes

FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models

Model ReleasesDGX agent

arXiv:2607.05711v1 Announce Type: cross Abstract: Diffusion models have become a dominant paradigm for high-quality generative modeling, while post-training is essential for adapting them to diverse d

Freqformer: Image-Demoireing Transformer via Effective Frequency Decomposition

Local AiDGX agent

arXiv:2505.19120v2 Announce Type: replace Abstract: Image demoireing remains a challenging task due to the complex interplay between texture corruption and color distortions caused by moire patterns.

From Pixels to Portraits: A Comprehensive Survey of Talking Head Generation Techniques and Applications

SafetyDGX agent

arXiv:2308.16041v2 Announce Type: replace Abstract: Talking head generation has progressed rapidly from landmark- and GAN-based facial animation to diffusion models, neural rendering, 3D-aware avatars

← Previous
1…4546474849…209
Next →