AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
30 Jul 2026

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

SafetyDGX agent

arXiv:2607.26170v1 Announce Type: new Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting imag

Anatomy Contextualized Adaption of CT Foundation Models

SafetyDGX agent

arXiv:2607.27154v1 Announce Type: new Abstract: CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume repres

Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.26647v1 Announce Type: new Abstract: While text-to-image diffusion models achieve impressive visual quality, they frequently struggle to maintain precise alignment with complex compositiona

BATS: Resource-Efficient Volumetric Segmentation with Boundary-Aware Mixed-Resolution Tokens

HardwareDGX agent

arXiv:2607.26829v1 Announce Type: new Abstract: Many high-performing volumetric segmentation models maintain dense multi-scale feature maps, leading to high activation memory and inference cost. We pr

BG-REAL: A Public Real-Data Anchored Benchmark for Background Manipulation Detection and Localization

Model ReleasesDGX agent

arXiv:2607.26232v1 Announce Type: new Abstract: Background manipulation is a practical but under-specified image-forensics setting: the manipulated evidence can sit outside the salient foreground obje

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

Model ReleasesDGX agent

arXiv:2606.19651v2 Announce Type: replace-cross Abstract: Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represented

Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization

TutorialsDGX agent

arXiv:2511.07210v3 Announce Type: replace Abstract: Clean-image backdoor attacks, which use only label manipulation in training datasets to compromise deep neural networks, pose a significant threat t

Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration

Model ReleasesDGX agent

arXiv:2603.24800v2 Announce Type: replace Abstract: In this paper, we uncover the hidden potential of Diffusion Transformers (DiTs) to significantly enhance generative tasks. Through an in-depth analy

CASIAL: Geometric Distortion Robust Image Watermarking

SafetyDGX agent

arXiv:2607.26729v1 Announce Type: new Abstract: Deep learning-based watermarking has shown strong robustness against non-geometric distortions, yet its performance under geometric transformations rema

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

SafetyDGX agent

arXiv:2607.26452v1 Announce Type: cross Abstract: World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually

CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents

SafetyDGX agent

arXiv:2607.26910v1 Announce Type: new Abstract: Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

SafetyDGX agent

arXiv:2607.26529v1 Announce Type: new Abstract: Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained contr

Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models

ResearchDGX agent

arXiv:2607.26580v1 Announce Type: new Abstract: With the increase in the number of cases related to respiratory diseases, there is an urgent need to detect them early and diagnose them accurately. Con

Clinical Graph-Mediated Distillation for Unpaired MRI-to-CFI Hypertension Prediction

ResearchDGX agent

arXiv:2603.21809v2 Announce Type: replace Abstract: Retinal fundus imaging enables low-cost and scalable hypertension (HTN) screening, but HTN-related retinal cues are subtle, yielding high-variance p

Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck Cancer

ResearchDGX agent

arXiv:2607.26276v1 Announce Type: new Abstract: Background: Early prediction of distant metastasis (DM) risk in head and neck cancer (HNC) can enable timely interventions that may improve treatment ou

ContactFlow: A video action conditioning that transfers across embodiments

ApplicationsDGX agent

arXiv:2607.26579v1 Announce Type: cross Abstract: World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. Howe

Context-measure: Contextualizing Metric for Camouflage

ResearchDGX agent

arXiv:2512.07076v4 Announce Type: replace Abstract: Camouflage relies heavily on context, but current metrics used in camouflaged object segmentation ignore contextual cues. We identify two major draw

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution

Model ReleasesDGX agent

arXiv:2607.26596v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified tran

Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge

ResearchDGX agent

arXiv:2603.07131v4 Announce Type: replace Abstract: Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis. However, their clinical deployment is severely hinde

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

SafetyDGX agent

arXiv:2607.26811v1 Announce Type: new Abstract: Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they t

DLAM: Distributional Latent Actions with Temporal Constraints

Local AiDGX agent

arXiv:2607.27138v1 Announce Type: cross Abstract: Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

SafetyDGX agent

arXiv:2607.26411v1 Announce Type: new Abstract: Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabi

Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

SafetyDGX agent

arXiv:2607.26735v1 Announce Type: new Abstract: Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target images without

DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving

AgentsDGX agent

arXiv:2607.26165v1 Announce Type: new Abstract: Safe autonomous navigation requires a holistic understanding of dynamic environments, necessitating the simultaneous estimation of metric depth, semanti

Eddeep: a deep-learning framework for fast eddy-current distortion correction in diffusion MRI

SafetyDGX agent

arXiv:2607.26292v1 Announce Type: new Abstract: Diffusion MRI (dMRI) relies on diffusion-weighted echo-planar imaging, which is highly susceptible to eddy-current-induced geometric distortions. These

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

Model ReleasesDGX agent

arXiv:2607.26518v1 Announce Type: new Abstract: Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic unc

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

SafetyDGX agent

arXiv:2607.27145v1 Announce Type: new Abstract: As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitorin

FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking

Local AiDGX agent

arXiv:2607.26641v1 Announce Type: new Abstract: Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Generative AI m

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

Model ReleasesDGX agent

arXiv:2607.26432v1 Announce Type: new Abstract: Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence f

FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows

SafetyDGX agent

arXiv:2607.26645v1 Announce Type: new Abstract: Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point clouds are cons

FreeShadow: Training-Free Shadow Removal via Illumination Transfer and Selective Content Preservation in Diffusion Models

ResearchDGX agent

arXiv:2607.26715v1 Announce Type: new Abstract: Existing supervised and unsupervised shadow removal methods often suffer from limited generalization due to the insufficient diversity of available trai

FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring

ResearchDGX agent

arXiv:2607.27110v1 Announce Type: new Abstract: Autoregressive video diffusion models enable real-time streaming video generation. However, errors introduced during self-rollout accumulate over long h

From Keypoints to Predictive Distributions: Post-Hoc Uncertainty for YOLO-Pose Models

Local AiDGX agent

arXiv:2607.26921v1 Announce Type: new Abstract: YOLO-Pose models provide efficient keypoint localization, but do not quantify the associated spatial uncertainty. We introduce a lightweight post-hoc pr

From Spatial Semantics to Temporal Context: Leveraging Gaze Trajectory for Weakly Supervised Medical Image Segmentation

ResearchDGX agent

arXiv:2607.26542v1 Announce Type: new Abstract: Medical image segmentation heavily depends on labor-intensive and time-consuming pixel-level annotations. Eye tracking offers a cost-effective solution

From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching

Local AiDGX agent

arXiv:2607.26817v1 Announce Type: cross Abstract: Visual Floorplan Localization (FLoc) has emerged as a promising solution for indoor localization by matching egocentric images against minimalist stru

Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation

HardwareDGX agent

arXiv:2607.26646v1 Announce Type: new Abstract: We address the problem of reconstructing a high-fidelity, freely navigable 3D scene from a single 360^irc panorama, without per-scene optimization or mu

HERMES: A Hybrid Ensemble for Head-and-Neck Tumor Segmentation, TN Staging, and Recurrence-Free Survival on PET/CT

ResearchDGX agent

arXiv:2607.26498v1 Announce Type: new Abstract: We present HERMES (Hybrid Ensemble for Radiotherapy-target segmentation, Malignancy staging, and Event-free Survival), a single containerized algorithm

HeteroPROPMT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework

AgentsDGX agent

arXiv:2607.26283v1 Announce Type: new Abstract: Collaborative Perception (CP) improves autonomous systems' awareness of their surroundings by sharing sensor data, intermediate features, and detection

HumanCLAW: Can Vision-Language Models Act Through a Body?

Model ReleasesDGX agent

arXiv:2607.27180v1 Announce Type: new Abstract: Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision wit

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

Model ReleasesDGX agent

arXiv:2607.26848v1 Announce Type: new Abstract: Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meani

Improving Knowledge Distillation Under Unknown Covariate Shift Through Confidence-Guided Data Augmentation

ResearchDGX agent

arXiv:2506.02294v3 Announce Type: replace Abstract: Large foundation models trained on extensive datasets demonstrate strong zero-shot capabilities in various domains. Knowledge distillation has becom

InkShield: Writing Style Protection Against Unauthorized Handwriting Mimicry

ResearchDGX agent

arXiv:2607.26976v1 Announce Type: cross Abstract: Recent handwritten text generators can reproduce a writer's style from publicly available references, posing risks of document forgery and identity mi

Interpretable Image-Level Acne Severity Grading via EfficientNet-B0 Transfer Learning and Grad-CAM

ResearchDGX agent

arXiv:2607.26461v1 Announce Type: new Abstract: Acne vulgaris affects most adolescents and many adults. Accurate severity grading guides treatment, monitoring, and clinical trial endpoints, but manual

JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2607.26600v1 Announce Type: new Abstract: Self-supervised monocular depth estimation typically relies on photometric reconstruction losses that couple depth, pose, and appearance assumptions. In

Kinetic Mining in Context: Few-Shot Action Synthesis via Text-to-Motion Distillation

ResearchDGX agent

arXiv:2512.11654v3 Announce Type: replace Abstract: The acquisition cost for large, annotated motion datasets remains a critical bottleneck for skeletal-based Human Activity Recognition (HAR). Althoug

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition

ResearchDGX agent

arXiv:2607.26097v1 Announce Type: new Abstract: Action recognition in complex scenes often involves multiple concurrent fine-grained actions, making it challenging to model internal action structures.

Lag-aware cross-hand alignment for dual-hand action segmentation

SafetyDGX agent

arXiv:2607.26215v1 Announce Type: new Abstract: Dual-hand action segmentation commonly fuses left- and right-hand representations at identical temporal indices, although coordinated hand transitions m

Level, Sharpness, and Corpus: Why Zero-Shot OOD Detector Rankings Do Not Transfer

Model ReleasesDGX agent

arXiv:2607.26582v1 Announce Type: new Abstract: Selecting a zero-shot out-of-distribution (OOD) detector for a new deployment is typically based on benchmark rankings, implicitly assuming that the hig

LiDARDraft: Generating LiDAR Point Cloud from Versatile Inputs

SafetyDGX agent

arXiv:2512.20105v2 Announce Type: replace Abstract: Generating realistic and diverse LiDAR point clouds is crucial for autonomous driving simulation. Although previous methods achieve LiDAR point clou

Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment

HardwareDGX agent

arXiv:2607.26238v1 Announce Type: new Abstract: We investigate lightweight raptor-species classification for real-time edge deployment in wind-turbine collision mitigation. Using DINOv2-L (304M parame

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency

ResearchDGX agent

arXiv:2512.01008v2 Announce Type: replace Abstract: Text-driven 3D reconstruction requires masks that understand free-form instructions and remain stable under viewpoint changes. We present LISA-3D, a

LLM-Grounded Dynamic Task Planning with Hierarchical Temporal Logic for Human-Aware Multi-Robot Handover

ResearchDGX agent

arXiv:2602.09472v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) enable non-experts to specify open-world multi-robot tasks, but the generated plans are often kinematically infea

Long-Tailed 3D Point Cloud Dataset Distillation

SafetyDGX agent

arXiv:2607.26763v1 Announce Type: new Abstract: Dataset distillation compresses large-scale datasets into compact synthetic sets while preserving their training utility, enabling efficient 3D point cl

Lottery Tickets Are Not Deployment Tickets

SafetyDGX agent

arXiv:2607.27031v1 Announce Type: cross Abstract: Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects

LumaGuide: Distribution Shaping for Training-Free HDR Generation in Diffusion Models

ResearchDGX agent

arXiv:2607.26237v1 Announce Type: new Abstract: Pretrained diffusion models generate realistic images but are constrained by the statistical biases of their training data, limiting their ability to pr

MedARC: Training-Free Adaptive Redundancy Compression of Visual Tokens for 3D Medical Vision-Language Models

Local AiDGX agent

arXiv:2607.26554v1 Announce Type: new Abstract: Integrating 3D medical images with vision-language models (VLMs) holds substantial promise for computer-aided diagnosis. However, volumetric images gene

Mitigating Compounding Error via Video Representation Regularization

AgentsDGX agent

arXiv:2607.27036v1 Announce Type: new Abstract: Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window

MoSAIC: Aligned Intervention Supervision for Part-Local Motion Style Transfer

Local AiDGX agent

arXiv:2607.26304v1 Announce Type: new Abstract: Editing character motion often requires transferring a gesture or gait from one or more reference motions while preserving the source action, timing, ro

Multimodal fusion of visual and morphometric features for avian bone classification

ResearchDGX agent

arXiv:2607.26743v1 Announce Type: new Abstract: Artificial intelligence has shown considerable potential for archaeological applications, yet its use in zooarchaeology remains limited, particularly fo

Neural Network Assisted Lifting Steps For Improved Fully Scalable Lossy Image Compression in JPEG 2000

ResearchDGX agent

arXiv:2403.01647v2 Announce Type: replace Abstract: This work proposes to augment the lifting steps of the conventional wavelet transform with additional neural network assisted lifting steps. These a

← Previous
1…2526272829…209
Next →