AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

Screening Is Effective for Visual Recognition

DGX agent

arXiv:2607.13983v1 Announce Type: new Abstract: Vision Transformer (ViT) has been widely used as a powerful framework for modeling global dependencies among image patches. However, its core component,

researcharxiv-cs-cv
16 Jul 2026
Model Releases

Self-Supervised Visual Representation Learning: Pretrain-Finetuning or Joint Training?

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2607.13192v1 Announce Type: new Abstract: Self-supervision is a powerful technique for learning visual representations from unlabeled data. Existing techniques primarily adopt a two-stage approa

model-releasesarxiv-cs-cv
16 Jul 2026
Safety

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

DGX agent

arXiv:2607.13931v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-languag

safetyarxiv-cs-cv
16 Jul 2026
Research

T3HG-Editor: Text-driven 3D Human Garment Editing with Body Priors Embedded in SMPL-X

DGX agent

arXiv:2607.13654v1 Announce Type: new Abstract: While 3D Gaussian Editing (3DGE) has seen substantial progress, text-driven 3D human garment editing remains largely underexplored. Existing 3DGE works

researcharxiv-cs-cv
16 Jul 2026
Research

Tactile Modality Fusion for Vision-Language-Action Models

DGX agent

arXiv:2603.14604v2 Announce Type: replace-cross Abstract: We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. Wh

researcharxiv-cs-cv
16 Jul 2026
Research

Task-Specific Feature Fusion Method for Multi-Task Affective Behavior Analysis

DGX agent

arXiv:2607.13986v1 Announce Type: new Abstract: The 11th Affective Behavior Analysis in-the-wild (ABAW11) Multi-Task Learning Challenge requires a unified system to predict valence-arousal, categorica

researcharxiv-cs-cv
16 Jul 2026
Tutorials

TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model

DGX agent

arXiv:2607.13812v1 Announce Type: cross Abstract: We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data.

tutorialsarxiv-cs-cv
16 Jul 2026
Hardware

Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation

DGX agent

arXiv:2607.13164v1 Announce Type: cross Abstract: Sign language is a primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly b

hardwarearxiv-cs-cv
16 Jul 2026
Research

The 2nd International StepUP Competition for Biometric Footstep Recognition: From Steps to Strides

DGX agent

arXiv:2607.13905v1 Announce Type: new Abstract: The International StepUP Competition Series was launched to advance research in pressure-based footstep biometrics through a standardized and challengin

researcharxiv-cs-cv
16 Jul 2026
Safety

ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

DGX agent

arXiv:2607.13539v1 Announce Type: new Abstract: While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based genera

safetyarxiv-cs-cv
16 Jul 2026
Model Releases

Thresholded Cross-Attention for Reliable Intensity-Chromaticity Fusion in Low-Light Image Enhancement

DGX agent

arXiv:2607.13925v1 Announce Type: new Abstract: Low-Light Image Enhancement (LLIE) requires a careful balance among noise suppression, color fidelity, and efficiency. Recent HVI-based methods alleviat

model-releasesarxiv-cs-cv
16 Jul 2026
Applications

Towards a Modular Bin-picking Framework for Handling Object Pose Uncertainties

DGX agent

arXiv:2607.13698v1 Announce Type: cross Abstract: In recent years, there has been growing interest in robust robotic systems for precise bin-picking applications. To achieve reliable performance, such

applicationsarxiv-cs-cv
16 Jul 2026
Research

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

DGX agent

arXiv:2607.13860v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in 2D medical image understanding, their extension to 3D volumetric

researcharxiv-cs-cv
16 Jul 2026
Model Releases

Towards Spatial Supersensing in the Wild

DGX agent

arXiv:2607.13681v1 Announce Type: new Abstract: Humans can efficiently parse continuous sensory streams, from hours to years, scaffolding an internal world model that grounds spatial reasoning and pre

model-releasesarxiv-cs-cv
16 Jul 2026
Research

TRACE-PCa: Predicting Prostate Cancer Progression from Longitudinal MRI During Active Surveillance

DGX agent

arXiv:2607.13506v1 Announce Type: new Abstract: Active surveillance (AS) is the preferred strategy for favorable-risk prostate cancer, yet current protocols rely on scheduled repeat biopsies, most of

researcharxiv-cs-cv
16 Jul 2026
Safety

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation

DGX agent

arXiv:2512.10607v2 Announce Type: replace Abstract: We present TCAM (Track and Caption Any Motion), a generative framework that watches a video and with no text query and no region prompt decides what

safetyarxiv-cs-cv
16 Jul 2026
Tutorials

TreeSRNF: Square-Root Normal Fields for Generative Modelling of the Geometric and Structural Variability in Tree-like 3D Objects

DGX agent

arXiv:2607.13456v1 Announce Type: new Abstract: We introduce a novel mathematical framework for analyzing and generating complex tree-shaped 3D objects, such as botanical trees and plants, which defor

tutorialsarxiv-cs-cv
16 Jul 2026
Model Releases

UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets

DGX agent

arXiv:2607.13586v1 Announce Type: new Abstract: Physically grounded 3D assets are increasingly important for embodied AI and robotic simulation. However, most existing 3D assets lack unified physical

model-releasesarxiv-cs-cv
16 Jul 2026
Model Releases

Unsupervised Detection of Entry and Exit Regions from Vehicle Trajectories for Camera-Agnostic Turning Movement Counts

DGX agent

arXiv:2607.10949v2 Announce Type: replace Abstract: Turning movement counts are essential for intersection-level traffic management, yet their collection remains predominantly manual due to the cost o

model-releasesarxiv-cs-cv
16 Jul 2026
Model Releases

VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation

DGX agent

arXiv:2607.13527v1 Announce Type: new Abstract: Recent video generation models (VGMs) have made substantial progress in visual fidelity, yet their ability to follow long, compositional instructions re

model-releasesarxiv-cs-cv
16 Jul 2026
Research

Video to All-in-focus Image Reconstruction Algorithm for Automated Microscopic Urinalysis

DGX agent

arXiv:2607.13601v1 Announce Type: cross Abstract: Microscopic urinalysis is a routine diagnostic test at hospitals. Recent studies have demonstrated the effectiveness of deep learning methods to autom

researcharxiv-cs-cv
16 Jul 2026
Model Releases

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

DGX agent

arXiv:2607.14088v1 Announce Type: new Abstract: Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimi

model-releasesarxiv-cs-cv
16 Jul 2026
Local Ai

Visual Place Recognition Using Rate-Encoded Spiking Neural Networks with Discrete STDP Learning

DGX agent

arXiv:2607.13584v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs) trained through unsupervised Spike-Timing-Dependent Plasticity (STDP) have been explored as solutions to visual loop cl

local-aiarxiv-cs-cv
16 Jul 2026
Local Ai

WAVE-Stereo: Warp-Aligned Volume Encoding for Stereo Matching

DGX agent

arXiv:2607.13674v1 Announce Type: new Abstract: Existing iterative stereo matching methods primarily adopt two types of correspondence representation: explicit matching search via correlation volumes

local-aiarxiv-cs-cv
16 Jul 2026
Model Releases

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

DGX agent

arXiv:2602.17659v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow langu

model-releasesarxiv-cs-cv
16 Jul 2026
Model Releases

A Calibrated Multimodal Ensemble for Ambivalence/Hesitancy Recognition: System Description and Private-Test Submission Strategy

DGX agent

arXiv:2607.12176v1 Announce Type: new Abstract: Ambivalence and hesitancy (A/H) undermine digital behaviour-change interventions, and recognizing them automatically from video is the goal of the ABAW

model-releasesarxiv-cs-cv
15 Jul 2026
Research

ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space

DGX agent

arXiv:2607.11673v2 Announce Type: replace Abstract: We present ABot-3DWorld 0, a universal multimodal 3D world model that turns text, image, and video inputs into high-fidelity, explorable 3D worlds.

researcharxiv-cs-cv
15 Jul 2026
Research

ACID: Adaptive Caching for vIDeo generation

DGX agent

arXiv:2607.12358v1 Announce Type: new Abstract: Video diffusion models produce high-quality generations but remain slow at inference due to their sequential denoising procedure. Caching-based accelera

researcharxiv-cs-cv
15 Jul 2026
Local Ai

ACZ-GSeg: Adaptive Concentric Zone-based Two-stage Ground Segmentation for LiDAR Point Clouds

DGX agent

arXiv:2607.12110v1 Announce Type: new Abstract: Ground segmentation is a fundamental prerequisite for autonomous navigation, environmental perception, and object detection in ground mobile platforms.

local-aiarxiv-cs-cv
15 Jul 2026
Model Releases

Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction

DGX agent

arXiv:2607.12293v1 Announce Type: new Abstract: Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs or

model-releasesarxiv-cs-cv
15 Jul 2026
Research

Adversarial Attacks on Online Handwriting using Salience-based Temporal Editing

DGX agent

arXiv:2607.12500v1 Announce Type: cross Abstract: Deep learning models for online handwriting recognition have been shown effective and are increasingly deployed in practical applications. However, th

researcharxiv-cs-cv
15 Jul 2026
Research

Affordance-Guided Diffusion Prior for 3D Hand Reconstruction

DGX agent

arXiv:2510.00506v2 Announce Type: replace Abstract: How can we reconstruct 3D hand poses when large portions of the hand are heavily occluded by itself or by objects? Humans often resolve such ambigui

researcharxiv-cs-cv
15 Jul 2026
Research

Analyzing Image Encoder Choices and Graph Homophily in GCN Frameworks for Breast Ultrasound Classification

DGX agent

arXiv:2607.12054v1 Announce Type: cross Abstract: Breast ultrasound is widely used for screening, yet automated analysis remains challenging due to speckle noise, acquisition variability, and weak sep

researcharxiv-cs-cv
15 Jul 2026
Tutorials

Anatomy-Privileged Distillation with Token Routing for MRI-Based Prediction of Perineural Invasion

DGX agent

arXiv:2607.11987v1 Announce Type: new Abstract: Perineural invasion (PNI) is associated with poor postoperative outcomes in intrahepatic cholangiocarcinoma, but it is confirmed by surgical pathology.

tutorialsarxiv-cs-cv
15 Jul 2026
Safety

Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity

DGX agent

arXiv:2607.11957v1 Announce Type: new Abstract: Maintenance of critical infrastructures, such as railways and power plants, is essential for ensuring operational safety and reliability. However, the d

safetyarxiv-cs-cv
15 Jul 2026
Safety

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment

DGX agent

arXiv:2510.13698v4 Announce Type: replace Abstract: Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety

safetyarxiv-cs-cv
15 Jul 2026
Model Releases

Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks

DGX agent

arXiv:2607.12278v1 Announce Type: new Abstract: Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answer

model-releasesarxiv-cs-cv
15 Jul 2026
Research

AVQ-Attention: Adaptive Vector-Quantized Attention

DGX agent

arXiv:2607.12789v1 Announce Type: cross Abstract: The O(N^2) complexity of attention over N tokens remains a computational bottleneck in transformer models. Vector-Quantized (VQ) attention reduces thi

researcharxiv-cs-cv
15 Jul 2026
Model Releases

AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

DGX agent

arXiv:2607.12820v1 Announce Type: new Abstract: Omni-modal video captioning is not merely combining visual captioning with audio transcription: a useful caption must describe how visual actions, speec

model-releasesarxiv-cs-cv
15 Jul 2026
Research

Beyond Perceptual Distance: Discrepancy Assessment on Deep Representation for Out-of-Distribution Detection with Diffusion Model

DGX agent

arXiv:2409.10094v3 Announce Type: replace Abstract: Out-of-Distribution (OoD) detection aims to justify whether a given sample is from the training distribution of the classifier-under-protection, i.e

researcharxiv-cs-cv
15 Jul 2026
Model Releases

Beyond Perfect Priors: Adaptive Gaussian Graph for 4D Driving Reconstruction in the Wild

DGX agent

arXiv:2607.12214v1 Announce Type: new Abstract: Reconstructing 4D driving scenes in the wild (e.g., internet and AI-generated videos) is critical for diverse autonomous driving simulation. While recen

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Breaking Deja Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

DGX agent

arXiv:2607.12818v1 Announce Type: new Abstract: Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop clos

model-releasesarxiv-cs-cv
15 Jul 2026
Safety

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation

DGX agent

arXiv:2601.08010v3 Announce Type: replace Abstract: Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasonin

safetyarxiv-cs-cv
15 Jul 2026
Applications

Causal Supervision of Attention for Affective Behaviour Analysis

DGX agent

arXiv:2607.12091v1 Announce Type: new Abstract: Affective Behaviour Analysis aims to enable machines to infer human affective states from behavioural signals, particularly facial expressions, in real-

applicationsarxiv-cs-cv
15 Jul 2026
Safety

CGRL: Concept-Guided Pruning and Representation Learning for Whole-Slide Image Classification

DGX agent

arXiv:2607.12556v1 Announce Type: new Abstract: Weakly supervised whole-slide image (WSI) classification is widely used in computational pathology because slide-level labels are easier to obtain than

safetyarxiv-cs-cv
15 Jul 2026
Applications

Color Pass-Through via Camera-Display Coupling

DGX agent

arXiv:2607.12746v1 Announce Type: new Abstract: When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often differs noticeably from the original scen

applicationsarxiv-cs-cv
15 Jul 2026
Safety

Compos3D: Interactive Part-Based Composition for Creative Control in Generative 3D Models

DGX agent

arXiv:2607.12193v1 Announce Type: cross Abstract: While generative AI has unlocked new opportunities for 3D content creation, current workflows often rely on multiple regenerations, which provides lim

safetyarxiv-cs-cv
15 Jul 2026
Applications

Contrastive-Augmented Flow Matching for Style-Content Disentanglement

DGX agent

arXiv:2607.12404v1 Announce Type: new Abstract: Learning representations that separate content and style is crucial for controllable generation and compositional generalization. However, diffusion and

applicationsarxiv-cs-cv
15 Jul 2026
← Previous
1…5051525354…261
Next →