AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

DGX agent

arXiv:2605.22109v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchm

model-releasesarxiv-cs-cv
22 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Physiology and Anatomy Aware Inverse Inference of Myocardial Infarction for Cardiac Digital Twin

DGX agent

arXiv:2605.22044v1 Announce Type: new Abstract: Accurate localization of myocardial infarction is essential for risk stratification. While LGE-MRI remains the gold standard, it is resource-intensive.

model-releasesarxiv-cs-cv
22 May 2026
Safety

PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects

DGX agent

arXiv:2605.21572v1 Announce Type: new Abstract: Simulation-ready physical 3D assets have emerged as a promising direction owing to their broad applicability in downstream tasks. However, most existing

safetyarxiv-cs-cv
22 May 2026
Research

PIU: Proximity-guided Identity Unlearning in ID-Conditioned Diffusion Models

DGX agent

arXiv:2605.22311v1 Announce Type: new Abstract: Identity-conditioned diffusion models enable high-quality and identity-consistent face generation, but they also raise severe privacy concerns, as model

researcharxiv-cs-cv
22 May 2026
Applications

PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-Thought

DGX agent

arXiv:2605.22013v1 Announce Type: new Abstract: Understanding 3D point clouds through language remains a fundamental challenge in computer graphics and visual computing, due to the irregular structure

applicationsarxiv-cs-cv
22 May 2026
Model Releases

Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

DGX agent

arXiv:2605.22446v1 Announce Type: new Abstract: While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deplo

model-releasesarxiv-cs-cv
22 May 2026
Safety

QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks

DGX agent

arXiv:2605.22351v1 Announce Type: new Abstract: Low-bit quantization is widely used to compress super-resolution (SR) models and reduce storage and computation costs for deployment on resource-limited

safetyarxiv-cs-cv
22 May 2026
Research

REACH: Hand Pose Estimation from Room Corners

DGX agent

arXiv:2605.22231v1 Announce Type: new Abstract: We introduce a novel 3D hand pose estimator that can accurately recover the shape and pose of people's hands in a room from afar, typically from fixed c

researcharxiv-cs-cv
22 May 2026
Model Releases

Rethinking Noise-Robust Training for Frozen Vision Foundation Models: A Cross-Dataset Benchmark with a Case Study of Small-Loss Failure

DGX agent

arXiv:2605.22591v1 Announce Type: new Abstract: Frozen Vision Foundation Models (VFMs) with lightweight classification heads are increasingly used in medical imaging because they offer efficient and r

model-releasesarxiv-cs-cv
22 May 2026
Research

Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness

DGX agent

arXiv:2605.22011v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) achieve superior image generation quality but suffer from quadratic computational complexity relative to token count. Whil

researcharxiv-cs-cv
22 May 2026
Research

Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning

DGX agent

arXiv:2602.23833v2 Announce Type: replace-cross Abstract: Automated identification of DICOM image series is essential for large-scale medical image analysis, quality control, protocol harmonization, a

researcharxiv-cs-cv
22 May 2026
Research

RiT: Vanilla Diffusion Transformers Suffice in Representation Space

DGX agent

arXiv:2605.21981v1 Announce Type: new Abstract: Flow matching with x-prediction -- regressing the clean data point rather than the ambient velocity -- is known to exploit low-dimensional manifold stru

researcharxiv-cs-cv
22 May 2026
Research

RobuQ: Pushing DiTs to W1.58A2 via Robust Activation Quantization

DGX agent

arXiv:2509.23582v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for image generation, demonstrating superior scalability and performance

researcharxiv-cs-cv
22 May 2026
Research

Robustness of breast lesion segmentation under MRI undersampling improves with k-space-aware deep learning

DGX agent

arXiv:2605.22327v1 Announce Type: new Abstract: Purpose: To assess whether breast lesion segmentation can be learned directly from acquired MRI k-space, and whether doing so improves robustness when d

researcharxiv-cs-cv
22 May 2026
Model Releases

SADGE: Structure and Appearance Domain Gap Estimation of Synthetic and Real Data

DGX agent

arXiv:2605.22467v1 Announce Type: new Abstract: We propose SADGE, a quantitative similarity metric that predicts the performance of synthetic image datasets for common computer vision tasks without do

model-releasesarxiv-cs-cv
22 May 2026
Local Ai

SceneAligner: 3D-Grounded Floorplan Localization in the Wild

DGX agent

arXiv:2605.22581v1 Announce Type: new Abstract: Many public buildings provide floorplans with a 'you are here' indicator to help visitors orient themselves. Floorplan localization seeks to computation

local-aiarxiv-cs-cv
22 May 2026
Model Releases

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching

DGX agent

arXiv:2605.21788v1 Announce Type: new Abstract: Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VL

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

DGX agent

arXiv:2605.21919v1 Announce Type: new Abstract: Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development

model-releasesarxiv-cs-cv
22 May 2026
Research

SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers

DGX agent

arXiv:2605.22668v1 Announce Type: new Abstract: Diffusion transformers (DiTs) have emerged as a dominant architecture for text-to-image generation, yet their performance drops when generating at resol

researcharxiv-cs-cv
22 May 2026
Local Ai

SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation

DGX agent

arXiv:2605.22658v1 Announce Type: new Abstract: While large language models provide strong compositional reasoning, existing reasoning segmentation pipelines fail to transparently connect this reasoni

local-aiarxiv-cs-cv
22 May 2026
Model Releases

SegGuidedNet: Sub-Region-Aware Attention Supervision for Interpretable Brain Tumor Segmentation

DGX agent

arXiv:2605.22572v1 Announce Type: new Abstract: Accurate segmentation of brain tumour sub-regions from multi-parametric MRI is critical for treatment planning yet remains challenging due to morphologi

model-releasesarxiv-cs-cv
22 May 2026
Tutorials

Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking

DGX agent

arXiv:2605.22538v1 Announce Type: new Abstract: Traditional visual object tracking (VOT) methods typically rely on task-specific supervised training, limiting their generalization to unseen objects an

tutorialsarxiv-cs-cv
22 May 2026
Model Releases

Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

DGX agent

arXiv:2605.21852v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret invo

model-releasesarxiv-cs-cv
22 May 2026
Agents

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving

DGX agent

arXiv:2605.22809v1 Announce Type: new Abstract: Robust training and validation of Autonomous Driving Systems (ADS) require massive, diverse datasets. Proprietary data collected by Autonomous Vehicle (

agentsarxiv-cs-cv
22 May 2026
Model Releases

SFN-YOLO: Towards Free-Range Poultry Detection via Scale-aware Fusion Networks

DGX agent

arXiv:2509.17086v2 Announce Type: replace Abstract: Detecting and localizing poultry is essential for advancing smart poultry farming. Despite the progress of detection-centric methods, challenges per

model-releasesarxiv-cs-cv
22 May 2026
Research

Skarimva: Skeleton-based Action Recognition is a Multi-view Application

DGX agent

arXiv:2602.23231v2 Announce Type: replace Abstract: Human action recognition plays an important role when developing intelligent interactions between humans and machines. While there is a lot of activ

researcharxiv-cs-cv
22 May 2026
Research

Slimmable ConvNeXt: Width-Adaptive Inference for Efficient Multi-Device Deployment

DGX agent

arXiv:2605.22677v1 Announce Type: new Abstract: Deploying vision models across devices with varying resource constraints, or even on a single device where available compute fluctuates due to battery s

researcharxiv-cs-cv
22 May 2026
Research

SO-Mamba: State-Ownership Mamba for Unrolled MRI Reconstruction

DGX agent

arXiv:2605.22031v1 Announce Type: new Abstract: Accelerated MRI reconstruction requires recovering missing details while preserving anatomically coherent structures across large spatial regions. State

researcharxiv-cs-cv
22 May 2026
Model Releases

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

DGX agent

arXiv:2511.07820v3 Announce Type: replace-cross Abstract: Despite the rise of billion-parameter foundation models trained across thousands of GPUs, similar scaling gains have not been shown for humano

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

DGX agent

arXiv:2512.10719v2 Announce Type: replace Abstract: End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual under

model-releasesarxiv-cs-cv
22 May 2026
Applications

Spectral Tail Auxiliary Learning for AI-Generated Image Detection

DGX agent

arXiv:2605.22751v1 Announce Type: new Abstract: As generative image models evolve rapidly, the perceptual gap between generated and real images continues to narrow, making AI-generated image detection

applicationsarxiv-cs-cv
22 May 2026
Research

SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents

DGX agent

arXiv:2603.08403v3 Announce Type: replace Abstract: Long-horizon action-conditioned video generation aims to synthesize temporally coherent videos that follow complex action instructions over extended

researcharxiv-cs-cv
22 May 2026
Research

ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs

DGX agent

arXiv:2605.22158v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) face significant computational overhead when processing long videos due to the massive number of visual token

researcharxiv-cs-cv
22 May 2026
Safety

Supervised Classification Heads as Semantic Prototypes: Unlocking Vision-Language Alignment via Weight Recycling

DGX agent

arXiv:2605.22484v1 Announce Type: new Abstract: Vision-Language Models (VLMs) excel at tasks like zero-shot classification and cross-modal retrieval by mapping images and text to a shared space, but t

safetyarxiv-cs-cv
22 May 2026
Research

Swift Sampling: Selecting Temporal Surprises via Taylor Series

DGX agent

arXiv:2605.22678v1 Announce Type: new Abstract: While most frames in long-form video are redundant, the critical information resides in temporal surprises: moments where the actual visual features dev

researcharxiv-cs-cv
22 May 2026
Applications

Synthetic Data Alone is Enough? Rethinking Data Scarcity in Pediatric Rare Disease Recognition

DGX agent

arXiv:2605.22767v1 Announce Type: new Abstract: Children with rare genetic diseases often exhibit distinctive facial phenotypes, yet developing computer vision systems for early diagnosis remains chal

applicationsarxiv-cs-cv
22 May 2026
Model Releases

Tackle CSM in JPEG Steganalysis with Data Adaptation

DGX agent

arXiv:2605.21523v1 Announce Type: cross Abstract: Steganalysis models excel on benchmark datasets but struggle in the wild when analyzed images are produced by a processing pipeline unseen during trai

model-releasesarxiv-cs-cv
22 May 2026
Tutorials

TextTeacher: What Can Language Teach About Images?

DGX agent

arXiv:2605.22098v1 Announce Type: new Abstract: The platonic representation hypothesis suggests that sufficiently large models converge to a shared representation geometry, even across modalities. Mot

tutorialsarxiv-cs-cv
22 May 2026
Research

The Neglected Baseline in Model Interpretation

DGX agent

arXiv:2605.22417v1 Announce Type: new Abstract: We observe that existing model interpretation methods generally ignore the baseline, and such neglect often results in imprecise or even incorrect inter

researcharxiv-cs-cv
22 May 2026
Model Releases

Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception

DGX agent

arXiv:2605.21882v1 Announce Type: new Abstract: Vision-language models (VLMs) often fail under low illumination because their visual grounding is learned predominantly from RGB imagery, whereas therma

model-releasesarxiv-cs-cv
22 May 2026
Research

Time-varying rPPG signal separation via block-sparse signal model

DGX agent

arXiv:2605.22425v1 Announce Type: cross Abstract: Remote photoplethysmography (rPPG) enables non-contact measurement of cardiac pulse signals by analyzing subtle color changes in facial videos. Nevert

researcharxiv-cs-cv
22 May 2026
Model Releases

Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence

DGX agent

arXiv:2605.22414v1 Announce Type: new Abstract: Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential f

model-releasesarxiv-cs-cv
22 May 2026
Research

Towards Initialization-free Calibrated Bundle Adjustment

DGX agent

arXiv:2506.23808v2 Announce Type: replace Abstract: A recent series of works has shown that initialization-free BA can be achieved using pseudo Object Space Error (pOSE) as a surrogate objective. The

researcharxiv-cs-cv
22 May 2026
Model Releases

Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images

DGX agent

arXiv:2511.02014v2 Announce Type: replace Abstract: The detection of Protected Health Information (PHI) in medical imaging is critical for safeguarding patient privacy and ensuring compliance with reg

model-releasesarxiv-cs-cv
22 May 2026
Research

Training-Free Fine-Grained Semantic Segmentations in Low Data Regimes: A FungiTastic Baseline

DGX agent

arXiv:2605.22492v1 Announce Type: new Abstract: Fine-grained semantic segmentation requires both precise localization and discrimination between visually similar classes. In FungiTastic, this problem

researcharxiv-cs-cv
22 May 2026
Research

Translating Signals to Languages for sEMG-Based Activity Recognition

DGX agent

arXiv:2605.22403v1 Announce Type: new Abstract: Surface electromyography (sEMG) signal-based activity recognition has attracted increasing research attention in recent years. To develop accurate sEMG

researcharxiv-cs-cv
22 May 2026
Model Releases

Transporting Task Vectors across Different Architectures without Training

DGX agent

arXiv:2602.12952v2 Announce Type: replace-cross Abstract: Adapting large pre-trained models to downstream tasks often produces task-specific parameter updates that are expensive to relearn for every m

model-releasesarxiv-cs-cv
22 May 2026
Research

TWINGS: Thin Plate Splines Warp-aligned Initialization for Sparse-View Gaussian Splatting

DGX agent

arXiv:2605.22069v1 Announce Type: new Abstract: Novel view synthesis from sparse-view inputs poses a significant challenge in 3D computer vision, particularly for achieving high-quality scene reconstr

researcharxiv-cs-cv
22 May 2026
← Previous
1…148149150151152…263
Next →