AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

AnyStyle: A Single LoRA is Sufficient for Image-Guided Style Transfer

DGX agent

arXiv:2607.04677v1 Announce Type: new Abstract: Image-guided style transfer aims to apply the artistic characteristics of a style image to a content image while preserving its semantic structure and l

researcharxiv-cs-cv
7 Jul 2026
Applications
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Aperture-aware Dispersion 5-D Light-field Imaging Spectrometer

DGX agent

arXiv:2607.04635v1 Announce Type: new Abstract: Enhancing perceptual dimensions while miniaturizing imaging systems presents significant challenges for high-dimensional visual sensing. Conventionally,

applicationsarxiv-cs-cv
7 Jul 2026
Agents

AppAgent: Multimodal Agents as Smartphone Users

DGX agent

arXiv:2312.13771v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper i

agentsarxiv-cs-cv
7 Jul 2026
Safety

AquaStereo: Enabling Underwater Stereo Matching via Depth-Conditioned Diffusion and Geometry Self-Distillation

DGX agent

arXiv:2607.04303v1 Announce Type: new Abstract: Learning-based stereo matching models struggle in underwater environments due to scarce in-domain data and the difficulty of extracting discriminative c

safetyarxiv-cs-cv
7 Jul 2026
Applications

Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation

DGX agent

arXiv:2512.18176v2 Announce Type: replace Abstract: Accurate segmentation of anatomical structures in medical images is essential for diagnosis and treatment planning. While recent interactive segment

applicationsarxiv-cs-cv
7 Jul 2026
Model Releases

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

DGX agent

arXiv:2606.12555v2 Announce Type: replace-cross Abstract: Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a

model-releasesarxiv-cs-cv
7 Jul 2026
Local Ai

AULLM++: Structured-Token-Conditioned Large Language Models for Micro-Expression Action Unit Detection

DGX agent

arXiv:2603.08387v2 Announce Type: replace Abstract: Micro-expression Action Unit (AU) detection identifies localized AUs from subtle facial muscle activations, providing a foundation for decoding affe

local-aiarxiv-cs-cv
7 Jul 2026
Safety

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

DGX agent

arXiv:2607.04311v1 Announce Type: new Abstract: Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity

safetyarxiv-cs-cv
7 Jul 2026
Research

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

DGX agent

arXiv:2607.02968v1 Announce Type: new Abstract: Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiT

researcharxiv-cs-cv
7 Jul 2026
Research

BAT3R: Bootstrapping Articulated 3D Reconstruction from 2D Image Collections

DGX agent

arXiv:2607.03891v1 Announce Type: new Abstract: 3D reconstruction of articulated objects from a single image is challenging because large training datasets with paired image and 3D supervision are dif

researcharxiv-cs-cv
7 Jul 2026
Research

Be Indiscrete: The Benefits of Learning Continuous Spine Degeneration Severity Scores

DGX agent

arXiv:2607.05090v1 Announce Type: new Abstract: Lumbar spine degeneration is a major contributor to chronic low back pain and is routinely assessed on MRI using ordinal grading systems, e.g. normal, m

researcharxiv-cs-cv
7 Jul 2026
Research

Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis

DGX agent

arXiv:2607.05348v1 Announce Type: new Abstract: Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring semantic knowledge from vision-language mo

researcharxiv-cs-cv
7 Jul 2026
Model Releases

Beyond Modality Fusion: Deep Ensembles for Multimodal Classification

DGX agent

arXiv:2607.05019v1 Announce Type: cross Abstract: In multimodal classification, late-fusion approaches classify concatenated modality-specific features extracted by unimodal neural networks. When moda

model-releasesarxiv-cs-cv
7 Jul 2026
Research

Beyond Post-Quantization: Native Hash Learning with a Dedicated HASH Token

DGX agent

arXiv:2607.03328v1 Announce Type: new Abstract: Efficient large-scale image retrieval requires compact representations that preserve semantic similarity under fast Hamming-space search. Deep hashing i

researcharxiv-cs-cv
7 Jul 2026
Safety

Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation

DGX agent

arXiv:2607.04249v1 Announce Type: new Abstract: Precise medical image segmentation is crucial for clinical diagnosis and treatment planning, yet relies heavily on expensive expert annotations. Semi-su

safetyarxiv-cs-cv
7 Jul 2026
Model Releases

Beyond Scene Priors: Fine-Grained Traffic Scene Reasoning with Benchmarking and Query-Guided Small-Object Focus

DGX agent

arXiv:2607.04149v1 Announce Type: new Abstract: In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language M

model-releasesarxiv-cs-cv
7 Jul 2026
Tutorials

Binary Iterative Method for Non-targeted Adversarial Attack

DGX agent

arXiv:2607.04145v1 Announce Type: cross Abstract: Adversarial attacks guide and provide additional training and test data for both adversarial training and adversarial robustness validation, and expos

tutorialsarxiv-cs-cv
7 Jul 2026
Research

Biomechanics-aware Multi-view Markerless Motion Capture of Dexterous Hand Movements

DGX agent

arXiv:2607.02796v1 Announce Type: new Abstract: Markerless motion capture (MMC) techniques have been widely beneficial in biomechanical analysis of human movement; however, application to complex moti

researcharxiv-cs-cv
7 Jul 2026
Safety

BiSLW: Bi-Spectral Latent Watermarking for Generative Diffusion Models

DGX agent

arXiv:2607.02643v1 Announce Type: new Abstract: Diffusion-based generative models have transformed visual content synthesis, yet they remain vulnerable to unauthorized usage and lack reliable attribut

safetyarxiv-cs-cv
7 Jul 2026
Research

BrainNormalizer: Anatomy-Informed Pseudo-Healthy Brain Reconstruction from Tumor MRI via Edge-Guided ControlNet

DGX agent

arXiv:2511.12853v2 Announce Type: replace-cross Abstract: Brain tumors induce complex structural deformations that obscure the patient' s original neuroanatomy, making it difficult to distinguish tumo

researcharxiv-cs-cv
7 Jul 2026
Research

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

DGX agent

arXiv:2607.03184v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive general capabilities, they struggle with fine-grained perception in ultra-high-res

researcharxiv-cs-cv
7 Jul 2026
Tutorials

C^3ASD: Multi-Level Consistency-Driven Representation Learning

DGX agent

arXiv:2607.03018v1 Announce Type: new Abstract: Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform wel

tutorialsarxiv-cs-cv
7 Jul 2026
Research

Cancelable Biometric Template Protection Based on Multi-Instance Fusion: A Contralateral Iris Approach

DGX agent

arXiv:2607.02860v1 Announce Type: new Abstract: Biometric templates are vulnerable to theft if stored without protection. Unlike passwords, a compromised iris cannot be reissued. Although existing can

researcharxiv-cs-cv
7 Jul 2026
Local Ai

Causal-RetiGraph: Cross-Cohort Retinal Support and Same-Subject Pathway Analysis for Diabetic Retinopathy

DGX agent

arXiv:2607.05204v1 Announce Type: new Abstract: Diabetic retinopathy (DR) is a local retinal lesion process and a visible manifestation of systemic microvascular injury. Modern retinal AI can grade im

local-aiarxiv-cs-cv
7 Jul 2026
Safety

CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation

DGX agent

arXiv:2607.04451v1 Announce Type: new Abstract: Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. However, generating con

safetyarxiv-cs-cv
7 Jul 2026
Research

CDST: Color Disentangled Style Transfer for Universal Style Reference Customization

DGX agent

arXiv:2506.13770v2 Announce Type: replace Abstract: We introduce Color Disentangled Style Transfer (CDST), a novel and efficient two-stream style transfer training paradigm which completely isolates c

researcharxiv-cs-cv
7 Jul 2026
Research

Cell as Point: One-Stage Framework for Efficient Cell Tracking

DGX agent

arXiv:2411.14833v4 Announce Type: replace-cross Abstract: Conventional multi-stage cell tracking approaches rely heavily on detection or segmentation in each frame as a prerequisite, requiring substan

researcharxiv-cs-cv
7 Jul 2026
Research

CenSynCMB: Centre Maps and Physics-Guided Synthesis for Microbleed Detection

DGX agent

arXiv:2607.05325v1 Announce Type: new Abstract: Cerebral microbleeds (CMBs) are MRI markers of small vessel disease and the microbleed component of amyloid related imaging abnormalities (ARIA-H), but

researcharxiv-cs-cv
7 Jul 2026
Model Releases

ChatImage: Navigating Long-Form LLM Answers through Interactive Images

DGX agent

arXiv:2607.05290v1 Announce Type: new Abstract: Large Language Models (LLMs) can produce detailed answers to complex queries, but these answers are typically presented as dense linear text, which make

model-releasesarxiv-cs-cv
7 Jul 2026
Applications

CIPHER: Causal Intervention Pathways for Healthcare Equity and Robustness

DGX agent

arXiv:2607.02596v1 Announce Type: new Abstract: Deep learning models for medical diagnosis frequently exhibit substantial performance disparities across sensitive subgroups (e.g., race, sex), even whe

applicationsarxiv-cs-cv
7 Jul 2026
Research

City landscape in sight: A crowdsourced framework for unlocking urban-scale window view perceptions from real estate imagery

DGX agent

arXiv:2606.15198v2 Announce Type: replace Abstract: City landscapes viewed through home windows influence quality of life, yet perceptions of actual window views at the urban scale remain understudied

researcharxiv-cs-cv
7 Jul 2026
Research

City-Level 3D Surface Reconstruction with Viewpoint Orientation Partitioning and Scene Completion

DGX agent

arXiv:2607.03771v1 Announce Type: new Abstract: Multi-view 3D surface reconstruction is a longstanding challenge in computer vision. Although recent large-scale reconstruction methods based on 3D Gaus

researcharxiv-cs-cv
7 Jul 2026
Model Releases

CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection

DGX agent

arXiv:2607.02930v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel in diverse vision tasks, but full-parameter retraining is computationally expensive as real-world knowled

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

CLABTOOLKIT: An Open-Source Toolkit for Routine Processing, Manipulation, and Visualization of Neuroimaging Data

DGX agent

arXiv:2607.02638v1 Announce Type: new Abstract: Neuroimaging research requires manipulating heterogeneous data structures, including raw MRI volumes, volumetric parcellations, cortical surface meshes,

model-releasesarxiv-cs-cv
7 Jul 2026
Safety

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

DGX agent

arXiv:2607.05150v1 Announce Type: new Abstract: In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinfor

safetyarxiv-cs-cv
7 Jul 2026
Research

CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space

DGX agent

arXiv:2512.08029v3 Announce Type: replace-cross Abstract: Clinical decision-making in oncology requires predicting dynamic disease evolution, a task current static AI predictors cannot perform. While

researcharxiv-cs-cv
7 Jul 2026
Applications

Classroom Behavior Monitoring with YOLO An Empirical Study in Higher Education Settings

DGX agent

arXiv:2607.02580v1 Announce Type: new Abstract: Classroom behavior monitoring plays a vital role in evaluating student engagement and improving teaching effectiveness. Traditional observation methods

applicationsarxiv-cs-cv
7 Jul 2026
Safety

CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving

DGX agent

arXiv:2607.02841v1 Announce Type: cross Abstract: End-to-end autonomous driving (E2E-AD) aims to directly map raw sensor information to driving actions. Recently, with the rapid advancement of multi-m

safetyarxiv-cs-cv
7 Jul 2026
Safety

CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization

DGX agent

arXiv:2509.15330v2 Announce Type: replace Abstract: Recent advances in pre-training vision-language models (VLMs), e.g., contrastive language-image pre-training (CLIP) methods, have shown great potent

safetyarxiv-cs-cv
7 Jul 2026
Agents

CogRad: A Cognitively-Inspired Multi-Agent Framework for Radiology Report Generation

DGX agent

arXiv:2607.03853v1 Announce Type: new Abstract: Automated radiology report generation (RRG) can ease radiologist workload, yet most existing systems produce a report in a single forward pass, with no

agentsarxiv-cs-cv
7 Jul 2026
Research

Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification

DGX agent

arXiv:2509.24181v2 Announce Type: replace Abstract: Active learning (AL) aims to build high-quality labeled datasets by iteratively selecting the most informative samples from an unlabeled pool under

researcharxiv-cs-cv
7 Jul 2026
Model Releases

ComplexMimic: Human-Scene Interaction Imitation in Complex 3D Environments

DGX agent

arXiv:2607.02034v2 Announce Type: replace Abstract: Physics-based Human-Scene Interaction (HSI) imitation learning is crucial for embodied intelligence as it bridges the gap between kinematic 3D motio

model-releasesarxiv-cs-cv
7 Jul 2026
Research

Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models

DGX agent

arXiv:2602.24264v2 Announce Type: replace Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems. Although mod

researcharxiv-cs-cv
7 Jul 2026
Research

CompressedVQA-AEV: Full-Reference and No-Reference Quality Assessment Models for Asymmetric Encoded Videos

DGX agent

arXiv:2607.04606v1 Announce Type: cross Abstract: This report presents our solutions to the QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos, comprising a full-refe

researcharxiv-cs-cv
7 Jul 2026
Research

Consistent and Editable: A Balanced Framework for Text-Guided Video Editing

DGX agent

arXiv:2607.05056v1 Announce Type: new Abstract: Recently, diffusion models have achieved considerable success in the text-guided video editing domain. However, existing works often struggle to balance

researcharxiv-cs-cv
7 Jul 2026
Model Releases

Continual Model Merging with Test-Time Adaptation for Whole-Slide Image Analysis

DGX agent

arXiv:2607.04755v1 Announce Type: new Abstract: Model merging offers a practical alternative to conventional continual learning by integrating independently fine-tuned models without retaining previou

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

ContiStain: Cross-Domain Relation-Preserving Distillation for Continual Multi-Domain Virtual IHC Staining

DGX agent

arXiv:2607.03851v1 Announce Type: new Abstract: A unified multiplex virtual staining model enables scalable and non-destructive multiplex analysis from H&E slides while promoting parameter efficiency,

model-releasesarxiv-cs-cv
7 Jul 2026
Research

ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering

DGX agent

arXiv:2509.21541v3 Announce Type: replace-cross Abstract: Hair simulation and rendering are challenging due to complex strand dynamics, diverse material properties, and intricate light-hair interactio

researcharxiv-cs-cv
7 Jul 2026
← Previous
1…6061626364…261
Next →