AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
12 Aug 2026

Logit Lens Supervision for Patch-Level Explanations in Vision-Language Models

ResearchDGX agent

arXiv:2602.01530v2 Announce Type: replace Abstract: Modern autoregressive Vision-Language Models (VLMs) can generate fluent answers while their visual-token representations become weakly tied to the i

Longitudinal 3D Foundation Modeling for Neoadjuvant Breast Cancer Response Prediction from Serial DCE-MRI

ResearchDGX agent

arXiv:2608.09991v1 Announce Type: cross Abstract: Pathologic complete response (pCR) is an important endpoint in neoadjuvant chemotherapy (NAC) for breast cancer, and predicting pCR from imaging durin

LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3

SafetyDGX agent

arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MAD-HOI: Masked Autoregressive Diffusion for Generating Articulated Hand Object Interactions from Text

Model ReleasesDGX agent

arXiv:2608.10162v1 Announce Type: new Abstract: Methods for text-based generation of hand-object interaction (HOI) sequences primarily focus on producing smooth, physically plausible trajectories. A t

MammoMix: Leveraging Mixture of Experts for Robust Mammogram Breast Detection

ApplicationsDGX agent

arXiv:2608.10437v1 Announce Type: new Abstract: Breast lesion detection in mammography remains a challenging task due to variations in image quality, lesion appearance, and population demographics acr

Mixture-of-Experts-based Entropy Model for Learned Image Compression

ResearchDGX agent

arXiv:2608.10947v1 Announce Type: new Abstract: Learned image compression has seen significant progress in recent years with the development of end-to-end learned models that achieve better compressio

MMArt A Multi-Perspective Multimodal Dataset for Visual Art Understanding

ResearchDGX agent

arXiv:2608.10706v1 Announce Type: new Abstract: Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface c

More Accurate, Less Human: Gestalt Grouping in Vision Models

Model ReleasesDGX agent

arXiv:2608.10195v1 Announce Type: new Abstract: Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into r

Motion Artifact-Aware Self-Supervised Representation Learning for 3D Brain MRI Motion Artifact Reduction

ResearchDGX agent

arXiv:2608.10170v1 Announce Type: new Abstract: Patient motion remains a source of image degradation in brain MRI, leading to signal loss, blurring, and geometric distortion that compromise quantitati

Multi-Level Evidence Aggregation for Robust Facial Phenotype Retrieval in Rare Genetic Disorder Prioritization

ResearchDGX agent

arXiv:2608.11037v1 Announce Type: new Abstract: AI-assisted facial phenotyping supports rare genetic disorder prioritization by retrieving visually similar diagnosed cases from facial image reference

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

SafetyDGX agent

arXiv:2608.10864v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile,

Multimodal Ambivalence and Hesitancy Recognition via Cross-Attention and Gated Fusion

ResearchDGX agent

arXiv:2607.15779v2 Announce Type: replace Abstract: We present a multimodal framework for Ambivalence/Hesitancy (A/H) recognition in video, developed for the ABAW11 challenge at ECCV 2026. The propose

Multiple Scale Latents for Learned Image Compression

ResearchDGX agent

arXiv:2608.10952v1 Announce Type: new Abstract: Most learned image compression systems rely on a single latent representation combined with a hyperprior, which limits their ability to efficiently capt

Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.10824v1 Announce Type: cross Abstract: Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer.

NullEdit: Stealthy Image Protection via VLM Condition Redirection

Model ReleasesDGX agent

arXiv:2608.10870v1 Announce Type: new Abstract: Modern image editors combine vision-language models (VLMs) with diffusion transformer backbones to modify a single reference image according to instruct

Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs

TutorialsDGX agent

arXiv:2608.10959v1 Announce Type: new Abstract: Existing vision-language model (VLM) backdoors are usually treated as static vulnerabilities: one-to-one and N-to-N attacks bind one or more triggers to

P3CA: Encoder-Agnostic Interpretation of Vision Foundation Model Embeddings via Spatial Probing

Local AiDGX agent

arXiv:2608.10131v1 Announce Type: new Abstract: Vision foundation models are increasingly used as reusable encoders in medical image computing, yet their high-dimensional spatial embeddings are diffic

PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2608.10985v1 Announce Type: new Abstract: Erasing concepts from large-scale text-to-image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infrin

PolyLayout: Hierarchical VLM-Guided Layout Generation Beyond Rectangular Rooms

ApplicationsDGX agent

arXiv:2608.10838v1 Announce Type: new Abstract: Generating physically plausible 3D room layouts is essential for home furnishing retail, enabling customers to visualize products in their own homes and

PolypVision: A Three-Stage Hierarchical Deep Learning Framework for Classification and Segmentation of Colorectal Polyps

ResearchDGX agent

arXiv:2608.10649v1 Announce Type: new Abstract: Colorectal cancer (CRC) remains one of the leading causes of cancer-related mortality worldwide, predominantly arising from precancerous polyps. Accurat

Pre- to Post-Contrast Synthesis of Breast DCE-MRI using Latent Bridge Matching

ResearchDGX agent

arXiv:2608.10000v1 Announce Type: cross Abstract: Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is central to breast cancer imaging, but gadolinium administration increases scan burde

Precise Top-Layer Fabric Segmentation for Fabric Destacking with Edge- and Shape-Aware Deep Networks

ApplicationsDGX agent

arXiv:2608.10648v1 Announce Type: new Abstract: Fabric destacking requires precise segmentation of the topmost fabric layer, a task complicated by subtle fabric boundaries and high visual similarity b

PRMU: A Corpus-Free Benchmark for Person-Centric Knowledge Unlearning in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.11149v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in storing and recalling rich person-related knowledge, raising incre

Protection Levels for Vision-Based Pose Estimation

AgentsDGX agent

arXiv:2608.10023v1 Announce Type: cross Abstract: Vision-based navigation complements Global Navigation Satellite Systems, but certification demands integrity guarantees that account for faulty measur

Rethinking Data Efficiency in Industrial Dense Prediction: Pretraining Coherence, Not Inductive Bias, Determines ViTs Low-Data Advantage

SafetyDGX agent

arXiv:2608.10590v1 Announce Type: new Abstract: Vision Transformers (ViTs) are widely believed to require more labeled data than CNNs for industrial dense prediction. Through controlled experiments on

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

Model ReleasesDGX agent

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets

ApplicationsDGX agent

arXiv:2608.10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing sin

Robustness of transferability estimation metrics for medical imaging

ResearchDGX agent

arXiv:2608.09999v1 Announce Type: cross Abstract: In transfer learning, the choice of source model largely influences the performance on a target dataset. Still, selecting a fitting source remains a c

SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense

Local AiDGX agent

arXiv:2608.10933v1 Announce Type: new Abstract: Text-to-Video (T2V) generative models are vulnerable to jailbreak attacks in real-world deployment, leading them to produce harmful or inappropriate con

SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception

SafetyDGX agent

arXiv:2608.10497v1 Announce Type: new Abstract: While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature ex

SAR2Agri: Learning SAR Intensity Representations for Agricultural Monitoring

Model ReleasesDGX agent

arXiv:2608.11142v1 Announce Type: new Abstract: Agricultural monitoring faces unique challenges, arising from the landscape's complex temporal, phenological, and climate dynamics, yet monitoring them

SceneNAT: Masked Generative Modeling for Language-Guided Indoor Scene Synthesis

TutorialsDGX agent

arXiv:2601.07218v2 Announce Type: replace Abstract: We present SceneNAT, a masked non-autoregressive Transformer for 3D indoor scene synthesis from natural language instructions. It generates complete

SeFaR: Semantic Feature-aware Robustness Testing of Deep Neural Networks

SafetyDGX agent

arXiv:2608.10289v1 Announce Type: new Abstract: Deep neural networks are increasingly deployed in safety-critical domains as perception modules, where failures are often caused due to rare and under-r

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

ResearchDGX agent

arXiv:2608.10708v1 Announce Type: new Abstract: Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-scene optimization, achieving stron

Sensor-Informed Per-Point Covariance for Structured-Light 3D Imaging

Local AiDGX agent

arXiv:2608.10888v1 Announce Type: new Abstract: Per-point uncertainty models are important in structured-light 3D reconstruction for probabilistic registration, fusion, and quality assessment. In prac

Significance and Stability Analysis of Gene-Environment Interaction using GxEStat

Model ReleasesDGX agent

arXiv:2604.03337v2 Announce Type: replace Abstract: Genotype-environment (GxE) interactions can influence the performance of genotypes across diverse environments, limiting the reliability of genotype

Signpost Watermarking: Joint Optimization for Visual Watermark Coexistence

ResearchDGX agent

arXiv:2608.10091v1 Announce Type: new Abstract: We present a method for training imperceptible visual watermarks to coexist with other such watermarks. Recent work has shown that independently trained

SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis

ResearchDGX agent

arXiv:2608.10519v1 Announce Type: new Abstract: InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids. Its changing scale and cross-clip context,

SpecF2M: A Spectral-Aware Multi-task Network Estimating Axial Length and Refractive Error from Pediatric Fundus Photographs

ResearchDGX agent

arXiv:2608.09994v1 Announce Type: cross Abstract: Spherical Equivalent Refraction (SER) and Axial Length (AL) are core indicators for pediatric myopia screening, yet their measurements require dedicat

Static in Frames, Dynamic in Events: Rethinking Features in Event Cameras as Motion Cues

Model ReleasesDGX agent

arXiv:2608.11075v1 Announce Type: new Abstract: Event cameras capture intensity changes asynchronously with high temporal resolution, requiring novel preprocessing methods for downstream tasks. Unlike

Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation

Model ReleasesDGX agent

arXiv:2608.10439v1 Announce Type: new Abstract: Streaming video generation holds strong potential for world modeling, where future frames must be inferred online sequentially to form a continuous vide

Structural Guidance for Unified Joint Demosaicing and Denoising

Local AiDGX agent

arXiv:2608.09995v1 Announce Type: cross Abstract: Joint demosaicing and denoising is a fundamental step in camera image signal processing, yet remains challenging because different Bayer-like color fi

SuperQuadricOcc: Real-Time Self-Supervised Semantic Occupancy Estimation with Superquadric Volume Rendering

AgentsDGX agent

arXiv:2511.17361v5 Announce Type: replace Abstract: Self-supervision for semantic occupancy estimation is appealing as it removes the labour-intensive manual annotation, thus allowing one to scale to

TEASR: Training-Efficient Any-Step Diffusion Transformer for Real-World Image Super-Resolution

Model ReleasesDGX agent

arXiv:2606.16188v2 Announce Type: replace Abstract: Diffusion models excel in Real-World Image Super-Resolution (Real-ISR) due to their powerful generative priors but suffer from slow iterative sampli

The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

SafetyDGX agent

arXiv:2608.10839v1 Announce Type: new Abstract: This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems traine

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes

Local AiDGX agent

arXiv:2608.10981v1 Announce Type: new Abstract: Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-languag

Towards Color-Faithful Low-Light Image Enhancement via Adaptive Color Debiasing and Saturation Rectification

SafetyDGX agent

arXiv:2608.10512v1 Announce Type: new Abstract: Low-light imaging often introduces color bias caused by the low signal-to-noise ratio and the image formation process. Although recent low-light image e

Towards Geometry-Grounded Dense Semantic Matching with VGGT Priors

ResearchDGX agent

arXiv:2509.21263v2 Announce Type: replace Abstract: Semantic matching aims to establish pixel-level correspondences between instances of the same category and represents a fundamental task in computer

TRACE-GS: On-Policy Trajectory Distillation with Privileged Geometric Conditioning for Sparse-View 3DGS Restoration

SafetyDGX agent

arXiv:2608.10286v1 Announce Type: new Abstract: We present TRACE-GS, an on-policy trajectory distillation framework that leverages privileged geometric conditioning at training time, thereby adapting

Transformer Geometry Observatory TGO-IV: Developmental Topology Observatory

TutorialsDGX agent

arXiv:2608.09997v1 Announce Type: cross Abstract: Transformers have had a profound impact on the world of language processing and computer vision. As efforts to answer the million-dollar question of `

UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment

SafetyDGX agent

arXiv:2608.10316v1 Announce Type: new Abstract: Multi-modal learning combining medical images and clinical text is promising for disease diagnosis. However, standard multi-modal training leads to shor

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

Local AiDGX agent

arXiv:2608.10835v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently hallucinate content unsupported by th

VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

Local AiDGX agent

arXiv:2608.11201v1 Announce Type: new Abstract: Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and auth

VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation

SafetyDGX agent

arXiv:2608.10903v1 Announce Type: new Abstract: Reliable clinical deployment of machine learning requires models that know when they are likely to fail, particularly for subgroups underrepresented in

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

Model ReleasesDGX agent

arXiv:2608.10682v1 Announce Type: new Abstract: Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning

SafetyDGX agent

arXiv:2608.11013v1 Announce Type: new Abstract: Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, lead

WaveInst: A Frequency-Domain Enhanced Network for Fine-Grained Thin Tree Trunk Extraction in Forest Scenes

ResearchDGX agent

arXiv:2505.01656v2 Announce Type: replace Abstract: Analyzing tree morphology, particularly trunk and branch extraction, is valuable for genetic breeding and forestry management. Existing image-based

What DINO saw: ALiBi positional encoding reduces positional bias in Vision Transformers

SafetyDGX agent

arXiv:2603.16840v2 Announce Type: replace Abstract: Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However

When Repository Labels Are Not Image-Level Truth: A Supervision Auditing Framework for Chest Radiograph AI

ApplicationsDGX agent

arXiv:2608.10084v1 Announce Type: cross Abstract: Public chest X-ray repositories are widely used to train medical AI systems, yet their labels are typically extracted from radiology reports rather th

When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs

Local AiDGX agent

arXiv:2608.10489v1 Announce Type: new Abstract: Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token p

← Previous
1234…207
Next →