AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
23 Jun 2026

TeleStyle V2: Beyond Content-Preserving Style Transfer with Self-Distillation and Distribution-Matching-Distillation

Model ReleasesDGX agent

arXiv:2606.20709v1 Announce Type: new Abstract: Given a content reference and a style reference, content-preserving style transfer requires the model to generate stylized outputs with content and styl

Temporally Aware Densification for Dynamic 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2606.23212v1 Announce Type: new Abstract: Despite modeling temporal motion, dynamic 3D Gaussian Splatting (3DGS) methods still inherit a static densification strategy that is ill-suited for dyna

Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2511.20889v2 Announce Type: replace Abstract: Test-time alignment (TTA) aims to adapt models to specific rewards during inference. However, existing methods tend to either under-optimise or over

Thalia: A Global, Multi-Modal Dataset for Volcanic Activity Monitoring

Model ReleasesDGX agent

arXiv:2505.17782v4 Announce Type: replace Abstract: Monitoring volcanic activity is of paramount importance to safeguarding lives, infrastructure, and ecosystems. However, only a small fraction of kno

The First Assessment of PhiSat-2 Imagery for Monocular Building Height Estimation

ResearchDGX agent

arXiv:2603.29245v3 Announce Type: replace Abstract: Monocular building height estimation from optical imagery is important for characterizing urban vertical structure, yet remains challenging due to t

The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production

ApplicationsDGX agent

arXiv:2606.22959v1 Announce Type: cross Abstract: Latent diffusion approaches to sign language production (SLP) rely on an initial stage that learns an encoding of sign pose sequences, enabling genera

The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction

Model ReleasesDGX agent

arXiv:2603.01250v3 Announce Type: replace Abstract: Breast cancer is the most frequently diagnosed malignancy among women worldwide and a leading cause of cancer-related mortality. Dynamic contrast-en

The Power of Light: Improving Synthetic-to-Real Domain Adaptation through Physically-Based Indirect Illumination

Model ReleasesDGX agent

arXiv:2606.22574v1 Announce Type: new Abstract: While synthetic data generation resolves the manual labeling bottleneck in computer vision, minimizing the syn-to-real domain gap requires optimizing re

The Scissors Effect: When Resize-Based Input Diversity Helps or Hurts Transfer Attacks

SafetyDGX agent

arXiv:2606.22516v1 Announce Type: cross Abstract: Input Diversity (DI), which applies random resizing and padding at each attack iteration, is a near-default ingredient of transfer-based adversarial a

The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection

ResearchDGX agent

arXiv:2606.21579v1 Announce Type: new Abstract: Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved signif

Three-Step Hierarchical Transformer for Multi-Pedestrian Trajectory Prediction

AgentsDGX agent

arXiv:2606.23058v1 Announce Type: new Abstract: Pedestrian trajectory prediction requires modeling temporal dynamics, multimodal cues, and social interactions in crowded environments. Existing methods

TIDY: Thermal Infrared Image Denoising via Wavelet Domain Entropy and Directional Stripe Index

ResearchDGX agent

arXiv:2606.19813v1 Announce Type: cross Abstract: Thermal infrared (TIR) imaging has been a popular choice for field robotics due to its robust perception capability under low light visual degradation

TooBad: Backdoor Diffusion Models with Ultra-Low Poison Rate and Imperceptible Trigger

ResearchDGX agent

arXiv:2606.23362v1 Announce Type: cross Abstract: Diffusion models (DMs), despite their impressive capabilities across a wide range of generative tasks, have been shown to be vulnerable to backdoor at

Topological summaries of fingerprint ridge patterns carry identity information

Model ReleasesDGX agent

arXiv:2606.22029v1 Announce Type: new Abstract: Fingerprints are the most widely deployed biometric. Verifying whether two impressions come from the same finger typically relies on minutiae, small lan

Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach

ApplicationsDGX agent

arXiv:2606.20886v1 Announce Type: new Abstract: As urban areas expand, automatic monitoring of parking lots becomes essential for efficient and sustainable cities. This work proposes a self-supervised

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning

ResearchDGX agent

arXiv:2606.22299v1 Announce Type: new Abstract: Idling Vehicle Detection (IVD) seeks to determine, at the final frame of a video clip, whether any vehicle is idling, meaning the vehicle is stationary

Towards Error-Free Long Video Generation

Model ReleasesDGX agent

arXiv:2606.22370v1 Announce Type: new Abstract: Recent advances in video generation have made minute-level synthesis possible; however, generating long videos remains challenging due to error accumula

Towards Practical Lossless Neural Compression for LiDAR Point Clouds

Model ReleasesDGX agent

arXiv:2603.25260v2 Announce Type: replace Abstract: LiDAR point clouds are fundamental to various applications, yet the extreme sparsity of high-precision geometric details hinders efficient context m

TraceMark-LDM: Authenticatable Watermarking for Latent Diffusion Models via Binary-Guided Rearrangement

ResearchDGX agent

arXiv:2503.23332v2 Announce Type: replace Abstract: Image generation algorithms are increasingly integral to diverse aspects of human society, driven by their practical applications. However, insuffic

Training-Free Semantic Correction for Autoregressive Visual Models

SafetyDGX agent

arXiv:2606.22550v1 Announce Type: new Abstract: Autoregressive visual models (AVMs) based on next-scale prediction have emerged as a prominent paradigm for image and video synthesis. However, decompos

Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories

Local AiDGX agent

arXiv:2606.22527v1 Announce Type: new Abstract: Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions a

Transfer learning-based method for automated ewaste recycling in smart cities

ApplicationsDGX agent

arXiv:2606.23286v1 Announce Type: new Abstract: Sorting a huge stream of waste accurately within a short period can be done with the support of digitalization, particularly Artificial Intelligence, in

Translating Inference-Time Control to Radiology Vision-Language Models: Activation Steering for Pneumonia Classification on Chest X-rays

Model ReleasesDGX agent

arXiv:2606.20852v1 Announce Type: new Abstract: Inference-time engineering can alter model behavior without fine-tuning. However, its utility for improving diagnostic performance in medical vision-lan

TriFlow: Generating Artist-Like 3D Mesh Topology via Nearest-Vertex Vector Fields

Local AiDGX agent

arXiv:2606.20131v2 Announce Type: replace Abstract: We present TriFlow, a new generative approach for producing compact 3D meshes with artist-like triangle topology directly from input geometry condit

TriMotion: Modality-Agnostic Camera Control for Video Generation

ResearchDGX agent

arXiv:2606.20774v1 Announce Type: new Abstract: Camera motion control is essential for directing viewpoint changes in generative systems. However, existing methods typically condition the generation p

Trustworthy MRI Reconstruction via Bayesian Uncertainty Quantification with Sparsity Prior Models

ResearchDGX agent

arXiv:2606.17343v2 Announce Type: replace Abstract: We propose a novel Bayesian framework for joint image reconstruction and uncertainty quantification from compressed sensing magnetic resonance imagi

TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation

SafetyDGX agent

arXiv:2606.13714v2 Announce Type: replace Abstract: Unsupervised video object-centric learning aims to decompose dynamic scenes into temporally persistent entity representations. Existing recurrent vi

UECP: Uncertainty-Enhanced Collaborative Perception

AgentsDGX agent

arXiv:2606.23046v1 Announce Type: new Abstract: Collaborative perception serves as a pivotal solution to enhance the perception capability of individual agents in autonomous driving, where a core chal

Uncertainty-Aware Domain Adaptation for Vitiligo Segmentation in Clinical Photographs

ResearchDGX agent

arXiv:2512.11791v2 Announce Type: replace Abstract: Accurately quantifying vitiligo extent in routine clinical photographs is crucial for longitudinal monitoring of treatment response. We propose a tr

UniSLAD: A Unified Framework for Structural and Logical Industrial Visual Anomaly Detection

ResearchDGX agent

arXiv:2606.20768v1 Announce Type: new Abstract: Visual anomaly detection is a fundamental task in industrial automation. While existing approaches have achieved notable progress in identifying structu

UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion

SafetyDGX agent

arXiv:2606.20971v1 Announce Type: new Abstract: We introduce UNITY, a Universal-to-Specialized adapter for efficient and scalable composite conditioning in diffusion based image generation. Unlike pri

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

Model ReleasesDGX agent

arXiv:2606.21661v1 Announce Type: new Abstract: Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist acros

UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation

ResearchDGX agent

arXiv:2606.23503v1 Announce Type: new Abstract: Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfer to Earth Observation (EO), where

Unlimited OCR Works

Model ReleasesDGX agent

arXiv:2606.23050v1 Announce Type: new Abstract: Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely held view is that employing a larg

Unmasking LAION-5B: Age, Gender, Race, and Emotion Biases in Large-Scale Image Datasets

ResearchDGX agent

arXiv:2606.23204v1 Announce Type: new Abstract: Large-scale image-text datasets, such as LAION-5B, are foundational to modern AI systems, yet their vast scale and uncurated nature raise significant co

Unsupervised Domain Adaptation for Sim-to-Real Object Pose Estimation with Contrastive Alignment and Pseudo-Label Refinement

SafetyDGX agent

arXiv:2606.21287v1 Announce Type: new Abstract: Unsupervised domain adaptation (UDA) enables robust transfer of knowledge from simulated to real environments while exploiting a subset of unlabeled tar

Unsupervised Susceptibility Distortion Correction of EPI without Calibration Scans via Image Translation-Based Registration

ResearchDGX agent

arXiv:2606.21588v1 Announce Type: cross Abstract: Functional magnetic resonance imaging (fMRI) utilizes echo-planar imaging (EPI) to capture blood-oxygen-level-dependent (BOLD) signals with high tempo

VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation

AgentsDGX agent

arXiv:2512.11061v2 Announce Type: replace Abstract: Generative video models, a leading approach to world modelling, face fundamental limitations. They often violate physical and logical rules, lack in

Venice-H1: Failure-Aware Query Re-Ranking with Multi-Scale Grid Signatures for Referring Image Segmentation

ResearchDGX agent

arXiv:2606.22546v1 Announce Type: new Abstract: Modern Referring Image Segmentation (RIS) systems generate multiple candidate masks per expression but rely on a simple heuristic--typically the argmax

Vera: A Layered Diffusion Model for Content-Preserving Video Editing

Model ReleasesDGX agent

arXiv:2606.23610v1 Announce Type: new Abstract: Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservation remains a core challenge: existin

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

Model ReleasesDGX agent

arXiv:2606.23543v1 Announce Type: cross Abstract: Scaling reinforcement learning for visual mathematical reasoning requires more than generating harder questions: as data volume grows, the reward labe

Video2Code: Generating Interactive Webpages from UI Videos via Action-Aware Revisit

ResearchDGX agent

arXiv:2606.20711v1 Announce Type: new Abstract: UI videos provide a natural input for generating interactive webpages, as they capture both webpage appearance and action-triggered state transitions. H

VideoAgent: All-in-One Framework for Video Understanding and Editing

Model ReleasesDGX agent

arXiv:2606.23327v1 Announce Type: new Abstract: Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-speci

VideoLatent: Video-Language Learning via Latent Self-Forcing

SafetyDGX agent

arXiv:2606.22870v1 Announce Type: new Abstract: Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal lar

Vision-language models for chest radiography do not always need the image

Model ReleasesDGX agent

arXiv:2606.17710v2 Announce Type: replace Abstract: Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That infe

Visual Geometry Transformer in the Wild: Distractor-Free 3D Reconstruction

ApplicationsDGX agent

arXiv:2606.22787v1 Announce Type: new Abstract: Current end-to-end multi-view 3D reconstruction methods achieve impressive results, but rely on a restrictive static assumption: the scenes is entire di

VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2606.21386v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) achieve state-of-the-art performance on many robotic manipulation tasks, yet they can still behave unpredictably

VolHuMe: a High-Resolution Large Scale Dataset of Volumetric Human Meshes

Model ReleasesDGX agent

arXiv:2606.23062v1 Announce Type: cross Abstract: We introduce VolHuMe, a dataset of high-quality 4D human scans captured with a state-of-the-art volumetric studio using 64 RGB and 32 depth cameras. V

VT-DUDA: Visual Token Conditioning for Diffusion-guided Unsupervised Domain Adaptation

TutorialsDGX agent

arXiv:2606.21700v1 Announce Type: new Abstract: Unsupervised domain adaptation (UDA) aims to learn a target-domain classifier from labeled source data and unlabeled target data under distribution shif

VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers

AgentsDGX agent

arXiv:2606.20728v1 Announce Type: new Abstract: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer

WebCryptoAgent: Agentic Crypto Trading with Web Informatics

AgentsDGX agent

arXiv:2601.04687v2 Announce Type: replace Abstract: Cryptocurrency trading increasingly depends on timely integration of heterogeneous web information and market microstructure signals to support shor

What if? Emulative Simulation with World Models for Situated Reasoning

SafetyDGX agent

arXiv:2603.06445v2 Announce Type: replace Abstract: Situated reasoning often relies on active exploration, yet in many real-world scenarios such exploration is infeasible due to physical constraints o

When Calibration Fails the Vulnerable Hospital: Federated Conformal Risk Control via Risk-Curve Shrinkage

ResearchDGX agent

arXiv:2606.20115v2 Announce Type: replace-cross Abstract: Conformal risk control (CRC) provides distribution-free guarantees on segmentation quality by calibrating a prediction-set threshold on held-o

When Confidence Lacks Concepts: Interpretable OOD Detection via Representation Perturbations

SafetyDGX agent

arXiv:2606.16196v2 Announce Type: replace-cross Abstract: Deep neural networks have achieved remarkable performance across medical imaging tasks, yet their tendency to overgeneralize under distributio

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR

ResearchDGX agent

arXiv:2606.22043v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is increasingly applied to large vision-language models (LVLMs), yet outcome-only optimization c

WildBox: A Dataset and Benchmark for Aerial Monocular 3D Detection of African Savanna Wildlife

Model ReleasesDGX agent

arXiv:2606.21309v1 Announce Type: new Abstract: We introduce WildBox, a dataset and benchmark for monocular 3D detection of wildlife from drone video, comprising 237,505 3D bounding box annotations ac

World Action Models: A Survey

ResearchDGX agent

arXiv:2606.20781v1 Announce Type: cross Abstract: World Action Models (WAMs) are embodied predictive-action models that make a forecast of the future available to action. Recent WAMs repurpose large v

XmoPipe: A Pipeline for Large-Scale In-the-Wild Human Motion Dataset Construction

ResearchDGX agent

arXiv:2606.20731v1 Announce Type: new Abstract: Large-scale human motion datasets are essential for training robust motion models for analysis, synthesis, and understanding. While marker-based motion

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Model ReleasesDGX agent

arXiv:2511.22699v4 Announce Type: replace Abstract: The landscape of high-performance image generation models is currently dominated by proprietary systems, such as Nano Banana Pro and Seedream 4.0. L

Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization

Model ReleasesDGX agent

arXiv:2606.21861v1 Announce Type: new Abstract: Automated classroom engagement recognition holds substantial promise for scalable learning analytics, yet the suitability of modern Vision-Language Mode

← Previous
1…8283848586…211
Next →