AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

TooBad: Backdoor Diffusion Models with Ultra-Low Poison Rate and Imperceptible Trigger

DGX agent

arXiv:2606.23362v1 Announce Type: cross Abstract: Diffusion models (DMs), despite their impressive capabilities across a wide range of generative tasks, have been shown to be vulnerable to backdoor at

researcharxiv-cs-cv
23 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Topological summaries of fingerprint ridge patterns carry identity information

DGX agent

arXiv:2606.22029v1 Announce Type: new Abstract: Fingerprints are the most widely deployed biometric. Verifying whether two impressions come from the same finger typically relies on minutiae, small lan

model-releasesarxiv-cs-cv
23 Jun 2026
Applications

Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach

DGX agent

arXiv:2606.20886v1 Announce Type: new Abstract: As urban areas expand, automatic monitoring of parking lots becomes essential for efficient and sustainable cities. This work proposes a self-supervised

applicationsarxiv-cs-cv
23 Jun 2026
Research

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning

DGX agent

arXiv:2606.22299v1 Announce Type: new Abstract: Idling Vehicle Detection (IVD) seeks to determine, at the final frame of a video clip, whether any vehicle is idling, meaning the vehicle is stationary

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Towards Error-Free Long Video Generation

DGX agent

arXiv:2606.22370v1 Announce Type: new Abstract: Recent advances in video generation have made minute-level synthesis possible; however, generating long videos remains challenging due to error accumula

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Towards Practical Lossless Neural Compression for LiDAR Point Clouds

DGX agent

arXiv:2603.25260v2 Announce Type: replace Abstract: LiDAR point clouds are fundamental to various applications, yet the extreme sparsity of high-precision geometric details hinders efficient context m

model-releasesarxiv-cs-cv
23 Jun 2026
Research

TraceMark-LDM: Authenticatable Watermarking for Latent Diffusion Models via Binary-Guided Rearrangement

DGX agent

arXiv:2503.23332v2 Announce Type: replace Abstract: Image generation algorithms are increasingly integral to diverse aspects of human society, driven by their practical applications. However, insuffic

researcharxiv-cs-cv
23 Jun 2026
Safety

Training-Free Semantic Correction for Autoregressive Visual Models

DGX agent

arXiv:2606.22550v1 Announce Type: new Abstract: Autoregressive visual models (AVMs) based on next-scale prediction have emerged as a prominent paradigm for image and video synthesis. However, decompos

safetyarxiv-cs-cv
23 Jun 2026
Local Ai

Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories

DGX agent

arXiv:2606.22527v1 Announce Type: new Abstract: Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions a

local-aiarxiv-cs-cv
23 Jun 2026
Applications

Transfer learning-based method for automated ewaste recycling in smart cities

DGX agent

arXiv:2606.23286v1 Announce Type: new Abstract: Sorting a huge stream of waste accurately within a short period can be done with the support of digitalization, particularly Artificial Intelligence, in

applicationsarxiv-cs-cv
23 Jun 2026
Model Releases

Translating Inference-Time Control to Radiology Vision-Language Models: Activation Steering for Pneumonia Classification on Chest X-rays

DGX agent

arXiv:2606.20852v1 Announce Type: new Abstract: Inference-time engineering can alter model behavior without fine-tuning. However, its utility for improving diagnostic performance in medical vision-lan

model-releasesarxiv-cs-cv
23 Jun 2026
Local Ai

TriFlow: Generating Artist-Like 3D Mesh Topology via Nearest-Vertex Vector Fields

DGX agent

arXiv:2606.20131v2 Announce Type: replace Abstract: We present TriFlow, a new generative approach for producing compact 3D meshes with artist-like triangle topology directly from input geometry condit

local-aiarxiv-cs-cv
23 Jun 2026
Research

TriMotion: Modality-Agnostic Camera Control for Video Generation

DGX agent

arXiv:2606.20774v1 Announce Type: new Abstract: Camera motion control is essential for directing viewpoint changes in generative systems. However, existing methods typically condition the generation p

researcharxiv-cs-cv
23 Jun 2026
Research

Trustworthy MRI Reconstruction via Bayesian Uncertainty Quantification with Sparsity Prior Models

DGX agent

arXiv:2606.17343v2 Announce Type: replace Abstract: We propose a novel Bayesian framework for joint image reconstruction and uncertainty quantification from compressed sensing magnetic resonance imagi

researcharxiv-cs-cv
23 Jun 2026
Safety

TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation

DGX agent

arXiv:2606.13714v2 Announce Type: replace Abstract: Unsupervised video object-centric learning aims to decompose dynamic scenes into temporally persistent entity representations. Existing recurrent vi

safetyarxiv-cs-cv
23 Jun 2026
Agents

UECP: Uncertainty-Enhanced Collaborative Perception

DGX agent

arXiv:2606.23046v1 Announce Type: new Abstract: Collaborative perception serves as a pivotal solution to enhance the perception capability of individual agents in autonomous driving, where a core chal

agentsarxiv-cs-cv
23 Jun 2026
Research

Uncertainty-Aware Domain Adaptation for Vitiligo Segmentation in Clinical Photographs

DGX agent

arXiv:2512.11791v2 Announce Type: replace Abstract: Accurately quantifying vitiligo extent in routine clinical photographs is crucial for longitudinal monitoring of treatment response. We propose a tr

researcharxiv-cs-cv
23 Jun 2026
Research

UniSLAD: A Unified Framework for Structural and Logical Industrial Visual Anomaly Detection

DGX agent

arXiv:2606.20768v1 Announce Type: new Abstract: Visual anomaly detection is a fundamental task in industrial automation. While existing approaches have achieved notable progress in identifying structu

researcharxiv-cs-cv
23 Jun 2026
Safety

UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion

DGX agent

arXiv:2606.20971v1 Announce Type: new Abstract: We introduce UNITY, a Universal-to-Specialized adapter for efficient and scalable composite conditioning in diffusion based image generation. Unlike pri

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

DGX agent

arXiv:2606.21661v1 Announce Type: new Abstract: Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist acros

model-releasesarxiv-cs-cv
23 Jun 2026
Research

UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation

DGX agent

arXiv:2606.23503v1 Announce Type: new Abstract: Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfer to Earth Observation (EO), where

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Unlimited OCR Works

DGX agent

arXiv:2606.23050v1 Announce Type: new Abstract: Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely held view is that employing a larg

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Unmasking LAION-5B: Age, Gender, Race, and Emotion Biases in Large-Scale Image Datasets

DGX agent

arXiv:2606.23204v1 Announce Type: new Abstract: Large-scale image-text datasets, such as LAION-5B, are foundational to modern AI systems, yet their vast scale and uncurated nature raise significant co

researcharxiv-cs-cv
23 Jun 2026
Safety

Unsupervised Domain Adaptation for Sim-to-Real Object Pose Estimation with Contrastive Alignment and Pseudo-Label Refinement

DGX agent

arXiv:2606.21287v1 Announce Type: new Abstract: Unsupervised domain adaptation (UDA) enables robust transfer of knowledge from simulated to real environments while exploiting a subset of unlabeled tar

safetyarxiv-cs-cv
23 Jun 2026
Research

Unsupervised Susceptibility Distortion Correction of EPI without Calibration Scans via Image Translation-Based Registration

DGX agent

arXiv:2606.21588v1 Announce Type: cross Abstract: Functional magnetic resonance imaging (fMRI) utilizes echo-planar imaging (EPI) to capture blood-oxygen-level-dependent (BOLD) signals with high tempo

researcharxiv-cs-cv
23 Jun 2026
Agents

VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation

DGX agent

arXiv:2512.11061v2 Announce Type: replace Abstract: Generative video models, a leading approach to world modelling, face fundamental limitations. They often violate physical and logical rules, lack in

agentsarxiv-cs-cv
23 Jun 2026
Research

Venice-H1: Failure-Aware Query Re-Ranking with Multi-Scale Grid Signatures for Referring Image Segmentation

DGX agent

arXiv:2606.22546v1 Announce Type: new Abstract: Modern Referring Image Segmentation (RIS) systems generate multiple candidate masks per expression but rely on a simple heuristic--typically the argmax

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Vera: A Layered Diffusion Model for Content-Preserving Video Editing

DGX agent

arXiv:2606.23610v1 Announce Type: new Abstract: Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservation remains a core challenge: existin

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

DGX agent

arXiv:2606.23543v1 Announce Type: cross Abstract: Scaling reinforcement learning for visual mathematical reasoning requires more than generating harder questions: as data volume grows, the reward labe

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Video2Code: Generating Interactive Webpages from UI Videos via Action-Aware Revisit

DGX agent

arXiv:2606.20711v1 Announce Type: new Abstract: UI videos provide a natural input for generating interactive webpages, as they capture both webpage appearance and action-triggered state transitions. H

researcharxiv-cs-cv
23 Jun 2026
Model Releases

VideoAgent: All-in-One Framework for Video Understanding and Editing

DGX agent

arXiv:2606.23327v1 Announce Type: new Abstract: Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-speci

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

VideoLatent: Video-Language Learning via Latent Self-Forcing

DGX agent

arXiv:2606.22870v1 Announce Type: new Abstract: Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal lar

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

Vision-language models for chest radiography do not always need the image

DGX agent

arXiv:2606.17710v2 Announce Type: replace Abstract: Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That infe

model-releasesarxiv-cs-cv
23 Jun 2026
Applications

Visual Geometry Transformer in the Wild: Distractor-Free 3D Reconstruction

DGX agent

arXiv:2606.22787v1 Announce Type: new Abstract: Current end-to-end multi-view 3D reconstruction methods achieve impressive results, but rely on a restrictive static assumption: the scenes is entire di

applicationsarxiv-cs-cv
23 Jun 2026
Applications

VLA-FAIL: Efficient Task Failure Detection for Finetuned Vision-Language-Action Models

DGX agent

arXiv:2606.21386v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) achieve state-of-the-art performance on many robotic manipulation tasks, yet they can still behave unpredictably

applicationsarxiv-cs-cv
23 Jun 2026
Model Releases

VolHuMe: a High-Resolution Large Scale Dataset of Volumetric Human Meshes

DGX agent

arXiv:2606.23062v1 Announce Type: cross Abstract: We introduce VolHuMe, a dataset of high-quality 4D human scans captured with a state-of-the-art volumetric studio using 64 RGB and 32 depth cameras. V

model-releasesarxiv-cs-cv
23 Jun 2026
Tutorials

VT-DUDA: Visual Token Conditioning for Diffusion-guided Unsupervised Domain Adaptation

DGX agent

arXiv:2606.21700v1 Announce Type: new Abstract: Unsupervised domain adaptation (UDA) aims to learn a target-domain classifier from labeled source data and unlabeled target data under distribution shif

tutorialsarxiv-cs-cv
23 Jun 2026
Agents

VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers

DGX agent

arXiv:2606.20728v1 Announce Type: new Abstract: Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer

agentsarxiv-cs-cv
23 Jun 2026
Agents

WebCryptoAgent: Agentic Crypto Trading with Web Informatics

DGX agent

arXiv:2601.04687v2 Announce Type: replace Abstract: Cryptocurrency trading increasingly depends on timely integration of heterogeneous web information and market microstructure signals to support shor

agentsarxiv-cs-cv
23 Jun 2026
Safety

What if? Emulative Simulation with World Models for Situated Reasoning

DGX agent

arXiv:2603.06445v2 Announce Type: replace Abstract: Situated reasoning often relies on active exploration, yet in many real-world scenarios such exploration is infeasible due to physical constraints o

safetyarxiv-cs-cv
23 Jun 2026
Research

When Calibration Fails the Vulnerable Hospital: Federated Conformal Risk Control via Risk-Curve Shrinkage

DGX agent

arXiv:2606.20115v2 Announce Type: replace-cross Abstract: Conformal risk control (CRC) provides distribution-free guarantees on segmentation quality by calibrating a prediction-set threshold on held-o

researcharxiv-cs-cv
23 Jun 2026
Safety

When Confidence Lacks Concepts: Interpretable OOD Detection via Representation Perturbations

DGX agent

arXiv:2606.16196v2 Announce Type: replace-cross Abstract: Deep neural networks have achieved remarkable performance across medical imaging tasks, yet their tendency to overgeneralize under distributio

safetyarxiv-cs-cv
23 Jun 2026
Research

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR

DGX agent

arXiv:2606.22043v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is increasingly applied to large vision-language models (LVLMs), yet outcome-only optimization c

researcharxiv-cs-cv
23 Jun 2026
Model Releases

WildBox: A Dataset and Benchmark for Aerial Monocular 3D Detection of African Savanna Wildlife

DGX agent

arXiv:2606.21309v1 Announce Type: new Abstract: We introduce WildBox, a dataset and benchmark for monocular 3D detection of wildlife from drone video, comprising 237,505 3D bounding box annotations ac

model-releasesarxiv-cs-cv
23 Jun 2026
Research

World Action Models: A Survey

DGX agent

arXiv:2606.20781v1 Announce Type: cross Abstract: World Action Models (WAMs) are embodied predictive-action models that make a forecast of the future available to action. Recent WAMs repurpose large v

researcharxiv-cs-cv
23 Jun 2026
Research

XmoPipe: A Pipeline for Large-Scale In-the-Wild Human Motion Dataset Construction

DGX agent

arXiv:2606.20731v1 Announce Type: new Abstract: Large-scale human motion datasets are essential for training robust motion models for analysis, synthesis, and understanding. While marker-based motion

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

DGX agent

arXiv:2511.22699v4 Announce Type: replace Abstract: The landscape of high-performance image generation models is currently dominated by proprietary systems, such as Nano Banana Pro and Seedream 4.0. L

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization

DGX agent

arXiv:2606.21861v1 Announce Type: new Abstract: Automated classroom engagement recognition holds substantial promise for scalable learning analytics, yet the suitability of modern Vision-Language Mode

model-releasesarxiv-cs-cv
23 Jun 2026
← Previous
1…103104105106107…263
Next →