AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
19 May 2026

Accelerating Rectified Flow Models via Trajectory-Aware Caching

Model ReleasesDGX agent

arXiv:2605.16789v1 Announce Type: new Abstract: Diffusion and rectified flow (RF) models generate high-fidelity images and videos, but their iterative velocity-field evaluations are computationally ex

Adaptive double-phase Rudin--Osher--Fatemi denoising model

ResearchDGX agent

arXiv:2510.04382v2 Announce Type: replace-cross Abstract: Even though more than 30 years have passed since the seminal Rudin--Osher--Fatemi (ROF) paper on total variation (TV) denoising, it remains re

Adaptive Fused Prior Transfer for Controllable Generative Image Compression

Model ReleasesDGX agent

arXiv:2605.16817v1 Announce Type: cross Abstract: Learned image compression has achieved competitive rate-distortion performance, but very-low-bitrate reconstruction remains difficult because the tran


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory

Model ReleasesDGX agent

arXiv:2605.18733v1 Announce Type: new Abstract: Autoregressive video generation has improved rapidly in visual fidelity and interactivity, but it still suffers from long-term inconsistency and memory

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

Model ReleasesDGX agent

arXiv:2605.17583v1 Announce Type: new Abstract: While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the

AIM: Adversarial Information Masking for Faithfulness Evaluation of Saliency Maps

SafetyDGX agent

arXiv:2605.16905v1 Announce Type: cross Abstract: Post-hoc saliency methods are widely used to interpret deep neural networks, but their faithfulness is difficult to evaluate reliably. Existing evalua

Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey

SafetyDGX agent

arXiv:2505.17352v2 Announce Type: replace Abstract: Diffusion models have become a central paradigm for image and multimodal generation, yet their deployment raises persistent questions about alignmen

An Efficient Streaming Video Understanding Framework with Agentic Control

SafetyDGX agent

arXiv:2605.17921v1 Announce Type: new Abstract: Streaming video requires handling dynamic information density under strict latency budgets. Yet, existing methods typically employ static strategies, su

Articulation in Prime: Primitive-Based Articulated Object Understanding from a Single Casual Video

ApplicationsDGX agent

arXiv:2605.18645v1 Announce Type: new Abstract: Retrieving the 3D kinematics of articulated objects from monocular video is a fundamental challenge in computer vision. Existing methods rely on complex

ArtMesh: Part-Aware Articulated Mesh Fields with Motion-Consistent Dynamics

Model ReleasesDGX agent

arXiv:2605.16582v1 Announce Type: new Abstract: We present ArtMesh, a mesh-native method for reconstructing articulated objects explicitly as connected triangle meshes with per-part rigid motion from

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents

ResearchDGX agent

arXiv:2605.17933v1 Announce Type: new Abstract: Vision-language model (VLM) agents increasingly rely on memory-augmented reinforcement learning to reuse experience across long-horizon tasks, yet most

AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling

Local AiDGX agent

arXiv:2605.16649v1 Announce Type: new Abstract: Recent diffusion-based video generators have achieved remarkable visual fidelity and prompt controllability, yet scaling them to ultra-high-resolution (

Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion

Local AiDGX agent

arXiv:2605.16579v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion is a powerful paradigm for streaming and interactive video generation. However, its reliance on softmax self-attenti

Attention-Aware Transformer-Based Aggregation Network for Video Periocular Recognition

ResearchDGX agent

arXiv:2605.16550v1 Announce Type: new Abstract: Video periocular recognition is the task of recognizing an individual's identity based on the region around an individual's eyes. The periocular area is

Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring

Model ReleasesDGX agent

arXiv:2605.16386v1 Announce Type: new Abstract: Multimodal large language models (LLMs) are increasingly explored as automated evaluators in clinical settings, yet their scoring behavior on ordinal cl

Aurora: Unified Video Editing with a Tool-Using Agent

Local AiDGX agent

arXiv:2605.18748v1 Announce Type: new Abstract: Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and ref

Axial-Relation Guided Fusion State Space Model for Optical-Elevation Sensing Image Segmentation

Local AiDGX agent

arXiv:2605.16768v1 Announce Type: new Abstract: Semantic segmentation of multi-source remote sensing images is a fundamental task for Earth observation applications. Existing methods often struggle wi

Benchmarking Recurrent Event-Based Object Detection for Industrial Multi-Class Recognition on MTevent

Model ReleasesDGX agent

arXiv:2603.21787v2 Announce Type: replace Abstract: Event cameras are attractive for industrial robotics because they provide high temporal resolution, high dynamic range, and reduced motion blur. How

Benchmarking transferability of SSL pretraining to same and different modality segmentation tasks

SafetyDGX agent

arXiv:2605.18491v1 Announce Type: new Abstract: Methods: Nine SSL methods spanning four pretext-task families were pretrained from scratch using the same 10{,}412 3D CT scans (1.89~M 2D axial slices)

Best Segmentation Buddies for Image-Shape Correspondence

ResearchDGX agent

arXiv:2605.18193v1 Announce Type: new Abstract: Finding correspondences is a fundamental and extensively researched problem in computer vision and graphics. In this work, we examine the underexplored

Better Together: Evaluating the Complementarity of Earth Embedding Models

ResearchDGX agent

arXiv:2605.18667v1 Announce Type: new Abstract: Earth embedding models transform Earth observation data into embeddings uniquely tied to locations on the Earth's surface. These models are typically ev

Beyond Detection: A Structure-Aware Framework for Scene Text Tracking

Model ReleasesDGX agent

arXiv:2605.17270v1 Announce Type: new Abstract: Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Althou

Beyond Euclidean Prototypes: Spectral Disentanglement and Geodesic Matching for Few-Shot Medical Image Segmentation

ResearchDGX agent

arXiv:2605.17904v1 Announce Type: new Abstract: Few-Shot Medical Image Segmentation (FSMIS) aims to delineate novel anatomical targets from one or a few annotated support images, addressing the annota

Beyond MMSE: Enhancing PnP Restoration with ProxiMAP

ResearchDGX agent

arXiv:2605.16396v1 Announce Type: new Abstract: Plug-and-Play (PnP) methods have become standard tools for solving imaging inverse problems by replacing the intractable maximum a posteriori (MAP) deno

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation

ResearchDGX agent

arXiv:2601.01593v2 Announce Type: replace Abstract: Manual font design is an intricate process that transforms a stylistic visual concept into a coherent glyph set. This challenge persists in automate

Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers

Model ReleasesDGX agent

arXiv:2605.16949v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) demonstrate that aligning noisy latent states with well-trained semantic features-as pioneered by Repre

BIDO: A Biometric Identity Online Authentication Framework

SafetyDGX agent

arXiv:2605.16908v1 Announce Type: cross Abstract: Security systems demand continuous, cryptograph- ically robust identity verification without requiring subjects to carry physical tokens, smart cards,

Bio-Inspired Event-Based Visual Servoing for Ground Robots

AgentsDGX agent

arXiv:2603.23672v2 Announce Type: replace-cross Abstract: Biological sensory systems are inherently adaptive, filtering out constant stimuli and prioritizing relative changes, likely enhancing computa

Brain-inspired spike-timing plasticity for reliable label-efficient event-camera vision

Model ReleasesDGX agent

arXiv:2605.17686v1 Announce Type: new Abstract: Deploying event-camera object detectors is constrained by per-frame labeling requirements and GPU compute demands. This work introduces three local spik

Breaking Annotation Barriers: Generalized Video Quality Assessment via Ranking-based Self-Supervision

Model ReleasesDGX agent

arXiv:2505.03631v4 Announce Type: replace Abstract: Video quality assessment (VQA) is essential for quantifying perceptual quality in various video processing workflows, spanning from camera capture s

Bridging Data Trials and Task Barriers: A Unified Framework for Sketch Biometric Identification

Model ReleasesDGX agent

arXiv:2605.17367v1 Announce Type: new Abstract: Different from existing cross-modality identification tasks (e.g., heterogeneous face recognition, sketch re-identification, etc.), we introduce a novel

Bridging the Intention-Expression Gap: Aligning Multi-Dimensional Preferences via Hierarchical Relevance Feedback in Text-to-Image Diffusion

SafetyDGX agent

arXiv:2603.14936v3 Announce Type: replace Abstract: Users often possess a clear visual intent but struggle to articulate it precisely in language. This intention-expression gap makes aligning generate

Bridging the Modality Bottleneck in Pathology MIL through Virtual Molecular Staining

ResearchDGX agent

arXiv:2605.16392v1 Announce Type: cross Abstract: Multiple instance learning (MIL) is the dominant framework for whole-slide image analysis in computational pathology, typically combining a frozen pat

Bundle Adjustment in the Eager Mode

HardwareDGX agent

arXiv:2409.12190v4 Announce Type: replace-cross Abstract: Bundle adjustment (BA) is a critical technique in various robotic applications such as simultaneous localization and mapping (SLAM), augmented

CAB: Accelerating Flow and Diffusion Sampling via Rectification and Corrected Adams-Bashforth

ResearchDGX agent

arXiv:2605.16736v1 Announce Type: new Abstract: Flow and diffusion models achieve high-fidelity, high-resolution image synthesis, but often require many function evaluations (NFEs) at sampling time. E

CADS: Conformal Adaptive Decision System for Cost-Efficient Image Classification

ResearchDGX agent

arXiv:2605.16401v1 Announce Type: new Abstract: While high-capacity AI models have advanced state-of-the-art performance, their practical deployment is often hindered by high inference costs, environm

Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate

Model ReleasesDGX agent

arXiv:2605.18754v1 Announce Type: new Abstract: Multiview 3D evaluation assumes that the images being scored are observations of one static 3D scene. This assumption can fail in NVS and sparse-view re

CanViT: Toward Active-Vision Foundation Models

Local AiDGX agent

arXiv:2603.22570v2 Announce Type: replace Abstract: Active computer vision promises efficient, biologically plausible perception through sequential, localized glimpses, but lacks scalable general-purp

CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model

ResearchDGX agent

arXiv:2605.16901v1 Announce Type: new Abstract: Segment Anything Models (SAMs) are extensively used in computer vision for universal image segmentation, but deploying them on resource-constrained devi

CATRF: Codec-Adaptive TriPlane Radiance Fields for Volumetric Content Delivery

ResearchDGX agent

arXiv:2605.18054v1 Announce Type: cross Abstract: Volumetric media promises next-generation content delivery applications, but its bandwidth demand remains a key bottleneck. Implicit and hybrid volume

Causal Attribution via Activation Patching

Local AiDGX agent

arXiv:2603.13652v2 Announce Type: replace Abstract: Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-l

ChronoSC: Task-Oriented Semantic Communication via Temporal-to-Color Encoding

ResearchDGX agent

arXiv:2605.16388v1 Announce Type: new Abstract: Semantic communication (SC) aims to reduce transmission overhead by conveying task-relevant information rather than raw data. However, existing SC appro

CineMatte: Background Matting for Virtual Production and Beyond

ApplicationsDGX agent

arXiv:2605.18328v1 Announce Type: new Abstract: LED Virtual Production (VP) uses large LED volumes to render backgrounds in real time, enabling in-camera visual effects but making post-shot changes la

CLEAR-HPV: Interpretable Concept Discovery for HPV-Associated Morphology in Whole-Slide Histologyhttps://arxiv.org/submit/7596892/preview

ResearchDGX agent

arXiv:2602.05126v2 Announce Type: replace Abstract: Human papillomavirus (HPV) status is a critical determinant of prognosis and treatment response in head and neck and cervical cancers. Although atte

CMAG: Concept-Scaffolded Retrieval for Marketplace Avatar Generation

Local AiDGX agent

arXiv:2605.18680v1 Announce Type: new Abstract: Metaverse platforms rely on creator-driven marketplaces where avatars are assembled from discrete, taxonomy-labeled 3D assets (e.g., tops, bottoms, shoe

Coarse Semantic Injection for LLM-Conditioned Structured Indoor Prediction

ResearchDGX agent

arXiv:2605.16832v1 Announce Type: new Abstract: Large language models (LLMs) have recently been used as structured decoders for indoor understanding from 3D point-token inputs. However, point cloud en

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis

Model ReleasesDGX agent

arXiv:2605.18451v1 Announce Type: new Abstract: Designing realistic and functional 3D indoor rooms is essential for a wide range of applications, including interior design, virtual reality, gaming, an

CogBlender: Towards Continuous Cognitive Intervention in Text-to-Image Generation

ResearchDGX agent

arXiv:2603.09286v2 Announce Type: replace Abstract: Beyond conveying semantic information, images also possess cognitive properties that elicit specific psychological responses from viewers, such as m

Collaborative Learning for Semi-Supervised LiDAR Semantic Segmentation

SafetyDGX agent

arXiv:2605.17135v1 Announce Type: new Abstract: Annotating large-scale LiDAR point clouds for 3D semantic segmentation is costly and time-consuming, which motivates the use of semi-supervised learning

Collision-Resistant Single-Pass Method for Unsupervised Fine-Grained Image Hashing

TutorialsDGX agent

arXiv:2605.18288v1 Announce Type: new Abstract: Unsupervised fine-grained image hashing aims to learn compact binary codes that preserve subtle visual differences among highly similar instances withou

Color as the Impetus: Transforming Few-Shot Learner

ResearchDGX agent

arXiv:2507.22136v3 Announce Type: replace Abstract: Humans possess innate meta-learning capabilities, partly attributable to their exceptional color perception. In this paper, we pioneer an innovative

CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects

Model ReleasesDGX agent

arXiv:2604.02060v2 Announce Type: replace Abstract: When told to 'cut the cake,' a robot must choose the knife over nearby scissors, despite both objects affording the same cutting function. In real-w

Compositional Adversarial Training for Robust Visual Watermarking

ResearchDGX agent

arXiv:2605.16720v1 Announce Type: new Abstract: Robust watermarking is typically trained with random post-processing augmentation, but random sampling under-covers the combinatorial space of realistic

Concepts Worth Having: Refining VLM-Guided Concept Bottleneck Models with Minimal Annotations

ResearchDGX agent

arXiv:2605.16405v1 Announce Type: new Abstract: Concept-bottleneck models (CBMs) are neural classifiers that compute predictions from high-level concepts extracted from the input. CBMs ensure stakehol

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

Model ReleasesDGX agent

arXiv:2508.04227v2 Announce Type: replace Abstract: Vision-language models (VLMs) and the recent surge of Multimodal Large Language Models (MLLMs) have revolutionized artificial intelligence with unpr

Contrastive-SDXL: Annotation-Preserving Night-Time Augmentation for Pedestrian Detection

SafetyDGX agent

arXiv:2605.16406v1 Announce Type: new Abstract: Night-time pedestrian detection remains challenging because labelled night-time data are limited and large illumination differences make daytime-only tr

Controlla: Learning Controllability via Graph-Constrained Latent Geometry

Model ReleasesDGX agent

arXiv:2605.16603v1 Announce Type: new Abstract: Controllable multimodal generation is commonly formulated as an inference-time conditioning problem using prompts, guidance, or auxiliary modules. While

Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities

SafetyDGX agent

arXiv:2605.16889v1 Announce Type: new Abstract: Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbala

Counting Machine Parts

ResearchDGX agent

arXiv:2605.17952v1 Announce Type: new Abstract: Counting objects in an image is a task applicable across many domains. For instance, crowd counting, inventory counting, and cell counting have been the

Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models

SafetyDGX agent

arXiv:2605.18413v1 Announce Type: new Abstract: Automated structural health monitoring is essential to prevent catastrophic infrastructure failures. Precise, pixel-level defect segmentation is needed

← Previous
1…126127128129130…211
Next →