AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
29 May 2026

Fairness Beyond Demographics: Optimizing Performance Across Appearance-Based Hidden Cohorts in Medical Imaging

Model ReleasesDGX agent

arXiv:2605.29827v1 Announce Type: new Abstract: Medical image analysis models can exhibit performance disparities across patient subgroups, threatening clinical safety and fairness. Existing methods t

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection

SafetyDGX agent

arXiv:2605.30062v1 Announce Type: new Abstract: The development of generative artificial intelligence technologies has propelled the visual realism of synthetic images to an unprecedented level. Altho

FedSmoothLoRA: Toward Smoother and Faster Convergence in Federated Low-Rank Adaptation

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.29460v1 Announce Type: new Abstract: Federated fine-tuning of foundation models with Low-Rank Adaptation (LoRA) provides an efficient solution for reducing communication and computation cos

Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using Language

SafetyDGX agent

arXiv:2605.29793v1 Announce Type: new Abstract: Given an untrimmed video and a sentence query, video moment retrieval using language (VMR) aims to locate a target query-relevant moment. Since the untr

FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation

SafetyDGX agent

arXiv:2605.29461v1 Announce Type: new Abstract: LLM-conditioned segmentation has recently advanced rapidly by coupling large language models with iterative mask generation frameworks. However, we iden

FreeForm: Reduced-Order Deformable Simulation from Particle-Based Skinning Eigenmodes

ResearchDGX agent

arXiv:2605.29318v1 Announce Type: cross Abstract: We present a novel formulation for mesh-free, reduced-order simulation of deformable hyperelastic objects. Existing work in reduced-order elastodynami

From General Vision to Reliable Traversability Estimation: Adapting Vision Foundation Models for Unstructured Outdoor Environments

SafetyDGX agent

arXiv:2605.29565v1 Announce Type: new Abstract: Vision-based approaches have become the dominant paradigm for traversability estimation in unstructured outdoor environments, typically adapting vision

FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views

AgentsDGX agent

arXiv:2605.29997v1 Announce Type: new Abstract: We present FRUC, a feed-forward 3D Gaussian splatting framework for dynamic scene reconstruction from uncalibrated collaborative driving views. Existing

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation

SafetyDGX agent

arXiv:2605.30083v1 Announce Type: new Abstract: Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on

Gaga: Group Any Gaussians via 3D-aware Memory Bank

ApplicationsDGX agent

arXiv:2404.07977v4 Announce Type: replace Abstract: We introduce Gaga, a framework that reconstructs and segments open-world 3D scenes by leveraging inconsistent 2D masks predicted by zero-shot class-

GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation

SafetyDGX agent

arXiv:2605.28995v1 Announce Type: new Abstract: Recent approaches integrating vision-language models (VLMs) as prompt encoders for generative model conditioning typically rely on expensive end-to-end

GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

SafetyDGX agent

arXiv:2602.17200v2 Announce Type: replace Abstract: Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In th

GenClaw: Code-Driven Agentic Image Generation

AgentsDGX agent

arXiv:2605.30248v1 Announce Type: new Abstract: Image generation models have evolved from text-conditioned pixel synthesis toward multimodal agents endowed with visual comprehension and tool invocatio

GenEraser: Generalizable Video Object Removal via Balanced Text-Mask Guidance and Decoupled Locator-Preserver

Model ReleasesDGX agent

arXiv:2605.30045v1 Announce Type: new Abstract: Video object removal frequently struggles to simultaneously eliminate target objects and their associated physical effects (e.g., smoke, reflections, li

GeoMag: Geometric-Aware Video Motion Magnification via State Space Model

ApplicationsDGX agent

arXiv:2605.29762v1 Announce Type: new Abstract: Video Motion Magnification (VMM) reveals imperceptible dynamics but often suffers from structural inconsistencies under complex geometric transformation

Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning

SafetyDGX agent

arXiv:2605.29661v1 Announce Type: new Abstract: Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object

Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence

TutorialsDGX agent

arXiv:2605.30093v1 Announce Type: new Abstract: Foundation features from self-supervised vision models and text-to-image diffusion models have proven effective for semantic correspondence estimation.

GeRaF: Neural Geometry Reconstruction from Radio Frequency Signals

ApplicationsDGX agent

arXiv:2605.29097v1 Announce Type: new Abstract: GeRaF is the first method to use neural implicit learning for near-range 3D geometry reconstruction from radio frequency (RF) signals. Unlike RGB or LiD

Getting to the Point: Pointing Improves LVLMs at Counting

ResearchDGX agent

arXiv:2603.21746v2 Announce Type: replace Abstract: Pointing-based methods decompose complex tasks as sequential grounding and reasoning steps. Given a query, the model first grounds the relevant obje

GMOS: Grounding Moving Object Segmentation in 3D Space and Time

ApplicationsDGX agent

arXiv:2605.30352v1 Announce Type: new Abstract: Moving Object Segmentation (MOS) aims to discover, segment, and track objects that move independently of the camera. Current MOS methods, however, exhib

Grounded 3D-Aware Spatial Vision-Language Modeling

SafetyDGX agent

arXiv:2605.30307v1 Announce Type: new Abstract: We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding,

Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization

SafetyDGX agent

arXiv:2605.29198v1 Announce Type: new Abstract: Group-advantage-based reinforcement learning methods, such as GRPO and DAPO, have demonstrated strong performance across diverse domains, including math

HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis

ResearchDGX agent

arXiv:2508.10566v3 Announce Type: replace Abstract: Audio-driven talking head generation faces a fundamental trade-off between personalization and generalization, limiting its practical application. I

How to Relieve Distribution Shifts in Semantic Segmentation for Off-Road Environments

AgentsDGX agent

arXiv:2605.29599v1 Announce Type: cross Abstract: Semantic segmentation is crucial for autonomous navigation in off-road environments, enabling precise classification of surroundings to identify trave

Improving Adversarial Robustness of Attribution via Implicit Regularization

Model ReleasesDGX agent

arXiv:2605.29983v1 Announce Type: cross Abstract: The adversarial robustness of attributions is a fundamental requirement for reliable explainability in deep learning, yet existing approaches typicall

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning

SafetyDGX agent

arXiv:2605.29776v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP demonstrate strong zero-shot generalization, but their performance significantly degrades in cross-domain sce

Inspectorch: Efficient rare event exploration in solar observations

ResearchDGX agent

arXiv:2602.20316v2 Announce Type: replace-cross Abstract: The Sun is observed in unprecedented detail, enabling studies of its activity on very small spatiotemporal scales. However, the large volume o

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

ResearchDGX agent

arXiv:2605.30230v1 Announce Type: new Abstract: With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still re

KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing

SafetyDGX agent

arXiv:2605.29509v1 Announce Type: new Abstract: In recent years, training-free video generation has progressed remarkably. However, when handling complex textual instructions, existing methods still s

Large Depth Completion Model from Sparse Observations

TutorialsDGX agent

arXiv:2605.30115v1 Announce Type: new Abstract: This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single-view metric depth estimation with spars

Learning Representations from 3D Gaussian Splats

Model ReleasesDGX agent

arXiv:2605.29549v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is a recent approach for scene rendering. Although primarily designed for view synthesis, its potential for scene understan

Lightweight Complementary-Cue Fusion for Robust Video Face Forgery Detection

ResearchDGX agent

arXiv:2605.29092v1 Announce Type: new Abstract: Current face video forgery detectors use wide or dual-stream backbones. We show that a single, lightweight fusion of two handcrafted cues can achieve hi

LiveSVG: Zero-Shot SVG Animation via Video Generation

Model ReleasesDGX agent

arXiv:2605.30174v1 Announce Type: new Abstract: We introduce LiveSVG, a zero-shot approach for generating Scalable Vector Graphics (SVG) animations using video diffusion models. Current SVG animation

Low-Magnification SEM May Suffice: Interpretable Deep Learning for Multi-Scale Fracture-Cause Classification in Zirconia-Toughened Alumina

SafetyDGX agent

arXiv:2605.29798v1 Announce Type: new Abstract: Reliable identification of fracture origins in alumina matrix composite hip and knee implants is critical for quality assurance and patient safety, yet

LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol

Model ReleasesDGX agent

arXiv:2603.14644v3 Announce Type: replace-cross Abstract: Publicly available full-field digital mammography (FFDM) datasets remain limited in size, clinical annotations, and vendor diversity, hinderin

MARTIAN: A Rendering Framework for Aerial Mars Imagery from HiRISE Orbital Data

ResearchDGX agent

arXiv:2605.29647v1 Announce Type: new Abstract: Aerial navigation on Mars requires vision-based pipelines that are robust to the diverse illumination conditions and terrain morphology of the Martian s

Masked Diffusion Vision-Language Models for Temporal Action Localization

ResearchDGX agent

arXiv:2605.29858v1 Announce Type: new Abstract: Temporal action localization (TAL) requires recognizing the target event and localizing its start and end times precisely in untrimmed videos. Recent vi

MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species

SafetyDGX agent

arXiv:2601.03729v2 Announce Type: replace Abstract: Fine-grained recognition of marine organisms is important for ecological research, biodiversity monitoring, habitat conservation, and evidence-based

Mesh-Aware Epipolar Matching for Multi-View Multi-Person 3D Pose Estimation in Basketball

ResearchDGX agent

arXiv:2605.29953v1 Announce Type: new Abstract: Multi-view multi-person 3D pose estimation in team sports scenarios remains challenging due to player occlusions, appearance similarity caused by team u

MetaRanker: Human-in-the-loop Active Ranking for Metalens Image Quality

SafetyDGX agent

arXiv:2605.29212v1 Announce Type: new Abstract: Image quality in modern imaging systems emerges from the coupled effects of the sensor, optics, and computational reconstruction. Ultra-thin metalenses

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

ResearchDGX agent

arXiv:2605.30263v1 Announce Type: new Abstract: Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning

SafetyDGX agent

arXiv:2605.29577v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a promising framework that unifies perception, reasoning, and control for robot manipulation by adap

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds

SafetyDGX agent

arXiv:2510.27391v2 Announce Type: replace Abstract: Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods e

MonoPhysics: Estimating Geometry, Appearance, and Physical Parameters from Monocular Videos

SafetyDGX agent

arXiv:2605.30320v1 Announce Type: new Abstract: Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D struc

Motion-guided sparse correction enables expert-quality point tracking across diverse microscopy regimes

ApplicationsDGX agent

arXiv:2605.29220v1 Announce Type: new Abstract: Tracking the dynamics of non-canonical biological systems in microscopy videos remains a persistent challenge. Both classical and learning-based tracker

Multi-level Collaborative Distillation Meets Global Workspace Model: A Unified Framework for OCIL

TutorialsDGX agent

arXiv:2508.08677v2 Announce Type: replace-cross Abstract: Online Class-Incremental Learning (OCIL) enables models to learn continuously from non-i.i.d. data streams. Since samples of the data streams

Multi-Scale Local Speculative Decoding for Image Generation

Local AiDGX agent

arXiv:2601.05149v2 Announce Type: replace Abstract: Autoregressive (AR) models have achieved remarkable success in image synthesis, yet their sequential nature imposes significant latency constraints.

Multi-Stage VLM Pipeline for Zero-Shot Traffic Accident Understanding

ResearchDGX agent

arXiv:2605.29325v1 Announce Type: new Abstract: We present the 1st-place solution to the ACCIDENT challenge at the CVPR 2026 AUTOPILOT Workshop, which asks for zero-shot prediction of accident timing,

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2602.14399v2 Announce Type: replace Abstract: Multi-turn jailbreak attacks have proven effective against text-only large language models (LLMs), where malicious content is gradually introduced t

Multimodal LLMs See Sentiment

Model ReleasesDGX agent

arXiv:2508.16873v3 Announce Type: replace Abstract: Understanding how visual content conveys sentiment is increasingly important in a digital landscape dominated by imagery. However, sentiment percept

Native Audio-Visual Alignment for Generation

SafetyDGX agent

arXiv:2605.30073v1 Announce Type: new Abstract: Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual-acoustic content. However, existing open-source

NeuROK: Generative 4D Neural Object Kinematics

TutorialsDGX agent

arXiv:2605.30347v1 Announce Type: new Abstract: Data-driven approaches have revolutionized 3D vision, enabling transformers to effectively reconstruct and generate static 3D objects. However, generati

Non-Forgetting Knowledge Allocation with Bi-level Competition for Class-Incremental Learning

ResearchDGX agent

arXiv:2605.29592v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) with pre-trained models (PTMs) aims to sequentially adapt PTMs to new categories without forgetting old knowledge. Buil

Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval Using Language

TutorialsDGX agent

arXiv:2605.29812v1 Announce Type: new Abstract: Video Moment Retrieval (VMR) targets to retrieve the specific moment corresponding to a sentence query from an untrimmed video. Although recent works ha

OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the Wild

TutorialsDGX agent

arXiv:2511.08423v3 Announce Type: replace Abstract: A truly universal AI-Generated Image (AIGI) detector must simultaneously generalize across diverse generative models and varied semantic content. Cu

OmniCD: A Foundational Framework for Remote Sensing Image Change Detection Guided by Multimodal Semantics

ResearchDGX agent

arXiv:2605.30168v1 Announce Type: new Abstract: Change detection (CD) in remote sensing is vital for applications such as urban monitoring and disaster assessment, yet traditional methods struggle wit

One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation

ResearchDGX agent

arXiv:2605.29429v1 Announce Type: new Abstract: Cell instance segmentation models trained on cell-specific datasets suffer severe performance drops on out-of-distribution cell types, while interactive

Optimizing Latent Representations for Robust Building Damage Assessment Onboard Earth Observation Satellites

Model ReleasesDGX agent

arXiv:2605.29575v1 Announce Type: new Abstract: Rapid identification of damaged buildings after natural disasters or on war areas is crucial to support emergency response and prioritize interventions.

Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation

SafetyDGX agent

arXiv:2605.29390v1 Announce Type: new Abstract: Text-to-image (T2I) models have become increasingly capable of generating high-quality images. Yet, enforcing the explicit absence of a specified object

Parameter-Efficient Subspace Decoupling ViT for Mitigating Multi-Task Negative Transfer in Histological Scoring

Model ReleasesDGX agent

arXiv:2605.29852v1 Announce Type: new Abstract: Histological scoring is essential for diagnosing Non-Alcoholic Fatty Liver Disease (NAFLD), yet its automation remains challenging due to the high annot

← Previous
1…108109110111112…211
Next →