AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
2 Jul 2026

Continuous Speculative Decoding for Autoregressive Image Generation

SafetyDGX agent

arXiv:2411.11925v3 Announce Type: replace Abstract: Continuous visual autoregressive (AR) models have demonstrated promising performance in image generation, but their inherently sequential nature res

CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild

TutorialsDGX agent

arXiv:2607.00321v1 Announce Type: new Abstract: Reconstructing high-fidelity 3D models of highly articulated animals, such as dogs, from a single in-the-wild image remains a formidable challenge. In t

CPDDNet: Color-Polarization Denoising and Demosaicking Network

Model ReleasesDGX agent

arXiv:2607.01100v1 Announce Type: new Abstract: Color-polarization imaging using a color-polarization filter array (CPFA) sensor captures both texture (color intensity) and physical (polarization) inf


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding

ResearchDGX agent

arXiv:2607.00672v1 Announce Type: new Abstract: Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing met

Dataset Biases and Shortcut Learning in Motion-Based AI-Generated Video Detection

SafetyDGX agent

arXiv:2607.00948v1 Announce Type: new Abstract: The visual quality of AI-generated videos has improved drastically in recent years, making it increasingly difficult for humans to distinguish between r

Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners

ResearchDGX agent

arXiv:2607.00125v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable abilities when analyzing images, yet translating these capabilities to few-shot im

Decoupled Guidance: Disentangling Subject and Context Pathways in Text-to-Image Personalization

ResearchDGX agent

arXiv:2607.00766v1 Announce Type: new Abstract: Text-to-image personalization aims to generate a user-provided subject in novel scenes described by text. However, most existing methods encode subject

Diffusion-Based Multi-Class Normality for OOD Detection: An Application to CDP Authentication

ResearchDGX agent

arXiv:2607.00609v1 Announce Type: new Abstract: Reconstruction-based generative models offer a natural framework for unsupervised out-of-distribution (OOD) detection, but multi-class normality modelli

Does Your ViT Still Need U-Net for Segmentation?

Model ReleasesDGX agent

arXiv:2607.00223v1 Announce Type: new Abstract: Medical image segmentation is dominated by U-Net-style encoder-decoder architectures. Vision Transformers (ViTs) overcome the limited receptive field of

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

SafetyDGX agent

arXiv:2607.00666v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models often fail to perform the same learned tasks under environmental shifts, such as changes in camera pose and shifts

DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation

TutorialsDGX agent

arXiv:2607.00183v1 Announce Type: new Abstract: Adapting pre-trained text-to-image diffusion models, whether to learn new visual concepts or erase unwanted ones, is routinely evaluated on its intended

DriveVA: Video Action Models are Zero-Shot Drivers

Model ReleasesDGX agent

arXiv:2604.04198v2 Announce Type: replace Abstract: Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor dom

DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.00399v1 Announce Type: new Abstract: End-to-end autonomous driving models often encounter performance bottlenecks, as training-time scaling leads to high computational costs and diminishing

DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images

Model ReleasesDGX agent

arXiv:2607.00338v1 Announce Type: new Abstract: Object detection for Unmanned Aerial Vehicles (UAVs) working in open and dynamic environments is a highly challenging task. While Vision-Language Models

DroneIQA-VLE: Multi-Task Drone Image Quality Assessment via Vision-Language Ensemble

ResearchDGX agent

arXiv:2607.00416v1 Announce Type: new Abstract: We present DroneIQA-VLE, our solution to the ICME 2026 Drone-IQA Grand Challenge on Target-aware Image Quality Assessment for Low-altitude UAV Images. T

E-TIDE: Fast, Structure-Preserving Motion Forecasting from Event Sequences

ResearchDGX agent

arXiv:2603.27757v2 Announce Type: replace Abstract: Event-based cameras capture visual information as asynchronous streams of per-pixel brightness changes, generating sparse, temporally precise data.

ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation

SafetyDGX agent

arXiv:2607.00545v1 Announce Type: new Abstract: Controllable traffic simulation is critical for testing autonomous driving systems, yet existing approaches often require retraining large generative mo

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection

Local AiDGX agent

arXiv:2607.00867v1 Announce Type: new Abstract: Long-video reasoning is fundamentally constrained by how models acquire and utilize visual evidence. Existing tool-augmented video frameworks often inte

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

Model ReleasesDGX agent

arXiv:2512.15702v2 Announce Type: replace Abstract: Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. Wh

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking

Model ReleasesDGX agent

arXiv:2412.20750v3 Announce Type: replace Abstract: Large-scale Vision-Language Models (VLMs) have achieved notable progress in aligning visual inputs with text. However, their ability to deeply under

EPO: Boosting 3D Foundation Models with Edge-based Pose Optimization

SafetyDGX agent

arXiv:2607.00579v1 Announce Type: new Abstract: We introduce extbf{Edge-based Pose Optimization (EPO)}, a trackless geometric optimization framework specifically designed to boost the Structure-from-M

EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation

SafetyDGX agent

arXiv:2607.01147v1 Announce Type: new Abstract: Text-to-image diffusion models power everyday creative tasks, but they still reproduce the demographic biases in their training data. On common prompts

Explainability in mulimodal deep transformation models for stroke outcome prediction

ResearchDGX agent

arXiv:2504.06299v2 Announce Type: replace-cross Abstract: Multimodal prediction models based on imaging and clinical data are increasingly used for clinical decision support, yet their interpretabilit

FCL-COD: Weakly Supervised Camouflaged Object Detection with Frequency-aware and Contrastive Learning

ResearchDGX agent

arXiv:2603.22969v2 Announce Type: replace Abstract: Existing camouflage object detection (COD) methods typically rely on fully-supervised learning guided by mask annotations. However, obtaining mask a

Foundation Model-driven Key Anatomy Frame Selection for Blind-sweep Ultrasound Fetal Birth Weight Estimation

ResearchDGX agent

arXiv:2607.00745v1 Announce Type: new Abstract: Accurate fetal birth weight (FBW) estimation shortly before delivery is clinically valuable yet challenging due to its reliance on operator expertise, p

Foundation Models vs. Radiomics for Lung Computed Tomography: A Benchmark of Feature Extractors, Classification Heads, and Segmentation Choices

Model ReleasesDGX agent

arXiv:2607.01001v1 Announce Type: new Abstract: Radiomics is the established approach for CT-based lung cancer phenotyping, yet comparisons with foundation models rarely isolate contributions of featu

FrameONE: Hierarchical Motion Modeling for Universal Multi-View Echocardiographic Keyframe Detection

SafetyDGX agent

arXiv:2607.00748v1 Announce Type: new Abstract: Accurate detection of end-systole (ES) and end-diastole (ED) frames is fundamental to echocardiographic assessment. Existing methods are typically devel

FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal

Model ReleasesDGX agent

arXiv:2603.19036v2 Announce Type: replace Abstract: Single image reflection removal (SIRR) is challenging in real scenes, where reflection strength varies spatially and reflection patterns are tightly

GADA: Geometry-Aware Deformable Aggregation for Image-Based Gaussian Splatting

ResearchDGX agent

arXiv:2607.00595v1 Announce Type: new Abstract: Gaussian Splatting has achieved significant improvements by incorporating warping-based techniques. However, such methods suffer from pixel-level inaccu

GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting

ResearchDGX agent

arXiv:2607.00959v1 Announce Type: new Abstract: Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avat

GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

TutorialsDGX agent

arXiv:2603.26661v2 Announce Type: replace Abstract: Most recent advances in 3D generative modeling rely on diffusion or flow-matching formulations. We instead explore a fully autoregressive alternativ

GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine

Model ReleasesDGX agent

arXiv:2607.00544v1 Announce Type: new Abstract: Reasoning segmentation requires localizing targets based on complex, implicit queries. Current end-to-end models typically entangle perception and deduc

GenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language Models

Local AiDGX agent

arXiv:2607.01049v1 Announce Type: new Abstract: Industrial inspection requires more than binary anomaly detection: a practical system should determine whether an anomaly exists, localize the defective

Generated Contents Enrichment

ResearchDGX agent

arXiv:2405.03650v4 Announce Type: replace Abstract: We study Generated Contents Enrichment (GCE), a conditional image-generation task in which a sparse scene description is first enriched through an e

GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models

TutorialsDGX agent

arXiv:2607.00492v1 Announce Type: new Abstract: We introduce GenSP, a data-driven framework that learns consistent spherical parameterizations across a collection of genus-0 shapes. Instead of optimiz

Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics

ApplicationsDGX agent

arXiv:2603.13859v2 Announce Type: replace Abstract: Intrinsic image decomposition aims to estimate physically based rendering (PBR) parameters such as albedo, roughness, and metallicity from images. W

Geometry-Aware Cross-Height Channel Knowledge Map Prediction for UAV-Assisted Communications With Uncertainty-Guided 3D Sensing

Model ReleasesDGX agent

arXiv:2607.00887v1 Announce Type: new Abstract: Low-altitude Unmanned Aerial Vehicles (UAVs) often need to infer channel knowledge across a range of heights from only sparse observations collected at

GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision

Model ReleasesDGX agent

arXiv:2607.01050v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have shown strong cross-modal understanding and coordinate generation abilities in visual grounding. How

GimbalDiffusion: Gravity-Aware Camera Control for Video Generation

ResearchDGX agent

arXiv:2512.09112v3 Announce Type: replace Abstract: Recent progress in text-to-video generation has achieved remarkable realism, yet fine-grained control over camera motion and orientation remains elu

GKDT: General Keypoint Detection Transformer

Model ReleasesDGX agent

arXiv:2607.00752v1 Announce Type: new Abstract: With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The con

GMO-E^2DIT: Grounded Multi-Operation Editing for E-Commerce Images

Model ReleasesDGX agent

arXiv:2607.00920v1 Announce Type: new Abstract: Real-world e-commerce image editing often requires multiple, localized, and auditable operations rather than global restyling. This compositional nature

GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition

Model ReleasesDGX agent

arXiv:2602.03370v2 Announce Type: replace Abstract: Handwritten mathematical expression recognition (HMER) requires reasoning over diverse symbols and structures, yet autoregressive models struggle wi

HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking

ResearchDGX agent

arXiv:2607.00494v1 Announce Type: new Abstract: Multi-animal tracking (MAT) is critical for wildlife monitoring and behavioral analysis, yet remains challenging due to uniform appearance, high density

High-dimensional Embedding Prior for Noisy K-space Domain MRIReconstruction

ResearchDGX agent

arXiv:2607.01176v1 Announce Type: new Abstract: Magnetic resonance imaging (MRI) reconstruction under realistic acquisition conditions can be fundamentally viewed as estimating the underlying k-space

Histopathology Multi-modal Embedding for Pathology Composed Retrieval

Model ReleasesDGX agent

arXiv:2502.07221v4 Announce Type: replace Abstract: To overcome the black-box nature of predictive AI and the hallucination risks of generative models, retrieval-based models offer an interpretable, e

Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation

ResearchDGX agent

arXiv:2607.01067v1 Announce Type: cross Abstract: As an essential modality for dexterous and contact-rich tasks, tactile sensing provides precise force feedback that cannot be reliably inferred from v

HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding

SafetyDGX agent

arXiv:2607.00428v1 Announce Type: new Abstract: CLIP (Contrastive Language-Image Pre-training) has become a de facto paradigm for image-text alignment, but it struggles with long-context descriptions

Imprint: Online Memory Compression for Long-Horizon Egocentric QA

Model ReleasesDGX agent

arXiv:2607.00696v1 Announce Type: new Abstract: Long-horizon egocentric question answering involves answering about events that have occurred hours or days in the past. This requires memory representa

Information-Regularized Attention for Visual-Centric Reasoning

Model ReleasesDGX agent

arXiv:2607.00434v1 Announce Type: new Abstract: Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, an

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models

ResearchDGX agent

arXiv:2607.01222v1 Announce Type: new Abstract: Recent 3D generative models can synthesize high-quality geometry but often struggle to reproduce intricate textures from reference images, largely due t

IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video

Model ReleasesDGX agent

arXiv:2603.16432v3 Announce Type: replace Abstract: Unsupervised physical parameter estimation from video lacks a common benchmark: existing methods evaluate on non-overlapping synthetic data, the sol

Joint Medical Image Enhancement and Segmentation with Diffusion-based Symbiotic Information Interaction

ResearchDGX agent

arXiv:2607.00058v1 Announce Type: new Abstract: Image quality is critical for accurate medical diagnosis. However, MRI, CT, and ultrasound images are often of low resolution and quality due to cost co

Learn Once, Edit Anywhere: Visual Direction Transfer for Diffusion Models

TutorialsDGX agent

arXiv:2403.19645v2 Announce Type: replace Abstract: The rapid advancement of diffusion models has enabled the generation of high-fidelity images from textual prompts, yet achieving precise, disentangl

Learning to Watch: Active Video Anomaly Understanding via Interleaved Policy Optimization

Local AiDGX agent

arXiv:2607.00622v1 Announce Type: new Abstract: Video anomaly understanding (VAU) relies on sparse, context-dependent cues. However, existing passive paradigms suffer from observational aliasing, wher

Linguistic Relative Policy Optimization for Video Anomaly Reasoning

Model ReleasesDGX agent

arXiv:2607.00654v1 Announce Type: new Abstract: Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale an

Linkify: Learning from Interface-Augmented Assembly Graphs

Model ReleasesDGX agent

arXiv:2607.01205v1 Announce Type: new Abstract: We present Linkify, a framework for learning from interface-augmented assembly graphs to enable context-aware part retrieval in mechanical assemblies. W

LIST3R: Long-sequence Instance-aware 3D Reconstruction

SafetyDGX agent

arXiv:2607.00375v1 Announce Type: new Abstract: We present LIST3R, an instance-aware framework for long-sequence 3D reconstruction inspired by the way humans organize spatial memory around stable and

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs

ResearchDGX agent

arXiv:2602.05275v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have shown immense promise in universal multimodal retrieval, which aims to find relevant items of various

MedCAGD: Context-Aware Gated Decoder for Efficient Medical Image Segmentation

SafetyDGX agent

arXiv:2607.00409v1 Announce Type: new Abstract: Medical image segmentation relies on the ability of encoder-decoder architectures to translate rich feature representations into accurate pixel-level pr

MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization

ResearchDGX agent

arXiv:2607.00902v1 Announce Type: new Abstract: Driven by Artificial Intelligence-Generated Content (AIGC), the authenticity of audio-visual content is facing severe challenges. Temporal Forgery Local

← Previous
1…5455565758…209
Next →