AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

AdaEraser: Training-Free Object Removal via Adaptive Attention Suppression

DGX agent

arXiv:2605.15921v1 Announce Type: new Abstract: Object removal aims to eliminate specified objects from images while plausibly inpainting the affected regions with background content. Current training

researcharxiv-cs-cv
18 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging

DGX agent

arXiv:2505.21698v3 Announce Type: replace Abstract: Vision-language foundation models achieve promising performance in natural image classification, yet their direct application to medical imaging is

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models

DGX agent

arXiv:2605.15584v1 Announce Type: new Abstract: Vision-language models like CLIP have demonstrated remarkable zero-shot transfer capabilities. However, their susceptibility to imperceptible adversaria

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

AnyAct: Towards Human Reenactment of Character Motion From Video

DGX agent

arXiv:2605.15497v1 Announce Type: new Abstract: We study the problem of directly deriving an initial human reenactment from a monocular video of a non-human character. Our goal is not to reconstruct t

model-releasesarxiv-cs-cv
18 May 2026
Research

ART: Articulated Reconstruction Transformer

DGX agent

arXiv:2512.14671v3 Announce Type: replace Abstract: We introduce ART, Articulated Reconstruction Transformer -- a category-agnostic, feed-forward model that reconstructs complete 3D articulated object

researcharxiv-cs-cv
18 May 2026
Agents

Attribute-Grounded Selective Reasoning for Artwork Emotion Understanding with Multimodal Large Language Models

DGX agent

arXiv:2605.15755v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can produce fluent artwork emotion explanations, but they often suffer from attribute flooding: they enumerate

agentsarxiv-cs-cv
18 May 2026
Research

BARRIER: Bounded Activation Regions for Robust Information Erasure

DGX agent

arXiv:2605.15737v1 Announce Type: new Abstract: Machine unlearning has reached a critical bottleneck. As traditional weight-space interventions focus primarily on erasing targeted concepts, they often

researcharxiv-cs-cv
18 May 2026
Model Releases

Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition

DGX agent

arXiv:2602.00841v4 Announce Type: replace Abstract: Visual Place Recognition (VPR) demands representations robust to drastic environmental and viewpoint shifts. Existing aggregation paradigms either d

model-releasesarxiv-cs-cv
18 May 2026
Safety

Beyond Performance Disparities: A Three-Level Audit of Representational Harm in CelebA

DGX agent

arXiv:2605.15312v1 Announce Type: cross Abstract: Large-scale facial datasets like CelebA are widely used in computer vision, yet the cultural biases embedded in their labels remain underexplored. Fai

safetyarxiv-cs-cv
18 May 2026
Research

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models

DGX agent

arXiv:2601.21798v2 Announce Type: replace Abstract: Large Language Models(LLMs) have revolutionized text generation and multimodal perception,but their capabilities in 3D content generation remain und

researcharxiv-cs-cv
18 May 2026
Model Releases

ChronoEarth-492K: A Large Scale and Long Horizon Spatiotemporal Hyperspectral Earth Observation Dataset and Benchmark

DGX agent

arXiv:2605.15666v1 Announce Type: new Abstract: Hyperspectral imaging (HSI) provides dense spectral information for the Earth's surface, enabling material-level understanding of land cover and ecosyst

model-releasesarxiv-cs-cv
18 May 2026
Tutorials

CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage

DGX agent

arXiv:2605.15597v1 Announce Type: new Abstract: Modern 3D visual learning relies on observations sampled from metric 3D assets, yet existing scans, meshes, point clouds, simulations, and reconstructio

tutorialsarxiv-cs-cv
18 May 2026
Research

Community-aware evaluation and threshold calibration for open-set plankton image recognition

DGX agent

arXiv:2605.15835v1 Announce Type: new Abstract: Automated plankton image recognition is increasingly used in aquatic ecosystem monitoring, but deployed classifiers inevitably encounter unseen taxa and

researcharxiv-cs-cv
18 May 2026
Model Releases

COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection

DGX agent

arXiv:2605.15325v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong performance in video anomaly detection (VAD) while providing interpretable predictions. However, existin

model-releasesarxiv-cs-cv
18 May 2026
Local Ai

Cross-Modal Registration Between 3D and 2D Fingerprints via Pose-Aware Unwrapping and Point-Cloud Fusion

DGX agent

arXiv:2605.15796v1 Announce Type: new Abstract: Three-dimensional (3D) fingerprints preserve global finger geometry and local ridge structure while avoiding contact-induced deformation, but they remai

local-aiarxiv-cs-cv
18 May 2026
Research

DealMaTe: Multi-Dimensional Material Transfer via Diffusion Transformer

DGX agent

arXiv:2605.15681v1 Announce Type: cross Abstract: Recently, diffusion-based material transfer methods rely on image fine-tuning or complex architectures with auxiliary networks but face challenges suc

researcharxiv-cs-cv
18 May 2026
Model Releases

Decentralized LoRA augmented transformer with multi-scale feature learning for secured eye diagnosis

DGX agent

arXiv:2505.06982v3 Announce Type: replace Abstract: Accurate and privacy-preserving diagnosis of ophthalmic diseases remains a critical challenge in medical imaging, particularly given the limitations

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

Deep Pre-Alignment for VLMs

DGX agent

arXiv:2605.15300v1 Announce Type: new Abstract: Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis sugge

model-releasesarxiv-cs-cv
18 May 2026
Research

Degradation-Aware Blur-Segmentation of Brain Tumor

DGX agent

arXiv:2605.15671v1 Announce Type: cross Abstract: Multimodal 3D MRI brain tumor segmentation is a pivotal step in radiotherapy target delineation, surgical planning and post-treatment assessment. Exis

researcharxiv-cs-cv
18 May 2026
Research

DIPA: Distilled Preconditioned Algorithms for Solving Imaging Inverse Problems

DGX agent

arXiv:2605.15456v1 Announce Type: cross Abstract: Solving imaging inverse problems has usually been addressed by designing proper prior models of the underlying signal. However, minimizing the data fi

researcharxiv-cs-cv
18 May 2026
Safety

Discretizing Group-Convolutional Neural Networks for 3D Geometry in Feature Space

DGX agent

arXiv:2605.15368v1 Announce Type: new Abstract: Group-convolutional neural networks (GCNNs) are among the most important methods for introducing symmetry as an inductive bias in deep learning: In each

safetyarxiv-cs-cv
18 May 2026
Safety

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

DGX agent

arXiv:2605.15855v1 Announce Type: new Abstract: Despite strong image-generation performance, diffusion models' reconstruction objectives limit alignment with human preferences. RL enables such alignme

safetyarxiv-cs-cv
18 May 2026
Local Ai

DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion Transformer

DGX agent

arXiv:2605.15682v1 Announce Type: new Abstract: Large-scale pre-trained diffusion models have been extensively adopted for real-world image Super-Resolution because of their powerful generative priors

local-aiarxiv-cs-cv
18 May 2026
Safety

DualReg: Dual-Space Filtering and Reinforcement for Rigid Registration

DGX agent

arXiv:2508.17034v2 Announce Type: cross Abstract: Noisy, partially overlapping data and the need for real-time processing pose major challenges for rigid registration. Considering that feature-based m

safetyarxiv-cs-cv
18 May 2026
Model Releases

Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation

DGX agent

arXiv:2605.16003v1 Announce Type: new Abstract: Autoregressive video diffusion models enable open-ended generation through local attention and KV caching. However, existing training-free long-video op

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos

DGX agent

arXiv:2603.03066v2 Announce Type: replace Abstract: Existing AI-generated video quality assessment (AIGVQA) methods mainly focus on global perceptual realism and coarse text-video alignment, while ove

model-releasesarxiv-cs-cv
18 May 2026
Research

Efficient Image Synthesis with Sphere Latent Encoder

DGX agent

arXiv:2605.15592v1 Announce Type: new Abstract: Few-step image generation has seen rapid progress, with consistency and meanflow-based methods significantly reducing the number of sampling steps. Desp

researcharxiv-cs-cv
18 May 2026
Safety

EgoExo-WM: Unlocking Exo Video for Ego World Models

DGX agent

arXiv:2605.15477v1 Announce Type: new Abstract: Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited avail

safetyarxiv-cs-cv
18 May 2026
Safety

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices

DGX agent

arXiv:2605.15684v1 Announce Type: new Abstract: The Diffusion Transformer (DiT) architecture is the state-of-the-art paradigm for high-fidelity image generation, underpinning models like Stable Diffus

safetyarxiv-cs-cv
18 May 2026
Model Releases

ELDOR: A Dataset and Benchmark for Illegal Gold Mining in the Amazon Rainforest

DGX agent

arXiv:2605.15397v1 Announce Type: new Abstract: Illegal gold mining in the Amazon rainforest causes deforestation, water contamination, and long-term ecosystem disruption, yet remains difficult to mon

model-releasesarxiv-cs-cv
18 May 2026
Safety

Embedding-perturbed Exploration Preference Optimization for Flow Models

DGX agent

arXiv:2605.15803v1 Announce Type: new Abstract: Recent advancements have established Reinforcement Learning (RL) as a pivotal paradigm for aligning generative models with human intent. However, group-

safetyarxiv-cs-cv
18 May 2026
Model Releases

End-to-end plaque counting and virus titration from laboratory plate images with deep learning

DGX agent

arXiv:2605.16008v1 Announce Type: new Abstract: Plaque assays remain the gold standard readout of virus infectivity; however, plaque counting from plate images is labor-intensive and prone to inter-op

model-releasesarxiv-cs-cv
18 May 2026
Research

EndoGSim: Physics-Aware 4D Dynamic Endoscopic Scene Simulations via MLLM-Guided Gaussian Splatting

DGX agent

arXiv:2605.16022v1 Announce Type: new Abstract: In robot-assisted minimally invasive surgery, high-fidelity dynamic endoscopic scene reconstruction and simulation are crucial to enhancing downstream t

researcharxiv-cs-cv
18 May 2026
Research

Enhancing Medical Image Segmentation via Heat Conduction Equation

DGX agent

arXiv:2511.03260v2 Announce Type: replace Abstract: Medical image segmentation models struggle to achieve efficient global context modeling and long-range dependency reasoning under practical computat

researcharxiv-cs-cv
18 May 2026
Model Releases

Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction

DGX agent

arXiv:2605.15466v1 Announce Type: new Abstract: Learning predictive world models from unlabelled video is a foundational challenge in artificial intelligence. While Joint Embedding Predictive Architec

model-releasesarxiv-cs-cv
18 May 2026
Safety

EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy

DGX agent

arXiv:2605.15711v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to backdoor attacks. Exi

safetyarxiv-cs-cv
18 May 2026
Research

Evaluation of Anatomical Shape Priors in Deep Learning-Based Cardiac Multi-Compartment Segmentation

DGX agent

arXiv:2605.15707v1 Announce Type: cross Abstract: Whole-heart multi-compartment CT segmentation is clinically important, but standard CNNs do not explicitly enforce anatomical plausibility. Based on s

researcharxiv-cs-cv
18 May 2026
Hardware

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

DGX agent

arXiv:2605.15824v1 Announce Type: new Abstract: Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot suppor

hardwarearxiv-cs-cv
18 May 2026
Model Releases

FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction

DGX agent

arXiv:2605.15320v1 Announce Type: cross Abstract: Avatar reconstruction has traditionally relied on per-subject optimization that requires hours of computation or on expensive preprocessing that limit

model-releasesarxiv-cs-cv
18 May 2026
Safety

FLASH: Efficient Visuomotor Policy via Sparse Sampling

DGX agent

arXiv:2605.15492v1 Announce Type: cross Abstract: Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative d

safetyarxiv-cs-cv
18 May 2026
Model Releases

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization

DGX agent

arXiv:2605.15980v1 Announce Type: new Abstract: Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computation

model-releasesarxiv-cs-cv
18 May 2026
Research

Frequency-domain Event-based Imaging for Selective Surveillance

DGX agent

arXiv:2605.15392v1 Announce Type: cross Abstract: Event-based cameras (EBCs) are an attractive sensing modality for surveillance due to their reporting of pixel-level radiance changes with microsecond

researcharxiv-cs-cv
18 May 2026
Research

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding

DGX agent

arXiv:2605.15951v1 Announce Type: new Abstract: Finetuning Large Vision-Language Models with reinforcement learning has emerged as a promising approach to enhance their capability in object-level grou

researcharxiv-cs-cv
18 May 2026
Safety

From Full and Partial Intraoral Scans to Crown Proposal: A Classification-Guided Restoration Assistance Pipeline

DGX agent

arXiv:2605.15241v1 Announce Type: cross Abstract: Single-unit crown restoration is among the most common procedures in clinical dentistry, with CAD/CAM workflows now designing crowns directly from int

safetyarxiv-cs-cv
18 May 2026
Safety

From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

DGX agent

arXiv:2602.20630v4 Announce Type: replace Abstract: Keypoint-based matching is a fundamental component of modern 3D vision systems, such as Structure-from-Motion (SfM) and SLAM. Most existing learning

safetyarxiv-cs-cv
18 May 2026
Research

GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D Reconstruction

DGX agent

arXiv:2605.15852v1 Announce Type: new Abstract: Streaming 3D reconstruction from long monocular video sequences requires maintaining a key-value (KV) cache that grows linearly with sequence length, cr

researcharxiv-cs-cv
18 May 2026
Safety

GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective

DGX agent

arXiv:2605.15723v1 Announce Type: cross Abstract: Multimodal alignment is commonly learned from isolated image-text pairs via CLIP-style dual encoders, leaving the relational context among entities la

safetyarxiv-cs-cv
18 May 2026
Local Ai

Grounded Reinforcement Learning for Visual Reasoning

DGX agent

arXiv:2505.23678v3 Announce Type: replace Abstract: While reinforcement learning (RL) over chains of thought has significantly advanced language models in tasks such as mathematics and coding, visual

local-aiarxiv-cs-cv
18 May 2026
← Previous
1…165166167168169…263
Next →