AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection

DGX agent

arXiv:2602.09524v3 Announce Type: replace Abstract: Unsupervised industrial anomaly detection (UAD) is essential for modern manufacturing inspection, where defect samples are scarce and reliable detec

safetyarxiv-cs-cv
12 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Tutorials

HPGN: Hybrid Priors-Guided Network for Compressed Low-Light Image Enhancement

DGX agent

arXiv:2504.02373v3 Announce Type: replace-cross Abstract: In practical applications, low-light images are often compressed for efficient storage and transmission. Most existing methods disregard compr

tutorialsarxiv-cs-cv
12 May 2026
Model Releases

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation

DGX agent

arXiv:2501.12202v4 Announce Type: replace Abstract: We present Hunyuan3D 2.0, an advanced large-scale 3D synthesis system for generating high-resolution textured 3D assets. This system includes two fo

model-releasesarxiv-cs-cv
12 May 2026
Safety

HyNeuralMap: Hyperbolic Mapping of Visual Semantics to Neural Hierarchies

DGX agent

arXiv:2605.09392v1 Announce Type: new Abstract: Understanding the intricate mappings between visual stimuli and neural responses is a fundamental challenge in cognitive neuroscience. While current app

safetyarxiv-cs-cv
12 May 2026
Research

Hypergraph-Enhanced Training-Free and Language-Free Few-Shot Anomaly Detection

DGX agent

arXiv:2605.10628v1 Announce Type: new Abstract: Few-shot anomaly detection (FSAD) has made significant strides, yet existing methods still face critical challenges: (i) dependence on task- or dataset-

researcharxiv-cs-cv
12 May 2026
Model Releases

Hystar: Hypernetwork-driven Style-adaptive Retrieval via Dynamic SVD Modulation

DGX agent

arXiv:2605.10009v1 Announce Type: new Abstract: Query-based image retrieval (QBIR) requires retrieving relevant images given diverse and often stylistically heterogeneous queries, such as sketches, ar

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models

DGX agent

arXiv:2605.08841v1 Announce Type: new Abstract: Vision-Language Models (VLMs) exhibit systematic bias toward visual illusions, recalling memorized facts rather than perceiving actual visual difference

model-releasesarxiv-cs-cv
12 May 2026
Research

Improved Mean Flows: On the Challenges of Fastforward Generative Models

DGX agent

arXiv:2512.02012v2 Announce Type: replace Abstract: MeanFlow (MF) has recently been established as a framework for one-step generative modeling. However, its ``fastforward'' nature introduces key chal

researcharxiv-cs-cv
12 May 2026
Tutorials

Improving Generative Adversarial Networks with Self-Distillation

DGX agent

arXiv:2605.08577v1 Announce Type: new Abstract: In modern GANs, maintaining an Exponential Moving Average (EMA) of the generator's weights is a standard practice, as such an averaged model consistentl

tutorialsarxiv-cs-cv
12 May 2026
Safety

Improving Human Image Animation via Semantic Representation Alignment

DGX agent

arXiv:2605.10523v1 Announce Type: new Abstract: The field of image-to-video generation has made remarkable progress. However, challenges such as human limb twisting and facial distortion persist, espe

safetyarxiv-cs-cv
12 May 2026
Research

Improving Temporal Action Segmentation via Constraint-Aware Decoding

DGX agent

arXiv:2605.10149v1 Announce Type: new Abstract: Temporal action segmentation (TAS) divides untrimmed videos into labeled action segments. While fully supervised methods have advanced the field, challe

researcharxiv-cs-cv
12 May 2026
Research

Increasing the Efficiency of DETR for Maritime High-Resolution Images

DGX agent

arXiv:2605.10269v1 Announce Type: new Abstract: Maritime object detection is critical for the safe navigation of unmanned surface vessels (USVs), requiring accurate recognition of obstacles from small

researcharxiv-cs-cv
12 May 2026
Research

INFANiTE: Implicit Neural representation for high-resolution Fetal brain spatio-temporal Atlas learNing from clinical Thick-slicE MRI

DGX agent

arXiv:2605.09977v1 Announce Type: new Abstract: Spatio-temporal fetal brain atlases are important for characterizing normative neurodevelopment and identifying congenital anomalies. However, existing

researcharxiv-cs-cv
12 May 2026
Local Ai

Initiation of Interaction Detection Framework using a Nonverbal Cue for Human-Robot Interaction

DGX agent

arXiv:2605.10087v1 Announce Type: new Abstract: This paper describes an initiation of interaction(IoI) detection framework without keywords for human-robot interaction(HRI) based on audio and vision s

local-aiarxiv-cs-cv
12 May 2026
Model Releases

IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts

DGX agent

arXiv:2605.08664v1 Announce Type: new Abstract: Current image quality assessment methods are heavily biased towards global distortions (e.g., noise, blur), neglecting local perceptual artifacts such a

model-releasesarxiv-cs-cv
12 May 2026
Safety

Is Class Signal Clustered or Routed in Task-Induced Implicit Neural Representation Weight Spaces?

DGX agent

arXiv:2605.08281v1 Announce Type: new Abstract: Implicit neural representations (INRs) encode images as neural-network weights, making image classification a problem of weight-space classifiability. A

safetyarxiv-cs-cv
12 May 2026
Model Releases

Is Your Driving World Model an All-Around Player?

DGX agent

arXiv:2605.10858v1 Announce Type: new Abstract: Today's driving world models can generate remarkably realistic dash-cam videos, yet no single model excels universally. Some generate photorealistic tex

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

JODA: Composable Joint Dynamics for Articulated Objects

DGX agent

arXiv:2605.09954v1 Announce Type: cross Abstract: Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamica

model-releasesarxiv-cs-cv
12 May 2026
Applications

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

DGX agent

arXiv:2601.22143v2 Announce Type: replace-cross Abstract: Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented abilit

applicationsarxiv-cs-cv
12 May 2026
Model Releases

KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection

DGX agent

arXiv:2605.09132v1 Announce Type: new Abstract: Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and

model-releasesarxiv-cs-cv
12 May 2026
Safety

KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes

DGX agent

arXiv:2512.11321v3 Announce Type: replace Abstract: Facial animation is a core component for creating digital characters in Computer Graphics (CG) industry. A typical production workflow relies on spa

safetyarxiv-cs-cv
12 May 2026
Applications

Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes

DGX agent

arXiv:2605.08635v1 Announce Type: new Abstract: Reconstructing dynamic 3D scenes from blurry monocular videos is challenging as motion-induced blur entangles object motion and geometry, hindering geom

applicationsarxiv-cs-cv
12 May 2026
Research

L2A: Learning to Accumulate Pose History for Accurate 3D Human Pose Estimation

DGX agent

arXiv:2605.08806v1 Announce Type: new Abstract: Existing 2D-3D lifting human pose estimation methods have achieved strong performance. But the utilization of historical pose representations across net

researcharxiv-cs-cv
12 May 2026
Agents

LCGNav: Local Candidate-Aware Geometric Enhancement for General Topological Planning in Vision-Language Navigation

DGX agent

arXiv:2605.09053v1 Announce Type: new Abstract: Online topological planning has become an effective paradigm for Vision-Language Navigation in Continuous Environments (VLN-CE), but existing methods st

agentsarxiv-cs-cv
12 May 2026
Applications

Learning-Augmented Scalable Linear Assignment Problem Optimization via Neural Dual Warm-Starts

DGX agent

arXiv:2605.09382v1 Announce Type: cross Abstract: The Linear Assignment Problem (LAP) is a fundamental combinatorial optimization task with applications ranging from computer vision to logistics. Clas

applicationsarxiv-cs-cv
12 May 2026
Safety

Learning to Align Generative Appearance Priors for Fine-grained Image Retrieval

DGX agent

arXiv:2605.09859v1 Announce Type: new Abstract: Fine-grained image retrieval (FGIR) typically relies on supervision from seen categories to learn discriminative embeddings for retrieving unseen catego

safetyarxiv-cs-cv
12 May 2026
Model Releases

Learning to Perceive 'Where': Spatial Pretext Tasks for Robust Self-Supervised Learning

DGX agent

arXiv:2605.09963v1 Announce Type: new Abstract: Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationshi

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LightAVSeg: Lightweight Audio-Visual Segmentation

DGX agent

arXiv:2605.08805v1 Announce Type: new Abstract: Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-mo

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency

DGX agent

arXiv:2605.10319v1 Announce Type: new Abstract: Layered image assets are widely used in real-world creative workflows, enabling non-destructive iteration and flexible re-composition. Recent advances i

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

DGX agent

arXiv:2605.08985v1 Announce Type: new Abstract: Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. T

model-releasesarxiv-cs-cv
12 May 2026
Research

Loom: Hybrid Retrieval-Scoring Outfit Recommendation with Semantic Material Compatibility and Occasion-Aware Embedding Priors

DGX agent

arXiv:2605.09830v1 Announce Type: cross Abstract: We present Loom, an outfit recommendation system that combines neural embedding retrieval with structured domain scoring to generate complete, coheren

researcharxiv-cs-cv
12 May 2026
Model Releases

Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

DGX agent

arXiv:2605.08787v1 Announce Type: new Abstract: Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical

model-releasesarxiv-cs-cv
12 May 2026
Research

Low-Cost Neural Radiance Fields

DGX agent

arXiv:2605.09312v1 Announce Type: new Abstract: Neural Radiance Fields (NeRF) achieve high-quality novel-view synthesis, but their long training times and reliance on dense input views limit accessibi

researcharxiv-cs-cv
12 May 2026
Agents

Low-Cost Stereo Vision for Robust 3D Positioning of Thin Radiata Pine Branches in Autonomous Drone Pruning

DGX agent

arXiv:2605.08213v1 Announce Type: new Abstract: Manual pruning of radiata pine, a species of major economic importance to New Zealand forestry, is hazardous, labour-intensive, and increasingly constra

agentsarxiv-cs-cv
12 May 2026
Model Releases

M^2E-UAV: A Benchmark and Analysis for Onboard Motion-on-Motion Event-Based Tiny UAV Detection

DGX agent

arXiv:2605.10496v1 Announce Type: new Abstract: Tiny UAV detection from an onboard event camera is difficult when the observer and target move at the same time. In this motion-on-motion regime, ego-mo

model-releasesarxiv-cs-cv
12 May 2026
Safety

Machine Unlearning on Pre-trained Models by Residual Feature Alignment Using LoRA

DGX agent

arXiv:2411.08443v2 Announce Type: replace-cross Abstract: Machine unlearning is an emerging technology that removes a subset of the training data from a trained model without significantly affecting t

safetyarxiv-cs-cv
12 May 2026
Safety

MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition

DGX agent

arXiv:2605.09418v1 Announce Type: new Abstract: Multi-modal cross-view place recognition remains a fundamental challenge in computer vision and robotics due to the severe viewpoint, modality, and spat

safetyarxiv-cs-cv
12 May 2026
Tutorials

Markerless Head Tracking for Accurate and Accessible Neuronavigation

DGX agent

arXiv:2602.07052v2 Announce Type: replace Abstract: Neuronavigation is widely used in biomedical research and interventions to guide the precise placement of instruments around the head to support pro

tutorialsarxiv-cs-cv
12 May 2026
Research

Masked Generative Transformer Is What You Need for Image Editing

DGX agent

arXiv:2605.10859v1 Announce Type: new Abstract: Diffusion models dominate image editing, yet their global denoising mechanism entangles edited regions with surrounding context, causing modifications t

researcharxiv-cs-cv
12 May 2026
Research

Measurement-Adapted Eigentask Representations for Photon-Limited Optical Readout

DGX agent

arXiv:2605.10008v1 Announce Type: cross Abstract: Optical readout in low-light imaging is fundamentally limited by measurement noise, including photon shot noise, detector noise, and quantization erro

researcharxiv-cs-cv
12 May 2026
Model Releases

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

DGX agent

arXiv:2605.10002v1 Announce Type: new Abstract: Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet inco

model-releasesarxiv-cs-cv
12 May 2026
Safety

MedFL-Stress: A Systematic Robustness Evaluation of Federated Brain Tumor Segmentation under Cross-Hospital MRI Appearance Shift

DGX agent

arXiv:2605.09025v1 Announce Type: new Abstract: Federated learning enables hospitals to collaboratively train segmentation models without sharing patient data. However, current evaluation protocols re

safetyarxiv-cs-cv
12 May 2026
Local Ai

MFVLR: Multi-domain Fine-grained Vision-Language Reconstruction for Generalizable Diffusion Face Forgery Detection and Localization

DGX agent

arXiv:2605.10071v1 Announce Type: new Abstract: The swift advancement in photo-realistic face generation technology has sparked considerable concerns across society and academia, emphasizing the requi

local-aiarxiv-cs-cv
12 May 2026
Research

MicroDiffuse3D: A Foundation Model for 3D Microscopy Imaging Restoration

DGX agent

arXiv:2605.08566v1 Announce Type: new Abstract: Chemical imaging enables label-free visualization of cells, tissues and living systems while providing direct biochemical information that is difficult

researcharxiv-cs-cv
12 May 2026
Research

MicroViTv2: Beyond the FLOPS for Edge Energy-Friendly Vision Transformers

DGX agent

arXiv:2605.10148v1 Announce Type: new Abstract: The Vision Transformer (ViT) achieves remarkable accuracy across visual tasks but remains computationally expensive for edge deployment. This paper pres

researcharxiv-cs-cv
12 May 2026
Research

ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality

DGX agent

arXiv:2605.09479v1 Announce Type: cross Abstract: We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information

researcharxiv-cs-cv
12 May 2026
Research

Model-based Dynamic 3D MRI Reconstructions using Neural Fields and Tensor Product Expansions

DGX agent

arXiv:2605.08275v1 Announce Type: cross Abstract: Conventional MRI reconstruction methods treat images and coil sensitivities as discrete objects, leading to high memory demands and limited structural

researcharxiv-cs-cv
12 May 2026
Applications

Modular Retrieval-Augmented Generalization for Human Action Recognition

DGX agent

arXiv:2605.08117v1 Announce Type: cross Abstract: Inertial Measurement Unit (IMU)-based Human Activity Recognition (HAR) aims to interpret and classify user behaviors from temporal motion signals. Rec

applicationsarxiv-cs-cv
12 May 2026
← Previous
1…182183184185186…263
Next →