AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Edge Detection for Organ Boundaries via Top Down Refinement and SubPixel Upsampling

DGX agent

arXiv:2508.06805v2 Announce Type: replace Abstract: Accurate localization of organ boundaries is critical in medical imaging for segmentation, registration, surgical planning, and radiotherapy. While

researcharxiv-cs-cv
11 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

DGX agent

arXiv:2605.07457v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as

local-aiarxiv-cs-cv
11 May 2026
Safety

EditTransfer++: Toward Faithful and Efficient Visual-Prompt-Guided Image Editing

DGX agent

arXiv:2605.07455v1 Announce Type: new Abstract: Visual-prompt-guided edit transfer aims to learn image transformations directly from example pairs, offering more precise and controllable editing than

safetyarxiv-cs-cv
11 May 2026
Research

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

DGX agent

arXiv:2605.07642v1 Announce Type: new Abstract: Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such a

researcharxiv-cs-cv
11 May 2026
Research

Enhancing Eye Movement Biometrics for User Authentication via Continuous Gaze Offset Score Fusion

DGX agent

arXiv:2605.06810v1 Announce Type: cross Abstract: Eye movement biometrics (EMB) use subject-specific gaze dynamics for user authentication and identification. Recent deep learning-based EMB systems ac

researcharxiv-cs-cv
11 May 2026
Research

Enhancing Federated Quadruplet Learning: Stochastic Client Selection and Embedding Stability Analysis

DGX agent

arXiv:2605.07888v1 Announce Type: cross Abstract: Federated Learning (FL) enables decentralised model training across distributed clients without requiring data centralisation. However, the generalisa

researcharxiv-cs-cv
11 May 2026
Local Ai

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency

DGX agent

arXiv:2605.06280v2 Announce Type: replace Abstract: Recent advancements in image animation have utilized diffusion models to breathe life into static images. However, existing controllable frameworks

local-aiarxiv-cs-cv
11 May 2026
Research

Explainable Part-Based Vehicle Classifier with Spatial Awareness

DGX agent

arXiv:2605.07831v1 Announce Type: new Abstract: In the area of Intelligent Transportation Systems (ITS), fine-grained vehicle classification systems play an essential role. Recently, the authors have

researcharxiv-cs-cv
11 May 2026
Research

EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding

DGX agent

arXiv:2605.07859v1 Announce Type: new Abstract: Driver cognitive distraction is a major cause of road collisions and remains difficult to detect. Unlike manual or visual distraction, cognitive distrac

researcharxiv-cs-cv
11 May 2026
Model Releases

Fine-tuning a vision-language model for fracture-surface morphology recognition

DGX agent

arXiv:2605.07145v1 Announce Type: cross Abstract: Vision-language models (VLMs) have shown strong potential for scientific image understanding, but general-purpose models often lack the domain-specifi

model-releasesarxiv-cs-cv
11 May 2026
Safety

Flatness and Gradient Alignment Are Both Necessary: Spectral-Aware Gradient-Aligned Exploration for Multi-Distribution Learning

DGX agent

arXiv:2605.07914v1 Announce Type: cross Abstract: Sharpness-aware and gradient-alignment methods have been shown to improve generalization, however each family of methods targets a single geometric pr

safetyarxiv-cs-cv
11 May 2026
Tutorials

From Pixels to Primitives: Scene Change Detection in 3D Gaussian Splatting

DGX agent

arXiv:2605.07203v1 Announce Type: new Abstract: Scene change detection methods built on Gaussian splatting universally follow a render-then-compare paradigm: the pre-change scene is rendered into 2D a

tutorialsarxiv-cs-cv
11 May 2026
Model Releases

From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data

DGX agent

arXiv:2605.07861v1 Announce Type: new Abstract: Makeup transfer aims to apply the makeup style of a reference portrait to a source portrait while preserving identity and background. Early methods form

model-releasesarxiv-cs-cv
11 May 2026
Hardware

Frozen Backpropagation: Relaxing Weight Symmetry in Deep Spiking Neural Networks

DGX agent

arXiv:2505.13741v2 Announce Type: replace Abstract: Direct training of Spiking Neural Networks (SNNs) on neuromorphic hardware can greatly reduce energy costs compared to GPU-based training. However,

hardwarearxiv-cs-cv
11 May 2026
Research

FS-I2P:A Hierarchical Focus-Sweep Registration Network with Dynamically Allocated Depth

DGX agent

arXiv:2605.07607v1 Announce Type: new Abstract: Image-to-point cloud registration is often challenged by viewpoint changes, cross-modal discrepancies, and repetitive textures, which induce scale ambig

researcharxiv-cs-cv
11 May 2026
Model Releases

GC-ART: Global Learnable Second-Order Rational Tone Curves for Illumination Robustness

DGX agent

arXiv:2605.07329v1 Announce Type: new Abstract: We introduce GC-ART (Global Curve Adaptive Rational Tone-mapping), a lightweight differentiable pre-processing module for robust image classification. G

model-releasesarxiv-cs-cv
11 May 2026
Agents

GEM: Generating LiDAR World Model via Deformable Mamba

DGX agent

arXiv:2605.07326v1 Announce Type: new Abstract: World models, which simulate environmental dynamics and generate sensor observations, are gaining increasing attention in autonomous driving. However, p

agentsarxiv-cs-cv
11 May 2026
Safety

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization

DGX agent

arXiv:2605.07399v1 Announce Type: new Abstract: Diffusion Vision-Language Models (dVLMs), built upon the non-causal foundations of Diffusion Large Language Models (dLLMs), have demonstrated remarkable

safetyarxiv-cs-cv
11 May 2026
Safety

GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval

DGX agent

arXiv:2509.23370v2 Announce Type: replace Abstract: The CLIP model has established itself as a cornerstone of large-scale retrieval systems. However, its performance often degrades under distributiona

safetyarxiv-cs-cv
11 May 2026
Research

GraphFusion3D: Dynamic Graph Attention Convolution with Adaptive Cross-Modal Transformer for 3D Object Detection

DGX agent

arXiv:2512.02991v2 Announce Type: replace Abstract: Despite significant progress in 3D object detection, point clouds remain challenging due to sparse data, incomplete structures, and limited semantic

researcharxiv-cs-cv
11 May 2026
Model Releases

Head Similarity: Modeling Structured Whole-Head Appearance Beyond Face Recognition

DGX agent

arXiv:2605.07766v1 Announce Type: new Abstract: Many vision applications require identity consistency beyond strict biometric recognition, especially under non-frontal views or when facial cues are mi

model-releasesarxiv-cs-cv
11 May 2026
Safety

HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models

DGX agent

arXiv:2605.07973v1 Announce Type: new Abstract: Text-to-image diffusion models can generate visually stunning images, yet, controlling what appears and how it appears, remains surprisingly difficult,

safetyarxiv-cs-cv
11 May 2026
Model Releases

Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models

DGX agent

arXiv:2605.07512v1 Announce Type: new Abstract: Class-incremental learning aims to continuously acquire new knowledge while preserving previously learned information, thereby mitigating catastrophic f

model-releasesarxiv-cs-cv
11 May 2026
Research

Hierarchical Perfusion Graphs for Tumor Heterogeneity Modeling in Glioma Molecular Subtyping

DGX agent

arXiv:2605.07156v1 Announce Type: new Abstract: Precise molecular subtyping of gliomas, including isocitrate dehydrogenase (IDH) mutation and 1p/19q codeletion, directly guides surgical and therapeuti

researcharxiv-cs-cv
11 May 2026
Research

High-Fidelity Surface Splatting-Based 3D Reconstruction from Multi-View Images

DGX agent

arXiv:2605.07254v1 Announce Type: new Abstract: Multi-view mesh reconstruction remains a core challenge in computer graphics and vision, especially for recovering high-frequency geometry from sparse o

researcharxiv-cs-cv
11 May 2026
Model Releases

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings

DGX agent

arXiv:2605.07492v1 Announce Type: new Abstract: The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanual

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

HumanNet: Scaling Human-centric Video Learning to One Million Hours

DGX agent

arXiv:2605.06747v1 Announce Type: new Abstract: Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, lea

model-releasesarxiv-cs-cv
11 May 2026
Research

ICDAR 2026 Competition on Writer Identification and Pen Classification from Hand-Drawn Circles

DGX agent

arXiv:2605.07816v1 Announce Type: new Abstract: This paper presents CircleID, a large-scale ICDAR 2026 competition on writer identification and pen classification from scanned hand-drawn circles. The

researcharxiv-cs-cv
11 May 2026
Research

ImplantMamba: Long-range Sequential Modeling Mamba For Dental Implant Position Prediction

DGX agent

arXiv:2605.07082v1 Announce Type: new Abstract: In the design of surgical guides for implant placement, determining the precise implant position is a critical step. However, the implant region itself

researcharxiv-cs-cv
11 May 2026
Tutorials

Implicit Multi-Camera System Calibration Using Gaussian Processes

DGX agent

arXiv:2605.07491v1 Announce Type: new Abstract: This paper proposes a novel framework for implicit multi-camera system calibration utilizing Gaussian Process (GP) regression. Conventional explicit cal

tutorialsarxiv-cs-cv
11 May 2026
Safety

InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-Localization

DGX agent

arXiv:2605.07099v1 Announce Type: new Abstract: Cross-view geo-localization (CVGL) is fundamental for precise localization and navigation in GPS-denied environments, aiming to match ground or UAV imag

safetyarxiv-cs-cv
11 May 2026
Research

InsHuman: Towards Natural and Identity-Preserving Human Insertion

DGX agent

arXiv:2605.07402v1 Announce Type: new Abstract: Human insertion aims to naturally place specific individuals into a target background. Although existing image editing models may have such ability, the

researcharxiv-cs-cv
11 May 2026
Safety

InterCoG: Towards Spatially Precise Image Editing with Interleaved Chain-of-Grounding Reasoning

DGX agent

arXiv:2603.01586v3 Announce Type: replace Abstract: Emerging unified editing models have demonstrated strong capabilities in general object editing tasks. However, it remains a significant challenge t

safetyarxiv-cs-cv
11 May 2026
Safety

Is the Future Compatible? Diagnosing Dynamic Consistency in World Action Models

DGX agent

arXiv:2605.07514v1 Announce Type: cross Abstract: World Action Models (WAMs) enable decision-making through imagined rollouts by predicting future observations and actions. However, the reliability of

safetyarxiv-cs-cv
11 May 2026
Applications

LAMES: A Large-Scale and Artisanal Mining Environmental Segmentation Dataset

DGX agent

arXiv:2605.07740v1 Announce Type: new Abstract: Mining operations are of utmost importance to the economy of some nations. However, such operations result in land-use change, very high energy consumpt

applicationsarxiv-cs-cv
11 May 2026
Applications

Large Video Planner Enables Generalizable Robot Control

DGX agent

arXiv:2512.15840v2 Announce Type: replace-cross Abstract: General-purpose robots require decision-making models that generalize across diverse tasks and environments. Recent works build robot foundati

applicationsarxiv-cs-cv
11 May 2026
Research

Learning Image-Adaptive Scale Fields for Metric Depth Recovery

DGX agent

arXiv:2605.07418v1 Announce Type: new Abstract: Monocular depth estimation (MDE) typically produces depth estimations that are defined up to an unknown scale or shift. When only sparse metric anchors

researcharxiv-cs-cv
11 May 2026
Safety

Learning to Track Instance from Single Nature Language Description

DGX agent

arXiv:2605.07064v1 Announce Type: new Abstract: How to achieve vision-language (VL) tracking using natural language descriptions from a video sequence extbf{without relying on any bounding-box ground

safetyarxiv-cs-cv
11 May 2026
Research

LENS: Low-Frequency Eigen Noise Shaping for Efficient Diffusion Sampling

DGX agent

arXiv:2605.07253v1 Announce Type: new Abstract: Distilled diffusion models accelerate image generation by reducing the number of denoising steps, but often suffer from degraded image quality. To mitig

researcharxiv-cs-cv
11 May 2026
Safety

Lightweight Unpaired Smartphone ISP Transfer with Semantic Pseudo-Pairing

DGX agent

arXiv:2605.07495v1 Announce Type: new Abstract: Unpaired smartphone ISP is a challenging problem due to the lack of scene and color alignment between RAW and target RGB images. Many existing methods e

safetyarxiv-cs-cv
11 May 2026
Local Ai

LoHGNet: Infrared Small Target Detection through Lorentz Geometric Encoding with High-Order Relation Learning

DGX agent

arXiv:2605.07213v1 Announce Type: new Abstract: Infrared small target detection (IRSTD) remains challenging due to the scarcity of useful target cues and the presence of severe background clutter. Mos

local-aiarxiv-cs-cv
11 May 2026
Tutorials

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute

DGX agent

arXiv:2605.06809v1 Announce Type: new Abstract: Transformers dominate video recognition. They split videos into tokens, and processing them has expensive superlinear computational cost. Yet videos are

tutorialsarxiv-cs-cv
11 May 2026
Research

Lossy Common Information in a Learnable Gray-Wyner Network

DGX agent

arXiv:2601.21424v3 Announce Type: replace-cross Abstract: Many computer vision tasks share substantial overlapping information, yet conventional codecs tend to ignore this, leading to redundant and in

researcharxiv-cs-cv
11 May 2026
Safety

Masks Can Talk: Extracting Structured Text Information from Single-Modal Images for Remote Sensing Change Detection

DGX agent

arXiv:2605.07178v1 Announce Type: new Abstract: Remote sensing change detection is pivotal for urban monitoring, disaster assessment, and environmental resource management. Yet, unimodal deep learning

safetyarxiv-cs-cv
11 May 2026
Research

Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers

DGX agent

arXiv:2605.06169v1 Announce Type: cross Abstract: Scaling Diffusion Transformers (DiTs) to hundreds of layers introduces a structural vulnerability: networks can enter a silent, mean-dominated collaps

researcharxiv-cs-cv
11 May 2026
Model Releases

MedVIGIL: Evaluating Trustworthy Medical VLMs Under Broken Visual Evidence

DGX agent

arXiv:2605.07919v1 Announce Type: new Abstract: Medical vision--language models (VLMs) are usually evaluated on intact image--question pairs, but trustworthy clinical use requires a stronger property:

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

MicroBi-ConvLSTM: An Ultra-Lightweight Efficient Model for Human Activity Recognition on Resource Constrained Devices

DGX agent

arXiv:2602.06523v2 Announce Type: replace Abstract: Human Activity Recognition (HAR) on resource constrained wearables requires models that balance accuracy against strict memory and computational bud

model-releasesarxiv-cs-cv
11 May 2026
Safety

Mind the Gap: Geometrically Accurate Generative Reconstruction from Disjoint Views

DGX agent

arXiv:2605.07550v1 Announce Type: new Abstract: 3D vision systems are fundamentally constrained by their reliance on visual overlap: reconstruction methods require it for geometric alignment, while ge

safetyarxiv-cs-cv
11 May 2026
← Previous
1…187188189190191…263
Next →