AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Gender Artifacts from Art History to Text-to-Image Generation

DGX agent

arXiv:2606.05829v1 Announce Type: new Abstract: Artistic styles are rooted in specific socio-historical contexts that encode social hierarchies, including distinct constructions of gender. Yet in AI r

researcharxiv-cs-cv
5 Jun 2026
Local Ai

GenTract: Generative Global Tractography

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2511.13183v2 Announce Type: replace Abstract: Tractography is the process of inferring the trajectories of white-matter pathways in the brain from diffusion magnetic resonance imaging (dMRI). Lo

local-aiarxiv-cs-cv
5 Jun 2026
Tutorials

Geodesic Flow Matching on a Riemannian Degradation Manifold for Blind Image Restoration

DGX agent

arXiv:2606.06278v1 Announce Type: new Abstract: Blind image restoration requires recovering clean images from observations corrupted by unknown and potentially mixed degradations. While recent determi

tutorialsarxiv-cs-cv
5 Jun 2026
Safety

Geometry-Aware Dataset Condensation for Diffusion Model Training

DGX agent

arXiv:2606.05883v1 Announce Type: new Abstract: Dataset condensation aims to construct compact datasets from real data via synthesis or selection. However, existing approaches are ill-suited for diffu

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework

DGX agent

arXiv:2603.08491v2 Announce Type: replace Abstract: Cross-modal Geo-localization (CMGL) matches ground-level text descriptions with geo-tagged aerial imagery, which is crucial for pedestrian navigatio

model-releasesarxiv-cs-cv
5 Jun 2026
Research

Global-Local Monte Carlo Tree Search in Vision-Language Models for Text-to-3D Indoor Scene Generation

DGX agent

arXiv:2606.06002v1 Announce Type: new Abstract: Large Vision-Language Models have achieved significant reasoning performance in various tasks.However, there are few studies on text-to-3D indoor scene

researcharxiv-cs-cv
5 Jun 2026
Research

GMBFormer: An NDVI-Guided Global Memory Bank Transformer for Urban Green-Space Extraction from Ultra-High-Resolution Imagery

DGX agent

arXiv:2606.06363v1 Announce Type: new Abstract: Urban green-space extraction from ultra-high-resolution (UHR) imagery is commonly performed patch by patch, which limits semantic reuse among spatially

researcharxiv-cs-cv
5 Jun 2026
Research

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention

DGX agent

arXiv:2606.06249v1 Announce Type: new Abstract: Transformer-based multimodal models rely on attention mechanisms to integrate information across heterogeneous modalities. Despite their success, existi

researcharxiv-cs-cv
5 Jun 2026
Hardware

GS-NFS: Bandwidth-adaptive Streaming of Dynamic Gaussian Splats and Point Clouds

DGX agent

arXiv:2606.05650v1 Announce Type: cross Abstract: Dynamic 3D Gaussian Splatting (3DGS) holds great promise as a 3D video streaming technology since it can represent complex 3D scenes with high fidelit

hardwarearxiv-cs-cv
5 Jun 2026
Model Releases

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

DGX agent

arXiv:2511.20158v2 Announce Type: replace Abstract: While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominan

model-releasesarxiv-cs-cv
5 Jun 2026
Research

HDST-GNN: Heterogeneous Dynamic Spatiotemporal Graph Neural Networks for Multi-Object Tracking in UAV Aerial Imagery

DGX agent

arXiv:2606.05587v1 Announce Type: new Abstract: Multi-object tracking (MOT) from UAV imagery presents unique challenges: altitude varies across sequences, objects are small and densely packed, and fre

researcharxiv-cs-cv
5 Jun 2026
Safety

HERO: Learning Humanoid End-Effector Control for Visual Whole-Body Open-Vocabulary Object Grasping

DGX agent

arXiv:2602.16705v3 Announce Type: replace-cross Abstract: Visual loco-manipulation of arbitrary in-the-wild objects requires accurate end-effector (EE) control and a generalizable understanding of the

safetyarxiv-cs-cv
5 Jun 2026
Research

Hierarchical Mask-Enhanced Dual Reconstruction Network for Few-Shot Fine-Grained Image Classification

DGX agent

arXiv:2506.20263v2 Announce Type: replace Abstract: Few-shot fine-grained image classification (FS-FGIC) is challenging as it requires distinguishing visually similar subclasses with extremely limited

researcharxiv-cs-cv
5 Jun 2026
Model Releases

HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps

DGX agent

arXiv:2601.02730v3 Announce Type: replace Abstract: Visual localization on standard-definition (SD) maps has emerged as a promising low-cost and scalable solution for autonomous driving. However, exis

model-releasesarxiv-cs-cv
5 Jun 2026
Research

HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes

DGX agent

arXiv:2606.06390v1 Announce Type: new Abstract: Indoor scene generation is crucial for robot simulation and modern interior design. However, complex layouts together with scarce 3D scene data make lea

researcharxiv-cs-cv
5 Jun 2026
Research

Horse Eye Blink Detection and Classification for Equine Affective State Assessment

DGX agent

arXiv:2606.05458v1 Announce Type: new Abstract: Automated detection of equine facial action units (AUs) is a promising yet under-explored avenue for pain and affective state assessment in horses. Half

researcharxiv-cs-cv
5 Jun 2026
Research

HyperVis: Continuous Latent Visual Relational Graphs on the Lorentz Hyperboloid for Compositional Reasoning

DGX agent

arXiv:2606.06100v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle with compositional reasoning that requires understanding inter-object relationships. A natural remedy is to injec

researcharxiv-cs-cv
5 Jun 2026
Research

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

DGX agent

arXiv:2606.05769v1 Announce Type: new Abstract: Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize inter

researcharxiv-cs-cv
5 Jun 2026
Applications

In-Context Multiple Instance Learning

DGX agent

arXiv:2606.06458v1 Announce Type: cross Abstract: Multiple Instance Learning (MIL) addresses problems where supervision is available at the level of bags of instances and has been successfully applied

applicationsarxiv-cs-cv
5 Jun 2026
Safety

Inverse Design of Realizable Metasurface based Absorbers using Improved Conditioning and Diversity Enhanced Progressively Growing GANs

DGX agent

arXiv:2606.05849v1 Announce Type: cross Abstract: Metasurfaces enable precise manipulation of electromagnetic waves for applications such as beam steering, sensing, and stealth technology. However, in

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing

DGX agent

arXiv:2606.05172v1 Announce Type: cross Abstract: Diffusion-based image editing has achieved strong visual fidelity under natural language instructions, yet most existing systems still operate at the

model-releasesarxiv-cs-cv
5 Jun 2026
Research

Knowledge Distillation for Visual Autoregressive Models

DGX agent

arXiv:2606.06078v1 Announce Type: new Abstract: Autoregressive (AR) image generation models are highly expressive but computationally intensive, motivating effective model compression. Knowledge disti

researcharxiv-cs-cv
5 Jun 2026
Model Releases

KV-Control: Parameter-Efficient K/V Injection for Trajectory-Controlled Text-to-Motion

DGX agent

arXiv:2606.05624v1 Announce Type: new Abstract: Text-conditioned 3D human motion models now synthesize plausible motions from prompts, but practical animation and embodied-agent workflows rarely stop

model-releasesarxiv-cs-cv
5 Jun 2026
Safety

LadderMan: Learning Humanoid Perceptive Ladder Climbing

DGX agent

arXiv:2606.05873v1 Announce Type: cross Abstract: Humanoid robots hold great promise for operating in human-centered environments, yet ladder climbing remains one of the most challenging tasks due to

safetyarxiv-cs-cv
5 Jun 2026
Research

Latent Implicit Visual Reasoning

DGX agent

arXiv:2512.21218v2 Announce Type: replace Abstract: While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning m

researcharxiv-cs-cv
5 Jun 2026
Applications

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

DGX agent

arXiv:2606.05833v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to m

applicationsarxiv-cs-cv
5 Jun 2026
Applications

Learning Predictive Visuomotor Coordination

DGX agent

arXiv:2503.23300v2 Announce Type: replace Abstract: Understanding and predicting human visuomotor coordination is crucial for applications in robotics, human-computer interaction, and assistive techno

applicationsarxiv-cs-cv
5 Jun 2026
Safety

Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation

DGX agent

arXiv:2606.06076v1 Announce Type: cross Abstract: While vision-language models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this to a perce

safetyarxiv-cs-cv
5 Jun 2026
Safety

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

DGX agent

arXiv:2606.05737v1 Announce Type: new Abstract: Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising. We argue that

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

LiAuto-GeoX: Efficient Grounded Driving Transformer

DGX agent

arXiv:2606.05774v1 Announce Type: new Abstract: Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for auton

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

LightVesselNet: An Ultra-Lightweight Sub-100K Parameter Network for Retinal Blood Vessel Segmentation

DGX agent

arXiv:2606.05354v1 Announce Type: new Abstract: Retinal blood vessel segmentation plays a vital role in the early detection of diabetic retinopathy and glaucoma. While recent deep learning models have

model-releasesarxiv-cs-cv
5 Jun 2026
Research

LLM-Conditioned Synthesis of Pathological Gaits via Structured Gait-Language Representations

DGX agent

arXiv:2606.06048v1 Announce Type: new Abstract: Pathological gait datasets remain scarce due to privacy, recruitment, cost, and movement variability. Our work presents a multimodal LLM-guided framewor

researcharxiv-cs-cv
5 Jun 2026
Model Releases

LLM-Guided ANN Index Optimization for Human-Object Interaction Retrieval

DGX agent

arXiv:2606.05489v1 Announce Type: new Abstract: Retrieval systems underpin modern AI applications -- spanning visual search, recommendation engines, and multi-modal question answering. Modern multi-st

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

DGX agent

arXiv:2606.06042v1 Announce Type: new Abstract: Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier fie

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

MAviS: A Multimodal Conversational Assistant For Avian Species

DGX agent

arXiv:2603.07294v2 Announce Type: replace Abstract: Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monit

model-releasesarxiv-cs-cv
5 Jun 2026
Research

Monte Carlo Steklov Operators for Large-Scale Geometry Processing in the Wild

DGX agent

arXiv:2606.05581v1 Announce Type: cross Abstract: Intrinsic methods fill the default toolbox for geometry processing on meshes. Intrinsic operators, in particular the Laplacian, underlie methods that

researcharxiv-cs-cv
5 Jun 2026
Research

MS-DKC: A Dataset Knowledge Card Framework for Designing and Adapting Medical Image Segmentation Models

DGX agent

arXiv:2606.06103v1 Announce Type: new Abstract: Medical image segmentation is often framed as a search for stronger architectures, but this can obscure a more fundamental question: what does the datas

researcharxiv-cs-cv
5 Jun 2026
Research

Multi-Task Crack Foundation Model for Engineering-Reliable Crack Representation and Topology Preservation in Civil Infrastructure

DGX agent

arXiv:2606.05641v1 Announce Type: new Abstract: Reliable crack assessment requires not only accurate pixel-level masks but also connected crack geometry and confidence estimates that remain stable und

researcharxiv-cs-cv
5 Jun 2026
Research

Multimodal Sexism Identification and Characterization using Large Language Models and Gradient Boosting

DGX agent

arXiv:2606.05997v1 Announce Type: new Abstract: We present the AILS-NTUA submission to the EXIST 2026 Lab at CLEF, addressing multimodal sexism identification and characterization in memes (Task 2) an

researcharxiv-cs-cv
5 Jun 2026
Research

Next-Generation Parallel Decoder for LPDR: Architectural Optimization and Class-Balanced GAN-Augmentation

DGX agent

arXiv:2606.05785v1 Announce Type: new Abstract: Real-Time License Plate Detection and Recognition (LPDR) forms the backbone of modern smart cities. Although the YOLOV5-PDLPR model substantially improv

researcharxiv-cs-cv
5 Jun 2026
Research

NIV: Neural Axis Variations for Variable Font Generation

DGX agent

arXiv:2606.05261v1 Announce Type: new Abstract: Variable fonts enable continuous variation of glyph geometry along semantic design axes such as weight, width, slant, and optical size. However, constru

researcharxiv-cs-cv
5 Jun 2026
Research

Noise-Adaptive Regularization for Robust Multi-Label Remote Sensing Image Classification

DGX agent

arXiv:2601.08446v2 Announce Type: replace Abstract: The development of reliable methods for multi-label classification (MLC) has become a prominent research direction in remote sensing (RS). As the sc

researcharxiv-cs-cv
5 Jun 2026
Model Releases

Noise-Aware Visual Representation Learning for Medical Visual Question Answering

DGX agent

arXiv:2606.05535v1 Announce Type: new Abstract: Medical visual question answering (Med-VQA) has strong potential for clinical decision support by enabling AI models to interpret medical images and ans

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Oklch+: A Three-Parameter Extension of Oklab for Improved Color Difference Prediction

DGX agent

arXiv:2606.05255v1 Announce Type: cross Abstract: Oklab and its cylindrical representation Oklch are widely adopted in interpolation and design workflows as perceptually motivated color spaces, but th

model-releasesarxiv-cs-cv
5 Jun 2026
Research

ORACLE-CT: Anatomy-Aware Support Pooling for CT Classification

DGX agent

arXiv:2606.05460v1 Announce Type: new Abstract: Abdominal CT disease classification is challenging because each scan is a large 3D volume with many possible findings, while diagnostic evidence is ofte

researcharxiv-cs-cv
5 Jun 2026
Research

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding

DGX agent

arXiv:2606.06485v1 Announce Type: new Abstract: Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual ques

researcharxiv-cs-cv
5 Jun 2026
Research

Parallel Jacobi Decoding for Fast Autoregressive Image Generation

DGX agent

arXiv:2606.05703v1 Announce Type: new Abstract: Autoregressive (AR) models have demonstrated remarkable performance in generating high-fidelity images. However, their inherently sequential next-token

researcharxiv-cs-cv
5 Jun 2026
Agents

PathWISE: Multi-Agent Cancer Pathway Triaging Ontology Learning from Clinical Flowcharts

DGX agent

arXiv:2605.25970v2 Announce Type: replace Abstract: Clinical pathways are disseminated as visual flowcharts where spatial topology, arrow direction, colour coding, and font weight encode critical tria

agentsarxiv-cs-cv
5 Jun 2026
← Previous
1…117118119120121…263
Next →