AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

DMDSC: A Dynamic-Margin Deep Simplex Classifier for Open-Set Recognition on Medical Image Datasets

DGX agent

arXiv:2605.00675v1 Announce Type: new Abstract: Medical imaging datasets are often characterized by extreme class imbalances, where rare pathologies are significantly underrepresented compared to comm

researcharxiv-cs-cv
4 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Hardware

DPU or GPU for Accelerating Neural Networks Inference -- Why not both? Split CNN Inference

DGX agent

arXiv:2605.00174v1 Announce Type: cross Abstract: Video and image streaming on edge devices requires low latency. To address this, Neural Networks (NNs) are widely used, and prior work mainly focuses

hardwarearxiv-cs-cv
4 May 2026
Model Releases

Driving with A Thousand Faces: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving

DGX agent

arXiv:2602.18757v2 Announce Type: replace Abstract: Human driving behavior is inherently diverse, yet most end-to-end autonomous driving (E2E-AD) systems learn a single average driving style, neglecti

model-releasesarxiv-cs-cv
4 May 2026
Model Releases

Efficient Spatio-Temporal Vegetation Pixel Classification with Vision Transformers

DGX agent

arXiv:2605.00296v1 Announce Type: new Abstract: Plant phenology-the study of recurrent life cycle events-is essential for understanding ecosystem dynamics and their responses to climate change impacts

model-releasesarxiv-cs-cv
4 May 2026
Research

Elimination Templates in Macaulay2

DGX agent

arXiv:2605.00278v1 Announce Type: cross Abstract: We introduce the package exttt{EliminationTemplates} for the Macaulay2 computer algebra system, which provides tools for constructing automatic solver

researcharxiv-cs-cv
4 May 2026
Research

End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer

DGX agent

arXiv:2605.00503v1 Announce Type: new Abstract: Autoregressive image modeling relies on visual tokenizers to compress images into compact latent representations. We design an end-to-end training pipel

researcharxiv-cs-cv
4 May 2026
Model Releases

Event-based Civil Infrastructure Visual Defect Detection: ev-CIVIL Dataset and Benchmark

DGX agent

arXiv:2504.05679v2 Announce Type: replace Abstract: Small unmanned aerial vehicle (UAV)-based visual inspections are a more efficient alternative to manual methods for examining civil structural defec

model-releasesarxiv-cs-cv
4 May 2026
Research

Exploring the Limits of End-to-End Feature-Affinity Propagation for Single-Point Supervised Infrared Small Target Detection

DGX agent

arXiv:2605.00722v1 Announce Type: new Abstract: Single-point supervised infrared small target detection (IRSTD) drastically reduces dense annotation costs. Current state-of-the-art (SOTA) methods achi

researcharxiv-cs-cv
4 May 2026
Model Releases

Faithful Extreme Image Rescaling with Learnable Reversible Transformation and Semantic Priors

DGX agent

arXiv:2605.00605v1 Announce Type: new Abstract: Most recent extreme rescaling methods struggle to preserve semantically consistent structures and produce realistic details, due to the severely ill-pos

model-releasesarxiv-cs-cv
4 May 2026
Local Ai

Federated Distillation for Whole Slide Image via Gaussian-Mixture Feature Alignment and Curriculum Integration

DGX agent

arXiv:2605.00578v1 Announce Type: new Abstract: Federated learning (FL) offers a promising framework for collaborative digital pathology by enabling model training across institutions. However, real-w

local-aiarxiv-cs-cv
4 May 2026
Applications

FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting

DGX agent

arXiv:2605.00177v1 Announce Type: cross Abstract: We consider the problem of synthesizing photorealistic, physically plausible combustion effects in in-the-wild 3D scenes. Traditional CFD and graphics

applicationsarxiv-cs-cv
4 May 2026
Research

Flow matching for Sentinel-2 super-resolution: implementation, application, and implications

DGX agent

arXiv:2605.00367v1 Announce Type: new Abstract: Developing robust techniques for super-resolution of satellite imagery involves navigating commonly observed trade-offs between spectral fidelity and pe

researcharxiv-cs-cv
4 May 2026
Safety

Foundation AI Models for Aerosol Optical Depth Estimation from PACE Satellite Data

DGX agent

arXiv:2605.00678v1 Announce Type: new Abstract: Aerosol Optical Depth (AOD) retrieval is essential for Earth observation, supporting applications from air quality monitoring to climate studies. Conven

safetyarxiv-cs-cv
4 May 2026
Safety

FreeRet: MLLMs as Training-Free Retrievers

DGX agent

arXiv:2509.24621v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are emerging as versatile foundations for mixed-modality retrieval. Yet, they often require heavy post-hoc

safetyarxiv-cs-cv
4 May 2026
Model Releases

From Images2Mesh: A 3D Surface Reconstruction Pipeline for Non-Cooperative Space Objects

DGX agent

arXiv:2605.00147v1 Announce Type: new Abstract: On-orbit inspection imagery is crucial as it enables characterization of non-cooperative resident space objects, providing the geometry and structural c

model-releasesarxiv-cs-cv
4 May 2026
Research

From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models

DGX agent

arXiv:2605.00474v1 Announce Type: new Abstract: Modern vision models achieve remarkable accuracy, but explaining where evidence arises, what the model encodes, and how internal computations assemble t

researcharxiv-cs-cv
4 May 2026
Research

GAFSV-Net: A Vision Framework for Online Signature Verification

DGX agent

arXiv:2605.00120v1 Announce Type: new Abstract: Online signature verification (OSV) requires distinguishing skilled forgeries from genuine samples under high intra-class variability and with very few

researcharxiv-cs-cv
4 May 2026
Local Ai

Gated Differential Linear Attention: A Linear-Time Decoder for High-Fidelity Medical Segmentation

DGX agent

arXiv:2603.02727v4 Announce Type: replace Abstract: Medical image segmentation requires models that preserve fine anatomical boundaries while remaining practical for clinical deployment. Transformers

local-aiarxiv-cs-cv
4 May 2026
Research

GMGaze: MoE-Based Context-Aware Gaze Estimation with CLIP and Multiscale Transformer

DGX agent

arXiv:2605.00799v1 Announce Type: new Abstract: Gaze estimation methods commonly use facial appearances to predict the direction of a person gaze. However, previous studies show three major challenges

researcharxiv-cs-cv
4 May 2026
Applications

GOR-IS: 3D Gaussian Object Removal in the Intrinsic Space

DGX agent

arXiv:2605.00498v1 Announce Type: new Abstract: Recent advances in Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have made it standard practice to reconstruct 3D scenes from multi-vie

applicationsarxiv-cs-cv
4 May 2026
Research

High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions

DGX agent

arXiv:2605.00496v1 Announce Type: new Abstract: Understanding human actions from visual observations is essential for human--robot interaction, particularly when semantic interpretation of unfamiliar

researcharxiv-cs-cv
4 May 2026
Model Releases

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

DGX agent

arXiv:2507.01955v3 Announce Type: replace Abstract: Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond que

model-releasesarxiv-cs-cv
4 May 2026
Applications

IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations

DGX agent

arXiv:2605.00526v1 Announce Type: new Abstract: Suspect face generation remains a technical challenge in crime investigations. Traditional sketch-drawing workflows suffer from low efficiency and quali

applicationsarxiv-cs-cv
4 May 2026
Research

Image Score: Learning and Evaluating Human Preferences for Mercari Search

DGX agent

arXiv:2408.11349v2 Announce Type: replace Abstract: Mercari is the largest C2C e-commerce marketplace in Japan, having more than 20 million active monthly users. Search being the fundamental way to di

researcharxiv-cs-cv
4 May 2026
Research

Information-geometric adaptive sampling for graph diffusion

DGX agent

arXiv:2605.00250v1 Announce Type: cross Abstract: Standard diffusion models for graph generation typically rely on uniform time-stepping, an approach that overlooks the non-homogeneous dynamics of dis

researcharxiv-cs-cv
4 May 2026
Safety

InpaintSLat: Inpainting Structured 3D Latents via Initial Noise Optimization

DGX agent

arXiv:2605.00664v1 Announce Type: new Abstract: We present a training-free approach for controllable 3D inpainting based on initial noise optimization. In the structured 3D latent diffusion framework,

safetyarxiv-cs-cv
4 May 2026
Research

Instance-Aware Pseudo-Labeling and Class-Focused Contrastive Learning for Weakly Supervised Domain Adaptive Segmentation of Electron Microscopy

DGX agent

arXiv:2510.16450v2 Announce Type: replace Abstract: Annotation-efficient segmentation of the numerous mitochondria instances from various electron microscopy (EM) images is highly valuable for biologi

researcharxiv-cs-cv
4 May 2026
Research

Intrinsic Gradient Suppression for Label-Noise Prompt Tuning in Vision-Language Models

DGX agent

arXiv:2605.00591v1 Announce Type: new Abstract: Contrastive vision-language models like CLIP exhibit remarkable zero-shot generalization. However, prompt tuning remains highly sensitive to label noise

researcharxiv-cs-cv
4 May 2026
Research

It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models

DGX agent

arXiv:2601.00090v2 Announce Type: replace Abstract: Contemporary text-to-image models exhibit a surprising degree of mode collapse, as can be seen when sampling several images given the same text prom

researcharxiv-cs-cv
4 May 2026
Model Releases

Jailbreaking Vision-Language Models Through the Visual Modality

DGX agent

arXiv:2605.00583v1 Announce Type: new Abstract: The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak atta

model-releasesarxiv-cs-cv
4 May 2026
Research

LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping

DGX agent

arXiv:2511.08156v2 Announce Type: replace Abstract: Land Use and Land Cover (LULC) mapping is a fundamental task in Earth Observation (EO). However, current LULC models are typically developed for a s

researcharxiv-cs-cv
4 May 2026
Safety

Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding

DGX agent

arXiv:2605.00642v1 Announce Type: cross Abstract: Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capabili

safetyarxiv-cs-cv
4 May 2026
Safety

Learning Coarse-to-Fine Osteoarthritis Representations under Noisy Hierarchical Labels

DGX agent

arXiv:2605.00718v1 Announce Type: new Abstract: Knee osteoarthritis (OA) assessment involves a natural but often underused label hierarchy: a coarse binary OA decision and a fine-grained Kellgren--Law

safetyarxiv-cs-cv
4 May 2026
Model Releases

Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis

DGX agent

arXiv:2605.00448v1 Announce Type: new Abstract: The deployment of artificial intelligence in medical imaging is hindered by high computational complexity and resource-intensive processing of volumetri

model-releasesarxiv-cs-cv
4 May 2026
Model Releases

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation

DGX agent

arXiv:2605.00051v1 Announce Type: new Abstract: Anticipating traffic accidents is a critical yet unresolved problem for autonomous driving, hindered by the inherent complexity of modeling interactions

model-releasesarxiv-cs-cv
4 May 2026
Model Releases

Learning Locally, Revising Globally: Global Reviser for Federated Learning with Noisy Labels

DGX agent

arXiv:2412.00452v2 Announce Type: replace-cross Abstract: Conventioanl federated learning (FL) heavily depends on high-quality labels, which are often impractical in the real world, leading to the fed

model-releasesarxiv-cs-cv
4 May 2026
Safety

Learning physically grounded traffic accident reconstruction from public accident reports

DGX agent

arXiv:2605.00050v1 Announce Type: cross Abstract: Traffic accidents are routinely documented in textual reports, yet physically grounded accident reconstruction remains difficult because detailed scen

safetyarxiv-cs-cv
4 May 2026
Research

Let ViT Speak: Generative Language-Image Pre-training

DGX agent

arXiv:2605.00809v1 Announce Type: new Abstract: In this paper, we present extbf{Gen}erative extbf{L}anguage-extbf{I}mage extbf{P}re-training (GenLIP), a minimalist generative pretraining framework for

researcharxiv-cs-cv
4 May 2026
Research

Leveraging Vision-Language Models as Weak Annotators in Active Learning

DGX agent

arXiv:2605.00480v1 Announce Type: new Abstract: Active learning aims to reduce annotation cost by selectively querying informative samples for supervision under a limited labeling budget. In this work

researcharxiv-cs-cv
4 May 2026
Applications

LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal Observations

DGX agent

arXiv:2605.00434v1 Announce Type: new Abstract: Real-world multimodal learning is often hindered by missing modalities. While Incomplete Multimodal Learning (IML) has gained traction, existing methods

applicationsarxiv-cs-cv
4 May 2026
Local Ai

Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation

DGX agent

arXiv:2605.00244v1 Announce Type: cross Abstract: We introduce Lucid-XR, a generative data engine for creating diverse and realistic-looking multi-modal data to train real-world robotic systems. At th

local-aiarxiv-cs-cv
4 May 2026
Tutorials

MAEPose: Self-Supervised Spatiotemporal Learning for Human Pose Estimation on mmWave Video

DGX agent

arXiv:2605.00242v1 Announce Type: new Abstract: Millimetre-wave (mmWave) radar offers a more privacy-preserving alternative to RGB-based human pose estimation. However, existing methods typically rely

tutorialsarxiv-cs-cv
4 May 2026
Model Releases

Make Your LVLM KV Cache More Lightweight

DGX agent

arXiv:2605.00789v1 Announce Type: new Abstract: Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency

model-releasesarxiv-cs-cv
4 May 2026
Agents

Map2World: Segment Map Conditioned Text to 3D World Generation

DGX agent

arXiv:2605.00781v1 Announce Type: new Abstract: 3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world gener

agentsarxiv-cs-cv
4 May 2026
Applications

MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video

DGX agent

arXiv:2605.00495v1 Announce Type: cross Abstract: Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound producti

applicationsarxiv-cs-cv
4 May 2026
Research

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

DGX agent

arXiv:2605.00431v1 Announce Type: cross Abstract: Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model ro

researcharxiv-cs-cv
4 May 2026
Research

Modeling Subjective Urban Perception with Human Gaze

DGX agent

arXiv:2605.00764v1 Announce Type: new Abstract: Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computationa

researcharxiv-cs-cv
4 May 2026
Local Ai

MSACT: Multistage Spatial Alignment for Stable Low-Latency Fine Manipulation

DGX agent

arXiv:2605.00475v1 Announce Type: cross Abstract: Real-world fine manipulation, particularly in bimanual manipulation, typically requires low-latency control and stable visual localization, while coll

local-aiarxiv-cs-cv
4 May 2026
← Previous
1…201202203204205…261
Next →