AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

Driving Intents Amplify Planning-Oriented Reinforcement Learning

DGX agent

arXiv:2605.12625v1 Announce Type: cross Abstract: Continuous-action policies trained on a single demonstrated trajectory per scene suffer from mode collapse: samples cluster around the demonstrated ma

safetyarxiv-cs-cv
14 May 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Dual-Pathway Circuits of Object Hallucination in Vision-Language Models

DGX agent

arXiv:2605.13156v1 Announce Type: new Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in bridging visual perception and natural language understanding, enabling a wid

researcharxiv-cs-cv
14 May 2026
Research

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation

DGX agent

arXiv:2605.13122v1 Announce Type: new Abstract: Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural langu

researcharxiv-cs-cv
14 May 2026
Model Releases

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

DGX agent

arXiv:2605.13062v1 Announce Type: new Abstract: Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, e

model-releasesarxiv-cs-cv
14 May 2026
Local Ai

EDITS: Enhancing Dataset Distillation with Implicit Textual Semantics

DGX agent

arXiv:2509.13858v2 Announce Type: replace Abstract: Dataset distillation aims to synthesize a compact dataset from the original large-scale one, enabling highly efficient learning while preserving com

local-aiarxiv-cs-cv
14 May 2026
Applications

EgoForce: Robust Online Egocentric Motion Reconstruction via Diffusion Forcing

DGX agent

arXiv:2605.13041v1 Announce Type: new Abstract: With recent advances in embodied agents and AR devices, egocentric observations are readily available as input for real-world interactive online applica

applicationsarxiv-cs-cv
14 May 2026
Model Releases

Energy Scaling Laws for Diffusion Models: Quantifying Compute in Image Generation

DGX agent

arXiv:2511.17031v2 Announce Type: replace-cross Abstract: The rapidly growing computational demands of diffusion models for image generation have raised significant concerns about energy consumption a

model-releasesarxiv-cs-cv
14 May 2026
Tutorials

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding

DGX agent

arXiv:2605.13803v1 Announce Type: new Abstract: Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the qu

tutorialsarxiv-cs-cv
14 May 2026
Model Releases

Exploring Multimodal LMMs for Online Episodic Memory Question Answering on the Edge

DGX agent

arXiv:2602.22455v2 Announce Type: replace Abstract: We investigate the feasibility of using Multimodal Large Language Models (MLLMs) for real-time online episodic memory question answering. While clou

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

Fast and Compact Graph Cuts for the Boykov-Kolmogorov Algorithm

DGX agent

arXiv:2605.13402v1 Announce Type: new Abstract: Computing a minimum s-t cut in a graph is a solution to a wide range of computer vision problems, and is often done using the Boykov-Kolmogorov (BK) alg

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

FedHPro: Federated Hyper-Prototype Learning via Gradient Matching

DGX agent

arXiv:2605.13475v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative training of distributed clients while protecting privacy. To enhance generalization capability in FL, prot

model-releasesarxiv-cs-cv
14 May 2026
Agents

FIKA-Bench: From Fine-grained Recognition to Fine-Grained Knowledge Acquisition

DGX agent

arXiv:2605.13193v1 Announce Type: new Abstract: Fine-grained recognition in everyday life is often not a closed-book classification problem: when encountering unfamiliar objects, humans actively searc

agentsarxiv-cs-cv
14 May 2026
Hardware

Flow Augmentation and Knowledge Distillation for Lightweight Face Presentation Attack Detection

DGX agent

arXiv:2605.13108v1 Announce Type: new Abstract: Face presentation attack detection (FacePAD) remains challenging under diverse spoofing representation, including 2D print and replay, 3D mask-based spo

hardwarearxiv-cs-cv
14 May 2026
Research

Flow Matching with Uncertainty Quantification and Guidance

DGX agent

arXiv:2602.10326v2 Announce Type: replace Abstract: Despite the remarkable success of sampling-based generative models such as flow matching, they can still produce samples of inconsistent or degraded

researcharxiv-cs-cv
14 May 2026
Applications

FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection

DGX agent

arXiv:2509.23056v2 Announce Type: replace Abstract: Remote sensing object detection is a critical technology for real-world applications such as natural resource monitoring, traffic management, and UA

applicationsarxiv-cs-cv
14 May 2026
Model Releases

From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning

DGX agent

arXiv:2603.26839v2 Announce Type: replace-cross Abstract: How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce ex

model-releasesarxiv-cs-cv
14 May 2026
Local Ai

GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose Estimation

DGX agent

arXiv:2605.13151v1 Announce Type: new Abstract: Category-agnostic pose estimation (CAPE) aims to localize keypoints on query images from arbitrary categories, using only a few annotated support exampl

local-aiarxiv-cs-cv
14 May 2026
Research

Generative Motion In-betweening by Diffusion over Continuous Implicit Representations

DGX agent

arXiv:2605.12778v1 Announce Type: cross Abstract: Recent advances in generative models have yielded impressive progress on motion in-betweening, allowing for more complex, varied, and realistic motion

researcharxiv-cs-cv
14 May 2026
Safety

Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception

DGX agent

arXiv:2605.13755v1 Announce Type: new Abstract: In recent years, autonomous driving has significantly in creased the demand for high-quality data to train 2D and 3D perception models for safety-critic

safetyarxiv-cs-cv
14 May 2026
Model Releases

GeomHair: Reconstruction of Hair Strands from Colorless 3D Scans

DGX agent

arXiv:2505.05376v3 Announce Type: replace Abstract: We propose a novel method that reconstructs hair strands directly from colorless 3D scans by leveraging multi-modal hair orientation extraction. Hai

model-releasesarxiv-cs-cv
14 May 2026
Tutorials

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

DGX agent

arXiv:2602.17555v3 Announce Type: replace Abstract: Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. C

tutorialsarxiv-cs-cv
14 May 2026
Agents

GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion

DGX agent

arXiv:2605.12957v1 Announce Type: new Abstract: Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains

agentsarxiv-cs-cv
14 May 2026
Model Releases

GuardMarkGS: Unified Ownership Tracing and Edit Deterrence for 3D Gaussian Splatting

DGX agent

arXiv:2605.12919v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is becoming a practical representation for novel view synthesis, but its growing adoption, together with rapid advances in

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

DGX agent

arXiv:2605.13632v1 Announce Type: cross Abstract: In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied

model-releasesarxiv-cs-cv
14 May 2026
Research

Guidestar-Free Adaptive Optics with Asymmetric Apertures

DGX agent

arXiv:2602.07029v3 Announce Type: replace-cross Abstract: This work introduces the first closed-loop adaptive optics (AO) system capable of optically correcting aberrations in real-time without a guid

researcharxiv-cs-cv
14 May 2026
Model Releases

HADAR-Based Thermal Infrared Hyperspectral Image Restoration

DGX agent

arXiv:2605.13664v1 Announce Type: new Abstract: Thermal-infrared (TIR) hyperspectral imagery (HSI) provides critical scene information for various applications. However, its practical utility is sever

model-releasesarxiv-cs-cv
14 May 2026
Research

HarmoGS: Robust 3D Gaussian Splatting in the Wild via Conflict-Aware Gradient Harmonization

DGX agent

arXiv:2605.13073v1 Announce Type: new Abstract: In-the-wild 3D Gaussian Splatting remains challenging due to transient distractors and illumination-induced cross-view appearance inconsistencies. Exist

researcharxiv-cs-cv
14 May 2026
Safety

HIR-ALIGN: Enhancing Hyperspectral Image Restoration via Diffusion-Based Data Generation

DGX agent

arXiv:2605.13581v1 Announce Type: new Abstract: Hyperspectral image (HSI) restoration is crucial for reliable analysis, as real HSIs suffer from degradations like noise, blur, and resolution loss. How

safetyarxiv-cs-cv
14 May 2026
Research

Human face perception reflects inverse-generative and naturalistic discriminative objectives

DGX agent

arXiv:2605.12619v1 Announce Type: cross Abstract: The perceptual representations supporting our ability to recognize faces remain a computational mystery. Deep neural networks offer mechanistic hypoth

researcharxiv-cs-cv
14 May 2026
Model Releases

ImageAttributionBench: How Far Are We from Generalizable Attribution?

DGX agent

arXiv:2605.12967v1 Announce Type: new Abstract: The rapid advancement of generative AI has enabled the creation of highly realistic and diverse synthetic images, posing critical challenges for image p

model-releasesarxiv-cs-cv
14 May 2026
Research

Img2CADSeq: Image-to-CAD Generation via Sequence-Based Diffusion

DGX agent

arXiv:2605.13293v1 Announce Type: new Abstract: Boundary Representation (BRep) is the standard format for Computer-Aided Design (CAD), yet reconstructing high-quality BReps from single-view images rem

researcharxiv-cs-cv
14 May 2026
Research

Improved monocular depth prediction using distance transform over pre-semantic contours with self-supervised neural networks

DGX agent

arXiv:2605.08320v1 Announce Type: cross Abstract: Monocular depth estimation (MDE) with self-supervised training approaches struggles in low-texture areas, where photometric losses may lead to ambiguo

researcharxiv-cs-cv
14 May 2026
Tutorials

Inference-Time Dynamic Modality Selection for Incomplete Multimodal Classification

DGX agent

arXiv:2601.22853v3 Announce Type: replace Abstract: Multimodal deep learning (MDL) has achieved remarkable success across various domains, yet its practical deployment is often hindered by incomplete

tutorialsarxiv-cs-cv
14 May 2026
Safety

Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models

DGX agent

arXiv:2605.12725v1 Announce Type: new Abstract: Recent video anomaly detection research has expanded rapidly with an emphasis on general models of normality intended to work across many different scen

safetyarxiv-cs-cv
14 May 2026
Research

JANUS: Anatomy-Conditioned Gating for Robust CT Triage Under Distribution Shift

DGX agent

arXiv:2605.13813v1 Announce Type: new Abstract: Automated CT triage requires models that are simultaneously accurate across diverse pathologies and reliable under institutional shift. While Vision Tra

researcharxiv-cs-cv
14 May 2026
Research

Just Ask for a Table: A Thirty-Token User Prompt Defeats Sponsored Recommendations in Twelve LLMs

DGX agent

arXiv:2605.12772v1 Announce Type: new Abstract: Wu et al. (2026) showed that most frontier large language models (LLMs) recommend a sponsored, roughly twice-as-expensive flight when their system promp

researcharxiv-cs-cv
14 May 2026
Model Releases

KamonBench: A Grammar-Based Dataset for Evaluating Compositional Factor Recovery in Vision-Language Models

DGX agent

arXiv:2605.13322v1 Announce Type: new Abstract: Kamon (family crests) are an important part of Japanese culture and a natural test case for compositional visual recognition: each crest combines a smal

model-releasesarxiv-cs-cv
14 May 2026
Research

Learning to Optimize Radiotherapy Plans via Fluence Maps Diffusion Model Generation and LSTM-based Optimization

DGX agent

arXiv:2605.13713v1 Announce Type: new Abstract: Volumetric Modulated Arc Therapy (VMAT) is a cornerstone of modern radiation therapy, enabling highly conformal tumor irradiation and healthy-tissue spa

researcharxiv-cs-cv
14 May 2026
Local Ai

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models

DGX agent

arXiv:2605.13080v1 Announce Type: new Abstract: When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their inten

local-aiarxiv-cs-cv
14 May 2026
Model Releases

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

DGX agent

arXiv:2505.15616v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to r

model-releasesarxiv-cs-cv
14 May 2026
Research

LEXI-SG: Monocular 3D Scene Graph Mapping with Room-Guided Feed-Forward Reconstruction

DGX agent

arXiv:2605.13741v1 Announce Type: cross Abstract: Scene graphs are becoming a standard representation for robot navigation, providing hierarchical geometric and semantic scene understanding. However,

researcharxiv-cs-cv
14 May 2026
Local Ai

LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters

DGX agent

arXiv:2605.13163v1 Announce Type: cross Abstract: Foundation models and low-rank adapters enable efficient on-device generative AI but raise risks such as intellectual property leakage and model recov

local-aiarxiv-cs-cv
14 May 2026
Research

M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement

DGX agent

arXiv:2605.12556v1 Announce Type: new Abstract: Low-light image enhancement is challenging due to complex degradations, including amplified noise, artifacts, and color distortion. While Retinex-based

researcharxiv-cs-cv
14 May 2026
Research

M3Net: A Macro-to-Meso-to-Micro Clinical-inspired Hierarchical 3D Network for Pulmonary Nodule Classification

DGX agent

arXiv:2605.12570v1 Announce Type: new Abstract: The accurate classification of benign and malignant pulmonary nodules in CT scans is critical for early lung cancer screening, yet remains challenging d

researcharxiv-cs-cv
14 May 2026
Research

Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters

DGX agent

arXiv:2512.16767v2 Announce Type: replace Abstract: Posing 3D characters is a fundamental task in computer graphics. However, existing paradigms, ranging from traditional auto-rigging to recent pose-c

researcharxiv-cs-cv
14 May 2026
Research

MambaPanoptic: A Vision Mamba-based Structured State Space Framework for Panoptic Segmentation

DGX agent

arXiv:2605.12640v1 Announce Type: new Abstract: Panoptic segmentation requires the simultaneous recognition of countable thing instances and amorphous stuff regions, placing joint demands on long-rang

researcharxiv-cs-cv
14 May 2026
Model Releases

MedCore: Boundary-Preserving Medical Core Pruning for MedSAM

DGX agent

arXiv:2605.13688v1 Announce Type: new Abstract: Medical segmentation foundation models such as SAM and MedSAM provide strong prompt-driven segmentation, but their image encoders are still too large fo

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

DGX agent

arXiv:2603.24649v2 Announce Type: replace Abstract: Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition.

model-releasesarxiv-cs-cv
14 May 2026
← Previous
1…172173174175176…263
Next →