AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Local Ai

SPIRONet: Spatial-Frequency Learning and Graph-based Channel Interaction Network for Vessel Segmentation

DGX agent

arXiv:2406.19749v2 Announce Type: replace-cross Abstract: Automatic vessel segmentation plays a pivotal role in the development of next-generation interventional navigation systems for surgical roboti

local-aiarxiv-cs-cv
9 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision Encoders

DGX agent

arXiv:2606.08634v1 Announce Type: new Abstract: The rapid advancement of generative models has blurred the boundary between synthetic and real imagery, creating an urgent need for reliable deepfake de

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization

DGX agent

arXiv:2606.09091v1 Announce Type: cross Abstract: On-policy distillation (OPD) has recently emerged as an important post-training paradigm. By using a stronger teacher model to provide dense, fine-gra

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

Stain-Aware Wavelet Regularization for Instant Adversarial Purification in Histopathology

DGX agent

arXiv:2606.08745v1 Announce Type: new Abstract: Deep learning has become prevalent in computational pathology pipelines that support tasks such as cancer screening and digital pathology analysis. Howe

safetyarxiv-cs-cv
9 Jun 2026
Research

Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation

DGX agent

arXiv:2606.07647v1 Announce Type: new Abstract: Large vision language models (LVLMs) have made rapid advancements and are deployed across various applications, yet hallucinations remain a major challe

researcharxiv-cs-cv
9 Jun 2026
Research

STGBD-Net: Spatio-temporal Gradient Basis Decomposition Network for Infrared Small Target Detection

DGX agent

arXiv:2512.03470v5 Announce Type: replace Abstract: A key challenge in infrared small target detection (IRSTD) is that weak target signal responses are easily obscured by strong background clutter, fr

researcharxiv-cs-cv
9 Jun 2026
Model Releases

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

DGX agent

arXiv:2606.09547v1 Announce Type: new Abstract: Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking

DGX agent

arXiv:2606.07689v1 Announce Type: new Abstract: Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with r

model-releasesarxiv-cs-cv
9 Jun 2026
Hardware

SwiftVR: Real-Time One-Step Generative Video Restoration

DGX agent

arXiv:2606.09516v1 Announce Type: new Abstract: Real-time video restoration (VR) for live streams requires high-resolution outputs under strict per-frame latency constraints. Existing one-step diffusi

hardwarearxiv-cs-cv
9 Jun 2026
Agents

Taming Perception Jitter: Uncertainty-Aware LiDAR Object Detection for Reliable Motion Classification

DGX agent

arXiv:2606.09350v1 Announce Type: cross Abstract: Reliable motion classification is critical for autonomous driving, as false dynamic predictions of static objects can cascade into unnecessary planner

agentsarxiv-cs-cv
9 Jun 2026
Applications

TBD-VLA: Temporal Block Diffusion Vision Language Action Model

DGX agent

arXiv:2606.07895v1 Announce Type: new Abstract: Discrete Vision-Language-Action (VLA) models typically formulate action generation as next-token prediction over discretized action spaces, conditioning

applicationsarxiv-cs-cv
9 Jun 2026
Local Ai

Temporal-Aware Reasoning Optimization for Video Temporal Grounding

DGX agent

arXiv:2606.09248v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) have achieved remarkable progress in video temporal grounding with reinforcement learning for generating reaso

local-aiarxiv-cs-cv
9 Jun 2026
Research

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning

DGX agent

arXiv:2606.08231v1 Announce Type: new Abstract: Test-time Scaling (TTS) has emerged as a pivotal research direction for enhancing model performance by dynamically allocating computational resources du

researcharxiv-cs-cv
9 Jun 2026
Local Ai

The Need for Neural ISP in the Small-Pixel Era: How Shrinking Pixels Push Optics to the Limit and Neural Restoration Pushes Back

DGX agent

arXiv:2606.07675v1 Announce Type: cross Abstract: Smartphone telephoto cameras are approaching a 'telephoto physics wall': as pixel pitches shrink toward sub-0.5 micron, the optics remain limited by g

local-aiarxiv-cs-cv
9 Jun 2026
Local Ai

Thinking Without Images: Internalizing Visual Manipulation with On-Policy Self-Distillation

DGX agent

arXiv:2606.08719v1 Announce Type: new Abstract: ''Thinking with Images'' has emerged as an effective paradigm for fine-grained visual reasoning: by explicitly zooming into relevant regions and reasoni

local-aiarxiv-cs-cv
9 Jun 2026
Research

TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation

DGX agent

arXiv:2606.08260v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separa

researcharxiv-cs-cv
9 Jun 2026
Research

Toward Scalable Co-located Practical Learning: Assisting with Computer Vision and Multimodal Analytics

DGX agent

arXiv:2603.13679v2 Announce Type: replace-cross Abstract: Co-located practical learning leaves evidence in visible actions around patients, task resources and room zones, but these traces are often re

researcharxiv-cs-cv
9 Jun 2026
Safety

Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction

DGX agent

arXiv:2606.08566v1 Announce Type: new Abstract: Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionally rich descriptions for videos. Existing E

safetyarxiv-cs-cv
9 Jun 2026
Safety

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

DGX agent

arXiv:2511.05017v2 Announce Type: replace Abstract: Hallucinations in Large Vision-Language Models (LVLMs) remain a persistent challenge, often stemming from inadequate integration of visual informati

safetyarxiv-cs-cv
9 Jun 2026
Model Releases

Training-Free Generalized Few-Shot Segmentation through Open-Vocabulary Semantic Arbitration

DGX agent

arXiv:2606.09474v1 Announce Type: new Abstract: Generalized Few-Shot Semantic Segmentation (GFSS) has traditionally been approached as a representation-learning problem, requiring task-specific adapta

model-releasesarxiv-cs-cv
9 Jun 2026
Local Ai

Trajectory Optimization in Single and Dual-UAV Bearing-Only Target Localization

DGX agent

arXiv:2606.09188v1 Announce Type: cross Abstract: Bearing-only target localization is a fundamental problem in optical measurement and finds extensive applications in unmanned aerial vehicle (UAV) tec

local-aiarxiv-cs-cv
9 Jun 2026
Research

Trustworthy Visual Predicates for Robust Manipulation Understanding under Degradation

DGX agent

arXiv:2606.08121v1 Announce Type: new Abstract: Manipulation understanding requires reliable relational evidence, such as contact, support, containment, motion coupling, grasp, release, and active-han

researcharxiv-cs-cv
9 Jun 2026
Hardware

TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolution

DGX agent

arXiv:2606.09608v1 Announce Type: new Abstract: Diffusion-based generative models have achieved remarkable success in real-world image super-resolution (SR). With tiled diffusion techniques, these mod

hardwarearxiv-cs-cv
9 Jun 2026
Research

TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding

DGX agent

arXiv:2606.08464v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models. However, when applied to multimodal LLMs (

researcharxiv-cs-cv
9 Jun 2026
Applications

UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models

DGX agent

arXiv:2602.18020v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models leverage pretrained Vision-Language Models (VLMs) as backbones to map images and instructions to actions, demons

applicationsarxiv-cs-cv
9 Jun 2026
Hardware

Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions

DGX agent

arXiv:2606.09150v1 Announce Type: new Abstract: While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., 480P), leaving e

hardwarearxiv-cs-cv
9 Jun 2026
Safety

Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment

DGX agent

arXiv:2603.01613v2 Announce Type: replace Abstract: Monocular re-localization enables robots to estimate camera poses from visual observations. However, many existing methods rely on dense maps or lar

safetyarxiv-cs-cv
9 Jun 2026
Research

UniADC: A Unified Framework for Anomaly Detection and Classification

DGX agent

arXiv:2511.06644v3 Announce Type: replace Abstract: In this paper, we introduce a novel task termed unified anomaly detection and classification, which aims to simultaneously detect anomalous regions

researcharxiv-cs-cv
9 Jun 2026
Research

vesselFM-CT: Segmenting All Blood Vessels in CT Images for System-Level Cardiovascular Analysis

DGX agent

arXiv:2606.09400v1 Announce Type: new Abstract: The vascular network in the human body is characterized by blood vessels exhibiting drastic structural variations in radius, length, topological propert

researcharxiv-cs-cv
9 Jun 2026
Model Releases

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

DGX agent

arXiv:2606.08091v1 Announce Type: new Abstract: Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video genera

model-releasesarxiv-cs-cv
9 Jun 2026
Agents

Virtual-point-based Solutions to Handle Generalized Absolute Pose Problem

DGX agent

arXiv:2606.09294v1 Announce Type: new Abstract: Multi-camera systems are increasingly adopted in robotics and autonomous navigation for their wide field of view, flexibility, and fault tolerance. Neve

agentsarxiv-cs-cv
9 Jun 2026
Safety

Vision-Language Asymmetry in Bistable Image Captioning

DGX agent

arXiv:2606.08031v1 Announce Type: new Abstract: Wittgenstein's duck-rabbit poses a question for vision-language models: when a model captions an ambiguous image, where in the model is the commitment t

safetyarxiv-cs-cv
9 Jun 2026
Tutorials

Vision-Language Guided Hyperspectral Object Tracking via Semantics Fusion and Contextual Template Updating

DGX agent

arXiv:2606.09167v1 Announce Type: new Abstract: Hyperspectral object tracking (HOT) leverages the rich spectral information provided by hyperspectral videos (HSVs), offering substantial potential for

tutorialsarxiv-cs-cv
9 Jun 2026
Safety

Vision-Language Work Zone Intelligence for Safety-Critical Speed Regulation of Mixed-Autonomy Vehicles in Dynamic Environments

DGX agent

arXiv:2606.08860v1 Announce Type: new Abstract: Temporary work-zone speed limits are communicated through visually inconsistent signage and are often missing from digital maps, creating safety risks f

safetyarxiv-cs-cv
9 Jun 2026
Safety

Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning

DGX agent

arXiv:2606.09290v1 Announce Type: new Abstract: Visual reasoning requires integrating evidence distributed across regions, attributes, and relations, making single-chain reasoning prone to early perce

safetyarxiv-cs-cv
9 Jun 2026
Model Releases

Visual Template Inference for Data Extraction from Documents

DGX agent

arXiv:2501.06659v2 Announce Type: replace-cross Abstract: Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, t

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?

DGX agent

arXiv:2606.07872v1 Announce Type: new Abstract: When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual e

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis

DGX agent

arXiv:2606.08670v1 Announce Type: new Abstract: Large and demographically balanced datasets are essential for reliable neuroimaging biomarkers. Full-resolution 3D brain MRI synthesis can support data

safetyarxiv-cs-cv
9 Jun 2026
Research

What neurosurgeons need to see: synthetic intra-operative MRI from ultrasound for brain-shift compensation in brain tumour surgery

DGX agent

arXiv:2606.07658v1 Announce Type: new Abstract: Maximal safe resection is the primary objective in glioma surgery. Neuronavigation guidance is progressively degraded by brain shift after dural opening

researcharxiv-cs-cv
9 Jun 2026
Local Ai

When Vision Misleads, Let Location Speak: A Worldwide Image Geo-Localization Method via Location Attention Mechanism and Large Multimodal Models

DGX agent

arXiv:2606.08918v1 Announce Type: new Abstract: Worldwide image geo-localization aims to determine the capture location of an image on a global scale. Existing methods often mislocalize images by matc

local-aiarxiv-cs-cv
9 Jun 2026
Model Releases

Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving

DGX agent

arXiv:2606.09644v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer accuracy alone does not indicate whether a

model-releasesarxiv-cs-cv
9 Jun 2026
Research

Where the Score Lives: A Wavelet View of Diffusion

DGX agent

arXiv:2606.08309v1 Announce Type: cross Abstract: Score-based generative models have had remarkable success over the last decade in generating a diverse set of visually plausible images. A variety of

researcharxiv-cs-cv
9 Jun 2026
Applications

Wispy to Voluminous: Prior-free Multi-view Capture of Strand-level Facial Hair

DGX agent

arXiv:2606.08041v1 Announce Type: cross Abstract: Facial hair is a defining trait of personal identity, yet remains a critical bottleneck for digital avatars. Recent volumetric methods achieve photore

applicationsarxiv-cs-cv
9 Jun 2026
Applications

X-Palm: Paired Multispectral-to-Smartphone Dataset for Cross-Domain Palmprint Authentication

DGX agent

arXiv:2606.08437v1 Announce Type: cross Abstract: Palmprint modality offers a privacy-preserving biometric solution, yet its deployment is hindered by the domain gap between controlled enrollment and

applicationsarxiv-cs-cv
9 Jun 2026
Model Releases

Zero-Parameter Geometric Gating for Temporally Stable Low-Altitude UAV Video Semantic Segmentation

DGX agent

arXiv:2606.09162v1 Announce Type: new Abstract: Video semantic segmentation for low-altitude UAVs requires temporal consistency, yet dense optical flow introduces spatially structured noise in the pla

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Zero-Shot Semantic Re-Identification for Autonomous Driving: A VLM Baseline Study

DGX agent

arXiv:2606.09362v1 Announce Type: new Abstract: Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cy

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

3DMorph: Single-Image-Guided Local 3D Shape Editing and Morphing

DGX agent

arXiv:2606.07115v1 Announce Type: new Abstract: Despite recent progress in 3D generation, intuitive editing of existing shapes remains limited. Unlike images, which benefit from well-established inpai

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

A Cross-view Fusion Framework for Robust 6-DoF Grasp Pose Estimation

DGX agent

arXiv:2606.06878v1 Announce Type: cross Abstract: In this paper, we propose a cross-view fusion framework that enhances the robustness of 6-DoF grasp pose estimation in corner views. Our framework all

model-releasesarxiv-cs-cv
8 Jun 2026
← Previous
1…113114115116117…263
Next →