AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

InstructMixup: Instruction-Guided Salient Patch Editing for Robust Data Augmentation

DGX agent

arXiv:2607.19324v1 Announce Type: new Abstract: In image and video technologies, data augmentation is widely used to improve the generalization of deep visual models, and mixup-based strategies that i

researcharxiv-cs-cv
23 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Interactive Medical-SAM2 GUI: A Napari-based semi-automatic annotation tool for medical images

DGX agent

arXiv:2602.22649v2 Announce Type: replace Abstract: Interactive Medical-SAM2 GUI is an open-source desktop application for semi-automatic annotation of 2D and 3D medical images. Built on the Napari mu

model-releasesarxiv-cs-cv
23 Jul 2026
Research

Internet-of-Things Architectures for Secure Cyber-Physical Spaces: the VISOR Experience Report

DGX agent

arXiv:2204.01531v2 Announce Type: replace-cross Abstract: Internet of things (IoT) technologies are becoming a more and more widespread part of civilian life in common urban spaces, which are rapidly

researcharxiv-cs-cv
23 Jul 2026
Model Releases

It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability

DGX agent

arXiv:2607.16292v2 Announce Type: replace Abstract: Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge. We ask whether th

model-releasesarxiv-cs-cv
23 Jul 2026
Research

Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

DGX agent

arXiv:2510.08532v2 Announce Type: replace Abstract: Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language. Yet, relying solely on text instr

researcharxiv-cs-cv
23 Jul 2026
Research

Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models

DGX agent

arXiv:2607.19120v1 Announce Type: new Abstract: Geometric foundation models, such as the Visual Geometry Grounded Transformer (VGGT), provide strong 3D priors from unposed images. However, such models

researcharxiv-cs-cv
23 Jul 2026
Local Ai

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

DGX agent

arXiv:2607.19889v1 Announce Type: new Abstract: Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language model

local-aiarxiv-cs-cv
23 Jul 2026
Research

LC-SLab -- An object-based deep learning framework for large-scale land cover classification from satellite imagery and sparse in-situ labels

DGX agent

arXiv:2509.15868v2 Announce Type: replace Abstract: Large-scale land cover maps generated using deep learning play a critical role across a wide range of Earth science applications. Open in-situ datas

researcharxiv-cs-cv
23 Jul 2026
Safety

Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2

DGX agent

arXiv:2607.19811v1 Announce Type: new Abstract: The Segment Anything Model 2 (SAM2) has advanced temporal promptable segmentation, yet its deployment remains hindered by heavy memory cross-attention o

safetyarxiv-cs-cv
23 Jul 2026
Safety

Learning Encoding-Decoding Direction Pairs to Unveil Concepts of Influence in Deep Vision Networks

DGX agent

arXiv:2509.23926v4 Announce Type: replace Abstract: Empirical evidence shows that deep vision networks often represent concepts as directions in latent space with concept information written along dir

safetyarxiv-cs-cv
23 Jul 2026
Model Releases

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

DGX agent

arXiv:2607.18924v1 Announce Type: new Abstract: Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward

model-releasesarxiv-cs-cv
23 Jul 2026
Tutorials

Learning Semantic-Robust Change Detection via Semantic-Invariant Self-Distillation

DGX agent

arXiv:2607.19000v1 Announce Type: new Abstract: Change detection aims to identify semantic changes between remote sensing images. However, features from models are easily disturbed by non-semantic var

tutorialsarxiv-cs-cv
23 Jul 2026
Model Releases

Local Label-Informed Feature Transfer for Generating Ground-Truth Medical Images: A Comparison of GAN- and Diffusion-Based Approaches

DGX agent

arXiv:2607.18882v1 Announce Type: new Abstract: Validating Explainable Artificial Intelligence (XAI) methods in medical imaging requires ground-truth data with known locations of informative features.

model-releasesarxiv-cs-cv
23 Jul 2026
Local Ai

Look Before You Edit: Attention-Guided Camera Placement and Multi-View Alignment for 3D Gaussian Splatting Editing

DGX agent

arXiv:2607.19777v1 Announce Type: new Abstract: Text-driven 3D scene editing with 3D Gaussian Splatting (3DGS) typically applies a 2D diffusion editor to views rendered from fixed training cameras, li

local-aiarxiv-cs-cv
23 Jul 2026
Tutorials

Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs

DGX agent

arXiv:2607.20357v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost,

tutorialsarxiv-cs-cv
23 Jul 2026
Model Releases

LoRFT: Benchmarking Long-Range Vehicle Trajectory Reconstruction from Fixed Highway Cameras

DGX agent

arXiv:2607.19911v1 Announce Type: new Abstract: Long-range vehicle trajectories provide important spatio-temporal evidence for traffic safety analysis, autonomous driving evaluation, and data-driven t

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation

DGX agent

arXiv:2607.14595v2 Announce Type: replace Abstract: Large-scale video diffusion models deliver strong generation performance, but full fine-tuning for downstream tasks incurs prohibitive computational

model-releasesarxiv-cs-cv
23 Jul 2026
Safety

Masked Visual Actions for Unified World Modeling

DGX agent

arXiv:2607.19343v1 Announce Type: new Abstract: Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world

safetyarxiv-cs-cv
23 Jul 2026
Safety

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

DGX agent

arXiv:2607.15273v2 Announce Type: replace Abstract: MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient genera

safetyarxiv-cs-cv
23 Jul 2026
Model Releases

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

DGX agent

arXiv:2607.19235v1 Announce Type: cross Abstract: Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challen

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts

DGX agent

arXiv:2607.18673v2 Announce Type: new Abstract: Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for

model-releasesarxiv-cs-cv
23 Jul 2026
Safety

Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

DGX agent

arXiv:2607.19027v1 Announce Type: new Abstract: Zero-shot video moment retrieval aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and it

safetyarxiv-cs-cv
23 Jul 2026
Applications

MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts

DGX agent

arXiv:2607.19826v1 Announce Type: new Abstract: Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' parad

applicationsarxiv-cs-cv
23 Jul 2026
Local Ai

MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement

DGX agent

arXiv:2607.17967v2 Announce Type: replace Abstract: Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notabl

local-aiarxiv-cs-cv
23 Jul 2026
Applications

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

DGX agent

arXiv:2607.18789v1 Announce Type: new Abstract: Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture

applicationsarxiv-cs-cv
23 Jul 2026
Safety

MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

DGX agent

arXiv:2607.19886v1 Announce Type: new Abstract: Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity deg

safetyarxiv-cs-cv
23 Jul 2026
Research

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

DGX agent

arXiv:2607.20284v1 Announce Type: new Abstract: The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU

researcharxiv-cs-cv
23 Jul 2026
Model Releases

MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction

DGX agent

arXiv:2607.19910v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by generating code directly from visual designs

model-releasesarxiv-cs-cv
23 Jul 2026
Research

MVGD-Net: A Novel Motion-aware Video Glass Surface Detection Method

DGX agent

arXiv:2601.13715v2 Announce Type: replace Abstract: Glass surface ubiquitous in both daily life and professional environments presents a potential threat to vision-based systems, such as robot and dro

researcharxiv-cs-cv
23 Jul 2026
Hardware

NGPS: GPS-Denied Aerial Geo-Localization and 2.5D Reconstruction via Deep Satellite Image Matching and Multi-Rate Sensor Fusion

DGX agent

arXiv:2607.18936v1 Announce Type: cross Abstract: We present NGPS (Next-Generation Positioning System), a visual geo-localization framework for high-altitude UAVs that provides GPS-free absolute posit

hardwarearxiv-cs-cv
23 Jul 2026
Safety

No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

DGX agent

arXiv:2607.19288v1 Announce Type: new Abstract: Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing

safetyarxiv-cs-cv
23 Jul 2026
Safety

Norm or Direction? Decoding Vision Mambas for High-Resolution Vision

DGX agent

arXiv:2607.18625v1 Announce Type: new Abstract: Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visual backbones.

safetyarxiv-cs-cv
23 Jul 2026
Model Releases

Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training

DGX agent

arXiv:2607.20238v1 Announce Type: new Abstract: Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrast

model-releasesarxiv-cs-cv
23 Jul 2026
Research

Now We Know? A Systematic Comparison of TerraMind and THOR

DGX agent

arXiv:2607.18504v1 Announce Type: cross Abstract: Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much

researcharxiv-cs-cv
23 Jul 2026
Research

Occlusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level Attention

DGX agent

arXiv:2607.18112v2 Announce Type: replace Abstract: Panoptic segmentation in complex scenes remains challenging because of occlusions, yet modern approaches often neglect occlusion modelling. In this

researcharxiv-cs-cv
23 Jul 2026
Model Releases

OffNadirLoc: Benchmark and Framework for Challenging UAV-to-Satellite Geo-Localization under Large Off-Nadir Views

DGX agent

arXiv:2607.19951v1 Announce Type: new Abstract: Cross-view geo-localization between UAV and satellite imagery remains a fundamental yet highly challenging task, especially under large off-nadir views

model-releasesarxiv-cs-cv
23 Jul 2026
Agents

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

DGX agent

arXiv:2607.19339v1 Announce Type: new Abstract: Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve wit

agentsarxiv-cs-cv
23 Jul 2026
Local Ai

OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation

DGX agent

arXiv:2607.18850v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgme

local-aiarxiv-cs-cv
23 Jul 2026
Model Releases

Open-Vocabulary Gaze Object Prediction: Benchmark and Method

DGX agent

arXiv:2607.18827v1 Announce Type: new Abstract: Gaze Object Prediction (GOP) aims to localize and recognize the objects humans attend to, a task crucial for understanding human-centric interactions. H

model-releasesarxiv-cs-cv
23 Jul 2026
Research

Pain in 3D: Generating Controllable Synthetic Faces for Automated Pain Assessment

DGX agent

arXiv:2509.16727v5 Announce Type: replace Abstract: Automated pain assessment from facial expressions is crucial for non-communicative patient. Progress has been limited by two challenges: (i) existin

researcharxiv-cs-cv
23 Jul 2026
Model Releases

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

DGX agent

arXiv:2607.19261v1 Announce Type: new Abstract: Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scal

model-releasesarxiv-cs-cv
23 Jul 2026
Safety

Pathologist Attention-Aligned Report Generation for Prostate Histopathology

DGX agent

arXiv:2607.19624v1 Announce Type: new Abstract: The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that critically shapes the information extracte

safetyarxiv-cs-cv
23 Jul 2026
Model Releases

PathReportEval: A Systematic Benchmark for Pathology Report Generation

DGX agent

arXiv:2607.18448v1 Announce Type: cross Abstract: Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, yet progress is difficult to measure beca

model-releasesarxiv-cs-cv
23 Jul 2026
Safety

PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration

DGX agent

arXiv:2607.16602v2 Announce Type: replace Abstract: Action-conditioned world models are a key component of embodied AI, serving as scalable policy evaluators that reduce reliance on expensive real-wor

safetyarxiv-cs-cv
23 Jul 2026
Tutorials

PC-Seg: Progressive Cross-View Consistency for 3D OCT Segmentation from Sparse 2D Annotations

DGX agent

arXiv:2607.17718v2 Announce Type: replace Abstract: Volumetric segmentation of optical coherence tomography (OCT) images is essential for diagnosing ocular diseases but requires labor-intensive voxel-

tutorialsarxiv-cs-cv
23 Jul 2026
Research

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

DGX agent

arXiv:2607.20389v1 Announce Type: new Abstract: Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal p

researcharxiv-cs-cv
23 Jul 2026
Agents

PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving

DGX agent

arXiv:2607.20175v1 Announce Type: new Abstract: Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-releva

agentsarxiv-cs-cv
23 Jul 2026
Local Ai

Physics Closure Matters for Machine Olfaction: A Maxwell--Stefan Graph Solver for Identifiable Dynamic Gas Unmixing

DGX agent

arXiv:2607.18544v1 Announce Type: new Abstract: Machine olfaction for gas unmixing is an underconstrained inverse problem in which gas compositions must be inferred from low-dimensional, delayed, and

local-aiarxiv-cs-cv
23 Jul 2026
← Previous
1…4647484950…261
Next →