AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

Location Is All You Need: Continuous Spatiotemporal Neural Representations of Earth Observation Data

DGX agent

arXiv:2604.07092v2 Announce Type: replace Abstract: In this work, we present LIANet (Location Is All You Need Network), a coordinate-based neural representation that models multi-temporal spaceborne E

researcharxiv-cs-cv
10 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification

DGX agent

arXiv:2604.08333v1 Announce Type: new Abstract: The rise of multimodal large language models (MLLMs) has sparked an unprecedented wave of applications in the field of medical imaging analysis. However

researcharxiv-cs-cv
10 Apr 2026
Model Releases

LPM 1.0: Video-based Character Performance Model

DGX agent

arXiv:2604.07823v1 Announce Type: new Abstract: Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Lear

model-releasesarxiv-cs-cv
10 Apr 2026
Research

LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models

DGX agent

arXiv:2512.17489v2 Announce Type: replace Abstract: Text-to-image (T2I) models have demonstrated remarkable progress in creative image generation, yet they still lack precise control over scene illumi

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation

DGX agent

arXiv:2604.06950v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are increasingly being deployed as automated content moderators. Within this landscape, we uncover a critic

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Mathematical Analysis of Image Matching Techniques

DGX agent

arXiv:2604.07574v1 Announce Type: new Abstract: Image matching is a fundamental problem in Computer Vision with direct applications in robotics, remote sensing, and geospatial data analysis. We presen

researcharxiv-cs-cv
10 Apr 2026
Safety

MCLR: Improving Conditional Modeling via Inter-Class Likelihood-Ratio Maximization and Unifying Classifier-Free Guidance with Alignment Objectives

DGX agent

arXiv:2603.22364v2 Announce Type: replace-cross Abstract: Diffusion models have achieved state-of-the-art performance in generative modeling, but their success often relies heavily on classifier-free

safetyarxiv-cs-cv
10 Apr 2026
Safety

MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning

DGX agent

arXiv:2604.08203v1 Announce Type: new Abstract: Medical Vision-Language Models (VLMs) hold immense promise for complex clinical tasks, but their reasoning capabilities are often constrained by text-on

safetyarxiv-cs-cv
10 Apr 2026
Research

MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping

DGX agent

arXiv:2604.08364v1 Announce Type: new Abstract: In this paper, we introduce MegaStyle, a novel and scalable data curation pipeline that constructs an intra-style consistent, inter-style diverse and hi

researcharxiv-cs-cv
10 Apr 2026
Local Ai

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models

DGX agent

arXiv:2601.04068v3 Announce Type: replace Abstract: Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization

local-aiarxiv-cs-cv
10 Apr 2026
Applications

Mitigating Domain Drift in Multi Species Segmentation with DINOv2: A Cross-Domain Evaluation in Herbicide Research Trials

DGX agent

arXiv:2508.07514v3 Announce Type: replace Abstract: Reliable plant species and damage segmentation for herbicide field research trials requires models that can withstand substantial real-world variati

applicationsarxiv-cs-cv
10 Apr 2026
Research

Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction

DGX agent

arXiv:2604.07914v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable success across cross-modal tasks but remain hindered by hallucinations, producing textual

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Mitigating Spurious Background Bias in Multimedia Recognition with Disentangled Concept Bottlenecks

DGX agent

arXiv:2510.15770v3 Announce Type: replace Abstract: Concept Bottleneck Models (CBMs) enhance interpretability by predicting human-understandable concepts as intermediate representations. However, exis

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework

DGX agent

arXiv:2509.23322v2 Announce Type: replace Abstract: With the continuous expansion of Large Language Models (LLMs) and advances in reinforcement learning, LLMs have demonstrated exceptional reasoning c

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models

DGX agent

arXiv:2412.20718v2 Announce Type: replace Abstract: The rapid integration of Large Vision-Language Models (LVLMs) into critical domains necessitates comprehensive moral evaluation to ensure their alig

model-releasesarxiv-cs-cv
10 Apr 2026
Agents

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

DGX agent

arXiv:2604.08516v1 Announce Type: new Abstract: Web agents--autonomous systems that navigate and execute tasks on the web on behalf of users--have the potential to transform how people interact with t

agentsarxiv-cs-cv
10 Apr 2026
Model Releases

Monocular Depth Estimation From the Perspective of Feature Restoration: A Diffusion Enhanced Depth Restoration Approach

DGX agent

arXiv:2604.07664v1 Announce Type: new Abstract: Monocular Depth Estimation (MDE) is a fundamental computer vision task with important applications in 3D vision. The current mainstream MDE methods empl

model-releasesarxiv-cs-cv
10 Apr 2026
Safety

MonoUNet: A Robust Tiny Neural Network for Automated Knee Cartilage Segmentation on Point-of-Care Ultrasound Devices

DGX agent

arXiv:2604.07780v1 Announce Type: cross Abstract: Objective: To develop a robust and compact deep learning model for automated knee cartilage segmentation on point-of-care ultrasound (POCUS) devices.

safetyarxiv-cs-cv
10 Apr 2026
Safety

MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models

DGX agent

arXiv:2604.07991v1 Announce Type: new Abstract: Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation f

safetyarxiv-cs-cv
10 Apr 2026
Safety

MSCT: Differential Cross-Modal Attention for Deepfake Detection

DGX agent

arXiv:2604.07741v1 Announce Type: new Abstract: Audio-visual deepfake detection typically employs a complementary multi-modal model to check the forgery traces in the video. These methods primarily ex

safetyarxiv-cs-cv
10 Apr 2026
Local Ai

MSGL-Transformer: A Multi-Scale Global-Local Transformer for Rodent Social Behavior Recognition

DGX agent

arXiv:2604.07578v1 Announce Type: new Abstract: Recognition of rodent behavior is important for understanding neural and behavioral mechanisms. Traditional manual scoring is time-consuming and prone t

local-aiarxiv-cs-cv
10 Apr 2026
Applications

MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation

DGX agent

arXiv:2603.11633v2 Announce Type: replace Abstract: Recent unified 3D generation models have made remarkable progress in producing high-quality 3D assets from a single image. Notably, layout-aware app

applicationsarxiv-cs-cv
10 Apr 2026
Research

MVOS_HSI: A Python Library for Preprocessing Agricultural Crop Hyperspectral Data

DGX agent

arXiv:2604.07656v1 Announce Type: cross Abstract: Hyperspectral imaging (HSI) allows researchers to study plant traits non-destructively. By capturing hundreds of narrow spectral bands per pixel, it r

researcharxiv-cs-cv
10 Apr 2026
Research

Nearest Neighbor Projection Removal Adversarial Training

DGX agent

arXiv:2509.07673v4 Announce Type: replace Abstract: Deep neural networks have exhibited impressive performance in image classification tasks but remain vulnerable to adversarial examples. Standard adv

researcharxiv-cs-cv
10 Apr 2026
Tutorials

Needle in a Haystack -- One-Class Representation Learning for Detecting Rare Malignant Cells in Computational Cytology

DGX agent

arXiv:2604.07722v1 Announce Type: new Abstract: In computational cytology, detecting malignancy on whole-slide images is difficult because malignant cells are morphologically diverse yet vanishingly r

tutorialsarxiv-cs-cv
10 Apr 2026
Research

Novel View Synthesis as Video Completion

DGX agent

arXiv:2604.08500v1 Announce Type: new Abstract: We tackle the problem of sparse novel view synthesis (NVS) using video diffusion models; given K (approx 5) multi-view images of a scene and their

researcharxiv-cs-cv
10 Apr 2026
Model Releases

NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results

DGX agent

arXiv:2604.06945v2 Announce Type: replace Abstract: This paper reports on the NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration (BSCVR). The challenge aims to advance research on recoverin

model-releasesarxiv-cs-cv
10 Apr 2026
Hardware

Object-Centric Stereo Ranging for Autonomous Driving: From Dense Disparity to Census-Based Template Matching

DGX agent

arXiv:2604.07980v1 Announce Type: new Abstract: Accurate depth estimation is critical for autonomous driving perception systems, particularly for long range vehicle detection on highways. Traditional

hardwarearxiv-cs-cv
10 Apr 2026
Tutorials

OceanMAE: A Foundation Model for Ocean Remote Sensing

DGX agent

arXiv:2604.08171v1 Announce Type: new Abstract: Accurate ocean mapping is essential for applications such as bathymetry estimation, seabed characterization, marine litter detection, and ecosystem moni

tutorialsarxiv-cs-cv
10 Apr 2026
Research

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering

DGX agent

arXiv:2604.08209v1 Announce Type: new Abstract: To extend the reinforcement learning post-training paradigm to omni-modal models for concurrently bolstering video-audio understanding and collaborative

researcharxiv-cs-cv
10 Apr 2026
Safety

On the Global Photometric Alignment for Low-Level Vision

DGX agent

arXiv:2604.08172v1 Announce Type: new Abstract: Supervised low-level vision models rely on pixel-wise losses against paired references, yet paired training sets exhibit per-pair photometric inconsiste

safetyarxiv-cs-cv
10 Apr 2026
Research

On the Uphill Battle of Image frequency Analysis

DGX agent

arXiv:2604.07563v1 Announce Type: new Abstract: This work is a follow up on the newly proposed clustering algorithm called The Inverse Square Mean Shift Algorithm. In this paper a special case of algo

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

DGX agent

arXiv:2604.08031v1 Announce Type: cross Abstract: Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an int

model-releasesarxiv-cs-cv
10 Apr 2026
Research

OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation

DGX agent

arXiv:2512.03532v2 Announce Type: replace Abstract: Generalizing open-vocabulary 3D instance segmentation (OV-3DIS) to diverse, unstructured, and mesh-free environments is crucial for robotics and AR/

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models

DGX agent

arXiv:2604.08266v1 Announce Type: new Abstract: Leveraging the general world knowledge of Large Language Models (LLMs) holds significant promise for improving the ability of autonomous driving systems

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

oslash Source Models Leak What They Shouldn't nrightarrow: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimization

DGX agent

arXiv:2604.08238v1 Announce Type: new Abstract: The increasing adaptation of vision models across domains, such as satellite imagery and medical scans, has raised an emerging privacy risk: models may

model-releasesarxiv-cs-cv
10 Apr 2026
Research

OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation

DGX agent

arXiv:2604.08110v1 Announce Type: new Abstract: Training-free open-vocabulary semantic segmentation(TF-OVSS) has recently attracted attention for its ability to perform dense prediction by leveraging

researcharxiv-cs-cv
10 Apr 2026
Safety

OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance

DGX agent

arXiv:2604.08461v1 Announce Type: new Abstract: Open-Vocabulary Segmentation (OVS) aims to segment image regions beyond predefined category sets by leveraging semantic descriptions. While CLIP based a

safetyarxiv-cs-cv
10 Apr 2026
Safety

OxEnsemble: Fair Ensembles for Low-Data Classification

DGX agent

arXiv:2512.09665v2 Announce Type: replace Abstract: We address the problem of fair classification in settings where data is scarce and unbalanced across demographic groups. Such low-data regimes are c

safetyarxiv-cs-cv
10 Apr 2026
Research

PANC: Prior-Aware Normalized Cut via Anchor-Augmented Token Graphs

DGX agent

arXiv:2602.06912v2 Announce Type: replace Abstract: Unsupervised segmentation from self-supervised ViT patches holds promise but lacks robustness: multi-object scenes confound saliency cues, and low-s

researcharxiv-cs-cv
10 Apr 2026
Research

PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation

DGX agent

arXiv:2604.07901v1 Announce Type: new Abstract: 360 video object segmentation (360VOS) aims to predict temporally-consistent masks in 360 videos, offering full-scene coverage, benefiting applications,

researcharxiv-cs-cv
10 Apr 2026
Agents

ParkSense: Where Should a Delivery Driver Park? Leveraging Idle AV Compute and Vision-Language Models

DGX agent

arXiv:2604.07912v1 Announce Type: new Abstract: Finding parking consumes a disproportionate share of food delivery time, yet no system addresses precise parking-spot selection relative to merchant ent

agentsarxiv-cs-cv
10 Apr 2026
Model Releases

ParseBench: A Document Parsing Benchmark for AI Agents

DGX agent

arXiv:2604.08538v1 Announce Type: new Abstract: AI agents are changing the requirements for document parsing. What matters is semantic correctness: parsed output must preserve the structure and

model-releasesarxiv-cs-cv
10 Apr 2026
Safety

Part^{2}GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting

DGX agent

arXiv:2506.17212v2 Announce Type: replace Abstract: Articulated objects are common in the real world, yet modeling their structure and motion remains a challenging task for 3D reconstruction methods.

safetyarxiv-cs-cv
10 Apr 2026
Safety

Personalizing Text-to-Image Generation to Individual Taste

DGX agent

arXiv:2604.07427v1 Announce Type: new Abstract: Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models opt

safetyarxiv-cs-cv
10 Apr 2026
Research

Phantasia: Context-Adaptive Backdoors in Vision Language Models

DGX agent

arXiv:2604.08395v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have greatly enhanced the integration of visual perception and linguistic reasoning, driving rapid prog

researcharxiv-cs-cv
10 Apr 2026
Applications

Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

DGX agent

arXiv:2604.08503v1 Announce Type: new Abstract: Recent advances in generative video modeling, driven by large-scale datasets and powerful architectures, have yielded remarkable visual realism. However

applicationsarxiv-cs-cv
10 Apr 2026
Model Releases

PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing

DGX agent

arXiv:2604.07230v2 Announce Type: replace Abstract: Achieving physically accurate object manipulation in image editing is essential for its potential applications in interactive world models. However,

model-releasesarxiv-cs-cv
10 Apr 2026
← Previous
1…255256257258259
Next →