AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

Towards Robust Text-to-Image Person Retrieval: Multi-View Reformulation for Semantic Compensation

DGX agent

arXiv:2604.18376v1 Announce Type: new Abstract: In text-to-image person retrieval tasks, the diversity of natural language expressions and the implicitness of visual semantics often lead to the proble

safetyarxiv-cs-cv
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Towards Symmetry-sensitive Pose Estimation: A Rotation Representation for Symmetric Object Classes

DGX agent

arXiv:2604.18208v1 Announce Type: new Abstract: Symmetric objects are common in daily life and industry, yet their inherent orientation ambiguities that impede the training of deep learning networks f

researcharxiv-cs-cv
21 Apr 2026
Safety

Towards Universal Skeleton-Based Action Recognition

DGX agent

arXiv:2604.17013v1 Announce Type: new Abstract: With the development of robotics, skeleton-based action recognition has become increasingly important, as human-robot interaction requires understanding

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

TowerDataset: A Heterogeneous Benchmark for Transmission Corridor Segmentation with a Global-Local Fusion Framework

DGX agent

arXiv:2604.16848v1 Announce Type: new Abstract: Fine-grained semantic segmentation of transmission-corridor point clouds is fundamental for intelligent power-line inspection. However, current progress

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Training-inference input alignment outweighs framework choice in longitudinal retinal image prediction

DGX agent

arXiv:2604.16955v1 Announce Type: new Abstract: Quantitative prediction of future retinal appearance from longitudinal imaging would support clinical decisions in progressive macular disease that curr

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

Tri-Modal Fusion Transformers for UAV-based Object Detection

DGX agent

arXiv:2604.16630v1 Announce Type: new Abstract: Reliable UAV object detection requires robustness to illumination changes, motion blur, and scene dynamics that suppress RGB cues. Thermal long-wave inf

model-releasesarxiv-cs-cv
21 Apr 2026
Local Ai

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization

DGX agent

arXiv:2603.08096v3 Announce Type: replace Abstract: Localizing objects and parts from natural language in 3D space is essential for robotics, AR, and embodied AI, yet existing methods face a trade-off

local-aiarxiv-cs-cv
21 Apr 2026
Model Releases

TriTS: Time Series Forecasting from a Multimodal Perspective

DGX agent

arXiv:2604.16748v1 Announce Type: new Abstract: Time series forecasting plays a pivotal role in critical sectors such as finance, energy, transportation, and meteorology. However, Long-term Time Serie

model-releasesarxiv-cs-cv
21 Apr 2026
Local Ai

Trustworthy Endoscopic Super-Resolution

DGX agent

arXiv:2604.18001v1 Announce Type: new Abstract: Super-resolution (SR) models are attracting growing interest for enhancing minimally invasive surgery and diagnostic videos under hardware constraints.

local-aiarxiv-cs-cv
21 Apr 2026
Research

TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents

DGX agent

arXiv:2603.19684v2 Announce Type: replace Abstract: Automatic tooth segmentation and identification from intra-oral scanned 3D models are fundamental problems in digital dentistry, yet most existing a

researcharxiv-cs-cv
21 Apr 2026
Model Releases

TSM-Pose: Topology-Aware Learning with Semantic Mamba for Category-Level Object Pose Estimation

DGX agent

arXiv:2604.16954v1 Announce Type: new Abstract: Category-level object pose estimation is fundamental for embodied intelligence, yet achieving robust generalization to unseen instances remains challeng

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

DGX agent

arXiv:2604.18518v1 Announce Type: new Abstract: Uniform Discrete Diffusion Model (UDM) has recently emerged as a promising paradigm for discrete generative modeling; however, its integration with rein

model-releasesarxiv-cs-cv
21 Apr 2026
Tutorials

UGD: An Unsupervised Geometric Distance for Evaluating Real-world Noisy Point Cloud Denoising

DGX agent

arXiv:2604.16976v1 Announce Type: new Abstract: Point cloud denoising is a fundamental and crucial challenge in real-world point cloud applications. Existing quantitative evaluation metrics for point

tutorialsarxiv-cs-cv
21 Apr 2026
Research

Understanding Counting Mechanisms in Large Language and Vision-Language Models

DGX agent

arXiv:2511.17699v2 Announce Type: replace Abstract: Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark

DGX agent

arXiv:2510.13759v3 Announce Type: replace Abstract: Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. E

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement

DGX agent

arXiv:2604.17850v1 Announce Type: new Abstract: Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, le

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion

DGX agent

arXiv:2512.20249v3 Announce Type: replace-cross Abstract: Multimodal brain decoding aims to reconstruct semantic information that is consistent with visual stimuli from brain activity signals such as

model-releasesarxiv-cs-cv
21 Apr 2026
Agents

Unified Ultrasound Intelligence Toward an End-to-End Agentic System

DGX agent

arXiv:2604.16914v1 Announce Type: new Abstract: Clinical ultrasound analysis demands models that generalize across heterogeneous organs, views, and devices, while supporting interpretable workflow-lev

agentsarxiv-cs-cv
21 Apr 2026
Research

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

DGX agent

arXiv:2604.17565v1 Announce Type: new Abstract: Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geomet

researcharxiv-cs-cv
21 Apr 2026
Research

UniMesh: Unifying 3D Mesh Understanding and Generation

DGX agent

arXiv:2604.17472v1 Announce Type: new Abstract: Recent advances in 3D vision have led to specialized models for either 3D understanding (e.g., shape classification, segmentation, reconstruction) or 3D

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Unveiling Deepfakes: A Frequency-Aware Triple Branch Network for Deepfake Detection

DGX agent

arXiv:2604.17477v1 Announce Type: new Abstract: Advanced deepfake technologies are blurring the lines between real and fake, presenting both revolutionary opportunities and alarming threats. While it

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning

DGX agent

arXiv:2402.13243v2 Announce Type: replace Abstract: Learning a human-like driving policy from large-scale driving demonstrations is promising, but the uncertainty and non-deterministic nature of plann

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Video Panels for Long Video Understanding

DGX agent

arXiv:2509.23724v2 Announce Type: replace Abstract: Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

VIDEOP2R: Video Understanding from Perception to Reasoning

DGX agent

arXiv:2511.11113v2 Announce Type: replace Abstract: Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promisin

safetyarxiv-cs-cv
21 Apr 2026
Local Ai

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

DGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

local-aiarxiv-cs-cv
21 Apr 2026
Model Releases

VIDS: A Verified Imaging Dataset Standard for Medical AI

DGX agent

arXiv:2604.17525v1 Announce Type: cross Abstract: Medical imaging AI development is fundamentally dependent on annotated datasets, yet no existing standard provides machine-enforceable validation acro

model-releasesarxiv-cs-cv
21 Apr 2026
Research

View-Consistent 3D Scene Editing via Dual-Path Structural Correspondense and Semantic Continuity

DGX agent

arXiv:2604.17801v1 Announce Type: new Abstract: Text-driven 3D scene editing has recently attracted increasing attention. Most existing methods follow a render-edit-optimize pipeline, where multi-view

researcharxiv-cs-cv
21 Apr 2026
Research

ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes

DGX agent

arXiv:2604.17623v1 Announce Type: new Abstract: Kinematic rigs provide a structured interface for articulating 3D meshes, but they lack an inherent representation of the plausible manifold of joint co

researcharxiv-cs-cv
21 Apr 2026
Research

Vision Language Models are Biased

DGX agent

arXiv:2505.23941v4 Announce Type: replace-cross Abstract: Large language models (LLMs) memorize a vast amount of prior knowledge from the Internet that helps them on downstream tasks but also may noto

researcharxiv-cs-cv
21 Apr 2026
Applications

Visual-RRT: Finding Paths toward Visual-Goals via Differentiable Rendering

DGX agent

arXiv:2604.16388v1 Announce Type: cross Abstract: Rapidly-exploring random trees (RRTs) have been widely adopted for robot motion planning due to their robustness and theoretical guarantees. However,

applicationsarxiv-cs-cv
21 Apr 2026
Research

ViT^3: Unlocking Test-Time Training in Vision

DGX agent

arXiv:2512.01643v2 Announce Type: replace Abstract: Test-Time Training (TTT) has recently emerged as a promising direction for efficient sequence modeling. TTT reformulates attention operation as an o

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Voronoi-guided Bilateral 2D Gaussian Splatting for Arbitrary-Scale Hyperspectral Image Super-Resolution

DGX agent

arXiv:2604.17727v1 Announce Type: new Abstract: Most existing hyperspectral image super-resolution methods require modifications for different scales, limiting their flexibility in arbitrary-scale rec

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Weakly-Supervised Referring Video Object Segmentation through Text Supervision

DGX agent

arXiv:2604.17797v1 Announce Type: new Abstract: Referring video object segmentation (RVOS) aims to segment the target instance in a video, referred by a text expression. Conventional approaches are mo

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

What's Left Unsaid? Detecting and Correcting Misleading Omissions in Multimodal News Previews

DGX agent

arXiv:2601.05563v2 Announce Type: replace Abstract: Even when factually correct, social-media news previews (image-headline pairs) can induce interpretation drift: by selectively omitting crucial cont

model-releasesarxiv-cs-cv
21 Apr 2026
Research

When Background Matters: Breaking Medical Vision Language Models by Transferable Attack

DGX agent

arXiv:2604.17318v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly used in clinical diagnostics, yet their robustness to adversarial attacks remains largely unexplored, pos

researcharxiv-cs-cv
21 Apr 2026
Model Releases

When Earth Foundation Models Meet Diffusion: An Application to Land Surface Temperature Super-Resolution

DGX agent

arXiv:2604.16841v1 Announce Type: new Abstract: Land surface temperature (LST) super-resolution is important for environmental monitoring. However, it remains challenging as coarse thermal observation

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators

DGX agent

arXiv:2602.19946v4 Announce Type: replace Abstract: Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following. But do they perform well as

model-releasesarxiv-cs-cv
21 Apr 2026
Research

When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models

DGX agent

arXiv:2507.13868v2 Announce Type: replace Abstract: Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their interna

researcharxiv-cs-cv
21 Apr 2026
Model Releases

When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models

DGX agent

arXiv:2604.17375v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have substantially enhanced their ability across multimodal video understanding benchmarks spanning tem

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations

DGX agent

arXiv:2603.22368v2 Announce Type: replace Abstract: Visualizations help communicate data insights, but deceptive data representations can distort their interpretation and propagate misinformation. Whi

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

When W4A4 Breaks Camouflaged Object Detection: Token-Group Dual-Constraint Activation Quantization

DGX agent

arXiv:2604.16855v1 Announce Type: new Abstract: Camouflaged object detection (COD) segments objects that intentionally blend with the background, so predictions depend on subtle texture and boundary c

model-releasesarxiv-cs-cv
21 Apr 2026
Local Ai

Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding

DGX agent

arXiv:2604.17422v1 Announce Type: new Abstract: Long video understanding remains a formidable challenge for Multimodal Large Language Models (MLLMs) due to the prohibitive computational cost of proces

local-aiarxiv-cs-cv
21 Apr 2026
Research

Why Training-Free Token Reduction Collapses: The Inherent Instability of Pairwise Scoring Signals

DGX agent

arXiv:2604.16745v1 Announce Type: cross Abstract: Training-free token reduction methods for Vision Transformers (ToMe, ToFu, PiToMe, and MCTF) employ different scoring mechanisms, yet they share a clo

researcharxiv-cs-cv
21 Apr 2026
Agents

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

DGX agent

arXiv:2604.18484v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex

agentsarxiv-cs-cv
21 Apr 2026
Research

ZSG-IAD: A Multimodal Framework for Zero-Shot Grounded Industrial Anomaly Detection

DGX agent

arXiv:2604.17949v1 Announce Type: new Abstract: Deep learning-based industrial anomaly detectors often behave as black boxes, making it hard to justify decisions with physically meaningful defect evid

researcharxiv-cs-cv
21 Apr 2026
Local Ai

A Single Image and Multimodality Is All You Need for Novel View Synthesis

DGX agent

arXiv:2602.17909v2 Announce Type: replace Abstract: Diffusion-based approaches have recently demonstrated strong performance for single-image novel view synthesis by conditioning generative models on

local-aiarxiv-cs-cv
20 Apr 2026
Local Ai

Adapting in the Dark: Efficient and Stable Test-Time Adaptation for Black-Box Models

DGX agent

arXiv:2604.15609v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) for black-box models accessible only via APIs remains a largely unexplored challenge. Existing approaches such as post-hoc

local-aiarxiv-cs-cv
20 Apr 2026
Local Ai

AdaVFM: Adaptive Vision Foundation Models for Edge Intelligence via LLM-Guided Execution

DGX agent

arXiv:2604.15622v1 Announce Type: new Abstract: Language-aligned vision foundation models (VFMs) enable versatile visual understanding for always-on contextual AI, but their deployment on edge devices

local-aiarxiv-cs-cv
20 Apr 2026
← Previous
1…232233234235236…261
Next →