AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

EventPrune: Cascaded Event-Assisted Token Pruning for Efficient First-Person Dynamic Spatial Reasoning

DGX agent

arXiv:2605.19506v1 Announce Type: new Abstract: First-person dynamic spatial reasoning requires models to track continuous motion and precise geometric structure, but the quadratic attention cost of T

model-releasesarxiv-cs-cv
20 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures

DGX agent

arXiv:2605.19478v1 Announce Type: cross Abstract: Existing ViT backdoor attacks based on backbone-overwriting full-tuning are computationally expensive and inflict performance degradation. This has fo

model-releasesarxiv-cs-cv
20 May 2026
Research

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models

DGX agent

arXiv:2605.19859v1 Announce Type: new Abstract: Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs

researcharxiv-cs-cv
20 May 2026
Research

Fast 4D Mesh Generation by Spatio-Temporal Attention Chains

DGX agent

arXiv:2605.19786v1 Announce Type: new Abstract: 4D mesh generation has recently emerged as a powerful paradigm for recovering dynamic 3D structure from videos, but existing methods remain slow, comput

researcharxiv-cs-cv
20 May 2026
Model Releases

Fast-BEV++: Fast by Algorithm, Deployable by Design

DGX agent

arXiv:2512.08237v3 Announce Type: replace Abstract: The advancement of vision-only Bird's-Eye-View (BEV) perception, a core paradigm for cost-effective autonomous driving, is hindered by the long-stan

model-releasesarxiv-cs-cv
20 May 2026
Safety

Feature-Space Smoothing: Certified Robustness of Deep Representations

DGX agent

arXiv:2601.16200v3 Announce Type: replace-cross Abstract: Modern deep learning models exhibit strong capabilities across diverse applications, yet remain vulnerable to malicious inputs that induce err

safetyarxiv-cs-cv
20 May 2026
Applications

Feed-Forward Gaussian Splatting from Sparse Aerial Views

DGX agent

arXiv:2605.19949v1 Announce Type: new Abstract: Reconstructing large-scale urban scenes from sparse aerial views is a crucial yet challenging task. Due to biased top-down and shallow-oblique camera po

applicationsarxiv-cs-cv
20 May 2026
Research

FGSVQA: Frequency-Guided Short-form Video Quality Assessment

DGX agent

arXiv:2605.20016v1 Announce Type: cross Abstract: Short-form video poses new challenges to the quality assessment of user-generated content (UGC) due to its complex generation pipeline, rapid content

researcharxiv-cs-cv
20 May 2026
Safety

FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models

DGX agent

arXiv:2605.19739v1 Announce Type: new Abstract: Recent advances in flow matching models have significantly improved text-to-image generation quality, but also introduce growing safety risks due to the

safetyarxiv-cs-cv
20 May 2026
Research

FPED: A Functional-Network Prior-Guided Mixture-of-Experts Framework for Interpretable Brain Decoding

DGX agent

arXiv:2605.19279v1 Announce Type: new Abstract: Visual image reconstruction from functional Magnetic Resonance Imaging (fMRI) is a fundamental task in brain decoding, providing a crucial pathway for u

researcharxiv-cs-cv
20 May 2026
Research

From Division to Decision: Leveraging Temporal Cell-Stage Segmentation for Embryo Transferability Prediction

DGX agent

arXiv:2605.18923v1 Announce Type: cross Abstract: Accurate selection of bovine embryos is a challenging task, as current practice relies on a single expert assessment on the seventh day after insemina

researcharxiv-cs-cv
20 May 2026
Model Releases

From Llama to Cria: Scaling Down Neural Networks via Neuron-Level Spectral Structural Importance Evaluation

DGX agent

arXiv:2605.18860v1 Announce Type: cross Abstract: This paper proposes a neuron pruning framework based on neuron-level spectral structural importance evaluation. Given a trained neural network, we rec

model-releasesarxiv-cs-cv
20 May 2026
Research

GeoMamba: A Geometry-driven MambaVision Framework and Dataset for Fine-grained Optical-SAR Object Retrieval

DGX agent

arXiv:2605.19734v1 Announce Type: new Abstract: Multi-source remote sensing enables complementary observation of ground objects, while cross-modal fine-grained object retrieval remains challenging, es

researcharxiv-cs-cv
20 May 2026
Research

GLUT: 3D Gaussian Lookup Table for Continuous Color Transformation

DGX agent

arXiv:2605.19889v1 Announce Type: cross Abstract: 3D Lookup Tables (3D LUTs) are widely used for color mapping, but their grid-based representation requires discretizing the RGB space, leading to a ca

researcharxiv-cs-cv
20 May 2026
Model Releases

GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation

DGX agent

arXiv:2605.19890v1 Announce Type: new Abstract: Test-time adaptation (TTA) enables a pre-trained model to adapt online to an unlabeled test stream under distribution shift. While most TTA research foc

model-releasesarxiv-cs-cv
20 May 2026
Local Ai

GRLoc: Geometric Representation Regression for Visual Localization

DGX agent

arXiv:2511.13864v2 Announce Type: replace Abstract: Absolute Pose Regression (APR) has emerged as a compelling paradigm for visual localization. However, APR models typically operate as black boxes, d

local-aiarxiv-cs-cv
20 May 2026
Safety

Hard-Label Black-Box Attacks on 3D Point Clouds

DGX agent

arXiv:2412.00404v2 Announce Type: replace Abstract: With the maturity of depth sensors in various 3D safety-critical applications, 3D point cloud models have been shown to be vulnerable to adversarial

safetyarxiv-cs-cv
20 May 2026
Model Releases

HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

DGX agent

arXiv:2605.19223v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over

model-releasesarxiv-cs-cv
20 May 2026
Agents

HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models

DGX agent

arXiv:2605.19631v1 Announce Type: cross Abstract: End-to-end autonomous driving has emerged as a compelling alternative to traditional modular pipelines by directly mapping raw sensor data to driving

agentsarxiv-cs-cv
20 May 2026
Local Ai

Hierarchical Schedule Optimization for Fast and Robust Diffusion Model Sampling

DGX agent

arXiv:2511.11688v3 Announce Type: replace-cross Abstract: Diffusion probabilistic models have set a new standard for generative fidelity but are hindered by a slow iterative sampling process. A powerf

local-aiarxiv-cs-cv
20 May 2026
Safety

HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance

DGX agent

arXiv:2506.07209v2 Announce Type: replace-cross Abstract: We present HOI-PAGE, a new approach that prioritizes part-level affordance reasoning to generate high-fidelity 4D human-object interactions (H

safetyarxiv-cs-cv
20 May 2026
Local Ai

iDiff: Interpretable Difference-aware Framework for Pairwise Image Quality Assessment

DGX agent

arXiv:2605.19522v1 Announce Type: new Abstract: Pairwise image quality assessment (IQA) in professional photography requires a model not only to identify the preferred image between two candidates, bu

local-aiarxiv-cs-cv
20 May 2026
Model Releases

iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models

DGX agent

arXiv:2605.19301v1 Announce Type: new Abstract: Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastroph

model-releasesarxiv-cs-cv
20 May 2026
Safety

Improved visual-information-driven model for crowd simulation and its modular application

DGX agent

arXiv:2504.03758v4 Announce Type: replace-cross Abstract: Crowd movement simulation is crucial for pedestrian safety management and facility design. Data-driven models offer the potential to improve r

safetyarxiv-cs-cv
20 May 2026
Tutorials

INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference

DGX agent

arXiv:2605.18853v1 Announce Type: cross Abstract: Edge deployment of Vision-Language Models (VLMs) faces a tradeoff between latency and accuracy: cloud execution provides high-quality predictions but

tutorialsarxiv-cs-cv
20 May 2026
Research

InterLight: Leveraging Intrinsic Illumination Priors for Low-Light Image Enhancement

DGX agent

arXiv:2605.19982v1 Announce Type: new Abstract: Low-Light Image Enhancement (LLIE) has long been a challenging problem in low-level vision, as insufficient illumination often leads to low contrast, de

researcharxiv-cs-cv
20 May 2026
Research

Interpretable Computer Vision for Defect Detection in X-ray Tomography of Aerospace SiC/SiC Composites

DGX agent

arXiv:2605.20159v1 Announce Type: new Abstract: Non-destructive testing of aerospace SiC/SiC composites via X-ray computed tomography (XCT) relies on expert visual assessment, with current workflows o

researcharxiv-cs-cv
20 May 2026
Safety

Inverse Design of Metasurface based Absorbers using Physics Guided Conditional Diffusion Models

DGX agent

arXiv:2605.19611v1 Announce Type: new Abstract: Inverse design of metasurfaces for specific electromagnetic responses requires generating geometries that satisfy stringent spectral constraints while m

safetyarxiv-cs-cv
20 May 2026
Applications

LaCoVL-FER: Landmark-Guided Contrastive Learning Network with Vision-Language Enhancement for Facial Expression Recognition

DGX agent

arXiv:2605.19821v1 Announce Type: new Abstract: Facial Expression Recognition (FER) in the wild is still challenging due to uncontrolled variations in pose, occlusion, and illumination. Most existing

applicationsarxiv-cs-cv
20 May 2026
Local Ai

Landscape-Awareness for Geometric View Diffusion Model

DGX agent

arXiv:2605.19865v1 Announce Type: new Abstract: Accurate camera viewpoint estimation under sparse-view conditions remains challenging, particularly in two-view scenarios. Recent approaches leverage di

local-aiarxiv-cs-cv
20 May 2026
Research

Landslide Detection and Mapping Using Deep Learning Across Multi-Source Satellite Data and Geographic Regions

DGX agent

arXiv:2507.01123v2 Announce Type: replace Abstract: Landslides pose severe threats to infrastructure, economies, and human lives, necessitating accurate detection and predictive mapping across diverse

researcharxiv-cs-cv
20 May 2026
Research

Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection

DGX agent

arXiv:2504.00470v2 Announce Type: replace-cross Abstract: To develop a trustworthy AI system, which aim to identify the input regions that most influence the models decisions. The primary task of exis

researcharxiv-cs-cv
20 May 2026
Model Releases

LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue

DGX agent

arXiv:2605.19390v1 Announce Type: new Abstract: Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spa

model-releasesarxiv-cs-cv
20 May 2026
Research

Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation

DGX agent

arXiv:2603.16284v2 Announce Type: replace Abstract: Despite the significant advancements in Large Vision-Language Models (LVLMs), their tendency to generate hallucinations undermines reliability and r

researcharxiv-cs-cv
20 May 2026
Safety

Low-Compute Watermark Removal via Dual-Domain Natural Projection

DGX agent

arXiv:2510.07538v2 Announce Type: replace Abstract: Effective removal of semantic watermarks requires balancing three competing objectives: high removal success, low perceptual distortion, and low com

safetyarxiv-cs-cv
20 May 2026
Model Releases

MAM-CLIP: Vision-Language Pretraining on Mammography Atlases for BI-RADS Classification

DGX agent

arXiv:2605.19359v1 Announce Type: new Abstract: Deep learning methods have demonstrated promising results in predicting BI-RADS scores from mammography images. However, the interpretation of these ima

model-releasesarxiv-cs-cv
20 May 2026
Research

MapAnything: Evaluating Monocular Metric Depth Models for 3D Urban Asset Localization

DGX agent

arXiv:2509.14839v2 Announce Type: replace Abstract: City administrations increasingly rely on comprehensive databases and urban digital twins of city assets, such as traffic signs and trees, as well a

researcharxiv-cs-cv
20 May 2026
Research

Matern Noise for Triangulation-Agnostic Flow Matching on Meshes

DGX agent

arXiv:2605.19305v1 Announce Type: cross Abstract: This paper tackles the task of learning to generate signals over triangle meshes in a triangulation-agnostic manner, meaning the trained model can be

researcharxiv-cs-cv
20 May 2026
Research

MatPhys: Learning Material-Aware Physics Parameters for Deformable Object Simulation from Videos

DGX agent

arXiv:2605.19386v1 Announce Type: new Abstract: Reconstructing simulation-ready deformable objects is important for vision, graphics, and robotics. Existing physics-driven methods can recover physical

researcharxiv-cs-cv
20 May 2026
Local Ai

Mechanisms of Object Localization in Vision-Language Models

DGX agent

arXiv:2605.19792v1 Announce Type: new Abstract: Visually-grounded language models (VLMs) are highly effective in linking visual and textual information, yet they often struggle with basic classificati

local-aiarxiv-cs-cv
20 May 2026
Model Releases

MedFM-Robust: Benchmarking Robustness of Medical Foundation Models

DGX agent

arXiv:2605.19027v1 Announce Type: new Abstract: Medical foundation models (MedFMs) have emerged as transformative tools in healthcare, demonstrating capabilities across diverse clinical applications.

model-releasesarxiv-cs-cv
20 May 2026
Research

MetaEarth-MM: Unified Multimodal Remote Sensing Image Generation with Scene-centered Joint Modeling

DGX agent

arXiv:2605.20090v1 Announce Type: new Abstract: Multi-modal remote sensing images are vital for Earth observation, yet complete paired observations are often scarce in practice. Existing generative me

researcharxiv-cs-cv
20 May 2026
Model Releases

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems

DGX agent

arXiv:2605.19307v1 Announce Type: new Abstract: Visual Question Answering (VQA), as the representative multimodal task, serves as a key benchmark for evaluating the reasoning capabilities of Multimoda

model-releasesarxiv-cs-cv
20 May 2026
Applications

Minimalist Visual Inertial Odometry

DGX agent

arXiv:2605.19990v1 Announce Type: cross Abstract: Visual-Inertial Odometry(VIO), which is critical to mobile robot navigation, uses cameras with a large number of pixels. Capturing and processing came

applicationsarxiv-cs-cv
20 May 2026
Model Releases

MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

DGX agent

arXiv:2510.25897v2 Announce Type: replace Abstract: The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one rew

model-releasesarxiv-cs-cv
20 May 2026
Research

MMGS: 10imes Compressed 3DGS through Optimal Transport Aggregation based on Multi-view Ranking

DGX agent

arXiv:2605.19304v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has revolutionized 3D reconstruction, it suffers from significant overhead due to massive redundant primitives. Exist

researcharxiv-cs-cv
20 May 2026
Research

Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion Generations

DGX agent

arXiv:2412.13111v2 Announce Type: replace Abstract: Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods of

researcharxiv-cs-cv
20 May 2026
Model Releases

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

DGX agent

arXiv:2605.18956v1 Announce Type: new Abstract: Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and

model-releasesarxiv-cs-cv
20 May 2026
← Previous
1…155156157158159…263
Next →