AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations

DGX agent

arXiv:2607.08398v1 Announce Type: cross Abstract: Standard pipelines for physics-ready 3D reconstruction rely on a decoupled two-stage paradigm: extracting surface geometry followed by an error-prone

researcharxiv-cs-cv
10 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition

DGX agent

arXiv:2607.08249v1 Announce Type: new Abstract: Slot attention is a powerful framework for object-centric learning, decomposing visual scenes into latent slots through iterative competitive attention.

safetyarxiv-cs-cv
10 Jul 2026
Model Releases

HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-Agent Forgery Rationales

DGX agent

arXiv:2607.08705v1 Announce Type: new Abstract: Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, posing unp

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning

DGX agent

arXiv:2603.12478v2 Announce Type: replace Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is

model-releasesarxiv-cs-cv
10 Jul 2026
Safety

LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding

DGX agent

arXiv:2604.01388v2 Announce Type: replace Abstract: Recent advancements in open-vocabulary 3D scene understanding heavily rely on 3D Gaussian Splatting (3DGS) to register vision-language features into

safetyarxiv-cs-cv
10 Jul 2026
Model Releases

LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting

DGX agent

arXiv:2607.08016v1 Announce Type: new Abstract: Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurat

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

LlamaSeg: Image Segmentation via Autoregressive Mask Generation

DGX agent

arXiv:2505.19422v2 Announce Type: replace Abstract: We present extbf{LlamaSeg}, a visual autoregressive framework that unifies multiple image segmentation tasks via natural language instructions. By r

model-releasesarxiv-cs-cv
10 Jul 2026
Tutorials

LOGOS: Language-guided Oriented Object Detection in Aerial Scenes

DGX agent

arXiv:2607.08004v1 Announce Type: new Abstract: Object detection in geospatial scenes, such as satellite and aerial imagery, poses significant challenges due to the varying orientations and densities

tutorialsarxiv-cs-cv
10 Jul 2026
Safety

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

DGX agent

arXiv:2607.08770v1 Announce Type: new Abstract: Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models

safetyarxiv-cs-cv
10 Jul 2026
Safety

LTM: Large-scale Terrain Model for Wildfire-prone Landscapes

DGX agent

arXiv:2607.08711v1 Announce Type: new Abstract: Accurate 3D terrain maps are essential for emergency response when assessing wildfire hazards. However, wildfire-prone regions often span vast areas whe

safetyarxiv-cs-cv
10 Jul 2026
Model Releases

LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression

DGX agent

arXiv:2607.08221v1 Announce Type: new Abstract: Large language model (LLM)-based lossless image compression methods typically represent pixel data through the native text interface of a pretrained mod

model-releasesarxiv-cs-cv
10 Jul 2026
Research

Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks

DGX agent

arXiv:2607.08203v1 Announce Type: new Abstract: Progress in colonoscopy polyp segmentation is routinely reported through leaderboard comparisons on a small set of public benchmarks. We argue that this

researcharxiv-cs-cv
10 Jul 2026
Model Releases

Mixture of Enhanced-View Experts for Multi-Query Vehicle ReID and A Large-Scale Benchmark

DGX agent

arXiv:2607.08085v1 Announce Type: new Abstract: Multi-query vehicle ReID aims to leverage complementary information from diverse views for robust feature learning. However, current methods suffer from

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Multi-Resolution Feature Stem for Diabetic Retinopathy lesion segmentation

DGX agent

arXiv:2607.08679v1 Announce Type: new Abstract: Diabetic Retinopathy (DR) is a leading cause of preventable blindness worldwide, requiring automated lesion segmentation using deep learning models for

model-releasesarxiv-cs-cv
10 Jul 2026
Research

Multimodal 3D LUT Generation via StatLUT with Statistical Features for Photorealistic Style Transfer

DGX agent

arXiv:2607.08227v1 Announce Type: new Abstract: Photorealistic Style Transfer (PST) aims to transfer the color and tonal style of a reference to a content image while strictly preserving its structura

researcharxiv-cs-cv
10 Jul 2026
Safety

Native Video-Action Pretraining for Generalizable Robot Control

DGX agent

arXiv:2607.08639v1 Announce Type: cross Abstract: The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed

safetyarxiv-cs-cv
10 Jul 2026
Local Ai

OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing

DGX agent

arXiv:2606.15920v2 Announce Type: replace Abstract: Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This cha

local-aiarxiv-cs-cv
10 Jul 2026
Applications

On the Design of Mixture-of-Experts for Dynamic Gaussian Splatting

DGX agent

arXiv:2607.08250v1 Announce Type: new Abstract: Dynamic scene reconstruction remains challenging due to the heterogeneous and spatially varying nature of real-world motion. Although recent 3D Gaussian

applicationsarxiv-cs-cv
10 Jul 2026
Safety

OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

DGX agent

arXiv:2607.08766v1 Announce Type: new Abstract: We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR v

safetyarxiv-cs-cv
10 Jul 2026
Safety

PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM

DGX agent

arXiv:2505.16456v3 Announce Type: replace Abstract: Recent advances in 3D content generation have amplified demand for dynamic models that are both visually realistic and physically consistent. Howeve

safetyarxiv-cs-cv
10 Jul 2026
Safety

Post-Training in End-to-End Autonomous Driving

DGX agent

arXiv:2607.08072v1 Announce Type: new Abstract: End-to-end models that map multimodal inputs directly to future trajectories/maneuvers have emerged as an increasingly prominent research paradigm in au

safetyarxiv-cs-cv
10 Jul 2026
Research

Predicting Viticulture Potential through an Ensemble of U-Net and a Geospatial Foundation Model

DGX agent

arXiv:2607.08449v1 Announce Type: new Abstract: Determining agricultural potential is fundamental to sustainable land management and agricultural planning. Remote sensing data is increasingly valuable

researcharxiv-cs-cv
10 Jul 2026
Research

PRGCN: A Graph Memory Network for Cross-Sequence Pattern Reuse in 3D Human Pose Estimation

DGX agent

arXiv:2510.19475v2 Announce Type: replace Abstract: Monocular 3D human pose estimation remains a fundamentally ill-posed inverse problem due to the inherent depth ambiguity in 2D-to-3D lifting. While

researcharxiv-cs-cv
10 Jul 2026
Local Ai

Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies

DGX agent

arXiv:2607.08270v1 Announce Type: new Abstract: Forecasting the future anatomy of slow-evolving neurodegenerative diseases could enable earlier, more targeted intervention and improve clinical trial d

local-aiarxiv-cs-cv
10 Jul 2026
Model Releases

Real-World Blind Super-Resolution via Feature Matching with Implicit High-Resolution Priors

DGX agent

arXiv:2202.13142v3 Announce Type: replace Abstract: A key challenge of real-world image super-resolution (SR) is to recover the missing details in low-resolution (LR) images with complex unknown degra

model-releasesarxiv-cs-cv
10 Jul 2026
Research

SAGA: Stable Acceleration Guidance for Autoregressive Video Generation

DGX agent

arXiv:2607.08020v1 Announce Type: new Abstract: Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context

researcharxiv-cs-cv
10 Jul 2026
Research

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

DGX agent

arXiv:2607.08688v1 Announce Type: new Abstract: Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable perform

researcharxiv-cs-cv
10 Jul 2026
Safety

SASGeo: Stability-Aware Semantic Map Localization for GNSS-Denied UAVs -- A Framework and Synthetic Proof of Concept

DGX agent

arXiv:2607.07737v1 Announce Type: cross Abstract: GNSS-denied unmanned aerial vehicles require occasional absolute position fixes to bound the drift of visual-inertial odometry. Cross-view image retri

safetyarxiv-cs-cv
10 Jul 2026
Safety

Search-based Testing of Vision Language Models for In-Car Scene Understanding

DGX agent

arXiv:2607.02300v2 Announce Type: replace Abstract: In the automotive domain, in-car scene understanding (ISU) enables the detection of safety-critical events, such as driver distraction, and supports

safetyarxiv-cs-cv
10 Jul 2026
Applications

SkelGen4D: Weakly-Supervised Skeleton-Based 4D Generation for Text-Driven Mesh Animation

DGX agent

arXiv:2607.08246v1 Announce Type: new Abstract: We study 4D generation to synthesize temporally coherent sequences of 3D geometry for animation and content creation. In contrast to existing SDS-based

applicationsarxiv-cs-cv
10 Jul 2026
Safety

SPHINX: First Explain, Then Explore

DGX agent

arXiv:2606.17482v2 Announce Type: replace Abstract: Generating adversarial driving scenarios is critical for evaluating and improving autonomous vehicle decision-making systems in simulation. Recent a

safetyarxiv-cs-cv
10 Jul 2026
Applications

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

DGX agent

arXiv:2507.05240v2 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low

applicationsarxiv-cs-cv
10 Jul 2026
Tutorials

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning

DGX agent

arXiv:2607.08572v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) often follow a fixed Think-then-Answer paradigm, which is inefficient in heterogeneous multitask settings becau

tutorialsarxiv-cs-cv
10 Jul 2026
Research

Systematic Evaluation of Learning Rate Scheduling Strategies Across Heterogeneous Architectures

DGX agent

arXiv:2607.08511v1 Announce Type: cross Abstract: Choosing a learning rate scheduling strategy is critical to neural network training, but manual selection is costly and rarely exhaustive. While class

researcharxiv-cs-cv
10 Jul 2026
Safety

Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception

DGX agent

arXiv:2607.08321v1 Announce Type: new Abstract: In computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment they

safetyarxiv-cs-cv
10 Jul 2026
Tutorials

TrackStudio: An Integrated Toolkit for Markerless Tracking

DGX agent

arXiv:2511.07624v3 Announce Type: replace Abstract: Markerless motion tracking has advanced rapidly in the past 10 years and currently offers powerful opportunities for behavioural, clinical, and biom

tutorialsarxiv-cs-cv
10 Jul 2026
Research

Transformed ell_1 Gradient Regularization for Image Denoising

DGX agent

arXiv:2511.15060v2 Announce Type: replace-cross Abstract: Total variation (TV) regularization is a classical edge-preserving technique widely used across image recovery and reconstruction problems; ho

researcharxiv-cs-cv
10 Jul 2026
Model Releases

TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading

DGX agent

arXiv:2607.08236v1 Announce Type: new Abstract: Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and mo

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation

DGX agent

arXiv:2607.08075v1 Announce Type: new Abstract: Unmanned Aerial Vehicle (UAV) videos are widely used in traffic monitoring, urban management, and emergency rescue. However, existing UAV video percepti

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

DGX agent

arXiv:2607.08127v1 Announce Type: new Abstract: Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these p

model-releasesarxiv-cs-cv
10 Jul 2026
Safety

Unified Face Attack Detection via Fine-Grained Semantic Guidance

DGX agent

arXiv:2607.08156v1 Announce Type: new Abstract: The growing applications of facial recognition systems are accompanied by increasingly diverse security threats. Existing datasets lack detailed textual

safetyarxiv-cs-cv
10 Jul 2026
Model Releases

UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery

DGX agent

arXiv:2607.08267v1 Announce Type: new Abstract: Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in comple

model-releasesarxiv-cs-cv
10 Jul 2026
Applications

Unpaired Joint Distribution Modeling via Multi-Scale Image Representations

DGX agent

arXiv:2607.08198v1 Announce Type: new Abstract: This paper studies the problem of learning a joint distribution from marginal observations, which is inherently ill-posed due to the ambiguity of feasib

applicationsarxiv-cs-cv
10 Jul 2026
Research

Vision-Language Memory for Spatial Reasoning

DGX agent

arXiv:2511.20644v2 Announce Type: replace Abstract: Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level perform

researcharxiv-cs-cv
10 Jul 2026
Model Releases

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness

DGX agent

arXiv:2607.08112v1 Announce Type: new Abstract: We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Trichogramma Wasps

DGX agent

arXiv:2607.08729v1 Announce Type: new Abstract: Multi-object tracking (MOT) has achieved strong performance on benchmarks dominated by short video sequences. However, such datasets do not adequately e

model-releasesarxiv-cs-cv
10 Jul 2026
Research

Wat3R: Underwater 3D Geometry Learning without Annotations

DGX agent

arXiv:2607.08772v1 Announce Type: new Abstract: Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-

researcharxiv-cs-cv
10 Jul 2026
Research

Whareformer: Learning to Track What is Where in Long Egocentric Videos

DGX agent

arXiv:2607.08537v1 Announce Type: new Abstract: The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wea

researcharxiv-cs-cv
10 Jul 2026
← Previous
1…5455565758…261
Next →