AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

Harnessing Streaming Video in the Wild

DGX agent

arXiv:2606.08615v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentar

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging

DGX agent
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

arXiv:2606.09110v1 Announce Type: new Abstract: Most existing multi-exposure HDR methods follow a fixed feed-forward reconstruction paradigm, making them prone to ghosting artifacts in complex dynamic

safetyarxiv-cs-cv
9 Jun 2026
Model Releases

HDSL: A Hierarchical Domain-Specific Language for Structured 3D Indoor Scene Generation and Localized Editing with LLM Agents

DGX agent

arXiv:2606.09738v1 Announce Type: new Abstract: Text-driven indoor scene generation and editing require an intermediate representation that language models can both produce and revise. Existing LLM-ba

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

HiMat: DiT-based Ultra-High Resolution SVBRDF Generation

DGX agent

arXiv:2508.07011v5 Announce Type: replace Abstract: Creating ultra-high-resolution spatially varying bidirectional reflectance functions (SVBRDFs) is critical for photorealistic 3D content creation, t

safetyarxiv-cs-cv
9 Jun 2026
Research

How Much MRI Preprocessing Is Enough? A Cost-Utility Study for Brain MRI Foundation Models

DGX agent

arXiv:2606.08164v1 Announce Type: new Abstract: MRI preprocessing defines the input distribution seen by brain MRI foundation models, yet it is usually treated as routine data cleaning rather than a m

researcharxiv-cs-cv
9 Jun 2026
Research

Hummus: A Dataset of Humorous Multimodal Metaphor Use

DGX agent

arXiv:2504.02983v3 Announce Type: replace-cross Abstract: Metaphor and humor share a lot of common ground, and metaphor is one of the most common humorous mechanisms. This study focuses on the humorou

researcharxiv-cs-cv
9 Jun 2026
Safety

Hyperspectral Smoke Segmentation via Mixture of Prototypes

DGX agent

arXiv:2602.10858v2 Announce Type: replace Abstract: Smoke segmentation is critical for wildfire management and industrial safety applications. Traditional visible-light-based methods face limitations

safetyarxiv-cs-cv
9 Jun 2026
Research

IB-HFN: Information Bottleneck-Driven SAR-Optical Fusion Network for High-Fidelity Cloud Removal

DGX agent

arXiv:2606.09347v1 Announce Type: new Abstract: Synthetic aperture radar (SAR)-assisted optical cloud removal aims to recover surface information obscured by clouds in optical remote sensing images by

researcharxiv-cs-cv
9 Jun 2026
Model Releases

IDDM: Identity-Decoupled Personalized Diffusion Models with a Tunable Privacy-Utility Trade-off

DGX agent

arXiv:2604.00903v2 Announce Type: replace Abstract: Personalized text-to-image diffusion models (e.g., DreamBooth, LoRA) enable users to synthesize high-fidelity avatars from a few reference photos fo

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation

DGX agent

arXiv:2601.04498v2 Announce Type: replace-cross Abstract: Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information

model-releasesarxiv-cs-cv
9 Jun 2026
Research

Illumination-Invariant Anomaly Detection for Sub-Canopy UAV Multispectral Point Clouds

DGX agent

arXiv:2606.09111v1 Announce Type: new Abstract: Unmanned Aerial Vehicle (UAV) multispectral point clouds (MPC) provide high-dimensional spatial-spectral data for sub-canopy target detection; however,

researcharxiv-cs-cv
9 Jun 2026
Applications

iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

DGX agent

arXiv:2606.09813v1 Announce Type: cross Abstract: Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environment simulation. However, conventio

applicationsarxiv-cs-cv
9 Jun 2026
Research

IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval

DGX agent

arXiv:2606.08144v1 Announce Type: new Abstract: Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing met

researcharxiv-cs-cv
9 Jun 2026
Model Releases

KITE: A Tri-Modal Transformer Integrating Text, Images, and Knowledge Graphs for Fake News Detection

DGX agent

arXiv:2606.07651v1 Announce Type: cross Abstract: Traditional fake news detection methods are falling behind as multimodal misinformation grows more advanced, seamlessly blending deceptive text, manip

model-releasesarxiv-cs-cv
9 Jun 2026
Research

Latent Spatial Memory for Video World Models

DGX agent

arXiv:2606.09828v1 Announce Type: new Abstract: Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space.

researcharxiv-cs-cv
9 Jun 2026
Tutorials

Learnable Token Sparsification for Efficient Gigapixel Whole Slide Image Reasoning

DGX agent

arXiv:2606.08641v1 Announce Type: new Abstract: The processing of gigapixel whole slide images within vision language models faces a major difficulty due to an excessive number of visual tokens. Exist

tutorialsarxiv-cs-cv
9 Jun 2026
Research

Learning a Semantic Calibration Network for Open-Vocabulary Semantic Segmentation

DGX agent

arXiv:2606.08001v1 Announce Type: new Abstract: Semantic image segmentation assigns a predefined category label to each pixel, has achieved significant progress lately. Open-Vocabulary Segmentation (O

researcharxiv-cs-cv
9 Jun 2026
Research

Learning to Solve Generative ODEs Beyond the Linear Span

DGX agent

arXiv:2606.08672v1 Announce Type: new Abstract: Diffusion and flow generative models sample by integrating a learned ODE, but high quality still requires many sequential model evaluations. Solver lear

researcharxiv-cs-cv
9 Jun 2026
Research

LEGS: Laplacian-Enhanced Gaussian Splatting with a Nonlinear Weighted Loss

DGX agent

arXiv:2606.07932v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has become an efficient explicit representation for radiance field reconstruction and real-time novel view synthesis. Howev

researcharxiv-cs-cv
9 Jun 2026
Local Ai

Less Is More: Training-Free Acceleration Framework of 3D Diffusion Models for Low-Count PET Denoising via Global-Local Trajectory Reduction

DGX agent

arXiv:2606.08751v1 Announce Type: new Abstract: Accurate quantification and uptake measurement in PET are critical for assessing disease progression and supporting clinical decision-making. While high

local-aiarxiv-cs-cv
9 Jun 2026
Tutorials

Leveraging Morphology for Historical Script Metrological Analysis

DGX agent

arXiv:2606.09446v1 Announce Type: new Abstract: Advances in handwritten text recognition have enabled large-scale transcription of historical documents, but still provide limited access to interpretab

tutorialsarxiv-cs-cv
9 Jun 2026
Research

Leveraging NeRF-Rendered Images for 3D Gaussian Splatting

DGX agent

arXiv:2606.09034v1 Announce Type: new Abstract: Neural radiance field (NeRF) and 3D Gaussian splatting (3DGS) are two mainstream approaches for novel view synthesis. They often show complementary perf

researcharxiv-cs-cv
9 Jun 2026
Safety

Light-WAM: Efficient World Action Models with State-Fusion Action Decoding

DGX agent

arXiv:2606.08242v1 Announce Type: new Abstract: World Action Models (WAMs) extend robot policy learning by incorporating future prediction as an additional training objective, encouraging the policy t

safetyarxiv-cs-cv
9 Jun 2026
Safety

LiteVSR: Lightweight Adaptation of Frozen Diffusion Transformers for Video Super-Resolution

DGX agent

arXiv:2606.09250v1 Announce Type: new Abstract: Adapting large-scale pre-trained video generators for Video Super-Resolution (VSR) in novel domains remains computationally prohibitive. Methods that re

safetyarxiv-cs-cv
9 Jun 2026
Research

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs

DGX agent

arXiv:2606.08511v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) face a significant inference bottleneck due to the quadratic computational cost of self-attention over long vis

researcharxiv-cs-cv
9 Jun 2026
Model Releases

MAGIS: Evidence-Based Multi-Agent Reasoning for Interpretable Strabismus Clinical Decision-Making

DGX agent

arXiv:2606.09249v1 Announce Type: new Abstract: Strabismus is a common ocular disorder that requires fine-grained subtype diagnosis for individualized treatment planning. However, existing deep learni

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training

DGX agent

arXiv:2606.08788v1 Announce Type: new Abstract: Representation alignment with pretrained vision models has recently shown strong potential for accelerating diffusion transformer training. By aligning

safetyarxiv-cs-cv
9 Jun 2026
Agents

MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding

DGX agent

arXiv:2606.09641v1 Announce Type: new Abstract: The dominant paradigm in video retrieval relies on embedding-based full-corpus scanning, which suffers from inherent computational inefficiency and the

agentsarxiv-cs-cv
9 Jun 2026
Local Ai

MB-Loc: Multi-planar Bird's-eye-view Localization in outdoor LiDAR scenes

DGX agent

arXiv:2606.08744v1 Announce Type: new Abstract: Global LiDAR localization is a fundamental task for autonomous navigation systems. Recent methods perform Scene Coordinate Regression (SCR) and achieve

local-aiarxiv-cs-cv
9 Jun 2026
Research

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models

DGX agent

arXiv:2606.09827v1 Announce Type: cross Abstract: Temporal modeling is essential for robotic manipulation, as effective control requires both memory of past interactions and imagination of future stat

researcharxiv-cs-cv
9 Jun 2026
Research

MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation

DGX agent

arXiv:2606.09056v1 Announce Type: new Abstract: Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames req

researcharxiv-cs-cv
9 Jun 2026
Applications

Minimal Solvers for Full-DoF Motion Estimation from Asynchronous Differential SfM

DGX agent

arXiv:2606.09218v1 Announce Type: new Abstract: As a bio-inspired intelligent sensor, event cameras have introduced a new paradigm in the intelligent perception of spatiotemporal information and visua

applicationsarxiv-cs-cv
9 Jun 2026
Agents

MinNav: Minimalist Navigation Using Optical Flow For Active Tiny Aerial Robots

DGX agent

arXiv:2606.07813v1 Announce Type: cross Abstract: Navigation using a monocular camera is pivotal for autonomous operation on tiny aerial robots due to their perfect balance of versatility, cost and ac

agentsarxiv-cs-cv
9 Jun 2026
Research

Mitigating Diffusion Model Hallucinations with Dynamic Guidance

DGX agent

arXiv:2510.05356v2 Announce Type: replace Abstract: Hallucinations in diffusion models are samples with structural inconsistencies that can emerge due to the excessive smoothing of the learned score f

researcharxiv-cs-cv
9 Jun 2026
Research

MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

DGX agent

arXiv:2410.21747v2 Announce Type: replace Abstract: Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging re

researcharxiv-cs-cv
9 Jun 2026
Research

MS-COOT: Comparing Morse-Smale Complexes with Co-Optimal Transport

DGX agent

arXiv:2606.08258v1 Announce Type: cross Abstract: Understanding and comparing structures in scalar fields is a central challenge in scientific visualization, with applications ranging from feature ana

researcharxiv-cs-cv
9 Jun 2026
Safety

Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training

DGX agent

arXiv:2601.03256v2 Announce Type: replace Abstract: We present Muses, the first training-free method for fantastic 3D creature generation in a feed-forward paradigm. Previous methods, which rely on pa

safetyarxiv-cs-cv
9 Jun 2026
Local Ai

Need We Teach Foundation Models What is a Generative Image? Gradient-Free Generative Artifact Detection via Analytic Spectral Adaptation

DGX agent

arXiv:2606.07660v1 Announce Type: new Abstract: Adapting foundation models to detect generative artifacts via gradient-based updates compromises their intrinsic representations. Under optimization on

local-aiarxiv-cs-cv
9 Jun 2026
Tutorials

Neural Field Tokenizations with Hierarchy and Spatial Locality Priors

DGX agent

arXiv:2606.08204v1 Announce Type: cross Abstract: Neural fields parameterize data as functions from coordinates to values, providing a unified framework for representation learning across modalities.

tutorialsarxiv-cs-cv
9 Jun 2026
Model Releases

NGram-MoSE: Efficient Remote Sensing Super-Resolution via N-Gram Context and Mixture-of-Experts

DGX agent

arXiv:2606.08535v1 Announce Type: new Abstract: Remote sensing applications for environmental monitoring and disaster management are frequently constrained by a spatial--temporal trade-off: imagery wi

model-releasesarxiv-cs-cv
9 Jun 2026
Safety

No Modality Left Behind: Adapting to Missing Modalities via Knowledge Distillation for Brain Tumor Segmentation

DGX agent

arXiv:2509.15017v2 Announce Type: replace Abstract: Accurate brain tumor segmentation is essential for preoperative evaluation and personalized treatment. Multi-modal MRI is widely used due to its abi

safetyarxiv-cs-cv
9 Jun 2026
Research

OctaOctree Neural Radiosity for Real-time Glossy Material Rendering

DGX agent

arXiv:2606.08469v1 Announce Type: cross Abstract: Modeling high-frequency outgoing radiance distributions remains a fundamental challenge in global illumination, especially for glossy and specular mat

researcharxiv-cs-cv
9 Jun 2026
Model Releases

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

DGX agent

arXiv:2606.08572v1 Announce Type: new Abstract: While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability t

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

OmniFaceRig: Fully Automatic Inner-Mouth-Aware Face Rigging Across Diverse 3D Character Topologies

DGX agent

arXiv:2606.08043v1 Announce Type: cross Abstract: Facial rigging - creating FACS-based blendshapes together with inner-mouth geometry (teeth, gums, and tongue) - remains a major bottleneck in 3D chara

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

OmniGen-AR: AutoRegressive Any-to-Image Generation

DGX agent

arXiv:2606.09156v1 Announce Type: new Abstract: Autoregressive (AR) models have demonstrated strong potential in visual generation, offering superior performance with simple architectures and optimiza

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

OmniTryOn: Video Try-On Anything at Once!

DGX agent

arXiv:2606.08514v1 Announce Type: new Abstract: Although video virtual try-on (VVT) has achieved significant progress, existing methods still exhibit two fundamental limitations: first, they are restr

model-releasesarxiv-cs-cv
9 Jun 2026
Research

One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling

DGX agent

arXiv:2606.08126v1 Announce Type: new Abstract: Vision-language models (VLMs) enable visual recognition from semantic class descriptions, which makes them attractive when target annotations are scarce

researcharxiv-cs-cv
9 Jun 2026
Applications

Optical Music Recognition for Real-World Manuscripts with Synthetic Data

DGX agent

arXiv:2606.09479v1 Announce Type: new Abstract: Optical Music Recognition (OMR) has seen major progress in model design, with end-to-end methods now capable of recognising notation at all levels of co

applicationsarxiv-cs-cv
9 Jun 2026
← Previous
1…111112113114115…263
Next →