AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
5 Aug 2026

CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation

SafetyDGX agent

arXiv:2608.03046v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by

Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models

ResearchDGX agent

arXiv:2608.03160v1 Announce Type: cross Abstract: When asked which of two events came first, video large language models can fail in two opposite ways: cave to a false claim, or reject a true one. Pri

Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.03100v1 Announce Type: new Abstract: Knowledge Distillation (KD) offers a promising yet underexplored path for compressing large action recognition models. However, existing KD methods suff

Clarity Contrast and Similarity Selection for Multi-Focus Image Fusion

ResearchDGX agent

arXiv:2608.03252v1 Announce Type: new Abstract: Multi-focus image fusion (MFIF) aims to generate an all-in-focus image from multiple images of the same scene focused at different regions. Most existin

Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray Interpretation

SafetyDGX agent

arXiv:2608.03016v1 Announce Type: new Abstract: Accurate chest X-ray interpretation is inherently hierarchical. Clinical decisions depend not only on what abnormality is present but where it is situat

CLIP4VI-ReID: Learning Modality-shared Representations via CLIP Semantic Bridge for Visible-Infrared Person Re-identification

SafetyDGX agent

arXiv:2511.10309v2 Announce Type: replace Abstract: This paper proposes a novel CLIP-driven modality-shared representation learning network named CLIP4VI-ReID for VI-ReID task, which consists of Text

Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities

ResearchDGX agent

arXiv:2608.03559v1 Announce Type: new Abstract: In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel-level performance wh

Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

Model ReleasesDGX agent

arXiv:2608.02790v1 Announce Type: new Abstract: Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evalu

CPrefix: A Combinatorial Tensor Framework for Structured Discrete Color Mappings

ResearchDGX agent

arXiv:2608.03863v1 Announce Type: new Abstract: Discrete multi-channel mappings are typically represented through sampled values, providing accurate evaluations but limited insight into their underlyi

CRIL-U-Net: Compact Ratio-Interaction Learning for Focal Cortical Dysplasia Segmentation from T1w and FLAIR MRI

TutorialsDGX agent

arXiv:2608.03185v1 Announce Type: new Abstract: Focal cortical dysplasia (FCD) type II is an important structural cause of drug-resistant focal epilepsy, but its small size, heterogeneous appearance,

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation

SafetyDGX agent

arXiv:2608.03147v1 Announce Type: new Abstract: Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SA

CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction

Model ReleasesDGX agent

arXiv:2608.03211v1 Announce Type: new Abstract: Visual world models typically learn future dynamics from a single observation stream, limiting their ability to model cooperative systems with multiple

Deep Frequency-Aware Functional Maps for Robust Shape Matching

ResearchDGX agent

arXiv:2402.03904v3 Announce Type: replace Abstract: Deep functional map frameworks are widely employed for 3D shape matching. However, most existing deep functional map methods cannot adaptively captu

Detecting Pose Estimation Failures via Keypoint Self-Consistency

ResearchDGX agent

arXiv:2608.03516v1 Announce Type: new Abstract: One common approach to pose estimation involves predicting object keypoints in an image, followed by using Perspective-n-Point algorithms to compute the

DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers

SafetyDGX agent

arXiv:2608.03082v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To fa

Double Down on Defense: Strengthening Deep Perceptual Hashes against Evasion Attacks without Retraining

SafetyDGX agent

arXiv:2608.03101v1 Announce Type: new Abstract: Near-duplicate image matching is crucial for trust and safety, provenance verification, copyright enforcement, and large-scale visual search. Modern pla

DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack

SafetyDGX agent

arXiv:2608.03207v1 Announce Type: new Abstract: Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been re

DriftWorld: Fast World Modeling through Drifting

SafetyDGX agent

arXiv:2607.15065v2 Announce Type: replace-cross Abstract: Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating man

DRPFNet: Dual-domain Residual Progressive Fusion Network for RGB-Thermal Object Detection

ResearchDGX agent

arXiv:2608.03370v1 Announce Type: new Abstract: RGB-thermal (RGB-T) object detection aims to fuse complementary information from visible and thermal modalities to achieve robust detection under varyin

Dual-domain U-Nets with embedded back projection operators for motion-resolved 4D CBCT reconstruction

ResearchDGX agent

arXiv:2608.03430v1 Announce Type: new Abstract: Four-dimensional cone beam CT (4D CBCT) is important for image-guided radiation therapy of thoracic cancers, but its use is limited by long scan times,

Earth Embeddings

ResearchDGX agent

arXiv:2608.03410v1 Announce Type: new Abstract: Earth observation is moving from foundation models that users must run themselves toward embedding products that package model feature outputs as reusab

EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation

Model ReleasesDGX agent

arXiv:2608.03179v1 Announce Type: new Abstract: Controllable local editing of 3D assets requires precise target localization and appropriate visual guidance. However, existing methods lack a simple ye

Enhanced Polarization Locking in VCSELs

SafetyDGX agent

arXiv:2604.01857v2 Announce Type: replace-cross Abstract: While optical injection locking (OIL) of vertical-cavity surface-emitting lasers (VCSELs) has been widely studied in the past, the polarizatio

FaithIR: Rethinking Infrared Image Super-Resolution from Perceptual Sharpness to Task Relevant Fidelity

ResearchDGX agent

arXiv:2608.03106v1 Announce Type: new Abstract: Infrared image super-resolution (IISR) is important for downstream tasks such as object detection and semantic segmentation. Existing IISR methods often

Fast Object Removal Attacks on Safety-Critical Video-based Perception Systems

SafetyDGX agent

arXiv:2608.02806v1 Announce Type: cross Abstract: By leveraging data from video-based perception systems, intelligent transportation systems (ITS) support safety-critical applications that improve roa

FreqAdapt: Frequency-Adaptive Processing for RAW Object Detection

Model ReleasesDGX agent

arXiv:2608.03385v1 Announce Type: new Abstract: Existing object detection methods predominantly utilize sRGB inputs, which are compressed from RAW sensor data using Image Signal Processors (ISP) origi

Frequency-Decorrelated Temporal Ensembles for EEG--fNIRS Imagined-Handwriting Decoding

ResearchDGX agent

arXiv:2608.03176v1 Announce Type: new Abstract: Imagined handwriting offers a temporally rich paradigm for non-invasive neural decoding, yet reliable recognition across unseen participants remains dif

From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

ResearchDGX agent

arXiv:2608.03508v1 Announce Type: new Abstract: Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-tr

From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation

SafetyDGX agent

arXiv:2608.03143v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) requires an agent to follow a route-level instruction by executing its constituent steps from egocentric visual obs

Frozen High-Resolution Inference for Cross-City Object Detection: An AI City Challenge 2026 Study

Model ReleasesDGX agent

arXiv:2608.03136v1 Announce Type: new Abstract: Cross-city object detection requires a detector trained in one city to generalize to an unlabeled target city. In AI City Challenge 2026 Track 6, we ana

Fusion-Poly: A Polyhedral Framework Based on Spatial-Temporal Fusion for 3D Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2603.08199v2 Announce Type: replace Abstract: LiDAR-camera 3D multi-object tracking (MOT) combines rich visual semantics with accurate depth cues to improve trajectory consistency and tracking r

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

Model ReleasesDGX agent

arXiv:2608.03826v1 Announce Type: new Abstract: Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations,

GeoMAR: Unleashing Geometrically Aligned Features for Masked Autoregressive Blind Face Restoration

ApplicationsDGX agent

arXiv:2608.03923v1 Announce Type: new Abstract: Codebook-based blind face restoration (BFR) often suffers from ambiguous conditioning features and a fragile prediction mechanism under severe degradati

Geospatial-Prior Guidance for 3D Semantic Scene Completion

ResearchDGX agent

arXiv:2608.03618v1 Announce Type: new Abstract: Inferring complete 3D geometry and semantics from onboard images remains challenging because occlusions and restricted fields of view leave large scene

Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation

ResearchDGX agent

arXiv:2608.03064v1 Announce Type: new Abstract: We study open-vocabulary 3D indoor layout generation, which synthesizes diverse and physically plausible scenes from unlabeled 3D assets using free-form

GVCCTurbo: Rate-Compute Quality Scheduling for Codebook Driven Generative Compression

TutorialsDGX agent

arXiv:2608.03517v1 Announce Type: new Abstract: Codebook-driven generative compression uses a pretrained image or video generator as a zero-shot visual prior and transmits compact codebook indices to

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation

ResearchDGX agent

arXiv:2608.03264v1 Announce Type: cross Abstract: Audio-visual instance segmentation (AVIS) requires accurately identifying and tracking individual sounding objects with pixel-level masks. Existing me

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding

SafetyDGX agent

arXiv:2608.03471v1 Announce Type: new Abstract: Generative Vision-Language Models (VLMs) commonly treat bounding-box coordinates as independent output symbols, leaving numerical order and axis semanti

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

ResearchDGX agent

arXiv:2608.02711v1 Announce Type: new Abstract: Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing.

HyperbolicDiffusion: Sharp & Scalable Tiled Generation on the Hyperbolic Plane

ResearchDGX agent

arXiv:2608.03422v1 Announce Type: new Abstract: Planar tiled diffusion denoises overlapping windows of one rectangular canvas. The hyperbolic plane has no such canvas, and its area grows exponentially

HyVIC: A Metric-Driven Spatio-Spectral Hyperspectral Image Compression Architecture Based on Variational Autoencoders

Model ReleasesDGX agent

arXiv:2603.26468v2 Announce Type: replace Abstract: The rapid growth of hyperspectral data archives in remote sensing (RS) necessitates effective compression methods for storage and transmission. Rece

iFAN: Inference-Aware Learning for Plain Mask Transformers

ResearchDGX agent

arXiv:2608.03216v1 Announce Type: new Abstract: Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of the final layer, yet this inferenc

IRIS: Visual-Semantic Binding for Forgery-Resistant Watermarking of Diffusion Images

ResearchDGX agent

arXiv:2608.03539v1 Announce Type: new Abstract: Most in-generation diffusion watermarks embed patterns independent of the image that carries them, and attackers transplant the marks onto images the ge

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Model ReleasesDGX agent

arXiv:2608.03974v1 Announce Type: new Abstract: Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term tempo

Keep the Needle, Prune the Haystack: Defect-Preserving Token Pruning for Efficient Zero-Shot Anomaly Detection

Local AiDGX agent

arXiv:2608.03681v1 Announce Type: new Abstract: Zero-shot visual anomaly detection has achieved remarkable progress, with recent vision-only approaches further improving performance while simplifying

Latent Reward Registers for Diffusion Preference Alignment

Model ReleasesDGX agent

arXiv:2608.03929v1 Announce Type: cross Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a sev

LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds

Model ReleasesDGX agent

arXiv:2608.03078v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong defect recognition capability in industrial anomaly detection. However, in lithography review,

Learning Attribute-aware Representations for Few-shot Scene Text Segmentation

SafetyDGX agent

arXiv:2504.11164v2 Announce Type: replace Abstract: Supervised scene text segmentation has achieved notable progress in recent years. However, its development is largely constrained by the scarcity of

Learning Biomechanically Plausible Human Motion from Sparse Radar Point Clouds

ResearchDGX agent

arXiv:2608.03637v1 Announce Type: new Abstract: Radar-based human pose estimation has focused on improving learning algorithms while representing the body as unconstrained keypoint coordinates. We add

Lightweight 3D Object Detection via Mamba-Based Knowledge Distillation

SafetyDGX agent

arXiv:2608.03490v1 Announce Type: cross Abstract: 3D object detection using light detection and ranging (LiDAR) sensors requires a balance between accuracy and computational efficiency for onboard per

LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation

ResearchDGX agent

arXiv:2608.03851v1 Announce Type: new Abstract: Real-time 3D perception is crucial for robotics, augmented reality, and embodied intelligence applications. Existing multi-view stereo (MVS) methods pri

Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery Previews

Model ReleasesDGX agent

arXiv:2608.02841v1 Announce Type: new Abstract: Ask a commercial image editor to preview a cosmetic procedure and it will often change more of the face than the request names: a nose edit can also smo

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

Model ReleasesDGX agent

arXiv:2608.03322v1 Announce Type: new Abstract: Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medica

Low-Dimensional High-Leverage Subspace Optimization: Beyond Full-Parameter Coupled Training for Neural Network Quantization

Model ReleasesDGX agent

arXiv:2608.03919v1 Announce Type: new Abstract: Low-bit quantization suffers severe accuracy degradation on compact networks, rooted in the dominant full-parameter coupled training paradigm that ignor

Material-Segmented Per-Pixel Emissivity Correction for Thermographic Anomaly Detection in Cultural Heritage Digital Twins

Model ReleasesDGX agent

arXiv:2608.02964v1 Announce Type: new Abstract: Quantitative longwave thermography of heritage surfaces is limited by the global-constant emissivity assumption in inverse-Planck temperature retrieval;

MaterialFusion: High-Quality, Zero-Shot, and Controllable Material Transfer with Diffusion Models

ApplicationsDGX agent

arXiv:2502.06606v3 Announce Type: replace Abstract: Manipulating the material appearance of objects in images is critical for applications like augmented reality, virtual prototyping, and digital cont

MeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusion

SafetyDGX agent

arXiv:2508.15169v4 Announce Type: replace Abstract: Mesh models have become increasingly accessible for numerous cities; however, the lack of realistic textures restricts their application in virtual

Micro-Segmentation Anomaly Detection in Zero-Trust Software-Defined Network Fabrics

ApplicationsDGX agent

arXiv:2608.02627v1 Announce Type: cross Abstract: Zero Trust Architecture (ZTA) principles need rigorous network segmentation and ongoing verification to reduce implicit trust and lateral threat propa

MinerU.Chem: A High-Precision System for Optical Chemical Structure and Reaction Recognition

Model ReleasesDGX agent

arXiv:2608.03525v1 Announce Type: new Abstract: In organic chemistry papers and patents, molecular structures, reaction schemes, and experimental conditions are often presented as molecular structure

Modeling Long-Term Memory and Temporal Attention Shifts for Video Salient Object Ranking with a New Benchmark

Model ReleasesDGX agent

arXiv:2203.17257v2 Announce Type: replace Abstract: Salient Object Ranking (SOR) aims to estimate the relative saliency order among multiple salient objects. While SOR has been extensively studied in

← Previous
1…1112131415…207
Next →