AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
11 Aug 2026

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

SafetyDGX agent

arXiv:2608.09448v1 Announce Type: cross Abstract: Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains d

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

Model ReleasesDGX agent

arXiv:2608.09573v1 Announce Type: new Abstract: Natural-language-driven 'vibe coding' enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of thei

View-Adaptive Renderer for View-Consistent 2D-to-3D Generation

ApplicationsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.09110v1 Announce Type: new Abstract: Reconstructing 3D shapes from a single image remains a fundamental yet challenging problem in computer vision. Traditional monocular 3D generation pipel

VIGIL: Tackling Hallucination Detection in Image Recontextualization

Model ReleasesDGX agent

arXiv:2602.14633v2 Announce Type: replace Abstract: We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), a benchmark dataset and framework that provides a fine-grained categoriz

Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties

ResearchDGX agent

arXiv:2608.07726v1 Announce Type: new Abstract: Estimating volumetric mechanical properties, including Young's modulus, Poisson's ratio, and density at each voxel, is intrinsically ambiguous from visi

VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs

ResearchDGX agent

arXiv:2510.16598v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) encounter significant computational and memory bottlenecks from the massive number of visual tokens generat

Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding

Local AiDGX agent

arXiv:2608.08832v1 Announce Type: new Abstract: Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nod

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

Model ReleasesDGX agent

arXiv:2608.08630v1 Announce Type: new Abstract: Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-atte

VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression

ResearchDGX agent

arXiv:2608.08366v1 Announce Type: new Abstract: Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few th

Warp-free Cross-view Geo-localization via Feature-space Consensus Mining

Local AiDGX agent

arXiv:2608.09321v1 Announce Type: new Abstract: Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imager

What Irregularity Costs: CUDA C++, Rust, and Triton on a Hash-Blocked GPU Workload

HardwareDGX agent

arXiv:2608.08287v1 Announce Type: new Abstract: GPU language comparisons are almost always run on tiled dense linear algebra, where every toolchain is good and the differences are small. We implement

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery

Model ReleasesDGX agent

arXiv:2608.08132v1 Announce Type: new Abstract: Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical informa

Where Is the Bee? Detecting Tiny Pollinators with a Single Collaborative-Head Transformer

ResearchDGX agent

arXiv:2608.08580v1 Announce Type: new Abstract: The CVPPA@ECCV 2026 BuzzSpot Challenge asks us to detect bees, bumblebees, hoverflies, and moths in 1920x1080 field keyframes. Its annotations carry 2 d

Wiener Representation Filtering for VLM Hallucination Suppression

Model ReleasesDGX agent

arXiv:2608.08167v1 Announce Type: new Abstract: Vision-language models (VLMs) excel at open-ended captioning and visual QA but often describe objects, attributes, or relations absent from the image, a

World Simulator: Queer Erotica and the Absurdity of AI Video Models That Promise the World

ResearchDGX agent

arXiv:2608.07510v1 Announce Type: cross Abstract: Increasingly, AI video models are marketed as 'world simulators,' suggesting their ability to model infinite realities. Despite such claims, these mod

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

SafetyDGX agent

arXiv:2608.09730v1 Announce Type: new Abstract: Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicit

XClipGS: Exact Half-Space Clipping for Medical Volume Gaussian Splatting

Local AiDGX agent

arXiv:2608.07760v1 Announce Type: new Abstract: Gaussian-splatting proxies enable interactive rendering of volumetric medical scans, but a clipping plane exposes anatomy not constrained by external-vi

XEns-CKD: An Explainable Ensemble-Based Approach for Chronic Kidney Disease Stage Detection

ResearchDGX agent

arXiv:2608.07561v1 Announce Type: new Abstract: Chronic kidney disease (CKD) is a silent disease. Its progression may not significantly hamper a person's daily routine. Human kidney function can be cl

XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher

Model ReleasesDGX agent

arXiv:2608.09519v1 Announce Type: new Abstract: We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images e

You Only Flow Once: Calibrated and Real-Time Radar Pose Estimation with Multi-Hypothesis Normalizing Flows

ResearchDGX agent

arXiv:2608.09579v1 Announce Type: new Abstract: Sparse and noisy millimeter-wave radar point cloud observations often correspond to multiple plausible human poses, making deterministic pose estimation

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No

Model ReleasesDGX agent

arXiv:2608.08315v1 Announce Type: new Abstract: Multimodal LLMs that recognise events reliably still fail to say when they happen. Prompted for timestamps, strong VLMs reach as little as 3.8% R@0.5 on

Zero-shot 2D Grounding with Novel Affordance Types

Model ReleasesDGX agent

arXiv:2608.08929v1 Announce Type: new Abstract: 2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types

Zero-Shot Traffic Accident Detection via a Coarse-to-Fine VLM-Tracking Pipeline

Model ReleasesDGX agent

arXiv:2608.08867v1 Announce Type: new Abstract: Traffic surveillance cameras capture accidents continuously, yet converting raw CCTV footage into structured event records that pinpoint when, where, an

ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models

ResearchDGX agent

arXiv:2608.08060v1 Announce Type: new Abstract: Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pa

10 Aug 2026

Addressable Memory for Video World Models

Model ReleasesDGX agent

arXiv:2608.07408v1 Announce Type: new Abstract: We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward p

AdvTiles: Physical Adversarial Camouflage Clothing against Person Detectors via Learnable Tiles

ApplicationsDGX agent

arXiv:2608.06801v1 Announce Type: new Abstract: Physical adversarial attacks against person detectors have evolved from localized patches to full-body textures. However, achieving both visual naturaln

An AI4AI Framework for Visual Token Pruning

SafetyDGX agent

arXiv:2608.07193v1 Announce Type: cross Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fix

AnyTrack: Unifying Visual Object Tracking with Any Modalities

ResearchDGX agent

arXiv:2608.06773v1 Announce Type: new Abstract: Visual object tracking aims to continuously locate specific targets within sequential frames, evolving from single-modal methods to multi-modal ones. Ho

Are Visual Place Recognition Models Recognizing Places or Conditions? Distractor-Augmented Evaluation and Condition Suppression

ResearchDGX agent

arXiv:2608.06847v1 Announce Type: cross Abstract: Long-term Visual Place Recognition (VPR) is typically evaluated by matching queries from one condition against a database from another. Crowdsourced m

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2608.06729v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially ob

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

Model ReleasesDGX agent

arXiv:2608.06930v1 Announce Type: new Abstract: Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main

Bend the Basics: Degradation-Aware Deformable Tokenization for All-in-One Image Restoration

ResearchDGX agent

arXiv:2608.06832v1 Announce Type: new Abstract: All-in-one image restoration seeks a single model that can recover images degraded by diverse and spatially non-uniform corruptions. However, many unifi

Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation

Model ReleasesDGX agent

arXiv:2608.07117v1 Announce Type: new Abstract: Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for

C2Dex: Contact-Consistent Reconstruction and Retargeting for Dexterous Manipulation from Monocular Video

TutorialsDGX agent

arXiv:2608.07045v1 Announce Type: cross Abstract: High-quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocular human videos provide a scalable so

CADSpotting: Robust Panoptic Symbol Spotting on Large-Scale CAD Drawings

Model ReleasesDGX agent

arXiv:2412.07377v5 Announce Type: replace Abstract: We introduce CADSpotting, an effective method for panoptic symbol spotting in large-scale architectural CAD drawings. Existing approaches often stru

CANIS: Generation-Assisted 3D Canonicalization via an Image-Semantic Bridge

ApplicationsDGX agent

arXiv:2608.07256v1 Announce Type: new Abstract: Canonicalizing 3D object orientation is fundamental to 3D understanding and analysis. Existing approaches often rely on geometric cues, although 3D cano

CAS2UML: A Handwritten Sketch-to-PlantUML Dataset for Class and Activity Diagrams

ResearchDGX agent

arXiv:2608.07036v1 Announce Type: cross Abstract: Automated UML generation from sketches and images is gaining renewed attention with the rise of large language models and multimodal AI. However, repr

Casting the Net! Revisiting MasterFace Impersonation Attacks

ApplicationsDGX agent

arXiv:2608.06952v1 Announce Type: cross Abstract: Impersonation is a fundamental security threat in face recognition systems (FRSs). While the security of FRSs has been challenged by various attack ve

CloudDiffusion: Diffusion-Based Scene Completion in the Point Cloud Domain

AgentsDGX agent

arXiv:2606.16048v3 Announce Type: replace Abstract: Reconstructing dense 3D scenes from sparse LiDAR point clouds (LiDAR scene completion) is a fundamental challenge in autonomous driving, where diffu

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

Model ReleasesDGX agent

arXiv:2608.06691v1 Announce Type: new Abstract: Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency

Conformal Coverage Guarantees for Any Video Temporal Grounder

Local AiDGX agent

arXiv:2608.07434v1 Announce Type: new Abstract: Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less t

ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE

Local AiDGX agent

arXiv:2608.06878v1 Announce Type: new Abstract: Layout-guided multi-instance generation is essential for controllable image synthesis in Multi-Modal Diffusion Transformers (MM-DiTs). However, integrat

Corrupting Attention: Evasion-Based Adversarial Attacks on Encoder Attention in Detection Transformers

SafetyDGX agent

arXiv:2608.06674v1 Announce Type: new Abstract: Adversarial vulnerabilities remain a major concern for the safe deployment of neural networks, particularly in object detection, a core task embedded in

DA-Cal: Towards Cross-Domain Calibration in Semantic Segmentation

SafetyDGX agent

arXiv:2602.20860v2 Announce Type: replace Abstract: While existing unsupervised domain adaptation (UDA) methods greatly enhance target domain performance in semantic segmentation, they often neglect n

Deep Evidential Regression for Sparse Forest Height Estimation from Multimodal Satellite Imagery

Model ReleasesDGX agent

arXiv:2608.06406v1 Announce Type: new Abstract: Accurate estimation of forest height from satellite imagery is essential for applications such as carbon accounting, biodiversity monitoring, and ecosys

Degradation-Aware Prompt Learning with Cross-Modal Compensation for Adverse Weather Removal

SafetyDGX agent

arXiv:2608.06939v1 Announce Type: new Abstract: Adverse weather causes diverse and complex image degradations, severely compromising the reliability of computer vision systems. Existing all-in-one res

Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model

ResearchDGX agent

arXiv:2608.07361v1 Announce Type: cross Abstract: Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself

Direct Visual Grounding by Directing Attention of Visual Tokens

ApplicationsDGX agent

arXiv:2511.12738v2 Announce Type: replace Abstract: Vision Language Models (VLMs) mix visual tokens and text tokens. A puzzling issue is the fact that visual tokens most related to the query receive l

DREAMS: Diverse Reactions of Engagement and Attention Mind States Dataset

ResearchDGX agent

arXiv:2608.06382v1 Announce Type: cross Abstract: Active attention and engagement are important in improving users' learning experiences. Engagement refers to the level of involvement and interest ind

Dual-Space Modality Consistency Learning for Universal Cross-Modal Re-Identification

SafetyDGX agent

arXiv:2608.06943v1 Announce Type: new Abstract: Cross-modal Re-Identification (ReID) aims to retrieve the same identity across heterogeneous imaging modalities and has been widely studied in visible-i

ECAD: Expanding Class-Agnostic Detection Beyond Thing-Centric Objectness

Model ReleasesDGX agent

arXiv:2608.06841v1 Announce Type: new Abstract: Object detection is a fundamental task in visual perception, providing structured region representations for recognition, grounding, reasoning, and inte

ELMZip: Onboard Satellite Image Compression via Extreme Learning Machines for Efficient Downlink

ResearchDGX agent

arXiv:2608.06942v1 Announce Type: cross Abstract: The acquisition of multispectral imagery via small satellites (e.g., CubeSats) presents significant data downlink challenges due to high data volumes

Explanation Stability of Test-Time Adaptation in Computational Pathology: A Large-Scale Benchmark

Model ReleasesDGX agent

arXiv:2608.07062v1 Announce Type: new Abstract: Test-time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in comput

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models

SafetyDGX agent

arXiv:2608.06768v1 Announce Type: new Abstract: Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distr

Flow-Corrected Shape Optimization: Taming Manifold Drift in High-Dimensional 3D Models

ResearchDGX agent

arXiv:2608.07199v1 Announce Type: new Abstract: Optimizing 3D shapes within the latent spaces of deep generative models is fundamental to computer assisted engineering, yet remains prone to a critical

Foundation Models Adaptation for Multi-View Multi-modal Cardiac MRI Segmentation and Direct Ejection Fraction Estimation

ResearchDGX agent

arXiv:2608.07291v1 Announce Type: new Abstract: Foundation models have shown strong transferability in cardiac MRI (CMR), but their effectiveness for heterogeneous multi-view and multi-sequence CMR an

Generative Embedding Benchmark: How Much Information Survives in a Dense Embedding?

Model ReleasesDGX agent

arXiv:2608.06972v1 Announce Type: new Abstract: Embeddings have emerged as a standard representational interface linking foundation models with downstream systems. Most embedding benchmarks assess rep

GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes

SafetyDGX agent

arXiv:2608.06836v1 Announce Type: new Abstract: We study the problem of inserting new furniture into indoor scene images. Under masked single-view 2D image-plane conditioning, however, the physical sc

Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP

SafetyDGX agent

arXiv:2502.18816v3 Announce Type: replace Abstract: Significant progress has been achieved on the improvement and downstream usages of the Contrastive Language-Image Pre-training (CLIP) vision-languag

GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.06769v1 Announce Type: new Abstract: Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision-language tasks, creating an urgent need for more

← Previous
1…56789…207
Next →