AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Applications

Structure-Aware Consistency Priors for Shape from Polarization in Complex Media

DGX agent

arXiv:2606.00509v1 Announce Type: new Abstract: Recovering surface normals from single view polarization images in complex media remains challenging. This paper focuses on ice as a representative comp

applicationsarxiv-cs-cv
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

DGX agent

arXiv:2606.00825v1 Announce Type: new Abstract: AI glasses present a compelling platform for AI agents to serve as personalized memory assistants. To be genuinely useful, such systems must move beyond

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models

DGX agent

arXiv:2601.22276v2 Announce Type: replace-cross Abstract: As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributor

safetyarxiv-cs-cv
2 Jun 2026
Safety

SWARD: Stochastic Window-Attention-Based Relational Distillation for Cross-Architectural Semantic Segmentation

DGX agent

arXiv:2606.00999v1 Announce Type: new Abstract: Large-scale vision foundation models have driven substantial gains on dense prediction tasks such as semantic segmentation, but their size makes deploym

safetyarxiv-cs-cv
2 Jun 2026
Agents

Symmetry-Aware 9D Pose Estimation with Sim(3)-Consistent Feature and Spherical Inception Convolution

DGX agent

arXiv:2606.02219v1 Announce Type: new Abstract: Object pose estimation is a fundamental problem for an agent system to perceive or manipulate objects in images or videos. However, current instance-lev

agentsarxiv-cs-cv
2 Jun 2026
Research

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining

DGX agent

arXiv:2606.00673v1 Announce Type: new Abstract: Thermal imaging offers a powerful alternative to visible-spectrum vision under challenging conditions such as low illumination and adverse weather, yet

researcharxiv-cs-cv
2 Jun 2026
Research

TAP-JEPA: Frozen Future-Latent Probing and Two-Stage Score Fusion for EPIC-KITCHENS-100 Action Anticipation

DGX agent

arXiv:2606.00662v1 Announce Type: new Abstract: This report presents TAP-JEPA, our runner-up submission to the EPIC-KITCHENS-100 (EK-100) Action Anticipation Challenge at EgoVis 2026. The task is to a

researcharxiv-cs-cv
2 Jun 2026
Research

Tempora: Characterising the Time-Contingent Utility of Online Test-Time Adaptation

DGX agent

arXiv:2602.06136v2 Announce Type: replace-cross Abstract: Test-time adaptation (TTA) offers a compelling remedy for machine learning (ML) models that degrade under domain shifts, improving generalisat

researcharxiv-cs-cv
2 Jun 2026
Research

Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA

DGX agent

arXiv:2606.01106v1 Announce Type: new Abstract: TimeLogicQA evaluates whether video question answering systems can reason over temporal relations such as event existence, ordering, persistence, bounda

researcharxiv-cs-cv
2 Jun 2026
Model Releases

TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge

DGX agent

arXiv:2605.24470v2 Announce Type: replace Abstract: Video-text retrieval has witnessed remarkable progress driven by large-scale vision-language pretraining, yet most existing approaches inherit an im

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

TextFake: Benchmarking AI-Generated Image Detection on Text-Rich Images

DGX agent

arXiv:2606.01050v1 Announce Type: new Abstract: Recent AI-generated image (AIGI) detectors perform well on natural-image benchmarks, but their behavior on text-rich forgeries, such as fabricated scree

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

The Harsh Truth: Segment-Level Analysis of Harsh Driving Events in Milan Using Large-Scale Telematics, Street Networks, and Google Street View

DGX agent

arXiv:2606.00261v1 Announce Type: new Abstract: Police-reported crash statistics remain the standard input for urban road-safety assessment, but their incompleteness and reporting lag limit their usef

safetyarxiv-cs-cv
2 Jun 2026
Research

The Right Inference Strategy Is All You Need: Nearly Training-Free Domain-Wise Inference for EgoCross Challenge

DGX agent

arXiv:2606.00829v1 Announce Type: new Abstract: EgoCross evaluates multimodal large language models on egocentric video question answering under substantial domain shift, where test videos come from s

researcharxiv-cs-cv
2 Jun 2026
Agents

Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models

DGX agent

arXiv:2606.02580v1 Announce Type: new Abstract: Inverse graphics is a longstanding and highly underconstrained problem that seeks to reconstruct images as editable 3D scenes which can be rendered, rel

agentsarxiv-cs-cv
2 Jun 2026
Model Releases

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs

DGX agent

arXiv:2606.00592v1 Announce Type: new Abstract: Effective visual communication stems from the harmony of multiple design principles, such as readability, contrast, alignment, overlap, and coherence, w

model-releasesarxiv-cs-cv
2 Jun 2026
Tutorials

TIDES: Time-Derivative Event Simulation via Deformable Reconstruction

DGX agent

arXiv:2606.02058v1 Announce Type: new Abstract: Event cameras emit asynchronous events in response to environmental appearance changes. The scarcity of real-world event datasets makes simulation essen

tutorialsarxiv-cs-cv
2 Jun 2026
Model Releases

TLG: Temporal-Logic Grounding for Video Question Answering via Source-Annotation Reconstruction and Category-Targeted Reasoning

DGX agent

arXiv:2606.01591v1 Announce Type: new Abstract: The TimeLogic Challenge evaluates formal temporal-logic reasoning over video - 16 operators (before, after, until, since, always, co-occur, ordering, ..

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

ToolFG: Towards Well-Grounded Fine-Grained Image Classification

DGX agent

arXiv:2606.02518v1 Announce Type: new Abstract: Fine-grained image classification (FGIC) has broad applications and has attracted significant research attention. In this paper, we explore a novel para

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification

DGX agent

arXiv:2509.16635v2 Announce Type: replace Abstract: In real applications, person re-identification (ReID) is expected to retrieve the target person at any time, including both daytime and nighttime, r

model-releasesarxiv-cs-cv
2 Jun 2026
Agents

Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends

DGX agent

arXiv:2606.01164v1 Announce Type: new Abstract: With rapid development of large language models and diffusion-based content generation, world modeling has attracted increasing research attention, bene

agentsarxiv-cs-cv
2 Jun 2026
Agents

Towards Sparse Video Understanding and Reasoning

DGX agent

arXiv:2602.13602v2 Announce Type: replace Abstract: We present revise (nderline{Re}asoning with nderline{Vi}deo nderline{S}parsity), a multi-round agent for video question answering (VQA). Instead of

agentsarxiv-cs-cv
2 Jun 2026
Research

Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning

DGX agent

arXiv:2606.02321v1 Announce Type: new Abstract: Recent advances in large vision-language models have expanded video retrieval from simple text-based search to more flexible scenarios, where users may

researcharxiv-cs-cv
2 Jun 2026
Applications

Training-Free Continuous Bitrate Control for Scalable Image Coding for Humans and Machines

DGX agent

arXiv:2606.00158v1 Announce Type: cross Abstract: Continuous variable-rate compression is highly demanded in real-world applications, but remains underexplored in scalable image coding for humans and

applicationsarxiv-cs-cv
2 Jun 2026
Research

Training-Free Coverless Multi-Image Steganography with Access Control

DGX agent

arXiv:2603.09390v2 Announce Type: replace Abstract: Coverless Image Steganography (CIS) hides information without explicitly modifying a cover image, providing strong imperceptibility and inherent rob

researcharxiv-cs-cv
2 Jun 2026
Safety

Training-free image inversion for one-step diffusion models

DGX agent

arXiv:2606.01380v1 Announce Type: new Abstract: In this work, we introduce a novel training-free inversion (TFinv) framework for one-step diffusion models,addressing key challenges in real image inver

safetyarxiv-cs-cv
2 Jun 2026
Research

Training-Free Object-Agnostic Jam Detection in Fulfillment Centers

DGX agent

arXiv:2606.00321v1 Announce Type: new Abstract: In fulfillment centers, diverse objects move continuously from inbound to outbound operations and can become jammed due to excessive conveyor friction,

researcharxiv-cs-cv
2 Jun 2026
Safety

TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos

DGX agent

arXiv:2606.02350v1 Announce Type: new Abstract: Reconstructing humans and their surrounding environments in a globally consistent 4D space is essential for comprehensive perception. However, prior wor

safetyarxiv-cs-cv
2 Jun 2026
Safety

Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval

DGX agent

arXiv:2606.01615v1 Announce Type: new Abstract: Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear

safetyarxiv-cs-cv
2 Jun 2026
Research

Two Datasets Are Better Than One: Method of Double Moments for 3-D Reconstruction in Cryo-EM

DGX agent

arXiv:2511.07438v3 Announce Type: replace Abstract: Cryo-electron microscopy (cryo-EM) is a powerful imaging technique for reconstructing three-dimensional molecular structures from noisy tomographic

researcharxiv-cs-cv
2 Jun 2026
Safety

Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking From Sparse Inertial Sensors and Ranging-Based Between-Sensor Distances

DGX agent

arXiv:2606.02153v1 Announce Type: new Abstract: Methods using inertial measurement units (IMUs) provide a wearable alternative to camera-based motion capture. To mitigate drift from inertial signals,

safetyarxiv-cs-cv
2 Jun 2026
Agents

Unified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning

DGX agent

arXiv:2606.01935v1 Announce Type: new Abstract: Discrete visual tokens should provide a compact representation for both token-based world modeling and planning in autonomous driving. However, most tok

agentsarxiv-cs-cv
2 Jun 2026
Applications

Unified Semantic Transformer for 3D Scene Understanding

DGX agent

arXiv:2512.14364v3 Announce Type: replace Abstract: Holistic 3D scene understanding involves capturing and parsing unstructured 3D environments. Due to the inherent complexity of the real world, exist

applicationsarxiv-cs-cv
2 Jun 2026
Research

UniVerse: A Unified Modulation Framework for Segmentation-Free,Disentangled Multi-Concept Personalization

DGX agent

arXiv:2606.00351v1 Announce Type: new Abstract: Personalized visual understanding has advanced significantly, yet existing approaches struggle to localize and extract specific concepts when input imag

researcharxiv-cs-cv
2 Jun 2026
Agents

Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing

DGX agent

arXiv:2606.01818v1 Announce Type: new Abstract: Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting

agentsarxiv-cs-cv
2 Jun 2026
Research

UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial Representations

DGX agent

arXiv:2510.13774v2 Announce Type: replace-cross Abstract: Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data.

researcharxiv-cs-cv
2 Jun 2026
Research

VEDAL: Variational Error-Driven Asynchronous Learning for 3D Gaussian Splatting Pruning

DGX agent

arXiv:2606.02346v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) achieves remarkable novel view synthesis quality with real-time rendering, yet suffers from excessive memory consumption du

researcharxiv-cs-cv
2 Jun 2026
Local Ai

VICR: Visual In-Context Restoration for Real-World Image Super-Resolution

DGX agent

arXiv:2606.00704v1 Announce Type: new Abstract: Real-world image super-resolution (Real-ISR) requires balancing structural fidelity to degraded observations with realistic detail synthesis. However, e

local-aiarxiv-cs-cv
2 Jun 2026
Agents

VideoBrain: Learning Adaptive Frame Sampling for Long Video Understanding

DGX agent

arXiv:2602.04094v2 Announce Type: replace Abstract: Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints an

agentsarxiv-cs-cv
2 Jun 2026
Research

Visible Light Positioning With Lame Curve LEDs: A Generic Approach for Camera Pose Estimation

DGX agent

arXiv:2602.01577v3 Announce Type: replace-cross Abstract: Camera-based visible light positioning (VLP) is a promising technique for accurate and low-cost indoor camera pose estimation (CPE). To reduce

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset

DGX agent

arXiv:2606.02273v1 Announce Type: new Abstract: Understanding subtle driver actions is essential for building reliable driver monitoring systems. Existing visionlanguage models (VLMs) are trained on g

model-releasesarxiv-cs-cv
2 Jun 2026
Research

VISReg: Variance-Invariance-Sketching Regularization for JEPA training

DGX agent

arXiv:2606.02572v1 Announce Type: new Abstract: Self-supervised learning methods prevent embedding collapse via modeling heuristics or explicit regularization of the embedding space. Among the latter,

researcharxiv-cs-cv
2 Jun 2026
Safety

Visualizing definitional divergence in high-dimensional data by manifold alignment: Application to 3D right ventricular strain computations

DGX agent

arXiv:2501.12178v2 Announce Type: replace Abstract: Medical imaging studies often rely on a single sample per subject, assuming it is representative of their physiological traits. However, variations

safetyarxiv-cs-cv
2 Jun 2026
Research

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

DGX agent

arXiv:2606.02564v1 Announce Type: new Abstract: The recent 'Reasoning with Video' paradigm utilizes Video Generation Models (VGMs) to generate temporally coherent visual trajectories to complete reaso

researcharxiv-cs-cv
2 Jun 2026
Applications

WALL-WM: Carving World Action Modeling at the Event Joints

DGX agent

arXiv:2606.01955v1 Announce Type: cross Abstract: WALL-WM is a World Action Model that shifts video-action learning from chunk-centric optimization to event-grounded Vision-Language-Action pretraining

applicationsarxiv-cs-cv
2 Jun 2026
Safety

Wavelet-Fusion Diffusion Model for Multimodal Brain MRI Synthesis with Modality and Metadata Conditioning

DGX agent

arXiv:2606.00689v1 Announce Type: new Abstract: Multimodal MRI provides complementary information for neuroimaging analysis, where different imaging modalities capture distinct anatomical, tissue, and

safetyarxiv-cs-cv
2 Jun 2026
Local Ai

WebSpline: Structure-Informed Splines for Real-Time 3D Gaussians from Monocular Videos

DGX agent

arXiv:2606.02096v1 Announce Type: new Abstract: Dynamic scene reconstruction from monocular videos remains highly challenging, as existing methods often struggle to balance global structural coherence

local-aiarxiv-cs-cv
2 Jun 2026
Safety

What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs

DGX agent

arXiv:2606.01624v1 Announce Type: new Abstract: Driving vision-language models (VLMs) must accurately understand scenes across diverse conditions defined by Operational Design Domains (ODDs), yet veri

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

DGX agent

arXiv:2606.01247v1 Announce Type: new Abstract: Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has la

model-releasesarxiv-cs-cv
2 Jun 2026
← Previous
1…130131132133134…263
Next →