AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
2 Jun 2026

Rank-Aware Quantile Activation for Motion-Robust Crop Segmentation in UAV Imagery

ResearchDGX agent

arXiv:2606.01118v1 Announce Type: new Abstract: Motion blur from high-speed UAV acquisition de-grades semantic segmentation on rare texture-dependent classes with high agronomic value. Standard CNNs r

RankByGene: Gene-Guided Histopathology Representation Learning Through Cross-Modal Ranking Consistency

SafetyDGX agent

arXiv:2411.15076v3 Announce Type: replace-cross Abstract: Spatial transcriptomics (ST) provides essential spatial context by mapping gene expression within tissue, enabling detailed study of cellular

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.01620v1 Announce Type: new Abstract: Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applic

Real-Time Physics Simulation with Dynamic Mesh-Gaussian Reconstructions

ResearchDGX agent

arXiv:2606.00444v1 Announce Type: new Abstract: Integrating dynamic 3D reconstructions into physics simulation requires fixed mesh topology for efficient collision detection, but state-of-the-art meth

Reason, Retrieve, Re-rank: A Zero-Shot Reasoning-Aware Framework for Composed Video Retrieval

ResearchDGX agent

arXiv:2606.00910v1 Announce Type: new Abstract: Composed Video Retrieval (CoVR) seeks the target video that results from applying a free-form textual modification to a reference video. We address the

Reason-Then-Retrieve for CoVR-R with Structured Edit Prompts and Dense-Sparse Fusion

ResearchDGX agent

arXiv:2606.02450v1 Announce Type: new Abstract: CoVR-R studies reason-aware composed video retrieval: given a reference video and an edit instruction, the system must retrieve the target video that sa

Recursive Vision Transformer with Dynamic Depth and Width Adjustment for Resource-Efficient Image Semantic Communication

Model ReleasesDGX agent

arXiv:2606.00114v1 Announce Type: new Abstract: Image semantic communication is a critical component in next-generation wireless communication systems. However, such systems typically suffer from larg

Relative Energy Learning for LiDAR Out-of-Distribution Detection

Model ReleasesDGX agent

arXiv:2511.06720v3 Announce Type: replace Abstract: Out-of-distribution (OOD) detection is a critical requirement for reliable autonomous driving, where safety depends on recognizing road obstacles an

RESBev: Making BEV Perception More Robust

SafetyDGX agent

arXiv:2603.09529v2 Announce Type: replace Abstract: Bird's-eye-view (BEV) perception has emerged as a cornerstone of autonomous driving systems, providing a structured, ego-centric representation crit

RescueBench: Can Embodied Agents Save Lives in the Wild ?

Model ReleasesDGX agent

arXiv:2606.01848v1 Announce Type: new Abstract: Search-and-rescue (SAR) requires embodied agents to explore unfamiliar environments under multimodal uncertainty, perform multi-stage interactions, and

Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering

Model ReleasesDGX agent

arXiv:2606.01911v1 Announce Type: new Abstract: Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong ove

Response-Aware Multimodal Learning for Post-Treatment Visual Acuity Forecasting

Local AiDGX agent

arXiv:2606.00588v1 Announce Type: new Abstract: Long-term visual acuity (VA) outcomes after anti-VEGF therapy are central to patient counseling, expectation setting, and follow-up planning in diabetic

Restoring Initial Noise Sensitivity in Text-to-Image Distillation via Geometric Alignment

SafetyDGX agent

arXiv:2606.01651v1 Announce Type: new Abstract: Generative distillation significantly accelerates text-to-image (T2I) generation by compressing multi-step trajectories into few-step student models whi

Rethinking Amortized Neural Representations for High-Resolution Terrain Elevation Data

Model ReleasesDGX agent

arXiv:2606.00404v1 Announce Type: new Abstract: Implicit neural representations (INRs) model a signal as a continuous coordinate-to-value function. For terrain elevation data, this supports analytic d

Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation

ResearchDGX agent

arXiv:2606.02479v1 Announce Type: new Abstract: Maintaining long-term geometric consistency remains challenging for long-horizon autoregressive video generation. Memory-augmented generative models add

Reusing Fusion-Time Spectral Reliability for Adaptive Fusion and Expert Routing in RGB-Infrared Object Detection

Model ReleasesDGX agent

arXiv:2606.01173v1 Announce Type: new Abstract: RGB-infrared detectors typically discard the statistics generated during cross-modal fusion, leaving downstream modules unaware of whether the current i

RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation

SafetyDGX agent

arXiv:2507.02792v5 Announce Type: replace Abstract: Text-to-image (T2I) diffusion models have shown remarkable success in generating high-quality images from text prompts. Recent efforts extend these

RoboDream: Compositional World Models for Scalable Robot Data Synthesis

SafetyDGX agent

arXiv:2606.02577v1 Announce Type: cross Abstract: Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive

RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes

Model ReleasesDGX agent

arXiv:2606.00828v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown strong visual understanding and are increasingly deployed in embodied AI systems, where reliable perception und

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search

Model ReleasesDGX agent

arXiv:2606.01825v1 Announce Type: new Abstract: Text-Based Person Search (TBPS) aims to retrieve pedestrian images using natural language queries. However, existing TBPS models, especially those based

RU4D-SLAM: Reweighting Uncertainty in Gaussian Splatting SLAM for 4D Scene Reconstruction

ResearchDGX agent

arXiv:2602.20807v2 Announce Type: replace Abstract: Combining 3D Gaussian splatting with Simultaneous Localization and Mapping (SLAM) has gained popularity as it enables continuous 3D environment reco

Safe2Drive: Evaluating Safe Driving Behaviors of E2E Autonomous Driving Models

SafetyDGX agent

arXiv:2606.00191v1 Announce Type: cross Abstract: Recent end-to-end (E2E) autonomous driving policies achieve high driving scores in closed-loop simulations. Yet it remains unclear whether these polic

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation

Model ReleasesDGX agent

arXiv:2606.01481v1 Announce Type: new Abstract: With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos fro

Saliency-Aware Model Merging

Model ReleasesDGX agent

arXiv:2606.00511v1 Announce Type: cross Abstract: Model merging aims to consolidate multiple task-specific models fine-tuned on different datasets into a unified architecture that performs cross-domai

SAVMap: Structure-Aided Visual Mapping of Large-Scale 2.5D Manhattan Wireframes from Panoramic Video

ApplicationsDGX agent

arXiv:2606.01939v1 Announce Type: new Abstract: Precise 3D representations of industrial environments enable tasks such as robot localization and digital twin generation. We propose SAVMap, a method f

Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders

Model ReleasesDGX agent

arXiv:2606.00746v1 Announce Type: new Abstract: Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale

Scaling Pre-training to One Hundred Billion Data for Vision Language Models

ResearchDGX agent

arXiv:2502.07617v2 Announce Type: replace Abstract: We provide an empirical investigation of the potential of pre-training vision-language models on an unprecedented scale: 100 billion examples. We fi

SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation

SafetyDGX agent

arXiv:2606.01940v1 Announce Type: new Abstract: Existing methods for category-level object articulation from a single 3D observation often rely on dense supervision, multi-frame inputs, or CAD templat

SCL: Towards Domain Generalization via Single-Temporal Multimodal Contrastive Learning for Remote Sensing Change Detection

TutorialsDGX agent

arXiv:2404.11326v5 Announce Type: replace Abstract: In recent years, change detection and anomaly detection models based on CNN and transformer have achieved remarkable success across various datasets

Score-Control for Hallucination Reduction in Diffusion Models

Model ReleasesDGX agent

arXiv:2606.00377v1 Announce Type: new Abstract: Diffusion models have emerged as the backbone of modern generative AI, powering advances in vision, language, audio and other modalities. Despite their

See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation

Model ReleasesDGX agent

arXiv:2603.09292v2 Announce Type: replace-cross Abstract: Measurement of task progress through explicit, actionable milestones is critical for robust robotic manipulation. This progress awareness enab

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Model ReleasesDGX agent

arXiv:2503.06520v3 Announce Type: replace Abstract: Traditional methods for reasoning segmentation rely on supervised fine-tuning with categorical labels and simple descriptions, limiting its out-of-d

Segmentation-Guided Spatial Indexing for Generalizable and Explainable Deepfake Detection

ResearchDGX agent

arXiv:2606.00098v1 Announce Type: new Abstract: We introduce segmentation-guided spatial indexing for generalizable and explainable deepfake detection. The key idea reverses the standard design order:

Self-Improving Small Object Grounding in LVLMs

ResearchDGX agent

arXiv:2606.01612v1 Announce Type: new Abstract: Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning? In this work, we provi

Semimage: HSV-Based Semantic Image Encoding for Disentangled Text Representation

ResearchDGX agent

arXiv:2512.00088v2 Announce Type: replace Abstract: We propose SemImage, a novel method for representing a text document as a two-dimensional semantic image to be processed by convolutional neural net

Sensitivity as a Double-Edged Sword: A Trade-off Between Discriminability and Adversarial Robustness

ResearchDGX agent

arXiv:2606.01746v1 Announce Type: new Abstract: Modern neural networks are highly susceptible to adversarial perturbations. In this work, we identify that part of this vulnerability stems from the sen

Seq-DeepIPC: Sequential Sensing for End-to-End Control in Legged Robot Navigation

AgentsDGX agent

arXiv:2510.23057v2 Announce Type: replace-cross Abstract: We present Seq-DeepIPC, a sequential end-to-end perception-to-control model for legged robot navigation in real-world environments. Seq-DeepIP

Shape-Prior-Based Point Cloud Completion for Single-Stage Fully Sparse 3D Object Detection

SafetyDGX agent

arXiv:2606.00688v1 Announce Type: new Abstract: Single-stage fully sparse 3D object detectors rely on point clouds data to detect objects in autonomous driving scenarios. However, the sparsity and inc

Shu Dao: A Calligraphy Score Framework Linking Calligraphy, Music, and Performance

ResearchDGX agent

arXiv:2606.00001v1 Announce Type: cross Abstract: This paper introduces Calligraphy Writing Score Representation (CWSR) and proposes Shu Dao as a framework that interprets East Asian calligraphy as a

Single-Channel Tissue Segmentation via Cross-Modal Distillation from Foundation Models

Model ReleasesDGX agent

arXiv:2606.00928v1 Announce Type: new Abstract: Multiplexed fluorescence microscopy improves tissue segmentation by providing complementary channels including nuclear (DAPI) and membrane (E-cadherin),

Single-Line Drawing Generation via Semantics-Driven Optimization

ResearchDGX agent

arXiv:2606.01910v1 Announce Type: cross Abstract: Line drawings are a highly expressive art form that requires the artist to abstract and distill the essence of their subject. We present the first sem

SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models

SafetyDGX agent

arXiv:2606.00664v1 Announce Type: cross Abstract: Embodied world models have emerged as a promising paradigm in robotics by predicting how robot actions affect the surrounding scene. However, the roll

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

AgentsDGX agent

arXiv:2512.04069v2 Announce Type: replace Abstract: Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required f

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation

SafetyDGX agent

arXiv:2606.02441v1 Announce Type: new Abstract: Identity-preserving video generation (IPVG) aims to synthesize high-fidelity videos that follow text prompts while faithfully preserving a reference ide

Spatio-Temporal Correlation Guided Geometric Partitioning for Versatile Video Coding

TutorialsDGX agent

arXiv:2606.01701v1 Announce Type: new Abstract: Geometric partitioning has attracted increasing attention by its remarkable motion field description capability in the hybrid video coding framework. Ho

SpikeReg: Energy-Efficient 3D Deformable Medical Image Registration with Spiking Neural Networks

ResearchDGX agent

arXiv:2605.25144v2 Announce Type: replace Abstract: Deformable medical image registration aligns anatomical structures across images but remains computationally dense at 3D resolution. Spiking neural

Splatshot: 3D Face Avatar Generation from a Single Unconstrained Photo

ResearchDGX agent

arXiv:2606.01493v1 Announce Type: new Abstract: Reconstructing a photorealistic 3D face avatar from a single unconstrained photograph is challenging: feed-forward 3D Gaussian Splatting (3DGS) models d

Stable Velocity: A Variance Perspective on Flow Matching

Model ReleasesDGX agent

arXiv:2602.05435v2 Announce Type: replace Abstract: While flow matching is elegant, its reliance on single-sample conditional velocities leads to high-variance training targets that destabilize optimi

Structure-Aware Consistency Priors for Shape from Polarization in Complex Media

ApplicationsDGX agent

arXiv:2606.00509v1 Announce Type: new Abstract: Recovering surface normals from single view polarization images in complex media remains challenging. This paper focuses on ice as a representative comp

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

Model ReleasesDGX agent

arXiv:2606.00825v1 Announce Type: new Abstract: AI glasses present a compelling platform for AI agents to serve as personalized memory assistants. To be genuinely useful, such systems must move beyond

SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models

SafetyDGX agent

arXiv:2601.22276v2 Announce Type: replace-cross Abstract: As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributor

SWARD: Stochastic Window-Attention-Based Relational Distillation for Cross-Architectural Semantic Segmentation

SafetyDGX agent

arXiv:2606.00999v1 Announce Type: new Abstract: Large-scale vision foundation models have driven substantial gains on dense prediction tasks such as semantic segmentation, but their size makes deploym

Symmetry-Aware 9D Pose Estimation with Sim(3)-Consistent Feature and Spherical Inception Convolution

AgentsDGX agent

arXiv:2606.02219v1 Announce Type: new Abstract: Object pose estimation is a fundamental problem for an agent system to perceive or manipulate objects in images or videos. However, current instance-lev

T-CLIP: Enabling Thermal Perception for Contrastive Language-Image Pretraining

ResearchDGX agent

arXiv:2606.00673v1 Announce Type: new Abstract: Thermal imaging offers a powerful alternative to visible-spectrum vision under challenging conditions such as low illumination and adverse weather, yet

TAP-JEPA: Frozen Future-Latent Probing and Two-Stage Score Fusion for EPIC-KITCHENS-100 Action Anticipation

ResearchDGX agent

arXiv:2606.00662v1 Announce Type: new Abstract: This report presents TAP-JEPA, our runner-up submission to the EPIC-KITCHENS-100 (EK-100) Action Anticipation Challenge at EgoVis 2026. The task is to a

Tempora: Characterising the Time-Contingent Utility of Online Test-Time Adaptation

ResearchDGX agent

arXiv:2602.06136v2 Announce Type: replace-cross Abstract: Test-time adaptation (TTA) offers a compelling remedy for machine learning (ML) models that degrade under domain shifts, improving generalisat

Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA

ResearchDGX agent

arXiv:2606.01106v1 Announce Type: new Abstract: TimeLogicQA evaluates whether video question answering systems can reason over temporal relations such as event existence, ordering, persistence, bounda

TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge

Model ReleasesDGX agent

arXiv:2605.24470v2 Announce Type: replace Abstract: Video-text retrieval has witnessed remarkable progress driven by large-scale vision-language pretraining, yet most existing approaches inherit an im

TextFake: Benchmarking AI-Generated Image Detection on Text-Rich Images

Model ReleasesDGX agent

arXiv:2606.01050v1 Announce Type: new Abstract: Recent AI-generated image (AIGI) detectors perform well on natural-image benchmarks, but their behavior on text-rich forgeries, such as fabricated scree

The Harsh Truth: Segment-Level Analysis of Harsh Driving Events in Milan Using Large-Scale Telematics, Street Networks, and Google Street View

SafetyDGX agent

arXiv:2606.00261v1 Announce Type: new Abstract: Police-reported crash statistics remain the standard input for urban road-safety assessment, but their incompleteness and reporting lag limit their usef

← Previous
1…103104105106107…211
Next →