AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Rank-Aware Quantile Activation for Motion-Robust Crop Segmentation in UAV Imagery

DGX agent

arXiv:2606.01118v1 Announce Type: new Abstract: Motion blur from high-speed UAV acquisition de-grades semantic segmentation on rare texture-dependent classes with high agronomic value. Standard CNNs r

researcharxiv-cs-cv
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

RankByGene: Gene-Guided Histopathology Representation Learning Through Cross-Modal Ranking Consistency

DGX agent

arXiv:2411.15076v3 Announce Type: replace-cross Abstract: Spatial transcriptomics (ST) provides essential spatial context by mapping gene expression within tissue, enabling detailed study of cellular

safetyarxiv-cs-cv
2 Jun 2026
Research

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

DGX agent

arXiv:2606.01620v1 Announce Type: new Abstract: Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applic

researcharxiv-cs-cv
2 Jun 2026
Research

Real-Time Physics Simulation with Dynamic Mesh-Gaussian Reconstructions

DGX agent

arXiv:2606.00444v1 Announce Type: new Abstract: Integrating dynamic 3D reconstructions into physics simulation requires fixed mesh topology for efficient collision detection, but state-of-the-art meth

researcharxiv-cs-cv
2 Jun 2026
Research

Reason, Retrieve, Re-rank: A Zero-Shot Reasoning-Aware Framework for Composed Video Retrieval

DGX agent

arXiv:2606.00910v1 Announce Type: new Abstract: Composed Video Retrieval (CoVR) seeks the target video that results from applying a free-form textual modification to a reference video. We address the

researcharxiv-cs-cv
2 Jun 2026
Research

Reason-Then-Retrieve for CoVR-R with Structured Edit Prompts and Dense-Sparse Fusion

DGX agent

arXiv:2606.02450v1 Announce Type: new Abstract: CoVR-R studies reason-aware composed video retrieval: given a reference video and an edit instruction, the system must retrieve the target video that sa

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Recursive Vision Transformer with Dynamic Depth and Width Adjustment for Resource-Efficient Image Semantic Communication

DGX agent

arXiv:2606.00114v1 Announce Type: new Abstract: Image semantic communication is a critical component in next-generation wireless communication systems. However, such systems typically suffer from larg

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Relative Energy Learning for LiDAR Out-of-Distribution Detection

DGX agent

arXiv:2511.06720v3 Announce Type: replace Abstract: Out-of-distribution (OOD) detection is a critical requirement for reliable autonomous driving, where safety depends on recognizing road obstacles an

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

RESBev: Making BEV Perception More Robust

DGX agent

arXiv:2603.09529v2 Announce Type: replace Abstract: Bird's-eye-view (BEV) perception has emerged as a cornerstone of autonomous driving systems, providing a structured, ego-centric representation crit

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

RescueBench: Can Embodied Agents Save Lives in the Wild ?

DGX agent

arXiv:2606.01848v1 Announce Type: new Abstract: Search-and-rescue (SAR) requires embodied agents to explore unfamiliar environments under multimodal uncertainty, perform multi-stage interactions, and

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text Rendering

DGX agent

arXiv:2606.01911v1 Announce Type: new Abstract: Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong ove

model-releasesarxiv-cs-cv
2 Jun 2026
Local Ai

Response-Aware Multimodal Learning for Post-Treatment Visual Acuity Forecasting

DGX agent

arXiv:2606.00588v1 Announce Type: new Abstract: Long-term visual acuity (VA) outcomes after anti-VEGF therapy are central to patient counseling, expectation setting, and follow-up planning in diabetic

local-aiarxiv-cs-cv
2 Jun 2026
Safety

Restoring Initial Noise Sensitivity in Text-to-Image Distillation via Geometric Alignment

DGX agent

arXiv:2606.01651v1 Announce Type: new Abstract: Generative distillation significantly accelerates text-to-image (T2I) generation by compressing multi-step trajectories into few-step student models whi

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

Rethinking Amortized Neural Representations for High-Resolution Terrain Elevation Data

DGX agent

arXiv:2606.00404v1 Announce Type: new Abstract: Implicit neural representations (INRs) model a signal as a continuous coordinate-to-value function. For terrain elevation data, this supports analytic d

model-releasesarxiv-cs-cv
2 Jun 2026
Research

Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation

DGX agent

arXiv:2606.02479v1 Announce Type: new Abstract: Maintaining long-term geometric consistency remains challenging for long-horizon autoregressive video generation. Memory-augmented generative models add

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Reusing Fusion-Time Spectral Reliability for Adaptive Fusion and Expert Routing in RGB-Infrared Object Detection

DGX agent

arXiv:2606.01173v1 Announce Type: new Abstract: RGB-infrared detectors typically discard the statistics generated during cross-modal fusion, leaving downstream modules unaware of whether the current i

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation

DGX agent

arXiv:2507.02792v5 Announce Type: replace Abstract: Text-to-image (T2I) diffusion models have shown remarkable success in generating high-quality images from text prompts. Recent efforts extend these

safetyarxiv-cs-cv
2 Jun 2026
Safety

RoboDream: Compositional World Models for Scalable Robot Data Synthesis

DGX agent

arXiv:2606.02577v1 Announce Type: cross Abstract: Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes

DGX agent

arXiv:2606.00828v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown strong visual understanding and are increasingly deployed in embodied AI systems, where reliable perception und

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search

DGX agent

arXiv:2606.01825v1 Announce Type: new Abstract: Text-Based Person Search (TBPS) aims to retrieve pedestrian images using natural language queries. However, existing TBPS models, especially those based

model-releasesarxiv-cs-cv
2 Jun 2026
Research

RU4D-SLAM: Reweighting Uncertainty in Gaussian Splatting SLAM for 4D Scene Reconstruction

DGX agent

arXiv:2602.20807v2 Announce Type: replace Abstract: Combining 3D Gaussian splatting with Simultaneous Localization and Mapping (SLAM) has gained popularity as it enables continuous 3D environment reco

researcharxiv-cs-cv
2 Jun 2026
Safety

Safe2Drive: Evaluating Safe Driving Behaviors of E2E Autonomous Driving Models

DGX agent

arXiv:2606.00191v1 Announce Type: cross Abstract: Recent end-to-end (E2E) autonomous driving policies achieve high driving scores in closed-loop simulations. Yet it remains unclear whether these polic

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation

DGX agent

arXiv:2606.01481v1 Announce Type: new Abstract: With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos fro

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Saliency-Aware Model Merging

DGX agent

arXiv:2606.00511v1 Announce Type: cross Abstract: Model merging aims to consolidate multiple task-specific models fine-tuned on different datasets into a unified architecture that performs cross-domai

model-releasesarxiv-cs-cv
2 Jun 2026
Applications

SAVMap: Structure-Aided Visual Mapping of Large-Scale 2.5D Manhattan Wireframes from Panoramic Video

DGX agent

arXiv:2606.01939v1 Announce Type: new Abstract: Precise 3D representations of industrial environments enable tasks such as robot localization and digital twin generation. We propose SAVMap, a method f

applicationsarxiv-cs-cv
2 Jun 2026
Model Releases

Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders

DGX agent

arXiv:2606.00746v1 Announce Type: new Abstract: Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale

model-releasesarxiv-cs-cv
2 Jun 2026
Research

Scaling Pre-training to One Hundred Billion Data for Vision Language Models

DGX agent

arXiv:2502.07617v2 Announce Type: replace Abstract: We provide an empirical investigation of the potential of pre-training vision-language models on an unprecedented scale: 100 billion examples. We fi

researcharxiv-cs-cv
2 Jun 2026
Safety

SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation

DGX agent

arXiv:2606.01940v1 Announce Type: new Abstract: Existing methods for category-level object articulation from a single 3D observation often rely on dense supervision, multi-frame inputs, or CAD templat

safetyarxiv-cs-cv
2 Jun 2026
Tutorials

SCL: Towards Domain Generalization via Single-Temporal Multimodal Contrastive Learning for Remote Sensing Change Detection

DGX agent

arXiv:2404.11326v5 Announce Type: replace Abstract: In recent years, change detection and anomaly detection models based on CNN and transformer have achieved remarkable success across various datasets

tutorialsarxiv-cs-cv
2 Jun 2026
Model Releases

Score-Control for Hallucination Reduction in Diffusion Models

DGX agent

arXiv:2606.00377v1 Announce Type: new Abstract: Diffusion models have emerged as the backbone of modern generative AI, powering advances in vision, language, audio and other modalities. Despite their

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation

DGX agent

arXiv:2603.09292v2 Announce Type: replace-cross Abstract: Measurement of task progress through explicit, actionable milestones is critical for robust robotic manipulation. This progress awareness enab

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

DGX agent

arXiv:2503.06520v3 Announce Type: replace Abstract: Traditional methods for reasoning segmentation rely on supervised fine-tuning with categorical labels and simple descriptions, limiting its out-of-d

model-releasesarxiv-cs-cv
2 Jun 2026
Research

Segmentation-Guided Spatial Indexing for Generalizable and Explainable Deepfake Detection

DGX agent

arXiv:2606.00098v1 Announce Type: new Abstract: We introduce segmentation-guided spatial indexing for generalizable and explainable deepfake detection. The key idea reverses the standard design order:

researcharxiv-cs-cv
2 Jun 2026
Research

Self-Improving Small Object Grounding in LVLMs

DGX agent

arXiv:2606.01612v1 Announce Type: new Abstract: Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning? In this work, we provi

researcharxiv-cs-cv
2 Jun 2026
Research

Semimage: HSV-Based Semantic Image Encoding for Disentangled Text Representation

DGX agent

arXiv:2512.00088v2 Announce Type: replace Abstract: We propose SemImage, a novel method for representing a text document as a two-dimensional semantic image to be processed by convolutional neural net

researcharxiv-cs-cv
2 Jun 2026
Research

Sensitivity as a Double-Edged Sword: A Trade-off Between Discriminability and Adversarial Robustness

DGX agent

arXiv:2606.01746v1 Announce Type: new Abstract: Modern neural networks are highly susceptible to adversarial perturbations. In this work, we identify that part of this vulnerability stems from the sen

researcharxiv-cs-cv
2 Jun 2026
Agents

Seq-DeepIPC: Sequential Sensing for End-to-End Control in Legged Robot Navigation

DGX agent

arXiv:2510.23057v2 Announce Type: replace-cross Abstract: We present Seq-DeepIPC, a sequential end-to-end perception-to-control model for legged robot navigation in real-world environments. Seq-DeepIP

agentsarxiv-cs-cv
2 Jun 2026
Safety

Shape-Prior-Based Point Cloud Completion for Single-Stage Fully Sparse 3D Object Detection

DGX agent

arXiv:2606.00688v1 Announce Type: new Abstract: Single-stage fully sparse 3D object detectors rely on point clouds data to detect objects in autonomous driving scenarios. However, the sparsity and inc

safetyarxiv-cs-cv
2 Jun 2026
Research

Shu Dao: A Calligraphy Score Framework Linking Calligraphy, Music, and Performance

DGX agent

arXiv:2606.00001v1 Announce Type: cross Abstract: This paper introduces Calligraphy Writing Score Representation (CWSR) and proposes Shu Dao as a framework that interprets East Asian calligraphy as a

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Single-Channel Tissue Segmentation via Cross-Modal Distillation from Foundation Models

DGX agent

arXiv:2606.00928v1 Announce Type: new Abstract: Multiplexed fluorescence microscopy improves tissue segmentation by providing complementary channels including nuclear (DAPI) and membrane (E-cadherin),

model-releasesarxiv-cs-cv
2 Jun 2026
Research

Single-Line Drawing Generation via Semantics-Driven Optimization

DGX agent

arXiv:2606.01910v1 Announce Type: cross Abstract: Line drawings are a highly expressive art form that requires the artist to abstract and distill the essence of their subject. We present the first sem

researcharxiv-cs-cv
2 Jun 2026
Safety

SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models

DGX agent

arXiv:2606.00664v1 Announce Type: cross Abstract: Embodied world models have emerged as a promising paradigm in robotics by predicting how robot actions affect the surrounding scene. However, the roll

safetyarxiv-cs-cv
2 Jun 2026
Agents

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

DGX agent

arXiv:2512.04069v2 Announce Type: replace Abstract: Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required f

agentsarxiv-cs-cv
2 Jun 2026
Safety

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation

DGX agent

arXiv:2606.02441v1 Announce Type: new Abstract: Identity-preserving video generation (IPVG) aims to synthesize high-fidelity videos that follow text prompts while faithfully preserving a reference ide

safetyarxiv-cs-cv
2 Jun 2026
Tutorials

Spatio-Temporal Correlation Guided Geometric Partitioning for Versatile Video Coding

DGX agent

arXiv:2606.01701v1 Announce Type: new Abstract: Geometric partitioning has attracted increasing attention by its remarkable motion field description capability in the hybrid video coding framework. Ho

tutorialsarxiv-cs-cv
2 Jun 2026
Research

SpikeReg: Energy-Efficient 3D Deformable Medical Image Registration with Spiking Neural Networks

DGX agent

arXiv:2605.25144v2 Announce Type: replace Abstract: Deformable medical image registration aligns anatomical structures across images but remains computationally dense at 3D resolution. Spiking neural

researcharxiv-cs-cv
2 Jun 2026
Research

Splatshot: 3D Face Avatar Generation from a Single Unconstrained Photo

DGX agent

arXiv:2606.01493v1 Announce Type: new Abstract: Reconstructing a photorealistic 3D face avatar from a single unconstrained photograph is challenging: feed-forward 3D Gaussian Splatting (3DGS) models d

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Stable Velocity: A Variance Perspective on Flow Matching

DGX agent

arXiv:2602.05435v2 Announce Type: replace Abstract: While flow matching is elegant, its reliance on single-sample conditional velocities leads to high-variance training targets that destabilize optimi

model-releasesarxiv-cs-cv
2 Jun 2026
← Previous
1…129130131132133…263
Next →