AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
7 Jul 2026

Reliability-Aware Monocular Depth Supervision for Sparse-View Neural Reconstruction

ResearchDGX agent

arXiv:2607.02554v1 Announce Type: new Abstract: Sparse-view neural reconstruction is challenging in outdoor driving scenes, where cameras usually move along a narrow forward-facing trajectory and prov

ReLo-IRR: Reflection-Guided LoRA Framework for Image Reflection Removal

Model ReleasesDGX agent

arXiv:2607.02957v1 Announce Type: new Abstract: Single-image reflection removal (SIRR) aims to recover the clean transmission layer from a reflection-contaminated image. Although recent methods achiev

RePos: Relative-to-Absolute Output Factorization for Cross-Environment WiFi-Based 3D Human Pose Estimation

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.02986v1 Announce Type: new Abstract: Device-free 3D human pose estimation using commodity WiFi Channel State Information (CSI) enables privacy-preserving and illumination-robust human sensi

Representation Recycling for Streaming Video Analysis

ResearchDGX agent

arXiv:2204.13492v5 Announce Type: replace Abstract: We present StreamDEQ, a method that aims to infer frame-wise representations on videos with minimal per-frame computation. Conventional deep network

Repurposing CLIP to Localize at Pixel Level

ResearchDGX agent

arXiv:2607.05253v1 Announce Type: new Abstract: Large-scale Vision-Language Models like CLIP have demonstrated impressive open-set localization capabilities at the image level. However, adapting this

Resolving Primitive-Sharing Ambiguity in Long-Tailed TLS-Based Industrial MEP Point Cloud Segmentation via Spatial Context Constraints

SafetyDGX agent

arXiv:2601.19128v2 Announce Type: replace Abstract: In terrestrial laser scanning (TLS)-based mechanical, electrical, and plumbing (MEP) point cloud segmentation, safety-critical components such as re

Rethinking Brain Decoding with CLIP: The Role of Adversarial Robustness

SafetyDGX agent

arXiv:2607.03165v1 Announce Type: new Abstract: Brain decoding aims to uncover neural mechanisms by inferring stimulus-related representations from brain signals. In fMRI studies, this is typically ac

Reward Lightning: Fast Video Generation via Homologous Preference Distillation

SafetyDGX agent

arXiv:2607.03960v1 Announce Type: new Abstract: Achieving simultaneous preference alignment and distillation acceleration in video diffusion models remains an open challenge. Existing methods optimize

RIGS-Refiner: Risk-Guided Recursive Refinement in Prediction Space for Colonoscopy Polyp Segmentation

Model ReleasesDGX agent

arXiv:2607.03058v1 Announce Type: new Abstract: Post-refinement can improve colonoscopy segmentation after host inference, but many designs still rely on extra correction heads or multi-stage pipeline

Road-Aware Anomaly Segmentation with Query-Guided Polygons and CLIP in Autonomous Driving

SafetyDGX agent

arXiv:2607.04304v1 Announce Type: new Abstract: Traditional semantic segmentation models operate under a closed-set assumption and struggle to recognize unknown or unexpected objects-an essential capa

Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

Model ReleasesDGX agent

arXiv:2510.17801v2 Announce Type: replace-cross Abstract: Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems

Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification

Model ReleasesDGX agent

arXiv:2607.03075v1 Announce Type: cross Abstract: Safety-critical applications require classifiers that are both robust and reliable. Adversarial training is a widely adopted defense for improving rob

RoMa v2: Harder Better Faster Denser Feature Matching

HardwareDGX agent

arXiv:2511.15706v3 Announce Type: replace Abstract: Dense feature matching aims to estimate all correspondences between two images of a 3D scene and has recently been established as the gold standard

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation

ResearchDGX agent

arXiv:2607.02584v1 Announce Type: new Abstract: In extbf{DiT-based video generation models equipped with 3D Rotary Position Embeddings (3D RoPE)}, the attention mechanism remains a primary computation

RSTNet: Enhancing Small-Target Recognition in Noisy SAR Imagery via Robust Feature Learning and Distribution-Aware Regression

ResearchDGX agent

arXiv:2602.23820v2 Announce Type: replace Abstract: SAR supports all-day-and-night oceanic observation, yet vessel identification from SAR images is hampered by speckle noise, intricate land-sea backg

SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection

ResearchDGX agent

arXiv:2601.03024v3 Announce Type: replace Abstract: We propose Self-Augmented Residual 3D Gaussian Splatting (SA-ResGS), a novel framework to stabilize uncertainty quantification and enhancing uncerta

SAD-LoRA: Spectral Alignment for Low-Rank Knowledge Distillation

Model ReleasesDGX agent

arXiv:2607.04306v1 Announce Type: cross Abstract: Distilling a fine-tuned teacher into a LoRA-adapted student is a standard recipe for parameter-efficient compression, but output-level KD does not exp

SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers

ResearchDGX agent

arXiv:2607.03612v1 Announce Type: new Abstract: Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success. However, scaling them to long image sequences remains chall

SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection

Model ReleasesDGX agent

arXiv:2607.03069v1 Announce Type: new Abstract: As video generation paradigms evolve from localized manipulation to full-scene synthesis, AI-generated video detection becomes increasingly challenging,

SAGE: Synchronized Action-Gaze Recognition and Anticipation for Human Behavior Understanding

Model ReleasesDGX agent

arXiv:2607.04017v1 Announce Type: new Abstract: Human object interaction (HOI), gaze pattern, and their anticipation are intricately linked, providing valuable insights into cognitive processes, inten

SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

ApplicationsDGX agent

arXiv:2508.03177v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucina

SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control

Model ReleasesDGX agent

arXiv:2607.04540v1 Announce Type: cross Abstract: Geometry-conditioned 3D scene generation enables the creation of 3D environments from user-provided geometry, offering direct control over scene struc

SE-UNet: Singular Equivariant Imaging for Real-World Constrained Generation

ApplicationsDGX agent

arXiv:2607.02628v1 Announce Type: new Abstract: While diffusion models have revolutionized image synthesis, their application to real-world inverse problems is often hampered by the need for massive d

See the Emotion: A Facial Emoji Proxy Modeling for EEG Emotion Recognition

ResearchDGX agent

arXiv:2607.02912v1 Announce Type: new Abstract: Despite the high accuracy of EEG-based emotion recognition, existing models remain opaque 'black boxes', lacking semantic grounding between abstract neu

Seeing Through WiFi: Lightweight Human Pose Estimation with Dynamic Kernel Attention

TutorialsDGX agent

arXiv:2607.03196v1 Announce Type: new Abstract: WiFi-based human pose estimation (HPE) enables the detection and interpretation of human body positions and movements without the need for wearable devi

Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration

ResearchDGX agent

arXiv:2606.12913v2 Announce Type: replace-cross Abstract: The rapid growth of modern training datasets has significantly increased computational cost, motivating dataset pruning~(DP) methods which ret

Selective Mask Propagation for Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2606.13033v2 Announce Type: replace Abstract: Multi-object tracking has a heavy-tailed difficulty distribution: most frames are easy for a lightweight base tracker, while a small fraction are in

Semantic Video Communication via Multi-Scale Convolution and Dynamic Routing for Next-Generation Networks

SafetyDGX agent

arXiv:2607.05093v1 Announce Type: new Abstract: The exponential growth of video traffic demands novel semantic communication paradigms that transmit meaning rather than raw bits. We present a generati

SGF-CDNet: A Consistency-Discrepancy Graph Network over Semantic-Geometric Fused Nodes for Face Forgery Detection

ResearchDGX agent

arXiv:2607.03883v1 Announce Type: new Abstract: The rapid advancement of deepfakes necessitates robust face forgery detection. Although forged faces may lack obvious artifacts, they often contain subt

SharpSplat: Edge-Regularized 3D Gaussian Splatting for High Fidelity Urban Building Reconstruction from UAV images

ResearchDGX agent

arXiv:2607.03872v1 Announce Type: new Abstract: Reconstructing high-fidelity 3D building models from UAV imagery is essential for large-scale digital twin development. However, existing 3D Gaussian Sp

Show Me Examples: Inferring Visual Concepts from Image Sets

TutorialsDGX agent

arXiv:2607.02402v2 Announce Type: replace Abstract: Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, curren

SiamJEPA: On the Role of Siamese Student Encoders in JEPA

SafetyDGX agent

arXiv:2607.04044v1 Announce Type: new Abstract: Recently, Joint Embedding Predictive Architectures (JEPAs) have attracted significant attention in the computer vision and machine learning communities

Signal Structure-Aware Gaussian Splatting for Large-Scale Scene Reconstruction

ResearchDGX agent

arXiv:2607.01698v2 Announce Type: replace Abstract: 3D Gaussian Splatting has demonstrated remarkable potential in novel view synthesis. In contrast to small-scale scenes, large-scale scenes inevitabl

SIMPLER: Efficient Foundation Model Adaptation via Similarity-Guided Layer Pruning for Earth Observation

Model ReleasesDGX agent

arXiv:2603.19873v2 Announce Type: replace Abstract: Fine-tuning foundation models for Earth Observation is computationally expensive, with high training time and memory demands for both training and d

SLAM: Structured and Localized Analytic Manifold Adaptation for Lifelong VPR

Model ReleasesDGX agent

arXiv:2607.04764v1 Announce Type: cross Abstract: Visual Place Recognition (VPR) in lifelong deployment requires continuous adaptation to new environments without catastrophic forgetting. In this pape

SMF-VO: Direct Ego-Motion Estimation via Sparse Motion Fields

Model ReleasesDGX agent

arXiv:2511.09072v2 Announce Type: replace-cross Abstract: Traditional Visual Odometry (VO) and Visual Inertial Odometry (VIO) methods rely on a 'pose-centric' paradigm, which computes absolute camera

SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices

Local AiDGX agent

arXiv:2601.08303v3 Announce Type: replace Abstract: Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on-device deployment due to

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

Model ReleasesDGX agent

arXiv:2607.04694v1 Announce Type: new Abstract: As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over giv

Sparse-View Surface Reconstruction using Gaussian Splatting through High-Confidence Depth Propagation with Normal Priors

ResearchDGX agent

arXiv:2607.03765v1 Announce Type: new Abstract: 3D reconstruction from sparse views is a challenging task in 3D computer vision. Recent studies on 3D Gaussian Splatting (3DGS) have achieved remarkable

Sparse4D-Radar: An Efficient and Robust Framework for Surround-View 3D Object Detection via 4D Radar-Camera Fusion

Local AiDGX agent

arXiv:2607.04098v1 Announce Type: new Abstract: In recent years, 4D imaging radar has gained wide attention in autonomous driving for its robustness against harsh weather and ability to output target

SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction

AgentsDGX agent

arXiv:2607.04732v1 Announce Type: new Abstract: Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty sp

Spatial Graph Representation and Morphometric Analysis of the Pulmonary Vascular Tree From Computed Tomography Using Multi-Scale Hessian-Based Filter Fusion and TEASAR Skeletonization

ResearchDGX agent

arXiv:2607.04457v1 Announce Type: new Abstract: Reconstructing the pulmonary vascular tree from computed tomography (CT) images is essential for quantitative lung analysis, vascular morphology assessm

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation

ResearchDGX agent

arXiv:2607.02922v1 Announce Type: new Abstract: Video reasoning segmentation demands pixel-accurate object tracking across hundreds of frames under complex natural language queries, producing dense sp

Stage-wise Attention-Guided Region Sequencing for Adversarial Attacks on Large Vision-Language Models

Local AiDGX agent

arXiv:2602.04356v2 Announce Type: replace Abstract: Targeted adversarial attacks on Large Vision-Language Models (LVLMs) test whether small image perturbations can steer model responses toward attacke

SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

Model ReleasesDGX agent

arXiv:2607.05264v1 Announce Type: new Abstract: Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test

Structure-Guided Self-Supervised Matching for One-Shot Medical Landmark Detection

SafetyDGX agent

arXiv:2203.01687v3 Announce Type: replace Abstract: Medical landmark detection usually requires accurate expert annotations, which are laborious and difficult to scale across anatomical regions. In th

StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation

Model ReleasesDGX agent

arXiv:2607.04612v1 Announce Type: cross Abstract: Graphic design editing requires precise manipulation of typography, layout, and visual hierarchy under strict design constraints. Following the introd

SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing

Model ReleasesDGX agent

arXiv:2603.08982v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) have become a leading backbone for video generation, yet their quadratic attention cost remains a major bottleneck. Sp

Symmetry-Structured Neural Completion of Islamic Geometric Patterns from Sparse Control Geometry

ResearchDGX agent

arXiv:2607.02573v1 Announce Type: new Abstract: Islamic geometric patterns are governed by exact rotational symmetry and strict construction rules. This paper treats these rules as formal geometric kn

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

ResearchDGX agent

arXiv:2607.05392v1 Announce Type: new Abstract: We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the abi

Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework

Model ReleasesDGX agent

arXiv:2606.19073v2 Announce Type: replace Abstract: Current image editing methods excel at static attributes but fail at complex Human-Object Interactions (HOI), a critical challenge unaddressed by ex

TemporalGS: Training-Free Plug-and-Play Acceleration for 3D Gaussian Splatting Rendering via Temporal Priors

ResearchDGX agent

arXiv:2607.03390v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has revolutionized novel-view synthesis with its fast and high-fidelity rendering. However, rendering at high FPS and low l

TESSERA v2: Scaling Pixel-wise Earth Foundation Models

Model ReleasesDGX agent

arXiv:2607.03949v1 Announce Type: new Abstract: Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spatial embeddings. However, how these

TestMate: Test-Time Domain Adaptation Aided by Lightweight Vision Foundation Model

Model ReleasesDGX agent

arXiv:2607.03810v1 Announce Type: new Abstract: Test-Time Domain Adaptation (TTDA) aims to adapt Deep Neural Networks to distribution shifts using only streaming, unlabeled test data in real time. Cur

Text-to-Image Generation for Projector-Camera System Registration

ApplicationsDGX agent

arXiv:2607.03046v1 Announce Type: new Abstract: Establishing correspondence between projector and camera images in a procam (projector + camera) system is essential for achieving high-resolution pixel

TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers

Model ReleasesDGX agent

arXiv:2601.02211v2 Announce Type: replace Abstract: Recent breakthroughs of transformer-based diffusion models, particularly with Multimodal Diffusion Transformers (MMDiT) driven models like FLUX and

TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

AgentsDGX agent

arXiv:2607.04812v1 Announce Type: new Abstract: Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the erro

The Good, the Bad, and the Brittle: Benchmarking Robustness and Generalisation of Histopathology Foundation Models

Model ReleasesDGX agent

arXiv:2607.04401v1 Announce Type: new Abstract: How robust and generalisable are pathology foundation models and have their scaling limites been reached? We benchmarked twelve pathology foundation mod

The Multipath Blind Spot: K-Agnostic Robust Calibration for Sparse-Anchor Metric Depth from Frozen Foundations

Model ReleasesDGX agent

arXiv:2607.04101v1 Announce Type: new Abstract: Monocular depth foundations predict domain-general relative depth but lack absolute scale; a handful of sparse metric anchors from a range sensor can ca

The P^3 Dataset: Pixels, Points and Polygons for Multimodal Building Vectorization

Model ReleasesDGX agent

arXiv:2505.15379v2 Announce Type: replace Abstract: We present the P^3 dataset, a large-scale multimodal benchmark for building vectorization, constructed from aerial LiDAR point clouds, high-resoluti

← Previous
1…5253545556…209
Next →