AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlog
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

SharpSplat: Edge-Regularized 3D Gaussian Splatting for High Fidelity Urban Building Reconstruction from UAV images

DGX agent

arXiv:2607.03872v1 Announce Type: new Abstract: Reconstructing high-fidelity 3D building models from UAV imagery is essential for large-scale digital twin development. However, existing 3D Gaussian Sp

researcharxiv-cs-cv
7 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Tutorials

Show Me Examples: Inferring Visual Concepts from Image Sets

DGX agent

arXiv:2607.02402v2 Announce Type: replace Abstract: Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, curren

tutorialsarxiv-cs-cv
7 Jul 2026
Safety

SiamJEPA: On the Role of Siamese Student Encoders in JEPA

DGX agent

arXiv:2607.04044v1 Announce Type: new Abstract: Recently, Joint Embedding Predictive Architectures (JEPAs) have attracted significant attention in the computer vision and machine learning communities

safetyarxiv-cs-cv
7 Jul 2026
Research

Signal Structure-Aware Gaussian Splatting for Large-Scale Scene Reconstruction

DGX agent

arXiv:2607.01698v2 Announce Type: replace Abstract: 3D Gaussian Splatting has demonstrated remarkable potential in novel view synthesis. In contrast to small-scale scenes, large-scale scenes inevitabl

researcharxiv-cs-cv
7 Jul 2026
Model Releases

SIMPLER: Efficient Foundation Model Adaptation via Similarity-Guided Layer Pruning for Earth Observation

DGX agent

arXiv:2603.19873v2 Announce Type: replace Abstract: Fine-tuning foundation models for Earth Observation is computationally expensive, with high training time and memory demands for both training and d

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

SLAM: Structured and Localized Analytic Manifold Adaptation for Lifelong VPR

DGX agent

arXiv:2607.04764v1 Announce Type: cross Abstract: Visual Place Recognition (VPR) in lifelong deployment requires continuous adaptation to new environments without catastrophic forgetting. In this pape

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

SMF-VO: Direct Ego-Motion Estimation via Sparse Motion Fields

DGX agent

arXiv:2511.09072v2 Announce Type: replace-cross Abstract: Traditional Visual Odometry (VO) and Visual Inertial Odometry (VIO) methods rely on a 'pose-centric' paradigm, which computes absolute camera

model-releasesarxiv-cs-cv
7 Jul 2026
Local Ai

SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices

DGX agent

arXiv:2601.08303v3 Announce Type: replace Abstract: Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on-device deployment due to

local-aiarxiv-cs-cv
7 Jul 2026
Model Releases

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

DGX agent

arXiv:2607.04694v1 Announce Type: new Abstract: As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over giv

model-releasesarxiv-cs-cv
7 Jul 2026
Research

Sparse-View Surface Reconstruction using Gaussian Splatting through High-Confidence Depth Propagation with Normal Priors

DGX agent

arXiv:2607.03765v1 Announce Type: new Abstract: 3D reconstruction from sparse views is a challenging task in 3D computer vision. Recent studies on 3D Gaussian Splatting (3DGS) have achieved remarkable

researcharxiv-cs-cv
7 Jul 2026
Local Ai

Sparse4D-Radar: An Efficient and Robust Framework for Surround-View 3D Object Detection via 4D Radar-Camera Fusion

DGX agent

arXiv:2607.04098v1 Announce Type: new Abstract: In recent years, 4D imaging radar has gained wide attention in autonomous driving for its robustness against harsh weather and ability to output target

local-aiarxiv-cs-cv
7 Jul 2026
Agents

SparseOcc++: Geometry-Aware Sparse Latent Representation for Semantic Occupancy Prediction

DGX agent

arXiv:2607.04732v1 Announce Type: new Abstract: Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty sp

agentsarxiv-cs-cv
7 Jul 2026
Research

Spatial Graph Representation and Morphometric Analysis of the Pulmonary Vascular Tree From Computed Tomography Using Multi-Scale Hessian-Based Filter Fusion and TEASAR Skeletonization

DGX agent

arXiv:2607.04457v1 Announce Type: new Abstract: Reconstructing the pulmonary vascular tree from computed tomography (CT) images is essential for quantitative lung analysis, vascular morphology assessm

researcharxiv-cs-cv
7 Jul 2026
Research

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation

DGX agent

arXiv:2607.02922v1 Announce Type: new Abstract: Video reasoning segmentation demands pixel-accurate object tracking across hundreds of frames under complex natural language queries, producing dense sp

researcharxiv-cs-cv
7 Jul 2026
Local Ai

Stage-wise Attention-Guided Region Sequencing for Adversarial Attacks on Large Vision-Language Models

DGX agent

arXiv:2602.04356v2 Announce Type: replace Abstract: Targeted adversarial attacks on Large Vision-Language Models (LVLMs) test whether small image perturbations can steer model responses toward attacke

local-aiarxiv-cs-cv
7 Jul 2026
Model Releases

SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

DGX agent

arXiv:2607.05264v1 Announce Type: new Abstract: Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test

model-releasesarxiv-cs-cv
7 Jul 2026
Safety

Structure-Guided Self-Supervised Matching for One-Shot Medical Landmark Detection

DGX agent

arXiv:2203.01687v3 Announce Type: replace Abstract: Medical landmark detection usually requires accurate expert annotations, which are laborious and difficult to scale across anatomical regions. In th

safetyarxiv-cs-cv
7 Jul 2026
Model Releases

StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation

DGX agent

arXiv:2607.04612v1 Announce Type: cross Abstract: Graphic design editing requires precise manipulation of typography, layout, and visual hierarchy under strict design constraints. Following the introd

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing

DGX agent

arXiv:2603.08982v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) have become a leading backbone for video generation, yet their quadratic attention cost remains a major bottleneck. Sp

model-releasesarxiv-cs-cv
7 Jul 2026
Research

Symmetry-Structured Neural Completion of Islamic Geometric Patterns from Sparse Control Geometry

DGX agent

arXiv:2607.02573v1 Announce Type: new Abstract: Islamic geometric patterns are governed by exact rotational symmetry and strict construction rules. This paper treats these rules as formal geometric kn

researcharxiv-cs-cv
7 Jul 2026
Research

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

DGX agent

arXiv:2607.05392v1 Announce Type: new Abstract: We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the abi

researcharxiv-cs-cv
7 Jul 2026
Model Releases

Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework

DGX agent

arXiv:2606.19073v2 Announce Type: replace Abstract: Current image editing methods excel at static attributes but fail at complex Human-Object Interactions (HOI), a critical challenge unaddressed by ex

model-releasesarxiv-cs-cv
7 Jul 2026
Research

TemporalGS: Training-Free Plug-and-Play Acceleration for 3D Gaussian Splatting Rendering via Temporal Priors

DGX agent

arXiv:2607.03390v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has revolutionized novel-view synthesis with its fast and high-fidelity rendering. However, rendering at high FPS and low l

researcharxiv-cs-cv
7 Jul 2026
Model Releases

TESSERA v2: Scaling Pixel-wise Earth Foundation Models

DGX agent

arXiv:2607.03949v1 Announce Type: new Abstract: Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spatial embeddings. However, how these

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

TestMate: Test-Time Domain Adaptation Aided by Lightweight Vision Foundation Model

DGX agent

arXiv:2607.03810v1 Announce Type: new Abstract: Test-Time Domain Adaptation (TTDA) aims to adapt Deep Neural Networks to distribution shifts using only streaming, unlabeled test data in real time. Cur

model-releasesarxiv-cs-cv
7 Jul 2026
Applications

Text-to-Image Generation for Projector-Camera System Registration

DGX agent

arXiv:2607.03046v1 Announce Type: new Abstract: Establishing correspondence between projector and camera images in a procam (projector + camera) system is essential for achieving high-resolution pixel

applicationsarxiv-cs-cv
7 Jul 2026
Model Releases

TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers

DGX agent

arXiv:2601.02211v2 Announce Type: replace Abstract: Recent breakthroughs of transformer-based diffusion models, particularly with Multimodal Diffusion Transformers (MMDiT) driven models like FLUX and

model-releasesarxiv-cs-cv
7 Jul 2026
Agents

TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

DGX agent

arXiv:2607.04812v1 Announce Type: new Abstract: Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the erro

agentsarxiv-cs-cv
7 Jul 2026
Model Releases

The Good, the Bad, and the Brittle: Benchmarking Robustness and Generalisation of Histopathology Foundation Models

DGX agent

arXiv:2607.04401v1 Announce Type: new Abstract: How robust and generalisable are pathology foundation models and have their scaling limites been reached? We benchmarked twelve pathology foundation mod

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

The Multipath Blind Spot: K-Agnostic Robust Calibration for Sparse-Anchor Metric Depth from Frozen Foundations

DGX agent

arXiv:2607.04101v1 Announce Type: new Abstract: Monocular depth foundations predict domain-general relative depth but lack absolute scale; a handful of sparse metric anchors from a range sensor can ca

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

The P^3 Dataset: Pixels, Points and Polygons for Multimodal Building Vectorization

DGX agent

arXiv:2505.15379v2 Announce Type: replace Abstract: We present the P^3 dataset, a large-scale multimodal benchmark for building vectorization, constructed from aerial LiDAR point clouds, high-resoluti

model-releasesarxiv-cs-cv
7 Jul 2026
Applications

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

DGX agent

arXiv:2602.06575v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models typically inject proprioception only as a late conditioning signal, preventing robot state from grounding

applicationsarxiv-cs-cv
7 Jul 2026
Research

TimeThink: Reasoning with Time for Video LLMs

DGX agent

arXiv:2607.05089v1 Announce Type: new Abstract: Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Vi

researcharxiv-cs-cv
7 Jul 2026
Model Releases

TiROD: Tiny Robotics Dataset and Benchmark for Continual Object Detection

DGX agent

arXiv:2409.16215v4 Announce Type: replace-cross Abstract: Detecting objects with visual sensors is crucial for numerous mobile robotics applications, from autonomous navigation to inspection. However,

model-releasesarxiv-cs-cv
7 Jul 2026
Tutorials

Token Communications: A Large Model-Driven Framework for Cross-modal Context-aware Semantic Communications

DGX agent

arXiv:2502.12096v5 Announce Type: replace-cross Abstract: In this paper, we introduce token communications (TokCom), a large model-driven framework to leverage cross-modal context information in gener

tutorialsarxiv-cs-cv
7 Jul 2026
Model Releases

Topology-Driven Transferability Estimation for 3D Medical Vision Foundation Models

DGX agent

arXiv:2607.04199v1 Announce Type: new Abstract: The growing number of medical vision foundation models highlights the need for effective model selection. However, mainstream selection methods rely on

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker

DGX agent

arXiv:2605.25706v2 Announce Type: replace Abstract: Referring expression comprehension (REC) aims to localize a target object within an image based on a given expression. Although recent advances in v

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Towards Realistic Remote Sensing Dataset Distillation with Discriminative Prototype-guided Diffusion

DGX agent

arXiv:2601.15829v2 Announce Type: replace Abstract: Recent years have witnessed the remarkable success of deep learning in remote sensing image interpretation, driven by the availability of large-scal

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Towards Standardized Light Field Quality Assessment: Hybrid Subjective Benchmarking and Objective Metric Evaluation

DGX agent

arXiv:2607.03494v1 Announce Type: new Abstract: Benchmarking immersive media coding solutions, especially in the standardization context, requires reliable and reproducible subjective quality assessme

model-releasesarxiv-cs-cv
7 Jul 2026
Local Ai

Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

DGX agent

arXiv:2607.02798v1 Announce Type: new Abstract: Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditi

local-aiarxiv-cs-cv
7 Jul 2026
Safety

Trajectory-Anchor Optimization for Overconfident Thermal Visual Place Recognition: Zero-Leakage OOD Auditing and Kidnapped-Robot Recovery

DGX agent

arXiv:2607.04745v1 Announce Type: cross Abstract: Modern thermal visual place recognition (TIR-VPR) frontends based on foundation models achieve remarkable closed-set retrieval but suffer from an over

safetyarxiv-cs-cv
7 Jul 2026
Model Releases

Triple-Phase Multimodal Knowledge Aggregation Framework for Microbial Keratitis Subtype Diagnosis on Slit-Lamp Photography

DGX agent

arXiv:2607.03740v1 Announce Type: cross Abstract: Microbial keratitis requires rapid pathogen identification to guide treatment, but culture- and PCR-based diagnostics are slow and resource-intensive.

model-releasesarxiv-cs-cv
7 Jul 2026
Agents

TRISTAR: Triple-Signal Stair Recognition and Vision-Only Indoor Navigation for Search-and-Rescue Micro-UAVs

DGX agent

arXiv:2607.03818v1 Announce Type: new Abstract: Indoor search-and-rescue (SAR) operations often require rapid situational awareness where GNSS signals are unavailable and human access is difficult or

agentsarxiv-cs-cv
7 Jul 2026
Research

TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction

DGX agent

arXiv:2607.04484v1 Announce Type: new Abstract: Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal rea

researcharxiv-cs-cv
7 Jul 2026
Research

TubeLite: Lightweight Multi-Actor Spatio-Temporal Action Detection

DGX agent

arXiv:2607.04684v1 Announce Type: new Abstract: Spatio-temporal action detection in videos requires jointly localizing actors in space and identifying action boundaries over time. A common challenge i

researcharxiv-cs-cv
7 Jul 2026
Safety

Uncertainty-Aware Last-Layer Adaptation of RETFound for Referable Diabetic Retinopathy Screening Under Dataset Shift

DGX agent

arXiv:2607.02569v1 Announce Type: new Abstract: This paper presents a safety-centered empirical evaluation of uncertainty-aware last-layer adaptation for referable diabetic retinopathy screening using

safetyarxiv-cs-cv
7 Jul 2026
Model Releases

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy

DGX agent

arXiv:2603.24690v2 Announce Type: replace Abstract: In-context learning (ICL) enables fast task adaptation from demonstrations without per-task parameter updates but remains highly sensitive to exampl

model-releasesarxiv-cs-cv
7 Jul 2026
Local Ai

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization

DGX agent

arXiv:2607.04498v1 Announce Type: new Abstract: With the proliferation of AI-generated content, sophisticated multimedia manipulation has raised critical concerns about malicious applications such as

local-aiarxiv-cs-cv
7 Jul 2026
← Previous
1…6869707172…263
Next →