AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Applications

See and Switch: Vision-Based Branching for Interactive Robot-Skill Programming

DGX agent

arXiv:2603.08057v2 Announce Type: replace-cross Abstract: Programming by demonstration (PbD) makes robot programming accessible to non-experts, but scaling it to real-world variability remains a chall

applicationsarxiv-cs-cv
30 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Local Ai

See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs

DGX agent

arXiv:2606.29847v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) excel at multimodal tasks but remain prone to object hallucinations. Prior training-free remedies often uniformly s

local-aiarxiv-cs-cv
30 Jun 2026
Research

Seed-to-Seed: Unpaired Image Translation in Diffusion Seed Space

DGX agent

arXiv:2409.00654v2 Announce Type: replace Abstract: We introduce Seed-to-Seed Translation (StS), a novel approach that combines GANs and diffusion models (DMs) for unpaired Image-to-Image Translation.

researcharxiv-cs-cv
30 Jun 2026
Safety

Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation

DGX agent

arXiv:2606.29941v1 Announce Type: cross Abstract: Visuo-Tactile policies leveraging optical tactile sensors have shown great promise in contact-rich manipulation. These sensors achieve high spatial re

safetyarxiv-cs-cv
30 Jun 2026
Agents

Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution

DGX agent

arXiv:2606.28971v1 Announce Type: new Abstract: Real-world image restoration (IR) remains challenging due to complex and coupled degradations. While recent agentic IR frameworks leverage Large Languag

agentsarxiv-cs-cv
30 Jun 2026
Research

Semantic-Aware Generative Image Transmission for Resource-Constrained Visual IoT Systems

DGX agent

arXiv:2606.28398v1 Announce Type: new Abstract: Resource-constrained visual Internet of Things (IoT) systems, such as edge cameras, unmanned sensing platforms, industrial inspection nodes, and remote

researcharxiv-cs-cv
30 Jun 2026
Research

Semantic Correspondence: Unified Benchmarking and a Strong Baseline

DGX agent

arXiv:2505.18060v4 Announce Type: replace Abstract: Establishing semantic correspondence is a challenging task in computer vision, aiming to match keypoints with the same semantic information across d

researcharxiv-cs-cv
30 Jun 2026
Model Releases

Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation

DGX agent

arXiv:2606.30244v1 Announce Type: new Abstract: Referring Remote Sensing Image Segmentation (RRSIS) seeks to localize and segment the target object or region specified by a natural language expression

model-releasesarxiv-cs-cv
30 Jun 2026
Research

SemConFlow: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching

DGX agent

arXiv:2603.26553v2 Announce Type: replace Abstract: While the field of co-speech gesture generation has seen significant advances, producing holistic, semantically grounded gestures remains a challeng

researcharxiv-cs-cv
30 Jun 2026
Agents

Shell-Supervised Gaussian Splatting for Urban Real-to-Sim Reconstruction

DGX agent

arXiv:2606.30014v1 Announce Type: new Abstract: Real-to-sim reconstruction for embodied AI requires geometry that is useful for collision reasoning, navigation, and agent-environment interaction, not

agentsarxiv-cs-cv
30 Jun 2026
Agents

SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset

DGX agent

arXiv:2606.30001v1 Announce Type: new Abstract: Recent co-speech gesture generation methods often overlook cultural differences, limiting their effectiveness in human-agent interaction. Moreover, cult

agentsarxiv-cs-cv
30 Jun 2026
Research

SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation

DGX agent

arXiv:2606.28626v1 Announce Type: new Abstract: Sign language translation (SLT) remains challenging due to its high spatio-temporal complexity, long sequences, and the need to model multiple articulat

researcharxiv-cs-cv
30 Jun 2026
Safety

SIR: Structured Image Representations for Explainable Robot Learning

DGX agent

arXiv:2606.30101v1 Announce Type: cross Abstract: Existing robot policies based on learned visual embeddings lack explicit structure and are sensitive to visual distractions. Thus, the representations

safetyarxiv-cs-cv
30 Jun 2026
Research

SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation

DGX agent

arXiv:2603.14152v2 Announce Type: replace Abstract: Native 3D generative models have achieved remarkable fidelity and speed, yet they suffer from a critical limitation: inability to prescribe precise

researcharxiv-cs-cv
30 Jun 2026
Model Releases

SkelEM: Training-Signal Decoupling of Skeleton and Diffusion for Self-supervised Axial Super-Resolution in Volume Microscopy

DGX agent

arXiv:2606.30012v1 Announce Type: new Abstract: Volume microscopy, including electron and light microscopy, suffers from severe anisotropic resolution due to physical axial sectioning. Existing self-s

model-releasesarxiv-cs-cv
30 Jun 2026
Research

SoccerNet 2026 Player-Centric Ball Action Spotting: Per-Player Attention with Agreement-Based Ensembling

DGX agent

arXiv:2606.28389v1 Announce Type: new Abstract: We present our submission to the SoccerNet 2026 Player-Centric Ball Action Spotting challenge, which uses a two-stage pipeline: a Track-Aware Action Det

researcharxiv-cs-cv
30 Jun 2026
Safety

SPARC: Scalable Path-Specific Counterfactual Fairness via Causal Conditional Independence

DGX agent

arXiv:2412.04739v2 Announce Type: replace Abstract: Deep learning models exhibit fairness concerns when predictions are inadvertently influenced by sensitive attributes. However, existing attempts to

safetyarxiv-cs-cv
30 Jun 2026
Applications

Spatially Localized Image Degradation Embeddings for Image Quality Assessment

DGX agent

arXiv:2606.29162v1 Announce Type: new Abstract: Self-supervised learning (SSL) currently drives state-of-the-art performance in no-reference image quality assessment (NR-IQA). However, standard SSL pi

applicationsarxiv-cs-cv
30 Jun 2026
Research

SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization

DGX agent

arXiv:2510.04961v2 Announce Type: replace Abstract: Tokenizers are a key component of state-of-the-art generative image models, extracting the most important features from the signal while reducing da

researcharxiv-cs-cv
30 Jun 2026
Model Releases

Stability and Concentration in Nonlinear Inverse Problems with Block-Structured Parameters: Lipschitz Geometry, Identifiability, and an Application to Gaussian Splatting

DGX agent

arXiv:2602.09415v2 Announce Type: replace Abstract: We develop an operator-theoretic framework for stability and statistical concentration in nonlinear inverse problems with block-structured parameter

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging

DGX agent

arXiv:2512.01461v2 Announce Type: replace-cross Abstract: Model merging has emerged as a promising paradigm for enabling multi-task capabilities without additional training. However, traditional basic

model-releasesarxiv-cs-cv
30 Jun 2026
Research

StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors

DGX agent

arXiv:2606.30545v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has achieved remarkable success in real-time novel view synthesis, yet it suffers from severe overfitting under sparse-view

researcharxiv-cs-cv
30 Jun 2026
Research

Stimulus Motion Perception Studies Imply Specific Neural Computations in Human Visual Stabilization

DGX agent

arXiv:2506.13506v3 Announce Type: replace Abstract: Even during fixation the human eye is constantly in low amplitude motion, jittering over small angles in random directions at up to 100Hz. This moti

researcharxiv-cs-cv
30 Jun 2026
Research

Stochastic Optimal Control Sampling for Diffusion Inverse Problems

DGX agent

arXiv:2606.28785v1 Announce Type: new Abstract: Benefiting from the strong ability to capture data distributions, diffusion models have become powerful tools for solving image inverse problems. The ke

researcharxiv-cs-cv
30 Jun 2026
Model Releases

StrucTab: A Structured Optimization Framework for Table Parsing

DGX agent

arXiv:2606.29905v1 Announce Type: new Abstract: Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial l

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

DGX agent

arXiv:2603.12703v3 Announce Type: replace Abstract: Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video u

model-releasesarxiv-cs-cv
30 Jun 2026
Tutorials

T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation

DGX agent

arXiv:2606.30147v1 Announce Type: new Abstract: Recent progress in Text-to-Image generation benefits from large-scale Text-Image pairs. However, the scarcity of Text-LiDAR pairs often causes over-smoo

tutorialsarxiv-cs-cv
30 Jun 2026
Safety

TerraDiT: Point-Conditioned Diffusion Transformer for Satellite Image Synthesis

DGX agent

arXiv:2603.02172v2 Announce Type: replace Abstract: We introduce TerraDiT, a diffusion transformer designed for text-to-satellite image generation with point-based control. Existing controlled satelli

safetyarxiv-cs-cv
30 Jun 2026
Research

Text-Conditioned Background Generation for Editable Multi-Layer Documents

DGX agent

arXiv:2512.17151v2 Announce Type: replace Abstract: We present a framework for document-centric background generation with multi-page editing and thematic continuity. To ensure text regions remain rea

researcharxiv-cs-cv
30 Jun 2026
Research

The Calibrated Deepfake Trust Score (CDTS): Competence-Coupled Trust Degradation Across Deepfake Detectors

DGX agent

arXiv:2606.29484v1 Announce Type: cross Abstract: Modern deepfake detectors are rarely consumed as bare classifiers. In moderation, provenance, and verification pipelines their output probability is r

researcharxiv-cs-cv
30 Jun 2026
Research

The Platonic Defense: Backdoor Defense for Self-Supervised Encoders in the Era of Large Scale Pre-training

DGX agent

arXiv:2606.29451v1 Announce Type: new Abstract: Self-supervised learning (SSL) pretrained models have become a dominant paradigm for visual representation learning, but they are vulnerable to backdoor

researcharxiv-cs-cv
30 Jun 2026
Research

The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction

DGX agent

arXiv:2606.30308v1 Announce Type: new Abstract: 4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-based pipelines depend on a detector

researcharxiv-cs-cv
30 Jun 2026
Model Releases

Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

DGX agent

arXiv:2606.30598v1 Announce Type: new Abstract: Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing lea

model-releasesarxiv-cs-cv
30 Jun 2026
Local Ai

Towards Long-Form Spatio-Temporal Video Grounding

DGX agent

arXiv:2602.23294v2 Announce Type: replace Abstract: In real scenarios, videos can span several minutes or even hours. However, existing research on spatio-temporal video grounding (STVG), given a text

local-aiarxiv-cs-cv
30 Jun 2026
Model Releases

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

DGX agent

arXiv:2512.13660v3 Announce Type: replace-cross Abstract: Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded

model-releasesarxiv-cs-cv
30 Jun 2026
Research

TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment

DGX agent

arXiv:2606.30313v1 Announce Type: new Abstract: Longitudinal glioblastoma response assessment requires comparing subtle tumor changes across MRI time points using structured clinical criteria such as

researcharxiv-cs-cv
30 Jun 2026
Tutorials

Traffic-CBM: A Structurally Interpretable Multimodal Framework for Encrypted Traffic Classification

DGX agent

arXiv:2606.29909v1 Announce Type: new Abstract: Encrypted traffic classification has achieved strong performance, but its decision process remains difficult to interpret. Existing methods usually comb

tutorialsarxiv-cs-cv
30 Jun 2026
Safety

TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation

DGX agent

arXiv:2606.29097v1 Announce Type: new Abstract: Recent research has investigated the use of large language models (LLMs) to generate traffic scenarios for autonomous driving. However, pretrained LLMs

safetyarxiv-cs-cv
30 Jun 2026
Model Releases

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision

DGX agent

arXiv:2606.30552v1 Announce Type: cross Abstract: Cross-embodiment transfer in vision-language-action (VLA) models remains challenging because low-level state and action spaces differ fundamentally ac

model-releasesarxiv-cs-cv
30 Jun 2026
Research

Transition-Aware best-of-N sampling for Longitudinal Chest X-ray Reports

DGX agent

arXiv:2606.28393v1 Announce Type: new Abstract: In longitudinal clinical practice, every chest X-ray is read in the context of the patients prior exam, and much of what the radiologist communicates is

researcharxiv-cs-cv
30 Jun 2026
Applications

TUGS: Physics-based Compact Representation of Underwater Scenes by Tensorized Gaussian

DGX agent

arXiv:2505.08811v3 Announce Type: replace Abstract: Underwater 3D scene reconstruction is crucial for multimedia applications in adverse environments, such as underwater robotic perception and navigat

applicationsarxiv-cs-cv
30 Jun 2026
Model Releases

Tumor-aware augmentation with task-guided attention analysis improves rectal cancer segmentation from magnetic resonance images

DGX agent

arXiv:2605.05522v2 Announce Type: replace-cross Abstract: Although self-supervised pretraining is expected to learn broadly transferable representations, its effectiveness across imaging modalities su

model-releasesarxiv-cs-cv
30 Jun 2026
Applications

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

DGX agent

arXiv:2602.22960v2 Announce Type: replace Abstract: World models based on video generation demonstrate remarkable potential for simulating interactive environments yet suffer from persistent difficult

applicationsarxiv-cs-cv
30 Jun 2026
Local Ai

UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention

DGX agent

arXiv:2510.16325v3 Announce Type: replace Abstract: Ultra-high-resolution text-to-image generation is increasingly vital for applications requiring fine-grained textures and global structural fidelity

local-aiarxiv-cs-cv
30 Jun 2026
Research

Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning

DGX agent

arXiv:2606.30020v1 Announce Type: new Abstract: Pathology foundation models (PFMs) offer generalizable representations for whole-slide image (WSI) analysis, yet their clinical adoption remains limited

researcharxiv-cs-cv
30 Jun 2026
Agents

UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image

DGX agent

arXiv:2606.30608v1 Announce Type: new Abstract: Articulated 3D objects are essential for interactive environments in embodied AI, robotics, and virtual reality, but reconstructing their structure and

agentsarxiv-cs-cv
30 Jun 2026
Model Releases

UniCA: Bi-directional Cross-Attention with Positive Similarity Loss for Robust Multi-Modal Retrieval

DGX agent

arXiv:2606.28350v1 Announce Type: cross Abstract: Multi-modal retrieval has become increasingly critical for handling the growing volume of integrated visual-textual data in real-world applications, b

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

DGX agent

arXiv:2606.30332v1 Announce Type: new Abstract: Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing app

safetyarxiv-cs-cv
30 Jun 2026
← Previous
1…8283848586…263
Next →