AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs

DGX agent

arXiv:2608.05616v1 Announce Type: new Abstract: Despite the remarkable progress of large vision language models (LVLMs), object hallucination remains a fundamental challenge that hinders their trustwo

researcharxiv-cs-cv
7 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

DGX agent

arXiv:2608.05597v1 Announce Type: new Abstract: Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based

applicationsarxiv-cs-cv
7 Aug 2026
Research

Universal Concept Disruption for SAM3 Image Segmentation

DGX agent

arXiv:2608.05983v1 Announce Type: new Abstract: SAM3 extends promptable segmentation from geometry-driven mask prediction to open-vocabulary concept segmentation, where a text-conditioned grounding mo

researcharxiv-cs-cv
7 Aug 2026
Local Ai

UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate Regression

DGX agent

arXiv:2608.06307v1 Announce Type: new Abstract: LiDAR-based Scene Coordinate Regression (SCR) maps point clouds directly to 3D scene coordinates, enabling precise 6-DoF localisation without explicit m

local-aiarxiv-cs-cv
7 Aug 2026
Research

URNet: A Unified Reparameterized Network for Efficient RGB-D Semantic Segmentation

DGX agent

arXiv:2608.05671v1 Announce Type: new Abstract: Previous RGB-D semantic segmentation methods commonly employ dual encoders to separately process RGB and depth inputs, followed by dedicated modules for

researcharxiv-cs-cv
7 Aug 2026
Research

Versatile Video Representation via Feed-Forward 2D Gaussian Splatting Tokenization

DGX agent

arXiv:2508.11183v2 Announce Type: replace Abstract: Recent video representation methods that rely on fixed-grid, patch-wise tokenization often exhibit limited versatility.Spatially, uniformly allocati

researcharxiv-cs-cv
7 Aug 2026
Model Releases

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing

DGX agent

arXiv:2608.05485v1 Announce Type: new Abstract: Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and edit

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

Visual Intention Grounding for Egocentric Assistants

DGX agent

arXiv:2504.13621v2 Announce Type: replace Abstract: Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object qu

model-releasesarxiv-cs-cv
7 Aug 2026
Tutorials

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

DGX agent

arXiv:2608.05215v1 Announce Type: cross Abstract: Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots ma

tutorialsarxiv-cs-cv
7 Aug 2026
Model Releases

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

DGX agent

arXiv:2608.05776v1 Announce Type: new Abstract: Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditi

model-releasesarxiv-cs-cv
7 Aug 2026
Research

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

DGX agent

arXiv:2608.05648v1 Announce Type: new Abstract: Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of

researcharxiv-cs-cv
7 Aug 2026
Research

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

DGX agent

arXiv:2608.05803v1 Announce Type: new Abstract: Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches

researcharxiv-cs-cv
7 Aug 2026
Local Ai

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

DGX agent

arXiv:2608.05663v1 Announce Type: new Abstract: Real-time long-form avatar audio--video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consis

local-aiarxiv-cs-cv
7 Aug 2026
Research

VSMP-IMU: Video-Grounded Semantic Motion Programs for Sensor-Aware Synthetic IMU Generation

DGX agent

arXiv:2608.05782v1 Announce Type: new Abstract: Wearable human activity recognition (HAR) is often limited by the scarcity of labeled sensor data, especially in low-resource, class-imbalanced, and sub

researcharxiv-cs-cv
7 Aug 2026
Research

Wan-Animate-2: Pushing the Application Boundaries of Character Animation

DGX agent

arXiv:2608.06009v1 Announce Type: new Abstract: Character image animation remains a foundational yet challenging task in computer vision. Existing approaches can be broadly categorized into three para

researcharxiv-cs-cv
7 Aug 2026
Model Releases

What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective

DGX agent

arXiv:2606.14299v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain v

model-releasesarxiv-cs-cv
7 Aug 2026
Local Ai

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

DGX agent

arXiv:2608.05369v1 Announce Type: cross Abstract: Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in r

local-aiarxiv-cs-cv
7 Aug 2026
Safety

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

DGX agent

arXiv:2608.05799v1 Announce Type: cross Abstract: Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to

safetyarxiv-cs-cv
7 Aug 2026
Applications

A GitOps-Driven Annotation Catalog for Fully Automatic Railway Operations

DGX agent

arXiv:2608.04724v1 Announce Type: cross Abstract: Automatic train operation (ATO) at grade of automation 3 and above (GoA3-GoA4) requires robust AI-based perception systems capable of reliably detecti

applicationsarxiv-cs-cv
6 Aug 2026
Research

A Multi-Sensor Dataset for Monitoring the Operational Environment of Rail Vehicles

DGX agent

arXiv:2608.04704v1 Announce Type: new Abstract: Reliable environment monitoring is essential for the safe and efficient operation of automated railway systems, covering all Grades of Automation (GoA),

researcharxiv-cs-cv
6 Aug 2026
Model Releases

ACA-GS: Adaptive-Capacity Anchored Gaussian Splatting for Compact Dynamic Radiance Fields

DGX agent

arXiv:2608.04581v1 Announce Type: new Abstract: Recent advances in 4D Gaussian Splatting (4DGS) enable high-fidelity, real-time spatiotemporal rendering, but expose a fundamental trade-off between mot

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

Advancing Utility Pole and Sign Detection Through Deep Learning

DGX agent

arXiv:2608.04061v1 Announce Type: new Abstract: Utility poles are an essential part of the infrastructure used to support power distribution systems and other critical public services. Their regular i

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle

DGX agent

arXiv:2608.04314v1 Announce Type: cross Abstract: Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Legal and regulatory remedies can add

model-releasesarxiv-cs-cv
6 Aug 2026
Research

Aligning Fetal Anatomy with Kinematic Tree Log-Euclidean PolyRigid Transforms

DGX agent

arXiv:2603.02371v2 Announce Type: replace Abstract: Automated analysis of articulated bodies is crucial in medical imaging. Existing surface-based models often ignore internal volumetric structures an

researcharxiv-cs-cv
6 Aug 2026
Model Releases

An active-learning framework for real-time depth perception from monocular vision streams

DGX agent

arXiv:2608.04917v1 Announce Type: new Abstract: Biological visual systems can perceive depth from monocular vision flow, continuously integrating temporal visual cues while maintaining a balance betwe

model-releasesarxiv-cs-cv
6 Aug 2026
Research

An Analysis and Implementation of Seam Carving for Content-Aware Image Resizing

DGX agent

arXiv:2608.04329v1 Announce Type: new Abstract: Seam carving is a classical content-aware image resizing operator that modifies the width or height of an image by repeatedly removing (or inserting) se

researcharxiv-cs-cv
6 Aug 2026
Model Releases

ArtChart: Faithful Artistic Chart Generation with Integrated Text Rendering

DGX agent

arXiv:2607.16060v2 Announce Type: replace Abstract: Artistic charts combine data visualization with expressive marks, textures, and typography, but they are difficult for image generators: an output i

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

Attention Fusion for Bridge Deck Delamination Detection

DGX agent

arXiv:2512.20113v4 Announce Type: replace Abstract: Subsurface delaminations in reinforced concrete bridge decks escape conventional visual inspection, and the two principal sensing techniques used to

safetyarxiv-cs-cv
6 Aug 2026
Research

Bag-of-Visual-Words for Spatial Mapping of Lung Adenocarcinoma Growth Patterns

DGX agent

arXiv:2608.05074v1 Announce Type: new Abstract: Spatial mapping of lung adenocarcinoma (LUAD) growth patterns across whole slide images (WSIs) requires resolving architectural context at the region le

researcharxiv-cs-cv
6 Aug 2026
Research

Beyond Boundary Frames: Talking-Head Inbetweening via Context-Aware Motion Modeling

DGX agent

arXiv:2512.03590v3 Announce Type: replace Abstract: Existing talking-head generation methods primarily target open-ended generation rather than bridging two existing video segments. In this paper, we

researcharxiv-cs-cv
6 Aug 2026
Research

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

DGX agent

arXiv:2608.04454v1 Announce Type: new Abstract: Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full exp

researcharxiv-cs-cv
6 Aug 2026
Local Ai

Beyond Motion Cues and Structural Sparsity: Revisiting Small Moving Target Detection

DGX agent

arXiv:2509.07654v2 Announce Type: replace Abstract: Small moving target detection is crucial for many defense applications but remains highly challenging due to low signal-to-noise ratios, ambiguous v

local-aiarxiv-cs-cv
6 Aug 2026
Research

Beyond Reprojection Error: Camera Calibration with 3D Targets

DGX agent

arXiv:2608.05066v1 Announce Type: new Abstract: In 3D reconstruction, camera calibration is an essential element for achieving high fidelity and accuracy of the reconstructed geometry. While existing

researcharxiv-cs-cv
6 Aug 2026
Model Releases

BIM-Native Tokenization for Constraint-Aware Room Layout Synthesis

DGX agent

arXiv:2512.04832v3 Announce Type: replace Abstract: We present a BIM-native tokenization for room-level layout synthesis in Building Information Modeling (BIM) scenes. The core contribution is represe

model-releasesarxiv-cs-cv
6 Aug 2026
Agents

Binding Biometrics with AI Agent Identifiers for Delegation of Authority

DGX agent

arXiv:2608.04292v1 Announce Type: new Abstract: The proliferation of agentic artificial intelligence (AI) systems has raised serious questions about the accountability for tasks performed by AI agents

agentsarxiv-cs-cv
6 Aug 2026
Safety

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

DGX agent

arXiv:2608.04302v1 Announce Type: new Abstract: Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate acc

safetyarxiv-cs-cv
6 Aug 2026
Model Releases

CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision

DGX agent

arXiv:2512.22969v2 Announce Type: replace Abstract: Conventional object detectors rely on cross-entropy classification, which can be vulnerable to class imbalance and label noise. We propose CLIP-Join

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

CoCo-IR: Contextual Composed Image Retrieval

DGX agent

arXiv:2608.05149v1 Announce Type: new Abstract: Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of compl

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention

DGX agent

arXiv:2608.04396v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override p

safetyarxiv-cs-cv
6 Aug 2026
Research

ColorFD: A Finite-Difference Guided Black-Box Physical Adversarial Attack for Remote Sensing Object Detection

DGX agent

arXiv:2608.04559v1 Announce Type: new Abstract: Although deep neural network-based remote sensing object detectors have achieved strong performance, they remain vulnerable to adversarial perturbations

researcharxiv-cs-cv
6 Aug 2026
Hardware

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

DGX agent

arXiv:2608.04956v1 Announce Type: new Abstract: Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate op

hardwarearxiv-cs-cv
6 Aug 2026
Model Releases

Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen

DGX agent

arXiv:2608.04865v1 Announce Type: new Abstract: Event cameras, also known as neuromorphic cameras, have gained significant attention in recent years due to their high temporal resolution, high dynamic

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

COSMO: Consensus-Driven Shift Modulation for Source-Free Domain Adaptation

DGX agent

arXiv:2608.04604v1 Announce Type: new Abstract: Source-free domain adaptation (SFDA) adapts a source-trained model to an unlabeled target domain without source data, a practical setting under privacy

safetyarxiv-cs-cv
6 Aug 2026
Model Releases

Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution

DGX agent

arXiv:2608.04525v1 Announce Type: new Abstract: Scene text image super-resolution (STISR) aims to recover visually plausible appearance while preserving character semantics from degraded inputs. Exist

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation

DGX agent

arXiv:2608.04622v1 Announce Type: new Abstract: AI agents have emerged as a powerful new paradigm in generative image synthesis, enabling systems to perform complex semantic reasoning rather than pass

safetyarxiv-cs-cv
6 Aug 2026
Tutorials

Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors

DGX agent

arXiv:2608.04737v1 Announce Type: new Abstract: Direct Time-of-Flight (dToF) sensors provide highly accurate metric depth and are more robust than indirect ToF systems in challenging real-world condit

tutorialsarxiv-cs-cv
6 Aug 2026
Model Releases

Differential 6-DOF Pose Estimation with Provable First-Order Immunity to Camera Calibration Errors

DGX agent

arXiv:2608.04673v1 Announce Type: new Abstract: Accurate six-degree-of-freedom (6-DOF) motion estimation is essential for robotic manipulation, autonomous systems, and structural displacement monitori

model-releasesarxiv-cs-cv
6 Aug 2026
Research

DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models

DGX agent

arXiv:2608.04496v1 Announce Type: new Abstract: Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottl

researcharxiv-cs-cv
6 Aug 2026
← Previous
1…1112131415…259
Next →