AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
6 Aug 2026

An active-learning framework for real-time depth perception from monocular vision streams

Model ReleasesDGX agent

arXiv:2608.04917v1 Announce Type: new Abstract: Biological visual systems can perceive depth from monocular vision flow, continuously integrating temporal visual cues while maintaining a balance betwe

An Analysis and Implementation of Seam Carving for Content-Aware Image Resizing

ResearchDGX agent

arXiv:2608.04329v1 Announce Type: new Abstract: Seam carving is a classical content-aware image resizing operator that modifies the width or height of an image by repeatedly removing (or inserting) se

ArtChart: Faithful Artistic Chart Generation with Integrated Text Rendering

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.16060v2 Announce Type: replace Abstract: Artistic charts combine data visualization with expressive marks, textures, and typography, but they are difficult for image generators: an output i

Attention Fusion for Bridge Deck Delamination Detection

SafetyDGX agent

arXiv:2512.20113v4 Announce Type: replace Abstract: Subsurface delaminations in reinforced concrete bridge decks escape conventional visual inspection, and the two principal sensing techniques used to

Bag-of-Visual-Words for Spatial Mapping of Lung Adenocarcinoma Growth Patterns

ResearchDGX agent

arXiv:2608.05074v1 Announce Type: new Abstract: Spatial mapping of lung adenocarcinoma (LUAD) growth patterns across whole slide images (WSIs) requires resolving architectural context at the region le

Beyond Boundary Frames: Talking-Head Inbetweening via Context-Aware Motion Modeling

ResearchDGX agent

arXiv:2512.03590v3 Announce Type: replace Abstract: Existing talking-head generation methods primarily target open-ended generation rather than bridging two existing video segments. In this paper, we

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

ResearchDGX agent

arXiv:2608.04454v1 Announce Type: new Abstract: Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full exp

Beyond Motion Cues and Structural Sparsity: Revisiting Small Moving Target Detection

Local AiDGX agent

arXiv:2509.07654v2 Announce Type: replace Abstract: Small moving target detection is crucial for many defense applications but remains highly challenging due to low signal-to-noise ratios, ambiguous v

Beyond Reprojection Error: Camera Calibration with 3D Targets

ResearchDGX agent

arXiv:2608.05066v1 Announce Type: new Abstract: In 3D reconstruction, camera calibration is an essential element for achieving high fidelity and accuracy of the reconstructed geometry. While existing

BIM-Native Tokenization for Constraint-Aware Room Layout Synthesis

Model ReleasesDGX agent

arXiv:2512.04832v3 Announce Type: replace Abstract: We present a BIM-native tokenization for room-level layout synthesis in Building Information Modeling (BIM) scenes. The core contribution is represe

Binding Biometrics with AI Agent Identifiers for Delegation of Authority

AgentsDGX agent

arXiv:2608.04292v1 Announce Type: new Abstract: The proliferation of agentic artificial intelligence (AI) systems has raised serious questions about the accountability for tasks performed by AI agents

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

SafetyDGX agent

arXiv:2608.04302v1 Announce Type: new Abstract: Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate acc

CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision

Model ReleasesDGX agent

arXiv:2512.22969v2 Announce Type: replace Abstract: Conventional object detectors rely on cross-entropy classification, which can be vulnerable to class imbalance and label noise. We propose CLIP-Join

CoCo-IR: Contextual Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2608.05149v1 Announce Type: new Abstract: Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of compl

CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention

SafetyDGX agent

arXiv:2608.04396v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override p

ColorFD: A Finite-Difference Guided Black-Box Physical Adversarial Attack for Remote Sensing Object Detection

ResearchDGX agent

arXiv:2608.04559v1 Announce Type: new Abstract: Although deep neural network-based remote sensing object detectors have achieved strong performance, they remain vulnerable to adversarial perturbations

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

HardwareDGX agent

arXiv:2608.04956v1 Announce Type: new Abstract: Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate op

Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen

Model ReleasesDGX agent

arXiv:2608.04865v1 Announce Type: new Abstract: Event cameras, also known as neuromorphic cameras, have gained significant attention in recent years due to their high temporal resolution, high dynamic

COSMO: Consensus-Driven Shift Modulation for Source-Free Domain Adaptation

SafetyDGX agent

arXiv:2608.04604v1 Announce Type: new Abstract: Source-free domain adaptation (SFDA) adapts a source-trained model to an unlabeled target domain without source data, a practical setting under privacy

Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution

Model ReleasesDGX agent

arXiv:2608.04525v1 Announce Type: new Abstract: Scene text image super-resolution (STISR) aims to recover visually plausible appearance while preserving character semantics from degraded inputs. Exist

DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation

SafetyDGX agent

arXiv:2608.04622v1 Announce Type: new Abstract: AI agents have emerged as a powerful new paradigm in generative image synthesis, enabling systems to perform complex semantic reasoning rather than pass

Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors

TutorialsDGX agent

arXiv:2608.04737v1 Announce Type: new Abstract: Direct Time-of-Flight (dToF) sensors provide highly accurate metric depth and are more robust than indirect ToF systems in challenging real-world condit

Differential 6-DOF Pose Estimation with Provable First-Order Immunity to Camera Calibration Errors

Model ReleasesDGX agent

arXiv:2608.04673v1 Announce Type: new Abstract: Accurate six-degree-of-freedom (6-DOF) motion estimation is essential for robotic manipulation, autonomous systems, and structural displacement monitori

DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models

ResearchDGX agent

arXiv:2608.04496v1 Announce Type: new Abstract: Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bottl

DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture

ResearchDGX agent

arXiv:2511.17354v4 Announce Type: replace Abstract: Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for lear

EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation

Model ReleasesDGX agent

arXiv:2608.04533v1 Announce Type: new Abstract: Part-level affordance grounding has advanced the localization of functional object regions associated with elemental actions. Extending this capability

Embedding Large Language Models into Flow Controls: An Agentic Framework for Adaptive and Trustworthy Automated Cooking

AgentsDGX agent

arXiv:2608.04768v1 Announce Type: new Abstract: Automated cooking robots have traditionally relied on predefined procedures and rule-based control, ensuring stable execution but offering limited perso

Enhancing Low Back Pain Assessment with Diffusion Models for Lumbar Spine MRI Segmentation

ResearchDGX agent

arXiv:2608.04906v1 Announce Type: new Abstract: This study introduces a diffusion-based framework for robust and accurate semantic segmentation of lumbar spine MRI scans from patients with low back pa

Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting

ResearchDGX agent

arXiv:2607.15890v2 Announce Type: replace Abstract: Perceiving multimodal cues and forecasting fine-grained actions from an egocentric (Ego) perspective is vital for applications like robot manipulati

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

Model ReleasesDGX agent

arXiv:2608.04404v1 Announce Type: new Abstract: World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approach

Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation

SafetyDGX agent

arXiv:2602.19161v2 Announce Type: replace Abstract: Latent diffusion models have enabled high-quality video synthesis, yet their inference remains costly and time-consuming. As diffusion transformers

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

SafetyDGX agent

arXiv:2608.04530v1 Announce Type: new Abstract: GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact so

Foreseeing the Invisible: Amodal Reconstruction of Leaf Fossil Images

Model ReleasesDGX agent

arXiv:2608.04423v1 Announce Type: new Abstract: Fossil leaves are rarely preserved whole -- sedimentary rock hides, breaks, and erodes the lamina, yet paleobotany depends on the complete shape and out

Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection

ResearchDGX agent

arXiv:2608.04394v1 Announce Type: new Abstract: Cross-Domain Few-Shot Object Detection (CDFSOD) aims to transfer knowledge from data-rich upstream generic domains to downstream expert domains using sc

From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs

SafetyDGX agent

arXiv:2601.03808v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved notable performance in code synthesis; however, data-aware augmentation remains a limiting factor, handle

From Understanding to Erasing: Towards Complete and Stable Video Object Removal

Local AiDGX agent

arXiv:2604.01693v2 Announce Type: replace Abstract: Video object removal aims to erase target objects while reconstructing visually plausible and temporally coherent content. However, target objects o

Generative neural physics enables quantitative volumetric ultrasound of tissue mechanics

ResearchDGX agent

arXiv:2508.12226v3 Announce Type: replace Abstract: Ultrasound Tomography (UT) is a radiation-free, high-resolution modality, but remains limited for musculoskeletal imaging due to the high computatio

Global Attention-Fused Image Cropping with Attention-Guided and Global-Aligned Crop Evaluator

ResearchDGX agent

arXiv:2608.04821v1 Announce Type: new Abstract: Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods

HelloWorld: Enabling Socially Interactive Characters in Video World Models

Model ReleasesDGX agent

arXiv:2608.05070v1 Announce Type: new Abstract: Despite the remarkable recent progress of video world models, social interaction between users and the characters within these worlds remains unsupporte

HexMIL: Hierarchical Attention MIL for Ante-Hoc Explainable Detection of AI-Manipulated CT Volumes

ResearchDGX agent

arXiv:2608.05101v1 Announce Type: new Abstract: The emergence of medical deepfakes, i.e., medical images manipulated by deep generative models, poses a significant threat to clinical workflows. Howeve

HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding

ResearchDGX agent

arXiv:2608.04610v1 Announce Type: new Abstract: 3D vision-language models (3D VLMs) enable spatial reasoning over multi-view scenes but suffer from substantial token redundancy due to duplicated obser

Industrial Synthetic Segment Pre-training

ApplicationsDGX agent

arXiv:2505.13099v3 Announce Type: replace Abstract: Vision Foundation Models (VFMs) have made remarkable progress and are increasingly being applied to segmentation tasks in real-world industrial sett

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers

Local AiDGX agent

arXiv:2608.05122v1 Announce Type: new Abstract: Vision transformers (ViTs) have become the de facto standard for image encoding across many perception tasks. Despite their empirical success, it remain

Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation

Model ReleasesDGX agent

arXiv:2607.10082v2 Announce Type: replace Abstract: Building pixel-level correspondence between event and image data is a fundamental task for multi-sensor systems. However, existing cross-modal match

Lesion Detection in CT with Frozen Self-Distilled Features: SALT, a Spatially Adaptive Label-Guided Temperature

SafetyDGX agent

arXiv:2608.05100v1 Announce Type: new Abstract: Self-supervised pretraining objectives are spatially uniform: the teacher temperature and the per-patch loss weight are identical everywhere in the imag

LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content

Model ReleasesDGX agent

arXiv:2410.10783v4 Announce Type: replace Abstract: The large-scale training of multi-modal models on data scraped from the web has shown outstanding utility in infusing these models with the required

LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching

Model ReleasesDGX agent

arXiv:2608.04106v1 Announce Type: new Abstract: Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending d

MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding

AgentsDGX agent

arXiv:2608.04587v1 Announce Type: new Abstract: Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in

MME: Mixture of Mesh Experts with Random Walk Transformer Gating

ResearchDGX agent

arXiv:2603.00828v2 Announce Type: replace Abstract: In recent years, various methods have been proposed for mesh analysis, each offering distinct advantages and often excelling on different object cla

MOAT: Model-Agnostic Randomized Transformations for preventing Efficiency Degradation Attacks on ViTs

ResearchDGX agent

arXiv:2608.04680v1 Announce Type: cross Abstract: To adopt the Vision Transformers (ViTs) in resource-constrained environment, token pruning is widely used to reduce computational cost without impacti

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

Model ReleasesDGX agent

arXiv:2608.04657v1 Announce Type: new Abstract: World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mob

Multi-View Face and Gesture Animation with Dynamic Gaussians

ResearchDGX agent

arXiv:2608.04722v1 Announce Type: new Abstract: Creating photorealistic 3D human avatars with realistic upper-body motion remains challenging. Existing approaches either focus on the head and overlook

muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards

SafetyDGX agent

arXiv:2608.04412v1 Announce Type: new Abstract: High-quality driving data are essential for autonomous-driving systems and generative world models. However, rare and safety-critical scenarios involvin

MVTOP: Multi-View Transformer-based Object Pose-Estimation

ResearchDGX agent

arXiv:2508.03243v2 Announce Type: replace Abstract: We present MVTOP, a novel transformer-based method for multi-view rigid object pose estimation. Through an early fusion of the view-specific feature

Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles

SafetyDGX agent

arXiv:2608.04483v1 Announce Type: new Abstract: Vision-language models (VLMs) process an image as a sequence of visual tokens, which creates a substantial computational bottleneck during inference. Re

Objects as Audio-Visual Modal Sound Fields

ApplicationsDGX agent

arXiv:2608.05145v1 Announce Type: new Abstract: While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical in

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

Model ReleasesDGX agent

arXiv:2608.05049v1 Announce Type: new Abstract: Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmar

OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

Model ReleasesDGX agent

arXiv:2608.04434v1 Announce Type: new Abstract: Recent large language models (LLMs) have demonstrated remarkable progress in constraint-aware navigation, maze reasoning, and graph reasoning. However,

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films

Model ReleasesDGX agent

arXiv:2608.04224v1 Announce Type: new Abstract: Historical films suffer from co-occurring visual and audio degradations---blur, noise, flicker, hiss, clipping, and dropout---yet existing methods resto

On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing

Model ReleasesDGX agent

arXiv:2608.04791v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative training of deep learning models across decentralized image archives without requiring data centralization

← Previous
1…910111213…207
Next →