AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning

DGX agent

arXiv:2605.15523v1 Announce Type: new Abstract: Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely s

researcharxiv-cs-cv
18 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Self-Supervised ImageNet Representations for In Vivo Confocal Microscopy: Tortuosity Grading without Segmentation Maps

DGX agent

arXiv:2603.15269v2 Announce Type: replace Abstract: The tortuosity of corneal nerve fibers are used as indication for different diseases. Current state-of-the-art methods for grading the tortuosity he

researcharxiv-cs-cv
18 May 2026
Safety

Self-Supervised Learning by Curvature Alignment

DGX agent

arXiv:2511.17426v2 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) has recently advanced through non-contrastive methods that couple an invariance term with variance, covariance,

safetyarxiv-cs-cv
18 May 2026
Safety

Semi-MedRef: Semi-Supervised Medical Referring Image Segmentation with Cross-Modal Alignment

DGX agent

arXiv:2605.15720v1 Announce Type: new Abstract: Medical referring image segmentation (MRIS) requires pixel-level masks aligned with textual descriptions of anatomical locations, making annotation cost

safetyarxiv-cs-cv
18 May 2026
Research

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting

DGX agent

arXiv:2511.18127v2 Announce Type: replace Abstract: Real-time 3D hand forecasting is a critical component for fluid human-computer interaction in applications like AR and assistive robotics. However,

researcharxiv-cs-cv
18 May 2026
Model Releases

SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization

DGX agent

arXiv:2603.08063v3 Announce Type: replace Abstract: Cross-view UAV geolocalization is fundamentally a challenging large-scale image retrieval task, aiming to determine the geographic coordinates of Un

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning

DGX agent

arXiv:2512.15693v2 Announce Type: replace Abstract: The misuse of AI-driven video generation technologies has raised serious social concerns, highlighting the urgent need for reliable AI-generated vid

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models

DGX agent

arXiv:2605.15424v1 Announce Type: new Abstract: Human trajectory forecasting is crucial for safe navigation in crowded environments, requiring models that balance accuracy with computational efficienc

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval

DGX agent

arXiv:2605.15868v1 Announce Type: new Abstract: In this work, we address the critical yet underexplored challenge of symmetric multimodal-to-multimodal (MM2MM) retrieval, where queries and contexts ar

model-releasesarxiv-cs-cv
18 May 2026
Local Ai

Sound Sparks Motion: Audio and Text Tuning for Video Editing

DGX agent

arXiv:2605.15307v1 Announce Type: cross Abstract: Motion-centric video editing remains difficult for large generative video models, which often respond well to appearance changes but struggle to produ

local-aiarxiv-cs-cv
18 May 2026
Model Releases

Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning

DGX agent

arXiv:2601.12894v2 Announce Type: replace-cross Abstract: Diffusion Policy has dominated action generation due to its strong capabilities for modeling multi-modal action distributions, but its multi-s

model-releasesarxiv-cs-cv
18 May 2026
Research

Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models

DGX agent

arXiv:2605.15961v1 Announce Type: new Abstract: Large-scale pre-trained vision-language models like CLIP demonstrate remarkable zero-shot performance across diverse tasks. However, fine-tuning these m

researcharxiv-cs-cv
18 May 2026
Safety

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System

DGX agent

arXiv:2605.16137v1 Announce Type: new Abstract: Generating simulation-ready tabletop scenes from task instructions is an intriguing and promising research direction in the field of Embodied AI. Howeve

safetyarxiv-cs-cv
18 May 2026
Model Releases

StippleDiffusion: Capacity-Constrained Stippling using Controlled Diffusion

DGX agent

arXiv:2605.15816v1 Announce Type: cross Abstract: Stipple patterns, point sets whose local density tracks a target image, are traditionally produced by per-density iterative optimizers, which are slow

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

SynthRender and IRIS: Open-Source Framework and Dataset for Bidirectional Sim-Real Transfer in Industrial Object Perception

DGX agent

arXiv:2602.21141v2 Announce Type: replace Abstract: Object perception is fundamental for tasks such as robotic material handling and quality inspection. However, modern supervised deep-learning models

model-releasesarxiv-cs-cv
18 May 2026
Research

Text-RSIR: A Text-Guided Framework for Efficient Remote Sensing Image Transmission and Reconstruction

DGX agent

arXiv:2605.15558v1 Announce Type: cross Abstract: High-resolution remote sensing imagery is critical for environmental monitoring, urban mapping, and land cover analysis, but its transmission is often

researcharxiv-cs-cv
18 May 2026
Model Releases

TSBOW -- Traffic Surveillance Benchmark for Occluded Vehicles Under Various Weather Conditions

DGX agent

arXiv:2602.05414v2 Announce Type: replace Abstract: Global warming has intensified the frequency and severity of extreme weather events, which degrade CCTV signal and video quality while disrupting tr

model-releasesarxiv-cs-cv
18 May 2026
Applications

TVRN: Invertible Neural Networks for Compression-Aware Temporal Video Rescaling

DGX agent

arXiv:2605.15579v1 Announce Type: cross Abstract: To fit diverse display and bandwidth constraints, high-frame-rate videos are temporally downscaled to low-frame-rate (LFR) and later upscaled, requiri

applicationsarxiv-cs-cv
18 May 2026
Research

U-SEG: Uncertainty in SEGmentation -- A systematic multi-variable exploration

DGX agent

arXiv:2605.15421v1 Announce Type: new Abstract: In this study, we explore in depth a few under-studied topics at the intersection of uncertainty estimation and segmentation. Prior work has shown that

researcharxiv-cs-cv
18 May 2026
Model Releases

Unlocking Dense Metric Depth Estimation in VLMs

DGX agent

arXiv:2605.15876v1 Announce Type: new Abstract: Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text

model-releasesarxiv-cs-cv
18 May 2026
Research

Unsupervised 3D Human Pose Estimation via Conditional Multi-view Ancestral Sampling

DGX agent

arXiv:2605.15583v1 Announce Type: new Abstract: We propose a method of estimating a 3D human pose from a single view without 3D supervision. The key to our method is to leverage the 2D diffusion prior

researcharxiv-cs-cv
18 May 2026
Safety

Video Models Can Reason with Verifiable Rewards

DGX agent

arXiv:2605.15458v1 Announce Type: new Abstract: Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generati

safetyarxiv-cs-cv
18 May 2026
Model Releases

VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?

DGX agent

arXiv:2510.08398v4 Announce Type: replace Abstract: The recent rapid advancement of Text-to-Video (T2V) generation technologies are engaging the trained models with more world model ability, making th

model-releasesarxiv-cs-cv
18 May 2026
Research

ViewBridge: Curriculum Knowledge Distillation for Activity View-Invariance Under Extreme Viewpoint Changes

DGX agent

arXiv:2504.05451v2 Announce Type: replace Abstract: Traditional methods for view-invariant learning rely on controlled multi-view training data with minimal scene clutter. However, they struggle with

researcharxiv-cs-cv
18 May 2026
Research

Visual Compositional Tuning

DGX agent

arXiv:2504.21850v3 Announce Type: replace Abstract: Visual instruction tuning (VIT) datasets have grown rapidly in scale, yet the informativeness of individual training samples has largely been overlo

researcharxiv-cs-cv
18 May 2026
Model Releases

WeatherOcc3D: VLM-Assisted Adverse Weather Aware 3D Semantic Occupancy Prediction

DGX agent

arXiv:2605.16127v1 Announce Type: new Abstract: While multi-modal 3D semantic occupancy prediction typically enhances robustness by fusing camera and LiDAR inputs, its effectiveness is fundamentally c

model-releasesarxiv-cs-cv
18 May 2026
Research

When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing

DGX agent

arXiv:2605.15484v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) networks promise favorable accuracy-compute trade-offs, yet practical vision deployments are hindered by expert collapse and li

researcharxiv-cs-cv
18 May 2026
Agents

Where to Perch in a Tree: Vision-Guidance for Tree-Grasping Drones

DGX agent

arXiv:2605.15430v1 Announce Type: cross Abstract: This study demonstrates a method to locate an ideal perch location on a tree for vision-guided autonomous tree-perching drones. Various image processi

agentsarxiv-cs-cv
18 May 2026
Agents

WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes

DGX agent

arXiv:2605.15843v1 Announce Type: new Abstract: Recent 3D world modeling systems based on generative scene synthesis, such as Marble, can create coherent and explorable 3D environments, yet their outp

agentsarxiv-cs-cv
18 May 2026
Agents

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

DGX agent

arXiv:2605.15964v1 Announce Type: cross Abstract: Aerial vision-language navigation (VLN) requires agents to follow natural-language instructions through closed-loop perception and action in 3D enviro

agentsarxiv-cs-cv
18 May 2026
Model Releases

3D Skew-Normal Splatting

DGX agent

arXiv:2605.15010v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as a leading representation for real-time novel view synthesis and been widely adopted in various downstream ap

model-releasesarxiv-cs-cv
15 May 2026
Research

A CUBS-Compatible Ultrasound Morphology and Uncertainty-Aware Baseline for Carotid Intima-Media Segmentation and Preliminary Risk Prediction

DGX agent

arXiv:2605.14949v1 Announce Type: new Abstract: Carotid atherosclerosis is a major contributor to ischemic stroke and transient ischemic attack. Conventional ultrasound assessment is commonly based on

researcharxiv-cs-cv
15 May 2026
Model Releases

ACE-LoRA: Adaptive Orthogonal Decoupling for Continual Image Editing

DGX agent

arXiv:2605.14948v1 Announce Type: new Abstract: State-of-the-art diffusion models often rely on parameter-efficient fine-tuning to perform specialized image editing tasks. However, real-world applicat

model-releasesarxiv-cs-cv
15 May 2026
Safety

Aligning Latent Geometry for Spherical Flow Matching in Image Generation

DGX agent

arXiv:2605.15193v1 Announce Type: new Abstract: Latent flow matching for image generation usually transports Gaussian noise to variational autoencoder latents along linear paths. Both endpoints, howev

safetyarxiv-cs-cv
15 May 2026
Research

Analogical Trajectory Transfer

DGX agent

arXiv:2605.14393v1 Announce Type: new Abstract: We study analogical trajectory transfer, where the goal is to translate motion trajectories in one 3D environment to a semantically analogous location i

researcharxiv-cs-cv
15 May 2026
Model Releases

AnchorRoute: Human Motion Synthesis with Interval-Routed Sparse Contro

DGX agent

arXiv:2605.14716v1 Announce Type: cross Abstract: Sparse anchors provide a compact interface for human motion authoring: users specify a few root positions, planar trajectory samples, or body-point ta

model-releasesarxiv-cs-cv
15 May 2026
Applications

Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

DGX agent

arXiv:2602.00807v2 Announce Type: replace Abstract: Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. H

applicationsarxiv-cs-cv
15 May 2026
Research

AnyBand-Diff: A Unified Remote Sensing Image Generation and Band Repair Framework with Spectral Priors

DGX agent

arXiv:2605.14341v1 Announce Type: new Abstract: Existing diffusion models have made significant progress in generating realistic images. However, their direct adaptation to remote sensing imagery ofte

researcharxiv-cs-cv
15 May 2026
Research

ArcGate: Adaptive Arctangent Gated Activation

DGX agent

arXiv:2605.14518v1 Announce Type: new Abstract: Activation functions are central to deep networks, influencing non-linearity, feature learning, convergence, and robustness. This paper proposes the Ada

researcharxiv-cs-cv
15 May 2026
Research

Architecture-Aware Explanation Auditing for Industrial Visual Inspection

DGX agent

arXiv:2605.14255v1 Announce Type: cross Abstract: Industrial visual inspection systems increasingly rely on deep classifiers whose heatmap explanations may appear visually plausible while failing to i

researcharxiv-cs-cv
15 May 2026
Research

Are Candidate Models Really Needed for Active Learning?

DGX agent

arXiv:2605.14689v1 Announce Type: new Abstract: Deep learning has profoundly impacted domains such as computer vision and natural language processing by uncovering complex patterns in vast datasets. H

researcharxiv-cs-cv
15 May 2026
Agents

Articraft: An Agentic System for Scalable Articulated 3D Asset Generation

DGX agent

arXiv:2605.15187v1 Announce Type: new Abstract: A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large lan

agentsarxiv-cs-cv
15 May 2026
Research

AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting

DGX agent

arXiv:2506.01015v2 Announce Type: replace Abstract: Segment Anything Model 2 (SAM2) exhibits strong generalisation for promptable segmentation in video clips; however, its integration with the audio m

researcharxiv-cs-cv
15 May 2026
Local Ai

Automatic Landmark-Based Segmentation of Human Subcortical Structures in MRI

DGX agent

arXiv:2605.14221v1 Announce Type: new Abstract: Precise segmentation of brain structures in magnetic resonance imaging (MRI) is essential for reliable neuroimaging analysis, yet voxel-wise deep models

local-aiarxiv-cs-cv
15 May 2026
Safety

AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving

DGX agent

arXiv:2603.14851v3 Announce Type: replace Abstract: Integrating vision-language models (VLMs) into end-to-end (E2E) autonomous driving (AD) systems has shown promise in improving scene understanding.

safetyarxiv-cs-cv
15 May 2026
Safety

Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control

DGX agent

arXiv:2605.14417v1 Announce Type: cross Abstract: Natural language is an intuitive interface for humanoid robots, yet streaming whole-body control requires control representations that are executable

safetyarxiv-cs-cv
15 May 2026
Safety

Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging

DGX agent

arXiv:2605.14654v1 Announce Type: new Abstract: Self-supervised pre-training methods in medical imaging typically treat each individual as an isolated instance, learning representations through augmen

safetyarxiv-cs-cv
15 May 2026
Model Releases

BioHuman: Learning Biomechanical Human Representations from Video

DGX agent

arXiv:2605.14772v1 Announce Type: new Abstract: Understanding human motion beyond surface kinematics is crucial for motion analysis, rehabilitation, and injury risk assessment. However, progress in th

model-releasesarxiv-cs-cv
15 May 2026
← Previous
1…167168169170171…263
Next →