AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
4 Aug 2026

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision

SafetyDGX agent

arXiv:2608.01392v1 Announce Type: new Abstract: Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets wit

Foveated Probes Recover Localized Binding Information in Vision Foundation Models

ResearchDGX agent

arXiv:2608.00726v1 Announce Type: new Abstract: Frozen vision foundation models are commonly evaluated through a single global image embedding, but this interface can conflate missing information with

FreqAnchorAD: Language-Free Zero-Shot Anomaly Detection via Frequency-Deviation Anchoring

Local AiDGX agent

arXiv:2608.00695v1 Announce Type: new Abstract: Zero-shot anomaly detection (ZSAD) aims to detect anomalies and localize defective regions in unseen target domains without target training data. Recent


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

From Forest to Future Capital: Tracking Land Cover Change in Ibu Kota Nusantara (IKN) from 2021 to 2026 with PlanetScope Imagery

ResearchDGX agent

arXiv:2608.01230v1 Announce Type: new Abstract: Indonesia's relocation of its political and administrative capital from Jakarta to Ibu Kota Nusantara (IKN) has been framed around a ``Forest City'' vis

From Patches to Evidence Balls: Class-Conditioned Evidence Retrieval for Few-Shot Whole Slide Image Classification

TutorialsDGX agent

arXiv:2608.01104v1 Announce Type: new Abstract: Whole slide image (WSI) classification is an evidence-driven task, where diagnostic cues are often sparse, spatially organized, and class-dependent. Exi

From Pixels to PCells: A Neurosymbolic Approach to Photonic Component Creation

ResearchDGX agent

arXiv:2608.00084v1 Announce Type: new Abstract: We present PixCell, a neurosymbolic system in which multimodal agents convert a visually presented photonic component into a parametric program over a s

From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving

Model ReleasesDGX agent

arXiv:2602.10719v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbones, yet it remains unclear how vis

Fruit-HSNet: A Machine Learning Approach for Hyperspectral Image-Based Fruit Ripeness Prediction

ApplicationsDGX agent

arXiv:2608.01202v1 Announce Type: new Abstract: Fruit ripeness prediction (FRP) is a classification-based agricultural computer vision task that has attracted much attention, thanks to its wide-rangin

FusionRS: A Large-Scale RGB-Infrared-Style Remote Sensing Dataset for Cross-Modal Vision-Language Learning

SafetyDGX agent

arXiv:2606.17020v2 Announce Type: replace Abstract: Remote sensing vision-language models have advanced Earth observation, but available large-scale vision-language resources remain RGB-centered, leav

G-Skin: Learning to Bind 3D Gaussians with Generative Visual Priors

ResearchDGX agent

arXiv:2608.01726v1 Announce Type: new Abstract: 3D Gaussian Splatting has achieved remarkable success in photorealistic and efficient rendering, leading to a rapid increase in 3D assets represented by

GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization

ApplicationsDGX agent

arXiv:2608.01492v1 Announce Type: new Abstract: Selecting a complete 3D object from a reconstructed scene with minimal user effort is essential for practical scene editing and embodied interaction. Ex

Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection

ResearchDGX agent

arXiv:2608.00716v1 Announce Type: new Abstract: Robust detection of generated images is critical to counter the misuse of generative models. Existing methods primarily depend on learning from human-an

Generative AI and Foundation Models in Medical Image

Model ReleasesDGX agent

arXiv:2608.01686v1 Announce Type: new Abstract: In recent years, generative AI has attracted significant public attention, and its use has been rapidly expanding across a wide range of domains. From c

Generative Brownian Bridge Diffusion In Motion Space For Enhanced Myocardial Strain Analysis

TutorialsDGX agent

arXiv:2608.01677v1 Announce Type: new Abstract: Myocardial strain analysis of cardiac magnetic resonance (CMR) images provides an important tool for evaluating cardiac function. However, current techn

GenPrior: Unleashing Text-to-Motion Generative Priors for Zero-Shot Skeleton-based Action Recognition

ResearchDGX agent

arXiv:2608.02236v1 Announce Type: new Abstract: Zero-shot skeleton-based action recognition (ZSAR) aims to recognize unseen action categories by aligning skeleton features with textual semantics. Howe

GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking

SafetyDGX agent

arXiv:2608.01410v1 Announce Type: cross Abstract: General-purpose humanoid trackers can execute diverse references, but their zero-shot coverage depends on large embodied corpora that are costly to ex

GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation

Model ReleasesDGX agent

arXiv:2608.01896v1 Announce Type: new Abstract: Existing generative models for earth observation (EO) predominantly rely on fine-tuning natural image priors, which limits their scalability and introdu

GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation

Model ReleasesDGX agent

arXiv:2608.02315v1 Announce Type: new Abstract: Geospatial foundation models aim to learn representations that transfer across regions and sensors, yet evaluating them on specific tasks requires large

Geometric-Topological Perception and Motion Prior for Real-Time Satellite Video Object Tracking

Local AiDGX agent

arXiv:2603.07564v2 Announce Type: replace Abstract: Satellite video object tracking (SVOT) remains fundamentally challenging due to texture scarcity, arbitrary rotation, aspect ratio changes, and seve

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

ResearchDGX agent

arXiv:2608.00663v1 Announce Type: new Abstract: Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle t

GIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2608.02068v1 Announce Type: new Abstract: Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hal

Gimbal360: Canonicalizing Planar Diffusion for Spherical Panorama Completion

ResearchDGX agent

arXiv:2603.23179v2 Announce Type: replace Abstract: Diffusion models provide powerful priors for 2D image completion, but these priors are learned on bounded planar images and do not transfer directly

Global-Scale Self-Supervised Spatiotemporal Learning for NDVI Time-Series Reconstruction

ApplicationsDGX agent

arXiv:2608.02322v1 Announce Type: new Abstract: Accurate and efficient reconstruction of cloud-contaminated and noise-corrupted NDVI time series remains a challenge in remote sensing. Deep learning pr

GraRe: Grasp Candidate Re-Ranking for Frozen 6-DoF Grasp Detectors

ResearchDGX agent

arXiv:2608.00946v1 Announce Type: cross Abstract: Existing 6-DoF grasp detectors typically rank grasp candidates by detector confidence. However, our analysis on GraspNet-1Billion shows that detector

Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering

ResearchDGX agent

arXiv:2608.01660v1 Announce Type: new Abstract: Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-to

Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment

Model ReleasesDGX agent

arXiv:2608.02470v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed as reasoning agents in real-world visual assessment pipelines, yet their spatial grounding remai

Grounding and Explaining Visual Evidence for AI-Generated Image Detection in Human-Centric Scenes

Model ReleasesDGX agent

arXiv:2608.01988v1 Announce Type: new Abstract: Rapid advances in image generation models call for interpretable AI-generated image detection methods that not only determine authenticity but also prov

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience

ResearchDGX agent

arXiv:2608.02392v1 Announce Type: new Abstract: A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing

GSRAIN: Physically Calibrated High-/Low-Frequency Rainfall Synthesis for 3D Gaussian Driving Scenes

AgentsDGX agent

arXiv:2608.02177v1 Announce Type: new Abstract: Existing rainfall simulation methods for autonomous driving remain limited in physical controllability and multi-view consistency. This paper presents G

GuideGround: VLM-guided Semantic Understanding and Viewpoint-aware Reasoning for 3D Visual Grounding

Model ReleasesDGX agent

arXiv:2608.00518v1 Announce Type: new Abstract: 3D visual grounding aims to localize the target object in a 3D scene from a natural language query, requiring both fine-grained semantic understanding a

HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts

Model ReleasesDGX agent

arXiv:2608.02252v1 Announce Type: new Abstract: Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the d

Harnessing Adversarial Distillation to Customise Debiased, Disease-Specific Pathology Foundation Models for Breast Cancer

Model ReleasesDGX agent

arXiv:2608.01356v1 Announce Type: new Abstract: Pathology foundation models (PFMs) provide strong tissue representations and have become central to digital pathology. However, deployment in disease-sp

Hermite Curves as Trajectory Priors for Vision-Language-Action Models

SafetyDGX agent

arXiv:2608.01265v1 Announce Type: cross Abstract: Despite recent progress in Vision-Language-Action (VLA) models for robotic manipulation, the action chunk remains a weakly structured interface. Exist

Hi-TOPS: Hierarchical Topology-aware Scoring Prior for 3D Part Decomposition

ResearchDGX agent

arXiv:2608.00767v1 Announce Type: cross Abstract: Accurate 3D part decomposition requires separating shapes into structurally meaningful components with precise boundaries while preserving articulatio

HiResNets: Native Full-HD Video Recognition with Foveal Residual Streams

TutorialsDGX agent

arXiv:2608.02140v1 Announce Type: new Abstract: Much of the recent progress in image and video recognition has come at the cost of memory: larger models, increased resolution, and longer temporal cont

HorusEye: Language as Dynamic Attention for Emergency Visual Analysis

Model ReleasesDGX agent

arXiv:2606.14741v2 Announce Type: replace Abstract: We introduce HorusEye, Language as Dynamic Attention for Emergency Visual Analysis. Our investigation followed five stages. The first one is benchma

Human-like working memory signatures emerge from intrinsically plastic artificial neurons for robust dynamic vision

AgentsDGX agent

arXiv:2512.15829v4 Announce Type: replace-cross Abstract: While the unsustainable energy cost of artificial intelligence necessitates physics-driven computing, its performance superiority over full-pr

Hybrid-Domain Posterior Sampling for Inverse Problems via Latent Flow Matching

SafetyDGX agent

arXiv:2608.00537v1 Announce Type: new Abstract: Latent Flow Models have revolutionized compressed-space image synthesis, yet their application to high-fidelity inverse problems remains bottlenecked. I

HyperGS: Fast and Generalizable Gaussian Video Representation

Model ReleasesDGX agent

arXiv:2607.11500v2 Announce Type: replace Abstract: Gaussian Splatting has emerged as an effective representation for video, but existing methods rely on per-video optimization. This leads to slow enc

IDraw: Artist Verification from Digital Drawing Images

ResearchDGX agent

arXiv:2608.01737v1 Announce Type: new Abstract: As digital drawings are increasingly shared online, reliable authorship verification has become important for protecting artists and resolving disputes.

Image-Space Rule Discovery

Model ReleasesDGX agent

arXiv:2608.00490v1 Announce Type: new Abstract: Can image-editing models discover visual rules in image space and complete problem-solving end-to-end? We tackle this question in the spirit of a human

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation

TutorialsDGX agent

arXiv:2511.20635v3 Announce Type: replace Abstract: Pre-trained video models learn powerful priors for generating high-quality, temporally coherent content. While these models excel at temporal cohere

Implicit Neural Representations for Multimodal Longitudinal Image Imputation and Interpolation

ApplicationsDGX agent

arXiv:2608.02324v1 Announce Type: new Abstract: Longitudinal multiparametric MRI is central to follow-up imaging in oncology, yet real-world clinical data are characterised by missing sequences, heter

InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

ResearchDGX agent

arXiv:2608.02437v1 Announce Type: new Abstract: Single-image feed-forward 3D Gaussian Splatting (3DGS) aims to directly generate a renderable 3D scene representation from one input image, avoiding the

InstancePin: Instance-Addressable Layout-to-Image Diffusion via Coordinate Pinning

ResearchDGX agent

arXiv:2608.00588v1 Announce Type: new Abstract: Layout-to-image diffusion models have achieved impressive semantic controllability by conditioning generation on category-level segmentation maps. Howev

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

Model ReleasesDGX agent

arXiv:2608.01157v1 Announce Type: new Abstract: Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by

Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers

TutorialsDGX agent

arXiv:2608.00264v1 Announce Type: new Abstract: Vision foundation models, such as DINOv2, learn highly expressive representations but rely on massive, opaque architectures that demand substantial comp

Investigating Social Bias in Narrative Image Generation

SafetyDGX agent

arXiv:2608.01780v1 Announce Type: new Abstract: Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how

Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents

SafetyDGX agent

arXiv:2608.02018v1 Announce Type: new Abstract: Computer-use agents (CUAs), which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to in

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models

ResearchDGX agent

arXiv:2607.15732v2 Announce Type: replace Abstract: Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs boundin

ISRS-DETR: Detection-Guided Click Propagation for Remote Sensing Interactive Segmentation

Model ReleasesDGX agent

arXiv:2608.02468v1 Announce Type: new Abstract: Interactive segmentation reduces the prohibitive cost of pixel-level annotation by allowing users to delineate objects with a few clicks. However, apply

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling

Model ReleasesDGX agent

arXiv:2608.01207v1 Announce Type: new Abstract: Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers poor

JADE-GS: Joint Allocation of Deblurring Evidence for Event-Assisted 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2607.14990v2 Announce Type: replace Abstract: Neural radiance fields and 3D Gaussian Splatting assume that each training image is a sharp and geometrically consistent observation of the scene. M

K-space Gaussian Representation for Parallel MRI

Local AiDGX agent

arXiv:2608.00075v1 Announce Type: new Abstract: Accelerated magnetic resonance imaging (MRI) aims to recover the k-space signal from acquired measurements, where accurate estimation of missing samples

Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving

Model ReleasesDGX agent

arXiv:2608.00237v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving, enabling agents to map multimodal inputs

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

ResearchDGX agent

arXiv:2608.00079v1 Announce Type: new Abstract: Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits strea

Learning How Much, Not Just What: Cross-Patient Burden Order for CT Vision-Language Pretraining

SafetyDGX agent

arXiv:2608.00231v1 Announce Type: new Abstract: Volumetric CT vision-language pretraining learns 3D representations from scan-report pairs, but global and anatomy-aware objectives supervise only corre

Learning to See Locally and Align Clinically with Pathology Semantics for Radiology Report Generation

SafetyDGX agent

arXiv:2608.00279v1 Announce Type: cross Abstract: Recent radiology-adapted vision-language models have achieved strong performance on standard report generation benchmarks, yet their robustness and ge

Learning to Tessellate: Point Cloud Generation via Recursive Spectral Partitioning

ResearchDGX agent

arXiv:2608.02432v1 Announce Type: new Abstract: Autoregressive models have emerged as an effective paradigm for point cloud generation. However, most existing approaches rely on heuristic tokenization

Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency

TutorialsDGX agent

arXiv:2608.01730v1 Announce Type: new Abstract: No-reference image quality assessment (NR IQA) has recently benefited from deep and multimodal models, yet many SOTA systems still violate at least one

← Previous
1…1516171819…207
Next →