AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
4 Aug 2026

Less is More: Compact-Token Masked Feature Prediction for Skeleton Representation Learning

ResearchDGX agent

arXiv:2603.10648v3 Announce Type: replace Abstract: Current skeleton representation learning paradigms face distinct limitations: Contrastive Learning (CL) often overlooks fine-grained motion details,

Lethe: How Hard Is It to Forget? A Benchmark for Federated Unlearning in Medical Imaging

Model ReleasesDGX agent

arXiv:2608.01094v1 Announce Type: new Abstract: Federated learning enables medical-imaging models to be trained across hospitals, and privacy law, most explicitly the GDPR ``right to be forgotten'', t

Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2602.04789v4 Announce Type: replace Abstract: Advanced autoregressive (AR) video generation models have improved visual fidelity and interactivity, but the quadratic complexity of attention rema

Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge

Model ReleasesDGX agent

arXiv:2608.01614v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face a critical computational bottleneck when processing high-resolution imagery due to the O(N^2) memory complexity of So

Linguistic Context Recodes Visual Representations in Vision-Language Models

ResearchDGX agent

arXiv:2608.00035v1 Announce Type: cross Abstract: Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categor

LiveLight: Real-time Streaming Video Relighting with Interactive Control

ApplicationsDGX agent

arXiv:2608.01771v1 Announce Type: new Abstract: We present LiveLight, the first diffusion-based framework for real-time streaming video relighting with interactive 3D lighting control. Achieving this

Local Margin Restoration for Test-Time Adaptation of Vision-Language Models

Local AiDGX agent

arXiv:2608.02216v1 Announce Type: new Abstract: Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected

Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models

Local AiDGX agent

arXiv:2608.00976v1 Announce Type: new Abstract: Fine-grained visual representations are essential for medical image analysis, particularly when diagnostically relevant evidence is subtle and spatially

Loggia dei Lanzi: AI Thermography Enhancement Comparisons through 3D Photogrammetry

Model ReleasesDGX agent

arXiv:2608.02404v1 Announce Type: new Abstract: The Loggia dei Lanzi in the Piazza della Signoria is one of Florence's most prominent structures visited by millions every year. Its construction histor

Logit-Origin Centering for Singleton Test-Time Adaptation

ApplicationsDGX agent

arXiv:2608.01074v1 Announce Type: cross Abstract: Tabular data is used extensively in many real-world use cases. Deep learning models have been developed to deal with tabular data, but generally perfo

Logographic Character Visual Pretraining via Semantic-based Contrastive Learning

ApplicationsDGX agent

arXiv:2608.00096v1 Announce Type: new Abstract: Current deep learning-based character vision studies, e.g., text recognition, character image denoising, and historical text completion, are offering ne

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Model ReleasesDGX agent

arXiv:2608.01964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interde

Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM

Model ReleasesDGX agent

arXiv:2608.00925v1 Announce Type: new Abstract: Monocular panoramic SLAM benefits from substantial visual overlap under large camera rotations, yet remains prone to errors caused by camera tilt, scale

Loop-Mamba: A Loop Mamba with Degradation-Aware and Shared Memory for Old Photo Restoration

Model ReleasesDGX agent

arXiv:2608.02346v1 Announce Type: new Abstract: Old photographs often suffer from multiple coupled degradations, including scratches, cracks, fading, blur, noise, and missing regions, severely degradi

LUT: Latent Utility Training for Visual Reasoning

SafetyDGX agent

arXiv:2608.00743v1 Announce Type: new Abstract: Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reason

Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models

Model ReleasesDGX agent

arXiv:2409.07163v3 Announce Type: replace-cross Abstract: Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing

Manifold-GS: Certified Hybrid Assets via Varifold-Conservative Gaussian Splatting

Model ReleasesDGX agent

arXiv:2608.00214v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) gives high-quality novel-view synthesis, but its adaptive radiance primitives are not directly usable as structured assets:

Mapping melliferous tree species in Kenya via one-class classification with hyperspectral unsupervised domain adaptation

ResearchDGX agent

arXiv:2608.02045v1 Announce Type: cross Abstract: The beekeeping sector holds significant potential for livelihood diversification among the agropastoral communities in Kenya. Melliferous tree species

MBO Scheme for Local Chan--Vese Segmentation

Local AiDGX agent

arXiv:2608.00893v1 Announce Type: new Abstract: Robust to intensity inhomogeneity, the local Chan--Vese (LCV) model extends the classical Chan--Vese (CV) image segmentation method by incorporating loc

MDTD-ArtIR: Benchmarking Image Editing and Restoration Models for Art Image Restoration under Texture-Overlay Degradations

Model ReleasesDGX agent

arXiv:2608.00736v1 Announce Type: new Abstract: Restoring severely degraded visual media still remains a formidable challenge, as existing methods often hallucinate unnatural textures and contents, st

MDWD: A Street-Level Dataset for Municipal Solid Waste Detection in Dense Urban Environments

Model ReleasesDGX agent

arXiv:2608.00257v1 Announce Type: new Abstract: Automated visual monitoring of urban environments is a growing Computer Vision research area, but municipal solid waste detection remains under-represen

Measuring Product Quality Using Images: The CLIP Q-Score and an Application to Real Estate

ResearchDGX agent

arXiv:2608.01544v1 Announce Type: cross Abstract: The CLIP Q-score is a novel, safe, fully reproducible, and computationally efficient method for extracting objective product quality metrics from visu

MedSAM2-Anatomy: Training-Free Inference-Time Optimization for Musculoskeletal Segmentation

SafetyDGX agent

arXiv:2608.00195v1 Announce Type: cross Abstract: High-resolution 3D segmentation of hip and shoulder anatomy from CT and MRI is essential for surgical planning, yet frozen segmentation models often f

Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression

ResearchDGX agent

arXiv:2608.02134v1 Announce Type: new Abstract: Modern vision language models (VLMs) turn high-resolution images into long sequences of visual tokens. Every token traverses the language decoder and pe

MIDAL: Math Image Descriptions for Accessible Learning

TutorialsDGX agent

arXiv:2608.00868v1 Announce Type: new Abstract: Many open educational resources are lacking in accessibility, especially in-depth image descriptions. In subjects like Science and Mathematics, however,

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

Model ReleasesDGX agent

arXiv:2608.02059v1 Announce Type: new Abstract: Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-

MiniWorld: Democratizing the Training of Video World Models from Scratch

HardwareDGX agent

arXiv:2608.01127v1 Announce Type: new Abstract: Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through auto

Mitigating Backdoors via Decoy Shortcuts and Knowledge Decoupling

Model ReleasesDGX agent

arXiv:2608.00732v1 Announce Type: cross Abstract: Backdoor attacks pose a serious threat to deep neural networks, especially when training relies on third-party data, allowing adversaries to inject ma

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning

SafetyDGX agent

arXiv:2608.01635v1 Announce Type: new Abstract: Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instructi

Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference

SafetyDGX agent

arXiv:2606.15782v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language response

MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling

SafetyDGX agent

arXiv:2604.02817v2 Announce Type: replace Abstract: Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel

MoCRA: Mixture of Compositional Rank-1 Atoms for 4K All-in-One Video Restoration

Model ReleasesDGX agent

arXiv:2608.01829v1 Announce Type: new Abstract: Real-world video arrives hazy, rainy, dark, or noisy, and a deployable restorer faces three demands at once: no degradation label, native 4K output, and

Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking

Model ReleasesDGX agent

arXiv:2608.00847v1 Announce Type: new Abstract: Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the

MonitorVLM-v2: A Deployed Vision-Language Framework for Real-Time Safety Violation Detection

SafetyDGX agent

arXiv:2608.00975v1 Announce Type: new Abstract: Large vision--language models (VLMs) can reason step by step about complex visual scenes, but this open-ended, autoregressive chain-of-thought (CoT) app

MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

Model ReleasesDGX agent

arXiv:2608.02449v1 Announce Type: new Abstract: Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compact

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations

ResearchDGX agent

arXiv:2608.01628v1 Announce Type: new Abstract: Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations largely rely on fixed structural corr

Move What Matters: Parameter-Efficient Domain Adaptation via Optimal Transport Flow for Collaborative Perception

Model ReleasesDGX agent

arXiv:2602.11565v5 Announce Type: replace Abstract: Efficient domain adaptation remains a fundamental challenge for deploying multi-agent systems across diverse environments in Vehicle-to-Everything (

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model

SafetyDGX agent

arXiv:2603.14686v2 Announce Type: replace Abstract: Human-Object Interaction (HOI) video reenactment aims to transfer the interaction dynamics of a source video to a novel target object while preservi

Neural Born Series Operator for Biomedical Ultrasound Computed Tomography

ResearchDGX agent

arXiv:2312.15575v2 Announce Type: replace-cross Abstract: Ultrasound Computed Tomography (USCT) provides a radiation-free option for high-resolution clinical imaging. Despite its potential, the comput

New York Smells: A Large Multimodal Dataset for Olfaction

Model ReleasesDGX agent

arXiv:2511.20544v2 Announce Type: replace Abstract: While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines. One key

NISF++: Geometrically-grounded implicit representations of 3D+time cardiac function from 2D short- and long-axis MR views

SafetyDGX agent

arXiv:2608.00752v1 Announce Type: new Abstract: Clinical acquisition in cardiac magnetic resonance (CMR) imaging involves obtaining cross-sectional planes of the heart along the radial and longitudina

Noise-Robust Conditional Flow Matching: Generating Clean Samples from Noisy Datasets

TutorialsDGX agent

arXiv:2608.00064v1 Announce Type: new Abstract: Generative models learn the statistical properties of their training data, so high-quality generation depends on clean and representative datasets. In s

On the Viability of Semi-Supervised Segmentation Methods for Statistical Shape Modeling

Model ReleasesDGX agent

arXiv:2407.15260v3 Announce Type: replace Abstract: Statistical Shape Models (SSMs) excel at identifying population level anatomical variations, which is at the core of various clinical and biomedical

Onboard Satellite Image Classification for Earth Observation: A Comparative Study of ViT Models

ResearchDGX agent

arXiv:2409.03901v4 Announce Type: replace Abstract: Remote sensing (RS) image classification is central to Earth observation, but onboard deployment requires models that are accurate, efficient, and r

One Query, Many Scales: Sparse Mixture-of-Experts for Efficient Hierarchical Cross-View Geo-Localization

Model ReleasesDGX agent

arXiv:2608.01060v1 Announce Type: new Abstract: Cross-view geo-localization (CVGL) retrieves geo-tagged satellite imagery for a ground-view query. Most systems exhaustively search a flat, fixed-resolu

One-Sided Quantile Coupling for Flow Matching

SafetyDGX agent

arXiv:2608.00978v1 Announce Type: cross Abstract: Flow Matching trains continuous-time generative models by regressing the velocity field of a probability path between a simple source distribution and

Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow

Local AiDGX agent

arXiv:2608.02258v1 Announce Type: new Abstract: Rapidly evolving Generative AI enables sophisticated visual text manipulations that increasingly evade current forensic detectors. Existing discriminati

Optical Flow from Photons

ApplicationsDGX agent

arXiv:2608.00499v1 Announce Type: new Abstract: Optical flow remains challenging in high-speed and low-light scenes, where the limited frame rate and sensitivity of conventional cameras lead to motion

ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression

Model ReleasesDGX agent

arXiv:2608.00345v1 Announce Type: new Abstract: A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this seque

ORCESTRA: VLM-driven Visual Robot programming in Mixed Reality

SafetyDGX agent

arXiv:2608.00775v1 Announce Type: cross Abstract: ORCESTRA is a mixed-reality system for programming robot digital twins through no-code waypoint teaching and language-guided control. In a passthrough

OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs

SafetyDGX agent

arXiv:2603.11804v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) adapted to remote sensing rely heavily on domain-specific image-text supervision, yet high-quality annotations for sat

OSSDD - a New Open Dataset for Sentinel-1 Ship Detection

ResearchDGX agent

arXiv:2608.01963v1 Announce Type: new Abstract: Ship detection in Synthetic Aperture Radar (SAR) images plays an important role for maritime situational awareness, especially with respect to different

PackingGPT: 3D Packing Agent for Real Furniture in Last-Mile Delivery

Model ReleasesDGX agent

arXiv:2608.01427v1 Announce Type: new Abstract: 3D bin packing rectangular items into standardised containers to maximise space utilisation under geometric shipping automation. Loading a furniture pur

Parameter-Dynamic Adaptive Fusion and Calibration Network for RGBT Tracking

Model ReleasesDGX agent

arXiv:2608.01807v1 Announce Type: new Abstract: Existing RGBT trackers typically employ fusion functions with fixed parameters across different targets and scenarios. Although dynamic-architecture met

Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization

Model ReleasesDGX agent

arXiv:2505.18819v2 Announce Type: replace Abstract: Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D underst

Partial FC: Training 10 Million Identities on a Single Machine

HardwareDGX agent

arXiv:2010.05222v3 Announce Type: replace Abstract: Training face recognition models with millions of identities is challenging because classifier storage, logit memory, and computation grow linearly

PartMat: Material-Aware 3D Part Decomposition with a Single Global Latent

TutorialsDGX agent

arXiv:2608.01825v1 Announce Type: new Abstract: Part-level 3D generation has recently attracted increasing attention for producing structured and editable 3D assets. However, existing methods typicall

PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding

Local AiDGX agent

arXiv:2601.02457v2 Announce Type: replace Abstract: Current foundation models for 3D shapes excel at global tasks (retrieval, classification) but transfer poorly to local part-level reasoning. Recent

PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos

ApplicationsDGX agent

arXiv:2608.00903v1 Announce Type: new Abstract: In animation production, paint-bucket colourisation for hand-drawn animation is a labour-intensive procedure that assigns each enclosed region in line s

PhenoStitch: Training-Free Panoptic Crop Mapping from Satellite Image Time Series

ResearchDGX agent

arXiv:2608.00870v1 Announce Type: new Abstract: Panoptic crop mapping requires both delineating individual agricultural parcels and assigning a crop type to each parcel from satellite image time serie

← Previous
1…1617181920…207
Next →