AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
28 Jul 2026

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models

Local AiDGX agent

arXiv:2607.24447v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories ge

RRTrack: Robust and Recoverable Object 6D Pose Tracking for Dynamic Scenes

Model ReleasesDGX agent

arXiv:2607.23669v1 Announce Type: new Abstract: Robust object 6D pose tracking is critical for robotic systems operating in dynamic and occluded scenes. Per-frame estimators are accurate but computati

SADe: Sparse-Atom Support Decontamination for Few-Shot Segmentation with Weak Support Annotations

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.24706v1 Announce Type: new Abstract: Few-shot segmentation (FSS) commonly assumes clean pixel-level support masks, yet practical support supervision often uses boxes, scribbles, coarse mask

Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations

ResearchDGX agent

arXiv:2607.22872v1 Announce Type: cross Abstract: Post-training quantization (PTQ) has become a practical solution for deploying deep learning models on resource-constrained edge devices by compressin

SARATR-X-v2: Scale-Aware Structural Pre-Training for SAR Foundation Models

ResearchDGX agent

arXiv:2607.23238v1 Announce Type: new Abstract: Masked image modeling has become a dominant paradigm for SAR pre-training, yet the design of the reconstruction target remains fundamentally unsettled.

Segmentation Robustness and Predictive Utility in Glioblastoma Radiomics: Evidence for a Trade-off in Survival Modelling

ResearchDGX agent

arXiv:2607.23626v1 Announce Type: cross Abstract: Radiomic biomarkers derived from magnetic resonance imaging (MRI) have been widely investigated as non-invasive tools for tumor characterization and p

Self-Distillation of Hidden Layers for Self-Supervised Representation Learning

ResearchDGX agent

arXiv:2603.15553v2 Announce Type: replace Abstract: The landscape of self-supervised learning (SSL) is currently dominated by generative approaches (e.g. MAE) that reconstruct raw low-level data, and

Shift-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment

SafetyDGX agent

arXiv:2501.19060v4 Announce Type: replace Abstract: Vision-language models (VLMs), such as CLIP, adapt effectively to downstream tasks through prompt tuning, but fine-tuning can misalign predictive co

SHReg: Strictly Rotation-Equivariant Point Cloud Registration via Spherical Harmonics

ResearchDGX agent

arXiv:2607.23096v1 Announce Type: new Abstract: Point cloud registration critically depends on local features that are both distinctive and robust to arbitrary 3D rotations. Existing learning-based me

SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth Estimation

Model ReleasesDGX agent

arXiv:2607.24249v1 Announce Type: new Abstract: Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navigation hazards. While specialized hardware

SimBEV2X: A Large-Scale Dataset and Data Generation Tool for Multi-Task Vehicle-to-Everything Cooperative Perception

AgentsDGX agent

arXiv:2607.23910v1 Announce Type: new Abstract: Cooperative perception through vehicle-to-everything (V2X) communication can overcome the inherent physical limitations of individual autonomous vehicle

SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion

Model ReleasesDGX agent

arXiv:2607.23580v1 Announce Type: new Abstract: Existing vector-sketch models treat recognition and generation as separate tasks, leaving a gap for streaming interfaces that must understand a drawing

SLAM-Former: Putting SLAM into One Transformer

ResearchDGX agent

arXiv:2509.16909v2 Announce Type: replace Abstract: We present SLAM-Former, a neural approach that integrates full SLAM capabilities into a single transformer. Similar to traditional SLAM systems, SLA

Small, Bias-Free, Blind and Convolutional Denoiser: A compact ConvNeXt U-Net for blind Gaussian color-image denoising

Model ReleasesDGX agent

arXiv:2607.22793v1 Announce Type: cross Abstract: We describe and evaluate BF-ConvUNeXt, a compact bias-free ConvNeXt U-Net for blind additive-white-Gaussian-noise color image denoising, combining fou

Small-Pollinator Detection in Cluttered Field Video

HardwareDGX agent

arXiv:2607.22913v1 Announce Type: new Abstract: Detecting pollinators in field video is challenging: targets are small, visually similar, and observed against cluttered vegetation under blur and occlu

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

ResearchDGX agent

arXiv:2607.24027v1 Announce Type: new Abstract: Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Traini

Source-Free Controlled Adaptation of Teachers for Continual Test-Time Adaptation

Model ReleasesDGX agent

arXiv:2607.23735v1 Announce Type: cross Abstract: In many real-world scenarios, encountering continual shifts in domain during inference is very common. Consequently, continual test-time adaptation (C

Spatially-Aware Class-Agnostic Object Counting

ResearchDGX agent

arXiv:2607.16826v2 Announce Type: replace Abstract: Generalised object counting aims to estimate the number of instances of an arbitrary object category from a single image, but many recent methods ca

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

Model ReleasesDGX agent

arXiv:2607.24701v1 Announce Type: new Abstract: Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Exist

Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions

ResearchDGX agent

arXiv:2607.22931v1 Announce Type: cross Abstract: Analytic Continual Learning (ACL) offers a computationally efficient alternative to gradient-based approaches. Recent ACL methods are based on Recursi

Stabilizing Deep Reconstruction Operators with Contractive Anchoring

ResearchDGX agent

arXiv:2607.23341v1 Announce Type: cross Abstract: Pretrained deep denoisers can be used to solve a wide range of model-based image reconstruction tasks via Plug-and-Play (PnP) and Regularization-by-De

StAR: Segment Anything Reasoner

Model ReleasesDGX agent

arXiv:2603.14382v2 Announce Type: replace Abstract: As AI systems are being integrated more rapidly into diverse and complex real-world environments, the ability to perform holistic reasoning over an

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

Model ReleasesDGX agent

arXiv:2607.22798v1 Announce Type: cross Abstract: Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screen

STEER: Steerable Dyadic Head Avatars

ResearchDGX agent

arXiv:2607.23840v1 Announce Type: new Abstract: Facial movement and expression are central to face-to-face communication, conveying turn-taking, attention, agreement, and engagement alongside speech.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design

Model ReleasesDGX agent

arXiv:2607.22708v1 Announce Type: new Abstract: Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the ac

Structural Loss Metrics for Tensor Approximation via Matrix Low-Rank Approximation

ResearchDGX agent

arXiv:2607.24009v1 Announce Type: new Abstract: Matricized low-rank approximation via SVD is a standard surrogate for tensor decompositions, but entry-wise reconstruction error fails to capture multiw

Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

AgentsDGX agent

arXiv:2603.27577v2 Announce Type: replace Abstract: Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typic

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs

SafetyDGX agent

arXiv:2607.23046v1 Announce Type: new Abstract: Recent high-resolution Multimodal Large Language Models (MLLMs) generate thousands of visual tokens per input, leading to a visual token explosion that

Superpixel-Based QUBO for Scalable Quantum-Enhanced Medical Image Segmentation

ResearchDGX agent

arXiv:2607.24288v1 Announce Type: new Abstract: Quadratic unconstrained binary optimization (QUBO) has emerged as a powerful framework for medical computing problems. Binary decision variables natural

Surgical Re-enactment for Operating Room Workflow Datasets

TutorialsDGX agent

arXiv:2607.24206v1 Announce Type: cross Abstract: The introduction of new technologies, such as surgical robots, is driving the vision of a connected, smart operating room (OR). However, realizing thi

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

Local AiDGX agent

arXiv:2607.24359v1 Announce Type: new Abstract: Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video

Test-Time Adaptation via Dual Distillation for Videos Under Severe Distribution Shifts

SafetyDGX agent

arXiv:2607.24611v1 Announce Type: new Abstract: Deep learning models have achieved state-of-the-art performance in several computer vision tasks. However, they experience severe performance degradatio

Text-based Tactile Graphics Generation for the Visually Impaired

ResearchDGX agent

arXiv:2607.22674v1 Announce Type: cross Abstract: Tactile graphics are a primary medium for blind and low-vision (BLV) individuals to access non-textual information. However, they are difficult to sca

The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.23335v1 Announce Type: new Abstract: Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that th

The Label Complexity of Class-Conditional Coverage under Distribution Shift

Model ReleasesDGX agent

arXiv:2607.18088v2 Announce Type: replace-cross Abstract: Conformal prediction certifies that a classifier's prediction sets cover the truth, and that certificate is marginal. Many recognition benchma

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

Model ReleasesDGX agent

arXiv:2607.23951v1 Announce Type: new Abstract: Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods

To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion

ResearchDGX agent

arXiv:2607.23492v1 Announce Type: new Abstract: Concept erasure techniques (CETs) edit text-to-image diffusion models to erase undesired targets such as NSFW content or copyrighted styles, while prese

TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians

ResearchDGX agent

arXiv:2607.22717v1 Announce Type: new Abstract: While Implicit Neural Representations (INRs) and dynamic 3D Gaussian Splatting (3DGS) achieve impressive results in video processing, they often fall sh

Toward Optimal Adenovirus Detection Using YOLO26

ResearchDGX agent

arXiv:2607.17799v2 Announce Type: replace Abstract: This study systematically benchmarks different data augmentation setups across the baseline YOLO26 model size variants to determine the most effecti

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

SafetyDGX agent

arXiv:2607.23181v1 Announce Type: new Abstract: Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen sc

Trainable Nonexpansive Denoisers for Contractive Image Reconstruction

ResearchDGX agent

arXiv:2607.23347v1 Announce Type: cross Abstract: Trainable denoisers with Lipschitz control have become central to convergent image reconstruction. However, training neural networks that simultaneous

TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation

ResearchDGX agent

arXiv:2607.24215v1 Announce Type: new Abstract: Although general text-to-image models excel in open-domain generation, their performance degrades significantly in specialized downstream domains, parti

Trustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation

Model ReleasesDGX agent

arXiv:2607.22727v1 Announce Type: new Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents

Model ReleasesDGX agent

arXiv:2508.00288v5 Announce Type: replace-cross Abstract: Aerial navigation is a fundamental yet underexplored capability in embodied intelligence, enabling agents to operate in large-scale, unstructu

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

Local AiDGX agent

arXiv:2607.23373v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge d

UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing

ResearchDGX agent

arXiv:2607.24298v1 Announce Type: new Abstract: Recent 3D foundation models can generate high-quality assets from a single image, but degrade markedly on unconstrained multi-image inputs, often produc

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

ResearchDGX agent

arXiv:2607.24157v1 Announce Type: new Abstract: Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled

URHead: A Unified UV-Space Representation for Joint Mesh-3DGS Optimization in Head Avatars

ResearchDGX agent

arXiv:2607.22673v1 Announce Type: cross Abstract: We present URHead, a unified representation for high-fidelity and animatable head avatars that fundamentally redefines mesh-Gaussian integration. Whil

VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent Mamba

ApplicationsDGX agent

arXiv:2603.00887v2 Announce Type: replace Abstract: Volume Electron Microscopy (VEM) is crucial for 3D tissue imaging but often produces anisotropic data with poor axial resolution, hindering visualiz

ViDS: Video Diffusion Shader using 3D Face Tracking

ResearchDGX agent

arXiv:2607.24124v1 Announce Type: new Abstract: We introduce ViDS, a Video Diffusion Shader that leverages 3D face tracking for expressive and identity-preserving portrait animation. We first reconstr

VIP: Finding Important People in Images

ResearchDGX agent

arXiv:1502.05678v3 Announce Type: replace Abstract: People preserve memories of events such as birthdays, weddings, or vacations by capturing photos, often depicting groups of people. Invariably, some

VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

TutorialsDGX agent

arXiv:2607.23472v1 Announce Type: new Abstract: Modern video generation models can synthesize visually compelling and temporally coherent clips, yet controlling their physical behavior remains difficu

Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models

ApplicationsDGX agent

arXiv:2607.22723v1 Announce Type: new Abstract: Visual information extraction (VIE) from visually rich documents remains challenging due to high layout variability and real-world impairments. Existing

Visual Token Compression Enhances Robustness of MLLMs

SafetyDGX agent

arXiv:2607.22716v1 Announce Type: new Abstract: In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vuln

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

ResearchDGX agent

arXiv:2607.23265v1 Announce Type: new Abstract: Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. Whil

Weakly Supervised Instance-Level Gleason Pattern Estimation Using Primary and Secondary Labels

ResearchDGX agent

arXiv:2607.23594v1 Announce Type: new Abstract: In prostate cancer histopathology, the Gleason Score is determined by the most frequent (Primary) and second most frequent (Secondary) Gleason patterns

WGDnet: Wishart-guided Geometric-aware Deep Network for PolSAR Image Classification

Local AiDGX agent

arXiv:2607.23638v1 Announce Type: new Abstract: Polarimetric Synthetic Aperture Radar (PolSAR) classification underpins all-weather Earth observation. Conventional Wishart methods depend on rigid hand

What Can I Edit? Open-Ended Strategy Discovery and the Emotion Editability Landscape

AgentsDGX agent

arXiv:2607.23920v1 Announce Type: new Abstract: Emotional image editing requires more than applying affective filters or modifying predefined visual factors: an effective edit must identify what a par

When Less Is More: A Controlled Benchmark of Lightweight CNNs for Satellite Land-Cover Segmentation on DeepGlobe

Model ReleasesDGX agent

arXiv:2607.23024v1 Announce Type: new Abstract: High-resolution satellite imagery is the backbone of good land-cover classification, and without that, environmental monitoring, urban planning, and sus

When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

Model ReleasesDGX agent

arXiv:2607.24077v1 Announce Type: new Abstract: Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged

← Previous
1…3132333435…209
Next →