AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
26 Jun 2026

Forget, Anticipate and Adapt: Test Time Training for Long Videos

TutorialsDGX agent

arXiv:2606.26515v1 Announce Type: new Abstract: Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating it

FracEvent: Event-Camera Simulation via Fractional-Relaxation Pixel Dynamics

ResearchDGX agent

arXiv:2606.26636v1 Announce Type: new Abstract: Event cameras asynchronously report brightness changes with microsecond-level temporal resolution, but real event data remain difficult to collect at sc

Full spectrum Unlearnable Examples via Spectral Equalization

ResearchDGX agent

arXiv:2606.26719v1 Announce Type: new Abstract: Unlearnable examples (UEs) protect training data by injecting imperceptible perturbations so that models fail to extract exploitable representations. In


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.26287v1 Announce Type: new Abstract: With the increase in model parameters and training data, the instruction following and generalization capabilities of Large VisionLanguage Models (LVLMs

Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval

TutorialsDGX agent

arXiv:2602.00813v5 Announce Type: replace Abstract: Composed Image Retrieval (CIR) is the task of retrieving a target image from a database using a multimodal query, which consists of a reference imag

Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning

ResearchDGX agent

arXiv:2606.26973v1 Announce Type: new Abstract: Open-set semi-supervised learning aims to leverage unlabeled data that may contain out-of-distribution outliers while maintaining performance on in-dist

Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification

ResearchDGX agent

arXiv:2606.20390v2 Announce Type: replace Abstract: Automated skin cancer classification from dermoscopic images remains challenging due to heterogeneous lesion structure, strong intra-class variabili

Hallucination in World Models is Predictable and Preventable

Model ReleasesDGX agent

arXiv:2606.27326v1 Announce Type: cross Abstract: Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fl

Identifying the Unknown: Prompt-Free Open Vocabulary Anomaly Recognition for Robot-Object Interaction

ApplicationsDGX agent

arXiv:2606.26829v1 Announce Type: new Abstract: Robots operating in real-world environments must in general be able to recognize previously unseen objects. As robotic systems move toward open-world au

Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision

SafetyDGX agent

arXiv:2606.26801v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervisio

Intracranial Aneurysm Classification and Segmentation via Tri-Axial ROI and Multi-Task Learning

ResearchDGX agent

arXiv:2606.26706v1 Announce Type: new Abstract: Intracranial aneurysms are often asymptomatic until rupture, which carries high mortality. Rupture risk assessment and treatment planning depend on both

Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models

Model ReleasesDGX agent

arXiv:2606.26379v1 Announce Type: new Abstract: Visual prompt tuning has emerged as a parameter-efficient fine-tuning approach for adapting large-scale Vision Transformers (ViTs) to downstream tasks.

LayersReg: A Layer-by-Layer Progressive Regressor for Reliable Intraoperative 3D/2D Registration

SafetyDGX agent

arXiv:2606.26647v1 Announce Type: new Abstract: 3D/2D registration serves as a cornerstone technique in surgical navigation. Traditional iterative optimization algorithms suffer from low efficiency an

LearniBridge: Learnable Calibration of Feature Caching for Diffusion Models Acceleration

ResearchDGX agent

arXiv:2606.26778v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have driven substantial progress in image and video generation but suffer from prohibitive computational costs. Feature ca

Learning Adversarial Augmentation Policies for Robust Garlic Seedling Detection

Local AiDGX agent

arXiv:2606.26828v1 Announce Type: new Abstract: Accurate seedling detection during early growth stages is essential for timely replanting and effective crop management in precision agriculture. Howeve

Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification

Model ReleasesDGX agent

arXiv:2601.12062v2 Announce Type: replace Abstract: The core of video-based visible-infrared person re-identification (VVI-ReID) lies in learning sequence-level modal-invariant representations across

Liquid Fusion of Heterogeneous Representations Towards General Salient Object Detection

Model ReleasesDGX agent

arXiv:2606.26849v1 Announce Type: new Abstract: General Salient Object Detection (SOD) aims to identify and segment visually interesting objects from uni-modality or multi-modality scenes, recently ad

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

SafetyDGX agent

arXiv:2606.27192v1 Announce Type: new Abstract: The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen pre

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

Model ReleasesDGX agent

arXiv:2606.26740v1 Announce Type: new Abstract: Streaming video editing has made rapid progress, yet practical deployment is still limited by two core issues: maintaining stable backgrounds and non-ed

LogicIR: Logic Gate Networks for Image Restoration

ResearchDGX agent

arXiv:2606.26609v1 Announce Type: new Abstract: Image restoration aims to reconstruct high-quality images from degraded low-quality inputs. As the computational demands of image restoration models con

Mask to Concept: Auto-Promptable SAM3 via Efficient Test-Time Concept Embedding Search for Few-Shot Annotation

Model ReleasesDGX agent

arXiv:2606.26711v1 Announce Type: new Abstract: Transforming foundation segmentation models from human-prompted tools into auto-promptable annotators is critical for scalable medical data annotation.

MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction

ResearchDGX agent

arXiv:2604.01958v2 Announce Type: replace Abstract: Infrared and visible video fusion combines the object saliency from infrared images with the texture details from visible images to produce semantic

Methane-Plume Segmentation From Hyperspectral Satellite Imagery Via Multimodal Deep Learning

ApplicationsDGX agent

arXiv:2606.26416v1 Announce Type: new Abstract: Efficient detection of methane plumes is crucial for understanding and mitigating global warming, as accurately identifying and segmenting them in earth

Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI

TutorialsDGX agent

arXiv:2606.26894v1 Announce Type: new Abstract: Brain MRI poses a fundamental challenge for machine learning: models must learn from high-dimensional 3D data spanning multiple co-registered modalities

Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction

ApplicationsDGX agent

arXiv:2606.26812v1 Announce Type: new Abstract: Multi-modality image fusion (MMIF) enhances scene representation by exploiting complementary cues from different modalities. Adverse weather, however, c

Neural Texture Compression using Hypernetworks

TutorialsDGX agent

arXiv:2606.26913v1 Announce Type: cross Abstract: Recent work on neural texture compression has demonstrated that it is possible to learn small, per-material texture representations (composed of laten

Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection

TutorialsDGX agent

arXiv:2606.26410v1 Announce Type: new Abstract: We present a self-supervised framework for learning implicit 3D physical dynamics directly from video-derived supervisory signals. While current generat

Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model

Local AiDGX agent

arXiv:2606.27325v1 Announce Type: new Abstract: Recent advances in action-conditioned world models show promising progress in modeling complex interactions and forecasting future states under diverse

OctoSense: Self-Supervised Learning for Multimodal Robot Perception

HardwareDGX agent

arXiv:2606.27317v1 Announce Type: new Abstract: We present OctoSense, an open-source sensor platform with stereo RGB and event cameras, LiDAR, a thermal camera, an inertial measurement unit, RTK-corre

Ordinal Neural Collapse as a Representation Prior for Visual Navigation

SafetyDGX agent

arXiv:2606.26839v1 Announce Type: cross Abstract: Learning robust navigation policies directly from visual observations remains a fundamental challenge in vision-based robotic navigation. In end-to-en

PanoImager: Geometry-Guided Novel View Synthesis and Reconstruction from Sparse Panoramic Views

ResearchDGX agent

arXiv:2606.27071v1 Announce Type: new Abstract: Panoramic sensing offers wide field-of-view coverage, yet 3D reconstruction from sparse panoramas remains challenging under rotation-dominant, weak-para

PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology

SafetyDGX agent

arXiv:2512.17621v2 Announce Type: replace Abstract: While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

SafetyDGX agent

arXiv:2606.27373v1 Announce Type: new Abstract: Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However,

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

Model ReleasesDGX agent

arXiv:2606.26551v1 Announce Type: new Abstract: While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing benchmarks lack a comprehensive ev

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

SafetyDGX agent

arXiv:2606.26694v1 Announce Type: new Abstract: Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to ex

PhysiFormer: Learning to Simulate Mechanics in World Space

ApplicationsDGX agent

arXiv:2606.27364v1 Announce Type: new Abstract: We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models that operate in view-dependent pixe

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2606.26916v1 Announce Type: new Abstract: Developing physically aware video generation models remains a significant challenge due to the difficulty in capturing diverse physical phenomena, such

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation

Model ReleasesDGX agent

arXiv:2606.26930v1 Announce Type: new Abstract: Reinforcement Learning like Group Relative Policy Optimization (GRPO) has significantly advanced text-to-image post-training. However, current methods o

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning

ResearchDGX agent

arXiv:2606.26631v1 Announce Type: new Abstract: Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically re

Predicting Fruit Quality with a Hybrid Machine Learning and Image Processing Approach

ResearchDGX agent

arXiv:2606.26165v1 Announce Type: new Abstract: Fruit spoilage is a significant issue in agriculture, leading to substantial economic losses. Addressing this, our study introduces a hybrid approach co

PressMimic: Pressure-Guided Motion Capture and Control for Humanoid Robot Imitation

SafetyDGX agent

arXiv:2606.26741v1 Announce Type: cross Abstract: Humanoid motion imitation requires not only accurate perception of human kinematics but also faithful reproduction of physical interactions with the e

PrivacyBench: Privacy Isn't Free in Hybrid Privacy-Preserving Vision Systems

AgentsDGX agent

arXiv:2602.18900v2 Announce Type: replace-cross Abstract: Privacy preserving machine learning deployments in sensitive deep learning applications; from medical imaging to autonomous systems; increasin

Probabilistic NDVI Forecasting from Sparse Satellite Time Series and Weather Covariates

ResearchDGX agent

arXiv:2602.17683v3 Announce Type: replace-cross Abstract: Short-term forecasting of vegetation dynamics is a key enabler for data-driven decision support in precision agriculture. Normalized Differenc

Proposal-Conditioned Latent Diffusion for Closed-Loop Traffic Scenario Generation

SafetyDGX agent

arXiv:2606.27123v1 Announce Type: cross Abstract: Closed-loop traffic simulation remains challenging because it must generate interactive multi-agent behaviors that are scene-consistent and controllab

ProtoKV: Streaming Video Understanding under Delayed Query with Summary-State Memory

HardwareDGX agent

arXiv:2606.26762v1 Announce Type: new Abstract: Streaming video understanding (SVU) must answer queries that arrive asynchronously while visual tokens stream continuously under strict GPU-memory and q

Pseudo-Text-Conditioned 3D Grounding DINO for Organ Localization in Abdominal CT

Local AiDGX agent

arXiv:2606.27084v1 Announce Type: new Abstract: Reliable organ localization in abdominal CT can provide spatial priors for downstream trauma analysis. We propose CT-3GDINO, a lightweight 3D detector t

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

Model ReleasesDGX agent

arXiv:2606.26907v1 Announce Type: new Abstract: While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or d

RayPE: Ray-Space Positional Encoding for 3D-Aware Video Generation

ResearchDGX agent

arXiv:2606.27345v1 Announce Type: new Abstract: Modern video diffusion transformers position their tokens through RoPE on the (u,v,t) axes -- a description of the camera's sampling grid that says noth

Rendering Novel Views of MRI Using 3D Gaussian Splatting

ResearchDGX agent

arXiv:2606.26236v1 Announce Type: cross Abstract: The objective of this paper is to improve radiological gradings measured on MRIs of spines, by resampling scans so that the new view planes are better

Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs

AgentsDGX agent

arXiv:2606.26424v1 Announce Type: cross Abstract: Trajectory forecasting for autonomous driving has advanced rapidly, yet representative models often produce uninformative posteriors over forecast mod

Revealing Mammographic Phenotypes in Deep Learning Breast Cancer Risk Models

ResearchDGX agent

arXiv:2606.26431v1 Announce Type: cross Abstract: Mammogram-based deep learning models have improved breast cancer risk prediction, but the learned imaging patterns remain underexplored. Existing inte

RIS-Assisted Proactive Handover for Reliable mmWave Wireless Networks

AgentsDGX agent

arXiv:2606.26885v1 Announce Type: new Abstract: Millimeter-wave (mmWave) networks are highly susceptible to line-of-sight (LoS) blockages. Vision-aided wireless communications (VAWC) enable proactive

Rolling Shutter Relative Pose Estimation Made Practical

Model ReleasesDGX agent

arXiv:2606.26863v1 Announce Type: new Abstract: Rolling shutter (RS) cameras equip virtually all consumer devices, yet RS-aware relative pose estimation has remained impractical: the state-of-the-art

RoPEMover: Depth-Aware Object Relocation via Positional Embeddings

Model ReleasesDGX agent

arXiv:2606.27332v1 Announce Type: new Abstract: Moving an object in a single image requires geometry-consistent spatial rearrangement, including handling occlusions, revealing previously unseen region

SAM2Matting: Generalized Image and Video Matting

ResearchDGX agent

arXiv:2606.27339v1 Announce Type: new Abstract: Despite impressive advances in image matting, video matting remains challenging due to the inherent gap between high-level tracking, which requires fram

SatSplatDiff: Geometry-preserving generative refinement for high-fidelity satellite Gaussian Splatting

TutorialsDGX agent

arXiv:2606.27223v1 Announce Type: new Abstract: Gaussian Splatting has been recently explored for satellite 3D reconstruction, demonstrating flexibility and efficiency in representing radiometrically

Sculpting NeRF Geometry: Human-Preference Fine-Tuning of a 3D-Aware Face GAN

ResearchDGX agent

arXiv:2606.27305v1 Announce Type: new Abstract: Reinforcement learning from human feedback (RLHF) for 3D generation is now established across a number of works, but most existing pipelines optimise ex

See & Sniff: Learning Visuo-Olfactory Representations

Model ReleasesDGX agent

arXiv:2606.27307v1 Announce Type: new Abstract: While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olf

Self-Supervised Tree-level Biomass Estimation in Urban Environments From Airborne LiDAR and Optical Observations

ResearchDGX agent

arXiv:2606.26194v1 Announce Type: new Abstract: Urban tree biomass remains less spatially explicitly quantified than biomass in managed forests because many estimates rely on inventories or coarse pro

SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning

ApplicationsDGX agent

arXiv:2603.10446v4 Announce Type: replace Abstract: Sign Language Production (SLP) faces a fundamental trade-off: direct text-to-pose models suffer from regression-to-the-mean effects, while dictionar

← Previous
1…6970717273…211
Next →