AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
26 Jun 2026

DeCoFlow: Structural Decomposition of Normalizing Flows for Continual Anomaly Detection

Model ReleasesDGX agent

arXiv:2606.26687v1 Announce Type: new Abstract: In industrial environments, new product categories arrive sequentially, requiring continual anomaly detection without access to past data. Normalizing F

Depth-Semantic Alignment and Affinity-Guided Fusion for Structured Radar Point Cloud Generation

SafetyDGX agent

arXiv:2606.26743v1 Announce Type: new Abstract: Point clouds are an important carrier of three-dimensional spatial information, and their quality directly affects the performance of downstream percept

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.26602v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive fine-grained perception capabilities. However, existing ben

DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception

ResearchDGX agent

arXiv:2606.26398v1 Announce Type: new Abstract: High-precision remote perception is often hindered by the severe bandwidth constraints of Vehicle-to-Everything (V2X) networks. We propose extit{DinoLin

DnA: Denoising Attention for Visual Tasks

ResearchDGX agent

arXiv:2606.27372v1 Announce Type: new Abstract: The softmax activation in multihead attention (MHA) is the de facto standard for attention-based models in visual perception tasks. However, standard so

Do Image Editing Models Understand Lighting?

Model ReleasesDGX agent

arXiv:2606.26738v1 Announce Type: new Abstract: While recent advancements in generative image editing models have achieved stunning visual fidelity, it remains an open question whether these systems p

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents

SafetyDGX agent

arXiv:2606.26122v1 Announce Type: new Abstract: Recent methods train search agents via reinforcement learning from (question, answer, evidence) tuples without requiring expert trajectories. The tuples

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance

SafetyDGX agent

arXiv:2606.27371v1 Announce Type: new Abstract: State-of-the-art flow models generate stunning images from text or image prompts. However, they suffer from diversity collapse when generating multiple

Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution

Model ReleasesDGX agent

arXiv:2606.26716v1 Announce Type: cross Abstract: Arbitrary slice super-resolution reconstructs isotropic volumes from anisotropic clinical acquisitions by synthesizing intermediate slices at arbitrar

DynFS-MoE: Dynamic Functional-Structural Mixture-of-Experts for Post-Traumatic Epilepsy Diagnosis

TutorialsDGX agent

arXiv:2606.16203v3 Announce Type: replace Abstract: Post-traumatic epilepsy (PTE) is a severe complication of traumatic brain injury (TBI). Yet, early identification remains challenging due to the com

EndoUFM: Utilizing Foundation Models for Monocular depth estimation of endoscopic images

Local AiDGX agent

arXiv:2508.17916v2 Announce Type: replace Abstract: Depth estimation is a foundational component for 3D reconstruction in minimally invasive endoscopic surgeries. However, existing monocular depth est

Event-based Gaze Control System for Accurate Real-time Spin Estimation in Professional Ball Games

HardwareDGX agent

arXiv:2606.26780v1 Announce Type: new Abstract: Spin plays a crucial role in many ball sports due to its effect on the trajectory of the ball. Vision-based estimation of the ball's spin during a game

Exact and Deterministic Patch Descriptor Retrieval via Hierarchical Normalization

ResearchDGX agent

arXiv:2606.27280v1 Announce Type: new Abstract: We present a patch descriptor retrieval method that returns the exact nearest neighbour -- provably identical to exhaustive full-vector search -- while

Extracting Neural Materials from Multi-view Images

ResearchDGX agent

arXiv:2606.26715v1 Announce Type: new Abstract: Neural materials can represent complex specular reflections and scattering effects in a compact, universal basis. However, acquiring and authoring such

Fast LeWorldModel

Local AiDGX agent

arXiv:2606.26217v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free vis

FlameVQA: A Physically-Grounded UAV Wildfire VQA Benchmark with Radiometric Thermal Supervision

Model ReleasesDGX agent

arXiv:2606.27128v1 Announce Type: new Abstract: Wildfire monitoring from UAVs requires reliable reasoning over complex aerial scenes, where smoke, scale variation, and occlusions often limit RGB-only

Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE

ResearchDGX agent

arXiv:2606.26938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling diffusion models in visual generation. Recent advancements have f

Forget, Anticipate and Adapt: Test Time Training for Long Videos

TutorialsDGX agent

arXiv:2606.26515v1 Announce Type: new Abstract: Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating it

FracEvent: Event-Camera Simulation via Fractional-Relaxation Pixel Dynamics

ResearchDGX agent

arXiv:2606.26636v1 Announce Type: new Abstract: Event cameras asynchronously report brightness changes with microsecond-level temporal resolution, but real event data remain difficult to collect at sc

Full spectrum Unlearnable Examples via Spectral Equalization

ResearchDGX agent

arXiv:2606.26719v1 Announce Type: new Abstract: Unlearnable examples (UEs) protect training data by injecting imperceptible perturbations so that models fail to extract exploitable representations. In

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.26287v1 Announce Type: new Abstract: With the increase in model parameters and training data, the instruction following and generalization capabilities of Large VisionLanguage Models (LVLMs

Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval

TutorialsDGX agent

arXiv:2602.00813v5 Announce Type: replace Abstract: Composed Image Retrieval (CIR) is the task of retrieving a target image from a database using a multimodal query, which consists of a reference imag

Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning

ResearchDGX agent

arXiv:2606.26973v1 Announce Type: new Abstract: Open-set semi-supervised learning aims to leverage unlabeled data that may contain out-of-distribution outliers while maintaining performance on in-dist

Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification

ResearchDGX agent

arXiv:2606.20390v2 Announce Type: replace Abstract: Automated skin cancer classification from dermoscopic images remains challenging due to heterogeneous lesion structure, strong intra-class variabili

Hallucination in World Models is Predictable and Preventable

Model ReleasesDGX agent

arXiv:2606.27326v1 Announce Type: cross Abstract: Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fl

Identifying the Unknown: Prompt-Free Open Vocabulary Anomaly Recognition for Robot-Object Interaction

ApplicationsDGX agent

arXiv:2606.26829v1 Announce Type: new Abstract: Robots operating in real-world environments must in general be able to recognize previously unseen objects. As robotic systems move toward open-world au

Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision

SafetyDGX agent

arXiv:2606.26801v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervisio

Intracranial Aneurysm Classification and Segmentation via Tri-Axial ROI and Multi-Task Learning

ResearchDGX agent

arXiv:2606.26706v1 Announce Type: new Abstract: Intracranial aneurysms are often asymptomatic until rupture, which carries high mortality. Rupture risk assessment and treatment planning depend on both

Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models

Model ReleasesDGX agent

arXiv:2606.26379v1 Announce Type: new Abstract: Visual prompt tuning has emerged as a parameter-efficient fine-tuning approach for adapting large-scale Vision Transformers (ViTs) to downstream tasks.

LayersReg: A Layer-by-Layer Progressive Regressor for Reliable Intraoperative 3D/2D Registration

SafetyDGX agent

arXiv:2606.26647v1 Announce Type: new Abstract: 3D/2D registration serves as a cornerstone technique in surgical navigation. Traditional iterative optimization algorithms suffer from low efficiency an

LearniBridge: Learnable Calibration of Feature Caching for Diffusion Models Acceleration

ResearchDGX agent

arXiv:2606.26778v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have driven substantial progress in image and video generation but suffer from prohibitive computational costs. Feature ca

Learning Adversarial Augmentation Policies for Robust Garlic Seedling Detection

Local AiDGX agent

arXiv:2606.26828v1 Announce Type: new Abstract: Accurate seedling detection during early growth stages is essential for timely replanting and effective crop management in precision agriculture. Howeve

Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification

Model ReleasesDGX agent

arXiv:2601.12062v2 Announce Type: replace Abstract: The core of video-based visible-infrared person re-identification (VVI-ReID) lies in learning sequence-level modal-invariant representations across

Liquid Fusion of Heterogeneous Representations Towards General Salient Object Detection

Model ReleasesDGX agent

arXiv:2606.26849v1 Announce Type: new Abstract: General Salient Object Detection (SOD) aims to identify and segment visually interesting objects from uni-modality or multi-modality scenes, recently ad

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

SafetyDGX agent

arXiv:2606.27192v1 Announce Type: new Abstract: The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen pre

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

Model ReleasesDGX agent

arXiv:2606.26740v1 Announce Type: new Abstract: Streaming video editing has made rapid progress, yet practical deployment is still limited by two core issues: maintaining stable backgrounds and non-ed

LogicIR: Logic Gate Networks for Image Restoration

ResearchDGX agent

arXiv:2606.26609v1 Announce Type: new Abstract: Image restoration aims to reconstruct high-quality images from degraded low-quality inputs. As the computational demands of image restoration models con

Mask to Concept: Auto-Promptable SAM3 via Efficient Test-Time Concept Embedding Search for Few-Shot Annotation

Model ReleasesDGX agent

arXiv:2606.26711v1 Announce Type: new Abstract: Transforming foundation segmentation models from human-prompted tools into auto-promptable annotators is critical for scalable medical data annotation.

MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction

ResearchDGX agent

arXiv:2604.01958v2 Announce Type: replace Abstract: Infrared and visible video fusion combines the object saliency from infrared images with the texture details from visible images to produce semantic

Methane-Plume Segmentation From Hyperspectral Satellite Imagery Via Multimodal Deep Learning

ApplicationsDGX agent

arXiv:2606.26416v1 Announce Type: new Abstract: Efficient detection of methane plumes is crucial for understanding and mitigating global warming, as accurately identifying and segmenting them in earth

Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI

TutorialsDGX agent

arXiv:2606.26894v1 Announce Type: new Abstract: Brain MRI poses a fundamental challenge for machine learning: models must learn from high-dimensional 3D data spanning multiple co-registered modalities

Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction

ApplicationsDGX agent

arXiv:2606.26812v1 Announce Type: new Abstract: Multi-modality image fusion (MMIF) enhances scene representation by exploiting complementary cues from different modalities. Adverse weather, however, c

Neural Texture Compression using Hypernetworks

TutorialsDGX agent

arXiv:2606.26913v1 Announce Type: cross Abstract: Recent work on neural texture compression has demonstrated that it is possible to learn small, per-material texture representations (composed of laten

Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection

TutorialsDGX agent

arXiv:2606.26410v1 Announce Type: new Abstract: We present a self-supervised framework for learning implicit 3D physical dynamics directly from video-derived supervisory signals. While current generat

Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model

Local AiDGX agent

arXiv:2606.27325v1 Announce Type: new Abstract: Recent advances in action-conditioned world models show promising progress in modeling complex interactions and forecasting future states under diverse

OctoSense: Self-Supervised Learning for Multimodal Robot Perception

HardwareDGX agent

arXiv:2606.27317v1 Announce Type: new Abstract: We present OctoSense, an open-source sensor platform with stereo RGB and event cameras, LiDAR, a thermal camera, an inertial measurement unit, RTK-corre

Ordinal Neural Collapse as a Representation Prior for Visual Navigation

SafetyDGX agent

arXiv:2606.26839v1 Announce Type: cross Abstract: Learning robust navigation policies directly from visual observations remains a fundamental challenge in vision-based robotic navigation. In end-to-en

PanoImager: Geometry-Guided Novel View Synthesis and Reconstruction from Sparse Panoramic Views

ResearchDGX agent

arXiv:2606.27071v1 Announce Type: new Abstract: Panoramic sensing offers wide field-of-view coverage, yet 3D reconstruction from sparse panoramas remains challenging under rotation-dominant, weak-para

PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology

SafetyDGX agent

arXiv:2512.17621v2 Announce Type: replace Abstract: While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

SafetyDGX agent

arXiv:2606.27373v1 Announce Type: new Abstract: Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However,

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

Model ReleasesDGX agent

arXiv:2606.26551v1 Announce Type: new Abstract: While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing benchmarks lack a comprehensive ev

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

SafetyDGX agent

arXiv:2606.26694v1 Announce Type: new Abstract: Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to ex

PhysiFormer: Learning to Simulate Mechanics in World Space

ApplicationsDGX agent

arXiv:2606.27364v1 Announce Type: new Abstract: We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models that operate in view-dependent pixe

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2606.26916v1 Announce Type: new Abstract: Developing physically aware video generation models remains a significant challenge due to the difficulty in capturing diverse physical phenomena, such

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation

Model ReleasesDGX agent

arXiv:2606.26930v1 Announce Type: new Abstract: Reinforcement Learning like Group Relative Policy Optimization (GRPO) has significantly advanced text-to-image post-training. However, current methods o

Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning

ResearchDGX agent

arXiv:2606.26631v1 Announce Type: new Abstract: Interleaved multimodal reasoning improves visual grounding by revisiting visual evidence during multi-step generation, yet existing methods typically re

Predicting Fruit Quality with a Hybrid Machine Learning and Image Processing Approach

ResearchDGX agent

arXiv:2606.26165v1 Announce Type: new Abstract: Fruit spoilage is a significant issue in agriculture, leading to substantial economic losses. Addressing this, our study introduces a hybrid approach co

PressMimic: Pressure-Guided Motion Capture and Control for Humanoid Robot Imitation

SafetyDGX agent

arXiv:2606.26741v1 Announce Type: cross Abstract: Humanoid motion imitation requires not only accurate perception of human kinematics but also faithful reproduction of physical interactions with the e

PrivacyBench: Privacy Isn't Free in Hybrid Privacy-Preserving Vision Systems

AgentsDGX agent

arXiv:2602.18900v2 Announce Type: replace-cross Abstract: Privacy preserving machine learning deployments in sensitive deep learning applications; from medical imaging to autonomous systems; increasin

Probabilistic NDVI Forecasting from Sparse Satellite Time Series and Weather Covariates

ResearchDGX agent

arXiv:2602.17683v3 Announce Type: replace-cross Abstract: Short-term forecasting of vegetation dynamics is a key enabler for data-driven decision support in precision agriculture. Normalized Differenc

← Previous
1…6768697071…209
Next →