AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
14 May 2026

Asymmetric Flow Models

ResearchDGX agent

arXiv:2605.12964v1 Announce Type: new Abstract: Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has s

Backbone is All You Need: Assessing Vulnerabilities of Frozen Foundation Models in Synthetic Image Forensics

ResearchDGX agent

arXiv:2605.13381v1 Announce Type: new Abstract: As AI-generated synthetic images become increasingly realistic, Vision Transformers (ViTs) have emerged as a cornerstone of modern deepfake detection. H

Bayesian In Vivo Tracking of Synapses using Joint Poisson Deconvolution and Diffeomorphic Registration

Model ReleasesDGX agent

arXiv:2605.13455v1 Announce Type: new Abstract: Synapses are densely packed submicron structures that dynamically reorganize during learning and memory formation. Longitudinal extit{in vivo} imaging o


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception

ResearchDGX agent

arXiv:2510.01502v2 Announce Type: replace-cross Abstract: Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social in

Bias In, Bias Out? Finding Unbiased Subnetworks in Vanilla Models

Model ReleasesDGX agent

arXiv:2603.05582v2 Announce Type: replace-cross Abstract: The issue of algorithmic biases in deep learning has led to the development of various debiasing techniques, many of which perform complex tra

BlitzGS: City-Scale Gaussian Splatting at Lightning Speed

SafetyDGX agent

arXiv:2605.13794v1 Announce Type: cross Abstract: We present BlitzGS, a distributed 3DGS framework that reduces active Gaussian workload for fast city-scale reconstruction. BlitzGS manages this worklo

Brain Tumor Classification in MRI Images: A Computationally Efficient Convolutional Neural Network

ResearchDGX agent

arXiv:2605.12560v1 Announce Type: cross Abstract: Improving patient outcomes depends on the prompt and accurate diagnosis of brain tumors, but manual MRI scan analysis is still time-consuming and unre

BrainAnytime: Anatomy-Aware Cross-Modal Pretraining for Brain Image Analysis with Arbitrary Modality Availability

ResearchDGX agent

arXiv:2605.13059v1 Announce Type: new Abstract: Clinical diagnostic workups typically follow a modality escalation pathway: after initial clinical evaluation, clinicians begin with routine structural

CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding

SafetyDGX agent

arXiv:2605.13544v1 Announce Type: new Abstract: Fine-grained Vision-Language Pre-training (FVLP) demonstrates significant potential in 3D medical image understanding by aligning anatomy-level visual r

Characterizing Universal Object Representations Across Vision Models

SafetyDGX agent

arXiv:2605.13675v1 Announce Type: new Abstract: Deep neural networks trained with different architectures, objectives, and datasets have been reported to converge on similar visual representations. Ho

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

Model ReleasesDGX agent

arXiv:2605.12882v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answ

Color Constancy in Hyperspectral Imaging via Reduced Spectral Spaces

ResearchDGX agent

arXiv:2605.13306v1 Announce Type: new Abstract: Illuminant estimation aims to infer scene illumination from image measurements despite intrinsic ambiguities between surface reflectance and lighting. M

Compact 3D Gaussian Splatting For Dense Visual SLAM

Model ReleasesDGX agent

arXiv:2403.11247v3 Announce Type: replace Abstract: Recent work has shown that 3D Gaussian-based SLAM enables high-quality reconstruction, accurate pose estimation, and real-time rendering of scenes.

Conditional Compatibility Learning for Context-Dependent Anomaly Detection

ApplicationsDGX agent

arXiv:2601.22868v3 Announce Type: replace Abstract: Anomaly detection usually assumes that abnormality is an intrinsic property of an observation. A defect is a defect, and a rare object is rare, rega

CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis

SafetyDGX agent

arXiv:2605.12650v1 Announce Type: new Abstract: Foundation diffusion models can generate photorealistic natural images, but adapting them to medical imaging remains challenging. In medical adaptation,

Data Agent: Learning to Select Data via End-to-End Dynamic Optimization

SafetyDGX agent

arXiv:2603.07433v2 Announce Type: replace-cross Abstract: Dynamic Data selection aims to accelerate training by prioritizing informative samples during online training. However, existing methods typic

Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation

ResearchDGX agent

arXiv:2605.12952v1 Announce Type: new Abstract: Grad-ECLIP is published at ICML 2024 and represents a new Transformer interpretation technical route (intermediate features-based). First, this paper de

DeepFilters: Scattering-Aware Pupil Engineering with Learned Digital Filter Reconstruction for Extended Depth of Field Microscopy

ResearchDGX agent

arXiv:2605.13619v1 Announce Type: cross Abstract: Extended depth of field microscopy encodes axial information into a single acquisition through engineered point spread functions, but conventional and

DiffST: Spatiotemporal-Aware Diffusion for Real-World Space-Time Video Super-Resolution

TutorialsDGX agent

arXiv:2605.13182v1 Announce Type: new Abstract: Diffusion-based models have shown strong performance in video super-resolution (VSR) and video frame interpolation (VFI). However, their role in the cou

DirectTryOn: One-Step Virtual Try-On via Straightened Conditional Transport

ResearchDGX agent

arXiv:2605.12939v1 Announce Type: new Abstract: Recent diffusion- and flow-based VTON methods achieve strong results with pretrained generative models, but their reliance on multi-step sampling incurs

DIVER:Diving Deeper into Distilled Data via Expressive Semantic Recovery

HardwareDGX agent

arXiv:2605.12649v1 Announce Type: new Abstract: Dataset distillation aims to synthesize a compact proxy dataset that is unreadable or non-raw from the original dataset for privacy protection and highl

DocAtlas: Multilingual Document Understanding Across 80+ Languages

Model ReleasesDGX agent

arXiv:2605.12623v1 Announce Type: cross Abstract: Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that p

Does Engram Do Memory Retrieval in Autoregressive Image Generation?

Local AiDGX agent

arXiv:2605.13179v1 Announce Type: new Abstract: The Engram module -- a hash-keyed, O(1) associative memory injected into Transformer layers -- was recently shown to improve large language model pretra

Drag within Prior Distribution: Text-Conditioned Point-Based Image Editing within Distribution Constraints

Local AiDGX agent

arXiv:2605.13349v1 Announce Type: new Abstract: Diffusion-based point editing methods have gained significant traction in image editing tasks due to their ability to manipulate image semantics and fin

Driving Intents Amplify Planning-Oriented Reinforcement Learning

SafetyDGX agent

arXiv:2605.12625v1 Announce Type: cross Abstract: Continuous-action policies trained on a single demonstrated trajectory per scene suffer from mode collapse: samples cluster around the demonstrated ma

Dual-Pathway Circuits of Object Hallucination in Vision-Language Models

ResearchDGX agent

arXiv:2605.13156v1 Announce Type: new Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in bridging visual perception and natural language understanding, enabling a wid

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation

ResearchDGX agent

arXiv:2605.13122v1 Announce Type: new Abstract: Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural langu

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling

Model ReleasesDGX agent

arXiv:2605.13062v1 Announce Type: new Abstract: Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, e

EDITS: Enhancing Dataset Distillation with Implicit Textual Semantics

Local AiDGX agent

arXiv:2509.13858v2 Announce Type: replace Abstract: Dataset distillation aims to synthesize a compact dataset from the original large-scale one, enabling highly efficient learning while preserving com

EgoForce: Robust Online Egocentric Motion Reconstruction via Diffusion Forcing

ApplicationsDGX agent

arXiv:2605.13041v1 Announce Type: new Abstract: With recent advances in embodied agents and AR devices, egocentric observations are readily available as input for real-world interactive online applica

Energy Scaling Laws for Diffusion Models: Quantifying Compute in Image Generation

Model ReleasesDGX agent

arXiv:2511.17031v2 Announce Type: replace-cross Abstract: The rapidly growing computational demands of diffusion models for image generation have raised significant concerns about energy consumption a

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding

TutorialsDGX agent

arXiv:2605.13803v1 Announce Type: new Abstract: Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the qu

Exploring Multimodal LMMs for Online Episodic Memory Question Answering on the Edge

Model ReleasesDGX agent

arXiv:2602.22455v2 Announce Type: replace Abstract: We investigate the feasibility of using Multimodal Large Language Models (MLLMs) for real-time online episodic memory question answering. While clou

Fast and Compact Graph Cuts for the Boykov-Kolmogorov Algorithm

Model ReleasesDGX agent

arXiv:2605.13402v1 Announce Type: new Abstract: Computing a minimum s-t cut in a graph is a solution to a wide range of computer vision problems, and is often done using the Boykov-Kolmogorov (BK) alg

FedHPro: Federated Hyper-Prototype Learning via Gradient Matching

Model ReleasesDGX agent

arXiv:2605.13475v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative training of distributed clients while protecting privacy. To enhance generalization capability in FL, prot

FIKA-Bench: From Fine-grained Recognition to Fine-Grained Knowledge Acquisition

AgentsDGX agent

arXiv:2605.13193v1 Announce Type: new Abstract: Fine-grained recognition in everyday life is often not a closed-book classification problem: when encountering unfamiliar objects, humans actively searc

Flow Augmentation and Knowledge Distillation for Lightweight Face Presentation Attack Detection

HardwareDGX agent

arXiv:2605.13108v1 Announce Type: new Abstract: Face presentation attack detection (FacePAD) remains challenging under diverse spoofing representation, including 2D print and replay, 3D mask-based spo

Flow Matching with Uncertainty Quantification and Guidance

ResearchDGX agent

arXiv:2602.10326v2 Announce Type: replace Abstract: Despite the remarkable success of sampling-based generative models such as flow matching, they can still produce samples of inconsistent or degraded

FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection

ApplicationsDGX agent

arXiv:2509.23056v2 Announce Type: replace Abstract: Remote sensing object detection is a critical technology for real-world applications such as natural resource monitoring, traffic management, and UA

From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning

Model ReleasesDGX agent

arXiv:2603.26839v2 Announce Type: replace-cross Abstract: How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce ex

GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose Estimation

Local AiDGX agent

arXiv:2605.13151v1 Announce Type: new Abstract: Category-agnostic pose estimation (CAPE) aims to localize keypoints on query images from arbitrary categories, using only a few annotated support exampl

Generative Motion In-betweening by Diffusion over Continuous Implicit Representations

ResearchDGX agent

arXiv:2605.12778v1 Announce Type: cross Abstract: Recent advances in generative models have yielded impressive progress on motion in-betweening, allowing for more complex, varied, and realistic motion

Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception

SafetyDGX agent

arXiv:2605.13755v1 Announce Type: new Abstract: In recent years, autonomous driving has significantly in creased the demand for high-quality data to train 2D and 3D perception models for safety-critic

GeomHair: Reconstruction of Hair Strands from Colorless 3D Scans

Model ReleasesDGX agent

arXiv:2505.05376v3 Announce Type: replace Abstract: We propose a novel method that reconstructs hair strands directly from colorless 3D scans by leveraging multi-modal hair orientation extraction. Hai

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

TutorialsDGX agent

arXiv:2602.17555v3 Announce Type: replace Abstract: Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. C

GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion

AgentsDGX agent

arXiv:2605.12957v1 Announce Type: new Abstract: Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains

GuardMarkGS: Unified Ownership Tracing and Edit Deterrence for 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2605.12919v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is becoming a practical representation for novel view synthesis, but its growing adoption, together with rapid advances in

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.13632v1 Announce Type: cross Abstract: In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied

Guidestar-Free Adaptive Optics with Asymmetric Apertures

ResearchDGX agent

arXiv:2602.07029v3 Announce Type: replace-cross Abstract: This work introduces the first closed-loop adaptive optics (AO) system capable of optically correcting aberrations in real-time without a guid

HADAR-Based Thermal Infrared Hyperspectral Image Restoration

Model ReleasesDGX agent

arXiv:2605.13664v1 Announce Type: new Abstract: Thermal-infrared (TIR) hyperspectral imagery (HSI) provides critical scene information for various applications. However, its practical utility is sever

HarmoGS: Robust 3D Gaussian Splatting in the Wild via Conflict-Aware Gradient Harmonization

ResearchDGX agent

arXiv:2605.13073v1 Announce Type: new Abstract: In-the-wild 3D Gaussian Splatting remains challenging due to transient distractors and illumination-induced cross-view appearance inconsistencies. Exist

HIR-ALIGN: Enhancing Hyperspectral Image Restoration via Diffusion-Based Data Generation

SafetyDGX agent

arXiv:2605.13581v1 Announce Type: new Abstract: Hyperspectral image (HSI) restoration is crucial for reliable analysis, as real HSIs suffer from degradations like noise, blur, and resolution loss. How

Human face perception reflects inverse-generative and naturalistic discriminative objectives

ResearchDGX agent

arXiv:2605.12619v1 Announce Type: cross Abstract: The perceptual representations supporting our ability to recognize faces remain a computational mystery. Deep neural networks offer mechanistic hypoth

ImageAttributionBench: How Far Are We from Generalizable Attribution?

Model ReleasesDGX agent

arXiv:2605.12967v1 Announce Type: new Abstract: The rapid advancement of generative AI has enabled the creation of highly realistic and diverse synthetic images, posing critical challenges for image p

Img2CADSeq: Image-to-CAD Generation via Sequence-Based Diffusion

ResearchDGX agent

arXiv:2605.13293v1 Announce Type: new Abstract: Boundary Representation (BRep) is the standard format for Computer-Aided Design (CAD), yet reconstructing high-quality BReps from single-view images rem

Improved monocular depth prediction using distance transform over pre-semantic contours with self-supervised neural networks

ResearchDGX agent

arXiv:2605.08320v1 Announce Type: cross Abstract: Monocular depth estimation (MDE) with self-supervised training approaches struggles in low-texture areas, where photometric losses may lead to ambiguo

Inference-Time Dynamic Modality Selection for Incomplete Multimodal Classification

TutorialsDGX agent

arXiv:2601.22853v3 Announce Type: replace Abstract: Multimodal deep learning (MDL) has achieved remarkable success across various domains, yet its practical deployment is often hindered by incomplete

Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models

SafetyDGX agent

arXiv:2605.12725v1 Announce Type: new Abstract: Recent video anomaly detection research has expanded rapidly with an emphasis on general models of normality intended to work across many different scen

JANUS: Anatomy-Conditioned Gating for Robust CT Triage Under Distribution Shift

ResearchDGX agent

arXiv:2605.13813v1 Announce Type: new Abstract: Automated CT triage requires models that are simultaneously accurate across diverse pathologies and reliable under institutional shift. While Vision Tra

Just Ask for a Table: A Thirty-Token User Prompt Defeats Sponsored Recommendations in Twelve LLMs

ResearchDGX agent

arXiv:2605.12772v1 Announce Type: new Abstract: Wu et al. (2026) showed that most frontier large language models (LLMs) recommend a sponsored, roughly twice-as-expensive flight when their system promp

← Previous
1…137138139140141…211
Next →