AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
11 Aug 2026

DiffSafeMerge: Mitigating Backdoor Inheritance in Diffusion Model Merging

ResearchDGX agent

arXiv:2608.09445v1 Announce Type: cross Abstract: Unconditional diffusion checkpoint merging assumes benign sources, yet a compromised public checkpoint can transfer a dormant backdoor while clean gen

Diffuse the object, keep its label: curating detector training data from a few unlabeled photographs via VLM-built 3D vegetation scenes

Model ReleasesDGX agent

arXiv:2608.09691v1 Announce Type: new Abstract: Labeled images of small objects hidden in vegetation are scarce, and detectors trained on them generalize poorly across sites. Rather than reusing label

Diffusion Image Editing via Asynchronous Token Decoding

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.09322v1 Announce Type: new Abstract: Text-guided diffusion image editing aims to modify semantic attributes of an image while preserving its identity, layout, and background. However, naive

DINO-3DRA: Leveraging 2D Foundation Model Semantics for 3D Cerebral Aneurysm Segmentation

ResearchDGX agent

arXiv:2608.07767v1 Announce Type: new Abstract: Accurate aneurysm segmentation in 3D rotational angiography (3DRA) is hindered by extreme class imbalance, morphological similarity to vessels, and abse

Disentangling Co-Occurring Retinal Pathologies with Saliency-Guided Sparse Expert Routing

Model ReleasesDGX agent

arXiv:2608.09752v1 Announce Type: new Abstract: Retinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation t

Distilling Physical Priors into Streaming World Models

ApplicationsDGX agent

arXiv:2608.07981v1 Announce Type: new Abstract: Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts of

DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning

Local AiDGX agent

arXiv:2608.09907v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domain

Diversity Matters: Distributional Feature Coverage Sample Selection for Data-Efficient Backdoor Attacks

ResearchDGX agent

arXiv:2608.09047v1 Announce Type: cross Abstract: Backdoor attacks compromise training data so that a model retains clean accuracy but predicts an attacker-chosen target on triggered inputs. At very l

DocPure: Prompt-Free Unified Document Restoration via Degradation-Aware Structure-Guided Wavelet Modulation

ResearchDGX agent

arXiv:2608.09536v1 Announce Type: new Abstract: High-quality document images are pivotal for information archiving and downstream automatic processing. However, they are frequently compromised by dive

DoRF++: Spherical Representation Learning over Doppler Radiance Fields for Robust Wi-Fi Sensing

ApplicationsDGX agent

arXiv:2608.08381v1 Announce Type: new Abstract: Motivated by the IEEE 802.11bf effort to standardize advanced WLAN sensing, interest in Wi-Fi Channel State Information (CSI) for passive, device-free,

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models

SafetyDGX agent

arXiv:2608.09233v1 Announce Type: cross Abstract: Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-trai

Drone-Assisted UAV-UGV Collaboration for Autonomous Navigation in Snow-Covered Terrain

AgentsDGX agent

arXiv:2608.07797v1 Announce Type: cross Abstract: This paper presents a collaborative UAV-UGV navigation framework for high-altitude, snow-covered terrain, where reduced visibility and unstable ground

eBIRD: Event-based Intensity Image Reconstruction Using Controllable Diffusion Models

ResearchDGX agent

arXiv:2608.08519v1 Announce Type: new Abstract: Intensity-image reconstruction from event streams remains a challenging problem due to the binary, sparse, and asynchronous nature of event data. This w

EFFEKT: Efficient Federated Knowledge Transfer to Foundation Models

SafetyDGX agent

arXiv:2608.08138v1 Announce Type: new Abstract: Recent data protection laws have accelerated the adoption of Federated Learning (FL) for privacy-preserving decentralized training. Nevertheless, increa

Efficient Fine-Tuning of DINOv3 Pretrained on Natural Images for Atypical Mitotic Figure Classification

Model ReleasesDGX agent

arXiv:2508.21041v4 Announce Type: replace-cross Abstract: Atypical mitotic figures (AMFs) indicate abnormal cell division associated with poor prognosis. Their detection remains difficult due to low p

Efficient Human-Contact Representation for Human-Scene Interaction

Model ReleasesDGX agent

arXiv:2608.09388v1 Announce Type: new Abstract: Human-scene interaction is an active research topic with several industrial applications in virtual reality, gaming, robotics, and surveillance. Despite

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

Local AiDGX agent

arXiv:2608.08285v1 Announce Type: new Abstract: We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs

EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization

TutorialsDGX agent

arXiv:2608.09656v1 Announce Type: new Abstract: Visual query localization (VQL) aims to retrieve and re-localize a queried object in egocentric videos, yet remains challenging when object boundaries a

EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking

AgentsDGX agent

arXiv:2608.08016v1 Announce Type: new Abstract: Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions

EndoMD-SLAM: Endoscopic Gaussian Splatting SLAM under Optical Degradation with Memory and Static-Transient Decomposition

Model ReleasesDGX agent

arXiv:2608.08949v1 Announce Type: new Abstract: Dense 3D reconstruction is critical for clinical endoscopic navigation and documentation. While Gaussian Splatting SLAM systems show promise in this dom

Ensemble learning of pathology foundation models for precision oncology

TutorialsDGX agent

arXiv:2508.16085v2 Announce Type: replace Abstract: Histopathology is essential for cancer diagnosis and treatment selection, and pathology foundation models learn visual representations from whole-sl

ERF-GS: Reconstructing Fast Motion from Disjoint Event-RGB Viewpoints

HardwareDGX agent

arXiv:2608.08531v1 Announce Type: new Abstract: Deep learning-driven representations such as neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS) have revolutionized the field of dynamic 3D

EvBS: Event-guided Blur Synthesis for Domain-adaptive Motion Deblurring

ApplicationsDGX agent

arXiv:2608.08066v1 Announce Type: new Abstract: Motion deblurring has achieved remarkable progress with deep learning, yet pre-trained deblurring models often suffer from performance degradation in re

EvTrajGS: Accurate and Efficient 3D Gaussian Splatting from Unposed Event Streams

ApplicationsDGX agent

arXiv:2608.08585v1 Announce Type: new Abstract: Event cameras, with high temporal resolution, high dynamic range, and asynchronous sensing characteristics, have shown great potential for dense 3D reco

FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2608.09591v1 Announce Type: cross Abstract: Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing met

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

Model ReleasesDGX agent

arXiv:2608.09474v1 Announce Type: new Abstract: Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained prima

Financial Numerical Prediction and Allocation as Token Generation

SafetyDGX agent

arXiv:2608.09880v1 Announce Type: new Abstract: Financial prediction typically relies on task-specific regression, ranking, or policy heads, separating the language model from the numerical object ult

FiRe: Fixed-Noise Refinement for Visual Counterfactual Explanations

Local AiDGX agent

arXiv:2608.08664v1 Announce Type: new Abstract: Visual counterfactual explanations aim to change classifier decisions through realistic and localized edits while preserving decision-irrelevant content

FlexSplat: Flexible Feed-Forward 3D Gaussian Splatting without Point Cloud Correspondence

ResearchDGX agent

arXiv:2608.07937v1 Announce Type: new Abstract: We present FlexSplat, a feed-forward framework for novel view synthesis (NVS) from uncalibrated, object-centric multi-view image collections. A recent l

Flow-based conditional cardiac anatomy generation for virtual cohorts

ResearchDGX agent

arXiv:2608.09460v1 Announce Type: cross Abstract: Cardiac digital twin research is moving from subject-specific anatomical replicas toward virtual cohorts that represent clinically relevant population

FlowErase-OPD: Multi-Concept Erasure via Anchored On-Policy Distillation in Flow Matching Models

SafetyDGX agent

arXiv:2608.07620v1 Announce Type: new Abstract: Recent advances in flow matching models have substantially improved the quality of text-to-image generation, but have also raised increasing safety conc

Forged Peer Judgments Mislead Multimodal LLM Judge Panels: Source-Blind Anchoring and Panel-Consensus Verification

ResearchDGX agent

arXiv:2608.07920v1 Announce Type: new Abstract: Multimodal LLM judge panels can cross-reference peers, but a quoted peer judgment may itself be untrusted. We expose source-blind anchoring as a text-le

Forgetting-Resistant and Lesion-Aware Source-Free Domain Adaptive Fundus Image Analysis with Vision-Language Model

Model ReleasesDGX agent

arXiv:2602.19471v2 Announce Type: replace Abstract: Source-free domain adaptation (SFDA) aims to adapt a model trained in the source domain to perform well in the target domain, with only unlabeled ta

Foundation Models are Implicit Deepfake Detectors

ResearchDGX agent

arXiv:2608.09427v1 Announce Type: new Abstract: Pretrained self-supervised representations have emerged as a core component of current deepfake detection methods, yet it remains unclear which of their

FreCast: Refining Radar Echo Intensity via Phase-Preserving Amplitude Residual Diffusion for Precipitation Nowcasting

TutorialsDGX agent

arXiv:2608.08436v1 Announce Type: new Abstract: Precipitation nowcasting predicts the spatiotemporal evolution of future radar echoes from historical radar echo sequences, thereby estimating the occur

From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.07770v1 Announce Type: cross Abstract: Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial con

From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing

Model ReleasesDGX agent

arXiv:2608.09842v1 Announce Type: new Abstract: Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on comp

From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition

SafetyDGX agent

arXiv:2308.04553v4 Announce Type: replace Abstract: Visual recognition models are prone to learning spurious correlations induced by a biased training set where certain conditions B (eg, Indoors) are

From Noise to Meaning: Meaningful Secret Sharing with Tamper Detection for Facial Recognition

ResearchDGX agent

arXiv:2608.08924v1 Announce Type: new Abstract: Popularity of AI-based face recognition system directly demands protection of sensitive biometric data used for training. Visual secret sharing is an in

From Visual to Multimodal: Systematic Ablation of Encoders and Fusion Strategies in Animal Identification

ResearchDGX agent

arXiv:2603.02270v2 Announce Type: replace Abstract: Automated animal identification is a practical task for reuniting lost pets with their owners, yet current systems often struggle due to limited dat

Gated Spatial Redundancy Projection for Pathology Transformer Attentions

ResearchDGX agent

arXiv:2608.08374v1 Announce Type: new Abstract: Transformer models are increasingly used for whole-slide image analysis in computational pathology. Yet, WSIs differ fundamentally from natural images:

GenTrack3: Hybrid Stochastic-Deterministic Online Multi-Object Tracking with Cluster-Aware Association

ResearchDGX agent

arXiv:2608.09581v1 Announce Type: new Abstract: Multi-object tracking (MOT) involves maintaining consistent target identities as objects dynamically enter and leave a scene. Deterministic approaches,

GeoAI-based post-segmentation quality validation of building footprints via spatial feature engineering

ApplicationsDGX agent

arXiv:2608.09048v1 Announce Type: new Abstract: Deep learning-based building footprint extraction from high-resolution imagery often produces topologically inconsistent vectors unfit for direct GIS da

GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

Model ReleasesDGX agent

arXiv:2608.09493v1 Announce Type: new Abstract: Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains chal

Ghost Features and Spooky Transfer Learning for Hypercomplex-Valued Neural Networks

TutorialsDGX agent

arXiv:2608.07735v1 Announce Type: new Abstract: Hypercomplex numbers extend the concept of complex numbers by introducing additional imaginary components. Besides increasing dimensionality, operations

Goal-oriented Navigation Instruction Generation with Tour Video Priors

Model ReleasesDGX agent

arXiv:2608.08596v1 Announce Type: new Abstract: Navigation Instruction Generation (NIG) aims to produce step-by-step natural language instructions for navigation guidance. Existing studies primarily t

HandSplatter: Automated Digital Goniometry from Neural Rendering

ResearchDGX agent

arXiv:2608.09735v1 Announce Type: new Abstract: Hand and finger disorders are leading contributors to musculoskeletal disability, creating a clinical need for precise methods to quantify joint motion.

HeatCast: A Benchmark for Neighborhood-Scale LST Forecasting across 124 U.S. Cities

Model ReleasesDGX agent

arXiv:2608.07640v1 Announce Type: new Abstract: Land Surface Temperature (LST) is a widely used satellite-derived measure of urban surface heat, but there is no shared benchmark for forecasting it at

High-Capacity Generalized Hopfield Networks

SafetyDGX agent

arXiv:2608.08226v1 Announce Type: cross Abstract: Generalized Hopfield networks are introduced where memories and neurons are continuous variables that lie on a Riemannian manifold. We explicitly focu

High-Quality Exposure Correction with Diffusion-Based Image Generation Priors

ResearchDGX agent

arXiv:2608.08720v1 Announce Type: new Abstract: Although most existing exposure correction methods achieve high fidelity, they often place excessive focus on overall pixel-wise accuracy, making it cha

HonestFace: Towards Honest Face Restoration with One-Step Diffusion Model

SafetyDGX agent

arXiv:2505.18469v2 Announce Type: replace Abstract: Face restoration has achieved significant advancements through the years of development. However, maintaining high fidelity and authenticity while a

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems

Model ReleasesDGX agent

arXiv:2608.07861v1 Announce Type: new Abstract: Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasse

HSMLA: Hierarchical Softmax Multi-scale Linear Attention for Efficient Vision Transformers

Local AiDGX agent

arXiv:2608.07616v1 Announce Type: new Abstract: Vision transformers face significant computational overheads in high-resolution dense prediction due to the quadratic complexity of self-attention. Line

IDATA: Scalable Invertible Diffusion for Unrestricted Adversarial Transfer Attack

ResearchDGX agent

arXiv:2608.08734v1 Announce Type: new Abstract: Unrestricted adversarial transfer attacks are important for evaluating the black-box robustness of deep visual models. Diffusion-based attacks have show

Impact of Dataset Composition on Embedded Real-Time UAV Wildfire Detection Using Compact YOLO Models

SafetyDGX agent

arXiv:2608.07554v1 Announce Type: new Abstract: The development of vision-based wildfire detection systems for unmanned aerial vehicles is constrained by the limited availability of diverse real-world

In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation

SafetyDGX agent

arXiv:2608.09244v1 Announce Type: new Abstract: Text-to-image diffusion models have achieved remarkable success in generating high-quality images from a given text prompt. Subject-driven generation ai

InstructionCrafter: Generating Consistent and High-Fidelity Visual Instructions

Model ReleasesDGX agent

arXiv:2608.08460v1 Announce Type: new Abstract: Given textual task instructions, generating step-by-step visual instructions as an image sequence requires the simultaneous satisfaction of multiple pro

IRPol-Fuse: Energy-structure coordination for infrared polarization fusion under low visibility

ResearchDGX agent

arXiv:2608.07848v1 Announce Type: new Abstract: Robust perception under low-visibility conditions requires fused imagery that jointly preserves infrared thermal saliency and polarization-derived struc

Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models

ResearchDGX agent

arXiv:2604.02048v2 Announce Type: replace Abstract: Developing vision-language models (VLMs) that generalize across diverse tasks requires large-scale training datasets with diverse content. In Englis

JSGS: JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views

ResearchDGX agent

arXiv:2608.08659v1 Announce Type: new Abstract: Standard 3D Gaussian Splatting (3DGS) assumes that every input image faithfully samples scene radiance. However, mixed-quality JPEG images violate this

← Previous
123456…207
Next →