AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
30 Jun 2026

Benchmark AUC Is Not Deployable Reliability: A Cross-Dataset Audit of Off-the-Shelf Features for Surveillance Video Anomaly Detection

Model ReleasesDGX agent

arXiv:2606.29506v1 Announce Type: new Abstract: Automated 'suspicious behavior' flagging is a headline promise of AI surveillance, and the field reports high frame-level ROC-AUC on standard video anom

Benchmarking Geospatial Foundation Models for Agriculture Applications

Model ReleasesDGX agent

arXiv:2606.29664v1 Announce Type: new Abstract: Geospatial foundation models pretrained on satellite imagery promise broad generalization across remote sensing tasks and regions, but their geographic

Beyond Backscatter: AlphaEarth Land-Cover Priors for Rapid SAR Flood Segmentation Across Foundation Backbones

Safety

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.29134v1 Announce Type: new Abstract: Rapid flood mapping is critical for emergency response, yet optical imagery is often unusable during major flooding and single-temporal SAR is ambiguous

Beyond Trajectory Matching: Reflow with Marginal Distribution Alignment

Model ReleasesDGX agent

arXiv:2606.29287v1 Announce Type: cross Abstract: Diffusion and continuous-flow generative models achieve high-quality generation, and their deterministic sampling can be formulated as solving learned

Bit-ViP: Leveraging Bit-planes to Preserve Visual Privacy in Images through Obfuscation

ResearchDGX agent

arXiv:2606.29417v1 Announce Type: new Abstract: The unprecedented growth of computer vision applications, such as surveillance systems and social media, raises security and visual privacy concerns, es

BLUE: A Stale-Pixel Optical-Flow Compositor for Entropy-Efficient Surveillance Video Encoding

ResearchDGX agent

arXiv:2606.28753v1 Announce Type: cross Abstract: Continuous-recording surveillance systems face a storage problem that codec tuning alone cannot fully solve: even at aggressive CRF settings, a static

BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

SafetyDGX agent

arXiv:2606.30319v1 Announce Type: new Abstract: Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscien

BrepLLM: Enabling Large Language Models to Understand Boundary Representations

Model ReleasesDGX agent

arXiv:2512.16413v2 Announce Type: replace Abstract: Current token-sequence-based Large Language Models (LLMs) struggle to directly process 3D Boundary Representation (B-rep) models that contain comple

Bricker to BRACE: A Bracket Exposure RAW Dataset and Restoration Model for Flicker-Banding

ResearchDGX agent

arXiv:2606.29845v1 Announce Type: new Abstract: Flicker-banding (FB), arises from temporal aliasing between a camera's rolling shutter and a display's brightness modulation, degrading screen-captured

Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance

Model ReleasesDGX agent

arXiv:2606.30049v1 Announce Type: new Abstract: Visibility distance is critical to maritime navigational safety because it determines the effective observation range of shipborne and shore-based monit

Building artificial intelligence virtual tissue (AIVT) for tissue state representation, feature prediction, and dynamic simulation

TutorialsDGX agent

arXiv:2606.29883v1 Announce Type: new Abstract: Modeling tissue states and their transitions is essential for understanding tissue homeostasis in health and pathological remodeling in disease. However

Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models

Model ReleasesDGX agent

arXiv:2606.28406v1 Announce Type: cross Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design sch

CellDETR: A Detection-Guided Framework for Scalable Cell Representation Learning from Histopathology Images

ResearchDGX agent

arXiv:2606.29463v1 Announce Type: new Abstract: Recent advances in pathology foundation models have substantially improved patch and slide level representation learning from whole-slide images (WSIs).

Character Recognition of Nepali Number Plate

ApplicationsDGX agent

arXiv:2606.28946v1 Announce Type: new Abstract: This paper presents a robust Automatic Number Plate Recognition (ANPR) system tailored for Nepali license plates written in Devanagari script. In this p

CLEAR-MoE: Shared-Basis Expert Extraction from Frozen Vision Transformers via Calibration-Driven Layer Selection

HardwareDGX agent

arXiv:2606.28516v1 Announce Type: new Abstract: We present CLEAR-MoE, a four-phase post-training pipeline that converts a frozen pretrained Vision Transformer (ViT) into a sparse Mixture-of-Experts (M

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation

SafetyDGX agent

arXiv:2606.29805v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are prone to hallucination as their generation preferences are insufficiently calibrated to visual evidence, ca

CLIMP: Contrastive Language-Image Mamba Pretraining

ResearchDGX agent

arXiv:2601.06891v2 Announce Type: replace Abstract: Contrastive Language-Image Pre-training (CLIP) relies on Vision Transformers whose attention mechanism is susceptible to spurious correlations, and

Clinical Risk-Aware Multi-Level Grading for Coronary Artery Stenosis through Curved Feature Reconstruction

ResearchDGX agent

arXiv:2606.30082v1 Announce Type: new Abstract: Developing a multi-level grading model for coronary artery stenosis holds great clinical significance for the diagnosis of coronary artery disease. Howe

ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation

Local AiDGX agent

arXiv:2512.02453v2 Announce Type: replace Abstract: Existing stylized motion generation models have shown their remarkable ability to understand specific style information from the style motion, and i

CoGS: Compositional Dynamic Human-Object Scenes Gaussian Splatting from Monocular Video

Model ReleasesDGX agent

arXiv:2606.28820v1 Announce Type: new Abstract: Reconstructing dynamic human--object interaction scenes from monocular video is difficult because the human, manipulated object, and background obey dif

CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion

ApplicationsDGX agent

arXiv:2606.30030v1 Announce Type: new Abstract: Blind image deblurring demands the recovery of high-fidelity details and coherent structures from complex, unknown degradations. Current blind image deb

CollabOD: Collaborative Multi-Backbone with Cross-scale Vision for UAV Small Object Detection

Local AiDGX agent

arXiv:2603.05905v2 Announce Type: replace Abstract: Small object detection in unmanned aerial vehicle (UAV) imagery is challenging because high-altitude viewpoints produce severe scale variation, weak

CoLR-Det: Collaborative Latent Restoration for Small Object Detection in Low-Resolution Remote Sensing Images

ResearchDGX agent

arXiv:2601.12507v2 Announce Type: replace Abstract: Low-resolution remote sensing small object detection is limited by both missing visual details and the ambiguity of how details serve detection. Exi

Complete virtual unwrapping and reading of a rolled Herculaneum papyrus

ResearchDGX agent

arXiv:2606.29085v1 Announce Type: cross Abstract: The carbonized papyri from Herculaneum preserve the only large-scale library to survive from classical antiquity, but many unopened rolls remain unrea

Concept Removal Guidance: Evidence-Calibrated Negative Guidance for Safe Diffusion Sampling

SafetyDGX agent

arXiv:2606.29801v1 Announce Type: new Abstract: Text-to-image diffusion models remain vulnerable to adversarial prompts that elicit disallowed content, motivating reliable inference-time controls. A p

Consensus Clustering of Free-Viewing Gaze Data: New Insights into Human-Information Interaction

ResearchDGX agent

arXiv:2606.30035v1 Announce Type: new Abstract: Free-viewing gaze data provides a rich, task-free window into human visual attention. Conventional exploratory data analysis of the data provides user a

Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning

SafetyDGX agent

arXiv:2606.29812v1 Announce Type: new Abstract: Inductive biases steer learning toward generalizable solutions by encoding task structure. In this work, we identify a crucial missing bias in MLLMs: cr

Contrastive vision-language learning with paraphrasing and negation

TutorialsDGX agent

arXiv:2511.16527v2 Announce Type: replace Abstract: Contrastive vision-language models continue to be the dominant approach for image-text retrieval. Contrastive Language-Image Pre-training (CLIP) tra

CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates

Model ReleasesDGX agent

arXiv:2512.10342v3 Announce Type: replace Abstract: Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual deci

CouCE: A Unified Causal Framework for Debiased Deep Metric Learning

ResearchDGX agent

arXiv:2606.30365v1 Announce Type: new Abstract: Deep Metric Learning (DML) often struggles with zero-shot generalization because standard objectives inherently capture what co-occurs rather than what

Cross-Modal Iteration Distillation for Robust IHD Screening: The IDNet Framework and A New Benchmark

Model ReleasesDGX agent

arXiv:2606.30027v1 Announce Type: new Abstract: Color Fundus Photography (CFP) offers a low-cost and non-invasive route for ischemic heart disease (IHD) screening, but current studies are limited by s

Cross-Resolution Semantic Transfer for Robust Text-to-Image Retrieval in Low-Resolution Surveillance

ApplicationsDGX agent

arXiv:2606.30458v1 Announce Type: new Abstract: Text-to-image person re-identification (TIPR) retrieves target persons using natural language descriptions. However, existing methods largely overlook r

CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation

ResearchDGX agent

arXiv:2604.02948v2 Announce Type: replace Abstract: Multimodal semantic segmentation has shown great potential in leveraging complementary information across diverse sensing modalities. However, exist

CylindTrack: Depth-Aware Cylindrical Motion Modeling for Panoramic Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2606.30097v1 Announce Type: new Abstract: Multi-Object Tracking (MOT) is a core capability for embodied perception, and panoramic cameras are attractive for embodied systems because their 360{eg

D^{2}R^{2}OSR: Degradation-Disentangled Representation for Real-World Omnidirectional Image Super-Resolution

ApplicationsDGX agent

arXiv:2606.29314v1 Announce Type: new Abstract: With the growing demand for immersive visual experiences, high-quality omnidirectional images (ODIs) have become increasingly important. However, limita

DCGrasp: Distance-aware Controllable Grasp Generation

ResearchDGX agent

arXiv:2606.29924v1 Announce Type: new Abstract: Generating 3D hand-object interactions is essential for applications in robotics, XR, and synthetic data generation, where flexible controllability and

DCSNet: Multiscale Feature Aggregation for Small Medical Object Segmentation with Detection-guided Hierarchical Cropping

ResearchDGX agent

arXiv:2606.28402v1 Announce Type: new Abstract: Small object segmentation in medical imaging is primarily hindered by class imbalance and inherent boundary complexity. Consequently, conventional globa

DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering

ResearchDGX agent

arXiv:2602.19323v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful paradigm for real-time and high-fidelity 3D reconstruction from posed images. However, recent

Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation

SafetyDGX agent

arXiv:2512.20117v2 Announce Type: replace Abstract: Audio-Visual Segmentation (AVS) aims to localize sound-producing objects at the pixel level by integrating auditory and visual cues. However, existi

DeVAR: Low-Dose CT Denoising via Visual Autoregressive Modeling

ResearchDGX agent

arXiv:2606.28453v1 Announce Type: cross Abstract: Computed tomography (CT) plays a crucial role in medical diagnosis, but minimizing radiation exposure while maintaining image quality remains a critic

DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent

ResearchDGX agent

arXiv:2606.28417v1 Announce Type: new Abstract: Recently, diffusion models have been widely adopted in generative modeling and have served as foundational models for many image generation tasks. To co

Distribution Matching Variational AutoEncoder

SafetyDGX agent

arXiv:2512.07778v2 Announce Type: replace Abstract: Most visual generative models compress images into a latent space before applying diffusion or autoregressive modelling. Yet, existing approaches su

DivAS: Interactive 3D Segmentation by Depth-Weighted Voxel Aggregation

ResearchDGX agent

arXiv:2601.04860v2 Announce Type: replace Abstract: Interactive 3D segmentation of a reconstructed scene should not require a representation-specific optimization loop. We observe that the recipe for

DLGStream: Dynamic Language-embedded Guassian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming

ResearchDGX agent

arXiv:2606.28840v1 Announce Type: new Abstract: 3D Gaussian Splatting~(3DGS) has emerged as a promising paradigm for reconstructing streamable free-viewpoint video~(FVV) from multi-view videos. Howeve

DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model

HardwareDGX agent

arXiv:2606.30292v1 Announce Type: cross Abstract: We present DreamForge-World 0.1 Preview, a preview foundational world model for real-time interactive world simulation. The system adapts the LongLive

DRESS: Disentangled Representation-based Self-Supervised Meta-Learning for Diverse Tasks

ResearchDGX agent

arXiv:2503.09679v2 Announce Type: replace-cross Abstract: Meta-learning represents a strong class of approaches for solving few-shot learning tasks. Nonetheless, recent research suggests that simply p

Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting

ResearchDGX agent

arXiv:2603.28049v3 Announce Type: replace Abstract: Autoregressive (AR)-Diffusion hybrid paradigms combine AR's structured semantic modeling with diffusion's high-fidelity synthesis, yet suffer from a

DrivenMorph: Bridging Attention Mechanism and Variational Image Registration via Difference Modeling

SafetyDGX agent

arXiv:2606.30183v1 Announce Type: new Abstract: Medical image registration benefits significantly from deep learning, yet existing approaches often lack physical explainability and fine-grained deform

DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution

SafetyDGX agent

arXiv:2606.29198v1 Announce Type: new Abstract: As the most perceptually powerful Face Video Super-Resolution (FVSR) method, existing works in Generative FVSR (GFVSR) mainly exploit the generative pri

Dynamic High-frequency Convolution for Infrared Small Target Detection

ResearchDGX agent

arXiv:2602.02969v2 Announce Type: replace Abstract: Infrared small targets are typically tiny and locally salient, which belong to high-frequency components (HFCs) in images. Single-frame infrared sma

Early Estimation of Language to Latent Alignment in Diffusion Models

Model ReleasesDGX agent

arXiv:2512.08505v2 Announce Type: replace Abstract: Conditional diffusion models frequently suffer from language-image misalignments. Due to the ambiguity of intermediate noise corrupted latents, asse

EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics

ResearchDGX agent

arXiv:2606.30557v1 Announce Type: new Abstract: DiT video generation is latency-intensive due to iterative full-frame denoising, while prior cloud-edge methods largely rely on static inter-step decoup

Efficient 3D Gaussian Splatting with Axis-Shared Rasterization and Order-independent Transmittance

AgentsDGX agent

arXiv:2506.07069v2 Announce Type: replace-cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, combining high-quality reconstruction with efficien

Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction

Model ReleasesDGX agent

arXiv:2606.29850v1 Announce Type: new Abstract: Visual pointing maps a language instruction to pixel co ordinates, a core skill for embodied AI. We describe our PointArena 2026 solution, which achieve

Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation

AgentsDGX agent

arXiv:2512.10310v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated significant promise in Vision-Language Navigation (VLN), existing agents remain hea

Emergence of a Shared Canonical Object Frame from In-the-Wild Videos

ResearchDGX agent

arXiv:2606.30058v1 Announce Type: new Abstract: Comparing object orientations and positions across different instances requires their poses to be expressed in a shared canonical frame. Establishing su

Empirical Evaluation of Multi-Modal Touch Detection in Over-the-Shoulder Video Surveillance

ResearchDGX agent

arXiv:2606.29504v1 Announce Type: new Abstract: Video Intelligence Surveillance (VIDINT) on over-the-shoulder footage is a proposed vector for monitoring human-computer interaction patterns without di

End-to-End Facial Expression Detection in Long Videos

ApplicationsDGX agent

arXiv:2504.07660v2 Announce Type: replace Abstract: Facial expression detection requires spotting when expressions occur and recognizing which emotional category they belong to. Despite their close re

Enhancing Layer Interaction Using Key-Correlated Layer Attention

ResearchDGX agent

arXiv:2606.28405v1 Announce Type: new Abstract: Recent advances in network architecture design have introduced layer attention to enhance inter-layer interactions. In such frameworks, each layer queri

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

ResearchDGX agent

arXiv:2606.29267v1 Announce Type: new Abstract: Visual grounding aims to associate free-form textual queries with specific regions in an image. While recent Multimodal Large Language Models (MLLMs) ha

← Previous
1…6061626364…209
Next →