AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners

DGX agent

arXiv:2607.00125v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable abilities when analyzing images, yet translating these capabilities to few-shot im

researcharxiv-cs-cv
2 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Decoupled Guidance: Disentangling Subject and Context Pathways in Text-to-Image Personalization

DGX agent

arXiv:2607.00766v1 Announce Type: new Abstract: Text-to-image personalization aims to generate a user-provided subject in novel scenes described by text. However, most existing methods encode subject

researcharxiv-cs-cv
2 Jul 2026
Research

Diffusion-Based Multi-Class Normality for OOD Detection: An Application to CDP Authentication

DGX agent

arXiv:2607.00609v1 Announce Type: new Abstract: Reconstruction-based generative models offer a natural framework for unsupervised out-of-distribution (OOD) detection, but multi-class normality modelli

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Does Your ViT Still Need U-Net for Segmentation?

DGX agent

arXiv:2607.00223v1 Announce Type: new Abstract: Medical image segmentation is dominated by U-Net-style encoder-decoder architectures. Vision Transformers (ViTs) overcome the limited receptive field of

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

DGX agent

arXiv:2607.00666v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models often fail to perform the same learned tasks under environmental shifts, such as changes in camera pose and shifts

safetyarxiv-cs-cv
2 Jul 2026
Tutorials

DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation

DGX agent

arXiv:2607.00183v1 Announce Type: new Abstract: Adapting pre-trained text-to-image diffusion models, whether to learn new visual concepts or erase unwanted ones, is routinely evaluated on its intended

tutorialsarxiv-cs-cv
2 Jul 2026
Model Releases

DriveVA: Video Action Models are Zero-Shot Drivers

DGX agent

arXiv:2604.04198v2 Announce Type: replace Abstract: Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor dom

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving

DGX agent

arXiv:2607.00399v1 Announce Type: new Abstract: End-to-end autonomous driving models often encounter performance bottlenecks, as training-time scaling leads to high computational costs and diminishing

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images

DGX agent

arXiv:2607.00338v1 Announce Type: new Abstract: Object detection for Unmanned Aerial Vehicles (UAVs) working in open and dynamic environments is a highly challenging task. While Vision-Language Models

model-releasesarxiv-cs-cv
2 Jul 2026
Research

DroneIQA-VLE: Multi-Task Drone Image Quality Assessment via Vision-Language Ensemble

DGX agent

arXiv:2607.00416v1 Announce Type: new Abstract: We present DroneIQA-VLE, our solution to the ICME 2026 Drone-IQA Grand Challenge on Target-aware Image Quality Assessment for Low-altitude UAV Images. T

researcharxiv-cs-cv
2 Jul 2026
Research

E-TIDE: Fast, Structure-Preserving Motion Forecasting from Event Sequences

DGX agent

arXiv:2603.27757v2 Announce Type: replace Abstract: Event-based cameras capture visual information as asynchronous streams of per-pixel brightness changes, generating sparse, temporally precise data.

researcharxiv-cs-cv
2 Jul 2026
Safety

ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation

DGX agent

arXiv:2607.00545v1 Announce Type: new Abstract: Controllable traffic simulation is critical for testing autonomous driving systems, yet existing approaches often require retraining large generative mo

safetyarxiv-cs-cv
2 Jul 2026
Local Ai

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection

DGX agent

arXiv:2607.00867v1 Announce Type: new Abstract: Long-video reasoning is fundamentally constrained by how models acquire and utilize visual evidence. Existing tool-augmented video frameworks often inte

local-aiarxiv-cs-cv
2 Jul 2026
Model Releases

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

DGX agent

arXiv:2512.15702v2 Announce Type: replace Abstract: Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. Wh

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking

DGX agent

arXiv:2412.20750v3 Announce Type: replace Abstract: Large-scale Vision-Language Models (VLMs) have achieved notable progress in aligning visual inputs with text. However, their ability to deeply under

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

EPO: Boosting 3D Foundation Models with Edge-based Pose Optimization

DGX agent

arXiv:2607.00579v1 Announce Type: new Abstract: We introduce extbf{Edge-based Pose Optimization (EPO)}, a trackless geometric optimization framework specifically designed to boost the Structure-from-M

safetyarxiv-cs-cv
2 Jul 2026
Safety

EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation

DGX agent

arXiv:2607.01147v1 Announce Type: new Abstract: Text-to-image diffusion models power everyday creative tasks, but they still reproduce the demographic biases in their training data. On common prompts

safetyarxiv-cs-cv
2 Jul 2026
Research

Explainability in mulimodal deep transformation models for stroke outcome prediction

DGX agent

arXiv:2504.06299v2 Announce Type: replace-cross Abstract: Multimodal prediction models based on imaging and clinical data are increasingly used for clinical decision support, yet their interpretabilit

researcharxiv-cs-cv
2 Jul 2026
Research

FCL-COD: Weakly Supervised Camouflaged Object Detection with Frequency-aware and Contrastive Learning

DGX agent

arXiv:2603.22969v2 Announce Type: replace Abstract: Existing camouflage object detection (COD) methods typically rely on fully-supervised learning guided by mask annotations. However, obtaining mask a

researcharxiv-cs-cv
2 Jul 2026
Research

Foundation Model-driven Key Anatomy Frame Selection for Blind-sweep Ultrasound Fetal Birth Weight Estimation

DGX agent

arXiv:2607.00745v1 Announce Type: new Abstract: Accurate fetal birth weight (FBW) estimation shortly before delivery is clinically valuable yet challenging due to its reliance on operator expertise, p

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Foundation Models vs. Radiomics for Lung Computed Tomography: A Benchmark of Feature Extractors, Classification Heads, and Segmentation Choices

DGX agent

arXiv:2607.01001v1 Announce Type: new Abstract: Radiomics is the established approach for CT-based lung cancer phenotyping, yet comparisons with foundation models rarely isolate contributions of featu

model-releasesarxiv-cs-cv
2 Jul 2026
Safety

FrameONE: Hierarchical Motion Modeling for Universal Multi-View Echocardiographic Keyframe Detection

DGX agent

arXiv:2607.00748v1 Announce Type: new Abstract: Accurate detection of end-systole (ES) and end-diastole (ED) frames is fundamental to echocardiographic assessment. Existing methods are typically devel

safetyarxiv-cs-cv
2 Jul 2026
Model Releases

FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal

DGX agent

arXiv:2603.19036v2 Announce Type: replace Abstract: Single image reflection removal (SIRR) is challenging in real scenes, where reflection strength varies spatially and reflection patterns are tightly

model-releasesarxiv-cs-cv
2 Jul 2026
Research

GADA: Geometry-Aware Deformable Aggregation for Image-Based Gaussian Splatting

DGX agent

arXiv:2607.00595v1 Announce Type: new Abstract: Gaussian Splatting has achieved significant improvements by incorporating warping-based techniques. However, such methods suffer from pixel-level inaccu

researcharxiv-cs-cv
2 Jul 2026
Research

GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting

DGX agent

arXiv:2607.00959v1 Announce Type: new Abstract: Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avat

researcharxiv-cs-cv
2 Jul 2026
Tutorials

GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

DGX agent

arXiv:2603.26661v2 Announce Type: replace Abstract: Most recent advances in 3D generative modeling rely on diffusion or flow-matching formulations. We instead explore a fully autoregressive alternativ

tutorialsarxiv-cs-cv
2 Jul 2026
Model Releases

GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine

DGX agent

arXiv:2607.00544v1 Announce Type: new Abstract: Reasoning segmentation requires localizing targets based on complex, implicit queries. Current end-to-end models typically entangle perception and deduc

model-releasesarxiv-cs-cv
2 Jul 2026
Local Ai

GenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language Models

DGX agent

arXiv:2607.01049v1 Announce Type: new Abstract: Industrial inspection requires more than binary anomaly detection: a practical system should determine whether an anomaly exists, localize the defective

local-aiarxiv-cs-cv
2 Jul 2026
Research

Generated Contents Enrichment

DGX agent

arXiv:2405.03650v4 Announce Type: replace Abstract: We study Generated Contents Enrichment (GCE), a conditional image-generation task in which a sparse scene description is first enriched through an e

researcharxiv-cs-cv
2 Jul 2026
Tutorials

GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models

DGX agent

arXiv:2607.00492v1 Announce Type: new Abstract: We introduce GenSP, a data-driven framework that learns consistent spherical parameterizations across a collection of genus-0 shapes. Instead of optimiz

tutorialsarxiv-cs-cv
2 Jul 2026
Applications

Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics

DGX agent

arXiv:2603.13859v2 Announce Type: replace Abstract: Intrinsic image decomposition aims to estimate physically based rendering (PBR) parameters such as albedo, roughness, and metallicity from images. W

applicationsarxiv-cs-cv
2 Jul 2026
Model Releases

Geometry-Aware Cross-Height Channel Knowledge Map Prediction for UAV-Assisted Communications With Uncertainty-Guided 3D Sensing

DGX agent

arXiv:2607.00887v1 Announce Type: new Abstract: Low-altitude Unmanned Aerial Vehicles (UAVs) often need to infer channel knowledge across a range of heights from only sparse observations collected at

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision

DGX agent

arXiv:2607.01050v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have shown strong cross-modal understanding and coordinate generation abilities in visual grounding. How

model-releasesarxiv-cs-cv
2 Jul 2026
Research

GimbalDiffusion: Gravity-Aware Camera Control for Video Generation

DGX agent

arXiv:2512.09112v3 Announce Type: replace Abstract: Recent progress in text-to-video generation has achieved remarkable realism, yet fine-grained control over camera motion and orientation remains elu

researcharxiv-cs-cv
2 Jul 2026
Model Releases

GKDT: General Keypoint Detection Transformer

DGX agent

arXiv:2607.00752v1 Announce Type: new Abstract: With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The con

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

GMO-E^2DIT: Grounded Multi-Operation Editing for E-Commerce Images

DGX agent

arXiv:2607.00920v1 Announce Type: new Abstract: Real-world e-commerce image editing often requires multiple, localized, and auditable operations rather than global restyling. This compositional nature

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition

DGX agent

arXiv:2602.03370v2 Announce Type: replace Abstract: Handwritten mathematical expression recognition (HMER) requires reasoning over diverse symbols and structures, yet autoregressive models struggle wi

model-releasesarxiv-cs-cv
2 Jul 2026
Research

HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking

DGX agent

arXiv:2607.00494v1 Announce Type: new Abstract: Multi-animal tracking (MAT) is critical for wildlife monitoring and behavioral analysis, yet remains challenging due to uniform appearance, high density

researcharxiv-cs-cv
2 Jul 2026
Research

High-dimensional Embedding Prior for Noisy K-space Domain MRIReconstruction

DGX agent

arXiv:2607.01176v1 Announce Type: new Abstract: Magnetic resonance imaging (MRI) reconstruction under realistic acquisition conditions can be fundamentally viewed as estimating the underlying k-space

researcharxiv-cs-cv
2 Jul 2026
Model Releases

Histopathology Multi-modal Embedding for Pathology Composed Retrieval

DGX agent

arXiv:2502.07221v4 Announce Type: replace Abstract: To overcome the black-box nature of predictive AI and the hallucination risks of generative models, retrieval-based models offer an interpretable, e

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation

DGX agent

arXiv:2607.01067v1 Announce Type: cross Abstract: As an essential modality for dexterous and contact-rich tasks, tactile sensing provides precise force feedback that cannot be reliably inferred from v

researcharxiv-cs-cv
2 Jul 2026
Safety

HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding

DGX agent

arXiv:2607.00428v1 Announce Type: new Abstract: CLIP (Contrastive Language-Image Pre-training) has become a de facto paradigm for image-text alignment, but it struggles with long-context descriptions

safetyarxiv-cs-cv
2 Jul 2026
Model Releases

Imprint: Online Memory Compression for Long-Horizon Egocentric QA

DGX agent

arXiv:2607.00696v1 Announce Type: new Abstract: Long-horizon egocentric question answering involves answering about events that have occurred hours or days in the past. This requires memory representa

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Information-Regularized Attention for Visual-Centric Reasoning

DGX agent

arXiv:2607.00434v1 Announce Type: new Abstract: Vision-language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, an

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models

DGX agent

arXiv:2607.01222v1 Announce Type: new Abstract: Recent 3D generative models can synthesize high-quality geometry but often struggle to reproduce intricate textures from reference images, largely due t

researcharxiv-cs-cv
2 Jul 2026
Model Releases

IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video

DGX agent

arXiv:2603.16432v3 Announce Type: replace Abstract: Unsupervised physical parameter estimation from video lacks a common benchmark: existing methods evaluate on non-overlapping synthetic data, the sol

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Joint Medical Image Enhancement and Segmentation with Diffusion-based Symbiotic Information Interaction

DGX agent

arXiv:2607.00058v1 Announce Type: new Abstract: Image quality is critical for accurate medical diagnosis. However, MRI, CT, and ultrasound images are often of low resolution and quality due to cost co

researcharxiv-cs-cv
2 Jul 2026
Tutorials

Learn Once, Edit Anywhere: Visual Direction Transfer for Diffusion Models

DGX agent

arXiv:2403.19645v2 Announce Type: replace Abstract: The rapid advancement of diffusion models has enabled the generation of high-fidelity images from textual prompts, yet achieving precise, disentangl

tutorialsarxiv-cs-cv
2 Jul 2026
← Previous
1…7071727374…263
Next →