AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

Expert Consensus on Criteria for the Automated Assessment of Laparoscopic Camera Navigation

DGX agent

arXiv:2606.23131v1 Announce Type: new Abstract: Background: Laparoscopic camera navigation (LCN) is a critical skill, yet its current assessment typically relies on manual rating systems which are tim

safetyarxiv-cs-cv
23 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders

DGX agent

arXiv:2606.21197v1 Announce Type: new Abstract: Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioni

safetyarxiv-cs-cv
23 Jun 2026
Research

Failure Analysis in Transition: An Industry Survey of Challenges, Priorities, and Standardization Needs in Advanced Packaging and Heterogeneous Integration

DGX agent

arXiv:2606.22149v1 Announce Type: cross Abstract: Failure analysis is being reshaped by heterogeneous integration, chiplet-based architectures, hybrid bonding, backside technologies, & increasingly bu

researcharxiv-cs-cv
23 Jun 2026
Research

Faithful Grounded Visual Reasoning via Learned Proxy-Tokens

DGX agent

arXiv:2606.23354v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their 'black-box' nature hinders deplo

researcharxiv-cs-cv
23 Jun 2026
Research

FCR: Investigating Generative AI models for Forensic Craniofacial Reconstruction

DGX agent

arXiv:2508.18031v2 Announce Type: replace Abstract: Craniofacial reconstruction in forensics is one of the processes to identify victims of crime and natural disasters. Identifying an individual from

researcharxiv-cs-cv
23 Jun 2026
Research

FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs

DGX agent

arXiv:2606.22875v1 Announce Type: new Abstract: Training Latent Diffusion Models (LDMs) within Federated Learning (FL) has attracted increasing attention due to its ability to combine the powerful gen

researcharxiv-cs-cv
23 Jun 2026
Research

Feed-forward Motion In-betweening for Any 4D

DGX agent

arXiv:2606.22131v1 Announce Type: new Abstract: 4D dynamics (3D geometry evolving over time) is a fundamental representation of the physical world and plays a crucial role in world modeling (e.g., ani

researcharxiv-cs-cv
23 Jun 2026
Research

FetSelect: Task-Specific Architectures and Self-Supervised Learning for Automated Fetal Ultrasound Frame Selection

DGX agent

arXiv:2606.22487v1 Announce Type: new Abstract: Automated frame selection for fetal biometry remains under addressed, with most prior work targeting generic quality assessment or downstream measuremen

researcharxiv-cs-cv
23 Jun 2026
Research

Few-Shot Hyperspectral Aphid Detection via FastGAN Synthetic Data Generation, Transformer-Based Classification and Explainable AI

DGX agent

arXiv:2606.21267v1 Announce Type: new Abstract: Early detection of aphid infestation in crops is essential for preventing yield loss and reducing unnecessary pesticide use. Hyperspectral imaging combi

researcharxiv-cs-cv
23 Jun 2026
Research

Fidelity- and Perception-Aware Local Implicit Attention for Arbitrary-Scale Image Super-Resolution

DGX agent

arXiv:2606.21910v1 Announce Type: new Abstract: Arbitrary-scale image super-resolution (ASISR) aims to reconstruct high-resolution images from low-resolution inputs over a continuous range of upscalin

researcharxiv-cs-cv
23 Jun 2026
Research

Fine-grained Human Motion Understanding with Language Models

DGX agent

arXiv:2606.20888v1 Announce Type: new Abstract: In this work, we propose methodname, an LLM-based model for fine-grained human motion understanding that represents motion as a sequence of skeletal pos

researcharxiv-cs-cv
23 Jun 2026
Local Ai

FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction

DGX agent

arXiv:2606.21373v1 Announce Type: new Abstract: Recent indoor occupancy prediction methods adopt Gaussian primitives as a sparse 3D representation for computational efficiency. However, their training

local-aiarxiv-cs-cv
23 Jun 2026
Local Ai

Flow6D: Discrete-to-Continuous Flow Matching for Efficient and Accurate Category-Level 6D Pose Estimation

DGX agent

arXiv:2606.23293v1 Announce Type: new Abstract: 6D pose estimation is a key task in computer vision and embodied AI, widely used in robotic manipulation, augmented reality, etc. Existing methods direc

local-aiarxiv-cs-cv
23 Jun 2026
Model Releases

FlowCodec: One-Step Flow Prior for Generative Image Compression

DGX agent

arXiv:2606.21030v1 Announce Type: cross Abstract: Diffusion-based image compression methods, leveraging powerful generative priors, have demonstrated remarkable perceptual quality at ultra-low bitrate

model-releasesarxiv-cs-cv
23 Jun 2026
Applications

FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation

DGX agent

arXiv:2606.22424v1 Announce Type: new Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions in unseen scenes. While Large

applicationsarxiv-cs-cv
23 Jun 2026
Safety

FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation

DGX agent

arXiv:2606.20867v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-sho

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

Following the Flow: Advection-Consistent Modeling for Event-based Small Object Detection

DGX agent

arXiv:2606.22378v1 Announce Type: new Abstract: Event cameras enable high-frequency visual perception with microsecond latency, offering advantages for dynamic scenes. However, event-based small objec

model-releasesarxiv-cs-cv
23 Jun 2026
Research

ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos

DGX agent

arXiv:2602.06226v2 Announce Type: replace Abstract: The ubiquity of monocular videos capturing daily hand-object interactions presents a valuable resource for embodied intelligence. While 3D hand reco

researcharxiv-cs-cv
23 Jun 2026
Model Releases

From Convolution to Transformer: A Comparative Study of U-Net Variants for Brain Tumor and Retinal Vessel Segmentation

DGX agent

arXiv:2606.22168v1 Announce Type: new Abstract: Medical image segmentation plays an important role in computer aided diagnosis, treatment planning, and disease monitoring. U-Net has been widely used f

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

From Driving Videos to Simulatable Scenarios

DGX agent

arXiv:2606.21993v1 Announce Type: cross Abstract: Autonomous vehicles (AVs) face driving scenarios ranging from routine traffic to rare events. To assess safety it is crucial to reproduce these scenar

safetyarxiv-cs-cv
23 Jun 2026
Safety

From Gradient Clipping to Structural Refinement: Improving DPSGD for Medical Image Segmentation

DGX agent

arXiv:2606.21763v1 Announce Type: new Abstract: Medical image segmentation is widely used for disease detection but relies on sensitive data, raising privacy concerns as trained models can leak inform

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

From Point Estimates to Distributions: GMM Pooling for MIL in Preterm Birth Prediction

DGX agent

arXiv:2606.23005v1 Announce Type: new Abstract: Preterm birth (PTB) prediction can enable targeted surveillance and timely intervention, yet most ultrasound-based models use a single selected transvag

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation

DGX agent

arXiv:2606.23486v1 Announce Type: new Abstract: Curvilinear object segmentation, including vessels and cracks, is challenging due to extreme spatial sparsity and topological fragility, where small loc

safetyarxiv-cs-cv
23 Jun 2026
Tutorials

From Uncertainty to Stability and Fidelity: Guiding Sparse-View 3D Gaussian Splatting with Fisher Information

DGX agent

arXiv:2606.20842v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as a promising technique for novel view synthesis. However, 3DGS requires dense input views to achieve high-qua

tutorialsarxiv-cs-cv
23 Jun 2026
Tutorials

Full-Body Golf Swing Kinematic Reconstruction From a Smartwatch IMU

DGX agent

arXiv:2606.22876v1 Announce Type: new Abstract: Quantitative measurement of the golf swing is critical for evaluating technique and enabling individualized feedback. However, existing methods are impr

tutorialsarxiv-cs-cv
23 Jun 2026
Model Releases

Fursee: Hybrid YOLO-DINOv3 Framework for Fursuit Identity Retrieval and Clustering

DGX agent

arXiv:2606.22872v1 Announce Type: new Abstract: Global furry conventions produce massive fursuit photographs, while manual sorting brings heavy labor costs and calls for automatic identity retrieval a

model-releasesarxiv-cs-cv
23 Jun 2026
Research

G-MASt3R-SfM: Graph-based View Pruning and Multi-stage Optimization for Robust SfM

DGX agent

arXiv:2606.22856v1 Announce Type: new Abstract: Structure from Motion (SfM) is essential for multi-view 3D reconstruction, however, its accuracy heavily relies on the accuracy of image matching. While

researcharxiv-cs-cv
23 Jun 2026
Tutorials

Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition

DGX agent

arXiv:2606.22416v1 Announce Type: new Abstract: We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video genera

tutorialsarxiv-cs-cv
23 Jun 2026
Model Releases

Generating Fearful Images: Investigating Potential Emotional Biases in Image-Generation Models

DGX agent

arXiv:2411.05985v3 Announce Type: replace-cross Abstract: This paper examines potential biases and inconsistencies in the emotions evoked by images produced by generative artificial intelligence (AI)

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Generative Relightable Avatars

DGX agent

arXiv:2606.22718v1 Announce Type: new Abstract: We present Generative Relightable Avatars (GRA), a person-specific method for photorealistic free-view rendering and environment-map relighting of full-

researcharxiv-cs-cv
23 Jun 2026
Agents

GenTrack: A New Generation of Multi-Object Tracking

DGX agent

arXiv:2510.24399v3 Announce Type: replace Abstract: This paper introduces a novel multi-object tracking (MOT) method, dubbed GenTrack, whose main contributions include: first-a hybrid tracking approac

agentsarxiv-cs-cv
23 Jun 2026
Model Releases

GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation

DGX agent

arXiv:2606.23669v1 Announce Type: new Abstract: Text-to-image models can generate visually plausible city streets, but whether their outputs correspond to a requested road segment rather than a generi

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

GEOPHYS: The Geometry of Physical Plausibility

DGX agent

arXiv:2606.20707v1 Announce Type: new Abstract: While humans can identify physically implausible events within milliseconds, machine learning approaches addressing the same problem are extremely slow

model-releasesarxiv-cs-cv
23 Jun 2026
Research

GIM-ENDO: A Multimodal Endoscopic Image and Video Dataset for Gastric Intestinal Metaplasia Morphology and Pathology

DGX agent

arXiv:2606.20919v1 Announce Type: new Abstract: Gastric intestinal metaplasia (GIM) is a precursor lesion to gastric dysplasia and adenocarcinoma whose early detection is crucial for intervening in th

researcharxiv-cs-cv
23 Jun 2026
Research

Go-with-the-Track: Video Compositing and Motion Control with Point Tracking

DGX agent

arXiv:2606.20891v1 Announce Type: new Abstract: Filmmaking demands precise motion control and reference image compositing -- capabilities that existing methods treat separately. Point-track-conditione

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Gold Points Sniper: Self-guided Visual Reasoning in VLM for Fine-grained Action Understanding

DGX agent

arXiv:2606.22409v1 Announce Type: new Abstract: Robots operating in everyday environments must understand fine-grained human actions, intentions, and contextual cues from broad views where people occu

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication

DGX agent

arXiv:2509.09597v3 Announce Type: replace-cross Abstract: Graph alignment, the problem of identifying corresponding nodes across multiple graphs, is fundamental to numerous applications. Most existing

safetyarxiv-cs-cv
23 Jun 2026
Safety

Graph-of-Differences: Anatomy-Structured Difference Alignment for Medical Image Re-Identification

DGX agent

arXiv:2606.21368v1 Announce Type: new Abstract: Medical image re-identification (MedReID) enables longitudinal patient linkage but remains vulnerable to shortcut learning and often produces decisions

safetyarxiv-cs-cv
23 Jun 2026
Local Ai

GreenRFM: Learning a resource-efficient radiology vision-language foundation model via supervision-centric pre-training

DGX agent

arXiv:2603.06467v2 Announce Type: replace Abstract: Radiology foundation models (RFMs) have largely inherited the scale-first recipe of natural-image vision--language pre-training. This recipe is diff

local-aiarxiv-cs-cv
23 Jun 2026
Model Releases

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

DGX agent

arXiv:2606.20799v1 Announce Type: new Abstract: Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate across shots, caus

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

GTA-Net: Cooperative Game Theory for Vision-Language Alignment in Chest X-Ray Report Generation

DGX agent

arXiv:2606.21915v1 Announce Type: new Abstract: Automated chest X-ray report generation requires precise cross-modal grounding to ensure clinically reliable descriptions. However, existing vision-lang

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

HaineiFRDM: Structure-Preserving Diffusion for Film Restoration under Fast Motion and Diverse Defects

DGX agent

arXiv:2512.24946v2 Announce Type: replace Abstract: Existing film-restoration methods frequently fail under fast motion, producing limb disappearance and structural distortion due to inaccurate motion

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

Happy Young Women, Grumpy Old Men? Emotion-Driven Demographic Biases in Synthetic Face Generation

DGX agent

arXiv:2602.00032v3 Announce Type: replace-cross Abstract: Synthetic faces from text-to-image (T2I) models pervade digital media, yet their demographic biases under emotionally conditioned prompts rema

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

Hedgementation = Hedgerow Segmentation: A Remote Sensing Benchmark

DGX agent

arXiv:2606.23615v1 Announce Type: new Abstract: We propose Hedgementation: a new benchmark to evaluate machine learning models for hedgerow mapping from remote sensing data at country scale and 10m^2

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

HEM: a margin-based loss for visual categorisation tasks

DGX agent

arXiv:2501.12191v2 Announce Type: replace-cross Abstract: Training deep neural networks (DNNs) on classification tasks can be performed with a number of different losses, but cross-entropy (CE) loss i

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

HERCULES: An Open-Source Simulation Framework for Heterogeneous Multi-Robot SLAM, Collaborative Perception, and Exploration

DGX agent

arXiv:2606.22756v1 Announce Type: cross Abstract: We present HERCULES, an open-source simulator and data-collection pipeline for heterogeneous multi-robot autonomy. Built upon the Unreal Engine 5 (UE5

model-releasesarxiv-cs-cv
23 Jun 2026
Research

HERMAN: Hierarchical Representation Matching for CLIP-based Class-Incremental Learning

DGX agent

arXiv:2509.22645v2 Announce Type: replace Abstract: Class-Incremental Learning (CIL) aims to endow models with the ability to continuously adapt to evolving data streams. Recent advances in pre-traine

researcharxiv-cs-cv
23 Jun 2026
Research

HERO: Hypothesis-Driven Evidence Retrieval from Omics for Multi-Task Breast Cancer Analysis

DGX agent

arXiv:2606.21174v1 Announce Type: new Abstract: Matched multi-omics can improve WSI-based biomarker and prognosis prediction, but most existing pipelines use omics as a paral lel feature stream or tex

researcharxiv-cs-cv
23 Jun 2026
← Previous
1…979899100101…263
Next →