AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression

DGX agent

arXiv:2607.12756v1 Announce Type: new Abstract: Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated

model-releasesarxiv-cs-cv
15 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors

DGX agent

arXiv:2512.15748v2 Announce Type: replace-cross Abstract: Visual Species Recognition (VSR) is a fundamental task in scientific disciplines that require species-level identification, including ecology,

researcharxiv-cs-cv
15 Jul 2026
Local Ai

WanToFight: Real-Time Generative Game Engine for Multi-Player Combat Interaction

DGX agent

arXiv:2607.12592v1 Announce Type: new Abstract: We present WanToFight, a generative game engine that simulates real-time, two-player The King of Fighters '97 (KOF~'97) gameplay from keyboard input. Pr

local-aiarxiv-cs-cv
15 Jul 2026
Model Releases

What Does a Temporal Benchmark Score Measure? Decomposing Channel Use in Video VLM Evaluation

DGX agent

arXiv:2607.12304v1 Announce Type: new Abstract: A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1.

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

DGX agent

arXiv:2607.12993v1 Announce Type: new Abstract: We present X-lens, a compact feed-forward model for metric depth estimation from a variable number of calibrated fisheye and pinhole views. To support r

model-releasesarxiv-cs-cv
15 Jul 2026
Applications

3D Reconstruction of deciduous Trees using low-cost UAV- and Crane-based Photogrammetry for Monitoring Shoot Elongation across entire Canopies

DGX agent

arXiv:2607.07905v1 Announce Type: new Abstract: Tree growth determines how much CO2 is sequestered from the atmosphere and temporarily stored in woody biomass. At the same time tree growth is affected

applicationsarxiv-cs-cv
10 Jul 2026
Local Ai

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding

DGX agent

arXiv:2512.21414v2 Announce Type: replace Abstract: Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tool

local-aiarxiv-cs-cv
10 Jul 2026
Research

Anatomically Guided Latent Diffusion for Brain MRI Progression Modeling

DGX agent

arXiv:2601.14584v2 Announce Type: replace Abstract: Accurately modeling longitudinal brain MRI progression is crucial for understanding neurodegenerative diseases and predicting individualized structu

researcharxiv-cs-cv
10 Jul 2026
Model Releases

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

DGX agent

arXiv:2607.08741v1 Announce Type: cross Abstract: Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While rece

model-releasesarxiv-cs-cv
10 Jul 2026
Research

Are Current Continual Learning Methods Truly Agnostic? Introducing OPRE, a Step Toward Agnostic Continual Learning

DGX agent

arXiv:2511.08226v2 Announce Type: replace-cross Abstract: In order to achieve Continual Learning (CL), the problem of catastrophic forgetting, one that has plagued neural networks since their inceptio

researcharxiv-cs-cv
10 Jul 2026
Hardware

ARGUS: Accelerated, Robust, General, and Unsupervised Cell Tracking Solutions

DGX agent

arXiv:2607.08297v1 Announce Type: new Abstract: Background and Objective: Quantitative analysis of cell dynamics is central to modern biological research, providing critical insights into immune cell

hardwarearxiv-cs-cv
10 Jul 2026
Local Ai

Asynchronous Federated Continual Segmentation with Evolving Clients and Label Spaces

DGX agent

arXiv:2503.15414v3 Announce Type: replace-cross Abstract: Federated learning seeks to foster collaboration among distributed clients while preserving the privacy of their local data. Traditional feder

local-aiarxiv-cs-cv
10 Jul 2026
Research

Attention-Based Segmentation of WMHs and Differentiation of Vascular vs. Demyelinating Lesions

DGX agent

arXiv:2607.08171v1 Announce Type: new Abstract: White Matter Hyperintensities (WMHs) are commonly observed in brain Magnetic Resonance Imaging (MRI) scans. They are associated with various neurologica

researcharxiv-cs-cv
10 Jul 2026
Model Releases

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation

DGX agent

arXiv:2607.08397v1 Announce Type: new Abstract: Referring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understand

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Benchmark Evaluation of Feredated Learning on Multi-organ Images

DGX agent

arXiv:2607.08219v1 Announce Type: new Abstract: The privacy requirements of medical data and its substantial variations across organs and modalities hinder the clinical implementation of medical AI. F

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Beyond wheelchairs and blindfolds: Investigating disability stereotypes in T2I models with INCLUDE-BENCH

DGX agent

arXiv:2607.08515v1 Announce Type: new Abstract: Text-to-image (T2I) models have been shown to exhibit social biases. Prior work has mainly focused on gender, skin tone, and cultural representation wit

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

BiasBench: A reproducible benchmark for tuning the biases of event cameras

DGX agent

arXiv:2504.18235v2 Announce Type: replace Abstract: Event-based cameras are bio-inspired sensors that detect light changes asynchronously for each pixel. They are increasingly used in fields like comp

model-releasesarxiv-cs-cv
10 Jul 2026
Safety

Borrowing from anything: A generalizable framework for reference-guided instance editing

DGX agent

arXiv:2512.15138v2 Announce Type: replace Abstract: Reference-guided instance editing is fundamentally limited by semantic entanglement, where a reference's intrinsic appearance is intertwined with it

safetyarxiv-cs-cv
10 Jul 2026
Model Releases

Classical versus Deep Mirror-Symmetry Scoring: A Benchmark of Thirteen Methods

DGX agent

arXiv:2607.08379v1 Announce Type: new Abstract: Quantifying how mirror-symmetric an image is about a given axis (symmetry scoring) underpins applications from visual aesthetics to medical imaging, yet

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Closing the Null Space: Guidance-Aware Quantization for Classifier-Free Diffusion

DGX agent

arXiv:2607.08241v1 Announce Type: new Abstract: Deploying classifier-free guidance (CFG) diffusion models under real-world compute budgets requires quantization, yet existing post-training quantizatio

model-releasesarxiv-cs-cv
10 Jul 2026
Agents

Computation, Condensation, and the Incompleteness Between Them: A Coupled Foundation of Intelligence

DGX agent

arXiv:2303.04203v4 Announce Type: replace-cross Abstract: The theory of computation was built to answer Turing's question: what is effectively calculable by an unbounded, immortal, disembodied agent f

agentsarxiv-cs-cv
10 Jul 2026
Research

ConRad: Efficient Conformal Prediction for Radiomics

DGX agent

arXiv:2607.08084v1 Announce Type: cross Abstract: Radiomic features derived from medical images and segmentation masks are used to support decision making in clinical imaging pipelines. In practice, t

researcharxiv-cs-cv
10 Jul 2026
Model Releases

Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions

DGX agent

arXiv:2607.08164v1 Announce Type: new Abstract: Deep neural nets achieve remarkable performance when training and test data share the same distribution, but this assumption frequently breaks in real-w

model-releasesarxiv-cs-cv
10 Jul 2026
Research

CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction

DGX agent

arXiv:2607.08503v1 Announce Type: new Abstract: Accurate prognosis prediction is important for treatment planning in lung cancer, but deep learning-driven survival modelling is often limited by the sc

researcharxiv-cs-cv
10 Jul 2026
Research

Data Alchemy: Mitigating Cross-Site Model Variability Through Test Time Data Calibration

DGX agent

arXiv:2407.13632v2 Announce Type: replace Abstract: Deploying deep learning-based imaging tools across various clinical sites poses significant challenges due to inherent domain shifts and regulatory

researcharxiv-cs-cv
10 Jul 2026
Safety

DeltaDeno: Zero-Shot Anomaly Generation via Delta-Denoising Attribution

DGX agent

arXiv:2511.16920v2 Announce Type: replace Abstract: Anomaly generation is often framed as few-shot fine-tuning with anomalous samples, which contradicts the scarcity that motivates generation and tend

safetyarxiv-cs-cv
10 Jul 2026
Model Releases

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models

DGX agent

arXiv:2607.08434v1 Announce Type: new Abstract: Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but t

model-releasesarxiv-cs-cv
10 Jul 2026
Safety

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models

DGX agent

arXiv:2511.19032v2 Announce Type: replace Abstract: Visual corruptions can change vision--language model (VLM) behavior in ways that top-1 accuracy does not capture. A model may keep the same answer w

safetyarxiv-cs-cv
10 Jul 2026
Model Releases

Dive Into the Implicit Biases of Low-rank Vision-language Alignment

DGX agent

arXiv:2607.08194v1 Announce Type: new Abstract: Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requi

model-releasesarxiv-cs-cv
10 Jul 2026
Research

Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?

DGX agent

arXiv:2607.08514v1 Announce Type: new Abstract: Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models

researcharxiv-cs-cv
10 Jul 2026
Model Releases

Do Transformations Reveal the Truth? Generative Residual Learning for Generalized AI-Generated Image Detection

DGX agent

arXiv:2607.08674v1 Announce Type: new Abstract: The rapid advancement of generative AI has enabled the creation of highly realistic deepfake media, posing significant threats, including misinformation

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark

DGX agent

arXiv:2607.08191v1 Announce Type: new Abstract: RGB-Thermal (RGBT) Video Object Detection (VOD) has gained significant traction due to its ability to overcome the limitations of conventional RGB-based

model-releasesarxiv-cs-cv
10 Jul 2026
Research

Effective Gaussian Management for High-fidelity Scene Reconstruction

DGX agent

arXiv:2509.12742v4 Announce Type: replace Abstract: This paper proposes an effective Gaussian management framework for high-fidelity scene reconstruction of both appearance and geometry. Unlike recent

researcharxiv-cs-cv
10 Jul 2026
Applications

Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding

DGX agent

arXiv:2512.14236v2 Announce Type: replace Abstract: The growing demand for immersive 3D content calls for automated monocular-to-stereo video conversion. We present Elastic3D, a controllable, direct e

applicationsarxiv-cs-cv
10 Jul 2026
Tutorials

Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

DGX agent

arXiv:2607.08765v1 Announce Type: new Abstract: In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream t

tutorialsarxiv-cs-cv
10 Jul 2026
Research

Enhancing the KidSat Model: Integrating Geographical Encoding and Data Quality Assessment for Childhood Poverty Prediction

DGX agent

arXiv:2607.08281v1 Announce Type: new Abstract: Accurate poverty mapping using satellite imagery is often hindered by (i) noisy and sparse survey-derived supervision, (ii) image quality issues such as

researcharxiv-cs-cv
10 Jul 2026
Model Releases

Equivariant Quantum Clustering with Differential Privacy: Parameter-Efficient Privacy-Preserving Analysis Across Heterogeneous Sensitive Datasets

DGX agent

arXiv:2607.08092v1 Announce Type: cross Abstract: Privacy-preserving clustering is critical for analyzing sensitive data in healthcare, cybersecurity, and enterprise applications, where maintaining da

model-releasesarxiv-cs-cv
10 Jul 2026
Hardware

EVIS: A Physics-Grounded Event Camera Plugin for NVIDIA Isaac Sim

DGX agent

arXiv:2607.08098v1 Announce Type: new Abstract: Event cameras offer microsecond temporal resolution, low latency, and high dynamic range, making them attractive for robotics. However, labeled event-ca

hardwarearxiv-cs-cv
10 Jul 2026
Model Releases

False Confidence: Automated Labels Confound Fairness Audits in Cervical Spine Segmentation

DGX agent

arXiv:2607.07852v1 Announce Type: cross Abstract: Automated segmentation of cervical-spine MRI is increasingly used in clinical workflows, yet no fairness audit exists for this anatomy. We show that a

model-releasesarxiv-cs-cv
10 Jul 2026
Local Ai

FedTR: Federated Learning Framework with Transfer Learning for Industrial Visual Inspection

DGX agent

arXiv:2607.08014v1 Announce Type: new Abstract: Federated learning (FL) is a collaborative learning scheme to train deep learning models, where collaborating parties can consolidate their models witho

local-aiarxiv-cs-cv
10 Jul 2026
Research

FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidance

DGX agent

arXiv:2502.20805v3 Announce Type: replace-cross Abstract: Hand-object interaction(HOI) is the fundamental link between human and environment, yet its dexterous and complex pose significantly challenge

researcharxiv-cs-cv
10 Jul 2026
Model Releases

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

DGX agent

arXiv:2512.01803v3 Announce Type: replace Abstract: Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elu

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Geometry and Gradient-based Partitioning for Panoramic Outdoor Reconstruction

DGX agent

arXiv:2607.08769v1 Announce Type: new Abstract: Scaling 3D Gaussian Splatting (3DGS) to large outdoor scenes is costly in both data acquisition and computation. Adopting panoramic images with equirect

model-releasesarxiv-cs-cv
10 Jul 2026
Applications

GERD: Geometric event response data generation

DGX agent

arXiv:2412.03259v3 Announce Type: replace Abstract: Event-based vision sensors offer high temporal resolution, high dynamic range, and low power consumption, yet event-based vision models lag behind c

applicationsarxiv-cs-cv
10 Jul 2026
Research

GIRAF: Towards Generalizable Human Interactions with Articulated Objects

DGX agent

arXiv:2607.07880v1 Announce Type: new Abstract: Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications

researcharxiv-cs-cv
10 Jul 2026
Research

GRE-Diff: Gaussian Room Embeddings for Structured Layout Diffusion

DGX agent

arXiv:2607.08086v1 Announce Type: new Abstract: Designing functional and aesthetically coherent floor plans requires exploring a vast space of possible room arrangements, a task that quickly becomes o

researcharxiv-cs-cv
10 Jul 2026
Research

GSurf: Learning Signed Distance Fields from Splatting Opaque Gaussians for High-quality 3D Reconstruction

DGX agent

arXiv:2411.15723v4 Announce Type: replace Abstract: High-fidelity surface reconstruction from multi-view images is a core problem in 3D computer vision. While neural implicit surfaces like SDFs offer

researcharxiv-cs-cv
10 Jul 2026
Safety

HairWeaver: Few-Shot Photorealistic Hair Motion Synthesis with Sim-to-Real Guided Video Diffusion

DGX agent

arXiv:2602.11117v2 Announce Type: replace Abstract: We present HairWeaver, a diffusion-based pipeline that animates a single human image with realistic and expressive hair dynamics. While existing met

safetyarxiv-cs-cv
10 Jul 2026
← Previous
1…5354555657…261
Next →