AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

On The Application of Linear Attention in Multimodal Transformers

DGX agent

arXiv:2604.10064v1 Announce Type: new Abstract: Multimodal Transformers serve as the backbone for state-of-the-art vision-language models, yet their quadratic attention complexity remains a critical b

researcharxiv-cs-cv
14 Apr 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

On the Effectiveness of Textual Prompting with Lightweight Fine-Tuning for SAM3 Remote Sensing Segmentation

DGX agent

arXiv:2512.15564v2 Announce Type: replace Abstract: Remote sensing (RS) image segmentation is constrained by the limited availability of annotated data and a gap between overhead imagery and natural i

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

Online Reasoning Video Object Segmentation

DGX agent

arXiv:2604.11411v1 Announce Type: new Abstract: Reasoning video object segmentation predicts pixel-level masks in videos from natural-language queries that may involve implicit and temporally grounded

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

Optimization-Guided Diffusion for Interactive Scene Generation

DGX agent

arXiv:2512.07661v3 Announce Type: replace Abstract: Realistic and diverse multi-agent driving scenes are crucial for evaluating autonomous vehicles, but safety-critical events which are essential for

safetyarxiv-cs-cv
14 Apr 2026
Hardware

PA-SFM: Tracker-free differentiable acoustic radiation for freehand 3D photoacoustic imaging

DGX agent

arXiv:2604.09643v1 Announce Type: new Abstract: Three-dimensional (3D) handheld photoacoustic tomography typically relies on bulky and expensive external positioning sensors to correct motion artifact

hardwarearxiv-cs-cv
14 Apr 2026
Safety

PACO: Proxy-Task Alignment and Online Calibration for On-the-Fly Category Discovery

DGX agent

arXiv:2604.11484v1 Announce Type: new Abstract: On-the-Fly Category Discovery (OCD) requires a model, trained on an offline support set, to recognize known classes while discovering new ones from an o

safetyarxiv-cs-cv
14 Apr 2026
Local Ai

Pair2Scene: Learning Local Object Relations for Procedural Scene Generation

DGX agent

arXiv:2604.11808v1 Announce Type: new Abstract: Generating high-fidelity 3D indoor scenes remains a significant challenge due to data scarcity and the complexity of modeling intricate spatial relation

local-aiarxiv-cs-cv
14 Apr 2026
Research

PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion

DGX agent

arXiv:2601.07447v2 Announce Type: replace Abstract: Existing image foundation models are not optimized for spherical images having been trained primarily on perspective images. PanoSAMic integrates th

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Parameter Efficient Fine-tuning for Domain-specific Gastrointestinal Disease Recognition

DGX agent

arXiv:2604.10451v1 Announce Type: new Abstract: Despite recent advancements in the field of medical image analysis with the use of pretrained foundation models, the issue of distribution shifts betwee

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

Particle Diffusion Matching: Random Walk Correspondence Search for the Alignment of Standard and Ultra-Widefield Fundus Images

DGX agent

arXiv:2604.10085v1 Announce Type: new Abstract: We propose a robust alignment technique for Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which are challenging to align due

safetyarxiv-cs-cv
14 Apr 2026
Tutorials

PASTA: Vision Transformer Patch Aggregation for Weakly Supervised Target and Anomaly Segmentation

DGX agent

arXiv:2604.09701v1 Announce Type: new Abstract: Detecting unseen anomalies in unstructured environments presents a critical challenge for industrial and agricultural applications such as material recy

tutorialsarxiv-cs-cv
14 Apr 2026
Research

PERCEPT-Net: A Perceptual Loss Driven Framework for Reducing MRI Artifact Tissue Confusion

DGX agent

arXiv:2604.10439v1 Announce Type: new Abstract: Purpose: Existing deep learning-based MRI artifact correction models exhibit poor clinical generalization due to inherent artifact-tissue confusion, fai

researcharxiv-cs-cv
14 Apr 2026
Safety

Perceptual Inductive Bias Is What You Need Before Contrastive Learning

DGX agent

arXiv:2506.01201v2 Announce Type: replace Abstract: David Marr's seminal theory of human perception stipulates that visual processing is a multi-stage process, prioritizing the derivation of boundary

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit--Explicit Optimization

DGX agent

arXiv:2604.10125v1 Announce Type: new Abstract: Existing single-image 3D indoor scene generators often produce results that look visually plausible but fail to obey real-world physics, limiting their

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Physics-Informed Synthetic Dataset and Denoising TIE-Reconstructed Phase Maps in Transient Flows Using Deep Learning

DGX agent

arXiv:2604.10610v1 Announce Type: cross Abstract: High-speed quantitative phase imaging enables non-intrusive visualization of transient compressible gas flows and energetic phenomena. However, phase

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers

DGX agent

arXiv:2604.10415v1 Announce Type: new Abstract: We present Point2Pose, a model-free method for causal 6D pose tracking of multiple rigid objects from monocular RGB-D video. Initialized only from spars

model-releasesarxiv-cs-cv
14 Apr 2026
Applications

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs

DGX agent

arXiv:2604.11627v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the ra

applicationsarxiv-cs-cv
14 Apr 2026
Research

PointSplat: Efficient Geometry-Driven Pruning and Transformer Refinement for 3D Gaussian Splatting

DGX agent

arXiv:2604.09903v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently unlocked real-time, high-fidelity novel view synthesis by representing scenes using explicit 3D primitives. Ho

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Precision Synthesis of Multi-Tracer PET via VLM-Modulated Rectified Flow for Stratifying Mild Cognitive Impairment

DGX agent

arXiv:2604.11176v1 Announce Type: new Abstract: The biological definition of Alzheimer's disease (AD) relies on multi-modal neuroimaging, yet the clinical utility of positron emission tomography (PET)

model-releasesarxiv-cs-cv
14 Apr 2026
Research

PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback

DGX agent

arXiv:2506.21834v2 Announce Type: replace Abstract: Inpainting, the process of filling missing or corrupted image parts, has broad applications in medical imaging. However, generating anatomically acc

researcharxiv-cs-cv
14 Apr 2026
Safety

Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation

DGX agent

arXiv:2603.20725v2 Announce Type: replace Abstract: Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on

safetyarxiv-cs-cv
14 Apr 2026
Tutorials

Preventing Latent Rehearsal Decay in Online Continual SSL with SOLAR

DGX agent

arXiv:2604.10586v1 Announce Type: cross Abstract: This paper explores Online Continual Self-Supervised Learning (OCSSL), a scenario in which models learn from continuous streams of unlabeled, non-stat

tutorialsarxiv-cs-cv
14 Apr 2026
Research

Prints in the Magnetic Dust: Robust Similarity Search in Legacy Media Images Using Checksum Count Vectors

DGX agent

arXiv:2604.09657v1 Announce Type: new Abstract: Digitizing magnetic media containing computer data is only the first step towards the preservation of early home computing era artifacts. The audio tape

researcharxiv-cs-cv
14 Apr 2026
Applications

ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction

DGX agent

arXiv:2604.02003v2 Announce Type: replace Abstract: Generating ground-level views and coherent 3D site models from aerial-only imagery is challenging due to extreme viewpoint changes, missing intermed

applicationsarxiv-cs-cv
14 Apr 2026
Research

Progressive Deep Learning for Automated Spheno-Occipital Synchondrosis Maturation Assessment

DGX agent

arXiv:2604.10945v1 Announce Type: new Abstract: Accurate assessment of spheno-occipital synchondrosis (SOS) maturation is a key indicator of craniofacial growth and a critical determinant for orthodon

researcharxiv-cs-cv
14 Apr 2026
Research

Progressively Texture-Aware Diffusion for Contrast-Enhanced Sparse-View CT

DGX agent

arXiv:2604.11559v1 Announce Type: new Abstract: Diffusion-based sparse-view CT (SVCT) imaging has achieved remarkable advancements in recent years, thanks to its more stable generative capability. How

researcharxiv-cs-cv
14 Apr 2026
Safety

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation

DGX agent

arXiv:2604.10030v1 Announce Type: new Abstract: Video diffusion models have achieved remarkable progress in generating high-quality videos. However, these models struggle to represent the temporal suc

safetyarxiv-cs-cv
14 Apr 2026
Local Ai

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

DGX agent

arXiv:2512.05564v2 Announce Type: replace Abstract: Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to pro

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

PSF-Med: Measuring and Explaining Paraphrase Sensitivity in Medical Vision Language Models

DGX agent

arXiv:2602.21428v2 Announce Type: replace Abstract: Medical Vision Language Models (VLMs) can change their answers when clinicians rephrase the same question, a failure mode that threatens deployment

model-releasesarxiv-cs-cv
14 Apr 2026
Applications

Quantization Robustness to Input Degradations for Object Detection

DGX agent

arXiv:2508.19600v2 Announce Type: replace Abstract: Post-training quantization (PTQ) is crucial for deploying efficient object detection models, like YOLO, on resource-constrained devices. However, th

applicationsarxiv-cs-cv
14 Apr 2026
Tutorials

Quantum-Gated Task-interaction Knowledge Distillation for Pre-trained Model-based Class-Incremental Learning

DGX agent

arXiv:2604.11112v1 Announce Type: cross Abstract: Class-incremental learning (CIL) aims to continuously accumulate knowledge from a stream of tasks and construct a unified classifier over all seen cla

tutorialsarxiv-cs-cv
14 Apr 2026
Research

R2E-VID: Two-Stage Robust Routing via Temporal Gating for Elastic Edge-Cloud Video Inference

DGX agent

arXiv:2604.09681v1 Announce Type: cross Abstract: With the rapid growth of large-scale video analytics applications, edge-cloud collaborative systems have become the dominant paradigm for real-time in

researcharxiv-cs-cv
14 Apr 2026
Research

RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation

DGX agent

arXiv:2604.11164v1 Announce Type: new Abstract: Deep learning has greatly advanced medical image segmentation, but its success relies heavily on fully supervised learning, which requires dense annotat

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Radiology Report Generation for Low-Quality X-Ray Images

DGX agent

arXiv:2604.10188v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have significantly advanced automated Radiology Report Generation (RRG). However, existing methods implicitly assume high-

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting

DGX agent

arXiv:2604.10259v1 Announce Type: new Abstract: We present a generalizable feed-forward Gaussian splatting framework for human 3D reconstruction and real-time animation that operates directly on multi

researcharxiv-cs-cv
14 Apr 2026
Research

ReaLiTy and LADS: A Unified Framework and Dataset Suite for LiDAR Adaptation Across Sensors and Adverse Weather Conditions

DGX agent

arXiv:2604.10213v1 Announce Type: cross Abstract: Reliable LiDAR perception requires robustness across sensors, environments, and adverse weather. However, existing datasets rarely provide physically

researcharxiv-cs-cv
14 Apr 2026
Safety

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought

DGX agent

arXiv:2506.16796v4 Announce Type: replace Abstract: Real-World Image Super-Resolution is one of the most challenging task in image restoration. However, existing methods struggle with an accurate unde

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

ReContraster: Making Your Posters Stand Out with Regional Contrast

DGX agent

arXiv:2604.10442v1 Announce Type: new Abstract: Effective poster design requires rapidly capturing attention and clearly conveying messages. Inspired by the ``contrast effects'' principle, we propose

model-releasesarxiv-cs-cv
14 Apr 2026
Local Ai

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models

DGX agent

arXiv:2604.10578v1 Announce Type: new Abstract: The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. Howev

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment

DGX agent

arXiv:2604.10789v1 Announce Type: new Abstract: Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D sc

model-releasesarxiv-cs-cv
14 Apr 2026
Agents

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction

DGX agent

arXiv:2604.11707v1 Announce Type: new Abstract: Accurate future video prediction requires both high visual fidelity and consistent scene semantics, particularly in complex dynamic environments such as

agentsarxiv-cs-cv
14 Apr 2026
Applications

RESP: Reference-guided Sequential Prompting for Visual Glitch Detection in Video Games

DGX agent

arXiv:2604.11082v1 Announce Type: new Abstract: Visual glitches in video games degrade player experience and perceived quality, yet manual quality assurance cannot scale to the growing test surface of

applicationsarxiv-cs-cv
14 Apr 2026
Applications

Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification

DGX agent

arXiv:2604.10695v1 Announce Type: new Abstract: Recent Audio-Visual Question Answering (AVQA) methods have advanced significantly. However, most AVQA methods lack effective mechanisms for handling mis

applicationsarxiv-cs-cv
14 Apr 2026
Model Releases

Revisiting the Scale Loss Function and Gaussian-Shape Convolution for Infrared Small Target Detection

DGX agent

arXiv:2604.09991v1 Announce Type: new Abstract: Infrared small target detection still faces two persistent challenges: training instability from non-monotonic scale loss functions, and inadequate spat

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

RL makes MLLMs see better than SFT

DGX agent

arXiv:2510.16333v2 Announce Type: replace Abstract: A dominant assumption in Multimodal Language Model (MLLM) research is that its performance is largely inherited from the LLM backbone, given its imm

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization

DGX agent

arXiv:2603.12639v2 Announce Type: replace Abstract: Scalable Embodied AI faces fundamental constraints due to prohibitive costs and safety risks of real-world interaction. While Embodied World Models

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

Robust Fair Disease Diagnosis in CT Images

DGX agent

arXiv:2604.09710v1 Announce Type: new Abstract: Automated diagnosis from chest CT has improved considerably with deep learning, but models trained on skewed datasets tend to perform unevenly across pa

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

RobustMedSAM: Degradation-Resilient Medical Image Segmentation via Robust Foundation Model Adaptation

DGX agent

arXiv:2604.09814v1 Announce Type: new Abstract: Medical image segmentation models built on Segment Anything Model (SAM) achieve strong performance on clean benchmarks, yet their reliability often degr

model-releasesarxiv-cs-cv
14 Apr 2026
← Previous
1…247248249250251…259
Next →