AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation

DGX agent

arXiv:2605.08050v1 Announce Type: new Abstract: Talking-head generation requires joint modeling of identity, head pose, facial expression, and mouth dynamics. Existing methods typically address only a

safetyarxiv-cs-cv
11 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation

DGX agent

arXiv:2411.16748v5 Announce Type: replace Abstract: Long-duration talking video synthesis faces enduring challenges in achieving high video quality, portrait consistency, temporal coherence, and compu

researcharxiv-cs-cv
11 May 2026
Research

Multimodal Latent Reasoning via Hierarchical Visual Cues Injection

DGX agent

arXiv:2602.05359v2 Announce Type: replace Abstract: The advancement of multimodal large language models (MLLMs) has enabled impressive perception capabilities. However, their reasoning process often r

researcharxiv-cs-cv
11 May 2026
Research

Multimodal Stepwise Clinically-Guided Attention Learning for Pathological Complete Response Prediction in Breast Cancer

DGX agent

arXiv:2605.07561v1 Announce Type: new Abstract: Pathological complete response (pCR) is a key prognostic factor in breast cancer patients undergoing neoadjuvant therapy, strongly associated with long-

researcharxiv-cs-cv
11 May 2026
Research

Multispectral Indices for Wildfire Management

DGX agent

arXiv:2309.01751v3 Announce Type: replace-cross Abstract: The increasing frequency and severity of wildfires necessitates advanced methods for effective surveillance and management, as traditional gro

researcharxiv-cs-cv
11 May 2026
Research

Normalizing Trajectory Models

DGX agent

arXiv:2605.08078v1 Announce Type: new Abstract: Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a

researcharxiv-cs-cv
11 May 2026
Research

Not All Tokens Need 40 Steps: Heterogeneous Step Allocation in Diffusion Transformers for Efficient Video Generation

DGX agent

arXiv:2605.06892v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have achieved state-of-the-art video generation quality, but they incur immense computational cost because standard infere

researcharxiv-cs-cv
11 May 2026
Model Releases

NS-Net: Decoupling CLIP Semantic Information through NULL-Space for Generalizable AI-Generated Image Detection

DGX agent

arXiv:2508.01248v4 Announce Type: replace Abstract: The rapid progress of generative models, such as GANs and diffusion models, has facilitated the creation of highly realistic images, raising growing

model-releasesarxiv-cs-cv
11 May 2026
Safety

Object Hallucination-Free Reinforcement Unlearning for Vision-Language Models

DGX agent

arXiv:2605.08031v1 Announce Type: new Abstract: Vision-language models (VLMs) raise growing concerns about privacy, copyright, and bias, motivating machine unlearning to remove sensitive knowledge. Ho

safetyarxiv-cs-cv
11 May 2026
Research

On the Role of Strain and Vorticity in Numerical Integration Error for Flow Matching

DGX agent

arXiv:2605.06680v1 Announce Type: cross Abstract: Flow matching generates data by integrating a learned velocity field, where the number of integration steps (NFE) directly determines inference cost.

researcharxiv-cs-cv
11 May 2026
Agents

One World, Dual Timeline: Decoupled Spatio-Temporal Gaussian Scene Graph for 4D Cooperative Driving Reconstruction

DGX agent

arXiv:2605.07910v1 Announce Type: new Abstract: Reconstructing dynamic scenes from Vehicle-to-Infrastructure Cooperative Autonomous Driving (VICAD) data is fundamentally complicated by temporal asynch

agentsarxiv-cs-cv
11 May 2026
Applications

OneViewAll: Semantic Prior Guided One-View 6D Pose Estimation for Novel Objects

DGX agent

arXiv:2605.07023v1 Announce Type: new Abstract: In many practical 6D object pose estimation scenarios, we often have access to only a single real-world RGB-D reference view per object, typically witho

applicationsarxiv-cs-cv
11 May 2026
Research

OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical Videos

DGX agent

arXiv:2605.07695v1 Announce Type: new Abstract: High-fidelity surgical video generation can greatly improve medical training and the development of AI, adapting these generative models for precise vid

researcharxiv-cs-cv
11 May 2026
Research

PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation

DGX agent

arXiv:2605.07252v1 Announce Type: cross Abstract: Co-speech gesture generation aims to synthesize realistic body movements that are semantically coherent with speech and faithful to a user-specified g

researcharxiv-cs-cv
11 May 2026
Research

PET-Adapter: Test-Time Domain Adaptation for Full and Limited-Angle PET Image Reconstruction

DGX agent

arXiv:2605.08030v1 Announce Type: new Abstract: Positron Emission Tomography (PET) image reconstruction is inherently challenged by Poisson noise and physical degradation factors, which are further ex

researcharxiv-cs-cv
11 May 2026
Research

PicoEyes: Unified Gaze Estimation Framework for Mixed Reality with a Large-Scale Multi-View Dataset

DGX agent

arXiv:2605.07188v1 Announce Type: new Abstract: We present PicoEyes, a unified gaze estimation framework that directly predicts all key attributes of gaze, including 3D eye parameters, eye-region segm

researcharxiv-cs-cv
11 May 2026
Model Releases

PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models

DGX agent

arXiv:2605.07574v1 Announce Type: new Abstract: Mainstream vision-language models (VLMs) fundamentally struggle with severe optical ambiguities, such as reflections and transparent objects, due to the

model-releasesarxiv-cs-cv
11 May 2026
Research

Pre-training Enables Extraordinary All-optical Image Denoising

DGX agent

arXiv:2605.07810v1 Announce Type: cross Abstract: Optical neural networks are emerging as powerful machine learning and information processing tools because of their potential advantages in speed and

researcharxiv-cs-cv
11 May 2026
Research

Pretty Good Measurement for Radiomics: A Quantum-Inspired Multi-Class Classifier for Lung Cancer Subtyping and Prostate Cancer Risk Stratification

DGX agent

arXiv:2603.00223v2 Announce Type: replace Abstract: We investigate a quantum-inspired approach to supervised multi-class classification based on the Pretty Good Measurement (PGM), viewed as an operato

researcharxiv-cs-cv
11 May 2026
Model Releases

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition

DGX agent

arXiv:2605.07154v1 Announce Type: new Abstract: Referring Audio-Visual Segmentation (Ref-AVS) seeks to localize and segment target objects in video frames based on visual, auditory, and textual referr

model-releasesarxiv-cs-cv
11 May 2026
Safety

Probabilistic Object Detection with Conformal Prediction

DGX agent

arXiv:2605.07549v1 Announce Type: new Abstract: Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a su

safetyarxiv-cs-cv
11 May 2026
Model Releases

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos

DGX agent

arXiv:2512.03479v2 Announce Type: replace Abstract: Procedural activities are fundamentally driven by object state transitions, yet existing instructional video benchmarks remain action-centric and ca

model-releasesarxiv-cs-cv
11 May 2026
Safety

Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment

DGX agent

arXiv:2605.08064v1 Announce Type: new Abstract: Spatial intelligence in vision-language models (VLMs) attracts research interest with the practical demand to reason in the 3D world.Despite promising r

safetyarxiv-cs-cv
11 May 2026
Safety

PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection

DGX agent

arXiv:2509.26272v3 Announce Type: replace Abstract: The rapid rise of synthetic media has made deepfake detection a critical challenge for online safety and trust. Progress remains constrained by the

safetyarxiv-cs-cv
11 May 2026
Safety

Radiologist-Guided Causal Concept Bottleneck Models for Chest X-Ray Interpretation

DGX agent

arXiv:2605.07785v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) in medical imaging aim to improve model interpretability by predicting intermediate clinical concepts before final diag

safetyarxiv-cs-cv
11 May 2026
Model Releases

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation

DGX agent

arXiv:2605.07334v1 Announce Type: new Abstract: Video Reasoning Segmentation (VRS) aims to segment target objects in videos based on implicit instructions that convey human intent and temporal logic.

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

Real-IAD MVN: A Multi-View Normal Vector Dataset and Benchmark for High-Fidelity Industrial Anomaly Detection

DGX agent

arXiv:2605.07149v1 Announce Type: new Abstract: Industrial Anomaly Detection (IAD) is critical for quality control, but existing methods struggle with subtle, geometric defects. Standard 2D (RGB) imag

model-releasesarxiv-cs-cv
11 May 2026
Safety

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning

DGX agent

arXiv:2605.07477v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have achieved remarkable progress, however, many edited results still suffer from artifacts, unintended mo

safetyarxiv-cs-cv
11 May 2026
Research

Rebalancing gradient to improve self-supervised co-training of depth, odometry and optical flow predictions

DGX agent

arXiv:2605.07945v1 Announce Type: new Abstract: We present CoopNet, an approach that improves the cooperation of co-trained networks by dynamically adapting the apportionment of gradient, to ensure eq

researcharxiv-cs-cv
11 May 2026
Safety

ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation

DGX agent

arXiv:2408.06747v4 Announce Type: replace Abstract: Recent works utilize CLIP to perform the challenging unsupervised semantic segmentation task where only images without annotations are available. Ho

safetyarxiv-cs-cv
11 May 2026
Applications

RECON: Robust symmetry discovery via Explicit Canonical Orientation Normalization

DGX agent

arXiv:2505.13289v5 Announce Type: replace-cross Abstract: Real world data often exhibits unknown, instance-specific symmetries that rarely exactly match a transformation group G fixed a priori. Class-

applicationsarxiv-cs-cv
11 May 2026
Model Releases

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion

DGX agent

arXiv:2503.06223v5 Announce Type: replace Abstract: Large Vision-Language Models (VLMs) are increasingly deployed in open-ended environments, where ensuring reliable safety under multimodal inputs is

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

Rethinking Dense Optical Flow without Test-Time Scaling

DGX agent

arXiv:2605.08000v1 Announce Type: new Abstract: Recent progress in dense optical flow has been driven by increasingly complex architectures and multi-step refinement for test-time scaling. While these

model-releasesarxiv-cs-cv
11 May 2026
Research

RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection

DGX agent

arXiv:2602.19974v2 Announce Type: replace Abstract: Recent advancements in image generation have achieved impressive results in producing high-quality images. However, existing image generation models

researcharxiv-cs-cv
11 May 2026
Model Releases

S2M-Net: Spectral-Spatial Mixing for Medical Image Segmentation with Morphology-Aware Adaptive Loss

DGX agent

arXiv:2601.01285v2 Announce Type: replace Abstract: Medical image segmentation requires balancing local precision for boundary-critical clinical applications, global context for anatomical coherence,

model-releasesarxiv-cs-cv
11 May 2026
Safety

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models

DGX agent

arXiv:2605.07800v1 Announce Type: new Abstract: Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions spe

safetyarxiv-cs-cv
11 May 2026
Model Releases

Sat3R: Satellite DSM Reconstruction via RPC-Aware Depth Fine-tuning

DGX agent

arXiv:2605.07264v1 Announce Type: new Abstract: Accurate Digital Surface Model (DSM) reconstruction from satellite imagery is critical for applications such as disaster response, urban planning, and l

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

SatSurfGS: Generalizable 2D Gaussian Splatting for Sparse-View Satellite Surface Reconstruction

DGX agent

arXiv:2605.07181v1 Announce Type: new Abstract: Sparse-view satellite image surface reconstruction remains highly challenging, fundamentally because the reliability of multi-view matching under satell

model-releasesarxiv-cs-cv
11 May 2026
Research

Saving Foundation Flow-Matching Priors for Inverse Problems

DGX agent

arXiv:2511.16520v2 Announce Type: replace-cross Abstract: Foundation flow-matching (FM) models promise a universal prior for solving inverse problems (IPs), yet today they trail behind domain-specific

researcharxiv-cs-cv
11 May 2026
Model Releases

Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts

DGX agent

arXiv:2602.03473v2 Announce Type: replace-cross Abstract: Continual learning, especially class-incremental learning (CIL), on the basis of a pre-trained model (PTM) has garnered substantial research i

model-releasesarxiv-cs-cv
11 May 2026
Agents

See Tomorrow, Act Today: Foresight-Driven Autonomous Driving

DGX agent

arXiv:2605.07195v1 Announce Type: new Abstract: Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actio

agentsarxiv-cs-cv
11 May 2026
Local Ai

Seeing Across Skies and Streets: Feedforward 3D Reconstruction from Satellite, Drone, and Ground Images

DGX agent

arXiv:2605.07978v1 Announce Type: new Abstract: Cross-view localization classically asks: where does this ground image lie on the satellite tile? Existing methods are typically limited to 3-DoF estima

local-aiarxiv-cs-cv
11 May 2026
Hardware

SemanticDialect: Semantic-Aware Mixed-Format Quantization for Video Diffusion Transformers

DGX agent

arXiv:2603.02883v3 Announce Type: replace Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art video generation quality, but their substantial memory and computational footprints hinder ed

hardwarearxiv-cs-cv
11 May 2026
Model Releases

Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models

DGX agent

arXiv:2603.14186v4 Announce Type: replace Abstract: State-of-the-art text-to-image models produce high-quality images, but inference remains expensive as generation requires several sequential ODE or

model-releasesarxiv-cs-cv
11 May 2026
Model Releases

ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs

DGX agent

arXiv:2605.07338v1 Announce Type: new Abstract: The decline of global shellfish biodiversity poses a severe threat to coastal ecosystems. Although artificial intelligence (AI) technologies show potent

model-releasesarxiv-cs-cv
11 May 2026
Applications

SIMI: Self-information Mining Network for Low-light Image Enhancement

DGX agent

arXiv:2605.07767v1 Announce Type: new Abstract: Poor lighting conditions significantly impact image quality, posing substantial challenges for image editing and visualization. Many existing enhancemen

applicationsarxiv-cs-cv
11 May 2026
Local Ai

SoftSAE: Dynamic Top-K Selection for Adaptive Sparse Autoencoders

DGX agent

arXiv:2605.06610v2 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) have become an important tool in mechanistic interpretability, helping to analyze internal representations in both

local-aiarxiv-cs-cv
11 May 2026
Research

SoLAR: Error-Resilient Streamable Long-Horizon Free-Viewpoint Video Reconstruction with Anchor Activation and Latent Recalibration

DGX agent

arXiv:2605.07346v1 Announce Type: new Abstract: Free-Viewpoint Video (FVV) has emerged as a cornerstone of next-generation immersive media systems and attracted widespread attention. Previous methods

researcharxiv-cs-cv
11 May 2026
← Previous
1…188189190191192…263
Next →