AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
4 Aug 2026

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph

ApplicationsDGX agent

arXiv:2608.01905v1 Announce Type: new Abstract: Hand-object interaction (HOI) is a fundamental human behavior with broad applications in AR/VR, digital humans, and embodied interaction. Existing metho

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

ResearchDGX agent

arXiv:2608.02150v1 Announce Type: new Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of

PhysAgent: A Multi-Agent Framework for Reliable Remote Heart Rate Estimation

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.00066v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) enables non-contact heart-rate estimation from facial videos, but its weak physiological signal is easily corrupted b

Pixel Ignores, Superpixel Sees: Adverse Weather Image Restoration via Semantic-Center SSM

ResearchDGX agent

arXiv:2608.01760v1 Announce Type: new Abstract: Adverse weather image restoration aims to recover clear visibility from degraded images in complex weather conditions. Existing works attempt to address

PixelSR: Efficient Screen Content Super-Resolution via Pixel Classification

Local AiDGX agent

arXiv:2608.00646v1 Announce Type: new Abstract: Screen content images are generally composed of texts and graphics. Compared to natural images, these man-made images contain a large quantity of sharp

PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle

TutorialsDGX agent

arXiv:2608.01354v1 Announce Type: new Abstract: Recent studies develop pixel-level multimodal large language models (MLLMs) that support both Region Segmentation and Region Understanding, extending mu

PlantRig - From Bones to Branches: Adaptation of Autoregressive Rigging Models for Plant Skeletal Reconstruction

SafetyDGX agent

arXiv:2608.01072v1 Announce Type: new Abstract: Autoregressive rigging models such as UniRig and SkinTokens perform well on articulated characters, but their ability to generalize to plant structures

PNEC-Mamba: Prototype-Guided Positive-Negative Evidence Calibration for Hyperspectral Image Classification

Model ReleasesDGX agent

arXiv:2608.01910v1 Announce Type: new Abstract: In real-world hyperspectral scenes, pixel representations are often ambiguous due to factors such as spectral similarity, mixed pixels, and local contex

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

ResearchDGX agent

arXiv:2608.00440v1 Announce Type: new Abstract: Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single succ

Practical Noise Modeling for SPAD Intensity Imaging

SafetyDGX agent

arXiv:2608.00489v1 Announce Type: new Abstract: Single-photon avalanche diode (SPAD) cameras are promising for low-light and high-dynamic-range intensity imaging, but their practical use is limited by

PRISM: Privileged Probabilistic Latent Supervision for End-to-End Autonomous Driving Motion Planning

AgentsDGX agent

arXiv:2608.01201v1 Announce Type: cross Abstract: End-to-end autonomous driving (E2E AD) systems integrate perception, prediction, and planning into a single differentiable architecture. While these m

Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers

TutorialsDGX agent

arXiv:2608.01495v1 Announce Type: new Abstract: Detection transformer models, including DETR and its extensions, learn to output a set of object-level embeddings that can be simultaneously decoded int

Prompt-Driven Simulation with Feature Perturbation for Cross-Domain Few-Shot Object Detection

ResearchDGX agent

arXiv:2608.01348v1 Announce Type: new Abstract: Data augmentation, which simulates diverse visual variations to expand the source distribution and induce synthetic domain shifts, is a simple yet effec

PromptPath: Prompt-Adaptive Computational Pathways for In-Context Learning

ResearchDGX agent

arXiv:2608.02129v1 Announce Type: new Abstract: In-context learning (ICL) has attracted increasing attention for enabling models to perform new tasks using only a few ``input--output'' prompt examples

Proteus: A Truncation-Robust Entropy Model for Progressive LiDAR Compression

SafetyDGX agent

arXiv:2608.00687v1 Announce Type: new Abstract: LiDAR point clouds provide explicit, deterministic physical boundaries critical for collaborative safety-critical perception. However, wireless channels

Protocol generalisation for brain tissue microstructure estimation via hypernetwork-controlled geometric deep learning

Model ReleasesDGX agent

arXiv:2608.02053v1 Announce Type: cross Abstract: Brain tissue microstructure estimation with machine learning provides higher computational efficiency than conventional fitting. However, machine lear

Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation

ResearchDGX agent

arXiv:2608.01978v1 Announce Type: new Abstract: Audio-driven portrait animation has advanced rapidly with diffusion-based generative models, yet real-time one-shot generation with expressive emotion c

Quaternion Tensor Modeling for Joint Color-Polarization Demosaicking

ResearchDGX agent

arXiv:2608.02144v1 Announce Type: new Abstract: Division-of-focal-plane (DoFP) color polarization cameras enable snapshot acquisition of color polarization mosaic images, but the inherently sparse sam

QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction

Model ReleasesDGX agent

arXiv:2608.01186v1 Announce Type: new Abstract: While feed-forward 3D Gaussian Splatting (3DGS) enables efficient 3D reconstruction, achieving high-fidelity rendering remains challenging. Existing pix

RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding

SafetyDGX agent

arXiv:2608.00147v1 Announce Type: new Abstract: Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a sing

RadYOLO: Computationally Efficient 3D Object Detection and Segmentation in CT and MRI

Local AiDGX agent

arXiv:2608.00508v1 Announce Type: new Abstract: Object detection and segmentation in three-dimensional medical images is a very active area of research. However, most proposed deep learning models car

Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for Infrared-Visible Fusion Assessment

Model ReleasesDGX agent

arXiv:2608.01301v1 Announce Type: new Abstract: Infrared-visible image fusion (IVIF) has no ideal fused reference, so fusion algorithms are routinely ranked by scalar objective metrics that formalize

ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models

ResearchDGX agent

arXiv:2608.01067v1 Announce Type: new Abstract: Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the

Real-Time Visual Obstruction Detection in Surgical Augmented Reality

Model ReleasesDGX agent

arXiv:2608.00232v1 Announce Type: new Abstract: Surgical augmented reality (AR) can provide contextual guidance by overlaying virtual annotations, tool cues, and procedural information onto the surgic

Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection

ResearchDGX agent

arXiv:2608.01930v1 Announce Type: new Abstract: Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficien

Reconstruction-Shift Discrimination via Mask-Guided Latent Diffusion for Medical Anomaly Detection

Local AiDGX agent

arXiv:2608.00444v1 Announce Type: new Abstract: Unsupervised medical anomaly detection learns normal anatomical patterns from healthy training images and identifies deviations at test time. Reconstruc

Recursive Vision Language Models for General Symbolic Reasoning

Model ReleasesDGX agent

arXiv:2608.01534v1 Announce Type: new Abstract: Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive reasoning, w

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2608.00574v1 Announce Type: new Abstract: Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this

Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning

ResearchDGX agent

arXiv:2608.01314v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengthe

ReMiX-MAE: Learning Missing-Channel Cross-Modal Representations from RGB-Only Clinical Facial Videos for Sympathetic-Mediated Pain Assessment

ApplicationsDGX agent

arXiv:2608.02561v1 Announce Type: new Abstract: Automated pain assessment in real clinics is limited by scarce clinically grounded facial video data with weak labels (often sequence-level self-report)

Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging

ResearchDGX agent

arXiv:2608.00586v1 Announce Type: new Abstract: Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pret

Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance

TutorialsDGX agent

arXiv:2510.21590v3 Announce Type: replace Abstract: Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image qu

Rethinking IRSTD: Single-Point Supervision Guided Encoder-only Framework is Enough for Infrared Small Target Detection

ResearchDGX agent

arXiv:2604.05363v2 Announce Type: replace Abstract: Infrared small target detection (IRSTD) aims to separate small targets from clutter backgrounds. Extensive research is dedicated to the pixel-level

Rethinking Pretraining for Specialized Design Data: Evidence from the JONES-19 Cultural Design Dataset

ResearchDGX agent

arXiv:2608.00135v1 Announce Type: cross Abstract: Design and architectural archives encode expert human knowledge in graphical formats, providing a critical testbed for design-inspired Machine Learnin

Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere

ResearchDGX agent

arXiv:2608.01271v1 Announce Type: new Abstract: Video large language models (Video-LLMs) represent videos as dense sequences of visual tokens, whose length grows with the temporal and spatial extent o

Retrieval-Based Cross-Domain Generalization in Optical Networks via Global Features

ResearchDGX agent

arXiv:2608.00044v1 Announce Type: cross Abstract: We propose a retrieval-based framework for crossdomain quality-of-transmission (QoT) estimation that leverages transferable feature representations wh

Robust Watermarks Meet Backdoored Models: Evading Diffusion Semantic Watermarks via Stealthy Backdoor

Model ReleasesDGX agent

arXiv:2608.00543v1 Announce Type: cross Abstract: Although semantic watermarking is considered a promising safeguard for images generated by Latent Diffusion Models (LDMs), the reliance of the waterma

Rolling Shutter Camera Self-Calibration

ResearchDGX agent

arXiv:2608.01509v1 Announce Type: new Abstract: Rolling shutter (RS) cameras are widely used in consumer devices, but their row-wise exposure causes distortions under motion, making geometric 3D visio

Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis

Local AiDGX agent

arXiv:2608.01973v1 Announce Type: cross Abstract: Existing indoor layout generators produce globally plausible layouts yet may retain local violations such as collisions, out-of-bounds placements, obs

RPL-UIE: Reliable Prior Learning for Underwater Image Enhancement

ApplicationsDGX agent

arXiv:2608.00137v1 Announce Type: cross Abstract: Underwater image enhancement (UIE) aims to recover clear images from observations affected by wavelength-dependent absorption, scattering, and spatial

RSC-GestureNet: Reliability-Aware Selective Causal Recognition of Chinese Traffic Police Gestures

Model ReleasesDGX agent

arXiv:2608.02200v1 Announce Type: new Abstract: Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous

RSRA: Training-Free Probing of Representation Sensitivity for Efficient LoRA Rank Allocation

Model ReleasesDGX agent

arXiv:2607.09757v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning enables large language models to adapt to downstream tasks with substantially lower computational and storage cost,

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?

Model ReleasesDGX agent

arXiv:2608.02039v1 Announce Type: new Abstract: Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, acti

SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining

Model ReleasesDGX agent

arXiv:2608.00068v1 Announce Type: new Abstract: Construction-safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only re

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression

SafetyDGX agent

arXiv:2608.02109v1 Announce Type: new Abstract: Vision-Text Compression (VTC) renders long texts into images and encodes them through the vision encoder (ViT), compressing thousands of text tokens int

SARe: Structure-Aware Generative 3D Fragment Reassembly

Local AiDGX agent

arXiv:2603.21611v2 Announce Type: replace Abstract: 3D fragment reassembly estimates the rigid pose of each fragment to recover a complete object from unordered point clouds or meshes. The task become

SCALP: Semi-Supervised Statistical Shape Modeling from Imperfect 3D Photogrammetry via Landmark-Anchored Spectral Warp

ApplicationsDGX agent

arXiv:2608.00187v1 Announce Type: new Abstract: Correspondence-based statistical shape modeling (SSM) is vital for population-level morphometric analysis, but conventional pipelines assume clean, full

Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

ApplicationsDGX agent

arXiv:2608.00463v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these world

Score-Based Turbo Message Passing for Plug-and-Play Compressive Imaging

ResearchDGX agent

arXiv:2512.14435v2 Announce Type: replace Abstract: Message-passing algorithms have been adapted for compressive imaging by incorporating various off-the-shelf image denoisers. However, these denoiser

SecondOpinion: Anatomy-Aware Gated Reasoning for Efficient Medical Image Analysis

ResearchDGX agent

arXiv:2608.01808v1 Announce Type: new Abstract: Deep learning models for medical image analysis typically apply a fixed amount of computation to every input, regardless of case difficulty. Anatomy-gui

Seeing the Unseen: Towards Training-Free Inspection for Wind Turbine Blades Using Knowledge-Augmented Vision Language Models

ResearchDGX agent

arXiv:2510.22868v2 Announce Type: replace Abstract: Wind turbine blades operate in harsh environments, making timely damage detection essential for preventing failures and optimizing maintenance. Dron

Self-supervised DXA representations encode multi-system disease risk, biological aging and heritability

ResearchDGX agent

arXiv:2608.02208v1 Announce Type: new Abstract: Whole-body dual-energy X-ray absorptiometry (DXA) scans are routinely acquired to measure bone density and regional body composition, leaving their spat

Semantic-Guided Cross-Sensor Super Resolution of Remote Sensing Images: A Gated Dual Conditioning Flow Matching Model

Model ReleasesDGX agent

arXiv:2510.23816v3 Announce Type: replace Abstract: High spatial resolution satellite imagery is critical for monitoring fine-scale Earth surface processes, but is often limited by cost and revisit ti

Semantically Calibrated Evidence Composition for CT Vision-Language Learning

SafetyDGX agent

arXiv:2608.00239v1 Announce Type: new Abstract: Learning transferable representations from CT-report pairs requires combining whole-volume context with anatomy-specific evidence. Existing methods typi

Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration

SafetyDGX agent

arXiv:2608.02285v1 Announce Type: new Abstract: We propose Sen-Cap, a Sensor-Flexible and Noise-Resilient 3D human motion Capture framework that integrates multi-modal data from LiDAR and camera. Whil

SG-Layout: Structured Scene Graph-Guided Layout Generation with LLMs

SafetyDGX agent

arXiv:2608.01106v1 Announce Type: new Abstract: Understanding and generating spatially coherent layouts from natural language remains a fundamental yet challenging task for large language models (LLMs

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

SafetyDGX agent

arXiv:2608.01397v1 Announce Type: cross Abstract: World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are model

Similarity Weighted Aggregation with Global Differential Privacy for Federated Brain Lesion Segmentation

ResearchDGX agent

arXiv:2608.00872v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative training of machine learning models across multiple institutions without sharing sensitive data, making

SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents

SafetyDGX agent

arXiv:2608.01306v1 Announce Type: new Abstract: Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods s

SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.00100v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly being evaluated for medical imaging, but many available benchmarks emphasize disease classification, repo

← Previous
1…1718192021…207
Next →