AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture

DGX agent

arXiv:2511.17354v4 Announce Type: replace Abstract: Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for lear

researcharxiv-cs-cv
6 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation

DGX agent

arXiv:2608.04533v1 Announce Type: new Abstract: Part-level affordance grounding has advanced the localization of functional object regions associated with elemental actions. Extending this capability

model-releasesarxiv-cs-cv
6 Aug 2026
Agents

Embedding Large Language Models into Flow Controls: An Agentic Framework for Adaptive and Trustworthy Automated Cooking

DGX agent

arXiv:2608.04768v1 Announce Type: new Abstract: Automated cooking robots have traditionally relied on predefined procedures and rule-based control, ensuring stable execution but offering limited perso

agentsarxiv-cs-cv
6 Aug 2026
Research

Enhancing Low Back Pain Assessment with Diffusion Models for Lumbar Spine MRI Segmentation

DGX agent

arXiv:2608.04906v1 Announce Type: new Abstract: This study introduces a diffusion-based framework for robust and accurate semantic segmentation of lumbar spine MRI scans from patients with low back pa

researcharxiv-cs-cv
6 Aug 2026
Research

Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting

DGX agent

arXiv:2607.15890v2 Announce Type: replace Abstract: Perceiving multimodal cues and forecasting fine-grained actions from an egocentric (Ego) perspective is vital for applications like robot manipulati

researcharxiv-cs-cv
6 Aug 2026
Model Releases

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

DGX agent

arXiv:2608.04404v1 Announce Type: new Abstract: World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approach

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation

DGX agent

arXiv:2602.19161v2 Announce Type: replace Abstract: Latent diffusion models have enabled high-quality video synthesis, yet their inference remains costly and time-consuming. As diffusion transformers

safetyarxiv-cs-cv
6 Aug 2026
Safety

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

DGX agent

arXiv:2608.04530v1 Announce Type: new Abstract: GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact so

safetyarxiv-cs-cv
6 Aug 2026
Model Releases

Foreseeing the Invisible: Amodal Reconstruction of Leaf Fossil Images

DGX agent

arXiv:2608.04423v1 Announce Type: new Abstract: Fossil leaves are rarely preserved whole -- sedimentary rock hides, breaks, and erodes the lamina, yet paleobotany depends on the complete shape and out

model-releasesarxiv-cs-cv
6 Aug 2026
Research

Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection

DGX agent

arXiv:2608.04394v1 Announce Type: new Abstract: Cross-Domain Few-Shot Object Detection (CDFSOD) aims to transfer knowledge from data-rich upstream generic domains to downstream expert domains using sc

researcharxiv-cs-cv
6 Aug 2026
Safety

From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs

DGX agent

arXiv:2601.03808v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved notable performance in code synthesis; however, data-aware augmentation remains a limiting factor, handle

safetyarxiv-cs-cv
6 Aug 2026
Local Ai

From Understanding to Erasing: Towards Complete and Stable Video Object Removal

DGX agent

arXiv:2604.01693v2 Announce Type: replace Abstract: Video object removal aims to erase target objects while reconstructing visually plausible and temporally coherent content. However, target objects o

local-aiarxiv-cs-cv
6 Aug 2026
Research

Generative neural physics enables quantitative volumetric ultrasound of tissue mechanics

DGX agent

arXiv:2508.12226v3 Announce Type: replace Abstract: Ultrasound Tomography (UT) is a radiation-free, high-resolution modality, but remains limited for musculoskeletal imaging due to the high computatio

researcharxiv-cs-cv
6 Aug 2026
Research

Global Attention-Fused Image Cropping with Attention-Guided and Global-Aligned Crop Evaluator

DGX agent

arXiv:2608.04821v1 Announce Type: new Abstract: Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods

researcharxiv-cs-cv
6 Aug 2026
Model Releases

HelloWorld: Enabling Socially Interactive Characters in Video World Models

DGX agent

arXiv:2608.05070v1 Announce Type: new Abstract: Despite the remarkable recent progress of video world models, social interaction between users and the characters within these worlds remains unsupporte

model-releasesarxiv-cs-cv
6 Aug 2026
Research

HexMIL: Hierarchical Attention MIL for Ante-Hoc Explainable Detection of AI-Manipulated CT Volumes

DGX agent

arXiv:2608.05101v1 Announce Type: new Abstract: The emergence of medical deepfakes, i.e., medical images manipulated by deep generative models, poses a significant threat to clinical workflows. Howeve

researcharxiv-cs-cv
6 Aug 2026
Research

HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding

DGX agent

arXiv:2608.04610v1 Announce Type: new Abstract: 3D vision-language models (3D VLMs) enable spatial reasoning over multi-view scenes but suffer from substantial token redundancy due to duplicated obser

researcharxiv-cs-cv
6 Aug 2026
Applications

Industrial Synthetic Segment Pre-training

DGX agent

arXiv:2505.13099v3 Announce Type: replace Abstract: Vision Foundation Models (VFMs) have made remarkable progress and are increasingly being applied to segmentation tasks in real-world industrial sett

applicationsarxiv-cs-cv
6 Aug 2026
Local Ai

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers

DGX agent

arXiv:2608.05122v1 Announce Type: new Abstract: Vision transformers (ViTs) have become the de facto standard for image encoding across many perception tasks. Despite their empirical success, it remain

local-aiarxiv-cs-cv
6 Aug 2026
Model Releases

Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation

DGX agent

arXiv:2607.10082v2 Announce Type: replace Abstract: Building pixel-level correspondence between event and image data is a fundamental task for multi-sensor systems. However, existing cross-modal match

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

Lesion Detection in CT with Frozen Self-Distilled Features: SALT, a Spatially Adaptive Label-Guided Temperature

DGX agent

arXiv:2608.05100v1 Announce Type: new Abstract: Self-supervised pretraining objectives are spatially uniform: the teacher temperature and the per-patch loss weight are identical everywhere in the imag

safetyarxiv-cs-cv
6 Aug 2026
Model Releases

LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content

DGX agent

arXiv:2410.10783v4 Announce Type: replace Abstract: The large-scale training of multi-modal models on data scraped from the web has shown outstanding utility in infusing these models with the required

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching

DGX agent

arXiv:2608.04106v1 Announce Type: new Abstract: Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending d

model-releasesarxiv-cs-cv
6 Aug 2026
Agents

MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding

DGX agent

arXiv:2608.04587v1 Announce Type: new Abstract: Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in

agentsarxiv-cs-cv
6 Aug 2026
Research

MME: Mixture of Mesh Experts with Random Walk Transformer Gating

DGX agent

arXiv:2603.00828v2 Announce Type: replace Abstract: In recent years, various methods have been proposed for mesh analysis, each offering distinct advantages and often excelling on different object cla

researcharxiv-cs-cv
6 Aug 2026
Research

MOAT: Model-Agnostic Randomized Transformations for preventing Efficiency Degradation Attacks on ViTs

DGX agent

arXiv:2608.04680v1 Announce Type: cross Abstract: To adopt the Vision Transformers (ViTs) in resource-constrained environment, token pruning is widely used to reduce computational cost without impacti

researcharxiv-cs-cv
6 Aug 2026
Model Releases

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

DGX agent

arXiv:2608.04657v1 Announce Type: new Abstract: World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mob

model-releasesarxiv-cs-cv
6 Aug 2026
Research

Multi-View Face and Gesture Animation with Dynamic Gaussians

DGX agent

arXiv:2608.04722v1 Announce Type: new Abstract: Creating photorealistic 3D human avatars with realistic upper-body motion remains challenging. Existing approaches either focus on the head and overlook

researcharxiv-cs-cv
6 Aug 2026
Safety

muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards

DGX agent

arXiv:2608.04412v1 Announce Type: new Abstract: High-quality driving data are essential for autonomous-driving systems and generative world models. However, rare and safety-critical scenarios involvin

safetyarxiv-cs-cv
6 Aug 2026
Research

MVTOP: Multi-View Transformer-based Object Pose-Estimation

DGX agent

arXiv:2508.03243v2 Announce Type: replace Abstract: We present MVTOP, a novel transformer-based method for multi-view rigid object pose estimation. Through an early fusion of the view-specific feature

researcharxiv-cs-cv
6 Aug 2026
Safety

Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles

DGX agent

arXiv:2608.04483v1 Announce Type: new Abstract: Vision-language models (VLMs) process an image as a sequence of visual tokens, which creates a substantial computational bottleneck during inference. Re

safetyarxiv-cs-cv
6 Aug 2026
Applications

Objects as Audio-Visual Modal Sound Fields

DGX agent

arXiv:2608.05145v1 Announce Type: new Abstract: While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical in

applicationsarxiv-cs-cv
6 Aug 2026
Model Releases

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

DGX agent

arXiv:2608.05049v1 Announce Type: new Abstract: Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmar

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

DGX agent

arXiv:2608.04434v1 Announce Type: new Abstract: Recent large language models (LLMs) have demonstrated remarkable progress in constraint-aware navigation, maze reasoning, and graph reasoning. However,

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films

DGX agent

arXiv:2608.04224v1 Announce Type: new Abstract: Historical films suffer from co-occurring visual and audio degradations---blur, noise, flicker, hiss, clipping, and dropout---yet existing methods resto

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing

DGX agent

arXiv:2608.04791v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative training of deep learning models across decentralized image archives without requiring data centralization

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

OutLangSplat: 3D Language Gaussian Splatting for UAV Outdoor Scenes

DGX agent

arXiv:2608.04560v1 Announce Type: new Abstract: 3D Language Gaussian Splatting embeds open-vocabulary language features into 3D Gaussian Splatting, providing an efficient explicit representation for t

safetyarxiv-cs-cv
6 Aug 2026
Safety

Overcoming Statistical Bias in Action-Controllable World Models

DGX agent

arXiv:2608.04653v1 Announce Type: new Abstract: Action-conditioned world models aim to predict how visual environments evolve under an agent's actions. Yet future frames are often highly predictable f

safetyarxiv-cs-cv
6 Aug 2026
Model Releases

PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images

DGX agent

arXiv:2608.04210v1 Announce Type: new Abstract: Pose-agnostic Anomaly Detection (PAD) remains challenging as anomalies can appear under arbitrary viewpoints, requiring methods to handle significant po

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

Persistent Object Narratives for Token-Efficient Video Language Models

DGX agent

arXiv:2608.04866v1 Announce Type: new Abstract: Video large language models (Video-LLMs) have made strong progress in open-ended video understanding. However, their visual interfaces remain token-inte

model-releasesarxiv-cs-cv
6 Aug 2026
Safety

Poly-OPD: Heterogeneous Multi-Teacher On-Policy Distillation for Capability-Selectable Flow Models

DGX agent

arXiv:2608.04349v1 Announce Type: new Abstract: Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional

safetyarxiv-cs-cv
6 Aug 2026
Research

Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous Time with Graph-Based Metric Tensor Embeddings

DGX agent

arXiv:2608.05132v1 Announce Type: new Abstract: Predicting how a subcortical structure's shape will evolve from a few prior scans could support prognosis and clinical-trial enrichment. Existing longit

researcharxiv-cs-cv
6 Aug 2026
Model Releases

Privacy-Preserving Action Recognition: Taxonomy, Methods, and Privacy-Utility Trade-offs

DGX agent

arXiv:2608.04501v1 Announce Type: new Abstract: Video surveillance in public safety, healthcare, and smart environments has made continuous human monitoring routine, raising real risks to personal ide

model-releasesarxiv-cs-cv
6 Aug 2026
Local Ai

Promptable Animal Pose Tracking Across Species

DGX agent

arXiv:2608.04995v1 Announce Type: new Abstract: Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated

local-aiarxiv-cs-cv
6 Aug 2026
Model Releases

Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

DGX agent

arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modali

model-releasesarxiv-cs-cv
6 Aug 2026
Research

RegisterBridgeMM: A Register-Centric Framework for RGB-Infrared Object Detection

DGX agent

arXiv:2608.04833v1 Announce Type: new Abstract: RGB-infrared (RGB-IR) object detection benefits from complementary visible and thermal cues, but effective fusion remains challenging under illumination

researcharxiv-cs-cv
6 Aug 2026
Model Releases

ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

DGX agent

arXiv:2608.04385v1 Announce Type: new Abstract: Vision-Language Models (VLMs) often lose visual grounding during multi-step reasoning: as reasoning chains grow longer, later inference steps rely incre

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

ResPlan: A Large-Scale Vector-Graph Dataset of 17,000 Residential Floor Plans

DGX agent

arXiv:2508.14006v2 Announce Type: replace Abstract: We introduce ResPlan, a dataset of 17,000 residential floor plans with vector geometry, room-connectivity graphs, and metric-scale coordinates. Each

model-releasesarxiv-cs-cv
6 Aug 2026
← Previous
1…1213141516…259
Next →