AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
9 Jun 2026

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs

ResearchDGX agent

arXiv:2606.08511v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) face a significant inference bottleneck due to the quadratic computational cost of self-attention over long vis

MAGIS: Evidence-Based Multi-Agent Reasoning for Interpretable Strabismus Clinical Decision-Making

Model ReleasesDGX agent

arXiv:2606.09249v1 Announce Type: new Abstract: Strabismus is a common ocular disorder that requires fine-grained subtype diagnosis for individualized treatment planning. However, existing deep learni

MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.08788v1 Announce Type: new Abstract: Representation alignment with pretrained vision models has recently shown strong potential for accelerating diffusion transformer training. By aligning

MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding

AgentsDGX agent

arXiv:2606.09641v1 Announce Type: new Abstract: The dominant paradigm in video retrieval relies on embedding-based full-corpus scanning, which suffers from inherent computational inefficiency and the

MB-Loc: Multi-planar Bird's-eye-view Localization in outdoor LiDAR scenes

Local AiDGX agent

arXiv:2606.08744v1 Announce Type: new Abstract: Global LiDAR localization is a fundamental task for autonomous navigation systems. Recent methods perform Scene Coordinate Regression (SCR) and achieve

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models

ResearchDGX agent

arXiv:2606.09827v1 Announce Type: cross Abstract: Temporal modeling is essential for robotic manipulation, as effective control requires both memory of past interactions and imagination of future stat

MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation

ResearchDGX agent

arXiv:2606.09056v1 Announce Type: new Abstract: Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames req

Minimal Solvers for Full-DoF Motion Estimation from Asynchronous Differential SfM

ApplicationsDGX agent

arXiv:2606.09218v1 Announce Type: new Abstract: As a bio-inspired intelligent sensor, event cameras have introduced a new paradigm in the intelligent perception of spatiotemporal information and visua

MinNav: Minimalist Navigation Using Optical Flow For Active Tiny Aerial Robots

AgentsDGX agent

arXiv:2606.07813v1 Announce Type: cross Abstract: Navigation using a monocular camera is pivotal for autonomous operation on tiny aerial robots due to their perfect balance of versatility, cost and ac

Mitigating Diffusion Model Hallucinations with Dynamic Guidance

ResearchDGX agent

arXiv:2510.05356v2 Announce Type: replace Abstract: Hallucinations in diffusion models are samples with structural inconsistencies that can emerge due to the excessive smoothing of the learned score f

MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

ResearchDGX agent

arXiv:2410.21747v2 Announce Type: replace Abstract: Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging re

MS-COOT: Comparing Morse-Smale Complexes with Co-Optimal Transport

ResearchDGX agent

arXiv:2606.08258v1 Announce Type: cross Abstract: Understanding and comparing structures in scalar fields is a central challenge in scientific visualization, with applications ranging from feature ana

Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training

SafetyDGX agent

arXiv:2601.03256v2 Announce Type: replace Abstract: We present Muses, the first training-free method for fantastic 3D creature generation in a feed-forward paradigm. Previous methods, which rely on pa

Need We Teach Foundation Models What is a Generative Image? Gradient-Free Generative Artifact Detection via Analytic Spectral Adaptation

Local AiDGX agent

arXiv:2606.07660v1 Announce Type: new Abstract: Adapting foundation models to detect generative artifacts via gradient-based updates compromises their intrinsic representations. Under optimization on

Neural Field Tokenizations with Hierarchy and Spatial Locality Priors

TutorialsDGX agent

arXiv:2606.08204v1 Announce Type: cross Abstract: Neural fields parameterize data as functions from coordinates to values, providing a unified framework for representation learning across modalities.

NGram-MoSE: Efficient Remote Sensing Super-Resolution via N-Gram Context and Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2606.08535v1 Announce Type: new Abstract: Remote sensing applications for environmental monitoring and disaster management are frequently constrained by a spatial--temporal trade-off: imagery wi

No Modality Left Behind: Adapting to Missing Modalities via Knowledge Distillation for Brain Tumor Segmentation

SafetyDGX agent

arXiv:2509.15017v2 Announce Type: replace Abstract: Accurate brain tumor segmentation is essential for preoperative evaluation and personalized treatment. Multi-modal MRI is widely used due to its abi

OctaOctree Neural Radiosity for Real-time Glossy Material Rendering

ResearchDGX agent

arXiv:2606.08469v1 Announce Type: cross Abstract: Modeling high-frequency outgoing radiance distributions remains a fundamental challenge in global illumination, especially for glossy and specular mat

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

Model ReleasesDGX agent

arXiv:2606.08572v1 Announce Type: new Abstract: While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability t

OmniFaceRig: Fully Automatic Inner-Mouth-Aware Face Rigging Across Diverse 3D Character Topologies

Model ReleasesDGX agent

arXiv:2606.08043v1 Announce Type: cross Abstract: Facial rigging - creating FACS-based blendshapes together with inner-mouth geometry (teeth, gums, and tongue) - remains a major bottleneck in 3D chara

OmniGen-AR: AutoRegressive Any-to-Image Generation

Model ReleasesDGX agent

arXiv:2606.09156v1 Announce Type: new Abstract: Autoregressive (AR) models have demonstrated strong potential in visual generation, offering superior performance with simple architectures and optimiza

OmniTryOn: Video Try-On Anything at Once!

Model ReleasesDGX agent

arXiv:2606.08514v1 Announce Type: new Abstract: Although video virtual try-on (VVT) has achieved significant progress, existing methods still exhibit two fundamental limitations: first, they are restr

One Stone, Three Birds: Self-adaptive Optimal Transport for Multi-VLM Selection, Adaptation, and Ensembling

ResearchDGX agent

arXiv:2606.08126v1 Announce Type: new Abstract: Vision-language models (VLMs) enable visual recognition from semantic class descriptions, which makes them attractive when target annotations are scarce

Optical Music Recognition for Real-World Manuscripts with Synthetic Data

ApplicationsDGX agent

arXiv:2606.09479v1 Announce Type: new Abstract: Optical Music Recognition (OMR) has seen major progress in model design, with end-to-end methods now capable of recognising notation at all levels of co

Optimizing Few-Step Generation with Adaptive Matching Distillation

ResearchDGX agent

arXiv:2602.07345v2 Announce Type: replace Abstract: Distribution Matching Distillation (DMD) is a powerful acceleration paradigm, yet its stability is often compromised in Forbidden Zone, regions wher

OrderDP: A Theoretically Guaranteed Lossless Dynamic Data Pruning Framework

SafetyDGX agent

arXiv:2606.08574v1 Announce Type: cross Abstract: Data pruning (DP), as an oft-stated strategy to alleviate heavy training burdens, reduces the volume of training samples according to a well-defined p

PairWise Image Finder: An Open-source Tool for Finding Visually Aligned Street-Level Image Pairs for Urban Perception Studies

SafetyDGX agent

arXiv:2606.08795v1 Announce Type: new Abstract: Change detection and scene recognition techniques have been widely applied to Street View Imagery (SVI) to understand changes in scenes across the years

Pantheon360: Taming Digital Twin Generation via 3D-Aware 360{eg} Video Diffusion

ResearchDGX agent

arXiv:2605.25449v2 Announce Type: replace Abstract: Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constr

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation

Model ReleasesDGX agent

arXiv:2510.20182v2 Announce Type: replace Abstract: Pedestrian simulation traditionally relies on expert-tuned, hand-crafted models that limit scalability and generalization. Meanwhile, large-scale vi

PereStruct: Multimodal Semantic Assembly for Robust Historical Document Parsing

Model ReleasesDGX agent

arXiv:2606.07661v1 Announce Type: new Abstract: Parsing historical documents with complex, non-standard layouts remains a fundamental bottleneck in large-scale archival digitization. Unlike modern typ

Phase Marginalization for Patch-Grid Instability in Vision Transformers

ResearchDGX agent

arXiv:2606.08132v1 Announce Type: new Abstract: Vision Transformers operate on fixed patch grids, which can introduce phase-dependent instability for dense prediction: changing the patch partition can

PhysAgent: Automating Physics-Based 4D Synthesis via Trajectory-Grounded Multi-Agent Feedback

Local AiDGX agent

arXiv:2606.08688v1 Announce Type: cross Abstract: Achieving fully automated, physically plausible 3D motion synthesis is a core objective in graphics and generative AI. However, configuring complex en

PhysGraph: A Physics-aware 3D Scene Graph for Perception and Reasoning

ApplicationsDGX agent

arXiv:2606.08655v1 Announce Type: cross Abstract: To perform a wide range of daily tasks, robots need to construct a 3D representation that is semantically rich, physically grounded, and structured en

PicoSAM3: Real-Time In-Sensor Region-of-Interest Segmentation

Local AiDGX agent

arXiv:2603.11917v2 Announce Type: replace Abstract: Real-time, on-device segmentation is critical for latency-sensitive and privacy-aware applications such as smart glasses and Internet-of-Things devi

Polaffini: A feature-based approach for robust affine and polyaffine image registration

SafetyDGX agent

arXiv:2602.17337v2 Announce Type: replace Abstract: In this work we present Polaffini, a robust and versatile framework for anatomically grounded registration. Medical image registration is dominated

POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction

Model ReleasesDGX agent

arXiv:2606.09788v1 Announce Type: new Abstract: Large-scale document processing requires contextually aware table extraction (TE) that is both accurate and efficient. Yet current approaches require bi

Prisma-World: Camera-Controllable Multi-Agent Video World Model

SafetyDGX agent

arXiv:2606.09507v1 Announce Type: new Abstract: Video world models have made rapid progress in generating controllable visual experiences, but most of them still simulate the world from a single obser

Programmable Silicon Retina on Pixel Processor Array

Model ReleasesDGX agent

arXiv:2606.08370v1 Announce Type: cross Abstract: Standard dynamic vision sensors approximate retinal processing by detecting temporal contrast changes, offering high speed and high dynamic range. In

Property-Informed Diffusion-Based Text-to-Microstructure Generation

SafetyDGX agent

arXiv:2606.08150v1 Announce Type: new Abstract: Designing 3D metamaterial microstructures that meet the intended functions remains a major challenge, as it typically requires domain expertise, iterati

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping

SafetyDGX agent

arXiv:2606.08708v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reasoning capability of Large Vision-Language M

Quantifying Noise of Dynamic Vision Sensor

TutorialsDGX agent

arXiv:2404.01948v3 Announce Type: replace Abstract: Dynamic visual sensors (DVS) are characterized by a large amount of background activity (BA) noise, which it is mixed with the original (cleaned) se

QuoVLA: Quotient Space for Vision-Language-Action Models

ResearchDGX agent

arXiv:2605.24890v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models commonly adapt pretrained Vision-Language Models (VLMs) to robot control by mapping visual observations and lang

RAD: A Dataset and Benchmark for Real-Life Anomaly Detection with Robotic Observations

Model ReleasesDGX agent

arXiv:2410.00713v4 Announce Type: replace Abstract: Anomaly detection is a core capability for robotic perception and industrial inspection, yet most existing benchmarks are collected under controlled

REACT 2026: The Fourth Multiple Appropriate Facial Reaction Generation Challenge: Personalised MAFRG and Appropriate EEG Reaction Prediction

ResearchDGX agent

arXiv:2606.07935v1 Announce Type: new Abstract: In dyadic interactions, various human facial reactions could be appropriate for responding to each human speaker behaviour. Following the successful org

Readable Yet Unpredictable: Rotated-Outcome Prediction in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.07641v1 Announce Type: new Abstract: Can vision-language models predict what a 180{eg} rotation would reveal from the original image alone? We study this ability through Rotated-Outcome Pre

Real-Time Industrial Defect Detection on Edge Hardware Using Fine-Tuned YOLOv8: A Systematic Benchmark on the NEU Surface Defect Database and MVTec AD with Automotive & Battery Manufacturing Extensions

Model ReleasesDGX agent

arXiv:2606.07659v1 Announce Type: new Abstract: Automated surface defect detection is critical for ensuring rigorous quality control in high-speed manufacturing environments. While deep learning model

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning

Model ReleasesDGX agent

arXiv:2606.09303v1 Announce Type: new Abstract: The rapid development of pretrained foundation models has enabled more general image segmentation. Multimodal large language models (MLLMs) have been wi

REFINE: Super-efficient 3D Gaussian Splatting Pruning via Rendering-Free Primitive Importance

Model ReleasesDGX agent

arXiv:2606.09074v1 Announce Type: new Abstract: Existing pruning methods for 3D Gaussian splatting (3DGS) suffer from either severe quality degradation or prohibitive computational overhead. In this p

Region-Wise Correspondence Prediction between Manga Line Art Images

SafetyDGX agent

arXiv:2509.09501v4 Announce Type: replace Abstract: Understanding region-wise correspondences between manga line art images is fundamental for high-level manga processing, supporting downstream tasks

Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning

SafetyDGX agent

arXiv:2606.08436v1 Announce Type: new Abstract: The task of temporal answer grounding in instructional video (TAGV), which aims to locate precise video segments that respond to natural language querie

Relational Epipolar Graphs for Robust Relative Camera Pose Estimation

Local AiDGX agent

arXiv:2604.04554v2 Announce Type: replace Abstract: A key component of Visual Simultaneous Localization and Mapping (VSLAM) is estimating relative camera poses using matched keypoints. Accurate estima

Remember with Confidence: Uncertainty Quantification for Spatio-temporal Memory with Probabilistic Guarantees

Model ReleasesDGX agent

arXiv:2606.08277v1 Announce Type: new Abstract: Long-horizon robot operation requires spatio-temporal memory to record the environment state and recall it for downstream reasoning. Scene graphs and re

Rethinking 3D Shape Generation: Diffusion over Superquadrics

ResearchDGX agent

arXiv:2606.08957v1 Announce Type: new Abstract: Diffusion models have advanced 3D shape generation, yet most methods still denoise in high-cardinality spaces (e.g., voxel/SDF grids, meshes, or point c

Revisiting Articulated Parts Perception in Robot Manipulation

SafetyDGX agent

arXiv:2606.08103v1 Announce Type: cross Abstract: We are surrounded by various objects with movable, articulated parts, e.g., box, handle, door. An accurate and generalizable perception of articulated

RGB-S: Image-Aligned Tactile Saliency for Robust Dexterous Manipulation

TutorialsDGX agent

arXiv:2606.08765v1 Announce Type: cross Abstract: Effective visuo-tactile integration is critical for robotic dexterous manipulation, especially when visual observations are unreliable or occluded. Ho

RT-SDGOD: Real-Time Single-Domain Generalized Object Detection

ApplicationsDGX agent

arXiv:2606.09367v1 Announce Type: new Abstract: In real-world deployment under strict real-time constraints, weather and imaging variations induce significant distribution shifts, severely degrading d

Scaling by Diversified Experience for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.09009v1 Announce Type: new Abstract: Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of high-level reasoning with low-level contro

SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

Model ReleasesDGX agent

arXiv:2602.09809v2 Announce Type: replace Abstract: Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorr

SDTrack: A Baseline for Event-based Tracking via Spiking Neural Networks

ResearchDGX agent

arXiv:2503.08703v4 Announce Type: replace-cross Abstract: Event cameras provide superior temporal resolution, dynamic range, energy efficiency, and pixel bandwidth. Spiking Neural Networks (SNNs) natu

Securing Self-supervised Data Curation for Foundation Models Robustness

ResearchDGX agent

arXiv:2606.09511v1 Announce Type: new Abstract: Self-supervised data curation provides a pathway to scaling and improving the generalization capabilities of machine learning models. By leveraging self

← Previous
1…8990919293…211
Next →