AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
21 May 2026

WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark from Wikipedia and Wikidata

Model ReleasesDGX agent

arXiv:2605.21479v1 Announce Type: new Abstract: Visual Question Answering (VQA) benchmarks have largely emphasized perception-based tasks that can be solved from visual content alone. In contrast, man

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

Model ReleasesDGX agent

arXiv:2605.20306v1 Announce Type: new Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous

Winfree Oscillatory Neural Network

Model ReleasesDGX agent

arXiv:2605.20922v1 Announce Type: cross Abstract: Oscillations and synchronization are widely believed to play a fundamental role in representation and computation. However, existing machine learning


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

You Don't Need Attention: Gated Convolutional Modeling for Watch-Based Fall Detection

Local AiDGX agent

arXiv:2605.20275v1 Announce Type: new Abstract: Existing deep learning approaches for wearable fall detection systems rely on self-attention mechanisms that impose quadratic computational overhead, di

20 May 2026

3D Modeling and Automated Measurement of Concrete Cracks via Segment Anything Refinement and Visual Inertial LiDAR Fusion

Local AiDGX agent

arXiv:2501.09203v2 Announce Type: replace Abstract: Visual-Spatial Systems has become increasingly essential in concrete crack inspection. However, existing methods often lacks adaptability to diverse

3DMambaComplete: Exploring Structured State Space Model for Point Cloud Completion

ResearchDGX agent

arXiv:2404.07106v2 Announce Type: replace Abstract: Point cloud completion aims to generate a complete and high-fidelity point cloud from an initially incomplete and low-quality input. A prevalent str

A Grid-Based Framework for E-Scooter Demand Representation and Temporal Input Design for Deep Learning: Evidence from Austin, Texas

ResearchDGX agent

arXiv:2603.13609v2 Announce Type: replace Abstract: Despite progress in deep learning for shared micromobility demand prediction, the systematic design and statistical validation of temporal input str

A Multi-Dimensional Clustering Approach for Identifying Inborn Errors of Immunity

TutorialsDGX agent

arXiv:2605.18880v1 Announce Type: cross Abstract: Rare diseases such as inborn errors of immunity (IEI) require early diagnosis to prevent end organ damage and improve quality of life. Hurdles in acce

A Nash Equilibrium Framework For Training-Free Multimodal Step Verification

ResearchDGX agent

arXiv:2605.20033v1 Announce Type: new Abstract: Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approache

A Systematic Failure Analysis of Vision Foundation Models for Open Set Iris Presentation Attack Detection

Model ReleasesDGX agent

arXiv:2605.19020v1 Announce Type: new Abstract: Vision foundation models have demonstrated strong transferability across diverse visual recognition tasks and are increasingly considered for biometric

Adapted Center and Scale Prediction: More Stable and More Accurate

Model ReleasesDGX agent

arXiv:2002.09053v3 Announce Type: replace Abstract: Pedestrian detection benefits from deep learning technology and gains rapid development in recent years. Most of detectors follow general object det

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls

Model ReleasesDGX agent

arXiv:2605.19728v1 Announce Type: new Abstract: Foundation video models produce visually impressive results, but their use in embodied AI remains limited because they are primarily trained on natural

AffectVerse: Emotional World Models for Multimodal Affective Computing

ResearchDGX agent

arXiv:2605.19950v1 Announce Type: new Abstract: Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large languag

An Automated Framework for Large-Scale Graph-Based Cerebrovascular Analysis

ApplicationsDGX agent

arXiv:2512.03869v4 Announce Type: replace Abstract: We present CaravelMetrics, a computational framework for automated cerebrovascular analysis that models vessel morphology through skeletonization-de

AnchorFlow: Editable SVG Reconstruction via Sparse Anchor Point Fields

ResearchDGX agent

arXiv:2605.19551v1 Announce Type: cross Abstract: Image-to-SVG reconstruction aims to produce vector graphics that are faithful to raster inputs and easy to edit. Existing methods face a structural tr

Are Watermarked Images Editable? SafeMark for Watermark-Preserving Text-Guided Image Editing

ResearchDGX agent

arXiv:2605.19511v1 Announce Type: new Abstract: This paper investigates a fundamental yet underexplored question: can watermarked images remain editable without compromising watermark integrity? We pr

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

Model ReleasesDGX agent

arXiv:2605.18984v1 Announce Type: new Abstract: Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inco

BabyMamba-HAR: Lightweight Selective State Space Models for Efficient Human Activity Recognition on Resource Constrained Devices

Local AiDGX agent

arXiv:2602.09872v2 Announce Type: replace Abstract: Human activity recognition (HAR) on resource constrained devices requires high accuracy across diverse sensor setups. Selective state space models (

Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

Model ReleasesDGX agent

arXiv:2605.19639v1 Announce Type: new Abstract: Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a sin

Beyond Binary Success: A Diagnostic Meta-Evaluation Framework for Fine-Grained Manipulation

Model ReleasesDGX agent

arXiv:2605.19986v1 Announce Type: cross Abstract: Fine-grained manipulation marks a regime where global scene context no longer suffices, and success hinges on the tight coupling of local attribute gr

Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives

Model ReleasesDGX agent

arXiv:2605.19771v1 Announce Type: cross Abstract: Existing imitation learning methods for end-to-end autonomous driving predominantly learn from successful demonstrations by minimizing geometric devia

Bezier Degradation Modeling for LiDAR-based Human Motion Capture

AgentsDGX agent

arXiv:2605.19620v1 Announce Type: new Abstract: LiDAR-based 3D human motion capture has broad applications in fields such as autonomous driving and robotics, where accurate motion reconstruction is cr

Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection

SafetyDGX agent

arXiv:2605.19532v1 Announce Type: new Abstract: Text-to-image diffusion models can synthesize high-quality images, yet the outcome is notoriously sensitive to the random seed: different initial seeds

CAD-Free Learning of Spacecraft Pose Estimators via NeRF-Based Augmentations

ResearchDGX agent

arXiv:2605.19649v1 Announce Type: new Abstract: Spacecraft pose estimation networks require tens of thousands of CAD-rendered images to be trained. This reliance on synthetic CAD data (i) limits appli

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models

ResearchDGX agent

arXiv:2605.20165v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong performance on spatial question answering benchmarks, yet it remains unclear whether such gains reflect gen

Cardiac fat segmentation using computed tomography and an image-to-image conditional generative adversarial neural network

AgentsDGX agent

arXiv:2605.20064v1 Announce Type: new Abstract: In recent years, research has highlighted the association between increased adipose tissue surrounding the human heart and elevated susceptibility to ca

Character-Centered Dialogue Generation from Scene-Level Prompts

TutorialsDGX agent

arXiv:2505.16819v4 Announce Type: replace Abstract: Recent advances in scene-based video generation enable coherent visual narratives from structured prompts, yet a key aspect of storytelling -- chara

Closed-Loop Hybrid Digital Twin Platform for Connected and Automated Vehicle Validation

ApplicationsDGX agent

arXiv:2605.19490v1 Announce Type: cross Abstract: Comprehensive and efficient validation of connected and automated vehicles (CAVs) is critical prior to real-world deployment. While simulation-based t

CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition

ApplicationsDGX agent

arXiv:2605.19995v1 Announce Type: new Abstract: Recent diffusion models achieve strong photorealism and fluency in video generation, yet remain fragile under abstract, sparse or complex conditions, le

Contextualized Visual Personalization in Vision-Language Models

ResearchDGX agent

arXiv:2602.03454v3 Announce Type: replace Abstract: Despite recent progress in vision-language models (VLMs), existing approaches often fail to generate personalized responses based on the user's spec

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models

ResearchDGX agent

arXiv:2605.19750v1 Announce Type: new Abstract: Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation. Despite their strong generative capabili

Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

Model ReleasesDGX agent

arXiv:2605.19656v1 Announce Type: new Abstract: We present Cross-View Splatter, a feed-forward method that predicts pixel-aligned Gaussian splats for outdoor scenes captured at ground level AND by sat

D-Convexity: A Unified Differentiable Convex Shape Prior via Quasi-Concavity for Data-driven Image Segmentation

ResearchDGX agent

arXiv:2605.19210v1 Announce Type: new Abstract: Convexity is a fundamental geometric prior that underlies many natural and man-made structures, yet remains challenging to impose effectively in end-to-

deadtrees.earth-aerial: A Multi-Resolution Aerial Image Dataset for Tree Cover and Mortality Detection

Model ReleasesDGX agent

arXiv:2605.19605v1 Announce Type: new Abstract: Forests worldwide are increasingly threatened by climate change and disturbances such as fire, pests, and pathogens, creating an urgent need for scalabl

Decentralized Direct Volume Rendering: A Browser-Native GPU Architecture for MRI Digital Twins in Resource-Constrained Settings

HardwareDGX agent

arXiv:2605.19737v1 Announce Type: cross Abstract: Digital Twin (DT) technology holds immense potential for surgical planning and personalized medicine. However, generating interactive, patient-specifi

Delta Attention Residuals

ResearchDGX agent

arXiv:2605.18855v1 Announce Type: cross Abstract: Attention Residuals replace standard additive residual connections with learned softmax attention over previous layer outputs, enabling selective cros

Depth2Pose: A Pose-Based Benchmark for Monocular Depth Estimation without Ground-Truth Depth

Model ReleasesDGX agent

arXiv:2605.19797v1 Announce Type: new Abstract: Monocular depth estimation has improved significantly in recent years, driven by increasingly powerful models and large-scale training data. Predicted d

Differential-Integral Neural Operator for Long-Term Turbulence Forecasting

Model ReleasesDGX agent

arXiv:2509.21196v3 Announce Type: replace-cross Abstract: Accurately forecasting the long-term evolution of turbulence represents a grand challenge in scientific computing and is crucial for applicati

Distribution Matching Distillation without Fake Score Network

ResearchDGX agent

arXiv:2605.19256v1 Announce Type: new Abstract: Distribution Matching Distillation (DMD) provides an effective distribution-level correction for few-step generation, while relying on an auxiliary fake

Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis

ResearchDGX agent

arXiv:2602.03139v2 Announce Type: replace Abstract: Distribution matching distillation (DMD) facilitates few-step image generation by aligning a distilled student with a reference multi-step teacher.

DocQT: Improving Document Forgery Localization Robustness via Diverse JPEG Quantization Tables

Model ReleasesDGX agent

arXiv:2605.19688v1 Announce Type: new Abstract: Document manipulation localization models achieve strong performance on public benchmarks yet fail to generalize to operational document workflows. We i

Dual-Prompt CLIP with Hybrid Visual Encoders for Occluded Person Re-Identification

Model ReleasesDGX agent

arXiv:2605.19527v1 Announce Type: new Abstract: Occluded person re-identification focuses on matching partially visible pedestrians across multiple camera views. However, occlusions disrupt body-regio

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs

SafetyDGX agent

arXiv:2605.19322v1 Announce Type: new Abstract: Recent advances in Video Large Language Models (Video-LLMs) have greatly expanded multimodal reasoning capabilities. However, the massive number of visu

Efficient coding along the visual hierarchy

Local AiDGX agent

arXiv:2605.19155v1 Announce Type: new Abstract: Biological visual systems learn from limited experience, unlike deep learning models that rely on millions of training images. What learning principles

Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention

ResearchDGX agent

arXiv:2605.19726v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) enable globally coherent, bidirectional, and controllable text generation, offering advantages over traditional autoreg

Efficient Transferable Optimal Transport via Min-Sliced Transport Plans

SafetyDGX agent

arXiv:2511.19741v3 Announce Type: replace Abstract: Optimal Transport (OT) offers a powerful framework for finding correspondences between distributions and addressing matching and alignment problems

EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction

Model ReleasesDGX agent

arXiv:2605.19004v1 Announce Type: new Abstract: Accurately forecasting human trajectories from an egocentric perspective plays a central role in applications such as humanoid robotics, wearable sensin

EpiDiffVO: Geometry-Aware Epipolar Diffusion for Robust Visual Odometry

ResearchDGX agent

arXiv:2605.19556v1 Announce Type: new Abstract: Estimating relative pose from image pairs fundamentally requires only a minimal subset of geometrically consistent correspondences. However, most learni

EventPrune: Cascaded Event-Assisted Token Pruning for Efficient First-Person Dynamic Spatial Reasoning

Model ReleasesDGX agent

arXiv:2605.19506v1 Announce Type: new Abstract: First-person dynamic spatial reasoning requires models to track continuous motion and precise geometric structure, but the quadratic attention cost of T

Exposing Functional Fusion: A New Class of Strategic Backdoor in Dynamic Prompt Architectures

Model ReleasesDGX agent

arXiv:2605.19478v1 Announce Type: cross Abstract: Existing ViT backdoor attacks based on backbone-overwriting full-tuning are computationally expensive and inflict performance degradation. This has fo

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models

ResearchDGX agent

arXiv:2605.19859v1 Announce Type: new Abstract: Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs

Fast 4D Mesh Generation by Spatio-Temporal Attention Chains

ResearchDGX agent

arXiv:2605.19786v1 Announce Type: new Abstract: 4D mesh generation has recently emerged as a powerful paradigm for recovering dynamic 3D structure from videos, but existing methods remain slow, comput

Fast-BEV++: Fast by Algorithm, Deployable by Design

Model ReleasesDGX agent

arXiv:2512.08237v3 Announce Type: replace Abstract: The advancement of vision-only Bird's-Eye-View (BEV) perception, a core paradigm for cost-effective autonomous driving, is hindered by the long-stan

Feature-Space Smoothing: Certified Robustness of Deep Representations

SafetyDGX agent

arXiv:2601.16200v3 Announce Type: replace-cross Abstract: Modern deep learning models exhibit strong capabilities across diverse applications, yet remain vulnerable to malicious inputs that induce err

Feed-Forward Gaussian Splatting from Sparse Aerial Views

ApplicationsDGX agent

arXiv:2605.19949v1 Announce Type: new Abstract: Reconstructing large-scale urban scenes from sparse aerial views is a crucial yet challenging task. Due to biased top-down and shallow-oblique camera po

FGSVQA: Frequency-Guided Short-form Video Quality Assessment

ResearchDGX agent

arXiv:2605.20016v1 Announce Type: cross Abstract: Short-form video poses new challenges to the quality assessment of user-generated content (UGC) due to its complex generation pipeline, rapid content

FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models

SafetyDGX agent

arXiv:2605.19739v1 Announce Type: new Abstract: Recent advances in flow matching models have significantly improved text-to-image generation quality, but also introduce growing safety risks due to the

FPED: A Functional-Network Prior-Guided Mixture-of-Experts Framework for Interpretable Brain Decoding

ResearchDGX agent

arXiv:2605.19279v1 Announce Type: new Abstract: Visual image reconstruction from functional Magnetic Resonance Imaging (fMRI) is a fundamental task in brain decoding, providing a crucial pathway for u

From Division to Decision: Leveraging Temporal Cell-Stage Segmentation for Embryo Transferability Prediction

ResearchDGX agent

arXiv:2605.18923v1 Announce Type: cross Abstract: Accurate selection of bovine embryos is a challenging task, as current practice relies on a single expert assessment on the seventh day after insemina

From Llama to Cria: Scaling Down Neural Networks via Neuron-Level Spectral Structural Importance Evaluation

Model ReleasesDGX agent

arXiv:2605.18860v1 Announce Type: cross Abstract: This paper proposes a neuron pruning framework based on neuron-level spectral structural importance evaluation. Given a trained neural network, we rec

← Previous
1…123124125126127…211
Next →