AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
5 Jun 2026

Deep Learning-based 3D Oral Cavity Reconstruction Using 2D Intraoral Images

ResearchDGX agent

arXiv:2606.05998v1 Announce Type: new Abstract: Oral 3D modelling is one of the most essential stages in dentistry, and many different approaches, such as impression taking and intraoral scanning, are

Diff-CA: Separating Common and Salient Factors with Diffusion Models

TutorialsDGX agent

arXiv:2606.06120v1 Announce Type: new Abstract: Contrastive Analysis aims to separate factors that are common between two data distributions from those that are salient to only one of them. Existing c

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments

Model ReleasesDGX agent

arXiv:2606.06217v1 Announce Type: new Abstract: When a disaster unfolds, responders must answer not only what is happening, but also why it is happening, what will happen next, and what to do now, oft


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Disentangled Fine-Grained Prototype Learning for Incomplete Image-Tabular Classification

SafetyDGX agent

arXiv:2606.05455v1 Announce Type: new Abstract: The missing-modality problem poses a significant challenge in image-tabular multimodal learning across a wide range of multimedia applications, includin

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

SafetyDGX agent

arXiv:2606.05290v1 Announce Type: new Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring ret

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models

ResearchDGX agent

arXiv:2606.05758v1 Announce Type: new Abstract: Many modern vision-language models (VLMs) build on autoregressive decoding of discrete tokens. While text-based output interfaces enable scalable pretra

Drishti AI-Event Guardian: An Intelligent Real-Time Crowd Monitoring and Emergency Response System for Mass Gathering Events

SafetyDGX agent

arXiv:2606.05185v1 Announce Type: cross Abstract: Mass gathering events are associated with critical safety incidents caused by insufficient crowd monitoring and inadequate emergency response coordina

Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving

Model ReleasesDGX agent

arXiv:2601.21288v2 Announce Type: replace-cross Abstract: Autonomous driving is an important and safety-critical task, and recent advances in LLMs/VLMs have opened new possibilities for reasoning and

Dual Feature Decoupling for Fine-Grained OOD Detection

ApplicationsDGX agent

arXiv:2606.05536v1 Announce Type: new Abstract: Out-of-distribution detection (OOD) is an indispensable technique when applying machine learning models to real-world scenarios. Most existing OOD detec

EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

Local AiDGX agent

arXiv:2606.06379v1 Announce Type: new Abstract: Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and report generatio

Efficient Mean Curvature Computation on High-Dimensional Data Manifolds

ApplicationsDGX agent

arXiv:2606.06329v1 Announce Type: cross Abstract: Estimating local mean curvature at each point of a high-dimensional dataset is a key ingredient of geometry-aware machine learning algorithms, such as

EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026

Local AiDGX agent

arXiv:2605.24496v2 Announce Type: replace Abstract: The EPIC-KITCHENS-100 Action Detection challenge evaluates whether a model can localize the start and end of each action in long untrimmed egocentri

EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge

Model ReleasesDGX agent

arXiv:2605.24500v2 Announce Type: replace Abstract: This technical report presents our solution, EgoAdapt (Egocentric Adaptation via Category, Calibration, and Consistency), to the CVPR 2026 HD-EPIC V

Emotion-Aware Image Generation from Korean Diary Text via LLM-based Prompt Translation and LoRA Fine-Tuning

ResearchDGX agent

arXiv:2606.05816v1 Announce Type: new Abstract: T2I models cannot effectively capture sentiment from various types of text, including diaries, as they primarily focus on visual object-related patterns

Entropy-Based Evaluation of AI Agents: A Lightweight Framework for Measuring Behavioral Patterns

AgentsDGX agent

arXiv:2606.05872v1 Announce Type: cross Abstract: AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of age

Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning

ApplicationsDGX agent

arXiv:2512.15153v2 Announce Type: replace Abstract: Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challengi

ExpSpeech-Net: Multimodal Fusion of Expression and Speech for Deepfake Detection

ResearchDGX agent

arXiv:2606.05760v1 Announce Type: new Abstract: Deepfake videos are increasingly challenging the credibility of online content. Many existing detection methodology relies on complex, resource-intensiv

Facial-R1: Aligning Reasoning and Recognition for Facial Emotion Analysis

Model ReleasesDGX agent

arXiv:2511.10254v2 Announce Type: replace Abstract: Facial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrat

Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models

Model ReleasesDGX agent

arXiv:2606.05949v1 Announce Type: new Abstract: Scientific illustrations are essential tools for communicating research findings, especially in natural science, where they visualize complex concepts a

FATE: Focal-modulated Attention Encoder for Multivariate Time-series Forecasting

Model ReleasesDGX agent

arXiv:2408.11336v3 Announce Type: replace-cross Abstract: Climate change stands as one of the most pressing global challenges of the twenty-first century, with far-reaching consequences such as rising

Flash-WAM: Modality-Aware Distillation for World Action Models

HardwareDGX agent

arXiv:2606.05254v1 Announce Type: cross Abstract: World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation b

FontFusion: Enhancing Generative Text in Diffusion Models with Typographic Conditioning

ResearchDGX agent

arXiv:2606.06066v1 Announce Type: new Abstract: Typography generation in diffusion models faces a persistent trade-off: enabling precise font control typically degrades text legibility, while maintain

Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based Learning

TutorialsDGX agent

arXiv:2606.05471v1 Announce Type: new Abstract: Learning semantics is essential for deep learning models to be interpretable and better aligned with human reasoning. Concept-based models approach this

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery

Model ReleasesDGX agent

arXiv:2602.19190v4 Announce Type: replace Abstract: Research on the intelligent interpretation of all-weather, all-time Synthetic Aperture Radar (SAR) is crucial for advancing remote sensing applicati

Gender Artifacts from Art History to Text-to-Image Generation

ResearchDGX agent

arXiv:2606.05829v1 Announce Type: new Abstract: Artistic styles are rooted in specific socio-historical contexts that encode social hierarchies, including distinct constructions of gender. Yet in AI r

GenTract: Generative Global Tractography

Local AiDGX agent

arXiv:2511.13183v2 Announce Type: replace Abstract: Tractography is the process of inferring the trajectories of white-matter pathways in the brain from diffusion magnetic resonance imaging (dMRI). Lo

Geodesic Flow Matching on a Riemannian Degradation Manifold for Blind Image Restoration

TutorialsDGX agent

arXiv:2606.06278v1 Announce Type: new Abstract: Blind image restoration requires recovering clean images from observations corrupted by unknown and potentially mixed degradations. While recent determi

Geometry-Aware Dataset Condensation for Diffusion Model Training

SafetyDGX agent

arXiv:2606.05883v1 Announce Type: new Abstract: Dataset condensation aims to construct compact datasets from real data via synthesis or selection. However, existing approaches are ill-suited for diffu

Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework

Model ReleasesDGX agent

arXiv:2603.08491v2 Announce Type: replace Abstract: Cross-modal Geo-localization (CMGL) matches ground-level text descriptions with geo-tagged aerial imagery, which is crucial for pedestrian navigatio

Global-Local Monte Carlo Tree Search in Vision-Language Models for Text-to-3D Indoor Scene Generation

ResearchDGX agent

arXiv:2606.06002v1 Announce Type: new Abstract: Large Vision-Language Models have achieved significant reasoning performance in various tasks.However, there are few studies on text-to-3D indoor scene

GMBFormer: An NDVI-Guided Global Memory Bank Transformer for Urban Green-Space Extraction from Ultra-High-Resolution Imagery

ResearchDGX agent

arXiv:2606.06363v1 Announce Type: new Abstract: Urban green-space extraction from ultra-high-resolution (UHR) imagery is commonly performed patch by patch, which limits semantic reuse among spatially

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention

ResearchDGX agent

arXiv:2606.06249v1 Announce Type: new Abstract: Transformer-based multimodal models rely on attention mechanisms to integrate information across heterogeneous modalities. Despite their success, existi

GS-NFS: Bandwidth-adaptive Streaming of Dynamic Gaussian Splats and Point Clouds

HardwareDGX agent

arXiv:2606.05650v1 Announce Type: cross Abstract: Dynamic 3D Gaussian Splatting (3DGS) holds great promise as a 3D video streaming technology since it can represent complex 3D scenes with high fidelit

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

Model ReleasesDGX agent

arXiv:2511.20158v2 Announce Type: replace Abstract: While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominan

HDST-GNN: Heterogeneous Dynamic Spatiotemporal Graph Neural Networks for Multi-Object Tracking in UAV Aerial Imagery

ResearchDGX agent

arXiv:2606.05587v1 Announce Type: new Abstract: Multi-object tracking (MOT) from UAV imagery presents unique challenges: altitude varies across sequences, objects are small and densely packed, and fre

HERO: Learning Humanoid End-Effector Control for Visual Whole-Body Open-Vocabulary Object Grasping

SafetyDGX agent

arXiv:2602.16705v3 Announce Type: replace-cross Abstract: Visual loco-manipulation of arbitrary in-the-wild objects requires accurate end-effector (EE) control and a generalizable understanding of the

Hierarchical Mask-Enhanced Dual Reconstruction Network for Few-Shot Fine-Grained Image Classification

ResearchDGX agent

arXiv:2506.20263v2 Announce Type: replace Abstract: Few-shot fine-grained image classification (FS-FGIC) is challenging as it requires distinguishing visually similar subclasses with extremely limited

HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps

Model ReleasesDGX agent

arXiv:2601.02730v3 Announce Type: replace Abstract: Visual localization on standard-definition (SD) maps has emerged as a promising low-cost and scalable solution for autonomous driving. However, exis

HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes

ResearchDGX agent

arXiv:2606.06390v1 Announce Type: new Abstract: Indoor scene generation is crucial for robot simulation and modern interior design. However, complex layouts together with scarce 3D scene data make lea

Horse Eye Blink Detection and Classification for Equine Affective State Assessment

ResearchDGX agent

arXiv:2606.05458v1 Announce Type: new Abstract: Automated detection of equine facial action units (AUs) is a promising yet under-explored avenue for pain and affective state assessment in horses. Half

HyperVis: Continuous Latent Visual Relational Graphs on the Lorentz Hyperboloid for Compositional Reasoning

ResearchDGX agent

arXiv:2606.06100v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle with compositional reasoning that requires understanding inter-object relationships. A natural remedy is to injec

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

ResearchDGX agent

arXiv:2606.05769v1 Announce Type: new Abstract: Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize inter

In-Context Multiple Instance Learning

ApplicationsDGX agent

arXiv:2606.06458v1 Announce Type: cross Abstract: Multiple Instance Learning (MIL) addresses problems where supervision is available at the level of bags of instances and has been successfully applied

Inverse Design of Realizable Metasurface based Absorbers using Improved Conditioning and Diversity Enhanced Progressively Growing GANs

SafetyDGX agent

arXiv:2606.05849v1 Announce Type: cross Abstract: Metasurfaces enable precise manipulation of electromagnetic waves for applications such as beam steering, sensing, and stealth technology. However, in

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing

Model ReleasesDGX agent

arXiv:2606.05172v1 Announce Type: cross Abstract: Diffusion-based image editing has achieved strong visual fidelity under natural language instructions, yet most existing systems still operate at the

Knowledge Distillation for Visual Autoregressive Models

ResearchDGX agent

arXiv:2606.06078v1 Announce Type: new Abstract: Autoregressive (AR) image generation models are highly expressive but computationally intensive, motivating effective model compression. Knowledge disti

KV-Control: Parameter-Efficient K/V Injection for Trajectory-Controlled Text-to-Motion

Model ReleasesDGX agent

arXiv:2606.05624v1 Announce Type: new Abstract: Text-conditioned 3D human motion models now synthesize plausible motions from prompts, but practical animation and embodied-agent workflows rarely stop

LadderMan: Learning Humanoid Perceptive Ladder Climbing

SafetyDGX agent

arXiv:2606.05873v1 Announce Type: cross Abstract: Humanoid robots hold great promise for operating in human-centered environments, yet ladder climbing remains one of the most challenging tasks due to

Latent Implicit Visual Reasoning

ResearchDGX agent

arXiv:2512.21218v2 Announce Type: replace Abstract: While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning m

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

ApplicationsDGX agent

arXiv:2606.05833v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to m

Learning Predictive Visuomotor Coordination

ApplicationsDGX agent

arXiv:2503.23300v2 Announce Type: replace Abstract: Understanding and predicting human visuomotor coordination is crucial for applications in robotics, human-computer interaction, and assistive techno

Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation

SafetyDGX agent

arXiv:2606.06076v1 Announce Type: cross Abstract: While vision-language models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this to a perce

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

SafetyDGX agent

arXiv:2606.05737v1 Announce Type: new Abstract: Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising. We argue that

LiAuto-GeoX: Efficient Grounded Driving Transformer

Model ReleasesDGX agent

arXiv:2606.05774v1 Announce Type: new Abstract: Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for auton

LightVesselNet: An Ultra-Lightweight Sub-100K Parameter Network for Retinal Blood Vessel Segmentation

Model ReleasesDGX agent

arXiv:2606.05354v1 Announce Type: new Abstract: Retinal blood vessel segmentation plays a vital role in the early detection of diabetic retinopathy and glaucoma. While recent deep learning models have

LLM-Conditioned Synthesis of Pathological Gaits via Structured Gait-Language Representations

ResearchDGX agent

arXiv:2606.06048v1 Announce Type: new Abstract: Pathological gait datasets remain scarce due to privacy, recruitment, cost, and movement variability. Our work presents a multimodal LLM-guided framewor

LLM-Guided ANN Index Optimization for Human-Object Interaction Retrieval

Model ReleasesDGX agent

arXiv:2606.05489v1 Announce Type: new Abstract: Retrieval systems underpin modern AI applications -- spanning visual search, recommendation engines, and multi-modal question answering. Modern multi-st

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

Model ReleasesDGX agent

arXiv:2606.06042v1 Announce Type: new Abstract: Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier fie

MAviS: A Multimodal Conversational Assistant For Avian Species

Model ReleasesDGX agent

arXiv:2603.07294v2 Announce Type: replace Abstract: Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monit

Monte Carlo Steklov Operators for Large-Scale Geometry Processing in the Wild

ResearchDGX agent

arXiv:2606.05581v1 Announce Type: cross Abstract: Intrinsic methods fill the default toolbox for geometry processing on meshes. Intrinsic operators, in particular the Laplacian, underlie methods that

← Previous
1…9394959697…211
Next →