AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing

DGX agent

arXiv:2606.07171v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) enable flexible instruction-driven image editing, but privacy risks arise when user images expose diverse and u

model-releasesarxiv-cs-cv
8 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

DGX agent

arXiv:2606.06538v1 Announce Type: new Abstract: In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types

model-releasesarxiv-cs-cv
8 Jun 2026
Research

Zero-Shot Polygon Matching with Pre-trained Models for Pose Estimation and Polygon Cloud from Challenging Stereo

DGX agent

arXiv:2511.05949v2 Announce Type: replace Abstract: While stereo matching has achieved maturity for 0D point and 1D line primitives, establishing correspondences for 2D polygons remains largely unexpl

researcharxiv-cs-cv
8 Jun 2026
Research

A Vision-language Framework for Comparative Reasoning in Radiology

DGX agent

arXiv:2606.06407v1 Announce Type: new Abstract: Medical imaging artificial intelligence has achieved strong performance in isolated image interpretation, but remains poorly aligned with radiological p

researcharxiv-cs-cv
5 Jun 2026
Applications

ActiveMimic: Egocentric Video Pretraining with Active Perception

DGX agent

arXiv:2606.06194v1 Announce Type: cross Abstract: Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform thos

applicationsarxiv-cs-cv
5 Jun 2026
Model Releases

Adaptive Tokenisation Via Temporal Redundancy Masking And Latent Inpainting

DGX agent

arXiv:2606.06158v1 Announce Type: new Abstract: Adaptive video tokenisation seeks to dynamically allocate token budgets based on the underlying visual complexity of a sequence. Current continuous-regi

model-releasesarxiv-cs-cv
5 Jun 2026
Safety

Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models

DGX agent

arXiv:2606.06186v1 Announce Type: new Abstract: Vision-Language Models (VLMs), such as CLIP, have shown strong zero-shot generalization but remain highly vulnerable to adversarial perturbations, posin

safetyarxiv-cs-cv
5 Jun 2026
Local Ai

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

DGX agent

arXiv:2606.06155v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robo

local-aiarxiv-cs-cv
5 Jun 2026
Applications

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue

DGX agent

arXiv:2509.15061v2 Announce Type: cross Abstract: The ultimate goal of embodied agents is to create collaborators that can interact with humans, not mere executors that passively follow instructions.

applicationsarxiv-cs-cv
5 Jun 2026
Research

ATT-CR: Adaptive Triangular Transformer for Cloud Removal

DGX agent

arXiv:2606.05999v1 Announce Type: new Abstract: Cloud removal aims to accurately reconstruct the ground objects obscured by clouds in remote sensing images. Existing Transformer-based methods utilizin

researcharxiv-cs-cv
5 Jun 2026
Research

Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment

DGX agent

arXiv:2606.05778v1 Announce Type: new Abstract: Traditional Image Aesthetic Assessment (IAA) methods mainly rely on regressing absolute Mean Opinion Scores (MOS). However, such a paradigm overlooks th

researcharxiv-cs-cv
5 Jun 2026
Model Releases

Biomazon: A Multimodal Dataset for 3D Forest Structure and Biomass Modeling in the Amazon Basin

DGX agent

arXiv:2606.05368v1 Announce Type: new Abstract: Accurate, spatially explicit characterization of tropical forest structure is essential for carbon accounting and ecosystem monitoring, yet most ML pipe

model-releasesarxiv-cs-cv
5 Jun 2026
Local Ai

BMCR: Adaptive Backbone Module Composition via Reinforcement Learning for Remote Sensing Object Detection

DGX agent

arXiv:2606.05586v1 Announce Type: new Abstract: In remote sensing object detection, Convolutional Neural Networks (CNNs) excel at capturing local details while Vision Transformers (ViTs) are better at

local-aiarxiv-cs-cv
5 Jun 2026
Research

BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain

DGX agent

arXiv:2512.08560v3 Announce Type: replace Abstract: Understanding how the human brain represents visual concepts, and in which brain regions these representations are encoded, remains a long-standing

researcharxiv-cs-cv
5 Jun 2026
Research

BRepCLIP: Contrastive Multimodal Pretraining on BRep Primitives for CAD Understanding

DGX agent

arXiv:2606.05515v1 Announce Type: new Abstract: Learning representations of CAD models is a largely open problem. While 3D representation learning has flourished around point clouds and meshes, the na

researcharxiv-cs-cv
5 Jun 2026
Model Releases

CamFlow+: Hybrid Motion Bases for 2D Camera Motion Estimation with Stabilization Applications

DGX agent

arXiv:2606.05915v1 Announce Type: new Abstract: Estimating 2D camera motion is fundamental to computer vision and computational photography. Existing homography-based methods work well for planar scen

model-releasesarxiv-cs-cv
5 Jun 2026
Tutorials

Can Language Models Learn to Listen?

DGX agent

arXiv:2308.10897v2 Announce Type: replace Abstract: We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given

tutorialsarxiv-cs-cv
5 Jun 2026
Research

Can We Predict The Human Preference For Text-to-Image Content Prior To Generation And Is It Even Useful To Do So?

DGX agent

arXiv:2606.05478v1 Announce Type: new Abstract: Diffusion Models (DM) have revolutionized text-driven generation by enabling the synthesis of high-quality, photorealistic visual content from user prom

researcharxiv-cs-cv
5 Jun 2026
Tutorials

CoFi-UCGen: Coarse-to-Fine Unsupervised Conditional Generation without Label Priors

DGX agent

arXiv:2606.05652v1 Announce Type: new Abstract: Unsupervised conditional image generation (UCGen) aims to control generation without relying on manually annotated labels, yet remains challenging due t

tutorialsarxiv-cs-cv
5 Jun 2026
Research

Comparison of Deep Learning Frameworks For Rice Disease Mapping From UAV Multispectral Imaging

DGX agent

arXiv:2606.06359v1 Announce Type: new Abstract: In this study, UAV multispectral imagery is used to segment the severity of bacterial leaf blight (BLB) in rice using convolutional neural networks (CNN

researcharxiv-cs-cv
5 Jun 2026
Local Ai

Complexity-Balanced Diffusion Splitting

DGX agent

arXiv:2606.06477v1 Announce Type: new Abstract: Standard continuous-time generative models rely on monolithic architectures that must navigate vastly different signal regimes, from isotropic noise to

local-aiarxiv-cs-cv
5 Jun 2026
Research

Computation-Aware Event-to-Frame Reconstruction via Selective Attention

DGX agent

arXiv:2606.06142v1 Announce Type: new Abstract: Event-to-frame (E2F) reconstruction bridges asynchronous event streams with frame-based vision pipelines, but existing methods often face a trade-off be

researcharxiv-cs-cv
5 Jun 2026
Safety

Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latents

DGX agent

arXiv:2606.05753v1 Announce Type: new Abstract: Latent visual reasoning (LVR) inserts supervised latent tokens between perception and answer generation in vision-language models (VLMs). The field uses

safetyarxiv-cs-cv
5 Jun 2026
Research

Deep Learning-assisted AMD Staging based on OCT and OCT Angiography

DGX agent

arXiv:2606.05379v1 Announce Type: new Abstract: To develop and evaluate deep learning models for automated grading of age-related macular degeneration (AMD) severity using optical coherence tomography

researcharxiv-cs-cv
5 Jun 2026
Research

Deep Learning-based 3D Oral Cavity Reconstruction Using 2D Intraoral Images

DGX agent

arXiv:2606.05998v1 Announce Type: new Abstract: Oral 3D modelling is one of the most essential stages in dentistry, and many different approaches, such as impression taking and intraoral scanning, are

researcharxiv-cs-cv
5 Jun 2026
Tutorials

Diff-CA: Separating Common and Salient Factors with Diffusion Models

DGX agent

arXiv:2606.06120v1 Announce Type: new Abstract: Contrastive Analysis aims to separate factors that are common between two data distributions from those that are salient to only one of them. Existing c

tutorialsarxiv-cs-cv
5 Jun 2026
Model Releases

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments

DGX agent

arXiv:2606.06217v1 Announce Type: new Abstract: When a disaster unfolds, responders must answer not only what is happening, but also why it is happening, what will happen next, and what to do now, oft

model-releasesarxiv-cs-cv
5 Jun 2026
Safety

Disentangled Fine-Grained Prototype Learning for Incomplete Image-Tabular Classification

DGX agent

arXiv:2606.05455v1 Announce Type: new Abstract: The missing-modality problem poses a significant challenge in image-tabular multimodal learning across a wide range of multimedia applications, includin

safetyarxiv-cs-cv
5 Jun 2026
Safety

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

DGX agent

arXiv:2606.05290v1 Announce Type: new Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring ret

safetyarxiv-cs-cv
5 Jun 2026
Research

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models

DGX agent

arXiv:2606.05758v1 Announce Type: new Abstract: Many modern vision-language models (VLMs) build on autoregressive decoding of discrete tokens. While text-based output interfaces enable scalable pretra

researcharxiv-cs-cv
5 Jun 2026
Safety

Drishti AI-Event Guardian: An Intelligent Real-Time Crowd Monitoring and Emergency Response System for Mass Gathering Events

DGX agent

arXiv:2606.05185v1 Announce Type: cross Abstract: Mass gathering events are associated with critical safety incidents caused by insufficient crowd monitoring and inadequate emergency response coordina

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving

DGX agent

arXiv:2601.21288v2 Announce Type: replace-cross Abstract: Autonomous driving is an important and safety-critical task, and recent advances in LLMs/VLMs have opened new possibilities for reasoning and

model-releasesarxiv-cs-cv
5 Jun 2026
Applications

Dual Feature Decoupling for Fine-Grained OOD Detection

DGX agent

arXiv:2606.05536v1 Announce Type: new Abstract: Out-of-distribution detection (OOD) is an indispensable technique when applying machine learning models to real-world scenarios. Most existing OOD detec

applicationsarxiv-cs-cv
5 Jun 2026
Local Ai

EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

DGX agent

arXiv:2606.06379v1 Announce Type: new Abstract: Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and report generatio

local-aiarxiv-cs-cv
5 Jun 2026
Applications

Efficient Mean Curvature Computation on High-Dimensional Data Manifolds

DGX agent

arXiv:2606.06329v1 Announce Type: cross Abstract: Estimating local mean curvature at each point of a high-dimensional dataset is a key ingredient of geometry-aware machine learning algorithms, such as

applicationsarxiv-cs-cv
5 Jun 2026
Local Ai

EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026

DGX agent

arXiv:2605.24496v2 Announce Type: replace Abstract: The EPIC-KITCHENS-100 Action Detection challenge evaluates whether a model can localize the start and end of each action in long untrimmed egocentri

local-aiarxiv-cs-cv
5 Jun 2026
Model Releases

EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge

DGX agent

arXiv:2605.24500v2 Announce Type: replace Abstract: This technical report presents our solution, EgoAdapt (Egocentric Adaptation via Category, Calibration, and Consistency), to the CVPR 2026 HD-EPIC V

model-releasesarxiv-cs-cv
5 Jun 2026
Research

Emotion-Aware Image Generation from Korean Diary Text via LLM-based Prompt Translation and LoRA Fine-Tuning

DGX agent

arXiv:2606.05816v1 Announce Type: new Abstract: T2I models cannot effectively capture sentiment from various types of text, including diaries, as they primarily focus on visual object-related patterns

researcharxiv-cs-cv
5 Jun 2026
Agents

Entropy-Based Evaluation of AI Agents: A Lightweight Framework for Measuring Behavioral Patterns

DGX agent

arXiv:2606.05872v1 Announce Type: cross Abstract: AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of age

agentsarxiv-cs-cv
5 Jun 2026
Applications

Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning

DGX agent

arXiv:2512.15153v2 Announce Type: replace Abstract: Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challengi

applicationsarxiv-cs-cv
5 Jun 2026
Research

ExpSpeech-Net: Multimodal Fusion of Expression and Speech for Deepfake Detection

DGX agent

arXiv:2606.05760v1 Announce Type: new Abstract: Deepfake videos are increasingly challenging the credibility of online content. Many existing detection methodology relies on complex, resource-intensiv

researcharxiv-cs-cv
5 Jun 2026
Model Releases

Facial-R1: Aligning Reasoning and Recognition for Facial Emotion Analysis

DGX agent

arXiv:2511.10254v2 Announce Type: replace Abstract: Facial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrat

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models

DGX agent

arXiv:2606.05949v1 Announce Type: new Abstract: Scientific illustrations are essential tools for communicating research findings, especially in natural science, where they visualize complex concepts a

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

FATE: Focal-modulated Attention Encoder for Multivariate Time-series Forecasting

DGX agent

arXiv:2408.11336v3 Announce Type: replace-cross Abstract: Climate change stands as one of the most pressing global challenges of the twenty-first century, with far-reaching consequences such as rising

model-releasesarxiv-cs-cv
5 Jun 2026
Hardware

Flash-WAM: Modality-Aware Distillation for World Action Models

DGX agent

arXiv:2606.05254v1 Announce Type: cross Abstract: World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation b

hardwarearxiv-cs-cv
5 Jun 2026
Research

FontFusion: Enhancing Generative Text in Diffusion Models with Typographic Conditioning

DGX agent

arXiv:2606.06066v1 Announce Type: new Abstract: Typography generation in diffusion models faces a persistent trade-off: enabling precise font control typically degrades text legibility, while maintain

researcharxiv-cs-cv
5 Jun 2026
Tutorials

Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based Learning

DGX agent

arXiv:2606.05471v1 Announce Type: new Abstract: Learning semantics is essential for deep learning models to be interpretable and better aligned with human reasoning. Concept-based models approach this

tutorialsarxiv-cs-cv
5 Jun 2026
Model Releases

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery

DGX agent

arXiv:2602.19190v4 Announce Type: replace Abstract: Research on the intelligent interpretation of all-weather, all-time Synthetic Aperture Radar (SAR) is crucial for advancing remote sensing applicati

model-releasesarxiv-cs-cv
5 Jun 2026
← Previous
1…116117118119120…263
Next →