AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
7 Jul 2026IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

arXiv:2511.17384v2 Announce Type: replace-cross Abstract: While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reas

→7 Jul 2026Evaluating Agentic Harness Systems for Autonomous Computational Pathology

arXiv:2607.02598v1 Announce Type: new Abstract: Autonomous computational pathology (ACP) converts high-level pathology analysis goals into executable, traceable and clinically bounded workflows. Reali

HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→
16 Jul 2026Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

arXiv:2607.10057v1 Announce Type: cross Abstract: Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circuit Vision,

→27 Jul 2026Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet

→27 Jul 2026LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models

arXiv:2606.16794v2 Announce Type: replace Abstract: This study proposes a domain-specific LLM-based Visual Explanation Evaluation Framework for assessing visual attention explanations in facial skin d

→28 Jul 2026StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

arXiv:2607.22798v1 Announce Type: cross Abstract: Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screen

→31 Jul 2026Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3

arXiv:2607.28287v1 Announce Type: cross Abstract: ARC-AGI-3 turns abstraction into an interactive problem of skill acquisition. A player must infer an unfamiliar game's rules, hidden state, and goal w

→4 Aug 2026LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

arXiv:2608.01964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interde

CompanyOpenAI8 recent entries
8 Jul 2026Abductive Corroboration of Probabilistic AI Models for Forensic Synthetic Media Detection

arXiv:2607.05434v1 Announce Type: cross Abstract: Artificial Intelligence (AI) models, at their core, apply general learnings from broad datasets to individual circumstances using probabilistic behavi

→9 Jul 2026InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models

arXiv:2607.07288v1 Announce Type: new Abstract: Infrared vision-language models are increasingly used for perception under low-light and adverse visual conditions, yet their robustness to localized st

→24 Jul 2026ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

arXiv:2607.21022v1 Announce Type: new Abstract: Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing t

→27 Jul 2026Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet

→3 Aug 2026Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

arXiv:2607.28696v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual

→4 Aug 2026XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs

arXiv:2603.28568v2 Announce Type: replace Abstract: Vision-language models (VLMs) share visual-textual representations across zero-shot classification, image captioning, and visual question answering

→4 Aug 2026Generative AI and Foundation Models in Medical Image

arXiv:2608.01686v1 Announce Type: new Abstract: In recent years, generative AI has attracted significant public attention, and its use has been rapidly expanding across a wide range of domains. From c

→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

CompanyGoogle8 recent entries
24 Jul 2026Do Pathology Vision-Language Models Truly See Pathology?

arXiv:2607.21065v1 Announce Type: new Abstract: Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. Howe

→27 Jul 2026Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet

→27 Jul 2026LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models

arXiv:2606.16794v2 Announce Type: replace Abstract: This study proposes a domain-specific LLM-based Visual Explanation Evaluation Framework for assessing visual attention explanations in facial skin d

→30 Jul 2026EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

arXiv:2607.26518v1 Announce Type: new Abstract: Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic unc

→31 Jul 2026Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

arXiv:2607.28227v1 Announce Type: cross Abstract: GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision a

→4 Aug 2026HorusEye: Language as Dynamic Attention for Emergency Visual Analysis

arXiv:2606.14741v2 Announce Type: replace Abstract: We introduce HorusEye, Language as Dynamic Attention for Emergency Visual Analysis. Our investigation followed five stages. The first one is benchma

→4 Aug 2026Generative AI and Foundation Models in Medical Image

arXiv:2608.01686v1 Announce Type: new Abstract: In recent years, generative AI has attracted significant public attention, and its use has been rapidly expanding across a wide range of domains. From c

→11 Aug 2026VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

arXiv:2608.09573v1 Announce Type: new Abstract: Natural-language-driven 'vibe coding' enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of thei

CompanyMeta8 recent entries
23 Jul 2026Development of an automated, reliable, and clinically meaningful artificial intelligence (AI) tool for diagnosing cardiac disease from conventional cardiovascular magnetic resonance (CMR) images

arXiv:2607.20087v1 Announce Type: new Abstract: Aims: Cardiovascular magnetic resonance (CMR) imaging enables non-invasive assessment of myocardial structure, function, and pathology, but requires sub

→28 Jul 2026Dementia Etiology Diagnosis via Collaborative Meta Knowledge Enhancement

arXiv:2607.22770v1 Announce Type: cross Abstract: Although artificial intelligence (AI) has shown promising performance in several medical tasks, accurate dementia etiology diagnosis with AI remains c

→30 Jul 2026TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models

arXiv:2607.26536v1 Announce Type: new Abstract: High-pressure prompts can push vision-language models (VLMs) into unsupported commitments, such as reading illegible text, reporting indeterminate times

→4 Aug 2026DS@GT ARC at MEDIQA-CORE-Task-1 2026: Trimodal Model Fusion with Task-Specific Gates for Brain Tumor Subtype Classification

arXiv:2608.00086v1 Announce Type: new Abstract: Brain tumor diagnosis is a time-sensitive process in which patients may wait weeks for a finalized pathology report. This problem motivates automated sy

→6 Aug 2026Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modali

→7 Aug 2026A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from d

→10 Aug 2026Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation

arXiv:2608.07117v1 Announce Type: new Abstract: Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for

→12 Aug 2026VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

arXiv:2608.11201v1 Announce Type: new Abstract: Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and auth

CompanyMistral8 recent entries
16 Apr 2026Synthesis: Arxiv-Cs-Cl

Auto-generated synthesis of 505 entries about arxiv-cs-cl

→21 Apr 2026From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration

arXiv:2604.16462v1 Announce Type: new Abstract: High-resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Ex

→29 Apr 2026One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations

arXiv:2604.25102v1 Announce Type: new Abstract: Typographic prompt injection exploits vision language models' (VLMs) ability to read text rendered in images, posing a growing threat as VLMs power auto

→7 May 2026Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs

arXiv:2605.04903v1 Announce Type: cross Abstract: Large language models (LLMs) show strong potential for neural architecture generation, yet existing approaches produce complete model implementations

→29 May 2026Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models

arXiv:2602.14399v2 Announce Type: replace Abstract: Multi-turn jailbreak attacks have proven effective against text-only large language models (LLMs), where malicious content is gradually introduced t

→25 Jun 2026Steering Vision-Language Models with Joint Sparse Autoencoders

arXiv:2606.25657v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representat

→4 Aug 2026RSRA: Training-Free Probing of Representation Sensitivity for Efficient LoRA Rank Allocation

arXiv:2607.09757v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning enables large language models to adapt to downstream tasks with substantially lower computational and storage cost,

→6 Aug 2026Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modali

CompanyxAI8 recent entries
2 Jul 2026Explainability in mulimodal deep transformation models for stroke outcome prediction

arXiv:2504.06299v2 Announce Type: replace-cross Abstract: Multimodal prediction models based on imaging and clinical data are increasingly used for clinical decision support, yet their interpretabilit

→9 Jul 2026Why Fake ? Unveiling the Semantic Vocabulary of Deepfake Detectors

arXiv:2607.07216v1 Announce Type: new Abstract: Deepfake (DF) technology poses a significant threat to information integrity, driving the need for robust detection methods. Most DF detectors only cons

→23 Jul 2026Local Label-Informed Feature Transfer for Generating Ground-Truth Medical Images: A Comparison of GAN- and Diffusion-Based Approaches

arXiv:2607.18882v1 Announce Type: new Abstract: Validating Explainable Artificial Intelligence (XAI) methods in medical imaging requires ground-truth data with known locations of informative features.

→24 Jul 2026Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification

arXiv:2607.21068v1 Announce Type: cross Abstract: Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reduci

→30 Jul 2026Comparing the Performance of Foundation Model Derived Embeddings with Traditional Approaches for Distant Metastasis Prediction in Head and Neck Cancer

arXiv:2607.26276v1 Announce Type: new Abstract: Background: Early prediction of distant metastasis (DM) risk in head and neck cancer (HNC) can enable timely interventions that may improve treatment ou

→4 Aug 2026Explainable Multimodal AI for Adaptive Calibration of Archaeological Sensing Workflows

arXiv:2608.00074v1 Announce Type: new Abstract: This paper presents a multimodal machine-learning framework for calibration monitoring, quality assessment, and adaptive acquisition support in archaeol

→4 Aug 2026Empirical investigation of 3D CT Foundation Models and Unsupervised Adaptation for Head and Neck Cancer Recurrence Prediction

arXiv:2608.00071v1 Announce Type: new Abstract: The rapid emergence of 3D CT foundation models has opened new avenues for predictive modeling from CT imaging, offering a compelling alternative to trad

→4 Aug 2026Does Explainability Transfer? A Controlled Benchmark of Attribution Methods on Vision Transformers and CNNs

arXiv:2608.02396v1 Announce Type: new Abstract: Most evidence on the effectiveness of explainable artificial intelligence (XAI) attribution methods has been established on convolutional neural network

CompanyDeepSeek8 recent entries
12 May 2026RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology

arXiv:2605.10761v1 Announce Type: new Abstract: Cancer screening is a reasoning task. A radiologist observes findings, compares them to prior scans, integrates clinical context, and reaches a diagnost

→14 May 2026DocAtlas: Multilingual Document Understanding Across 80+ Languages

arXiv:2605.12623v1 Announce Type: cross Abstract: Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that p

→10 Jun 2026Kwai Keye-VL-2.0 Technical Report

arXiv:2606.10651v1 Announce Type: new Abstract: We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding

→23 Jun 2026Unlimited OCR Works

arXiv:2606.23050v1 Announce Type: new Abstract: Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely held view is that employing a larg

→8 Jul 2026Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models

arXiv:2607.06012v1 Announce Type: new Abstract: Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.e., legal status, structural

→8 Jul 2026LLM-Driven Neural Network Generation with Same-Family Architecture Guidance: Disentangling Transfer and Adaptation

arXiv:2607.05704v1 Announce Type: cross Abstract: Large language models (LLMs) can generate neural-network modifications, but unrestricted generation is often invalid or harmful. This paper studies a

→4 Aug 2026Decoupling semantics from vision: A framework for faithful visual-text compression evaluation

arXiv:2608.01848v1 Announce Type: new Abstract: Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks

→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

CompanyNVIDIA8 recent entries
31 Jul 2026Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles

arXiv:2607.28483v1 Announce Type: new Abstract: Real-time anomaly segmentation is essential for the safety of autonomous systems. Although recent approaches offer high accuracy, their computational co

→3 Aug 2026TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking

arXiv:2604.01207v2 Announce Type: replace Abstract: Existing 3D Gaussian Splatting (3DGS) editing methods primarily focus on appearance modification and often struggle to support flexible geometry edi

→3 Aug 2026Inference-time Trajectory Optimization for Structure-Preserving Manga Image Editing

arXiv:2603.27790v2 Announce Type: replace Abstract: We present a lightweight, training-free trajectory correction method that adapts a pretrained image editing model to each input manga image using on

→5 Aug 2026JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

arXiv:2608.03974v1 Announce Type: new Abstract: Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term tempo

→10 Aug 2026Summarize First, Download Later: Onboard VLMs for Bandwidth-Efficient Earth Observation

arXiv:2608.06959v1 Announce Type: new Abstract: Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downl

→11 Aug 2026What Irregularity Costs: CUDA C++, Rust, and Triton on a Hash-Blocked GPU Workload

arXiv:2608.08287v1 Announce Type: new Abstract: GPU language comparisons are almost always run on tiled dense linear algebra, where every toolchain is good and the differences are small. We implement

→11 Aug 2026ERF-GS: Reconstructing Fast Motion from Disjoint Event-RGB Viewpoints

arXiv:2608.08531v1 Announce Type: new Abstract: Deep learning-driven representations such as neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS) have revolutionized the field of dynamic 3D

→11 Aug 2026City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

arXiv:2608.08887v1 Announce Type: new Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems o