AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Teach-to-Reason: Competition-Guided Reasoning with a Self-Improving Teacher

DGX agent

arXiv:2606.25407v1 Announce Type: new Abstract: Chest X-ray visual question answering (CXR VQA) requires models not only to predict correct answers, but also to produce reliable medical reasoning. How

researcharxiv-cs-cv
25 Jun 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

DGX agent

arXiv:2606.25975v1 Announce Type: cross Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight

model-releasesarxiv-cs-cv
25 Jun 2026
Research

TensorLDM: A Component-Wise Latent Diffusion Model for Volumetric DTI Reconstruction from Sparse DWIs

DGX agent

arXiv:2606.25545v1 Announce Type: new Abstract: Reconstructing diffusion tensors from sparse DWIs is critical for accelerating Diffusion Tensor Imaging (DTI) in clinical settings, yet current deep lea

researcharxiv-cs-cv
25 Jun 2026
Applications

Test-Time Adaptation in Optical Coherence Tomography Using Trajectory-Aligned Time-Independent Flow

DGX agent

arXiv:2606.18876v2 Announce Type: replace Abstract: Optical coherence tomography (OCT) is essential in ophthalmology, but inconsistent image quality especially in low-cost devices hinders automated an

applicationsarxiv-cs-cv
25 Jun 2026
Agents

To View Transform or Not to View Transform: NeRF-based Pre-training Perspective

DGX agent

arXiv:2603.28090v2 Announce Type: replace Abstract: Neural radiance fields (NeRFs) have emerged as a prominent pre-training paradigm for vision-centric autonomous driving, which enhances 3D geometry a

agentsarxiv-cs-cv
25 Jun 2026
Model Releases

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding

DGX agent

arXiv:2606.25160v1 Announce Type: cross Abstract: The rapid rise of Vision-Language Models (VLMs) in egocentric visual understanding has made low-latency inference in human-robot collaborative (HRC) t

model-releasesarxiv-cs-cv
25 Jun 2026
Research

Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding

DGX agent

arXiv:2606.25658v1 Announce Type: new Abstract: Currently, streaming video understanding is still a daunting task for existing multimodal large language models (MLLMs). Its difficulties not only lie i

researcharxiv-cs-cv
25 Jun 2026
Research

Transferable Attack against Face Swapping in an Extended Space

DGX agent

arXiv:2606.25376v1 Announce Type: new Abstract: Although deep Face Swapping (FS) models may benefit the entertainment industry, they pose severe threats to privacy and security. Existing protections,

researcharxiv-cs-cv
25 Jun 2026
Model Releases

TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs

DGX agent

arXiv:2606.26029v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under co

model-releasesarxiv-cs-cv
25 Jun 2026
Research

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

DGX agent

arXiv:2606.26092v1 Announce Type: new Abstract: While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynamic subjects, existing paradigms rem

researcharxiv-cs-cv
25 Jun 2026
Safety

TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D Reconstruction

DGX agent

arXiv:2601.22615v3 Announce Type: replace Abstract: Streaming recurrent models enable efficient 3D reconstruction by maintaining persistent state representations. However, they suffer from catastrophi

safetyarxiv-cs-cv
25 Jun 2026
Agents

UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving

DGX agent

arXiv:2606.25736v1 Announce Type: new Abstract: Diffusion models have shown strong potential for multi-modal planning in end-to-end autonomous driving. However, most existing methods confine diffusion

agentsarxiv-cs-cv
25 Jun 2026
Model Releases

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning

DGX agent

arXiv:2606.25880v1 Announce Type: new Abstract: Embodied Visual Tracking (EVT) requires an agent to continuously follow a specified target while actively moving through dynamic environments. However,

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

DGX agent

arXiv:2606.25319v1 Announce Type: new Abstract: Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in

model-releasesarxiv-cs-cv
25 Jun 2026
Research

VENI: Variational Encoder for Natural Illumination

DGX agent

arXiv:2601.14079v2 Announce Type: replace Abstract: Inverse rendering is an ill-posed problem, but priors such as illumination priors can help simplify it. Existing work either disregards the spherica

researcharxiv-cs-cv
25 Jun 2026
Safety

VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction

DGX agent

arXiv:2509.19297v3 Announce Type: replace Abstract: Feed-forward 3D Gaussian Splatting (3DGS) has emerged as a highly effective solution for novel view synthesis. Existing methods predominantly rely o

safetyarxiv-cs-cv
25 Jun 2026
Model Releases

VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

DGX agent

arXiv:2606.25592v1 Announce Type: new Abstract: Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfac

model-releasesarxiv-cs-cv
25 Jun 2026
Research

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

DGX agent

arXiv:2606.25041v1 Announce Type: new Abstract: We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex

researcharxiv-cs-cv
25 Jun 2026
Model Releases

What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment

DGX agent

arXiv:2606.25718v1 Announce Type: new Abstract: Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging

model-releasesarxiv-cs-cv
25 Jun 2026
Safety

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

DGX agent

arXiv:2606.25034v1 Announce Type: new Abstract: General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversaria

safetyarxiv-cs-cv
25 Jun 2026
Local Ai

3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy

DGX agent

arXiv:2606.23964v1 Announce Type: cross Abstract: Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We prese

local-aiarxiv-cs-cv
24 Jun 2026
Model Releases

3DCarGen: Scalable 3D Car Generation via 3D-consistent Multi-view Synthesis

DGX agent

arXiv:2606.24257v1 Announce Type: new Abstract: High-quality 3D vehicle assets are essential for autonomous driving simulation. Although multi-view diffusion-based paradigms enable controllable single

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

A Benchmark of State-Space Models vs. Transformers and BiLSTM-based Models for Historical Newspaper OCR

DGX agent

arXiv:2604.00725v2 Announce Type: replace Abstract: End-to-end OCR for historical newspapers remains challenging, as models must handle long text sequences, degraded print quality, and complex layouts

model-releasesarxiv-cs-cv
24 Jun 2026
Research

A Dual Edge Spatial Jacobian Image Graph for Interpretable Diabetic Retinopathy Grading

DGX agent

arXiv:2606.24168v1 Announce Type: cross Abstract: Automated diabetic retinopathy (DR) grading from colour fundus photographs can achieve strong predictive performance, but clinical interpretation requ

researcharxiv-cs-cv
24 Jun 2026
Safety

A Geometry-Informed Computer Vision Method for Detecting and Examining Overtaking Vehicles From A Bicycle

DGX agent

arXiv:2606.23699v1 Announce Type: new Abstract: Instrumented bicycle studies have produced direct field evidence on vehicle passing behavior, but extracting overtaking events from continuous rear-faci

safetyarxiv-cs-cv
24 Jun 2026
Research

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP

DGX agent

arXiv:2603.05962v2 Announce Type: replace Abstract: To address the limitations of existing open-vocabulary object recognition methods, including high system complexity, substantial training costs, and

researcharxiv-cs-cv
24 Jun 2026
Model Releases

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

DGX agent

arXiv:2606.23835v1 Announce Type: new Abstract: ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generati

model-releasesarxiv-cs-cv
24 Jun 2026
Research

Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction

DGX agent

arXiv:2606.24156v1 Announce Type: new Abstract: Visual token reduction has emerged as an effective strategy for accelerating Multimodal Large Language Models (MLLMs). Many existing methods prune token

researcharxiv-cs-cv
24 Jun 2026
Local Ai

ActiveScope: Actively Seeking and Correcting Perception for MLLMs

DGX agent

arXiv:2606.24292v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated impressive vision-language understanding, yet still struggle with fine-grained perception in

local-aiarxiv-cs-cv
24 Jun 2026
Research

Adaptive Hebbian Memory Routing in Vision Transformers for Few-Shot Learning

DGX agent

arXiv:2606.24756v1 Announce Type: new Abstract: Few-shot image recognition requires models to adapt to new classes from a small labeled support set. Hebbian fast-weight memory can provide temporary as

researcharxiv-cs-cv
24 Jun 2026
Research

Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods

DGX agent

arXiv:2606.24484v1 Announce Type: new Abstract: WordArt (artistic text) features highly customized fonts, textures, and layouts, making WordArt-oriented scene TExt Recognition (WATER) substantially mo

researcharxiv-cs-cv
24 Jun 2026
Research

AerialFusionMapNet: Online HD Map Construction with Aerial-Onboard BEV Fusion

DGX agent

arXiv:2606.24784v1 Announce Type: new Abstract: High-resolution aerial imagery has recently emerged as a complementary modality for automated driving perception and has shown potential to improve bird

researcharxiv-cs-cv
24 Jun 2026
Agents

Agentic Collaborative Cognition for Zero-Shot 3D Understanding

DGX agent

arXiv:2606.24649v1 Announce Type: new Abstract: Recent advancements have explored agentic zero-shot 3D understanding by reformulating it as video keyframe understanding with Multimodal Large Language

agentsarxiv-cs-cv
24 Jun 2026
Local Ai

An LMM for Precisely Grounding Elements in Documents

DGX agent

arXiv:2606.24118v1 Announce Type: new Abstract: Visual grounding in documents is a crucial ability for Large Multimodal Models (LMMs) in areas such as document understanding, deep research and documen

local-aiarxiv-cs-cv
24 Jun 2026
Model Releases

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

DGX agent

arXiv:2606.24548v1 Announce Type: new Abstract: Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it rem

model-releasesarxiv-cs-cv
24 Jun 2026
Applications

ArtiTwinSplat: Interactable Digital Twin Reconstruction via Gaussian Splatting from RGB-D videos

DGX agent

arXiv:2606.24628v1 Announce Type: cross Abstract: Deploying robots in unstructured real-world environments needs accurate, interactive models of the objects. Constructing these models at scale remains

applicationsarxiv-cs-cv
24 Jun 2026
Research

Automated Residual Plot Assessment With the R Package autovi and the Shiny Application autovi.web

DGX agent

arXiv:2606.24236v1 Announce Type: cross Abstract: Visual assessment of residual plots is a common approach for diagnosing linear models, but it relies on manual evaluation, which does not scale well a

researcharxiv-cs-cv
24 Jun 2026
Agents

Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models

DGX agent

arXiv:2606.24152v1 Announce Type: new Abstract: Existing literature claims that video generation essentially is world modelling. On the one hand, the claim is productive because it pushes generative A

agentsarxiv-cs-cv
24 Jun 2026
Model Releases

BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases

DGX agent

arXiv:2606.24883v1 Announce Type: new Abstract: Artificial intelligence (AI) has achieved remarkable success in medical imaging, but it is widely recognized that these models often perform inconsisten

model-releasesarxiv-cs-cv
24 Jun 2026
Research

Bengal-HP_RU: A Dataset of Bengal People For Head Pose Estimation

DGX agent

arXiv:2606.24122v1 Announce Type: new Abstract: Existing head pose datasets predominantly feature subjects of Western or East Asian origin, leaving South Asian populations, particularly Bengali indivi

researcharxiv-cs-cv
24 Jun 2026
Model Releases

BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming

DGX agent

arXiv:2606.24740v1 Announce Type: new Abstract: Recent advances in vision-language models (VLMs) such as CLIP have demonstrated strong generalization across natural-image domains. However, adapting th

model-releasesarxiv-cs-cv
24 Jun 2026
Safety

Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation

DGX agent

arXiv:2606.24464v1 Announce Type: new Abstract: Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing mod

safetyarxiv-cs-cv
24 Jun 2026
Safety

Bridging the Manifold Gap: Riemannian Residual Line Search for One-Step Image Editing

DGX agent

arXiv:2606.24844v1 Announce Type: new Abstract: One-step diffusion editors are fast because they avoid inversion and iterative optimization, but a single transport update must be aggressive enough to

safetyarxiv-cs-cv
24 Jun 2026
Model Releases

CanadaFireSat: Toward high-resolution wildfire forecasting with multiple modalities

DGX agent

arXiv:2506.08690v3 Announce Type: replace Abstract: Canada experienced in 2023 one of the most severe wildfire seasons in recent history, causing damage across ecosystems, destroying communities, and

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Compact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual Relocalization

DGX agent

arXiv:2606.24767v1 Announce Type: new Abstract: Indoor visual relocalization plays a critical role in emerging spatial and embodied AI applications. However, prior research was predominantly devoted t

model-releasesarxiv-cs-cv
24 Jun 2026
Research

Configurable Holography: Towards Display and Scene Adaptation

DGX agent

arXiv:2405.01558v4 Announce Type: replace Abstract: Rendering holograms for holographic displays is often an iterative and computationally costly process. Emerging learned holography methods have alle

researcharxiv-cs-cv
24 Jun 2026
Model Releases

Counting Trees from Satellite Imagery with Noisy Supervision

DGX agent

arXiv:2606.24786v1 Announce Type: new Abstract: Counting individual trees is a fundamental task for environmental monitoring, yet remains largely unexplored with satellite imagery. At these resolution

model-releasesarxiv-cs-cv
24 Jun 2026
Research

CrossFusion: A Multi-Scale Cross-Attention Convolutional Fusion Model for Cancer Survival Prediction

DGX agent

arXiv:2503.02064v2 Announce Type: replace-cross Abstract: Cancer survival prediction from whole slide images (WSIs) is a challenging task in computational pathology due to the large size, irregular sh

researcharxiv-cs-cv
24 Jun 2026
← Previous
1…9192939495…263
Next →