AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

DGX agent

arXiv:2608.02437v1 Announce Type: new Abstract: Single-image feed-forward 3D Gaussian Splatting (3DGS) aims to directly generate a renderable 3D scene representation from one input image, avoiding the

researcharxiv-cs-cv
4 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

InstancePin: Instance-Addressable Layout-to-Image Diffusion via Coordinate Pinning

DGX agent

arXiv:2608.00588v1 Announce Type: new Abstract: Layout-to-image diffusion models have achieved impressive semantic controllability by conditioning generation on category-level segmentation maps. Howev

researcharxiv-cs-cv
4 Aug 2026
Model Releases

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

DGX agent

arXiv:2608.01157v1 Announce Type: new Abstract: Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by

model-releasesarxiv-cs-cv
4 Aug 2026
Tutorials

Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers

DGX agent

arXiv:2608.00264v1 Announce Type: new Abstract: Vision foundation models, such as DINOv2, learn highly expressive representations but rely on massive, opaque architectures that demand substantial comp

tutorialsarxiv-cs-cv
4 Aug 2026
Safety

Investigating Social Bias in Narrative Image Generation

DGX agent

arXiv:2608.01780v1 Announce Type: new Abstract: Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how

safetyarxiv-cs-cv
4 Aug 2026
Safety

Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents

DGX agent

arXiv:2608.02018v1 Announce Type: new Abstract: Computer-use agents (CUAs), which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to in

safetyarxiv-cs-cv
4 Aug 2026
Research

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models

DGX agent

arXiv:2607.15732v2 Announce Type: replace Abstract: Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs boundin

researcharxiv-cs-cv
4 Aug 2026
Model Releases

ISRS-DETR: Detection-Guided Click Propagation for Remote Sensing Interactive Segmentation

DGX agent

arXiv:2608.02468v1 Announce Type: new Abstract: Interactive segmentation reduces the prohibitive cost of pixel-level annotation by allowing users to delineate objects with a few clicks. However, apply

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling

DGX agent

arXiv:2608.01207v1 Announce Type: new Abstract: Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers poor

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

JADE-GS: Joint Allocation of Deblurring Evidence for Event-Assisted 3D Gaussian Splatting

DGX agent

arXiv:2607.14990v2 Announce Type: replace Abstract: Neural radiance fields and 3D Gaussian Splatting assume that each training image is a sharp and geometrically consistent observation of the scene. M

model-releasesarxiv-cs-cv
4 Aug 2026
Local Ai

K-space Gaussian Representation for Parallel MRI

DGX agent

arXiv:2608.00075v1 Announce Type: new Abstract: Accelerated magnetic resonance imaging (MRI) aims to recover the k-space signal from acquired measurements, where accurate estimation of missing samples

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving

DGX agent

arXiv:2608.00237v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving, enabling agents to map multimodal inputs

model-releasesarxiv-cs-cv
4 Aug 2026
Research

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

DGX agent

arXiv:2608.00079v1 Announce Type: new Abstract: Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits strea

researcharxiv-cs-cv
4 Aug 2026
Safety

Learning How Much, Not Just What: Cross-Patient Burden Order for CT Vision-Language Pretraining

DGX agent

arXiv:2608.00231v1 Announce Type: new Abstract: Volumetric CT vision-language pretraining learns 3D representations from scan-report pairs, but global and anatomy-aware objectives supervise only corre

safetyarxiv-cs-cv
4 Aug 2026
Safety

Learning to See Locally and Align Clinically with Pathology Semantics for Radiology Report Generation

DGX agent

arXiv:2608.00279v1 Announce Type: cross Abstract: Recent radiology-adapted vision-language models have achieved strong performance on standard report generation benchmarks, yet their robustness and ge

safetyarxiv-cs-cv
4 Aug 2026
Research

Learning to Tessellate: Point Cloud Generation via Recursive Spectral Partitioning

DGX agent

arXiv:2608.02432v1 Announce Type: new Abstract: Autoregressive models have emerged as an effective paradigm for point cloud generation. However, most existing approaches rely on heuristic tokenization

researcharxiv-cs-cv
4 Aug 2026
Tutorials

Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency

DGX agent

arXiv:2608.01730v1 Announce Type: new Abstract: No-reference image quality assessment (NR IQA) has recently benefited from deep and multimodal models, yet many SOTA systems still violate at least one

tutorialsarxiv-cs-cv
4 Aug 2026
Research

Less is More: Compact-Token Masked Feature Prediction for Skeleton Representation Learning

DGX agent

arXiv:2603.10648v3 Announce Type: replace Abstract: Current skeleton representation learning paradigms face distinct limitations: Contrastive Learning (CL) often overlooks fine-grained motion details,

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Lethe: How Hard Is It to Forget? A Benchmark for Federated Unlearning in Medical Imaging

DGX agent

arXiv:2608.01094v1 Announce Type: new Abstract: Federated learning enables medical-imaging models to be trained across hospitals, and privacy law, most explicitly the GDPR ``right to be forgotten'', t

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention

DGX agent

arXiv:2602.04789v4 Announce Type: replace Abstract: Advanced autoregressive (AR) video generation models have improved visual fidelity and interactivity, but the quadratic complexity of attention rema

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge

DGX agent

arXiv:2608.01614v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face a critical computational bottleneck when processing high-resolution imagery due to the O(N^2) memory complexity of So

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Linguistic Context Recodes Visual Representations in Vision-Language Models

DGX agent

arXiv:2608.00035v1 Announce Type: cross Abstract: Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categor

researcharxiv-cs-cv
4 Aug 2026
Applications

LiveLight: Real-time Streaming Video Relighting with Interactive Control

DGX agent

arXiv:2608.01771v1 Announce Type: new Abstract: We present LiveLight, the first diffusion-based framework for real-time streaming video relighting with interactive 3D lighting control. Achieving this

applicationsarxiv-cs-cv
4 Aug 2026
Local Ai

Local Margin Restoration for Test-Time Adaptation of Vision-Language Models

DGX agent

arXiv:2608.02216v1 Announce Type: new Abstract: Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected

local-aiarxiv-cs-cv
4 Aug 2026
Local Ai

Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models

DGX agent

arXiv:2608.00976v1 Announce Type: new Abstract: Fine-grained visual representations are essential for medical image analysis, particularly when diagnostically relevant evidence is subtle and spatially

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

Loggia dei Lanzi: AI Thermography Enhancement Comparisons through 3D Photogrammetry

DGX agent

arXiv:2608.02404v1 Announce Type: new Abstract: The Loggia dei Lanzi in the Piazza della Signoria is one of Florence's most prominent structures visited by millions every year. Its construction histor

model-releasesarxiv-cs-cv
4 Aug 2026
Applications

Logit-Origin Centering for Singleton Test-Time Adaptation

DGX agent

arXiv:2608.01074v1 Announce Type: cross Abstract: Tabular data is used extensively in many real-world use cases. Deep learning models have been developed to deal with tabular data, but generally perfo

applicationsarxiv-cs-cv
4 Aug 2026
Applications

Logographic Character Visual Pretraining via Semantic-based Contrastive Learning

DGX agent

arXiv:2608.00096v1 Announce Type: new Abstract: Current deep learning-based character vision studies, e.g., text recognition, character image denoising, and historical text completion, are offering ne

applicationsarxiv-cs-cv
4 Aug 2026
Model Releases

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

DGX agent

arXiv:2608.01964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interde

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM

DGX agent

arXiv:2608.00925v1 Announce Type: new Abstract: Monocular panoramic SLAM benefits from substantial visual overlap under large camera rotations, yet remains prone to errors caused by camera tilt, scale

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Loop-Mamba: A Loop Mamba with Degradation-Aware and Shared Memory for Old Photo Restoration

DGX agent

arXiv:2608.02346v1 Announce Type: new Abstract: Old photographs often suffer from multiple coupled degradations, including scratches, cracks, fading, blur, noise, and missing regions, severely degradi

model-releasesarxiv-cs-cv
4 Aug 2026
Safety

LUT: Latent Utility Training for Visual Reasoning

DGX agent

arXiv:2608.00743v1 Announce Type: new Abstract: Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reason

safetyarxiv-cs-cv
4 Aug 2026
Model Releases

Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models

DGX agent

arXiv:2409.07163v3 Announce Type: replace-cross Abstract: Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Manifold-GS: Certified Hybrid Assets via Varifold-Conservative Gaussian Splatting

DGX agent

arXiv:2608.00214v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) gives high-quality novel-view synthesis, but its adaptive radiance primitives are not directly usable as structured assets:

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Mapping melliferous tree species in Kenya via one-class classification with hyperspectral unsupervised domain adaptation

DGX agent

arXiv:2608.02045v1 Announce Type: cross Abstract: The beekeeping sector holds significant potential for livelihood diversification among the agropastoral communities in Kenya. Melliferous tree species

researcharxiv-cs-cv
4 Aug 2026
Local Ai

MBO Scheme for Local Chan--Vese Segmentation

DGX agent

arXiv:2608.00893v1 Announce Type: new Abstract: Robust to intensity inhomogeneity, the local Chan--Vese (LCV) model extends the classical Chan--Vese (CV) image segmentation method by incorporating loc

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

MDTD-ArtIR: Benchmarking Image Editing and Restoration Models for Art Image Restoration under Texture-Overlay Degradations

DGX agent

arXiv:2608.00736v1 Announce Type: new Abstract: Restoring severely degraded visual media still remains a formidable challenge, as existing methods often hallucinate unnatural textures and contents, st

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

MDWD: A Street-Level Dataset for Municipal Solid Waste Detection in Dense Urban Environments

DGX agent

arXiv:2608.00257v1 Announce Type: new Abstract: Automated visual monitoring of urban environments is a growing Computer Vision research area, but municipal solid waste detection remains under-represen

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Measuring Product Quality Using Images: The CLIP Q-Score and an Application to Real Estate

DGX agent

arXiv:2608.01544v1 Announce Type: cross Abstract: The CLIP Q-score is a novel, safe, fully reproducible, and computationally efficient method for extracting objective product quality metrics from visu

researcharxiv-cs-cv
4 Aug 2026
Safety

MedSAM2-Anatomy: Training-Free Inference-Time Optimization for Musculoskeletal Segmentation

DGX agent

arXiv:2608.00195v1 Announce Type: cross Abstract: High-resolution 3D segmentation of hip and shoulder anatomy from CT and MRI is essential for surgical planning, yet frozen segmentation models often f

safetyarxiv-cs-cv
4 Aug 2026
Research

Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression

DGX agent

arXiv:2608.02134v1 Announce Type: new Abstract: Modern vision language models (VLMs) turn high-resolution images into long sequences of visual tokens. Every token traverses the language decoder and pe

researcharxiv-cs-cv
4 Aug 2026
Tutorials

MIDAL: Math Image Descriptions for Accessible Learning

DGX agent

arXiv:2608.00868v1 Announce Type: new Abstract: Many open educational resources are lacking in accessibility, especially in-depth image descriptions. In subjects like Science and Mathematics, however,

tutorialsarxiv-cs-cv
4 Aug 2026
Model Releases

MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing

DGX agent

arXiv:2608.02059v1 Announce Type: new Abstract: Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-

model-releasesarxiv-cs-cv
4 Aug 2026
Hardware

MiniWorld: Democratizing the Training of Video World Models from Scratch

DGX agent

arXiv:2608.01127v1 Announce Type: new Abstract: Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through auto

hardwarearxiv-cs-cv
4 Aug 2026
Model Releases

Mitigating Backdoors via Decoy Shortcuts and Knowledge Decoupling

DGX agent

arXiv:2608.00732v1 Announce Type: cross Abstract: Backdoor attacks pose a serious threat to deep neural networks, especially when training relies on third-party data, allowing adversaries to inject ma

model-releasesarxiv-cs-cv
4 Aug 2026
Safety

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning

DGX agent

arXiv:2608.01635v1 Announce Type: new Abstract: Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instructi

safetyarxiv-cs-cv
4 Aug 2026
Safety

Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference

DGX agent

arXiv:2606.15782v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language response

safetyarxiv-cs-cv
4 Aug 2026
Safety

MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling

DGX agent

arXiv:2604.02817v2 Announce Type: replace Abstract: Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel

safetyarxiv-cs-cv
4 Aug 2026
← Previous
1…2223242526…261
Next →