AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
4 Aug 2026

Decoding Children's Gait Behavior

ResearchDGX agent

arXiv:2608.00371v1 Announce Type: new Abstract: We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We speci

DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing

AgentsDGX agent

arXiv:2608.01761v1 Announce Type: new Abstract: End-to-end (E2E) autonomous driving algorithms require rigorous closed-loop validation in simulation environments offering high visual fidelity, strong

Decoupling semantics from vision: A framework for faithful visual-text compression evaluation

Model ReleasesDGX agent

arXiv:2608.01848v1 Announce Type: new Abstract: Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Deep Learning CNN and Recurrence Analysis for Alpha Gamma EEG Biomarkers in Fragile X Syndrome

TutorialsDGX agent

arXiv:2608.00835v1 Announce Type: cross Abstract: Fragile X Syndrome (FXS) is a neurodevelopmental disorder caused by reduced expression of fragile X mental retardation protein (FMRP), leading to disr

Deep Learning for Retinal Degeneration Assessment: A Comprehensive Analysis of the MARIO Challenge

Model ReleasesDGX agent

arXiv:2506.02976v4 Announce Type: replace Abstract: The MARIO challenge, held at MICCAI 2024, focused on advancing the automated detection and monitoring of age-related macular degeneration (AMD) thro

Deep Multimodal Fusion Detection through Spatial Mask and Channel Fusion

TutorialsDGX agent

arXiv:2608.02092v1 Announce Type: new Abstract: Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics. However, existing feature-level fusi

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

AgentsDGX agent

arXiv:2608.01827v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to

DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization

ResearchDGX agent

arXiv:2608.02099v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor ar

Deja Cue: Localizing States in Object Histories via Vocabulary-Relative Coordinates

ResearchDGX agent

arXiv:2608.02044v1 Announce Type: new Abstract: Tracking links observations of the same object through visual change, yet cannot by itself determine when the object is empty or filled, intact or cut.

DerainSplat: Feed-Forward Clean 3D Gaussian Splatting from Sparse Rainy Views

AgentsDGX agent

arXiv:2608.02191v1 Announce Type: new Abstract: Although image deraining has advanced substantially, existing methods mainly focus on 2D image restoration. As spatial intelligence applications such as

Detail Continuation over a Trustworthy Coarse Scale for Autoregressive Super-Resolution

Model ReleasesDGX agent

arXiv:2608.01823v1 Announce Type: new Abstract: Hallucination remains a persistent challenge in generative super-resolution (GSR), where reconstructed results may contain visually plausible yet weakly

Device-First Feedback: Toward Mobile-Native LLM-Driven Neural Architecture Search

Local AiDGX agent

arXiv:2608.00078v1 Announce Type: new Abstract: Deploying convolutional neural networks generated by large language models (LLMs) on real mobile hardware requires more than GPU validation accuracy: IN

DeVIT: Low-Power Vision Transformer Acceleration Using Delta Computation

ResearchDGX agent

arXiv:2608.01343v1 Announce Type: new Abstract: The emergence of transformer-based deep learning models has brought unprecedented performance across various domains, particularly in natural language p

DexMani: Human-Derived Manipulability Guidance for Dexterous Rotation

TutorialsDGX agent

arXiv:2608.00554v1 Announce Type: cross Abstract: Dexterous object rotation is a sequential contact problem: each support, release, and re-contact decision must both produce the desired object motion,

DF^3: World Modeling via Decoder-Free Feature Forecasting in Autonomous Navigation

AgentsDGX agent

arXiv:2608.02428v1 Announce Type: new Abstract: Forecasting future states from video sequences is a critical challenge for autonomous robotic systems and a fundamental objective of world modeling. Pri

Diagnosing Under-Development of Irreversible Processes in Video Generation

ResearchDGX agent

arXiv:2608.00617v1 Announce Type: new Abstract: Many physical attributes are irreversible: ice melts but does not re-freeze, paper chars but does not un-burn. Do video generators respect this? We show

DiffPhysCam: Differentiable Physics-Based Camera Simulation for Inverse Rendering and Embodied AI

AgentsDGX agent

arXiv:2508.08831v2 Announce Type: replace-cross Abstract: Generating synthetic images that closely mimic those from real cameras is instrumental in training visual models and enabling end-to-end visuo

DiffPrune: differentiable information throttling for token pruning in vision-language models

TutorialsDGX agent

arXiv:2608.01985v1 Announce Type: new Abstract: Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score th

DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning

SafetyDGX agent

arXiv:2608.00540v1 Announce Type: new Abstract: Tool-integrated vision-language agents have made remarkable progress on compositional and multi-step visual reasoning. Yet their outputs frequently exhi

Direct and Adaptable Mesh-Gaussian Scene Reconstruction from Multi-View Images

AgentsDGX agent

arXiv:2405.06945v4 Announce Type: replace Abstract: Jointly recovering explicit surface geometry and high-quality appearance from multi-view images remains challenging. This capability is essential fo

Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models

ResearchDGX agent

arXiv:2608.00110v1 Announce Type: new Abstract: 3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient

Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework

ResearchDGX agent

arXiv:2607.15933v2 Announce Type: replace Abstract: The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretiz

Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding

Model ReleasesDGX agent

arXiv:2607.17999v2 Announce Type: replace-cross Abstract: Spatial understanding is crucial for foundation models (FMs), and maps have long helped humans organize and reason about geographic informatio

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards

SafetyDGX agent

arXiv:2608.00536v1 Announce Type: new Abstract: Reinforcement learning (RL) for document parsing often relies on reference-based rewards rooted in edit distance (e.g., tree edit distance), yet it rema

DODA: A Database of Datasets for Aesthetics Research

ResearchDGX agent

arXiv:2608.00089v1 Announce Type: new Abstract: With rapid growth in the fields of empirical and computational aesthetics we have seen a vast increase in large image datasets annotated for aesthetics.

Does Explainability Transfer? A Controlled Benchmark of Attribution Methods on Vision Transformers and CNNs

Model ReleasesDGX agent

arXiv:2608.02396v1 Announce Type: new Abstract: Most evidence on the effectiveness of explainable artificial intelligence (XAI) attribution methods has been established on convolutional neural network

DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable

Model ReleasesDGX agent

arXiv:2608.00548v1 Announce Type: new Abstract: Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their ras

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

ResearchDGX agent

arXiv:2608.00486v1 Announce Type: new Abstract: Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: a

DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving

AgentsDGX agent

arXiv:2603.00919v3 Announce Type: replace Abstract: Large language models (LLMs) have shown great promise for autonomous driving. However, discretizing numbers into tokens limits precise numerical rea

Driver2Map: Imitating Human Driving for Online High-Definition Map Construction

SafetyDGX agent

arXiv:2608.01338v1 Announce Type: new Abstract: High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition

DS@GT ARC at MEDIQA-CORE-Task-1 2026: Trimodal Model Fusion with Task-Specific Gates for Brain Tumor Subtype Classification

Model ReleasesDGX agent

arXiv:2608.00086v1 Announce Type: new Abstract: Brain tumor diagnosis is a time-sensitive process in which patients may wait weeks for a finalized pathology report. This problem motivates automated sy

DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation

Local AiDGX agent

arXiv:2608.02495v1 Announce Type: new Abstract: Despite the remarkable progress over the past decades, accurately identifying small objects remains challenging because of their insufficient visual cue

DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction

AgentsDGX agent

arXiv:2608.01178v1 Announce Type: new Abstract: We present DynActiveGS, a dynamic-aware active reconstruction framework based on 3D Gaussian Splatting (3DGS) for autonomous exploration in dynamic envi

Dynamic Resolution Routing for Efficient Egocentric Grounding

Local AiDGX agent

arXiv:2608.01638v1 Announce Type: new Abstract: Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain

DynamicManip: Enabling Dynamic Manipulation from a Single Static Demonstration

Model ReleasesDGX agent

arXiv:2608.01452v1 Announce Type: cross Abstract: Dynamic manipulation is a critical capability for robots operating in complex and dynamic environments, where robots must interact with objects that a

E2Pano: Learning Event-to-Panorama Image Reconstruction

Model ReleasesDGX agent

arXiv:2608.00694v1 Announce Type: new Abstract: Event cameras offer microsecond-level temporal resolution and high dynamic range, potentially facilitating motion-blur-free panoramic imaging from fast

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation

Model ReleasesDGX agent

arXiv:2608.02474v1 Announce Type: new Abstract: Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its infer

EEG-FM-Compass: Progress, Benchmarking, and Future Directions for EEG Foundation Models

Model ReleasesDGX agent

arXiv:2601.17883v3 Announce Type: replace-cross Abstract: Electroencephalography (EEG) foundation models (FMs) have recently emerged as a promising paradigm for brain-computer interfaces, aiming to le

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next

Model ReleasesDGX agent

arXiv:2603.12147v2 Announce Type: replace Abstract: Egocentric video provides a natural modality for studying human behavior, but conventional visual understanding captures mainly observable scenes, o

ELECTRIC: Evidential Learning-Enhanced CT Reconstruction via Iterative Correction

ResearchDGX agent

arXiv:2608.00060v1 Announce Type: new Abstract: Here we introduce ELECTRIC (Evidential Learning-Enhanced CT Reconstruction via Iterative Correction), a physics-guided Bayesian formulation. An evidenti

Element-Aware Group Learning for E-Commerce Image Generation

SafetyDGX agent

arXiv:2608.00584v1 Announce Type: new Abstract: Recent advances in image generation and editing have made prompt quality a key bottleneck for e-commerce creatives. Vision-language models (VLMs) can ge

EmoScene: A Dual-space Dataset for Controllable Affective Image Generation

ResearchDGX agent

arXiv:2604.00933v2 Announce Type: replace Abstract: Text-to-image diffusion models achieve high visual fidelity, yet fine-grained affective control remains difficult because textual emotion cues often

Empirical investigation of 3D CT Foundation Models and Unsupervised Adaptation for Head and Neck Cancer Recurrence Prediction

ResearchDGX agent

arXiv:2608.00071v1 Announce Type: new Abstract: The rapid emergence of 3D CT foundation models has opened new avenues for predictive modeling from CT imaging, offering a compelling alternative to trad

Enhancing Visual Perception in Foggy Conditions via Multiclass Fog Density Modeling

AgentsDGX agent

arXiv:2608.01572v1 Announce Type: new Abstract: Autonomous driving (AD) systems have advanced rapidly over the past decade; however, robust perception under adverse weather conditions remains a major

Entity-Aware Sequence Transduction for Player-Centric Ball Action Spotting

SafetyDGX agent

arXiv:2608.01696v1 Announce Type: new Abstract: Player-centric ball action spotting requires temporally precise event detection together with actor attribution in crowded, partially observed multi-age

EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass

ResearchDGX agent

arXiv:2608.02284v1 Announce Type: new Abstract: Open-vocabulary segmentation identifies and segments objects from arbitrary textual descriptions. SAM 3 supports noun-phrase-guided segmentation and ach

Estimating SSIM from MSE for DCT-Based Compressed Images

ResearchDGX agent

arXiv:2608.02549v1 Announce Type: cross Abstract: Efficient and perceptually meaningful quality assessment is a fundamental requirement for image and video processing, compression, and streaming syste

Event ActivityNet: A Large-Scale Simulated-Event Benchmark for Untrimmed Action Understanding

Model ReleasesDGX agent

arXiv:2608.01948v1 Announce Type: new Abstract: Long-horizon event-based action understanding remains underexplored because existing datasets largely comprise short, trimmed clips, while collecting na

Explainable Multimodal AI for Adaptive Calibration of Archaeological Sensing Workflows

ResearchDGX agent

arXiv:2608.00074v1 Announce Type: new Abstract: This paper presents a multimodal machine-learning framework for calibration monitoring, quality assessment, and adaptive acquisition support in archaeol

Extended Field of View Analysis for VideoGAN-based Trajectory Generation

ResearchDGX agent

arXiv:2608.02289v1 Announce Type: new Abstract: Realistic and diverse trajectory generation is central to enabling higher levels of vehicle automation. While rule-based and classical learning-based me

Extended KAFR: A kinematic-adaptive paradigm for the efficient analysis of surgical video

Model ReleasesDGX agent

arXiv:2608.01058v1 Announce Type: new Abstract: Artificial Intelligence is increasingly applied to surgical video analysis for phase segmentation, skill assessment, and workflow optimization. A key ch

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

AgentsDGX agent

arXiv:2608.01049v1 Announce Type: cross Abstract: World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this e

FairForensics: Seeing Expressions and Parsing Demographics via Vision-Language Modeling for Generalizable Fair Deepfake Detection

Model ReleasesDGX agent

arXiv:2608.01661v1 Announce Type: new Abstract: The challenge of fair deepfake detection (FDD) has attracted increasing attention. Existing fairness-enhanced detectors often suffer from suboptimal gen

FAST-GS: Frequency Aware Space-time Gaussian Splatting for Photorealistic Dynamic Novel View Synthesis

Local AiDGX agent

arXiv:2608.01958v1 Announce Type: new Abstract: 4D Gaussian Splatting (4DGS) excels in dynamic 3D reconstruction and real-time novel view synthesis via efficient 4D Gaussian representations and parall

Fast Trainable Multilinear Bases for Image Compression

ResearchDGX agent

arXiv:2608.00053v1 Announce Type: cross Abstract: The Discrete Fourier Transform, the Discrete Cosine Transform, and their block-wise variants underpin most deployed image and video codecs. Their effe

FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering

ResearchDGX agent

arXiv:2608.01664v1 Announce Type: new Abstract: We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answer

FDIR: Harmonizing Fidelity and Human-Machine Preference in Lossy Compression Image Restoration

ResearchDGX agent

arXiv:2608.00111v1 Announce Type: cross Abstract: Image restoration quality can be evaluated along three complementary facets: pixel-level fidelity, human perception, and downstream machine preference

FeDepth: Federated Learning for Depth Estimation under Robot Heterogeneity

ResearchDGX agent

arXiv:2608.01129v1 Announce Type: cross Abstract: Although recent robot perception research emphasizes training on data from diverse environments to improve generalization, most existing methods still

Fermat Active Laplace Learning for Semi-Supervised Hyperspectral Image Classification

ResearchDGX agent

arXiv:2608.02483v1 Announce Type: new Abstract: Two active learning algorithms for hyperspectral image (HSI) classification are proposed that combine density-aware Fermat distances with Poisson-reweig

Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding

ResearchDGX agent

arXiv:2608.01663v1 Announce Type: new Abstract: Promptable segmentation foundation models (FMs) such as SAM3 and Medical SAM3 promise few-shot, interactively-specified segmentation for medical imaging

← Previous
1…1415161718…207
Next →