AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “synthesis”

GridTimelineEvolution
2,844 results
Safety

KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta

DGX agent

arXiv:2512.23236v4 Announce Type: replace-cross Abstract: Making deep learning recommendation model (DLRM) training and inference fast and efficient is important. However, this presents three key syst

safetyarxiv-cs-ai
8 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

MorphGS: Morphology-Adaptive Articulated 3D Motion Transfer from Videos

DGX agent

arXiv:2601.02716v3 Announce Type: replace Abstract: Transferring articulated motion from monocular videos to rigged 3D characters is challenging due to pose ambiguity in 2D observations and morphologi

applicationsarxiv-cs-cv
8 Jul 2026
Model Releases

PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation

DGX agent

arXiv:2607.06440v1 Announce Type: new Abstract: Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personaliz

model-releasesarxiv-cs-cv
8 Jul 2026
Research

RayRoPE: Projective Ray Positional Encoding for Multi-view Attention

DGX agent

arXiv:2601.15275v3 Announce Type: replace Abstract: We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes pa

researcharxiv-cs-cv
8 Jul 2026
Safety

Recovering Cloud Microstructures with Cascaded Diffusion Inversion

DGX agent

arXiv:2607.05637v1 Announce Type: new Abstract: High-resolution satellite imagery is critical for observing fine-scale cloud structures that inform weather modification strategies like cloud seeding f

safetyarxiv-cs-cv
8 Jul 2026
Research

SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation

DGX agent

arXiv:2607.05994v1 Announce Type: new Abstract: Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenu

researcharxiv-cs-cv
8 Jul 2026
Safety

Volumetric Directional Diffusion: Anchoring Uncertainty Quantification in Anatomical Consensus for Ambiguous Medical Image Segmentation

DGX agent

arXiv:2603.04024v2 Announce Type: replace-cross Abstract: Ambiguous 3D medical image segmentation often involves boundaries where different expert delineations are non-identical yet clinically plausib

safetyarxiv-cs-ai
8 Jul 2026
Safety

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

DGX agent

arXiv:2607.06461v1 Announce Type: cross Abstract: While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit e

safetyarxiv-cs-cl
8 Jul 2026
Applications

ASSEMCAD: Production-Ready CAD Assembly Generation from Natural Language

DGX agent

arXiv:2607.05123v1 Announce Type: new Abstract: Recent advances in large language models and programmatic CAD have significantly improved Text-to-CAD generation for individual parts. However, producti

applicationsarxiv-cs-ai
7 Jul 2026
Model Releases

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

DGX agent

arXiv:2606.12555v2 Announce Type: replace-cross Abstract: Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a

model-releasesarxiv-cs-cv
7 Jul 2026
Research

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

DGX agent

arXiv:2607.02968v1 Announce Type: new Abstract: Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiT

researcharxiv-cs-cv
7 Jul 2026
Research

City-Level 3D Surface Reconstruction with Viewpoint Orientation Partitioning and Scene Completion

DGX agent

arXiv:2607.03771v1 Announce Type: new Abstract: Multi-view 3D surface reconstruction is a longstanding challenge in computer vision. Although recent large-scale reconstruction methods based on 3D Gaus

researcharxiv-cs-cv
7 Jul 2026
Applications

Deep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical Translation

DGX agent

arXiv:2607.05311v1 Announce Type: new Abstract: Male infertility contributes substantially to the global infertility burden, and sperm analysis remains central to diagnosis, treatment planning, and as

applicationsarxiv-cs-cv
7 Jul 2026
Local Ai

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech

DGX agent

arXiv:2607.04140v1 Announce Type: cross Abstract: Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness by

local-aiarxiv-cs-cl
7 Jul 2026
Model Releases

ELiTeFormer: An Efficient Transformer for FPGAs

DGX agent

arXiv:2607.03652v1 Announce Type: cross Abstract: Transformer blocks are prevalent in large language model (LLM) but present deployment challenges due to their challenging computational and memory dem

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Enhancing Facial Expression Recognition in Head-Mounted Displays with Synthetic Data

DGX agent

arXiv:2607.04490v1 Announce Type: new Abstract: Facial expression recognition (FER) is crucial for social interaction in mixed reality environments that employ head-mounted displays (HMD). However, co

safetyarxiv-cs-cv
7 Jul 2026
Safety

Few-Shot Demonstration-Driven Task Coordination and Trajectory Execution for Multi-Robot Systems

DGX agent

arXiv:2510.15686v2 Announce Type: replace Abstract: Learning coordinated behaviors for multi-robot systems from only a few demonstrations is difficult because temporal task dependencies and spatial tr

safetyarxiv-cs-ro
7 Jul 2026
Safety

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

DGX agent

arXiv:2607.03509v1 Announce Type: new Abstract: Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inferen

safetyarxiv-cs-cv
7 Jul 2026
Agents

FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents

DGX agent

arXiv:2607.04718v1 Announce Type: new Abstract: Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form reports. This work

agentsarxiv-cs-ai
7 Jul 2026
Model Releases

GenShin: Guiding Rational Liposome Design by Ranking Liposomal Protein Corona through a Docking-Pose-Free GNN

DGX agent

arXiv:2504.13853v2 Announce Type: replace-cross Abstract: Rational design of lipid nanoparticles (LNPs) for tissue-specific delivery critically depends on predicting the composition of the protein cor

model-releasesarxiv-cs-ai
7 Jul 2026
Research

InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem

DGX agent

arXiv:2512.05672v2 Announce Type: replace-cross Abstract: Recent approaches in controllable novel view video generation often rely on fine-tuning pre-trained Video Diffusion Models (VDMs). This domina

researcharxiv-cs-ai
7 Jul 2026
Model Releases

iVISION-2DCD: A Long-Term Change Detection Dataset for Large-Scale Outdoor Construction Monitoring

DGX agent

arXiv:2607.03553v1 Announce Type: new Abstract: Automation in construction is essential for reducing costs and human errors in large-scale projects. We approach the construction progress monitoring fr

model-releasesarxiv-cs-cv
7 Jul 2026
Applications

Measuring the Robustness of Audio Deepfake Detection under Real-World Corruption

DGX agent

arXiv:2503.17577v2 Announce Type: replace-cross Abstract: Deepfakes have emerged as a widespread and rapidly escalating concern in generative AI, spanning images, audio, and videos. Among these, audio

applicationsarxiv-cs-ai
7 Jul 2026
Research

Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules

DGX agent

arXiv:2606.16337v3 Announce Type: replace Abstract: Predictive modeling for clinical decision support requires not only strong predictive performance but also transparent decision logic. Although deep

researcharxiv-cs-ai
7 Jul 2026
Research

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space

DGX agent

arXiv:2607.02907v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step

researcharxiv-cs-cl
7 Jul 2026
Research

Sparse-View Surface Reconstruction using Gaussian Splatting through High-Confidence Depth Propagation with Normal Priors

DGX agent

arXiv:2607.03765v1 Announce Type: new Abstract: 3D reconstruction from sparse views is a challenging task in 3D computer vision. Recent studies on 3D Gaussian Splatting (3DGS) have achieved remarkable

researcharxiv-cs-cv
7 Jul 2026
Model Releases

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

DGX agent

arXiv:2511.07403v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial rea

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers

DGX agent

arXiv:2601.02211v2 Announce Type: replace Abstract: Recent breakthroughs of transformer-based diffusion models, particularly with Multimodal Diffusion Transformers (MMDiT) driven models like FLUX and

model-releasesarxiv-cs-cv
7 Jul 2026
Safety

TokAN: Accent Normalization Using Self-Supervised Speech Tokens

DGX agent

arXiv:2607.03928v1 Announce Type: cross Abstract: Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity. The current te

safetyarxiv-cs-ai
7 Jul 2026
Research

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

DGX agent

arXiv:2607.04515v1 Announce Type: new Abstract: Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresen

researcharxiv-cs-cl
7 Jul 2026
Model Releases

UniVideo: Unified Understanding, Generation, and Editing for Videos

DGX agent

arXiv:2510.08377v4 Announce Type: replace Abstract: Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain.

model-releasesarxiv-cs-cv
7 Jul 2026
Research

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

DGX agent

The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories fo

researchapple-ml-research
7 Jul 2026
Model Releases

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

DGX agent

arXiv:2607.03562v1 Announce Type: new Abstract: As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limi

model-releasesarxiv-cs-cv
7 Jul 2026
Hardware

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

DGX agent

arXiv:2607.02119v1 Announce Type: cross Abstract: While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This is severe

hardwarearxiv-cs-ai
3 Jul 2026
Safety

CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

DGX agent

arXiv:2603.22435v2 Announce Type: replace-cross Abstract: 'Code-as-Policy' considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as

safetyarxiv-cs-ai
3 Jul 2026
Safety

DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving

DGX agent

arXiv:2603.18315v2 Announce Type: replace-cross Abstract: Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse collision signals, which fail to capture the ric

safetyarxiv-cs-ai
3 Jul 2026
Applications

Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

DGX agent

arXiv:2607.01590v1 Announce Type: new Abstract: Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually navigate impl

applicationsarxiv-cs-ai
3 Jul 2026
Research

How Indian Dermatologists are Utilizing Artificial Intelligence for Clinical Practice and Workflow Management: A Nationwide Survey with a Special Focus on atopic dermatitis

DGX agent

arXiv:2607.01252v1 Announce Type: cross Abstract: Background: Dermatology AI has mainly focused on image-based diagnosis, while chronic disease workflows have received less attention. We surveyed Indi

researcharxiv-cs-ai
3 Jul 2026
Model Releases

LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models

DGX agent

arXiv:2504.02327v2 Announce Type: replace Abstract: Natural Language to SQL (NL2SQL) aims to translate natural language queries into executable SQL statements, offering non-expert users intuitive acce

model-releasesarxiv-cs-cl
3 Jul 2026
Tutorials

Learning 3D-Gaussian Simulators from RGB Videos

DGX agent

arXiv:2503.24009v3 Announce Type: replace-cross Abstract: Realistic simulation is critical for applications ranging from robotics to animation. Learned simulators have emerged as a possibility to capt

tutorialsarxiv-cs-ai
3 Jul 2026
Tutorials

LEFT: Learnable Fusion of Tri-view Tokens for Unsupervised Time Series Anomaly Detection

DGX agent

arXiv:2602.08638v2 Announce Type: replace-cross Abstract: As a fundamental data mining task, unsupervised time series anomaly detection (TSAD) aims to build a model for identifying abnormal timestamps

tutorialsarxiv-cs-ai
3 Jul 2026
Model Releases

Office Comprehension Benchmark

DGX agent

arXiv:2607.01245v1 Announce Type: cross Abstract: We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration

DGX agent

arXiv:2607.01531v1 Announce Type: new Abstract: Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networ

model-releasesarxiv-cs-ai
3 Jul 2026
Agents

Traceable Fault Diagnosis for Battery Energy Storage Systems via Retrieval-Augmented Multi-Agent O&M Assistant

DGX agent

arXiv:2607.01992v1 Announce Type: new Abstract: Large-scale battery energy storage systems (BESSs) require O&M decisions that combine alarms, cell-level measurements, device topology, diagnostic table

agentsarxiv-cs-ai
3 Jul 2026
Model Releases

VisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval

DGX agent

arXiv:2607.02371v1 Announce Type: cross Abstract: Over 285 million people worldwide live with a visual impairment, for whom everyday tasks such as avoiding obstacles, locating personal belongings, rec

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

An LLM-Based Framework for Intent-Driven Network Topology Design

DGX agent

arXiv:2607.00292v1 Announce Type: cross Abstract: Designing deployable and resilient network topologies from natural language requirements remains a challenging problem in network automation. This wor

model-releasesarxiv-cs-ai
2 Jul 2026
Safety

ASPIRE: Agentic /Skills Discovery for Robotics

DGX agent

arXiv:2607.00272v1 Announce Type: cross Abstract: Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling divers

safetyarxiv-cs-ai
2 Jul 2026
Safety

ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation

DGX agent

arXiv:2607.00545v1 Announce Type: new Abstract: Controllable traffic simulation is critical for testing autonomous driving systems, yet existing approaches often require retraining large generative mo

safetyarxiv-cs-cv
2 Jul 2026
← Previous
1…3435363738…60
Next →