AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “synthesis”

GridTimelineEvolution
2,844 results
8 Apr 2026

PS: I finally got around to trying out @randal_olson 's Tufte Test tool to prettify the benchmark plot. Great tool 👌! https://www.goodeyela…

Model ReleasesDGX agent

Sebastian Raschka (rasbt) used Randal Olson's Tufte Test tool, developed by Goodeye Labs, to improve the visual quality of a machine learning benchmark plot. The Tufte Test encodes seven of Tufte'...

13 Aug 2026

A Generalized Theory of Load Distribution in Redundantly-actuated Robotic Systems

ResearchDGX agent

arXiv:2603.11431v2 Announce Type: replace Abstract: This paper presents a generalized theory which describes how applied loads are distributed within rigid bodies handled by redundantly-actuated robot

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2608.12290v1 Announce Type: cross Abstract: Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and rel

D3D-GEN: Robot-Aware Domain-Grounded Interactive 3D World Generation for Social Robotics

AgentsDGX agent

arXiv:2608.11876v1 Announce Type: new Abstract: Training and validation of Embodied AI for social navigation critically depends on realistic simulation environments, yet many current approaches fail t

Diffusion Probe: Generated Image Result Prediction Using CNN Probes

AgentsDGX agent

arXiv:2602.23783v5 Announce Type: replace Abstract: Text-to-image (T2I) diffusion models lack an efficient mechanism for early quality assessment, leading to costly trial-and-error in multi-generation

Explainability in Practice: A Survey of Explainable NLP Across Various Domains

Model ReleasesDGX agent

arXiv:2502.00837v3 Announce Type: replace-cross Abstract: Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management, whe

First-order friction models with bristle dynamics: lumped and distributed formulations

ResearchDGX agent

arXiv:2602.09429v3 Announce Type: replace-cross Abstract: Dynamic models, particularly rate-dependent models, have proven effective in capturing the key phenomenological features of frictional process

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

ResearchDGX agent

arXiv:2608.11219v1 Announce Type: new Abstract: Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a seg

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

ResearchDGX agent

arXiv:2608.12209v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent

GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings

AgentsDGX agent

arXiv:2608.12133v1 Announce Type: new Abstract: Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images. Existing LLM and VLM sys

How Organizations Use AI: Evidence from ChatGPT

TutorialsDGX agent

arXiv:2608.12236v1 Announce Type: cross Abstract: We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

Model ReleasesDGX agent

arXiv:2608.11616v1 Announce Type: new Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to

Multi-Agent Target-Existence Verification and Learned Mask Geometry Refinement: Winning Report of the MeViS-Text Track at the 8th LSVOS Challenge 2026

AgentsDGX agent

arXiv:2608.11458v1 Announce Type: new Abstract: We present the first-place solution to the MeViS-Text track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge 2026: referring video obj

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation

ApplicationsDGX agent

arXiv:2606.17598v2 Announce Type: replace-cross Abstract: Humans naturally leverage diverse sensing modalities to interact with the physical world, while most Vision-Language-Action (VLA) models for r

ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference

SafetyDGX agent

arXiv:2608.12232v1 Announce Type: new Abstract: Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, tempora

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

ApplicationsDGX agent

arXiv:2608.12314v1 Announce Type: new Abstract: Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refi

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

SafetyDGX agent

arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. Howeve

Understanding Why Foundation Models Work for Diffusion-Generated Image Detection

Local AiDGX agent

arXiv:2608.12155v1 Announce Type: new Abstract: Vision foundation models have recently emerged as powerful feature extractors for detecting AI-generated images, achieving strong generalization across

12 Aug 2026

AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations

ResearchDGX agent

arXiv:2608.11123v1 Announce Type: new Abstract: Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

SafetyDGX agent

arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu

DreamOmni3: Scribble-based Editing and Generation

Model ReleasesDGX agent

arXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text

DSAR: Dual-Stream Autoregressive Modeling of Temporal Cloth Dynamics for Photorealistic Animatable Avatars

TutorialsDGX agent

arXiv:2608.10500v1 Announce Type: new Abstract: Creating photorealistic and temporally coherent animatable human avatars from RGB videos remains challenging. Current methods struggle to capture realis

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

Local AiDGX agent

arXiv:2608.10756v1 Announce Type: cross Abstract: Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before ex

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

Model ReleasesDGX agent

arXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive fie

FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

SafetyDGX agent

arXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc

Introspective Attention Modulation for Safe Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr

LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

AgentsDGX agent

arXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deploymen

LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3

SafetyDGX agent

arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int

MAD-HOI: Masked Autoregressive Diffusion for Generating Articulated Hand Object Interactions from Text

Model ReleasesDGX agent

arXiv:2608.10162v1 Announce Type: new Abstract: Methods for text-based generation of hand-object interaction (HOI) sequences primarily focus on producing smooth, physically plausible trajectories. A t

Motion Artifact-Aware Self-Supervised Representation Learning for 3D Brain MRI Motion Artifact Reduction

ResearchDGX agent

arXiv:2608.10170v1 Announce Type: new Abstract: Patient motion remains a source of image degradation in brain MRI, leading to signal loss, blurring, and geometric distortion that compromise quantitati

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

Local AiDGX agent

arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image repr

Rationale-Guided Learning for Multimodal Emotion Recognition

TutorialsDGX agent

arXiv:2608.10448v1 Announce Type: new Abstract: Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most exis

Real-World Cooperative Bimanual Dexterous Grasp of Large Objects from Single-View Observations

ApplicationsDGX agent

arXiv:2608.10383v1 Announce Type: new Abstract: Bimanual dexterous grasping of large objects is a critical challenge in robotic manipulation. However, most existing studies focus on sequential manipul

Situation Graph Prediction for User Perspective Modeling

Model ReleasesDGX agent

arXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data

The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

SafetyDGX agent

arXiv:2608.10839v1 Announce Type: new Abstract: This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems traine

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

Model ReleasesDGX agent

arXiv:2608.10682v1 Announce Type: new Abstract: Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing

What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research

SafetyDGX agent

arXiv:2608.10431v1 Announce Type: cross Abstract: Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpret, implemen

11 Aug 2026

A continually expandable foundation model for brain MRI

ResearchDGX agent

arXiv:2608.08319v1 Announce Type: new Abstract: Brain magnetic resonance imaging (MRI) is central to neuroscience and clinical assessment, but models are commonly developed for individual diseases, po

Accelerate PostgreSQL migrations using Gemini in Database Migration Service

Model ReleasesDGX agent

Imagine this scenario: Your team decides to migrate a core application from an existing commercial database like Oracle or SQL Server to open source PostgreSQL or a fully managed service such as Alloy

AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2603.05868v2 Announce Type: replace Abstract: Despite remarkable progress in Vision-Language-Action models (VLAs) for robot manipulation, these large pre-trained models require fine-tuning to be

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

Model ReleasesDGX agent

arXiv:2608.08459v1 Announce Type: cross Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domain

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

Model ReleasesDGX agent

arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

TutorialsDGX agent

arXiv:2608.08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitatio

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

ResearchDGX agent

arXiv:2608.08067v1 Announce Type: cross Abstract: Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

SafetyDGX agent

arXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent

Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial Intelligence

HardwareDGX agent

arXiv:2608.08761v1 Announce Type: cross Abstract: In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor manufacturin

Efficient Human-Contact Representation for Human-Scene Interaction

Model ReleasesDGX agent

arXiv:2608.09388v1 Announce Type: new Abstract: Human-scene interaction is an active research topic with several industrial applications in virtual reality, gaming, robotics, and surveillance. Despite

EsaacSim: A Multimodal Event Camera Add-on for NVIDIA Isaac Sim

HardwareDGX agent

arXiv:2608.08522v1 Announce Type: new Abstract: Event-based vision is becoming an increasingly important sensing paradigm for robotics, yet its adoption remains limited by sensor availability and the

GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum

AgentsDGX agent

arXiv:2603.28533v3 Announce Type: replace Abstract: Agentic knowledge graph question answering (KGQA) requires an agent to iteratively interact with knowledge graphs (KGs), posing challenges in both t

HandSplatter: Automated Digital Goniometry from Neural Rendering

ResearchDGX agent

arXiv:2608.09735v1 Announce Type: new Abstract: Hand and finger disorders are leading contributors to musculoskeletal disability, creating a clinical need for precise methods to quantify joint motion.

IndexTTS 2.5 Technical Report

Model ReleasesDGX agent

arXiv:2601.03888v4 Announce Type: replace-cross Abstract: In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-base

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

Model ReleasesDGX agent

arXiv:2608.09818v1 Announce Type: cross Abstract: Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-

North Africa's Missing Framework: NLP-Driven Mental Healthcare in Algeria and Implications for Low-resource Settings

SafetyDGX agent

arXiv:2608.08607v1 Announce Type: new Abstract: Mental health disorders are a leading cause of disability worldwide, yet Natural Language Processing (NLP) research for mental healthcare has remained c

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Model ReleasesDGX agent

arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing

Progressive Learned Image Compression for Machine Perception

ApplicationsDGX agent

arXiv:2512.20070v2 Announce Type: replace Abstract: Recent advances in learned image codecs have extended from human perception toward machine perception However, progressive image compression with fi

RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing

Model ReleasesDGX agent

arXiv:2608.09186v1 Announce Type: new Abstract: Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geom

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

Model ReleasesDGX agent

arXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance,

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Model ReleasesDGX agent

arXiv:2608.09873v1 Announce Type: cross Abstract: We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It co

Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence

ApplicationsDGX agent

arXiv:2608.08075v1 Announce Type: cross Abstract: Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archi

Sekai2: From World Exploration to Interactive World Modeling

Model ReleasesDGX agent

arXiv:2608.09449v1 Announce Type: new Abstract: Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefor

← Previous
1…2223242526…48
Next →