AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “synthesis”

GridTimelineEvolution
2,818 results
12 May 2026

SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness

Model ReleasesDGX agent

arXiv:2603.14889v2 Announce Type: replace-cross Abstract: The rapid evolution of end-to-end spoken dialogue systems demands transcending mere textual semantics to incorporate paralinguistic nuances an

SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making

SafetyDGX agent

arXiv:2605.09343v1 Announce Type: new Abstract: Decision making in large-scale complaint handling systems increasingly relies on heterogeneous evidence, including complaint narratives, screenshots, or

SmartEval: A Benchmark for Evaluating LLM-Generated Smart Contracts from Natural Language Specifications

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.09610v1 Announce Type: cross Abstract: We introduce SmartEval, a benchmark for systematically evaluating the quality of Solidity smart contracts generated by large language models (LLMs) fr

Stable Long-Horizon PDE Forecasting via Latent Structured Spectral Propagators

SafetyDGX agent

arXiv:2605.10154v1 Announce Type: new Abstract: Long-horizon forecasting of time-dependent partial differential equations (PDEs) is critical for characterizing the sustained evolution of physical syst

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation

HardwareDGX agent

arXiv:2605.06356v2 Announce Type: replace Abstract: High-resolution image-to-video (I2V) generation aims to synthesize realistic temporal dynamics while preserving fine-grained appearance details of t

The EDA Primer: From RTL to Silicon

HardwareDGX agent

The EDA Primer covers the semiconductor design and manufacturing workflow, explaining how Electronic Design Automation tools transform Register Transfer Level (RTL) code into physical silicon through

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence

ResearchDGX agent

arXiv:2605.10588v1 Announce Type: new Abstract: Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confi

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

Model ReleasesDGX agent

arXiv:2605.10832v1 Announce Type: new Abstract: Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visua

VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation

Model ReleasesDGX agent

arXiv:2605.08553v1 Announce Type: cross Abstract: Large language models can generate useful code from natural language, but their outputs come without correctness guarantees. Verifiable code generatio

VISTA: A Generative Egocentric Video Framework for Daily Assistance

SafetyDGX agent

arXiv:2605.10579v1 Announce Type: new Abstract: Training AI agents to proactively assist humans in daily activities, from routine household tasks to urgent safety situations, requires large-scale visu

What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers

ResearchDGX agent

arXiv:2605.10180v1 Announce Type: new Abstract: The rise of text-to-image (T2I) models has increasingly raised concerns regarding the generation of risky content, such as sexual, violent, and copyrigh

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

SafetyDGX agent

arXiv:2605.05611v2 Announce Type: replace-cross Abstract: In this paper, we present X-Voice, a 0.4B multilingual zero-shot voice cloning model that clones arbitrary voices and enables everyone to spea

11 May 2026

3DSS: 3D Surface Splatting for Inverse Rendering

ResearchDGX agent

arXiv:2605.05876v2 Announce Type: replace-cross Abstract: We present 3D Surface Splatting (3DSS), the first differentiable surface splatting renderer for physically-based inverse rendering from multi-

Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph

SafetyDGX agent

arXiv:2605.08037v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) aligns language models using pairwise preference comparisons, offering a simple and effective alternative to Rein

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring

Model ReleasesDGX agent

arXiv:2605.07415v1 Announce Type: cross Abstract: Referring expression grounding is a core problem in visual grounding and is widely used as a diagnostic of spatial grounding and reasoning in vision a

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation

Model ReleasesDGX agent

arXiv:2605.07247v1 Announce Type: new Abstract: Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments

From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms

AgentsDGX agent

arXiv:2605.06716v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents have fundamentally reshaped artificial intelligence by integrating external tools and planning capabilities. Whi

HiDream-O1-Dev vs ZImage Base (style comparison)

Local AiDGX agent

A Reddit discussion comparing HiDream-O1-Image-Dev (the newer, 8B pixel-native distilled model) with ZImage Base, examining stylistic and performance differences between the two text-to-image generati

HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization

Local AiDGX agent

arXiv:2605.07214v1 Announce Type: new Abstract: Large Language Models have recently emerged as a promising paradigm for automated heuristic design for NP-hard combinatorial optimization problems. Desp

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

Model ReleasesDGX agent

arXiv:2605.07177v1 Announce Type: cross Abstract: Existing multimodal search agents process target entities sequentially, issuing one tool call per entity and accumulating redundant interaction rounds

Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features

ResearchDGX agent

arXiv:2603.03096v2 Announce Type: replace-cross Abstract: How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information

Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents

Model ReleasesDGX agent

arXiv:2605.06957v1 Announce Type: new Abstract: We present a dynamic policy-learning approach that combines generalized planning and hierarchical task decomposition for LLM-based agents. Our method, H

LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

AgentsDGX agent

arXiv:2605.08083v1 Announce Type: new Abstract: Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during infe

// LLMs Improving LLMs // Interesting progress the past of couple of weeks around self-improving AI agents. If autoresearch was interesting,…

TutorialsDGX agent

// LLMs Improving LLMs // Interesting progress the past of couple of weeks around self-improving AI agents. If autoresearch was interesting, you will like this read. (bookmark it) We've been hand-tuni

Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate

Model ReleasesDGX agent

arXiv:2605.07342v1 Announce Type: cross Abstract: Compile-pass rate is the dominant evaluation signal for LLM code generation, yet for multi-component domain-specific artifacts it can be actively misl

On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics

ResearchDGX agent

arXiv:2605.06835v1 Announce Type: cross Abstract: Tabular data plays an important role in many fields and industries, including those with elevated privacy considerations and risks. As such, there is

Open-Ended Task Discovery via Bayesian Optimization

TutorialsDGX agent

arXiv:2605.07572v1 Announce Type: new Abstract: When applying Bayesian optimization (BO) to scientific workflow, a major yet often overlooked source of uncertainty is the task itself -- namely, what t

Pretraining a Foundation Model for Small-Molecule Natural Products

ResearchDGX agent

arXiv:2503.17656v4 Announce Type: replace-cross Abstract: Natural products, as metabolites from microorganisms, animals, or plants, exhibit diverse biological activities, making them crucial for drug

ReasonSTL: Bridging Natural Language and Signal Temporal Logic via Tool-Augmented Process-Rewarded Learning

Model ReleasesDGX agent

arXiv:2605.06483v2 Announce Type: replace Abstract: Signal Temporal Logic (STL) is an expressive formal language for specifying spatio-temporal requirements over real-valued, real-time signals. It has

TextLDM: Language Modeling with Continuous Latent Diffusion

SafetyDGX agent

arXiv:2605.07748v1 Announce Type: new Abstract: Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next st

Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimization

ResearchDGX agent

arXiv:2605.08054v1 Announce Type: new Abstract: Generating human motion that satisfies customized zero-shot goal functions, enabling applications such as controllable character animation and behavior

TRACE: Tourism Recommendation with Accountable Citation Evidence

ResearchDGX agent

arXiv:2605.07677v1 Announce Type: cross Abstract: Tourism is a high-stakes setting for conversational recommender systems (CRS): a plausible-sounding suggestion can waste real money and trip time once

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

ApplicationsDGX agent

arXiv:2605.06761v1 Announce Type: new Abstract: The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection at

7 May 2026

A Skill-Based AI Agentic Pipeline for Library of Congress Subject Indexing

SafetyDGX agent

arXiv:2605.03537v1 Announce Type: cross Abstract: This paper presents a modular AI agentic skill pipeline for automating subject indexing with Library of Congress Subject Headings (LCSH). Subject inde

Advancing voice intelligence with new models in the API

Model ReleasesDGX agent

OpenAI announced new voice intelligence models available through its API, expanding capabilities for developers to integrate advanced voice processing and understanding features into their application

Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

ResearchDGX agent

arXiv:2605.04128v1 Announce Type: cross Abstract: We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing

Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs

Model ReleasesDGX agent

arXiv:2605.04903v1 Announce Type: cross Abstract: Large language models (LLMs) show strong potential for neural architecture generation, yet existing approaches produce complete model implementations

Empirical Study of Pop and Jazz Mix Ratios for Genre-Adaptive Chord Generation

Model ReleasesDGX agent

arXiv:2605.04998v1 Announce Type: cross Abstract: Chord progression generation is practically important but understudied. Most large-scale symbolic music systems target melody, multi-track arrangement

MedFabric and EtHER: A Data-Centric Framework for Word-Level Fabrication Generation and Detection in Medical LLMs

SafetyDGX agent

arXiv:2605.04180v1 Announce Type: new Abstract: Large Language Models exhibit strong reasoning and semantic understanding capabilities but often hallucinate in domains that require expert knowledge, a

MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents

Model ReleasesDGX agent

arXiv:2605.03675v1 Announce Type: new Abstract: Long-running autonomous AI agents suffer from a well-documented memory coherence problem: tool-execution success rates degrade 14 percentage points over

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

AgentsDGX agent

arXiv:2605.05185v1 Announce Type: new Abstract: Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence v

Stream-T1: Test-Time Scaling for Streaming Video Generation

TutorialsDGX agent

arXiv:2605.04461v1 Announce Type: new Abstract: While Test-Time Scaling (TTS) offers a promising direction to enhance video generation without the surging costs of training, current test-time video ge

6 May 2026

A Few-Step Generative Model on Cumulative Flow Maps

Local AiDGX agent

arXiv:2605.03623v1 Announce Type: new Abstract: We propose a unified, few-step generative modeling framework based on cumulative flow maps for long-range transport in probability space, inspired by fl

Audio-Visual Intelligence in Large Foundation Models

SafetyDGX agent

arXiv:2605.04045v1 Announce Type: new Abstract: Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

Model ReleasesDGX agent

arXiv:2605.03637v1 Announce Type: new Abstract: Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans an

Can synthetic data reproduce real-world findings in epidemiology? A replication study using adversarial random forests

ApplicationsDGX agent

arXiv:2508.14936v3 Announce Type: replace-cross Abstract: Synthetic data holds substantial potential to address practical challenges in epidemiology due to restricted data access and privacy concerns.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding

ResearchDGX agent

arXiv:2605.02290v1 Announce Type: new Abstract: Distilling large reasoning models is essential for making Long-CoT reasoning practical, as full-scale inference remains computationally prohibitive. Exi

HistCAD: A Constraint-Aware Parametric History-Based CAD Representation, Dataset, and Benchmark with Industrial Complexity

Model ReleasesDGX agent

arXiv:2602.19171v2 Announce Type: replace-cross Abstract: Parametric CAD sequences are reusable because dimensional and geometric constraints govern how parameter changes propagate. Existing CAD gener

Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects

AgentsDGX agent

arXiv:2511.20657v2 Announce Type: replace-cross Abstract: The development of agents with emotional intelligence is becoming increasingly vital due to their significant role in human-computer interacti

Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling

ResearchDGX agent

arXiv:2602.00594v2 Announce Type: replace Abstract: A good language model starts with a good tokenizer. Tokenization is especially important for speech modeling, which must handle continuous signals t

Large Language Models are Universal Reasoners for Visual Generation

SafetyDGX agent

arXiv:2605.04040v1 Announce Type: new Abstract: Text-to-image generation has advanced rapidly with diffusion models, progressing from CLIP and T5 conditioning to unified systems where a single LLM bac

MindMelody: A Closed-Loop EEG-Driven System for Personalized Music Intervention

Local AiDGX agent

arXiv:2605.01235v1 Announce Type: cross Abstract: Driven by the escalating global burden of mental health conditions, music-based interventions have attracted significant attention as a non-invasive,

OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories

AgentsDGX agent

arXiv:2605.04036v1 Announce Type: cross Abstract: Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominat

Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

ResearchDGX agent

arXiv:2605.03079v1 Announce Type: cross Abstract: Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake

Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems

Model ReleasesDGX agent

arXiv:2605.04018v1 Announce Type: new Abstract: Reasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity. This capabilit

Safety-critical Control Under Partial Observability: Reach-Avoid POMDP meets Belief Space Control

SafetyDGX agent

arXiv:2603.10572v2 Announce Type: replace Abstract: Partially Observable Markov Decision Processes (POMDPs) provide a principled framework for robot decision-making under uncertainty. Solving reach-av

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning

ApplicationsDGX agent

arXiv:2601.04809v5 Announce Type: replace Abstract: Reinforcement learning (RL) offers a principled way to enhance the reasoning capabilities of large language models, yet its effectiveness hinges on

The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail

Model ReleasesDGX agent

arXiv:2605.03073v1 Announce Type: new Abstract: Niche-domain Indic ASR -- digit strings, currency amounts, addresses, brand names, English/Indic codemix -- is under-served by both open-source SOTA and

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation

Model ReleasesDGX agent

arXiv:2605.01809v1 Announce Type: cross Abstract: Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive medi

5 May 2026

ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control

ResearchDGX agent

arXiv:2510.18232v2 Announce Type: replace Abstract: Generating high-quality synthetic text under differential privacy (DP) is critical for training and evaluating language models without compromising

← Previous
1…3738394041…47
Next →