AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “synthesis”

GridTimelineEvolution
2,737 results
Applications

MuseVLA: An Adaptive Multimodal Sensing Vision-Language-Action Model for Robotic Manipulation

DGX agent

arXiv:2606.17598v2 Announce Type: replace-cross Abstract: Humans naturally leverage diverse sensing modalities to interact with the physical world, while most Vision-Language-Action (VLA) models for r

applicationsarxiv-cs-cv
13 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference

DGX agent

arXiv:2608.12232v1 Announce Type: new Abstract: Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, tempora

safetyarxiv-cs-cv
13 Aug 2026
Applications

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

DGX agent

arXiv:2608.12314v1 Announce Type: new Abstract: Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refi

applicationsarxiv-cs-cv
13 Aug 2026
Safety

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

DGX agent

arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. Howeve

safetyarxiv-cs-cl
13 Aug 2026
Local Ai

Understanding Why Foundation Models Work for Diffusion-Generated Image Detection

DGX agent

arXiv:2608.12155v1 Announce Type: new Abstract: Vision foundation models have recently emerged as powerful feature extractors for detecting AI-generated images, achieving strong generalization across

local-aiarxiv-cs-cv
13 Aug 2026
Research

AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations

DGX agent

arXiv:2608.11123v1 Announce Type: new Abstract: Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for

researcharxiv-cs-cv
12 Aug 2026
Safety

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

DGX agent

arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

DreamOmni3: Scribble-based Editing and Generation

DGX agent

arXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text

model-releasesarxiv-cs-cv
12 Aug 2026
Tutorials

DSAR: Dual-Stream Autoregressive Modeling of Temporal Cloth Dynamics for Photorealistic Animatable Avatars

DGX agent

arXiv:2608.10500v1 Announce Type: new Abstract: Creating photorealistic and temporally coherent animatable human avatars from RGB videos remains challenging. Current methods struggle to capture realis

tutorialsarxiv-cs-cv
12 Aug 2026
Local Ai

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

DGX agent

arXiv:2608.10756v1 Announce Type: cross Abstract: Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before ex

local-aiarxiv-cs-cv
12 Aug 2026
Model Releases

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

DGX agent

arXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive fie

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

DGX agent

arXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc

safetyarxiv-cs-cl
12 Aug 2026
Model Releases

Introspective Attention Modulation for Safe Text-to-Image Generation

DGX agent

arXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr

model-releasesarxiv-cs-cv
12 Aug 2026
Agents

LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

DGX agent

arXiv:2608.09934v1 Announce Type: cross Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deploymen

agentsarxiv-cs-ai
12 Aug 2026
Safety

LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3

DGX agent

arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int

safetyarxiv-cs-cv
12 Aug 2026
Model Releases

MAD-HOI: Masked Autoregressive Diffusion for Generating Articulated Hand Object Interactions from Text

DGX agent

arXiv:2608.10162v1 Announce Type: new Abstract: Methods for text-based generation of hand-object interaction (HOI) sequences primarily focus on producing smooth, physically plausible trajectories. A t

model-releasesarxiv-cs-cv
12 Aug 2026
Research

Motion Artifact-Aware Self-Supervised Representation Learning for 3D Brain MRI Motion Artifact Reduction

DGX agent

arXiv:2608.10170v1 Announce Type: new Abstract: Patient motion remains a source of image degradation in brain MRI, leading to signal loss, blurring, and geometric distortion that compromise quantitati

researcharxiv-cs-cv
12 Aug 2026
Local Ai

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

DGX agent

arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image repr

local-aiarxiv-cs-cl
12 Aug 2026
Tutorials

Rationale-Guided Learning for Multimodal Emotion Recognition

DGX agent

arXiv:2608.10448v1 Announce Type: new Abstract: Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most exis

tutorialsarxiv-cs-ai
12 Aug 2026
Applications

Real-World Cooperative Bimanual Dexterous Grasp of Large Objects from Single-View Observations

DGX agent

arXiv:2608.10383v1 Announce Type: new Abstract: Bimanual dexterous grasping of large objects is a critical challenge in robotic manipulation. However, most existing studies focus on sequential manipul

applicationsarxiv-cs-ro
12 Aug 2026
Model Releases

Situation Graph Prediction for User Perspective Modeling

DGX agent

arXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

DGX agent

arXiv:2608.10839v1 Announce Type: new Abstract: This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems traine

safetyarxiv-cs-cv
12 Aug 2026
Model Releases

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

DGX agent

arXiv:2608.10682v1 Announce Type: new Abstract: Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing

model-releasesarxiv-cs-cv
12 Aug 2026
Safety

What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research

DGX agent

arXiv:2608.10431v1 Announce Type: cross Abstract: Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpret, implemen

safetyarxiv-cs-ai
12 Aug 2026
Research

A continually expandable foundation model for brain MRI

DGX agent

arXiv:2608.08319v1 Announce Type: new Abstract: Brain magnetic resonance imaging (MRI) is central to neuroscience and clinical assessment, but models are commonly developed for individual diseases, po

researcharxiv-cs-cv
11 Aug 2026
Model Releases

AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models

DGX agent

arXiv:2603.05868v2 Announce Type: replace Abstract: Despite remarkable progress in Vision-Language-Action models (VLAs) for robot manipulation, these large pre-trained models require fine-tuning to be

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

DGX agent

arXiv:2608.08459v1 Announce Type: cross Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domain

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

DGX agent

arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise

model-releasesarxiv-cs-ai
11 Aug 2026
Tutorials

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

DGX agent

arXiv:2608.08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitatio

tutorialsarxiv-cs-ai
11 Aug 2026
Research

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

DGX agent

arXiv:2608.08067v1 Announce Type: cross Abstract: Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to

researcharxiv-cs-ai
11 Aug 2026
Safety

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

DGX agent

arXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent

safetyarxiv-cs-ai
11 Aug 2026
Hardware

Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial Intelligence

DGX agent

arXiv:2608.08761v1 Announce Type: cross Abstract: In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor manufacturin

hardwarearxiv-cs-ai
11 Aug 2026
Model Releases

Efficient Human-Contact Representation for Human-Scene Interaction

DGX agent

arXiv:2608.09388v1 Announce Type: new Abstract: Human-scene interaction is an active research topic with several industrial applications in virtual reality, gaming, robotics, and surveillance. Despite

model-releasesarxiv-cs-cv
11 Aug 2026
Hardware

EsaacSim: A Multimodal Event Camera Add-on for NVIDIA Isaac Sim

DGX agent

arXiv:2608.08522v1 Announce Type: new Abstract: Event-based vision is becoming an increasingly important sensing paradigm for robotics, yet its adoption remains limited by sensor availability and the

hardwarearxiv-cs-ro
11 Aug 2026
Agents

GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum

DGX agent

arXiv:2603.28533v3 Announce Type: replace Abstract: Agentic knowledge graph question answering (KGQA) requires an agent to iteratively interact with knowledge graphs (KGs), posing challenges in both t

agentsarxiv-cs-cl
11 Aug 2026
Research

HandSplatter: Automated Digital Goniometry from Neural Rendering

DGX agent

arXiv:2608.09735v1 Announce Type: new Abstract: Hand and finger disorders are leading contributors to musculoskeletal disability, creating a clinical need for precise methods to quantify joint motion.

researcharxiv-cs-cv
11 Aug 2026
Model Releases

IndexTTS 2.5 Technical Report

DGX agent

arXiv:2601.03888v4 Announce Type: replace-cross Abstract: In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-base

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

DGX agent

arXiv:2608.09818v1 Announce Type: cross Abstract: Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

North Africa's Missing Framework: NLP-Driven Mental Healthcare in Algeria and Implications for Low-resource Settings

DGX agent

arXiv:2608.08607v1 Announce Type: new Abstract: Mental health disorders are a leading cause of disability worldwide, yet Natural Language Processing (NLP) research for mental healthcare has remained c

safetyarxiv-cs-cl
11 Aug 2026
Model Releases

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

DGX agent

arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing

model-releasesarxiv-cs-cl
11 Aug 2026
Applications

Progressive Learned Image Compression for Machine Perception

DGX agent

arXiv:2512.20070v2 Announce Type: replace Abstract: Recent advances in learned image codecs have extended from human perception toward machine perception However, progressive image compression with fi

applicationsarxiv-cs-cv
11 Aug 2026
Model Releases

RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing

DGX agent

arXiv:2608.09186v1 Announce Type: new Abstract: Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geom

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

DGX agent

arXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance,

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

DGX agent

arXiv:2608.09873v1 Announce Type: cross Abstract: We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It co

model-releasesarxiv-cs-ai
11 Aug 2026
Applications

Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence

DGX agent

arXiv:2608.08075v1 Announce Type: cross Abstract: Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archi

applicationsarxiv-cs-cv
11 Aug 2026
Model Releases

Sekai2: From World Exploration to Interactive World Modeling

DGX agent

arXiv:2608.09449v1 Announce Type: new Abstract: Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefor

model-releasesarxiv-cs-cv
11 Aug 2026
Research

Task-Adaptive 3D Cross-Field MRI Translation via Field-Conditioned Content-Style Pretraining

DGX agent

arXiv:2608.09264v1 Announce Type: new Abstract: Magnetic field strength is a major source of domain shift in magnetic resonance imaging (MRI), affecting signal-to-noise ratio, tissue contrast, spatial

researcharxiv-cs-cv
11 Aug 2026
Local Ai

UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

DGX agent

arXiv:2608.09143v1 Announce Type: new Abstract: Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodifi

local-aiarxiv-cs-cv
11 Aug 2026
← Previous
1…2728293031…58
Next →