AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
Human
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
84,561 results
4 Aug 2026

DiffPrune: differentiable information throttling for token pruning in vision-language models

TutorialsDGX agent

arXiv:2608.01985v1 Announce Type: new Abstract: Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score th

DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning

SafetyDGX agent

arXiv:2608.00540v1 Announce Type: new Abstract: Tool-integrated vision-language agents have made remarkable progress on compositional and multi-step visual reasoning. Yet their outputs frequently exhi

Diffusion-Based Body Schema Learning Enabling Abnormal-State Adaptation in Musculoskeletal Robots

TutorialsDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.01029v1 Announce Type: new Abstract: Musculoskeletal robots require an internal body schema that remains consistent under a wide range of physical state changes, including abnormalities suc

Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning

SafetyDGX agent

arXiv:2608.02332v1 Announce Type: new Abstract: In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous Q-value estimation,

DiffusionGemma Technical Report

Model ReleasesDGX agent

arXiv:2608.00146v1 Announce Type: new Abstract: We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rathe

Direct and Adaptable Mesh-Gaussian Scene Reconstruction from Multi-View Images

AgentsDGX agent

arXiv:2405.06945v4 Announce Type: replace Abstract: Jointly recovering explicit surface geometry and high-quality appearance from multi-view images remains challenging. This capability is essential fo

Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts

TutorialsDGX agent

arXiv:2608.01740v1 Announce Type: new Abstract: Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps. Recent work has mainly focused on de

Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding

ResearchDGX agent

arXiv:2608.01560v1 Announce Type: new Abstract: Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what

Disentangled Contrastive Learning for Zero-Shot Multilingual Dense Retrieval

SafetyDGX agent

arXiv:2608.02189v1 Announce Type: cross Abstract: Multilingual dense retrieval aims to handle queries and documents across different languages based on a unified retriever model. The challenge lies in

Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models

Model ReleasesDGX agent

arXiv:2608.00547v1 Announce Type: new Abstract: Contact-rich manipulation remains challenging because successful control depends on physical interaction cues that are often weakly observable from visi

Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models

ResearchDGX agent

arXiv:2608.00110v1 Announce Type: new Abstract: 3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient

Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models

SafetyDGX agent

arXiv:2608.01263v1 Announce Type: new Abstract: On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance

SafetyDGX agent

arXiv:2608.00782v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relativ

Distilling Drifting Transformers with Representation Autoencoders

ResearchDGX agent

arXiv:2606.15553v2 Announce Type: replace Abstract: Despite the significant training acceleration and promising performance, Representation Autoencoders (RAEs) are mainly criticized for poor distillat

Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework

ResearchDGX agent

arXiv:2607.15933v2 Announce Type: replace Abstract: The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretiz

Divergent large language model predictions from convergent representations in ambiguous word pairs

Model ReleasesDGX agent

arXiv:2608.01816v1 Announce Type: new Abstract: In this work we investigate how decoder-only transformers resolve lexical ambiguity through layer-by-layer analysis of three models spanning three param

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

Model ReleasesDGX agent

arXiv:2608.00011v1 Announce Type: new Abstract: Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale model

Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding

Model ReleasesDGX agent

arXiv:2607.17999v2 Announce Type: replace-cross Abstract: Spatial understanding is crucial for foundation models (FMs), and maps have long helped humans organize and reason about geographic informatio

Do Neural Networks Really Beat the Curse of Dimensionality? A Bit-Complexity View

ResearchDGX agent

arXiv:2608.01357v1 Announce Type: new Abstract: Traditional approximation theory measures convergence rates in terms of the number of parameters or degrees of freedom. However, practical computation o

Do Static Embeddings Add Value to Hybrid Dutch Retrieval?

Model ReleasesDGX agent

arXiv:2608.02112v1 Announce Type: new Abstract: Embedding benchmarks measure standalone model quality, but they do not establish whether a low-cost retriever contributes complementary ranking informat

DocNavRAG: Document-Structured Graph RAG with Stateful Evidence Construction for Complex Document Question Answering

AgentsDGX agent

arXiv:2608.01565v1 Announce Type: new Abstract: Answering complex questions over large document collections requires assembling complementary evidence across sections and documents. GraphRAG offers st

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards

SafetyDGX agent

arXiv:2608.00536v1 Announce Type: new Abstract: Reinforcement learning (RL) for document parsing often relies on reference-based rewards rooted in edit distance (e.g., tree edit distance), yet it rema

DODA: A Database of Datasets for Aesthetics Research

ResearchDGX agent

arXiv:2608.00089v1 Announce Type: new Abstract: With rapid growth in the fields of empirical and computational aesthetics we have seen a vast increase in large image datasets annotated for aesthetics.

Does Accuracy Equal Evidence? Reasoning Faithfulness under KV Cache Compression

ResearchDGX agent

arXiv:2608.01631v1 Announce Type: new Abstract: KV cache compression is commonly evaluated by final-answer accuracy, implicitly assuming that preserving the answer also preserves the reasoning that su

Does Explainability Transfer? A Controlled Benchmark of Attribution Methods on Vision Transformers and CNNs

Model ReleasesDGX agent

arXiv:2608.02396v1 Announce Type: new Abstract: Most evidence on the effectiveness of explainable artificial intelligence (XAI) attribution methods has been established on convolutional neural network

Does Machine 'know' interpersonal pragmatics? Evidence from MARBERT's learning of emoji pragmatics in Arabic digital discourse

TutorialsDGX agent

arXiv:2608.01174v1 Announce Type: new Abstract: This study examines Transformer-based models' ability to learn emoji pragmatics in Arabic digital discourse (ADD), providing evidence from MARBERT's beh

Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result

ApplicationsDGX agent

arXiv:2608.01559v1 Announce Type: cross Abstract: Adversarial self-play is an appealing recipe for legal reasoning: have a student model draft an argument, have an adversary attack it, and reward the

Domain-Generalized Adaptive Semantic Communication for Collaborative Perception

SafetyDGX agent

arXiv:2608.00056v1 Announce Type: cross Abstract: We propose RSTA, a domain-generalized semantic communication framework enabling source-free V2X collaborative perception under both observation-domain

Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

Model ReleasesDGX agent

arXiv:2608.02235v1 Announce Type: new Abstract: Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages. However

Dominant Arm Identification with Mixing and Recycling Observed Samples

ResearchDGX agent

arXiv:2608.01545v1 Announce Type: cross Abstract: We study the problem of identifying the dominant arm in multi-armed bandits, where the objective is to find the action with the highest probability of

Don't Judge a Book by its Cover: Testing LLMs' Robustness Under Logical Obfuscation

Model ReleasesDGX agent

arXiv:2602.01132v2 Announce Type: replace Abstract: Tasks such as solving arithmetic equations, evaluating truth tables, and completing syllogisms are handled well by large language models (LLMs) in t

Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale

ApplicationsDGX agent

arXiv:2608.01050v1 Announce Type: cross Abstract: Production LLM agents that select from large skill libraries face a limitation that semantic relevance alone cannot resolve: a skill may match a user'

Douyin Multimodal Embedding Model Technical Report

ApplicationsDGX agent

arXiv:2608.02148v1 Announce Type: cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search

DPL: Depth-only Perceptive Humanoid Locomotion via Realistic Depth Synthesis and Cross-Attention Terrain Reconstruction

SafetyDGX agent

arXiv:2510.07152v3 Announce Type: replace Abstract: Recent advancements in legged robot perceptive locomotion have shown promising progress. However, terrain-aware humanoid locomotion remains largely

DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable

Model ReleasesDGX agent

arXiv:2608.00548v1 Announce Type: new Abstract: Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their ras

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

ResearchDGX agent

arXiv:2608.00486v1 Announce Type: new Abstract: Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: a

DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation

SafetyDGX agent

arXiv:2608.01381v1 Announce Type: new Abstract: Mobile manipulation requires a robot to coordinate base and arm motion under continuously changing viewpoints and contact conditions, within an action s

DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving

AgentsDGX agent

arXiv:2603.00919v3 Announce Type: replace Abstract: Large language models (LLMs) have shown great promise for autonomous driving. However, discretizing numbers into tokens limits precise numerical rea

Driver2Map: Imitating Human Driving for Online High-Definition Map Construction

SafetyDGX agent

arXiv:2608.01338v1 Announce Type: new Abstract: High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition

DSETA: A Dual-Stage Continual Learning Framework for Travel Time Prediction in Dynamic Traffic Environments

ApplicationsDGX agent

arXiv:2608.00402v1 Announce Type: new Abstract: Estimated Time of Arrival (ETA) prediction is a core component of intelligent transportation systems. As traffic congestion patterns become increasingly

DS@GT ARC at MEDIQA-CORE-Task-1 2026: Trimodal Model Fusion with Task-Specific Gates for Brain Tumor Subtype Classification

Model ReleasesDGX agent

arXiv:2608.00086v1 Announce Type: new Abstract: Brain tumor diagnosis is a time-sensitive process in which patients may wait weeks for a finalized pathology report. This problem motivates automated sy

DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation

Local AiDGX agent

arXiv:2608.02495v1 Announce Type: new Abstract: Despite the remarkable progress over the past decades, accurately identifying small objects remains challenging because of their insufficient visual cue

DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction

AgentsDGX agent

arXiv:2608.01178v1 Announce Type: new Abstract: We present DynActiveGS, a dynamic-aware active reconstruction framework based on 3D Gaussian Splatting (3DGS) for autonomous exploration in dynamic envi

Dynamic Resolution Routing for Efficient Egocentric Grounding

Local AiDGX agent

arXiv:2608.01638v1 Announce Type: new Abstract: Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain

Dynamic UAV-based search operations using probabilistic diffusion modeling of Man Overboard incident victims

ResearchDGX agent

arXiv:2608.02093v1 Announce Type: new Abstract: More than 70% of the people that fell overboard cruise ships in the period 2010-2019 lost their lives. This paper presents a strategy for reliably predi

DynamicManip: Enabling Dynamic Manipulation from a Single Static Demonstration

Model ReleasesDGX agent

arXiv:2608.01452v1 Announce Type: cross Abstract: Dynamic manipulation is a critical capability for robots operating in complex and dynamic environments, where robots must interact with objects that a

DynamicWAM: Dual-Path Motion Conditioning for World-Action Models in Dynamic Manipulation

ApplicationsDGX agent

arXiv:2608.00793v1 Announce Type: new Abstract: Dynamic manipulation requires robots to infer target motion and respond promptly, yet existing World-Action Models (WAMs) typically condition only on th

DynImmune-BERT: Dynamic Immune Repertoire Modeling with Neural ODE Driven Continuous Transformers

Model ReleasesDGX agent

arXiv:2607.17244v2 Announce Type: replace Abstract: Longitudinal T cell receptor repertoires contain signals of clonal expansion, contraction, disappearance, and reappearance after immune perturbation

E2Pano: Learning Event-to-Panorama Image Reconstruction

Model ReleasesDGX agent

arXiv:2608.00694v1 Announce Type: new Abstract: Event cameras offer microsecond-level temporal resolution and high dynamic range, potentially facilitating motion-blur-free panoramic imaging from fast

early sparks of rsi? Ali Taha @waterloo_intern explains how http://Z.ai's GLM-5.2 (@Zai_org) profiled its SGLang serving path and rewrote bo…

HardwareDGX agent

early sparks of rsi? Ali Taha @waterloo_intern explains how http://Z.ai's GLM-5.2 (@Zai_org) profiled its SGLang serving path and rewrote bottlenecked GPU kernels, (still lacks reliable judgment) grea

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation

Model ReleasesDGX agent

arXiv:2608.02474v1 Announce Type: new Abstract: Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its infer

EEG-FM-Compass: Progress, Benchmarking, and Future Directions for EEG Foundation Models

Model ReleasesDGX agent

arXiv:2601.17883v3 Announce Type: replace-cross Abstract: Electroencephalography (EEG) foundation models (FMs) have recently emerged as a promising paradigm for brain-computer interfaces, aiming to le

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs

Model ReleasesDGX agent

arXiv:2511.00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work sho

Efficient nonlinear flame response modeling for propulsion thermoacoustic analysis using limited numerical data

Local AiDGX agent

arXiv:2409.05885v2 Announce Type: replace Abstract: Characterizing nonlinear flame response is critical for predicting thermoacoustic instabilities in propulsion combustors, yet obtaining a comprehens

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

ResearchDGX agent

arXiv:2608.02580v1 Announce Type: new Abstract: Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next

Model ReleasesDGX agent

arXiv:2603.12147v2 Announce Type: replace Abstract: Egocentric video provides a natural modality for studying human behavior, but conventional visual understanding captures mainly observable scenes, o

EHR2Path: Comprehensive Pathway-Level Modeling of Longitudinal Patient Trajectories from Multimodal Electronic Health Records

ResearchDGX agent

arXiv:2506.04831v3 Announce Type: replace-cross Abstract: Forecasting how a patient's condition is likely to evolve, including possible deterioration, recovery, treatment needs, and care transitions,

Eigenvalues as a Metric for Memory Dynamics in Sequence Models

ResearchDGX agent

arXiv:2510.09379v2 Announce Type: replace Abstract: While softmax attention drives state-of-the-art performance in sequence modeling, its quadratic complexity motivates linear alternatives such as sta

ELECTRIC: Evidential Learning-Enhanced CT Reconstruction via Iterative Correction

ResearchDGX agent

arXiv:2608.00060v1 Announce Type: new Abstract: Here we introduce ELECTRIC (Evidential Learning-Enhanced CT Reconstruction via Iterative Correction), a physics-guided Bayesian formulation. An evidenti

Element-Aware Group Learning for E-Commerce Image Generation

SafetyDGX agent

arXiv:2608.00584v1 Announce Type: new Abstract: Recent advances in image generation and editing have made prompt quality a key bottleneck for e-commerce creatives. Vision-language models (VLMs) can ge

← Previous
1…112113114115116…1410
Next →