AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
Human
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
59,295 results
11 Aug 2026

Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation

SafetyDGX agent

arXiv:2608.08494v1 Announce Type: new Abstract: Despite significant advances in Medical Report Generation (MRG), the reliability remains constrained by the prevalence of factual errors. While Direct P

LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding

Model ReleasesDGX agent

arXiv:2608.07596v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

Model ReleasesDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.09435v1 Announce Type: new Abstract: Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet exi

Listwise Cross-Encoder Fine-Tuning vs. Agentic Instruction Tuning for LLM Rerankers: A Systematic Study in Medical Procedure Reranking

Model ReleasesDGX agent

arXiv:2608.09650v1 Announce Type: cross Abstract: Reranking medical procedures against patient queries is a critical component of health insurance information retrieval, complicated by a substantial l

LITEWAY: LIghtweight HAR via Temporal Efficient highWAY

ResearchDGX agent

arXiv:2608.09421v1 Announce Type: cross Abstract: Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learning models on resource-limi

LLM-Based Embeddings for Program Analysis and Optimization

ResearchDGX agent

arXiv:2608.07894v1 Announce Type: new Abstract: Recent advances have highlighted the potential of machine learning, particularly Large Language Models (LLMs), for analyzing and optimizing programs. We

LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

Model ReleasesDGX agent

arXiv:2607.15509v2 Announce Type: replace-cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture desig

LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling

Model ReleasesDGX agent

arXiv:2608.09343v1 Announce Type: new Abstract: Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a black box: aggreg

LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing

SafetyDGX agent

arXiv:2608.08889v1 Announce Type: new Abstract: Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large L

LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs

Model ReleasesDGX agent

arXiv:2608.08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequentl

LLMs Remember First, Forget Last: Dual-Process Interference in Large Language Models

SafetyDGX agent

arXiv:2603.00270v3 Announce Type: replace-cross Abstract: Large language models can process millions of tokens, yet how they handle conflicting information within context remains poorly understood. Fr

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

Model ReleasesDGX agent

arXiv:2608.08382v1 Announce Type: new Abstract: As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Enabling fractio

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

Model ReleasesDGX agent

arXiv:2608.07943v1 Announce Type: new Abstract: Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's co

Logarithmic-Free Moment and Generalization Bounds for Uniformly Stable Algorithms

ResearchDGX agent

arXiv:2608.09870v1 Announce Type: cross Abstract: Uniform stability is a classical tool for controlling the generalization error of a learning algorithm. Bousquet, Klochkov, and Zhivotovskiy (2020) sh

LogicIF: Towards Complex Logic Instruction Following

Model ReleasesDGX agent

arXiv:2508.09125v4 Announce Type: replace Abstract: Instruction following has catalyzed the recent era of Large Language Models (LLMs) and is the foundational skill underpinning more advanced capabili

LogiShot: Logically Coherent Cross-Shot Video Generation

Model ReleasesDGX agent

arXiv:2608.08820v1 Announce Type: new Abstract: Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such

Long SKILL Compliance as Logical Reasoning: Closure-Grounded Detection with Scaling-Guided On-Policy Distillation

Model ReleasesDGX agent

arXiv:2608.08146v1 Announce Type: new Abstract: The increasing complexity of enterprise business scenarios has promoted the widespread adoption of long SKILL documents in agent systems, posing new cha

LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection

ResearchDGX agent

arXiv:2608.09723v1 Announce Type: new Abstract: Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models

ResearchDGX agent

arXiv:2607.16305v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (

LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection

ResearchDGX agent

arXiv:2608.09633v1 Announce Type: new Abstract: Face presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. While PAD methods achieve strong performance with

LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks

Model ReleasesDGX agent

arXiv:2608.07749v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a si

Loss-Resilient Wireless Video Token Communication over Block Fading Channels

ResearchDGX agent

arXiv:2608.08698v1 Announce Type: new Abstract: Video token communication represents video content as discrete tokens that differ in their importance to reconstruction and exhibit temporal dependencie

LUCID: Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation

SafetyDGX agent

arXiv:2608.07746v1 Announce Type: new Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making. Existing methods often c

Lying mirror using structured surfaces

ResearchDGX agent

arXiv:2410.15521v2 Announce Type: replace-cross Abstract: We introduce an all-optical system, termed the 'lying mirror', to hide input information by transforming it into misleading, ordinary-looking

M^3Prune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2511.19969v2 Announce Type: replace Abstract: Recent advancements in multi-modal retrieval-augmented generation (mRAG), which enhance multi-modal large language models (MLLMs) with external know

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

Local AiDGX agent

arXiv:2608.09819v1 Announce Type: cross Abstract: Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after de

Machine Learning and Data Analysis Using Posets: A Survey

SafetyDGX agent

arXiv:2404.03082v3 Announce Type: replace Abstract: Partially ordered sets (posets) are discrete mathematical structures that formalize the notion of comparison without forcing every pair of objects t

Machine-Learning-Based Diagnostic Framework for Passive Ultrasonic Detection of Railway Wheel Defects

SafetyDGX agent

arXiv:2608.08301v1 Announce Type: new Abstract: Reliable identification of railway wheel defects is important for safety and maintenance. This study develops a machine-learning-based diagnostic framew

Machine Learning for Inverse Problems and Data Assimilation

ResearchDGX agent

arXiv:2410.10523v3 Announce Type: replace-cross Abstract: The aim of this book is to demonstrate the potential for ideas in machine learning to impact on the fields of inverse problems and data assimi

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

Model ReleasesDGX agent

arXiv:2608.09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a real

MAGIC-SSCIL: Manifold Anchoring and Geometric Incremental Calibration for Semi-Supervised Class Incremental Learning

SafetyDGX agent

arXiv:2608.07586v1 Announce Type: new Abstract: Semi-supervised Class Incremental Learning (SSCIL) is a severe challenge for neural networks, and it is hardest in the exemplar-free setting where no pa

'Many Are My Names': The Anatomy of the Assistant and Its Personas via Sparse Autoencoders

ResearchDGX agent

arXiv:2608.07852v1 Announce Type: new Abstract: How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexp

MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning

SafetyDGX agent

arXiv:2608.09130v1 Announce Type: cross Abstract: Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but its required

Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking

Model ReleasesDGX agent

arXiv:2608.09613v1 Announce Type: new Abstract: Existing unified 4D reconstruction and point tracking approaches typically rely on heuristic interpolations or just predict at integer timestamps, lacki

MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures

Model ReleasesDGX agent

arXiv:2608.07556v1 Announce Type: cross Abstract: Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their original a

Mask-aware inference with State-Space Models

ApplicationsDGX agent

arXiv:2603.04568v2 Announce Type: replace Abstract: Many real-world computer vision tasks, such as depth completion, must handle inputs with arbitrarily shaped regions of missing or invalid data. For

Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training

Model ReleasesDGX agent

arXiv:2604.01499v2 Announce Type: replace Abstract: Evolution Strategies (ES) have emerged as a scalable gradient-free alternative to reinforcement learning based LLM fine-tuning, but it remains uncle

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.08176v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-disti

MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks

Model ReleasesDGX agent

arXiv:2507.03162v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These mod

Math-Vision Diagrams: A Comprehensive Benchmark for Evaluating LLM Mathematical Diagram Generation Capabilities

Model ReleasesDGX agent

arXiv:2608.08964v1 Announce Type: new Abstract: The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Models

MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models

Model ReleasesDGX agent

arXiv:2608.08503v1 Announce Type: new Abstract: Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) su

Matrix Completion via Nonsmooth Regularization of Fully Connected Neural Networks

ResearchDGX agent

arXiv:2403.10232v2 Announce Type: replace-cross Abstract: Conventional matrix completion methods approximate the missing values by assuming the matrix to be low-rank, which leads to a linear approxima

Matrix-free Neural Preconditioner for the Dirac Operator in Lattice Gauge Theory

Model ReleasesDGX agent

arXiv:2509.10378v2 Announce Type: replace-cross Abstract: Linear systems arise in generating samples and in calculating observables in lattice quantum chromodynamics~(QCD). Solving the Hermitian posit

Matryoshka Language Model Suites

Model ReleasesDGX agent

arXiv:2608.09703v1 Announce Type: new Abstract: Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inferen

Mawqif-v2: An Arabic Benchmark Dataset for Cross-Target Stance Detection

Model ReleasesDGX agent

arXiv:2608.09539v1 Announce Type: new Abstract: Publicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This pa

MaxModShift: Model Privacy via Designed Shifts

TutorialsDGX agent

arXiv:2608.09328v1 Announce Type: new Abstract: Model learning by an eavesdropper is treated as an estimation problem in a federated environment. The Fisher Information Matrix for the eavesdropper's e

MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks

Model ReleasesDGX agent

arXiv:2507.19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a s

MDB-Link: Hierarchical Schema Linking for Multi-Database Text-to-SQL

ResearchDGX agent

arXiv:2608.09588v1 Announce Type: new Abstract: Traditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, het

MeanSR: Restoration Trajectory Learning for One-Step Perceptual Super-Resolution

ApplicationsDGX agent

arXiv:2608.09405v1 Announce Type: new Abstract: Diffusion-based super-resolution (SR) achieves strong perceptual quality but requires costly iterative denoising. Existing one-step distillation methods

Measuring and Reducing WebGPU Dispatch Overhead for LLM Inference

ResearchDGX agent

arXiv:2608.08730v1 Announce Type: new Abstract: Large Language Models are deployed to multiple types of environments, from internet browsers to edge devices, and WebGPU serves as a modern cross-platfo

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

Model ReleasesDGX agent

arXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure doe

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

Model ReleasesDGX agent

arXiv:2608.09624v1 Announce Type: cross Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones.

Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision-Language Models for Centimeter-Level Flood Depth Estimation

Model ReleasesDGX agent

arXiv:2608.07562v1 Announce Type: new Abstract: Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estim

Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets

ResearchDGX agent

arXiv:2512.01045v2 Announce Type: replace Abstract: Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning

SafetyDGX agent

arXiv:2608.08623v1 Announce Type: new Abstract: In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated usi

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

Model ReleasesDGX agent

arXiv:2608.09818v1 Announce Type: cross Abstract: Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-

MELLON - Multimodal Enhanced LLM for Online Navigation

Model ReleasesDGX agent

arXiv:2608.09121v1 Announce Type: new Abstract: Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are either unimodal or

MemeMind: Reference-Guided Trace Construction for Offline Context Optimization

Model ReleasesDGX agent

arXiv:2608.09316v1 Announce Type: new Abstract: Offline context optimization improves an agent by revising its instructions and examples while keeping the model frozen. This approach learns from rollo

Memorization Dynamics in Knowledge Distillation for Language Models

Model ReleasesDGX agent

arXiv:2601.15394v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improv

Memory-Efficient Activation Checkpointing with Sliding Window and Hirschberg's Algorithm for 0/1 Knapsack Solving in PyTorch

Model ReleasesDGX agent

arXiv:2608.08740v1 Announce Type: new Abstract: Activation checkpointing minimizes the runtime of neural networks under a given memory budget, by selecting which intermediate tensors to store and whic

← Previous
1…2930313233…989
Next →