AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

87,573Total entries
1Added by human
87,572Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,952 results
1 May 2026

GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs

Model ReleasesDGX agent

arXiv:2603.25385v2 Announce Type: replace-cross Abstract: Quantization techniques such as BitsAndBytes, AWQ, and GPTQ are widely used as a standard method in deploying large language models but often

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions

Model ReleasesDGX agent

arXiv:2604.27763v1 Announce Type: new Abstract: The emergence of Large Language Models (LLMs) offers a transformative interface for Web3, yet existing benchmarks fail to capture the complexity of tran

InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.27419v1 Announce Type: new Abstract: With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent

Making Logic a First-Class Citizen in Generative ML for Networking

ApplicationsDGX agent

arXiv:2506.23964v3 Announce Type: replace-cross Abstract: Generative ML models are increasingly popular in networking for tasks such as telemetry imputation, prediction, and synthetic trace generation

MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

Model ReleasesDGX agent

arXiv:2604.27819v1 Announce Type: new Abstract: Multi-server MCP agents create an information-flow control problem: faithful tool composition can turn individually benign read/write permissions into c

ORFS-agent: Tool-Using Agents for Chip Design Optimization

Model ReleasesDGX agent

arXiv:2506.08332v3 Announce Type: replace Abstract: Machine learning has been widely used to optimize complex engineering workflows across numerous domains. In integrated circuit design, modern flows

Position-Aware Drafting for Inference Acceleration in LLM-Based Generative List-Wise Recommendation

ApplicationsDGX agent

arXiv:2604.27747v1 Announce Type: cross Abstract: Large language model (LLM)-based generative list-wise recommendation has advanced rapidly, but decoding remains sequential and thus latency-prone. To

Post-Optimization Adaptive Rank Allocation for LoRA

Model ReleasesDGX agent

arXiv:2604.27796v1 Announce Type: new Abstract: Exponential growth in the scale of modern foundation models has led to the widespread adoption of Low-Rank Adaptation (LoRA) as a parameter-efficient fi

Predicting Covariate-Driven Spatial Deformation for Nonstationary Gaussian Processes

ApplicationsDGX agent

arXiv:2604.27280v1 Announce Type: new Abstract: Nonstationary Gaussian processes (GPs) are essential for modeling complex, locally heterogeneous spatial data. A common modeling approach is the spatial

Reinforced Agent: Inference-Time Feedback for Tool-Calling Agents

Model ReleasesDGX agent

arXiv:2604.27233v1 Announce Type: new Abstract: Tool-calling agents are evaluated on tool selection, parameter accuracy, and scope recognition, yet LLM trajectory assessments remain inherently post-ho

Sequential Inference for Gaussian Processes: A Signal Processing Perspective

TutorialsDGX agent

arXiv:2604.28163v1 Announce Type: cross Abstract: The proliferation of capable and efficient machine learning (ML) models marks one of the strongest methodological shifts in signal processing (SP) in

Step-level Optimization for Efficient Computer-use Agents

Model ReleasesDGX agent

arXiv:2604.27151v1 Announce Type: new Abstract: Computer-use agents provide a promising path toward general software automation because they can interact directly with arbitrary graphical user interfa

30 Apr 2026

AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents

Model ReleasesDGX agent

arXiv:2603.16496v2 Announce Type: replace Abstract: Large language model (LLM) agents increasingly rely on external memory to support long-horizon interaction, personalized assistance, and multi-step

Associative-State Universal Transformers: Sparse Retrieval Meets Structured Recurrence

Model ReleasesDGX agent

arXiv:2604.25930v1 Announce Type: new Abstract: We study whether a structured recurrent state can serve as a compact associative backbone for language modeling while still supporting exact retrieval.

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch

ResearchDGX agent

arXiv:2601.13606v2 Announce Type: replace Abstract: Chart reasoning is a critical capability for Vision Language Models (VLMs). However, the development of open-source models is severely hindered by t

ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation

Model ReleasesDGX agent

arXiv:2604.26923v1 Announce Type: cross Abstract: LLMs have achieved strong results on both function-level code synthesis and repository-level code modification, yet a capability that falls between th

Demis says he wants to see a Western open source AI stack and that we’re losing to China. He also says Google doesn’t have enough compute to…

Model ReleasesDGX agent

Demis says he wants to see a Western open source AI stack and that we’re losing to China. He also says Google doesn’t have enough compute to build two frontier (open and closed) models, which is why G

EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation

SafetyDGX agent

arXiv:2604.26170v1 Announce Type: new Abstract: Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge. Such adaptation often requires ite

FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving

ResearchDGX agent

arXiv:2604.26881v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models offer high capacity with efficient inference cost by activating a small subset of expert models per input. However, de

Human-in-the-Loop Benchmarking of Heterogeneous LLMs for Automated Competency Assessment in Secondary Level Mathematics

Model ReleasesDGX agent

arXiv:2604.26607v1 Announce Type: new Abstract: As Competency-Based Education (CBE) is gaining traction around the world, the shift from marks-based assessment to qualitative competency mapping is a m

HumanOmni-Speaker: Identifying Who said What and When

Model ReleasesDGX agent

arXiv:2603.21664v2 Announce Type: replace Abstract: While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human intera

Learning Neural Operator Surrogates for the Black Hole Accretion Code

Model ReleasesDGX agent

arXiv:2604.25985v1 Announce Type: cross Abstract: General-relativistic magnetohydrodynamic (GR-MHD) simulations are essential for studying black hole accretion, relativistic jets, and magnetic reconne

Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection

ApplicationsDGX agent

arXiv:2603.04337v2 Announce Type: replace-cross Abstract: Constructing computer-aided design (CAD) models is labor-intensive but essential for engineering and manufacturing. Recent advances in Large L

Privacy-Preserving Federated Learning Framework for Distributed Chemical Process Optimization

Local AiDGX agent

arXiv:2604.26073v1 Announce Type: cross Abstract: Industrial chemical plants often operate under strict data confidentiality constraints, making centralized data-driven process modeling difficult. Fed

Reasoning Gets Harder for LLMs Inside A Dialogue

Model ReleasesDGX agent

arXiv:2603.20133v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that d

SciMDR: Advancing Scientific Multimodal Document Reasoning

Model ReleasesDGX agent

arXiv:2603.12249v2 Announce Type: replace-cross Abstract: Constructing scientific multimodal document reasoning datasets for foundation model training involves an inherent trade-off among scale, faith

The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences

Model ReleasesDGX agent

arXiv:2509.11295v2 Announce Type: replace Abstract: Developing effective prompts demands significant cognitive investment to generate reliable, high-quality responses from Large Language Models (LLMs)

Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall

ResearchDGX agent

arXiv:2505.13963v3 Announce Type: replace Abstract: Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's

tweeted about this yesterday and Cursor already dropped the alpha today! 🚀very cool to see how us, them, and others have converged on good …

Model ReleasesDGX agent

tweeted about this yesterday and Cursor already dropped the alpha today! 🚀very cool to see how us, them, and others have converged on good design patterns in Agent + Harness Engineering: 1. Tuning dif

Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness

Model ReleasesDGX agent

arXiv:2512.03992v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are essential for embodied AI and safety-critical applications, such as robotics and autonomous systems. However

29 Apr 2026

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

Model ReleasesDGX agent

arXiv:2604.25850v1 Announce Type: new Abstract: Harnesses have become a central determinant of coding-agent performance, shaping how models interact with repositories, tools, and execution environment

BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate

SafetyDGX agent

arXiv:2604.25203v1 Announce Type: new Abstract: Deploying guardrails for custom policies remains challenging, as generic safety models fail to capture task-specific requirements, while prompting LLMs

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding

Model ReleasesDGX agent

arXiv:2512.12087v3 Announce Type: replace Abstract: The growing demand for long-context inference capabilities in Large Language Models (LLMs) has intensified the computational and memory bottlenecks

Conditional Flow Matching for Probabilistic Downscaling of Maximum 3-day Snowfall in Alaska

ResearchDGX agent

arXiv:2604.25172v1 Announce Type: cross Abstract: Precipitation in complex terrain is governed by orographic processes operating at scales of a few kilometers, yet climate models typically run at reso

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

TutorialsDGX agent

arXiv:2604.25477v1 Announce Type: new Abstract: Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance t

DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference

ResearchDGX agent

arXiv:2510.19669v4 Announce Type: replace Abstract: Recent reasoning Large Language Models (LLMs) demonstrate remarkable problem-solving abilities but often generate long thinking traces whose utility

Doing More With Less: Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling

Model ReleasesDGX agent

arXiv:2604.25098v1 Announce Type: cross Abstract: While current Large Language Models (LLMs) exhibit remarkable reasoning capabilities through test-time compute scaling (TTS), their massive parameter

FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices

Model ReleasesDGX agent

arXiv:2604.25421v1 Announce Type: new Abstract: Federated fine-tuning provides a practical route to adapt large language models (LLMs) on edge devices without centralizing private data, yet in mobile

Golden RPG: Confidence-Adaptive Region-Aware Noise for Compositional Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2604.25314v1 Announce Type: new Abstract: Compositional text-to-image (T2I) generation requires a model to honour multiple sub-prompts that describe distinct image regions. Recent work shows tha

I released LLM 0.32a0 this morning, a major backwards-compatible refactor of my LLM Python library and CLI tool for working with language mo…

Model ReleasesDGX agent

I released LLM 0.32a0 this morning, a major backwards-compatible refactor of my LLM Python library and CLI tool for working with language models - the new changes should help LLM work better with reas

Improving LLM Predictions via Inter-Layer Structural Encoders

Model ReleasesDGX agent

arXiv:2603.22665v2 Announce Type: replace Abstract: The standard practice in Large Language Models (LLMs) is to base predictions on final-layer representations. However, intermediate layers encode com

Limited Linguistic Diversity in Embodied AI Datasets

ResearchDGX agent

arXiv:2601.03136v2 Announce Type: replace Abstract: Language plays a critical role in Vision-Language-Action (VLA) models, yet the linguistic characteristics of the datasets used to train and evaluate

LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation

Model ReleasesDGX agent

arXiv:2604.25665v1 Announce Type: new Abstract: Reliable evaluation of large language model (LLM)-generated summaries remains an open challenge, particularly across heterogeneous domains and document

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition

Model ReleasesDGX agent

arXiv:2512.07348v2 Announce Type: replace Abstract: In controllable image generation, synthesizing coherent and consistent images from multiple reference inputs, i.e., Multi-Image Composition (MICo),

Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study

Model ReleasesDGX agent

arXiv:2602.17262v2 Announce Type: replace Abstract: Human self-report questionnaires are increasingly used in NLP to benchmark and audit large language models (LLMs), from persona consistency to safet

RCProb: Probabilistic Rule Extraction for Efficient Simplification of Tree Ensembles

Model ReleasesDGX agent

arXiv:2604.25304v1 Announce Type: new Abstract: Tree ensembles are widely used in industrial machine learning due to their strong predictive performance and efficient training procedures. However, as

Relational In-Context Learning via Synthetic Pre-training with Structural Prior

ApplicationsDGX agent

arXiv:2603.03805v2 Announce Type: replace Abstract: Relational Databases (RDBs) are the backbone of modern business, yet they lack foundation models comparable to those in text or vision. A key obstac

ReSim: Reliable World Simulation for Autonomous Driving

SafetyDGX agent

arXiv:2506.09981v2 Announce Type: replace Abstract: How can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusivel

Sensitivity-Based Tube NMPC for Cooperative Aerial Structures Under Parametric Uncertainty

ResearchDGX agent

arXiv:2604.25766v1 Announce Type: new Abstract: This paper presents a sensitivity-based tube Nonlinear Model Predictive Control (NMPC) framework for cooperative aerial chains under bounded parametric

TopoMamba: Topology-Aware Scanning and Fusion for Segmenting Heterogeneous Medical Visual Media

ResearchDGX agent

arXiv:2604.25545v1 Announce Type: new Abstract: Visual state-space models (SSMs) have shown strong potential for medical image segmentation, yet their effectiveness is often limited by two practical i

When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs

Model ReleasesDGX agent

arXiv:2510.07499v2 Announce Type: replace Abstract: Recent Long-Context Language Models (LCLMs) can process hundreds of thousands of tokens in a single prompt, enabling new opportunities for knowledge

Xiaomi MiMo-V2.5-Pro achieves multiple breakthroughs in the latest Arena rankings (Apr 26, 2026) 🔥 🏆 Text Arena (Expert) — #6 globally | #…

ApplicationsDGX agent

Xiaomi MiMo-V2.5-Pro achieves multiple breakthroughs in the latest Arena rankings (Apr 26, 2026) 🔥 🏆 Text Arena (Expert) — #6 globally | #1 open-source model Also #1 among Chinese models, with Xiaomi

28 Apr 2026

Aligning with Your Own Voice: Self-Corrected Preference Learning for Hallucination Mitigation in LVLMs

SafetyDGX agent

arXiv:2604.24395v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) frequently suffer from hallucinations. Existing preference learning-based approaches largely rely on proprietary mo

amazing! it’s like talking to an ai from the past

AgentsDGX agent

amazing! it’s like talking to an ai from the past Announcing Talkie: a new, open-weight historical LLM! We trained and finetuned a 13B model on a newly-curated dataset of only pre-1930 data. Try it be

Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis

SafetyDGX agent

arXiv:2604.23072v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet t

AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation

Model ReleasesDGX agent

arXiv:2604.24086v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have been demonstrated possessing strong zero-shot generalization for robot control, their massive parameter

Benchmarking Source-Sensitive Reasoning in Turkish: Humans and LLMs under Evidential Trust Manipulation

Local AiDGX agent

arXiv:2604.24665v1 Announce Type: cross Abstract: This paper investigates whether source trustworthiness shapes Turkish evidential morphology and whether large language models (LLMs) track this sensit

Beyond Local vs. External: A Game-Theoretic Framework for Trustworthy Knowledge Acquisition

Model ReleasesDGX agent

arXiv:2604.23413v1 Announce Type: new Abstract: Cloud-hosted Large Language Models (LLMs) offer unmatched reasoning capabilities and dynamic knowledge, yet submitting raw queries to these external ser

BIR-Adapter: A parameter-efficient diffusion adapter for blind image restoration

Model ReleasesDGX agent

arXiv:2509.06904v3 Announce Type: replace Abstract: We introduce the BIR-Adapter, a parameter-efficient diffusion adapter for blind image restoration. Diffusion-based restoration methods have demonstr

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Model ReleasesDGX agent

arXiv:2604.23781v1 Announce Type: new Abstract: Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surroundi

← Previous
1…361362363364365…1050
Next →