AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
5 May 2026

Submodular Benchmark Selection

Model ReleasesDGX agent

arXiv:2605.02209v1 Announce Type: cross Abstract: Evaluating large language models across many benchmarks is expensive, yet many benchmarks are highly correlated. We formalize the selection of a small

SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

Model ReleasesDGX agent

arXiv:2508.15658v5 Announce Type: replace Abstract: The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show pr

Time-series forecasting through the lens of dynamics

TutorialsDGX agent

arXiv:2507.15774v2 Announce Type: replace Abstract: While deep learning is facing an homogenization across modalities led by Transformers, they are still challenged by shallow linear models in the tim

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Toward a Scientific Discovery Engine for Weather and Climate Data: A Visual Analytics Workbench for Embedding-Based Exploration

SafetyDGX agent

arXiv:2605.00972v1 Announce Type: cross Abstract: Earth system science is producing increasingly large, high-dimensional datasets from physics based Earth system models to AI-based weather and climate

Unsupervised full-field Bayesian inference of orthotropic hyperelasticity from a single biaxial test: a myocardial case study

Model ReleasesDGX agent

arXiv:2510.09498v3 Announce Type: replace-cross Abstract: Cardiac muscle tissue exhibits highly non-linear hyperelastic and orthotropic material behavior during passive deformation. Traditional consti

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

Model ReleasesDGX agent

arXiv:2605.02735v1 Announce Type: new Abstract: Continuous latent-space reasoning offers a compact alternative to textual chain-of-thought for multimodal models, enabling high-dimensional visual evide

4 May 2026

AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs

Model ReleasesDGX agent

arXiv:2605.00539v1 Announce Type: new Abstract: Quantization is a key method for reducing the GPU memory requirement of training large language models (LLMs). Yet, current approaches are ineffective f

BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social Interaction

Model ReleasesDGX agent

arXiv:2603.15949v3 Announce Type: replace Abstract: Large Language Models have demonstrated strong multilingual fluency, yet fluency alone does not guarantee socially appropriate language use. In high

Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs

Model ReleasesDGX agent

arXiv:2407.10853v5 Announce Type: replace Abstract: Bias and fairness risks in Large Language Models (LLMs) vary substantially across deployment contexts, yet existing approaches lack systematic guida

Can Coding Agents Reproduce Findings in Computational Materials Science?

Model ReleasesDGX agent

arXiv:2605.00803v1 Announce Type: cross Abstract: Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering be

ControBench: An Interaction-Aware Benchmark for Controversial Discourse Analysis on Social Networks

Model ReleasesDGX agent

arXiv:2605.00513v1 Announce Type: new Abstract: Understanding how people argue across ideological divides online is important for studying political polarization, misinformation, and content moderatio

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

SafetyDGX agent

arXiv:2506.00166v2 Announce Type: replace-cross Abstract: Existing paradigms for ensuring AI safety, such as guardrail models and alignment training, often compromise either inference efficiency or de

Documentation https://docs.ollama.com/integrations/claude-desktop

Model ReleasesDGX agent

This documentation page describes how to integrate Ollama with Claude Desktop, enabling users to run local language models through the Anthropic Claude interface. The integration allows Claude Desktop

Granite 4.1 3B SVG Pelican Gallery

Model ReleasesDGX agent

Granite 4.1 3B SVG Pelican Gallery IBM released their Granite 4.1 family of LLMs a few days ago. They're Apache 2.0 licensed and come in 3B, 8B and 30B sizes. Granite 4.1 LLMs: How They’re Built by Gr

How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses

Model ReleasesDGX agent

arXiv:2605.00113v1 Announce Type: new Abstract: We examine if frontier chat-based large language models (LLMs) adjust their outputs based on neurodivergence (ND) context in system prompts and describe

LWiAI Podcast #243 - GPT 5.5, DeepSeek V4, AI safety sabotage

Model ReleasesDGX agent

This podcast episode from Last Week in AI discusses recent developments in large language models, including updates on GPT 5.5 and DeepSeek V4, while also covering concerns about potential sabotage or

Make Your LVLM KV Cache More Lightweight

Model ReleasesDGX agent

arXiv:2605.00789v1 Announce Type: new Abstract: Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency

Minimizing Human Intervention in Online Classification

Model ReleasesDGX agent

arXiv:2510.23557v2 Announce Type: replace-cross Abstract: Training or fine-tuning large language model (LLM)-based systems often requires costly human feedback, yet there is limited understanding of h

PEACE: Cross-modal Enhanced Pediatric-Adult ECG Alignment for Robust Pediatric Diagnosis

Model ReleasesDGX agent

arXiv:2605.00647v1 Announce Type: new Abstract: Automated pediatric electrocardiogram (ECG) diagnosis remains challenging because models trained predominantly on adult data suffer from substantial cro

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs

Model ReleasesDGX agent

arXiv:2605.00814v1 Announce Type: new Abstract: While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a 'Visual Signal Dilution' p

Reasoning-Intensive Regression

Model ReleasesDGX agent

arXiv:2508.21762v3 Announce Type: replace Abstract: AI researchers and practitioners increasingly apply large language models (LLMs) to what we call reasoning-intensive regression (RiR), i.e., deducin

RouteProfile: Elucidating the Design Space of LLM Profiles for Routing

ResearchDGX agent

arXiv:2605.00180v1 Announce Type: cross Abstract: As the large language model (LLM) ecosystem expands, individual models exhibit varying capabilities across queries, benchmarks, and domains, motivatin

Smart Profit-Aware Crop Advisory System: Kisan AI

Model ReleasesDGX agent

arXiv:2605.00133v1 Announce Type: new Abstract: Modern crop advisory systems exhibit a critical limitation termed extit{economic blindness}. These systems primarily optimize for biological yield, ofte

Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium

SafetyDGX agent

arXiv:2503.10990v2 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with diverse human preferences is critical for ensuring fairness and informed outcomes when deploying th

The Power of Order: Fooling LLMs with Adversarial Table Permutations

ApplicationsDGX agent

arXiv:2605.00445v1 Announce Type: new Abstract: Large Language Models have achieved remarkable success and are increasingly deployed in critical applications involving tabular data, such as Table Ques

Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions

Model ReleasesDGX agent

arXiv:2605.00226v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly tasked with strategic decision-making under incomplete information, such as in negotiation and policymakin

3 May 2026

I am not sure I would agree with all of this, but the relationship between Anthropic and Claude is quite different than the relationship bet…

Model ReleasesDGX agent

I am not sure I would agree with all of this, but the relationship between Anthropic and Claude is quite different than the relationship between other labs and their models. And that shows up in lots

Its getting hard to benchmark frontier agent performance on longer tasks. Repeated measurement is very expensive and there are differences b…

Model ReleasesDGX agent

Its getting hard to benchmark frontier agent performance on longer tasks. Repeated measurement is very expensive and there are differences between using models in harnesses versus via APIs. I suspect

Public interpretability dataset and benchmark library for a novel transformer architecture [R]

Model ReleasesDGX agent

This work presents an explainability library for transformer models that provides tools for understanding transformer behavior through attributions and concept-based explanations . The resource likely

SULPHUR 2 RELEASED

Model ReleasesDGX agent

Stable Diffusion 2.0 is an open-source text-to-image model that includes improved text-to-image capabilities using the OpenCLIP encoder, generating higher quality images at resolutions of 512x512 and

2 May 2026

(Sorry, after seeing so many of these, could not resist): 🚨 BREAKING: Google just dropped a NEW paper that completely deletes RNNs from exi…

Model ReleasesDGX agent

(Sorry, after seeing so many of these, could not resist): 🚨 BREAKING: Google just dropped a NEW paper that completely deletes RNNs from existence. No recurrence. No convolutions. Nothing. Just one mec

1 May 2026

BatteryPass-12K: The First Dataset for the Novel Digital Battery Passport Conformance Task

Model ReleasesDGX agent

arXiv:2604.26986v1 Announce Type: new Abstract: We introduce a novel task of digital battery passport (DBP) conformance classification and introduce the first public benchmark for the task: BatteryPas

Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows

Model ReleasesDGX agent

arXiv:2604.28139v1 Announce Type: cross Abstract: LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks

DEFault++: Automated Fault Detection, Categorization, and Diagnosis for Transformer Architectures

Model ReleasesDGX agent

arXiv:2604.28118v1 Announce Type: cross Abstract: Transformer models are widely deployed in critical AI applications, yet faults in their attention mechanisms, projections, and other internal componen

Exploring Interaction Paradigms for LLM Agents in Scientific Visualization

Model ReleasesDGX agent

arXiv:2604.27996v1 Announce Type: new Abstract: This paper examines how different types of large language model (LLM) agents perform on scientific visualization (SciVis) tasks, where users generate vi

From Test-taking to Cognitive Scaffolding: A Pedagogical Diagnostic Benchmark for LLMs on English Standardized Tests

Model ReleasesDGX agent

arXiv:2505.17056v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly integrated into educational tools, current evaluations on standardized tests predominantly fo

Generative Human Geometry Distribution

ResearchDGX agent

arXiv:2503.01448v5 Announce Type: replace Abstract: Realistic human geometry generation is an important yet challenging task, requiring both the preservation of fine clothing details and the accurate

GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs

Model ReleasesDGX agent

arXiv:2603.25385v2 Announce Type: replace-cross Abstract: Quantization techniques such as BitsAndBytes, AWQ, and GPTQ are widely used as a standard method in deploying large language models but often

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions

Model ReleasesDGX agent

arXiv:2604.27763v1 Announce Type: new Abstract: The emergence of Large Language Models (LLMs) offers a transformative interface for Web3, yet existing benchmarks fail to capture the complexity of tran

InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?

Model ReleasesDGX agent

arXiv:2604.27419v1 Announce Type: new Abstract: With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent

Making Logic a First-Class Citizen in Generative ML for Networking

ApplicationsDGX agent

arXiv:2506.23964v3 Announce Type: replace-cross Abstract: Generative ML models are increasingly popular in networking for tasks such as telemetry imputation, prediction, and synthetic trace generation

MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents

Model ReleasesDGX agent

arXiv:2604.27819v1 Announce Type: new Abstract: Multi-server MCP agents create an information-flow control problem: faithful tool composition can turn individually benign read/write permissions into c

ORFS-agent: Tool-Using Agents for Chip Design Optimization

Model ReleasesDGX agent

arXiv:2506.08332v3 Announce Type: replace Abstract: Machine learning has been widely used to optimize complex engineering workflows across numerous domains. In integrated circuit design, modern flows

Position-Aware Drafting for Inference Acceleration in LLM-Based Generative List-Wise Recommendation

ApplicationsDGX agent

arXiv:2604.27747v1 Announce Type: cross Abstract: Large language model (LLM)-based generative list-wise recommendation has advanced rapidly, but decoding remains sequential and thus latency-prone. To

Post-Optimization Adaptive Rank Allocation for LoRA

Model ReleasesDGX agent

arXiv:2604.27796v1 Announce Type: new Abstract: Exponential growth in the scale of modern foundation models has led to the widespread adoption of Low-Rank Adaptation (LoRA) as a parameter-efficient fi

Predicting Covariate-Driven Spatial Deformation for Nonstationary Gaussian Processes

ApplicationsDGX agent

arXiv:2604.27280v1 Announce Type: new Abstract: Nonstationary Gaussian processes (GPs) are essential for modeling complex, locally heterogeneous spatial data. A common modeling approach is the spatial

Reinforced Agent: Inference-Time Feedback for Tool-Calling Agents

Model ReleasesDGX agent

arXiv:2604.27233v1 Announce Type: new Abstract: Tool-calling agents are evaluated on tool selection, parameter accuracy, and scope recognition, yet LLM trajectory assessments remain inherently post-ho

Sequential Inference for Gaussian Processes: A Signal Processing Perspective

TutorialsDGX agent

arXiv:2604.28163v1 Announce Type: cross Abstract: The proliferation of capable and efficient machine learning (ML) models marks one of the strongest methodological shifts in signal processing (SP) in

Step-level Optimization for Efficient Computer-use Agents

Model ReleasesDGX agent

arXiv:2604.27151v1 Announce Type: new Abstract: Computer-use agents provide a promising path toward general software automation because they can interact directly with arbitrary graphical user interfa

30 Apr 2026

AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents

Model ReleasesDGX agent

arXiv:2603.16496v2 Announce Type: replace Abstract: Large language model (LLM) agents increasingly rely on external memory to support long-horizon interaction, personalized assistance, and multi-step

Associative-State Universal Transformers: Sparse Retrieval Meets Structured Recurrence

Model ReleasesDGX agent

arXiv:2604.25930v1 Announce Type: new Abstract: We study whether a structured recurrent state can serve as a compact associative backbone for language modeling while still supporting exact retrieval.

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch

ResearchDGX agent

arXiv:2601.13606v2 Announce Type: replace Abstract: Chart reasoning is a critical capability for Vision Language Models (VLMs). However, the development of open-source models is severely hindered by t

ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation

Model ReleasesDGX agent

arXiv:2604.26923v1 Announce Type: cross Abstract: LLMs have achieved strong results on both function-level code synthesis and repository-level code modification, yet a capability that falls between th

Demis says he wants to see a Western open source AI stack and that we’re losing to China. He also says Google doesn’t have enough compute to…

Model ReleasesDGX agent

Demis says he wants to see a Western open source AI stack and that we’re losing to China. He also says Google doesn’t have enough compute to build two frontier (open and closed) models, which is why G

EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation

SafetyDGX agent

arXiv:2604.26170v1 Announce Type: new Abstract: Adapting large language models (LLMs) to a targeted task efficiently and effectively remains a fundamental challenge. Such adaptation often requires ite

FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving

ResearchDGX agent

arXiv:2604.26881v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models offer high capacity with efficient inference cost by activating a small subset of expert models per input. However, de

Human-in-the-Loop Benchmarking of Heterogeneous LLMs for Automated Competency Assessment in Secondary Level Mathematics

Model ReleasesDGX agent

arXiv:2604.26607v1 Announce Type: new Abstract: As Competency-Based Education (CBE) is gaining traction around the world, the shift from marks-based assessment to qualitative competency mapping is a m

HumanOmni-Speaker: Identifying Who said What and When

Model ReleasesDGX agent

arXiv:2603.21664v2 Announce Type: replace Abstract: While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human intera

Learning Neural Operator Surrogates for the Black Hole Accretion Code

Model ReleasesDGX agent

arXiv:2604.25985v1 Announce Type: cross Abstract: General-relativistic magnetohydrodynamic (GR-MHD) simulations are essential for studying black hole accretion, relativistic jets, and magnetic reconne

Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection

ApplicationsDGX agent

arXiv:2603.04337v2 Announce Type: replace-cross Abstract: Constructing computer-aided design (CAD) models is labor-intensive but essential for engineering and manufacturing. Recent advances in Large L

← Previous
1…358359360361362…1044
Next →