AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
27 May 2026

WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization

Model ReleasesDGX agent

arXiv:2605.26660v1 Announce Type: new Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in

26 May 2026

Acting on the Unseen: Communication-Free Collaborative Filtering for Decentralized Multi-Robot Task Allocation

Model ReleasesDGX agent

arXiv:2605.25584v1 Announce Type: cross Abstract: Multi-robot task allocation usually assumes some combination of communication, known task models, or a coordinator. We study the opposite extreme, a r

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2605.23974v1 Announce Type: new Abstract: Current language models create two safety challenges: risk must be detected early enough to avoid exposing harmful continuation, and the harmfulness its

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

Model ReleasesDGX agent

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digita

An Interactive Paradigm for Deep Research

Model ReleasesDGX agent

arXiv:2605.24266v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled deep research systems that synthesize comprehensive, report-style answers to open-ended q

BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning

ResearchDGX agent

arXiv:2511.12046v2 Announce Type: replace-cross Abstract: Knowledge Distillation (KD) is essential for compressing large models, yet relying on pre-trained 'teacher' models downloaded from third-party

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

Model ReleasesDGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?

Model ReleasesDGX agent

arXiv:2605.23913v1 Announce Type: cross Abstract: Cloud-hosted large language models (LLMs) commonly rely on LoRA for domain adaptation, yet domain data are distributed across multiple edge devices an

Chain-of-Thought Hijacking

Model ReleasesDGX agent

arXiv:2510.26418v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reas

ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale

Model ReleasesDGX agent

arXiv:2605.24305v1 Announce Type: cross Abstract: Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

Model ReleasesDGX agent

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Y

Committed SAE-Feature Traces for Audited-Session Substitution Detection in Hosted LLMs

Model ReleasesDGX agent

arXiv:2604.18179v2 Announce Type: replace-cross Abstract: Hosted-LLM providers have a silent-substitution incentive: advertise a stronger model while serving cheaper replies. Probe-after-return scheme

Context-Instrumental Data Distillation for Kubernetes Manifest Generation: Method and Experimental Evaluation

Model ReleasesDGX agent

arXiv:2605.25835v1 Announce Type: cross Abstract: This paper examines the specialization of Small Language Models (SLMs) with up to 4 billion parameters for generating artifacts in domain-specific lan

Continual Speaker Identity Unlearning with Minimal Interference

Model ReleasesDGX agent

arXiv:2605.25962v1 Announce Type: cross Abstract: Machine unlearning removes designated concepts or knowledge from pre-trained models. Recent work has extended this paradigm to speaker identity unlear

Counterfactual Explanations for Hypergraph Neural Networks

Model ReleasesDGX agent

arXiv:2602.04360v2 Announce Type: replace-cross Abstract: Hypergraph neural networks (HGNNs) effectively model higher-order interactions in many real-world systems but remain difficult to interpret, l

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy

ResearchDGX agent

arXiv:2605.25603v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves the problem-solving ability of large language models (LLMs), but generated reasoning traces may not faithfully

E = T*H/(O+B): A Dimensionless Control Parameter for Mixture-of-Experts Ecology

Model ReleasesDGX agent

arXiv:2605.06415v2 Announce Type: replace-cross Abstract: We introduce E = T*H/(O+B), a dimensionless control parameter that predicts whether Mixture-of-Experts (MoE) models will develop a healthy exp

ExplainReduce: Generating global explanations from many local explanations

Local AiDGX agent

arXiv:2502.10311v3 Announce Type: replace-cross Abstract: Most commonly used non-linear machine learning methods are closed-box models, uninterpretable to humans. The field of explainable artificial i

Fine-Tuning and Serving Gemma 4 31B on Google Cloud TPU: A Technical Comparison with GPU Baselines

Model ReleasesDGX agent

arXiv:2605.25645v1 Announce Type: cross Abstract: We present the first end-to-end demonstration of fine-tuning and serving Google's Gemma 4 31B model on TPU hardware, providing an empirical comparison

Fine-Tuning Masked Diffusion for Provable Self-Correction

ResearchDGX agent

arXiv:2510.01384v4 Announce Type: replace Abstract: A natural desideratum for generative models is self-correction--detecting and revising low-quality tokens at inference. While Masked Diffusion Model

Forgettable Federated Linear Learning with Certified Data Unlearning

ResearchDGX agent

arXiv:2306.02216v3 Announce Type: replace Abstract: Federated Learning (FL) enables collaborative model training across distributed clients while preserving user privacy. Recently, Federated Unlearnin

From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression

Model ReleasesDGX agent

arXiv:2605.24316v1 Announce Type: new Abstract: Scaling laws provide compact descriptions of how prediction error varies with compute, model size, and data, but existing theory mainly treats single-sa

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

Model ReleasesDGX agent

arXiv:2605.24636v1 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios re

Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers

Model ReleasesDGX agent

arXiv:2605.24518v1 Announce Type: cross Abstract: The quadratic complexity of self-attention in Transformer models remains a significant bottleneck for processing long sequences and deploying large la

How does Bayesian Sampling help Membership Inference Attacks?

ResearchDGX agent

arXiv:2503.07482v2 Announce Type: replace-cross Abstract: Membership Inference Attacks (MIAs) aim to estimate whether a specific data point was used in the training of a given model. Existing state-of

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference

Model ReleasesDGX agent

arXiv:2605.25475v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that grows line

Inference-Time Backdoors via Chat Templates: From LLM Supply Chains to Agentic System Compromise

AgentsDGX agent

arXiv:2602.04653v4 Announce Type: replace-cross Abstract: Open-weight language models are increasingly used in production settings, raising new security challenges. One prominent threat is backdoor at

Inference Time Optimization with Confidence Dynamics

Model ReleasesDGX agent

arXiv:2605.25244v1 Announce Type: new Abstract: Inference time optimization techniques, such as repeated sampling, have significantly advanced the reasoning capabilities of Large Language Models (LLMs

LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

Model ReleasesDGX agent

arXiv:2605.25415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adve

LLMTabBench: Evaluating LLMs on Binary Tabular Classification From Zero to Few Shots

Model ReleasesDGX agent

arXiv:2605.24417v1 Announce Type: new Abstract: Supervised classification for tabular data remains a core machine learning task, yet its reliance on large labeled datasets limits applicability in data

MimirRAG: A Multi-Agent RAG Framework for Financial Data Retrieval with Metadata Integration

Model ReleasesDGX agent

arXiv:2605.25030v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems offer a promising approach to reduce hallucinations and improve answer accuracy in large language models (L

Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent

Model ReleasesDGX agent

arXiv:2605.25590v1 Announce Type: cross Abstract: We study nonstationary generalized linear bandits (GLBs), where the expected reward is modeled through a nonlinear link function with an unknown time-

OrpQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization

Model ReleasesDGX agent

arXiv:2605.26092v1 Announce Type: cross Abstract: The deployment of Large Language Models (LLMs) and Vision Transformers (ViTs) on edge devices is significantly constrained by memory limitations and t

Partition of Unity Neural Networks for Interpretable Classification with Explicit Class Regions

Model ReleasesDGX agent

arXiv:2602.00511v2 Announce Type: replace Abstract: Despite their empirical success, neural network classifiers remain difficult to interpret. In softmax-based models, class regions are defined implic

PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation

Model ReleasesDGX agent

arXiv:2605.25572v1 Announce Type: cross Abstract: The growing complexity of quantum programming frameworks has exposed a critical limitation in existing large language model (LLM)-based code assistant

Qwen 3.7 Max is now supported in Hermes Agent

Model ReleasesDGX agent

Nous Research has added support for Qwen 3.7 Max, a large language model, within their Hermes Agent framework. This integration enables users to leverage Qwen 3.7 Max's capabilities when building or d

Reward-free Alignment for Conflicting Objectives

Model ReleasesDGX agent

arXiv:2602.02495v3 Announce Type: replace-cross Abstract: Direct alignment methods are increasingly used to align large language models (LLMs) with human preferences. However, many real-world alignmen

SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation

Model ReleasesDGX agent

arXiv:2605.25984v1 Announce Type: cross Abstract: Ensuring safe and contextually appropriate behaviour in Large Language Models (LLMs) remains a critical challenge for real-world deployment. We presen

SemanticZip: A Pilot Framework for Lossy Text Compression with LLMs as Semantic Decompressors

Model ReleasesDGX agent

arXiv:2605.24541v1 Announce Type: cross Abstract: Text compression for large language model (LLM) systems is usually framed as token deletion, retrieval, summarization, or exact reconstruction. We stu

Stochastic Linear Bandits with Parameter Noise

Model ReleasesDGX agent

arXiv:2601.23164v2 Announce Type: replace Abstract: We study the stochastic linear bandits with parameter noise model, in which the reward of action a is a^op heta where heta is sampled i.i.d. We show

STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media

Model ReleasesDGX agent

arXiv:2605.25162v1 Announce Type: cross Abstract: Large language models for vertical domains are bottlenecked by the scarcity of complex, domain-specific task-oriented dialogues. Existing data acquisi

StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios

Model ReleasesDGX agent

arXiv:2605.25758v1 Announce Type: new Abstract: Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the re

The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench

Model ReleasesDGX agent

arXiv:2605.24782v1 Announce Type: new Abstract: While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than u

TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

Model ReleasesDGX agent

arXiv:2605.25850v1 Announce Type: cross Abstract: This paper investigates large language model (LLM) abstention learning, specifically using ternary reward, which incentivize truthfulness in large lan

TopoAlign: Topology-Aware Visual Representation Alignment

SafetyDGX agent

arXiv:2605.25541v1 Announce Type: cross Abstract: Neural networks encode inputs as high-dimensional vectors, known as representations, that capture how models process data by encoding task-relevant st

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild

ResearchDGX agent

arXiv:2605.24213v1 Announce Type: cross Abstract: Evaluation harnesses are software systems that orchestrate model evaluation by managing model invocation, data loading, metric computation, and result

Truthful Online Preference Aggregation for LLM Fine-Tuning in Mobile Crowdsourcing

Model ReleasesDGX agent

arXiv:2605.24052v1 Announce Type: cross Abstract: To better serve users' demands in mobile applications (e.g., navigation), mobile crowdsourcing platforms can iteratively align large language model (L

Understanding Conversational Patterns in Multi-agent Programming: A Case Study on Fibonacci Game Development

Model ReleasesDGX agent

arXiv:2605.24138v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly applied to software engineering (SE), yet their potential for autonomous, role-oriented collaboration re

25 May 2026

A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek Parliamentary Text

Model ReleasesDGX agent

arXiv:2605.22978v1 Announce Type: new Abstract: Katharevousa Greek remains poorly served by contemporary NLP pipelines despite its importance for legal, administrative, and parliamentary archives. We

Asymmetric Scaling Laws from Sparse Features

ResearchDGX agent

arXiv:2605.23591v1 Announce Type: cross Abstract: We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never o

Benchmarking and Enhancing VLM for Compressed Image Understanding

Model ReleasesDGX agent

arXiv:2512.20901v2 Announce Type: replace Abstract: With the rapid development of Vision-Language Models (VLMs) and the growing demand for their applications, efficient compression of the image inputs

Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2509.26383v5 Announce Type: replace-cross Abstract: Knowledge-graph retrieval-augmented generation (KG-RAG) couples large language models (LLMs) with structured, verifiable knowledge graphs (KGs

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation

Model ReleasesDGX agent

arXiv:2510.08945v3 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) has emerged as a promising paradigm for improving factual accuracy in large language models (LLMs). We introduc

Is TabPFN the Silver Bullet for Insurance Pricing?

ResearchDGX agent

arXiv:2605.22892v1 Announce Type: cross Abstract: Modelling claim frequency and severity for non-life insurance pricing predominantly relies on generalised linear models, with gradient-boosted machine

Memorization Dynamics of Fill-in-the-Middle Pretraining

Model ReleasesDGX agent

arXiv:2605.22981v1 Announce Type: cross Abstract: Fill-in-the-middle (FIM) is a pretraining objective widely used to equip causal language models with infilling ability, yet its effect on verbatim mem

Mitigating Object Hallucinations via Sentence-Level Early Intervention

ResearchDGX agent

arXiv:2507.12455v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have revolutionized cross-modal understanding but continue to struggle with hallucinations - fabricated con

ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU

Model ReleasesDGX agent

arXiv:2605.23057v1 Announce Type: cross Abstract: ModeSwitch-LLM is a lightweight request-boundary controller for improving single-GPU large language model inference efficiency by routing each request

Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection

Model ReleasesDGX agent

arXiv:2605.23036v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based lang

Product >> Brand

IndustryDGX agent

Product >> Brand 2 years ago: 'Elon destroyed the Tesla brand globally! It'll never recover!' Today: #1 selling EV in California: Tesla Model Y #1 selling EV in the U.S.: Tesla Model Y #1 selling EV i

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

Model ReleasesDGX agent

arXiv:2601.12805v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. Howeve

← Previous
1…349350351352353…1044
Next →