AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

86,457Total entries
1Added by human
86,456Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,039 results
4 Jun 2026

OckBench: Measuring the Efficiency of LLM Reasoning

Model ReleasesDGX agent

arXiv:2511.05722v3 Announce Type: replace-cross Abstract: Large language models (LLMs) such as GPT-5 and Gemini 3 have pushed the frontier of automated reasoning and code generation. Yet current bench

Prediction Under Imperfect Compression: A Theory of Approximate MDL

ResearchDGX agent

arXiv:2606.04834v1 Announce Type: new Abstract: Minimum Description Length (MDL) formalizes the principle of Occam's razor by optimizing the total description length: L(model)+L(data | model). For seq

RIDE: An Open Dataset and Benchmark for Train Delay Prediction

Model ReleasesDGX agent

arXiv:2606.05070v1 Announce Type: new Abstract: Train delay prediction is an important problem for both passengers and railway operators, yet progress in the field remains difficult to assess due to t

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

Model ReleasesDGX agent

arXiv:2603.10044v2 Announce Type: replace-cross Abstract: A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never

Stepwise Reasoning Enhancement for LLMs via External Subgraph Generation

Model ReleasesDGX agent

arXiv:2606.04454v1 Announce Type: new Abstract: Large language models have shown strong performance in natural language generation and downstream reasoning tasks, but they still struggle with logical

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

Model ReleasesDGX agent

arXiv:2606.05165v1 Announce Type: cross Abstract: Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data. The gold standard for TDA relies on causal interventio

SymTRELLIS: Symmetry-Enforced Voxel Latents for 3D Generation

Model ReleasesDGX agent

arXiv:2606.04108v1 Announce Type: cross Abstract: Single-view 3D generative models have achieved impressive visual quality, yet they are not designed to satisfy structural or functional requirements,

3 Jun 2026

A Quantitative Approximation Framework for Flow Distillation in Diffusion Models

ResearchDGX agent

arXiv:2606.03820v1 Announce Type: cross Abstract: We develop a quantitative approximation framework for diffusion distillation, viewing few-step sampling as error propagation under compositions of lea

Capability Advertisement as a Market for Lemons: A Trust Layer for Heterogeneous Agent Networks

AgentsDGX agent

arXiv:2606.03034v1 Announce Type: cross Abstract: Large language model (LLM) agents have begun to delegate work to one another. Protocols such as the Model Context Protocol (MCP) and the Agent2Agent p

EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction

Model ReleasesDGX agent

arXiv:2606.02971v1 Announce Type: new Abstract: Extracting reporting obligations from EU legislation is critical for assessing and reducing regulatory reporting burden. However, distinguishing reporti

Evaluating LLMs' Effectiveness on Real-World Consumer Device Repair Questions

Model ReleasesDGX agent

arXiv:2606.03331v1 Announce Type: cross Abstract: Consumer device repair is an important but underexplored testbed for large language models (LLMs). Repair tasks require reasoning over incomplete prob

Fast Unlearning at Scale via Margin Self-Correction

ResearchDGX agent

arXiv:2606.02920v1 Announce Type: new Abstract: Language-model unlearning updates a trained model to behave as if it had not seen selected training examples, while preserving utility and avoiding cost

Gemma 4 12B is here! Dense, mid-sized Gemma that fits right on your laptop - released by @google under Apache 2.0 Available now in LM Studio…

Model ReleasesDGX agent

Gemma 4 12B is a newly released dense language model from Google available under the Apache 2.0 open license, designed for mid-sized computing resources that can run locally on personal computers. The

Inference Cost Attacks for Retrieval-Augmented Large Language Models

SafetyDGX agent

arXiv:2606.02643v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG)-enhanced LLM systems, while powerful, introduce substantial inference costs due to the inclusion of an extra mult

Introducing new capabilities to GPT-Rosalind

Model ReleasesDGX agent

GPT-Rosalind is an OpenAI model with newly introduced capabilities designed to enhance its performance in specific tasks or domains. The update likely expands the model's functionality in areas such a

Language Bias under Conflicting Information in Multilingual LLMs

Model ReleasesDGX agent

arXiv:2604.07123v2 Announce Type: replace Abstract: Large Language Models (LLMs) have been shown to contain biases in the process of integrating conflicting information when answering questions. Here

LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

Model ReleasesDGX agent

arXiv:2606.03303v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong informal mathematical reasoning but struggle to generate mechanically verifiable proofs in formal languages

PSViT: A Methodology for Structurally Pruning Spiking Vision Transformers

ResearchDGX agent

arXiv:2606.03257v1 Announce Type: cross Abstract: Spiking Vision Transformer (SViT) models are promising low-power ViT models for solving vision-based tasks with state-of-the-art performance. However,

Quadratic integrate-and-fire neurons exhibit less fragmented loss landscapes and outperform leaky integrate-and-fire neurons in spike-based gradient descent

Model ReleasesDGX agent

arXiv:2606.03935v1 Announce Type: cross Abstract: The ability to train spiking neural networks is essential for modeling biological neural networks as well as for neuromorphic computing. However, for

Relational Linearity is a Predictor of Hallucinations

Model ReleasesDGX agent

arXiv:2601.11429v2 Announce Type: replace-cross Abstract: Hallucination is a central failure mode of language models (LMs). We focus on hallucinations in response to questions like: 'Which instrument

ReLoRA: Knowledge-Reusing Adaptation for Fast Rollout of Evolving LLM Services

ResearchDGX agent

arXiv:2606.02606v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as continuously evolving services, where frequent base-model updates may invalidate previously

State-Coupled Volatility in Latent Dynamical Systems: Recovery Under Partial Observation

Model ReleasesDGX agent

arXiv:2606.02664v1 Announce Type: cross Abstract: Latent state-space models are widely used to study partially observed dynamical systems, yet most formulations assume that process variability is inde

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks

Model ReleasesDGX agent

arXiv:2606.03606v1 Announce Type: cross Abstract: Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate

v0.30.4: llama-server: fix gemma4 patch wiring (#16477)

Model ReleasesDGX agent

Ollama v0.30.4 is a patch release addressing a bug in the llama-server component related to incorrect parameter wiring in the Gemma 4 model implementation. This fix ensures Gemma 4 models operate corr

VistaHop: Benchmarking Multi-hop Visual Reasoning for Visual DeepSearch

Model ReleasesDGX agent

arXiv:2606.03273v1 Announce Type: cross Abstract: Visual DeepSearch requires multimodal large reasoning model (MLRM) agents to answer complex visual queries by repeatedly inspecting image regions, gro

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?

Model ReleasesDGX agent

arXiv:2606.03837v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it a

Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

Model ReleasesDGX agent

arXiv:2602.08873v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now used for academic expert recommendation. Existing audits typically evaluate such recommendations in isola

2 Jun 2026

A Biconvex Formulation for Stable Transport of Mixture Models with a Unique Solution

ApplicationsDGX agent

arXiv:2606.02515v1 Announce Type: new Abstract: Optimal transport (OT) provides a principled framework for mapping between probability distributions. Despite extensive progress, applying OT to large-s

A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization

ResearchDGX agent

arXiv:2606.00230v1 Announce Type: new Abstract: Grokking, the phenomenon in which neural networks generalize long after fitting their training data, has been studied in supervised settings on many epo

AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

Model ReleasesDGX agent

arXiv:2603.14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reason

[AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark

Model ReleasesDGX agent

NVIDIA announced three new offerings: Cosmos 3, an advanced video generation model; Nemotron 3 Ultra, an upgraded language model; and RTX Spark, likely a tool or framework for developers. These releas

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate

Model ReleasesDGX agent

arXiv:2606.00257v1 Announce Type: cross Abstract: Token-level credit assignment for language-model reinforcement learning is usually formulated as if the policy were fully trainable, while practical L

Before the Model Learns the Bug:Fuzzing RLVR Verifiers

TutorialsDGX agent

arXiv:2606.01066v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) replaces human preference labels with executable reward functions such as math answer checkers, JS

Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing

Model ReleasesDGX agent

arXiv:2606.01079v1 Announce Type: new Abstract: Image compositing aims to seamlessly insert a foreground object into a background image, and recent advances in diffusion models have significantly enha

Collaborative and Efficient Fine-tuning: Leveraging Task Similarity

Model ReleasesDGX agent

arXiv:2602.07218v2 Announce Type: replace-cross Abstract: Adaptability has been regarded as a central feature in the foundation models, enabling them to effectively acclimate to unseen downstream task

CoMIC: Collaborative Memory and Insights Circulation for Long-Horizon LLM Agents in Cloud-Edge Systems

Model ReleasesDGX agent

arXiv:2606.00756v1 Announce Type: new Abstract: Deploying lightweight Large Language Model (LLM) agents on edge servers can reduce latency and move agentic services closer to users, but resource-const

CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2606.01879v1 Announce Type: new Abstract: Existing research largely reduces cultural intelligence in LLMs to a knowledge-level problem, overlooking whether models can effectively utilize their a

Easy, robust approximate message passing for planted spike models

ResearchDGX agent

arXiv:2606.00500v1 Announce Type: cross Abstract: We present a simple and efficient algorithm for robust approximate message passing (AMP) in the spiked matrix setting. In particular, let arepsilon be

FACT: A Simple and Efficient Framework for Active Finetuning

Model ReleasesDGX agent

arXiv:2606.02079v1 Announce Type: new Abstract: The main goal of active finetuning is to improve a pretrained model's performance on a specific task or domain by finetuning it with carefully selected

Feature to Dynamics: Feature-space to Autoregression strategy for Zero-shot Time Series Forecasting

Model ReleasesDGX agent

arXiv:2606.01289v1 Announce Type: new Abstract: Zero-shot time series forecasting aims to predict future values for previously unseen series, requiring models to generalize temporal dynamics beyond th

From Scaling to Structured Expressivity: Rethinking Transformers for CTR Prediction

Model ReleasesDGX agent

arXiv:2511.12081v2 Announce Type: replace-cross Abstract: Despite massive investments in scale, deep models for click-through rate (CTR) prediction often exhibit rapidly diminishing returns -- a stark

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling

ResearchDGX agent

arXiv:2606.00110v1 Announce Type: new Abstract: Achieving robust generalization from limited data is a central challenge in embodied intelligence. Prevailing methods fail by regressing absolute coordi

Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation

SafetyDGX agent

arXiv:2510.18439v3 Announce Type: replace Abstract: Hallucination, where models generate fluent text unsupported by visual evidence, remains a major flaw in vision-language models and is particularly

HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

Model ReleasesDGX agent

arXiv:2606.01132v1 Announce Type: new Abstract: Understanding chart and table images is essential for applying vision-language models (VLMs) to real-world document understanding. While English benchma

Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs

ResearchDGX agent

arXiv:2606.00642v1 Announce Type: new Abstract: Reasoning traces have become a valuable form of learning signals for improving and transferring the capabilities of large language models. In particular

How to Correctly Report LLM-as-a-Judge Evaluations

Model ReleasesDGX agent

arXiv:2511.21140v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensiti

Improving IoT Intrusion Detection Through SMOTE-Based Oversampling and Extended Multi-Model Evaluation on Side-Channel Power Data

ResearchDGX agent

arXiv:2606.00161v1 Announce Type: cross Abstract: The detection of intrusions in IoT-based networks poses challenges that cannot be overcome using traditional machine learning methods. Perhaps the big

Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

SafetyDGX agent

arXiv:2606.00334v1 Announce Type: cross Abstract: Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models

Logit Distillation on Manifolds: Mapping by Learning

ResearchDGX agent

arXiv:2606.00771v1 Announce Type: cross Abstract: A simple way to improve the performance of almost any machine learning model is not to train a single but several models with diverse algorithms which

MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation

Model ReleasesDGX agent

arXiv:2606.02470v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has emerged as a transformative standard for connecting large language models (LLMs) with external data sources and too

PaintBench: Deterministic Evaluation of Precise Visual Editing

Model ReleasesDGX agent

arXiv:2606.00188v1 Announce Type: cross Abstract: While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle. To p

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs

Local AiDGX agent

arXiv:2606.00104v1 Announce Type: cross Abstract: Foundation models are increasingly used to drive autonomous systems, yet existing approaches either keep the model in a tight control loop, raising la

ProductWebGen: Benchmarking Multimodal Product Webpage Generation

Model ReleasesDGX agent

arXiv:2606.01022v1 Announce Type: cross Abstract: Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value f

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression

Model ReleasesDGX agent

arXiv:2606.00494v1 Announce Type: new Abstract: Post-Training Quantization (PTQ) and Low-Rank Adaptation (LoRA) constitute the standard pipeline for efficient Large Language Model (LLM) deployment. Ho

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

Model ReleasesDGX agent

arXiv:2601.04946v3 Announce Type: replace-cross Abstract: Automatic metrics are widely used to evaluate text-to-image models, often replacing human judgment in benchmarking, model selection, and large

Riemannian Optimization for Hadamard Products of Low-Rank Matrices

Model ReleasesDGX agent

arXiv:2606.01216v1 Announce Type: new Abstract: The elementwise Hadamard product of two low-rank matrices provides a parameter-efficient model for data with multiplicative structure, but its modeling

Structure Enables Effective Self-Localization of Errors in LLMs

Local AiDGX agent

arXiv:2602.02416v2 Announce Type: replace Abstract: Self-correction in language models remains elusive. In this work, we explore whether language models can explicitly localize errors in incorrect rea

Subliminal Learning Is Steering Vector Distillation

ResearchDGX agent

arXiv:2606.00995v1 Announce Type: new Abstract: Subliminal learning refers to a student language model acquiring a teacher's traits (e.g. a system-prompted preference for owls) when fine-tuned on the

The Right Inference Strategy Is All You Need: Nearly Training-Free Domain-Wise Inference for EgoCross Challenge

ResearchDGX agent

arXiv:2606.00829v1 Announce Type: new Abstract: EgoCross evaluates multimodal large language models on egocentric video question answering under substantial domain shift, where test videos come from s

Toward accurate RUL and SoH estimation using reinforced graph-based physics-informed neural networks enhanced with dynamic weights

Model ReleasesDGX agent

arXiv:2507.09766v2 Announce Type: replace-cross Abstract: Accurate estimation of Remaining Useful Life (RUL) and State of Health (SoH) is essential for reliable Prognostics and Health Management (PHM)

← Previous
1…279280281282283…1034
Next →