AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

86,965Total entries
1Added by human
86,964Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,458 results
19 May 2026

Gemini 3.5 Flash might be fast enough for gen AI to make sense

Model ReleasesDGX agent

Gemini 3.5 Flash runs 4x faster than other frontier models in output tokens per second while delivering frontier-level intelligence, proving that speed and quality no longer require trade-offs. The mo

General Preference Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.18721v1 Announce Type: cross Abstract: Post-training has split large language model (LLM) alignment into two largely disconnected tracks. Online reinforcement learning (RL) with verifiable

HPC-LLM: Practical Domain Adaptation and Retrieval-Augmented Generation for HPC Support

Model ReleasesDGX agent

arXiv:2605.16347v1 Announce Type: new Abstract: Modern scientific research increasingly depends on High-Performance Computing (HPC) infrastructures, yet many researchers face significant operational b

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

iMiGUE-3K: A Large-Scale Benchmark for Micro-Gesture Analysis with Self-Supervised Learning

Model ReleasesDGX agent

arXiv:2605.17179v1 Announce Type: new Abstract: Emotion understanding is a fundamental challenge in affective computing and artificial intelligence. While existing approaches predominantly focus on fa

Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization

Model ReleasesDGX agent

arXiv:2605.17379v1 Announce Type: cross Abstract: Large language models pretrained on general-domain corpora often exhibit tokenization inefficiencies when applied to specialized domains. Although con

LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations

Model ReleasesDGX agent

arXiv:2605.16538v1 Announce Type: cross Abstract: This paper examines the opportunities, limitations, and practical considerations associated with the use of large language models (LLMs) in qualitativ

MANTA: Multi-turn Assessment for Nonhuman Thinking & Alignment

Model ReleasesDGX agent

arXiv:2605.16301v1 Announce Type: cross Abstract: Single-turn benchmarks such as AnimalHarmBench (AHB) have established important baselines for measuring animal welfare alignment in large language mod

Mixture of Experts for Low-Resource LLMs

Model ReleasesDGX agent

arXiv:2605.17598v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures enable efficient model scaling, yet expert routing behavior across underrepresented languages remains poorly unde

NewsLens: A Multi-Agent Framework for Adversarial News Bias Navigation

Model ReleasesDGX agent

arXiv:2605.17364v1 Announce Type: new Abstract: Media bias detection has predominantly been framed as a classification task: assign a political label to an article or outlet. We argue this framing is

PESD-TSF: A Period-Aware and Explicit Structured Decomposition Framework for Long-Term Time Series Forecasting

Model ReleasesDGX agent

arXiv:2605.16449v1 Announce Type: cross Abstract: Deep forecasting models often suffer from attenuated periodic perception and entangled trend-noise representations as network depth increases. Moreove

PRISMat: Policy-Driven, Permutation-Invariant Autoregressive Material Generation

Model ReleasesDGX agent

arXiv:2605.16612v1 Announce Type: new Abstract: Rapid identification of candidate materials with target properties has become a key task in materials science. Machine learning has emerged as an altern

ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge

Model ReleasesDGX agent

arXiv:2510.18941v2 Announce Type: replace-cross Abstract: Evaluating progress in large language models (LLMs) is often constrained by the challenge of verifying responses, limiting assessments to task

RAP: Runtime Adaptive Pruning for LLM Inference

Model ReleasesDGX agent

arXiv:2505.17138v5 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at language understanding and generation, but their enormous computational and memory requirements hinder d

Scaling Laws for Code: A More Data-Hungry Regime

Model ReleasesDGX agent

arXiv:2510.08702v2 Announce Type: replace Abstract: Code Large Language Models (LLMs) are revolutionizing software engineering. However, scaling laws that guide the efficient training are predominantl

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science

Model ReleasesDGX agent

arXiv:2605.18630v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as scientific AI as- sistants, and a growing body of benchmarks evaluates their capabilities acro

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning

Model ReleasesDGX agent

arXiv:2605.18209v1 Announce Type: cross Abstract: Spatial question answering over egocentric video is a challenging task that requires Vision-Language Models (VLMs) to reason about 3D object positions

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning

SafetyDGX agent

arXiv:2510.16416v4 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown remarkable abilities by integrating large language models with visual inputs. However, they often fai

Tensor Channel Equivariant Graph Neural Networks for Molecular Polarizability Prediction

Model ReleasesDGX agent

arXiv:2605.16891v1 Announce Type: new Abstract: We introduce a tensor-channel equivariant graph neural network for direct prediction of molecular polarizability tensors. Building on the efficient PaiN

Text2CAD-Bench: A Benchmark for LLM-based Text-to-Parametric CAD Generation

Model ReleasesDGX agent

arXiv:2605.18430v1 Announce Type: new Abstract: Text-to-CAD generation aims to create parametric CAD models from natural language, enabling rapid prototyping and intuitive design workflows. However, e

'The Whole Is Greater Than the Sum of Its Parts': A Compatibility-Aware Multi-Teacher CoT Distillation Framework

Model ReleasesDGX agent

arXiv:2601.13992v2 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) reasoning empowers Large Language Models (LLMs) with remarkable capabilities but typically requires prohibitive paramet

TPV: Parameter Perturbations Through the Lens of Test Prediction Variance

Model ReleasesDGX agent

arXiv:2512.11089v4 Announce Type: replace-cross Abstract: We introduce test prediction variance (TPV)--the first-order sensitivity of a trained model's outputs to parameter perturbations--as a unifyin

UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation

Model ReleasesDGX agent

arXiv:2605.17140v1 Announce Type: cross Abstract: Brain tumor diagnosis is largely dependent on Magnetic Resonance Imaging (MRI) evaluation, which requires radiologists to synthesize thousands of imag

Weak-to-Strong Elicitation via Mismatched Wrong Drafts

SafetyDGX agent

arXiv:2605.17314v1 Announce Type: cross Abstract: We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-tuning (e.g.

When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack

ResearchDGX agent

arXiv:2605.17288v1 Announce Type: cross Abstract: Large Language Model (LLM) cascade systems are designed to balance efficiency and performance by processing queries with lightweight models while sele

XCTFormer: Leveraging Cross-Channel and Cross-Time Dependencies for Enhanced Time-Series Analysis

ApplicationsDGX agent

arXiv:2605.18534v1 Announce Type: new Abstract: Multivariate time-series analysis involves extracting informative representations from sequences of multiple interdependent variables, supporting tasks

ZeroSiam: An Efficient Asymmetry for Test-Time Entropy Optimization without Collapse

SafetyDGX agent

arXiv:2509.23183v3 Announce Type: replace Abstract: Test-time entropy minimization helps adapt a model to novel environments and incentivize its reasoning capability, unleashing the model's potential

18 May 2026

A Reproducible and Physically Feasible Dynamic Parameter Identification Framework for a Low-Cost Robot Arm

Model ReleasesDGX agent

arXiv:2605.15949v1 Announce Type: new Abstract: This paper presents a reproducible and physically feasible dynamic parameter identification framework for CRANE-X7, a low-cost robot arm driven by modul

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation

ResearchDGX agent

arXiv:2605.15761v1 Announce Type: new Abstract: Evaluation leaderboards such as LMArena play a central role in benchmarking large language models by aggregating pairwise human preferences into model r

Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design

Model ReleasesDGX agent

arXiv:2605.15871v1 Announce Type: new Abstract: Toward recursive self-improvement, we investigate LLM agents autonomously designing foundation models beyond standard Transformers. We introduce a dual-

Deep Pre-Alignment for VLMs

Model ReleasesDGX agent

arXiv:2605.15300v1 Announce Type: new Abstract: Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis sugge

DimMem: Dimensional Structuring for Efficient Long-Term Agent Memory

Model ReleasesDGX agent

arXiv:2605.15759v1 Announce Type: new Abstract: Large language model (LLM) agents require long-term memory to leverage information from past interactions. However, existing memory systems often face a

Federated Learning of Spiking Neural Networks under Heterogeneous Temporal Resolutions

Model ReleasesDGX agent

arXiv:2605.15355v1 Announce Type: new Abstract: Spiking neural networks (SNNs) are biologically inspired energy-efficient models that use sparse binary spike-based communication between neurons, makin

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization

Model ReleasesDGX agent

arXiv:2605.15980v1 Announce Type: new Abstract: Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computation

FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast

Model ReleasesDGX agent

arXiv:2605.16233v1 Announce Type: new Abstract: Can LLM agents improve decision-making through self-generated memory without gradient updates? We propose FORGE (Failure-Optimized Reflective Graduation

Fortress: A Case Study in Stabilizing Search Recommendations via Temporal Data Augmentation and Feature Pruning

ApplicationsDGX agent

arXiv:2605.15299v1 Announce Type: cross Abstract: In search and recommendation systems, predictive models often suffer from temporal instability when certain input features introduce volatility in out

Grounded Reinforcement Learning for Visual Reasoning

Local AiDGX agent

arXiv:2505.23678v3 Announce Type: replace Abstract: While reinforcement learning (RL) over chains of thought has significantly advanced language models in tasks such as mathematics and coding, visual

How to Choose Your Teacher for Fine Grained Image Recognition

TutorialsDGX agent

arXiv:2605.15689v1 Announce Type: new Abstract: Fine-grained image recognition classifies subcategories such as bird species or car models. While state-of-the-art (SOTA) models are accurate, they are

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling

ResearchDGX agent

arXiv:2605.15393v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed to perform tasks with minimal human oversight, it is crucial that these models operate robustl

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints

Model ReleasesDGX agent

arXiv:2504.11320v3 Announce Type: replace-cross Abstract: Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day. Each request requires toke

PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization

Model ReleasesDGX agent

arXiv:2605.15222v1 Announce Type: cross Abstract: Large language models (LLMs) can often generate functionally correct code, but their ability to produce efficient implementations for performance-crit

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.15394v1 Announce Type: cross Abstract: Joint-embedding predictive architectures (JEPAs) propose that a model should learn more useful abstractions when trained to predict latent representat

SkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces

Model ReleasesDGX agent

arXiv:2605.15215v1 Announce Type: new Abstract: Recently, skills have been widely adopted in large language model (LLM)-based agent systems across various domains. In existing frameworks, skills are t

16 May 2026

I like what Langchain recently released in their deepagents harness, which is an adapter to modify the syntax of primitive file system comma…

Model ReleasesDGX agent

I like what Langchain recently released in their deepagents harness, which is an adapter to modify the syntax of primitive file system commands depending on the model Claude likes “Bash”, Gemini likes

15 May 2026

Cattle Trade: A Multi-Agent Benchmark for LLM Bluffing, Bidding, and Bargaining

Model ReleasesDGX agent

arXiv:2605.14537v1 Announce Type: new Abstract: We introduce extsc{Cattle Trade, a multi-agent benchmark for evaluating large language models (LLMs) as agents in strategic reasoning under imperfect in

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

Model ReleasesDGX agent

arXiv:2605.14133v1 Announce Type: new Abstract: Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend

Communication-Efficient Federated Fine-Tuning

Model ReleasesDGX agent

arXiv:2505.04535v3 Announce Type: replace Abstract: Federated Learning (FL) enables the utilization of vast, previously inaccessible data sources. At the same time, pre-trained Language Models (LMs) h

Descriptor: Distance-Annotated Traffic Perception Question Answering (DTPQA)

Model ReleasesDGX agent

arXiv:2511.13397v2 Announce Type: replace-cross Abstract: The remarkable progress of Vision-Language Models (VLMs) on a variety of tasks has raised interest in their application to automated driving.

Do-Undo Bench: Reversibility for Action Understanding in Image Generation

Model ReleasesDGX agent

arXiv:2512.13609v2 Announce Type: replace Abstract: We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transf

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring

Model ReleasesDGX agent

arXiv:2605.14589v1 Announce Type: new Abstract: Extending the context window of large language models typically requires training on sequences at the target length, incurring quadratic memory and comp

ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

Model ReleasesDGX agent

arXiv:2605.14153v1 Announce Type: cross Abstract: Exploitation is not a binary event. It is a ladder of acquiring progressive capabilities, from executing a single buggy line of code to taking full co

Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution

Model ReleasesDGX agent

arXiv:2605.15138v1 Announce Type: cross Abstract: Standard unlearning evaluations measure behavioral suppression in full precision, immediately after training, despite every deployed language model be

HDRFace: Rethinking Face Restoration with High-Dimensional Representation

Model ReleasesDGX agent

arXiv:2605.14821v1 Announce Type: new Abstract: Face restoration under complex degradations still remains an ill-posed inverse problem due to severe information loss. Although diffusion models benefit

HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning

Model ReleasesDGX agent

arXiv:2605.15024v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) aims to achieve high-level semantic understanding of genuine changes occurring between bi-temporal images

LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling

Model ReleasesDGX agent

arXiv:2605.14186v1 Announce Type: new Abstract: Large language models (LLMs) often expose useful signals of self-monitoring: before solving a problem, they can estimate whether they are likely to succ

NodeSynth: Socially Aligned Synthetic Data for AI Evaluation

Model ReleasesDGX agent

arXiv:2605.14381v1 Announce Type: cross Abstract: Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, thes

Test-Time Learning with an Evolving Library

Model ReleasesDGX agent

arXiv:2605.14477v1 Announce Type: new Abstract: We introduce EvoLib, a test-time learning framework that enables large language models to accumulate, reuse, and evolve knowledge across problem instanc

Text Knows What, Tables Know When: Clinical Timeline Reconstruction via Retrieval-Augmented Multimodal Alignment

Model ReleasesDGX agent

arXiv:2605.15168v1 Announce Type: cross Abstract: Reconstructing precise clinical timelines is essential for modeling patient trajectories and forecasting risk in complex, heterogeneous conditions lik

TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale

Model ReleasesDGX agent

arXiv:2605.15053v1 Announce Type: cross Abstract: Continually pre-training a large language model on heterogeneous text domains, without replay or task labels, has remained an unsolved architectural p

What Makes Words Hard? Sakura at BEA 2026 Shared Task on Vocabulary Difficulty Prediction

ApplicationsDGX agent

arXiv:2605.14257v1 Announce Type: new Abstract: We describe two types of models for vocabulary difficulty prediction: a high-accuracy black-box model, which achieved the top shared task result in the

14 May 2026

Characteristic Root Analysis and Regularization for Linear Time Series Forecasting

TutorialsDGX agent

arXiv:2509.23597v5 Announce Type: replace-cross Abstract: Time series forecasting remains a critical challenge across numerous domains, yet the effectiveness of complex models often varies unpredictab

← Previous
1…313314315316317…1041
Next →