AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,929 results
13 May 2026

Probabilistic Calibration Is a Trainable Capability in Language Models

Model ReleasesDGX agent

arXiv:2605.11845v1 Announce Type: new Abstract: Language models are increasingly used in settings where outputs must satisfy user-specified randomness constraints, yet their generation probabilities a

StoicLLM: Preference Optimization for Philosophical Alignment in Small Language Models

Model ReleasesDGX agent

arXiv:2605.11483v1 Announce Type: new Abstract: While large language models excel at factual adaptation, their ability to internalize nuanced philosophical frameworks under severe data constraints rem

World Action Models: The Next Frontier in Embodied AI

SafetyDGX agent

arXiv:2605.12090v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
12 May 2026

A Geometric Perspective on Next-Token Prediction in Large Language Models: Three Emerging Phases

Model ReleasesDGX agent

arXiv:2605.09011v1 Announce Type: cross Abstract: We investigate the geometry of predictive information across the layers of large language models (LLMs). We repurpose representation lenses-learned af

CLEF: EEG Foundation Model for Learning Clinical Semantics

Model ReleasesDGX agent

arXiv:2605.10817v1 Announce Type: new Abstract: Clinical EEG interpretation requires reasoning over full EEG sessions and integrating signal patterns with clinical context. Existing EEG foundation mod

Continuity Laws for Sequential Models

SafetyDGX agent

arXiv:2605.08539v1 Announce Type: cross Abstract: Inductive biases influence the behavior and performance of sequential models. In this work, we study an underexplored inductive bias in sequential mod

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification

ResearchDGX agent

arXiv:2605.09269v1 Announce Type: new Abstract: Aligning Multimodal Large Language Models (MLLMs) requires reliable reward models, yet existing single-step evaluators can suffer from lazy judging, exp

Efficient Estimation of Kernel Surrogate Models for Task Attribution

TutorialsDGX agent

arXiv:2602.03783v2 Announce Type: replace-cross Abstract: Modern AI agents such as large language models are trained on diverse tasks -- translation, code generation, mathematical reasoning, and text

Evaluating Pragmatic Reasoning in Large Language Models: Evidence from Scalar Diversity

ResearchDGX agent

arXiv:2605.09042v1 Announce Type: new Abstract: Evaluating pragmatic reasoning in large language models (LLMs) remains challenging because model behavior can vary depending on evaluation methods. Prev

FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models

Model ReleasesDGX agent

arXiv:2605.10141v1 Announce Type: new Abstract: Recent neural theorem provers use reinforcement learning with verifiable rewards (RLVR), where proof assistants provide binary correctness signals. Whil

H-POPE: Hierarchical Polling-based Probing Evaluation of Hallucinations in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2411.04077v2 Announce Type: replace Abstract: By leveraging both texts and images, large vision language models (LVLMs) have shown significant progress in various multi-modal tasks. Nevertheless

LaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual Observations

SafetyDGX agent

arXiv:2605.08279v1 Announce Type: cross Abstract: Learning predictive world models from visual observations is a core problem in embodied AI, with applications to model-based reinforcement learning an

Modeling Atomic Conformational Ensembles of Proteins via Test-Time Supervision of Boltz-2 on Cryo-EM Density Maps

ResearchDGX agent

arXiv:2605.09832v1 Announce Type: new Abstract: Knowledge of a protein's atomic conformational ensemble is critical to determining its function, yet state-of-the-art ensemble prediction models are lim

PhyGround: Benchmarking Physical Reasoning in Generative World Models

Model ReleasesDGX agent

arXiv:2605.10806v1 Announce Type: cross Abstract: Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern re

Position: AI Security Policy Should Target Systems, Not Models

Model ReleasesDGX agent

arXiv:2605.09504v1 Announce Type: cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared memory, paral

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models

ResearchDGX agent

arXiv:2605.09630v1 Announce Type: new Abstract: Tokenizer-free language models eliminate the tokenizer step of the language modeling pipeline by operating directly on bytes; patch-based variants furth

SMIXAE: Towards Unsupervised Manifold Discovery in Language Models

Model ReleasesDGX agent

arXiv:2605.09224v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have been used widely to decompose and interpret neural network activations, especially those of transformer language models.

SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?

Model ReleasesDGX agent

arXiv:2602.03916v3 Announce Type: replace-cross Abstract: Spatial reasoning is a fundamental aspect of human cognition, yet it remains a major challenge for contemporary vision-language models (VLMs).

TabPFN-3 just released: a pre-trained tabular foundation model for up to 1M rows [R][N]

Model ReleasesDGX agent

TabPFN-3 is a pre-trained tabular foundation model that supports datasets up to 1,000,000 rows × 200 features , representing a significant scaling improvement for the TabPFN family. The model delivers

TIDES: Implicit Time-Awareness in Selective State Space Models

Model ReleasesDGX agent

arXiv:2605.09742v1 Announce Type: cross Abstract: Selective state space models (SSMs), such as Mamba, achieve strong per-token expressivity by making the time discretization step Tilde{Delta} a learne

Towards Generative Predictive Display for Vision-Based Teleoperation: A Zero-Shot Benchmark of Off-the-Shelf Video Models

Model ReleasesDGX agent

arXiv:2605.09670v1 Announce Type: cross Abstract: Teleoperation systems are fundamentally limited by communication latency, which degrades situational awareness and control performance. Predictive dis

Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models

Model ReleasesDGX agent

arXiv:2601.22478v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models

Model ReleasesDGX agent

arXiv:2603.18113v2 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly shape content generation, interaction, and decision-making across the Web, aligning them with hum

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.10485v1 Announce Type: new Abstract: Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominan

Weighted Rules under the Stable Model Semantics

ResearchDGX agent

arXiv:2605.09519v1 Announce Type: new Abstract: We introduce the concept of weighted rules under the stable model semantics following the log-linear models of Markov Logic. This provides versatile met

XPERT: Expert Knowledge Transfer for Effective Training of Language Models

ResearchDGX agent

arXiv:2605.08842v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models organize knowledge into explicitly routed expert modules, making expert-level representations traceable and ana

11 May 2026

Continually Evolving Skill Knowledge in Vision Language Action Model

Model ReleasesDGX agent

arXiv:2511.18085v4 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains chal

Coupling Models for One-Step Discrete Generation

ResearchDGX agent

arXiv:2605.07193v1 Announce Type: new Abstract: Generative modeling over discrete structures underpins applications across deep learning, from biological sequence design and code generation to large l

From Pixels to Prompts: Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.07544v1 Announce Type: new Abstract: When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago. Teachi

Generative Modeling with Flux Matching

ResearchDGX agent

arXiv:2605.07319v1 Announce Type: cross Abstract: We introduce Flux Matching, a new paradigm for generative modeling that generalizes existing score-based models to a broader family of vector fields t

Give our early preview of Computer Use (with ANY model) a try today! Built into the latest Hermes Agent and powered by @trycua - opens the d…

AgentsDGX agent

Give our early preview of Computer Use (with ANY model) a try today! Built into the latest Hermes Agent and powered by @trycua - opens the door to any model, not just the frontier models in special mo

Inference of Qualitative Models from Steady-State Data via Weighted MaxSMT

ResearchDGX agent

arXiv:2605.07433v1 Announce Type: cross Abstract: Qualitative models provide crucial instruments for modelling complex biological systems. While advances in automated reasoning and symbolic encodings

Model-Driven Policy Optimization in Differentiable Simulators via Stochastic Exploration

Model ReleasesDGX agent

arXiv:2605.07520v1 Announce Type: new Abstract: Differentiable planning enables gradient-based optimization of decision-making problems by leveraging differentiable models of system dynamics. However,

OmicsLM: A Multimodal Large Language Model for Multi-Sample Omics Reasoning

Model ReleasesDGX agent

arXiv:2605.06728v1 Announce Type: cross Abstract: Interpreting transcriptomic data is one of the most common analytical tasks in modern biology. Yet most current models either consume expression profi

Self-Programmed Execution for Language-Model Agents

SafetyDGX agent

arXiv:2605.06898v1 Announce Type: new Abstract: At the heart of existing language model agents is a fixed orchestrator program responsible for the state transition between consecutive turns. This pape

TAP: Two-Stage Adaptive Personalization of Multi-Task and Multi-Modal Foundation Models in Federated Learning

Local AiDGX agent

arXiv:2509.26524v3 Announce Type: replace-cross Abstract: In federated learning (FL), local personalization of models has received significant attention, yet personalized fine-tuning of foundation mod

Teaching Language Models to Think in Code

Model ReleasesDGX agent

arXiv:2605.07237v1 Announce Type: new Abstract: Tool-integrated reasoning (TIR) has emerged as a dominant paradigm for mathematical problem solving in language models, combining natural language (NL)

Tool Calling is Linearly Readable and Steerable in Language Models

Model ReleasesDGX agent

arXiv:2605.07990v1 Announce Type: cross Abstract: When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed. Probing 12 ins

Uneven Evolution of Cognition Across Generations of Generative AI Models

Model ReleasesDGX agent

arXiv:2605.06815v1 Announce Type: new Abstract: The pursuit of artificial general intelligence necessitates robust methods for evaluating the cognitive capabilities of models beyond narrow task perfor

When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models

Model ReleasesDGX agent

arXiv:2605.07260v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models route each token to a small subset of experts, but whether the routes selected by a trained top-k router are

7 May 2026

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model

ResearchDGX agent

arXiv:2603.11911v3 Announce Type: replace Abstract: We present InSpatio-WorldFM, an open-source real-time frame model for spatial intelligence. Unlike video-based world models that rely on sequential

Not All That Is Fluent Is Factual: Investigating Hallucinations of Large Language Models in Academic Writing

Model ReleasesDGX agent

arXiv:2605.04171v1 Announce Type: new Abstract: Large Language models (LLMs) show extraordinary abilities, but they are still prone to hallucinations, especially when we use them for generating Academ

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism

HardwareDGX agent

arXiv:2605.05049v1 Announce Type: cross Abstract: Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE mo

6 May 2026

Conservative quantum offline model-based optimization

Model ReleasesDGX agent

arXiv:2506.19714v2 Announce Type: replace-cross Abstract: Offline model-based optimization (MBO) refers to the task of optimizing a black-box objective function using only a fixed set of prior input-o

ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity

Model ReleasesDGX agent

arXiv:2605.03667v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable capabilities, but their immense computational demands during training remain a critical bottleneck

Evaluating Reasoning Models for Queries with Presuppositions

ResearchDGX agent

arXiv:2605.03050v1 Announce Type: new Abstract: Millions of users turn to AI models for their information needs. It is conceivable that a large number of user queries contain assumptions that may be f

🧠 Introducing NeuralBench: a unified, open-source framework to benchmark NeuroAI models. v1.0: 36 EEG tasks, 94 datasets, task-specific + f…

Model ReleasesDGX agent

🧠 Introducing NeuralBench: a unified, open-source framework to benchmark NeuroAI models. v1.0: 36 EEG tasks, 94 datasets, task-specific + foundation models. MEG/fMRI ready. MIT-licensed, FAIR's Brain

ISAAC: Auditing Causal Reasoning in Deep Models for Drug-Target Interaction

Model ReleasesDGX agent

arXiv:2605.02962v1 Announce Type: new Abstract: Deep learning models for drug--target interaction (DTI) prediction often achieve strong benchmark performance without necessarily relying on mechanistic

RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models

Model ReleasesDGX agent

arXiv:2605.03821v1 Announce Type: new Abstract: Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly ali

5 May 2026

Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis

Model ReleasesDGX agent

arXiv:2506.08849v4 Announce Type: replace Abstract: Vision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models

Model ReleasesDGX agent

arXiv:2506.09082v5 Announce Type: replace Abstract: The rise of vision foundation models (VFMs) calls for systematic evaluation. A common approach pairs VFMs with large language models (LLMs) as gener

Focus on the Core: Empowering Diffusion Large Language Models by Self-Contrast

Model ReleasesDGX agent

arXiv:2605.01373v1 Announce Type: new Abstract: The iterative denoising paradigm of Diffusion Large Language Models (DLMs) endows them with a distinct advantage in global context modeling. However, cu

Grounding Synthetic Data Generation With Vision and Language Models

Model ReleasesDGX agent

arXiv:2603.09625v2 Announce Type: replace Abstract: Deep learning models benefit from increasing data diversity and volume, motivating synthetic data augmentation to improve existing datasets. However

Learning a Stochastic Differential Equation Model of Tropical Cyclone Intensification from Reanalysis and Observational Data

ResearchDGX agent

arXiv:2601.08116v2 Announce Type: replace Abstract: Tropical cyclones are dangerous natural hazards, but their hazard is challenging to quantify directly from historical datasets due to limited datase

LVLM-Aided Alignment of Task-Specific Vision Models

SafetyDGX agent

arXiv:2512.21985v2 Announce Type: replace Abstract: In high-stakes domains, small task-specific vision models are crucial due to their low computational requirements and the availability of numerous m

MolmoAct2: Action Reasoning Models for Real-world Deployment

Model ReleasesDGX agent

arXiv:2605.02881v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models aim to provide a single generalist controller for robots, but today's systems fall short on the criteria that matter

Open models should compete on cost and specialization, not frontier benchmarks @natolambert puts it well: the right benchmark is savings in …

Model ReleasesDGX agent

Open models should compete on cost and specialization, not frontier benchmarks @natolambert puts it well: the right benchmark is savings in compute and time, especially for repetitive agent tasks deep

OphMAE: Bridging Volumetric and Planar Imaging with a Foundation Model for Adaptive Ophthalmological Diagnosis

Model ReleasesDGX agent

arXiv:2605.02714v1 Announce Type: new Abstract: The advent of foundation models has heralded a new era in medical artificial intelligence (AI), enabling the extraction of generalizable representations

Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models

Model ReleasesDGX agent

arXiv:2511.21086v2 Announce Type: replace Abstract: Large language models must satisfy hard orthographic constraints during controlled text generation, yet systematic cross-family evaluation remains l

Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm

Model ReleasesDGX agent

arXiv:2602.11543v2 Announce Type: replace Abstract: Pretraining large language models (LLMs) typically requires centralized clusters with thousands of high-memory GPUs (e.g., H100/A100). Recent decent

← Previous
1…5354555657…999
Next →