AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,904 results
11 Jun 2026

Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View

ResearchDGX agent

arXiv:2603.05573v2 Announce Type: replace Abstract: Scalable sequence models, such as Transformer variants and structured state-space models, often trade expressivity power for sequence-level parallel

10 Jun 2026

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Model ReleasesDGX agent

arXiv:2606.11063v1 Announce Type: new Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This part

Divide-and-Conquer Modeling for the CTF-4-Science Lorenz Benchmark

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2606.10084v1 Announce Type: cross Abstract: This work presents a divide-and-conquer modeling strategy for the CTF-4-Science Lorenz benchmark, which evaluates chaotic-system prediction across twe

MARCH: Model-Assisted Reinforcement Learning for the Perceptive Control of Humanoids over Sparse Footholds

SafetyDGX agent

arXiv:2606.10288v1 Announce Type: new Abstract: Perceptive bipedal locomotion over sparse terrain remains a difficult challenge: model-based methods are precise but brittle to uncertainty, while model

Parallel Causal Associative Fields: Gated Sparse Memory for Long-Context Language Modeling

Model ReleasesDGX agent

arXiv:2606.10435v1 Announce Type: cross Abstract: Transformers achieve strong language modeling performance by providing direct token-to-token communication paths, but causal self-attention scales qua

When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models

SafetyDGX agent

arXiv:2512.06343v3 Announce Type: replace-cross Abstract: Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling

WorldOlympiad: Can Your World Model Survive a Triathlon?

Model ReleasesDGX agent

arXiv:2606.11129v1 Announce Type: new Abstract: We introduce WorldOlympiad, a benchmark for diagnosing video-based world models across physical faithfulness, geometric consistency, and interaction fid

9 Jun 2026

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

SafetyDGX agent

arXiv:2606.09811v1 Announce Type: cross Abstract: World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical

BREAKING: Anthropic just dropped Claude Fable 5—this is Mythos, made safe for public release. It is the best coding model in the world. We'v…

Model ReleasesDGX agent

BREAKING: Anthropic just dropped Claude Fable 5—this is Mythos, made safe for public release. It is the best coding model in the world. We've been testing it internally @every for the last week or so

CoVEBench: Can Video Editing Models Handle Complex Instructions?

Model ReleasesDGX agent

arXiv:2606.08415v1 Announce Type: cross Abstract: While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly

Domain-Adapted Small Language Models with Hybrid Post-Processing: Achieving Cost-Efficient, Low-Latency Multi-Label Structured Prediction via LoRA Fine-Tuning on Scarce Data

Model ReleasesDGX agent

arXiv:2606.05781v2 Announce Type: replace Abstract: Deploying frontier large language models (LLMs) for domain-specific structured evaluation tasks incurs prohibitive latency, cost, and data-privacy o

Executable World Models for ARC-AGI-3 in the Era of Coding Agents

Model ReleasesDGX agent

arXiv:2605.05138v2 Announce Type: replace Abstract: We evaluate an initial coding-agent system for ARC-AGI-3 in which the agent maintains an executable Python world model, verifies it against previous

In-Context Learning of Temporal Point Processes with Foundation Inference Models

Model ReleasesDGX agent

arXiv:2509.24762v3 Announce Type: replace Abstract: Modeling event sequences of multiple event types with marked temporal point processes (MTPPs) provides a principled way to uncover governing dynamic

Introducing North Mini Code: Cohere’s First Model For Developers

ToolsDGX agent

Cohere Labs introduced North Mini Code, a compact code generation model designed specifically for developers, marking Cohere's entry into specialized coding models. The model is optimized for efficien

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models

SafetyDGX agent

arXiv:2606.07996v1 Announce Type: cross Abstract: Pretraining is fundamental to the development of Large Language Models (LLMs), yet the opacity of pretraining data complicates model analysis and rais

Prescriptive Scaling Reveals the Evolution of Language Model Capabilities

Model ReleasesDGX agent

arXiv:2602.15327v2 Announce Type: replace-cross Abstract: Machine learning model performance improvements tend to arise from competition and application. For deployment, we consider prescriptive scali

Readable Yet Unpredictable: Rotated-Outcome Prediction in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.07641v1 Announce Type: new Abstract: Can vision-language models predict what a 180{eg} rotation would reveal from the original image alone? We study this ability through Rotated-Outcome Pre

Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models

Model ReleasesDGX agent

arXiv:2603.19183v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, little research has mechan

This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great…

Model ReleasesDGX agent

This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *

UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2602.18020v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models leverage pretrained Vision-Language Models (VLMs) as backbones to map images and instructions to actions, demons

Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks

SafetyDGX agent

arXiv:2606.08775v1 Announce Type: cross Abstract: Visual world models have shown great potential in learning complex system dynamics. Recent advancements leverage these models as transition functions

8 Jun 2026

Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

Model ReleasesDGX agent

arXiv:2606.06523v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence. Despite recen

MACD: Model-Aware Contrastive Decoding via Counterfactual Data

Model ReleasesDGX agent

arXiv:2602.01740v3 Announce Type: replace Abstract: Video language models (Video-LLMs) are prone to hallucinations, generating plausible but ungrounded content when visual evidence is weak, ambiguous,

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.06696v1 Announce Type: cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.06853v1 Announce Type: cross Abstract: The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLM

The Geometry of Representational Failures in Vision Language Models

Model ReleasesDGX agent

arXiv:2602.07025v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) exhibit puzzling failures in multi-object visual tasks, such as hallucinating non-existent elements or failing t

TokaMind: A Multi-Modal Transformer Foundation Model for Tokamak Plasma Dynamics

Model ReleasesDGX agent

arXiv:2602.15084v2 Announce Type: replace-cross Abstract: We present TokaMind, to our knowledge the first open-source foundation model for tokamak plasma dynamics, based on a Multi-Modal Transformer (

6 Jun 2026

Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning

Model ReleasesDGX agent

arXiv:2507.12612v3 Announce Type: replace-cross Abstract: Supervised fine-tuning performance for large language models depends strongly on how training budget is distributed across a heterogeneous set

LLMCodec: Adapting Video Codecs for Efficient Weight Compression of Large Language Models

Model ReleasesDGX agent

arXiv:2606.05861v1 Announce Type: cross Abstract: The rapid development of large language models(LLMs) has led to remarkable advances in natural language processing. However, the increasing scale of t

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models

Model ReleasesDGX agent

arXiv:2606.05434v1 Announce Type: cross Abstract: Group Relative Policy Optimisation (GRPO) has emerged as an effective reinforcement-learning algorithm for aligning language models on reasoning tasks

Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation

ResearchDGX agent

arXiv:2606.05510v1 Announce Type: new Abstract: Telehealth systems have become increasingly important for delivering accessible and timely medical information. Existing large language models often str

5 Jun 2026

Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2606.05874v1 Announce Type: new Abstract: Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-n

Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models

Model ReleasesDGX agent

arXiv:2606.05949v1 Announce Type: new Abstract: Scientific illustrations are essential tools for communicating research findings, especially in natural science, where they visualize complex concepts a

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

Model ReleasesDGX agent

arXiv:2606.05177v1 Announce Type: new Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and

Measuring the sensitivity of LLM-based structured extraction to prompt, model, and schema choices in clinical discharge summaries

ResearchDGX agent

arXiv:2606.05970v1 Announce Type: new Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream con

Nemotron 3 Ultra is now available for Pro and Max subscribers on Perplexity and Computer. It's @nvidia's new open model built for long-runni…

Model ReleasesDGX agent

Nvidia's Nemotron 3 Ultra, an open-source model designed for long-context tasks, is now available to Pro and Max subscribers on Perplexity and Computer. This new model represents Nvidia's latest contr

The Granularity Gap: A Multi-Dimensional Longitudinal Audit of Sycophancy in Gemini Models

Model ReleasesDGX agent

arXiv:2606.05183v1 Announce Type: new Abstract: Large language models are increasingly deployed as high-stakes advisors, yet standard alignment benchmarks treat sycophancy as a binary failure mode. We

Using Large Language Models to Support High Volume Application Review for an Undergraduate Research Program

Model ReleasesDGX agent

arXiv:2606.05564v1 Announce Type: new Abstract: Undergraduate research programs such as the Summer Undergraduate Research Fellowship (SURF) at Purdue University receive thousands of applications every

Value-and-Structure Alignment for Routing-Consistent Quantization of Mixture-of-Experts Models

SafetyDGX agent

arXiv:2606.05688v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale foundation models efficiently by activating only a subset of experts for each token, but their large number of exp

What Makes Two Language Models Think Alike?

ResearchDGX agent

arXiv:2406.12620v3 Announce Type: replace Abstract: Do architectural and training differences influence the way models represent and process language? Traditional similarity metrics tell us whether tw

4 Jun 2026

A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and Human Brain Networks

SafetyDGX agent

arXiv:2510.24342v2 Announce Type: replace Abstract: Prior brain-AI alignment studies are typically constrained by specific inputs and tasks, limiting their ability to capture organizational properties

Can Large Language Models Generalize Procedures Across Representations?

Model ReleasesDGX agent

arXiv:2602.03542v2 Announce Type: replace Abstract: Large language models (LLMs) are trained and tested extensively on symbolic representations such as code and graphs, yet real-world user tasks are o

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

Model ReleasesDGX agent

arXiv:2602.12215v2 Announce Type: replace Abstract: Recent robot foundation models largely rely on large-scale behavior cloning, which imitates expert actions but discards transferable dynamics knowle

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training

Model ReleasesDGX agent

arXiv:2506.05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention. Although widely adopted, transfo

Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from k-Parity

Model ReleasesDGX agent

arXiv:2601.22450v2 Announce Type: replace-cross Abstract: Masked Diffusion Language Models have recently emerged as a powerful generative paradigm, yet their generalization properties remain understud

3 Jun 2026

Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement

SafetyDGX agent

arXiv:2412.01282v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) bring powerful understanding and reasoning capabilities to multimodal tasks. Meanwhile, the great need for capab

Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams

Model ReleasesDGX agent

arXiv:2603.19250v2 Announce Type: replace Abstract: Evaluating language models in streaming environments is critical, yet underexplored. Existing benchmarks either focus on single complex events or pr

Coupled Local and Global World Models for Efficient First Order RL

Local AiDGX agent

arXiv:2602.06219v2 Announce Type: replace-cross Abstract: World models offer a promising avenue for more faithfully capturing complex dynamics, including contacts and non-rigidity, as well as complex

.@GoogleDeepMind's Gemma 4 - 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes --model gemma4:1…

Model ReleasesDGX agent

.@GoogleDeepMind's Gemma 4 - 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes --model gemma4:12b-mlx Claude Code: ollama launch claude --model gemma4:12b-

Grounding Functional Similarity by Invariance-Aware Model Stitching

TutorialsDGX agent

arXiv:2505.20142v2 Announce Type: replace Abstract: In deep learning, functional similarity evaluation quantifies the extent to which independently trained models learn similar input--output relations

Investigating Adversarial Robustness of Multi-modal Large Language Models

Model ReleasesDGX agent

arXiv:2606.03713v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder

QUIVER: Quantum-Informed Views for Enhanced Representations in Large ML Models

Model ReleasesDGX agent

arXiv:2606.02785v1 Announce Type: new Abstract: Large machine learning models benefit substantially from multimodal inputs that provide a complementary view of the same example. We introduce QUIVER (Q

Self-Soupervision: Cooking Model Soups without Labels

ResearchDGX agent

arXiv:2602.02890v2 Announce Type: replace Abstract: Model soups are strange and strangely effective combinations of parameters. They take a model (the stock), fine-tune it into multiple models (the in

Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs

Model ReleasesDGX agent

arXiv:2606.03026v1 Announce Type: cross Abstract: Spiking language models expose activation sparsity that dense Transformer runtimes do not directly exploit. This paper studies that property from a sy

2 Jun 2026

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

Model ReleasesDGX agent

arXiv:2606.00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or

Automatic behind the scene routing in user interfaces (instead of model picker) will redistribute value capture and usage towards many more …

IndustryDGX agent

Automatic behind the scene routing in user interfaces (instead of model picker) will redistribute value capture and usage towards many more models than just frontier ones (especially towards open-sour

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models

Model ReleasesDGX agent

arXiv:2606.00658v1 Announce Type: cross Abstract: Large video diffusion models achieve strong visual quality but remain expensive to deploy because each sample requires many denoising steps and a larg

Computation-Aware Kalman Filtering with Model Selection for Neural Dynamics

ResearchDGX agent

arXiv:2606.01468v1 Announce Type: cross Abstract: Due to their explicit priors and ability to model uncertainty, Bayesian methods have played a major role in dynamical latent variable modeling of sing

EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models

ResearchDGX agent

arXiv:2606.00722v1 Announce Type: cross Abstract: Controlling language model outputs is essential for ensuring structural validity, reliability, and downstream usability, and diffusion language models

Evaluating the Reversal Curse in Model Editing

Model ReleasesDGX agent

arXiv:2310.10322v3 Announce Type: replace Abstract: Large language models (LLMs) are prone to hallucinate unintended text due to false or outdated knowledge. Since retraining LLMs is resource intensiv

← Previous
1…4950515253…999
Next →