AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,841 results
30 Jun 2026

MACROCAST: A Vintage-Consistent Time Series Foundation Model for Real-Time Macroeconomic Forecasting

Model ReleasesDGX agent

arXiv:2606.28670v1 Announce Type: cross Abstract: We introduce MACROCAST, a lightweight Time Series Foundation Model (TSFM) for real-time macroeconomic forecasting. Existing TSFMs suffer from data lea

On Test-Time Scaling for Vision-Language Models

ResearchDGX agent

arXiv:2606.28864v1 Announce Type: new Abstract: Test-time scaling is a paradigm where large models use additional compute at inference to achieve better performance, without changing model weights. Wh

Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

Model ReleasesDGX agent

arXiv:2606.29196v1 Announce Type: cross Abstract: Do language models know when they are being tested? This question matters for AI safety: a model that recognises an evaluation context could alter its

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models

Model ReleasesDGX agent

arXiv:2606.29815v1 Announce Type: new Abstract: Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by e

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

Model ReleasesDGX agent

arXiv:2606.28325v1 Announce Type: cross Abstract: Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a se

29 Jun 2026

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

Model ReleasesDGX agent

arXiv:2602.10179v2 Announce Type: replace-cross Abstract: Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user int

27 Jun 2026

Model page: https://ollama.com/library/ornith

Local AiDGX agent

Ornith is a model available through the Ollama library, a platform for running large language models locally. The specific capabilities and parameters of this model can be found on its dedicated model

26 Jun 2026

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.26566v1 Announce Type: cross Abstract: Adversarial evaluation of AI systems has matured along four largely disconnected tracks: diffusion-based attacks on text and large language models (LL

Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models

Model ReleasesDGX agent

arXiv:2606.14122v2 Announce Type: replace Abstract: Byte-level tokenization enables language models to handle any Unicode input, but models can generate invalid UTF-8 sequences when encountering rare

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

Model ReleasesDGX agent

arXiv:2606.26429v1 Announce Type: cross Abstract: Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.26287v1 Announce Type: new Abstract: With the increase in model parameters and training data, the instruction following and generalization capabilities of Large VisionLanguage Models (LVLMs

Hallucination in World Models is Predictable and Preventable

Model ReleasesDGX agent

arXiv:2606.27326v1 Announce Type: cross Abstract: Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fl

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

Model ReleasesDGX agent

arXiv:2606.27187v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing ha

Learning to Recover Task Experts from a Multi-Task Merged Model

Model ReleasesDGX agent

arXiv:2606.26902v1 Announce Type: new Abstract: Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter

RSPC: A Benchmark for Modeling Stress and Psychiatric Conditions in Digitally Mediated Relationships using Psychiatrist Annotations

Model ReleasesDGX agent

arXiv:2606.27247v1 Announce Type: new Abstract: In NLP, mental health conditions are often modeled as isolated phenomena, without interpersonal context. We use Reddit posts about long-distance relatio

Sampling sea state using a diffusion model

ResearchDGX agent

arXiv:2606.26389v1 Announce Type: cross Abstract: Sea state prediction is essential for operational maritime applications and coupled earth system modeling, yet current spectral wave models remain com

The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report

Model ReleasesDGX agent

arXiv:2606.26529v1 Announce Type: cross Abstract: AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one specified. We

25 Jun 2026

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

Model ReleasesDGX agent

arXiv:2606.25476v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in high-stakes appl

DRM: Diffusion-based Reward Model With Step-wise Guidance

SafetyDGX agent

arXiv:2605.25661v2 Announce Type: replace Abstract: Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward model

Internal Data Repetition Destroys Language Models

Model ReleasesDGX agent

arXiv:2606.24998v1 Announce Type: new Abstract: Language models are running out of high-quality training data, and even aggressively deduplicated corpora retain some amount of repetition. Earlier cont

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

Model ReleasesDGX agent

arXiv:2606.26071v1 Announce Type: new Abstract: A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But beh

Small edits, large models: How Wikipedia advocacy shapes LLM values

Model ReleasesDGX agent

arXiv:2606.24890v1 Announce Type: new Abstract: Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in near

24 Jun 2026

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

Local AiDGX agent

arXiv:2606.24506v1 Announce Type: cross Abstract: Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU memory pro

Experiments with Optimal Model Trees

Model ReleasesDGX agent

arXiv:2503.12902v4 Announce Type: replace Abstract: Model trees provide an appealing way to perform interpretable machine learning for both classification and regression problems. In contrast to ``cla

Grounded Chess Reasoning in Language Models via Master Distillation

Model ReleasesDGX agent

arXiv:2603.20510v2 Announce Type: replace Abstract: Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel. We introd

OpenThoughts-Agent: Data Recipes for Agentic Models

Model ReleasesDGX agent

arXiv:2606.24855v1 Announce Type: new Abstract: Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable ag

Rapid FinFET Modelling Using an Autoencoder

Model ReleasesDGX agent

arXiv:2606.24046v1 Announce Type: cross Abstract: This work presents a machine learning framework that leverages an autoencoder (AE) for the efficient modeling of FinFET. We first calibrated a BSIM-CM

23 Jun 2026

Beyond the Next Step: Variable-Length Latent World Models for Long-Horizon Planning

ResearchDGX agent

arXiv:2606.21775v1 Announce Type: new Abstract: Recently, world models have emerged as a promising paradigm for building intelligent agents by learning predictive models that estimate future environme

Discretizing Reward Models

ResearchDGX agent

arXiv:2606.21795v1 Announce Type: new Abstract: Despite their widespread use, the role of reward models in shaping reinforcement learning is poorly understood. Reward models offer a tempting promise:

Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback

Model ReleasesDGX agent

arXiv:2510.02561v2 Announce Type: replace Abstract: Recent advances in large video-language models (VLMs) rely on extensive fine-tuning techniques that strengthen alignment between textual and visual

Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models

Model ReleasesDGX agent

arXiv:2606.23057v1 Announce Type: cross Abstract: Large language models now mediate how buyers discover products and services, making the competitive structure of AI-generated recommendations a strate

20 Jun 2026

I ran GLM 5.2 with OpenCode harness against Claude Opus this week deployed locally. Bottom line: It is a real frontier coding model and insa…

Model ReleasesDGX agent

I ran GLM 5.2 with OpenCode harness against Claude Opus this week deployed locally. Bottom line: It is a real frontier coding model and insanely good for the price (free). Open source model + open sou

19 Jun 2026

Deepagents code is sick cuz you can just use the best model as it comes out

Model ReleasesDGX agent

Deepagents code is sick cuz you can just use the best model as it comes out it is indeed quite good! don't try it in claude code/codex - those harnesses are overly tuned for their proprietary models d

11 Jun 2026

ICA Lens: Interpreting Language Models Without Training Another Dictionary

Model ReleasesDGX agent

arXiv:2606.11722v1 Announce Type: cross Abstract: Finding interpretable directions in language-model representations is critical for understanding and controlling model behavior. Sparse autoencoders (

10 Jun 2026

BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression

ApplicationsDGX agent

arXiv:2606.10135v1 Announce Type: cross Abstract: Transitioning bidirectional video diffusion models into an autoregressive paradigm improves the interactivity of video world models, but existing caus

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

Model ReleasesDGX agent

arXiv:2606.11187v1 Announce Type: new Abstract: Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow trainin

9 Jun 2026

Calibration of Structured Ignorance Certificates for Diagnosing Unknown Unknowns in Reasoning Models

Model ReleasesDGX agent

arXiv:2606.08571v1 Announce Type: cross Abstract: Large language models frequently fail in a characteristic way: rather than acknowledging ignorance, they produce fluent but incorrect answers to quest

Component Ablation for Efficient Hybrid Language Model Architectures: Performance, Resilience, and Compression Implications

Model ReleasesDGX agent

arXiv:2603.22473v2 Announce Type: replace-cross Abstract: Hybrid language models combine softmax attention with linear-time sequence mechanisms such as state-space or linear-attention layers, but the

DisCo: World Models with Discrete Camera Motion Control

Model ReleasesDGX agent

arXiv:2606.07967v1 Announce Type: new Abstract: Controllable video world models target interactive world exploration, where models must faithfully execute explicit action commands while preserving vis

From inverse problems to neural operators: prediction, mechanism, and generalization of data-driven models

TutorialsDGX agent

arXiv:2606.08956v1 Announce Type: new Abstract: Scientists have historically relied on mathematical models based on differential equations to relate system inputs -- forces, fluxes, or heat sources --

How Small Can You Go? LoRA Fine-Tuning 270M-8B Models for Merchant Information Extraction in Financial Transactions

Model ReleasesDGX agent

arXiv:2606.08051v1 Announce Type: new Abstract: Financial transaction processing requires extracting structured merchant information from noisy, abbreviated bank transaction strings at scale. Our curr

IDDM: Identity-Decoupled Personalized Diffusion Models with a Tunable Privacy-Utility Trade-off

Model ReleasesDGX agent

arXiv:2604.00903v2 Announce Type: replace Abstract: Personalized text-to-image diffusion models (e.g., DreamBooth, LoRA) enable users to synthesize high-fidelity avatars from a few reference photos fo

I'm really not interested in testing a new closed-weight model in the expensive tier. What's the point? I know what will happen: 1. They pre…

SafetyDGX agent

I'm really not interested in testing a new closed-weight model in the expensive tier. What's the point? I know what will happen: 1. They present the model. 2. The model beats the competition on benchm

Lost in the Non-convex Loss Landscape: How to Fine-tune the Large Time Series Model?

Model ReleasesDGX agent

arXiv:2606.08578v1 Announce Type: new Abstract: Recently, large time series models (LTSMs) have gained increasing attention due to their similarities to large language models, including flexible conte

PRISM: PRior-guided Imagination Sampling in world Models

Model ReleasesDGX agent

arXiv:2606.07974v1 Announce Type: cross Abstract: A learned world model provides a powerful physical intuition for evaluating future states. But its effectiveness in continuous control also depends cr

When Do Local Score Models Extrapolate Across Size? A Diagnostic Theory and Benchmark

Model ReleasesDGX agent

arXiv:2606.09705v1 Announce Type: new Abstract: Scientific generative modeling often requires size transfer, where models trained on small systems are evaluated on larger ones. While translation-invar

Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models

Model ReleasesDGX agent

arXiv:2606.07808v1 Announce Type: new Abstract: Reasoning language models deployed in agentic workflows must follow an instruction hierarchy: when instructions from different sources conflict, the mod

8 Jun 2026

Closed-Form Spectral Regularization for Multi-Task Model Merging

Model ReleasesDGX agent

arXiv:2606.07289v1 Announce Type: cross Abstract: Model merging combines several independently fine-tuned experts into a single multi-task model without any training data, reducing the storage, servin

Drifting Models for Surrogate Flow Modeling

ResearchDGX agent

arXiv:2606.07481v1 Announce Type: new Abstract: While Computational Fluid Dynamics (CFD) provides high-fidelity flow fields for optimizing indoor environments, its computational cost limits rapid expl

Position: A Dynamical Systems Perspective is Needed to Advance Time Series Modeling

ResearchDGX agent

arXiv:2602.16864v2 Announce Type: replace-cross Abstract: Time series (TS) modeling has come a long way from early statistical, mainly linear, approaches to the current trend in TS foundation models.

4 Jun 2026

Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases

Model ReleasesDGX agent

arXiv:2606.05112v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed as clinical agents, yet static, single-turn benchmarks cannot capture how a model dynamically del

GENEB: Why Genomic Models Are Hard to Compare

Model ReleasesDGX agent

arXiv:2606.04525v1 Announce Type: new Abstract: Progress in genomic foundation models is difficult to assess due to fragmented benchmarks, incompatible evaluation protocols, and task-specific reportin

Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use

Model ReleasesDGX agent

arXiv:2603.03205v2 Announce Type: replace Abstract: Agentic language models operate in a fundamentally different safety regime than chat models: they must plan, call tools, and execute long-horizon ac

3 Jun 2026

Conditional Latent Diffusion Model with Fourier-based Motion Modelling for Virtual Population Synthesis

ResearchDGX agent

arXiv:2606.03827v1 Announce Type: cross Abstract: In-silico trials of medical devices require the generation of virtual populations of anatomies. In cardiovascular applications, virtual anatomy is typ

Echelon: Auditable Aggregate-Only Language-Model Adaptation Across Privacy Boundaries

Model ReleasesDGX agent

arXiv:2606.02958v1 Announce Type: cross Abstract: Cross-organization language-model adaptation increasingly faces hard governance constraints: in many deployments, device-level model state-parameters,

From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2606.03660v1 Announce Type: new Abstract: Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical

Greener Than Humans? Environmental Attitudes in Large Language Models

Model ReleasesDGX agent

arXiv:2606.02741v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in sustainability-related decision support, reporting, and public communication, yet little systemati

2 Jun 2026

Comprehensive AI governance requires addressing non-model gains

AgentsDGX agent

arXiv:2606.00047v1 Announce Type: cross Abstract: Frontier AI governance often centres on the model-level governance paradigm, which assumes that a model's capability profile is primarily a function o

Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis

SafetyDGX agent

arXiv:2606.00005v1 Announce Type: new Abstract: We present the Consilium Protocol, a Byzantine Fault Tolerance-derived architecture for structured multi-model AI deliberation that treats inter-model d

Geometry-Aware Implicit Memory for Video World Models

Model ReleasesDGX agent

arXiv:2606.02436v1 Announce Type: new Abstract: Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations lea

← Previous
1…2829303132…998
Next →