AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,859 results
Applications

Cross-device Collaborative Test-time Adaptation with Zeroth-order Optimization and Model Merging

DGX agent

arXiv:2607.02988v1 Announce Type: new Abstract: Test-time adaptation (TTA) mitigates domain shifts by using incoming test data to update a model on the fly. The majority of TTA methods require resourc

applicationsarxiv-cs-cv
7 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

EvoXplain: When Machine Learning Models Agree on Predictions but Disagree on Why -- Measuring Mechanistic Multiplicity Across Training Runs

DGX agent

arXiv:2512.22240v5 Announce Type: replace-cross Abstract: Machine learning models are primarily judged by predictive performance, especially in applied genomics, where explanations are read as biologi

researcharxiv-cs-ai
7 Jul 2026
Model Releases

Multiplayer Interactive World Models with Representation Autoencoders

DGX agent

arXiv:2607.05352v1 Announce Type: cross Abstract: We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

No Reliable Evidence of Self-Reported Sentience in Small Large Language Models

DGX agent

arXiv:2601.15334v2 Announce Type: replace-cross Abstract: Whether language models possess sentience has no empirical answer. But whether they believe themselves to be sentient can, in principle, be te

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Spectral Rewiring for Exploration, Purification, and Model Merging

DGX agent

arXiv:2607.03065v1 Announce Type: cross Abstract: Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-re

model-releasesarxiv-cs-ai
7 Jul 2026
Local Ai

Trust-Region Noise Search for Black-Box Alignment of Diffusion and Flow Models

DGX agent

arXiv:2603.14504v2 Announce Type: replace-cross Abstract: Optimizing the noise samples of diffusion and flow models is an increasingly popular approach to align these models to target rewards at infer

local-aiarxiv-cs-ai
7 Jul 2026
Model Releases

Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns

DGX agent

arXiv:2607.00048v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in exam- and certification-style question answering tasks, where their ability to retrieve, interpr

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

Steal the Patch Size: Adversarially Manipulate Vision-Language Models

DGX agent

arXiv:2607.00174v1 Announce Type: new Abstract: We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including

model-releasesarxiv-cs-cv
2 Jul 2026
Research

Utilizing Earth Foundation Models to Enhance the Simulation Performance of Hydrological Models with AlphaEarth Embeddings

DGX agent

arXiv:2601.01558v2 Announce Type: replace-cross Abstract: Predicting river flow in places without streamflow records is challenging because basins respond differently to climate, terrain, vegetation,

researcharxiv-cs-ai
2 Jul 2026
Model Releases

Run NVIDIA Nemotron and OpenAI GPT OSS models on Amazon Bedrock in AWS GovCloud (US)

DGX agent

We're excited to introduce US-based frontier open-weight models in AWS GovCloud (US). With this release, Amazon Bedrock now supports OpenAI’s open-weight GPT OSS models (120B and 20B) and NVIDIA Nemot

model-releasesaws-ml-blog
1 Jul 2026
Model Releases

China’s Meituan open-sources massive LongCat-2.0 AI model, saying it was trained on domestic chips

DGX agent

Beijing, China-based Meituan Inc. today debuted its next-generation LongCat-2.0 open-source large language model, stating that the company trained the 1.6-trillion-parameter model on domestic Chinese

model-releasessiliconangle
30 Jun 2026
Safety

How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position

DGX agent

arXiv:2606.15032v2 Announce Type: replace Abstract: World models have become a central abstraction in modern AI. The term now refers to several different objects: action-conditioned environment models

safetyarxiv-cs-lg
30 Jun 2026
Model Releases

MACROCAST: A Vintage-Consistent Time Series Foundation Model for Real-Time Macroeconomic Forecasting

DGX agent

arXiv:2606.28670v1 Announce Type: cross Abstract: We introduce MACROCAST, a lightweight Time Series Foundation Model (TSFM) for real-time macroeconomic forecasting. Existing TSFMs suffer from data lea

model-releasesarxiv-cs-ai
30 Jun 2026
Research

On Test-Time Scaling for Vision-Language Models

DGX agent

arXiv:2606.28864v1 Announce Type: new Abstract: Test-time scaling is a paradigm where large models use additional compute at inference to achieve better performance, without changing model weights. Wh

researcharxiv-cs-cv
30 Jun 2026
Model Releases

Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

DGX agent

arXiv:2606.29196v1 Announce Type: cross Abstract: Do language models know when they are being tested? This question matters for AI safety: a model that recognises an evaluation context could alter its

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models

DGX agent

arXiv:2606.29815v1 Announce Type: new Abstract: Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by e

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

DGX agent

arXiv:2606.28325v1 Announce Type: cross Abstract: Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a se

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

DGX agent

arXiv:2602.10179v2 Announce Type: replace-cross Abstract: Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user int

model-releasesarxiv-cs-ai
29 Jun 2026
Local Ai

Model page: https://ollama.com/library/ornith

DGX agent

Ornith is a model available through the Ollama library, a platform for running large language models locally. The specific capabilities and parameters of this model can be found on its dedicated model

local-aiollama--x
27 Jun 2026
Model Releases

Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models

DGX agent

arXiv:2606.26566v1 Announce Type: cross Abstract: Adversarial evaluation of AI systems has matured along four largely disconnected tracks: diffusion-based attacks on text and large language models (LL

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models

DGX agent

arXiv:2606.14122v2 Announce Type: replace Abstract: Byte-level tokenization enables language models to handle any Unicode input, but models can generate invalid UTF-8 sequences when encountering rare

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

DGX agent

arXiv:2606.26429v1 Announce Type: cross Abstract: Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models

DGX agent

arXiv:2606.26287v1 Announce Type: new Abstract: With the increase in model parameters and training data, the instruction following and generalization capabilities of Large VisionLanguage Models (LVLMs

model-releasesarxiv-cs-cv
26 Jun 2026
Model Releases

Hallucination in World Models is Predictable and Preventable

DGX agent

arXiv:2606.27326v1 Announce Type: cross Abstract: Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fl

model-releasesarxiv-cs-cv
26 Jun 2026
Model Releases

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

DGX agent

arXiv:2606.27187v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing ha

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

Learning to Recover Task Experts from a Multi-Task Merged Model

DGX agent

arXiv:2606.26902v1 Announce Type: new Abstract: Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

RSPC: A Benchmark for Modeling Stress and Psychiatric Conditions in Digitally Mediated Relationships using Psychiatrist Annotations

DGX agent

arXiv:2606.27247v1 Announce Type: new Abstract: In NLP, mental health conditions are often modeled as isolated phenomena, without interpersonal context. We use Reddit posts about long-distance relatio

model-releasesarxiv-cs-lg
26 Jun 2026
Research

Sampling sea state using a diffusion model

DGX agent

arXiv:2606.26389v1 Announce Type: cross Abstract: Sea state prediction is essential for operational maritime applications and coupled earth system modeling, yet current spectral wave models remain com

researcharxiv-cs-ai
26 Jun 2026
Model Releases

The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report

DGX agent

arXiv:2606.26529v1 Announce Type: cross Abstract: AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one specified. We

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

DGX agent

arXiv:2606.25476v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in high-stakes appl

model-releasesarxiv-cs-cl
25 Jun 2026
Safety

DRM: Diffusion-based Reward Model With Step-wise Guidance

DGX agent

arXiv:2605.25661v2 Announce Type: replace Abstract: Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward model

safetyarxiv-cs-cv
25 Jun 2026
Model Releases

Internal Data Repetition Destroys Language Models

DGX agent

arXiv:2606.24998v1 Announce Type: new Abstract: Language models are running out of high-quality training data, and even aggressively deduplicated corpora retain some amount of repetition. Earlier cont

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

DGX agent

arXiv:2606.26071v1 Announce Type: new Abstract: A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But beh

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Small edits, large models: How Wikipedia advocacy shapes LLM values

DGX agent

arXiv:2606.24890v1 Announce Type: new Abstract: Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in near

model-releasesarxiv-cs-cl
25 Jun 2026
Local Ai

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

DGX agent

arXiv:2606.24506v1 Announce Type: cross Abstract: Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU memory pro

local-aiarxiv-cs-ai
24 Jun 2026
Model Releases

Experiments with Optimal Model Trees

DGX agent

arXiv:2503.12902v4 Announce Type: replace Abstract: Model trees provide an appealing way to perform interpretable machine learning for both classification and regression problems. In contrast to ``cla

model-releasesarxiv-cs-lg
24 Jun 2026
Model Releases

Grounded Chess Reasoning in Language Models via Master Distillation

DGX agent

arXiv:2603.20510v2 Announce Type: replace Abstract: Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel. We introd

model-releasesarxiv-cs-ai
24 Jun 2026
Model Releases

OpenThoughts-Agent: Data Recipes for Agentic Models

DGX agent

arXiv:2606.24855v1 Announce Type: new Abstract: Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable ag

model-releasesarxiv-cs-ai
24 Jun 2026
Model Releases

Rapid FinFET Modelling Using an Autoencoder

DGX agent

arXiv:2606.24046v1 Announce Type: cross Abstract: This work presents a machine learning framework that leverages an autoencoder (AE) for the efficient modeling of FinFET. We first calibrated a BSIM-CM

model-releasesarxiv-cs-ai
24 Jun 2026
Research

Beyond the Next Step: Variable-Length Latent World Models for Long-Horizon Planning

DGX agent

arXiv:2606.21775v1 Announce Type: new Abstract: Recently, world models have emerged as a promising paradigm for building intelligent agents by learning predictive models that estimate future environme

researcharxiv-cs-lg
23 Jun 2026
Research

Discretizing Reward Models

DGX agent

arXiv:2606.21795v1 Announce Type: new Abstract: Despite their widespread use, the role of reward models in shaping reinforcement learning is poorly understood. Reward models offer a tempting promise:

researcharxiv-cs-lg
23 Jun 2026
Model Releases

Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback

DGX agent

arXiv:2510.02561v2 Announce Type: replace Abstract: Recent advances in large video-language models (VLMs) rely on extensive fine-tuning techniques that strengthen alignment between textual and visual

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models

DGX agent

arXiv:2606.23057v1 Announce Type: cross Abstract: Large language models now mediate how buyers discover products and services, making the competitive structure of AI-generated recommendations a strate

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

I ran GLM 5.2 with OpenCode harness against Claude Opus this week deployed locally. Bottom line: It is a real frontier coding model and insa…

DGX agent

I ran GLM 5.2 with OpenCode harness against Claude Opus this week deployed locally. Bottom line: It is a real frontier coding model and insanely good for the price (free). Open source model + open sou

model-releasesclem-delangue--x
20 Jun 2026
Model Releases

Deepagents code is sick cuz you can just use the best model as it comes out

DGX agent

Deepagents code is sick cuz you can just use the best model as it comes out it is indeed quite good! don't try it in claude code/codex - those harnesses are overly tuned for their proprietary models d

model-releasesharrison-chase--x
19 Jun 2026
Model Releases

ICA Lens: Interpreting Language Models Without Training Another Dictionary

DGX agent

arXiv:2606.11722v1 Announce Type: cross Abstract: Finding interpretable directions in language-model representations is critical for understanding and controlling model behavior. Sparse autoencoders (

model-releasesarxiv-cs-ai
11 Jun 2026
Applications

BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression

DGX agent

arXiv:2606.10135v1 Announce Type: cross Abstract: Transitioning bidirectional video diffusion models into an autoregressive paradigm improves the interactivity of video world models, but existing caus

applicationsarxiv-cs-ai
10 Jun 2026
Model Releases

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

DGX agent

arXiv:2606.11187v1 Announce Type: new Abstract: Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow trainin

model-releasesarxiv-cs-cv
10 Jun 2026
← Previous
1…3536373839…1248
Next →