AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,515 results
7 May 2026

Agent-Based Modeling of Low-Emission Fertilizer Adoption for Dairy Farm Decarbonisation using Empirical Farm Data

SafetyDGX agent

arXiv:2605.03648v1 Announce Type: new Abstract: To understand complex system dynamics in dairy farming, it is essential to use modeling tools that capture farm heterogeneity, social interactions, and

Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models

ApplicationsDGX agent

arXiv:2605.05090v1 Announce Type: new Abstract: We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base mode

Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2512.06393v5 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve high accuracy on many reasoning benchmarks but remain brittle under structural perturbations of rule-base

Counterfactual identifiability beyond global monotonicity: non-monotone triangular structural causal models

Local AiDGX agent

arXiv:2605.04413v1 Announce Type: new Abstract: Structural causal models provide a unified semantics for interventions and counterfactuals, but most identifiability results rely on restrictive assumpt

Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation

ResearchDGX agent

arXiv:2512.14954v2 Announce Type: replace Abstract: Computing next-token likelihood ratios between two language models (LMs) is a standard task in training paradigms such as knowledge distillation. Si

Deep Wave Network for Modeling Multi-Scale Physical Dynamics

HardwareDGX agent

arXiv:2605.04198v1 Announce Type: new Abstract: Performance of deep learning models is strongly governed by architectural capacity, with width and depth as primary controls. However, in physical-scien

I don't really ever trust benchmarks, so ocassionally I'm stress-vibe-testing a bunch of new models on some very complex agent work (hundred…

Model ReleasesDGX agent

I don't really ever trust benchmarks, so ocassionally I'm stress-vibe-testing a bunch of new models on some very complex agent work (hundreds of tools, not your simple coding agent stuff) and to my su

Lookahead Drifting Model

ResearchDGX agent

arXiv:2605.04060v1 Announce Type: cross Abstract: Recently, a new paradigm named drifting model has been proposed for mapping distributions, which achieves the SOTA image generation performance over I

Multi Language Models for On-the-Fly Syntax Highlighting

TutorialsDGX agent

arXiv:2510.04166v2 Announce Type: replace-cross Abstract: Syntax highlighting is a critical feature in modern software development environments, enhancing code readability and developer productivity.

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations…

Model ReleasesDGX agent

New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read

Open Models Make Agentic Batch Processing Economically Viable A lot of world’s work looks like “Do X for EVERY Y” - read every trace - respo…

AgentsDGX agent

Open Models Make Agentic Batch Processing Economically Viable A lot of world’s work looks like “Do X for EVERY Y” - read every trace - respond to every email - deep dive into every document - enrich e

PSK at SemEval-2026 Task 9: Multilingual Polarization Detection Using Ensemble Gemma Models with Synthetic Data Augmentation

Model ReleasesDGX agent

arXiv:2605.05159v1 Announce Type: new Abstract: We present our system for SemEval-2026 Task 9: Multilingual Polarization Detection, a binary classification task spanning 22 languages. Our approach fin

Right Model, Right Time: Real-Time Cascaded-Fidelity MPC for Bipedal Walking

ResearchDGX agent

arXiv:2605.04607v1 Announce Type: new Abstract: This paper presents a multi-phase whole-body model predictive control approach for bipedal walking, combining a detailed whole-body model in the near ho

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

Model ReleasesDGX agent

arXiv:2605.04539v1 Announce Type: new Abstract: Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard preference si

So Mythos was, indeed, not marketing hype. Remember this is a general purpose model that just happens to be good at finding exploits because…

ApplicationsDGX agent

So Mythos was, indeed, not marketing hype. Remember this is a general purpose model that just happens to be good at finding exploits because good models are good at lots of things. Expect similar from

TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference

Model ReleasesDGX agent

arXiv:2603.26498v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) power platforms like ChatGPT, Gemini, and Copilot, enabling richer interactions with text, images, an

The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors

Model ReleasesDGX agent

arXiv:2602.02315v2 Announce Type: replace Abstract: Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these b

Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models

SafetyDGX agent

arXiv:2605.04874v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) b

We recently found some instances of CoT grading during the training of previously deployed models after building a system that scans all Ope…

Model ReleasesDGX agent

We recently found some instances of CoT grading during the training of previously deployed models after building a system that scans all OpenAI RL runs for accidental CoT grading. We did not find clea

6 May 2026

A Few-Step Generative Model on Cumulative Flow Maps

Local AiDGX agent

arXiv:2605.03623v1 Announce Type: new Abstract: We propose a unified, few-step generative modeling framework based on cumulative flow maps for long-range transport in probability space, inspired by fl

A Unified Framework for Tabular Generative Modeling: Loss Functions, Benchmarks, and Improved Multi-objective Bayesian Optimization Approaches

ApplicationsDGX agent

arXiv:2405.16971v2 Announce Type: replace Abstract: Deep learning (DL) models require extensive data to achieve strong performance and generalization. Deep generative models (DGMs) offer a solution by

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

SafetyDGX agent

arXiv:2507.07982v2 Announce Type: replace Abstract: Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw v

Google Chrome silently installs a ~4GB Gemini Nano model on desktop devices; Google says it has been offered since 2024 and users can remove it via settings (Ben Schoon/9to5Google)

Model ReleasesDGX agent

Ben Schoon / 9to5Google: Google Chrome silently installs a ~4GB Gemini Nano model on desktop devices; Google says it has been offered since 2024 and users can remove it via settings — The ongoing marc

GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning

SafetyDGX agent

arXiv:2605.03403v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has recently shown strong performance in post-training large language models and vision-language models. It ra

Hybrid Models for Natural Language Reasoning: The Case of Syllogistic Logic

Model ReleasesDGX agent

arXiv:2510.09472v2 Announce Type: replace Abstract: Despite the remarkable progress in neural models, their ability to generalize, a cornerstone for applications such as logical reasoning, remains a c

Joint Relational Database Generation via Graph-Conditional Diffusion Models

ApplicationsDGX agent

arXiv:2505.16527v2 Announce Type: replace Abstract: Building generative models for relational databases (RDBs) is important for many applications, such as privacy-preserving data release and augmentin

Khala: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation

SafetyDGX agent

arXiv:2605.01790v1 Announce Type: cross Abstract: A common design pattern in high-quality music generation is to handle structure and fidelity in different representation spaces: a generator first mod

Latent State Design for World Models under Sufficiency Constraints

AgentsDGX agent

arXiv:2605.01694v1 Announce Type: new Abstract: A world model matters to an agent only through the state it constructs. That state must preserve some information, discard other information, and suppor

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models

Model ReleasesDGX agent

arXiv:2605.03485v1 Announce Type: new Abstract: Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchma

MRC is already deployed across all of OpenAI’s largest supercomputers that we use to train frontier models, including our site with @Oracle …

Model ReleasesDGX agent

MRC is already deployed across all of OpenAI’s largest supercomputers that we use to train frontier models, including our site with @Oracle Cloud Infrastructure (OCI) in Abilene, Texas, and in @Micros

SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering

TutorialsDGX agent

arXiv:2605.02819v1 Announce Type: new Abstract: Large language models excel at complex reasoning, yet evaluating their intermediate steps remains challenging. Although process reward models provide st

SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification

Local AiDGX agent

arXiv:2605.03301v1 Announce Type: new Abstract: De-identification of clinical text remains essential for secondary use of electronic health records (EHRs), yet public benchmarks such as i2b2 2006/2014

Tencent is about to release an anime video model (AniMatrix).

Local AiDGX agent

Tencent launched Hunyuan Video in December 2024, an open-source AI video generation model with 13 billion parameters that supports text-to-video and image-to-video conversion. The model brings unique

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development

Model ReleasesDGX agent

arXiv:2603.04601v2 Announce Type: replace-cross Abstract: Code generation has emerged as one of AI's highest-impact use cases, yet existing benchmarks measure isolated tasks rather than the complete '

5 May 2026

Break the Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models via Monotonic Entropy Descent with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.02263v1 Announce Type: new Abstract: Recent diffusion large language models (dLLMs) have demonstrated both effectiveness and efficiency in reasoning via a block-based semi-autoregressive ge

Counting as a minimal probe of language model reliability

ResearchDGX agent

arXiv:2605.02028v1 Announce Type: new Abstract: Large language models perform strongly on benchmarks in mathematical reasoning, coding and document analysis, suggesting a broad ability to follow instr

Empowering Heterogeneous Graph Foundation Models via Decoupled Relation Alignment

Model ReleasesDGX agent

arXiv:2605.00731v1 Announce Type: cross Abstract: While Graph Foundation Models (GFMs) have achieved remarkable success in homogeneous graphs, extending them to multi-domain heterogeneous graphs (MDHG

GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models

Model ReleasesDGX agent

arXiv:2605.00371v1 Announce Type: cross Abstract: In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understandin

GenRecEdit: Adapting Model Editing for Generative Recommendation with Cold-Start Items

TutorialsDGX agent

arXiv:2603.14259v2 Announce Type: replace-cross Abstract: Generative recommendation (GR) has shown strong potential for sequential recommendation in an end-to-end generation paradigm. However, existin

Google, Microsoft, and xAI will allow the US government to review their new AI models

Model ReleasesDGX agent

Google DeepMind, Microsoft, and Elon Musk's xAI have agreed to allow the US government to review new AI models before they're released to the public. In an announcement on Tuesday, the Commerce Depart

GPT-5.5 Instant is rolling out over the next two days as the default model to all ChatGPT users, and as ‘gpt-5.5-chat-latest’ in the API. Pe…

Model ReleasesDGX agent

GPT-5.5 Instant is rolling out over the next two days as the default model to all ChatGPT users, and as ‘gpt-5.5-chat-latest’ in the API. Personalization improvements are rolling out to Plus and Pro u

Importance-Guided Basis Selection for Low-Rank Decomposition of Large Language Models

Model ReleasesDGX agent

arXiv:2605.01627v1 Announce Type: new Abstract: Low-rank decomposition is a compelling approach for compressing large language models, but its effectiveness hinges on selecting which singular-vector b

Molecular Representations for Large Language Models

Model ReleasesDGX agent

arXiv:2605.01822v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly being used to support scientific discovery. In chemistry, tasks such as reaction prediction and structure

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models

Model ReleasesDGX agent

arXiv:2605.00877v1 Announce Type: cross Abstract: The vast and underexplored ocean plays a critical role in regulating global climate and supporting marine biodiversity, yet artificial intelligence ha

Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval

Model ReleasesDGX agent

arXiv:2605.02623v1 Announce Type: new Abstract: Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a sing

Shadow-Loom: Causal Reasoning over Graphical World Model of Narratives

Model ReleasesDGX agent

arXiv:2605.02475v1 Announce Type: cross Abstract: Stories hold a reader's attention because they have causes, secrets, and consequences. Shadow-Loom is an experimental open-source framework that turns

Skipping the Zeros in Diffusion Models for Sparse Data Generation

ResearchDGX agent

arXiv:2605.01817v1 Announce Type: new Abstract: Diffusion models (DMs) excel on dense continuous data, but are not designed for sparse continuous data. They do not model exact zeros that represent the

SlimDiffSR: Toward Lightweight and Efficient Remote Sensing Image Super-Resolution via Diffusion Model Distillation

ApplicationsDGX agent

arXiv:2605.02198v1 Announce Type: new Abstract: Diffusion models have recently achieved remarkable performance in image super-resolution (SR), but their high computational cost limits practical deploy

SpecTM: Spectral Targeted Masking for Trustworthy Foundation Models

TutorialsDGX agent

arXiv:2603.22097v2 Announce Type: replace-cross Abstract: Foundation models are now increasingly being developed for Earth observation (EO), yet they often rely on stochastic masking that do not expli

SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?

Model ReleasesDGX agent

arXiv:2605.01911v1 Announce Type: new Abstract: Purpose: Vision-language models (VLMs) have shown promising performance in surgical visual question answering (VQA). However, existing surgical VQA data

TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models

Model ReleasesDGX agent

arXiv:2504.20605v2 Announce Type: replace Abstract: Moral stories are a time-tested vehicle for transmitting values, yet modern NLP lacks a large, structured corpus that couples coherent narratives wi

Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data

ResearchDGX agent

arXiv:2605.00902v1 Announce Type: new Abstract: Foundation models are reshaping computational histopathology, yet their value for whole-slide image retrieval relative to strong patch-based and supervi

we have very efficient models, especially for their capability level happy codexing

Model ReleasesDGX agent

we have very efficient models, especially for their capability level happy codexing yo, i'm actually worried. codex limits are genuinely insane so it's sus af .. i feel this is an intentional move for

we shipped gpt-5.5 instant today to chat; it's rolling out over the next couple days to everyone. for this model, we focused on factuality, …

Model ReleasesDGX agent

we shipped gpt-5.5 instant today to chat; it's rolling out over the next couple days to everyone. for this model, we focused on factuality, crushing hacks, and improving the baseline intelligence. 5.5

4 May 2026

A challenge with AI regulation and vetting is how bad our benchmarks of AI model performance and risks are. There is no benchmark for risks …

Model ReleasesDGX agent

A challenge with AI regulation and vetting is how bad our benchmarks of AI model performance and risks are. There is no benchmark for risks and red-teaming requires experiments from dedicated speciali

Been noticing a lot of 'slow responses' today: models do not inherently more slow, rate limiting more likely.

Local AiDGX agent

A Reddit discussion from the Ollama community addresses reports of slow model responses, clarifying that the models themselves are not inherently slower but that rate limiting is a more likely cause o

Being-H0.7: A Latent World-Action Model from Egocentric Videos

ApplicationsDGX agent

arXiv:2605.00078v1 Announce Type: cross Abstract: Visual-Language-Action models (VLAs) have advanced generalist robot control by mapping multimodal observations and language instructions directly to a

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

ResearchDGX agent

arXiv:2605.00607v1 Announce Type: new Abstract: Probing is widely used to study which features can be decoded from language model representations. However, the common decoding probe approach has two l

Bias in Large Language Models: Origin, Evaluation, and Mitigation

SafetyDGX agent

arXiv:2411.10915v2 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized natural language processing, but their susceptibility to biases poses significant challenges. This

Can Small Language Models Handle Context-Summarized Multi-Turn Customer-Service QA? A Synthetic Data-Driven Comparative Evaluation

SafetyDGX agent

arXiv:2602.00665v3 Announce Type: replace Abstract: Customer-service question answering (QA) systems increasingly rely on conversational language understanding. While Large Language Models (LLMs) achi

← Previous
1…115116117118119…1009
Next →