AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,860 results
29 Jul 2026

TaylorPODA: A Taylor Expansion-Based Method to Improve Post-Hoc Attributions for Opaque Models

Local AiDGX agent

arXiv:2507.10643v4 Announce Type: replace-cross Abstract: Post-hoc model-agnostic local attribution (LA) methods have been widely adopted to explain opaque AI models by quantifying feature-wise contri

28 Jul 2026

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

Model ReleasesDGX agent

arXiv:2607.23124v1 Announce Type: new Abstract: Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settin

AIFL: A Global Daily Streamflow Forecasting Model Using a Deterministic LSTM Pre-trained on ERA5-Land and Fine-tuned on IFS

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
ResearchDGX agent

arXiv:2602.16579v2 Announce Type: replace-cross Abstract: Reliable global streamflow forecasting is essential for flood preparedness and water resource management, yet data-driven models often suffer

Algorithmic Blindness in Large Language Models: A Calibration Study of Performance Prediction

Model ReleasesDGX agent

arXiv:2602.21947v5 Announce Type: replace Abstract: Large language models (LLMs) demonstrate remarkable breadth of knowledge, yet their ability to reason about computational processes remains poorly u

Alibaba has released Qwen Audio 3.0 Realtime, with the Plus variant debuting as the new #1 model on the Artificial Analysis Speech to Speech…

Model ReleasesDGX agent

Alibaba has released Qwen Audio 3.0 Realtime, with the Plus variant debuting as the new #1 model on the Artificial Analysis Speech to Speech Index at 84.1%, ahead of GPT-Realtime-2.1 High at 79.1% Rel

AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models

AgentsDGX agent

arXiv:2603.28963v2 Announce Type: replace-cross Abstract: Simulation with realistic traffic agents is essential for validating autonomous driving systems. Existing data-driven simulators learn agent b

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning

ResearchDGX agent

arXiv:2607.22769v1 Announce Type: cross Abstract: The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data. Existing dynamic d

Embodied GPT-5.1: Evidence of a World Model?

Model ReleasesDGX agent

arXiv:2607.23899v1 Announce Type: cross Abstract: This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mobile robot

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models

ResearchDGX agent

arXiv:2607.22646v1 Announce Type: new Abstract: Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but

Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling

Model ReleasesDGX agent

arXiv:2607.22847v1 Announce Type: new Abstract: Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typica

Generative Artificial Intelligence (GenAI) to convert images of queuing networks into verifiable simulation models: an open-weight LLM workflow approach

ResearchDGX agent

arXiv:2607.24259v1 Announce Type: new Abstract: Recent work has explored the use of Large Language Models (LLMs) to automate simulation model building, typically by generating executable code directly

GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.23913v1 Announce Type: new Abstract: Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substant

Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families

Model ReleasesDGX agent

arXiv:2607.24339v1 Announce Type: new Abstract: Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stu

Infinite-Precision Autoregressive Modeling for Vector Graphics and Layouts

Model ReleasesDGX agent

arXiv:2601.05680v2 Announce Type: replace-cross Abstract: While Transformer-based autoregressive models excel in data generation, their token discretization strategy inherently limits their precision

INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models

Model ReleasesDGX agent

arXiv:2607.24273v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential in financial reasoning, but existing benchmarks often evaluate domain knowledge, numerical reas

Interesting. Grok 4.6 releases around August 7. This will be the 1.5T model with significantly improved SFT & RL. Grok 4.7 will be the 2.1T …

Model ReleasesDGX agent

Interesting. Grok 4.6 releases around August 7. This will be the 1.5T model with significantly improved SFT & RL. Grok 4.7 will be the 2.1T model released a few weeks later. This will be better than 4

Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning

Model ReleasesDGX agent

arXiv:2603.09255v2 Announce Type: replace-cross Abstract: Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars. These techniques play

Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models

Model ReleasesDGX agent

arXiv:2601.02580v2 Announce Type: replace-cross Abstract: Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing

SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion

Model ReleasesDGX agent

arXiv:2607.23580v1 Announce Type: new Abstract: Existing vector-sketch models treat recognition and generation as separate tasks, leaving a gap for streaming interfaces that must understand a drawing

Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.23474v1 Announce Type: new Abstract: This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions

Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It

Model ReleasesDGX agent

arXiv:2607.23893v1 Announce Type: cross Abstract: Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the questio

27 Jul 2026

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

Model ReleasesDGX agent

arXiv:2607.21988v1 Announce Type: new Abstract: Self-harm content is particularly challenging to detect using NLP techniques, and is also a high-stakes task which requires the highest accuracy to enab

Autoregressive EHR Foundation Models with Multimodal Inputs

SafetyDGX agent

arXiv:2607.22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on st

Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination

Model ReleasesDGX agent

arXiv:2607.22067v1 Announce Type: new Abstract: The integration of large language models (LLMs) into the nuclear power industry requires outputs grounded in domain-specific knowledge. This study evalu

Correlation-Aware and Gaussianity-Preserving Robust Latent Angular Watermarking for Diffusion Models

ResearchDGX agent

arXiv:2607.22386v1 Announce Type: new Abstract: Latent domain watermarking for diffusion models embeds watermarks directly into the latent prior, enjoying non-intrusiveness to model parameters and sea

InnoText: A Unified Model for Visual Text Generation and Editing

ResearchDGX agent

arXiv:2607.22101v1 Announce Type: new Abstract: Diffusion models have recently achieved remarkable success in high-fidelity image synthesis, yet their application to visual text generation and editing

Kimi K3 is now available on Ollama’s cloud. To use it with Claude Code, run: ollama launch claude --model kimi-k3:cloud Currently Kimi K3 re…

Model ReleasesDGX agent

Kimi K3 is now available on Ollama’s cloud. To use it with Claude Code, run: ollama launch claude --model kimi-k3:cloud Currently Kimi K3 requires a Pro or Max subscription, and consumes extra usage c

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

ResearchDGX agent

arXiv:2607.21655v1 Announce Type: cross Abstract: Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is co

24 Jul 2026

A Sovereign, Open-Source Foundation Model for German and English

Model ReleasesDGX agent

arXiv:2607.09424v3 Announce Type: replace-cross Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English

An LLM-Driven Workflow for Automated Process Control Strategy Generation and Tuning from Dynamic Process Models

Model ReleasesDGX agent

arXiv:2607.21292v1 Announce Type: new Abstract: We present a structured large-language-model-driven workflow for automated multi-variable control design from dynamic process models. The workflow decom

Anthropic launches Claude Opus 5, which it says comes close to Fable 5 performance at half the price; it is the new default model on Claude Max (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic launches Claude Opus 5, which it says comes close to Fable 5 performance at half the price; it is the new default model on Claude Max — Claude Opus 5 is available today. It's a th

Expectation Alignment of Language Models for Real-World User Expectations

Model ReleasesDGX agent

arXiv:2607.20485v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet

HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.21105v1 Announce Type: new Abstract: Although Multimodal Large Language Models have achieved strong performance across a wide range of vision-language tasks, they still suffer from hallucin

Incomplete Prompt Jailbreaks in Large Language Models

Model ReleasesDGX agent

arXiv:2607.20473v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion

Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing

Model ReleasesDGX agent

arXiv:2607.20433v1 Announce Type: cross Abstract: While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alternative to ful

Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most excitin…

Model ReleasesDGX agent

Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt inject

Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents

Model ReleasesDGX agent

arXiv:2602.10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyper

Statistical Inference for Generative Model Comparison

Model ReleasesDGX agent

arXiv:2501.18897v4 Announce Type: replace-cross Abstract: Generative models have achieved remarkable success across a range of applications, yet their evaluation still lacks principled uncertainty qua

23 Jul 2026

Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models

ResearchDGX agent

arXiv:2607.19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and th

Comparing Model-agnostic Feature Selection Methods through Relative Efficiency

ResearchDGX agent

arXiv:2508.14268v2 Announce Type: replace-cross Abstract: Feature selection and importance estimation in a model-agnostic setting is an ongoing challenge of significant interest. Wrapper methods are c

Germany's Black Forest Labs launches Flux 3 and Flux-mimic, its first models for robotics, as it expands from generative AI into physical AI (Yazhou Sun/Bloomberg)

Model ReleasesDGX agent

Yazhou Sun / Bloomberg: Germany's Black Forest Labs launches Flux 3 and Flux-mimic, its first models for robotics, as it expands from generative AI into physical AI — Black Forest Labs, a German artif

Good Practice Guide for quantifying uncertainties for machine learning models applied to photoplethysmography signals

Model ReleasesDGX agent

arXiv:2607.19999v1 Announce Type: new Abstract: This Good Practice Guide presents work done in the QUMPHY project (Uncertainty quantification for machine learning models applied to photoplethysmograph

KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding

Model ReleasesDGX agent

arXiv:2607.19876v1 Announce Type: new Abstract: Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-loop evaluation via simulator rollouts off

Masked Visual Actions for Unified World Modeling

SafetyDGX agent

arXiv:2607.19343v1 Announce Type: new Abstract: Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

Model ReleasesDGX agent

arXiv:2607.20058v1 Announce Type: new Abstract: Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics.

When Visual Evidence is Ambiguous: Pareidolia as a Diagnostic Probe for Vision Models

Model ReleasesDGX agent

arXiv:2603.03989v3 Announce Type: replace-cross Abstract: When visual evidence is ambiguous, vision models must decide how to interpret face-like patterns. Face pareidolia, the perception of faces in

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Model ReleasesDGX agent

arXiv:2607.15330v2 Announce Type: replace Abstract: We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a

21 Jul 2026

Reproducing OpenAI’s “persistently beneficial models” - GRPO trait install barely moves. Ideas? [P] [R]

Model ReleasesDGX agent

TL;DR: I’m reproducing the trait-persistence result from arXiv:2606.24014 on one RTX 3090. Before I can test persistence I need to install a trait via RL — and my GRPO run moves the trait only +2.4 po

16 Jul 2026

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

Model ReleasesDGX agent

arXiv:2603.00546v2 Announce Type: replace Abstract: Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging paradigm acr

Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering

Model ReleasesDGX agent

arXiv:2607.13568v1 Announce Type: cross Abstract: Can a language model estimate its familiarity with an entity before generating an answer? We study activations at the final prompt token in twelve ins

MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model

Model ReleasesDGX agent

arXiv:2607.13763v1 Announce Type: cross Abstract: Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the lowest in-

S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.13926v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated remarkable potential for high-level reasoning in autonomous driving, yet they fundamentally struggle to

Self-Improving is Often Sudden: Enlightenment-style Finetuning for Large-Scale Models

ResearchDGX agent

arXiv:2607.13395v1 Announce Type: new Abstract: The pursuit of autonomously self-improving models has attracted growing interest in the era of large-scale foundation models. Drawing inspiration from t

15 Jul 2026

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

AgentsDGX agent

arXiv:2607.11183v2 Announce Type: replace Abstract: Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plau

Belief-reality separation lives in routing over a shared value slot in language models

Model ReleasesDGX agent

arXiv:2607.11945v1 Announce Type: new Abstract: Capable language models hold what a character believes apart from what is true: told 'Anna believes the cup is blue; in reality it is red,' they answer

Can a Language Model Learn Facts Continually in Its Weights?

TutorialsDGX agent

arXiv:2607.11020v2 Announce Type: replace Abstract: Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. Whether wei

Saturation Makes Quantization Error Additive: A Coverage Model with a Certificate

ResearchDGX agent

arXiv:2607.12266v1 Announce Type: new Abstract: Mixed-precision quantization must decide which parts of a model to keep at higher precision. A common premise, shared by sensitivity-based methods such

So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis

Model ReleasesDGX agent

arXiv:2607.11890v1 Announce Type: cross Abstract: Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditiona

Thinking Machines Lab debuts Inkling, an open-weight MoE model with 975B total and 41B active parameters, trained to be broad rather than optimized for one area (Thinking Machines Lab)

IndustryDGX agent

Thinking Machines Lab: Thinking Machines Lab debuts Inkling, an open-weight MoE model with 975B total and 41B active parameters, trained to be broad rather than optimized for one area — Try on Tinker

Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage Control

Model ReleasesDGX agent

arXiv:2607.12856v1 Announce Type: new Abstract: Buildings are expected to shift cooling loads in response to grid conditions. Thermal energy storage (TES) enables this shift, but scheduling it well re

← Previous
1…3536373839…998
Next →