AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,490 results
11 May 2026

OpenAI launches Daybreak, a cybersecurity initiative integrating AI models and Codex Security to help organizations patch vulnerabilities (Alexey Shabanov/TestingCatalog AI News)

Model ReleasesDGX agent

Alexey Shabanov / TestingCatalog AI News: OpenAI launches Daybreak, a cybersecurity initiative integrating AI models and Codex Security to help organizations patch vulnerabilities — OpenAI launches Da

Predictive but Not Plannable: RC-aux for Latent World Models

Local AiDGX agent

arXiv:2605.07278v1 Announce Type: cross Abstract: A latent world model may achieve accurate short-horizon prediction while still inducing a latent space that is poorly aligned with planning. A key iss

Self-Consolidating Language Models: Continual Knowledge Incorporation from Context

ResearchDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.07076v1 Announce Type: new Abstract: Large language models (LLMs) increasingly receive information as streams of passages, conversations, and long-context workflows. While longer context wi

SOD: Step-wise On-policy Distillation for Small Language Model Agents

SafetyDGX agent

arXiv:2605.07725v1 Announce Type: cross Abstract: Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model

ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation

Local AiDGX agent

arXiv:2605.07390v1 Announce Type: new Abstract: Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spati

The Convergence Gap: Instruction-Tuned Language Models Stabilize Later in the Forward Pass

Model ReleasesDGX agent

arXiv:2605.07282v1 Announce Type: new Abstract: Final outputs hide when a checkpoint commits to its next-token prediction. We introduce the convergence gap, a model-diffing diagnostic that decodes eac

Tracing Uncertainty in Language Model 'Reasoning'

Model ReleasesDGX agent

arXiv:2605.07776v1 Announce Type: cross Abstract: Language model (LM) 'reasoning', commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics u

TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models

Model ReleasesDGX agent

arXiv:2601.18744v2 Announce Type: replace Abstract: Time series are ubiquitous in real-world scenarios and crucial for applications ranging from energy management to traffic control. Consequently, the

Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions

Model ReleasesDGX agent

arXiv:2605.07984v1 Announce Type: cross Abstract: We study planning site formation in language models -- where internal representations of structurally-constrained future tokens form during the forwar

10 May 2026

Detailed review and guide from my testing of local ollama setup with DeepSeek models (Ryzen APU's only)

Model ReleasesDGX agent

This post provides a detailed review and practical guide for setting up and testing Ollama with DeepSeek models specifically on Ryzen APU systems. It likely covers performance benchmarks, configuratio

9 May 2026

100% agree on the Context Hub. Developers constantly tweak their approaches to manage context in their prompts with each new model and tool …

Model ReleasesDGX agent

100% agree on the Context Hub. Developers constantly tweak their approaches to manage context in their prompts with each new model and tool suite release. I would even say that the core problem we fac

Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers — Last year, we released a case study on

8 May 2026

Chrome's 4GB AI model isn't new, but you're not wrong for being confused

Local AiDGX agent

Google has offered Gemini Nano for Chrome since 2024 as a lightweight, on-device model , but users reasonably expect the visible AI Mode to use the on-device model with queries staying local, when in

CyberSecQwen-4B: Why Defensive Cyber Needs Small, Specialized, Locally-Runnable Models

ToolsDGX agent

CyberSecQwen-4B is a specialized 4-billion parameter language model designed for cybersecurity defense tasks that can run locally on standard hardware. The model addresses the need for small, efficien

7 May 2026

Advancing voice intelligence with new models in the API

Model ReleasesDGX agent

OpenAI announced new voice intelligence models available through its API, expanding capabilities for developers to integrate advanced voice processing and understanding features into their application

Anthropic researchers detail 'natural language autoencoders', which convert LLM activations, the numbers encoding a model's thoughts, into natural language text (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic researchers detail “natural language autoencoders”, which convert LLM activations, the numbers encoding a model's thoughts, into natural language text — When you talk to an AI mod

Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

Model ReleasesDGX agent

arXiv:2605.05092v1 Announce Type: cross Abstract: Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models for

Efficient Model-Based Reinforcement Learning for Robot Control via Online Optimization

SafetyDGX agent

arXiv:2510.18518v2 Announce Type: replace Abstract: We present an online model-based reinforcement learning algorithm suitable for controlling complex robotic systems directly in the real world. Unlik

External Validation of Deep Learning Models for BI-RADS Breast Density Prediction from Ultrasound Images

ResearchDGX agent

arXiv:2605.05082v1 Announce Type: cross Abstract: We externally validated three deep learning models (DenseNet121, ViT-B/32, and ResNet50) for predicting mammographic breast density from breast ultras

Full-chip CMP modelling based on Fully Convolutional Network leveraging White Light Interferometry

ResearchDGX agent

arXiv:2605.05062v1 Announce Type: new Abstract: As time-to-market is crucial in the Integrated Circuit (IC) industry, speeding up layout manufacturability verifi-cation is essential. Chemical-Mechanic

Introducing GPT-Realtime-2 in the API: our most intelligent voice model yet, bringing GPT-5-class reasoning to voice agents. Voice agents ar…

Model ReleasesDGX agent

Introducing GPT-Realtime-2 in the API: our most intelligent voice model yet, bringing GPT-5-class reasoning to voice agents. Voice agents are now real-time collaborators that can listen, reason, and s

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)

Model ReleasesDGX agent

arXiv:2605.05187v1 Announce Type: new Abstract: This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and

Manifold of Failure: Behavioral Attraction Basins in Language Models

Model ReleasesDGX agent

arXiv:2602.22291v3 Announce Type: replace Abstract: While prior work has focused on projecting adversarial examples back onto the manifold of natural data to restore safety, we argue that a comprehens

Multi-site modelling and reconstruction of past extreme skew surges along the French Atlantic coast

ResearchDGX agent

arXiv:2505.00835v2 Announce Type: replace-cross Abstract: Appropriate modelling of extreme skew surges is crucial, particularly for coastal risk management. Our study focuses on modelling extreme skew

Saw this and thought 'yes! ChatGPT voice mode is going to stop acting like a two-year-model' but that upgrade hasn't shipped just yet

Model ReleasesDGX agent

Saw this and thought 'yes! ChatGPT voice mode is going to stop acting like a two-year-model' but that upgrade hasn't shipped just yet Introducing GPT-Realtime-2 in the API: our most intelligent voice

6 May 2026

A TLDR on Harness Profiles: ✅ Model-specific profiles to adjust prompts, tools, and middleware. 📦 Profiles for @OpenAI, @Anthropic, and @Go…

AgentsDGX agent

A TLDR on Harness Profiles: ✅ Model-specific profiles to adjust prompts, tools, and middleware. 📦 Profiles for @OpenAI, @Anthropic, and @Google models out of the box. 📈 A 10–20 point jump on a subset

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing

Model ReleasesDGX agent

arXiv:2605.03903v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capabi

Code World Model Preparedness Report

ResearchDGX agent

arXiv:2605.00932v1 Announce Type: cross Abstract: This report documents the preparedness assessment of Code World Model (CWM), a model for code generation and reasoning about code from Meta. We conduc

Complexity Horizons of Compressed Models in Analog Circuit Analysis

AgentsDGX agent

arXiv:2605.02285v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) for specialized engineering domains, such as circuit analysis, often faces a trade-off between reasoning

Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models

Model ReleasesDGX agent

arXiv:2605.03547v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs), trained on web-scale data, risk memorizing and regenerating copyrighted visual content such as characters and logo

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior

ResearchDGX agent

arXiv:2605.03855v1 Announce Type: new Abstract: Human-AI collaboration requires AI agents to understand human behavior for effective coordination. While advances in foundation models show promising ca

FinSTaR: Towards Financial Reasoning with Time Series Reasoning Models

Model ReleasesDGX agent

arXiv:2605.03460v1 Announce Type: cross Abstract: Time series (TS) reasoning models (TSRMs) have shown promising capabilities in general domains, yet they consistently fail on financial domain, which

Google releases Multi-Token Prediction drafters for its Gemma 4 models, which use a form of speculative decoding to guess future tokens for faster inference (Ryan Whitwam/Ars Technica)

Model ReleasesDGX agent

Ryan Whitwam / Ars Technica: Google releases Multi-Token Prediction drafters for its Gemma 4 models, which use a form of speculative decoding to guess future tokens for faster inference — Google launc

GRIFDIR: Graph Resolution-Invariant FEM Diffusion Models in Function Spaces over Irregular Domains

ResearchDGX agent

arXiv:2605.03497v1 Announce Type: new Abstract: Score-based diffusion models in infinite-dimensional function spaces provide a mathematically principled framework for modelling function-valued data, o

Hierarchical Memorization in Large Language Models: Evidence from Citation Generation

Model ReleasesDGX agent

arXiv:2511.08877v2 Announce Type: replace Abstract: Large language models (LLMs) generate fluent text across a wide range of tasks, but the fabrication of non-existent academic citations remains a cri

Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models

Model ReleasesDGX agent

arXiv:2605.03438v1 Announce Type: new Abstract: Pre-trained 3D point cloud foundation models (PFMs) have demonstrated strong transferability across diverse downstream tasks. However, full fine-tuning

Mitigating Frequency Learning Bias in Quantum Models via Multi-Stage Residual Learning

Model ReleasesDGX agent

arXiv:2603.10083v2 Announce Type: replace-cross Abstract: Quantum machine learning models based on parameterized circuits can be viewed as Fourier series approximators. However, they often struggle to

ORPilot: A Production-Oriented Agentic LLM-for-OR Tool for Optimization Modeling

Model ReleasesDGX agent

arXiv:2605.02728v1 Announce Type: new Abstract: This paper presents ORPilot, an open-source agentic AI system that translates real-world business problems into solver-ready optimization models. Unlike

Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

SafetyDGX agent

arXiv:2605.01147v1 Announce Type: new Abstract: As large language models are increasingly deployed as interacting agents in high-stakes decisions, the AI safety community assumes that safety propertie

Safety and accuracy follow different scaling laws in clinical large language models

Model ReleasesDGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

Should We Still Pretrain Encoders with Masked Language Modeling?

ResearchDGX agent

arXiv:2507.00994v4 Announce Type: replace Abstract: Learning high-quality text representations is fundamental to a wide range of NLP tasks. While encoder pretraining has traditionally relied on Masked

SPRINT: Robust Model Attribution of Generated Images via Secret Pixel Reconstruction

ResearchDGX agent

arXiv:2508.05691v3 Announce Type: replace-cross Abstract: Detecting the source model of AI-generated images is a growing accountability problem. AI fingerprinting techniques address this by detecting

Towards Understanding Specification Gaming in Reasoning Models

Model ReleasesDGX agent

arXiv:2605.02269v1 Announce Type: new Abstract: Specification gaming is a critical failure mode of LLM agents. Despite this, there has been little systematic research into when it arises and what driv

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

Model ReleasesDGX agent

arXiv:2605.03276v1 Announce Type: new Abstract: Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage i

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.01449v1 Announce Type: cross Abstract: Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, sug

We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-paramet…

Model ReleasesDGX agent

We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-parameter LLMs. With CuTeDSL integrated into our inference engine,

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

Model ReleasesDGX agent

arXiv:2605.03475v1 Announce Type: new Abstract: Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal t

5 May 2026

A Systematic Benchmark of Machine Transliteration Models for the Tajik-Farsi Language Pair: A Comparative Study from Rule-Based to Transformer Architectures

Model ReleasesDGX agent

arXiv:2605.02270v1 Announce Type: new Abstract: This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic sc

ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts

Model ReleasesDGX agent

arXiv:2605.00245v1 Announce Type: new Abstract: Large language models (LLMs) are now being explored for defense applications that require reliable and legally compliant decision support. They also hol

Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models

SafetyDGX agent

arXiv:2605.01451v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly being integrated into high-stakes public safety systems, including emergency call triage and dispatch decision

Controlled Paraphrase Geometry in Sentence Embedding Space: Local Manifold Modeling and Latent Probing

ResearchDGX agent

arXiv:2605.01073v1 Announce Type: new Abstract: The paper studies the local geometry of embedding clouds induced by controlled local classes of semantically close sentences. The central question is ho

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models

SafetyDGX agent

arXiv:2605.01896v1 Announce Type: new Abstract: Emerging multi-modal world models attempt to jointly generate videos across diverse modalities (e.g., RGB, depth, and mask), yet they fail to fully expl

Earth System Foundation Model (ESFM): A unified framework for heterogeneous data integration and forecasting

TutorialsDGX agent

arXiv:2605.00850v1 Announce Type: cross Abstract: Foundation models (FMs) for the Earth system learn statistical relationships between physical variables across massive datasets to enable versatile do

Geospatial foundation-model embeddings improve population estimation unevenly across space and scale

Model ReleasesDGX agent

arXiv:2605.01650v1 Announce Type: new Abstract: Reliable subnational population estimates are essential for applications, yet remain difficult where censuses are sparse, outdated or spatially coarse.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models

SafetyDGX agent

arXiv:2605.02626v1 Announce Type: new Abstract: Preference optimization has become a central paradigm for aligning large language models with human feedback. Direct Preference Optimization (DPO) simpl

llms are getting expensive why we need oss models

AgentsDGX agent

Large language models are becoming increasingly costly to develop and operate, creating a need for open-source alternatives that reduce dependency on expensive proprietary models and make AI more acce

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice

Model ReleasesDGX agent

arXiv:2605.01333v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have emerged as a promising paradigm for dental image analysis. However, their ability to capture the multi-lev

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

Model ReleasesDGX agent

arXiv:2603.27259v2 Announce Type: replace Abstract: Long video understanding (LVU) remains a core challenge in multimodal learning. Although recent vision-language models (VLMs) have made notable prog

SteeringDiffusion: A Bottlenecked Activation Control Interface for Diffusion Models

Model ReleasesDGX agent

arXiv:2605.01653v1 Announce Type: new Abstract: We introduce SteeringDiffusion, a bottlenecked activation-level control interface for diffusion models that exposes a smooth, monotonic, and runtime-adj

TRAP: Tail-aware Ranking Attack for World-Model Planning

SafetyDGX agent

arXiv:2605.01950v1 Announce Type: new Abstract: World models enable long-horizon planning by internally generating and evaluating imagined trajectories, making them a promising foundation for generali

← Previous
1…9091929394…1009
Next →