AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,521 results
Model Releases

Tracing Uncertainty in Language Model 'Reasoning'

DGX agent

arXiv:2605.07776v1 Announce Type: cross Abstract: Language model (LM) 'reasoning', commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics u

model-releasesarxiv-cs-ai
11 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models

DGX agent

arXiv:2601.18744v2 Announce Type: replace Abstract: Time series are ubiquitous in real-world scenarios and crucial for applications ranging from energy management to traffic control. Consequently, the

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions

DGX agent

arXiv:2605.07984v1 Announce Type: cross Abstract: We study planning site formation in language models -- where internal representations of structurally-constrained future tokens form during the forwar

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Detailed review and guide from my testing of local ollama setup with DeepSeek models (Ryzen APU's only)

DGX agent

This post provides a detailed review and practical guide for setting up and testing Ollama with DeepSeek models specifically on Ryzen APU systems. It likely covers performance benchmarks, configuratio

model-releasesr-ollama
10 May 2026
Model Releases

100% agree on the Context Hub. Developers constantly tweak their approaches to manage context in their prompts with each new model and tool …

DGX agent

100% agree on the Context Hub. Developers constantly tweak their approaches to manage context in their prompts with each new model and tool suite release. I would even say that the core problem we fac

model-releasesharrison-chase--x
9 May 2026
Model Releases

Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers (Anthropic)

DGX agent

Anthropic: Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers — Last year, we released a case study on

model-releasestechmeme
9 May 2026
Local Ai

Chrome's 4GB AI model isn't new, but you're not wrong for being confused

DGX agent

Google has offered Gemini Nano for Chrome since 2024 as a lightweight, on-device model , but users reasonably expect the visible AI Mode to use the on-device model with queries staying local, when in

local-aiars-technica
8 May 2026
Tools

CyberSecQwen-4B: Why Defensive Cyber Needs Small, Specialized, Locally-Runnable Models

DGX agent

CyberSecQwen-4B is a specialized 4-billion parameter language model designed for cybersecurity defense tasks that can run locally on standard hardware. The model addresses the need for small, efficien

toolshugging-face
8 May 2026
Model Releases

Advancing voice intelligence with new models in the API

DGX agent

OpenAI announced new voice intelligence models available through its API, expanding capabilities for developers to integrate advanced voice processing and understanding features into their application

model-releasesopenai
7 May 2026
Model Releases

Anthropic researchers detail 'natural language autoencoders', which convert LLM activations, the numbers encoding a model's thoughts, into natural language text (Anthropic)

DGX agent

Anthropic: Anthropic researchers detail “natural language autoencoders”, which convert LLM activations, the numbers encoding a model's thoughts, into natural language text — When you talk to an AI mod

model-releasestechmeme
7 May 2026
Model Releases

Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

DGX agent

arXiv:2605.05092v1 Announce Type: cross Abstract: Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models for

model-releasesarxiv-cs-cv
7 May 2026
Safety

Efficient Model-Based Reinforcement Learning for Robot Control via Online Optimization

DGX agent

arXiv:2510.18518v2 Announce Type: replace Abstract: We present an online model-based reinforcement learning algorithm suitable for controlling complex robotic systems directly in the real world. Unlik

safetyarxiv-cs-ro
7 May 2026
Research

External Validation of Deep Learning Models for BI-RADS Breast Density Prediction from Ultrasound Images

DGX agent

arXiv:2605.05082v1 Announce Type: cross Abstract: We externally validated three deep learning models (DenseNet121, ViT-B/32, and ResNet50) for predicting mammographic breast density from breast ultras

researcharxiv-cs-cv
7 May 2026
Research

Full-chip CMP modelling based on Fully Convolutional Network leveraging White Light Interferometry

DGX agent

arXiv:2605.05062v1 Announce Type: new Abstract: As time-to-market is crucial in the Integrated Circuit (IC) industry, speeding up layout manufacturability verifi-cation is essential. Chemical-Mechanic

researcharxiv-cs-lg
7 May 2026
Model Releases

Introducing GPT-Realtime-2 in the API: our most intelligent voice model yet, bringing GPT-5-class reasoning to voice agents. Voice agents ar…

DGX agent

Introducing GPT-Realtime-2 in the API: our most intelligent voice model yet, bringing GPT-5-class reasoning to voice agents. Voice agents are now real-time collaborators that can listen, reason, and s

model-releasesopenai--x
7 May 2026
Model Releases

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)

DGX agent

arXiv:2605.05187v1 Announce Type: new Abstract: This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Manifold of Failure: Behavioral Attraction Basins in Language Models

DGX agent

arXiv:2602.22291v3 Announce Type: replace Abstract: While prior work has focused on projecting adversarial examples back onto the manifold of natural data to restore safety, we argue that a comprehens

model-releasesarxiv-cs-lg
7 May 2026
Research

Multi-site modelling and reconstruction of past extreme skew surges along the French Atlantic coast

DGX agent

arXiv:2505.00835v2 Announce Type: replace-cross Abstract: Appropriate modelling of extreme skew surges is crucial, particularly for coastal risk management. Our study focuses on modelling extreme skew

researcharxiv-cs-lg
7 May 2026
Model Releases

Saw this and thought 'yes! ChatGPT voice mode is going to stop acting like a two-year-model' but that upgrade hasn't shipped just yet

DGX agent

Saw this and thought 'yes! ChatGPT voice mode is going to stop acting like a two-year-model' but that upgrade hasn't shipped just yet Introducing GPT-Realtime-2 in the API: our most intelligent voice

model-releasessimon-willison--x
7 May 2026
Agents

A TLDR on Harness Profiles: ✅ Model-specific profiles to adjust prompts, tools, and middleware. 📦 Profiles for @OpenAI, @Anthropic, and @Go…

DGX agent

A TLDR on Harness Profiles: ✅ Model-specific profiles to adjust prompts, tools, and middleware. 📦 Profiles for @OpenAI, @Anthropic, and @Google models out of the box. 📈 A 10–20 point jump on a subset

agentsharrison-chase--x
6 May 2026
Model Releases

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing

DGX agent

arXiv:2605.03903v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capabi

model-releasesarxiv-cs-cl
6 May 2026
Research

Code World Model Preparedness Report

DGX agent

arXiv:2605.00932v1 Announce Type: cross Abstract: This report documents the preparedness assessment of Code World Model (CWM), a model for code generation and reasoning about code from Meta. We conduc

researcharxiv-cs-ai
6 May 2026
Agents

Complexity Horizons of Compressed Models in Analog Circuit Analysis

DGX agent

arXiv:2605.02285v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) for specialized engineering domains, such as circuit analysis, often faces a trade-off between reasoning

agentsarxiv-cs-ai
6 May 2026
Model Releases

Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models

DGX agent

arXiv:2605.03547v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs), trained on web-scale data, risk memorizing and regenerating copyrighted visual content such as characters and logo

model-releasesarxiv-cs-cv
6 May 2026
Research

Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior

DGX agent

arXiv:2605.03855v1 Announce Type: new Abstract: Human-AI collaboration requires AI agents to understand human behavior for effective coordination. While advances in foundation models show promising ca

researcharxiv-cs-ro
6 May 2026
Model Releases

FinSTaR: Towards Financial Reasoning with Time Series Reasoning Models

DGX agent

arXiv:2605.03460v1 Announce Type: cross Abstract: Time series (TS) reasoning models (TSRMs) have shown promising capabilities in general domains, yet they consistently fail on financial domain, which

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

Google releases Multi-Token Prediction drafters for its Gemma 4 models, which use a form of speculative decoding to guess future tokens for faster inference (Ryan Whitwam/Ars Technica)

DGX agent

Ryan Whitwam / Ars Technica: Google releases Multi-Token Prediction drafters for its Gemma 4 models, which use a form of speculative decoding to guess future tokens for faster inference — Google launc

model-releasestechmeme
6 May 2026
Research

GRIFDIR: Graph Resolution-Invariant FEM Diffusion Models in Function Spaces over Irregular Domains

DGX agent

arXiv:2605.03497v1 Announce Type: new Abstract: Score-based diffusion models in infinite-dimensional function spaces provide a mathematically principled framework for modelling function-valued data, o

researcharxiv-cs-lg
6 May 2026
Model Releases

Hierarchical Memorization in Large Language Models: Evidence from Citation Generation

DGX agent

arXiv:2511.08877v2 Announce Type: replace Abstract: Large language models (LLMs) generate fluent text across a wide range of tasks, but the fabrication of non-existent academic citations remains a cri

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models

DGX agent

arXiv:2605.03438v1 Announce Type: new Abstract: Pre-trained 3D point cloud foundation models (PFMs) have demonstrated strong transferability across diverse downstream tasks. However, full fine-tuning

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Mitigating Frequency Learning Bias in Quantum Models via Multi-Stage Residual Learning

DGX agent

arXiv:2603.10083v2 Announce Type: replace-cross Abstract: Quantum machine learning models based on parameterized circuits can be viewed as Fourier series approximators. However, they often struggle to

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

ORPilot: A Production-Oriented Agentic LLM-for-OR Tool for Optimization Modeling

DGX agent

arXiv:2605.02728v1 Announce Type: new Abstract: This paper presents ORPilot, an open-source agentic AI system that translates real-world business problems into solver-ready optimization models. Unlike

model-releasesarxiv-cs-ai
6 May 2026
Safety

Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment

DGX agent

arXiv:2605.01147v1 Announce Type: new Abstract: As large language models are increasingly deployed as interacting agents in high-stakes decisions, the AI safety community assumes that safety propertie

safetyarxiv-cs-ai
6 May 2026
Model Releases

Safety and accuracy follow different scaling laws in clinical large language models

DGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

model-releasesarxiv-cs-cl
6 May 2026
Research

Should We Still Pretrain Encoders with Masked Language Modeling?

DGX agent

arXiv:2507.00994v4 Announce Type: replace Abstract: Learning high-quality text representations is fundamental to a wide range of NLP tasks. While encoder pretraining has traditionally relied on Masked

researcharxiv-cs-cl
6 May 2026
Research

SPRINT: Robust Model Attribution of Generated Images via Secret Pixel Reconstruction

DGX agent

arXiv:2508.05691v3 Announce Type: replace-cross Abstract: Detecting the source model of AI-generated images is a growing accountability problem. AI fingerprinting techniques address this by detecting

researcharxiv-cs-lg
6 May 2026
Model Releases

Towards Understanding Specification Gaming in Reasoning Models

DGX agent

arXiv:2605.02269v1 Announce Type: new Abstract: Specification gaming is a critical failure mode of LLM agents. Despite this, there has been little systematic research into when it arises and what driv

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

DGX agent

arXiv:2605.03276v1 Announce Type: new Abstract: Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage i

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models

DGX agent

arXiv:2605.01449v1 Announce Type: cross Abstract: Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, sug

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-paramet…

DGX agent

We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-parameter LLMs. With CuTeDSL integrated into our inference engine,

model-releasesperplexity--x
6 May 2026
Model Releases

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

DGX agent

arXiv:2605.03475v1 Announce Type: new Abstract: Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal t

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

A Systematic Benchmark of Machine Transliteration Models for the Tajik-Farsi Language Pair: A Comparative Study from Rule-Based to Transformer Architectures

DGX agent

arXiv:2605.02270v1 Announce Type: new Abstract: This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic sc

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts

DGX agent

arXiv:2605.00245v1 Announce Type: new Abstract: Large language models (LLMs) are now being explored for defense applications that require reliable and legally compliant decision support. They also hol

model-releasesarxiv-cs-ai
5 May 2026
Safety

Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models

DGX agent

arXiv:2605.01451v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly being integrated into high-stakes public safety systems, including emergency call triage and dispatch decision

safetyarxiv-cs-cl
5 May 2026
Research

Controlled Paraphrase Geometry in Sentence Embedding Space: Local Manifold Modeling and Latent Probing

DGX agent

arXiv:2605.01073v1 Announce Type: new Abstract: The paper studies the local geometry of embedding clouds induced by controlled local classes of semantically close sentences. The central question is ho

researcharxiv-cs-cl
5 May 2026
Safety

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models

DGX agent

arXiv:2605.01896v1 Announce Type: new Abstract: Emerging multi-modal world models attempt to jointly generate videos across diverse modalities (e.g., RGB, depth, and mask), yet they fail to fully expl

safetyarxiv-cs-cv
5 May 2026
Tutorials

Earth System Foundation Model (ESFM): A unified framework for heterogeneous data integration and forecasting

DGX agent

arXiv:2605.00850v1 Announce Type: cross Abstract: Foundation models (FMs) for the Earth system learn statistical relationships between physical variables across massive datasets to enable versatile do

tutorialsarxiv-cs-lg
5 May 2026
Model Releases

Geospatial foundation-model embeddings improve population estimation unevenly across space and scale

DGX agent

arXiv:2605.01650v1 Announce Type: new Abstract: Reliable subnational population estimates are essential for applications, yet remain difficult where censuses are sparse, outdated or spatially coarse.

model-releasesarxiv-cs-lg
5 May 2026
← Previous
1…113114115116117…1261
Next →