TechniqueRLHF / Alignment1 recent entries22 Jul 2026How to measure human-LLM judge alignmentNo single metric proves an LLM judge is trustworthy. This field guide shows how to measure human–human agreement, compare it to LLM–human agreement, and diagnose errors with precision, recall, and F1.
TechniqueAgents8 recent entries27 Jul 2026From traditional ML to AI agents: How Booking.com scales AI observability with ArizeHow Booking.com built a unified AI observability stack with Arize for agentic GenAI workflows and traditional ML — from telemetry collection and PII redaction to latency monitors and evaluations. The
TechniqueFine-tuning3 recent entries27 May 2026From production traces to better AI agents: Automating the LLMOps feedback loopProduction AI traces are the raw material for better evals, prompts, datasets, and fine-tuned models. This post shows how the Arize AX Airflow Provider turns that feedback loop into scheduled, monitor→2 Jun 2026The end of fine-tuning: Why evals, context, and traces matter moreFine-tuning isn't dead, but the way most teams iterate on AI products has split in two. A tiny fraction run continuous RL against their own environments; everyone else has moved the iteration loop out→13 Jul 2026How do you make an LLM, anyway? Microsoft just published a textbook.Microsoft published a 109-page technical report on MAI-Thinking-1. Here’s the abbreviated version of how a modern lab actually trains a frontier reasoning model — from scraping the web to reinforcemen
TechniqueMultimodal1 recent entries4 Jun 2026Building the AI factory for self-improving agents: What’s new in Arize AXArize AX is adding managed agents, full-agent experimentation, expanded multimodal support, and Harness-as-a-Judge to help teams observe, evaluate, and improve production agents. The post Building the
TechniqueSafety3 recent entries14 Apr 2026Building smarter AI agents: architecture, evals, and lessons from the fieldShipping an AI agent is easy. Understanding whether it actually works in production is not. That was the common thread across two AI Builders events in San Francisco at GitHub... The post Building sma→5 May 2026AI agent evaluation: How to test, debug, and improve agents in productionAI agents require specialized testing and debugging approaches that differ from traditional software due to their non-deterministic behavior and complex decision-making processes. This entry likely co→22 Jul 2026How to measure human-LLM judge alignmentNo single metric proves an LLM judge is trustworthy. This field guide shows how to measure human–human agreement, compare it to LLM–human agreement, and diagnose errors with precision, recall, and F1.