Agents

Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation

arXiv:2608.02345v1 Announce Type: new Abstract: A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering e

DGX agentpaper
agentsarxiv-cs-cl

arXiv:2608.02345v1 Announce Type: new Abstract: A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time. Can AI agents---conditioned on behavioral profiles and contextual descriptions of the intervention---simulate outcomes accurately enough to vet candidate treatments before committing live traffic? We formalize this question as a Simulated Randomized Controlled Trial (S-RCT) and derive a two-layer error decomposition that separates agent approximation error from subsampling error, enabling targeted improvements to each. The framework is agent-agnostic: any behavioral model---from a fine-tuned specialist to a general-purpose foundation model---can serve as the simulation engine. Validated on 67 historical marketing A/B tests, a baseline S-RCT using an off-the-shelf foundation model captures directional signal (sign overlap 0.70) but systematically overshoots effect magnitudes. A two-phase pre-period calibration protocol reduces the squared prediction error (after removing irreducible measurement noise) by {sim}77imes; a within-subject design---where each agent is exposed to both arms---reduces standard errors by {sim}2.4imes. We discuss limitations of the current approach and identify applications where experimenters stand to benefit from agentic signals.

Related

Source: arXiv cs.CL | 2026-08-04

Loading related sources…