Safety
Research Design Tracking and Assessment for the Social Sciences
arXiv:2608.27049v1 Announce Type: new Abstract: Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on ma
arXiv:2608.27049v1 Announce Type: new Abstract: Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis. We introduce Automated Research Design Tracking and Assessment (ARDTrA), a task that involves detecting the research design used in a paper and assessing the quality of its application. We create an expert-annotated dataset of papers covering six families of counterfactual research designs and evaluate the task using a multi-turn RAG-based conversational pipeline. Across four retrieval strategies, four LLMs and six embedding models, we find that passage length is the main driver of performance, explaining 52-66% of the variance. A per-research-design analysis also shows that human and machine difficulty do not align: the designs that prove hardest for the system are not those on which expert annotators disagree most, pointing to two independent sources of task difficulty.
Related
- From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
- Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction
Source: arXiv cs.CL | 2026-08-28