Demos · Hands-on

▁▂▃▅▃▂▁

Orchestration Lab

Run a real analysis through four ways of using frontier models, then score your run against a pre-built answer key. Most cells cost $1–5 and finish in minutes.

Six real social-science briefs, four ways to run a frontier model on them. The four arms are skills I wrote for my own work — three for Claude Code, one for the Codex CLI, all in the Open Science Skills toolkit. Every brief has a reference solution and a rubric written before any model ran. Pick a brief, pick an arm, paste one command, and grade what comes back against the answer key and our committed runs.

One draw. A captured run is one sample from a non-deterministic process, not a benchmark. Our numbers below are specimens. Re-run the briefs and expect yours to differ.

Step 1Set up

You need Claude Code with the Open Science Skills plugin, R with the data packages, and the repo itself. That assumes Node and git on your machine, plus a paid claude.ai login. The Codex arm additionally needs the Codex CLI with a ChatGPT login.

npm install -g @anthropic-ai/claude-code
claude plugin marketplace add scdenney/open-science-skills
claude plugin install oss@open-science-skills
Rscript -e 'install.packages(c("projoint","ivdoctr","AER","car","causaldata","MatchIt","did","fixest","staggered","bacondecomp"))'
git clone https://github.com/scdenney/ai-for-research.git
cd ai-for-research/demos/orchestration-lab

# Codex arm only — the Codex CLI plus the toolkit's Codex-side skills:
npm install -g @openai/codex
git clone https://github.com/scdenney/open-science-skills.git ~/open-science-skills
cd ~/open-science-skills
python3 plugin/scripts/install-codex.py --all --dry-run
python3 plugin/scripts/install-codex.py --all

Every brief uses public data that ships with an R package — nothing to download. The runs call hosted models and cost real money: in our captures, $0.98 to $4.99 per cell on the Claude arms. Keep ANTHROPIC_API_KEY unset so headless runs bill your claude.ai plan, not an API account.

Step 2Pick a brief

The six briefs sit on a four-rung difficulty ladder, and each rung is defined by what kind of problem it is, not just how much statistics it takes to solve it:

RungWhat it testsWhy it's on the ladder
easy–moderateMechanical due diligence → standard estimation → an original judgment callThree briefs, same data, escalating from "get the facts right" to "make a defensible call nobody hands you the answer to."
hard–very hardCanonical, decades-old disputes with literature-settled verdictsHarder statistics (IV diagnostics, matching estimators), but every capable pipeline has seen a worked version. All four arms hit the top band on both.
extremeA 2020s-vintage methods debate with a genuine, checkable trapBuilt to test whether the high/very-high convergence was about the arms or about the tasks being over-rehearsed. It wasn't differentiating either way — see the results below.

Each brief links to the exact prompt the models get.

Step 3Pick an arm and run it

Choose an arm and a brief. The command below updates; it writes into a fresh myruns/ leaf so the committed runs stay untouched. Full protocol, including the excerpt and hashing rules we used for the captured runs: run.md.

Arm Brief

What you'll get. The run writes its deliverables into your myruns/ leaf — typically script.R, a figure or two, and the brief's memo or summary — plus, on the headless Claude arms, claude-envelope.json, whose total_cost_usd and duration_ms fields are your cost and wall-clock. Grade the deliverables against the checklist below, then compare your cell against the committed run of the same cell under runs/.

Step 4Score your run

Every brief is graded on six yes-or-no items: four core facts a competent run must get right, one judgment call the answer key hands to the analyst, and one completeness credit for work beyond a correct answer. Pass means all four core items. Pass+ adds the judgment. Distinction means all six. Miss any core item and the run fails, whatever else it does well. Definitions and every captured run's item-by-item score: SCORING.md; executed reference solutions: reference/.

Describe — checklist
  • coreStates 400 respondents, 8 tasks each, 2 profiles per task, 6,400 profile rows.
  • coreNames all 7 attributes in readable form, not att1..att7.
  • corePer-attribute level counts read 3, 3, 4, 2, 4, 6, 2.
  • coreIdentifies the repeated, flipped task 1 and does not count it as a 9th task.
  • judgmentFlags Total Daily Driving Time as the lone imbalance; makes no "perfect balance" claim.
  • completenessReports the exact max deviation from uniform (~1.94 pp), not only a min-max spread.
Estimate — checklist
  • coreViolent Crime Rate is the largest |AMCE|, magnitude right for its stated scale (~25.1 pp corrected / ~16.5 uncorrected, never mislabeled).
  • coreEvery large effect signs correctly; attribute ordering matches the key.
  • coreStandard errors clustered on respondent id.
  • coreEvery attribute's reference level fixed at zero; estimates presented as AMCEs.
  • judgmentStates corrected or uncorrected as a deliberate, labeled choice.
  • completenessNames the profile-level estimand and the IRR mechanism (tau ~0.17, ×1.52), explained not asserted.
Reviewer reply — checklist
  • coreConcedes multi-level AMCEs shift under relabeling, backed by a number from these data.
  • coreStates crime is binary, so a baseline flip only flips the sign — |AMCE| invariant.
  • coreComputes marginal means (.626 / .374) and uses the baseline-free MM range as the ranking currency.
  • coreCaps the claim: crime is the largest single driver, commute-comparable — never "dominates."
  • judgmentFlags the ~1.4 pp crime-vs-commute gap as within noise — a statistical tie.
  • completenessReports both magnitudes, uncorrected 16.5 and corrected 25.1 pp.
IV replication — checklist
  • coreReproduces the headline: 2SLS 0.944, OLS 0.522, first-stage F 22.95.
  • coreRuns the stress specs: controls keep F > 10; dropping the neo-Europes pushes F to 8.65; Africa-only collapses to F = 0.30.
  • coreFlags the weak-instrument specs as weak.
  • coreNo overclaim about what the replication proves.
  • judgmentStates the two-sided ceiling: a collapsed first stage neither confirms nor overturns the original result.
  • completenessOne unified table across all five specifications.
Methods dispute — checklist
  • coreAnchors exact: experimental benchmark +$1,794; naive CPS comparison −$8,498.
  • coreRuns the specification curve including pre-earnings matching (demographics-only must fail; pre-earnings 1-NN lands near the benchmark).
  • coreBenchmark-referenced results table.
  • coreNo overclaim in the adjudication.
  • judgmentLands the "helps but does not settle" verdict — favorable-specification recovery, not universal recovery.
  • completenessBenchmark-referenced figure.
Staggered-DiD reconciliation — checklist
  • coreAt least four ATT estimates, all clustering at −0.037 to −0.047 log points.
  • coreCorrectly codes (or explicitly notes) that the staggered package needs never-treated units as Inf, not the 0 the raw data and did use — the trap.
  • coreRuns a Goodman-Bacon decomposition and reports the correct ~5.4% negative-weight-risk share, with the right inference (small, not zero).
  • coreAn event-study / pre-trend check covering at least two pre-periods.
  • judgmentLands the calibrated middle: the TWFE critique is real in general and empirically small in this dataset — not "always fine" and not "always use the modern estimator."
  • completenessFull comparison table across every estimator, plus the event-study time path (effect grows after impact).

What our runs showed

Thirty captured cells (six briefs × five arms), graded blind against the rubrics, all in one sitting under one consistent settings pass. Two findings matter more than any single band. First: on the easy brief, both orchestrated Claude leads (Fable, Opus) landed at Pass rather than the higher band their historical captures had reached, because neither write-up explains the conjoint design's repeated reliability check — even though both arms' analysis code still handles it correctly (6,400 rows, not an inflated 7,200). The miss is in the diagnostic write-up, not the computation or the core design facts, and it is the clearest demonstration yet of this project's standing caveat: a captured run is one draw, not a benchmark. Second: the new Extreme brief, built specifically to test whether frontier models would separate from competent ones on harder, less-rehearsed work, did not show that separation across these five arms — every one reached the top band, including the historical weak point (the Codex headless fallback).

A sixth arm, not part of the interactive picker below, tells a different story on that same Extreme brief. An all-Codex advisor (a plain gpt-5.6-terra solve, one read-only gpt-5.6-sol consult, a gpt-5.6-terra revise) failed it — and the reason is worth stating plainly rather than folding into a single Pass/Fail count. The pre-revision draft correctly handled the brief's central trap, then the consult raised a separate, defensible methodological point that led the model to remove the exact code demonstrating that awareness during revision. A substantively reasonable second opinion cost the run the one thing its rubric was built to check for. Full detail is in the report; the run-by-run matrix, including this arm, is in RESULTS.md.

Four small panels, one per arm, each plotting rubric items met across the six briefs against dotted threshold lines labeled Pass, Pass plus, and Distinction.
Rubric items met across the six briefs, one panel per arm, read against the same three band thresholds. Our captures: one draw each.
SetupWhen to reach for it
Fable leadWhen the task is too big to hold in one context — a broad migration, a multi-file audit, a many-source survey. That regime is untested here. Cheapest and fastest Claude orchestrator on these briefs, but the easy brief landed at Pass rather than a higher band, a reminder that speed doesn't buy the extra diagnostic thoroughness a second read supplies.
Opus leadRarely, on work this size. Ties or beats the Fable lead on the harder rungs at higher cost; its edge is write-up thoroughness, not correctness — and it hit the same easy-brief miss Fable did. Reserve it for a task hard enough to strain a top model.
Codex leadA cross-vendor read on the same brief. The historical Sol-lead capture reaches Distinction on five of six briefs at xhigh effort with a genuine Fable-5 cross-vendor peer, missing only completeness on the standard-estimation brief. The current skill accepts either an active Astra or Sol lead: Astra keeps compact hard reasoning in-session, while Sol escalates unusually difficult units to Astra. Reach for it when you want the second opinion to come from a different vendor entirely.
AdvisorThe default for a single, well-scoped analysis. Its second read re-computes the work rather than eyeballing it, and it's the only arm that caught a real analytical error (a misread comparison-type label) on the extreme brief before it shipped. Reach for it first.

Our captures are one draw each. Re-run a cell and expect yours to differ — that's not a caveat, it's the point.