Six real social-science briefs, four ways to run a frontier model on them. The four arms are skills I wrote for my own work — three for Claude Code, one for the Codex CLI, all in the Open Science Skills toolkit. Every brief has a reference solution and a rubric written before any model ran. Pick a brief, pick an arm, paste one command, and grade what comes back against the answer key and our committed runs.
Read the report
Four ways to run a frontier model — why the leaderboards don't answer this, the full findings, what each setup caught and missed, and where it broke. On Pixels & Patterns.
Step 1Set up
You need Claude Code with the Open Science Skills plugin, R with the data packages, and the repo itself. That assumes Node and git on your machine, plus a paid claude.ai login. The Codex arm additionally needs the Codex CLI with a ChatGPT login.
npm install -g @anthropic-ai/claude-code
claude plugin marketplace add scdenney/open-science-skills
claude plugin install oss@open-science-skills
Rscript -e 'install.packages(c("projoint","ivdoctr","AER","car","causaldata","MatchIt","did","fixest","staggered","bacondecomp"))'
git clone https://github.com/scdenney/ai-for-research.git
cd ai-for-research/demos/orchestration-lab
# Codex arm only — the Codex CLI plus the toolkit's Codex-side skills:
npm install -g @openai/codex
git clone https://github.com/scdenney/open-science-skills.git ~/open-science-skills
cd ~/open-science-skills
python3 plugin/scripts/install-codex.py --all --dry-run
python3 plugin/scripts/install-codex.py --all
Every brief uses public data that ships with an R package — nothing to download. The runs call hosted models and cost real money: in our captures, $0.98 to $4.99 per cell on the Claude arms. Keep ANTHROPIC_API_KEY unset so headless runs bill your claude.ai plan, not an API account.
Step 2Pick a brief
The six briefs sit on a four-rung difficulty ladder, and each rung is defined by what kind of problem it is, not just how much statistics it takes to solve it:
| Rung | What it tests | Why it's on the ladder |
|---|---|---|
| easy–moderate | Mechanical due diligence → standard estimation → an original judgment call | Three briefs, same data, escalating from "get the facts right" to "make a defensible call nobody hands you the answer to." |
| hard–very hard | Canonical, decades-old disputes with literature-settled verdicts | Harder statistics (IV diagnostics, matching estimators), but every capable pipeline has seen a worked version. All four arms hit the top band on both. |
| extreme | A 2020s-vintage methods debate with a genuine, checkable trap | Built to test whether the high/very-high convergence was about the arms or about the tasks being over-rehearsed. It wasn't differentiating either way — see the results below. |
Each brief links to the exact prompt the models get.
- easyDescribe — document a community-choice conjoint design (projoint's
exampleData1) and check randomization balance. - standardEstimate — fit the AMCEs, cluster correctly, plot them with reference levels at zero.
- moderateReviewer reply — answer a referee who calls the headline crime finding an artifact of the AMCE baseline.
- hardIV replication — reproduce Acemoglu-Johnson-Robinson's colonial-origins 2SLS result (via ivdoctr) and stress it until the instrument breaks.
- very hardMethods dispute — adjudicate Dehejia-Wahba vs Smith-Todd on whether matching recovers the LaLonde experimental benchmark (via causaldata).
- extremeStaggered-DiD reconciliation — reconcile four modern difference-in-differences estimators on a real staggered minimum-wage rollout (Callaway-Sant'Anna's own
mpdta), and catch a genuine package-convention trap along the way.
Step 3Pick an arm and run it
Choose an arm and a brief. The command below updates; it writes into a fresh myruns/ leaf so the committed runs stay untouched. Full protocol, including the excerpt and hashing rules we used for the captured runs: run.md.
What you'll get. The run writes its deliverables into your myruns/ leaf — typically script.R, a figure or two, and the brief's memo or summary — plus, on the headless Claude arms, claude-envelope.json, whose total_cost_usd and duration_ms fields are your cost and wall-clock. Grade the deliverables against the checklist below, then compare your cell against the committed run of the same cell under runs/.
Step 4Score your run
Every brief is graded on six yes-or-no items: four core facts a competent run must get right, one judgment call the answer key hands to the analyst, and one completeness credit for work beyond a correct answer. Pass means all four core items. Pass+ adds the judgment. Distinction means all six. Miss any core item and the run fails, whatever else it does well. Definitions and every captured run's item-by-item score: SCORING.md; executed reference solutions: reference/.
Describe — checklist
- coreStates 400 respondents, 8 tasks each, 2 profiles per task, 6,400 profile rows.
- coreNames all 7 attributes in readable form, not
att1..att7. - corePer-attribute level counts read 3, 3, 4, 2, 4, 6, 2.
- coreIdentifies the repeated, flipped task 1 and does not count it as a 9th task.
- judgmentFlags Total Daily Driving Time as the lone imbalance; makes no "perfect balance" claim.
- completenessReports the exact max deviation from uniform (~1.94 pp), not only a min-max spread.
Estimate — checklist
- coreViolent Crime Rate is the largest |AMCE|, magnitude right for its stated scale (~25.1 pp corrected / ~16.5 uncorrected, never mislabeled).
- coreEvery large effect signs correctly; attribute ordering matches the key.
- coreStandard errors clustered on respondent id.
- coreEvery attribute's reference level fixed at zero; estimates presented as AMCEs.
- judgmentStates corrected or uncorrected as a deliberate, labeled choice.
- completenessNames the profile-level estimand and the IRR mechanism (tau ~0.17, ×1.52), explained not asserted.
Reviewer reply — checklist
- coreConcedes multi-level AMCEs shift under relabeling, backed by a number from these data.
- coreStates crime is binary, so a baseline flip only flips the sign — |AMCE| invariant.
- coreComputes marginal means (.626 / .374) and uses the baseline-free MM range as the ranking currency.
- coreCaps the claim: crime is the largest single driver, commute-comparable — never "dominates."
- judgmentFlags the ~1.4 pp crime-vs-commute gap as within noise — a statistical tie.
- completenessReports both magnitudes, uncorrected 16.5 and corrected 25.1 pp.
IV replication — checklist
- coreReproduces the headline: 2SLS 0.944, OLS 0.522, first-stage F 22.95.
- coreRuns the stress specs: controls keep F > 10; dropping the neo-Europes pushes F to 8.65; Africa-only collapses to F = 0.30.
- coreFlags the weak-instrument specs as weak.
- coreNo overclaim about what the replication proves.
- judgmentStates the two-sided ceiling: a collapsed first stage neither confirms nor overturns the original result.
- completenessOne unified table across all five specifications.
Methods dispute — checklist
- coreAnchors exact: experimental benchmark +$1,794; naive CPS comparison −$8,498.
- coreRuns the specification curve including pre-earnings matching (demographics-only must fail; pre-earnings 1-NN lands near the benchmark).
- coreBenchmark-referenced results table.
- coreNo overclaim in the adjudication.
- judgmentLands the "helps but does not settle" verdict — favorable-specification recovery, not universal recovery.
- completenessBenchmark-referenced figure.
Staggered-DiD reconciliation — checklist
- coreAt least four ATT estimates, all clustering at −0.037 to −0.047 log points.
- coreCorrectly codes (or explicitly notes) that the
staggeredpackage needs never-treated units asInf, not the0the raw data anddiduse — the trap. - coreRuns a Goodman-Bacon decomposition and reports the correct ~5.4% negative-weight-risk share, with the right inference (small, not zero).
- coreAn event-study / pre-trend check covering at least two pre-periods.
- judgmentLands the calibrated middle: the TWFE critique is real in general and empirically small in this dataset — not "always fine" and not "always use the modern estimator."
- completenessFull comparison table across every estimator, plus the event-study time path (effect grows after impact).
What our runs showed
Thirty captured cells (six briefs × five arms), graded blind against the rubrics, all in one sitting under one consistent settings pass. Two findings matter more than any single band. First: on the easy brief, both orchestrated Claude leads (Fable, Opus) landed at Pass rather than the higher band their historical captures had reached, because neither write-up explains the conjoint design's repeated reliability check — even though both arms' analysis code still handles it correctly (6,400 rows, not an inflated 7,200). The miss is in the diagnostic write-up, not the computation or the core design facts, and it is the clearest demonstration yet of this project's standing caveat: a captured run is one draw, not a benchmark. Second: the new Extreme brief, built specifically to test whether frontier models would separate from competent ones on harder, less-rehearsed work, did not show that separation across these five arms — every one reached the top band, including the historical weak point (the Codex headless fallback).
A sixth arm, not part of the interactive picker below, tells a different story on that same Extreme brief. An all-Codex advisor (a plain gpt-5.6-terra solve, one read-only gpt-5.6-sol consult, a gpt-5.6-terra revise) failed it — and the reason is worth stating plainly rather than folding into a single Pass/Fail count. The pre-revision draft correctly handled the brief's central trap, then the consult raised a separate, defensible methodological point that led the model to remove the exact code demonstrating that awareness during revision. A substantively reasonable second opinion cost the run the one thing its rubric was built to check for. Full detail is in the report; the run-by-run matrix, including this arm, is in RESULTS.md.

| Setup | When to reach for it |
|---|---|
| Fable lead | When the task is too big to hold in one context — a broad migration, a multi-file audit, a many-source survey. That regime is untested here. Cheapest and fastest Claude orchestrator on these briefs, but the easy brief landed at Pass rather than a higher band, a reminder that speed doesn't buy the extra diagnostic thoroughness a second read supplies. |
| Opus lead | Rarely, on work this size. Ties or beats the Fable lead on the harder rungs at higher cost; its edge is write-up thoroughness, not correctness — and it hit the same easy-brief miss Fable did. Reserve it for a task hard enough to strain a top model. |
| Codex lead | A cross-vendor read on the same brief. The historical Sol-lead capture reaches Distinction on five of six briefs at xhigh effort with a genuine Fable-5 cross-vendor peer, missing only completeness on the standard-estimation brief. The current skill accepts either an active Astra or Sol lead: Astra keeps compact hard reasoning in-session, while Sol escalates unusually difficult units to Astra. Reach for it when you want the second opinion to come from a different vendor entirely. |
| Advisor | The default for a single, well-scoped analysis. Its second read re-computes the work rather than eyeballing it, and it's the only arm that caught a real analytical error (a misread comparison-type label) on the extreme brief before it shipped. Reach for it first. |
Our captures are one draw each. Re-run a cell and expect yours to differ — that's not a caveat, it's the point.