2 · Statistical thinking
Two agents can agree on a correlation without ever checking whether it holds. Have one agent analyse, a second agent attack, and then decide for yourself what the data actually supports.
At a glance
Section titled “At a glance”| Time | ~80 minutes, in pairs (Part 1: 25 · Part 2: 25 · Part 3: 30) |
| Work with | A partner |
| You need | Two agent sessions — Claude or Manus, the same platform twice or a fresh session for the second run |
| You hand in | Three claim critiques, one paragraph each (the Statistical Claim Audit) |
| Graded on | Accuracy, scepticism, clarity, judgement — not whether the verdict is definitive |
The scenario
Section titled “The scenario”Three claims are circulating from Kopi Kembali’s last marketing review, and a budget decision hangs on each. The evidence file is kk-halfyear-performance.csv: six months (January to June 2026) of mixed observations, distinguished by an observation_type column:
social_postrows — one row per post across four platforms (platform, day, time slot, impressions, engagements)site_sessionsrows — daily cohorts for three devices (sessions, average browse minutes, conversion rate)blog_dailyrows — daily reads and engagement for two blog sections, with apaywall_activeflag
Data dictionary (columns populate differently depending on observation_type)
| Column | Meaning |
|---|---|
date, day_of_week, time_slot |
The grain — time_slot only applies to social posts |
observation_type |
social_post, site_sessions, or blog_daily |
channel |
Platform (social posts) or blog section (blog rows) |
device |
mobile / desktop / tablet — site-session rows only |
impressions, engagements |
Social post metrics |
sessions_or_reads |
Session count (site rows) or read count (blog rows) |
avg_browse_min, conversion_rate_pct |
Site-session metrics |
paywall_active |
Whether the blog section’s paywall was active that day |
The claims
- Claim A. “Social media engagement is highest on Friday afternoons.” Before you schedule everything for Friday: does it hold within each platform, or only in the pooled average? Who posts on Friday afternoons? How many posts is the claim standing on?
- Claim B. “Customers who browse longer make fewer purchases, so we should streamline the site.” The pooled correlation is real. Is it causal? Split by device before you answer, and remember: could something else cause both?
- Claim C. “Removing the paywall from Blog-Recipes on 1 June lifted reads 18% and engagement 12%.” Check the arithmetic against the file. Then check the section whose paywall never changed. Then count the weeks on each side of the comparison.
Part 1 — First agent analysis · 25 min
Section titled “Part 1 — First agent analysis · 25 min”Analyse this claim: [CLAIM].
Use the attached dataset to investigate.Report the following:1. What summary statistics or patterns in the data back up orcontradict the claim?2. What is the sample size? Who is in it? Are any groups missing?3. Could a confounding variable explain the pattern instead of whatthe claim says?4. On a scale of "definitely true," "probably true," "unclear,""probably false," and "definitely false," where does this claim land?
Be specific. Show the numbers. Do not assume the claim is true; testit.Read the report carefully. Note the key numbers, any warnings raised, and whether the agent made a causation claim (“this caused that”) or only described a pattern (“this moved with that”). Save or screenshot the response — you will hand it to the second agent next.
Part 2 — The challenge · 25 min
Section titled “Part 2 — The challenge · 25 min”I have just read this analysis of a claim:
[PASTE THE FIRST AGENT'S REPORT HERE]
Your task is to argue against this conclusion. Be sceptical. Do notjust agree.
Answer:1. What are three ways this analysis could be wrong?2. What confounding variables or sampling problems might the firstagent have missed?3. What data would you need to see to be more confident in theconclusion?4. What does this analysis not tell us?
Be specific. Give examples.Note where the second agent disagreed with the first, and which confounders or sampling issues it named. Compare the two reports — which sounds more careful, and why?
Part 3 — Your verdict · 30 min
Section titled “Part 3 — Your verdict · 30 min”Write one paragraph per claim, 150–200 words, covering:
- What the data shows — one or two key numbers or patterns.
- What the data does not show — confounders you can name, sample limitations, causation overreach.
- Your verdict — supported, partly supported, unclear, or not supported. Pick one.
- A note on the agents — did the first agent get it right? What did the second catch that the first missed?
Use the three questions from class: could something else cause both? Is the sample big enough and unbiased? What did the analysis actually prove, versus claim? Consider whether the pattern holds in subgroups (Simpson’s Paradox) or might be regression to the mean, and sanity-check the reported numbers against what’s plausible.
Be specific about numbers — “the sample was 47 people,” not “the sample was small.” Name confounders concretely — “high-spending customers might be repeat buyers who browse less because they know the site,” not “there could be another explanation.” “Unclear” is a valid verdict if the evidence doesn’t give you enough to decide.
Verify & submit
Section titled “Verify & submit”- Every verdict cites at least one number with its group size
- Claim A is answered within platforms, not only pooled
- Claim B is split by device before any causal language is used
- Claim C is compared against the unchanged section and the claimed percentages are checked against the file
- Each critique names at least one confounder or sampling issue, and notes whether the two agents agreed or disagreed
- At least one critique checks for Simpson’s Paradox or regression to the mean, and at least one includes a sanity check on the numbers
- No paragraph says “proves” where the honest word is “is consistent with”
This is the Statistical Claim Audit — evidence section 4 of your Individual Analytics Portfolio, separate from Knowledge Check 9. Add the three critiques to your individual Google Doc and submit the Unit 8 xSiTe checkpoint.
Capstone connection
Section titled “Capstone connection”Apply this same thinking to Stage 3 (explore and analyse) of your capstone project. When the agent reports a correlation, ask: could something else explain it? Is the sample representative? Document this in your orchestration log under Stage 3.