What does the evidence already say?
A living review of academic and working papers on AI-moderated interviews — what holds up, what's contested, and the questions no one has answered yet. It's the map the rest of this roadmap is drawn against.
Observation Log — a living map
Every serious research team, sooner or later, asks the same handful of things about AI-moderated work: does it hold up, what do people actually say to it, how much sample is enough, and how much does the moderator itself shape the answer.
This is where each of those questions stands in our work — some already charted, some still rising. It's grounded in the field's open questions, which we track in the literature review.
A living review of academic and working papers on AI-moderated interviews — what holds up, what's contested, and the questions no one has answered yet. It's the map the rest of this roadmap is drawn against.
A study of disclosure — the topics, candor, and detail people bring to a moderator that isn't a person. Where the absence of a human in the room helps, and where it doesn't.
Where qualitative signal saturates as you add participants — and where extra sample stops buying insight. The old benchmarks put "good" qual near 4–10 per segment and academic-grade quant at 300–800; precision only scales with the square root of n, so halving your uncertainty roughly quadruples the cost. Episode 1 tests what actually holds for AI-moderated qual.
Re-analyzing a fielded dataset to ask whether a stronger or weaker moderator meaningfully moves the quality of what respondents give back. Early signal: less than we'd expect. If that holds, it reframes where the leverage in an interview really sits.
Same script, three moderators — George, Mia, Riley — differing in perceived voice and gender. Whether who appears to be asking moves what people are willing to say. Baked into a current product study so the comparison rides along with real fieldwork.
A replication study: take findings from published, human-led interviews and re-run them with an AI moderator, then compare what comes back. Replication is how a young method earns trust — and it's the gap the literature flags most.
Run the same questions through more than one AI-moderated platform and measure how much the tool, rather than the respondent, shapes the result. A reliability check the whole category needs and few are willing to run.