Turn Documents into Data with AI – September 2026

Event Phone: 1-610-715-0115

Details Price Qty
Regular Admissionshow details + $995.00 USD  ea 

Upcoming Dates

  • 10
    Sep
    Turn Documents into Data with AI
    10:00 AM
    -
    3:30 PM
Cancellation Policy: If you cancel your registration two weeks or more before the course is scheduled to begin, you are entitled to receive your choice of either a credit for a future seminar (which can be applied toward any of our courses) or a refund of the registration fee (minus a processing fee of $50). 
In the unlikely event that Statistical Horizons LLC must cancel a seminar, we will do our best to inform you as soon as possible of the cancellation. You would then have the option of receiving a full refund of the seminar fee or a credit towards another seminar. In no event shall Statistical Horizons LLC be liable for any incidental or consequential damages that you may incur because of the cancellation.
A 3-Day Livestream Seminar Taught by Karl Rohe, Ph.D.

Whether you call it data extraction, document coding, content analysis, structured abstracting or systematic review, the underlying work is the same: turning a body of literature into structured, analyzable, trustworthy data.

For a century, statistical methodology has systematized the final steps of inference—modeling and testing—but not the miles of reading and extraction that come before. That work has been done by humans, one paper at a time, with criteria written in a codebook and no measure of external validity. AI can change this, but not how you might think.

The stance of this course is that AI is a poor judge, but an excellent reader. Most researchers don’t trust AI for their research—and they shouldn’t, at least not blindly. But you can build your trust in AI the same way you can build it with a research assistant: give it a clear task, audit the first outputs, see where it stumbles, refine the instructions, and avoid assigning tasks it can’t do. Then, iterate. Instead of asking “Is the assistant trustworthy?” ask “Which tasks can I entrust to it? What should I ask so it understands? How do I know it did the job well?’

The craft lies in decomposing research judgment into reading tasks that AI can do reliably—and the codebook is the instrument for that decomposition. This course teaches you how to create codebooks with and for AI.

Each session combines lecture and live demonstration: building codebooks, running extractions, diagnosing failures, and refining outputs. At times, you will watch the work unfold on the instructor’s screen; at others, you will make the same moves on your own, with supervision and feedback.

Along the way you will learn the transferable craft: how to spot where the AI is guessing rather than reading, how to refine a field to remove codebook ambiguity, and how to know when the answers are trustworthy enough to publish.

The course is built around four commitments that, together, provide an audit trail for trustworthy use of AI:

  • The codebook is the durable artifact. It is the document that holds your operationalization—every choice you made about what counts and what doesn’t, written down so that reviewers (and your future self) can see them. A good codebook outlasts the model that ran it.
  • AI is a reader, not a judge. Decomposing a research question into reading tasks an AI can do reliably is the core engineering move. In this course, you’ll learn decomposition as a craft, with concrete refinement techniques and a vocabulary for diagnosing failures.
  • Disagreement is the diagnostic. When multiple AI readers run the same codebook and disagree, those are the cells worth auditing first. Multi-reader agreement is treated as publication-defensible evidence of reliability, but not a stamp of approval.
  • Data are shared with codebook and reasoning. Your spreadsheet gets a URL that you can share. You can click any cell to see the codebook item, key quotes from your document, and the AI’s reasoning for the value in the cell.

The codebook can be reused on different documents by you or others and revised or refined as needed. The disagreements reveal where the AI readers interpret the codebook differently, or places where they make errors.

Venue: