Can You Trust Your AI Research Assistant? Evaluating LLMs and AI Agents – January 2027
Event Phone: 1-610-715-0115
Upcoming Dates
-
27JanCan You Trust Your AI Research Assistant? Evaluating LLMs and AI Agents10:00 AM-3:30 PM
Cancellation Policy: If you cancel your registration two weeks or more before the course is scheduled to begin, you are entitled to receive your choice of either a credit for a future seminar (which can be applied toward any of our courses) or a refund of the registration fee (minus a processing fee of $50).
In the unlikely event that Statistical Horizons LLC must cancel a seminar, we will do our best to inform you as soon as possible of the cancellation. You would then have the option of receiving a full refund of the seminar fee or a credit towards another seminar. In no event shall Statistical Horizons LLC be liable for any incidental or consequential damages that you may incur because of the cancellation.
A 3-Day Livestream Seminar Taught by Anupam Singh, Ph.D.
LLMs and AI agents are now doing real research work: coding survey responses and interview transcripts, screening abstracts for systematic reviews, extracting evidence from policy documents and clinical guidelines, and chaining these steps together with little supervision. The outputs are fast, fluent, and often impressive. They are also, in ways that are easy to miss, unstable. Run the task again and the codes shift; reword the prompt and the results move; switch models and the story changes. Once those outputs become a variable in your analysis or a finding in your review, the question a reviewer will ask is not whether you used AI, but whether you can show it was reliable enough to trust.
This seminar teaches you how to answer that question. Most AI courses teach researchers how to use LLMs and agents; this one teaches you how to evaluate them, in R, with your own research materials. You will benchmark an LLM against human judgement, test whether its results survive repeated runs, alternative prompts, and different models, audit an LLM that scores or ranks your material, inspect what an AI agent actually did rather than what it reported, and document all of it in a record that reviewers, editors, and ethics boards will accept. You will leave with reusable R notebooks and templates, and an evaluation plan for a project of your own.
The seminar follows one evaluation workflow, from research task to reproducibility record, applied to three running examples. First, an LLM codes open-ended text and is validated against expert human coders. Second, that pipeline is stress-tested, and an LLM is then used as a judge applying a scoring rubric, with planted fabrications it must catch. Third, a transparent R research agent assembles cited evidence from a document corpus and is deliberately sabotaged to see whether it notices.
Along the way you will learn the measurement ideas that make an evaluation defensible: why agreement between human coders sets the ceiling for any model, which agreement statistics fit which tasks, how to tell systematic failure from random noise, and why an agent’s process matters as much as its answer. The emphasis is on repeatable evaluation experiments rather than informal spot-checks. This is not a prompt-engineering course, and it does not teach you to build agent frameworks; it teaches you to audit what such tools do and to decide, on the evidence, whether their output can count as research evidence.
Everything is done in R with the ellmer package and familiar tidyverse tools, so each method transfers directly to your own data. You will leave with reusable notebooks, benchmark and audit templates, an error taxonomy, a human-review protocol, and a reporting template. The final session is a guided clinic, so bring a dataset, a corpus, or a research question you are considering using an LLM for.
Venue: Livestream Seminar