Recent work

A sample of the questions I’ve studied. Each uses published research, public data or my own testing, and each says what it can and can’t show.

Can AI write accurate law school practice problems?

Two current AI models wrote 100 practice problems on hearsay, each with an analysis and full case citations. Every citation was checked against the published opinion or rule, and every analysis against current federal law. All 272 citations named real cases with the correct name, volume, reporter, court and year. The errors were in what the law was said to hold: 24 of the 100 analyses had a problem, and some cases were cited for points they don’t make. Those errors look right on the page and can mostly be caught only by reading the source. The ratings came from AI checkers working from the sources they fetched; no person reviewed them.

Does using AI weaken students’ thinking?

A review of the research, with each claim checked against its source. Using AI to get answers lowers performance once the AI is taken away. Access to AI is not harmful in itself: across 14 randomized studies that tested students without AI afterward, the average effect was a small gain. People do follow wrong AI answers and misjudge how much they have learned. Several widely cited studies turned out weaker than reported, and one has been retracted.

Does grading toward the skills practice values make grades less reliable?

A statistical model, built on Kane and Case’s 2004 framework, of what happens when part of a course grade shifts to skills that are harder to measure. The answer turns on how well the new parts are measured and how much testing time they get. It also shows that a one-section pilot of about 80 students does little to settle the question, while about 240 students does much more.

Every school is starting from a different place. Tell me where yours is.

Start a conversation