← Back to Projects

Human-Verified Dataset Builder

A lightweight human-in-the-loop pipeline for turning raw documents into clean, labeled datasets: an LLM proposes a grounded label for each source chunk, demonstrated as Q&A extraction against UK annual report excerpts, and a human reviews it through a local Streamlit UI (accept, edit, reject, or skip), with every decision logged alongside full provenance: what was proposed, what was confirmed, the edit distance between them, and how long review took. Multiple reviewers can work the same items independently, with agreement between them measured via decision match rate and Cohen’s kappa, and a paginated, filterable browse mode makes it possible to search past decisions at scale rather than only stepping through the confidence-sorted queue.

Corrections don’t just get exported, they get reused: accepted edits are fed back into the generation prompt as few-shot examples on every subsequent batch, and the effect is measured rather than assumed: one real correction, tested against an unrelated topic it hadn’t seen before, taught the model to catch the same class of mistake unprompted. A second layer of heuristics flags reviews that may deserve a second look, a decision made suspiciously fast, or a human overriding a high-confidence proposal, treating the reviewer, not just the model, as a source of error worth checking.

Python, Streamlit, Anthropic API (Claude), SQLite, uv, pandas