Mercor
Mercor supplies frontier labs with the expert-generated data their models post-train on. I ran quality on the coding side, where the product is LLM-generated agent trajectories and the question is whether the reasoning inside them is any good.
I spearheaded quality optimization across those trajectories, evaluating agent reasoning and decision-making patterns over 1,000+ open-source GitHub repositories to improve what the models learn downstream.
I built and ran the end-to-end evaluation pipeline behind that work: LLM code generation, expert human annotation, multi-tier review, and final delivery to leading research labs, with quality benchmarks held consistent across task cycles.
Running it meant directing a distributed team of 50 to 100 domain experts through Airtable workflow orchestration. On-time sprint delivery stayed above 95%, review turnaround fell 35%, and I identified top performers for promotion into senior reviewer roles.