Snorkel AI — Frontier Model Evals
Software Engineer, AI Evaluation (Contract) — building and reviewing SWE evals with deterministic grading and anti-cheat validation
- 50+ evals built
- 200+ evals reviewed
- Reward-hacking resistant
Read the case study
Impact
- Designed and delivered 50+ software-engineering evaluations for frontier AI coding models, each with a deterministic reference solution and automated grading.
- Reviewed 200+ evaluations for correctness, test coverage, reproducibility, and resistance to reward hacking.
- Strengthened model-assessment signal by applying AI-assisted code review and systematic failure analysis.
Problem
Evaluating whether a frontier model can actually ship working code requires tasks that are unambiguous, reproducible, and impossible to game. Weak specs, flaky tests, or exploitable graders produce misleading scores that overstate model capability.
Approach
- Authored deterministic reference solutions and pytest grading harnesses for each task.
- Packaged every evaluation in a reproducible container so results are environment-independent.
- Designed anti-cheat validation to catch reward hacking and shortcut solutions.
- Reviewed peers' evaluations for verifier strength, instruction/test alignment, and reproducibility.
- Used AI-assisted code review and failure analysis to raise task quality and code quality.
Results
- 50+ evaluations delivered with deterministic grading and anti-cheat validation.
- 200+ evaluations reviewed, improving correctness, coverage, and reproducibility across the set.
- Clearer, harder-to-game tasks that produce more trustworthy model-capability signal.