Evaluation systems
Deterministic graders, realistic terminal tasks, pinned environments, and anti-cheat validation that preserve signal.
Software engineer · Salt Lake City
Backend engineer focused on data pipelines, APIs, and cloud infrastructure, with an emphasis on observability, reproducibility, and operational reliability. Currently developing AI model evaluation systems at Snorkel AI
01 / Selected evidence
Built and reviewed software-engineering evaluations that measure whether frontier AI models can ship correct code.
Patched connector to enable Snowflake sync; migration unblocked.
Modernized docket notices to Kotlin with deterministic PDF generation.

Personal build · real-world outcome
An audio-first citizenship study app I built for my mom.
Bilingual, audio-first exam prep with adaptive testing—built for one real learner, and it helped her pass.
Read the story02 / Operating range
The common thread is not a framework. It is turning ambiguous, failure-prone systems into something a team can test, observe, and trust.
Deterministic graders, realistic terminal tasks, pinned environments, and anti-cheat validation that preserve signal.
Schema mapping, reconciliation, connector diagnostics, live sync, and careful production migration at enterprise scale.
Telemetry pipelines, backend APIs, observability, CI/CD, and reproducible systems that help teams move with confidence.
03 / Open channel
I’m open to software engineering roles and focused contract work. A concise email is the fastest way to start.