About the Role
This is a senior, hands-on leadership role at an early-stage AI/ML startup building infrastructure for reinforcement learning environments and post-training data. You'll own the strategy and systems that measure, improve, and scale training data quality for frontier agents — shaping both the technical direction and the internal culture around what makes agent training data genuinely useful.
What You'll Do
- Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data, benchmarks, and domain-specific workflows.
- Define data quality strategy by building QC systems, enforcing standards, and designing experiments to grade agent outputs.
- Develop and implement methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing.
- Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows.
- Turn qualitative research insights into production systems — internal tools, dashboards, validation pipelines, and feedback loops.
- Mentor research engineers to maintain a high bar for technical rigor, clarity, and execution speed.
What We're Looking For
- 5+ years of experience in research or data quality engineering, specifically building systems for AI/ML data evaluation.
- Demonstrated track record leading technical teams or projects on ambiguous problems, from definition through implementation and iteration.
- Advanced proficiency in Python, Docker, and Linux environments.
- Deep intuition for what makes AI agent training data realistic, learnable, diverse, reliable, and useful.
- Experience translating research insights into production-grade tools and pipelines.
- Ability to design metrics, experiments, and QA/QC processes — not just execute them.
- Experience working with subject-matter experts to capture domain judgment and convert it into scalable review or generation systems.
- Strong written communication skills; able to explain methodology clearly to technical and non-technical audiences alike.
- Comfort navigating complex systems involving domain experts, vendors, model outputs, graders, and infrastructure.
- Prior early-stage startup experience; self-directed and effective in fast-paced, resource-constrained environments.
Compensation & Benefits
Base salary $150,000 – $250,000 USD annually. Visa sponsorship is available.
Location
On-site in San Francisco, CA. This role is not remote.