Turing is collaborating with NVIDIA to develop one of the most rigorous STEM AI training datasets in the industry through the SciCode project. This initiative focuses on creating high-quality scientific coding tasks that help train and evaluate advanced AI models. As a SciCode Trainer, you will play a vital role in cutting-edge AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines. Your work will directly impact the development of frontier AI technologies by ensuring the quality and accuracy of scientific coding challenges.
Key Responsibilities
- Write detailed scientific problem specifications that include one main problem and at least three logically connected sub-problems, designed to progressively build toward the solution of the main problem.
- Implement verified golden solutions in Python, ensuring complete unit test coverage for reliability and accuracy.
- Design discriminative test cases that effectively distinguish correct model outputs from incorrect ones.
- Conduct quality control (QC) validation checks on the Turing Central Task Platform (CTP), including Tier 1 structural checks and Tier 2 quality rubrics.
- Iterate on tasks based on QC feedback to meet Pass@K evaluation criteria across multiple large language model (LLM) judges such as GPT, Gemini, and Nemotron.
- Maintain high output quality with a low rework rate, aiming for consistent first-time approval (L1) on submissions.
- Participate in regular sync calls for reviews, feedback sessions, and project standups during overlapping hours with the team.
Required Qualifications
- Master’s or PhD degree in Biology.
- Strong Python programming skills, particularly in scientific computing environments.
- Proven ability to write rigorous and well-posed scientific problems with clear constraints and expected outputs.
- Exceptional attention to detail to ensure tasks meet strict rubrics for well-posedness, test case discriminativeness, scientific correctness, and determinism.
- Prior experience in AI data annotation, research, or scientific writing.
- Familiarity with LLM evaluation frameworks or coding benchmarks.
- Experience with scientific libraries such as NumPy, SciPy, SymPy, or other domain-specific scientific tools.
- Published research or academic project experience in a STEM domain demonstrating quality standards.
Preferred Qualifications and Offer Details
- Commitment to a 40-hour workweek with a required 4-hour overlap with Pacific Standard Time (PST).
- Engagement is on a contractor/freelancer basis without medical benefits or paid leave.
- The contract duration is 8 weeks, providing a focused, project-based opportunity to contribute to advanced AI research.
This role offers a unique chance to work at the intersection of scientific research and artificial intelligence, contributing to the development of datasets that push the boundaries of AI capabilities in STEM fields.