Turing is developing one of the most rigorous STEM AI training datasets in the industry through its SciCode project. This initiative focuses on creating high-quality scientific coding tasks designed to train and evaluate advanced AI models. As a SciCode Trainer, you will play a vital role in advancing AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines. Your contributions will directly impact the quality and effectiveness of frontier AI systems.
Key Responsibilities:
- Write detailed scientific problem specifications that include one main problem and at least three logically connected sub-problems, progressively building toward the main solution.
- Implement verified golden solutions in Python, ensuring complete unit test coverage for accuracy and reliability.
- Design discriminative test cases that effectively distinguish correct from incorrect outputs generated by AI models.
- Conduct quality control (QC) validation checks on the Turing Central Task Platform (CTP), including Tier 1 structural checks and Tier 2 quality rubric assessments.
- Iterate on tasks based on QC feedback to meet Pass@K evaluation criteria across multiple large language model (LLM) judges such as GPT, Gemini, and Nemotron.
- Maintain a high standard of output quality with a low rate of rework, targeting consistent first-time approval at Level 1 (L1).
- Participate in regular sync calls for reviews, feedback sessions, and project standups during designated overlap hours.
Required Qualifications:
- Master’s or PhD degree in Mathematics.
- Strong proficiency in Python programming, particularly with experience in scientific computing.
- Ability to craft rigorous, well-posed scientific problems with clear constraints and expected outputs.
- Exceptional attention to detail to ensure tasks meet strict rubrics for well-posedness, test case discriminativeness, scientific correctness, and determinism.
- Prior experience in AI data annotation, research, or scientific writing.
- Familiarity with LLM evaluation frameworks or coding benchmarks.
- Experience working with scientific libraries such as NumPy, SciPy, SymPy, or other domain-specific scientific tools.
- Demonstrated research or academic project experience in a STEM field.
Engagement Details:
This is a contractor/freelancer role requiring a commitment of 40 hours per week with a mandatory 4-hour overlap with Pacific Standard Time (PST). Please note that this engagement does not include medical benefits or paid leave.
This position offers a unique opportunity to contribute to cutting-edge AI research by leveraging your expertise in mathematics and scientific computing. If you are passionate about advancing AI through high-quality STEM problem creation and evaluation, this role will allow you to make a significant impact.