Turing is developing one of the most rigorous STEM AI training datasets in the industry through its SciCode project. This initiative focuses on creating high-quality scientific coding tasks designed to train and evaluate advanced AI models. As a SciCode Trainer, you will play a vital role in cutting-edge AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines. Your work will directly impact the quality and effectiveness of AI training datasets used by leading models.
Key Responsibilities
- Develop scientific problem specifications that include one main problem along with at least three logically connected sub-problems, each progressively building toward solving the main challenge.
- Implement verified golden solutions in Python, ensuring complete unit test coverage for accuracy and reliability.
- Design discriminative test cases that effectively differentiate between correct and incorrect outputs generated by AI models.
- Conduct quality control (QC) validation checks on the Turing Central Task Platform (CTP), including Tier 1 structural checks and Tier 2 quality rubrics.
- Iterate on tasks based on QC feedback to meet Pass@K evaluation criteria across multiple large language model (LLM) judges such as GPT, Gemini, and Nemotron.
- Maintain consistently high output quality with a low rework rate, aiming for first-submission approval at Level 1 (L1).
- Participate in regular sync calls for reviews, feedback sessions, and project standups during designated overlap hours.
Required Qualifications
- Master’s degree or PhD in Mathematics.
- Strong proficiency in Python programming, especially in scientific computing contexts.
- Demonstrated ability to write rigorous, well-posed scientific problems with clear constraints and expected outputs.
- Exceptional attention to detail to ensure tasks meet strict rubrics related to well-posedness, test case discriminativeness, scientific correctness, and determinism.
- Prior experience in AI data annotation, research, or scientific writing.
- Familiarity with LLM evaluation frameworks or coding benchmarks.
- Experience using scientific computing libraries such as NumPy, SciPy, SymPy, or other domain-specific tools.
- Published research or academic project experience in a STEM field is highly valued.
Preferred Qualifications and Offer Details
- Commitment to a 40-hour workweek with a required 4-hour overlap with Pacific Standard Time (PST).
- Engagement is on a contractor or freelancer basis, without medical benefits or paid leave.
This role offers a unique opportunity to contribute to pioneering AI research by shaping the datasets that train next-generation models. If you have a strong STEM background, excellent Python skills, and a passion for scientific problem-solving, this position allows you to work at the intersection of AI and scientific innovation.