Turing is developing one of the most rigorous STEM AI training datasets in the industry through its SciCode project. This initiative focuses on creating high-quality scientific coding tasks designed to train and evaluate advanced AI models. As a SciCode Trainer, you will play a vital role in cutting-edge AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines, particularly physics. Your work will directly impact the quality and effectiveness of AI model training.
Key Responsibilities:
- Write detailed scientific problem specifications that include one main problem and at least three logically connected sub-problems, progressively building toward the solution of the main problem.
- Implement verified "golden" solutions in Python, ensuring complete unit test coverage for accuracy and reliability.
- Design discriminative test cases that effectively distinguish between correct and incorrect AI model outputs.
- Conduct quality control (QC) validation checks on the Turing Central Task Platform (CTP), including Tier 1 structural checks and Tier 2 quality rubrics.
- Iterate on tasks based on QC feedback to meet Pass@K evaluation criteria across multiple large language model (LLM) judges such as GPT, Gemini, and Nemotron.
- Maintain a high standard of output quality with a low rework rate, aiming for consistent first-time approval (L1) on submissions.
- Participate in regular sync calls for reviews, feedback sessions, and project standups during overlapping hours.
Required Qualifications:
- Master’s or PhD degree in Physics.
- Strong Python programming skills with experience in scientific computing.
- Ability to write rigorous, well-posed scientific problems with clear constraints and expected outputs.
- Exceptional attention to detail to ensure tasks meet strict rubrics for well-posedness, test case discriminativeness, scientific correctness, and determinism.
- Prior experience in AI data annotation, scientific research, or scientific writing.
- Familiarity with LLM evaluation frameworks or coding benchmarks.
- Experience with scientific computing libraries such as NumPy, SciPy, SymPy, or other domain-specific tools.
- Published research or academic project experience in a STEM domain is highly desirable.
Preferred Qualifications and Offer Details:
- Commitment to a 40-hour workweek with a required 4-hour overlap with Pacific Standard Time (PST).
- Engagement type is a contractor/freelancer role without medical benefits or paid leave.
- The contract duration is 8 weeks.
This role offers a unique opportunity to contribute to the forefront of AI research by developing and validating scientific coding tasks that help shape the future of intelligent systems. If you have a strong background in physics, programming, and scientific problem formulation, this position will allow you to apply your expertise in a dynamic and impactful environment.