Turing is a rapidly growing AI company focused on accelerating the development and deployment of advanced AI systems. The company collaborates with leading AI labs to push the boundaries of model capabilities in areas such as reasoning, coding, multilinguality, and STEM knowledge. These innovations are then applied to build real-world AI solutions that address critical business challenges. This role offers the opportunity to work on fine-tuning large language models by applying strong analytical and mathematical expertise. The ideal candidate will have a solid foundation in mathematics, ranging from advanced engineering entrance concepts to graduate and PhD-level topics, and the ability to clearly articulate complex ideas through step-by-step reasoning. You will contribute directly to identifying model limitations and leveraging cutting-edge AI tools, helping to future-proof your career in this evolving field.
Key Responsibilities:
- Problem Creation & Solutioning: Design and solve challenging physics problems that push the limits of large language models.
- Authoring Gold-Standard Data: Develop clear, high-quality solutions with detailed, articulated reasoning presented step-by-step.
- Research Collaboration: Work closely with LLM researchers to align task designs with evaluation objectives, focusing on areas where models typically struggle, such as symbolic manipulation, abstraction, and multi-step reasoning.
- Benchmark Definition: Assist in defining and building new evaluation benchmarks covering physics topics from early undergraduate to PhD levels.
Required Qualifications:
- Strong analytical, research, and problem-solving skills paired with excellent English comprehension.
- Exceptional written communication skills, capable of providing structured, constructive feedback and detailed annotations in a remote work environment.
- Demonstrated lateral thinking ability to create novel scenarios and evaluate complex reasoning pathways.
- Self-motivated and able to work independently within a fast-paced, remote-first setup.
- Access to a personal desktop or laptop with a stable, high-speed internet connection.
- Educational background or Doctorate in Mathematics or an equivalent technical field.
Preferred Qualifications and Benefits:
- Experience in AI evaluation, data annotation, content review, quality assurance, or related analytical roles is advantageous but not mandatory.
- Commitment of at least 4 hours per day, up to 40 hours per week, with a minimum of 4 hours overlapping with Pacific Standard Time (PST).
- Engagement type is contractor-based, with a contract length of 12 weeks.
Candidates who are shortlisted will receive a Job Interest Form as part of the evaluation process. Those selected for final consideration will be contacted regarding next steps, including onboarding requirements. This role offers a unique chance to contribute to setting benchmarks for AI capabilities in advanced physical sciences while developing expertise in frontier AI technologies.