We are looking for experts in Bioinformatics and Computational Genomics to develop complex and agentic genomics tasks along with reference solutions for GeneBench Pro. This role involves creating realistic computational biology problems that challenge AI agents to interpret biological data, write and execute code, navigate scientific files, and generate objectively verifiable results. The tasks should mirror authentic genomics and bioinformatics workflows rather than isolated biology questions, ensuring relevance and rigor in scientific analysis.
Key Responsibilities:
- Design innovative and model-challenging tasks in computational genomics and bioinformatics.
- Develop tasks that require advanced scientific judgment, including managing ambiguous or noisy data, detecting artifacts, selecting suitable analytical methods, revising assumptions based on intermediate findings, and deciding when conclusions are ready for decision-making.
- Construct tasks using realistic scientific datasets such as FASTA/FASTQ, VCF, BAM/SAM, BED, TSV/CSV, sequence annotations, expression data, and germline or somatic variant data.
- Create multi-step analyses involving Python programming, command-line tools, and established bioinformatics libraries.
- Prepare clear task specifications, input datasets, expected output formats, and deterministic or objectively verifiable ground truths.
- Develop expert reference solutions and reproducible computational workflows compatible with the provided Python-based environment.
- Validate that tasks are scientifically accurate, solvable with the provided information, and sufficiently challenging for cutting-edge AI models.
- Design robust grading criteria that differentiate scientifically valid solutions from superficially plausible ones.
- Ensure all deliverables are well documented, reproducible, and ready for client use.
- Maintain high quality and throughput while integrating feedback from reviewers.
- Communicate progress, challenges, and scientific or technical needs effectively to project leads and reviewers.
Required Qualifications:
- Ph.D., postdoctoral experience, or equivalent research background in Bioinformatics, Computational Biology, Genomics, Computational Genetics, or a closely related field.
- Strong hands-on programming skills in Python.
- Experience analyzing biological sequence or genomics datasets.
- Proficiency working in Linux and command-line computational environments.
Preferred Qualifications and Benefits:
- Experience with genomics workflows such as variant analysis, transcriptomics, sequence analysis, phylogenetics, population genetics, functional genomics, or clinical genomics.
- Familiarity with common bioinformatics libraries and tools including Biopython, pandas, NumPy, SciPy, samtools, bcftools, BLAST, PLINK, or equivalents.
- Experience developing reproducible scientific pipelines.
- Background in evaluating AI or large language models (LLMs) on computational scientific tasks.
- Strong understanding of the experimental and biological context underlying computational analyses.
Additional Advantages:
- Experience working with AI agents or coding agents.
- Experience designing benchmark datasets or automated grading systems.
- Publications in computational genomics or bioinformatics.
- Familiarity with Docker or containerized scientific workflows.
Engagement Details:
This is a contractor/freelancer assignment with a commitment of 40 hours per week and a required overlap of 4 hours with Pacific Standard Time (PST). The contract duration is six weeks. Please note that this is a contractor role without medical benefits or paid leave.