Senior Data Engineer with 5 to 8 years of experience is required to design, develop, and optimize large-scale data platforms and ETL/ELT pipelines. The ideal candidate will have strong hands-on expertise in PySpark, AWS Glue, Amazon EMR, Amazon Redshift, and SQL-based data warehousing. Proven experience in performance tuning and data optimization is essential. This role involves close collaboration with a UAE-based customer, requiring availability during UAE business hours and adherence to UAE public holidays. The position is hybrid based in Pakistan, and the candidate must be comfortable working with distributed teams across different locations.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using PySpark, AWS Glue, and Amazon EMR.
- Build and optimize data ingestion, transformation, and processing frameworks for both structured and semi-structured data.
- Develop and maintain enterprise data warehouse solutions leveraging Amazon Redshift.
- Write complex SQL queries, stored procedures, and data transformations to support analytics and reporting needs.
- Implement efficient ETL/ELT processes to move data across multiple systems and platforms.
- Perform performance tuning and optimization of Spark jobs, ETL pipelines, SQL queries, and Redshift workloads.
- Ensure data quality, integrity, security, and governance across all data platforms.
- Troubleshoot production issues and conduct root cause analysis for data-related incidents.
- Collaborate with business stakeholders, analysts, architects, and engineering teams to understand and fulfill data requirements.
- Participate in code reviews, technical design discussions, and implement best practices.
- Monitor data pipelines proactively and identify opportunities for performance improvements and automation.
- Create and maintain technical documentation, data models, and operational procedures.
Required Qualifications
- 5 to 8 years of experience in Data Engineering and Data Warehousing.
- Strong hands-on experience with PySpark.
- Extensive experience working with AWS Glue.
- Proven expertise in building and managing workloads on Amazon EMR.
- Strong knowledge of Amazon Redshift and data warehouse solutions.
- Excellent SQL development skills, including query optimization.
- Deep understanding of Data Warehousing concepts, dimensional modeling, and ETL/ELT processes.
- Experience in performance tuning of Spark jobs, SQL queries, ETL pipelines, and data warehouse workloads.
- Ability to handle large-scale datasets and distributed data processing environments.
- Strong debugging, troubleshooting, and analytical skills.
Preferred Qualifications and Additional Skills
- Experience with other AWS services such as S3, IAM, CloudWatch, Lambda, and Step Functions.
- Familiarity with CI/CD pipelines and DevOps practices related to data platforms.
- Experience with workflow orchestration tools.
- Knowledge of data governance, security, and compliance standards.
- Exposure to Agile/Scrum development methodologies.
- Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
- Relevant AWS certifications are considered an advantage.
Soft Skills
- Strong communication and stakeholder management capabilities.
- Ability to work independently in a remote or hybrid environment.
- Excellent problem-solving and analytical thinking skills.
- Effective collaboration with cross-functional and geographically distributed teams.
- Ownership mindset with a commitment to delivering high-quality solutions.
This role offers the opportunity to work on cutting-edge data engineering projects within a dynamic and collaborative environment, supporting a UAE-based client while maintaining a flexible hybrid work arrangement in Pakistan.