EY is looking for a motivated Staff Data Engineer with 1 to 3 years of experience to join their Global Delivery Services team. This role involves developing, testing, and maintaining data pipelines primarily using Databricks and Azure services. The candidate will support data quality and governance initiatives while working closely with cross-functional teams to deliver reliable and scalable data solutions. The position offers a competitive salary and the flexibility to work from Hyderabad or any location within the country. EY fosters a globally connected and inclusive culture that encourages professional growth and development within a diverse, technology-driven environment.
Key Responsibilities
Develop, test, and maintain ETL/ELT pipelines using Databricks, PySpark, SQL, and Azure services under the guidance of senior team members. Support data ingestion, transformation, and processing of both structured and semi-structured data sources. Implement data quality checks, validation rules, and reconciliation controls in line with established standards. Assist in designing and maintaining logical and physical data models that align with business requirements. Ensure compliance with enterprise data governance, security, and data management policies. Support performance tuning efforts by identifying bottlenecks, documenting findings, and applying improvements. Collaborate with business analysts, data scientists, and downstream consumers to understand data needs and deliver appropriate solutions. Troubleshoot routine data pipeline issues and escalate complex problems with clear documentation. Prepare and maintain technical documentation, runbooks, and operational procedures. Leverage Databricks Genie Spaces and Genie Code features to enhance productivity and support AI-assisted development workflows.
Required Qualifications
Hands-on experience with Databricks is essential. Proficiency in PySpark and SQL for data processing and pipeline development. Basic understanding of Azure Data Platform services, including data ingestion and analytics components. Experience in developing or supporting ETL/ELT pipelines in cloud environments. Knowledge of data quality management concepts and validation practices. Familiarity with relational database management systems and fundamental data modeling techniques such as Star Schema, Snowflake Schema, and Dimensional Modeling. Awareness of data governance frameworks, metadata management, and data security principles. Ability to follow coding standards, maintain documentation, and incorporate feedback from code reviews. A Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related field. Between 1 to 3 years of relevant experience in data engineering and cloud-based data platforms.
Preferred Qualifications and Benefits
Python programming skills are advantageous. Exposure to Databricks Genie Spaces or Genie Code for AI-assisted analytics and development is a plus. Familiarity with Delta Lake, Lakehouse architecture, and CI/CD practices will be beneficial. Knowledge of Azure Data Factory, Azure Synapse, or other Azure analytics services is preferred. Experience with Git and basic DevOps practices is desirable.
EY offers future-focused skill development, world-class experiences, and a flexible, inclusive work environment. The company is committed to diversity, equity, and inclusion, fostering a culture where diverse perspectives are valued and employees can thrive. Benefits include comprehensive health and wellness packages, rewards, and continuous learning opportunities to support both professional and personal growth.
Candidates who are passionate about contributing to a better working world and align with EY’s values are encouraged to apply. EY is an equal opportunity employer dedicated to creating an equitable recruitment experience and providing accommodations as needed.