We are looking for a skilled Databricks Data Engineer with 5 to 10 years of experience to design, develop, and optimize scalable data engineering solutions on the Databricks Lakehouse Platform. The ideal candidate will have deep expertise in Apache Spark, PySpark, SQL, Delta Lake, and cloud platforms, along with a strong understanding of enterprise data architecture. This role involves building robust data pipelines, managing data workflows, and ensuring data quality and security within a collaborative, fast-paced environment.
Key Responsibilities
- Design and develop scalable ETL/ELT data pipelines using Databricks, Apache Spark, and PySpark.
- Develop and maintain Lakehouse architecture leveraging Delta Lake and Bronze, Silver, Gold data layers.
- Build batch and real-time data processing and streaming solutions to support business needs.
- Develop and manage Databricks Jobs, Workflows, and Delta Live Tables (DLT) for efficient data orchestration.
- Implement data governance, security, and access controls using Unity Catalog to ensure compliance and data protection.
- Optimize Spark jobs, SQL queries, and data pipelines for improved performance and cost efficiency.
- Integrate Databricks with enterprise data sources and cloud services to enable seamless data flow.
- Implement frameworks for data quality, validation, monitoring, and error handling to maintain data integrity.
- Utilize Git, CI/CD, and DevOps practices to automate deployments and streamline development processes.
- Troubleshoot production issues promptly to ensure reliability and availability of data platforms.
- Collaborate closely with Data Architects, Developers, Analysts, and business stakeholders to deliver effective data solutions.
Required Qualifications
- 5 to 10 years of hands-on experience in Data Engineering with strong expertise in Databricks.
- Advanced knowledge of Databricks, Apache Spark, PySpark, and SQL for data processing and analytics.
- Proven experience with Delta Lake and Lakehouse architecture concepts.
- Hands-on experience managing Unity Catalog, Databricks Workflows, and Delta Live Tables (DLT).
- Solid background in batch and streaming data processing techniques.
- Strong understanding of data modeling, ETL/ELT processes, and data warehousing principles.
- Experience working with at least one major cloud platform such as Azure, AWS, or GCP.
- Familiarity with CI/CD pipelines, Git version control, and DevOps methodologies.
- Excellent troubleshooting, analytical, and problem-solving skills to handle complex data challenges.
Preferred Qualifications and Certifications
- Databricks Certified Data Engineer – Associate (mandatory).
- Databricks Certified Data Engineer – Professional (preferred).
- Cloud-specific Data Engineering certifications from Azure, AWS, or GCP (preferred).
- Bachelor’s degree in Computer Science, Information Technology, Data Engineering, or a related field.
This position offers the opportunity to work with cutting-edge data technologies in a dynamic environment, contributing to enterprise-wide data initiatives and innovation. Candidates who meet the qualifications and are passionate about data engineering and cloud platforms are encouraged to apply.