We are looking for a skilled Databricks Engineer to design, develop, and optimize scalable data engineering solutions using the Databricks Lakehouse Platform. The ideal candidate will possess strong expertise in Databricks, Apache Spark, PySpark, SQL, Delta Lake, cloud platforms, and enterprise data pipelines. This role requires a hands-on professional who can build and maintain robust data architectures, ensuring efficient data processing and governance within a cloud environment.
Key Responsibilities
- Design and develop scalable ETL/ELT pipelines utilizing Databricks, Spark, and PySpark technologies.
- Develop and maintain Lakehouse architecture by implementing Delta Lake with Bronze, Silver, and Gold data layers.
- Build both batch and real-time data processing and streaming solutions to support diverse business needs.
- Manage Databricks Jobs, Workflows, and Delta Live Tables (DLT) to ensure smooth and automated data operations.
- Implement data governance and access controls using Unity Catalog to maintain data security and compliance.
- Optimize Spark jobs, SQL queries, data pipelines, and cloud resources for maximum performance and cost efficiency.
- Integrate Databricks with various enterprise data sources and cloud platforms such as Azure, AWS, or GCP.
- Implement data quality checks, validation, monitoring, and error-handling processes to ensure data reliability.
- Support Git, CI/CD, and DevOps practices to enable automated deployments and continuous integration.
- Troubleshoot production issues promptly to maintain reliable and stable data platform operations.
- Collaborate closely with Data Architects, Developers, Analysts, and business stakeholders to deliver effective data solutions.
Required Qualifications
- 5 to 10 years of experience in Data Engineering with strong hands-on expertise in Databricks.
- Advanced knowledge of Databricks, Apache Spark, PySpark, and SQL for data processing and analytics.
- Solid experience with Delta Lake and Lakehouse architecture concepts.
- Hands-on experience managing Unity Catalog, Workflows, and Delta Live Tables (DLT).
- Proven ability to work with both batch and streaming data processing frameworks.
- Strong understanding of ETL/ELT processes, data modeling, and data warehousing principles.
- Experience working with cloud platforms such as Azure, AWS, or Google Cloud Platform (GCP).
- Familiarity with Git, CI/CD pipelines, and DevOps methodologies for streamlined development and deployment.
- Excellent analytical, troubleshooting, and problem-solving skills to handle complex data challenges.
Certifications
Candidates must hold at least one of the following certifications:
- Databricks Certified Data Engineer – Associate
- Databricks Certified Data Engineer – Professional
- Azure, AWS, or GCP Data Engineering certification
Education
A Bachelor’s degree in Computer Science, Information Technology, Data Engineering, or a related field is required.
This role offers an excellent opportunity to work with cutting-edge data technologies in a dynamic environment, contributing to the development of scalable and efficient data solutions that drive business success.