OdoCore is seeking an experienced Data Engineer to contribute to a significant data platform modernization project. This role involves hands-on work migrating a legacy Hadoop-based data warehouse to a modern lakehouse architecture. The successful candidate will work directly on production-critical data pipelines that support essential business reporting and analytics. Practical, real-world experience with the relevant tools and technologies is crucial for this position.
Key Responsibilities
- Lead the migration from Spark 2 to Spark 3, ensuring seamless transition and performance optimization.
- Migrate data processing workflows from Apache Hive to Apache Iceberg, focusing on table formats, partitioning, and schema evolution.
- Transition orchestration pipelines from Oozie to Apache Airflow to improve workflow management and scalability.
- Offload workloads from IBM Netezza and IBM DataStage onto a Spark 3 and Iceberg-based lakehouse platform.
- Collaborate closely with team members to maintain and enhance production-critical pipelines that power core business analytics.
- Analyze and optimize complex legacy SQL queries to improve performance and maintainability.
- Work within a Linux/Unix environment using Git for version control and collaboration.
- Engage with Cloudera CDP platform components such as Impala and Ranger, or quickly ramp up on these technologies as needed.
Required Qualifications
- Minimum of 3 years of professional experience in data engineering roles.
- Demonstrated experience in at least one major data migration, ETL modernization, or lakehouse architecture build.
- Strong proficiency in SQL, with the ability to read, understand, and optimize complex and poorly written legacy queries.
- Hands-on experience with Apache Spark (both Spark 2 and Spark 3), using PySpark or Scala.
- Solid knowledge of Apache Hive and Apache Iceberg, including table formats, partitioning strategies, and schema evolution concepts.
- Familiarity with ETL and data warehousing fundamentals, including medallion architecture (bronze/silver/gold) and dimensional modeling.
- Experience with orchestration tools such as Oozie and/or Apache Airflow.
- Comfortable working in Linux/Unix environments and using Git for source control.
- Exposure to Cloudera CDP platform components, particularly Impala and Ranger, or the ability to quickly learn these technologies.
Preferred Qualifications and Benefits
This role is ideal for mid-to-senior individual contributors who thrive in a fast-paced environment and have a strong technical background in data engineering. While not explicitly stated, candidates can expect to work on cutting-edge technologies within a collaborative team focused on large-scale data modernization efforts, offering opportunities for professional growth and skill development.
OdoCore values practical expertise and a proactive approach to problem-solving, making this an excellent opportunity for data engineers looking to deepen their experience in modern data architectures and migration projects.