We are looking for a Senior AI Data Engineer to design, build, and maintain the critical data infrastructure that supports our AI systems. This role is responsible for managing the flow, modeling, and quality of data from source systems through the data warehouse and into retrieval and feature layers that enable large language model (LLM) pipelines, agentic workflows, and analytical products. The ideal candidate is primarily a skilled software engineer who approaches data engineering with rigor—developing well-structured warehouses, writing maintainable production Python code, and treating data pipelines as versioned, tested, and observable software rather than simple scripts. This position works closely with AI/ML Data Scientists, where the engineer owns the pipeline, schema, and data guarantees, while the scientists focus on algorithms, prompts, and evaluation.

Key Responsibilities

Data Warehousing & Dimensional Modeling
Design and evolve dimensional models using Kimball methodology, including star schemas, conformed dimensions, and enterprise bus matrices. Implement fact tables such as transaction, periodic snapshot, and accumulate snapshot with clear grain definitions. Manage slowly changing dimensions (Types 1, 2, 3, and hybrids) with effective dating, surrogate keys, and late-arriving data handling. Model semi-structured and unstructured data sources—documents, transcripts, event streams—into queryable formats while preserving provenance. Maintain a governed semantic layer to ensure consistent metrics for both human analysts and automated consumers.

Pipeline & ETL/ELT Engineering
Build and operate batch and streaming pipelines using orchestration tools like Airflow, Prefect, or Dagster, supporting backfill, replay, and idempotent retries. Develop ELT transformation layers with well-tested, documented, and version-controlled SQL. Own data contracts between producers and consumers, managing schema evolution, compatibility, and breaking-change procedures. Instrument pipelines for observability, monitoring freshness, volume, distribution, and schema drift, with alerting that distinguishes real incidents from normal variance. Maintain end-to-end data lineage from source to warehouse to retrieval.

Data Engineering for AI
Develop embedding generation and refresh pipelines, including chunk materialization, embedding jobs, incremental re-embedding, and index lifecycle management. Operate vector stores such as Pinecone, Weaviate, Chroma, Milvus, or pgvector as production systems, managing capacity, indexing strategies, upsert/delete correctness, and staleness SLAs. Build extraction pipelines for unstructured data using OCR, document parsers, and vision-language models, treating extraction confidence as a key data attribute. Implement preprocessing and feature pipelines that support model training and inference, ensuring train/serve consistency. Provide data access to AI services via well-defined FastAPI or gRPC interfaces.

Engineering Craft
Apply SOLID principles and domain-driven design to create bounded contexts aligned with business domains, using ubiquitous language and separating domain logic from transport and persistence layers. Maintain comprehensive test coverage, including unit, contract, and data quality assertions, treating untested pipelines as incomplete. Own CI/CD pipelines, containerization, and environment promotion for data services. Contribute to code reviews, architectural decisions, and internal standards.

Governance & Cost Management
Implement access control, PII handling, data retention, and audit requirements at the data layer. Manage warehouse and pipeline costs through partitioning, clustering, materialization strategies, and storage tiering.

Collaboration
Translate business problems into data models in partnership with product and client stakeholders. Mentor junior and mid-level data engineers to foster team growth.

Required Qualifications
- 5 to 10+ years of experience in software or data engineering, with significant production data platform work.
- Deep expertise in Kimball dimensional modeling, including grain selection, handling many-to-many relationships, and denormalization judgment. Familiarity with alternative modeling approaches and their tradeoffs.
- Expert-level SQL skills, including window functions, CTEs, query plan analysis, and performance tuning on columnar warehouses.
- Advanced Python programming skills with production-grade practices such as typing, packaging, dependency management, and testing.
- Strong understanding and application of SOLID principles and domain-driven design in real-world systems.
- Experience with orchestration frameworks like Airflow, Prefect, or Dagster in production environments.
- Expert knowledge of cloud platforms (AWS, Azure, or GCP), including storage, compute, IAM, networking, and cost management.
- Proficiency with containerization, Kubernetes (EKS/AKS/GKE), and CI/CD pipelines.
- Experience working with Lakehouse table formats such as Iceberg, Delta Lake, or Hudi, including compaction, snapshot expiry, and schema/partition evolution.

Preferred Qualifications
- Experience building data layers for production Retrieval-Augmented Generation (RAG) systems, including hybrid search infrastructure and index freshness guarantees.
- Familiarity with streaming technologies such as Kafka, Kinesis, Flink, or Spark Structured Streaming.
- Experience with dbt or equivalent transformation and testing frameworks.
- Knowledge of data quality tools like Great Expectations or Soda, and data catalog/lineage platforms.
- Understanding of ML lifecycle tools (MLflow, Weights & Biases) and feature stores to collaborate effectively with data scientists.
- Proficiency in a second programming language such as TypeScript, Java, Go, Scala, or Rust.
- Experience with AI security, governance, and compliance frameworks.
- Contributions to open-source data or AI infrastructure projects.

This role offers the opportunity to shape the data foundation of cutting-edge AI systems in a collaborative and technically rigorous environment.

Job Details

Total Positions:
1 Post
Job Shift:
First Shift (Day)
Job Type:
Job Location:
Gender:
No Preference
Age:
18 - 65 Years
Career Level:
Manager
Maximum Experience:
5 Years
Apply Before:
Oct 10, 2026
Posting Date:
Oct 04, 2026

Strategic Systems International

· 11-50 employees - Lahore

What is your Competitive Advantage?

Get quick competitive analysis and professional insights about yourself
Talk to our expert team of counsellors to improve your CV!
Try Rozee Premium

Similar Job Titles

I found a job on Rozee!