Design, build, and maintain scalable batch and real-time ETL/ELT pipelines on enterprise data platforms, including Databricks, Snowflake, Microsoft Fabric, Cloudera, Informatica IDMC, and Oracle
Develop lakehouse and data warehouse solutions using Medallion (Bronze/Silver/Gold) architecture on Delta Lake, Apache Iceberg, OneLake, Snowflake, and Oracle Autonomous Data Warehouse (ADW)
Build and orchestrate data workflows using Databricks Lakeflow, Fabric Data Factory, Azure Data Factory, Informatica Cloud Data Integration, Snowflake Streams & Tasks, and Apache Airflow
Implement Change Data Capture (CDC) and streaming ingestion using Oracle GoldenGate, Apache Kafka, and Spark Structured Streaming
Apply dimensional data modelling, including Kimball star schemas, to deliver analytics-ready data marts
Develop Power BI semantic models, including Direct Lake, DAX, and row-level security, in partnership with BI and analytics teams
Implement data governance, security, data quality, and lineage using Databricks Unity Catalog, Microsoft Purview, Cloudera SDX, and Informatica Data Quality
Prepare governed, high-quality data for AI and Machine Learning use cases, including feature pipelines and RAG-ready datasets using vector search capabilities on Databricks, Snowflake Cortex, and Oracle AI Vector Search
Apply DataOps practices, including Git-based version control, CI/CD for data pipelines, automated testing, and Infrastructure as Code
Monitor, troubleshoot, and optimize production pipelines for performance and cloud cost, supporting the practice's 99.90% uptime SLA commitment
Work directly with client stakeholders across the delivery lifecycle, including requirements gathering, data model validation, UAT, Go-Live, and post-Go-Live SLA support
Mentor junior engineers and contribute to internal engineering standards, reusable pipeline frameworks, and technical documentation
Person Specification
Possess a Bachelor's Degree in Data Science or a higher qualification, such as an MSc in Data Science, Data Engineering, or Artificial Intelligence, from a recognized university
Have 2–5 years of professional experience in building and operating enterprise data pipelines, data warehouses, or lakehouses
Possess hands-on experience with at least two of the following platforms: Databricks, Snowflake, Microsoft Fabric/Azure Data Services, Cloudera, Informatica (IDMC/PowerCenter), or Oracle (ADW/Exadata/ODI)
Demonstrate strong experience with Apache Spark and distributed data processing at scale
Possess a solid understanding of data modelling, data quality, and data governance principles
Have experience developing Power BI semantic models and reports
Demonstrate strong communication skills and the ability to work directly with client stakeholders
Professional certifications such as Databricks Certified Data Engineer (Associate/Professional), SnowPro Core or SnowPro Advanced: Data Engineer, Microsoft Certified: Fabric Data Engineer Associate (DP-700) or Fabric Analytics Engineer Associate (DP-600), Informatica IDMC, or Oracle Autonomous Database certifications will be considered an added advantage
Experience in migrating legacy ETL platforms such as Informatica PowerCenter, SSIS, or ODI, or on-premises data warehouses to modern cloud lakehouse platforms will be considered an added advantage
Experience with real-time streaming and CDC tools, including Kafka and Oracle GoldenGate, will be considered an added advantage
Exposure to GenAI data engineering, including RAG pipelines, vector databases, and LLM-ready data preparation, will be considered an added advantage
Experience with dbt, Terraform, Azure DevOps, or GitHub Actions will be considered an added advantage
Prior experience in banking, telecommunications, or public-sector data projects will be considered an added advantage
Demonstrate strong SQL and Python (PySpark) skills; knowledge of Scala or Java will be considered an added advantage