Available for data engineering opportunities

Peter Opapa

Data Engineer & Cloud Architect building reliable, cloud-native pipelines and real-time systems.

I design end-to-end data platforms on Azure, from ingestion to analytics-ready models using tools like Kafka, Spark, dbt, and Snowflake.

9+
Certifications
9+
Projects
5+
Awards
4+
Clients
9End-to-end builds
Cloud + real-timeCore focus areas
Open sourceGitHub-linked work
Data systems / 01 Online
Peter Opapa

Turning complex data into

clear, useful decisions.

Azure Kafka Spark dbt
About / 01

Building data systems that turn complexity into measurable impact.

I'm a Data Engineer and pan-African innovator with a track record of building data systems that matter. I've represented Kenya at Carnegie Mellon University's AI programme, led cross-border teams across the Afretec network, and earned recognition — including a Silver Award, seed funding, and a CMU-Africa incubation invitation — for work at the intersection of data engineering and social impact.

Technically, I specialise in scalable cloud-native pipelines on Azure — from raw ingestion through real-time streaming to analytics-ready models — using tools like Databricks, Kafka, Spark, dbt, and Snowflake. I care about systems that are dependable, observable, and built to last.

Peter Opapa - Data Engineer

Peter Opapa

Data Engineer · Nairobi, Kenya

Based in

Nairobi, Kenya

Speciality

Cloud data platforms

9+

Engineering projects

5+

Awards & recognitions

4+

Clients supported

#1

CMU-Africa project rank

Cloud-native systems

Reliable platforms designed for scale, security, and repeatable delivery.

AzureDatabricksADLS

Real-time data

Streaming pipelines that make fresh, trusted data available when it matters.

KafkaSparkAirflow

Analytics engineering

Well-modeled data products that help teams move from questions to decisions.

dbtSnowflakePower BI
My approach

From raw signals to useful decisions.

A dependable data lifecycle, designed with clarity at every stage.

01

Ingest

Capture signals

02

Transform

Shape the data

03

Model

Create meaning

04

Observe

Build trust

05

Deliver

Enable action

Education / 02

A strong engineering foundation for building systems that scale.

From electrical and electronics engineering to cloud-native data platforms, every stage has strengthened how I solve complex problems.

Completed degree

BSc. Electrical and Electronics Engineering

University of Nairobi

2021 – 2026

A rigorous engineering programme that developed my foundation in systems thinking, analytical problem-solving, programming, electronics, and designing reliable solutions.

First Class Honours 2nd Best 4th Year Student
Secondary education

Kenya Certificate of Secondary Education (KCSE)

Homabay High School

2017 – 2020
Grade: A Plain — 82 Points
Experience / 03

Turning technical depth into meaningful outcomes.

A timeline of engineering, innovation, and collaboration across data, cloud, IoT, and pan-African problem-solving.

Current role

AI & Data Engineer

Sisili Agritech

Present
  • Design and develop AI-enabled data solutions that support smarter, more efficient agricultural operations.
  • Build and integrate IoT data pipelines that transform connected-farm signals into reliable, actionable insights.
  • Collaborate across data, software, and domain teams to apply machine learning and analytics to real-world agritech challenges.
AI-enabledOperational solutions
IoT dataConnected agriculture
Impact-ledAgritech innovation
Featured milestone

Finalist, Case Competitions

McKinsey & Company — Nairobi, Kenya

July 2026
  • Selected as 1 of 5 finalist teams in Nairobi, Kenya, from a highly competitive applicant pool screened on CV, GPA, and other criteria.
  • Developed and presented a data-driven strategic solution to a real-world business challenge before a panel of McKinsey consultants.
  • Sharpened structured problem-solving, hypothesis-driven analysis, and client-facing communication skills valued in top-tier consulting.
1 of 5Finalist teams
Data-ledStrategic solution
McKinseyConsultant panel

KamiLimu Fellow

KamiLimu Limited — Nairobi, Kenya (Hybrid)

Mar 2026 – Present

Selected as 1 of 40 mentees in the KamiLimu Mentorship Program, focusing on Data Engineering and professional development.

Tariff Equity — EPRA Hackathon 2026

EPRA — Nairobi, Kenya

April 2026
  • Competed in the EPRA Research Week Youth Summit Hackathon 2026 under Optimal Tariff Design for Energy Affordability and Equity.
  • Led the data science component of FairGrid — a regulatory platform helping EPRA reclassify Kenya's 8.8M residential electricity customers using 7 socioeconomic variables and ML clustering, reducing simulated subsidy leakage from 43.2% to 11.4% and improving the equity index from 0.41 to 0.74.
  • Outcome: 2nd Runners Up + Best Use of Data Award — judged by EPRA regulators and energy sector experts.

Silver Medallist & Innovator – Financial Inclusion

C4DLab / Afretec Network

Feb 2026 – Mar 2026
  • Led a Pan-African team of 7 researchers to engineer blockchain-driven fintech solutions for Libex, raising blockchain adoption for underserved communities from 30% to 60%.
  • Awarded Silver Award (2nd Place) out of 7 teams; secured exclusive invitation and seed funding for CMU-Africa Business Incubation Program in Kigali, Rwanda.

Participant – AI & Digital Manufacturing Track

African Inclusive Digital Industries, Carnegie Mellon University

Dec 2025
  • Selected to represent Kenya in a competitive, fully funded pan-African programme focused on AI, digital manufacturing, IoT, and optimization.
  • Designed and implemented an edge computing–based health monitoring system to reduce unnecessary hospital visits and enable continuous remote patient care.
  • Project ranked #1 among all participating teams.

Freelance Junior Data Engineer

Upwork — Remote

Mar 2024 – May 2025
  • Delivered scalable data engineering solutions for 4 clients; designed and optimised ETL workflows for structured and unstructured data.
  • Implemented cloud-based data infrastructure and translated client requirements into actionable analytics pipelines ensuring data quality, security, and compliance.

Electronics Engineering Intern (IoT)

Ubuntu Water Hub

May 2025 – Aug 2025
  • Assisted in the design and simulation of IoT devices for real-time water pump monitoring and control.
  • Evaluated prototype performance focusing on sensor data accuracy, reliability, and signal integrity.
  • Analysed real-time data (flow rate, status, pressure) to inform system design and evaluation.
Toolkit / 04

The tools behind dependable data products.

A practical toolkit for building cloud platforms, streaming systems, analytics models, and delivery workflows from end to end.

Cloud Data Platforms

Designing secure, scalable foundations for ingestion, processing, and analytics.

Azure Databricks Snowflake Oracle Cloud

Big Data Frameworks

Moving data in real time and orchestrating reliable batch and streaming workflows.

Apache Spark Apache Flink Apache Kafka Apache Airflow dbt Core

Programming & Databases

Writing production-minded code and shaping reliable storage layers for data products.

Python SQL PostgreSQL MongoDB SQL Server

DevOps & Tools

Automating repeatable delivery with version control, containers, infrastructure, and CI/CD.

Docker Git & GitHub GitHub Actions Terraform

Analytics & Visualization

Turning well-prepared data into clear reporting, exploration, and business insight.

Power BI Excel
Reliable by design
Automated for repeatability
Built for useful decisions
Credentials / 05

Proof of continuous learning.

Industry credentials across streaming, cloud architecture, analytics engineering, and modern data platforms.

Recognition / 06

Work that earned a seat at the table.

Recognition for building useful technology, leading ambitious teams, and solving problems with measurable social and technical impact.

Silver Award (2nd Place) – C4DLab Makerthon

Afretec Network Pan-African Competition

Awarded 2nd Place out of 7 Pan-African teams for engineering blockchain-driven fintech solutions for financial inclusion.

Mar 2026

Top Project Award – African Inclusive Digital Industries

Carnegie Mellon University, Kigali

Represented Kenya at the African Inclusive Digital Industries Programme. Group project on edge computing health monitoring ranked #1 among all participating teams.

Dec 2025

2nd Runners Up + Best Use of Data

EPRA Research Week Youth Summit Hackathon 2026

Recognised for data-driven regulatory platform FairGrid, judged by EPRA regulators and energy sector experts.

2026

Second Best 4th Year Student

University of Nairobi – Faculty of Engineering

Recognised as the second best 4th Year student in the Faculty of Engineering for the academic year 2024/2025.

2024/2025
5+Recognitions
3Pan-African or international
#1Top project ranking
Selected work / 07

Systems built to move data forward.

A selection of cloud, streaming, analytics, and infrastructure projects that show how I turn architecture into working systems.

Architecture diagram for a Medallion data warehouse
Featured build

Data Warehouse: Medallion Architecture

Built a production-ready SQL Server data warehouse using the Medallion architecture and T-SQL ETL. Features automated loading, transformation scripts, and data quality checks. The final output is a star-schema optimized for BI, serving as a reusable reference for SQL data pipelines.

Microsoft SQL Server Powershell Data Modelling ETL Git
View on GitHub
Architecture diagram for a real-time data pipeline with Kafka and Spark

Real-time Data Pipeline

A real-time streaming pipeline ingesting people's profile data from an API, orchestrated by Airflow. Kafka decouples data, which is backed up in PostgreSQL. Apache Spark processes the stream and stores enriched data in Cassandra. The entire system is containerized with Docker for easy deployment and scalability.

Apache Airflow Apache Kafka Apache Spark Apache Cassandra Docker PostgreSQL Python
View on GitHub
Financial transactions pipeline with Flink and Datadog

Financial Transactions Pipeline

Built a real-time financial analytics pipeline using Kafka, Flink SQL, and PostgreSQL. It ingests events from a Python producer, performs event-time aggregations in Apache Flink, and stores results in PostgreSQL Database. Datadog monitors the pipeline system health in real time. The entire system is containerized with Docker for easy deployment.

Python Apache Kafka Apache Flink PostgreSQL Datadog Docker
View on GitHub
Azure and Terraform infrastructure diagram

Azure Infrastructure Deployment Using Terraform

This project automates the deployment and management of Azure Kubernetes Service (AKS) infrastructure using Terraform, with CI/CD pipelines powered by GitHub Actions/ Azure DevOps. The setup includes a AKS, Azure Active Directory, Azure Resource Manager, Storage Accounts, and Azure Key Vault, ensuring a consistent and repeatable infrastructure deployment process.

Terraform Azure CLI GitHub Actions Git Azure Kubernetes Service Azure Key Vault Azure Active Directory
View on GitHub
Jumia ELT web scraping pipeline architecture

Jumia ELT Pipeline

This project implements a production-grade ELT data pipeline that scrapes laptop product data from Jumia Kenya using BeautifulSoup and Requests, then processes it through a medallion architecture using PostgreSQL stored procedures. Apache Airflow orchestrates scraping, loading, and transformation tasks. The entire pipeline is containerized with Docker for portability and scalability and integrated with GitHub Actions for CI/CD.

Python(Pandas, Requests, Beautiful Soup) PostgreSQL Docker
View on GitHub
dbt and Snowflake ELT pipeline architecture

MovieLens Data Pipeline - ADLS Gen2 + dbt + Snowflake

This project showcases a modern ELT pipeline. It extracts the raw MovieLens 20M dataset from Azure Data Lake Storage Gen2, loads it into Snowflake data warehouse, and then dbt connects to Snowflake to perform data modeling and transformations. This process creates a dimensional model, with auto-generated documentation, optimized for analytics and ML applications.

Data Build Tool(dbt) Snowflake Azure Data Lake Storage Gen2 Git Github Actions Github Pages
View on GitHub
Azure weather streaming pipeline architecture

Weather Streaming

This project demonstrates a comprehensive real-time weather data streaming pipeline using Azure cloud services. The system ingests weather data from external APIs, processes it through Azure Event Hubs, and visualizes real-time insights through Power BI dashboards. The project showcases cost optimization strategies by providing both Databricks and Azure Functions implementations.

Databricks Azure Functions Azure Event Hub Azure Stream Analytics Power BI Azure Key Vault Azure Cost Management
View on GitHub
Azure enterprise-scale medallion architecture

Azure Data Engineering Project: Enterprise-Scale Medallion Architecture

This comprehensive Azure data engineering solution demonstrates the implementation of a production-ready Medallion Architecture (Bronze → Silver → Gold) pattern. The project leverages Microsoft Azure's native cloud services to create a scalable, maintainable, and cost-effective data pipeline that processes AdventureWorks business data from ingestion through to business intelligence reporting.

Pyspark Azure Data Factory Azure Data Lake Storage Databricks Azure Synapse Analytics Power BI
View on GitHub
Azure Databricks and Unity Catalog architecture

Azure-Databricks Project with Unity Catalog

This project uses Azure Data Factory to ingest data from GitHub into ADLS Gen2. It follows a medallion architecture: Databricks Autoloader streams raw data to the Bronze layer, Delta Lake tables clean it for the Silver layer, and the Gold layer provides aggregated data for analytics.

Azure Data Factory Azure Data Lake Databricks Git Delta Lake Pyspark
View on GitHub
Let's connect / 08

Have a data challenge in mind?

Let's connect and discuss how we can work together on your next data project.

Open to meaningful opportunities

Let's build something useful.

Whether you are designing a data platform, exploring a collaboration, or looking for an engineer who thinks in systems, I would love to hear from you.

Start a conversation