SaraZarei · GitHub
Skip to content
View SaraZarei's full-sized avatar

Block or report SaraZarei

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SaraZarei/README.md

👋 Hi, I'm Sara Zarei

🔍 About Me

I’m passionate about building reliable data systems that turn raw, scattered data into clean, trusted, and analytics-ready datasets. My main focus is Data Engineering and Analytics Engineering, with a strong interest in data quality, pipeline reliability, scalable data models, and KPI-ready reporting layers.

I also have knowledge of Data Analytics, Data Science, and Machine Learning, including exploratory analysis, feature preparation, classification, regression, clustering, and preparing model-ready datasets.


💼 What I Do

  • Build end-to-end ETL/ELT pipelines
  • Design clean and reliable data models
  • Create staging, intermediate, canonical, and mart layers
  • Develop KPI-ready tables for dashboards and reporting
  • Perform data quality checks, validation, deduplication, and reconciliation
  • Work with structured, semi-structured, and API-based data
  • Transform raw data into trusted datasets for analytics and decision-making
  • Prepare clean datasets for analytics, reporting, and machine learning workflows.

🛠️ Technologies & Tools

  • Programming: Python, SQL
  • Databases & Warehouses: PostgreSQL, BigQuery, MySQL, MongoDB
  • Data Engineering: ETL/ELT, REST APIs, incremental loads, data validation, data modeling
  • Orchestration & Cloud: Apache Airflow, AWS Lambda, EventBridge, CloudWatch, S3
  • Analytics & BI: Power BI, Domo, KPI reporting, dashboard-ready marts
  • Data Science & ML: pandas, exploratory data analysis, feature preparation, classification, regression, clustering

📫 Let's Connect


"Turning raw data into robust systems and actionable insights is what excites me the most!"

Pinned Loading

  1. bookstore-normalized-data-pipeline bookstore-normalized-data-pipeline Public

    Bookstore relational data model — raw Excel data → 3NF PostgreSQL schema via a staging-to-production pipeline with automated PK/FK validation, rejected-record logging, and Python-based data loading.

    Python 1

  2. cross-brand-marketing-lakehouse-aws-bigquery-dbt-airflow cross-brand-marketing-lakehouse-aws-bigquery-dbt-airflow Public

    Production-grade ELT Lakehouse unifying GA4, Meta Ads, Google Ads, and CMS data across 4 brands into BigQuery — Python extraction via AWS Lambda + EventBridge, dbt transformation (47 models, 112 te…

    Python 1

  3. ecommerce-orders-canonical-dbt-bigquery-pipeline ecommerce-orders-canonical-dbt-bigquery-pipeline Public

    Multi-source ecommerce pipeline: Amazon + Shopify → BigQuery canonical model + KPI marts, orchestrated with Airflow and transformed with dbt

    Python 1

  4. news-airflow-mongodb-etl news-airflow-mongodb-etl Public

    Daily news ETL pipeline: NewsAPI → Apache Airflow (3-task DAG with XCom) → MongoDB. Includes transformation layer with title cleaning, deduplication, and encoding validation, plus full unit test co…

    Python 1

  5. openlibrary-mongodb-etl-pipeline openlibrary-mongodb-etl-pipeline Public

    End-to-end ETL pipeline: Open Library REST API → pre-load DQ checks (schema validation, flatten, nulls, duplicates) → MongoDB ingestion → post-load aggregation queries. Python · MongoDB · pymongo ·…

    Python 1