Gaurav Gurjar — Senior Data Engineer & AI Data Platform Architect
Gaurav Gurjar

Senior data engineer & AI data platform architect · Dubai

Reliable data
and governed AI
systems.

From architecture through production. I help teams modernize data platforms, automate high-value workflows, and build policy-aware AI systems that remain reliable in production.

Discuss your data challenge →

300+production pipelines built and operated

2M+people reached by a public-health application

7+ yearsacross data engineering and AI

300+

production pipelines built and operated across regulated, geospatial, API and PDF sources

2M+

people reached by a public-health risk application with serverless data services

7+ years

across data engineering and AI, insurance, public health, capital markets, government research

Make complex systems dependable.

Three capabilities, one standard: what gets built stays governable, testable, and operable after handoff.

01 Data platforms

Ingestion to warehouse, with discipline.

Cloud ingestion, dimensional modeling, data quality, incremental processing, and production operations. Tested fact and dimension models, not just pipelines that run.

Python · SQL · AWS · Redshift · dbt · PySpark · Kafka
02 Governed AI systems

Policy before orchestration.

Knowledge ingestion, policy processing, grounding, routing, provenance, PII controls, and runtime telemetry. The system's decisions live in one place and its actions in another.

FastAPI · Dagster · Vector stores · Grounding · Provenance
03 Delivery leadership

Architecture through implementation.

Stakeholder communication, testing, release discipline, and team guidance. Guidance of junior engineers where the platform demands it.

Architecture · Reviews · Mentorship · Release discipline

Regulated-industry analytics platform

Regulated analytics300+ production pipelines

Cloud ingestion and transformation across regulated-business, geospatial, API, and PDF sources.

Runs as the governed ingestion standard behind 300+ production pipelines — auditable, reproducible, and no longer dependent on one-off scripts.

View case study →

Public-health risk data services

Public health2M+ people reached

Statistical components, pipelines, and backend data services for a public risk application.

Backs a public risk application serving 2M+ people, with serverless services that hold up under real public-load spikes.

View case study →

Insurance analytics data platform

InsuranceDaily and monthly incremental snapshots

Tested warehouse models, compliance datasets, analytical marts, and scheduled snapshots.

Gives compliance and actuarial teams reconciled, trusted marts on a fixed daily and monthly cadence instead of manual extracts.

View case study →

More in the ledger

Capital markets · Veterinary · Government research

Batch and streaming signals in capital markets, validated ETL across 46 veterinary clinics, and a 2 to 20 PB genetics-data blueprint for precision-medicine research.

Each with the same standard: tested, documented, handed off operable.

View all case studies →

Maintained projects

Independent work, shipped and maintained.

Independent work

Maintained projects live outside client engagements. Same standard, public by default.

Setu

A Rust data-activation engine that reacts to PostgreSQL change streams and delivers matching events to webhooks, Slack, or Telegram without polling or middleware.

Rust · PostgreSQL · Webhooks

Rental Market Dynamics Dubai

An automated pipeline for extracting Dubai rent-contract data, transforming it to Parquet, and publishing analysis-ready releases and property-usage reports.

Python · Parquet · Automated releases

Hermes Google Sheets

A maintained Hermes Agent plugin for reading, searching, and updating Google Sheets through focused spreadsheet tools.

Hermes · Google Sheets API

Open-source contributions

PUDL and sportsdataverse-py: test modernization, AssetSpec migration, cache handling, and analysis examples. Reviewed upstream.

PUDL · sportsdataverse-py

Trusted in the work.

Three voices, not mine. Each one names a different thing: technical range, collegiality, delivery under difficulty.

“I have worked with Gaurav for close to a year and he is a very well rounded, skilled, and innovative data scientist. He has helped me in the development of statistical methods, backend server infrastructure, and data science tasks. Gaurav is not only a great developer but also a great communicator. He has always been very prompt, responsive, and completes tasks on time. He goes above and beyond to ensure that the customer requirements and needs are met. I recommend him to anyone seeking expert level data science services.”
Benjamin Harvey, Ph.D. · Founder of AI Squared
“Gaurav is a thoughtful person with a very creative mind. He is intellectually curious and looks for efficient solutions to any problems. I enjoyed my time working with him and appreciate the collegial relationship we developed.”
Ivette Basterrechea · Department of Justice
“I worked with Gaurav on several projects, he has strong technical skills and is also a good team player. He was able to deliver high-quality work and found solutions to difficult problems.”
Le Zhang · Google

Field notes

View all insights →

Technical depth, delivery focus.

Dubai-based senior data engineer and AI data platform architect with 7+ years across governed AI, insurance, public health, regulated analytics, veterinary healthcare, capital markets, and government research.

UAE Golden Visa holder, available for remote global consulting engagements.

Core technology: Python · SQL · AWS · Redshift · dbt · PySpark · Kafka · FastAPI · Dagster

Start a conversation

Have a data challenge that has outgrown quick fixes?

Discuss your data challenge →

15 minutes, no pitch deck. Bring the messy source.

Gaurav Gurjar · Senior data engineering and governed AI consulting, from architecture through production.

© 2026 Gaurav Gurjar · Built on paper, shipped on the web.