I build data platforms end to end — streaming ingestion, distributed processing, orchestration, and the AI layer on top. Most of my work lives at the point where data engineering meets applied AI: real-time pipelines feeding analytics, and retrieval systems that make that data useful to people and agents.
Everything below is a system I designed, built, and ran — not a tutorial follow-along.
Real-time e-commerce analytics on a lambda architecture. Kafka feeds four independent Spark Structured Streaming jobs into PostgreSQL and a MinIO data lake, reconciled nightly by Airflow. 27 Docker services · 36 automated tests · 9 bugs caught in live validation View repository → |
End-to-end recruitment intelligence platform. Aggregates job offers via APIs and scrapers, extracts skills from CVs with NLP, and ranks candidate matches using embedding similarity. Medallion architecture · Airflow DAGs · FastAPI + Power BI surface View repository → |
Retrieval-augmented search over an enterprise knowledge base. Returns grounded answers with inline citations, plus a knowledge-base manager and admin analytics for retrieval quality. Cited answers · hybrid retrieval · admin analytics View repository → |
An MCP server that lets AI agents understand a dataset without reading it. Returns a bounded JSON profile — types, ranges, null rates, quality flags — instead of pasted rows. Published on PyPI · 645 MB → 13 KB profile · CSV, Parquet, JSON, Excel View repository → |
|
WhatBreaks Static breaking-change analysis for dbt. Column-level blast radius in CI, with no warehouse or credentials required. Python SQLGlot dbt
|
Doc Doctor A GitHub Action that executes the code examples in your docs and fails the PR when they break. TypeScript GitHub Actions
|
Procurement Pipeline Big-data procurement analytics pipeline built on a Hadoop and Presto stack, orchestrated with Airflow. Hadoop Presto Airflow
|
| Languages |
|
| Data Engineering |
|
| AI & ML |
|
| Storage |
|
| Infrastructure |
|
| Observability |
|




