Data Engineer · Data Platforms · Lakehouse · Analytics Engineering
I build reliable and maintainable data systems — from ingestion and orchestration to lakehouse architecture, analytics-ready models, observability, and AI-powered data products.
Currently working as a Data Engineer at Timo Digital Bank, where I build data pipelines and analytics solutions using AWS, Airflow, dbt, Lightdash, and modern data platform technologies.
- Data Engineer at Timo Digital Bank by BVBank
- B.Sc. in Data Science, University of Science, VNU-HCM
- Interested in Data Platform Engineering, Lakehouse Architecture, Cloud Infrastructure, and Analytics Engineering
- Exploring the intersection of Data Engineering and AI, including MCP, conversational analytics, and AI-assisted research systems
- I enjoy building systems that are reliable, observable, maintainable, and easy to extend
Data Engineering Python · SQL · PySpark · Airflow · dbt · PyIceberg · Polars · Dagster
Lakehouse & Storage Apache Iceberg · AWS Glue Data Catalog · Amazon S3 · PostgreSQL · MariaDB · MongoDB
Cloud & Infrastructure AWS · Amazon EMR · AWS Glue · Amazon Athena · Terraform · Docker · OCI · Cloudflare Zero Trust
Analytics & Observability Lightdash · Streamlit · SigNoz
AI & Data Applications Model Context Protocol (MCP) · FastMCP · PydanticAI · Semantic Layers
Built a production-oriented AWS lakehouse using Apache Iceberg, Spark, dbt, Airflow, and AWS Glue Data Catalog. Infrastructure is managed with Terraform across S3, IAM, KMS, ECR, and EMR Serverless, while Airflow orchestrates Spark, dbt, and OCR workloads. The platform also includes GitHub activity analytics and is deployed behind Cloudflare Zero Trust.
Built a Vietnamese web research and content pipeline that uses LLM-generated queries across multiple search providers, then deduplicates, reranks, and filters retrieved sources before generating source-backed content. The application uses PydanticAI for typed outputs and validation, with both CLI and Streamlit interfaces.
Focused on data platform engineering, cloud-native infrastructure, Apache Iceberg, distributed data processing, observability, Infrastructure as Code, and AI-native data tooling.

