I am a data engineer and machine learning practitioner with 10+ years of experience building large-scale, production-grade data pipelines.
I hold a PhD in astrophysics, where I worked with terabyte-scale data from gamma-ray observatories including VERITAS and HAWC.
I have since applied that same rigor to industry-focused data problems. I design end-to-end pipelines (Airflow, HPC/Slurm, Docker), build ML models for anomaly detection and classification on high-dimensional, noisy datasets, and develop LLM/RAG systems for semantic search and retrieval.
I am seeking Data engineering / Analytics engineering roles in industry.
Over the course of my career, I have developed strong expertise in:
- Data Engineering — End-to-end pipeline design, Data modeling, workflow automation (Airflow), distributed processing on HPC/Slurm, containerization (Docker)
- Machine Learning — Anomaly detection, classification, and clustering on high-dimensional, imbalanced data
- LLM & RAG Pipelines — Semantic search, dense embeddings, retrieval systems for large-scale text corpora
- Data Infrastructure — Relational databases (PostgreSQL, MySQL), vector databases (ChromaDB), Git-based workflows
- Python Stack — NumPy, SciPy, pandas, scikit-learn, PyTorch
- Statistical Modeling — Maximum likelihood estimation, Bayesian inference, hypothesis testing
- Visualization — Matplotlib, Plotly, Streamlit dashboards, Power BI
- Leadership & Communication — Led cross-institutional research teams; 50+ peer-reviewed publications, 20+ conference talks