I'm a data scientist who likes models with receipts: a clear baseline, a fair evaluation, and an explanation someone can challenge. I have an M.S. in Bioinformatics and enjoy applying machine learning to different kinds of real-world data. I'm open to Data Scientist and Applied ML roles across industries.
| Project | What you'll find | |
|---|---|---|
| 🏦 | Banking Product Propensity Modeling | Transaction-sequence transformer vs. gradient boosting on anonymized banking data, with client-disjoint evaluation and a local Streamlit dashboard. Results and limitations. |
| 🚚 | Supplier Risk Intelligence | Supplier performance + disruption-news NLP → calibrated, explainable risk ranking by expected loss. Synthetic demo benchmarks, SHAP, and Streamlit. |
| 🧪 | Clinical Safety Intelligence | FAERS safety-signal ranking with ROR/PRR, followed by quality-gated FDA, PubMed, and trial evidence synthesis using LangGraph and Gemini. |
| 💬 | Text2SQL Studio | Natural-language questions → read-only SQL, query validation, and an evaluation workflow. |
- Data + modeling:
Python·SQL·Pandas·NumPy·scikit-learn·XGBoost·PyTorch - NLP + LLM workflows:
Hugging Face Transformers·LangGraph·Gemini API - Apps + evaluation:
Streamlit·Plotly·SHAP·Pytest·Git
Curious about a project? Start with its README—or say hello above.