I'm a Computer Science M.S. candidate at NJIT (graduating May 2026) who builds and deploys machine learning on large-scale clinical and financial data end to end from SQL feature engineering to AWS deployment - plus LLM / GraphRAG systems (LangChain, Neo4j, vector + graph retrieval). I care a lot about rigorous, leak-free evaluation (walk-forward validation, AUC-ROC, calibration) and reproducible, version-controlled workflows. I also enjoy translating technical results into something both clinical and non-technical stakeholders can actually act on.
Graduate Researcher @ NJIT β Clinical Data Analysis, PLCO Cancer-Screening Trial (Advisor: Prof. Arashdeep Kaur) Analyzing the large-scale PLCO cancer-screening dataset in Python β running EDA and statistical testing to rank candidate predictors, building preprocessing and feature-engineering pipelines (cut redundant variables ~30% via data-driven selection), and evaluating with AUC / F1 / Brier score. All under strict privacy and governance standards for confidential patient health and medical-imaging data, through a standardized, reproducible workflow.
-
GraphRAG Β· live demo β a knowledge-graph Q&A system over 16 peer-reviewed cancer-prediction papers (prostate, lung, colorectal, ovarian). Extracts entities and relationships into a Neo4j graph with Claude via LangChain's LLMGraphTransformer, then combines vector search (Gemini embeddings) with one-hop graph traversal to answer multi-hop questions plain RAG can't. Full-stack (FastAPI + React/TypeScript), deployed on Render/Vercel/Neo4j Aura. Every answer ships with the exact
Entity β RELATIONSHIP β Entityfacts behind it β explainable, not a black box. -
Lung Cancer Risk Prediction β a super-stacking ensemble (6 base learners + logistic-regression meta-learner) predicting lung cancer risk from 15 clinically interpretable PLCO variables. A data-efficiency study (100% β 10% training data, 5 seeds per fraction, metrics as mean Β± std) shows AUC-ROC holding near ~0.836 even at 10% training data, with threshold tuning on validation only to avoid test-set leakage.
-
VolGuard Β· live dashboard β an end-to-end ML pipeline forecasting 5-day realized stock volatility: raw prices β SQL feature engineering β walk-forward-validated model β AWS Lambda β live Tableau dashboard. Features engineered entirely in SQL with strict chronological splits (no look-ahead); shipped the simplest model (RMSE ~20% below baseline) to AWS via Terraform.
Languages β Python, SQL, PL/SQL, TypeScript ML & GenAI β scikit-learn, XGBoost, LightGBM, CatBoost, ensemble / stacking Β· ROC-AUC / model evaluation, feature selection Β· LangChain, RAG / GraphRAG, embeddings, Claude / Gemini APIs Data β Pandas, NumPy, Matplotlib Β· EDA, statistical analysis, cross-validation, hyperparameter tuning Databases β MySQL, PostgreSQL, MongoDB, SQLite, Neo4j (Cypher) Cloud & DevOps β AWS (Lambda, S3), Terraform, Docker, Vercel, Render Web β FastAPI, React, Vite, Tailwind CSS Tools β Git, GitHub, Jupyter, VS Code, Linux CLI, Tableau, Excel
- π§ Email: pvraj0019@gmail.com
- πΌ LinkedIn: linkedin.com/in/vraj-s-patel
Comfortable across the full ML lifecycle β and happiest when the pipeline is reproducible and the evaluation is honest.
