Skip to content
View Vraj2712's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report Vraj2712

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
vraj2712/README.md

Hi, I'm Vraj πŸ‘‹

Typing SVG

πŸ§‘β€πŸ’» Who am I?

I'm a Computer Science M.S. candidate at NJIT (graduating May 2026) who builds and deploys machine learning on large-scale clinical and financial data end to end from SQL feature engineering to AWS deployment - plus LLM / GraphRAG systems (LangChain, Neo4j, vector + graph retrieval). I care a lot about rigorous, leak-free evaluation (walk-forward validation, AUC-ROC, calibration) and reproducible, version-controlled workflows. I also enjoy translating technical results into something both clinical and non-technical stakeholders can actually act on.

πŸ”¬ Research

Graduate Researcher @ NJIT β€” Clinical Data Analysis, PLCO Cancer-Screening Trial (Advisor: Prof. Arashdeep Kaur) Analyzing the large-scale PLCO cancer-screening dataset in Python β€” running EDA and statistical testing to rank candidate predictors, building preprocessing and feature-engineering pipelines (cut redundant variables ~30% via data-driven selection), and evaluating with AUC / F1 / Brier score. All under strict privacy and governance standards for confidential patient health and medical-imaging data, through a standardized, reproducible workflow.

πŸš€ Projects

  • GraphRAG Β· live demo β€” a knowledge-graph Q&A system over 16 peer-reviewed cancer-prediction papers (prostate, lung, colorectal, ovarian). Extracts entities and relationships into a Neo4j graph with Claude via LangChain's LLMGraphTransformer, then combines vector search (Gemini embeddings) with one-hop graph traversal to answer multi-hop questions plain RAG can't. Full-stack (FastAPI + React/TypeScript), deployed on Render/Vercel/Neo4j Aura. Every answer ships with the exact Entity β†’ RELATIONSHIP β†’ Entity facts behind it β€” explainable, not a black box.

  • Lung Cancer Risk Prediction β€” a super-stacking ensemble (6 base learners + logistic-regression meta-learner) predicting lung cancer risk from 15 clinically interpretable PLCO variables. A data-efficiency study (100% β†’ 10% training data, 5 seeds per fraction, metrics as mean Β± std) shows AUC-ROC holding near ~0.836 even at 10% training data, with threshold tuning on validation only to avoid test-set leakage.

  • VolGuard Β· live dashboard β€” an end-to-end ML pipeline forecasting 5-day realized stock volatility: raw prices β†’ SQL feature engineering β†’ walk-forward-validated model β†’ AWS Lambda β†’ live Tableau dashboard. Features engineered entirely in SQL with strict chronological splits (no look-ahead); shipped the simplest model (RMSE ~20% below baseline) to AWS via Terraform.

πŸ› οΈ What tools do I use?

Languages β€” Python, SQL, PL/SQL, TypeScript ML & GenAI β€” scikit-learn, XGBoost, LightGBM, CatBoost, ensemble / stacking Β· ROC-AUC / model evaluation, feature selection Β· LangChain, RAG / GraphRAG, embeddings, Claude / Gemini APIs Data β€” Pandas, NumPy, Matplotlib Β· EDA, statistical analysis, cross-validation, hyperparameter tuning Databases β€” MySQL, PostgreSQL, MongoDB, SQLite, Neo4j (Cypher) Cloud & DevOps β€” AWS (Lambda, S3), Terraform, Docker, Vercel, Render Web β€” FastAPI, React, Vite, Tailwind CSS Tools β€” Git, GitHub, Jupyter, VS Code, Linux CLI, Tableau, Excel

πŸ“« How to reach me

Comfortable across the full ML lifecycle β€” and happiest when the pipeline is reproducible and the evaluation is honest.

Pinned Loading

  1. graphrag-neo4j graphrag-neo4j Public

    GraphRAG knowledge-graph Q&A over 16 cancer-prediction research papers β€” Neo4j + Claude + Gemini embeddings, FastAPI backend, React frontend. Live demo included.

    TypeScript

  2. VolGuard VolGuard Public

    End-to-end ML pipeline forecasting stock volatility: SQL features, walk-forward-validated model, AWS Lambda deployment via Terraform, and a live Tableau dashboard.

    Python

  3. lung-cancer-risk-ml-PLCO lung-cancer-risk-ml-PLCO Public

    Early lung cancer risk prediction using machine-learning models trained on the PLCO dataset. End-to-end pipeline including data cleaning, feature preprocessing, multiple ML baselines, super-stackin…

    Jupyter Notebook

  4. data-efficient-lung-cancer-prediction data-efficient-lung-cancer-prediction Public

    Evaluates data efficiency in lung cancer risk prediction using a super-stacking ensemble. Trains models on progressively reduced fractions of the PLCO dataset while keeping a fixed test set, analyz…

    Jupyter Notebook

  5. shlok695/Repopilot shlok695/Repopilot Public

    AI-assisted repo scanner that generates README docs, vulnerability reports, bug findings, and developer onboarding summaries.

    TypeScript