Skip to content
View KyleZ8's full-sized avatar

Highlights

  • Pro

Block or report KyleZ8

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
KyleZ8/README.md

Kyle Zhu

Data Analyst — SQL · Python · Experimentation · Decision Analytics

MS in Business Analytics, Carlson School of Management, University of Minnesota Minneapolis, Minnesota · Open to Data Analyst roles in any industry

I turn messy data into decisions. Most analyses stop at a chart; I try to finish the sentence — what changed, why, what it's worth, and what we should do about it.

What I enjoy most is the part before the model: finding out that a spike was duplicate records, that a feature was only knowable after the fact, or that a metric everyone quotes was never defined the same way twice. Getting that right is usually worth more than a better algorithm, and it's the difference between an analysis people act on and one they quietly ignore.

How I work: validate before explaining · write the definition down · answer in decisions, not dashboards · state what would break the conclusion.


Featured projects

Project Question it answers Result
MetricGuard AI Why did the dispute rate spike? A third of a 32.8% spike was duplicate data; the real driver was travel on mobile
Retail Customer Analytics · live dashboard Where should the retention budget go? 335 slipping high-value customers, £2.0M of 12-month revenue at stake
Fair & Explainable Credit Risk Who should get credit, and is the model defensible? Cost-based threshold cut the approved-book default rate from 15.7% to 9.4%
Bank Marketing Targeting Which customers should the call center call? 87% of conversions from 60% of calls; 4,117 calls saved
Experimentation & Causal Impact When is the obvious read wrong? Paid-search ROI fell from 320% to 80% once organic substitution was removed
Restaurant Recommender Which recommender should ship? Matrix factorization for warm users, popularity fallback for cold start

Each repo runs from a clean clone, has automated tests and continuous integration, and states its own limitations.


Experience

Data Analytics Consultant — Carlson Analytics Lab

Oct 2025 – Aug 2026
C.H. Robinson (Fortune 500 logistics)

  • Built a KNN matching model estimating relationship health for the 54% of accounts that never respond to the customer survey, extending churn early-warning coverage to all 22K active accounts.
  • Designed a two-axis prioritization matrix combining revenue-weighted growth and comment sentiment, flagging ~$6B of current revenue as satisfied-but-shrinking for retention outreach.
  • Delivered a SQL-to-Python pipeline and dashboard giving VPs and marketing and operations stakeholders a standing retention-priority view; served as Scrum Master and Product Owner for a 5-person team.

Oct 2025 – Aug 2026
4Mativ (school-transportation technology)

  • Engineered a GPS data-quality layer over 2.35M pings and 10,800 trips, scoring connection health with per-route Isolation Forest models and isolating a device fault affecting 19% of one provider's trips against under 2% elsewhere.
  • Surfaced 1,191 detour and 248 wrong-route trips with DBSCAN corridor modeling, concentrating 64% of service failures in 3 of 16 vendors.
  • Built a vendor reliability matrix pairing GPS health against operational execution, separating fleets with broken tracking from fleets with real service failures so each got the right fix.

Oct 2025 – Aug 2026
Central Specialties (Midwest infrastructure)

  • Constructed the bid-level dataset behind a competitive-intelligence engagement — 2,736 bids, 629 projects, 419 rival firms — and defined 5 competitor KPIs across market presence, win share, success rate, pricing aggression, and winning margin.
  • Sized $15.6M of margin left on won projects and modeled a 3% price reduction that would raise win rate 43% across $223M of near-miss contract value.
  • Presented an interactive Tableau dashboard with county-level competitor mapping, letting the bidding team price against each rival's historical behavior before committing estimating resources.

Research Assistant, Pharmaceuticals — Kaiyuan Securities

Jun–Sep 2024

  • Extracted company financials through the Wind API and consolidated three vendor sources into 12 analysis-ready tables covering 100+ manufacturers, building the comparative valuation base for a new coverage segment.
  • Built top-down market-sizing models under multiple growth scenarios, producing the revenue forecasts behind the firm's first two published reports on the segment.

Compliance Analyst — Agricultural Bank of China

Jun–Sep 2023

  • Screened 200+ employee financial-disclosure and account records against internal compliance rules, reconciling data across systems to identify undisclosed holdings and conflicts of interest.
  • Coordinated across retail banking, compliance, and corporate relationship teams to onboard payroll accounts for 5 enterprise clients and 300+ employees, performing KYC verification for fraud and account-misuse risk.

🛠️ Tools

Languages: SQL · Python · R
Analysis: A/B testing · causal inference (DiD, propensity matching) · segmentation & RFM · cohort and retention analysis · classification · lift and ROI analysis · anomaly detection · model explainability and fairness auditing
Stack: pandas · DuckDB · scikit-learn · XGBoost · statsmodels · Tableau · Streamlit · Plotly · Git · GitHub Actions · pytest


📫 Contact

LinkedIn · kyle2601701@gmail.com

Pinned Loading

  1. bank-marketing-targeting bank-marketing-targeting Public

    Which customers should a bank call? Campaign analysis in SQL, a leakage-aware response model, lift/gains analysis, and campaign ROI (Python, DuckDB)

    Jupyter Notebook

  2. experimentation-causal-impact experimentation-causal-impact Public

    A/B testing and causal inference casebook: three business decisions where naive analysis is wrong, and the method that gets it right (Python, SQL, statsmodels)

    Jupyter Notebook

  3. fair-explainable-credit-risk fair-explainable-credit-risk Public

    Credit default model built for model risk review: WoE scorecard vs. gradient boosting, SHAP reason codes, fairness audit, and a cost-based approval threshold (Python)

    Jupyter Notebook

  4. metricguard-ai metricguard-ai Public

    KPI monitoring and root-cause analytics for a credit card portfolio: data quality validation, anomaly detection, segment driver analysis, and verified explanations (Python, DuckDB, Streamlit)

    Python

  5. restaurant-recommender restaurant-recommender Public

    Which recommender should a restaurant platform ship? Model comparison, segment-level and cold-start evaluation, popularity bias, and an A/B launch plan (Python, SQL)

    Jupyter Notebook

  6. retail-customer-analytics retail-customer-analytics Public

    Retail customer analytics in SQL and Python: data cleaning, RFM segmentation, cohort retention, CLV, and a revenue-at-risk targeting list with an interactive dashboard

    HTML