I build and operate data platforms, connecting ingestion and lakehouse modeling with business metrics, application data, and AI-assisted analytics.
My work spans system design, implementation, production operations, and the analysis that makes the data useful.
Website & writing ↗ All repositories ↗
Build and operate a layered lakehouse with Dagster, dbt, Spark, and Iceberg, turning application and payment data into reusable analytical models.
The work spans incremental processing, partitioned backfills, data quality checks, and production delivery on Kubernetes through GitOps.
DAGSTER / DBT / SPARK / ICEBERG / KUBERNETES
Develop PostgreSQL CDC ingestion and publish warehouse-derived user attributes to application-facing data stores.
Handle snapshot/change ordering, deletion semantics, and event deduplication. Deliver attributes to DynamoDB through validated imports and versioned table cutovers, retaining the previous version for recovery.
POSTGRESQL / AWS DMS / KINESIS / DYNAMODB
Define subscription, retention, and revenue metrics across billing sources, reconciling currency units, transaction timing, and payment states.
Investigate churn through cohort and behavioral analysis, with complete observation windows and explicit leakage checks. Deliver traceable datasets, dashboards, and written findings for product, finance, and audit work.
SQL / DATA MODELING / RECONCILIATION / COHORT ANALYSIS
Deploy and extend Superset for shared analytics, with MCP access tied to individual user identities and existing role-based permissions.
Implement credential issuance, expiry, and revocation, and validate identity isolation across concurrent requests. Agent access follows the same Superset permissions as interactive use.
SUPERSET / MCP / RBAC / PYTHON / ARGO CD
| Area | What I work with |
|---|---|
| Data platforms | Python, SQL, Dagster, dbt, Spark, Iceberg · orchestration, incremental models, data quality |
| Ingestion & serving | PostgreSQL, DMS, Kinesis, DynamoDB · change capture, state reconstruction, application data delivery |
| Analytics & AI | Superset, MCP, cohort analysis · metric definitions, reconciliation, governed data access |
| Platform operations | Kubernetes, Helm, Argo CD, GitHub Actions · GitOps, deployment checks, environment isolation |
I write about data systems and share what I learn while building. Explore the notes and experiments on my personal site.
Based in Shanghai · Building across data, AI, and the interface between them.


