|
I'm a System Architect & Founder at Scaibu, based in Bengaluru, India. I design and engineer mission-critical distributed backends, high-throughput streaming topologies, production AI/ML pipelines with Vector Databases, and self-healing cloud infrastructure across Google Cloud, TensorFlow, and Terraform.
π’ Currently open for System Architect roles, Advisory & Consulting β Remote Worldwide! |
|
| Metric / Dimension | Target / Benchmark | Architectural Implementation |
|---|---|---|
| π Peak Ingestion & Scale | 1,000,000+ Requests in 3 mins | Asynchronous non-blocking I/O event loops, connection pooling, and multi-threaded stream workers. |
| β±οΈ Latency Budget (p99) | < 15ms | In-memory Redis caching layers, zero-copy serialization, and kernel-level socket optimizations. |
| π‘οΈ Reliability & SLO | 99.99% High Availability | 4-stage progressive canary deployment (5% β 25% β 50% β 100%) with automatic sub-5s rollbacks. |
| π Event Streaming Flow | 50k+ msgs / sec | Partition-aware Apache Kafka pipelines with idempotent consumer offsets and zero-data-loss guarantees. |
| π Vector Search Retrieval | < 20ms p95 | Hierarchical semantic chunking with HNSW indexed vector spaces across Pinecone, Qdrant & pgvector. |
| π¦ Modular Reusability | 90+ Composable Packages | Schema-driven anti-corruption adapters and generic data engines for instant plug-and-play reuse. |
- Weighted Traffic Shifting: Integrated Argo Rollouts and Traefik TrafficSplit CRDs to gradually promote new binaries across 4 structured soak phases (
5% β 25% β 50% β 100%). - Automated Rollback Safeguards: Prometheus metrics continuously evaluate p99 latency ceilings and HTTP 5xx error thresholds, automatically aborting unhealthy rollouts in under 5 seconds.
- High-Throughput Partitioning: Designed streaming pipelines to absorb sudden traffic spikes (such as flash sales or real-time telemetry) by sharding workloads across dynamically-rebalanced Kafka partitions.
- Distributed Concurrency Primitives: Engineered custom high-performance async mutex locking (
scaibu_mutex_lock) to eliminate race conditions without sacrificing throughput.
- Contract-Driven Anti-Corruption Layer: Universal
fromApi/toApitransform pipelines that isolate backend contract changes from UI and business domains. - Generic Adaptor & Saga Engines: Reusable CRUD adapters, Redux-Saga workers, and rules engines that eliminate hand-rolled repetitive logic across 90+ microservices.
I write production-depth technical articles on distributed systems, backend architecture, RAG, & AI β 300+ published on Medium.
| Article | Tag | Read | |
|---|---|---|---|
| π₯ | Why Replication Is One of the Hardest Problems in Distributed Systems | Distributed Systems |
20 min |
| π₯ | The Physics of Payment Systems: Why Exactly-Once Semantics Fail in Practice | Backend |
15 min |
| π₯ | From Minutes to Milliseconds: Docker Build Optimization | DevOps |
117 min |
| π₯ | Hierarchical Semantic Chunking | AI / RAG / Vectors |
15 min |
| π | Retry, Error Handling & Idempotency: The Hidden Science Behind Reliable Distributed Systems | Distributed Systems |
48 min |
| π | Stop Building Slow Systems: Master Advanced Queuing & Flow Control | Backend |
41 min |
| Project | Description | Stack |
|---|---|---|
| π llm-observability-platform | LLM monitoring & observability dashboard | Python TypeScript Vector DB |
| π ProcureIQ | AI-powered enterprise procurement platform | Next.js Python LangChain |
| π kafka-messaging-pipeline | High-throughput event-driven microservice pipeline | Node.js Kafka Docker |
| π a2a-demo | Google A2A Protocol demo with LangGraph & AI Agents | Python Google Cloud Gemini |
| π scaibu_mutex_lock | High-performance async concurrency lock for Dart/Flutter | Dart Flutter |
I'm available for the following opportunities:
β System Architect Β |Β β Distributed Systems & AI Consulting Β |Β β High-Scale Engineering Advisory Β |Β β Remote Worldwide
Deterministic Finite State Machine (FSM) governing progressive delivery phases, automated soak windows, Prometheus metric gates, and instant rollback paths.
stateDiagram-v2
[*] --> HealthyStable : Normal Operations (100% Stable)
HealthyStable --> RolloutInitiated : New Pod Template (Image Tag Bump)
state RolloutInitiated {
[*] --> CreatingCanaryRS
CreatingCanaryRS --> AwaitingProbes : Pods Scheduled & Started
AwaitingProbes --> CanaryReady : Readiness Probe Passed
}
RolloutInitiated --> Step1_Weight5 : Apply Step 1 (Weight = 5%)
state Step1_Weight5 {
[*] --> Timer120s_1
Timer120s_1 --> Analyzing1 : Scrape Prometheus Every 30s
Analyzing1 --> Step1_Passed : Error Rate < 0.5% & P99 < 250ms
}
Step1_Weight5 --> Step2_Weight25 : Step 1 Complete (Promote to 25%)
state Step2_Weight25 {
[*] --> Timer120s_2
Timer120s_2 --> Analyzing2 : Scrape Prometheus Every 30s
Analyzing2 --> Step2_Passed : Error Rate < 0.5% & P99 < 250ms
}
Step2_Weight25 --> Step3_Weight50 : Step 2 Complete (Promote to 50%)
state Step3_Weight50 {
[*] --> Timer120s_3
Timer120s_3 --> Analyzing3 : Scrape Prometheus Every 30s
Analyzing3 --> Step3_Passed : Parity Validated
}
Step3_Weight50 --> FullPromotion_Weight100 : Final Step Complete
state FullPromotion_Weight100 {
[*] --> CutoverTraffic : Set Weight = 100%
CutoverTraffic --> DrainOldStable : Wait terminationGracePeriod (30s)
DrainOldStable --> PromoteRS : Label Canary RS as New Stable
}
FullPromotion_Weight100 --> HealthyStable : Rollout Complete
%% Error & Abort Transitions
Step1_Weight5 --> Aborted : Analysis Failure OR Manual Abort
Step2_Weight25 --> Aborted : Analysis Failure OR Manual Abort
Step3_Weight50 --> Aborted : Analysis Failure OR Manual Abort
state Aborted {
[*] --> InstantTrafficZero : Reset TrafficSplit (Stable=100%, Canary=0%)
InstantTrafficZero --> TerminateCanary : Scale Canary RS to 0 Replicas
TerminateCanary --> PostIncidentAlert : Emit CloudEvent / Slack Alert
}
Aborted --> HealthyStable : Manual Retry or Rollback Spec
| Reference | Domain / Focus | Engineering Specification & Design | Architectural Strategy |
|---|---|---|---|
| Spec-0017 | Progressive Delivery | Canary Deployment & Progressive Delivery Architecture | 4-Stage Traffic Shift (5% β 25% β 50% β 100%) with automated sub-5s rollback |
| Spec-0021 | Cloud Autoscaling | Stateless Compute Autoscaling & Stateful Decoupling | Split-brain elimination with decoupled persistent data plane |
| Spec-0016 | Orchestration | Kubernetes Migration & CI/CD Pipeline Architecture | Container-native Kubernetes workload manifests with CSI storage binding |
| Spec-0014 | Ingress & Security | Traefik Edge Proxy Gateway & Centralized Logging | Edge TLS termination, rate-limiting, and middleware filter pipeline |
| Spec-0018 | Delivery Automation | CI/CD Pipeline Architecture & Validation Tiers | 5-Tier automated validation gates & GitOps deployment flow |
| Spec-0013 | Observability | OpenTelemetry Collector & Memory Protection | Bounded memory allocator with backpressure flow control |
| Spec-0010 | High Availability | Active-Passive Zero-Downtime Failover & Fallback | Automated health monitoring with active fallback triggers |
| Spec-0020 | Data Persistence | Persistent Storage Lifecycle & Data Protection | Immutable PVC host mounts with atomic backup pipelines |




