Google (via Vaco Binary Semantics)
|Software Engineer
Gurugram, Haryana, India
Summary
Spearheaded the development and optimization of high-performance data pipelines and real-time streaming solutions for critical business and blockchain data across Google Cloud Platform.
Highlights
Optimized a modular batch pipeline for Amazon SP API reports, processing over 10K reports daily and reducing end-to-end latency by 70% through parallelized Pub/Sub and autoscaling Dataflow jobs.
Converted complex transformation logic into reusable Dataflow Flex Templates with config-driven schema mapping and parameterization, enabling rapid onboarding of new sellers and marketplaces without code changes.
Designed and implemented a low-latency real-time pipeline to ingest, transform, and load 10M+ Ethereum transactions daily into Google's proprietary graph database, leveraging Kafka and PySpark Structured Streaming.
Reduced data propagation latency by 65% in blockchain streaming through optimized micro-batch intervals, checkpoint tuning, partitioning, and parallelized graph API ingestion.
Enhanced a global AQI data pipeline, increasing daily data coverage by 60% to over 20M records across 120+ countries, by integrating enrichment, geospatial tagging, and deduplication components.
Improved pipeline uptime to 99.8% and reduced ingestion failures by 90% in the AQI pipeline by implementing robust validation rules and schema drift handling.