Lyft
|Software Engineer - Data Platform
Kyiv, Kyiv, Ukraine
Summary
Independently architected and managed Lyft’s Data Lifecycle management system for all offline storage, covering 300K tables, 900M partitions, billions of S3 objects, and hundreds of petabytes of data. Designed and implemented the next-generation lifecycle framework from scratch for the new Databricks platform. Simultaneously optimized retention, maintenance, and compaction mechanisms for legacy systems across the entire Data Platform.
Highlights
* Designed and implemented a self-service Data Lifecycle platform for Databricks to manage TTL retention and external Delta table maintenance across all Lyft offline tables, purging 10PB of redundant data daily. Projected Impact: This architecture enabled a shadow migration strategy that is forecast to save $4.5M, with ongoing automated maintenance expected to yield an additional $2.7M.
Identified and resolved architectural bottlenecks in an Iceberg Maintenance DAG, resulting in a $350,000 annual infrastructure cost reduction.
Engineered a custom cold data detection engine that delivered an immediate $500,000 one-time storage savings alongside ongoing annual reductions of at least $500,000.
Resolved a critical scalability bottleneck in the legacy offline table retention framework, increasing throughput by 30x, guaranteeing a strict 10-hour SLA, and generating $415,000 in annual storage savings.
Led Hive Metastore to AWS Glue data catalog migration efforts; implemented a user action ingestion pipeline that accelerated table user detection by 5x and eliminated metadata bottlenecks, yielding an 86x write and 100x read performance improvement.
Developed an end-to-end testing framework for table retention using Docker, reducing the feature-to-production feedback loop by 90% (from 15 days to under 2 days) while safely de-risking legacy system modernizations.
Designed and deployed a table-level cost attribution engine that scans S3 objects using custom heuristics, providing precise compute and storage cost visibility across the entire Data Platform to streamline cross-functional budgeting.