Senior - Data Engineer, Data Plaftform
Introduction
MoMo's data platform processes petabytes of data through thousands of pipelines. It powers product analytics, business performance deep dives, insight validation, and machine learning models for fraud detection, credit scoring, and personalization. Its reliability directly affects business decisions and product performance.
We are moving from a platform built primarily on BigQuery toward a hybrid lakehouse across cloud and on-premise infrastructure. The shift addresses growing demands for stronger data control, lower cost, and regulatory compliance. It also creates engineering problems at every layer: ingestion and replication across environments, data models that deploy to more than one platform, distributed storage and compute on Kubernetes and bare metal, and the separation of workloads by data sensitivity.
We are growing our Data Engineering team at middle and senior levels across these layers. You do not need experience in every area below. We are looking for engineers who have built or run production data systems and want to work on problems at this scale.
Where you could work
Data ingestion and modeling. Batch and streaming ingestion from databases, backend services, and Kafka, replication between BigQuery and the lakehouse, and layered data models in dbt and Spark that serve reporting, analytics, and machine learning.
Data infrastructure. Lakehouse storage on Apache Iceberg and object storage, compute engines such as Spark and StarRocks on Kubernetes, multi-tenancy, performance, and cost.
Data platform engineering. Reliability, observability, incident response, and the tooling that automates platform operations across cloud and on-premise.
Data governance. Access control, PII separation, data contracts, and the tools that keep the platform compliant.
Mô tả công việc
Build production pipelines. Design and run batch and streaming pipelines across cloud and on-premise, with incremental loads, safe re-runs, and backfills.
Model data that lasts. Design data models and storage layouts that hold up as data volume and business requirements grow.
Improve performance and cost. Tune Spark and other distributed engines, and fix inefficient jobs rather than adding more resources.
Improve reliability. Build monitoring, alerts, and runbooks, and lead root-cause analysis so the same incident does not happen twice.
Migrate production workloads. Move pipelines and models from BigQuery to the new lakehouse without breaking what is already in use.
Build compliant data boundaries. Help separate PII and PII-free workloads with appropriate access controls.
Automate recurring work. Build tools and frameworks that other engineers and teams use to ship data faster.
Yêu cầu công việc
Must have
2+ years of experience in Data Engineering, Data Platform Engineering, or a related backend or infrastructure role, with production systems you built or ran.
Strong SQL and hands-on experience with at least one of Python, Java, or Scala.
Hands-on experience with at least one distributed data technology such as Apache Spark, Apache Flink, Kafka, BigQuery, or an MPP engine.
Experience building and running scheduled or streaming pipelines in production, with Airflow or a similar orchestrator.
A solid grasp of data fundamentals: data modeling, partitioning, incremental loads, and why a pipeline can produce a wrong result.
Git, code review, and CI/CD as normal practice.
Nice to have
Experience with dbt or data quality and testing frameworks.
Experience with open table formats such as Apache Iceberg, Delta Lake, or Apache Hudi.
Experience with CDC and streaming ingestion tools such as Debezium or Kafka Connect.
Experience with Kubernetes, Docker, and infrastructure as code (Terraform, Pulumi, Helm).
Experience with monitoring and SRE practice, such as Prometheus, Grafana, SLOs, and on-call.
Experience with GCP, AWS, or hybrid cloud and on-premise environments.
Experience with payment, fintech, or other high-volume transaction data.
Experience with data governance, access control, or data privacy requirements.
In your application, tell us which of the areas above you would most like to work on.