Expert - Data Engineering

Trung tâm Công nghệ Thông tin
Hồ Chí Minh
26-ITC-0566

Introduction


MoMo's data platform processes petabytes of data through thousands of pipelines. It powers product analytics, business performance analysis, insight validation, and machine learning models for fraud detection, credit scoring, and personalization. Its reliability directly affects business decisions and product performance.


We are moving from a platform built primarily on BigQuery toward a hybrid lakehouse across cloud and on-premises infrastructure. The shift addresses growing demands for data control, lower cost, and regulatory compliance. Running our own infrastructure brings new demands: storage must stay fast and durable as workloads grow, while capacity and performance problems must be caught early.


As our Expert System Engineer, you will own the storage and systems foundation of the lakehouse. You will build out and operate our new Ceph object storage cluster, defining how storage, networking, and operating systems work together as the platform grows. Your decisions will shape its performance, cost, and ability to recover from failure.


Mô tả công việc

What you will do

  • Set the storage architecture. Define the storage topology and its network and operating system requirements. Evaluate open source, commercial, cloud, and hybrid options against production workloads, weighing cost, performance, and operational effort.

  • Keep object storage dependable. Build and operate the Ceph cluster. Design pools and data placement, choose between replication and erasure coding, and balance usable capacity with durability.

  • Keep data moving. Trace slow reads and writes across compute, clients, network, and storage. Tune concurrency, object size, client configuration, and access patterns using measurements from production workloads.

  • Stay ahead of capacity limits. Forecast growth, set tiering and retention policies, and plan hardware expansion before capacity constrains workloads.

  • Recover from failure. Handle disk and node loss, degraded states, and rebalancing under load. Plan upgrades, test backup restores, and rehearse disaster recovery.

  • Strengthen the systems layer. Diagnose and tune Linux, filesystems, disks, and networking beneath the lakehouse. Work with network and infrastructure teams on topology, bandwidth, and hardware profiles.

  • Make operations repeatable and visible. Automate provisioning, configuration, patching, and releases with tools such as Pulumi, Helm, and CI/CD. Instrument the storage layer so the team can spot degradation, explain slow jobs, and act before workloads fail.

Yêu cầu công việc

Must have

  • 7+ years in systems, infrastructure, or storage engineering, including production systems you have designed, built, and operated at scale.

  • Production experience with storage or other stateful distributed systems. You can explain architecture choices, diagnose performance, plan capacity, and recover from failures.

  • Strong Linux systems knowledge. You can trace a bottleneck through the I/O path, filesystem, memory, and kernel, then verify that your change improved it.

  • Data center networking experience. You can trace a throughput problem across NICs, switches, and network topology.

  • Hands-on experience with bare metal and on-premises infrastructure, including hardware sizing, provisioning, and patching.

  • Clear technical writing. You can turn design decisions, benchmark results, and recovery procedures into documents other engineers can use.


We also look for strong potential in engineers who have built and operated complex infrastructure. You do not need hands-on experience with Ceph or the other storage systems listed above if you can show how you learn unfamiliar systems and solve production problems.

Nice to have

  • Hands-on experience with Ceph, MinIO, GlusterFS, HDFS, or a commercial object store.

  • Experience with Apache Iceberg or another open table format. You understand how small files, compaction, and partitioning affect storage behavior.

  • Experience with Apache Spark. You understand how its read and write patterns affect object storage performance.