Senior - AI Engineer, Speech

Trung tâm Công nghệ Thông tin
Hồ Chí Minh
26-ITC-0268
If you love building big systems, solving hard technical problems, and creating products that millions of people use every day, MoMo is the place for you. We are looking for a Senior Software Engineer to lead the architecture and development of our next-generation Real-Time Voice AI Agent Platform, connecting conversational AI directly with telecommunication networks, contact centers, and core financial services.

Mô tả công việc

Job Description

1. Voice Agent Platform & Real-Time Communication (Core Focus)

  • Real-Time Architecture: Design and build a low-latency (< 1s) real-time audio/voice processing infrastructure based on the LiveKit Framework and WebRTC protocol.

  • Voice AI Pipeline Integration: Directly integrate and optimize the end-to-end Voice AI pipeline: STT (Whisper, Deepgram), LLM Context & Reasoning (OpenAI, Claude), TTS (Cartesia, ElevenLabs).

  • Telephony & Call Center Integration: Develop flexible connectivity between the Voice Agent and existing Call Center/PBX infrastructure via SIP Trunking, Twilio, and telecom service gateways.

2. Protocol Engineering & Session Management

  • Protocol Mastery: Master real-time transmission protocols (WebRTC, WebSocket, SIP, RTP/RTCP) to ensure stable and reliable audio signal streaming.

  • Session Optimization: Design and optimize session management mechanisms for hundreds of concurrent connections, handling reconnections efficiently, maintaining state, and recovering sessions during network disruptions.

  • Function Calling & Integrations: Design APIs/Microservices that connect the Voice Agent with core data systems (Core Banking, CRM, Database, Tool lookup) in real time during live calls.

  • Performance Tuning: Identify and resolve latency bottlenecks, optimizing media streaming flows and system concurrency.

3. Technical Leadership & System Quality

  • Code Quality & Architecture: Establish clean code standards, conduct rigorous code reviews, and propose robust distributed system architectures.

  • Observability & Monitoring: Build a specialized monitoring system for Voice AI (tracking Audio Latency, Packet Loss, Token usage, STT/TTS Drift, and Call drop rates).

  • Mentorship: Guide and mentor junior engineers on the team while actively sharing expertise in WebRTC and AI Systems


Yêu cầu công việc

  • Qualifications

    • Experience: 4+ years of experience in Backend Engineering / System Architecture; at least 1–2 years of hands-on experience directly working with WebRTC, LiveKit, or real-time Voice/Media Streaming systems.

    • Engineering Mindset: Distributed systems design mindset, ability to master open-source codebases, and a product-driven mindset prioritizing the end-user experience.

    Must-Have

    • Proficiency: Highly proficient in at least one of the following languages: Go, C++, Java, or Python (AsyncIO).

    • Media & Real-Time Protocols: Deep understanding of WebRTC, WebSocket, RTP/RTCP, gRPC, and audio stream processing techniques (PCM, Opus).

    • AI/LLM Ecosystem: Practical experience working with LLM Frameworks, Prompt Engineering, Function Calling/Tool Use, and integrating STT/TTS APIs.

    • Infrastructure & Database: Proficient with Docker, Kubernetes, Caching (Redis), Messaging Queues (Kafka/PubSub), and databases such as PostgreSQL and MongoDB.

    Tech Stack

    • Languages: Go, Java/Kotlin, Python (AsyncIO)

    • Real-Time & Telecom: LiveKit, WebRTC, SIP, Twilio

    • AI Stack: OpenAI APIs, Anthropic, Deepgram, ElevenLabs, Cartesia, LangChain/LlamaIndex

    • Data & Messaging: Kafka, Redis, PostgreSQL, MongoDB

    • DevOps & Infrastructure: Kubernetes, Docker, Helm, Prometheus, Grafana, AWS/GCP