Senior - AI Engineer, Speech
Mô tả công việc
Job Description
1. Voice Agent Platform & Real-Time Communication (Core Focus)
Real-Time Architecture: Design and build a low-latency (< 1s) real-time audio/voice processing infrastructure based on the LiveKit Framework and WebRTC protocol.
Voice AI Pipeline Integration: Directly integrate and optimize the end-to-end Voice AI pipeline: STT (Whisper, Deepgram), LLM Context & Reasoning (OpenAI, Claude), TTS (Cartesia, ElevenLabs).
Telephony & Call Center Integration: Develop flexible connectivity between the Voice Agent and existing Call Center/PBX infrastructure via SIP Trunking, Twilio, and telecom service gateways.
2. Protocol Engineering & Session Management
Protocol Mastery: Master real-time transmission protocols (WebRTC, WebSocket, SIP, RTP/RTCP) to ensure stable and reliable audio signal streaming.
Session Optimization: Design and optimize session management mechanisms for hundreds of concurrent connections, handling reconnections efficiently, maintaining state, and recovering sessions during network disruptions.
Function Calling & Integrations: Design APIs/Microservices that connect the Voice Agent with core data systems (Core Banking, CRM, Database, Tool lookup) in real time during live calls.
Performance Tuning: Identify and resolve latency bottlenecks, optimizing media streaming flows and system concurrency.
3. Technical Leadership & System Quality
Code Quality & Architecture: Establish clean code standards, conduct rigorous code reviews, and propose robust distributed system architectures.
Observability & Monitoring: Build a specialized monitoring system for Voice AI (tracking Audio Latency, Packet Loss, Token usage, STT/TTS Drift, and Call drop rates).
Mentorship: Guide and mentor junior engineers on the team while actively sharing expertise in WebRTC and AI Systems
Yêu cầu công việc
Qualifications
Experience: 4+ years of experience in Backend Engineering / System Architecture; at least 1–2 years of hands-on experience directly working with WebRTC, LiveKit, or real-time Voice/Media Streaming systems.
Engineering Mindset: Distributed systems design mindset, ability to master open-source codebases, and a product-driven mindset prioritizing the end-user experience.
Must-Have
Proficiency: Highly proficient in at least one of the following languages: Go, C++, Java, or Python (AsyncIO).
Media & Real-Time Protocols: Deep understanding of WebRTC, WebSocket, RTP/RTCP, gRPC, and audio stream processing techniques (PCM, Opus).
AI/LLM Ecosystem: Practical experience working with LLM Frameworks, Prompt Engineering, Function Calling/Tool Use, and integrating STT/TTS APIs.
Infrastructure & Database: Proficient with Docker, Kubernetes, Caching (Redis), Messaging Queues (Kafka/PubSub), and databases such as PostgreSQL and MongoDB.
Tech Stack
Languages: Go, Java/Kotlin, Python (AsyncIO)
Real-Time & Telecom: LiveKit, WebRTC, SIP, Twilio
AI Stack: OpenAI APIs, Anthropic, Deepgram, ElevenLabs, Cartesia, LangChain/LlamaIndex
Data & Messaging: Kafka, Redis, PostgreSQL, MongoDB
DevOps & Infrastructure: Kubernetes, Docker, Helm, Prometheus, Grafana, AWS/GCP