AI Engineer II (Speech)
Research, develop, and improve machine learning models for speech and audio understanding, including speech recognition, speech synthesis, voice activity detection, speaker modeling, speech enhancement, and spoken-language understanding.
Explore advanced modeling approaches for robust and natural speech processing across diverse accents, speaking styles, acoustic environments, and application domains.
Design rigorous experiments, datasets, and evaluation frameworks to measure model accuracy, robustness, generalization, latency, and perceptual quality.
Collaborate with research, product, and engineering teams to translate emerging speech technologies into practical, high-impact applications.
Mô tả công việc
Research and develop models for ASR, TTS, VAD, speaker recognition, speech enhancement, and related speech-processing tasks.
Train, fine-tune, and adapt speech models using internal and public datasets, with a focus on Vietnamese speech, regional accents, code switching, and noisy audio.
Explore modern architectures such as Transformers, Conformers, self-supervised speech models, speech-language models, diffusion models, and neural audio codecs.
Build reproducible pipelines for data preparation, model training, evaluation, ablation studies, and error analysis.
Design benchmarks and evaluate models using metrics such as WER, CER, semantic accuracy, perceptual quality, speaker similarity, robustness, and latency.
Investigate transfer learning, semi-supervised learning, distillation, quantization, and parameter-efficient fine-tuning.
Review and reproduce recent research, propose modeling improvements, and communicate findings through technical reports and presentations.
Collaborate with engineering and product teams to validate and integrate research outcomes.
Yêu cầu công việc
Bachelor’s degree in Computer Science, Engineering, Mathematics, Data Science, or a related field.
Strong Python skills and hands-on experience with PyTorch, TensorFlow, JAX, or similar frameworks.
Solid understanding of machine learning, deep learning, sequence modeling, and Transformer-based architectures.
Experience in at least one speech or audio area, such as ASR, TTS, VAD, speaker recognition, diarization, speech enhancement, or audio classification.
Familiarity with speech-processing tools such as torchaudio, librosa, Kaldi, ESPnet, NeMo, SpeechBrain, Whisper, or WeNet.
Experience preparing speech datasets and conducting controlled experiments, ablation studies, and model-error analysis.
Understanding of common speech evaluation metrics, including WER, CER, perceptual quality, and speaker similarity.
Ability to read, reproduce, and extend recent research papers.
Strong analytical, problem-solving, and cross-functional collaboration skills.