AI Engineer II (Speech)

Trung tâm Công nghệ Thông tin
Hồ Chí Minh
26-ITC-0155
Position Overivew

  • Research, develop, and improve machine learning models for speech and audio understanding, including speech recognition, speech synthesis, voice activity detection, speaker modeling, speech enhancement, and spoken-language understanding.

  • Explore advanced modeling approaches for robust and natural speech processing across diverse accents, speaking styles, acoustic environments, and application domains.

  • Design rigorous experiments, datasets, and evaluation frameworks to measure model accuracy, robustness, generalization, latency, and perceptual quality.

  • Collaborate with research, product, and engineering teams to translate emerging speech technologies into practical, high-impact applications.

Mô tả công việc

  • Research and develop models for ASR, TTS, VAD, speaker recognition, speech enhancement, and related speech-processing tasks.

  • Train, fine-tune, and adapt speech models using internal and public datasets, with a focus on Vietnamese speech, regional accents, code switching, and noisy audio.

  • Explore modern architectures such as Transformers, Conformers, self-supervised speech models, speech-language models, diffusion models, and neural audio codecs.

  • Build reproducible pipelines for data preparation, model training, evaluation, ablation studies, and error analysis.

  • Design benchmarks and evaluate models using metrics such as WER, CER, semantic accuracy, perceptual quality, speaker similarity, robustness, and latency.

  • Investigate transfer learning, semi-supervised learning, distillation, quantization, and parameter-efficient fine-tuning.

  • Review and reproduce recent research, propose modeling improvements, and communicate findings through technical reports and presentations.

  • Collaborate with engineering and product teams to validate and integrate research outcomes.

Yêu cầu công việc

  • Bachelor’s degree in Computer Science, Engineering, Mathematics, Data Science, or a related field.

  • Strong Python skills and hands-on experience with PyTorch, TensorFlow, JAX, or similar frameworks.

  • Solid understanding of machine learning, deep learning, sequence modeling, and Transformer-based architectures.

  • Experience in at least one speech or audio area, such as ASR, TTS, VAD, speaker recognition, diarization, speech enhancement, or audio classification.

  • Familiarity with speech-processing tools such as torchaudio, librosa, Kaldi, ESPnet, NeMo, SpeechBrain, Whisper, or WeNet.

  • Experience preparing speech datasets and conducting controlled experiments, ablation studies, and model-error analysis.

  • Understanding of common speech evaluation metrics, including WER, CER, perceptual quality, and speaker similarity.

  • Ability to read, reproduce, and extend recent research papers.

  • Strong analytical, problem-solving, and cross-functional collaboration skills.