42dot
42dot
51 – 200 Employees
AutomotiveConsultingLogistics
42dot is a mobility AI company working at the intersection of automotive technology and logistics. Its work focuses on software-defined vehicle technologies, autonomous electric vehicles, and the software and data infrastructure needed for scalable urban transportation. The company develops vehicle software, fleet data platforms, and AI models for perception and control, alongside TAP!, its mobility operating system for deploying autonomous services. Through research and collaboration on vehicle electronics and software-first architectures, 42dot aims to make mobility systems more autonomous, adaptable, and easier to operate.

Deep Learning Engineer, Text-to-Speech Development

Optimize text-to-speech models for Gleo AI’s in-vehicle voice experiences. Build real-time, multilingual speech synthesis for on-device and server environments.

Description

  • Optimize TTS models for on-device and server environments
  • Port and optimize models for NPU and other AI accelerator environments
  • Reduce model size and improve performance through quantization
  • Develop real-time streaming speech synthesis systems and optimize latency
  • Develop and deploy multilingual, multi-speaker TTS models
  • Research and develop advanced TTS models based on LLMs and flow matching
  • Build speech synthesis datasets with generative models and improve data quality
  • Analyze and resolve TTS model and service issues in production environments
  • Integrate researched and developed TTS technologies into vehicle products and services

Requirements

  • At least 3 years of relevant TTS experience
  • Foundational knowledge of speech synthesis and TTS concepts
  • Experience with model compression and computational optimization
  • Experience optimizing quantization, memory usage, and latency
  • Experience porting and optimizing models for NPU and other AI accelerator environments
  • Experience optimizing models with TensorRT and ONNX
  • Understanding of audio signal processing
  • Development and debugging experience in Linux environments
  • Programming proficiency in Python, C++, or shell
  • Proficiency with open-source deep learning frameworks such as PyTorch
  • Experience developing multilingual and multi-speaker TTS models is preferred
  • Research or development experience with speech or audio tokenization technologies is preferred
  • Publication as an author in a leading conference or journal in speech synthesis, machine learning, or artificial intelligence is preferred
  • A relevant master's or doctoral degree, or completion of such a program, is preferred
  • Korean proficiency information must be provided

Benefits

  • Applicants eligible for national merit or employment protection are preferred
  • Applicants with a registered disability certificate are preferred
  • A three-month probationary period may apply
  • A reference check may be conducted with the applicant’s consent after the interview process is complete

Related Jobs