Unity
Unity
Unity develops tools and services for creating games, applications, and immersive real-time 3D experiences. Its platform includes Unity Engine for cross-platform content creation, Unity Cloud for development workflows, and Unity Grow for user acquisition and game monetization. Unity serves creators and teams in gaming as well as enterprise fields such as manufacturing, combining graphics technology, AI capabilities, developer resources, community support, and professional training to help bring interactive ideas into production.

Senior Machine Learning Infrastructure Engineer

Build distributed training systems, data pipelines, and model-serving infrastructure at Unity. Support the company’s game engine and 3D creation platform for games and industrial applications.

Description

  • Create and maintain data and feature pipelines for model training and experimentation.
  • Build and improve infrastructure for distributed training.
  • Operate multi-stage machine learning pipelines with workflow orchestration tools.
  • Support model serving and deployment so models move reliably from training into production.
  • Strengthen reproducibility and reliability with monitoring, alerts, dataset validation, and testing.
  • Lead manager-scoped projects from design through rollout.
  • Diagnose unusual or inconsistent pipeline and training issues, then turn recurring patterns into improvements.
  • Meet near-term delivery needs while advancing the platform’s broader direction.
  • Collaborate with machine learning engineers, researchers, and colleagues across teams.
  • Communicate progress and tradeoffs to managers and directors.
  • Adapt to shifting team priorities by learning new tools and areas of the stack.

Requirements

  • Production experience building machine learning infrastructure, data platforms, or distributed systems.
  • Strong Python skills and experience handling data-intensive workloads.
  • Practical experience with machine learning frameworks such as PyTorch and distributed computing tools such as Ray or Spark.
  • Experience with data pipelines, model training workflows, or large datasets outside academic settings.
  • Ability to work across the machine learning stack and quickly learn unfamiliar tools and systems.
  • Skill in finding patterns in ambiguous or inconsistent problems and translating them into practical technical needs.
  • Clear communicator who works collaboratively with peers, managers, and partner teams.
  • Bachelor’s degree in computer science, machine learning, systems, or a related field, or equivalent practical experience.
  • Professional proficiency in spoken and written English.
  • Preferred: Experience with Kubernetes and cloud infrastructure such as GCP or AWS.
  • Preferred: Familiarity with observability tools such as Prometheus or Grafana.
  • Preferred: Experience with workflow orchestration systems such as Airflow or Prefect.
  • Preferred: Experience with model-serving frameworks such as Triton, KServe, or Ray Serve.
  • Preferred: Exposure to large-scale data platforms, including data lakes, warehouses, and streaming systems such as Kafka.
  • Preferred: Familiarity with feature stores or experimentation platforms.

Benefits

  • Health, life, and disability insurance.
  • Commute subsidy.
  • Employee stock ownership.
  • Retirement and pension plans.
  • Vacation and personal days.
  • Parental leave and family-care programs.
  • Food and snacks at the office.
  • Mental health and wellbeing programs and support.
  • Employee resource groups.
  • Global employee assistance program.
  • Training and development programs.
  • Volunteering and donation matching.
  • Equity awards.
  • Participation in company incentive plans, including annual discretionary bonuses or sales commissions.

Related Jobs