Stack AV
Stack AV
Stack AV dezvoltă sisteme pentru camioane autonome bazate pe inteligență artificială, destinate industriilor transporturilor și logisticii. Activitatea companiei combină inteligența artificială, învățarea automată și tehnologiile cloud pentru a sprijini operațiuni de transport rutier mai sigure, mai fiabile și mai eficiente. Compania se concentrează pe îmbunătățirea inteligenței lanțului de aprovizionare, a performanței livrărilor și a rezultatelor de afaceri prin tehnologia vehiculelor autonome, având siguranța în centrul abordării sale inginerești. Echipa sa multidisciplinară activează în domenii precum logistica, producția și consultanța, dezvoltând soluții inteligente pentru transportul modern de mărfuri.

Staff Software Engineer, ML Training Infrastructure at Stack AV (Remote, Pennsylvania)

Lead the design of scalable machine learning training infrastructure and end-to-end model pipelines. Help advance Stack AV’s autonomous trucking systems through platform engineering, performance optimization, and cross-functional collaboration.

Descriere

  • Build and refine high-performance training platform components spanning orchestration, training abstractions, control-plane services, observability, and performance optimization.
  • Develop complete machine learning pipelines covering log processing, feature extraction, dataset schemas and storage, model configuration, training, profiling, and acceleration.
  • Evaluate training infrastructure to uncover and address performance constraints.
  • Promote system abstractions and developer tools that help machine learning engineers iterate quickly on models.
  • Adopt open-source technologies that let machine learning engineers independently profile and improve their workflows.
  • Uphold strong engineering standards and help foster a team culture centered on technical excellence.
  • Lead the architecture and implementation of a high-performance, multi-tenant AI training platform.
  • Partner with teams across ML Platform, Infrastructure, Autonomy, and Safety Evaluation.

Cerințe

  • Bachelor’s or master’s degree in computer science, engineering, or a related discipline.
  • At least six years of experience developing ML platforms and machine learning applications.
  • Advanced programming ability in Python, C++, or an equivalent language.
  • Experience with Lance, PyTorch, Ray Data, or comparable technologies.
  • Demonstrated success building scalable, dependable infrastructure in a fast-moving environment and partnering with machine learning engineers across modeling teams.
  • Strong grasp of system design tradeoffs, with the communication skills to build alignment among cross-functional teams.
  • Background in model training, model optimization, or large-scale data-processing pipelines.
  • Strong analytical reasoning and problem-solving ability.
  • Clear written and verbal communication skills, including the ability to explain complex technical ideas to nontechnical stakeholders.
  • May need to verify residence, U.S. person status, and/or citizenship to meet U.S. national security and export-control requirements.

Beneficii

  • An equal-opportunity workplace focused on inclusion, entrepreneurship, and innovation across gender, race, age, sexual orientation, religion, disability, and identity.

Locuri de muncă similare

Terumo Medical Corporation

Territory Manager, Interventional Systems — Philadelphia

Terumo Medical Corporation

Lead sales of Terumo Interventional Systems devices to hospitals and outpatient facilities in the Philadelphia area. Build accounts, support procedures, educate clinicians, and meet territory sales goals.

Deschide
Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Angajați
B2BInteligență artificială

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Deschide