Capital One
Capital One
Capital One is a financial services and fintech company offering credit cards, auto loans, banking, and savings products. Its work combines financial expertise with technology and digital tools to improve how customers manage money. The company’s careers focus includes building customer-oriented solutions within a diverse and inclusive workforce.

Capital One Machine Learning Engineer Onsite

Build and operate production machine learning systems at Capital One. Work across models, cloud infrastructure, data pipelines, and responsible AI applications.

Description

  • Build and deliver machine learning models and components that address practical business needs
  • Partner with Product and Data Science teams throughout development
  • Guide ML infrastructure choices using modeling methods, data and feature selection, training, tuning, dimensionality, bias-variance analysis, and validation
  • Develop and test application code, validate ML models, and automate testing and deployment workflows
  • Work with cross-functional Agile teams to build and improve big data and machine learning applications
  • Retrain, maintain, and monitor models running in production
  • Use or create cloud architectures, platforms, and technologies to deliver efficient ML models at scale
  • Build optimized data pipelines that supply machine learning models
  • Implement continuous integration and delivery practices with automated testing and monitoring
  • Maintain code that limits vulnerabilities and supports model governance and responsible, explainable AI
  • Program with languages such as Python, Scala, or Java

Requirements

  • Bachelor’s degree or higher in computer science, machine learning, or a related quantitative discipline
  • At least four years of programming experience with Python, Java, Golang, or C++
  • At least four years of machine learning experience with PyTorch or TensorFlow and libraries such as Pandas, NumPy, and Scikit-learn
  • At least four years of experience using and operating large-scale distributed systems, including Spark or Ray, to prepare AI or machine learning data
  • At least two years of experience deploying and operating machine learning solutions in production, along with production cloud services on AWS, GCP, or Azure
  • At least two years of experience using Kubernetes to manage large-scale, containerized machine learning software systems
  • New employment authorization sponsorship and immigration-related support are not available
  • Preferred: master’s or doctoral degree in computer science, electrical engineering, mathematics, or a related field
  • Preferred: three or more years optimizing ML algorithms, configurations, and infrastructure
  • Preferred: three or more years applying software development practices such as source control, testing, code reviews, and CI/CD
  • Preferred: three or more years building resilient software with pre-production testing, advanced deployment methods, monitoring, alarms, and incident response planning
  • Preferred: three or more years working with machine learning techniques, model types, architectures, training concepts, and evaluation methods
  • Preferred: three or more years designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models
  • Preferred: one or more years serving as a technical lead for machine learning solutions
  • Preferred: authorship or co-authorship of a paper covering an ML technique, model, or proof of concept

Benefits

  • Performance-based incentive compensation may include cash bonuses and/or long-term incentives
  • Comprehensive, competitive, and inclusive health, financial, and other benefits designed to support overall well-being
  • Equal opportunity employer committed to nondiscrimination
  • Reasonable accommodations are available for applicants who require them

Related Jobs

Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Employees
Artificial IntelligenceB2B

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Open