Cerebras
Cerebras
Cerebras develops wafer-scale AI hardware and integrated software for high-performance model training and inference. Its Wafer-Scale Engine powers the CS-3 system and Cerebras Inference Cloud, which supports OpenAI API-compatible workflows and deployment across on-premise, cloud, and edge environments. The company’s platform is designed for applications including model serving, fine-tuning, pre-training, real-time AI agents, genomics, drug discovery, cybersecurity, and enterprise search. Cerebras brings together specialized chips, systems, and software for organizations building and operating demanding artificial intelligence workloads, with a team of 501–1,000 employees.

Machine Learning Engineer, Model Bring-Up

Bring transformer models onto Cerebras AI accelerator hardware and optimize their inference performance. Develop MLIR compiler paths while improving model correctness, latency, throughput, and memory efficiency.

Description

  • Bring new models into production by analyzing architectures, converting weights, implementing execution paths, and validating results against reference implementations.
  • Lower models to Cerebras hardware with MLIR dialects, graph transformations, lowering passes, and hardware-specific mappings.
  • Support attention, matrix multiplication, normalization, positional embeddings, and other operators through compiler and kernel development.
  • Optimize operator fusion, tensor layouts, tiling, memory allocation, data movement, and parallel execution.
  • Improve prefill and decode performance by tuning KV-cache management, batching, and quantization for better latency, throughput, and memory use.
  • Analyze numerical discrepancies and assess how precision changes and compiler optimizations affect accuracy.
  • Use profilers, execution traces, and hardware counters to locate compute, memory, communication, and runtime bottlenecks.
  • Partner with hardware, compiler, kernel, and runtime teams to deliver dependable model support and reproducible performance benchmarks.

Requirements

  • Proficiency in C++ and Python programming.
  • Hands-on experience bringing up and debugging machine learning models in PyTorch or a comparable framework.
  • Practical MLIR experience with dialects, rewrite patterns, transformation passes, and lowering pipelines.
  • Knowledge of compiler fundamentals, including intermediate representations, dataflow analysis, and code generation.
  • Understanding of transformer architectures, attention mechanisms, tensor operations, and numerical precision.
  • Experience profiling and optimizing workloads on GPUs or other AI accelerators.
  • Ability to diagnose correctness and performance issues across model code, compiler output, kernels, and runtime execution.
  • Experience with large language model inference, including GQA, sliding-window attention, mixture-of-experts, KV caching, and speculative decoding.
  • Experience with FP16, BF16, FP8, or low-bit quantization and their accuracy and performance tradeoffs.
  • Background developing accelerator kernels or hardware-specific compiler backends.
  • Familiarity with distributed execution, model parallelism, and accelerator memory hierarchies.
  • Contributions to MLIR, LLVM, inference frameworks, or related open-source projects.

Benefits

  • Stable employment in a startup-paced environment.
  • Opportunities to publish and open-source advanced AI research.
  • Work with one of the world's fastest AI supercomputers.
  • A straightforward, non-corporate culture that respects individual beliefs.
  • Ongoing learning, professional growth, and support.
  • An equitable and diverse workplace.

Related Jobs

Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Employees
Artificial IntelligenceB2B

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Open