Cerebras
Cerebras
Cerebras dezvoltă hardware de inteligență artificială la nivel de wafer și software integrat pentru antrenarea și inferența modelelor de înaltă performanță. Wafer-Scale Engine al companiei alimentează sistemul CS-3 și Cerebras Inference Cloud, care acceptă fluxuri de lucru compatibile cu OpenAI API și implementarea în medii locale, cloud și edge. Platforma companiei este concepută pentru aplicații precum servirea modelelor, ajustarea fină, preantrenarea, agenții AI în timp real, genomica, descoperirea de medicamente, securitatea cibernetică și căutarea enterprise. Cerebras reunește cipuri, sisteme și software specializate pentru organizații care construiesc și operează sarcini de lucru solicitante de inteligență artificială, având o echipă de 501–1.000 de angajați.

Machine Learning Engineer, Model Bring-Up

Bring transformer models onto Cerebras AI accelerator hardware and optimize their inference performance. Develop MLIR compiler paths while improving model correctness, latency, throughput, and memory efficiency.

Descriere

  • Bring new models into production by analyzing architectures, converting weights, implementing execution paths, and validating results against reference implementations.
  • Lower models to Cerebras hardware with MLIR dialects, graph transformations, lowering passes, and hardware-specific mappings.
  • Support attention, matrix multiplication, normalization, positional embeddings, and other operators through compiler and kernel development.
  • Optimize operator fusion, tensor layouts, tiling, memory allocation, data movement, and parallel execution.
  • Improve prefill and decode performance by tuning KV-cache management, batching, and quantization for better latency, throughput, and memory use.
  • Analyze numerical discrepancies and assess how precision changes and compiler optimizations affect accuracy.
  • Use profilers, execution traces, and hardware counters to locate compute, memory, communication, and runtime bottlenecks.
  • Partner with hardware, compiler, kernel, and runtime teams to deliver dependable model support and reproducible performance benchmarks.

Cerințe

  • Proficiency in C++ and Python programming.
  • Hands-on experience bringing up and debugging machine learning models in PyTorch or a comparable framework.
  • Practical MLIR experience with dialects, rewrite patterns, transformation passes, and lowering pipelines.
  • Knowledge of compiler fundamentals, including intermediate representations, dataflow analysis, and code generation.
  • Understanding of transformer architectures, attention mechanisms, tensor operations, and numerical precision.
  • Experience profiling and optimizing workloads on GPUs or other AI accelerators.
  • Ability to diagnose correctness and performance issues across model code, compiler output, kernels, and runtime execution.
  • Experience with large language model inference, including GQA, sliding-window attention, mixture-of-experts, KV caching, and speculative decoding.
  • Experience with FP16, BF16, FP8, or low-bit quantization and their accuracy and performance tradeoffs.
  • Background developing accelerator kernels or hardware-specific compiler backends.
  • Familiarity with distributed execution, model parallelism, and accelerator memory hierarchies.
  • Contributions to MLIR, LLVM, inference frameworks, or related open-source projects.

Beneficii

  • Stable employment in a startup-paced environment.
  • Opportunities to publish and open-source advanced AI research.
  • Work with one of the world's fastest AI supercomputers.
  • A straightforward, non-corporate culture that respects individual beliefs.
  • Ongoing learning, professional growth, and support.
  • An equitable and diverse workplace.

Locuri de muncă similare

Terumo Medical Corporation

Territory Manager, Interventional Systems — Philadelphia

Terumo Medical Corporation

Lead sales of Terumo Interventional Systems devices to hospitals and outpatient facilities in the Philadelphia area. Build accounts, support procedures, educate clinicians, and meet territory sales goals.

Deschide
Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Angajați
B2BInteligență artificială

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Deschide
General Motors

Summer AI Research Intern, Embodied AI

General Motors

Research machine learning, robotics, and multimodal models for autonomous driving at General Motors. Run experiments and collaborate with research and engineering teams in a hybrid internship based in Sunnyvale, California.

Deschide
Mobileye

Senior Deep Learning Researcher, Autonomous Vehicles

Mobileye

Research and develop deep learning systems for Mobileye’s autonomous vehicles, including large-scale neural networks built for its EyeQ chip. Improve model performance and bring research into production.

Deschide