Cerebras
Cerebras
Cerebras өнімділігі жоғары модельдерді оқытуға және инференске арналған пластина ауқымындағы жасанды интеллект аппараттық құралдары мен біріктірілген бағдарламалық жасақтаманы әзірлейді. Оның Wafer-Scale Engine жүйесі CS-3 жүйесі мен OpenAI API үйлесімді жұмыс процестерін қолдайтын Cerebras Inference Cloud платформасының негізі болып табылады; бұл шешімдер жергілікті инфрақұрылымда, бұлтта және шеткі ортада орналастыруға мүмкіндік береді. Компания платформасы модельдерді ұсыну, дәлдеп баптау, алдын ала оқыту, нақты уақыттағы AI агенттері, геномика, дәрі-дәрмек жасау, киберқауіпсіздік және кәсіпорындық іздеу сияқты қолданбаларға арналған. Cerebras күрделі жасанды интеллект жұмыс жүктемелерін жасайтын және іске қосатын ұйымдарға арналған мамандандырылған чиптерді, жүйелерді және бағдарламалық жасақтаманы біріктіреді; компанияда 501–1 000 қызметкер жұмыс істейді.

Machine Learning Engineer, Model Bring-Up

Bring transformer models onto Cerebras AI accelerator hardware and optimize their inference performance. Develop MLIR compiler paths while improving model correctness, latency, throughput, and memory efficiency.

Сипаттама

  • Bring new models into production by analyzing architectures, converting weights, implementing execution paths, and validating results against reference implementations.
  • Lower models to Cerebras hardware with MLIR dialects, graph transformations, lowering passes, and hardware-specific mappings.
  • Support attention, matrix multiplication, normalization, positional embeddings, and other operators through compiler and kernel development.
  • Optimize operator fusion, tensor layouts, tiling, memory allocation, data movement, and parallel execution.
  • Improve prefill and decode performance by tuning KV-cache management, batching, and quantization for better latency, throughput, and memory use.
  • Analyze numerical discrepancies and assess how precision changes and compiler optimizations affect accuracy.
  • Use profilers, execution traces, and hardware counters to locate compute, memory, communication, and runtime bottlenecks.
  • Partner with hardware, compiler, kernel, and runtime teams to deliver dependable model support and reproducible performance benchmarks.

Талаптар

  • Proficiency in C++ and Python programming.
  • Hands-on experience bringing up and debugging machine learning models in PyTorch or a comparable framework.
  • Practical MLIR experience with dialects, rewrite patterns, transformation passes, and lowering pipelines.
  • Knowledge of compiler fundamentals, including intermediate representations, dataflow analysis, and code generation.
  • Understanding of transformer architectures, attention mechanisms, tensor operations, and numerical precision.
  • Experience profiling and optimizing workloads on GPUs or other AI accelerators.
  • Ability to diagnose correctness and performance issues across model code, compiler output, kernels, and runtime execution.
  • Experience with large language model inference, including GQA, sliding-window attention, mixture-of-experts, KV caching, and speculative decoding.
  • Experience with FP16, BF16, FP8, or low-bit quantization and their accuracy and performance tradeoffs.
  • Background developing accelerator kernels or hardware-specific compiler backends.
  • Familiarity with distributed execution, model parallelism, and accelerator memory hierarchies.
  • Contributions to MLIR, LLVM, inference frameworks, or related open-source projects.

Артықшылықтар

  • Stable employment in a startup-paced environment.
  • Opportunities to publish and open-source advanced AI research.
  • Work with one of the world's fastest AI supercomputers.
  • A straightforward, non-corporate culture that respects individual beliefs.
  • Ongoing learning, professional growth, and support.
  • An equitable and diverse workplace.

Ұқсас жұмыс орындары

Artemys

IT Talent Acquisition Specialist

Artemys

Recruit IT infrastructure professionals for Artemys, a digital transformation company. Lead sourcing and hiring while contributing to employer branding initiatives.

Ашу
4M Analytics

Senior Computer Vision Algorithm Engineer

4M Analytics

Develop computer vision and machine learning systems for 4M Analytics’ subsurface utility mapping platform. Own algorithms from research through production in partnership with data engineering and product teams.

Ашу
Joom

Enterprise Customer Success Manager, Brazil (Remote)

Joom

Lead renewals, retention, and account growth for JoomPulse enterprise customers. Turn Mercado Livre and Shopee analytics into practical recommendations for marketplace sellers.

Ашу
Cyera

Regional Sales Director, Identity Security

Cyera

Lead Cyera’s Pacific Northwest sales team, grow the territory, and exceed quotas for its AI and data security platform.

Ашу
Terumo Medical Corporation

Region Manager, Terumo Interventional Systems Sales

Terumo Medical Corporation

Lead medical device sales across a North Central New Jersey region, managing field teams and hospital relationships. Drive regional revenue, sales performance, and compliant promotion of Terumo Interventional Systems products.

Ашу
Dream

Senior Program Manager, Sovereign AI (Hybrid, Israel)

Dream

Lead Dream’s Sovereign AI roadmap for government and critical infrastructure customers, coordinating research, engineering, product, launches, and AI-enabled operations.

Ашу
Claroty

QA Automation Engineer

Claroty

Develop API and UI automation and CI/CD quality gates for Claroty’s cyber-physical security platform. Help protect critical infrastructure through reliable testing.

Ашу
SKF Group

Global Product Engineer, Bearings — SKF, Bengaluru

SKF Group

Design bearing products, models, drawings, and variants at SKF, a global provider of rotating equipment solutions. Support modular design, testing, and product lifecycle management.

Ашу
Webbing

Mobile Fraud Analyst – Bucharest Hybrid

Webbing

Investigate and prevent telecom fraud across Webbing’s global MVNO connectivity and IoT services. Analyze traffic, roaming, SIM, voice, SMS, and data activity for anomalies.

Ашу