Cerebras
Cerebras
Cerebras उच्च-प्रदर्शन वाले मॉडल प्रशिक्षण और इन्फरेंस के लिए वेफर-स्केल AI हार्डवेयर और एकीकृत सॉफ़्टवेयर विकसित करता है। इसका Wafer-Scale Engine CS-3 सिस्टम और Cerebras Inference Cloud को शक्ति देता है, जो OpenAI API-संगत वर्कफ़्लो तथा ऑन-प्रिमाइसेस, क्लाउड और एज वातावरणों में परिनियोजन का समर्थन करता है। कंपनी का प्लेटफ़ॉर्म मॉडल सर्विंग, फ़ाइन-ट्यूनिंग, प्री-ट्रेनिंग, रीयल-टाइम AI एजेंट, जीनोमिक्स, दवा खोज, साइबर सुरक्षा और एंटरप्राइज़ सर्च जैसे अनुप्रयोगों के लिए बनाया गया है। Cerebras चुनौतीपूर्ण कृत्रिम बुद्धिमत्ता वर्कलोड बनाने और संचालित करने वाले संगठनों के लिए विशेषज्ञ चिप, सिस्टम और सॉफ़्टवेयर को एक साथ लाता है तथा इसमें 501–1,000 कर्मचारी हैं।

Machine Learning Engineer, Model Bring-Up

Bring transformer models onto Cerebras AI accelerator hardware and optimize their inference performance. Develop MLIR compiler paths while improving model correctness, latency, throughput, and memory efficiency.

विवरण

  • Bring new models into production by analyzing architectures, converting weights, implementing execution paths, and validating results against reference implementations.
  • Lower models to Cerebras hardware with MLIR dialects, graph transformations, lowering passes, and hardware-specific mappings.
  • Support attention, matrix multiplication, normalization, positional embeddings, and other operators through compiler and kernel development.
  • Optimize operator fusion, tensor layouts, tiling, memory allocation, data movement, and parallel execution.
  • Improve prefill and decode performance by tuning KV-cache management, batching, and quantization for better latency, throughput, and memory use.
  • Analyze numerical discrepancies and assess how precision changes and compiler optimizations affect accuracy.
  • Use profilers, execution traces, and hardware counters to locate compute, memory, communication, and runtime bottlenecks.
  • Partner with hardware, compiler, kernel, and runtime teams to deliver dependable model support and reproducible performance benchmarks.

आवश्यकताएँ

  • Proficiency in C++ and Python programming.
  • Hands-on experience bringing up and debugging machine learning models in PyTorch or a comparable framework.
  • Practical MLIR experience with dialects, rewrite patterns, transformation passes, and lowering pipelines.
  • Knowledge of compiler fundamentals, including intermediate representations, dataflow analysis, and code generation.
  • Understanding of transformer architectures, attention mechanisms, tensor operations, and numerical precision.
  • Experience profiling and optimizing workloads on GPUs or other AI accelerators.
  • Ability to diagnose correctness and performance issues across model code, compiler output, kernels, and runtime execution.
  • Experience with large language model inference, including GQA, sliding-window attention, mixture-of-experts, KV caching, and speculative decoding.
  • Experience with FP16, BF16, FP8, or low-bit quantization and their accuracy and performance tradeoffs.
  • Background developing accelerator kernels or hardware-specific compiler backends.
  • Familiarity with distributed execution, model parallelism, and accelerator memory hierarchies.
  • Contributions to MLIR, LLVM, inference frameworks, or related open-source projects.

लाभ

  • Stable employment in a startup-paced environment.
  • Opportunities to publish and open-source advanced AI research.
  • Work with one of the world's fastest AI supercomputers.
  • A straightforward, non-corporate culture that respects individual beliefs.
  • Ongoing learning, professional growth, and support.
  • An equitable and diverse workplace.

संबंधित नौकरियाँ

WeAssist.io

Remote Operations Manager (Philippines)

WeAssist.io
51 – 200 कर्मचारी
B2BSaaSउत्पादकता

Lead cross-functional operations across Sales, Marketing, Customer Success, and Product for an AI business. Manage executive support, systems, KPIs, client communications, and global events.

खोलें
Cape

GRC Engineer

Cape

Build automated controls and compliance programs to secure Cape’s privacy-focused mobile carrier. Lead SOC 2 and CMMC efforts, risk management, audits, and customer security reviews.

खोलें
Fortinet

Fortinet Accountant

Fortinet

Support Fortinet’s Canadian general accounting operations through reconciliations, journal entries, month-end close, audits, procurement, and operating expense accounting.

खोलें
AXON Networks

Data Engineer, Databricks and PySpark

AXON Networks

Build and support Databricks, PySpark, and SQL pipelines for AXON Networks’ AI-driven telecom orchestration platform. Help keep lakehouse datasets traceable, tested, and production-ready in Madrid.

खोलें
Sleek

Senior Performance Marketing Lead

Sleek

Lead performance marketing across Sleek’s markets, optimizing Google Search campaigns, attribution, and ROAS for its AI-powered back-office services. Drive revenue growth across Asia-Pacific through data-led campaign strategy.

खोलें
Nooks

Commercial Contracts Manager

Nooks

Manage negotiations and contract playbooks for Nooks’ classified infrastructure business, overseeing agreements, lifecycle records, and deal-cycle reporting.

खोलें
Runway

Remote Product Marketing Copywriter (United States)

Runway

Write conversion-focused copy across acquisition, onboarding, lifecycle, and product experiences. Help users discover and adopt Runway’s AI simulation and creative tools.

खोलें
SupportNinja

Remote MLS Coordinator (Colombia)

SupportNinja
1,001 – 5,000 कर्मचारी
B2BSaaS

SupportNinja’s MLS Coordinator provides remote member assistance and Tier 1 systems support while handling listing, membership, compliance, and data operations from Colombia.

खोलें
OpenArt AI

Senior Enterprise Content Marketing Writer

OpenArt AI

Shape enterprise marketing for OpenArt’s AI creative platform through research, thought leadership, customer stories, and buyer education.

खोलें