Walmart
Walmart
Walmart este o companie globală de retail și comerț electronic care deservește consumatorii prin hipermarketuri, magazine universale cu discount, magazine alimentare și o piață online. Oferta sa largă de produse este susținută de o rețea extinsă de magazine fizice, operațiuni de vânzare directă către consumatori și sistemele de logistică și de lanț de aprovizionare necesare pentru circulația produselor în întreaga companie.

Senior Data Engineer - Bangalore, India (Onsite)

Build scalable batch and streaming pipelines and GenAI data products for Walmart’s post-payment audit operations. Improve payment accuracy, reconciliation, recovery processes, and finance analytics.

Descriere

  • Build and maintain scalable batch, streaming, and near-real-time pipelines for post-payment audit, payment integrity, supplier reconciliation, recovery operations, and finance analytics.
  • Develop curated datasets, semantic models, feature-ready tables, and reusable data products from invoice, purchase order, receiving, supplier, claims, pricing, and payment systems.
  • Create ELT, ETL, and streaming workflows with SQL, Python, PySpark, Apache Flink, dbt, orchestration tools, event-streaming platforms, and cloud data technologies.
  • Develop streaming and change-data-capture pipelines with event-time processing, late-data handling, replayability, idempotency, backfills, and failure recovery.
  • Design and operate lakehouse data products using Apache Hudi, Apache Iceberg, or Delta Lake.
  • Establish data quality checks, reconciliation controls, schema validation, lineage, monitoring, alerting, and observability practices.
  • Work with Finance, Audit, Operations, Product, Data Science, and Engineering teams to deliver production-ready data solutions.
  • Enable anomaly detection, exception prioritization, audit rules, machine-learning scoring, and GenAI workflows.
  • Create automation and data services that help audit teams summarize evidence, triage exceptions, provide investigation context, and reduce manual review.
  • Implement GenAI data engineering patterns such as retrieval-ready document stores, embedding pipelines, vector database integrations, metadata enrichment, and governed knowledge retrieval.
  • Improve batch and streaming workloads for reliability, freshness, latency, state management, compute efficiency, and downstream usability.
  • Investigate data discrepancies, pipeline failures, reconciliation mismatches, and source-system issues through root-cause analysis.

Cerințe

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, Data Engineering, Analytics, Mathematics, Statistics, or a related discipline plus 6+ years of relevant experience; alternatively, a master’s degree with 4+ years or a PhD with 3+ years of relevant experience.
  • Demonstrated hands-on ability with SQL, Python, PySpark, Apache Flink or a comparable stream-processing framework, and large-scale distributed data processing.
  • Experience with dbt or a comparable transformation framework, Airflow or similar orchestration tools, Kafka or an equivalent event-streaming platform, and cloud or enterprise data technologies such as BigQuery, Hive, or Spark.
  • Track record of designing, deploying, and supporting production batch and streaming pipelines, ELT/ETL workflows, data models, and reusable data products.
  • Solid knowledge of data modeling, partitioning, performance tuning, incremental processing, CDC, schema evolution, data contracts, metadata management, lineage, and lakehouse formats including Apache Hudi, Apache Iceberg, or Delta Lake.
  • Experience delivering data quality checks, reconciliation logic, validation frameworks, observability, monitoring, alerting, and incident resolution for production pipelines.
  • Experience developing streaming solutions with event-time processing, stateful transformations, checkpointing, watermarking, exactly-once or effectively-once processing, and replay or backfill strategies.
  • Experience handling large structured and semi-structured datasets from multiple source systems, including operational data with inconsistent business definitions.
  • Strong collaboration skills and the ability to work across Finance, Audit, Operations, Product, Data Science, and Engineering.
  • Minimum qualification alternatives include a bachelor’s degree in Computer Science with 3 years of software engineering or related experience; 5 years of software engineering or related experience; or a master’s degree in Computer Science with 1 year of software engineering or related experience.
  • At least 2 years of experience in data engineering, database engineering, business intelligence, or business analytics.

Beneficii

  • Performance-based incentive awards.
  • Maternity and parental leave.
  • Paid time off.
  • Health benefits.

Locuri de muncă similare

Ambush

Senior Machine Learning and AI Engineer

Ambush

Build production-grade generative AI and agent workflows for Ambush’s financial services clients. Develop Python and FastAPI services, SQL data workflows, retrieval-augmented generation, and cloud-based AI systems.

Deschide
GFN

Senior Software Engineer, Machine Learning

GFN
201 – 500 Angajați
Comerț electronicFintechSaaS

Build backend infrastructure and AI-enabled learning products for GFN, a German education provider. Develop scalable educational platforms with a focus on robust software architecture and practical machine learning.

Deschide
TEKsystems

Remote Data Program Project Manager — Washington

TEKsystems

Coordinate enterprise data initiatives, report migrations, and adoption of new processes and reporting solutions for TEKsystems. Track project progress, risks, deliverables, and communications with business and technical stakeholders.

Deschide