Walmart
Walmart
Walmart is a global retail and eCommerce company serving consumers through hypermarkets, discount department stores, grocery locations, and an online marketplace. Its broad product offering is supported by a large physical-store network, direct-to-consumer operations, and the logistics and supply-chain systems needed to keep products moving across the business.

Senior Data Engineer - Bangalore, India (Onsite)

Build scalable batch and streaming pipelines and GenAI data products for Walmart’s post-payment audit operations. Improve payment accuracy, reconciliation, recovery processes, and finance analytics.

Description

  • Build and maintain scalable batch, streaming, and near-real-time pipelines for post-payment audit, payment integrity, supplier reconciliation, recovery operations, and finance analytics.
  • Develop curated datasets, semantic models, feature-ready tables, and reusable data products from invoice, purchase order, receiving, supplier, claims, pricing, and payment systems.
  • Create ELT, ETL, and streaming workflows with SQL, Python, PySpark, Apache Flink, dbt, orchestration tools, event-streaming platforms, and cloud data technologies.
  • Develop streaming and change-data-capture pipelines with event-time processing, late-data handling, replayability, idempotency, backfills, and failure recovery.
  • Design and operate lakehouse data products using Apache Hudi, Apache Iceberg, or Delta Lake.
  • Establish data quality checks, reconciliation controls, schema validation, lineage, monitoring, alerting, and observability practices.
  • Work with Finance, Audit, Operations, Product, Data Science, and Engineering teams to deliver production-ready data solutions.
  • Enable anomaly detection, exception prioritization, audit rules, machine-learning scoring, and GenAI workflows.
  • Create automation and data services that help audit teams summarize evidence, triage exceptions, provide investigation context, and reduce manual review.
  • Implement GenAI data engineering patterns such as retrieval-ready document stores, embedding pipelines, vector database integrations, metadata enrichment, and governed knowledge retrieval.
  • Improve batch and streaming workloads for reliability, freshness, latency, state management, compute efficiency, and downstream usability.
  • Investigate data discrepancies, pipeline failures, reconciliation mismatches, and source-system issues through root-cause analysis.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, Data Engineering, Analytics, Mathematics, Statistics, or a related discipline plus 6+ years of relevant experience; alternatively, a master’s degree with 4+ years or a PhD with 3+ years of relevant experience.
  • Demonstrated hands-on ability with SQL, Python, PySpark, Apache Flink or a comparable stream-processing framework, and large-scale distributed data processing.
  • Experience with dbt or a comparable transformation framework, Airflow or similar orchestration tools, Kafka or an equivalent event-streaming platform, and cloud or enterprise data technologies such as BigQuery, Hive, or Spark.
  • Track record of designing, deploying, and supporting production batch and streaming pipelines, ELT/ETL workflows, data models, and reusable data products.
  • Solid knowledge of data modeling, partitioning, performance tuning, incremental processing, CDC, schema evolution, data contracts, metadata management, lineage, and lakehouse formats including Apache Hudi, Apache Iceberg, or Delta Lake.
  • Experience delivering data quality checks, reconciliation logic, validation frameworks, observability, monitoring, alerting, and incident resolution for production pipelines.
  • Experience developing streaming solutions with event-time processing, stateful transformations, checkpointing, watermarking, exactly-once or effectively-once processing, and replay or backfill strategies.
  • Experience handling large structured and semi-structured datasets from multiple source systems, including operational data with inconsistent business definitions.
  • Strong collaboration skills and the ability to work across Finance, Audit, Operations, Product, Data Science, and Engineering.
  • Minimum qualification alternatives include a bachelor’s degree in Computer Science with 3 years of software engineering or related experience; 5 years of software engineering or related experience; or a master’s degree in Computer Science with 1 year of software engineering or related experience.
  • At least 2 years of experience in data engineering, database engineering, business intelligence, or business analytics.

Benefits

  • Performance-based incentive awards.
  • Maternity and parental leave.
  • Paid time off.
  • Health benefits.

Related Jobs