Itix
Itix
51 – 200 Employees
ConsultingHealthcareLogistics
Itix is a Brazilian technology and consulting company that helps organizations modernize operations through custom software platforms, dedicated development teams, and technology consulting. Its services also include support and analytics or business intelligence, enabling enterprise customers to improve processes and use data more effectively. Itix works with organizations in areas including industry, finance, healthcare, logistics, insurance, energy, chemical, and mining, supported by distributed teams across Brazil.

Mid-Level Data Engineer at Itix (Remote, Brazil)

Join Itix as a Data Engineer building Lakehouse pipelines and RAG solutions. Work with Databricks, Delta Lake, embeddings, data quality, and governance.

Description

  • Build ETL/ELT pipelines in Databricks or open-source frameworks within a Delta Lake medallion architecture
  • Process and standardize PDF and image data with associated metadata
  • Curate and anonymize data under Brazil’s LGPD for RAG, few-shot learning, and test datasets
  • Develop and maintain vector indexes with chunking, embeddings, incremental updates, and version control
  • Create versioned golden datasets for accuracy measurement and benchmarking
  • Design feedback-loop persistence and supply data to the Quality and SLA Dashboard
  • Connect databases with inference services and external APIs while managing performance, cost, monitoring, and alerts
  • Automate workflows through Jobs, CI/CD pipelines, and data quality testing

Requirements

  • Advanced proficiency in Python and SQL
  • Experience with data modeling and medallion or Lakehouse architectures
  • Experience building unstructured data pipelines for PDFs, images, and text, including preparation for RAG
  • Familiarity with vector databases and embeddings such as Databricks Vector Search or pgvector
  • Experience with data quality, testing, and version control
  • Practical Databricks experience with Delta Lake, Workflows or Jobs, notebooks, PySpark, and Spark SQL
  • Experience with Git and CI/CD
  • Experience working in Azure, AWS, or GCP
  • Knowledge of LGPD and sensitive-data handling
  • Unity Catalog experience in a governed environment
  • Experience with MLflow, Databricks Model Serving, or Mosaic AI
  • Experience with Delta Live Tables or Lakeflow and Auto Loader
  • Experience in insurance or financial services
  • Experience with Terraform or Databricks Asset Bundles
  • Databricks Data Engineer Associate or Professional certification

Benefits

  • Geographic flexibility to work from the location that best fits the project
  • Discount partnerships
  • Rewarded employee referral program
  • Birthday day off
  • TotalPass or Gympass access
  • Health and wellness initiatives
  • Access to office relaxation spaces with beanbags, a pool table, video games, a lounge, a fully equipped kitchen, and coffee

Related Jobs