Jimmy Technologies
Jimmy Technologies
Jimmy Technologies develops technology solutions for the mobility industry, working across automotive, logistics, and fintech. Its team supports clients from early product concepts through launch with end-to-end product development, cloud services, and technical consulting. The company combines in-house components through its Jimmy Framework with tailored engineering work to help businesses build and launch digital products more efficiently.

Senior Data Engineer, OCR and AI Pipelines (Remote, Czechia)

Build OCR and document-ingestion pipelines that transform unstructured insurance records into structured data for retrieval-augmented generation and anti-financial-crime AI models.

Description

  • Develop scalable ingestion pipelines for large volumes of unstructured insurance documents
  • Create connectors for enterprise sources including SharePoint and email
  • Configure and optimise OCR and document-parsing tools to capture accurate text and document layouts
  • Automate text cleaning, normalisation, semantic chunking, and metadata enrichment
  • Design vector-storage schemas and reliable retrieval systems for downstream RAG models
  • Maintain enterprise security and low-latency SLA standards across document-processing pipelines
  • Implement error monitoring and validation processes to identify low-confidence OCR results
  • Follow robust engineering practices, including version control, CI/CD, and automated testing
  • Convert extensive insurance-document collections into structured, high-quality data for AI systems
  • Support a Dutch insurance client through an AI, automation, advanced analytics, and anti-financial-crime consultancy

Requirements

  • Five to ten years of professional data engineering experience
  • Demonstrated ability to build data-processing and document-ingestion pipelines on public cloud platforms
  • Practical experience handling unstructured files, including PDFs, Word documents, spreadsheets, presentations, scans, and emails
  • Experience developing connectors for enterprise systems such as SharePoint and email
  • Hands-on experience with document-extraction or OCR technologies, such as AWS Textract or an equivalent tool
  • Strong software-engineering discipline with Git, CI/CD, and automated testing
  • Python and SQL are required
  • AWS experience with S3, Step Functions, and CloudWatch is required
  • AWS Textract or equivalent OCR and document-extraction experience is required
  • Experience with unstructured-document processing and ingestion pipelines is required
  • Git, CI/CD, and testing experience are required
  • Banking or insurance experience is advantageous
  • Experience with vector databases and RAG architectures is desirable
  • Azure and Databricks experience is desirable
  • Financial-services or insurance-domain knowledge is desirable
  • Located in Europe; candidates based in Central and Eastern Europe are preferred
  • Availability to start within 30 days

Benefits

  • Long-term remote-first contract engagement
  • Contract through July 2027, with potential for extension
  • Remote work within Europe, with Central and Eastern Europe preferred

Related Jobs