Mindrift
Mindrift
Mindrift is an artificial intelligence talent marketplace and SaaS platform that connects professional contributors and domain experts with AI training and evaluation projects from technology companies. As part of Toloka, it works with writers, editors, QA specialists, scientists, and other subject-matter experts who help generate training data, assess model safety and reasoning, evaluate media quality, and create real-world scenarios for improving AI performance and reliability. Mindrift recruits and vets contributors through assessments and identity verification, then offers flexible paid project opportunities across AI training and evaluation work.

Senior Python Data Scraping Engineer — Freelance, Poland Remote

Mindrift is seeking a senior Python engineer to build dependable web-scraping workflows for AI projects. The role focuses on extracting, validating, and processing structured data from complex, dynamic websites.

Description

  • Manage complete extraction workflows across complex websites
  • Deliver comprehensive, accurate, and dependable structured datasets
  • Apply tools and tailored workflows to speed up collection, validation, and task execution
  • Extract data reliably from interactive and dynamic web sources
  • Adjust scraping methods for JavaScript-rendered content and evolving site behavior
  • Maintain data quality through validation, cross-source checks, formatting standards, and systematic review
  • Handle large-scale scraping with efficient batching and parallel processing
  • Track extraction failures and preserve stability through minor website structure changes
  • Apply Apify, OpenRouter, and related technologies to web extraction and data processing

Requirements

  • 5+ years of relevant experience in data engineering, web scraping, automation, or software development
  • A bachelor’s or master’s degree in engineering, applied mathematics, computer science, or a related technical discipline is a plus
  • Strong practical foundation in scripting, automation, and data extraction workflows
  • Advanced Python scraping experience with BeautifulSoup, Selenium or comparable tools, JavaScript and AJAX content, infinite scrolling, and proxy-based APIs
  • Ability to extract information from complex hierarchies, archived pages, and inconsistent HTML
  • Experience cleaning, normalizing, and validating data for delivery in CSV, JSON, and Google Sheets
  • Experience working with anti-bot systems and changing website structures at scale
  • Practical experience using AWS or equivalent cloud infrastructure and Docker in production workflows
  • Hands-on use of LangChain, OpenRouter, or similar LLM frameworks for automation
  • High attention to detail and a strong focus on data accuracy
  • Independent working style with the ability to investigate and resolve issues without close supervision
  • Upper-intermediate English proficiency (B2) or higher
  • A GitHub profile is a plus

Benefits

  • Freelance project engagement
  • Part-time workload of approximately 10–20 hours per week during active project phases
  • Remote work from Poland
  • Compensation of up to $40 per hour equivalent, depending on experience level and contribution pace

Related Jobs