Mindrift
Mindrift
Mindrift is an artificial intelligence talent marketplace and SaaS platform that connects professional contributors and domain experts with AI training and evaluation projects from technology companies. As part of Toloka, it works with writers, editors, QA specialists, scientists, and other subject-matter experts who help generate training data, assess model safety and reasoning, evaluate media quality, and create real-world scenarios for improving AI performance and reliability. Mindrift recruits and vets contributors through assessments and identity verification, then offers flexible paid project opportunities across AI training and evaluation work.

Senior Python Data Scraping Engineer — Remote in Japan

Build reliable Python web-scraping workflows for Mindrift’s Tendem AI project. Extract, validate, and scale structured datasets from complex, dynamic websites.

Description

  • Lead end-to-end data extraction across complex websites
  • Deliver complete, accurate, and reliable structured datasets
  • Use available tools and custom workflows to accelerate collection, validation, and task execution
  • Extract data from dynamic and interactive sources, including JavaScript-rendered content
  • Adjust scraping methods as website behavior changes
  • Maintain data quality through validation, cross-source checks, formatting standards, and systematic verification
  • Scale large-volume scraping with efficient batching or parallel processing
  • Monitor failures and preserve stability through minor site structure changes
  • Apply Apify, OpenRouter, and other technologies to web extraction and processing

Requirements

  • At least 5 years of relevant experience in data engineering, web scraping, automation, or software development
  • A Bachelor’s or Master’s degree in Engineering, Applied Mathematics, Computer Science, or a related technical field is a plus
  • Strong Python scraping experience with BeautifulSoup, Selenium or similar tools, dynamic content such as JavaScript, AJAX, and infinite scroll, and APIs accessed through proxies
  • Ability to extract data from complex structures, including hierarchies, archived pages, and inconsistent HTML
  • Strong background in data cleaning, normalization, and validation, with experience delivering CSV, JSON, or Google Sheets datasets
  • Experience managing anti-bot mechanisms and dynamic site structures at scale
  • Experience with AWS or equivalent cloud infrastructure
  • Experience using Docker for containerization
  • Hands-on experience with LLM frameworks such as LangChain, OpenRouter, or similar tools for automation
  • Strong attention to detail and commitment to data accuracy
  • Self-directed approach with the ability to troubleshoot independently
  • Upper-intermediate English proficiency (B2) or higher
  • A GitHub link is a plus

Benefits

  • Freelance opportunity
  • Part-time work estimated at approximately 10–20 hours per week during active phases
  • Remote work
  • Opportunity to contribute to innovative technology and AI development projects

Related Jobs

Creative information Technology

Senior Java Developer, Onsite in Virginia

Creative information Technology
11 – 50 Employees
B2BHardwareSecurity

Build scalable Java microservices for CITI, an IT services provider serving government and commercial clients. Develop secure applications for identity, credentials, and access management.

Open
CrowdStrike

Engineer III, Cloud Authentication and Authorization

CrowdStrike

Build authentication and authorization services for CrowdStrike’s global cybersecurity platform. Lead distributed backend development using Go, AWS, Kafka, and related technologies.

Open