Mindrift
Mindrift
Mindrift is an artificial intelligence talent marketplace and SaaS platform that connects professional contributors and domain experts with AI training and evaluation projects from technology companies. As part of Toloka, it works with writers, editors, QA specialists, scientists, and other subject-matter experts who help generate training data, assess model safety and reasoning, evaluate media quality, and create real-world scenarios for improving AI performance and reliability. Mindrift recruits and vets contributors through assessments and identity verification, then offers flexible paid project opportunities across AI training and evaluation work.

Senior Python Data Scraping Engineer in Mexico (Remote)

Build and scale dependable web-scraping workflows for Mindrift’s AI platform. Extract, validate, and deliver structured data from complex, dynamic websites.

Description

  • Manage end-to-end extraction workflows across complex websites, delivering complete, accurate, and reliable structured datasets.
  • Use available tools and custom workflows to accelerate collection, validation, and task execution against defined requirements.
  • Extract data reliably from dynamic and interactive sources by adapting to JavaScript-rendered content and evolving site behavior.
  • Maintain data quality through validation checks, cross-source consistency reviews, formatting requirements, and systematic pre-delivery verification.
  • Scale large-volume scraping through efficient batching and parallel processing.
  • Track extraction failures and preserve operational stability when site structures change slightly.
  • Work with Tendem Agents that perform repetitive tasks within Mindrift’s hybrid AI and human operating model.
  • Apply tools including Apify, OpenRouter, and other technologies alongside technical expertise and tailored methods.
  • Use web-scraping, data-extraction, and data-processing expertise to produce accurate, dependable, high-quality results.

Requirements

  • At least five years of relevant experience in data engineering, web scraping, automation, or software development is required.
  • A bachelor’s or master’s degree in engineering, applied mathematics, computer science, or a related technical discipline is advantageous.
  • Strong Python scraping experience with BeautifulSoup, Selenium, or comparable tools, including JavaScript, AJAX, infinite scroll, and proxied APIs.
  • Ability to extract information from complex structures such as hierarchies, archived pages, and inconsistent HTML.
  • Practical experience cleaning, normalizing, and validating data for structured delivery in CSV, JSON, or Google Sheets.
  • Experience managing anti-bot measures and changing website structures at scale.
  • Hands-on experience with AWS or comparable cloud infrastructure and Docker in production workflows.
  • Experience applying LLM frameworks such as LangChain or OpenRouter to automation tasks.
  • Exceptional attention to detail and a consistent focus on data accuracy.
  • Ability to work independently, investigate issues, and troubleshoot effectively.
  • Upper-intermediate English proficiency at B2 level or higher is required.
  • A GitHub profile link is advantageous.

Benefits

  • Freelance engagement.
  • Part-time remote work.
  • An estimated 10–20 hours per week during active project phases.

Related Jobs

Talkdesk

Senior Solution Consultant, CXA and CCaaS - United States

Talkdesk

Lead strategic implementations of Talkdesk CXA and CCaaS solutions, helping customers integrate the technology and achieve lasting adoption. Partner with customers, executives, and internal teams to guide implementation strategy and long-term success.

Open
Naveera Technology LLC

Senior .NET Core Engineer – React and Claude AI

Naveera Technology LLC
201 – 500 Employees
ConsultingHealthcareLogistics

Lead full-stack engineering work across insurance and payments platforms using .NET, React, AWS, and Claude AI. Drive architecture, batch processing, and distributed-systems initiatives.

Open