Mercor
Mercor
51 – 200 Employees
Mercor’s available profile does not include enough verified information about its products, services, customers, sector, or hiring focus to support an accurate company overview. More source material is needed before describing the business or its opportunities.

AI Safety Expert - English and Punjabi

Remote contract role for English- and Punjabi-fluent AI safety experts supporting Mercor’s conversational AI red-teaming projects. Test jailbreaks, prompt injections, harmful-content risks, and model vulnerabilities while documenting reproducible findings.

Description

  • Conduct adversarial testing of conversational AI models and agents using jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation
  • Create human-generated data by annotating failures, categorizing vulnerabilities, and identifying systemic risks
  • Apply established taxonomies, benchmarks, and playbooks to maintain consistent evaluations
  • Prepare reproducible reports, datasets, and attack cases for customers
  • Assess AI outputs for bias, misinformation, harmful behavior, and other sensitive-content risks
  • Identify weaknesses that automated evaluations may overlook
  • Broaden evaluation coverage to help limit unexpected issues in production
  • Improve customer AI systems through structured adversarial testing

Requirements

  • Native fluency in both English and Punjabi
  • Sound judgment when evaluating language and content
  • Ability to determine whether AI responses are accurate, complete, and appropriate, and explain the assessment
  • Strong attention to subtle errors, inconsistencies, and omissions
  • Consistent adherence to guidelines and quality standards
  • Clear communication of reasoning to technical and non-technical audiences
  • Ability to adapt across projects, task formats, and customers
  • Independent contractor status
  • Must not be an H1-B or STEM OPT candidate
  • Preferred: adversarial machine learning experience with jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction
  • Preferred: cybersecurity experience in penetration testing, exploit development, or reverse engineering
  • Preferred: socio-technical risk experience, including harassment or disinformation probing, abuse analysis, or conversational AI testing
  • Preferred: psychology, acting, or writing experience that supports unconventional adversarial thinking

Benefits

  • Fully remote work
  • Flexible scheduling with the ability to work on your own schedule
  • Weekly payments through Stripe or Wise
  • Projects may be extended, shortened, or ended early based on business needs and performance
  • Optional participation in higher-sensitivity projects
  • Content guidelines and wellness resources provided
  • Competitive pay
  • Reasonable accommodations available upon request
  • Referral payments of up to $90 per successful referral, subject to referral limits

Related Jobs

Kreato Global | BPO and Language Solutions

Remote English-Spanish OPI/VRI Interpreter

Kreato Global | BPO and Language Solutions
201 – 500 Employees
HealthcareHospitalityLogistics

Interpret remotely between English- and Spanish-speaking people in medical, financial, social service, and customer care settings. Provide language support for Kreato Global across Latin America.

Open
SPERTON - Where Great People Meet

Sales Executive, Elevators and Car Parking Systems

SPERTON - Where Great People Meet
51 – 200 Employees

Drive elevator and car parking system sales across Mumbai’s Western Region. Build client and dealer relationships, and manage deals from initial enquiry through project execution.

Open