Back to the board
MercorVerified· Posted 19d ago

SWE-Bench Task Auditor

Commitment
40 hrs/week
Location
Remote
Availability
3 spots left
$70–90 / hr
Apply on Mercor
Applying through our link may earn us a small commission — at no extra cost to you.

About this role

Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.

Basic Qualifications

  • 3+ years professional software engineering
  • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles)
  • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking
  • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)

Preferred Qualifications

  • Familiarity with SWE-Bench (Verified) or similar repository benchmarks
  • Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.)
  • Prior code-review or task-grading experience
About Mercor

The strongest pipeline for credentialed experts. Contracts are clear, and workers report payouts landing on schedule.

Trust score 9.2/10Read our Mercor review →

AITrainerGigs aggregates this listing from Mercor.