QA/Test Engineer
Mercor / QA/Test Engineer
RATE
$60-$90/HR
LOCATION
UNITED STATES
DESCRIPTION
A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models, and complex multi-step tasks are only useful if they are airtight: unambiguous, correctly graded, and robust to shortcuts. We are seeking experienced QA and test engineers to be the quality backbone of this benchmark — designing the test cases and review processes that guarantee every task measures what it claims to measure. Each task under review represents one to two days of expert effort and spans multiple technical skills, so quality review here means genuinely understanding the task: running it, probing its edge cases, and debugging its environment. You will work in a tight feedback loop with the lab's researchers and task authors. This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. 2. Key Responsibilities Design checks: Create test cases that confirm each task works as intended — including the tricky edge cases. Review tasks: Give tasks and reference solutions a careful read before they're finalized, catching ambiguity and gaps early. Debug: Roll up your sleeves in Python when a task or its checks don't behave the way they should. Shape the process: Help build simple, repeatable quality checklists, and share feedback authors can act on right away. Protect the results: Watch for shortcuts and grading gaps in AI agent runs so benchmark scores stay trustworthy. 3. Core
REQUIREMENTS
- ▸MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain.
- ▸1+ years of experience in test engineering, quality assurance, or a research/software engineering role with strong quality ownership.
- ▸Demonstrated skill designing test cases and quality-review processes, and debugging complex systems end-to-end.
- ▸Working proficiency in Python and Git, and comfort navigating unfamiliar codebases and environments.
- ▸Exceptional attention to detail and clear written documentation habits.
- ▸Past experience in AI training, model evaluation, or quality review of AI-generated work is preferred.
- ▸A perfectionist mindset: creativity in finding what others missed, and the ability to work independently through ambiguous, open-ended problems.
- ▸Ability to engage reliably for approximately 35 hours per week.
REFERRAL LINK · OPENS MERCOR · WE MAY EARN A REFERRAL
Or create a free Data Label Jobs profile to save roles.
More Mercor contracts
Cybersecurity Research Expert – Offensive Security & Vulnerability Research
Mercor / Cybersecurity Research Expert – Offensive Security & Vulnerability Research
Finance Specialist — CFA/ACA/ACCA/CPA Required
Mercor / Finance Specialist
Architecture Expert
Mercor / Architecture Expert
Fraud Analyst – Content & Reviews Abuse
Mercor / Fraud Analyst – Content & Reviews Abuse