Machine Learning Engineer — Model Evaluation & Experimentation
Mercor / Machine Learning Engineer
RATE
$60-$90/HR
LOCATION
UNITED STATES
DESCRIPTION
A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation. You will author complex, multi-step ML tasks — for example, taking a vague research idea like "modify how an RL reward is computed," implementing the change, running the training experiment, and analyzing the results to determine success — and verify exactly where frontier models fall short. Each task represents one to two days of continuous, focused effort and spans multiple technical skills: implementation, experiment setup and execution, and rigorous analysis. You will work in a tight feedback loop with the lab's researchers. This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. 2. Key Responsibilities Design tasks: Turn real ML research ideas — like tweaking how an RL reward is computed — into well-defined, multi-step tasks. Run experiments: Implement changes, run training experiments, and analyze the results to show what a correct solution looks like. Explore RL ideas: Build some of your tasks around reinforcement-learning basics such as reward functions and training behavior. Evaluate models: See how frontier models handle your tasks, and note where and why they fall short. Work as a team: Compare notes with researchers and fellow experts so tasks stay consistent, rigorous, and fair. 3. Core
REQUIREMENTS
- ▸MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain.
- ▸1+ years of experience in a research or research-engineering role.
- ▸Hands-on experience training and evaluating ML models and running experiments end-to-end — experiment setup, execution, and analysis — not just using ML libraries superficially.
- ▸Strong familiarity with large language models: their capabilities, limitations, and evaluation techniques.
- ▸Working proficiency in Python and Git, with comfort in both scripting and notebook environments.
- ▸Basic understanding of reinforcement learning (reward functions, policy training) is preferred.
- ▸Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
- ▸A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems.
REFERRAL LINK · OPENS MERCOR · WE MAY EARN A REFERRAL
Or create a free Data Label Jobs profile to save roles.
More Mercor contracts
Cybersecurity Research Expert – Offensive Security & Vulnerability Research
Mercor / Cybersecurity Research Expert – Offensive Security & Vulnerability Research
Finance Specialist — CFA/ACA/ACCA/CPA Required
Mercor / Finance Specialist
Architecture Expert
Mercor / Architecture Expert
Fraud Analyst – Content & Reviews Abuse
Mercor / Fraud Analyst – Content & Reviews Abuse