Skip to content
DATA LABEL JOBS
MERCORREF#R-72332H AGO

Machine Learning Engineer — Model Evaluation & Experimentation

Mercor / Machine Learning Engineer

RATE

$60-$90/HR

LOCATION

UNITED STATES

DESCRIPTION

A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation. You will author complex, multi-step ML tasks — for example, taking a vague research idea like "modify how an RL reward is computed," implementing the change, running the training experiment, and analyzing the results to determine success — and verify exactly where frontier models fall short. Each task represents one to two days of continuous, focused effort and spans multiple technical skills: implementation, experiment setup and execution, and rigorous analysis. You will work in a tight feedback loop with the lab's researchers. This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week. 2. Key Responsibilities Design tasks: Turn real ML research ideas — like tweaking how an RL reward is computed — into well-defined, multi-step tasks. Run experiments: Implement changes, run training experiments, and analyze the results to show what a correct solution looks like. Explore RL ideas: Build some of your tasks around reinforcement-learning basics such as reward functions and training behavior. Evaluate models: See how frontier models handle your tasks, and note where and why they fall short. Work as a team: Compare notes with researchers and fellow experts so tasks stay consistent, rigorous, and fair. 3. Core

REQUIREMENTS

  • MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain.
  • 1+ years of experience in a research or research-engineering role.
  • Hands-on experience training and evaluating ML models and running experiments end-to-end — experiment setup, execution, and analysis — not just using ML libraries superficially.
  • Strong familiarity with large language models: their capabilities, limitations, and evaluation techniques.
  • Working proficiency in Python and Git, with comfort in both scripting and notebook environments.
  • Basic understanding of reinforcement learning (reward functions, policy training) is preferred.
  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems.
Apply on Mercor

REFERRAL LINK · OPENS MERCOR · WE MAY EARN A REFERRAL

Or create a free Data Label Jobs profile to save roles.

More Mercor contracts