Welcome to Silicon Jobs, your connection to the AI Economy > 

Search

Reinforcement Learning Environment Engineer (Contract)

PublishedPublished: 6/14/2022

Job Description

About the role:

\n

Cobalt is seeking people who can build the environments frontier labs train and evaluate agents in: tasks with real difficulty, unambiguous success conditions, and scoring that survives contact with a capable model.

\n

This opportunity is suited to reinforcement learning researchers, research engineers, simulation and tooling engineers, and people who have built serious benchmarks, competition problems, or training environments. Depth in RL is valuable, but so is the engineering discipline required to make an environment reproducible and hard to game.

\n

You do not need prior experience in data annotation. What matters is that you can take a domain, decide what a meaningful task in it looks like, and build something that measures it correctly.

\n


\n

What you'll do:

\n

Depending on the project, you may:

\n

    \n
  • Design and build task environments with programmatic success criteria, including multi-step and tool-using tasks that cannot be solved by a shortcut
  • \n

  • Specify reward functions and partial-credit schemes, and stress-test them for the ways a capable agent would exploit them rather than solve the task
  • \n

  • Produce written reasoning traces and reference solutions showing how a competent human works through the tasks you build
  • \n

  • Evaluate agent trajectories, identifying the specific step at which behavior goes wrong, and classifying failures into a consistent taxonomy
  • \n

  • Assess whether a scored result reflects genuine task completion, and flag cases where the environment or the metric is measuring the wrong thing
  • \n

\n

Projects follow their own guidelines, formatting conventions, and quality standards, and you will work with feedback from reviewers and lab research teams.

\n


\n

Required qualifications:

\n

    \n
  • Direct experience with reinforcement learning, agent evaluation, simulation, or benchmark and environment construction, whether in research, industry, or substantial open-source work
  • \n

  • Strong software engineering ability in Python, sufficient to build reproducible environments, harnesses, and automated scoring
  • \n

  • A PhD in a quantitative discipline, or equivalent depth demonstrated through published work, open-source contributions, or production systems
  • \n

  • Understanding of reward hacking and specification gaming, and the instinct to look for them in your own designs before someone else does
  • \n

  • Ability to explain each step of your reasoning and design decisions clearly in writing
  • \n

\n


\n

Why join Cobalt AI:

\n

    \n
  • Advance frontier AI where it counts. Apply your expertise to data that frontier labs cannot obtain any other way, where your reasoning directly shapes how the next generation of models works through technical problems.
  • \n

  • Grow professionally. Expand your influence through evaluation projects, advisory roles, and research collaborations, while developing a working understanding of how frontier models are trained and assessed.
  • \n

  • Work with a top-tier network. Collaborate with researchers and engineers from leading institutions and labs on high-impact, flexible work.
  • \n

  • Set your own schedule. Flexible 10 to 40 hour weeks that fit around your existing work and your life.
  • \n

  • Competitive pay. Rates vary by project and are determined by a number of factors, including scope, skillset, and experience.
  • \n

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...