Welcome to Silicon Jobs, your connection to the AI Economy > 

Search

Research Scientist (Post-Training)

PublishedPublished: 6/14/2022
Science

Job Description

Post-Training Researcher

\n

Bay Area | Individual Contributor | Growth-Stage AI Company

\n


\n

About the role

\n

A growth-stage AI company is building a decision agent that reasons over structured enterprise data, connecting to existing data platforms and answering complex operational questions with recommended actions, not just retrieving reports. The research team is small and IC-driven, with real ownership over technical direction. They're hiring one strong post-training specialist to help shape how the system reasons and adapts, particularly over structured/relational data.

\n


\n

What you'll do

\n

    \n
  • Design and/or build post-training workflows, RLHF, DPO, GRPO/RFT, or preference optimization, for models that reason over structured, relational data
  • \n

  • Work closely with a small team of ICs on training approach, evaluation, and model behavior
  • \n

  • Help define what "good reasoning" looks like for an enterprise decision agent - not a general-purpose chatbot
  • \n

\n


\n

What we're looking for

\n

    \n
  • Hands-on experience with modern post-training methods (RLHF/DPO/GRPO/RFT, SFT, preference data, reward modeling)
  • \n

  • Strong research judgment on training/alignment approach
  • \n

  • Comfort in a small, fast-moving team with direct influence over research direction, vs. a large-lab structure
  • \n

  • Bonus: experience with structured/tabular/graph data, symbolic reasoning, or knowledge representation
  • \n

\n


\n

Compensation: $200k-$300k base, plus equity

\n


Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...