Search

Principal Research Scientist Speech Voice Foundation Models

PublishedPublished: 6/14/2022
Healthcare

Job Description

Principal Research Scientist Speech & Audio Foundation Models

\n

Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.

\n


\n

$270,000–$500,000 base plus bonus, equity and benefits (US).

\n

Relocation assistance available. Visa transfer supported

\n

San Francisco on-site preferred | Remote considered in the US, UK and parts of Europe.

\n

Permanent, full-time.

\n


\n


\n

A top end research lab building realtime voice models text-to-speech, speech-to-text and speech-to-speech delivered as an API. The models run in production behind consumer applications used at very large scale, across health, learning, therapy, companionship, media and gaming.

\n


\n

Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.

\n


\n


\n

The role

\n


\n

Build the models that are the product! This is full-stack research ownership: you frame the question, run the experiments, and ship the result. Research is only finished when it is in production and measurable.

\n


\n


\n

Responsibilities

\n


\n


\n

    \n
  • Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.
  • \n

  • Build and improve voice and speech models across TTS, STT and speech-to-speech.
  • \n

  • Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis. Evaluation is treated as a research product in its own right, not as a pre-launch checkbox.
  • \n

  • Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.
  • \n

  • Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.
  • \n

\n


\n


\n

Essential

\n


\n


\n

    \n
  • Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling. Fine-tuning or building on top of someone else's model is a different discipline and is not what this role is.
  • \n

  • Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count. Text-only research does not transfer.
  • \n

  • Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.
  • \n

\n


\n


\n

Desirable

\n


\n


\n

    \n
  • Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.
  • \n

  • Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.
  • \n

  • PhD in ML or NLP, or equivalent practical experience you can point to.
  • \n

  • Frontier exposure: multimodal, agents, tool use, test-time compute.
  • \n

  • Public work: side projects, open-source, technical write-ups.
  • \n

\n

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...