Job Description
Job DescriptionAbout the Role
This is a foundational engineering role at an early-stage AI consumer hardware and software startup, where you will own the transcription pipeline end-to-end. You will work hands-on with product and general management leadership to build, tune, and ship a cloud-based ASR system with a narrowly scoped on-device component. Your work directly shapes how well the core product experience feels to real users.
What You'll Do
-
Build and iterate on the cloud-based ASR pipeline, from audio capture through post-processing, running in production at scale.
-
Own ASR quality and reliability end-to-end, shipping measurable improvements across latency, small-word accuracy, and voice-print reliability.
-
Work across data preparation, model training and fine-tuning, evaluation, and deployment to translate product feedback into shipped pipeline changes.
-
Collaborate with a Partner Product Engineer on shared backend and pipeline surfaces.
-
Coordinate across time zones with R&D, hardware, and supply-chain teams based in China.
-
Operate with minimal specification, turning informal asks into concrete, shipped improvements.
What We're Looking For
-
3 or more years building and tuning transcription and ASR pipelines end-to-end in production, primarily in cloud-based settings.
-
Demonstrated ownership of production ASR systems across the full lifecycle: data preparation, model training and fine-tuning, evaluation, and deployment.
-
Experience building and optimizing latency-sensitive or streaming audio and ASR pipelines.
-
Track record of making latency, accuracy, and reliability tradeoffs based on real user feedback.
-
Experience debugging and tuning transcription quality issues in production environments.
-
Comfort shipping in early-stage or founding engineering environments with small teams and limited specification.
-
On-device or embedded ML experience using frameworks such as Core ML or TensorFlow Lite.
-
Prior experience with wearable, hardware, or robotics device products.
-
Background at AI-native consumer applications focused on transcription or audio.
-
Experience building agent or LLM-based product features including tool use, memory, or retrieval systems.
-
Ability to work hybrid three days per week in the San Francisco Bay Area.
-
Ability to collaborate asynchronously with international teams across time zones.
Compensation & Benefits
Salary range: $150,000 to $200,000 USD annually. Visa sponsorship is not available for this role.
Location
Hybrid, three days per week on-site in the San Francisco Bay Area, California, United States.
