Computer Vision Engineer (PhD) -- Foundation Models & Edge ML
Job Description
STEALTH ROBOTICS STARTUP — HIRING
\n
\n
Computer Vision Engineer (PhD) — Foundation Models & Edge ML
\n
Remote · Part-time to start (10-20 hours/week), scaling to full-time as we grow
\n
Compensation: $2500 - 5,000 / month + 1-4% equity
\n
\n
\n
ABOUT US
\n
\n
We're a stealth robotics startup building real-time spatial intelligence for physical operations — turning multi-camera and on-robot sensor streams into a live, queryable 3D scene graph of a space, running on NVIDIA edge compute (Jetson-class) with TensorRT foundation models. Small, exceptional, AI-augmented team, targeting a first deployment in 2026. Full product and architecture shared under NDA once we're talking.
\n
\n
You'll work on real sensor data and its messiness — calibration drift, depth noise, motion blur, multi-view ambiguity — and ship code that runs in production on the edge, not just in a notebook.
\n
\n
Critical requirement: you ship production-grade, tested code that runs in real time on real sensor data at the edge — not research notebooks. If it isn't reproducible and benchmarked on target hardware, it isn't done.
\n
\n
\n
THE ROLE — MAKE LARGE MODELS RUN REAL-TIME ON EDGE-DEVICES (e.g NVIDIA AGX JETSONs)
\n
\n
You take modern vision foundation models and make them hit hard latency budgets on edge hardware by giving up the least accuracy possible — the difference between a demo and a shipping robot.
\n
\n
Focus: open-vocabulary detection and promptable segmentation, learned metric depth, vision-language embeddings, and TensorRT/BF16 deployment — ONNX graph-level optimization, custom CUDA plugins, tiling/preprocessing for high-res sources — hitting latency budgets on edge hardware.
\n
\n
\n
WHAT YOU'LL DO
\n
\n
• Design, implement, and productionize perception components that run in real time on edge hardware.
\n
• Own the path from research idea to tested, benchmarked, deployed module (we take testing and reproducibility seriously).
\n
• Improve on existing model benchmarks — find the weak slices (rare classes, hard lighting, motion blur, high-res tiles) and close them with a targeted synthetic-data mix and light fine-tuning, without blowing the latency budget.
\n
• Work across the boundary with robotics, ML-deployment, and platform — integration is where the hard problems live.
\n
• Validate on real rigs (fixed multi-camera rigs + on-robot depth/IMU) and close the loop on real-world failure modes.
\n
\n
\n
MUST HAVE
\n
\n
• ONNX graph-level optimization — graph editing/surgery, operator fusion, and model compression at the graph level (conversion itself is table stakes).
\n
• Model speedup with TensorRT — precision (BF16/INT8), layer and profile tuning, Nsight-class profiling to hit hard latency budgets on edge hardware.
\n
• Custom TensorRT CUDA plugin development and familiarity with CUDA kernels (writing, profiling, debugging).
\n
• Rigorous latency/accuracy benchmarking on target hardware — and improving weak slices with a synthetic-data mix, not just reporting the gap.
\n
\n
Fine-tuning/adapting vision foundation models in PyTorch is useful for patching weak slices — but it is not the core of this role, and large-scale/multi-GPU training is out of scope.
\n
\n
\n
BASELINE REQUIREMENTS
\n
\n
• PhD in EE, Robotics, or ML — or equivalent depth via a strong record of shipped, real-world CV systems.
\n
• Fresh PhDs welcome. Non-PhD candidates need an equivalent record of deployed systems. We calibrate level (early-career to senior) to the person, not to a year count.
\n
• Demonstrated experience: shipped systems - has owned a component end-to-end.
\n
• Expert Python; strong C++ or Rust; genuine software-engineering discipline (tests, code review, real repos — not throwaway scripts).
\n
• Fluency with real sensor data and comfort debugging the physical-world edge cases synthetic data never shows.
\n
\n
Level: some prior production / TensorRT deployment experience preferred (this track benefits from scar tissue).
\n
\n
\n
KEY DELIVERABLES
\n
\n
• PyTorch to ONNX to TensorRT/BF16 engines for the model set.
\n
• Quantized Aware Training / Finetuning of production-ready models.
\n
• Finetune, Prune, re-design, Re-Finetune VLMs for smaller weight targets.
\n
• A reproducible conversion + benchmark harness.
\n
• A latency/accuracy conformance report on target hardware (server + edge).
\n
• Measurable accuracy recovery on identified weak slices (via synthetic-data mix + fine-tuning) that holds inside the latency budget.
\n
\n
\n
BONUS
\n
\n
Deformable-attention / other exotic fused plugins, INT8 PTQ & QAT, structured sparsity/pruning, TensorRT-LLM, open-vocabulary/promptable models, dataset/labeling pipelines.
\n
\n
\n
STACK
\n
\n
Python · Rust (client) · C++ (perf paths) · PyTorch · TensorRT / ONNX (BF16) · OpenCV · ROS 2 / DDS · GPU volumetric fusion · NVIDIA Jetson-class edge compute · stereo/depth cameras · AprilTag/Kalibr.
\n
\n
\n
We optimize for exceptional cross-disciplinary engineers over headcount — if you also span multi-view geometry or scene-graph reasoning deeply, tell us.
\n
\n
How to apply: email jaribido@imasiv.ai with your CV and a short note. Referrals welcome.
