DECISION 01
Self-hosting AI inference
Third-party inference was too expensive and slow at scale.
I evaluated third-party API inference at approximately $0.075 per generation and 25 seconds. I designed a self-hosted asynchronous inference pipeline on AWS EC2 GPU infrastructure, bringing inference to approximately $0.0095 per generation and 16 seconds.
- 87% lower inference cost
- 36% lower generation latency
- ~30K virtual try-ons/month
