Scalable Python microservices on FastAPI, AsyncIO and gRPC inside an event-driven architecture — sustaining high-volume traffic with reliable uptime while cutting response latency.
- Integrated ML and generative-AI endpoints for content recommendation, moderation and metadata summarization into backend services over REST, streamlining inference handling across creator-facing apps.
- Built caching and persistence layers with Redis, PostgreSQL and Cloud Storage for high-volume retrieval APIs — less database load and lower average query latency.
- Automated cloud-native delivery on GCP with Docker, Kubernetes and GitHub Actions, shortening release cycles and hardening rollbacks.
- Rolled out distributed tracing with OpenTelemetry and Grafana Tempo, reducing time to diagnosis during peak traffic.
- Hardened REST and LLM inference endpoints through PyTest suites, profiling and code review, improving code quality.