Skip to content
View Jkudjo's full-sized avatar
🎯
Focusing
🎯
Focusing

Organizations

@EddieHubCommunity

Block or report Jkudjo

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jkudjo/README.md

Platform Reliability Engineer

AWS EKS | Terraform | Kubernetes | Aurora | Redis | OpenSearch | Prometheus | Grafana


I operate production SaaS infrastructure on AWS at scale.

My focus is Kubernetes reliability, cloud cost reduction, incident response, performance systems, and production-grade observability.

I Build and maintain systems across EKS, Aurora, Redis, OpenSearch, and GitHub Actions.


Proof of work

Area Outcome
Cloud cost optimization Worked with a Qubole team to reduce AWS spend from $375k → $70k through rightsizing, idle resource cleanup, and architecture changes
Availability Maintained 95% +uptime on production SaaS workloads
Data platform Operated Aurora, Redis, and OpenSearch under production load


How I work in incidents

  1. Stabilize — stop the bleed, rollback or scale, protect data
  2. Communicate — clear status, ETA, single owner
  3. Diagnose — metrics, logs, traces, recent changes
  4. Fix — minimal safe change with verification
  5. Prevent — action items, runbook updates, monitoring gaps closed

Contact

Remote · UTC+0 · Open to Senior SRE / Platform Reliability / AWS Infrastructure roles

Pinned Loading

  1. nat-gateway-oauth-incident nat-gateway-oauth-incident Public

  2. All-Things-DevOps All-Things-DevOps Public

    Forked from joseeden/All-Things-DevOps

    All-in-one stop for all my DevOps notes and Projects.

  3. asynqmon asynqmon Public

    Forked from hibiken/asynqmon

    Web UI for Asynq task queue

    TypeScript