Metrics
Datadog · Prometheus · Grafana
CPU · Memory · Latency · Error Rate · SLO
I build, observe and automate production systems across cloud infrastructure, distributed applications and Kubernetes environments.
Site Reliability Engineer focused on AWS, Kubernetes, DevOps, observability, incident management and automation. My work connects application behavior, infrastructure signals and operational response into one reliable production loop.
Metrics, logs and traces form one operational picture. Instrument once, correlate signals, and shorten the path from alert to root cause.
Datadog · Prometheus · Grafana
CPU · Memory · Latency · Error Rate · SLO
Splunk · Kibana · ELK
Application · Infrastructure · Errors · Correlation
OpenTelemetry · Datadog APM
Service A → Service B → Service C → Database
Select any stage to see a realistic command used in a production-style delivery workflow.
git add . && git commit -m "fix: update payment API"
A realistic payment-service incident: detect → investigate → mitigate → recover → prevent.
Manual monitoring → manual investigation → manual deployment → long MTTR.
Python · Bash · Terraform · Ansible · AWS
Observability → correlation → automation → automated response.
Open to opportunities across SRE, DevOps, Cloud Infrastructure and Observability.