Streaming SRE team (via Insight Global), keeping large-scale streaming infrastructure reliable.
$ whoami
Lucas Oliveira
Site Reliability Engineer
$ systemctl status availability ● active · Available for opportunities
$ cat about.md
SRE with 7+ years building and running reliable, cost-efficient infrastructure with Kubernetes, Terraform, GCP and AWS. I care about clean Infrastructure-as-Code, solid SLOs and calm on-call.
about
I'm an SRE with 7+ years across DevOps and SRE roles, focused on building and running reliable, cost-efficient systems. My core stack is Kubernetes, Terraform, GCP, AWS and Linux. Along the way I've cut cloud costs (COGS), improved resource allocation and made infrastructure more repeatable and automated.
Much of that came from committing to Infrastructure-as-Code. Keeping Terraform state in sync, speeding up plan runs and avoiding drift-related incidents made day-to-day operations calmer and safer. I rely on well-defined SLOs and SLIs to keep performance aligned with what the business needs.
I work closely with other engineers and share what I learn. I run on Agile with Jira, keep things pragmatic, and try to leave every system a bit more boring and predictable than I found it.
stack
Orchestration & IaC
Cloud
Observability & Reliability
CI/CD & Foundations
experience
Owned shared engineering infrastructure and reliability across staging and production.
- Ran Kubernetes clusters for staging and production: RBAC access control, Velero backup and restore, and Kong/Nginx API gateways.
- Managed Terraform and Ansible IaC, keeping state in sync, speeding up plan execution and cutting drift-related incidents.
- Built and optimized CI/CD pipelines with CircleCI, Jenkins, Cloud Build and Spinnaker.
- Owned observability and on-call with Datadog, Prometheus, Grafana, PagerDuty/Opsgenie and incident.io.
- Administered MongoDB, Postgres and Elasticsearch for IaaS/SaaS platforms on GCP and AWS.
- Cut COGS, optimized Kubernetes resource usage, and replicated the full infrastructure through automation.
Built and managed cloud infrastructure and monitoring for the engineering team.
- Created and managed cloud infrastructure on Google Cloud to improve performance, scalability and availability.
- Ran Kubernetes, Docker, Helm and Drone CI/CD.
- Built monitoring and tooling with Grafana, Prometheus and Stackdriver.
Early-career internships across DevOps and web development on Linux.
education
selected work
Drove Infrastructure-as-Code adoption end to end. Synchronized Terraform state, sped up plan execution and cut drift-related incidents, making infrastructure repeatable and automated.
Cut cloud COGS and improved Kubernetes resource allocation, backed by well-defined SLOs and SLIs that keep performance aligned with business needs.
Multi-environment Kubernetes platform (Kind and GKE Autopilot) provisioned with Terraform and delivered via GitOps: policy-as-code, SLOs, backups and a self-service internal developer platform.
$ ./reach-out.sh --role sre
Let's build reliable systems together
Open to SRE / Platform Engineering roles. Fastest path to me is email.