Engineering for Reality: Cloud Automation, Production Failures & Resilience
Key Topics
Infrastructure as Code: Automation Without Losing Grip
“Create another environment exactly like production.” Which version of production?
An urgent console fix, a copied configuration and a script that stops halfway can quickly turn infrastructure into a guessing game. Automation helps—but how do we make changes repeatable without losing visibility, ownership and control?
This session follows the journey from manual work → scripts → Terraform → Terragrunt → a full GitOps ecosystem. Through practical examples and guided walkthroughs, we’ll explore reusable infrastructure, environment management, drift detection and controlled execution with OCI Resource Manager. We’ll then connect infrastructure provisioning with Helm and Argo CD, examining how GitOps changes deployment, reconciliation and recovery.
We’ll also discuss Karpenter for OCI and Crossplane, focusing on the ownership boundaries needed when multiple controllers manage infrastructure.
Participants will learn how to:
- Build reusable environments while keeping differences explicit.
- Distinguish configuration drift from changes awaiting deployment.
- Define clear ownership across infrastructure tools and Kubernetes controllers.
- Review changes, limit their impact and plan recovery beyond reverting a commit.
Audience: DevOps and platform engineers, SREs, cloud architects and developers responsible for infrastructure. Familiarity with cloud fundamentals is recommended; Terraform and Kubernetes knowledge is helpful.
Failure Occurs Only in Production
The servers are green. The customer is still waiting. What do you investigate first?
Production brings together traffic, dependencies, configuration and timing in ways that testing rarely reproduces. When something fails, more dashboards do not automatically lead to a better explanation.
Using an illustrative production incident, this session connects metrics, logs and distributed traces to investigate competing explanations, identify bottlenecks and verify recovery. We’ll examine how an apparently slow database operation can actually be time spent waiting for a shared resource—and what evidence separates those possibilities.
Product walkthroughs and screenshots show where OCI Application Performance Monitoring (APM), Oracle Log Analytics and OCI Monitoring fit, alongside OpenTelemetry, Prometheus, Grafana, Loki, Tempo and Jaeger. We’ll also translate user expectations into service indicators, service targets and error budgets.
Participants will learn how to:
- Choose evidence that tests a failure hypothesis.
- Correlate telemetry while accounting for sampling and coverage gaps.
- Define a bounded intervention and a clear recovery test.
- Turn incident findings into reusable operational knowledge.
Audience: Developers, DevOps engineers, SREs, DBAs and architects involved in diagnosing production systems. Basic familiarity with APIs and cloud applications is recommended.
Resilience by Design: Building Resilient Applications on OCI
When critical applications go down, the impact goes far beyond IT. Resilient platforms keep critical services running, protect data and enable rapid recovery when disruption occurs.
This seminar explores how resilience is designed across the OCI stack, covering high availability, networking, data protection, RPO/RTO, and cross-region disaster recovery, followed by a live OCI Full Stack DR demonstration.
Participants will learn how to:
- Design highly available architectures across Compute, OKE, Networking and Storage.
- Protect critical data with backup, recovery and cross-region replication.
- Translate business requirements into the right RPO and RTO.
- Build resilient hybrid connectivity and network architectures.
- Orchestrate cross-region application recovery with OCI Full Stack DR.
- Monitor, test and automate recovery to ensure your DR strategy works when it matters.
Audience: Cloud and infrastructure architects, application architects, database administrators, DevOps and platform engineers, and IT professionals responsible for service availability, business continuity, and disaster recovery.
No mandatory prerequisites; basic familiarity with cloud concepts is required.
Schedule
Seminar Program
Infrastructure as Code: Automation Without Losing Grip
“Create another environment exactly like production.” Which version of production?
An urgent console fix, a copied configuration and a script that stops halfway can quickly turn infrastructure into a guessing game. Automation helps—but how do we make changes repeatable without losing visibility, ownership and control?
This session follows the journey from manual work → scripts → Terraform → Terragrunt → a full GitOps ecosystem. Through practical examples and guided walkthroughs, we’ll explore reusable infrastructure, environment management, drift detection and controlled execution with OCI Resource Manager. We’ll then connect infrastructure provisioning with Helm and Argo CD, examining how GitOps changes deployment, reconciliation and recovery.
We’ll also discuss Karpenter for OCI and Crossplane, focusing on the ownership boundaries needed when multiple controllers manage infrastructure.
Audience: DevOps and platform engineers, SREs, cloud architects and developers responsible for infrastructure. Familiarity with cloud fundamentals is recommended; Terraform and Kubernetes knowledge is helpful.
Failure Occurs Only in Production
The servers are green. The customer is still waiting. What do you investigate first?
Production brings together traffic, dependencies, configuration and timing in ways that testing rarely reproduces. When something fails, more dashboards do not automatically lead to a better explanation.
Using an illustrative production incident, this session connects metrics, logs and distributed traces to investigate competing explanations, identify bottlenecks and verify recovery.
We’ll examine how an apparently slow database operation can actually be time spent waiting for a shared resource - and what evidence separates those possibilities.
Product walkthroughs and screenshots show where OCI Application Performance Monitoring (APM), Oracle Log Analytics and OCI Monitoring fit, alongside OpenTelemetry, Prometheus, Grafana, Loki, Tempo and Jaeger. We’ll also translate user expectations into service indicators, service targets and error budgets.
Audience: Developers, DevOps engineers, SREs, DBAs and architects involved in diagnosing production systems. Basic familiarity with APIs and cloud applications is recommended.
Resilience by Design: Building Resilient Applications on OCI
When critical applications go down, the impact goes far beyond IT. Resilient platforms keep critical services running, protect data, and enable rapid recovery when disruption occurs.
This seminar explores how resilience is designed across the OCI stack, covering high availability, networking, data protection, RPO/RTO, and cross-region disaster recovery, followed by a live OCI Full Stack DR demonstration.
Audience: Cloud and infrastructure architects, application architects, database administrators, DevOps and platform engineers, and IT professionals responsible for service availability, business continuity, and disaster recovery.
No mandatory prerequisites; basic familiarity with cloud concepts is required.


