Failure Occurs Only in Production

Main Speaker

Learning Tracks

Course ID

4394

Date

16.11.2026

Time

Daily seminar
9:00-16:30

Location

Daniel Hotel, 60 Ramat Yam st. Herzliya

Overview

In distributed systems, identifying the root cause of a problem can be challenging-is it the application code, the database, the network, an external service, or simply a configuration issue? This session explores how to build observability into the application and infrastructure from the start, giving development and operations teams the context they need to troubleshoot complex environments quickly and effectively. We’ll cover practical approaches for combining logs, metrics, and distributed tracing, propagating correlation IDs across services, monitoring the actual user experience, and identifying bottlenecks and slow transactions. The session will also demonstrate how OpenTelemetry can provide a consistent observability framework across modern architectures and how SLI, SLO, and error budgets can help teams move from reactive monitoring to measurable, service-level reliability.

Who Should Attend

Prerequisites

Course Contents

Software Engineering 16.11.2026 סמינר 4394
פתוח להרשמה

Failure Occurs Only in Production

רכשו אונליין

About the Seminar

In distributed systems, identifying the root cause of a problem can be challenging-is it the application code, the database, the network, an external service, or simply a configuration issue?
This session explores how to build observability into the application and infrastructure from the start, giving development and operations teams the context they need to troubleshoot complex environments quickly and effectively.
We’ll cover practical approaches for combining logs, metrics, and distributed tracing, propagating correlation IDs across services, monitoring the actual user experience, and identifying bottlenecks and slow transactions.
The session will also demonstrate how OpenTelemetry can provide a consistent observability framework across modern architectures and how SLI, SLO, and error budgets can help teams move from reactive monitoring to measurable, service-level reliability.

Schedule

Seminar Program

1 Lecture
11:3013:00

Failure Occurs Only in Production

Or Ben David מק״ט 4394
03-7100780 וואטסאפ