The training platform built for modern SRE roles

Most SREs are strong at incidents and operations. Today's market also expects software engineering instinct, deep observability thinking, and automation judgment. Here you work production incidents, production readiness reviews, and observability gaps — the situations Japan SRE interviews are built around.

Try a scenario

The sample scenario opens without an account.

The gap

Companies now want SREs who can think and build — not just monitor and react

The market has shifted. SRE job descriptions now list software engineering, observability design, toil reduction, and system judgment alongside traditional operations skills.

SLO arithmetic and error budget calls under pressure
Reading OpenTelemetry traces and Prometheus alert rules
Python automation for kubectl loops and log parsing
Postmortem facilitation and reliability ROI to leadership

What a session looks like

A practice loop that builds real confidence

01

See a real scenario

The situation, a strip of live metrics, and where it matters the config or log itself — Terraform state, an Istio VirtualService, a Kafka consumer group — with the problems in it called out.

02

Make decisions and take action

Five to seven tasks: single-answer calls, multi-select that awards partial credit, and a written answer you compose yourself. A hint is there when you want it.

03

Get expert-level feedback

A score per task, then what a strong answer contains, why companies ask this, and what it adds to your range. Your confidence moves in the skill areas the scenario touched.

Skill areas

9 skill areas covering every dimension of modern SRE practice

Confidence grows automatically as you complete scenarios.

Core

SLO Monitoring

High demand

Observability Engineering

Response

Incident Response

Bridge

Engineering for SRE

Resilience

Disaster Recovery

Validation

Chaos Testing

Efficiency

Automation & Toil

PRR

Production Readiness

Scale

Capacity Planning

What changes

You will be able to

Debug a Kubernetes DNS failure from resolv.conf and ndots alone
Set SLOs for a service with no historical data
Run a blameless postmortem when the team is blaming each other
Read a Prometheus alert rule and know whether it will fire
Hold an error budget freeze conversation with a VP of Product
Automate a kubectl toil loop in Python with subprocess

Built for

Ops-heavy SRE

Strong in incidents and production support, wants more engineering confidence.

DevOps → SRE transition

Needs broader reliability practice and stronger market-facing SRE language.

Interview preparation

Preparing for senior SRE roles or lateral moves to stronger positions.

Team levelling up

Reliability chapter leads building broader skills across their squad.

Start where you are

Open the SLO scenario without an account. Or create one, and confidence tracking starts across all nine skill areas.

Try a scenario

Scenarios run 10 to 20 minutes each.