About

Every stage taught me a different way to ask "what actually happened?"

I began my career in software testing, trying to break systems before customers did.
Every step since—from automation and reliability engineering to incident engineering—has been driven by the same curiosity: understanding what actually happened, why it happened, and how better decisions can prevent it happening again.

Software Testing

Learning to break things on purpose

Testing taught me that a system's behaviour under pressure is never quite what the documentation says. I got good at finding the gap between the two.

Automation

Making the breaking repeatable

Automation turned one-off investigations into something that could run every day, at scale — and taught me how much trust an automated system has to earn before people rely on it.

Operational Thinking

Moving from "does it work" to "how does it fail"

Somewhere in here the question changed. It stopped being about whether a system passed a test, and started being about how it behaved once real operators and real load got involved.

Reliability Engineering

Designing for the failure, not just the feature

Reliability work is where operational thinking gets formalised — readiness reviews, failure patterns, the unglamorous work of making sure a system degrades gracefully instead of catastrophically.

Incident Engineering

Treating the incident itself as the subject

Eventually I stopped treating incidents as interruptions and started treating them as the richest source of information a team has. Incident engineering is the discipline of actually using it.

Forward Deployed Engineering

Working alongside the people solving the problem

Incident engineering taught me how organisations respond under pressure. Forward Deployed Engineering extends that thinking by working directly with customers and engineering teams to solve operational challenges in real environments, where technology, people and business priorities meet.

AI-assisted Operational Judgement

Helping humans make better operational decisions

The newest stage: using AI to help investigate faster and surface patterns a tired on-call engineer might miss — without letting it make the call.
The judgement stays human. The investigation gets faster.

Why you care

Why Bernalo Exists.

Most organisations learn something from every production incident. The challenge is turning those lessons into better operational judgement rather than another document that nobody reads.
Bernalo exists to help engineering teams investigate incidents, improve decision-making, and build more reliable systems by treating operational judgement as a capability that can be developed—not simply an instinct gained through experience.