Service Outage Analytics on Databricks Challenge

East European Bank
0 +

Core banking processes brought under unified monitoring

- 0 %

Post-incident analysis reduction time

0 %

Of the outage-related data consolidated into a single Gold layer

Challenge

A large East European bank operating a broad portfolio of digital and branch-based banking services struggled to measure and understand the business impact of service outages in a consistent, reliable way. When disruptions occurred — whether affecting payment processing, digital channels, or core banking operations — different business units were assessing impact using different data sources, different logic, and different definitions of key metrics. This made it nearly impossible to produce a unified picture of how an outage affected customers, transaction volumes, or overall service availability. There was no standardised model for outage analysis, no agreed set of KPIs, and no governed data layer to support cross-functional reporting. As a result, post-incident reviews were slow, inconsistent, and difficult to act on — and the organisation lacked the visibility needed to prioritise remediation efforts based on actual business impact.

Solution

Squery contributed data engineering and architecture expertise as a supporting partner in the client’s initiative to build a unified outage analytics solution on Databricks. Working alongside the client’s internal team, Squery helped design the outage analysis data model and assisted in implementing the full data pipeline within the Databricks environment. The agreed business KPIs — including impacted transaction volume, number of affected customers, and service availability metrics — were translated into a structured, layered data architecture, with PySpark and SQL used to ingest, transform, and consolidate data through to a Gold layer where unified business logic was applied consistently across all monitored banking processes. Throughout the engagement, Squery supported the adoption of CI/CD standards and engineering best practices, including version-controlled notebooks, automated testing, and deployment pipelines, helping the client build a solution that was maintainable and scalable by their own teams going forward.

Benefits

The initiative delivered a single, authoritative framework for measuring service outages and their business impact across the bank’s core processes. With all relevant banking services evaluated against the same KPIs using consistent data and logic, the cross-team inconsistencies that had previously undermined post-incident analysis were eliminated. Business stakeholders gained access to accurate, timely outage impact data, enabling faster reviews and better-informed decisions around service reliability priorities. The Gold layer’s centralised business logic reduced duplication of analytical effort across departments, and the CI/CD-compliant architecture ensured the solution could accommodate new processes or evolving KPI definitions without significant rework. Squery’s involvement helped the client accelerate delivery while embedding sound engineering practices that the internal team could own and build on independently.