← All case studies
Case studies · Insurance

From Azure Synapse to Databricks in three months

How the entire decision package was produced in days rather than months, with quality proven object by object.

A leading Nordic insurer needed to migrate its analytics platform, 16 source systems, more than 350 million rows per day and 124 Power BI models, from Azure Synapse to Databricks. The contractual requirement: one hundred per cent report correctness, without a single report changing for its users.

About the engagement

Client
Leading Nordic insurer, anonymised
Industry
Insurance
Engagement
Analytics platform migration, Azure Synapse to Databricks
Scope
16 source systems, more than 350 million rows per day, 124 Power BI models
Duration
Three-month planned migration, five people at roughly half time
Tech stack
Azure Synapse, Databricks, dbt, Power BI
Status
Planning and design complete and operator-reviewed. Delivery under way
Method
RAID
The challenge

Nobody knew exactly what was still running

The client's data platform had grown over many years: 400 pipelines, around 730 SQL scripts, 165 notebooks and a reporting layer where 91 of 124 Power BI models carried business logic of their own. The platform was to move to Databricks with dbt, with no new functionality, no changes for report users, and with parallel running and business sign-off per increment.

The risk profile was the classic one for legacy migrations: nobody knew exactly what was still running, which dependencies existed, or where the history actually lived.

The solution

Map everything, before anything is built

RAID's first phase is an exhaustive, AI-driven assessment of the current state, not a sample. Every pipeline, every SQL script and every report model was inventoried and cross-referenced: more than 3,500 dependencies between scripts and sources, a further 1,200 between reports and tables. The mapping found what manual reviews usually miss:

  • 187 of 400 pipelines were dead, so almost half the migration surface could be cut before a single line was built.
  • One of the source systems turned out to be four separate subsystems on different schedules, which reshaped an entire migration phase into four smaller ones, each with its own cutover decision.
  • The history existed only in the legacy platform and could not be recreated from the sources, so it had to be copied. Without that discovery, every historical report figure would have been wrong.
  • Parallel validation with two independent source reads would have produced false failures because of timing differences between the reads. The design was reworked to a single-extract principle where both platforms validate exactly the same data.
  • Five report models pointed at an already decommissioned component, were confirmed as decommissioned and removed from scope.

Every finding was reviewed and decided by the client's operator at a formal gate. The AI delivers the material, the human owns the decision.

The delivery

A complete decision package in days

In days rather than months, RAID produced and revised a complete, internally consistent decision package: a current-state analysis in five versions as new facts were confirmed, an architecture with 14 documented decisions including mechanism diagrams, a migration plan with 56 units where each unit carries its own quality checks and a definition of done, a schedule packaged into the client's own two-week sprints, and a data migration and validation design.

All of it versioned, traceable and operator-reviewed.

The plan

Quality proven, not promised

The quality floor is untouched in the plan: every migrated object is compared row by row against legacy, every increment runs in parallel with business sign-off, and every cutover is a separate human decision backed by evidence. The speed comes from never having to redo the work.

In brief
  1. 01

    Three-month planned migration, against roughly 13.5 months for the same work done manually

  2. 02

    Around 900 hours of human working time, five people at roughly half time

  3. 03

    One hundred per cent report correctness, proven row by row per object

Facing a similar challenge?

What would a mapping find in your platform?