← All articles
// ARTICLE Updated By Steven Robillart · Founder

Data migration plan: the four lines everyone leaves out

PlanningData migrationMethodologyERP

The plan was thirty-four tasks, resource-loaded, with dependencies drawn and a critical path highlighted in red. It had been reviewed by three people and signed off by the steering committee. One line read: “Data cleansing: 3 weeks, migration team.”

That line was wrong in a way no amount of reviewing would have caught. Cleansing was not three weeks of work for the migration team. It was eleven weeks of decisions by four site engineering teams who had never been told they were on the project, about 9,000 duplicate parts where no rule could pick a winner. The migration team could prepare the decisions, route them, and record them. It could not make them.

The plan was not incomplete. It was complete about the wrong things.

What every plan has

Open any data migration plan and you will find the same skeleton: extraction, mapping, transformation rules, load, testing, cutover, hypercare. Durations on each. It is a reasonable description of the work, and it is the part that planning software is good at.

These tasks have their own uncertainties: performance at real volume, target limits, technical dependencies. At least they have a task box where a measurement and a margin can be recorded. Other decisive work often remains buried inside a cleansing or testing line.

Here are four lines to make visible so the date rests on more than technical task durations.

1. Arbitration time, which is not development time

At the end of the data audit you can estimate what share of records pass on a validated deterministic rule and what share needs a human decision. For the calculation, take a 95/5 split on a given scope.

The 95% is engineering. The effort can be estimated and shared where tasks can run in parallel. Adding people will not accelerate a blocked dependency or a target already at its throughput limit.

The 5% is different in kind. Each case needs someone with authority to say which value is correct, and that someone is a business person with a day job. They may only have a few hours a week available. Engineers can prepare the cases, but they cannot replace that authority. Without reserved business capacity, the arbitration queue can become the critical path.

Put it in the plan as its own workstream, with three things attached:

  • A volume, in decisions, not in records. Ten thousand duplicates grouped into 900 decision sets is a 900-item queue, and that is the number that matters.
  • A named owner per data family, on the business side, with the authority to decide. Not a committee. A migration blocked on arbitration is blocked on the org chart, not on technology.
  • A rate, in decisions per week, agreed with those people before the date is committed. In this calculation example, separate from the opening case, the four teams process 80 decisions a week in total. The 900-item queue therefore takes 900 / 80 = 11.25 weeks: it finishes during week twelve, at a steady rate with no new cases. Allocation between teams and absences can extend that duration.

The failure mode here is subtle. Nobody refuses to do arbitration. They just do it at the speed of a secondary task, which is the speed the plan never assumed.

2. The delta window

Every plan contains the word “freeze”. Very few explain what happens on either side of it.

Source data keeps living until writes stop. The delta covers inserts, updates and deletions from the last data position actually covered by extraction through to the freeze. That position needs a consistent marker, such as a log position or extraction watermark; it is not simply the end time of the last load. A missed delta can leave a part absent, a supplier outdated or an order still open in the target.

Your plan needs to answer four questions explicitly:

  • When does the source freeze, and for how long? “The weekend” is not an answer if three plants are in different time zones.
  • What interruption to writes can the business tolerate? Continuous capture can shorten the window, but the transfer of write authority still needs organising. If the measured window does not fit, revisit the mechanism and the cutover strategy before the dry run.
  • How are changes covered from the last extracted position through to the freeze? Through a change log, a reliable timestamp or a consistent full re-extraction and comparison. A timestamp on surviving rows does not reveal physical deletions; a mechanism must retain or detect them. Microsoft distinguishes watermark-based copying from capturing inserts, updates and deletions.
  • Who applies and reconciles the final delta, and by when? Assign an owner, a deadline and evidence that every change through the freeze marker was processed or placed on an explicitly accepted exception list. Any writes exceptionally authorised after the freeze must be recorded and reconciled separately before final acceptance.

The delta can reuse the initial load’s business rules in either a batch or an asynchronous chain. It also requires complete change capture, update and dependency ordering, deletion handling and idempotent replay. Per-record status helps track the result; it does not detect the delta. These targeted recovery capabilities need preparation and testing before cutover.

3. The rollback call

Ask any programme manager what triggers a rollback and you will get a version of “if it goes badly, we roll back”. That is not a decision rule. It is an intention, and at 4am, with a partial load and a business waiting, an intention produces paralysis.

The rollback decision rule needs at least three elements:

  • A person, by name, who makes the call, plus a named substitute. Reachable, awake, and with the authority to overrule the room.
  • A threshold, measurable and validated for the scope. Not “too many errors”. For example: more than 2% of parts rejected after replay, or any rejection in the open orders family. These illustrate a rule to agree, not a universal tolerance.
  • A time, on the clock. “If we have not reached the go/no-go checkpoint by 05:00, we roll back” is a decision that can be made calmly on Tuesday and executed mechanically on Sunday.

This rule avoids discovering at 4am who can stop the cutover. It must refer to a tested, timed recovery procedure. The deadline must allow enough time to restore service. After writes in the target, it must also explain how to preserve or reconcile them; AWS explicitly distinguishes this from rolling back without new data.

4. Targeted iterations and full rehearsals

If I were handed a plan to review, I would look for both types of run, their objectives and the time reserved for corrections.

Run / Fix loops can start on a functional scope as soon as its rules are available. They resolve rejections and verify corrections with their dependencies. Full rehearsals then exercise the entire chain at real volume: initial load, final delta, business controls and cutover duration. They need reserved windows, environments and participants.

There is no universal threshold of twenty or thirty passes. Two or three full rehearsals may be sufficient on a scope already qualified through many targeted iterations; they may also leave major risks open. Judge the coverage achieved, result stability, ability to meet the window and remaining anomalies. Our article on what a passing test actually proves explains these criteria.

Measure a full rehearsal and its correction time to check what actually fits the calendar. Targeted recovery reduces the cost of intermediate iterations, but does not replace rehearsing the complete path. If the plan leaves no window to correct and recheck a late discovery, it lacks a concrete margin.

The plan is written after the audit, or it is provisional

A date committed before the audit is a date committed to an unmeasured scope. It is not a plan, it is a hope with a Gantt chart around it.

This is uncomfortable because programmes want a date early, and the audit sits at the start when nobody wants to spend three weeks measuring. The workable compromise is to say it out loud: publish a provisional plan with the phases and the dependencies, mark every duration as unqualified, and commit dates only after the audit returns volumes, anomaly rates and the arbitration queue size. A plan that says “eleven weeks of arbitration, measured” survives a steering committee. A plan that says “three weeks of cleansing, assumed” survives until the first real load.

Plan and checklist are not the same document

Worth separating, because they get conflated. The checklist is what you verify: extraction complete, duplicates counted, history scope decided, rollback threshold defined. There is one at the end of the ERP data migration guide, and it is the right thing to run against your scope before committing.

The plan is what you sequence: who does what, in which order, with what duration and which dependency. A checklist item like “reserve arbitration time from the business” becomes, in the plan, a named workstream with a volume, an owner and a weekly rate.

Running the checklist tells you whether your plan is missing something. It does not write the plan. The order is: audit, then checklist, then plan, then dates. Programmes that go audit, plan, dates, checklist find out in month four that the checklist would have moved the date, and by then the date is a commitment rather than an estimate.

The full five-phase sequence these fit into is on our methodology page.

// END_OF_DOCUMENT Discuss a migration project→