Define the downtime and consistency boundaries first
Do not begin by choosing a migration tool. Begin with what the business can tolerate. The two useful boundaries are the recovery time objective, which limits how long the service may be unavailable, and the recovery point objective, which limits how much data may be lost. An order, payment, or inventory system may require every committed transaction to survive. A reporting database or historical archive may tolerate a longer gap. Those requirements determine whether an offline migration is sufficient or whether continuous replication is necessary.
The downtime window must cover more than copying bytes. It includes stopping application writes, draining queues and scheduled jobs, applying the final changes, updating connection settings, restarting services, clearing stale connections, validating business functions, and restoring the old environment if the cutover fails. If the approved window is two hours, the transfer cannot consume the entire two hours. Validation and rollback need reserved time, with explicit checkpoints for continuing or stopping.
- Business impact: Identify which workflows must stop, which can operate in read-only mode, and which batch jobs can wait.
- Data tolerance: Specify acceptable loss, replication delay, and whether transaction ordering must be preserved.
- Dependencies: Inventory ERP, CRM, LINE services, reporting, ETL, schedulers, and external systems that connect directly.
- Authority: Name the people who approve the cutover, accept validation results, and order a rollback.
Choose a migration pattern based on change rate, not fashion
An offline migration is straightforward: stop writes, export or back up the source, transfer the data, restore it in the cloud, and redirect applications. This is appropriate when the database is modest, downtime is acceptable, or the system is not operationally critical. Its main risk is an optimistic schedule. Export time, network throughput, restore time, index construction, statistics updates, and database recovery all matter. A rehearsal must measure the time until the target can serve production queries, not merely the time required to upload a backup.
When a long outage is unacceptable, teams commonly perform an initial full load and then keep the target current through change data capture, transaction-log shipping, or native replication. At cutover, they briefly stop source writes, wait for replication lag to reach the agreed checkpoint, confirm the final log position, and redirect the applications. This shortens the outage but increases preparation work. Schema changes, unsupported data types, large transactions, sequence values, replication interruptions, and network instability all require documented handling.
Dual writing is sometimes proposed as the default answer to zero downtime, but it moves consistency risk into the application. One write can succeed while the other fails; retries can create duplicates; transactions may arrive in different orders; and an older service may continue writing only to the on-premises database. Unless the application already supports idempotency, reconciliation, and compensating actions, a short read-only period with continuous replication is often safer than introducing dual writes specifically for the migration.
Turn the cutover into an executable runbook
Rehearse the entire procedure with production-like volume before the actual event. The runbook should list every step in time order, its owner, the command or control surface, the expected result, the verification method, the maximum wait time, and the response to failure. Avoid vague instructions such as “confirm the data.” Define concrete checks: latest transaction timestamps, row counts for critical tables, sampled primary keys, financial or inventory totals, referential integrity, sequence positions, and the results of important queries.
On cutover day, stop nonessential batch jobs, ETL processes, and third-party writers first. Put the core application into maintenance or read-only mode, drain active connections, record the source log position, and wait for the target to catch up. After changing the connection, use controlled accounts to perform low-risk reads and writes before restoring background jobs and customer traffic. Connection pools may retain old sessions, so changing an environment variable is not enough. Confirm service restarts, secrets, routing, firewall rules, and DNS or client-side caches.
- Before cutover: Freeze schema changes, prove that backups restore, inspect replication health, and notify dependent teams.
- During cutover: Stop writes, record the final checkpoint, complete synchronization, update connections, and run smoke tests.
- After cutover: Watch errors, lock waits, connection counts, slow queries, resource pressure, and critical business transactions.
- Evidence: Preserve timestamps, validation output, change records, and approval decisions for later investigation.
Set rollback thresholds before the move becomes irreversible
A rollback plan must exist before the window opens. Define objective triggers such as failed core writes, inconsistent validation results, latency beyond the service requirement, or a critical integration that cannot connect. Also establish a latest rollback time. Once the cloud database accepts new production writes, pointing applications back to the old database may discard those transactions. At that moment, rollback changes from a connection update into a reverse-synchronization and reconciliation exercise.
A common safe posture is to keep the on-premises database intact but read-only during early validation, together with its final checkpoint and a tested backup. If the target fails before meaningful new writes occur, the old connection can be restored quickly. If new transactions already exist in the cloud, the team needs a prepared method such as reverse replication, event replay, or controlled manual reconciliation. A procedure that cannot be completed reliably inside the available window is not a usable rollback plan.
Do not decommission the source immediately after a successful switch. Observe at least a complete operational cycle and verify backups, monitoring, access controls, audit records, scheduled jobs, integrations, and disaster recovery before retirement. The migration is complete only when normal operations and failure recovery have both been demonstrated. Where several enterprise systems are involved, an integration team can help align application, network, database, and business owners around one runbook and one set of acceptance criteria.