The migrated pipeline finishes successfully. Every task is green. Yet the finance dashboard shows a different revenue total from the old warehouse. Is the difference a bug, a source defect the migration uncovered, or a deliberate change in the calculation?
Until someone can answer that question with evidence, the workload is not ready for cutover. Code conversion and a successful run prove that the target executes.
Migration acceptance requires more: correct data and business outputs, acceptable performance, reliable incremental behaviour and approval from the people accountable for the result.
Databricks’ SQL conversion guidance makes the same distinction between syntax checks and comparing the meaning of the original and converted queries.
The unit of acceptance is usually a workload, including its inputs, transformations, downstream consumers and operating requirement. A platform-level green status leaves an untested month-end report or an unproven streaming recovery path unresolved.
For the wider build and deployment sequence, see Dateonic’s Databricks implementation step-by-step guide. This guide begins once a target workload exists and needs to be validated.
For each workload, follow a traceable sequence:
Baseline → test case → execution → comparison → tolerance → exception → evidence → acceptance decision

Establish a baseline you can actually compare
Record the source workload’s outputs and behaviour before using it as an oracle. Identify the relevant datasets, schema versions, historical coverage, refresh cadence, downstream reports and business KPIs.
Capture ordinary and peak volumes, important reporting periods, query and job runtimes, freshness, and how the workload recovers after failures.
Choose cases that expose the workload’s real behaviour: a month-end close may matter more than a random weekday sample. Align source and target to the same logical snapshot or processing window.
Record the input cutoff, time zone, late-arrival policy, code versions and comparison timestamp. Otherwise, the systems may simply have observed different data.
The legacy result is evidence, not an infallible truth. If the source contains a known duplicate or a broken rule, decide whether the target is expected to preserve or correct it. Document the desired output and who approved that choice before labeling every difference a migration defect.
Test structure, then reconcile the data
Start with the contract consumers rely on: expected objects, column names and types, nullability, decimal precision and scale, keys, timestamp and time-zone semantics, and historical partitions where relevant.
A table can contain the right number of rows and still break a report through a changed type or missing period.
Reconciliation then moves from cheap coverage checks toward business meaning.
A Databricks Community technical article on migration validation discusses checks across ingestion and transformed layers; AWS’s data parity example separates historical, incremental and functional-query comparisons.
| Comparison layer | Example test | What it can reveal | What it cannot prove alone |
|---|---|---|---|
| Volume | Counts by table, date and business partition | Missing or unexpectedly added batches | Matching counts can conceal changed values or offsetting omissions and duplicates. |
| Aggregate | Sums, min/max, null rates and distinct counts by relevant group | Material shifts in measures or distribution | Opposing errors may cancel out in a total. |
| Record | Compare stable business keys and canonicalized values; inspect missing, extra, changed and duplicate records | Which records differ and how | Raw hashes are misleading if ordering, formatting, null handling or precision differ legitimately. |
| Semantic | Compare invoice totals, inventory availability or named report cells | Whether consumers receive the intended result | A matching headline KPI does not prove every underlying record is correct. |
Apply the deeper checks according to materiality and feasibility. For a critical financial output, exact key-level or value-level evidence may be required.
For a large analytical workload, full population counts plus targeted record comparisons and business-output checks may be a practical design, if the owner agrees to that coverage. Sampling still leaves the completeness of unexamined records unproven.
Compare like with like. Normalize keys and data types deliberately, sort or compare by key rather than row order, and define how nulls, decimal scales and timestamps are represented before calculating hashes.
Preserve the query, input window and results so another reviewer can reproduce a failure.
Put tolerances in the test design, not in the explanation after a failure
Some fields should match exactly: for example, a stable order identifier and a legally required amount at agreed precision. Elsewhere, a bounded difference may be justified by floating-point arithmetic, rounding, late records or an approved change to transformation logic.
These are possibilities to investigate, not blanket excuses.
A tolerance needs a unit, scope and owner. “Within 1%” is meaningless without specifying which measure, population, time window and treatment of small denominators.
Technical tolerance and business tolerance can differ: a small aggregate variance may still affect a regulated report or a single high-value customer account.
For each material comparison, store a record such as this illustrative example, not a default standard:
| Test | Baseline | Target | Approved tolerance | Result | Exception | Owner |
|---|---|---|---|---|---|---|
| Closed-month invoice total, EUR | 2,400,000.00 | 2,400,000.00 | Exact at cents | Pass | None | Finance data owner |
| Open-day order count at 12:00 UTC | 8,412 | 8,406 | Compare again after agreed late-arrival window; no unclassified missing orders | Pending | Six source events still in transit | Commerce owner |
The six-order gap remains pending despite its known timing explanation; the comparison needs a retest. Approval belongs to the workload or data owner, not solely to the engineer implementing the comparison.
Separate reconciliation from data-quality rules
Reconciliation asks whether target and expected results agree. Data-quality testing asks whether the resulting dataset satisfies rules for its intended use. Source and target might match while both violate a business rule.
Correcting a source defect, meanwhile, may create a legitimate difference.
Test relevant rules for completeness, uniqueness, valid domains, referential integrity, cross-table consistency and freshness. Examples include a non-null order ID, one current record per business key, a valid currency code and a product reference that resolves to the agreed catalog.
Specify which violations stop acceptance, which are quarantined for review and which are measured against an owner-approved threshold.
Databricks Lakeflow pipeline expectations can report violations and, depending on the chosen rule, warn, drop records or fail an update. Use them to execute quality rules while retaining separate source-to-target comparisons and business approval.
Keep the tested rule and its result in the acceptance evidence regardless of implementation tool.
Prove incremental, stateful and failure behaviour
After a clean historical reload, exercise the changes that the workload must handle: new records, updates, deletes, duplicate inputs, late arrivals, a backfill, a replay and a schema change if those occur in the real source.
Then simulate a partial failure and retry. Verify whether the resulting business keys and outputs are complete and free of unintended duplicates.
For a streaming or CDC workload, record the input range and observed output after a restart. Test a recovery path that matches the deployed design, including checkpoint compatibility and the sink’s behaviour on replay.
Databricks documents checkpoint state and restart constraints; its foreachBatch guidance specifies at-least-once write guarantees unless the implementation handles deduplication or idempotency appropriately.
A green streaming query still leaves downstream loss or duplication to be checked.
Negative tests matter too. A malformed record, missing source file, upstream delay, denied permission or unavailable downstream sink should produce a known outcome.
Record whether the workload stops, retries, quarantines data, alerts the responsible team and recovers without silent loss.
Benchmark against the operating requirement
Compare the source and target under stated conditions: data volume, concurrency, compute configuration, warm or cold start, measurement window and workload version.
Track batch completion, query latency, dashboard response, throughput, freshness or streaming delay as appropriate. One fast query on a small sample says little about the month-end processing window.
The acceptance question is whether the target meets the agreed service requirement and any minimum parity requirement. Correct results delivered after the reporting deadline still fail that requirement.
A target can also run slower than the source and remain acceptable under an approved operating requirement. Optimization after functional validation is a separate activity.
Store repeat runs and the conditions behind them. If a query meets its latency target only with an unusually large temporary cluster, the result needs review against the intended production configuration and cost constraints.
Avoid declaring success from the best run alone.
Make UAT a business decision with recorded scenarios
Business users need to verify that the new system supports the work they actually do. Invite the report owner, analysts, operations users or other consumers who can judge the workload.
Give each participant a representative scenario, a known period, expected output, a place to record differences and a named decision authority.
For example, a finance owner may verify the month-end report and its drill-down; an inventory user may check the refresh after returns and cancellations.
If the numbers differ, log the affected report, filters, timestamps and business rule. Investigate the difference before asking for approval. A meeting where users see a dashboard and say it looks fine is weak UAT evidence.
A short test may miss a meaningful business cycle. A parallel run can strengthen the evidence: keep source and target operating on comparable inputs, compare results repeatedly and exercise incremental changes across that cycle.
Set its duration from workload criticality, cycle length, failure consequence and the coverage needed, rather than a default number of days. Align snapshots so each comparison measures the same inputs.
The time reserved for user review and parallel evidence can determine how long a Databricks migration takes. Here the question is what those days actually prove.
Record exceptions and decide workload by workload
A failed comparison follows a controlled path: detect, classify, investigate, decide, remediate or accept, retest, and record.
Classify the cause as a target defect, a source defect exposed by the migration, an intentional transformation change, an approved technical variance or an unresolved discrepancy. Link the decision to evidence and a named owner.
An aggregate project status must not hide a blocked workload. If a critical discrepancy remains unexplained, the affected workload stays out of cutover even if other workloads are accepted.
The Databricks migration risk register is the place to track the broader exposure, trigger and contingency; the test record shows exactly what failed and whether acceptance criteria were met.
The evidence varies by workload. A batch pipeline needs output reconciliation, completion-window results and retry evidence. BI needs metric and report equivalence, concurrent query performance and UAT.
Streaming needs event completeness, duplicate handling, latency and restart/replay evidence. An ML feature pipeline needs feature consistency, freshness and downstream compatibility.
The decision about which workloads to move belongs to migration planning; these tests apply after a target implementation exists.
Use an acceptance matrix and a test pack
The matrix below is an illustrative template. Replace its rules and owners with the workload’s agreed requirements before execution.
| Validation area | What is tested | Evidence | Example acceptance rule | Acceptance owner |
|---|---|---|---|---|
| Structure | Schema, types and coverage | Automated schema diff and scope record | No unexplained material contract difference | Data engineering lead |
| Reconciliation | Volumes, records and business outputs | Versioned comparison results | Within approved test-specific tolerance; exceptions classified | Data owner |
| Data quality | Named domain and freshness rules | Rule results, rejected-record record | All agreed blocking rules pass | Data owner |
| Performance | Batch window, latency and concurrency | Repeatable benchmark, configuration and input size | Meets agreed workload SLA/SLO under intended configuration | Workload/platform owner |
| UAT | Representative business scenarios | Scenarios, results and issue decisions | Named business owner accepts required scenarios | Business owner |
| Recovery | Failure, retry and replay | Test run, logs and resulting output comparison | Recovers within agreed requirement without unexplained loss or duplication | Engineering/operations owner |
Keep one test pack per workload: its ID and owners; source baseline and input windows; target version and environment; test datasets and cases; reconciliation outputs; quality rules; performance conditions and results; UAT records; approved tolerances; defects and exceptions; evidence links; and the named sign-off authority.
Include the final disposition and date. That pack should let a reviewer understand what was compared, what passed, what remains open and who made the decision.
The exit gate has four possible outcomes: accepted for cutover, accepted with documented exceptions, remediation and retest required, or blocked.
Before accepting, confirm that required comparisons and quality rules pass, performance meets the agreed requirement, representative UAT is complete, necessary recovery cases have been exercised, exceptions have named approval and evidence is reviewable.
No composite score should conceal a failed blocking test.
After cutover, the workload needs continuing pipeline monitoring in Databricks. Pre-cutover tests establish that the new implementation can be accepted; production monitoring checks whether it keeps meeting its obligations.
If your team is planning a migration and needs to define validation and sign-off for each workload, Dateonic’s Databricks migration services can help map the current estate, build the target and agree what evidence is required before consumers switch.
