Author:

Kamil Klepusewicz

Software Engineer

Date:

Table of Contents

A Databricks proof of concept (PoC) can end with a polished demo and still leave the investment decision no clearer. A notebook ran. A dashboard loaded. A sample query was faster.

 

None of that, on its own, shows that the proposed workload will be correct, governable, recoverable, affordable or operable under the conditions that matter to the enterprise.

 

The purpose of a PoC is to resolve a material uncertainty before the organization makes a larger commitment.

 

That means translating an assumption into a bounded experiment, collecting evidence under defined conditions and deciding what the result permits the organization to do next.

 

The working logic is:

 

uncertainty → hypothesis → baseline → representative test → metric → threshold → evidence → result → decision

 

Full implementation is one possible outcome. A changed architecture, a narrower follow-up test or a decision to stop may be equally valid.

 

What should a Databricks PoC prove?

 

The useful question is not whether Databricks supports a feature. It is whether the selected workload can satisfy an important requirement under conditions that resemble the decision being made.

 

That might mean establishing whether:

 

  • a representative pipeline can complete within its batch window;
  • source and target results reconcile to an acceptable tolerance;
  • required access controls can be enforced and audited;
  • a difficult source can be integrated without an unacceptable workaround;
  • a streaming workload can recover without lost or duplicated output;
  • measured consumption fits the agreed range for the test workload;
  • the intended owner can understand, support and diagnose the solution.

 

If the unresolved question is still whether Databricks belongs on the shortlist, begin with the broader Databricks go/no-go decision framework.

 

If workloads, constraints and unknowns have not yet been established, that belongs in Databricks discovery. The PoC begins where an important assumption remains and documents or architecture review cannot settle it with enough confidence.

 

Its scope should be the smallest one capable of answering that decision. Recreating the future platform wastes time; removing every difficult condition produces reassuring but weak evidence.

 

PoC, prototype, pilot and MVP are not the same thing

 

These terms are often used loosely. Agreeing what they mean prevents a demonstration from being assessed as if it were a production trial.

 

Term Primary purpose Typical question
Prototype Explore or demonstrate an approach, interaction or design. Can we show how this might work?
Proof of concept Test a material assumption under defined conditions. Does the evidence support this technical or operational hypothesis?
Pilot Operate a limited solution with representative users, data or processes. Does it work acceptably in a restricted real operating context?
MVP Deliver the smallest usable product that creates value and can evolve. What is the minimum usable release worth operating and improving?

 

A PoC may contain prototype code, and a pilot may follow it. An MVP may reuse some experimental assets. Those relationships do not make the terms interchangeable.

 

Start with the uncertainty, not a feature list

 

“Test Databricks” is too broad to be useful. No bounded experiment can validate every future workload or operating condition.

 

Name the proposed decision and the uncertainty blocking it:

 

We are considering moving the overnight order-processing workload to Databricks. We do not know whether it can process representative peak-day volume, reconcile the required finance outputs and finish before the 06:00 reporting deadline within the acceptable consumption range.

 

This statement creates an experiment. “Build a Bronze, Silver and Gold pipeline” merely describes work.

 

Each material criterion needs nine fields:

 

  1. Question being tested — what uncertainty are we resolving?
  2. Hypothesis — what do we expect to be true?
  3. Baseline — what does the current system or process achieve today?
  4. Metric — what exactly will be measured?
  5. Test conditions — what volume, file distribution, concurrency, integrations and constraints will apply?
  6. Threshold — what result counts as acceptable?
  7. Evidence source — which records will make the finding reviewable and reproducible?
  8. Owner — who validates the result, and who has authority to accept it?
  9. Outcome — pass, fail or inconclusive.

 

Agree the threshold before running the test. Otherwise, a team can reinterpret an attractive result as success after the fact.

 

Use the current state as a real baseline

 

Improvement claims need a credible comparison. Record how the existing process behaves for the same business scope: its runtime, latency, throughput, failure pattern, reconciliation result, operating effort or consumption, whichever measures matter to the hypothesis.

 

The baseline does not have to be good. It does have to be measured consistently. Comparing a peak-volume Databricks run with an average day on the current platform, or including support effort on one side but not the other, produces a number without a defensible conclusion.

 

When no baseline exists, say so. The criterion can still test an absolute requirement, such as completing before 06:00, but it cannot support a comparative improvement claim.

 

Select the smallest representative workload

 

The easiest workload is often a poor PoC because it avoids the scale, skew, integration, governance or recovery problem that created the uncertainty.

 

Representativeness depends on the hypothesis. It may require:

 

  • realistic data structure, quality and history;
  • production-like volume, file sizes and skew when scale is in question;
  • the transformations that are hardest to reproduce or validate;
  • the source or downstream integration most likely to constrain the design;
  • expected concurrency or arrival patterns;
  • sensitive-data controls and ownership boundaries;
  • late, duplicate, malformed or corrupted records;
  • a dependency failure or recovery path that matters to the workload.

 

This does not mean connecting every future source, reproducing every report or building every production control. Include an element because it could change the result, not because it appears on a generic implementation checklist.

 

Production-style engineering follows the same rule. Repeatable deployment and monitoring may be essential when the uncertainty concerns operability. They add little to an experiment whose only question is whether an unusual file format can be parsed correctly.

 

The engineering effort should match the claim the PoC is expected to support.

 

Define success criteria across the risks that matter

 

Query speed is rarely the whole decision. Choose criteria from the areas below only where the result could support or invalidate the hypothesis.

 

Functional correctness

 

Completion without an exception is not enough. The output must be correct for the agreed business purpose under representative data conditions.

 

Useful evidence includes source-to-target reconciliation, control totals, record-completeness checks, transformation tests, explained exceptions, output equivalence and, where relevant, model-quality measures.

 

Interpret tolerances in context: a finance balance, clickstream aggregate and probabilistic model should not share the same acceptance rule.

 

Performance and scale

 

Measure the requirement the workload must meet runtime, end-to-end latency, throughput, concurrency, refresh SLA or batch window, and disclose the conditions behind the result.

 

Data volume, file layout, skew, compute configuration, cache state, competing workloads and repetitions can all affect the finding.

 

A single fast run is weak evidence. If the hypothesis concerns repeatable performance, test enough controlled runs to show whether the threshold holds and retain job or query history with the configuration record.

 

Cost and consumption

 

Capture what the representative test actually consumed. Depending on scope, that may include Databricks usage, cloud infrastructure, storage, network traffic and workload attribution.

 

The system.billing.usage system table provides billable usage records, while query history can supply execution evidence for supported workloads. Those records support a bounded statement:

 

Under the documented test conditions, this workload consumed the measured resources and met, or missed, the agreed range.

 

They do not, by themselves, forecast enterprise-wide production spend. Passed measurements can later inform the Databricks migration business case, where scale, operating model and wider costs belong.

 

Integration feasibility

 

Test the integration most capable of invalidating the design. The criterion might cover authentication, private connectivity, source limits, schema behaviour, API reliability, throughput, downstream consumption or an operational dependency.

 

Evidence should show what happened through the organization’s actual network and identity path, not merely that a connector exists. A workaround counts as a finding; its security, support and maintenance consequences determine whether the hypothesis still holds.

 

Governance and security

 

Where a control is decision-critical, exercise both the permitted and prohibited path. Confirm that the approved identity can perform the intended action, that an unapproved identity is denied and that the event is visible to the control owner.

 

Evidence may include permission-test results, audit records, lineage, policy output and service-identity behaviour. Unity Catalog capabilities matter only insofar as the PoC demonstrates that the required control can be enforced and reviewed under the proposed arrangement.

 

Reliability and recovery

 

Relevant failure modes should be introduced deliberately. Stop a job, replay late or duplicate data, make a dependency unavailable or interrupt a stream, then measure whether processing restarts from the expected point and whether output remains correct.

 

For streaming workloads, checkpoints retain processing state and processed-record information.

 

The existence of that mechanism is not the result: recovery still has to be verified for the actual source, sink, stateful logic and operating procedure in scope. Job history, checkpoint state, reconciliation output and an incident timeline make the finding reviewable.

 

Operability

 

When operability is part of the hypothesis, ask someone other than the author to run, observe or recover the workload. Evidence should show whether configuration and logs are understandable, the failure can be located and the agreed recovery action can be completed.

 

Full CI/CD is unnecessary unless deployment repeatability is itself being tested.

 

Team and operating-model fit

 

Limit this criterion to the ownership assumption attached to the selected workload. Can the intended engineers review a change, interpret the performance and consumption evidence and respond to the tested failure without permanent reliance on the PoC authors?

 

A supervised exercise and named sign-off provide stronger evidence than a general skills assessment.

 

Business acceptance

 

For a business-facing use case, name the person authorized to accept the outcome and define what they will inspect: reconciled figures, report usability, freshness, model quality or the effect on a decision workflow.

 

Attendance at a demo is not acceptance unless the acceptance question and authority were agreed in advance.

 

Example Databricks PoC scorecard

 

The values below are illustrative examples, not Databricks benchmarks or universal recommendations. Thresholds must come from the organization’s baseline, workload requirements and risk tolerance.

 

Validation question Baseline Success threshold Test conditions Evidence Owner Result
Can the migrated order pipeline reproduce the approved finance output? Current warehouse is the approved source for 18 control totals. All 18 totals reconcile; row-level variance stays within the agreed 0.1% tolerance and every exception is explained. Ninety representative days including corrections, cancellations, duplicates and late records. Reconciliation report, exception records and finance sign-off. Finance data owner PASS
Can the daily workload meet the reporting window? Median 142 minutes; current deadline is 06:00. Each of five controlled runs completes within 90 minutes; p95 task latency remains inside agreed component limits. Production-like volume, file distribution and skew; fixed compute configuration; cold and warm runs reported separately. Job-run history, Spark metrics and configuration record. Platform lead PASS
Can expected BI concurrency be supported? Current system supports 25 users with p95 response of 11 seconds. At 40 simulated users, agreed critical queries stay below 8 seconds p95 and error rate remains below 1%. Representative query mix, cache state disclosed, three repeated test windows. Query history export and load-test report. BI lead INCONCLUSIVE — test mix omitted two critical queries
Is workload consumption inside the PoC range? Comparable current run costs €31 in platform and infrastructure consumption. Databricks usage remains within the agreed equivalent range of €25–€35 per run under the stated configuration. Same input period and output scope; five runs; discounts and idle time treated consistently. Billing system table, cloud bill export, workload tags and calculation sheet. FinOps owner PASS
Can restricted customer data be governed as required? Existing policy permits only the Customer Risk group. Approved group can query permitted columns; an unapproved group and user are denied; access and changes are auditable. Enterprise identities, representative sensitive columns and service principal included. Permission-test report, audit records and lineage evidence. Security owner PASS
Can the streaming workload recover without corrupting output? Current process requires manual replay after interruption. After a forced interruption, processing resumes from the expected checkpoint, no accepted event is lost or duplicated, and backlog clears within 30 minutes. Peak event rate, late and duplicate events, plus a simulated sink outage. Checkpoint state, progress metrics, target reconciliation and incident timeline. Data engineering lead FAIL — duplicate output after sink recovery
Can the support team operate the workload? Only the PoC engineer knows the current process. A designated support engineer identifies the failed stage, follows the runbook and restores a test run within 20 minutes without author intervention. Failure injected without advance disclosure of its exact location. Job logs, runbook record and operator debrief. Service owner PASS

 

Do not turn the table into one average score. A failed mandatory security, correctness or reliability condition can block continuation even when every performance measure passes.

 

Classify each criterion before testing:

 

  • Mandatory gate: failure blocks the proposed next step.
  • Trade-off criterion: a miss may be accepted with an explicit cost, design change or remediation owner.
  • Exploratory criterion: the result informs later design but does not decide the PoC alone.

 

Interpret PASS, FAIL and INCONCLUSIVE separately

 

 

PASS

 

The agreed evidence supports the hypothesis under the stated conditions. It does not support claims beyond those conditions.

 

FAIL

 

Reliable evidence contradicts the hypothesis, a mandatory threshold is missed or a required control cannot be demonstrated. That may justify stopping, changing the architecture, narrowing the workload or testing a materially different approach.

 

A failed hypothesis is useful if it prevents a larger commitment based on a false assumption.

 

INCONCLUSIVE

 

The experiment did not produce evidence strong enough to decide. Perhaps the data was unrepresentative, a critical workload condition was omitted, configurations changed between runs, the result could not be reproduced or the authorized owner was unavailable to validate it.

 

An inconclusive result cannot become a pass because the timebox expired. If the uncertainty is still material and testable, define a narrower follow-up. Otherwise, carry it forward explicitly as an unresolved assumption or risk.

 

When is the Databricks PoC actually finished?

 

Completion is an evidence condition, not a calendar event. The PoC is finished when the decision-makers can see what was tested, judge the strength of the evidence and act without hidden assumptions.

 

The exit review should contain:

 

  • the decision the experiment was designed to support;
  • hypotheses, criteria and thresholds agreed before testing;
  • the current-state baseline;
  • the workload and conditions actually tested;
  • measured results with retained evidence;
  • every pass, fail and inconclusive result;
  • deviations from the planned experiment;
  • mandatory criteria not met;
  • untested areas, limitations and remaining assumptions;
  • new risks and their owners;
  • deadlines and validation methods for unresolved questions;
  • implications for the proposed architecture, scope and next investment.

 

That package can support five legitimate outcomes:

 

  1. Proceed to implementation — mandatory criteria passed and the remaining uncertainty is acceptable.
  2. Proceed conditionally — the evidence supports continuation, with named remediation or unresolved questions carried into the approved scope.
  3. Run another focused validation — a material assumption remains inconclusive and can be resolved through a narrower experiment.
  4. Change the architecture or workload — the original hypothesis failed, but the evidence supports a materially different option.
  5. Stop the initiative — a critical assumption was contradicted or the expected benefit no longer justifies the next investment.

 

The final review is therefore a decision gate, not a ceremony for endorsing work already completed.

 

What a successful PoC does not prove

 

The decision record should place the limits of the evidence next to the results. Even a well-designed PoC does not establish that:

 

  • every future workload will behave like the one tested;
  • short-term execution proves long-term reliability;
  • one scale and concurrency profile predicts performance at every production level;
  • measured PoC consumption equals production TCO;
  • a working technical integration guarantees organizational adoption or long-term supportability;
  • an experimental workspace or its code is ready for production;
  • implementation risk has been eliminated.

 

The experiment reduces named uncertainties. It does not remove the engineering, governance and organizational work required after approval.

 

Carry the evidence into implementation

 

If the decision is to continue, hand over the evidence rather than only the demo and code. The implementation team needs:

 

  • measured results and the underlying records;
  • workload and test-condition definitions;
  • architecture decisions made during the experiment;
  • unresolved assumptions, failed criteria and known limitations;
  • integrations, identities and controls tested;
  • consumption observations;
  • risks and production requirements discovered;
  • reusable code, with its quality and constraints stated;
  • named owners for remediation and acceptance.

 

Experimental code may need to be redesigned, secured, automated, tested and documented rather than promoted unchanged. Once implementation has been authorized, the Databricks implementation checklist can be used to review delivery coverage.

 

The Databricks implementation cost guide addresses how validated scope translates into effort and budget.

 

Make the next investment an evidence-based decision

 

A credible Databricks PoC resolves a defined uncertainty. Before testing begins, record the hypothesis, baseline, conditions, metric, threshold, evidence source and owner.

 

At the exit gate, distinguish pass from fail and inconclusive, disclose what remains unknown and authorize only the next step supported by the findings.

 

Dateonic can help design a representative PoC, test a material technical assumption and assess whether the evidence supports full implementation. See Databricks Migration Services to discuss a focused validation.