Data contracts should fail in CI, not production
Why I am exploring data contract testing as an earlier feedback loop for breaking changes in data pipelines.
- data-engineering
- data-contracts
- ci-cd
- testing
A data pipeline can be green while the data flowing through it is already broken.
The job ran. The container exited with code 0. The scheduler is happy. Then a dashboard goes blank, a transformation starts producing nulls, or a downstream consumer quietly changes behavior because a column changed shape.
That gap is the part I am interested in.
The problem is not only pipeline failure
Traditional CI is very good at checking code. We lint it, type-check it, run unit tests, and block a merge when the program violates expectations.
Data systems have another interface to protect: the data contract between producers and consumers.
A producer might still return valid JSON while renaming a field. A table might still exist while changing a column from integer to string. A dataset can pass through every infrastructure check and still violate what downstream jobs assume.
If those assumptions only live in people's heads or inside the consumer implementation, feedback arrives too late.
What I want a data contract to do
For the kind of systems I am exploring, a useful contract should be executable.
It should describe things such as:
- required fields and columns,
- expected types,
- nullable versus non-nullable values,
- compatibility rules for schema changes,
- and business-level invariants that are important enough to block a release.
The important part is not the document format. The important part is where the contract is enforced.
My preferred failure point is CI.
Move the failure to the pull request
Imagine a producer changes:
customer_id: integer
to:
customer_id: string
That change may be completely intentional. The question is whether consumers are ready for it.
Instead of discovering the mismatch after deployment, CI can compare the proposed schema against the contract and classify the change. Compatible changes continue. Breaking changes fail with a useful explanation.
That turns a production incident into a pull-request conversation.
The pipeline becomes roughly:
code change -> build -> unit tests -> data contract validation -> compatibility check -> deploy
The contract check is not replacing integration or production monitoring. It is moving one class of failure earlier, when the cost of fixing it is much lower.
What makes this harder than normal schema validation
Checking whether a JSON document matches a schema is the easy part.
The more interesting questions are about evolution:
- Is adding a nullable field breaking?
- What if a field changes from integer to number?
- Can a producer remove a field that no current consumer uses?
- Who owns the contract when several consumers depend on the same dataset?
- How should CI report a failure so the developer can actually act on it?
Those questions turn a validator into a workflow problem.
That is why my current research is less about inventing another schema language and more about how contract testing fits into CI/CD for real data pipelines.
The direction I am testing
My current prototype focuses on a small loop:
- define a machine-readable contract,
- validate a candidate schema or sample dataset,
- compare it with the accepted contract,
- classify incompatible changes,
- return a CI-friendly result and explanation.
I want the interface to stay small enough that it can run locally and in a GitHub Actions job without needing an entire data platform around it.
If this works well, the useful outcome is not a fancy contract file. It is a boring, predictable failure before merge.
That is exactly where I want this class of problem to become boring.