What does good data quality mean for your task?
There is no universal standard of “clean data”. A dataset is good enough when it supports the intended decision with limitations you can accept. So start from the purpose: which fields are required for this report or process, which formats are valid, which identifiers link the records and how fresh the data must be. Write these requirements down for each report or process, because the same table can be perfectly adequate for a monthly overview and quite inadequate for an automated invoice run. Setting such requirements is where data analytics services should begin.
Not all gaps are equal. An empty comment field and a missing order number are problems of very different scale: the first barely matters, while the second can break a whole calculation. Ranking issues by their impact on the task stops the team from spending a week polishing fields nobody uses.
Keep accuracy, completeness and timeliness apart in the review. A record can be correct but late, or complete but inaccurate. When each type of issue is named separately, it is clear who should fix it and how. A single combined quality score hides that distinction and gives the team nothing concrete to act on.
Where do records break when systems are combined?
Most problems appear where data from different systems meets, so database development and integration work starts at those joins. Look at how records are matched, by which key and under which rules, and how updates change their status as an order moves from placed to paid to shipped, or is cancelled.
Test the typical trouble spots on purpose: duplicated identifiers, references that point to nothing, cancelled transactions that still appear in totals and data that arrives after the report has been built. A database audit checks these trouble spots systematically. Reconcile a representative sample with the people responsible for the source systems; they often know about quirks that are not documented anywhere. Late data deserves special attention: if a report is built before the last updates arrive, decide whether it is rebuilt later or marked as provisional, so readers know which figures may still change.
Records that fail the checks should go into an explicit review flow. The worst outcome is invalid data being silently turned into zeros or dropped from totals without a documented rule. The report then looks tidy, but nobody knows what was removed or why, and the figures cannot be explained when someone asks.
How do you build checks into everyday work?
A one-off clean-up helps for a while, but without checks in the process the same errors come back. Place validations at the points where data crosses a boundary: when it is collected, when it is transferred between systems and when it enters a report.
For every check, decide who investigates a failure, how the data is corrected and at what point downstream use should pause; sometimes it is better to delay a report by a day than to publish a wrong figure. Record exceptions and track which defects recur, because that pattern shows where the real cause lies.
To stop errors returning after a clean-up, you need a person responsible and a rule that prevents repeats, not just the correction itself. Otherwise a dashboard or an AI workflow will faithfully reproduce the same inconsistency on every run, and the effort spent on cleaning will have been wasted.
Checklist: data quality
- Define the required identifiers, valid formats and how fresh the data must be for this task.
- Test duplicates, record matching between systems and late updates.
- Reconcile a sample with the people responsible for the source systems.
- Decide who handles failed checks and who is responsible for preventing repeats.
Example: duplicated customer keys in a sales report
A sales report combines orders with customer data, but one import has created duplicate customer keys, so some orders could be counted against the wrong customer. The team moves the ambiguous rows into a separate review set, confirms the matching rule with the system owner and reruns the reconciliation before publishing the figure.
A documented validation step is then added to the import, so the next batch is checked automatically and the same defect does not reappear.


