Skip to content

Checking the quality of your figures

What the Data quality page finds in the platform figures stored for a workspace, why every finding states how it was detected, and why "nothing has been ingested" is not the same as "everything is fine".

Updated

Data quality looks at the platform figures stored against one workspace and tells you what is known to be wrong with them. It finds ten kinds of fault, and every one of them states the basis it was detected on, so you can check the reasoning rather than take it on trust.

Detection is arithmetic over the rows already stored. No language model decides whether a figure is faulty, and nothing on the page is a score. A quality score would be a judgement compressed into a number, which is the one thing this page exists to stop.

BEFORE YOU START

  • The "view" permission for the workspace, and access to the workspace itself. If you can read the workspace report, you can read this.
  • Nothing else. The page works before anything has been ingested, and says so honestly rather than showing an empty screen.

STEPS

  1. Open Data quality for the workspace, or go to /app/workspaces/{workspace}/data-quality. There is also a link from the workspace report.
  2. Read the verdict at the top. It is one of four states and they mean different things - see below.
  3. Read the Connections table. Each connected account shows what it is actually doing, and a connection that is working and reporting no activity is shown as exactly that rather than as a failure.
  4. Read the findings. Each one shows what it means for a number, which account and period it concerns, how it was detected, and what to do about it.
  5. Choose "Show the evidence" on a finding to see the raw observations behind it - the missing dates, the row counts, the values.
  6. Read the restatement ledger at the bottom to see any figure the platform has changed since it was first read, with the value it replaced.
  7. Set "From" and "To" and choose "Apply period" to narrow the assessment to the window you are looking at.

WHAT YOU SHOULD SEE

One of four verdicts, and the difference between them matters:

  • "assessed clean" - rows were read, every check ran, and none found a fault. This is the only good news on the list.
  • "no data ingested" - nothing has been read from any platform, so there is nothing to assess. This is an EMPTY result, not a clean one, and the page will not describe it as fine.
  • "findings" - one or more faults, each with its basis.
  • "not assessed" - the assessment did not run. Nothing is claimed, including that your data is sound.

WHAT THIS WILL NOT DO

  • It will not tell you anything. Nothing here is pushed, emailed or notified. Findings wait until somebody opens this page or calls the API.
  • It will not fix anything. There is no repair button and no delete. Which of two conflicting rows is right is a question about the platform, not about this database, and a control that could delete one could destroy the evidence you need to answer it.
  • It will not show a trend. Findings are worked out fresh each time you open the page, so there is no history of how quality has changed. The restatement ledger is the only running record, and it only covers figures that changed.
  • It will not tell you what your numbers MEAN. A figure that looks surprisingly high is not a fault; this page is about whether the data is sound, not about what it says.

WORTH KNOWING

A reporting gap is deliberately ambiguous, and the finding says so. A day with no stored row is either a day that was not read or a day the platform reported nothing for, and the stored data cannot tell those apart. It is reported anyway, because the alternative - reading the absence as zero - is invisible: a chart draws a flat line and nobody asks.

A finding that says "not a total" is different from one that says "total is qualified". A gap makes a total lower than the truth, which is still a real total you can use once you know which way it leans. Money in two currencies, or two platforms counting their days against different clocks, means the number is not a total of anything - so the report shows no figure for that column at all rather than a misleading one.

A platform restating a figure within two days of the period ending is normal and is not reported as a fault. Platforms keep counting for a while as late events land. What is reported is a figure that changed AFTER that window closed, because you may already have sent that number to somebody.

IF IT DOES NOT WORK

  • "This is not a clean result - it is an empty one." means nothing has been ingested yet. Check the Connections table for why.
  • "This deployment holds no application credentials for the provider" means an operator has to configure them. No action you take will change it, and you should not be asked to reconnect.
  • "This organisation's connection can no longer be used to read; it needs reconnecting." means you can fix it, from the Connections centre.
  • "The provider has not granted the access this read needs." means the platform itself is the blocker. Nobody here can hurry it and no date can be promised.
  • "No platform connection has ever attempted a read for this workspace" means there is no sync set up at all - which is why there are no figures, rather than the platforms having reported none.

COMMON QUESTIONS

Why is there no overall quality score?

Because a score would be a made-up number in a feature whose whole purpose is stopping made-up numbers. The findings say what is wrong, what it does to a figure and what to do; a percentage would say none of that while sounding more precise.

Why does a finding say the amount missing is unknown instead of estimating it?

Because an estimated shortfall inside a report about missing data is the same fabrication one level up. Where the platform states how many rows it held back, the figure is shown. Where it does not, the page says the amount is unknown.

Should I delete the old rows for a post that no longer exists?

No. A post that has been deleted still had the engagement it had on the days it existed, and removing those rows would falsify every past total that included them. The finding says this explicitly, because the risk with it is not that you ignore it - it is that you tidy up the data it points at.

  • Reading your reporting

    What the workspace report totals, why a figure can say "Not reported" instead of zero, and the two ways a figure reaches it.

  • Attributing revenue to channels

    Reading first-touch, last-touch and linear credit over your recorded touchpoints and conversions, why some conversions are unattributed, and why two currencies are never added together.