Is your business data ready for AI?

Check whether your business data is ready for AI: quality, access, location, permissions, personal data, source of truth and duplicates, step by step.

Open a support ticket

When people ask whether their data is ready for AI, they usually mean something simple: if we connect an AI tool to our information, will the results be useful and safe?

The answer depends less on the AI and more on the data. An AI step can only work with what it receives. If customer records are duplicated, product information is out of date, or the right information is scattered across inboxes and spreadsheets, the AI will reflect those problems. Sometimes it will hide them behind confident, well-written output.

This is not a reason to wait until everything is perfect. No business has perfect data. The goal is to know what you have, where the weak points are, and whether they matter for the specific process you want to improve.

Data readiness is also not a single yes or no. Your data may be ready for one use case, such as summarizing support tickets, and not ready for another, such as automatically updating your CRM.

This guide explains:

  • what "ready for AI" actually means
  • the main areas to check: quality, access, location, permissions, personal data, source of truth and duplicates
  • how to decide whether your data is good enough for a first step
  • common mistakes when preparing data
  • a practical checklist you can complete yourself

What does "data ready for AI" mean?

Data is ready for AI when it is good enough, accessible enough and safe enough for a specific task.

That definition has three parts:

  • good enough: accurate, complete and consistent for the decisions the AI will support
  • accessible enough: the workflow can reach the data in a reliable, controlled way
  • safe enough: you understand what data is involved, who may see it and where it will be sent

"Good enough" is important. Data that works for a draft reply checked by a person may not be good enough for an action that happens automatically. The higher the consequence of an error, the higher the bar.

Why does data readiness matter so much?

Most disappointing AI pilots fail because of data, not because of the model.

Typical symptoms include:

  • answers that are correct in form but wrong in content
  • different results for what should be the same customer or product
  • a workflow that works on test examples and breaks on real records
  • AI output based on old information that someone forgot to update
  • personal data appearing in places where it should not

Checking data before you build saves time, because you fix the cause once instead of correcting the output again and again.

Is your data accurate and complete?

Start with quality. You do not need a formal audit. A small, honest sample is usually revealing.

Accuracy

Accuracy means the data reflects reality. Ask:

  • Are contact details, prices, stock levels or statuses correct today?
  • When was this information last updated, and by whom?
  • Are there fields that people fill in carelessly because "nobody uses them"?

Completeness

Completeness means the information the AI needs is actually there. Ask:

  • Which fields are often empty?
  • Is important context stored in free-text notes, attachments or email threads instead of fields?
  • Do older records follow different rules from newer ones?

Consistency

Consistency means the same thing is recorded the same way. Look for:

  • different spellings of the same company or product
  • dates and numbers in mixed formats
  • categories or tags that overlap or mean different things to different people
  • statuses that are used differently by different teams

Inconsistency is often more harmful than missing data, because it is harder to notice.

Can the workflow access the data?

Data that exists is not always data that a workflow can use.

Check how each source can be reached:

  • Does the tool offer an API, an export or an integration with platforms such as n8n or Make?
  • Does access require a personal login, or can you create a dedicated account or key for the workflow?
  • Are there limits on how much data can be read, or how often?
  • Is the data only available as PDFs, scanned documents or images?

Data that is only available through manual exports is a warning sign. It often leads to workflows that depend on someone remembering to download a file.

If access is the main obstacle, the work may be more about integration than about AI. D4Hub's API integrations and automations service covers this kind of work.

Where does the data live?

Location matters for two reasons: practical and legal.

Practical location

List where the relevant data is stored today. In many small and medium businesses the answer includes:

  • a CRM or helpdesk
  • an ecommerce platform
  • shared spreadsheets
  • cloud drives and shared folders
  • email inboxes
  • accounting or ERP software
  • notes in project tools such as Notion or Airtable

The more places involved, the more connections the workflow needs, and the more points where things can break.

Geographic and contractual location

It is also worth knowing where data is processed and stored, especially when an AI step sends it to an external service. Check the terms and settings of each provider: where data is processed, whether it may be retained, and whether it may be used to improve the provider's services. These details vary between providers, plans and settings, so read the current documentation rather than relying on assumptions.

Who has permission to see and change the data?

AI workflows can accidentally widen access to information.

For example, a chatbot connected to a shared drive may be able to answer questions using documents that only some employees should see. An automation running with an administrator account may be able to change far more than it needs to.

Check:

  • who can currently read and edit each data source
  • which account the workflow will use, and what that account can do
  • whether the workflow needs to write data, or only read it
  • whether the output will be visible to people who should not see the original data

A good principle is to give the workflow the minimum access it needs and nothing more. A dedicated account with limited permissions is usually safer than a personal or administrator login.

Does the data include personal information?

Many business processes involve personal data: names, emails, phone numbers, addresses, order history, support conversations, employee details.

If you operate in the European Union or handle data of people in the EU, the GDPR applies to personal data, whether or not AI is involved. Adding an AI step does not change that, but it may add new questions, for example about which providers receive the data and on what basis.

Practical steps that usually help:

  • identify which fields and documents contain personal data
  • check whether the AI step actually needs that data, or whether it can work with less
  • remove or mask personal details where they are not necessary
  • check the data processing terms of every external service involved
  • make sure logs and test copies do not keep personal data longer than needed
  • involve the person responsible for privacy in your organization early

This guide is not legal advice. For questions about your specific obligations, consult a qualified advisor. Depending on the use case, other rules such as the EU AI Act may also be relevant. The AI compliance page explains how D4Hub can support the technical and documentation side.

Which system is the source of truth?

A source of truth is the system that holds the official, current version of a piece of information.

Problems start when the same information exists in several places and nobody knows which one is correct. For example:

  • customer details in the CRM, the ecommerce platform and a spreadsheet, each slightly different
  • product descriptions in a shared document, the website and a supplier file
  • prices in an internal list that does not match the store

For each type of data the AI will use, decide which system is the source of truth. The workflow should read from that system, and other copies should be updated from it, not edited separately.

If you cannot agree on a source of truth, that is an important finding. It usually needs to be resolved before an AI workflow can be reliable.

How serious are duplicates?

Duplicates are one of the most common data problems and one of the most damaging for AI and automation.

They cause:

  • the same customer being contacted twice, or with conflicting information
  • reports that count the same thing more than once
  • AI summaries that merge or confuse different records
  • automations that update one copy and leave another unchanged

Check for:

  • the same email address or company appearing in several records
  • the same product with slightly different names or codes
  • contacts created by different forms or imports without a matching rule

Cleaning duplicates is often worth doing before any AI work, because it improves every process, not only the new one. Decide on a matching rule first, and back up the data before merging anything, since merges are often hard to reverse.

How do you decide if your data is good enough?

You do not need to fix everything. Focus on the data used by the specific process you want to improve.

A practical way to decide:

  1. Choose one process and list the data it uses.
  2. Take a small sample of real records.
  3. Check each record for the issues described in this guide.
  4. Ask the process owner: if the AI used these records, would the output be acceptable?
  5. Decide whether to proceed, proceed with a human review step, or fix the data first.

Often the answer is to proceed with a limited pilot and a human review step, while fixing the most serious data issues in parallel. The guide on what an AI integration roadmap should include explains how to structure that pilot.

What are the most common mistakes?

  • Assuming AI will clean the data. AI can help spot inconsistencies, but it should not silently decide which version is correct.
  • Testing only with good examples. Real data contains the problems that matter.
  • Connecting everything at once. Start with the minimum data the process needs.
  • Using personal or admin accounts for workflows. This widens access and makes changes hard to trace.
  • Ignoring free-text fields. Notes and email threads often contain the most important context and the most personal data.
  • Forgetting about logs and copies. Test exports and workflow logs can contain sensitive data long after the test is over.
  • Leaving privacy questions to the end. They are easier to answer during design.

What does good look like?

Data that is ready for a first AI use case usually has these characteristics:

  • the process owner knows which data the workflow uses
  • each type of data has an agreed source of truth
  • the workflow can access the data through an API or a stable integration
  • the workflow uses a dedicated account with limited permissions
  • personal data is identified and reduced to what is necessary
  • duplicates in the relevant data have been reviewed
  • the provider terms for external services have been checked
  • there is a plan to keep the data accurate after launch

None of this requires a large project. For one process, it may take a few focused sessions.

How D4Hub can help

D4Hub can help you understand whether your data can support the AI workflow you have in mind, and what to fix first.

Depending on your situation, D4Hub can help with:

  • mapping where the relevant data lives and how it moves between tools
  • checking which systems offer APIs or reliable integrations
  • defining a source of truth for each type of data
  • identifying duplicates and inconsistencies, and planning a safe cleanup
  • setting up dedicated accounts and limited permissions for workflows
  • reducing the personal data sent to external services
  • designing workflows in n8n, Make or custom code that read from the right sources
  • preparing technical documentation to support privacy and compliance work

You can ask for support at any stage, whether you are planning a first pilot or reviewing a workflow that is already running.

Open a support ticket

For hands-on users: a data readiness checklist

Use this checklist for one process at a time. It is designed to be completed with a spreadsheet or a notepad, without technical tools.

Before you start, work on copies or read-only views. Do not merge, delete or bulk-edit records during the check. If you export data, store the file securely and delete it when you are done.

Step 1: list the data sources

For the process you have chosen, write down every place the data comes from. For each source note:

  • the tool or file name
  • who owns it
  • how it can be accessed (API, integration, export, manual copy)
  • whether it contains personal data

Step 2: take a small sample

Pick a small number of recent, real records from the main source, for example a few dozen. Include some older records too, because they often follow different rules.

Step 3: check quality

For each record in the sample, mark:

  • fields that are empty but should be filled
  • information that is clearly out of date
  • values that use a different format or spelling from the others
  • important information that exists only in notes or attachments

Count how many records have at least one problem. This gives you your own baseline, without relying on external figures.

Step 4: check for duplicates

In the sample and in the full list if possible:

  • sort by email, company name or product code
  • look for entries that are the same thing recorded twice
  • note how the duplicates were probably created (form, import, manual entry)

Step 5: agree the source of truth

For each type of data (customers, products, prices, orders, documents), write which system holds the official version. If the team does not agree, mark it as an open issue.

Step 6: check permissions

For each source, answer:

  • Which account would the workflow use?
  • Does it need to read only, or also write?
  • Could the output reveal information to people who should not see it?

Step 7: check personal data

List the personal data fields involved and, for each one, answer:

  • Does the AI step really need it?
  • Can it be removed, shortened or masked?
  • Which external services would receive it?

Share the open questions with the person responsible for privacy in your organization.

Step 8: decide

Based on the checklist, choose one of three options:

  • proceed with a pilot
  • proceed with a pilot and a mandatory human review step
  • fix specific data issues first, then reassess

If you share the results with D4Hub, send the summary, not the raw data, and never include passwords or API keys.

FAQFrequently asked questions

No. It needs to be good enough for the specific task, with known weak points. Many businesses start with a limited pilot and a human review step while they improve the data.

AI can help find likely duplicates, inconsistent formats or missing fields, and it can suggest corrections. The final decision on which version is correct should normally stay with a person, and changes should be backed up and traceable.

It depends on the provider, the plan, the settings and your obligations. Read the provider's current data processing terms, send only the data that is necessary and involve your privacy contact. This guide is not legal advice.

That is common and workable. It usually means the first steps are about structure and access: deciding on a source of truth, moving key information into fields, and creating reliable connections. D4Hub can help plan that work.

Pick the system where the data is created or most carefully maintained, and where the team already looks for the official version. Then make other systems read from it rather than keeping separate copies.

GDPR does not prohibit AI. It sets rules on how personal data is processed, which apply with or without AI. Reducing the personal data involved and checking provider terms are good starting points, together with advice suited to your situation.

Yes. A first review can often start from a description of your tools, a list of data sources and a few anonymized examples. Further access can be agreed later, with dedicated accounts and limited permissions.

Related resources

Related services and technologies