LTK Soft

Data Engineering

Is Your Data Ready for AI? A 10-Point Readiness Checklist

Most AI projects don’t fail at the model — they fail at the data. Use this checklist to find out where you stand before you commit budget, and what to fix first.

LTK

LTK Soft Team

9 min read

Scattered data points organised into clean layers, with a readiness gauge
Key takeaways
  • Gartner has predicted that through 2026, organisations will abandon 60% of AI projects that aren’t supported by AI-ready data.
  • AI-ready doesn’t mean perfect. It means the data behind your first use case is findable, trustworthy, permissioned, and accessible.
  • Documents and permissions matter most for AI assistants; clean history and outcomes matter most for predictive models.
  • Score yourself on the ten points below and fix the gaps for one use case first, not the whole company.

When an AI initiative stalls, the post-mortem rarely blames the model. It finds that the data was spread across systems nobody could connect, that two departments disagreed on what a field meant, or that the documents the assistant relied on were out of date. These are solvable problems — but far cheaper to solve before the project starts than halfway through.

Why data decides AI success

Gartner has predicted that through 2026, organisations will abandon 60% of AI projects that are not supported by AI-ready data. Whether you are building a private knowledge assistant, automating document processing, or forecasting demand, the model can only be as good as what it is given.

The 10-point AI data readiness checklist

Give yourself one point for each statement that is true for the data behind your priority use case.

01You know where your data lives

What good looks likeA current inventory of systems, databases, file shares, and SaaS tools, with what each contains
Red flagNobody can list every place customer data is stored

02Every key dataset has a business owner

What good looks likeA named person who can answer “is this right?” and approve access
Red flagData issues bounce between IT and the business with no decision

03Quality is measured, not assumed

What good looks likeKnown rates for missing values, duplicates, and errors in the fields that matter
Red flag“The data is fine” — but reports from different teams never match

04Definitions are shared

What good looks likeOne agreed meaning for terms like “active customer,” “claim closed,” or “revenue”
Red flagEach department calculates the same metric differently

05Documents are findable and current

What good looks likePolicies, procedures, and contracts in managed repositories with clear versions
Red flagTen copies of the same policy across drives and inboxes, some years out of date

06Permissions are documented and enforceable

What good looks likeAccess is role-based and managed through a central identity provider
Red flagAccess is granted ad hoc and nobody knows who can see what

07Sensitive data is classified

What good looks likePersonal, health, financial, and confidential client data is tagged and handled by policy
Red flagSensitive data appears in shared spreadsheets and test environments

08Pipelines are automated

What good looks likeData moves between systems on schedules or events, with monitoring and alerts
Red flagCritical reports depend on someone exporting and emailing a spreadsheet

09History and outcomes are retained

What good looks likePast cases are kept with their results — approved, denied, resolved, churned
Red flagOld records are purged or outcomes were never recorded

10Core systems are reachable by API

What good looks likeKey systems can be read from, and written to, programmatically
Red flagThe only way into the main system is through its screens

Scoring your results

ScoreWhat it meansRecommended next step
8–10Ready for a production-minded pilotPick the highest-value workflow and run a measured pilot
5–7Ready with targeted fixesFix the specific gaps behind your first use case while the pilot is scoped
0–4Foundation work needed firstStart with data integration, ownership, and access before investing in AI

Fixing the common gaps

  • Scattered data: build automated pipelines into a central, governed store — a warehouse or lakehouse for structured data, a managed repository for documents.
  • Conflicting definitions: agree a short business glossary for the metrics your first use case depends on, and encode it in shared data models.
  • Unclear permissions: move access to role-based groups in your identity provider so AI systems can enforce the same rules automatically.
  • Systems without APIs: wrap them with an integration layer, or plan their modernization.
Don’t boil the ocean

You do not need a perfect enterprise data platform before starting with AI. Choose one valuable use case, fix the data it depends on, and let each project leave the foundation a little stronger for the next.

Find out where to start in 2–3 weeksWorkflow and data review, compliance constraints, and a prioritised list of AI use cases with the data work each needs.
See the AI Readiness Assessment

Need to build the foundation itself? Our data engineering team designs and builds the pipelines, models, and governance that AI depends on.

Frequently asked questions

Do we need a data warehouse before we can use AI?

Not always. Knowledge assistants built on retrieval (RAG) depend mostly on well-organised documents and permissions. Predictive models, forecasting, and analytics-driven AI do need clean, consolidated, historical data — which is where a warehouse or lakehouse earns its keep.

How long does it take to become AI-ready?

It depends on the starting point and the use case. You rarely need to fix everything first: targeted work on the data behind one priority use case can often be done in weeks, while a broader data foundation is built in parallel.

What does an AI readiness assessment include?

Typically interviews with each department about workflows and pain points, a review of systems, data sources, and data quality, a check of compliance constraints, and a prioritised list of AI use cases with the data work each one requires. Ours is a focused two-to-three-week engagement.

Is unstructured data like PDFs and emails useful for AI?

Very. Modern language models can read contracts, emails, reports, and scanned forms, which means a large share of company knowledge that was previously unusable becomes valuable — provided it can be found, accessed securely, and is reasonably current.

Not sure whether your data is ready?

Our AI Readiness Assessment reviews your workflows, data, and constraints in two to three weeks and tells you exactly where to start.