What Open Data Stewards Can Learn from Corporate Guidance on AI-Ready Data

For two decades, the open data community has worked to articulate what it means to make data fit for use, converging on the now-canonical FAIR principles: Findable, Accessible, Interoperable, and Reusable. As we argue in Moving Toward the FAIR-R Principles, this framework, though foundational, is no longer sufficient in an era of artificial intelligence. It must be extended to include a fifth commitment: Ready for AI.

The conversation over data readiness for the AI age is not confined to the public-interest domain. Parallel discussions have unfolded in the corporate sector, where the imperatives of AI readiness have been confronted at scale, and it is worth asking what the open data movement might learn from them.

McKinsey’s recent article, AI Data Readiness: The Key to Scaling Impact, is instructive in this regard. On its surface, the article reads as a memo to chief data officers diagnosing why corporate AI pilots so often stall before reaching production. It also offers some valuable lessons and principles for data stewards operating in the public interest, and who may be considering how to apply FAIR-R principles in practice.

In what follows, we treat this private-sector-focused article as a source of transferable knowledge while still remaining attentive to its limits. The caveat about limits is consequential, and we return to it below: in general, while enterprises optimize primarily for reliable outputs and cost avoidance, public-interest stewardship must also account for equity, openness, and public accountability. Some lessons therefore transfer cleanly, while others may require translation (or simply not apply at all).

To be clear, our premise is not that public-interest data work should become more corporate. It is, rather, that firms have been contending for some time with the mechanics of AI-ready data (at scale, under audit, and with material consequences for failure) and that the open data movement stands to learn from at least some of these experiences. While this piece focuses on a specific article from McKinsey, many of these observations are more generalizable to the private sector and the way it is engaging with data more broadly.

Six lessons worth borrowing

1. Govern the fragment, not only the file, because AI dissolves the dataset.

This is perhaps the single most valuable idea to import. FAIR-R, like the FAIR tradition it extends, largely treats the dataset as the object that needs to be made ready. McKinsey’s article, on the other hand, observes that the dataset is no longer the operative unit within an AI pipeline. A single PDF or data set fragments into extracted text, parsed tables, image summaries, chunks, embeddings, and index entries, each a derived artifact capable of independently shaping an output. A document may be wholly accurate in its entirety yet yield an erroneous answer when an outdated or incomplete fragment is retrieved in isolation.

The implication for open data stewards is significant: provenance, sensitivity labels, and usage rules applied at the level of the dataset no longer survive the ingestion into an AI system. It therefore implies that every chunk carries its lineage, not merely every file.

2. Define “good enough” for each use case, and treat quality as a continuous process.

The article emphasizes that companies have learned to resist the instinct to perfect data before deployment. McKinsey’s framing calibrates quality to the use case and its attendant risk profile: the governance appropriate to an internal drafting tool differs markedly from that required for a regulated decision, where full lineage and auditability become essential (see illustration below).

For open data stewards, this reframing is methodologically valuable, in that it replaces the untenable mandate of universal data quality with a defensible, risk-tiered standard (while still countering the emerging assumption that data quality is of diminished importance in AI environments). The corollary is that quality assurance shifts from a one-time gate applied at publication to a continuous discipline exercised across extraction, retrieval, and generation.

Companies can define what

From McKinsey Article: AI data readiness: The key to scaling impact June 23, 2026

In most cases, McKinsey’s metric of success is financial. For example, the report documents how, in the absence of shared infrastructure, teams often rebuild their own extraction logic and retrieval configurations, producing inconsistent outputs and duplicating cost. By contrast, McKinsey shows how a financial services firm converted reusable pipelines into an estimated $10–20 million in cost savings as use cases multiplied.

Translated into the open data ecosystem, this mechanism for cost savings is more than simply an argument for reusable infrastructure; it is an argument for reuse as an organizing principle, and it resonates with a commitment the open data world already holds, for example, in its championing of the data commons. The commons is not simply a shared asset to be drawn upon; appropriately designed, it can be a governed arrangement built deliberately to enable equitable, accountable reuse across many actors and purposes. While FAIR-R frames the commons primarily as a mechanism for equity, the corporate evidence demonstrates that the same architecture can also yield efficiency. For under-resourced public institutions, this dual justification is strategically consequential: shared, governed pipelines are both fairer and less costly, an argument that carries particular weight with those who control budgets.

4. Make traceability infrastructure.

When a single answer stitches together fragments from many documents, the ability to explain where the answer came from (i.e., the output of an AI system) collapses. McKinsey calls the result “indefensible” outputs, which are problematic in regulatory audits or legal discovery.

Public-interest deployments face an even sharper version of this problem: humanitarian, health, and civic uses require accountability to the people the data describes. According to McKinsey, corporations treat lineage as an essential characteristic with owners, versions, and retirement dates. Open data stewards should borrow that seriousness. A traceable chain of custody from source through every derived artifact to the final output is the deliverable, not a nice-to-have appended to it.

5. Expect the data steward’s role to broaden.

McKinsey describes an expanding mandate for chief data officers, from owning pipelines and warehouses to owning the standards and conditions. The article is explicit that this is a shift in kind rather than degree: the CDO becomes responsible not for the data alone but for the conditions under which data can be reused, traced, and governed consistently across the enterprise; linking structured and unstructured content, anchoring both to shared business entities, and ensuring that the tools and skills AI systems rely on behave predictably. As that mandate widens, McKinsey notes, the surrounding roles begin to blend, with data, engineering, product, and governance functions converging and organizations obliged to cultivate multifaceted skill sets rather than treating these as separate domains.

Our work on FAIR-R implies the same trajectory for public-interest data stewards. The lesson is to anticipate the shift deliberately: the data steward of the near future architects the conditions for systematic, sustainable, and responsible reuse rather than curating static holdings.

6. Extend readiness to unstructured and machine-generated data.

The FAIR tradition developed around structured, discrete datasets, yet the McKinsey report locates the AI-readiness problem primarily in unstructured content (documents, emails, transcripts, images, and video), which now drives a growing share of consequential decisions and which does not “stay whole” as it passes through extraction, chunking, and embedding. McKinsey also observes that AI systems do not merely consume data but generate it continuously: prompts, responses, summaries, and decisions accumulate and frequently flow back into core systems, where they seed further outputs and create feedback loops over time.

For open data stewards, this reality carries two implications. First, readiness cannot be confined to the tidy, tabular holdings the FAIR principles were first designed around; the harder and more valuable work increasingly concerns unstructured sources, whose provenance and sensitivity are far less legible. Second, the movement must begin to reckon with synthetic and machine-generated data as objects of stewardship in their own right (a matter we raised earlier), since data produced by AI systems can propagate error, especially if reused without lineage or quality controls. Readiness, in other words, is not a property to be established once at publication but a discipline that must follow data through its generation as well as its consumption.

What not to borrow

Borrowing is not copying. As we stated at the outset, we are interested in translation rather than repetition. Two differences in particular are important to highlight.

First, corporations govern risk to protect the institution: its outputs, its liability, its costs. Public-interest stewardship also governs risk, but to protect data subjects and the public. Our FAIR-R framework is more pronounced on this, insisting, for instance, that data be evaluated for bias and that a dataset that cannot be de-biased simply should not be used. McKinsey’s risk framing, managing exposure after AI accesses data, and applying human oversight proportional to business risk supply excellent mechanics, but the threshold for what counts as unacceptable risk must be set by public-interest values, not solely for-profit business ones.

Second, the corporate default is enclosure: governed and controlled, with data treated as a proprietary asset. The open data default, by contrast, is responsible openness. The reusable-foundation lesson is sound, but data stewards must resist letting “governed” quietly become fully “closed.” The goal is governed openness: commons that are equitable and accountable and that guard against extraction.

To conclude, McKinsey’s report can be of value to the open data world not as a mirror but as a field report from those who are rapidly experimenting with what AI-readiness means. Its transferable lessons provide valuable frontline experience for our FAIR-R framework and perhaps add some cautionary lessons as well.

Thanks to Adam Zable and Akash Kapur for editorial review.

Stefaan Verhulst headshot

Author

Stefaan Verhulst

Course Lead · Data Stewards Founder

Dr. Stefaan Verhulst is Co-Founder of the DataTank and The GovLab and the main lecturer of the data stewardship academy. In addition, he is a Research Professor at the Center for Urban Science and Progress at the Tandon School of Engineering of New York University; and a Senior Advisor to the Markle Foundation where he spent more than a decade as Chief of Research. He is also the Editor-in-Chief of the open-access journal Data & Policy (Cambridge University Press); the Research Director of the MacArthur Research Network on Opening Governance; Chair of the Data for Children Collaborative with Unicef; a member of the High-Level Expert Group to the European Commission on Business-to-Government Data Sharing; and of the Expert Group to Eurostat on using Private Sector data for Official Statistics. In addition he is also a member of the UNESCO Information Ethics Working Group; Researcher at the ISI Foundation (Torino, Italy); Senior Researcher at SMIT (Studies in Media, Innovation and Technology) at the Free University of Brussels (VUB) . In 2018 he was recognized as one of the 10 Most Influential Academics in Digital Government globally (by the global policy platform Apolitical). Previously at Oxford University, he was the UNESCO Chairholder in Communications Law and Policy and co-founded and was the Head of the Program in Comparative Media Law and Policy at the Center for Socio-Legal Studies. He was the Socio-Legal Fellow at Wolfson College, and is still an emeritus fellow at Oxford. He also taught for several years at the London School of Economics and was Co-Founder and Co-Director of the International Media and Info-Comms Policy and Law Studies (IMPS) at the University of Glasgow School of Law. He has published widely - including seven books- and his writings and work have appeared in the Harvard Business Review, Stanford Social Innovation Review, Project Syndicate, Wall Street Journal, and The Conversation (among many other outlets). He is asked regularly to present at international conferences including, for instance, TED, Collision, and the UN World Data Forum. Numerous organizations have sought his counsel on a variety of topics including data and AI governance - including the WorldBank; IDB, CAP, USAID, DFID, IDRC, AFP, the European Commission, Council of Europe, the World Economic Forum, UNICEF, OECD, UN-OCHA, UNDP, UNESCO and several other international and national private and public organizations. He is also a Linkedin Learning instructor seeking to democratize the practice of data stewardship globally.

AIDataData Stewardship