Led by Stefaan Verhulst, The Data Tank and The GovLab organized a summer session of the Data Stewardship Trends to Watch series. 

Trends to Watch brings together our global data stewards alumni every two months to explore and discuss the latest developments shaping the data stewardship landscape.

In this session, we focused on the ways that AI is accelerating the data winter phenomenon and new mechanisms to address this phenomenon—data commons and other intermediary data ecosystems. Participants also considered the increased importance and focus on non-traditional data. 

Check out a recording of the full conversation here. The following resources were discussed during the session:

Trends: 

Data Policy Trends 

  • Anticipating Data Policy in the Age of AI is a blog based on two forecasting studios with data governance experts convened by the GovLab. This piece from Stefaan Verhulst and Adam Zable identifies seven signals shaping the future of data access and reuse. The signals are summarized below: 

Screenshot 2026 08 04 at 16.28.20

AI and Data Winter

  • Google Was a Lifeline for Publishers. Now Some Are Thinking of Cutting It Off. This piece reports on how Reddit, the online message board that powers many Google AI search results, has discussed shutting off Google’s access to its content for AI use. This is part of a larger discussion of online media companies expressing concerns with the way their data is being used for AI. 

  • The Attribution Crisis in LLM Search Results is a working paper series that touches on the ‘attribution gap’ which is the difference between relevant URLs read by web-enabled LLMs and those actually cited. The paper documents three exploitation patterns: 1) No Search: 34% of Google Gemini and 24% of OpenAI GPT-4o responses are generated without explicitly fetching any online content; 2) No citation: Gemini provides no clickable citation source in 92% of answers; 3) High-volume, low-credit: Perplexity’s Sonar visits approximately 10 relevant pages per query but cites only three to four.

  • AI Vision by GPTZero is a visualization of millions of posts daily scans to provide actual information on AI-generated content on major platforms. A concerning figure is the estimation that by 2031 100% of posts on Linkedin will most likely be AI-generated. 

  • AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop captures another concerning phenomenon where as AI companies are searching for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. 

  • What U.S. Restrictions on Satellite Imagery Mean for Iran Reporting touches on another aspect of data winter where the U.S. satellite providers cut off access to high-resolution images of Iran and surrounding countries shortly after the war began. Satellite imagery has historically been used as a great source for investigative journalists and this turn poses many threats.

  • White House deletes thousands of webpages about energy conservation as heatwave slams US. This piece reports on The US Department of Energy reportedly deleting about 6,000 pages related to energy conservation as a historic heatwave tears across the country. It serves as another testament that the data winter is real. 

  • Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident provides a detailed overview of the recent Open AI agent incident with Hugging Face. The pieces describes howOpenAI researchers had given this unreleased AI a series of challenging test questions. When unable to find the solution the agent broke out of its testbed and raided Hugging Face’s servers in search of the answer key. 

  • How AI Companies Can Pay Fair Rates for the Content They Need is another piece that touches upon the fight over publishers and the data that trains artificial intelligence becoming one of the defining economic conflicts of the decade. The writing offers a set of potential solutions to this fight. 

  • Attribution in the Age of AI provides an optimistic approach that despite a recent study by the AI Disclosures Project finding that more than 30% of responses by leading LLMs provided no attribution whatsoever, there are still some reasons to be optimistic that AI can attribute the sources they use. The piece outlines the reasons. 

  • Project Provenance addresses how people understand and use signals about the origin and history of digital content, focusing on end-user understanding through human-centered design and empirical study. In this regard, Microsoft co-founded C2PA (Coalition for Content Provenance and Authenticity), an industry standards effort in this space. 

  • ITU-T Focus Group on Trust and Identity for Humans and Agentic AI (FG-TIDA) is a focus group that will address trust management and interoperable digital identity infrastructure for humans and for agentic AI, supporting the development of secure, trustworthy digital ecosystems in which humans and agentic AI can safely interact and collaborate. 

Data Commons

  • Fed up with Big Tech, communities turn to data collectives for control is a piece that shines a light on the growing pushback against big tech companies, and greater awareness of the value of data which is spurring interest in data collectives and cooperatives. Mozilla Data Collective is mentioned as one of the examples that communities trust to share their data as they can determine the terms and conditions of how their data is used by big tech companies moving forward. 

  • Feasibility study European Books Data Commons is a study that paves the way for European libraries to make the full text of millions of digitised books available for research, innovation and the development of trustworthy AI applications. 

  • Newcommons.ai is another initiative that offers a six-month capacity-building incubator program focused on supporting indigenous-led teams in developing a fundable data commons proposal. Data commons are collaboratively governed data ecosystems designed to pool and provide responsible (and governed) access to diverse, high-quality datasets from one or multiple sectors to enable the development and deployment of generative AI applications that address public-interest challenges.

  • Empowering people through data intermediaries is another open consultation initiative launched by the UK Government that emphasizes the need for data intermediaries amidst growing concerns around data that feeds AI. 

AI-Ready Data

  • AI data readiness: The key to scaling impact is a piece that’s built on the argument that only 7 percent of companies have fully scaled AI across their organizations, and data is the bottleneck. It notes that more than two-thirds of high-performing companies cite data as their primary obstacle. The piece names four shifts stalling scaling: unstructured data losing traceability as it's chunked and recombined, expanding risk from AI assembling context in real time, AI generating data faster than governance can track, and fragmentation pushing chief data officers into a more central role. 

  • What Open Data Stewards Can Learn from Corporate Guidance on AI-Ready Data is a piece by Stefaan Verhulst that reads McKinsey’s AI data readiness through the lens of FAIR-R—extended FAIR principles (Findable, Accessible, Interoperable, Reusable) with a fifth commitment: Ready for AI. The blog extracts six transferable lessons for open data stewards in this aspect. 

  • Agentic AI Is Rewriting the Rules of Data Risk Management is another piece that focuses on how Agentic AI is shifting how data is created, accessed, and acted upon across the enterprise. It argues that a fundamental reframing of data risk itself is needed. It calls for companies needing a shared data risk taxonomy that aligns privacy, cybersecurity, data governance, and business leaders around a common view of exposure, among others. 

  • Pokémon Go data trained AI that could assist military drones in war zones reports on a concerning phenomenon on an AI model being trained on data collected from users of Pokémon Go which will help military drones find their location in war zones. 

Non Traditional Data 

  • The Re-Use of Non-Traditional Data for Public Interest Purposes is a curated compilation of 100 case use cases (2024–2026). The list divides the cases across 8 themes: Digital Communication and Online Interaction Data, Mobility and Geolocation Data, Health, Wellness, and Biomedical Data, Financial, Commercial, and Consumption Data, Work, Education, and Skill Data, In-Home and IoT Device Data, Media, Entertainment, and Advertising Interaction Data, and Environmental, Geospatial, and Infrastructure Data. It ends with providing a set of reflections. 

  • Design and Implementation of Mobile Phone Data Initiatives is a recently published practical manual that provides an in-depth guide to planning and sustaining a Mobile Phone Data (MPD) initiative, with a primary focus on the use of Call Detail Records (CDRs) for public policy, statistical, and development purposes, including operational decision-making.

  • Non-Traditional Data and the Challenge of Measurement in the United States is a piece focusing on the challenge of measurement in the United States. The main argument is that the ground beneath official statistics is changing with survey response rates falling, collection costs rising, privacy concerns mounting, and public trust in government information eroding. The idea is that while statistical systems were built for a slower more uniform economy, non-traditional data is a viable addition, if the re-use of it becomes more adopted and reliable. 

Other Updates

Finally, in conclusion of the session, the group was updated with some upcoming events and opportunities: 

Interested to learn more about data stewardship trends? Sign up for the Data Stewards Network mailing at this link. Or contact contact@datastewards.net to learn more.

Author

Paulina Behluli & Andrew Zahuranec

datastewardshiptrendsAIdatastewards