We examined how media content and journalistic data could be used in AI development in a way that is lawful, transparent, and respectful of intellectual property rights.

It quickly became clear that there is currently no simple, legally robust way to license media content for AI use.

This work led to a pilot in which Finnish news media organisations are testing content use and pricing models with AI developers.

What Did We Do?

Sitra’s project set out to build a shared understanding of how media content is used in AI development and to define the conditions for a marketplace where journalistic content could be licensed for AI training and AI-powered services.

The organisations that participated in the project include major media outlets in Finland: Alma Media, Kaleva, Keskisuomalainen, MTV, Otavamedia, Sanoma, Yleisradio, and the Finnish News Media Association.

The work was grounded in a legal and business background report commissioned from Geradin Partners. The report reviewed AI use cases, copyright implications, and the European regulatory framework. We then held a series of workshops to examine specific content use cases, marketplace implementation options, stakeholder roles, and possible governance models.

This work resulted in a pilot in which participating media companies, supported by copyright management organisation Kopiosto and News Media Finland, are seeking AI developers to help shape a licensing framework for lawful media content use. The pilot also tests what kind of marketplace could best support copyright-respecting AI development in practice.

Starting Points from the Media Industry’s Perspective

Locking everything behind a fence and tightening restrictions is not necessarily the right approach, but openness requires clear rules of the game.

Discussions among media companies showed a strong interest in controlling who uses media content for AI development and how. The use of articles, videos, and audio produced by media organisations should be transparent, clearly bounded, and governed by contracts. Protecting journalistic values, information traceability, and brand visibility emerged as core priorities.

New AI-driven business models for media have not yet taken shape, and pricing professionally produced content for AI training remains difficult. Media companies want any compensation to reflect the human labour involved in producing the content.

Early discussions with AI developers suggest that training material is most valuable when it is available in large volumes and usable formats. From their perspective, whether any single carefully crafted article is included in the dataset matters relatively little.

Revenue-sharing models drew strong interest among publishers, but also raised questions about long-term consequences. Publishers want to take part in AI development, but they also see a clear risk in using their own content to strengthen potential competitors.

Key Lessons

As the project advanced, several key lessons emerged:

  • Content can be used for AI training and services in many ways. Defining use cases – what content is used for and what that means technically – helped structure the discussion and clarify the copyright aspects in each scenario. 
  • The value of media data in AI stems from both quality and timeliness, not volume alone. 
  • There is currently no simple, legal route for licensing media content directly from publishers for AI use. Intermediaries have emerged to acquire rights, especially for image and audio material, which they organise by theme or use case and sell in bundles. A more flexible model that preserves content brands and is built for European legislation has yet to emerge. 
  • A centralised data intermediary model can support new data products, but only if content governance and licensing are trusted. 
  • The marketplace must address competition law risks. Pricing coordination and the exchange of commercially sensitive information between competitors must be strictly avoided. 
  • Finland alone is too small a market. Success will require collaboration and consistent rules of the game at a minimum across the Nordic countries, and ideally at the European level.  

What Comes Next?

Based on this exploratory work, a pilot was proposed to test content use and pricing with AI developers. Its aim is not to settle every legal or economic question, but to assess demand and model how content could be governed and commercialised in the AI era.

The pilot, which will finish by the end of 2026, is seeking companies developing AI services that are interested in licensing media content. During the trial, participating companies can license media content for AI training or for use within their services.

All media companies involved in Sitra’s project are taking part in the pilot. Their main goals are to understand:

  • which content types are most useful, and for which training purposes; 
  • how many potential users exist in Finland and internationally; 
  • how AI developers assess the value of the material. 

For AI developers, the pilot provides a direct channel to key Finnish media organisations and an opportunity to shape the operating model and data terms of a future marketplace.

Why Does This Matter?

How media data is used in AI is a major societal question. High-quality journalism is a fundamental part of reliable information provision and democratic debate. At the same time, it provides valuable supplementary information and training data for AI systems.

Without workable rules, value may flow away from rights holders while technological development advances at the expense of the media sector. Without fair value distribution or new business models, media organisations may lose the economic basis for producing human-generated content. In the worst case, the consequences could be deeply concerning.

Further Reading

The background report commissioned for this project has been published on the website of our partner, Geradin Partners.

Contact

Anssi Komulainen

Senior Lead, Democracy Innovations Programme

See also