A "Principle Code" asking generative AI companies to voluntarily disclose how they handle training data and intellectual property was finalized on August 25, 2026. It asks companies to publish the types of data they train on and the names of their crawlers on their own websites, and to respond to inquiries from rights holders and users. There is no penalty for ignoring it, but a company that declines to follow a principle is expected to explain why.

Not a Punishment Mechanism, but an Explanation Mechanism

The full title is the Principle Code on the Protection of Intellectual Property and Transparency for the Appropriate Use of Generative AI. It was drafted following discussions at the Study Group on Intellectual Property Rights in the Age of AI, with the Intellectual Property Strategy Promotion Secretariat of the Cabinet Office acting as the point of contact.

The document carries no legal force. Instead it adopts comply-or-explain, an approach long used in corporate governance. A company either implements a principle or explains in sufficient detail why it does not. Choosing the latter is not a violation.

The quality of that explanation, however, is scrutinized. Simply stating that the internal structure for handling inquiries is not yet in place is treated as insufficient; the company is expected to indicate when that structure will be completed. If terms of service contain a clause that limits the application of the principles, pointing to the existence of the clause is not enough either. The company must explain why the clause was introduced.

The code explicitly states that it does not compel disclosure of sensitive information such as trade secrets or details related to safety and security.

What Companies Are Asked to Publish

Principle 1 covers disclosure on a corporate site or an equivalent website that anyone can view. It has two halves: transparency measures and intellectual property protection measures.

On the transparency side, the listed items include the model name and version, its history including release dates, architecture and design specifications, intended and prohibited uses, and the training method. For training data, the scope covers the types of data such as text and images, whether web crawling is performed, how non-public datasets obtained from third parties and public datasets are handled, and whether synthetic data is used and for what purpose. Crawler names, identifiers and collection periods are included, as are the names of any third-party crawlers in use.

The intellectual property side reads as a list of practical measures. Adopt crawlers that respect paywalls and machine-readable instructions such as robots.txt, publish that policy per user agent, and give notice when it changes. Avoid crawling pirate sites. Retain training logs for a set period. Implement technical measures to prevent infringing outputs, and where feasible adopt provenance technologies such as digital watermarking and C2PA. Set up a contact point for rights holders, clarify the requirements for filing a request, and keep records of responses. Companies are asked to disclose how far they have gone on each of these.

The secretariat also released a companion set of concrete examples on the same day, showing the level of detail expected for each item. Entries such as Transformer architecture, obtained from a data vendor, and bot name: collected continuously since a given date appear as samples, which should keep disclosure formats from varying too widely between companies.

Rights Holders Can Ask Whether Their Work Was Used

Principle 2 covers responses to inquiries from rights holders. The scenario in mind is a creator who publishes work on a website and then finds an output that is identical or similar to it. The creator can present the URL of the page hosting the work and ask whether that domain was included in crawling, or whether it was among the sources of training data acquired from third parties.

Not everyone can ask. The right is limited to people who are actually pursuing or preparing litigation, mediation, or ADR (alternative dispute resolution), along with attorneys they have retained. The requester must show reasons supporting that status, state the intended use of the answer, pledge not to use it for any other purpose, and identify the URL to be checked along with the grounds for the request.

The scope of the answer is narrow: whether the presented URL is included. The code notes that this is not a mechanism for asking comprehensively whether a work itself is part of the training set. If a generative AI provider cannot answer, it responds with the name of the company that developed the model embedded in the service.

Users Who Generated the Content Can Ask the Same Question

Principle 3 covers inquiries from users who generated content with a service. The example given is a user who created an image and then found a closely similar image on some website.

Here there is no requirement to be preparing legal action. In exchange, the user must present the generated output and the prompt used to create it, and must pledge not to use the answer for litigation, mediation, or ADR filings. The design positions this as a window for confirming whether your own output is safe to use, rather than a tool for asserting rights.

For both Principle 2 and Principle 3, the code allows fees or limits on the number of requests within a given period to prevent abuse. It also cautions against settings that would discourage people from asking at all.

Who Is Covered and Who Is Not

The code applies to generative AI developers who build systems and make them available to the public, and to generative AI providers who offer services built on top of them. Companies without a head office or principal place of business in Japan are still covered if their services are offered to the Japanese market, so overseas firms are not exempt.

The exclusions are spelled out just as clearly. Subcontractors who handle only part of the pipeline, such as data collection or training, under commission are in principle not covered. Neither are those doing research and development without offering anything to the public. A company that builds an internal generative AI system using only one client's data and supplies it to that client alone falls outside the scope. So do systems where the risk of producing infringing output is extremely low, such as those whose outputs amount to statistical data or inference results.

Participating Companies Will Be Listed Publicly

A company that accepts the principles publishes its acceptance, the measures it implements, and the reasons for any principle it does not implement on its own site, then files a notification with the Intellectual Property Strategy Promotion Secretariat using the prescribed form. The content is reviewed annually in principle, and updates are announced when made.

The secretariat plans to publish a list of the companies that have filed, along with links to their disclosure pages. It states plainly that it will not examine the content of those filings and will not respond to third-party inquiries about them. Its role ends at putting the information side by side; evaluation is left to the market. The start date for accepting filings will be announced separately.

Summary

What goes into training data has so far been left to each company's discretion. The Principle Code introduces a choice into that space, not a penalty: disclose, or explain why you will not. Its effectiveness will depend on how many companies end up on the filing list and how specific their explanations become. A provisional English translation was released alongside the Japanese text, so how far it reaches overseas firms is worth watching as well.