Draft:International Standard Content Code
Review waiting, please be patient.
This may take 4 months or more, since drafts are reviewed in no specific order. There are 4,566 pending submissions waiting for review.
Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
Reviewer tools
|
Submission declined on 8 June 2026 by Stuartyeates (talk). The text as is stands is so one-sided it borders on misrepresentation. For example, when is says "Traditional identifiers are assigned by registration authorities and identify abstract works or specific editions" it ignores at least 60 years of the use of hashes for identifiers.
Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
This draft has been resubmitted and is currently awaiting re-review. |
Submission declined on 9 February 2026 by ChrysGalley (talk). This draft appears to contain text generated by a large language model (such as ChatGPT). You cannot use LLMs to generate article content.
Declined by ChrysGalley 5 months ago.LLM-generated pages with certain obvious signs of being machine generated may be deleted without notice. These tools are prone to specific issues that violate our policies:
Instead, only summarize in your own words a range of independent, reliable, published sources that discuss the subject. See the advice page on large language models for more information. |
Comment: Notability basis: independent significant coverage in NISO/Carpenter (2024), the joint IEC-ISO-ITU AMAS technical report (2025), and the peer-reviewed Fraunhofer paper in Electronic Imaging (2025). The subject also meets WP:NSTANDARD as a published ISO standard (ISO 24138:2024) that has been adopted nationally as AS/NZS ISO 24138:2025. Etma 1222 (talk) 10:42, 4 June 2026 (UTC)
Comment: In accordance with the Wikimedia Foundation's Terms of Use, I disclose that I have been paid by my employer for my contributions to this article.Correction: this disclosure was added automatically by the Article Wizard in error. I am not paid to edit Wikipedia. I do have a conflict of interest as an unpaid board member of the ISCC Foundation, which is separately disclosed on my user page and on the talk page per WP:COI. Etma 1222 (talk) 11:12, 23 April 2026 (UTC)
Comment: Sources do not support the text. E.g. source 13, https://iscc.codes/conceptdoes not support base32. The CODE and UNIT composition is also not directly mentioned. ChrysGalley (talk) 19:44, 9 February 2026 (UTC)
Comment: Thanks for the review. I've revised the draft to address the points raised:* The lead and a new Background section now explicitly cover the long history of hash-based file identification (cryptographic hashes, version control, content-addressable storage) and perceptual fingerprinting systems (PhotoDNA, Chromaprint, Content ID), and position the ISCC relative to both rather than presenting it in isolation.* The comparison table now includes cryptographic hashes and perceptual fingerprints alongside registration-based identifiers, instead of comparing the ISCC only against registry identifiers.* Self-published sources (iscc.codes history pages, IEP specifications) have been removed; technical claims now rest on the ISO standard and the NISO article.* Promotional or weakly sourced material (reception quotes, repository tables, blog-sourced sections) has been removed.Happy to make further changes. Etma 1222 (talk) 22:56, 10 June 2026 (UTC)
| International Standard Content Code (ISCC) | |
|---|---|
| Abbreviation | ISCC |
| Status | Published |
| Year started | 2016 |
| First published | 15 May 2024 |
| Latest version | ISO 24138:2024 |
| Organization | ISO/TC 46/SC 9 |
| Domain | Digital media content identification |
| Website | iscc |
The International Standard Content Code (ISCC) is a content-derived identifier for digital media assets, standardized as ISO 24138:2024.[1] An ISCC is computed algorithmically from a media file, covering text, images, audio, or video, rather than assigned by a registration authority. Because it is derived from the content, independent parties can compute the same code from the same file without central coordination.[2]
Identifying digital files by a value computed from their data is a long-established practice; cryptographic hashes have been used this way for decades, for example to address content in version control systems and content-addressable storage. Such hashes change completely when a single bit changes. The ISCC differs in that several of its component codes are similarity-preserving, so that perceptually similar content yields numerically similar codes, and in that it combines multiple content-derived codes (covering metadata, content, data, and a cryptographic checksum) into a single composite identifier defined by an international standard.[1]
Background
[edit]Persistent identifiers for creative works, such as the ISBN for books or the DOI for scholarly documents, are assigned by registration authorities and refer to an abstract work or a specific edition. Several of these schemes are maintained within ISO/TC 46, the ISO technical committee for identification and description.[2]
A separate practice identifies a digital file by computing a value directly from its data. Cryptographic hashes are widely used for this purpose but, by design, produce an unrelated output for any change to the input. To recognize content that has been re-encoded, resized, or otherwise modified, media-matching systems use perceptual hashing and fingerprinting techniques instead, including acoustic fingerprinting for audio and perceptual hashes for images. Examples of such systems include PhotoDNA, Chromaprint, and Content ID.
The ISCC draws on both practices. It is computed from content like a hash, several of its components are similarity-preserving like a perceptual fingerprint, and it is published as an international standard intended for use alongside existing registration-based identifiers.[1]
History
[edit]Development
[edit]The ISCC was created by Titusz Pan (ORCID 0000-0002-0521-4214), a software developer working on content identification.[3][4] Development began in 2016 within the Content Blockchain Project, an initiative funded by the Google Digital News Initiative and undertaken by a consortium of publishing, legal, and technology organizations including Zeit Online, Golem.de, and dpa.[5]
The project examined how content identification could support journalism and digital media authenticity. In 2019 it received one of Germany's inaugural Digital Publishing Awards at the Leipzig Book Fair; the jury described it as "one of the few blockchain projects to have an immediate utility for the publishing industry."[6]
Standardization
[edit]In May 2019, the International Organization for Standardization (ISO) established Working Group 18 within Technical Committee 46, Subcommittee 9 (TC 46/SC 9, Identification and description) to develop the ISCC as an international standard.[7]
ISO 24138:2024 Information and documentation - International Standard Content Code (ISCC) was published in May 2024 following review by national standards bodies. The standard defines the syntax, structure, and algorithms for generating ISCCs, and describes their use together with existing identifier schemes including the DOI, ISAN, ISBN, ISRC, ISSN, and ISWC.[1][2]
Adoption
[edit]The ISCC was added to ONIX, the book-trade metadata standard maintained by EDItEUR, in July 2020, allowing publishers to include a content-derived identifier alongside product identifiers such as the ISBN.[8]
In May 2024, the ISCC was added to the approved soft-binding algorithm list of the Coalition for Content Provenance and Authenticity (C2PA) as the first fingerprint-based algorithm on the list. C2PA soft-binding algorithms support recovery of content-provenance manifests after metadata has been removed from a media asset.[9]
Design
[edit]An ISCC-CODE is a composite code assembled from one or more ISCC-UNITs, each derived from a different aspect of the digital content.[1]
Units
[edit]The standard defines five unit types:[1]
- Meta-Code: A similarity hash derived from referent metadata (name, and optionally description and other metadata), supporting clustering of assets with similar metadata.
- Semantic-Code: Reserved in ISO 24138 but not standardized. Experimental implementations use deep learning embeddings to capture the semantic meaning of text and images.
- Content-Code: A modality-specific perceptual fingerprint:
- Text-Code: based on normalized character n-grams
- Image-Code: based on perceptual hashing
- Audio-Code: based on Chromaprint acoustic fingerprinting
- Video-Code: based on MPEG-7 frame signatures
- Data-Code: A similarity-preserving hash of the file's bytes using content-defined chunking with MinHash.
- Instance-Code: A BLAKE3 cryptographic hash for exact data-integrity verification.
Each unit can be used on its own or combined into an ISCC-CODE. Unit bodies support variable bit-lengths in 32-bit increments, from 32 up to 256 bits, with a default of 64 bits, while remaining prefix-compatible. A minimal valid ISCC-CODE requires the Data-Code and Instance-Code; the other units are optional.[1]
Format
[edit]Each ISCC-UNIT consists of a variable-length header encoding the main type, subtype, version, and bit length (two bytes for all currently defined units), followed by a variable-length body containing the fingerprint data. Units are concatenated and encoded with Base32 for the canonical string representation.[1] An ISCC-CODE that combines Meta-Code, text Content-Code, Data-Code, and Instance-Code units appears as:
ISCC:KACT4EBWK27737D2AYCJRAL5Z36G76RFRMO4554RU26HZ4ORJGIVHDI
Similarity preservation
[edit]Several ISCC units are similarity-preserving. In contrast to a cryptographic hash, where any change to the input produces an unrelated output, the Hamming distance between two such ISCC units correlates with the similarity of the underlying content, which supports applications such as near-duplicate detection and clustering.[1]
Relationship to other identifiers
[edit]The ISCC belongs to the family of content-derived identifiers rather than registration-based ones. Unlike a plain cryptographic hash, several of its units are similarity-preserving, and unlike most fingerprinting systems it is defined by an international standard.
| Approach | Examples | Basis | Similarity-aware |
|---|---|---|---|
| Registration-based identifiers | ISBN, DOI, ISRC, ISWC | Assigned by an authority for a work or edition | No |
| Cryptographic hashes | SHA-256, BLAKE3 | Computed from file data | No |
| Perceptual hashes and fingerprints | PhotoDNA, Chromaprint, Content ID | Computed from perceptual features | Yes, within a modality |
| ISCC | — | Composite of content-derived units, standardized in ISO 24138 | Yes |
Registration-based and content-derived identifiers serve different purposes, and ISO 24138 positions the ISCC as complementary to existing schemes rather than a replacement, describing its use together with the DOI, ISAN, ISBN, ISRC, ISSN, and ISWC.[1] Because the ISCC identifies a digital manifestation, a single book edition (one ISBN) may correspond to many ISCCs representing different digital formats, compression levels, or excerpts.[2]
Applications
[edit]Reported and proposed applications of the ISCC include:
- Content deduplication: identification of exact and near-duplicate content.
- Integrity verification: detection of unauthorized modification using the Instance-Code.
- Version tracking: relating different versions of a work.
- AI training-data management: researchers at the Fraunhofer Institute for Secure Information Technology (Fraunhofer SIT) proposed the ISCC as a robust hashing method for EU AI Act compliance, including the identification of AI-generated content and the handling of text and data mining opt-out declarations.[10]
CommonsDB
[edit]CommonsDB is a European Commission-funded pilot that uses the ISCC as the primary identifier for a public registry of rights information for public domain and openly licensed works.[11] Led by Open Future with the Europeana Foundation, Wikimedia Sverige, and the Institute for Information Law, the pilot lets users check the rights status of content by generating an ISCC from a file and searching for matching declarations.[12]
Implementation
[edit]The reference implementation of ISO 24138 is published as open-source software by the ISCC Foundation, comprising a core algorithm library (iscc-core), a Python software development kit, and a web service used for a public demonstration.[13] A third-party TypeScript port, iscc-core-ts, provides an implementation for JavaScript environments.[13]
ISCC Foundation
[edit]The ISCC Foundation (Stichting ISCC) is a nonprofit organization based in Hengelo, Netherlands, established as a Dutch Stichting.[14] It maintains the reference implementation and coordinates standardization and adoption activities related to the ISCC. Its advisory board includes representatives from the National Information Standards Organization (NISO), the publishing industry, and library science.[14]
See also
[edit]- Content ID
- Digital object identifier
- Perceptual hashing
- Acoustic fingerprint
- Content Authenticity Initiative
- Cryptographic hash function
References
[edit]- ^ a b c d e f g h i j "ISO 24138:2024 Information and documentation - International Standard Content Code (ISCC)". International Organization for Standardization. 2024-05-15. Retrieved 2026-06-11.
- ^ a b c d Carpenter, Todd A. (2024-06-04). "Introducing the Newest ISO Identifier Standard". National Information Standards Organization. Retrieved 2026-06-11.
- ^ Nawotka, Ed (2025-10-15). "Frankfurt Book Fair 2025: Identity Stamps". Publishers Weekly. Retrieved 2026-06-11.
The ISCC framework was developed by Titusz Pan, Amlet's founder and CTO and the code's inventor.
- ^ Pan, Titusz (2024). "ISCC: Neue Perspektiven für die KI-gesteuerte Identifikation von Inhalten". Information - Wissenschaft & Praxis. 75 (5–6). De Gruyter: 253–259. doi:10.1515/iwp-2024-2032.
- ^ "Content Blockchain Project". Credibility Coalition. Retrieved 2026-06-11.
- ^ Anderson, Porter (2019-03-25). "Content Blockchain Project Wins One of Germany's Digital Publishing Awards". Publishing Perspectives. Retrieved 2026-06-11.
- ^ "ISO establishes Working Group for ISCC". ISCC Foundation. 2019-05-23. Retrieved 2026-06-11.
- ^ "ONIX for Books Codelists Issue 50". EDItEUR. 2020-07-09. Retrieved 2026-06-11.
- ^ "Soft Binding Algorithm List". C2PA. Retrieved 2026-06-11.
- ^ Heeger, Julian; Berchtold, Waldemar; Bugert, Simon; Steinebach, Martin (2025). "EU AI-Act: Tagging GenAI Content". Electronic Imaging. 37 (4). Society for Imaging Science and Technology: MWSF-301. doi:10.2352/EI.2025.37.4.MWSF-301. Retrieved 2026-06-11.
- ^ "CommonsDB". Europeana. Retrieved 2026-06-11.
- ^ "CommonsDB". Wikimedia Meta-Wiki. Retrieved 2026-06-11.
- ^ a b "ISCC Foundation". GitHub. Retrieved 2026-06-11.
- ^ a b "Foundation". ISCC Foundation. Retrieved 2026-06-11.

