Overview
This generated pilot indexes the marker/PyTorch markdown extraction layer without promoting every paper into curated source pages. It is the bridge between the 23,260 raw markdown documents and the public wiki source layer.
Counts
- Raw markdown documents scanned: 23,260.
- Candidate markdown files screened: 2739.
- Machine-tagged food/heavy-metal candidates: 2707.
- Unique candidate FM records after de-duplication: 2596.
- Pilot records emitted: 150.
- Pilot records already promoted to curated source pages: 6.
How ChatGPT Should Use This Layer
- Search or filter the corpus catalog to find potentially relevant papers.
- Treat corpus tags as machine-extracted candidates, not final claims.
- Promote only load-bearing sources into
wiki/sources/. - Put finished-product values on product pages and ingredient-only values on ingredient pages.
- Preserve metal species, units, basis, matrix, geography, method, and review state before using values for HMTc standards logic.
Pilot Indexes
- By metal: Corpus By Metal - Al, Corpus By Metal - Cd, Corpus By Metal - Cr, Corpus By Metal - MeHg, Corpus By Metal - Ni, Corpus By Metal - Pb, Corpus By Metal - Sn, Corpus By Metal - U, Corpus By Metal - iAs, Corpus By Metal - tAs, Corpus By Metal - tHg.
- By product row: Corpus By Product - baby-cereals-dry-non-rice, Corpus By Product - baby-cereals-dry-rice-based, Corpus By Product - fruit-purees, Corpus By Product - infant-formula-powder-non-soy, Corpus By Product - infant-formula-powder-soy-based, Corpus By Product - plant-milks-non-soy-non-rice, Corpus By Product - plant-milks-rice-based, Corpus By Product - plant-milks-soy-based, Corpus By Product - root-vegetable-purees.
- Promotion queue: Corpus Promotion Queue.
Data Files
data/corpus/markdown-corpus-pilot-catalog.ndjsondata/corpus/markdown-corpus-pilot-summary.json