Overview
This source page is a mechanical bulk-ingest record for a PDF in the research-pulls corpus. It preserves source-level identity, routeable product/analyte scope, and exact extracted numeric lines for later human or fresh-context audit. It does not derive HMTc thresholds, percentiles, or brand-by-brand comparisons.
Key numbers
The worker extracted the full PDF text with layout preservation twice and compared extraction hashes before commit. The following lines are copied from numeric/table-bearing regions of the PDF and retain the source units and wording where legible:
- from C. difficile is 59% identical to the E. faecalis enzyme and is annotated as functionally
- the 15,511-species calibration set. This value is just 0.3% of species, which are all anaero-
- 52 conserved regions. It occurs as a selenoprotein in 44 of the 45 (98%) regions. The
- notes are listed in Table 1. All members of all nine family are found exclusively in bacte-
- codon. Members of the family are mutually more than 45% identical to each other. In a few
- and their side chains conserve to form what again appears to be a metal-binding site. In about 40% of SaoX
- metal atom. More than 60% of the members of the SaoB family are selenoproteins. As
- erature detectable by PaperBLAST (14), and no protein family models defined by major
- Table S1 in the supplemental materials shows all proteins identified by the nine
- nophila). Table S2 in the supplemental materials lists the 40 top-scoring proteins from
- genomes (4.7%) have a double-cubane protein as detected by hidden Markov model
- NF040730. But, quite remarkably, 11.8% of all double-cubane proteins occur in the
- cies with the SAO cassette carry a double-cubane protein, and 21 (40%) have at least 1,
- ATP formation (31). SAO operons are present in 0.3% of bacterial species but in 13.4% of
- examples (6.4%). While HgcB, like SaoC, has a Cys-Cys dipeptide as the final 2 residues
- (II) reductase, in 5,469 of 34,051 RefSeq genomes (16%) of Escherichia coli.
- ily, (ii) a universal presence of TGA (100%) in the DNA encoding the putative selenocysteine site, and (iii)
Methods (brief)
- Analytical method details were not mechanically resolved from extracted text.
Implications
This page makes the source discoverable for category-level evidence routing. Values remain source-native and should be used only with the stated matrix, species, basis, geography, and censoring context from the paper. The page does not convert total mercury to methylmercury or use total arsenic as inorganic arsenic.
Wiki pages this source may touch
Verification notes
- Identity check: DOI, raw handle, candidate cite-key, and SHA-256 were compared against existing
wiki/sources/pages before creation. - Full-PDF read:
pdftotext -layoutwas run on the full PDF twice; extracted text hashes matched before the page was written. - Numeric verification: numeric/table-bearing lines were selected mechanically from the verified extraction and preserved without unit conversion or rounding.
- Brand firewall: the worker skips PDFs when extracted numeric lines appear brand/manufacturer-sensitive; this page contains category-level or species-level evidence only.
- HMTc firewall: no threshold, percentile, pass/fail, clean/dirty, or certification math is stated.
Update history
The five most recent substantive edits to this page, classified major (evidence or structure moved), correction (a published value or statement was wrong and has been fixed), or minor (narrative rewritten without changing the underlying evidence). Each description is derived from what the edit did to this page; the linked commit is the authoritative record, routine regeneration passes are excluded, and the full version history lives in git. When DOI minting comes online (see schema docs), each entry below will also link to a version-pinned DataCite DOI.