Natural-product machine-learning models do not encounter molecules, spectra, organisms, or biological responses directly; they encounter database records produced through choices about what is deposited, normalized, linked, corrected, omitted, and retained. This article develops an original data-governance architecture for natural-product modeling that treats these choices as part of the inferential system rather than as background data-management operations. The analysis integrates structural identity, spectral evidence, bioassay context, taxonomy, provenance, missingness, duplication, record quality, correction, versioning, and negative results while preserving distinctions between chemical identity and material provenance, analytical association and structural proof, assay observations and context-independent biological properties, and observed negatives and untested or generated examples. The central contribution is a proposed governance architecture in which record lineage, evidence state, contextual metadata, correction history, and task-specific fitness remain separable but interoperable. Quality is therefore treated not as a single database-wide score but as a conditional relationship between a record state and the modeling task for which it is used. The framework further argues that model evaluation should account for coverage, redundancy, source dependence, assay context, and version structure because nominally independent test data may reproduce the same curation choices as training data. The architecture is conceptual rather than prospectively validated. Its practical value therefore depends on record-level auditability, reproducible state assignment, cross-resource reconciliation, and future benchmarking against external and temporally separated data.