Turning scattered literature references into a single, quality-controlled source of truth for molecule-level biodegradability data
Overview
Excelra partnered with a specialty chemistry innovator to structure, curate, and aggregate biodegradability data scattered across multiple literature and reference sources. The engagement combined Excelra’s chemistry data curation expertise with public chemical-structure repositories and internationally recognized test guidelines to convert inconsistent source data into a clean, analysis-ready dataset. Under a rigorous three-level quality control process, the work was structured across two phases: cleaning and mapping the source data, then aggregating it into a single, molecule-level record for each attribute.
Our client
The client is a chemistry-driven organization developing biodegradable molecules and materials, working to strengthen the scientific data foundation behind its research and product development. As part of this effort, the client needed to consolidate biodegradability evidence drawn from multiple internal and literature references into a structured format suitable for downstream analysis and AI/ML applications.
Client’s challenge
Ahead of the engagement, the client’s biodegradability data existed in a form that limited its usefulness for analysis and modeling:
- Source data spread across multiple files and reference documents with no shared structure
- A meaningful share of samples missing molecular structure identifiers (SMILES) or CAS registry number
- Biodegradability outcomes recorded as labels needing consistent interpretation against a recognized test standard
- Multiple, sometimes conflicting, data points per molecule per attribute across sources, with no single value to work from
- No consolidated, long-format dataset that could feed directly into analysis or AI/ML pipelines
Client’s goals
The client set out to achieve the following through the engagement:
- Consolidate data from all shared source references into one clean, long-format dataset with full source traceability
- Fill gaps in molecular structure data by resolving CAS registry numbers and SMILES through trusted public chemistry databases
- Apply a consistent, defensible interpretation of biodegradability labels based on an internationally recognized test guideline
- Aggregate the cleaned data down to a single value per molecule for each attribute
- Receive a documented account of the curation process, including challenges encountered and insights gained
Our Approach
Excelra structured the engagement as a two-phase program: cleaning and mapping the source data, then consolidating it into a single, trustworthy record per molecule.
Phase 1 — Cleaning and mapping
- Ingested and cleaned data from every source file and reference shared by the client
- Resolved missing SMILES and CAS registry numbers through trusted public chemistry databases and PubChem
- Standardized biodegradability labels by referencing documented results for each assay type against the OECD 301 Ready Biodegradability test guideline
- Mapped every compound to its attributes in a long-format structure, with one row per data point and full source and reference traceability
Phase 2 — Aggregation and consolidation
- Canonicalized chemical structures across all sources so the same molecule could be reliably matched regardless of how it was named or formatted in the original reference
- Built a rules-based conflict-resolution framework to reduce multiple, sometimes disagreeing, source values down to one: OECD-compliant and more recent studies were weighted above non-standard or dated ones, and quantitative results such as percentage biodegradation were consolidated using statistical consensus rather than simple averaging
- Applied a defined hierarchy of trust to reconcile disagreements in categorical biodegradability labels between sources
- Ran every aggregated record back through independent QC to confirm one unique, defensible value per molecule per attribute before final sign-off
Figure 1: End-to-end curation and aggregation workflow
Our solution
A long-format, source-traceable curated dataset
- Minimum required fields captured for every data point: SMILES, CAS registry number, biodegradability label, source, and source identifier
- Additional fields captured where available: dose and dose unit, compliance standard and year, percentage biodegradation, study duration, 95% confidence interval bounds, and inoculum type and details
- Lower-priority, day-wise biodegradation results (e.g., day 7, 14, 21, 28) captured wherever source data allowed
Structure enrichment from trusted public sources
- Samples without SMILES were resolved using their CAS registry number via trusted public chemistry databases and PubChem
- Biodegradability labels requiring interpretation were mapped against documented OECD 301 assay results rather than assumed
Three-level quality control
Every stage of curation and aggregation moved through data curation, independent review, and a final QC analyst audit, targeting greater than 99% data accuracy on delivery.
Figure: Three-level quality control process applied throughout
A single, molecule-level source of truth
- A final aggregated table with one row per unique molecule, carrying SMILES, CAS number, and a validated biodegradability label
- Every aggregated value traceable back to the source records and resolution rule that produced it
- A final report documenting the curation and aggregation methodology, challenges encountered, and insights gained
Conclusion
Excelra’s structured, two-phase curation and aggregation process gave the client a single, defensible source of truth for biodegradability data that had previously been fragmented across disconnected references. Every structural gap was resolved against trusted public chemistry sources, every label was interpreted consistently against the OECD 301 standard, and every aggregated value can be traced back to the source and rule that produced it. The result is a clean, molecule-level analysis-ready dataset the client can build analysis and AI/ML applications on with confidence.
| >99% | 10 weeks | 2 phases | 1:1 |
| Data accuracy through 3-level QC | Phase 1 delivery timeline | Cleaning & mapping, then aggregation | Unique, validated record per molecule |
