Methods and limitations
Spec section 7 calls this page non-negotiable "given how easily this kind of tool is over-read". This project has more to disclose than most.
Trained on the measured set alone, the activity head reaches AUC 0.790 (± 0.010) across 126 independent sequence clusters at 30% identity, against 0.565 for a classifier using amino-acid composition alone. It clears that baseline by a real margin.
It does not clear the one that matters. Retrieval — scoring a sequence by its similarity to the nearest enzyme already known to degrade PET, outside its own cluster — reaches 0.787. The learned head trained on measured labels comes in 0.003 against it. On labels somebody actually measured, the model does not beat looking the answer up. Trained on the annotated set instead it reaches 0.859 and clears retrieval by +0.072 — on labels a computer assigned by similarity, which is the rule retrieval already encodes.
These figures are read from the evaluation run's own output rather than written here, because this paragraph previously quoted an AUC of 0.976 against a composition baseline of 0.829 for some time after the protocol had been re-run and returned different numbers. Prose does not get recomputed.
Measured negatives, and why they changed everything
The hardest question in this project is not "is this a polyesterase" — that separates at AUC 0.975 — but "does this polyesterase degrade PET". Answering it needs enzymes that were tested and found not to, and for most of this project's life there were 29 of those, all on the weakest possible basis: PAZy records only positive substrate associations, so an enzyme missing from its PET list might have been assayed and failed, or might never have been assayed at all. Those two are not the same claim and the database cannot tell them apart.
Two 2025 papers screened panels under a single protocol and, crucially, reported what did not work. Both deposited their source data openly. Ingesting them takes the measured negative class from nothing to 150.
| Screen | Assayed | Active | Measured inactive | How the line was drawn |
|---|---|---|---|---|
| ACS Catalysis 2025 Machine Learning-Guided Identification of PET Hydrolases from Natural Diversity |
216 | 126 | 88 | Percent depolymerisation at a stated 0.1% floor, not a per-milligram rate: dividing by a tiny enzyme loading turns noise into a large normalised number. |
| Science 2025 Landscape profiling of PET depolymerases using a natural sequence cluster framework |
183 | 102 | 69 | Against each entry's own replicate standard deviation — active above 2 SD, inactive at or below 1 SD. The 12 in between are labelled neither way and excluded from training, because they are exactly the enzymes a threshold would be tuned on. |
Both screens also measured melting temperatures, and the ACS screen assayed at 40 °C as well as 60 °C. Almost every optimum in this catalogue sits between 50 and 78 °C; the therapeutic target is 37 °C, so a screen that reports what happens near body temperature is worth more here than its enzyme count suggests.
And the same question asked of geometry
Active-site geometry was the one signal here measured off coordinates rather than inherited from somebody's annotation, so it was the obvious thing to re-test once real labels existed. On the inferred negatives it looked convincing: cleft depth separated PET-active from PET-inactive at AUC 0.808, p 1.7×10−7, falling to 0.534 under cluster-grouped splitting — which was read as an underpowered evaluation waiting for more data.
Asked of 237 PET-active against 139 measured-inactive enzymes, all ESMFold on both sides so no coordinate-source confound is in play, the answer is AUC 0.398 (0.360–0.453 across 10 random splits).
Below a coin flip, and consistently so — every one of the 10 splits lands under 0.5. That is worth reading carefully, because it does not mean the classifier is useful inverted. It means whatever relationship the features carry inside the training clusters reverses in held-out ones, which is the signature of fitting lineage rather than activity: the geometry says which family an enzyme belongs to, and family does not say whether it degrades PET.
The trajectory is the strongest part of the evidence. As the measured negative set filled in, the result went 0.507 at 95 negatives, 0.498 at 125, 0.459 at 139, and 0.398 with the positive set complete. A real effect emerging from a thin sample strengthens and narrows around a non-null value; moving toward and past chance as the sample grows is a confound being diluted. That distinction is exactly the one that let the earlier reading survive as "underpowered" for as long as it did.
The raw features are where it is clearest. Not one of the nine survives at p < 0.01, and cleft depth — the 0.808 above — comes in at 0.547, p 0.12. That signal was never measuring PET activity. It was measuring whatever separates the enzymes a database happens to list from the ones it does not, and it disappeared the moment the labels came from somebody running the assay.
Two limits worth stating with it. Only 8 clusters contain both classes, so the grouped estimate rests on few independent lineages — though its stability across seeds argues the answer is real. And every structure here is a prediction: if ESMFold idealises the active site, which the paired crystal-versus-model comparison above shows it does, a genuine difference could be flattened by the modelling rather than absent from the enzymes.
What the rest of that supplement held
The Science paper ships seven data files. Only two of the remaining three were worth ingesting, and saying which is part of the point.
| File | Records | Disposition |
|---|---|---|
| S2 — library sequences | 2,064 | Skipped. Every accession and every sequence is byte-identical to the column already ingested from S3. Checked rather than assumed — all 2,064 matched — because ingesting it would have created nothing and risked duplicating rows whose counts this project quotes as evidence. |
| S1 — homologue search space | 25,041 | Sequences with no activity measured on any of them. Held in a separate table, deliberately outside the catalogue: nothing here has been characterised, and 25,000 unlabelled homologues inside the table whose totals are quoted would corrupt every count on the site. Useful as an unlabelled pool; it cannot add a measured negative, which is the thing that binds. |
| S7 — ecology per cluster | 170 | Isolation source, biome, habitat, geography and the temperature the source organism was isolated at — not the enzyme's optimum, and stored under a name that says so. Kept at cluster level because that is the level the paper reports it at; attaching a temperature to one enzyme would invent a precision the source does not have. 354 catalogued enzymes are linked to their cluster. |
One result falls out of S7 directly. Under the paper's own sequence-cluster framework, 14 clusters contain both a PET-active and a measured-inactive enzyme, against 8 under the 30% identity clustering used for the splits here. The shortage of lineages measured on both sides is real under either definition, and slightly less severe under theirs.
What 123 more negatives actually bought
Not what was expected, and the disappointment is the finding. The within-family contrast — measured PET-active against measured PET-inactive, both polyesterases — now runs on 149 negatives instead of 26. The number that governs whether it can be evaluated barely moved: clusters holding both an active and an inactive enzyme went from 7 to 8. The constraint was never how many negatives exist. It is how many independent lineages anyone has measured on both sides, and five times the data bought one more.
| Contrast | Pos | Neg | Shared clusters | Head AUC | Composition |
|---|---|---|---|---|---|
| Out of family — other α/β-hydrolase folds | 308 | 108 | 2 | 0.989 | 0.919 |
| Near miss — esterases on soluble substrates | 308 | 251 | 9 | 0.769 | 0.693 |
| Within family, inferred negatives — PAZy "not reported active" | 308 | 26 | 7 | 0.808 | 0.547 |
| Within family, MEASURED negatives — expressed, assayed, no product | 308 | 149 | 8 | 0.648 | 0.644 |
| Within family, MEASURED, mixed clusters only — the hardest honest test | 250 | 149 | 8 | 0.625 | 0.519 |
Two readings, and both matter. On the full measured contrast the head scores 0.648 and amino-acid composition scores 0.644 — no advantage at all over counting residues. Restricted to the clusters that contain both classes, the head reaches 0.625 against composition's 0.519, which is a real margin.
The clearest gain is in the error bars. That hardest test previously read 0.435 ± 0.301, an interval so wide it excluded nothing; it now reads 0.625 ± 0.081. The model did not get better. The experiment got powered enough to give a stable answer, and the stable answer is a modest signal that is not composition. That is a smaller claim than "we solved it" and the first one here that survives its own error bar.
A protein is only labelled from a condition it was actually assayed under. An
enzyme absent from a condition is not a zero in it, so comparable_group_id keys on
substrate, temperature and pH together and a crystalline-powder result at 60 °C is never pooled
with an amorphous-film result at 40 °C.
Why none of it transfers: the determinants are lineage-specific
Three independent methods had failed the same way — sequence embeddings, a learned head, and active-site geometry all worked out-of-family and collapsed within it. The usual reading of that is too little data. This test asks a different question, and it needs no train/test split at all, so it cannot be leakage in disguise.
Fit the same model separately inside each of the three lineages large enough to support it, then compare the directions the fits point in. A single global rule for what makes a PETase can only exist if those directions broadly agree.
The comparison needs its own reference, and that is the part that makes it a test rather than a number. A cosine similarity of 0.3 between two coefficient vectors means nothing on its own, because with a few dozen enzymes the vectors are noisy and even two samples of an identical process would not agree perfectly. So the reference is not zero: it is how well two bootstrap resamples of the same lineage agree with each other — an empirical ceiling on what agreement looks like when the true direction is the same by construction.
| Features | Same lineage the ceiling |
Different lineages | Pairs pointing opposite ways |
|---|---|---|---|
| Active-site geometry 9 features | +0.727 | −0.318 to +0.336 | 2 of 3 |
| ESM-2 embeddings 20 components | +0.696 | −0.086 to +0.225 | 2 of 3 |
Two fits of the same lineage agree at about +0.70. Two different lineages agree at roughly zero, and in two of three pairs they point in opposite directions — what predicts activity in one family predicts inactivity in another. The same pattern appears in two feature sets that share nothing: one derived from 3D coordinates, one from a language model over sequence.
That the ceiling is high is what makes this a result rather than a shrug. The within-lineage directions are well determined; it is specifically the comparison between lineages that fails, so this is not the test being too noisy to answer.
It also explains the below-chance results. A model trained across lineages learns a direction that reverses on a lineage it has not seen, which is exactly how geometry reached 0.398 and why an eighteen-fold larger language model changed nothing. The problem was never capacity, and it was never only sample size: there may be no single global rule to learn. Established on three lineages, so it is a strong indication rather than a settled law — but it is the first explanation offered here that accounts for every failure at once.
Where the signal holds, and what to ship instead
With a global rule ruled out, one question remains and it is the one a user would actually ask: given a candidate resembling enzymes we have characterised, how close does it have to be before the prediction means anything? The 262-member lineage — 171 measured-active, 91 measured-inactive — is the only place with enough data to answer. Re-cluster inside it at successively stricter identities, hold out whole sub-clusters, and read where the curve meets chance.
| Identity | Sub-clusters scored | Head | Retrieval | Composition | Head − retrieval, paired |
|---|---|---|---|---|---|
| 95% | 4 | 0.500 | 0.500 | 0.750 | too few to score |
| 80% | 15 | 0.561 | 0.550 | 0.615 | +0.011 [−0.288, +0.285] |
| 70% | 17 | 0.703 | 0.544 | 0.625 | +0.159 [−0.263, +0.390] |
| 60% | 12 | 0.684 | 0.624 | 0.612 | +0.060 [−0.223, +0.436] |
| 50% | 5 | 0.550 | 0.518 | 0.577 | too few to score |
Read the point estimates alone and there is an operating band: 0.703 against retrieval's 0.544 at 70% identity. Read the intervals and it is not there. The paired lead — head minus retrieval on the same held-out sub-clusters — is +0.159 [−0.263, +0.390], spanning zero. Seventeen sub-clusters cannot resolve a difference that size, and the curve is not even monotonic: amino-acid composition beats both at 80%.
One thing does survive. At 70% identity the head's own interval, [0.521, 0.798], excludes 0.5 — so within a lineage, among relatives sharing about 70% of their sequence, there is weak evidence of better-than-chance discrimination. It just cannot be shown to beat looking up the nearest known PETase.
So the honest deliverable is retrieval, not a classifier. Nearest-neighbour similarity to a measured-active enzyme is simple, needs no training, carries no risk of learning a lineage direction that reverses, and nothing built here has been shown to beat it on labels somebody measured. That is a smaller claim than this project set out to make and it is the one the evidence supports.
Positives by evidence tier
| Tier | n |
|---|---|
| EC-auto-annotated | 446 |
| PAZy-measured | 312 |
| ESTHER-family-predicted | 50 |
| Science-landscape-measured | 30 |
| ACS-screen-measured | 28 |
| ESTHER-family-protein-evidence | 14 |
| EC-experimental | 7 |
| UniProt | 6 |
| HGMP-measured | 5 |
| Tournier et al. 2020, Nature | 1 |
| Son et al. 2019, ACS Catal. | 1 |
| Shi et al. 2023, Angew. Chem. Int. Ed. | 1 |
| PDB-construct | 1 |
| Orr et al. 2024, Biotechnol. J. | 1 |
| Oda et al. 2018, Appl. Microbiol. Biotechnol. | 1 |
| Lu et al. 2022, Nature (MutCompute) | 1 |
| Li et al. 2025, Int. J. Biol. Macromol. | 1 |
| J. Hazard. Mater. 2023 | 1 |
| Cui et al. 2024, Nat. Commun. 15:1417 | 1 |
| Cui et al. 2021, ACS Catal. | 1 |
| Austin et al. 2018, PNAS | 1 |
Training runs
| Head | Pos | Neg | Clusters | AUC | Composition baseline | Evidence |
|---|---|---|---|---|---|---|
| pet_activity | 308 | 385 | None | 0.774 | 0.565 | measured-only |
| pet_activity | 308 | 385 | 59 | 0.845 | 0.565 | measured-only |
| pet_activity | 852 | 385 | 97 | 0.835 | 0.565 | mixed |
| pet_activity | 308 | 385 | None | 0.739 | 0.565 | measured-only |
| pet_activity | 308 | 385 | 48 | 0.790 | 0.565 | measured-only |
| pet_activity | 852 | 385 | 65 | 0.859 | 0.565 | mixed |
| pet_activity | 302 | 241 | None | 0.845 | 0.803 | measured-only |
| pet_activity | 302 | 241 | 60 | 0.947 | 0.803 | measured-only |
| pet_activity | 789 | 241 | 97 | 0.972 | 0.803 | mixed |
| pet_activity | 302 | 241 | None | 0.857 | 0.803 | measured-only |
| pet_activity | 302 | 241 | 46 | 0.949 | 0.803 | measured-only |
| pet_activity | 789 | 241 | 61 | 0.981 | 0.803 | mixed |
| pet_activity | 304 | 241 | None | 0.860 | 0.736 | measured-only |
| pet_activity | 304 | 241 | 61 | 0.942 | 0.736 | measured-only |
| pet_activity | 791 | 241 | 97 | 0.975 | 0.736 | mixed |
| pet_activity | 304 | 241 | None | 0.782 | 0.736 | measured-only |
| pet_activity | 304 | 241 | 46 | 0.921 | 0.736 | measured-only |
| pet_activity | 791 | 241 | 60 | 0.967 | 0.736 | mixed |
| pet_activity | 304 | 241 | None | 0.860 | 0.736 | measured-only |
| pet_activity | 304 | 241 | 61 | 0.942 | 0.736 | measured-only |
| pet_activity | 791 | 241 | 97 | 0.975 | 0.736 | mixed |
| pet_activity | 304 | 241 | None | 0.782 | 0.736 | measured-only |
| pet_activity | 304 | 241 | 46 | 0.921 | 0.736 | measured-only |
| pet_activity | 791 | 241 | 60 | 0.967 | 0.736 | mixed |
| pet_activity | 304 | 241 | None | 0.860 | 0.736 | measured-only |
| pet_activity | 304 | 241 | 61 | 0.942 | 0.736 | measured-only |
| pet_activity | 791 | 241 | 97 | 0.975 | 0.736 | mixed |
| pet_activity | 304 | 241 | None | 0.782 | 0.736 | measured-only |
| pet_activity | 304 | 241 | 46 | 0.921 | 0.736 | measured-only |
| pet_activity | 791 | 241 | 60 | 0.967 | 0.736 | mixed |
| pet_activity | 304 | 241 | None | 0.860 | 0.736 | measured-only |
| pet_activity | 304 | 241 | 61 | 0.942 | 0.736 | measured-only |
| pet_activity | 791 | 241 | 97 | 0.975 | 0.736 | mixed |
| pet_activity | 304 | 241 | None | 0.782 | 0.736 | measured-only |
| pet_activity | 304 | 241 | 46 | 0.921 | 0.736 | measured-only |
| pet_activity | 791 | 241 | 60 | 0.967 | 0.736 | mixed |
| pet_activity | 152 | 26 | 7 | 0.493 | 0.398 | measured-only |
| pet_activity | 305 | 26 | 44 | 0.850 | 0.651 | measured-only |
| pet_activity | 300 | 220 | 45 | 0.976 | 0.829 | measured-only |
| pet_activity | 13 | 220 | 1 | not evaluable | - | experimental |
| pet_activity | 500 | 220 | 25 | 1.000 | 0.778 | mixed |
The composition baseline uses amino-acid fractions and length only. Any model claim must clear it as well as the retrieval baseline.
Pipeline runs
| Stage | Label | Status | In | Out | Discarded | Started |
|---|---|---|---|---|---|---|
| reference_structures | pazy-bulk | done | 5 | 5 | 0 | 2026-08-07T22:12:35Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-07T19:57:54Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-07T15:55:12Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-07T13:11:35Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-07T10:42:05Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-07T07:37:58Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-07T04:30:10Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-07T01:40:17Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-06T22:57:12Z |
| science_supplements | s1-s7 | done | 25,411 | 25053 | 358 | 2026-08-06T22:30:15Z |
| embed | esm2-t12-35M | done | 1,677 | 1677 | 0 | 2026-08-06T22:14:32Z |
| science_landscape | v1 | done | 183 | 183 | 12 | 2026-08-06T22:08:07Z |
| acs_screen | v1 | done | 477 | 214 | 263 | 2026-08-06T22:04:06Z |
| reference_structures | deposit-relink | done | 48 | 48 | 0 | 2026-08-06T12:12:17Z |
| reference_structures | pazy-bulk | done | 13 | 13 | 0 | 2026-08-06T11:44:36Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-06T10:49:50Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-06T09:31:15Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-06T08:21:34Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-06T06:48:42Z |
| reference_structures | pazy-bulk | done | 25 | 25 | 0 | 2026-08-06T04:35:13Z |
Data sources
| Source | Version | Records | Licence | Retrieved |
|---|---|---|---|---|
| ACS-screen | Norton-Baker, Komp, Gado et al. 2025, ACS Catal. 15:16070-16083 | 114 | CC BY 4.0 (author-deposited Supporting Information) | 2026-08-06T22:10:47Z |
| HGMP-SciDB | PMID 39551294 | 5 | see publication | 2026-08-04T22:32:55Z |
| PAZy | Buchholz et al. 2022, Proteins 90(7):1443-1456 | 320 | see publication | 2026-08-06T00:11:50Z |
| Science-landscape | Landscape profiling of PET depolymerases, Science 2025, Data S3 | 94 | see publication | 2026-08-06T22:10:47Z |
| Science-landscape-S1 | Science 2025 Data S1, homologue search space | 25041 | see publication | 2026-08-06T22:31:18Z |
| Science-landscape-S7 | Science 2025 Data S7, per-cluster ecological context | 170 | see publication | 2026-08-06T22:31:18Z |
| UniProt | REST | 6 | CC BY 4.0 | 2026-08-05T23:41:18Z |
Limitations
- Positives number in the low hundreds and most carry family annotation rather than measured PET activity.
- Published activity data is not harmonised across assay formats. Absolute rate predictions should not be trusted.
- Crystalline PET degradation at 37 °C by any known enzyme is slow. PANTS ranks relative promise, not therapeutic viability.
- Predicted structures are predictions. Cleft geometry on a metagenomic sequence with no close homologue carries real uncertainty, and a brittle version of that measurement scored IsPETase's own prediction at half its crystal value before being fixed.
- Predicted and experimental coordinates are not interchangeable. Measured on the same protein both ways (51 enzymes with a crystal structure and an AlphaFold model), the oxyanion hole differs systematically: the second donor's angle is 23.6° in the crystal against 15.6° in the model, p 1.1×10−6. Predictions build a tighter, more idealised active site than the protein has. Cleft depth is the one geometric feature that survives the comparison, and it is also the strongest activity feature. Pooling sources without accounting for this dropped a geometry model from 0.749 to 0.553, which is why every structure here records where its coordinates came from.
- Leave-one-family-out cannot be run at all. All 13 ESTHER families in this catalogue are wholly positive or wholly negative, so holding one out removes an entire class. That is not a gap in the protocol; it is the clearest statement of the problem, which is that in this dataset the label is family membership.
- That AUC is measured against hard negatives from other α/β-hydrolase families. Tested against the near misses instead it reaches 1.000, which is a warning rather than a result: every near miss is a single ESTHER family (Cutinase), and one homogeneous family separates as a block (composition alone scores 0.910). Neither contrast shows that the head ranks PET activity within the polyesterase family.
- The within-family question has no dataset. It needs PET-inactive polyesterases, measured and published, and databases record what works: nobody systematically publishes a failed assay. That is publication bias, not a curation gap, and more database mining will not fix it.
- Nothing here addresses delivery, immunogenicity, biodistribution, or what happens to liberated TPA and EG in vivo.
- Metagenomic candidates may come from unculturable organisms, may not express in a standard host, and may be fragments or misassemblies.
- Cleft width currently separates a fungal cutinase from the polyesterase family, but does not order that family by PET activity. Family-level separation is a much easier task than the one that matters.
Recent manifests
| Stage | Model | Schema | Commit | Wall (s) | Written |
|---|---|---|---|---|---|
| reference_structures | - | 16 | dc28608 | 565.9 | 2026-08-07T22:22:09Z |
| reference_structures | - | 16 | dc28608 | 4766.9 | 2026-08-07T22:12:35Z |
| reference_structures | - | 16 | dc28608 | 9078.5 | 2026-08-07T19:57:54Z |
| reference_structures | - | 16 | 9379723 | 4818.1 | 2026-08-07T15:55:12Z |
| reference_structures | - | 16 | e5e10c2 | 5799.3 | 2026-08-07T13:11:35Z |
| reference_structures | - | 16 | e5e10c2 | 8503.9 | 2026-08-07T10:42:05Z |
| reference_structures | - | 16 | d5b68f4 | 8208.9 | 2026-08-07T07:37:58Z |
| reference_structures | - | 16 | d5b68f4 | 7998.4 | 2026-08-07T04:30:10Z |
| reference_structures | - | 16 | d5b68f4 | 7244.8 | 2026-08-07T01:40:17Z |
| science_supplements | - | 16 | cc463cd | 4.4 | 2026-08-06T22:30:19Z |
| embed | facebook/esm2_t12_35M_UR50D | 15 | 26a50f4 | 80.4 | 2026-08-06T22:15:52Z |
| science_landscape | - | 15 | 26a50f4 | 3.0 | 2026-08-06T22:08:10Z |