########     ###    ##    ## ########  ######  
##     ##   ## ##   ###   ##    ##    ##    ## 
##     ##  ##   ##  ####  ##    ##    ##       
########  ##     ## ## ## ##    ##     ######  
##        ######### ##  ####    ##          ## 
##        ##     ## ##   ###    ##    ##    ## 
##        ##     ## ##    ##    ##     ######  

PETase ANnotation and Triage System

Nature's solution to a human-made health problem

Live instrumentation. Everything this project has read, produced and is running on, read straight from the database and the machine.

Nothing here is written by hand. Each figure is a query or a reading, and the page refreshes when the data actually changes.

updated 2026-09-20 23:22:27 UTC · database last written 62801.6 min ago

The funnel

How 14,804,920 sequences become a shortlist, and what that shortlist is judged against. Each of the first steps is a filter, and the one that matters most is the last of them: three catalytic residues can be present in a sequence and still not touch each other in the folded protein, and only a structure can tell you which.

Sequences read
14,804,920
every predicted protein in the assemblies
Survived recall
439
matched a polyesterase profile and carry a complete triad
Given a structure
416
folded with ESMFold, superposed onto IsPETase
Triad confirmed in 3D
403
the three residues actually meet in space
Too long to fold
24
over 450 residues; deferred, not discarded
Reference structures
559
the known enzymes a candidate is compared against, not part of the catch
Measured positives
342
reference enzymes whose activity someone actually assayed and published
Activity measurements
2759
individual published values behind those enzymes, each with its citation
Unlabelled sequences
25,041
homologues nobody has assayed, held apart from the catalogue on purpose so they cannot inflate a count that is quoted as evidence
Sequence clusters
170
groups of related sequences carrying their own ecology — where the source organism was found, and at what temperature
Registered sources
7
public databases and published screens this catalogue is built from, each with its retrieval date and licence

Where the sequences came from

Metagenomics reads the DNA of a whole microbial community straight from the environment, without culturing anything. Each row is one such environment, chosen because it has had decades of plastic exposure and every reason to have evolved a use for it. Yield per million is the comparable figure: raw candidate counts only tell you which environment was sequenced most deeply.

EnvironmentSequences readCandidates Per millionScan runs
Human gut 12,584,458 311 24.7 58
Compost 1,020,575 69 67.6 3
Marine plastisphere 737,027 44 59.7 2
Landfill 436,229 15 34.4 1
Wastewater 26,631 0 0.0 1

Protein structures

A structure is either measured or predicted, and the difference matters more than any single number on this page. An X-ray crystal structure is experimental evidence. AlphaFold and ESMFold produce a hypothesis — usually a good one for the fold, least reliable exactly where this project looks, in the flexible loops that gate the active site.

60
X-ray crystal structures

Deposited experimental structures. Best resolution here: 0.91 Å — lower is sharper, and below 1 Å individual atoms are resolved.

104
AlphaFold models

Predicted from an alignment of related sequences. Downloaded where a UniProt accession exists.

395
ESMFold predictions

Folded here, from sequence alone, for everything with no deposit and no accession. Mean confidence 91.0 pLDDT (0–100; above 90 the fold is reliable).

547
with a measured catalytic triad

Ser, His and Asp found close enough to pass charge between them. 1 deposits were caught as catalytic knockouts: inactivated constructs crystallographers make to trap substrate, which carry the right name and the wrong chemistry.

What the model learns from

The distinction this whole project turns on. Measured means a published experiment stands behind the label — someone ran the assay. Annotated only means a computer assigned it from sequence similarity and nobody measured anything; a model trained on those is largely rediscovering the rule that produced them, which is why the two are never pooled here.

342
measured positives

Mostly from PAZy, which lists an enzyme because activity on a plastic was measured and published, with the DOI attached.

568
annotated only

Labelled by similarity. Counted, reported, and kept out of any headline claim.

131
hard negatives

Other α/β-hydrolases: the same fold, no PET activity.

179
within-family negatives

Polyesterases that do not degrade PET — the scarcest and most valuable class here, because it is the only one that makes the hard question answerable. 150 of them were expressed and assayed and released no product; the rest rest on a database recording only what worked, which cannot separate tested-and-failed from never-tested.

Provenance of every positiveCountEvidence
EC-auto-annotated446 inferred
PAZy-measured312 measured
ESTHER-family-predicted50 inferred
Science-landscape-measured30 inferred
ACS-screen-measured28 inferred
ESTHER-family-protein-evidence14 inferred
EC-experimental7 measured
UniProt6 measured
HGMP-measured5 measured
Tournier et al. 2020, Nature1 measured
Son et al. 2019, ACS Catal.1 measured
Shi et al. 2023, Angew. Chem. Int. Ed.1 measured
PDB-construct1 measured
Orr et al. 2024, Biotechnol. J.1 measured
Oda et al. 2018, Appl. Microbiol. Biotechnol.1 measured
Lu et al. 2022, Nature (MutCompute)1 measured
Li et al. 2025, Int. J. Biol. Macromol.1 measured
J. Hazard. Mater. 20231 measured
Cui et al. 2024, Nat. Commun. 15:14171 measured
Cui et al. 2021, ACS Catal.1 measured
Austin et al. 2018, PNAS1 measured

Evidence and citations

Every measured value in this database points at the paper that reported it. 2757 of 2759 measurements carry a DOI or PubMed identifier, drawn from 38 distinct sources.

What was measured

QuantityValues
percent depolymerization 2142
Tm
the temperature at which the fold falls apart
359
Product release
how much PET breakdown product appeared
183
KM
how much substrate it takes to half-saturate the enzyme
21
Topt
the temperature at which the enzyme works fastest
20
Performance claim
a stated result from the paper — "6-fold faster than the wild type" — kept as text because it has no single unit
11
Catalytic activity
product formed under the assay conditions the paper used
10
pHopt
the acidity at which it works fastest
8
Ordinal activity
a rank rather than a number: this enzyme beat that one, with no value given
5

Strength of evidence

Evidence codeValues
Experimental, from a publication ECO:0000269 2737
Inferred by a curator ECO:0000305 22

Evidence Ontology codes. The first means somebody ran the assay; the second that a curator collated it from a review.

Models and software

Recorded per pipeline run rather than declared here, so this is what actually produced the data rather than what a README says it did.

Machine-learning models

ESM-2 t12-35M
a protein language model. Turns a sequence into 480 numbers capturing what it has learned about proteins; frozen, never fine-tuned here
embedding
ESMFold v1
predicts a 3D structure from sequence alone, no alignment needed. 8.4 GB of weights, run on CPU
folding
AlphaFold DB
precomputed structures, fetched rather than run
folding

No large language model is involved in producing any number on this site.

Tool versions in use

biotite1.7.1
gemmi0.7.5
hmmerHMMER 3.4 (Aug 2023); http://hmmer.org/
mmseqs218-8cc5c
numpy2.5.1
python3.14.3
sklearn1.9.0
torch2.13.0
transformers5.14.1

Schema v17 · code at commit dc28608

External databases

Public resources this project reads from, with the date each was last retrieved. Nothing is scraped continuously: a source is pulled when a stage runs, and the date below is when that last happened.

SourceWhat it providesRecords RetrievedLicence
ACS-screen Norton-Baker, Komp, Gado et al. 2025, ACS Catal. 15:16070-16083 114 2026-08-06 CC BY 4.0 (author-deposited Supporting Information)
HGMP-SciDB human gut metagenome polyesterases from the deposit accompanying the paper 5 2026-08-04 see publication
PAZy the plastics-active enzyme database: an enzyme is listed because activity on a plastic was measured and published 320 2026-08-06 see publication
Science-landscape Landscape profiling of PET depolymerases, Science 2025, Data S3 94 2026-08-06 see publication
Science-landscape-S1 Science 2025 Data S1, homologue search space 25041 2026-08-06 see publication
Science-landscape-S7 Science 2025 Data S7, per-cluster ecological context 170 2026-08-06 see publication
UniProt curated protein sequences, annotations and catalytic-site positions 6 2026-08-05 CC BY 4.0
RCSB PDBexperimental crystal structures and their metadata60 per structure buildCC0
AlphaFold DBpredicted structures by UniProt accession104 per structure buildCC BY 4.0
CrossrefDOI verification — every citation on this site was resolved and its title, journal and year checked 38 on curationopen
MGnifymetagenomic assemblies: the raw sequence space searched, counted here as environments rather than records5 per scanopen

Pipeline activity

Every stage writes a record of what went in, what came out, which tool versions ran and how long it took. 124 runs and 121 manifests so far, totalling 32.2 hours of recorded compute.

StageWhat it doesRunsIn OutLast run
recall search metagenome proteins for the polyesterase fold, then keep only those whose catalytic residues meet in space 6514,804,920 520 2026-08-05 05:10
reference_structures fetch or fold a structure for each characterised enzyme and superpose it onto IsPETase 33672 672 2026-08-07 22:22
seeds load the curated wild types and derive each engineered variant from its parent plus a checked mutation list 10168 152 2026-08-05 23:41
pazy import enzymes with published, measured activity on a plastic 41,936 679 2026-08-06 00:11
embed turn every sequence into 480 numbers with a frozen protein language model 44,002 4,002 2026-08-06 22:15
structure fold candidates with ESMFold and measure the active site 2128 88 2026-08-04 23:30
science_supplements 125,411 25,053 2026-08-06 22:30
science_landscape 1183 183 2026-08-06 22:08
positives_family add the wider enzyme family as similarity-labelled positives 183 79 2026-08-04 12:39
negatives collect hard negatives: the same fold, no PET activity 13,985 256 2026-08-04 12:41
harmonise reconcile activity values reported under different assays 1459 47 2026-08-04 12:39
acs_screen 1477 214 2026-08-06 22:04

Model evaluations

Every training run ever recorded, including the ones that failed. AUC is the chance the model ranks a true PET enzyme above a non-enzyme: 1.0 is perfect, 0.5 is a coin flip. The composition baseline beside it is what you get from amino-acid counts alone — a score only means something if it clears that.

Trained onPositivesNegatives ClustersAUCBaselineReading
measured-only
pet_activity · 2026-08-06
308385 0.774 0.565 clears the composition baseline by a real margin
measured-only
pet_activity · 2026-08-06
308385 59 0.845 0.565 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
852385 97 0.835 0.565 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
308385 0.739 0.565 clears the composition baseline by a real margin
measured-only
pet_activity · 2026-08-06
308385 48 0.790 0.565 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
852385 65 0.859 0.565 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
302241 0.845 0.803 barely clears amino-acid composition
measured-only
pet_activity · 2026-08-06
302241 60 0.947 0.803 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
789241 97 0.972 0.803 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
302241 0.857 0.803 barely clears amino-acid composition
measured-only
pet_activity · 2026-08-06
302241 46 0.949 0.803 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
789241 61 0.981 0.803 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
304241 0.860 0.736 clears the composition baseline by a real margin
measured-only
pet_activity · 2026-08-06
304241 61 0.942 0.736 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
791241 97 0.975 0.736 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
304241 0.782 0.736 barely clears amino-acid composition
measured-only
pet_activity · 2026-08-06
304241 46 0.921 0.736 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
791241 60 0.967 0.736 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
304241 0.860 0.736 clears the composition baseline by a real margin
measured-only
pet_activity · 2026-08-06
304241 61 0.942 0.736 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
791241 97 0.975 0.736 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
304241 0.782 0.736 barely clears amino-acid composition
measured-only
pet_activity · 2026-08-06
304241 46 0.921 0.736 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
791241 60 0.967 0.736 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
304241 0.860 0.736 clears the composition baseline by a real margin
measured-only
pet_activity · 2026-08-06
304241 61 0.942 0.736 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
791241 97 0.975 0.736 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
304241 0.782 0.736 barely clears amino-acid composition
measured-only
pet_activity · 2026-08-06
304241 46 0.921 0.736 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
791241 60 0.967 0.736 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
304241 0.860 0.736 clears the composition baseline by a real margin
measured-only
pet_activity · 2026-08-06
304241 61 0.942 0.736 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
791241 97 0.975 0.736 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-06
304241 0.782 0.736 barely clears amino-acid composition
measured-only
pet_activity · 2026-08-06
304241 46 0.921 0.736 clears the composition baseline by a real margin
mixed
pet_activity · 2026-08-06
791241 60 0.967 0.736 trained on similarity-derived labels
measured-only
pet_activity · 2026-08-05
15226 7 0.493 0.398 at chance — the hardest contrast, and the honest answer
measured-only
pet_activity · 2026-08-05
30526 44 0.850 0.651 clears the composition baseline by a real margin
measured-only
pet_activity · 2026-08-05
300220 45 0.976 0.829 clears the composition baseline by a real margin
experimental
pet_activity · 2026-08-04
13220 1 not evaluable — too few independent clusters to split on
mixed
pet_activity · 2026-08-04
500220 25 1.000 0.778 perfect, and meaningless: it reproduced the similarity rule that made the labels

The machine this runs on

A single DigitalOcean droplet shared with five other applications, which is why the web layer carries no scientific libraries at all: every structure and every measurement was computed offline and is served here as a precomputed file.

79.6%
disk used

16.7 GB free of 82.1 GB

70.9%
memory used

1.2 GB available of 4.11 GB. Memory is the binding constraint here, not disk or speed.

0.01
load average

across 2 cores. Above the core count means work is queueing.

20.6 MB
database

SQLite, one file, everything in it. 832 candidate and 559 reference structure files alongside it, totalling 139.5 + 117.6 MB.

Linux x86_64 · Python 3.12.3 · web process up 6.67 h · host up 1635.9 h · data version 0.3.0