Live instrumentation. Everything this project has read, produced and is running on, read straight from the database and the machine.
Nothing here is written by hand. Each figure is a query or a reading, and the page refreshes when the data actually changes.
updated 2026-09-20 23:22:27 UTC · database last written 62801.6 min ago
The funnel
How 14,804,920 sequences become a shortlist, and what that shortlist is judged against. Each of the first steps is a filter, and the one that matters most is the last of them: three catalytic residues can be present in a sequence and still not touch each other in the folded protein, and only a structure can tell you which.
- Sequences read
- 14,804,920
- every predicted protein in the assemblies
- Survived recall
- 439
- matched a polyesterase profile and carry a complete triad
- Given a structure
- 416
- folded with ESMFold, superposed onto IsPETase
- Triad confirmed in 3D
- 403
- the three residues actually meet in space
- Too long to fold
- 24
- over 450 residues; deferred, not discarded
- Reference structures
- 559
- the known enzymes a candidate is compared against, not part of the catch
- Measured positives
- 342
- reference enzymes whose activity someone actually assayed and published
- Activity measurements
- 2759
- individual published values behind those enzymes, each with its citation
- Unlabelled sequences
- 25,041
- homologues nobody has assayed, held apart from the catalogue on purpose so they cannot inflate a count that is quoted as evidence
- Sequence clusters
- 170
- groups of related sequences carrying their own ecology — where the source organism was found, and at what temperature
- Registered sources
- 7
- public databases and published screens this catalogue is built from, each with its retrieval date and licence
Where the sequences came from
Metagenomics reads the DNA of a whole microbial community straight from the environment, without culturing anything. Each row is one such environment, chosen because it has had decades of plastic exposure and every reason to have evolved a use for it. Yield per million is the comparable figure: raw candidate counts only tell you which environment was sequenced most deeply.
| Environment | Sequences read | Candidates | Per million | Scan runs |
|---|---|---|---|---|
| Human gut | 12,584,458 | 311 | 24.7 | 58 |
| Compost | 1,020,575 | 69 | 67.6 | 3 |
| Marine plastisphere | 737,027 | 44 | 59.7 | 2 |
| Landfill | 436,229 | 15 | 34.4 | 1 |
| Wastewater | 26,631 | 0 | 0.0 | 1 |
Protein structures
A structure is either measured or predicted, and the difference matters more than any single number on this page. An X-ray crystal structure is experimental evidence. AlphaFold and ESMFold produce a hypothesis — usually a good one for the fold, least reliable exactly where this project looks, in the flexible loops that gate the active site.
Deposited experimental structures. Best resolution here: 0.91 Å — lower is sharper, and below 1 Å individual atoms are resolved.
Predicted from an alignment of related sequences. Downloaded where a UniProt accession exists.
Folded here, from sequence alone, for everything with no deposit and no accession. Mean confidence 91.0 pLDDT (0–100; above 90 the fold is reliable).
Ser, His and Asp found close enough to pass charge between them. 1 deposits were caught as catalytic knockouts: inactivated constructs crystallographers make to trap substrate, which carry the right name and the wrong chemistry.
What the model learns from
The distinction this whole project turns on. Measured means a published experiment stands behind the label — someone ran the assay. Annotated only means a computer assigned it from sequence similarity and nobody measured anything; a model trained on those is largely rediscovering the rule that produced them, which is why the two are never pooled here.
Mostly from PAZy, which lists an enzyme because activity on a plastic was measured and published, with the DOI attached.
Labelled by similarity. Counted, reported, and kept out of any headline claim.
Other α/β-hydrolases: the same fold, no PET activity.
Polyesterases that do not degrade PET — the scarcest and most valuable class here, because it is the only one that makes the hard question answerable. 150 of them were expressed and assayed and released no product; the rest rest on a database recording only what worked, which cannot separate tested-and-failed from never-tested.
| Provenance of every positive | Count | Evidence |
|---|---|---|
| EC-auto-annotated | 446 | inferred |
| PAZy-measured | 312 | measured |
| ESTHER-family-predicted | 50 | inferred |
| Science-landscape-measured | 30 | inferred |
| ACS-screen-measured | 28 | inferred |
| ESTHER-family-protein-evidence | 14 | inferred |
| EC-experimental | 7 | measured |
| UniProt | 6 | measured |
| HGMP-measured | 5 | measured |
| Tournier et al. 2020, Nature | 1 | measured |
| Son et al. 2019, ACS Catal. | 1 | measured |
| Shi et al. 2023, Angew. Chem. Int. Ed. | 1 | measured |
| PDB-construct | 1 | measured |
| Orr et al. 2024, Biotechnol. J. | 1 | measured |
| Oda et al. 2018, Appl. Microbiol. Biotechnol. | 1 | measured |
| Lu et al. 2022, Nature (MutCompute) | 1 | measured |
| Li et al. 2025, Int. J. Biol. Macromol. | 1 | measured |
| J. Hazard. Mater. 2023 | 1 | measured |
| Cui et al. 2024, Nat. Commun. 15:1417 | 1 | measured |
| Cui et al. 2021, ACS Catal. | 1 | measured |
| Austin et al. 2018, PNAS | 1 | measured |
Evidence and citations
Every measured value in this database points at the paper that reported it. 2757 of 2759 measurements carry a DOI or PubMed identifier, drawn from 38 distinct sources.
What was measured
| Quantity | Values |
|---|---|
| percent depolymerization | 2142 |
| Tm the temperature at which the fold falls apart |
359 |
| Product release how much PET breakdown product appeared |
183 |
| KM how much substrate it takes to half-saturate the enzyme |
21 |
| Topt the temperature at which the enzyme works fastest |
20 |
| Performance claim a stated result from the paper — "6-fold faster than the wild type" — kept as text because it has no single unit |
11 |
| Catalytic activity product formed under the assay conditions the paper used |
10 |
| pHopt the acidity at which it works fastest |
8 |
| Ordinal activity a rank rather than a number: this enzyme beat that one, with no value given |
5 |
Strength of evidence
| Evidence code | Values |
|---|---|
| Experimental, from a publication ECO:0000269 | 2737 |
| Inferred by a curator ECO:0000305 | 22 |
Evidence Ontology codes. The first means somebody ran the assay; the second that a curator collated it from a review.
Models and software
Recorded per pipeline run rather than declared here, so this is what actually produced the data rather than what a README says it did.
Machine-learning models
| ESM-2 t12-35M a protein language model. Turns a sequence into 480 numbers capturing what it has learned about proteins; frozen, never fine-tuned here | embedding |
| ESMFold v1 predicts a 3D structure from sequence alone, no alignment needed. 8.4 GB of weights, run on CPU |
folding |
| AlphaFold DB precomputed structures, fetched rather than run | folding |
No large language model is involved in producing any number on this site.
Tool versions in use
| biotite | 1.7.1 |
| gemmi | 0.7.5 |
| hmmer | HMMER 3.4 (Aug 2023); http://hmmer.org/ |
| mmseqs2 | 18-8cc5c |
| numpy | 2.5.1 |
| python | 3.14.3 |
| sklearn | 1.9.0 |
| torch | 2.13.0 |
| transformers | 5.14.1 |
Schema v17 · code at commit dc28608
External databases
Public resources this project reads from, with the date each was last retrieved. Nothing is scraped continuously: a source is pulled when a stage runs, and the date below is when that last happened.
| Source | What it provides | Records | Retrieved | Licence |
|---|---|---|---|---|
| ACS-screen | Norton-Baker, Komp, Gado et al. 2025, ACS Catal. 15:16070-16083 | 114 | 2026-08-06 | CC BY 4.0 (author-deposited Supporting Information) |
| HGMP-SciDB | human gut metagenome polyesterases from the deposit accompanying the paper | 5 | 2026-08-04 | see publication |
| PAZy | the plastics-active enzyme database: an enzyme is listed because activity on a plastic was measured and published | 320 | 2026-08-06 | see publication |
| Science-landscape | Landscape profiling of PET depolymerases, Science 2025, Data S3 | 94 | 2026-08-06 | see publication |
| Science-landscape-S1 | Science 2025 Data S1, homologue search space | 25041 | 2026-08-06 | see publication |
| Science-landscape-S7 | Science 2025 Data S7, per-cluster ecological context | 170 | 2026-08-06 | see publication |
| UniProt | curated protein sequences, annotations and catalytic-site positions | 6 | 2026-08-05 | CC BY 4.0 |
| RCSB PDB | experimental crystal structures and their metadata | 60 | per structure build | CC0 |
| AlphaFold DB | predicted structures by UniProt accession | 104 | per structure build | CC BY 4.0 |
| Crossref | DOI verification — every citation on this site was resolved and its title, journal and year checked | 38 | on curation | open |
| MGnify | metagenomic assemblies: the raw sequence space searched, counted here as environments rather than records | 5 | per scan | open |
Pipeline activity
Every stage writes a record of what went in, what came out, which tool versions ran and how long it took. 124 runs and 121 manifests so far, totalling 32.2 hours of recorded compute.
| Stage | What it does | Runs | In | Out | Last run |
|---|---|---|---|---|---|
| recall | search metagenome proteins for the polyesterase fold, then keep only those whose catalytic residues meet in space | 65 | 14,804,920 | 520 | 2026-08-05 05:10 |
| reference_structures | fetch or fold a structure for each characterised enzyme and superpose it onto IsPETase | 33 | 672 | 672 | 2026-08-07 22:22 |
| seeds | load the curated wild types and derive each engineered variant from its parent plus a checked mutation list | 10 | 168 | 152 | 2026-08-05 23:41 |
| pazy | import enzymes with published, measured activity on a plastic | 4 | 1,936 | 679 | 2026-08-06 00:11 |
| embed | turn every sequence into 480 numbers with a frozen protein language model | 4 | 4,002 | 4,002 | 2026-08-06 22:15 |
| structure | fold candidates with ESMFold and measure the active site | 2 | 128 | 88 | 2026-08-04 23:30 |
| science_supplements | — | 1 | 25,411 | 25,053 | 2026-08-06 22:30 |
| science_landscape | — | 1 | 183 | 183 | 2026-08-06 22:08 |
| positives_family | add the wider enzyme family as similarity-labelled positives | 1 | 83 | 79 | 2026-08-04 12:39 |
| negatives | collect hard negatives: the same fold, no PET activity | 1 | 3,985 | 256 | 2026-08-04 12:41 |
| harmonise | reconcile activity values reported under different assays | 1 | 459 | 47 | 2026-08-04 12:39 |
| acs_screen | — | 1 | 477 | 214 | 2026-08-06 22:04 |
Model evaluations
Every training run ever recorded, including the ones that failed. AUC is the chance the model ranks a true PET enzyme above a non-enzyme: 1.0 is perfect, 0.5 is a coin flip. The composition baseline beside it is what you get from amino-acid counts alone — a score only means something if it clears that.
| Trained on | Positives | Negatives | Clusters | AUC | Baseline | Reading |
|---|---|---|---|---|---|---|
| measured-only
pet_activity · 2026-08-06 |
308 | 385 | — | 0.774 | 0.565 | clears the composition baseline by a real margin |
| measured-only
pet_activity · 2026-08-06 |
308 | 385 | 59 | 0.845 | 0.565 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
852 | 385 | 97 | 0.835 | 0.565 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
308 | 385 | — | 0.739 | 0.565 | clears the composition baseline by a real margin |
| measured-only
pet_activity · 2026-08-06 |
308 | 385 | 48 | 0.790 | 0.565 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
852 | 385 | 65 | 0.859 | 0.565 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
302 | 241 | — | 0.845 | 0.803 | barely clears amino-acid composition |
| measured-only
pet_activity · 2026-08-06 |
302 | 241 | 60 | 0.947 | 0.803 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
789 | 241 | 97 | 0.972 | 0.803 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
302 | 241 | — | 0.857 | 0.803 | barely clears amino-acid composition |
| measured-only
pet_activity · 2026-08-06 |
302 | 241 | 46 | 0.949 | 0.803 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
789 | 241 | 61 | 0.981 | 0.803 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | — | 0.860 | 0.736 | clears the composition baseline by a real margin |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | 61 | 0.942 | 0.736 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
791 | 241 | 97 | 0.975 | 0.736 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | — | 0.782 | 0.736 | barely clears amino-acid composition |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | 46 | 0.921 | 0.736 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
791 | 241 | 60 | 0.967 | 0.736 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | — | 0.860 | 0.736 | clears the composition baseline by a real margin |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | 61 | 0.942 | 0.736 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
791 | 241 | 97 | 0.975 | 0.736 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | — | 0.782 | 0.736 | barely clears amino-acid composition |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | 46 | 0.921 | 0.736 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
791 | 241 | 60 | 0.967 | 0.736 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | — | 0.860 | 0.736 | clears the composition baseline by a real margin |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | 61 | 0.942 | 0.736 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
791 | 241 | 97 | 0.975 | 0.736 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | — | 0.782 | 0.736 | barely clears amino-acid composition |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | 46 | 0.921 | 0.736 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
791 | 241 | 60 | 0.967 | 0.736 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | — | 0.860 | 0.736 | clears the composition baseline by a real margin |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | 61 | 0.942 | 0.736 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
791 | 241 | 97 | 0.975 | 0.736 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | — | 0.782 | 0.736 | barely clears amino-acid composition |
| measured-only
pet_activity · 2026-08-06 |
304 | 241 | 46 | 0.921 | 0.736 | clears the composition baseline by a real margin |
| mixed
pet_activity · 2026-08-06 |
791 | 241 | 60 | 0.967 | 0.736 | trained on similarity-derived labels |
| measured-only
pet_activity · 2026-08-05 |
152 | 26 | 7 | 0.493 | 0.398 | at chance — the hardest contrast, and the honest answer |
| measured-only
pet_activity · 2026-08-05 |
305 | 26 | 44 | 0.850 | 0.651 | clears the composition baseline by a real margin |
| measured-only
pet_activity · 2026-08-05 |
300 | 220 | 45 | 0.976 | 0.829 | clears the composition baseline by a real margin |
| experimental
pet_activity · 2026-08-04 |
13 | 220 | 1 | — | — | not evaluable — too few independent clusters to split on |
| mixed
pet_activity · 2026-08-04 |
500 | 220 | 25 | 1.000 | 0.778 | perfect, and meaningless: it reproduced the similarity rule that made the labels |
The machine this runs on
A single DigitalOcean droplet shared with five other applications, which is why the web layer carries no scientific libraries at all: every structure and every measurement was computed offline and is served here as a precomputed file.
16.7 GB free of 82.1 GB
1.2 GB available of 4.11 GB. Memory is the binding constraint here, not disk or speed.
across 2 cores. Above the core count means work is queueing.
SQLite, one file, everything in it. 832 candidate and 559 reference structure files alongside it, totalling 139.5 + 117.6 MB.
Linux x86_64 · Python 3.12.3 · web process up 6.67 h · host up 1635.9 h · data version 0.3.0