Data curation
Despite the vast economic and time resources devoted over the last three decades in actively investigating treatments for neurodegenerative amyloidosis, their success remains scarce. While none of the 14 different compounds that targetting amyloid-β that made it through phase III clinical trials made it through, two anti-amyloid antibodies have received regulatory approval in Europe (lecanemab and donanemab), and a third one in the United States (aducanumab), for treating Alzheimer’s disease (AD). Lecanemab and donanemab have shown to slow cognitive decline by around 30%. Long-term clinical evidence support the therapeutic concept of antibody-removal of amyloid deposits, as demonstrated in post-mortem evaluations of Alzheimer’s disease patients 15 years post-immunization”. Additionally, mAbs are in clinical trials for various amyloidoses with no current disease-modifying therapies, including Parkinson’s Disease (PD), Transthyretin amyloidosis or Amyloid light chain (AL) amyloidosis. Instead, multiple mAbs have shown insufficient results or are only useful for certain disease stages, genotypes or phenotypes. Understanding the molecular mechanisms of why certain antibody therapeutics succeed and others do not could allow for increasing the success rate and decreasing economic burden for these incurable conditions.
The AmyloGraph Antibody Database collects peer-reviewed experimental information on the impact of amyloid-targeting antibodies have on the amyloid aggregation.
We establish controlled-vocabulary to describe the interaction, considering:
No effect when the antibody does not affect the amyloid aggregation.
Slower aggregation (partial inhibition) when the antibody kinetically slows down the deposition process.
Diminished amyloid quantity at the final point (partial inhibition) when the effect of the antibody is reflexed as a diminished amyloid quantity.
No aggregation (complete inhibition) when the amyloid quantity remaining lies below the quantitative threshold of the method.
Amyloid proteins
Amyloid proteins were obtained according to the latest update of the International Society of Amyloidosis (ISA) Nomenclature Committee, from the amyloid fibrol ptoreins and precursors in human (Aβ, AαSyn, ATau, APrP, ATTR, etc), as well as the Pathological intracellular protein aggregates which display at least some typical amyloid fibril properties (TARDBP, TIA1, SOD1, C9orf72, etc). These were supplemented with bacterial and yeast amyloids (CsgA, Sup35, etc). Finally, mammalian homologues of amyloid proteins (mice Aβ, APrP, ATau) were accepted as amyloid proteins if proof of amyloid aggregation was shown on the respective publication. The basic amyloid proteins initially considered for the AmyloGraph Antibody Database are reported in the following tables.
Human amyloid proteins
| Amyloid Species | Canonical Protein Entry | UniProt Link | DisGeNET Link |
|---|---|---|---|
| A𝛽-42 | P05067 | UniProt Entry | DisGeNET Entry |
| A𝛽-40 | P05067 | UniProt Entry | DisGeNET Entry |
| pE3-Aβ | P05067 | UniProt Entry | DisGeNET Entry |
| ɑ-Synuclein | P37840 | UniProt Entry | DisGeNET Entry |
| Tau | P10636 | UniProt Entry | |
| Immunoglobulin light chain | Multi-entry | Variable dependent on clonal sequence | Variable dependent on clonal sequence |
| Immunoglobulin heavy chain | Multi-entry | Variable dependent on clonal sequence | Variable dependent on clonal sequence |
| (Apo) Serum amyloid A | P0DJI8 | UniProt Entry | DisGeNET Entry |
| Transthyretin | P02766 | UniProt Entry | DisGeNET Entry |
| 𝛽2-microglobulin | P61769 | UniProt Entry | DisGeNET Entry |
| Apolipoprotein A-I | P02647 | UniProt Entry | DisGeNET Entry |
| Apolipoprotein A-II | P02652 | UniProt Entry | DisGeNET Entry |
| Apolipoprotein A-IV | P06727 | UniProt Entry | DisGeNET Entry |
| Apolipoprotein C-II | P02655 | UniProt Entry | DisGeNET Entry |
| Apolipoprotein C-III | P02656 | UniProt Entry | DisGeNET Entry |
| Gelsolin | P06396 | UniProt Entry | DisGeNET Entry |
| Lysozyme | P61626 | UniProt Entry | DisGeNET Entry |
| Leukocyte chemotactic factor 2 | O14960 | UniProt Entry | DisGeNET Entry |
| Fibrinogen ɑ | P02671 | UniProt Entry | DisGeNET Entry |
| Cystatin C | P01034 | UniProt Entry | DisGeNET Entry |
| ABriPP | Q9Y287 | UniProt Entry | DisGeNET Entry |
| ADanPP | Q9Y287 | UniProt Entry | DisGeNET Entry |
| Prion protein | P04156 | UniProt Entry | DisGeNET Entry |
| Transmembrane protein 106B | Q9NUM4 | UniProt Entry | DisGeNET Entry |
| (Pro) Calcitonin | P01258 | UniProt Entry | DisGeNET Entry |
| Islet amyloid polypeptide | P10997 | UniProt Entry | DisGeNET Entry |
| Atrial natriuretic peptide | P01160 | UniProt Entry | DisGeNET Entry |
| Prolactin | P01236 | UniProt Entry | DisGeNET Entry |
| (Pro) Somatostatin | P61278 | UniProt Entry | DisGeNET Entry |
| Glucagon | P01275 | UniProt Entry | DisGeNET Entry |
| Parathyroid hormone | P01270 | UniProt Entry | DisGeNET Entry |
| Lung surfactant protein C | P11686 | UniProt Entry | DisGeNET Entry |
| Corneodesmosin | Q15517 | UniProt Entry | DisGeNET Entry |
| Lactadherin (Medin) | Q08431 | UniProt Entry | DisGeNET Entry |
| Kerato-epithelin (TGFBI) | Q15582 | UniProt Entry | DisGeNET Entry |
| Lactoferrin | P02788 | UniProt Entry | DisGeNET Entry |
| Odontogenic ameloblast-associated protein | A1E959 | UniProt Entry | DisGeNET Entry |
| Semenogelin-1 | P04279 | UniProt Entry | DisGeNET Entry |
| Cathepsin K | P43235 | UniProt Entry | DisGeNET Entry |
| EGF-containing fibulin-like extracellular matrix protein 1 | Q12805 | UniProt Entry | DisGeNET Entry |
| Cytokeratin | P02533 | UniProt Entry | DisGeNET Entry |
Human proteins forming inclusions with amloid properties
| Amyloid-like Species | Canonical Protein Entry | UniProt Link | DisGeNET Link |
|---|---|---|---|
| p53 | P04637 | UniProt Entry | DisGeNET Entry |
| Desmin | P17661 | UniProt Entry | DisGeNET Entry |
| Galectin 7 | P47929 | UniProt Entry | DisGeNET Entry |
| Superoxide dismutase [SOD1] | P00441 | UniProt Entry | DisGeNET Entry |
| TAR DNA-binding protein 43 [TDP-43] | Q13148 | UniProt Entry | DisGeNET Entry |
| Huntingtin | P42858 | UniProt Entry | DisGeNET Entry |
| Androgen receptor | P10275 | UniProt Entry | DisGeNET Entry |
Microbial Amyloid Proteins
| Amyloid Species | Specie - Genus | UniProt Link |
|---|---|---|
| CsgA | E. coli / Aeromonas | UniProt Entry |
| CsgB | E. coli / Aeromonas | UniProt Entry |
| FapB | Pseudomonas | UniProt Entry |
| FapC | Pseudomonas | UniProt Entry |
| SSB | Campylobacter hominis | UniProt Entry |
| Transcription termination factor Rho | Clostridium botulinum | UniProt Entry |
| Sup35 | Saccharomyces cerevisiae | UniProt Entry |
| Ure2 | Saccharomyces cerevisiae | UniProt Entry |
| HET-s | Podospora anserina | UniProt Entry |
Note: different microbial species and strains have shown to form amyloid aggregates. Therefore, beyond this list, specific proteins of these families are accepted.
Antibodies and antibody-based designs
The database considers the effect of a single monoclonal antibody on the amyloid aggregation of an amyloid protein. Different species IgGs, orthodox, IgG inspired, and heterodox, novel multimeric, antibody formats were accepted and named according to the original publication, or according to the following classification.
Antibody-conjugated nanoparticles and nanoparticles engineered with an antibody derivative, on the other hand, were consdidered separately. Following the World Health Organization (WHO) International Nonproprietary Names (INN) naming convention, immunoglobulin fusions with peptides were as well considered an antibody if both domains have immunoglobulin derived variable domains. Conjugates with small molecules or radioisotopes as well. Larger nanoparticles than 50 nm, approximately 5 times a typical IgG, were discarded.
Therapeutic monoclonal antibodies evaluated in clinical trials
Tha AmyloGraph Antibody Database compiles data for most monoclonal antibodies evaluated in clinical trials. These species are catalogues in the table below.
| INN or development name | IMGT DB | PubChem | Thera-SAbDab | YABS DB | Other IDs | Parental/Heterologous Antibody ID |
|---|---|---|---|---|---|---|
| Aducanumab | IMGT Entry | PubChem | SAbDab | YABS DB | BART, BIIB-037, BIIB037, aducanumab-avwa, Aduhelm | None |
| Gantenerumab | IMGT Entry | PubChem | SAbDab | YABS DB | R1450, RG-1450, RG1450, Ro-4909832, RO4909832 | None |
| Donanemab | IMGT Entry | PubChem | SAbDab | YABS DB | LY3002813, donanemab-azbt | mE8 |
| Crenezumab | IMGT Entry | PubChem | SAbDab | N/A | MABT5102A, RG7412, MABT-5102A, RG-7412 | mC2 |
| Solanezumab | IMGT Entry | PubChem | SAbDab | YABS DB | LY2062430 | m266 |
| Lecanemab | IMGT Entry | PubChem | SAbDab | YABS DB | BAN-2401, BAN2401, lecanemab-irmb, Leqembi, Leqembi™ | mAb158 |
| Bapineuzumab | IMGT Entry | PubChem | SAbDab | YABS DB | AAB-001 | 3D6 |
| Cliramitug | IMGT Entry | PubChem | SAbDab | N/A | ALXN-2220, ALXN2220 | NI006 |
| Anselamimab | IMGT Entry | PubChem | SAbDab | YABS DB | CAEL-101, CAEL101 | 11-1F4 |
| Birtamimab | IMGT Entry | PubChem | SAbDab | YABS DB | NEOD001, NEOD-001, ELT1-01 | 2A4 |
| Etalanetug | IMGT Entry | PubChem | SAbDab | YABS DB | E2814 | 7G6 |
| SAR228810 | N/A | N/A | N/A | N/A | 13C3, SAR255952 | N/A |
| Prasinezumab | IMGT Entry | PubChem | SAbDab | N/A | PRX002, RG-7935, RG7935 | 9E4 |
| Cinpanemab | IMGT Entry | PubChem | SAbDab | N/A | BIIB054 | 12F4 |
| Amlenetug | IMGT Entry | PubChem | SAbDab | YABS DB | Lu AF82422, GM37 | GM37 |
| Exidavnemab | IMGT Entry | PubChem | SAbDab | N/A | BAN0805 | Ab47 |
| Indenebart | IMGT Entry | PubChem | SAbDab | N/A | MEDI1341, TAK-341 | None |
| Ponezumab | IMGT Entry | PubChem | SAbDab | N/A | PF-04360365 | 2H6 |
| Gosuranemab | IMGT Entry | PubChem | SAbDab | N/A | BIIB092 | IPN002 |
The AmyloGraph Antidoby Database Curation Pipeline
Data acquisition
We designed keyword-based queries to obtain peer-reviewed publications reporting the effect of antibody and antibodies or nanobody and nanobodies, on amyloid aggregation.
- "amyloid"[Title/Abstract] AND "antibod*"[Title/Abstract]
- "amyloid"[Title/Abstract] AND "nanobod*"[Title/Abstract]
This query was executed on October 24, 2024. After removing duplicates, a total of 4023 unique publications were identified. An additional 58 publications were identified through manual screening of reference lists and included in the final analysis.
Initial curation
After first assessment on the query and non-query publications, 6.81% (209) publications were deemed as “Useful”, 167 corresponding to the query and 42 to the non-query bibliographic addition.
After the secondary curation, 8 (3.82%) of the initially taged as “Useful” publications were re-avaluated as “Not Useful”, while 1 publicaiton was retracted during the curation process.
Development of a Machine Learning method to improve curation queues
The large body of publications to initially review presented a significant bottleneck in terms of human and time resources. In parallel, we developed a Machine Learning methodology which could help prioritize curation queues. Once 25% of the publications were assessed, we used the Title, Abstracts and references of the positive “Useful” and negative “Rejected” publications to train this algorithm. The subsequent iterations of it were fed with the increasing number of reviewed publications in a continuous loop. Additionally, the revisions were tagged with “Reason for Rejection” (as presented in the Inclusion-Exclusion Criteria & Curation Tags section), to improve the ML curation queues development.
The curation of the AmyloGraph Antibody DB, and development of a ML curation queue algorithm, were considered as a topic for the German ELIXIR BioHackathon, 2025.
Two-step indepedent data curation
To ensure maintaining curation consistency across time and different curators, we implemented various overlapping strategies. We scheduled regular meetings between the Leading Curators and the Junior Curators, a one-on-one introduction meeting with each Junior Curator, provided a Curator handbook with definitions, FAQs and curation examples, and set-up a Slack channel for curation-related discussions.
Curator handbook with examples
Curators were provided with a reference guide featuring clear definitions FAQs, and real world curation examples. Definitions include what is considered an antibody, an amyloid, the biophysical, biomedical and biochemical basis of the most employed techniques to report changes in amyloid aggregation. FAQs include how to report non-human versions of the amyloid proteins, how to consider ambiguous entities (such as publications reporting “amyloid plaques” instead of a unique Aβ₄₂ or Aβ₄₀ species), how to annotate Immunoglobulin light chain (AL) amyloidosis when no unnequivocal sequential data is provided, etc. This trianing resource ensures maximum reproducibility and consistency across all annotators of the AmyloGraph Antibody Database.
First step: standardising annotation and description of the data.
To ensure maximum granularity and data integrity, our curation process divides each publication across three interconnected modules: Antibody Properties, Amyloid Targets, and Interaction Profiles. To maximize data standardization, these modules are pre-populated (in the forms of Google Forms), with acceptable answers to publication doi, the amyloid protein, the antibody, antibody type, impact of the antibody on the aggregation, etc.
- If a publication reports the effect on multiple (>1) amyloids, the curator will fill in one “Amyloid form” for each species.
- If a publication reports several antibodies, the curator will fill in one “Antibody form” for each species.
Finally, the curator will fill one “Interaction form” for each antibody-amyloid pair.
Second, independent, machine-assisted curation
Secondary curation was assisted by means of two python libraries: MarkItDown and LangExtract to accelerate curation and minimize data leakage.
- MarkItDown, developed by Microsoft, converts different document formats—including into clean, structured Markdown text.
MarkItDown was implemented for PDF-only publications, to generate Markdown files that could be processed by LangExtract.
- LangExtract, developed by Google, uses LLMs to read text-based publications, highlighting the rellevant information and develop human-friendly htmls with the rellevant text highlighted.
Text data from the publications were processed by LangExtract coupled with Ollama running the Gemma3:4b model.
Further considerations
Amyloid plaques
Amyloid plaques are composed of different proteotypes of the amyloid βeta peptide, such as Aβ₄₂, Aβ₄₀, Aβ₃₈, and Pyroglutamate-Aβ (pE3-Aβ). Despite this, the overwhelingly major constituent is the Aβ₄₂ peptide (42 amino acids). Therefore, for in vivo experiments reporting changes in amyloid plaques (plaque number, plaque density, Positron Emission Tomography uptake values…) we annotate Aβ₄₂ as the amyloid interactor.
Considerations on experimental amyloid reporters
The ISA recommends the use of Serum Free Light Chains (sFLC) as a the biomarker to follow Immunoglobulin light chain (AL) amyloidosis disease progression and therapeutic response number. sFLC reports the disease progression, but is an indirect reporter of the amyloid state, the amyloid deposits. Therefore, we have not included it as evidence of changes in amyloid aggregation. In a similar fashion, we have considered “indirect” other reporters such as clinical dementia rating, cellular cytotoxicity, Morris-water maze performance, brain soluble Aβ₄₀ or Aβ₄₂ levels, microglia uptake, or % truncated α-Synuclein.
Contact
:::