Skip to content

Add crosslinking proteomics SDRF annotations (batch 06/15, 50 datasets) - #544

Merged
ypriverol merged 4 commits into
mainfrom
add/crosslinking-batch-06
Sep 17, 2026
Merged

ypriverol merged 4 commits into
mainfrom
add/crosslinking-batch-06

Conversation

@ypriverol

Copy link
Copy Markdown
Contributor

Part of the crosslinking proteomics SDRF annotation effort (splits PR #537 into batches of 50). Adds 50 new datasets in flat datasets/<accession>/ layout. Accessions: PXD021708,PXD021709,PXD021770,PXD021809,PXD021822,PXD021831,PXD021870,PXD021923,PXD022119,PXD022279,PXD022335,PXD022440,PXD022443,PXD022608,PXD022690,PXD022772,PXD022785,PXD022861,PXD022991,PXD023072,PXD023164,PXD023221,PXD023239,PXD023277,PXD023522,PXD023525,PXD023542,PXD023577,PXD023814,PXD024010,PXD024065,PXD024131,PXD024160,PXD024253,PXD024335,PXD024366,PXD024367,PXD024399,PXD024822,PXD024946,PXD025066,PXD025099,PXD025172,PXD025208,PXD025220,PXD025357,PXD025581,PXD025662,PXD025843,PXD026037

Copilot AI balanced review requested due to automatic review settings September 17, 2026 04:39

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Sep 17, 2026

Copy link
Copy Markdown

Important

Review skipped

Review was skipped due to path filters

⛔ Files ignored due to path filters (50)
  • datasets/PXD021708/PXD021708.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD021709/PXD021709.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD021770/PXD021770.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD021809/PXD021809.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD021822/PXD021822.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD021831/PXD021831.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD021870/PXD021870.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD021923/PXD021923.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022119/PXD022119.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022279/PXD022279.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022335/PXD022335.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022440/PXD022440.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022443/PXD022443.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022608/PXD022608.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022690/PXD022690.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022772/PXD022772.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022785/PXD022785.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022861/PXD022861.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD022991/PXD022991.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023072/PXD023072.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023164/PXD023164.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023221/PXD023221.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023239/PXD023239.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023277/PXD023277.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023522/PXD023522.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023525/PXD023525.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023542/PXD023542.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023577/PXD023577.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD023814/PXD023814.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024010/PXD024010.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024065/PXD024065.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024131/PXD024131.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024160/PXD024160.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024253/PXD024253.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024335/PXD024335.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024366/PXD024366.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024367/PXD024367.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024399/PXD024399.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024822/PXD024822.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD024946/PXD024946.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD025066/PXD025066.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD025099/PXD025099.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD025172/PXD025172.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD025208/PXD025208.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD025220/PXD025220.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD025357/PXD025357.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD025581/PXD025581.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD025662/PXD025662.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD025843/PXD025843.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD026037/PXD026037.sdrf.tsv is excluded by !**/*.tsv

CodeRabbit blocks several paths by default. You can override this behavior by explicitly including those paths in the path filters. For example, including **/dist/** will override the default block on the dist directory, by removing the pattern from both the lists.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 73f87bec-a15d-4d75-a09f-bbdba0fb3d9e

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@qodo-code-review

qodo-code-review Bot commented Sep 17, 2026

Copy link
Copy Markdown

PR Summary by Qodo

Add crosslinking proteomics SDRF annotations for 50 datasets

✨ Enhancement 🐞 Bug fix 🕐 40+ Minutes

Grey Divider

AI Description

• Adds SDRF annotations for 50 crosslinking proteomics datasets.
• Supplies required organism-template fields with spec-compliant unavailable values.
• Corrects PXD022861 instrument assignments for Orbitrap crosslinking runs.
Diagram

graph TD
  A["PRIDE datasets"] --> B["50 SDRF files"] --> C["Sample metadata"] --> D["Crosslink metadata"] --> E["Assay mappings"]
  C --> F["Organism templates"]
  F --> D
Loading
High-Level Assessment

The flat accession layout and 50-dataset batch are appropriate for independently reviewable, curated SDRF metadata. Automated generation was considered, but heterogeneous source records and dataset-specific instrument and crosslinker interpretation still require manual curation.

Files changed (50) +1187 / -0

Enhancement (49) +1173 / -0
PXD021708.sdrf.tsvAdd PXD021708 DSSO SDRF annotation +5/-0

Add PXD021708 DSSO SDRF annotation

• Adds four human DSSO crosslinking assay annotations with Orbitrap Fusion Lumos acquisition metadata and SDRF 1.1 templates.

datasets/PXD021708/PXD021708.sdrf.tsv

PXD021709.sdrf.tsvAdd PXD021709 BioID SDRF annotation +9/-0

Add PXD021709 BioID SDRF annotation

• Adds eight fraction-level human BioID assay mappings with Orbitrap Fusion Lumos metadata.

datasets/PXD021709/PXD021709.sdrf.tsv

PXD021770.sdrf.tsvAdd PXD021770 TurboID SDRF annotation +15/-0

Add PXD021770 TurboID SDRF annotation

• Adds fourteen human TurboID entries covering timsTOF Pro runs and associated submission files.

datasets/PXD021770/PXD021770.sdrf.tsv

PXD021809.sdrf.tsvAdd PXD021809 TurboID SDRF annotation +57/-0

Add PXD021809 TurboID SDRF annotation

• Adds 56 human TurboID assay and support-file mappings acquired with a timsTOF Pro.

datasets/PXD021809/PXD021809.sdrf.tsv

PXD021822.sdrf.tsvAdd PXD021822 bacterial crosslinking annotation +46/-0

Add PXD021822 bacterial crosslinking annotation

• Adds 45 Deinococcus radiodurans assay and result-file mappings with Q-TOF acquisition metadata.

datasets/PXD021822/PXD021822.sdrf.tsv

PXD021831.sdrf.tsvAdd PXD021831 yeast DSS SDRF annotation +80/-0

Add PXD021831 yeast DSS SDRF annotation

• Adds 79 Saccharomyces cerevisiae DSS assays and includes required developmental-stage and strain fields.

datasets/PXD021831/PXD021831.sdrf.tsv

PXD021870.sdrf.tsvAdd PXD021870 E. coli crosslinking annotation +49/-0

Add PXD021870 E. coli crosslinking annotation

• Adds 48 E. coli crosslinking assays acquired by Q Exactive HF-X across protease conditions.

datasets/PXD021870/PXD021870.sdrf.tsv

PXD021923.sdrf.tsvAdd PXD021923 TurboID SDRF annotation +25/-0

Add PXD021923 TurboID SDRF annotation

• Adds 24 human TurboID assays acquired using an Orbitrap Exploris 480.

datasets/PXD021923/PXD021923.sdrf.tsv

PXD022119.sdrf.tsvAdd PXD022119 BS3 SDRF annotation +2/-0

Add PXD022119 BS3 SDRF annotation

• Adds one human BS3 crosslinking run with Orbitrap Fusion Lumos metadata.

datasets/PXD022119/PXD022119.sdrf.tsv

PXD022279.sdrf.tsvAdd PXD022279 crosslinking SDRF annotation +26/-0

Add PXD022279 crosslinking SDRF annotation

• Adds 25 human crosslinking assays acquired with a Q Exactive HF.

datasets/PXD022279/PXD022279.sdrf.tsv

PXD022335.sdrf.tsvAdd PXD022335 mouse APEX2 annotation +53/-0

Add PXD022335 mouse APEX2 annotation

• Adds 52 mouse APEX2 assays with Orbitrap Eclipse metadata and vertebrate-template fields.

datasets/PXD022335/PXD022335.sdrf.tsv

PXD022440.sdrf.tsvAdd PXD022440 formaldehyde SDRF annotation +14/-0

Add PXD022440 formaldehyde SDRF annotation

• Adds thirteen Xenopus laevis formaldehyde-crosslinking submission files acquired with a TripleTOF 4600.

datasets/PXD022440/PXD022440.sdrf.tsv

PXD022443.sdrf.tsvAdd PXD022443 DSS SDRF annotation +7/-0

Add PXD022443 DSS SDRF annotation

• Adds six human DSS crosslinking assays acquired with an Orbitrap Fusion Lumos.

datasets/PXD022443/PXD022443.sdrf.tsv

PXD022608.sdrf.tsvAdd PXD022608 BioID SDRF annotation +21/-0

Add PXD022608 BioID SDRF annotation

• Adds twenty fractionated human BioID assays acquired with a Q Exactive.

datasets/PXD022608/PXD022608.sdrf.tsv

PXD022690.sdrf.tsvAdd PXD022690 yeast SDA annotation +30/-0

Add PXD022690 yeast SDA annotation

• Adds 29 yeast SDA crosslinking runs with Orbitrap Fusion Lumos and invertebrate-template metadata.

datasets/PXD022690/PXD022690.sdrf.tsv

PXD022772.sdrf.tsvAdd PXD022772 DSSO SDRF annotation +17/-0

Add PXD022772 DSSO SDRF annotation

• Adds sixteen Streptococcus pyogenes DSSO crosslinking acquisition and analysis-file mappings for timsTOF Pro.

datasets/PXD022772/PXD022772.sdrf.tsv

PXD022785.sdrf.tsvAdd PXD022785 Toxoplasma BioID annotation +11/-0

Add PXD022785 Toxoplasma BioID annotation

• Adds ten Toxoplasma gondii BioID acquisition, result, and reference-file mappings.

datasets/PXD022785/PXD022785.sdrf.tsv

PXD022991.sdrf.tsvAdd PXD022991 crosslinking SDRF annotation +25/-0

Add PXD022991 crosslinking SDRF annotation

• Adds 24 human crosslinking assays with Q Exactive HF acquisition metadata.

datasets/PXD022991/PXD022991.sdrf.tsv

PXD023072.sdrf.tsvAdd PXD023072 mouse PhoX annotation +56/-0

Add PXD023072 mouse PhoX annotation

• Adds 55 mouse PhoX assay and analysis-file mappings with Orbitrap Fusion Lumos metadata.

datasets/PXD023072/PXD023072.sdrf.tsv

PXD023164.sdrf.tsvAdd PXD023164 yeast DSSO annotation +7/-0

Add PXD023164 yeast DSSO annotation

• Adds six yeast DSSO crosslinking replicates acquired with an Orbitrap Fusion.

datasets/PXD023164/PXD023164.sdrf.tsv

PXD023221.sdrf.tsvAdd PXD023221 BS3 SDRF annotation +12/-0

Add PXD023221 BS3 SDRF annotation

• Adds eleven human BS3 assay and archive mappings acquired with a SYNAPT G2-Si.

datasets/PXD023221/PXD023221.sdrf.tsv

PXD023239.sdrf.tsvAdd PXD023239 BioID SDRF annotation +64/-0

Add PXD023239 BioID SDRF annotation

• Adds 63 human BioID assays acquired with a Q Exactive Plus.

datasets/PXD023239/PXD023239.sdrf.tsv

PXD023277.sdrf.tsvAdd PXD023277 BioID SDRF annotation +41/-0

Add PXD023277 BioID SDRF annotation

• Adds forty human BioID assays and fraction mappings acquired with a Q Exactive Plus.

datasets/PXD023277/PXD023277.sdrf.tsv

PXD023522.sdrf.tsvAdd PXD023522 crosslinking SDRF annotation +11/-0

Add PXD023522 crosslinking SDRF annotation

• Adds ten human crosslinking assay mappings with ultraflex instrument metadata.

datasets/PXD023522/PXD023522.sdrf.tsv

PXD023525.sdrf.tsvAdd PXD023525 DSBU SDRF annotation +7/-0

Add PXD023525 DSBU SDRF annotation

• Adds six E. coli DSBU acquisition and result-file mappings for Q Exactive HF data.

datasets/PXD023525/PXD023525.sdrf.tsv

PXD023542.sdrf.tsvAdd PXD023542 DSS SDRF annotation +28/-0

Add PXD023542 DSS SDRF annotation

• Adds 27 human DSS crosslinking and control assays acquired with a Q Exactive HF.

datasets/PXD023542/PXD023542.sdrf.tsv

PXD023577.sdrf.tsvAdd PXD023577 PIR SDRF annotation +19/-0

Add PXD023577 PIR SDRF annotation

• Adds eighteen fractionated human PIR crosslinking assays acquired with a Q Exactive HF.

datasets/PXD023577/PXD023577.sdrf.tsv

PXD023814.sdrf.tsvAdd PXD023814 APEX SDRF annotation +51/-0

Add PXD023814 APEX SDRF annotation

• Adds fifty human APEX proximity-labeling assays acquired with an Orbitrap Fusion Lumos.

datasets/PXD023814/PXD023814.sdrf.tsv

PXD024010.sdrf.tsvAdd PXD024010 crosslinking SDRF annotation +30/-0

Add PXD024010 crosslinking SDRF annotation

• Adds 29 human crosslinking assays acquired with an Orbitrap Fusion Lumos.

datasets/PXD024010/PXD024010.sdrf.tsv

PXD024065.sdrf.tsvAdd PXD024065 yeast crosslinking annotation +5/-0

Add PXD024065 yeast crosslinking annotation

• Adds four yeast crosslinking acquisition and result-file mappings with LTQ Orbitrap Elite metadata.

datasets/PXD024065/PXD024065.sdrf.tsv

PXD024131.sdrf.tsvAdd PXD024131 yeast DSS annotation +5/-0

Add PXD024131 yeast DSS annotation

• Adds four yeast DSS crosslinking assays with required invertebrate-template characteristics.

datasets/PXD024131/PXD024131.sdrf.tsv

PXD024160.sdrf.tsvAdd PXD024160 yeast crosslinking annotation +10/-0

Add PXD024160 yeast crosslinking annotation

• Adds nine yeast crosslinking acquisition, result, and checksum mappings with Q Exactive HF metadata.

datasets/PXD024160/PXD024160.sdrf.tsv

PXD024253.sdrf.tsvAdd PXD024253 E. coli crosslinking annotation +8/-0

Add PXD024253 E. coli crosslinking annotation

• Adds seven E. coli crosslinking acquisition and result-file mappings with Orbitrap Fusion metadata.

datasets/PXD024253/PXD024253.sdrf.tsv

PXD024335.sdrf.tsvAdd PXD024335 APEX2 SDRF annotation +7/-0

Add PXD024335 APEX2 SDRF annotation

• Adds six human APEX2 proximity-labeling assays acquired with a Q Exactive HF.

datasets/PXD024335/PXD024335.sdrf.tsv

PXD024366.sdrf.tsvAdd PXD024366 mouse crosslinking annotation +49/-0

Add PXD024366 mouse crosslinking annotation

• Adds 48 mouse crosslinking RAW and mzXML mappings with Orbitrap Fusion Lumos metadata.

datasets/PXD024366/PXD024366.sdrf.tsv

PXD024367.sdrf.tsvAdd PXD024367 mouse crosslinking annotation +25/-0

Add PXD024367 mouse crosslinking annotation

• Adds 24 fractionated mouse crosslinking RAW and mzXML mappings with vertebrate-template metadata.

datasets/PXD024367/PXD024367.sdrf.tsv

PXD024399.sdrf.tsvAdd PXD024399 BioID SDRF annotation +61/-0

Add PXD024399 BioID SDRF annotation

• Adds sixty human SCYL1 BioID assays covering wild-type and variant conditions.

datasets/PXD024399/PXD024399.sdrf.tsv

PXD024822.sdrf.tsvAdd PXD024822 DMTMM SDRF annotation +25/-0

Add PXD024822 DMTMM SDRF annotation

• Adds 24 human DMTMM crosslinking assays acquired with an Orbitrap Fusion Lumos.

datasets/PXD024822/PXD024822.sdrf.tsv

PXD024946.sdrf.tsvAdd PXD024946 BS3 SDRF annotation +15/-0

Add PXD024946 BS3 SDRF annotation

• Adds fourteen human BS3 crosslinking gel assays acquired with a Q Exactive.

datasets/PXD024946/PXD024946.sdrf.tsv

PXD025066.sdrf.tsvAdd PXD025066 rabbit crosslinking annotation +7/-0

Add PXD025066 rabbit crosslinking annotation

• Adds six rabbit crosslinking acquisition and result-file mappings with Orbitrap Fusion Lumos metadata.

datasets/PXD025066/PXD025066.sdrf.tsv

PXD025099.sdrf.tsvAdd PXD025099 DSS SDRF annotation +31/-0

Add PXD025099 DSS SDRF annotation

• Adds thirty human DSS crosslinking and control assays acquired with a Q Exactive HF.

datasets/PXD025099/PXD025099.sdrf.tsv

PXD025172.sdrf.tsvAdd PXD025172 yeast DSSO annotation +21/-0

Add PXD025172 yeast DSSO annotation

• Adds twenty yeast DSSO TRAPPIII and HILIC assays with Orbitrap Fusion Lumos metadata.

datasets/PXD025172/PXD025172.sdrf.tsv

PXD025208.sdrf.tsvAdd PXD025208 BioID SDRF annotation +11/-0

Add PXD025208 BioID SDRF annotation

• Adds ten human BioID assays acquired with an Orbitrap Fusion Lumos.

datasets/PXD025208/PXD025208.sdrf.tsv

PXD025220.sdrf.tsvAdd PXD025220 BS3 SDRF annotation +3/-0

Add PXD025220 BS3 SDRF annotation

• Adds two Trypanosoma brucei BS3 crosslinking assays acquired with a Q Exactive.

datasets/PXD025220/PXD025220.sdrf.tsv

PXD025357.sdrf.tsvAdd PXD025357 BioID SDRF annotation +10/-0

Add PXD025357 BioID SDRF annotation

• Adds nine Trypanosoma brucei BioID acquisition and support-file mappings.

datasets/PXD025357/PXD025357.sdrf.tsv

PXD025581.sdrf.tsvAdd PXD025581 BioID SDRF annotation +16/-0

Add PXD025581 BioID SDRF annotation

• Adds fifteen human BioID assays acquired with an LTQ Orbitrap Elite.

datasets/PXD025581/PXD025581.sdrf.tsv

PXD025662.sdrf.tsvAdd PXD025662 DSS SDRF annotation +17/-0

Add PXD025662 DSS SDRF annotation

• Adds sixteen human DSS crosslinking assays acquired with an LTQ Orbitrap Elite.

datasets/PXD025662/PXD025662.sdrf.tsv

PXD025843.sdrf.tsvAdd PXD025843 yeast DMTMM annotation +23/-0

Add PXD025843 yeast DMTMM annotation

• Adds 22 yeast DMTMM crosslinking assays with required developmental-stage and strain characteristics.

datasets/PXD025843/PXD025843.sdrf.tsv

PXD026037.sdrf.tsvAdd PXD026037 horse crosslinking annotation +6/-0

Add PXD026037 horse crosslinking annotation

• Adds five horse crosslinking acquisition and result-file mappings with Q Exactive metadata.

datasets/PXD026037/PXD026037.sdrf.tsv

Bug fix (1) +14 / -0
PXD022861.sdrf.tsvAdd PXD022861 SDRF with corrected instruments +14/-0

Add PXD022861 SDRF with corrected instruments

• Adds thirteen human BS3/DSS and HX-MS entries. Assigns the two RAW crosslinking runs to Orbitrap Fusion Lumos while retaining TripleTOF 5600 for the eleven WIFF runs.

datasets/PXD022861/PXD022861.sdrf.tsv

@qodo-code-review

qodo-code-review Bot commented Sep 17, 2026

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (8) 📘 Rule violations (9) 📜 Skill insights (0)

⚠️ 14 lower-priority findings omitted to fit the comment size limit; re-run the review or view the findings in the Qodo portal.

Grey Divider


Action required

1. Identification results become assays 📘 Rule violation ≡ Correctness ⭐ New
Description
Rows 2–3 assign two .mzid.gz identification-result files as assay names and data files with
Orbitrap acquisition metadata. Genuine .raw acquisitions occur in rows 4–7, so the result files
add two artificial experimental fractions.
Code

datasets/PXD025066/PXD025066.sdrf.tsv[R2-3]

+PXD025066-sample	oryctolagus sp. 'rabbit_od'	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	3690_LM_C4H2O2_TrypAspN_Xi1.7.6.1.mzid.gz	proteomic profiling by mass spectrometry	3690_LM_C4H2O2_TrypAspN_Xi1.7.6.1.mzid.gz	1	1	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion Lumos;AC=MS:1002732	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
+PXD025066-sample	oryctolagus sp. 'rabbit_od'	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	3690_LM_C4H2O2_TrypChymo_Xi1.7.6.1.mzid.gz	proteomic profiling by mass spectrometry	3690_LM_C4H2O2_TrypChymo_Xi1.7.6.1.mzid.gz	1	2	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion Lumos;AC=MS:1002732	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
Rule 4 requires file mappings to represent the archive evidence accurately. Rows 2–3 treat
identification outputs as assays, while rows 4–7 separately identify the instrument-generated raw
files.

AGENTS.md: Align SDRF Metadata with Archive Evidence
datasets/PXD025066/PXD025066.sdrf.tsv[2-7]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Compressed mzIdentML identification results are modeled as mass-spectrometry acquisitions even though corresponding raw acquisitions are listed separately.

## Fix Focus Areas
- datasets/PXD025066/PXD025066.sdrf.tsv[2-3]

## Recommended Fix
Remove the `.mzid.gz` rows and retain only genuine acquisition files as SDRF assays.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Bovine albumin has wrong species 🐞 Bug ≡ Correctness ⭐ New
Description
The BSA-named assays in PXD024253.sdrf.tsv set characteristics[organism] to escherichia coli.
Those four BSA acquisitions are therefore indexed as bacterial material while the separate E. coli
assay group is explicitly identifiable as Ecoli in its names.
Code

datasets/PXD024253/PXD024253.sdrf.tsv[2]

+PXD024253-sample	escherichia coli	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	117-pDSBE-BSA2-_1_.mzML	proteomic profiling by mass spectrometry	117-pDSBE-BSA2-_1_.mzML	1	1	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion;AC=MS:1002416	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
Rows 2–5 consistently name the material BSA2 while declaring E. coli, whereas rows 6–8 use an
explicitly E. coli-named acquisition group under the same organism value.

datasets/PXD024253/PXD024253.sdrf.tsv[2-8]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The `117-pDSBE-BSA2` acquisition group is bovine serum albumin material but is annotated as `escherichia coli` in the organism characteristic.

## Fix Focus Areas
- datasets/PXD024253/PXD024253.sdrf.tsv[2-5]

## Recommended Fix
Change `characteristics[organism]` for the four `117-pDSBE-BSA2` rows to the appropriate bovine organism value. Retain the existing E. coli organism annotation for the `155-pDSBE-Ecoli-Z2` rows.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. Three result archives become assays 📘 Rule violation ≡ Correctness ⭐ New
Description
Rows 2, 9, and 10 of PXD025357.sdrf.tsv assign andromeda.zip, search.zip, and text.zip Q
Exactive acquisition metadata and separate fraction identifiers. Because the six intervening .raw
records are the actual instrument runs, treating these packaged search outputs as acquisitions adds
three artificial assays.
Code

datasets/PXD025357/PXD025357.sdrf.tsv[R9-10]

+PXD025357-sample	trypanosoma brucei	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	search.zip	proteomic profiling by mass spectrometry	search.zip	1	8	AC=MS:1002038;NT=label free sample	NT=Q Exactive;AC=MS:1001911	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=BioID;AC=XLMOD:02250	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
+PXD025357-sample	trypanosoma brucei	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	text.zip	proteomic profiling by mass spectrometry	text.zip	1	9	AC=MS:1002038;NT=label free sample	NT=Q Exactive;AC=MS:1001911	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=BioID;AC=XLMOD:02250	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
Lines 3–8 map six acquisition stems to genuine Q Exactive .raw files, while lines 2, 9, and 10
assign the same acquisition metadata and distinct fractions to generic Andromeda, search, and text
result archives, despite the requirement that assay mappings reflect archive evidence.

AGENTS.md: Align SDRF Metadata with Archive Evidence
datasets/PXD025357/PXD025357.sdrf.tsv[2-10]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Three packaged search-result archives are assigned instrument and acquisition metadata as independent experimental assays alongside the dataset's actual raw files.

## Fix Focus Areas
- datasets/PXD025357/PXD025357.sdrf.tsv[2-10]

## Recommended Fix
Remove the `andromeda.zip`, `search.zip`, and `text.zip` assay rows, retaining only the six `.raw` acquisition mappings.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View high (14)
4. Known chemistry remains unidentified 🐞 Bug ≡ Correctness ⭐ New
Description
PXD023522.sdrf.tsv records NT=unknown crosslinker for acquisitions whose assay and data-file
names explicitly identify DSG. This affects the DSG10, DSG20, and DSG50 rows, preventing structured
consumers from recognizing their cross-linking chemistry.
Code

datasets/PXD023522/PXD023522.sdrf.tsv[R4-7]

+PXD023522-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	LEDG_DSG10_11580.mzXML	proteomic profiling by mass spectrometry	LEDG_DSG10_11580.mzXML	1	3	AC=MS:1002038;NT=label free sample	NT=ultraflex;AC=MS:1000201	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
+PXD023522-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	LEDG_DSG20_11581.mzXML	proteomic profiling by mass spectrometry	LEDG_DSG20_11581.mzXML	1	4	AC=MS:1002038;NT=label free sample	NT=ultraflex;AC=MS:1000201	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
+PXD023522-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	Nkrp1B_CE_14N15N_DSG20_4007.mzXML	proteomic profiling by mass spectrometry	Nkrp1B_CE_14N15N_DSG20_4007.mzXML	1	5	AC=MS:1002038;NT=label free sample	NT=ultraflex;AC=MS:1000201	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
+PXD023522-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	Nkrp1B_CE_14N15N_DSG50_4008.mzXML	proteomic profiling by mass spectrometry	Nkrp1B_CE_14N15N_DSG50_4008.mzXML	1	6	AC=MS:1002038;NT=label free sample	NT=ultraflex;AC=MS:1000201	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
Evidence
Rows 4–7 identify DSG10, DSG20, or DSG50 in both filename fields while retaining the
unknown-crosslinker term; rows 9–10 repeat the same contradiction.

datasets/PXD023522/PXD023522.sdrf.tsv[4-10]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Multiple acquisition names explicitly identify DSG, but their structured cross-linker field says the chemistry is unknown.

## Fix Focus Areas
- datasets/PXD023522/PXD023522.sdrf.tsv[4-10]

## Recommended Fix
Replace `NT=unknown crosslinker;AC=XLMOD:00000` with the appropriate controlled DSG cross-linker term on every DSG-named row; review control rows separately rather than applying the replacement globally.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


5. Isotope-labelled runs appear label-free 🐞 Bug ≡ Correctness ⭐ New
Description
PXD023522.sdrf.tsv assigns AC=MS:1002038;NT=label free sample to assays whose names and data
files explicitly contain 14N15N. The four affected acquisitions are consequently indistinguishable
from genuinely label-free runs for consumers of the structured label field.
Code

datasets/PXD023522/PXD023522.sdrf.tsv[6]

+PXD023522-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	Nkrp1B_CE_14N15N_DSG20_4007.mzXML	proteomic profiling by mass spectrometry	Nkrp1B_CE_14N15N_DSG20_4007.mzXML	1	5	AC=MS:1002038;NT=label free sample	NT=ultraflex;AC=MS:1000201	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
Evidence
The changed rows directly pair 14N15N in both the assay and data-file fields with the label-free
ontology value; this applies to four independent acquisitions.

datasets/PXD023522/PXD023522.sdrf.tsv[1-10]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Rows whose assay and data-file values contain `14N15N` are recorded as label-free even though those names identify nitrogen-isotope-labelled acquisitions.

## Fix Focus Areas
- datasets/PXD023522/PXD023522.sdrf.tsv[6-10]

## Recommended Fix
Replace the label-free value on each `14N15N` row with the appropriate controlled-vocabulary isotope-labelling annotation. Leave the genuinely label-free rows unchanged.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


6. Digests are labelled as Lys-C 🐞 Bug ≡ Correctness ⭐ New
Description
PXD025066.sdrf.tsv records Lys-C as the second cleavage agent for TrypAspN and TrypChymo
assays. Each AspN- or chymotrypsin-containing run therefore has digestion metadata that conflicts
with its own acquisition identifier.
Code

datasets/PXD025066/PXD025066.sdrf.tsv[R2-3]

+PXD025066-sample	oryctolagus sp. 'rabbit_od'	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	3690_LM_C4H2O2_TrypAspN_Xi1.7.6.1.mzid.gz	proteomic profiling by mass spectrometry	3690_LM_C4H2O2_TrypAspN_Xi1.7.6.1.mzid.gz	1	1	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion Lumos;AC=MS:1002732	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
+PXD025066-sample	oryctolagus sp. 'rabbit_od'	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	3690_LM_C4H2O2_TrypChymo_Xi1.7.6.1.mzid.gz	proteomic profiling by mass spectrometry	3690_LM_C4H2O2_TrypChymo_Xi1.7.6.1.mzid.gz	1	2	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion Lumos;AC=MS:1002732	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
The header declares two cleavage-agent columns, and all six data rows pair a TrypAspN or TrypChymo
identifier with the same Trypsin/Lys-C annotation.

datasets/PXD025066/PXD025066.sdrf.tsv[1-7]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The second cleavage-agent column says Lys-C for assays named `TrypAspN` and `TrypChymo`, which identify AspN and chymotrypsin digestions instead.

## Fix Focus Areas
- datasets/PXD025066/PXD025066.sdrf.tsv[2-7]

## Recommended Fix
Replace the second cleavage-agent annotation with the appropriate AspN value on every `TrypAspN` row and the appropriate chymotrypsin value on every `TrypChymo` row. Keep Trypsin as the first agent if the samples were digested in combination with trypsin.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


7. A BS3 run is labeled as DSS 📘 Rule violation ≡ Correctness ⭐ New
Description
Row 27 maps the 2020_11_12_04_NSP1_strep_BS3_P2 assay and raw file to NT=DSS;AC=XLMOD:02001. The
explicit BS3 filename conflicts with that cross-linker while neighboring DSS-named acquisitions
use the DSS term consistently.
Code

datasets/PXD023542/PXD023542.sdrf.tsv[27]

+PXD023542-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	2020_11_12_04_NSP1_strep_BS3_P2	proteomic profiling by mass spectrometry	2020_11_12_04_NSP1_strep_BS3_P2.raw	1	26	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF;AC=MS:1002523	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=DSS;AC=XLMOD:02001	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
Evidence
Rule 4 requires annotations to agree with archive file mappings. The same added row identifies the
acquisition as BS3 but assigns the distinct DSS cross-linker term.

AGENTS.md: Align SDRF Metadata with Archive Evidence
datasets/PXD023542/PXD023542.sdrf.tsv[27-27]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A run explicitly identified as BS3 is annotated with the DSS cross-linker ontology term.

## Fix Focus Areas
- datasets/PXD023542/PXD023542.sdrf.tsv[27-27]

## Recommended Fix
Replace the DSS cross-linker value on this row with the appropriate BS3 ontology term, after confirming it against the archive metadata.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


8. Search bundles inflate the run count 🐞 Bug ≡ Correctness ⭐ New
Description
PXD021809.sdrf.tsv treats andromeda.7z and txt.7z as timsTOF assays with fraction identifiers
55 and 56. The preceding records are run-specific .d.7z acquisitions, so including these generic
result archives incorrectly extends the acquisition series.
Code

datasets/PXD021809/PXD021809.sdrf.tsv[R56-57]

+PXD021809-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	andromeda.7z	proteomic profiling by mass spectrometry	andromeda.7z	1	55	AC=MS:1002038;NT=label free sample	NT=timsTOF Pro;AC=MS:1003005	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=TurboID;AC=XLMOD:02251	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
+PXD021809-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	txt.7z	proteomic profiling by mass spectrometry	txt.7z	1	56	AC=MS:1002038;NT=label free sample	NT=timsTOF Pro;AC=MS:1003005	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=TurboID;AC=XLMOD:02251	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
Evidence
Lines 2–55 contain named Bruker acquisition archives, whereas lines 56–57 assign identical
acquisition metadata and new fractions to Andromeda and text-output archives.

datasets/PXD021809/PXD021809.sdrf.tsv[52-57]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Two generic search-result archives are annotated as timsTOF acquisitions and add spurious fractions to the dataset.

## Fix Focus Areas
- datasets/PXD021809/PXD021809.sdrf.tsv[56-57]

## Recommended Fix
Delete the `andromeda.7z` and `txt.7z` assay rows, leaving the run-specific `.d.7z` acquisition mappings intact.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


9. Peptide results become an assay 📘 Rule violation ≡ Correctness ⭐ New
Description
Row 6 of PXD026037.sdrf.tsv assigns the post-acquisition peptide-analysis result
iprophet-xl.pep.xml as both an assay and data file with Q Exactive acquisition metadata and
fraction identifier 5. Because rows 2–5 already map the actual same-stem .mzXML or .raw
acquisition files, this result row creates a fifth fraction that the instrument did not acquire.
Code

datasets/PXD026037/PXD026037.sdrf.tsv[6]

+PXD026037-sample	equus caballus	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	iprophet-xl.pep.xml	proteomic profiling by mass spectrometry	iprophet-xl.pep.xml	1	5	AC=MS:1002038;NT=label free sample	NT=Q Exactive;AC=MS:1001911	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
Rows 2–5 map the same-stem mass-spectrometry .mzXML and .raw files, while row 6 maps an
explicitly named iProphet .pep.xml peptide result and assigns it the same instrument and
acquisition fields plus a new fraction identifier. This shows that the assay mapping does not agree
with the archive evidence and treats a post-acquisition result as an additional instrument run.

AGENTS.md: Align SDRF Metadata with Archive Evidence
datasets/PXD026037/PXD026037.sdrf.tsv[2-6]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
An iProphet peptide-analysis XML result is represented as a Q Exactive acquisition and an independent fraction alongside the actual raw and mzXML acquisition files.

## Fix Focus Areas
- datasets/PXD026037/PXD026037.sdrf.tsv[6-6]

## Recommended Fix
Delete the `iprophet-xl.pep.xml` row from the SDRF and retain only mappings for genuine acquired or converted mass-spectrometry data files.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


10. Search archives become assays 📘 Rule violation ≡ Correctness ⭐ New
Description
Rows 14–15 of PXD021770.sdrf.tsv assign andromeda.7z and txt.7z assay names, data files, and
instrument, acquisition, digestion, and fraction metadata as though they were mass-spectrometry
runs. These search-output archives follow the twelve genuine Bruker .d.7z timsTOF acquisitions in
rows 2–13, creating two nonexistent experimental fractions.
Code

datasets/PXD021770/PXD021770.sdrf.tsv[R14-15]

+PXD021770-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	andromeda.7z	proteomic profiling by mass spectrometry	andromeda.7z	1	13	AC=MS:1002038;NT=label free sample	NT=timsTOF Pro;AC=MS:1003005	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=TurboID;AC=XLMOD:02251	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
+PXD021770-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	txt.7z	proteomic profiling by mass spectrometry	txt.7z	1	14	AC=MS:1002038;NT=label free sample	NT=timsTOF Pro;AC=MS:1003005	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=TurboID;AC=XLMOD:02251	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
Evidence
Rule 4 requires file mappings and relationships to agree with archive evidence: lines 2–13 map
run-specific Bruker .d.7z files as the dataset's instrument acquisitions, while lines 14–15 switch
to generic Andromeda and text-result archives but assign them the same acquisition metadata and new
fraction identifiers.

AGENTS.md: Align SDRF Metadata with Archive Evidence
datasets/PXD021770/PXD021770.sdrf.tsv[2-15]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description

`andromeda.7z` and `txt.7z` are search-output archives but are represented as mass-spectrometry assays and distinct fractions, creating records for files that were not acquired by the instrument.

## Fix Focus Areas

- datasets/PXD021770/PXD021770.sdrf.tsv[14-15]

## Recommended Fix

Remove the `andromeda.7z` and `txt.7z` rows from the SDRF, retaining only rows that map genuine mass-spectrometry instrument acquisitions.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


11. Checksum manifest becomes an assay 🐞 Bug ≡ Correctness
Description
PXD024160.sdrf.tsv represents checksum.txt as a mass-spectrometry assay with a Q Exactive
instrument and data-dependent acquisition. This adds a nonexistent ninth experimental fraction to
the dataset.
Code

datasets/PXD024160/PXD024160.sdrf.tsv[10]

+PXD024160-sample	saccharomyces cerevisiae	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	checksum.txt	proteomic profiling by mass spectrometry	checksum.txt	1	9	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF;AC=MS:1002523	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=invertebrates;VV=v1.1.0
Evidence
The row assigns checksum.txt mass-spectrometry technology, a Q Exactive instrument, data-dependent
acquisition, digestion, modifications, and fraction identifier 9.

datasets/PXD024160/PXD024160.sdrf.tsv[10-10]
CONTRIBUTING.md[63-70]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A checksum manifest is represented as a proteomics acquisition and assigned fraction nine.

## Fix Focus Areas
- datasets/PXD024160/PXD024160.sdrf.tsv[10-10]

## Recommended Fix
Delete the `checksum.txt` row and retain only genuine assay or supported assay-data records.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


12. DSBU runs are labeled as DSSO 📘 Rule violation ≡ Correctness
Description
Rows 2–3 and 11 onward in PXD022772.sdrf.tsv identify DSBU in assay, raw-acquisition, and
derived-result filenames but assign the DSSO ontology term in comment[cross-linker]. This conflict
affects every DSBU record while adjacent DSSO-named rows use the same DSSO term consistently, so
consumers cannot reliably distinguish the two cross-linking chemistries.
Code

datasets/PXD022772/PXD022772.sdrf.tsv[R2-3]

+PXD022772-sample	streptococcus pyogenes abc020006030	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	2020-06-05_RSLC8_capLC_XL-PASEF_Stepped-15per_CCSMR_polygon_DSBU_RH11_1_2957.d.zip	proteomic profiling by mass spectrometry	2020-06-05_RSLC8_capLC_XL-PASEF_Stepped-15per_CCSMR_polygon_DSBU_RH11_1_2957.d.zip	1	1	AC=MS:1002038;NT=label free sample	NT=timsTOF Pro;AC=MS:1003005	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
+PXD022772-sample	streptococcus pyogenes abc020006030	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	2020-06-05_RSLC8_capLC_XL-PASEF_Stepped-15per_CCSMR_polygon_DSBU_RH11_2_2958.d.zip	proteomic profiling by mass spectrometry	2020-06-05_RSLC8_capLC_XL-PASEF_Stepped-15per_CCSMR_polygon_DSBU_RH11_2_2958.d.zip	1	2	AC=MS:1002038;NT=label free sample	NT=timsTOF Pro;AC=MS:1003005	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
Lines 2–3 and 11–15 explicitly contain DSBU in their filenames while declaring
NT=DSSO;AC=XLMOD:02010; adjacent DSSO-named rows consistently use that same DSSO term. This
mismatch shows that the file mappings and chemistry metadata do not agree with the archive evidence.

AGENTS.md: Align SDRF Metadata with Public Archive Evidence
datasets/PXD022772/PXD022772.sdrf.tsv[2-3]
datasets/PXD022772/PXD022772.sdrf.tsv[11-13]
datasets/PXD022772/PXD022772.sdrf.tsv[2-4]
datasets/PXD022772/PXD022772.sdrf.tsv[11-15]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
DSBU-named assays, raw acquisitions, and derived result files are annotated with the DSSO cross-linker term, creating contradictory chemistry metadata.

## Fix Focus Areas
- datasets/PXD022772/PXD022772.sdrf.tsv[2-3]
- datasets/PXD022772/PXD022772.sdrf.tsv[11-17]

## Recommended Fix
Assign the appropriate DSBU controlled term, accession, and parameters to every DSBU assay and derived file while retaining DSSO only for DSSO-named records.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


13. Non-cross-linked controls claim DSS 📘 Rule violation ≡ Correctness
Description
Rows explicitly named NoXL in PXD023542.sdrf.tsv and PXD025099.sdrf.tsv still declare a
chemical cross-linking experiment and the DSS cross-linker. This contradiction affects multiple
control runs across both datasets, causing structured analyses to treat them like genuinely
cross-linked DSS samples.
Code

datasets/PXD023542/PXD023542.sdrf.tsv[8]

+PXD023542-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	2020_08_19_04_NSP2_24h_NoXL_Strep_p2	proteomic profiling by mass spectrometry	2020_08_19_04_NSP2_24h_NoXL_Strep_p2.raw	1	7	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF;AC=MS:1002523	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=DSS;AC=XLMOD:02001	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
Evidence
The cited assay and raw-file names identify the affected rows as NoXL, but their metadata declares
both the chemical-crosslinking method and NT=DSS;AC=XLMOD:02001; specifically, this appears on
lines 10, 21, 24, and 25 of PXD025099.sdrf.tsv, with additional affected rows in
PXD023542.sdrf.tsv. Thus, the sample relationships and mappings do not agree with the public file
evidence.

AGENTS.md: Align SDRF Metadata with Public Archive Evidence
datasets/PXD023542/PXD023542.sdrf.tsv[8-8]
datasets/PXD023542/PXD023542.sdrf.tsv[18-24]
datasets/PXD025099/PXD025099.sdrf.tsv[10-10]
datasets/PXD025099/PXD025099.sdrf.tsv[21-25]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Multiple runs explicitly named `NoXL` are annotated as DSS chemical-crosslinking experiments despite their names identifying them as non-cross-linked controls.

## Fix Focus Areas
- datasets/PXD023542/PXD023542.sdrf.tsv[8-8]
- datasets/PXD023542/PXD023542.sdrf.tsv[18-24]
- datasets/PXD025099/PXD025099.sdrf.tsv[10-10]
- datasets/PXD025099/PXD025099.sdrf.tsv[21-25]

## Recommended Fix
Represent each `NoXL` row as a non-cross-linked control using the repository-supported values for the cross-linking method and cross-linker fields, and retain DSS only on actual DSS runs.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


14. Final template column lacks heading 🐞 Bug ≡ Correctness
Description
The header in PXD023525.sdrf.tsv declares two template columns, but every data row supplies three
template values. Parsers therefore receive more fields than headings and cannot reliably associate
the final crosslinking template with a column.
Code

datasets/PXD023525/PXD023525.sdrf.tsv[1]

+source name	characteristics[organism]	characteristics[organism part]	characteristics[cell type]	characteristics[disease]	characteristics[biological replicate]	characteristics[material type]	characteristics[sample type]	characteristics[enrichment process]	characteristics[crosslink distance]	characteristics[crosslinking reaction time]	characteristics[crosslinking temperature]	assay name	technology type	comment[data file]	comment[technical replicate]	comment[fraction identifier]	comment[label]	comment[instrument]	comment[proteomics data acquisition method]	comment[cleavage agent details]	comment[cleavage agent details]	comment[modification parameters]	comment[modification parameters]	comment[dissociation method]	comment[collision energy]	comment[precursor mass tolerance]	comment[fragment mass tolerance]	comment[fractionation method]	comment[chemical cross-linking coupled with ms]	comment[cross-linker]	comment[crosslink enrichment method]	comment[crosslinker concentration]	comment[quenching reagent]	comment[reduction reagent]	comment[alkylation reagent]	comment[sdrf version]	comment[sdrf template]	comment[sdrf template]
Evidence
The header ends with only two template headings, while each row ends with the three values for
versioned mass-spectrometry, crosslinking, and associated template metadata.

datasets/PXD023525/PXD023525.sdrf.tsv[1-2]
datasets/PXD023542/PXD023542.sdrf.tsv[1-2]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The header has one fewer `comment[sdrf template]` column than every data row, leaving the final crosslinking template without a heading.

## Fix Focus Areas
- datasets/PXD023525/PXD023525.sdrf.tsv[1-7]

## Recommended Fix
Append a third `comment[sdrf template]` heading so the header and all data rows have identical field counts.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


15. One raw-file mapping has two suffixes 📘 Rule violation ≡ Correctness
Description
The assay 2020_02_06_13_DSS_CCT_IP maps to the malformed data-file value
2020_02_06_13_DSS_CCT_IP.raw.raw in comment[data file]. Because the extensionless assay stem and
surrounding acquisition mappings use a single .raw suffix, exact lookup against the corresponding
deposited archive file will fail for this row.
Code

datasets/PXD025099/PXD025099.sdrf.tsv[17]

+PXD025099-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	2020_02_06_13_DSS_CCT_IP	proteomic profiling by mass spectrometry	2020_02_06_13_DSS_CCT_IP.raw.raw	1	16	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF;AC=MS:1002523	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSS;AC=XLMOD:02001	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
Evidence
The cited row maps the extensionless assay name 2020_02_06_13_DSS_CCT_IP to a data-file name with
a duplicated .raw.raw suffix, while adjacent acquisition mappings consistently append .raw only
once. This inconsistency shows that the row does not agree with the expected archive naming
evidence.

AGENTS.md: Align SDRF Metadata with Public Archive Evidence
datasets/PXD025099/PXD025099.sdrf.tsv[16-18]
datasets/PXD025099/PXD025099.sdrf.tsv[17-18]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
One assay maps to a data-file name ending in `.raw.raw`, while its extensionless assay stem and neighboring records indicate a single `.raw` suffix, preventing an exact match to the deposited archive filename.

## Fix Focus Areas
- datasets/PXD025099/PXD025099.sdrf.tsv[17-17]

## Recommended Fix
Verify the exact deposited archive filename, then update the data-file value to match it, removing the duplicated `.raw` suffix if the archive confirms the expected single `.raw` extension.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


16. Glu-C runs claim Lys-C 🐞 Bug ≡ Correctness
Description
Glu-C-named runs in PXD021870.sdrf.tsv assign their second cleavage-agent column to Lys-C rather
than Glu-C. Searches or processing driven by the structured digestion metadata will therefore use an
enzyme that contradicts the run identifiers.
Code

datasets/PXD021870/PXD021870.sdrf.tsv[2]

+PXD021870-sample	escherichia coli	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	SurApaf105_OmpA_Trypsin_GluC_1	proteomic profiling by mass spectrometry	SurApaf105_OmpA_Trypsin_GluC_1.raw	1	1	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF-X;AC=MS:1002877	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
The assay and raw filename on line 2 contain Trypsin_GluC, but the two cleavage fields contain
Trypsin and Lys-C; the pattern recurs on numerous Glu-C rows.

datasets/PXD021870/PXD021870.sdrf.tsv[1-2]
datasets/PXD021870/PXD021870.sdrf.tsv[6-6]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Runs explicitly identifying Trypsin and Glu-C digestion are annotated as Trypsin and Lys-C.

## Fix Focus Areas
- datasets/PXD021870/PXD021870.sdrf.tsv[2-49]

## Recommended Fix
For each `_GluC_` run, replace the erroneous Lys-C cleavage-agent value with the controlled Glu-C term while preserving Trypsin.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


17. CDI runs are labeled as DSBU 📘 Rule violation ≡ Correctness
Description
The E1oE2oCDI result and raw-file rows in PXD023525.sdrf.tsv assign NT=DSBU;AC=XLMOD:02043
rather than the CDI chemistry named by both files. The neighboring DSBU-specific rows use the same
term, collapsing two explicitly distinct experimental groups into one cross-linker annotation.
Code

datasets/PXD023525/PXD023525.sdrf.tsv[R4-5]

+PXD023525-sample	escherichia coli	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	E1oE2oCDI.mzid.gz	proteomic profiling by mass spectrometry	E1oE2oCDI.mzid.gz	1	3	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF;AC=MS:1002523	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
+PXD023525-sample	escherichia coli	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	E1oE2oCDI	proteomic profiling by mass spectrometry	E1oE2oCDI.raw	1	4	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF;AC=MS:1002523	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
Rule 4 requires metadata mappings to agree with archive file evidence. Both cited rows explicitly
contain CDI in their assay and file names but declare DSBU as the cross-linker.

AGENTS.md: Align SDRF Metadata with Public Archive Evidence
datasets/PXD023525/PXD023525.sdrf.tsv[2-7]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The CDI-named result and raw files are annotated with the DSBU ontology term, contradicting their file mappings.

## Fix Focus Areas
- datasets/PXD023525/PXD023525.sdrf.tsv[4-5]

## Recommended Fix
Replace DSBU on the two CDI rows with the appropriate CDI cross-linker ontology term and parameters, preserving DSBU on the separately named DSBU rows.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context sources
Review mode: 🧠 Deep: This adds 50 independent SDRF datasets with substantial structured metadata and many previously identified, easy-to-miss annotation defects, making redundant review materially valuable.

Grey Divider

Tip of the day
💡 Did you know, you can enable the Remediation agent and Qodo fixes findings in a dedicated fix PR

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread datasets/PXD021831/PXD021831.sdrf.tsv Outdated
Comment on lines +2 to +3
PXD022772-sample streptococcus pyogenes abc020006030 not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available 2020-06-05_RSLC8_capLC_XL-PASEF_Stepped-15per_CCSMR_polygon_DSBU_RH11_1_2957.d.zip proteomic profiling by mass spectrometry 2020-06-05_RSLC8_capLC_XL-PASEF_Stepped-15per_CCSMR_polygon_DSBU_RH11_1_2957.d.zip 1 1 AC=MS:1002038;NT=label free sample NT=timsTOF Pro;AC=MS:1003005 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD022772-sample streptococcus pyogenes abc020006030 not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available 2020-06-05_RSLC8_capLC_XL-PASEF_Stepped-15per_CCSMR_polygon_DSBU_RH11_2_2958.d.zip proteomic profiling by mass spectrometry 2020-06-05_RSLC8_capLC_XL-PASEF_Stepped-15per_CCSMR_polygon_DSBU_RH11_2_2958.d.zip 1 2 AC=MS:1002038;NT=label free sample NT=timsTOF Pro;AC=MS:1003005 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

2. Dsbu runs are labeled as dsso 📘 Rule violation ≡ Correctness

Rows 2–3 and 11 onward in PXD022772.sdrf.tsv identify DSBU in assay, raw-acquisition, and
derived-result filenames but assign the DSSO ontology term in comment[cross-linker]. This conflict
affects every DSBU record while adjacent DSSO-named rows use the same DSSO term consistently, so
consumers cannot reliably distinguish the two cross-linking chemistries.
Agent Prompt
## Issue description
DSBU-named assays, raw acquisitions, and derived result files are annotated with the DSSO cross-linker term, creating contradictory chemistry metadata.

## Fix Focus Areas
- datasets/PXD022772/PXD022772.sdrf.tsv[2-3]
- datasets/PXD022772/PXD022772.sdrf.tsv[11-17]

## Recommended Fix
Assign the appropriate DSBU controlled term, accession, and parameters to every DSBU assay and derived file while retaining DSSO only for DSSO-named records.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment thread datasets/PXD022861/PXD022861.sdrf.tsv Outdated
@@ -0,0 +1,14 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[age] characteristics[sex] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD022861-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available 30 Å not available not available PRC2-BS3 proteomic profiling by mass spectrometry PRC2-BS3.raw 1 1 AC=MS:1002038;NT=label free sample NT=TripleTOF 5600;AC=MS:1000932 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=BS3;AC=XLMOD:02000 not available not available conditions not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD022861-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available 30 Å not available not available PRC2-DSS proteomic profiling by mass spectrometry PRC2-DSS.raw 1 2 AC=MS:1002038;NT=label free sample NT=TripleTOF 5600;AC=MS:1000932 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=BS3;AC=XLMOD:02000 not available not available conditions not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

3. A dss assay is labeled as bs3 📘 Rule violation ≡ Correctness

The PRC2-DSS assay and PRC2-DSS.raw file are assigned the BS3 term NT=BS3;AC=XLMOD:02000 in
comment[cross-linker]. This occurs beside the correctly named BS3 run, making the distinct DSS and
BS3 acquisitions indistinguishable in the structured metadata despite their explicit filenames.
Agent Prompt
## Issue description
The run explicitly named `PRC2-DSS`, including its `PRC2-DSS.raw` data file, is annotated with the BS3 controlled term, making the DSS and BS3 acquisitions indistinguishable in the structured cross-linker metadata.

## Fix Focus Areas
- datasets/PXD022861/PXD022861.sdrf.tsv[2-3]

## Recommended Fix
Replace `NT=BS3;AC=XLMOD:02000` on the `PRC2-DSS` row with the appropriate controlled DSS cross-linker term, accession, and parameters.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment on lines +4 to +5
PXD023525-sample escherichia coli not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available E1oE2oCDI.mzid.gz proteomic profiling by mass spectrometry E1oE2oCDI.mzid.gz 1 3 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD023525-sample escherichia coli not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available E1oE2oCDI proteomic profiling by mass spectrometry E1oE2oCDI.raw 1 4 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

4. Cdi runs are labeled as dsbu 📘 Rule violation ≡ Correctness

The E1oE2oCDI result and raw-file rows in PXD023525.sdrf.tsv assign NT=DSBU;AC=XLMOD:02043
rather than the CDI chemistry named by both files. The neighboring DSBU-specific rows use the same
term, collapsing two explicitly distinct experimental groups into one cross-linker annotation.
Agent Prompt
## Issue description
The CDI-named result and raw files are annotated with the DSBU ontology term, contradicting their file mappings.

## Fix Focus Areas
- datasets/PXD023525/PXD023525.sdrf.tsv[4-5]

## Recommended Fix
Replace DSBU on the two CDI rows with the appropriate CDI cross-linker ontology term and parameters, preserving DSBU on the separately named DSBU rows.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

PXD023542-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 2020_06_29_01_C-143_Strep_IP_No_Vec_p2 proteomic profiling by mass spectrometry 2020_06_29_01_C-143_Strep_IP_No_Vec_p2.raw 1 4 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=DSS;AC=XLMOD:02001 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023542-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 2020_06_29_03_C-143_Strep_IP_No_XL_p2 proteomic profiling by mass spectrometry 2020_06_29_03_C-143_Strep_IP_No_XL_p2.raw 1 5 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=DSS;AC=XLMOD:02001 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023542-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 2020_06_29_07_C-143_Strep_IP_DSS_p2 proteomic profiling by mass spectrometry 2020_06_29_07_C-143_Strep_IP_DSS_p2.raw 1 6 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=DSS;AC=XLMOD:02001 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023542-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 2020_08_19_04_NSP2_24h_NoXL_Strep_p2 proteomic profiling by mass spectrometry 2020_08_19_04_NSP2_24h_NoXL_Strep_p2.raw 1 7 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=DSS;AC=XLMOD:02001 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

5. Non-cross-linked controls claim dss 📘 Rule violation ≡ Correctness

Rows explicitly named NoXL in PXD023542.sdrf.tsv and PXD025099.sdrf.tsv still declare a
chemical cross-linking experiment and the DSS cross-linker. This contradiction affects multiple
control runs across both datasets, causing structured analyses to treat them like genuinely
cross-linked DSS samples.
Agent Prompt
## Issue description
Multiple runs explicitly named `NoXL` are annotated as DSS chemical-crosslinking experiments despite their names identifying them as non-cross-linked controls.

## Fix Focus Areas
- datasets/PXD023542/PXD023542.sdrf.tsv[8-8]
- datasets/PXD023542/PXD023542.sdrf.tsv[18-24]
- datasets/PXD025099/PXD025099.sdrf.tsv[10-10]
- datasets/PXD025099/PXD025099.sdrf.tsv[21-25]

## Recommended Fix
Represent each `NoXL` row as a non-cross-linked control using the repository-supported values for the cross-linking method and cross-linker fields, and retain DSS only on actual DSS runs.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment thread datasets/PXD022440/PXD022440.sdrf.tsv Outdated
@@ -0,0 +1,14 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD022440-sample xenopus laevis not applicable not applicable not applicable 1 synthetic reference not available 2 Å not available not available IP_15min_Plk1.wiff proteomic profiling by mass spectrometry IP_15min_Plk1.wiff 1 1 AC=MS:1002038;NT=label free sample NT=TripleTOF 4600;AC=MS:1002583 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=formaldehyde;AC=XLMOD:02006 not available not available by not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=vertebrates;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

16. Quenching reagent is stray prose 🐞 Bug ≡ Correctness

Every row in PXD022440.sdrf.tsv stores the bare word by in comment[quenching reagent]. The
field consequently supplies neither a reagent annotation nor an unavailable sentinel.
Agent Prompt
## Issue description
The quenching-reagent column contains `by`, which does not identify a reagent.

## Fix Focus Areas
- datasets/PXD022440/PXD022440.sdrf.tsv[2-14]

## Recommended Fix
Replace `by` with the source-backed quenching reagent or `not available`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment thread datasets/PXD022690/PXD022690.sdrf.tsv Outdated
@@ -0,0 +1,30 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD022690-sample saccharomyces cerevisiae not applicable not applicable not applicable 1 synthetic reference not available 18 Å not available not available OrbitrapLUMOS_20200215_02 proteomic profiling by mass spectrometry OrbitrapLUMOS_20200215_02.raw 1 1 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=SDA;AC=XLMOD:02171 not available not available the not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=invertebrates;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

17. Quenching reagent is stray prose 🐞 Bug ≡ Correctness

Every row in PXD022690.sdrf.tsv stores the bare word the in comment[quenching reagent]. This
fragment does not identify a reagent and corrupts the structured reaction metadata.
Agent Prompt
## Issue description
The quenching-reagent column contains the prose fragment `the` instead of a reagent.

## Fix Focus Areas
- datasets/PXD022690/PXD022690.sdrf.tsv[2-30]

## Recommended Fix
Replace `the` with the verified quenching reagent or `not available`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment thread datasets/PXD022861/PXD022861.sdrf.tsv Outdated
@@ -0,0 +1,14 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[age] characteristics[sex] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD022861-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available 30 Å not available not available PRC2-BS3 proteomic profiling by mass spectrometry PRC2-BS3.raw 1 1 AC=MS:1002038;NT=label free sample NT=TripleTOF 5600;AC=MS:1000932 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=BS3;AC=XLMOD:02000 not available not available conditions not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

18. Quenching reagent is stray prose 🐞 Bug ≡ Correctness

Every row in PXD022861.sdrf.tsv stores the bare word conditions in comment[quenching reagent].
The metadata therefore supplies a sentence fragment rather than a reagent or explicit unknown value.
Agent Prompt
## Issue description
The quenching-reagent column contains `conditions`, which is not a reagent annotation.

## Fix Focus Areas
- datasets/PXD022861/PXD022861.sdrf.tsv[2-14]

## Recommended Fix
Replace `conditions` with the source-backed quenching reagent or `not available`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@@ -0,0 +1,12 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[age] characteristics[sex] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD023221-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available 30 Å not available not available G2TGM201125_02.zip proteomic profiling by mass spectrometry G2TGM201125_02.raw.zip 1 1 AC=MS:1002038;NT=label free sample NT=SYNAPT G2-Si;AC=MS:1002726 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=BS3;AC=XLMOD:02000 not available not available for not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

19. Quenching reagent is stray prose 🐞 Bug ≡ Correctness

Every row in PXD023221.sdrf.tsv stores the bare word for in comment[quenching reagent]. This
preposition cannot identify how the BS3 reaction was quenched.
Agent Prompt
## Issue description
The quenching-reagent column contains the preposition `for` instead of a reagent annotation.

## Fix Focus Areas
- datasets/PXD023221/PXD023221.sdrf.tsv[2-12]

## Recommended Fix
Replace `for` with the verified quenching reagent or `not available`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@@ -0,0 +1,51 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[age] characteristics[sex] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD023814-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available a12900 proteomic profiling by mass spectrometry a12900.raw 1 1 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=APEX;AC=XLMOD:02252 not available not available by not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

20. Quenching reagent is stray prose 🐞 Bug ≡ Correctness

Every row in PXD023814.sdrf.tsv stores the bare word by in comment[quenching reagent]. The
APEX reaction metadata consequently contains an uninterpretable prose fragment instead of a reagent
value.
Agent Prompt
## Issue description
The quenching-reagent column contains `by`, which does not identify a reagent.

## Fix Focus Areas
- datasets/PXD023814/PXD023814.sdrf.tsv[2-51]

## Recommended Fix
Replace `by` with the verified quenching reagent or `not available`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

…lumns

Required by the vertebrates/invertebrates/plants SDRF templates; value set
to the spec-compliant reserved word 'not available' where the field was
not previously populated. human-only files are unaffected (field optional
in that template).
The 2 .raw files (DSS/BS3 crosslinking runs) were tagged with the AB Sciex
TripleTOF 5600 used for the separate HX-MS .wiff runs. Per the PRIDE
submission's sample processing protocol, the crosslinking data was
acquired on an Orbitrap Fusion Lumos (Thermo); the 11 .wiff HX-MS files
are unaffected.
@ypriverol ypriverol closed this Sep 17, 2026
@ypriverol ypriverol reopened this Sep 17, 2026
@github-actions

github-actions Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

SDRF change report

50 new · 0 modified · 0 deleted · highest risk: none

⚠️ Needs attention

Backed by this report's own checks of the SDRF data: high-risk changes and AI reviewer findings the data confirms.

  • PXD021770 · Search archives become assays: non-acquisition data files: txt.7z (qodo-code-review[bot])
  • PXD022440 · Checksum manifest becomes an assay: non-acquisition data files: checksum.txt (qodo-code-review[bot])
  • PXD024160 · Checksum manifest becomes an assay: non-acquisition data files: 20191112_F1_Ag5_Steig002_SA_SMC_E5-_2__Crosslinks.txt, 20191115_F1_Ag5_Steig002_SA_SMC_E1_2_Crosslinks.txt, 20210108_coreQEP2_BaSt_LC12-16_SA_Orbi3748_E_Sample1_5uL_Crossl… (qodo-code-review[bot])
  • PXD025357 · Three result archives become assays: non-acquisition data files: search.zip (qodo-code-review[bot])
  • PXD026037 · Peptide results become an assay: non-acquisition data files: iprophet-xl.pep.xml (qodo-code-review[bot])

External reviewer notes

Quoted from AI review bots on this PR. Not verified by this report unless marked as also flagged.

  • PXD021770 · qodo-code-review[bot]: Search archives become assays. Rows 14–15 of PXD021770.sdrf.tsv assign andromeda.7z and txt.7z assay names, data files, and instrument, acquisition, digestion, and fraction metadata as though they were mass-spectrometry runs. (source) · ✓ confirmed by data: non-acquisition data files: txt.7z
  • PXD022440 · qodo-code-review[bot]: Checksum manifest becomes an assay. PXD022440.sdrf.tsv represents checksum.txt as a mass-spectrometry assay with instrument, acquisition, digestion, and crosslinking metadata. A checksum manifest is not an acquisition, so this creates a nonexistent experimental run. (source) · ✓ confirmed by data: non-acquisition data files: checksum.txt
  • PXD024160 · qodo-code-review[bot]: Checksum manifest becomes an assay. PXD024160.sdrf.tsv represents checksum.txt as a mass-spectrometry assay with a Q Exactive instrument and data-dependent acquisition. This adds a nonexistent ninth experimental fraction to the dataset. (source) · ✓ confirmed by data: non-acquisition data files: 20191112_F1_Ag5_Steig002_SA_SMC_E5-_2__Crosslinks.txt, 20191115_F1_Ag5_Steig002_SA_SMC_E1_2_Crosslinks.txt, 20210108_coreQEP2_BaSt_LC12-16_SA_Orbi3748_E_Sample1_5uL_Crossl…
  • PXD025357 · qodo-code-review[bot]: Three result archives become assays. Rows 2, 9, and 10 of PXD025357.sdrf.tsv assign andromeda.zip , search.zip , and text.zip Q Exactive acquisition metadata and separate fraction identifiers. (source) · ✓ confirmed by data: non-acquisition data files: search.zip
  • PXD026037 · qodo-code-review[bot]: Peptide results become an assay. Row 6 of PXD026037.sdrf.tsv assigns the post-acquisition peptide-analysis result iprophet-xl.pep.xml as both an assay and data file with Q Exactive acquisition metadata and fraction identifier 5. (source) · ✓ confirmed by data: non-acquisition data files: iprophet-xl.pep.xml
  • PXD021809 · qodo-code-review[bot]: Search bundles inflate the run count. PXD021809.sdrf.tsv treats andromeda.7z and txt.7z as timsTOF assays with fraction identifiers 55 and 56. The preceding records are run-specific .d.7z acquisitions, so including these generic result archives incorrectly extends the acquisition series. (source)
  • PXD021822 · qodo-code-review[bot]: Quenching reagent is stray prose. Every row in PXD021822.sdrf.tsv stores the bare word by in comment[quenching reagent] . The field therefore identifies no reagent and cannot support structured interpretation of the reaction protocol. (source)
  • PXD021831, PXD023164, PXD024065, PXD024131, PXD024160, PXD025172, PXD025843 · qodo-code-review[bot]: Yeast samples use an animal template. Seven added SDRFs pair saccharomyces cerevisiae with the invertebrates template instead of a template appropriate for fungi. The mismatch occurs throughout PXD021831, PXD023164, PXD024065, PXD024131, PXD024160, PXD025172, and PXD025843, affecting every sample… (source)
  • PXD021870 · qodo-code-review[bot]: Glu-C runs claim Lys-C. Glu-C-named runs in PXD021870.sdrf.tsv assign their second cleavage-agent column to Lys-C rather than Glu-C. Searches or processing driven by the structured digestion metadata will therefore use an enzyme that contradicts the run identifiers. (source)
  • PXD022440 · qodo-code-review[bot]: Quenching reagent is stray prose. Every row in PXD022440.sdrf.tsv stores the bare word by in comment[quenching reagent] . The field consequently supplies neither a reagent annotation nor an unavailable sentinel. (source)
  • PXD022608 · qodo-code-review[bot]: Tagged runs appear label-free. PXD022608.sdrf.tsv marks every run as a label-free sample even though every assay and data-file name identifies a ten-channel tandem-mass-tag experiment. Consumers using the label column will consequently treat tagged quantification data as label-free. (source)
  • PXD022690 · qodo-code-review[bot]: Quenching reagent is stray prose. Every row in PXD022690.sdrf.tsv stores the bare word the in comment[quenching reagent] . This fragment does not identify a reagent and corrupts the structured reaction metadata. (source)
  • PXD022772 · qodo-code-review[bot]: DSBU runs are labeled as DSSO. Rows 2–3 and 11 onward in PXD022772.sdrf.tsv identify DSBU in assay, raw-acquisition, and derived-result filenames but assign the DSSO ontology term in comment[cross-linker] . (source)
  • PXD022772 · qodo-code-review[bot]: Sequence database becomes an assay. PXD022772.sdrf.tsv represents cas9_crapome.fasta as a mass-spectrometry assay with acquisition and instrument metadata. The sequence database is an analysis input rather than an acquired run, so the row invents an experimental fraction. (source)
  • PXD022785 · qodo-code-review[bot]: Two databases become assays. PXD022785.sdrf.tsv represents two FASTA sequence databases as mass-spectrometry assays. Both rows receive instrument, acquisition, digestion, and fraction metadata despite not being experimental runs. (source)
  • PXD022861 · qodo-code-review[bot]: A dss assay is labeled as bs3. The PRC2-DSS assay and PRC2-DSS.raw file are assigned the BS3 term NT=BS3;AC=XLMOD:02000 in comment[cross-linker] . This occurs beside the correctly named BS3 run, making the distinct DSS and BS3 acquisitions indistinguishable in the structured metadata despi… (source)
  • PXD022861 · qodo-code-review[bot]: Quenching reagent is stray prose. Every row in PXD022861.sdrf.tsv stores the bare word conditions in comment[quenching reagent] . The metadata therefore supplies a sentence fragment rather than a reagent or explicit unknown value. (source)
  • PXD023072 · qodo-code-review[bot]: Protein database becomes an assay. PXD023072.sdrf.tsv represents proteins.fasta as a mass-spectrometry assay. The sequence database is consequently assigned an instrument, acquisition method, digestion, and experimental fraction that it cannot possess. (source)
  • PXD023221 · qodo-code-review[bot]: Quenching reagent is stray prose. Every row in PXD023221.sdrf.tsv stores the bare word for in comment[quenching reagent] . This preposition cannot identify how the BS3 reaction was quenched. (source)
  • PXD023522 · qodo-code-review[bot]: Known chemistry remains unidentified. PXD023522.sdrf.tsv records NT=unknown crosslinker for acquisitions whose assay and data-file names explicitly identify DSG. This affects the DSG10, DSG20, and DSG50 rows, preventing structured consumers from recognizing their cross-linking chemistry. (source)
  • PXD023522 · qodo-code-review[bot]: Isotope-labelled runs appear label-free. PXD023522.sdrf.tsv assigns AC=MS:1002038;NT=label free sample to assays whose names and data files explicitly contain 14N15N . The four affected acquisitions are consequently indistinguishable from genuinely label-free runs for consumers of the structured lab… (source)
  • PXD023525, PXD023542 · qodo-code-review[bot]: Final template column lacks heading. The header in PXD023525.sdrf.tsv declares two template columns, but every data row supplies three template values. Parsers therefore receive more fields than headings and cannot reliably associate the final crosslinking template with a column. (source)
  • PXD023525 · qodo-code-review[bot]: CDI runs are labeled as DSBU. The E1oE2oCDI result and raw-file rows in PXD023525.sdrf.tsv assign NT=DSBU;AC=XLMOD:02043 rather than the CDI chemistry named by both files. (source)
  • PXD023542 · qodo-code-review[bot]: A BS3 run is labeled as DSS. Row 27 maps the 2020_11_12_04_NSP1_strep_BS3_P2 assay and raw file to NT=DSS;AC=XLMOD:02001 . The explicit BS3 filename conflicts with that cross-linker while neighboring DSS-named acquisitions use the DSS term consistently. (source)
  • PXD023542, PXD025099 · qodo-code-review[bot]: Non-cross-linked controls claim DSS. Rows explicitly named NoXL in PXD023542.sdrf.tsv and PXD025099.sdrf.tsv still declare a chemical cross-linking experiment and the DSS cross-linker. (source)
  • PXD023814 · qodo-code-review[bot]: Quenching reagent is stray prose. Every row in PXD023814.sdrf.tsv stores the bare word by in comment[quenching reagent] . The APEX reaction metadata consequently contains an uninterpretable prose fragment instead of a reagent value. (source)
  • PXD024253 · qodo-code-review[bot]: Bovine albumin has wrong species. The BSA-named assays in PXD024253.sdrf.tsv set characteristics[organism] to escherichia coli . Those four BSA acquisitions are therefore indexed as bacterial material while the separate E. coli assay group is explicitly identifiable as Ecoli in its names. (source)
  • PXD024399 · qodo-code-review[bot]: Variants cannot be grouped reliably. PXD024399.sdrf.tsv has no genotype, mutation, construct, or factor column even though its assay names encode WT, S754A, and S754E groups. (source)
  • PXD025066 · qodo-code-review[bot]: Identification results become assays. Rows 2–3 assign two .mzid.gz identification-result files as assay names and data files with Orbitrap acquisition metadata. Genuine .raw acquisitions occur in rows 4–7, so the result files add two artificial experimental fractions. (source)
  • PXD025066 · qodo-code-review[bot]: Digests are labelled as Lys-C. PXD025066.sdrf.tsv records Lys-C as the second cleavage agent for TrypAspN and TrypChymo assays. Each AspN- or chymotrypsin-containing run therefore has digestion metadata that conflicts with its own acquisition identifier. (source)
  • PXD025099 · qodo-code-review[bot]: One raw-file mapping has two suffixes. The assay 2020_02_06_13_DSS_CCT_IP maps to the malformed data-file value 2020_02_06_13_DSS_CCT_IP.raw.raw in comment[data file] . Because the extensionless assay stem and surrounding acquisition mappings use a single .raw suffix, exact lookup against the corr… (source)
New datasets (50)

parse_sdrf validation of new datasets is reported by the SDRF review gate check.

Dataset Rows Defects
PXD021708 4 no_factor_value: 1
PXD021709 8 no_factor_value: 1
PXD021770 14 no_factor_value: 1
PXD021809 56 no_factor_value: 1
PXD021822 45 no_factor_value: 1, peak_list_data_file: 30
PXD021831 79 no_factor_value: 1
PXD021870 48 no_factor_value: 1
PXD021923 24 no_factor_value: 1
PXD022119 1 no_factor_value: 1
PXD022279 25 no_factor_value: 1
PXD022335 52 no_factor_value: 1
PXD022440 13 no_factor_value: 1, peak_list_data_file: 4
PXD022443 6 no_factor_value: 1
PXD022608 20 no_factor_value: 1
PXD022690 29 no_factor_value: 1
PXD022772 16 no_factor_value: 1
PXD022785 10 no_factor_value: 1
PXD022861 13 no_factor_value: 1
PXD022991 24 no_factor_value: 1
PXD023072 55 no_factor_value: 1
PXD023164 6 no_factor_value: 1
PXD023221 11 no_factor_value: 1
PXD023239 63 no_factor_value: 1
PXD023277 40 no_factor_value: 1
PXD023522 10 no_factor_value: 1, peak_list_data_file: 10
PXD023525 6 no_factor_value: 1
PXD023542 27 no_factor_value: 1
PXD023577 18 no_factor_value: 1
PXD023814 50 no_factor_value: 1
PXD024010 29 no_factor_value: 1
PXD024065 4 no_factor_value: 1
PXD024131 4 no_factor_value: 1
PXD024160 9 no_factor_value: 1
PXD024253 7 no_factor_value: 1, peak_list_data_file: 3
PXD024335 6 no_factor_value: 1
PXD024366 48 no_factor_value: 1, peak_list_data_file: 24
PXD024367 24 no_factor_value: 1, peak_list_data_file: 12
PXD024399 60 no_factor_value: 1
PXD024822 24 no_factor_value: 1
PXD024946 14 no_factor_value: 1
PXD025066 6 no_factor_value: 1
PXD025099 30 no_factor_value: 1
PXD025172 20 no_factor_value: 1
PXD025208 10 no_factor_value: 1
PXD025220 2 no_factor_value: 1
PXD025357 9 no_factor_value: 1
PXD025581 15 no_factor_value: 1
PXD025662 16 no_factor_value: 1
PXD025843 22 no_factor_value: 1
PXD026037 5 no_factor_value: 1, peak_list_data_file: 2

Advisory report built from 9973fff. Risk labels do not block merging.

@github-actions github-actions Bot added the sdrf:new SDRF PR adds new datasets label Sep 17, 2026
Comment on lines +14 to +15
PXD021770-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available andromeda.7z proteomic profiling by mass spectrometry andromeda.7z 1 13 AC=MS:1002038;NT=label free sample NT=timsTOF Pro;AC=MS:1003005 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=TurboID;AC=XLMOD:02251 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD021770-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available txt.7z proteomic profiling by mass spectrometry txt.7z 1 14 AC=MS:1002038;NT=label free sample NT=timsTOF Pro;AC=MS:1003005 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=TurboID;AC=XLMOD:02251 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

1. Search archives become assays 📘 Rule violation ≡ Correctness

Rows 14–15 of PXD021770.sdrf.tsv assign andromeda.7z and txt.7z assay names, data files, and
instrument, acquisition, digestion, and fraction metadata as though they were mass-spectrometry
runs. These search-output archives follow the twelve genuine Bruker .d.7z timsTOF acquisitions in
rows 2–13, creating two nonexistent experimental fractions.
Agent Prompt
## Issue description

`andromeda.7z` and `txt.7z` are search-output archives but are represented as mass-spectrometry assays and distinct fractions, creating records for files that were not acquired by the instrument.

## Fix Focus Areas

- datasets/PXD021770/PXD021770.sdrf.tsv[14-15]

## Recommended Fix

Remove the `andromeda.7z` and `txt.7z` rows from the SDRF, retaining only rows that map genuine mass-spectrometry instrument acquisitions.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

PXD023542-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 2020_10_29_04_NoVec_NoXL_Exp_22_10_20_MS2 proteomic profiling by mass spectrometry 2020_10_29_04_NoVec_NoXL_Exp_22_10_20_MS2.raw 1 23 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=DSS;AC=XLMOD:02001 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023542-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 2020_10_29_05_NoVec_DSS_Exp_22_10_20_MS2 proteomic profiling by mass spectrometry 2020_10_29_05_NoVec_DSS_Exp_22_10_20_MS2.raw 1 24 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=DSS;AC=XLMOD:02001 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023542-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 2020_11_04_03_NSP1_DSS_strep_P2 proteomic profiling by mass spectrometry 2020_11_04_03_NSP1_DSS_strep_P2.raw 1 25 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=DSS;AC=XLMOD:02001 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023542-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 2020_11_12_04_NSP1_strep_BS3_P2 proteomic profiling by mass spectrometry 2020_11_12_04_NSP1_strep_BS3_P2.raw 1 26 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=DSS;AC=XLMOD:02001 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

2. A bs3 run is labeled as dss 📘 Rule violation ≡ Correctness

Row 27 maps the 2020_11_12_04_NSP1_strep_BS3_P2 assay and raw file to NT=DSS;AC=XLMOD:02001. The
explicit BS3 filename conflicts with that cross-linker while neighboring DSS-named acquisitions
use the DSS term consistently.
Agent Prompt
## Issue description
A run explicitly identified as BS3 is annotated with the DSS cross-linker ontology term.

## Fix Focus Areas
- datasets/PXD023542/PXD023542.sdrf.tsv[27-27]

## Recommended Fix
Replace the DSS cross-linker value on this row with the appropriate BS3 ontology term, after confirming it against the archive metadata.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment on lines +2 to +3
PXD025066-sample oryctolagus sp. 'rabbit_od' not applicable not applicable not applicable 1 synthetic reference not available not available not available not available 3690_LM_C4H2O2_TrypAspN_Xi1.7.6.1.mzid.gz proteomic profiling by mass spectrometry 3690_LM_C4H2O2_TrypAspN_Xi1.7.6.1.mzid.gz 1 1 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD025066-sample oryctolagus sp. 'rabbit_od' not applicable not applicable not applicable 1 synthetic reference not available not available not available not available 3690_LM_C4H2O2_TrypChymo_Xi1.7.6.1.mzid.gz proteomic profiling by mass spectrometry 3690_LM_C4H2O2_TrypChymo_Xi1.7.6.1.mzid.gz 1 2 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

3. Identification results become assays 📘 Rule violation ≡ Correctness

Rows 2–3 assign two .mzid.gz identification-result files as assay names and data files with
Orbitrap acquisition metadata. Genuine .raw acquisitions occur in rows 4–7, so the result files
add two artificial experimental fractions.
Agent Prompt
## Issue description
Compressed mzIdentML identification results are modeled as mass-spectrometry acquisitions even though corresponding raw acquisitions are listed separately.

## Fix Focus Areas
- datasets/PXD025066/PXD025066.sdrf.tsv[2-3]

## Recommended Fix
Remove the `.mzid.gz` rows and retain only genuine acquisition files as SDRF assays.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment on lines +9 to +10
PXD025357-sample trypanosoma brucei not applicable not applicable not applicable 1 synthetic reference not available not available not available not available search.zip proteomic profiling by mass spectrometry search.zip 1 8 AC=MS:1002038;NT=label free sample NT=Q Exactive;AC=MS:1001911 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=BioID;AC=XLMOD:02250 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD025357-sample trypanosoma brucei not applicable not applicable not applicable 1 synthetic reference not available not available not available not available text.zip proteomic profiling by mass spectrometry text.zip 1 9 AC=MS:1002038;NT=label free sample NT=Q Exactive;AC=MS:1001911 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=BioID;AC=XLMOD:02250 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

4. Three result archives become assays 📘 Rule violation ≡ Correctness

Rows 2, 9, and 10 of PXD025357.sdrf.tsv assign andromeda.zip, search.zip, and text.zip Q
Exactive acquisition metadata and separate fraction identifiers. Because the six intervening .raw
records are the actual instrument runs, treating these packaged search outputs as acquisitions adds
three artificial assays.
Agent Prompt
## Issue description
Three packaged search-result archives are assigned instrument and acquisition metadata as independent experimental assays alongside the dataset's actual raw files.

## Fix Focus Areas
- datasets/PXD025357/PXD025357.sdrf.tsv[2-10]

## Recommended Fix
Remove the `andromeda.zip`, `search.zip`, and `text.zip` assay rows, retaining only the six `.raw` acquisition mappings.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

PXD026037-sample equus caballus not applicable not applicable not applicable 1 synthetic reference not available not available not available not available 111820_6plex_protein_mix_sample_1_1 proteomic profiling by mass spectrometry 111820_6plex_protein_mix_sample_1_1.raw 1 2 AC=MS:1002038;NT=label free sample NT=Q Exactive;AC=MS:1001911 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD026037-sample equus caballus not applicable not applicable not applicable 1 synthetic reference not available not available not available not available 111829_6plex_protein_mix_sample_2_1.mzXML proteomic profiling by mass spectrometry 111829_6plex_protein_mix_sample_2_1.mzXML 1 3 AC=MS:1002038;NT=label free sample NT=Q Exactive;AC=MS:1001911 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD026037-sample equus caballus not applicable not applicable not applicable 1 synthetic reference not available not available not available not available 111829_6plex_protein_mix_sample_2_1 proteomic profiling by mass spectrometry 111829_6plex_protein_mix_sample_2_1.raw 1 4 AC=MS:1002038;NT=label free sample NT=Q Exactive;AC=MS:1001911 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD026037-sample equus caballus not applicable not applicable not applicable 1 synthetic reference not available not available not available not available iprophet-xl.pep.xml proteomic profiling by mass spectrometry iprophet-xl.pep.xml 1 5 AC=MS:1002038;NT=label free sample NT=Q Exactive;AC=MS:1001911 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

5. Peptide results become an assay 📘 Rule violation ≡ Correctness

Row 6 of PXD026037.sdrf.tsv assigns the post-acquisition peptide-analysis result
iprophet-xl.pep.xml as both an assay and data file with Q Exactive acquisition metadata and
fraction identifier 5. Because rows 2–5 already map the actual same-stem .mzXML or .raw
acquisition files, this result row creates a fifth fraction that the instrument did not acquire.
Agent Prompt
## Issue description
An iProphet peptide-analysis XML result is represented as a Q Exactive acquisition and an independent fraction alongside the actual raw and mzXML acquisition files.

## Fix Focus Areas
- datasets/PXD026037/PXD026037.sdrf.tsv[6-6]

## Recommended Fix
Delete the `iprophet-xl.pep.xml` row from the SDRF and retain only mappings for genuine acquired or converted mass-spectrometry data files.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment on lines +4 to +7
PXD023522-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available LEDG_DSG10_11580.mzXML proteomic profiling by mass spectrometry LEDG_DSG10_11580.mzXML 1 3 AC=MS:1002038;NT=label free sample NT=ultraflex;AC=MS:1000201 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023522-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available LEDG_DSG20_11581.mzXML proteomic profiling by mass spectrometry LEDG_DSG20_11581.mzXML 1 4 AC=MS:1002038;NT=label free sample NT=ultraflex;AC=MS:1000201 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023522-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available Nkrp1B_CE_14N15N_DSG20_4007.mzXML proteomic profiling by mass spectrometry Nkrp1B_CE_14N15N_DSG20_4007.mzXML 1 5 AC=MS:1002038;NT=label free sample NT=ultraflex;AC=MS:1000201 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023522-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available Nkrp1B_CE_14N15N_DSG50_4008.mzXML proteomic profiling by mass spectrometry Nkrp1B_CE_14N15N_DSG50_4008.mzXML 1 6 AC=MS:1002038;NT=label free sample NT=ultraflex;AC=MS:1000201 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

7. Known chemistry remains unidentified 🐞 Bug ≡ Correctness

PXD023522.sdrf.tsv records NT=unknown crosslinker for acquisitions whose assay and data-file
names explicitly identify DSG. This affects the DSG10, DSG20, and DSG50 rows, preventing structured
consumers from recognizing their cross-linking chemistry.
Agent Prompt
## Issue description
Multiple acquisition names explicitly identify DSG, but their structured cross-linker field says the chemistry is unknown.

## Fix Focus Areas
- datasets/PXD023522/PXD023522.sdrf.tsv[4-10]

## Recommended Fix
Replace `NT=unknown crosslinker;AC=XLMOD:00000` with the appropriate controlled DSG cross-linker term on every DSG-named row; review control rows separately rather than applying the replacement globally.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

PXD023522-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available LEDG_Ctrl.mzXML proteomic profiling by mass spectrometry LEDG_Ctrl.mzXML 1 2 AC=MS:1002038;NT=label free sample NT=ultraflex;AC=MS:1000201 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023522-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available LEDG_DSG10_11580.mzXML proteomic profiling by mass spectrometry LEDG_DSG10_11580.mzXML 1 3 AC=MS:1002038;NT=label free sample NT=ultraflex;AC=MS:1000201 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023522-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available LEDG_DSG20_11581.mzXML proteomic profiling by mass spectrometry LEDG_DSG20_11581.mzXML 1 4 AC=MS:1002038;NT=label free sample NT=ultraflex;AC=MS:1000201 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0
PXD023522-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available Nkrp1B_CE_14N15N_DSG20_4007.mzXML proteomic profiling by mass spectrometry Nkrp1B_CE_14N15N_DSG20_4007.mzXML 1 5 AC=MS:1002038;NT=label free sample NT=ultraflex;AC=MS:1000201 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

8. Isotope-labelled runs appear label-free 🐞 Bug ≡ Correctness

PXD023522.sdrf.tsv assigns AC=MS:1002038;NT=label free sample to assays whose names and data
files explicitly contain 14N15N. The four affected acquisitions are consequently indistinguishable
from genuinely label-free runs for consumers of the structured label field.
Agent Prompt
## Issue description
Rows whose assay and data-file values contain `14N15N` are recorded as label-free even though those names identify nitrogen-isotope-labelled acquisitions.

## Fix Focus Areas
- datasets/PXD023522/PXD023522.sdrf.tsv[6-10]

## Recommended Fix
Replace the label-free value on each `14N15N` row with the appropriate controlled-vocabulary isotope-labelling annotation. Leave the genuinely label-free rows unchanged.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@@ -0,0 +1,8 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template]
PXD024253-sample escherichia coli not applicable not applicable not applicable 1 synthetic reference not available not available not available not available 117-pDSBE-BSA2-_1_.mzML proteomic profiling by mass spectrometry 117-pDSBE-BSA2-_1_.mzML 1 1 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion;AC=MS:1002416 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

9. Bovine albumin has wrong species 🐞 Bug ≡ Correctness

The BSA-named assays in PXD024253.sdrf.tsv set characteristics[organism] to escherichia coli.
Those four BSA acquisitions are therefore indexed as bacterial material while the separate E. coli
assay group is explicitly identifiable as Ecoli in its names.
Agent Prompt
## Issue description
The `117-pDSBE-BSA2` acquisition group is bovine serum albumin material but is annotated as `escherichia coli` in the organism characteristic.

## Fix Focus Areas
- datasets/PXD024253/PXD024253.sdrf.tsv[2-5]

## Recommended Fix
Change `characteristics[organism]` for the four `117-pDSBE-BSA2` rows to the appropriate bovine organism value. Retain the existing E. coli organism annotation for the `155-pDSBE-Ecoli-Z2` rows.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment on lines +1 to +2
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[age] characteristics[sex] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD024399-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 20201028_DS_SCYL1BioID_WTandS754A_HML_R1_1 proteomic profiling by mass spectrometry 20201028_DS_SCYL1BioID_WTandS754A_HML_R1_1.raw 1 1 AC=MS:1002038;NT=label free sample NT=Q Exactive HF;AC=MS:1002523 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=BioID;AC=XLMOD:02250 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

25. Variants cannot be grouped reliably 🐞 Bug ⚙ Maintainability

PXD024399.sdrf.tsv has no genotype, mutation, construct, or factor column even though its assay
names encode WT, S754A, and S754E groups. Because all structured sample fields remain the same
across these groups, downstream users must parse filenames to separate the experimental conditions.
Agent Prompt
## Issue description
The distinct WT, S754A, and S754E experimental conditions exist only in assay names, not in a structured SDRF characteristic or factor.

## Fix Focus Areas
- datasets/PXD024399/PXD024399.sdrf.tsv[1-61]

## Recommended Fix
Add an appropriate structured genotype, mutation, construct, or factor column and populate it for every row so WT, S754A, and S754E assays can be grouped without parsing their filenames.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment on lines +2 to +3
PXD025066-sample oryctolagus sp. 'rabbit_od' not applicable not applicable not applicable 1 synthetic reference not available not available not available not available 3690_LM_C4H2O2_TrypAspN_Xi1.7.6.1.mzid.gz proteomic profiling by mass spectrometry 3690_LM_C4H2O2_TrypAspN_Xi1.7.6.1.mzid.gz 1 1 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD025066-sample oryctolagus sp. 'rabbit_od' not applicable not applicable not applicable 1 synthetic reference not available not available not available not available 3690_LM_C4H2O2_TrypChymo_Xi1.7.6.1.mzid.gz proteomic profiling by mass spectrometry 3690_LM_C4H2O2_TrypChymo_Xi1.7.6.1.mzid.gz 1 2 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

10. Digests are labelled as lys-c 🐞 Bug ≡ Correctness

PXD025066.sdrf.tsv records Lys-C as the second cleavage agent for TrypAspN and TrypChymo
assays. Each AspN- or chymotrypsin-containing run therefore has digestion metadata that conflicts
with its own acquisition identifier.
Agent Prompt
## Issue description
The second cleavage-agent column says Lys-C for assays named `TrypAspN` and `TrypChymo`, which identify AspN and chymotrypsin digestions instead.

## Fix Focus Areas
- datasets/PXD025066/PXD025066.sdrf.tsv[2-7]

## Recommended Fix
Replace the second cleavage-agent annotation with the appropriate AspN value on every `TrypAspN` row and the appropriate chymotrypsin value on every `TrypChymo` row. Keep Trypsin as the first agent if the samples were digested in combination with trypsin.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@qodo-code-review

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 8dbd0ca

comment[sdrf template] declared NT=invertebrates;VV=v1.1.0 (an animal-only
template) for 8 Saccharomyces cerevisiae datasets. Removing the mismatched
template column; ms-proteomics and crosslinking layers are unaffected.

Confirmed by qodo-code-review[bot] and this report's own data check.
@ypriverol
ypriverol merged commit d034881 into main Sep 17, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

sdrf:new SDRF PR adds new datasets

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants