Skip to content

Add crosslinking proteomics SDRF annotations (batch 13/15, 50 datasets) - #551

Merged
ypriverol merged 3 commits into
mainfrom
add/crosslinking-batch-13
Sep 17, 2026
Merged

ypriverol merged 3 commits into
mainfrom
add/crosslinking-batch-13

Conversation

@ypriverol

Copy link
Copy Markdown
Contributor

Part of the crosslinking proteomics SDRF annotation effort (splits PR #537 into batches of 50). Adds 50 new datasets in flat datasets/<accession>/ layout. Accessions: PXD051014,PXD051047,PXD051143,PXD051261,PXD051348,PXD051405,PXD051493,PXD051557,PXD051602,PXD051693,PXD051742,PXD051886,PXD051971,PXD052310,PXD052552,PXD052623,PXD052624,PXD052637,PXD052687,PXD052694,PXD052745,PXD052746,PXD052801,PXD052821,PXD052825,PXD052867,PXD052917,PXD052923,PXD052926,PXD052930,PXD053010,PXD053341,PXD053452,PXD053489,PXD053494,PXD053509,PXD053578,PXD053607,PXD053636,PXD053760,PXD053832,PXD053924,PXD053984,PXD054003,PXD054140,PXD054141,PXD054249,PXD054551,PXD054616,PXD054720

Copilot AI balanced review requested due to automatic review settings September 17, 2026 04:40
@coderabbitai

coderabbitai Bot commented Sep 17, 2026

Copy link
Copy Markdown

Important

Review skipped

Review was skipped due to path filters

⛔ Files ignored due to path filters (50)
  • datasets/PXD051014/PXD051014.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051047/PXD051047.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051143/PXD051143.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051261/PXD051261.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051348/PXD051348.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051405/PXD051405.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051493/PXD051493.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051557/PXD051557.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051602/PXD051602.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051693/PXD051693.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051742/PXD051742.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051886/PXD051886.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD051971/PXD051971.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052310/PXD052310.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052552/PXD052552.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052623/PXD052623.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052624/PXD052624.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052637/PXD052637.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052687/PXD052687.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052694/PXD052694.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052745/PXD052745.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052746/PXD052746.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052801/PXD052801.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052821/PXD052821.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052825/PXD052825.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052867/PXD052867.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052917/PXD052917.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052923/PXD052923.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052926/PXD052926.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD052930/PXD052930.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053010/PXD053010.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053341/PXD053341.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053452/PXD053452.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053489/PXD053489.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053494/PXD053494.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053509/PXD053509.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053578/PXD053578.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053607/PXD053607.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053636/PXD053636.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053760/PXD053760.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053832/PXD053832.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053924/PXD053924.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD053984/PXD053984.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD054003/PXD054003.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD054140/PXD054140.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD054141/PXD054141.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD054249/PXD054249.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD054551/PXD054551.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD054616/PXD054616.sdrf.tsv is excluded by !**/*.tsv
  • datasets/PXD054720/PXD054720.sdrf.tsv is excluded by !**/*.tsv

CodeRabbit blocks several paths by default. You can override this behavior by explicitly including those paths in the path filters. For example, including **/dist/** will override the default block on the dist directory, by removing the pattern from both the lists.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 1ebeb579-f7b6-408d-a04f-34079e85ee93

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@qodo-code-review

qodo-code-review Bot commented Sep 17, 2026

Copy link
Copy Markdown

PR Summary by Qodo

Add batch 13 crosslinking proteomics SDRF annotations

✨ Enhancement 📝 Documentation 🕐 40+ Minutes

Grey Divider

AI Description

• Adds standardized SDRF annotations for 50 crosslinking proteomics datasets.
• Captures sample, acquisition, crosslinker, modification, and template metadata.
• Supplies required organism-template fields with spec-compliant unavailable values.
Diagram

graph TD
  A["PRIDE datasets"] --> B["50 SDRF files"] --> C["Template validation"] --> D["Dataset catalog"]
  E["SDRF templates"] --> C
Loading
High-Level Assessment

The flat accession-based layout and standard SDRF templates match the repository's established data model. Automated generation was considered, but accession-specific biological and crosslinking metadata still requires curated values, so this batched approach is appropriate.

Files changed (50) +1203 / -0

Enhancement (50) +1203 / -0
PXD051014.sdrf.tsvAdd PXD051014 BioID SDRF annotation +84/-0

Add PXD051014 BioID SDRF annotation

• Adds human BioID assay mappings and standardized sample, acquisition, processing, and template metadata for PXD051014.

datasets/PXD051014/PXD051014.sdrf.tsv

PXD051047.sdrf.tsvAdd PXD051047 yeast SDRF annotation +75/-0

Add PXD051047 yeast SDRF annotation

• Adds yeast crosslinking assay mappings for PXD051047. Includes required developmental-stage and strain fields for the invertebrates template.

datasets/PXD051047/PXD051047.sdrf.tsv

PXD051143.sdrf.tsvAdd PXD051143 BioID SDRF annotation +7/-0

Add PXD051143 BioID SDRF annotation

• Adds human BioID assay and raw-file mappings with standardized crosslinking and acquisition metadata.

datasets/PXD051143/PXD051143.sdrf.tsv

PXD051261.sdrf.tsvAdd PXD051261 bacterial SDRF annotation +100/-0

Add PXD051261 bacterial SDRF annotation

• Adds extensive Streptococcus pyogenes assay and data-file mappings with crosslinking proteomics metadata.

datasets/PXD051261/PXD051261.sdrf.tsv

PXD051348.sdrf.tsvAdd PXD051348 BioID SDRF annotation +78/-0

Add PXD051348 BioID SDRF annotation

• Adds human BioID annotations covering timsTOF runs, analysis artifacts, and standardized template metadata.

datasets/PXD051348/PXD051348.sdrf.tsv

PXD051405.sdrf.tsvAdd PXD051405 BioID SDRF annotation +94/-0

Add PXD051405 BioID SDRF annotation

• Adds human BioID assay mappings across control and cell-line samples with Orbitrap acquisition metadata.

datasets/PXD051405/PXD051405.sdrf.tsv

PXD051493.sdrf.tsvAdd PXD051493 plant XL-MS annotation +58/-0

Add PXD051493 plant XL-MS annotation

• Adds Chlamydomonas DSS crosslinking annotations, including crosslink distance and required developmental-stage metadata.

datasets/PXD051493/PXD051493.sdrf.tsv

PXD051557.sdrf.tsvAdd PXD051557 APEX2 SDRF annotation +13/-0

Add PXD051557 APEX2 SDRF annotation

• Adds human APEX2 proximity-labeling runs with assay, replicate, instrument, and template metadata.

datasets/PXD051557/PXD051557.sdrf.tsv

PXD051602.sdrf.tsvAdd PXD051602 crosslinking SDRF annotation +9/-0

Add PXD051602 crosslinking SDRF annotation

• Adds human crosslinking assay mappings and Orbitrap Exploris acquisition metadata for PXD051602.

datasets/PXD051602/PXD051602.sdrf.tsv

PXD051693.sdrf.tsvAdd PXD051693 crosslinking SDRF annotation +15/-0

Add PXD051693 crosslinking SDRF annotation

• Adds human histone-group assay mappings with standardized acquisition and crosslinking metadata.

datasets/PXD051693/PXD051693.sdrf.tsv

PXD051742.sdrf.tsvAdd PXD051742 DSSO SDRF annotation +23/-0

Add PXD051742 DSSO SDRF annotation

• Adds human DSSO crosslinking runs with controlled crosslinker properties, quenching details, and raw-file mappings.

datasets/PXD051742/PXD051742.sdrf.tsv

PXD051886.sdrf.tsvAdd PXD051886 SDA SDRF annotation +39/-0

Add PXD051886 SDA SDRF annotation

• Adds human SDA crosslinking annotations with distance, Orbitrap Eclipse, and assay-file metadata.

datasets/PXD051886/PXD051886.sdrf.tsv

PXD051971.sdrf.tsvAdd PXD051971 yeast DSBU annotation +71/-0

Add PXD051971 yeast DSBU annotation

• Adds yeast DSBU crosslinking assay mappings. Includes developmental-stage and strain fields required by the organism template.

datasets/PXD051971/PXD051971.sdrf.tsv

PXD052310.sdrf.tsvAdd PXD052310 APEX2 SDRF annotation +52/-0

Add PXD052310 APEX2 SDRF annotation

• Adds Mycobacterium tuberculosis APEX2 assay and data-file mappings with standardized proteomics metadata.

datasets/PXD052310/PXD052310.sdrf.tsv

PXD052552.sdrf.tsvAdd PXD052552 bovine SDRF annotation +21/-0

Add PXD052552 bovine SDRF annotation

• Adds bovine crosslinking assay mappings and supplies required developmental-stage metadata for the vertebrates template.

datasets/PXD052552/PXD052552.sdrf.tsv

PXD052623.sdrf.tsvAdd PXD052623 APEX2 SDRF annotation +25/-0

Add PXD052623 APEX2 SDRF annotation

• Adds human APEX2 proximity-labeling runs with Q Exactive HF acquisition and template metadata.

datasets/PXD052623/PXD052623.sdrf.tsv

PXD052624.sdrf.tsvAdd PXD052624 APEX2 SDRF annotation +22/-0

Add PXD052624 APEX2 SDRF annotation

• Adds human APEX2 assay mappings with standardized raw-file, acquisition, and processing annotations.

datasets/PXD052624/PXD052624.sdrf.tsv

PXD052637.sdrf.tsvAdd PXD052637 BioID SDRF annotation +9/-0

Add PXD052637 BioID SDRF annotation

• Adds human BioID assay mappings and Q Exactive HF acquisition metadata for PXD052637.

datasets/PXD052637/PXD052637.sdrf.tsv

PXD052687.sdrf.tsvAdd PXD052687 TurboID SDRF annotation +8/-0

Add PXD052687 TurboID SDRF annotation

• Adds human TurboID assay and raw-file mappings with Orbitrap Exploris metadata.

datasets/PXD052687/PXD052687.sdrf.tsv

PXD052694.sdrf.tsvAdd PXD052694 mouse TurboID annotation +25/-0

Add PXD052694 mouse TurboID annotation

• Adds mouse TurboID data-file mappings and required developmental-stage metadata for the vertebrates template.

datasets/PXD052694/PXD052694.sdrf.tsv

PXD052745.sdrf.tsvAdd PXD052745 TurboID SDRF annotation +58/-0

Add PXD052745 TurboID SDRF annotation

• Adds human TurboID assay mappings with Orbitrap Fusion acquisition and standardized template metadata.

datasets/PXD052745/PXD052745.sdrf.tsv

PXD052746.sdrf.tsvAdd PXD052746 BioID SDRF annotation +46/-0

Add PXD052746 BioID SDRF annotation

• Adds human BioID assay and raw-file mappings with Q Exactive acquisition metadata.

datasets/PXD052746/PXD052746.sdrf.tsv

PXD052801.sdrf.tsvAdd PXD052801 PIR SDRF annotation +16/-0

Add PXD052801 PIR SDRF annotation

• Adds human PIR crosslinking assay mappings covering mzML files and standardized acquisition metadata.

datasets/PXD052801/PXD052801.sdrf.tsv

PXD052821.sdrf.tsvAdd PXD052821 DSSO SDRF annotation +19/-0

Add PXD052821 DSSO SDRF annotation

• Adds human DSSO crosslinking runs with crosslink distance, cleavable-linker properties, and instrument metadata.

datasets/PXD052821/PXD052821.sdrf.tsv

PXD052825.sdrf.tsvAdd PXD052825 TurboID SDRF annotation +10/-0

Add PXD052825 TurboID SDRF annotation

• Adds human TurboID assay mappings and LTQ Orbitrap acquisition metadata for PXD052825.

datasets/PXD052825/PXD052825.sdrf.tsv

PXD052867.sdrf.tsvAdd PXD052867 SDA SDRF annotation +40/-0

Add PXD052867 SDA SDRF annotation

• Adds human SDA crosslinking data-file mappings with controlled crosslink-distance and instrument annotations.

datasets/PXD052867/PXD052867.sdrf.tsv

PXD052917.sdrf.tsvAdd PXD052917 BioID SDRF annotation +46/-0

Add PXD052917 BioID SDRF annotation

• Adds human BioID assay and raw-file mappings with standardized processing and template metadata.

datasets/PXD052917/PXD052917.sdrf.tsv

PXD052923.sdrf.tsvAdd PXD052923 BS3 SDRF annotation +2/-0

Add PXD052923 BS3 SDRF annotation

• Adds a Trypanosoma brucei BS3 crosslinking assay with Orbitrap Fusion Lumos acquisition metadata.

datasets/PXD052923/PXD052923.sdrf.tsv

PXD052926.sdrf.tsvAdd PXD052926 canine SDRF annotation +27/-0

Add PXD052926 canine SDRF annotation

• Adds canine crosslinking assay and fraction mappings with Orbitrap Fusion acquisition metadata.

datasets/PXD052926/PXD052926.sdrf.tsv

PXD052930.sdrf.tsvAdd PXD052930 APEX2 SDRF annotation +49/-0

Add PXD052930 APEX2 SDRF annotation

• Adds human APEX2 assay mappings with Orbitrap Fusion and standardized crosslinking metadata.

datasets/PXD052930/PXD052930.sdrf.tsv

PXD053010.sdrf.tsvAdd PXD053010 EDC SDRF annotation +20/-0

Add PXD053010 EDC SDRF annotation

• Adds human EDC crosslinking runs with crosslink distance, target residues, and raw-file mappings.

datasets/PXD053010/PXD053010.sdrf.tsv

PXD053341.sdrf.tsvAdd PXD053341 SDA SDRF annotation +19/-0

Add PXD053341 SDA SDRF annotation

• Adds human SDA crosslinking assay mappings with Orbitrap Eclipse acquisition metadata.

datasets/PXD053341/PXD053341.sdrf.tsv

PXD053452.sdrf.tsvAdd PXD053452 BioID SDRF annotation +11/-0

Add PXD053452 BioID SDRF annotation

• Adds Toxoplasma gondii BioID assay mappings with Q Exactive HF acquisition and processing metadata.

datasets/PXD053452/PXD053452.sdrf.tsv

PXD053489.sdrf.tsvAdd PXD053489 BioID SDRF annotation +12/-0

Add PXD053489 BioID SDRF annotation

• Adds human BioID runs with Orbitrap Fusion Lumos acquisition and standardized template metadata.

datasets/PXD053489/PXD053489.sdrf.tsv

PXD053494.sdrf.tsvAdd PXD053494 crosslinking SDRF annotation +0/-0

Add PXD053494 crosslinking SDRF annotation

• Populates the previously empty accession file with human assays, raw-file mappings, and crosslinking metadata.

datasets/PXD053494/PXD053494.sdrf.tsv

PXD053509.sdrf.tsvAdd PXD053509 APEX2 SDRF annotation +0/-0

Add PXD053509 APEX2 SDRF annotation

• Populates the accession with human APEX2 timsTOF runs and associated analysis-file metadata.

datasets/PXD053509/PXD053509.sdrf.tsv

PXD053578.sdrf.tsvAdd PXD053578 APEX2 SDRF annotation +0/-0

Add PXD053578 APEX2 SDRF annotation

• Populates the accession with human APEX2 assay mappings and Orbitrap Fusion Lumos metadata.

datasets/PXD053578/PXD053578.sdrf.tsv

PXD053607.sdrf.tsvAdd PXD053607 SMCC SDRF annotation +0/-0

Add PXD053607 SMCC SDRF annotation

• Populates the accession with a human SMCC crosslinking assay and Orbitrap Fusion Lumos metadata.

datasets/PXD053607/PXD053607.sdrf.tsv

PXD053636.sdrf.tsvAdd PXD053636 yeast PhoX annotation +0/-0

Add PXD053636 yeast PhoX annotation

• Adds yeast PhoX crosslinking runs and related result files. Includes required developmental-stage and strain fields.

datasets/PXD053636/PXD053636.sdrf.tsv

PXD053760.sdrf.tsvAdd PXD053760 BioID SDRF annotation +0/-0

Add PXD053760 BioID SDRF annotation

• Populates the accession with Trypanosoma brucei BioID runs and TripleTOF acquisition metadata.

datasets/PXD053760/PXD053760.sdrf.tsv

PXD053832.sdrf.tsvAdd PXD053832 TurboID SDRF annotation +0/-0

Add PXD053832 TurboID SDRF annotation

• Populates the accession with human TurboID assay mappings and Orbitrap Fusion Lumos metadata.

datasets/PXD053832/PXD053832.sdrf.tsv

PXD053924.sdrf.tsvAdd PXD053924 BioID SDRF annotation +0/-0

Add PXD053924 BioID SDRF annotation

• Populates the accession with human BioID replicates and Q Exactive HF acquisition metadata.

datasets/PXD053924/PXD053924.sdrf.tsv

PXD053984.sdrf.tsvAdd PXD053984 DSSO SDRF annotation +0/-0

Add PXD053984 DSSO SDRF annotation

• Populates the accession with human DSSO runs, cleavable-linker properties, and Orbitrap metadata.

datasets/PXD053984/PXD053984.sdrf.tsv

PXD054003.sdrf.tsvAdd PXD054003 BioID SDRF annotation +0/-0

Add PXD054003 BioID SDRF annotation

• Populates the accession with human BioID assay mappings and Q Exactive acquisition metadata.

datasets/PXD054003/PXD054003.sdrf.tsv

PXD054140.sdrf.tsvAdd PXD054140 DSSO SDRF annotation +0/-0

Add PXD054140 DSSO SDRF annotation

• Populates the accession with human DSSO concentration runs and controlled crosslinker properties.

datasets/PXD054140/PXD054140.sdrf.tsv

PXD054141.sdrf.tsvAdd PXD054141 DSSO SDRF annotation +0/-0

Add PXD054141 DSSO SDRF annotation

• Populates the accession with human DSSO protein-complex runs and Orbitrap Fusion Lumos metadata.

datasets/PXD054141/PXD054141.sdrf.tsv

PXD054249.sdrf.tsvAdd PXD054249 BioID SDRF annotation +0/-0

Add PXD054249 BioID SDRF annotation

• Populates the accession with human BioID assay mappings and Q Exactive acquisition metadata.

datasets/PXD054249/PXD054249.sdrf.tsv

PXD054551.sdrf.tsvAdd PXD054551 crosslinking SDRF annotation +0/-0

Add PXD054551 crosslinking SDRF annotation

• Populates the accession with human Orbitrap Fusion Lumos runs and standardized crosslinking metadata.

datasets/PXD054551/PXD054551.sdrf.tsv

PXD054616.sdrf.tsvAdd PXD054616 TurboID SDRF annotation +0/-0

Add PXD054616 TurboID SDRF annotation

• Populates the accession with human TurboID assay replicates and Orbitrap Fusion metadata.

datasets/PXD054616/PXD054616.sdrf.tsv

PXD054720.sdrf.tsvAdd PXD054720 bacterial DSSO annotation +0/-0

Add PXD054720 bacterial DSSO annotation

• Populates the accession with Escherichia coli DSSO runs, identification files, and controlled crosslinker metadata.

datasets/PXD054720/PXD054720.sdrf.tsv

@qodo-code-review

qodo-code-review Bot commented Sep 17, 2026

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (3) 📘 Rule violations (4) 📜 Skill insights (0)

Grey Divider


Action required

1. Cattle samples omit required breed 🐞 Bug ≡ Correctness ⭐ New
Description
The PXD052552 and PXD052694 headers declare the vertebrates template but omit
characteristics[strain or breed] between their sample characteristics. Every added Bos taurus and
Mus musculus row therefore lacks the required field and cannot record breed or strain information,
even with the reserved not available value.
Code

datasets/PXD052552/PXD052552.sdrf.tsv[1]

+source name	characteristics[organism]	characteristics[organism part]	characteristics[cell type]	characteristics[disease]	characteristics[developmental stage]	characteristics[biological replicate]	characteristics[material type]	characteristics[sample type]	characteristics[enrichment process]	characteristics[crosslink distance]	characteristics[crosslinking reaction time]	characteristics[crosslinking temperature]	assay name	technology type	comment[data file]	comment[technical replicate]	comment[fraction identifier]	comment[label]	comment[instrument]	comment[proteomics data acquisition method]	comment[cleavage agent details]	comment[cleavage agent details]	comment[modification parameters]	comment[modification parameters]	comment[dissociation method]	comment[collision energy]	comment[precursor mass tolerance]	comment[fragment mass tolerance]	comment[fractionation method]	comment[chemical cross-linking coupled with ms]	comment[cross-linker]	comment[crosslink enrichment method]	comment[crosslinker concentration]	comment[quenching reagent]	comment[reduction reagent]	comment[alkylation reagent]	comment[sdrf version]	comment[sdrf template]	comment[sdrf template]	comment[sdrf template]
Evidence
PXD052552 and PXD052694 identify their samples as Bos taurus and Mus musculus, respectively, and
both select NT=vertebrates, but each header proceeds directly from developmental stage to
biological replicate. The established vertebrate SDRF example includes developmental stage and
strain or breed as distinct columns, demonstrating the omission in both datasets.

datasets/PXD052552/PXD052552.sdrf.tsv[1-3]
datasets/PXD010957/PXD010957.sdrf.tsv[1-2]
datasets/PXD052694/PXD052694.sdrf.tsv[1-3]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The PXD052552 and PXD052694 vertebrate SDRFs omit the required `characteristics[strain or breed]` column and corresponding row values.

## Fix Focus Areas
- datasets/PXD052552/PXD052552.sdrf.tsv[1-21]
- datasets/PXD052694/PXD052694.sdrf.tsv[1-25]

## Recommended Fix
Add `characteristics[strain or breed]` to each header alongside the other organism characteristics and add a value to every row. Use `not available` when the cattle breed or mouse strain is unknown.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. EDC runs are indexed as diazirine 🐞 Bug ≡ Correctness ⭐ New
Description
PXD053010 labels every EDC-named assay as NT=EDC but assigns AC=XLMOD:02009, an accession the
repository associates with NT=diazirine. All 19 added EDC assay rows therefore expose diazirine
chemistry to consumers filtering or interpreting cross-linker annotations.
Code

datasets/PXD053010/PXD053010.sdrf.tsv[2]

+PXD053010-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	11.4 Å	not available	not available	Zou_Rappsilber_AC_NF90-NF45-RNA_EDC_S1_R1	proteomic profiling by mass spectrometry	Zou_Rappsilber_AC_NF90-NF45-RNA_EDC_S1_R1.raw	1	1	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion Lumos;AC=MS:1002732	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=EDC;AC=XLMOD:02009;CL=no;TA=K,D,E	not available	not available	using	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
Evidence
The new EDC assay row contains both the EDC name in its run identifier and NT=EDC;AC=XLMOD:02009
in its cross-linker field. Elsewhere in the repository, the same accession is explicitly annotated
as diazirine, demonstrating that this accession does not identify EDC.

datasets/PXD053010/PXD053010.sdrf.tsv[2-2]
datasets/PXD004009/PXD004009.sdrf.tsv[2-2]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
PXD053010 assigns the diazirine accession `XLMOD:02009` to assays explicitly identified as EDC, so downstream consumers receive the wrong cross-linker identity.

## Fix Focus Areas
- datasets/PXD053010/PXD053010.sdrf.tsv[2-20]

## Recommended Fix
Replace the cross-linker annotation in every PXD053010 row with the validated XLMOD accession and associated attributes for EDC. Recheck the linked cross-link distance and target-residue metadata so they describe the corrected reagent consistently.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. Known crosslinkers become unknown 🐞 Bug ≡ Correctness
Description
PXD051261 assigns NT=unknown crosslinker;AC=XLMOD:00000 to files whose names explicitly identify
DSS and DSG. The mismatch affects both reagent groups and prevents consumers from selecting the
corresponding crosslinking chemistry.
Code

datasets/PXD051261/PXD051261.sdrf.tsv[68]

+PXD051261-sample	streptococcus pyogenes mgas315	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	P_2S_tPA_1mM_DSS_1_DT_C2203_101	proteomic profiling by mass spectrometry	P_2S_tPA_1mM_DSS_1_DT_C2203_101.raw	1	67	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF-X;AC=MS:1002877	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
The DSS and DSG identifiers occur in the assay and data filenames while the same rows declare an
unknown crosslinker.

datasets/PXD051261/PXD051261.sdrf.tsv[68-68]
datasets/PXD051261/PXD051261.sdrf.tsv[86-86]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
PXD051261 marks DSS and DSG runs as using an unknown crosslinker even though their filenames identify the reagents.

## Fix Focus Areas
- datasets/PXD051261/PXD051261.sdrf.tsv[68-99]

## Recommended Fix
Replace the unknown-crosslinker values on DSS and DSG rows with the appropriate XLMOD terms and reagent-specific metadata, using archive evidence to separate the two groups.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View high (5)
4. Cell-line acquisitions are mislabeled 📘 Rule violation ≡ Correctness
Description
The comment[proteomics data acquisition method] values in PXD052801, PXD052821, PXD052867, and
PXD053509 declare Data-dependent acquisition even though the corresponding assay, raw-file, or
quantitative-file names identify DIA or SpDIA runs. The conflict affects the HEK293, HeLa, Lumos,
and added quantitative-file rows, causing consumers of the acquisition-method field to classify
independent-acquisition experiments as dependent-acquisition data.
Code

datasets/PXD052801/PXD052801.sdrf.tsv[2]

+PXD052801-sample	homo sapiens	not applicable	not applicable	not applicable	not available	not applicable	1	synthetic	reference	not available	not available	not available	not available	062123_cell_lines_DIA_24mz_HEK293_1.mzML	proteomic profiling by mass spectrometry	062123_cell_lines_DIA_24mz_HEK293_1.mzML	1	1	AC=MS:1002038;NT=label free sample	NT=Q Exactive;AC=MS:1001911	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=PIR;AC=XLMOD:02014	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=human;VV=v1.1.0
Evidence
Each cited range contains rows pairing DIA- or SpDIA-identifying filenames with `NT=Data-dependent
acquisition;AC=PRIDE:0000449`; this conflicts with the archive-derived file evidence and with an
existing DIA SDRF that uses the independent-acquisition term. The repeated mismatch across the cited
datasets demonstrates that the acquisition metadata does not reflect the public file evidence
required by Rule 3.

AGENTS.md: Align SDRF Metadata with Public Archive Evidence
datasets/PXD052801/PXD052801.sdrf.tsv[2-16]
datasets/PXD052801/PXD052801.sdrf.tsv[2-2]
datasets/PXD052821/PXD052821.sdrf.tsv[2-3]
datasets/PXD052867/PXD052867.sdrf.tsv[3-3]
datasets/PXD053509/PXD053509.sdrf.tsv[2-2]
datasets/PXD049208/PXD049208.sdrf.tsv[2-2]
datasets/PXD052821/PXD052821.sdrf.tsv[2-19]
datasets/PXD052867/PXD052867.sdrf.tsv[3-23]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Four datasets contain files identified by assay, raw-file, or quantitative-file names as DIA or SpDIA runs, but their SDRF acquisition-method fields annotate them as data-dependent acquisitions.

## Fix Focus Areas
- datasets/PXD052801/PXD052801.sdrf.tsv[2-16]
- datasets/PXD052821/PXD052821.sdrf.tsv[2-19]
- datasets/PXD052867/PXD052867.sdrf.tsv[3-40]
- datasets/PXD053509/PXD053509.sdrf.tsv[2-35]

## Recommended Fix
Replace the data-dependent acquisition mapping with the repository's supported data-independent acquisition ontology mapping for each row confirmed by its filename as a DIA or SpDIA run. Limit the update to confirmed DIA or SpDIA rows and verify every changed mapping against the public archive metadata.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


5. EDC assays are annotated as DSBU 📘 Rule violation ≡ Correctness
Description
The EDC-named assay rows in PXD051971 assign the DSBU term NT=DSBU;AC=XLMOD:02043, target
residues, and mass values as their cross-linker parameters. This mismatch affects EDC Trp, Ubiq,
GluC, and GC assays and their associated files, exposing parameters for a different cross-linking
reaction to downstream searches.
Code

datasets/PXD051971/PXD051971.sdrf.tsv[6]

+PXD051971-sample	saccharomyces cerevisiae	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	20210519_H2A_H2B_Ubp10_EDC_Trp_1	proteomic profiling by mass spectrometry	20210519_H2A_H2B_Ubp10_EDC_Trp_1.raw	1	5	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF-X;AC=MS:1002877	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03	not available	not available	the	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=invertebrates;VV=v1.1.0
Evidence
Rule 3 requires sample and raw-file metadata to agree with archive evidence, but the assay and file
names identify these runs as EDC while their cross-linker fields identify DSBU. DSBU-named and
EDC-named rows carry identical DSBU annotations, showing that the chemistry was copied without
accounting for the EDC group.

AGENTS.md: Align SDRF Metadata with Public Archive Evidence
datasets/PXD051971/PXD051971.sdrf.tsv[6-7]
datasets/PXD051971/PXD051971.sdrf.tsv[21-22]
datasets/PXD051971/PXD051971.sdrf.tsv[32-33]
datasets/PXD051971/PXD051971.sdrf.tsv[46-47]
datasets/PXD051971/PXD051971.sdrf.tsv[2-9]
datasets/PXD051971/PXD051971.sdrf.tsv[62-67]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
PXD051971 assigns DSBU chemistry, including its ontology term, target residues, and mass values, to runs explicitly identified as EDC experiments.

## Fix Focus Areas
- datasets/PXD051971/PXD051971.sdrf.tsv[6-67]

## Recommended Fix
Identify every EDC row and replace the DSBU cross-linker term, DSBU-specific mass values, and target attributes with the correct EDC annotation supported by the study metadata. Retain DSBU only for assays whose archive names and evidence identify DSBU.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


6. Yeast assay rows have invalid metadata 📘 Rule violation ≡ Correctness
Description
Eight added SDRFs populate comment[quenching reagent] with unsupported fragments or generic terms
such as the, by, of, solution, and media, while PXD051971 and PXD053636 also pair
Saccharomyces cerevisiae with the animal-specific invertebrates template. The malformed
annotations repeat across the affected assays—including APEX2, PhoX, and DSSO runs—and, for the two
yeast accessions, their associated result-file rows, with the sole PXD053607 assay reaching the
canonical dataset unchanged.
Code

datasets/PXD051971/PXD051971.sdrf.tsv[2]

+PXD051971-sample	saccharomyces cerevisiae	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	20210519_H2A_H2B_Ubp10_DSBU_Trp_1	proteomic profiling by mass spectrometry	20210519_H2A_H2B_Ubp10_DSBU_Trp_1.raw	1	1	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF-X;AC=MS:1002877	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03	not available	not available	the	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=invertebrates;VV=v1.1.0
Evidence
The cited headers place the reported tokens in comment[quenching reagent], and representative rows
show them repeated after the crosslinker-concentration field across each affected accession. In
particular, PXD051971 and PXD053636 pair Saccharomyces cerevisiae, a yeast, with
NT=invertebrates;VV=v1.1.0, while using the and by as reagent values; the other cited rows
likewise use fragments or generic terms such as of, solution, and media. None identifies a
standardized quenching reagent or accepted missing value, demonstrating both incoherent template
selection and nonstandard column values under Rule 4.

AGENTS.md: Use Ontology Terms and Columns Consistent with the Selected SDRF Template
datasets/PXD051971/PXD051971.sdrf.tsv[1-2]
datasets/PXD051971/PXD051971.sdrf.tsv[2-2]
datasets/PXD053636/PXD053636.sdrf.tsv[2-2]
datasets/PXD053636/PXD053636.sdrf.tsv[1-2]
datasets/PXD052310/PXD052310.sdrf.tsv[1-2]
datasets/PXD052930/PXD052930.sdrf.tsv[1-2]
datasets/PXD053607/PXD053607.sdrf.tsv[1-2]
datasets/PXD053984/PXD053984.sdrf.tsv[1-2]
datasets/PXD054551/PXD054551.sdrf.tsv[1-2]
datasets/PXD053578/PXD053578.sdrf.tsv[1-2]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Eight datasets place sentence fragments or generic terms such as `the`, `by`, `of`, `solution`, and `media` in `comment[quenching reagent]`; PXD051971 and PXD053636 additionally assign yeast samples to the invertebrates template.

## Fix Focus Areas
- datasets/PXD051971/PXD051971.sdrf.tsv[2-71]
- datasets/PXD052310/PXD052310.sdrf.tsv[2-52]
- datasets/PXD052930/PXD052930.sdrf.tsv[2-49]
- datasets/PXD053578/PXD053578.sdrf.tsv[2-47]
- datasets/PXD053607/PXD053607.sdrf.tsv[2-2]
- datasets/PXD053636/PXD053636.sdrf.tsv[2-35]
- datasets/PXD053984/PXD053984.sdrf.tsv[2-23]
- datasets/PXD054551/PXD054551.sdrf.tsv[2-36]

## Recommended Fix
Replace every malformed quenching-reagent value with the actual archive-supported reagent and applicable ontology mapping based on study evidence, or use `not available` when no reagent can be established. For PXD051971 and PXD053636, remove the invertebrates template and select the repository-supported fungal or yeast template when applicable, then validate every SDRF row and associated result-file row.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


7. Ancillary files become acquisition assays 📘 Rule violation ≡ Correctness
Description
Multiple SDRFs map FASTA databases, identification results, checksums, spreadsheets, archives, and
documents into both assay name and comment[data file] instead of limiting those fields to
deposited raw-category runs. Whenever these ancillary files are included, each is treated as a
separate proteomics acquisition assay, creating nonexistent experimental runs and incorrect
sample-to-run relationships alongside the actual raw files.
Code

datasets/PXD054720/PXD054720.sdrf.tsv[R5-8]

+PXD054720-sample	escherichia coli	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	ABRF_iPRG_XL_2023.fasta	proteomic profiling by mass spectrometry	ABRF_iPRG_XL_2023.fasta	1	4	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion Lumos;AC=MS:1002732	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
+PXD054720-sample	escherichia coli	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	F001234.mzid.gz	proteomic profiling by mass spectrometry	F001234.mzid.gz	1	5	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion Lumos;AC=MS:1002732	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
+PXD054720-sample	escherichia coli	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	F001235.mzid.gz	proteomic profiling by mass spectrometry	F001235.mzid.gz	1	6	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion Lumos;AC=MS:1002732	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
+PXD054720-sample	escherichia coli	not applicable	not applicable	not applicable	1	synthetic	reference	not available	26.4 Å	not available	not available	F001236.mzid.gz	proteomic profiling by mass spectrometry	F001236.mzid.gz	1	7	AC=MS:1002038;NT=label free sample	NT=Orbitrap Fusion Lumos;AC=MS:1002732	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	chemical cross-linking coupled with mass spectrometry proteomics	NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0
Evidence
The cited rows use unmistakable support, reference, result, checksum, archive, or document file
extensions as assay data in both assay name and comment[data file]. This conflicts with Rule 3
and the repository criteria requiring raw-file mappings and assay relationships to reflect archive
evidence and comment[data file] entries to correspond to deposited raw-category runs.

AGENTS.md: Align SDRF Metadata with Public Archive Evidence
datasets/PXD054720/PXD054720.sdrf.tsv[2-10]
datasets/PXD051047/PXD051047.sdrf.tsv[71-71]
datasets/PXD051348/PXD051348.sdrf.tsv[47-47]
datasets/PXD051971/PXD051971.sdrf.tsv[66-71]
datasets/PXD052310/PXD052310.sdrf.tsv[50-52]
datasets/PXD052694/PXD052694.sdrf.tsv[3-4]
datasets/PXD052926/PXD052926.sdrf.tsv[26-26]
datasets/PXD053636/PXD053636.sdrf.tsv[29-34]
datasets/PXD054720/PXD054720.sdrf.tsv[5-10]
docs/heart-tier2/CRITERIA.md[84-86]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Several SDRFs incorrectly represent support, database, result, checksum, spreadsheet, archive, and document files as independent mass-spectrometry acquisition assays rather than limiting assay rows to deposited experimental run files.

## Fix Focus Areas
- datasets/PXD051047/PXD051047.sdrf.tsv[71-71]
- datasets/PXD051348/PXD051348.sdrf.tsv[47-47]
- datasets/PXD051971/PXD051971.sdrf.tsv[66-71]
- datasets/PXD052310/PXD052310.sdrf.tsv[50-52]
- datasets/PXD052694/PXD052694.sdrf.tsv[3-24]
- datasets/PXD052926/PXD052926.sdrf.tsv[26-26]
- datasets/PXD053636/PXD053636.sdrf.tsv[3-34]
- datasets/PXD054720/PXD054720.sdrf.tsv[5-10]

## Recommended Fix
Remove rows whose data-file values are ancillary files rather than deposited experimental runs. Retain one row per valid selected-format run, preserve the correct relationships to the raw acquisitions, and verify that every retained run maps to its correct sample.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


8. Yeast samples use an animal template ✓ Resolved 📘 Rule violation ≡ Correctness
Description
The rows declare saccharomyces cerevisiae while selecting NT=invertebrates as the organism
template. Every assay in the accession inherits this fungal-organism and animal-template
contradiction.
Code

datasets/PXD051047/PXD051047.sdrf.tsv[2]

+PXD051047-sample	saccharomyces cerevisiae	not applicable	not applicable	not applicable	1	synthetic	reference	not available	not available	not available	not available	HF1_MS_16313_E6_3_02062022	proteomic profiling by mass spectrometry	HF1_MS_16313_E6_3_02062022.raw	1	1	AC=MS:1002038;NT=label free sample	NT=Q Exactive HF;AC=MS:1002523	NT=Data-dependent acquisition;AC=PRIDE:0000449	NT=Trypsin;AC=MS:1001251	NT=Lys-C;AC=MS:1001309	NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed	NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable	NT=HCD;AC=PRIDE:0000590	not available	not available	not available	not available	cross-linking mass spectrometry	NT=unknown crosslinker;AC=XLMOD:00000	not available	not available	not available	not available	not available	v1.1.0	NT=ms-proteomics;VV=v1.1.0	NT=crosslinking;VV=v1.0.0	NT=invertebrates;VV=v1.1.0
Evidence
Rule 4 requires ontology terms and template layers to match the dataset, while the added row
combines a fungal organism with the invertebrates template.

AGENTS.md: Use Ontology Terms and Columns Consistent with the Selected SDRF Template
datasets/PXD051047/PXD051047.sdrf.tsv[2-2]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The dataset identifies its samples as yeast but assigns the invertebrates template, producing internally inconsistent metadata.

## Fix Focus Areas
- datasets/PXD051047/PXD051047.sdrf.tsv[2-75]

## Recommended Fix
Replace the invertebrates template annotation with the applicable fungal or default template supported by the project, then validate the complete SDRF.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context sources
Review mode: 🧠 Deep: This adds 50 independently authored SDRF datasets with substantial metadata logic, and prior reviews already reveal multiple distinct annotation defects across unrelated accessions and fields.

Grey Divider

Tip of the day
💡 Did you know, you can enable the Remediation agent and Qodo fixes findings in a dedicated fix PR

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread datasets/PXD051047/PXD051047.sdrf.tsv Outdated
Comment thread datasets/PXD051971/PXD051971.sdrf.tsv Outdated
PXD051971-sample saccharomyces cerevisiae not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available 20210519_H2A_H2B_Ubp10_DSBU_Trp_1.zhrm proteomic profiling by mass spectrometry 20210519_H2A_H2B_Ubp10_DSBU_Trp_1.zhrm 1 2 AC=MS:1002038;NT=label free sample NT=Q Exactive HF-X;AC=MS:1002877 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03 not available not available the not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=invertebrates;VV=v1.1.0
PXD051971-sample saccharomyces cerevisiae not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available 20210519_H2A_H2B_Ubp10_DSBU_Trp_2 proteomic profiling by mass spectrometry 20210519_H2A_H2B_Ubp10_DSBU_Trp_2.raw 1 3 AC=MS:1002038;NT=label free sample NT=Q Exactive HF-X;AC=MS:1002877 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03 not available not available the not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=invertebrates;VV=v1.1.0
PXD051971-sample saccharomyces cerevisiae not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available 20210519_H2A_H2B_Ubp10_DSBU_Trp_2.zhrm proteomic profiling by mass spectrometry 20210519_H2A_H2B_Ubp10_DSBU_Trp_2.zhrm 1 4 AC=MS:1002038;NT=label free sample NT=Q Exactive HF-X;AC=MS:1002877 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03 not available not available the not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=invertebrates;VV=v1.1.0
PXD051971-sample saccharomyces cerevisiae not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available 20210519_H2A_H2B_Ubp10_EDC_Trp_1 proteomic profiling by mass spectrometry 20210519_H2A_H2B_Ubp10_EDC_Trp_1.raw 1 5 AC=MS:1002038;NT=label free sample NT=Q Exactive HF-X;AC=MS:1002877 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03 not available not available the not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=invertebrates;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

2. Edc assays are annotated as dsbu 📘 Rule violation ≡ Correctness

The EDC-named assay rows in PXD051971 assign the DSBU term NT=DSBU;AC=XLMOD:02043, target
residues, and mass values as their cross-linker parameters. This mismatch affects EDC Trp, Ubiq,
GluC, and GC assays and their associated files, exposing parameters for a different cross-linking
reaction to downstream searches.
Agent Prompt
## Issue description
PXD051971 assigns DSBU chemistry, including its ontology term, target residues, and mass values, to runs explicitly identified as EDC experiments.

## Fix Focus Areas
- datasets/PXD051971/PXD051971.sdrf.tsv[6-67]

## Recommended Fix
Identify every EDC row and replace the DSBU cross-linker term, DSBU-specific mass values, and target attributes with the correct EDC annotation supported by the study metadata. Retain DSBU only for assays whose archive names and evidence identify DSBU.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment thread datasets/PXD051971/PXD051971.sdrf.tsv Outdated
@@ -0,0 +1,71 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD051971-sample saccharomyces cerevisiae not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available 20210519_H2A_H2B_Ubp10_DSBU_Trp_1 proteomic profiling by mass spectrometry 20210519_H2A_H2B_Ubp10_DSBU_Trp_1.raw 1 1 AC=MS:1002038;NT=label free sample NT=Q Exactive HF-X;AC=MS:1002877 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSBU;AC=XLMOD:02043;CL=yes;TA=K,S,T,Y,nterm;MH=85.05;ML=111.03 not available not available the not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=invertebrates;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

3. Yeast assay rows have invalid metadata 📘 Rule violation ≡ Correctness

Eight added SDRFs populate comment[quenching reagent] with unsupported fragments or generic terms
such as the, by, of, solution, and media, while PXD051971 and PXD053636 also pair
Saccharomyces cerevisiae with the animal-specific invertebrates template. The malformed
annotations repeat across the affected assays—including APEX2, PhoX, and DSSO runs—and, for the two
yeast accessions, their associated result-file rows, with the sole PXD053607 assay reaching the
canonical dataset unchanged.
Agent Prompt
## Issue description
Eight datasets place sentence fragments or generic terms such as `the`, `by`, `of`, `solution`, and `media` in `comment[quenching reagent]`; PXD051971 and PXD053636 additionally assign yeast samples to the invertebrates template.

## Fix Focus Areas
- datasets/PXD051971/PXD051971.sdrf.tsv[2-71]
- datasets/PXD052310/PXD052310.sdrf.tsv[2-52]
- datasets/PXD052930/PXD052930.sdrf.tsv[2-49]
- datasets/PXD053578/PXD053578.sdrf.tsv[2-47]
- datasets/PXD053607/PXD053607.sdrf.tsv[2-2]
- datasets/PXD053636/PXD053636.sdrf.tsv[2-35]
- datasets/PXD053984/PXD053984.sdrf.tsv[2-23]
- datasets/PXD054551/PXD054551.sdrf.tsv[2-36]

## Recommended Fix
Replace every malformed quenching-reagent value with the actual archive-supported reagent and applicable ontology mapping based on study evidence, or use `not available` when no reagent can be established. For PXD051971 and PXD053636, remove the invertebrates template and select the repository-supported fungal or yeast template when applicable, then validate every SDRF row and associated result-file row.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@@ -0,0 +1,16 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[age] characteristics[sex] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD052801-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available not available not available not available 062123_cell_lines_DIA_24mz_HEK293_1.mzML proteomic profiling by mass spectrometry 062123_cell_lines_DIA_24mz_HEK293_1.mzML 1 1 AC=MS:1002038;NT=label free sample NT=Q Exactive;AC=MS:1001911 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=PIR;AC=XLMOD:02014 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

4. Cell-line acquisitions are mislabeled 📘 Rule violation ≡ Correctness

The comment[proteomics data acquisition method] values in PXD052801, PXD052821, PXD052867, and
PXD053509 declare Data-dependent acquisition even though the corresponding assay, raw-file, or
quantitative-file names identify DIA or SpDIA runs. The conflict affects the HEK293, HeLa, Lumos,
and added quantitative-file rows, causing consumers of the acquisition-method field to classify
independent-acquisition experiments as dependent-acquisition data.
Agent Prompt
## Issue description
Four datasets contain files identified by assay, raw-file, or quantitative-file names as DIA or SpDIA runs, but their SDRF acquisition-method fields annotate them as data-dependent acquisitions.

## Fix Focus Areas
- datasets/PXD052801/PXD052801.sdrf.tsv[2-16]
- datasets/PXD052821/PXD052821.sdrf.tsv[2-19]
- datasets/PXD052867/PXD052867.sdrf.tsv[3-40]
- datasets/PXD053509/PXD053509.sdrf.tsv[2-35]

## Recommended Fix
Replace the data-dependent acquisition mapping with the repository's supported data-independent acquisition ontology mapping for each row confirmed by its filename as a DIA or SpDIA run. Limit the update to confirmed DIA or SpDIA rows and verify every changed mapping against the public archive metadata.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment on lines +5 to +8
PXD054720-sample escherichia coli not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available ABRF_iPRG_XL_2023.fasta proteomic profiling by mass spectrometry ABRF_iPRG_XL_2023.fasta 1 4 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD054720-sample escherichia coli not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available F001234.mzid.gz proteomic profiling by mass spectrometry F001234.mzid.gz 1 5 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD054720-sample escherichia coli not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available F001235.mzid.gz proteomic profiling by mass spectrometry F001235.mzid.gz 1 6 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD054720-sample escherichia coli not applicable not applicable not applicable 1 synthetic reference not available 26.4 Å not available not available F001236.mzid.gz proteomic profiling by mass spectrometry F001236.mzid.gz 1 7 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=DSSO;AC=XLMOD:02010;CL=yes;TA=K,S,T,Y,nterm;MH=54.01;ML=85.98 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

5. Ancillary files become acquisition assays 📘 Rule violation ≡ Correctness

Multiple SDRFs map FASTA databases, identification results, checksums, spreadsheets, archives, and
documents into both assay name and comment[data file] instead of limiting those fields to
deposited raw-category runs. Whenever these ancillary files are included, each is treated as a
separate proteomics acquisition assay, creating nonexistent experimental runs and incorrect
sample-to-run relationships alongside the actual raw files.
Agent Prompt
## Issue description
Several SDRFs incorrectly represent support, database, result, checksum, spreadsheet, archive, and document files as independent mass-spectrometry acquisition assays rather than limiting assay rows to deposited experimental run files.

## Fix Focus Areas
- datasets/PXD051047/PXD051047.sdrf.tsv[71-71]
- datasets/PXD051348/PXD051348.sdrf.tsv[47-47]
- datasets/PXD051971/PXD051971.sdrf.tsv[66-71]
- datasets/PXD052310/PXD052310.sdrf.tsv[50-52]
- datasets/PXD052694/PXD052694.sdrf.tsv[3-24]
- datasets/PXD052926/PXD052926.sdrf.tsv[26-26]
- datasets/PXD053636/PXD053636.sdrf.tsv[3-34]
- datasets/PXD054720/PXD054720.sdrf.tsv[5-10]

## Recommended Fix
Remove rows whose data-file values are ancillary files rather than deposited experimental runs. Retain one row per valid selected-format run, preserve the correct relationships to the raw acquisitions, and verify that every retained run maps to its correct sample.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

PXD051261-sample streptococcus pyogenes mgas315 not applicable not applicable not applicable 1 synthetic reference not available not available not available not available DT_M2109_366 proteomic profiling by mass spectrometry DT_M2109_366.raw 1 64 AC=MS:1002038;NT=label free sample NT=Q Exactive HF-X;AC=MS:1002877 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD051261-sample streptococcus pyogenes mgas315 not applicable not applicable not applicable 1 synthetic reference not available not available not available not available DT_M2109_367 proteomic profiling by mass spectrometry DT_M2109_367.raw 1 65 AC=MS:1002038;NT=label free sample NT=Q Exactive HF-X;AC=MS:1002877 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD051261-sample streptococcus pyogenes mgas315 not applicable not applicable not applicable 1 synthetic reference not available not available not available not available MSinjection_rawfile_to_biological_sample_information.csv proteomic profiling by mass spectrometry MSinjection_rawfile_to_biological_sample_information.csv 1 66 AC=MS:1002038;NT=label free sample NT=Q Exactive HF-X;AC=MS:1002877 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0
PXD051261-sample streptococcus pyogenes mgas315 not applicable not applicable not applicable 1 synthetic reference not available not available not available not available P_2S_tPA_1mM_DSS_1_DT_C2203_101 proteomic profiling by mass spectrometry P_2S_tPA_1mM_DSS_1_DT_C2203_101.raw 1 67 AC=MS:1002038;NT=label free sample NT=Q Exactive HF-X;AC=MS:1002877 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available chemical cross-linking coupled with mass spectrometry proteomics NT=unknown crosslinker;AC=XLMOD:00000 not available not available not available not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

6. Known crosslinkers become unknown 🐞 Bug ≡ Correctness

PXD051261 assigns NT=unknown crosslinker;AC=XLMOD:00000 to files whose names explicitly identify
DSS and DSG. The mismatch affects both reagent groups and prevents consumers from selecting the
corresponding crosslinking chemistry.
Agent Prompt
## Issue description
PXD051261 marks DSS and DSG runs as using an unknown crosslinker even though their filenames identify the reagents.

## Fix Focus Areas
- datasets/PXD051261/PXD051261.sdrf.tsv[68-99]

## Recommended Fix
Replace the unknown-crosslinker values on DSS and DSG rows with the appropriate XLMOD terms and reagent-specific metadata, using archive evidence to separate the two groups.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

…lumns

Required by the vertebrates/invertebrates/plants SDRF templates; value set
to the spec-compliant reserved word 'not available' where the field was
not previously populated. human-only files are unaffected (field optional
in that template).
@ypriverol ypriverol closed this Sep 17, 2026
@ypriverol ypriverol reopened this Sep 17, 2026
@github-actions

github-actions Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

SDRF change report

50 new · 0 modified · 0 deleted · highest risk: none

External reviewer notes

Quoted from AI review bots on this PR. Not verified by this report unless marked as also flagged.

  • PXD051047, PXD051348, PXD051971, PXD052310, PXD052694, PXD052926, PXD053636, PXD054720 · qodo-code-review[bot]: Ancillary files become acquisition assays. Multiple SDRFs map FASTA databases, identification results, checksums, spreadsheets, archives, and documents into both assay name and comment[data file] instead of limiting those fields to deposited raw-category runs. (source)
  • PXD051261 · qodo-code-review[bot]: Known crosslinkers become unknown. PXD051261 assigns NT=unknown crosslinker;AC=XLMOD:00000 to files whose names explicitly identify DSS and DSG. The mismatch affects both reagent groups and prevents consumers from selecting the corresponding crosslinking chemistry. (source)
  • PXD051971 · qodo-code-review[bot]: EDC assays are annotated as DSBU. The EDC-named assay rows in PXD051971 assign the DSBU term NT=DSBU;AC=XLMOD:02043 , target residues, and mass values as their cross-linker parameters. (source)
  • PXD051971, PXD052310, PXD052930, PXD053578, PXD053607, PXD053636, PXD053984, PXD054551 · qodo-code-review[bot]: Yeast assay rows have invalid metadata. Eight added SDRFs populate comment[quenching reagent] with unsupported fragments or generic terms such as the , by , of , solution , and media , while PXD051971 and PXD053636 also pair Saccharomyces cerevisiae with the animal-specific invertebrates template. (source)
  • PXD052552, PXD052694 · qodo-code-review[bot]: Cattle samples omit required breed. The PXD052552 and PXD052694 headers declare the vertebrates template but omit characteristics[strain or breed] between their sample characteristics. (source)
  • PXD052801, PXD052821, PXD052867, PXD053509 · qodo-code-review[bot]: Cell-line acquisitions are mislabeled. The comment[proteomics data acquisition method] values in PXD052801, PXD052821, PXD052867, and PXD053509 declare Data-dependent acquisition even though the corresponding assay, raw-file, or quantitative-file names identify DIA or SpDIA runs. (source)
  • PXD053010 · qodo-code-review[bot]: EDC runs are indexed as diazirine. PXD053010 labels every EDC-named assay as NT=EDC but assigns AC=XLMOD:02009 , an accession the repository associates with NT=diazirine . (source)
New datasets (50)

parse_sdrf validation of new datasets is reported by the SDRF review gate check.

Dataset Rows Defects
PXD051014 83 no_factor_value: 1
PXD051047 74 no_factor_value: 1
PXD051143 6 no_factor_value: 1
PXD051261 99 no_factor_value: 1
PXD051348 77 no_factor_value: 1
PXD051405 93 no_factor_value: 1
PXD051493 57 no_factor_value: 1
PXD051557 12 no_factor_value: 1
PXD051602 8 no_factor_value: 1
PXD051693 14 no_factor_value: 1
PXD051742 22 no_factor_value: 1
PXD051886 38 no_factor_value: 1
PXD051971 70 no_factor_value: 1
PXD052310 51 no_factor_value: 1
PXD052552 20 no_factor_value: 1
PXD052623 24 no_factor_value: 1
PXD052624 21 no_factor_value: 1
PXD052637 8 no_factor_value: 1
PXD052687 7 no_factor_value: 1
PXD052694 24 no_factor_value: 1, peak_list_data_file: 12
PXD052745 57 no_factor_value: 1
PXD052746 45 no_factor_value: 1
PXD052801 15 no_factor_value: 1, peak_list_data_file: 15
PXD052821 18 no_factor_value: 1
PXD052825 9 no_factor_value: 1
PXD052867 39 no_factor_value: 1
PXD052917 45 no_factor_value: 1
PXD052923 1 no_factor_value: 1
PXD052926 26 no_factor_value: 1
PXD052930 48 no_factor_value: 1
PXD053010 19 no_factor_value: 1
PXD053341 18 no_factor_value: 1
PXD053452 10 no_factor_value: 1
PXD053489 11 no_factor_value: 1
PXD053494 88 no_factor_value: 1
PXD053509 33 no_factor_value: 1
PXD053578 45 no_factor_value: 1
PXD053607 1 no_factor_value: 1
PXD053636 33 no_factor_value: 1, peak_list_data_file: 9
PXD053760 3 no_factor_value: 1
PXD053832 14 no_factor_value: 1
PXD053924 7 no_factor_value: 1
PXD053984 21 no_factor_value: 1
PXD054003 30 no_factor_value: 1
PXD054140 2 no_factor_value: 1
PXD054141 6 no_factor_value: 1
PXD054249 5 no_factor_value: 1
PXD054551 35 no_factor_value: 1
PXD054616 6 no_factor_value: 1
PXD054720 9 no_factor_value: 1

Advisory report built from 4311088. Risk labels do not block merging.

@github-actions github-actions Bot added the sdrf:new SDRF PR adds new datasets label Sep 17, 2026
@@ -0,0 +1,21 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[developmental stage] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

1. Cattle samples omit required breed 🐞 Bug ≡ Correctness

The PXD052552 and PXD052694 headers declare the vertebrates template but omit
characteristics[strain or breed] between their sample characteristics. Every added Bos taurus and
Mus musculus row therefore lacks the required field and cannot record breed or strain information,
even with the reserved not available value.
Agent Prompt
## Issue description
The PXD052552 and PXD052694 vertebrate SDRFs omit the required `characteristics[strain or breed]` column and corresponding row values.

## Fix Focus Areas
- datasets/PXD052552/PXD052552.sdrf.tsv[1-21]
- datasets/PXD052694/PXD052694.sdrf.tsv[1-25]

## Recommended Fix
Add `characteristics[strain or breed]` to each header alongside the other organism characteristics and add a value to every row. Use `not available` when the cattle breed or mouse strain is unknown.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@@ -0,0 +1,20 @@
source name characteristics[organism] characteristics[organism part] characteristics[cell type] characteristics[disease] characteristics[age] characteristics[sex] characteristics[biological replicate] characteristics[material type] characteristics[sample type] characteristics[enrichment process] characteristics[crosslink distance] characteristics[crosslinking reaction time] characteristics[crosslinking temperature] assay name technology type comment[data file] comment[technical replicate] comment[fraction identifier] comment[label] comment[instrument] comment[proteomics data acquisition method] comment[cleavage agent details] comment[cleavage agent details] comment[modification parameters] comment[modification parameters] comment[dissociation method] comment[collision energy] comment[precursor mass tolerance] comment[fragment mass tolerance] comment[fractionation method] comment[chemical cross-linking coupled with ms] comment[cross-linker] comment[crosslink enrichment method] comment[crosslinker concentration] comment[quenching reagent] comment[reduction reagent] comment[alkylation reagent] comment[sdrf version] comment[sdrf template] comment[sdrf template] comment[sdrf template]
PXD053010-sample homo sapiens not applicable not applicable not applicable not available not applicable 1 synthetic reference not available 11.4 Å not available not available Zou_Rappsilber_AC_NF90-NF45-RNA_EDC_S1_R1 proteomic profiling by mass spectrometry Zou_Rappsilber_AC_NF90-NF45-RNA_EDC_S1_R1.raw 1 1 AC=MS:1002038;NT=label free sample NT=Orbitrap Fusion Lumos;AC=MS:1002732 NT=Data-dependent acquisition;AC=PRIDE:0000449 NT=Trypsin;AC=MS:1001251 NT=Lys-C;AC=MS:1001309 NT=Carbamidomethyl;AC=UNIMOD:4;TA=C;MT=Fixed NT=Oxidation;AC=UNIMOD:35;TA=M;MT=Variable NT=HCD;AC=PRIDE:0000590 not available not available not available not available cross-linking mass spectrometry NT=EDC;AC=XLMOD:02009;CL=no;TA=K,D,E not available not available using not available not available v1.1.0 NT=ms-proteomics;VV=v1.1.0 NT=crosslinking;VV=v1.0.0 NT=human;VV=v1.1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

2. Edc runs are indexed as diazirine 🐞 Bug ≡ Correctness

PXD053010 labels every EDC-named assay as NT=EDC but assigns AC=XLMOD:02009, an accession the
repository associates with NT=diazirine. All 19 added EDC assay rows therefore expose diazirine
chemistry to consumers filtering or interpreting cross-linker annotations.
Agent Prompt
## Issue description
PXD053010 assigns the diazirine accession `XLMOD:02009` to assays explicitly identified as EDC, so downstream consumers receive the wrong cross-linker identity.

## Fix Focus Areas
- datasets/PXD053010/PXD053010.sdrf.tsv[2-20]

## Recommended Fix
Replace the cross-linker annotation in every PXD053010 row with the validated XLMOD accession and associated attributes for EDC. Recheck the linked cross-link distance and target-residue metadata so they describe the corrected reagent consistently.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@qodo-code-review

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 2f4ccff

comment[sdrf template] declared NT=invertebrates;VV=v1.1.0 (an animal-only
template) for 3 Saccharomyces cerevisiae datasets. Removing the mismatched
template column; ms-proteomics and crosslinking layers are unaffected.

Confirmed by qodo-code-review[bot] and this report's own data check.
@ypriverol
ypriverol merged commit a5f4e28 into main Sep 17, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

sdrf:new SDRF PR adds new datasets

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants