conversion run

auris-scf1

paper-to-galaxy — 2026-09-16 16:09 to 2026-09-18 16:41, 61.8 MB on disk.

paper-to-galaxyrev 2pipeline-paper-to-galaxyreconstructed--feedback
run
running
the record last said running
phases
10/12
furthest reached was phase 11
artifacts
13/15
1 declared artifact(s) due but absent
obligations
15 open
13 resolved, 0 surrendered, 0 blocking
feedback
61
1 blocker, 46 major, 14 minor
unmapped
11
61.2 MB no Mold declares

Phases

where the run got to
  1. 01summarize-paperdone · from feedback ledger
  2. 02freeform-summary-to-galaxy-interfacedone · from feedback ledger
  3. 03freeform-summary-to-galaxy-data-flowdone · from feedback ledger
  4. 04compare-against-iwc-exemplardone · from feedback ledger
  5. 05freeform-summary-to-galaxy-templatedone · from feedback ledger
  6. 06advance-galaxy-draft-steploop ×26done · from feedback ledger
  7. 07test-data-resolution→ paper-to-test-datadone · from feedback ledger
  8. 08freeform-summary-to-galaxy-test-plandone · from feedback ledger
  9. 09implement-galaxy-workflow-testdone · from feedback ledger
  10. 10validate-galaxy-workflowdone · from feedback ledger
  11. 11run-workflow-testrunning · from feedback ledger
  12. 12debug-galaxy-workflow-outputpending · from feedback ledger

Artifacts

what it declared and what is on disk
phaseartifactfilekindpresencesizemodifiedcopies
phase 1freeform-summaryfreeform-summary.mdmarkdownpresent15.1 KB2026-09-16 16:09
phase 2freeform-galaxy-interfacefreeform-galaxy-interface.mdmarkdownpresent14.5 KB2026-09-16 16:15
phase 2open-requirements-ledgeropen-requirements.ledger.ymlyamlpresent71.1 KB2026-09-17 21:06
phase 3freeform-galaxy-data-flowfreeform-galaxy-data-flow.mdmarkdownpresent28.9 KB2026-09-16 16:25
phase 4iwc-comparison-notesiwc-comparison-notes.mdmarkdownpresent24.0 KB2026-09-16 16:38
phase 4iwc-exemplar-gxformat2iwc-exemplar.gxwf.ymlyamlpresent20.0 KB2026-09-16 16:36
phase 5galaxy-workflow-draftgalaxy-workflow-draft.gxwf.ymlyamlpresent103.9 KB2026-09-16 19:36
phase 6galaxy-workflowgalaxy-workflow.gxwf.ymlyamlpresent42.6 KB2026-09-17 14:19
phase 7test-data-refstest-data-refs.jsonjsonpresent29.2 KB2026-09-16 19:56
phase 8galaxy-test-plangalaxy-test-plan.ymlyamlpresent71.6 KB2026-09-16 21:16
phase 9galaxy-workflow-testgalaxy-workflow.gxwf-tests.ymlyamlpresent32.3 KB2026-09-16 21:41
phase 10galaxy-workflow-validation-resultgalaxy-workflow-validation-result.jsonjsonpresent20.1 KB2026-09-17 22:31
phase 11workflow-test-resultworkflow-test-result.jsonjsonmissing
phase 12workflow-debug-reportworkflow-debug-report.mdmarkdownnot-yet-due
phase —foundry-feedback-ledgerfoundry-feedback.ledger.ymlyamlpresent169.9 KB2026-09-17 21:05
phase —foundry-run-manifestfoundry-run.ymlyamloptional-absent
foundry-feedback-ledgerpresentfoundry-feedback.ledger.yml

Runtime artifact initialized by the harness ([[foundry-feedback-ledger]]).

declared by
— (phase —)
consumed at
nothing downstream reads it
schema
none declared
sha256
f6c7c3e19455434cba4e11c2e99c06952897cc85a41cd17585c31e527587e18a
run:
  pipeline: paper-to-galaxy
  run_slug: auris-scf1
  status: running
  phases:
    - { n: 1, kind: mold, skill: summarize-paper, status: done, feedback_checked: true }
    - { n: 2, kind: mold, skill: freeform-summary-to-galaxy-interface, status: done, feedback_checked: true }
    - { n: 3, kind: mold, skill: freeform-summary-to-galaxy-data-flow, status: done, feedback_checked: true }
    - { n: 4, kind: mold, skill: compare-against-iwc-exemplar, status: done, feedback_checked: true }
    - { n: 5, kind: mold, skill: freeform-summary-to-galaxy-template, status: done, feedback_checked: true }
    - { n: 6, kind: mold, skill: advance-galaxy-draft-step, status: done, feedback_checked: true, iterations: 26 }
    - { n: 7, kind: branch, pattern: test-data-resolution, status: done, feedback_checked: true, selected: paper-to-test-data }
    - { n: 8, kind: mold, skill: freeform-summary-to-galaxy-test-plan, status: done, feedback_checked: true }
    - { n: 9, kind: mold, skill: implement-galaxy-workflow-test, status: done, feedback_checked: true }
    - { n: 10, kind: mold, skill: validate-galaxy-workflow, status: done, feedback_checked: true }
    - { n: 11, kind: mold, skill: run-workflow-test, status: running }
    - { n: 12, kind: mold, skill: debug-galaxy-workflow-output, status: pending }
entries:
  - id: summarize-paper-declares-no-freeform-summary-shape
    raised_by: summarize-paper
    observed_in:
      mold: summarize-paper
      path: content/molds/summarize-paper/index.md
      revision: 2
      content_hash: ac601516a655a5a97bf9cc6d865839622efd509a295f99410c9e47e3952f269d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: summarize-paper
      locator: content/molds/summarize-paper/index.md
      content_hash: ac601516a655a5a97bf9cc6d865839622efd509a295f99410c9e47e3952f269d
    kind: gap
    severity: major
    what: >-
      The Mold produces the shared freeform-summary handoff but specifies no shape for it.
      The procedure says only "a free-form Markdown summary capturing the workflow's steps,
      tools, parameters, and sample/reference-data leads", the cast bundle packages no
      template, example, or reference, and the runtime notes forbid reading Foundry source
      at runtime. Four downstream Molds consume this artifact
      (freeform-summary-to-galaxy-interface, -data-flow, -template, -test-plan), so the
      section structure they can rely on was guessed at, not settled by the instructions.
    expected: >-
      Package a freeform-summary skeleton or worked example in the cast bundle, or state in
      the procedure the minimum sections every consumer can assume (source identity,
      candidate workflow scope, per-step tool/parameter table, reference data, sample data
      accessions, assumptions, open questions). interview-to-freeform-summary emits the
      same artifact id and should share whatever contract is written.
    evidence: >-
      Bundle carries SKILL.md plus _feedback/_provenance/_verify only; refs is empty and
      "Load Upfront"/"Load On Demand" are both "None declared". Output section names the
      filename and format but no structure.
    status: open
    issue: null
  - id: summarize-paper-silent-on-supplementary-methods-retrieval
    raised_by: summarize-paper
    observed_in:
      mold: summarize-paper
      path: content/molds/summarize-paper/index.md
      revision: 2
      content_hash: ac601516a655a5a97bf9cc6d865839622efd509a295f99410c9e47e3952f269d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: summarize-paper
      locator: content/molds/summarize-paper/index.md
      content_hash: ac601516a655a5a97bf9cc6d865839622efd509a295f99410c9e47e3952f269d
    kind: gap
    severity: major
    what: >-
      The Mold's entire job is to extract methods from a paper, but it declares no required
      tools and describes no procedure for actually obtaining the text. For this run the
      article's main text contained no Methods section at all: every computational detail
      lived only in a supplementary PDF. The publisher page returned HTTP 403, the article
      is not in the Europe PMC open-access set, and a single fetch of the PMC article page
      yielded a tool list with no parameters and no pipeline. The retrieval route that
      worked was improvised and is not described anywhere in the bundle.
    expected: >-
      State in the procedure that Science/Nature-style papers carry their Methods in
      supplementary materials and that the main text alone is not a sufficient source, and
      package a reference covering the retrieval routes worth trying in order (NCBI eutils
      efetch db=pmc for JATS full text, the supplementary-material file list inside that
      XML, Europe PMC, preprint and institutional-repository copies), plus the instruction
      to report inaccessible Methods as an open question rather than summarize around them.
    evidence: >-
      Main-text JATS body ended at a bare "Supplementary Material" label; the full methods,
      tool parameters, and reference-assembly accession came only from the supplement PDF
      listed in that XML. Artifact open question 10 records the same for the next reader.
    status: open
    issue: null
  - id: testability-note-silent-on-sample-sheet-test-fixtures
    raised_by: freeform-summary-to-galaxy-interface
    observed_in:
      mold: freeform-summary-to-galaxy-interface
      path: content/molds/freeform-summary-to-galaxy-interface/index.md
      revision: 3
      content_hash: de0cb02532205ad18f4aa1ab2881cac725a3191b448468a8a11d0663ff1ec97d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-workflow-testability-design
      locator: content/research/galaxy-workflow-testability-design/index.md
      content_hash: 81b01801ec3d14810d570dd2d7815da2bb11e41a562af9182c29fd2974f4b855
    kind: gap
    severity: major
    what: >-
      Section 5 ("Design inputs with fixtures in mind") instructs the designer to match workflow
      input collection types to realistic fixture shapes and to choose labels readable as test
      `job:` keys, and its evidence covers `list`, `list:paired`, data, string, boolean and int
      inputs. It says nothing about the `sample_sheet` family, even though the sibling note
      galaxy-sample-sheet-collections is packaged in the same bundle and is what pushes the
      designer toward that shape. Nothing in the bundle states whether a `sample_sheet:paired`
      workflow input — element identifiers plus per-row typed `columns` plus collection-level
      `column_definitions` — can be expressed in a Planemo/IWC `-tests.yml` job block at all. This
      run's most consequential interface decision, the primary reads input, had to be made without
      knowing whether the same run's own test phase can express it.
    expected: >-
      Add a rule and evidence to section 5 covering sample_sheet-family inputs: either a corpus or
      Planemo-schema citation showing the job-block fixture syntax for `column_definitions` and
      per-row `columns`, or an explicit statement that no such fixture form exists yet and that a
      sample_sheet input therefore trades testability for metadata carriage. Either answer settles
      the decision; the absence of both leaves it a guess. The sample_sheet note's "Edges to flag"
      list would be the natural second home for the same statement.
    evidence: >-
      Bundle packages galaxy-sample-sheet-collections (which documents the workflow YAML
      `column_definitions` form) and galaxy-workflow-testability-design (which documents fixture
      design) with no overlap on this point; the note that would cover it,
      iwc-test-data-conventions, is cross-referenced from section 5 but is not packaged. The
      interface brief at <run>/freeform-galaxy-interface.md settles on `sample_sheet:paired` and
      carries open-requirements entry `sample-sheet-input-test-fixture-expressibility` plus a named
      `list:paired` fallback precisely because the question could not be answered.
    status: open
    issue: null
  - id: paper-to-galaxy-path-has-no-reference-data-owner
    raised_by: freeform-summary-to-galaxy-interface
    observed_in:
      mold: freeform-summary-to-galaxy-interface
      path: content/molds/freeform-summary-to-galaxy-interface/index.md
      revision: 3
      content_hash: de0cb02532205ad18f4aa1ab2881cac725a3191b448468a8a11d0663ff1ec97d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: pipeline
      label: paper-to-galaxy
      locator: content/pipelines/paper-to-galaxy/index.md
      content_hash: null
    kind: gap
    severity: major
    what: >-
      The Nextflow path has a dedicated Mold for deciding the Galaxy-side shape of external
      reference data (nextflow-summary-to-galaxy-reference-data, listed among the producers of the
      open-requirements ledger). The paper path has no equivalent, and this run's phase roster
      contains no reference-data phase. The source paper nonetheless pins a specific non-model
      assembly (GCA_002759435.2, C. auris B8441) and requires an annotation file, so the decision
      between a built-in index, a data-table string, and a portable history dataset had to be made
      somewhere. The interface Mold made it, unprompted: nothing in its procedure or references
      assigns reference-data shape to the interface tier.
    expected: >-
      Either add a reference-data phase to the paper-to-galaxy pipeline (a freeform sibling of
      nextflow-summary-to-galaxy-reference-data, since the input is a narrative accession rather
      than an igenomes-style parameter), or state explicitly in
      freeform-summary-to-galaxy-interface that reference-data delivery shape is its call and
      package the guidance that decision needs. Today the decision is unowned, which means it gets
      made by whichever Mold notices, with no shared basis.
    evidence: >-
      Run roster phases 1-12 in this ledger contain no `*-to-galaxy-reference-data` step. The
      interface brief settles the genome as a history fasta on portability grounds and carries
      open-requirements entry `reference-genome-delivery-shape-unverified` recording that no
      built-in-index check was performed, because no phase owns performing one. The subject's
      content hash is null: this Mold's runtime notes forbid reading Foundry source, so the
      pipeline note could not be hashed from inside the run.
    status: open
    issue: null
  - id: interface-mold-gives-no-rule-for-parameter-exposure
    raised_by: freeform-summary-to-galaxy-interface
    observed_in:
      mold: freeform-summary-to-galaxy-interface
      path: content/molds/freeform-summary-to-galaxy-interface/index.md
      revision: 3
      content_hash: de0cb02532205ad18f4aa1ab2881cac725a3191b448468a8a11d0663ff1ec97d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: freeform-summary-to-galaxy-interface
      locator: content/molds/freeform-summary-to-galaxy-interface/index.md
      content_hash: de0cb02532205ad18f4aa1ab2881cac725a3191b448468a8a11d0663ff1ec97d
    kind: gap
    severity: minor
    what: >-
      The procedure names inputs, outputs, labels, collection shapes and checkpoints as the things
      to settle, but gives no rule for which source-stated parameters become exposed typed workflow
      inputs and which are baked into step defaults. The only nearby guidance is one line in the
      packaged testability note ("Keep typed parameters explicit when tests need to set them"),
      which answers the test-facing half and not the design half. For a paper source the
      distinction is sharp and recurring: a value the paper pins is settled and can be baked, while
      a value inferred from a kit name or an accession list is exactly what a reviewer must be able
      to see and change. That rule was invented here, not read.
    expected: >-
      State the rule in the procedure: expose as a typed workflow parameter any value the source
      leaves unstated or that this brief infers, so the inference is visible and overridable, plus
      any value a test must set; bake values the source pins. Applies equally to the Nextflow and
      CWL interface Molds, whose sources pin parameters far more often than a paper does.
    evidence: >-
      Interface brief section 2.3 has to justify its own exposure policy from scratch: strandedness
      exposed because inferred, Cutadapt `-q 20` and RNA STAR defaults baked because stated,
      significance thresholds exposed because stated but test-relevant. Three different
      justifications for one decision class, none of them sourced from the bundle.
    status: open
    issue: null
  - id: design-briefs-duplicate-open-questions-into-ledger-entries
    raised_by: freeform-summary-to-galaxy-interface
    observed_in:
      mold: freeform-summary-to-galaxy-interface
      path: content/molds/freeform-summary-to-galaxy-interface/index.md
      revision: 3
      content_hash: de0cb02532205ad18f4aa1ab2881cac725a3191b448468a8a11d0663ff1ec97d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: open-requirements-ledger
      locator: content/research/open-requirements-ledger/index.md
      content_hash: e1d5e3a4d6d650af5b2efc4703fe6f79bb4d0cc12e6cd14e841c5b91df122ea0
    kind: friction
    severity: minor
    what: >-
      The Mold's output contract requires an "open questions" section in the Markdown brief and a
      ledger of open entries, with no rule for which destination an unresolved choice belongs in.
      Every unresolved item in this run was genuinely both an obligation a later Mold must
      discharge and a thing a human reviewer should see, so all ten were written twice, in two
      different shapes, and cross-referenced by entry id by hand to stop them drifting. Confirming
      a real-run instance of something the note already lists under Open work
      ("Reconcile the design-tier briefs' free-text 'open questions' sections with the ledger");
      filed as corroboration rather than as a new finding.
    expected: >-
      Pick one source of truth and say so. The cheapest version that would have helped here: make
      the ledger authoritative and specify the brief's open-questions section as a rendered index
      of the entries it raised, one line each keyed by entry id, rather than an independently
      authored list.
    evidence: >-
      Interface brief section 5 carries ten numbered open questions, each ending in the entry id it
      mirrors; the ledger carries the same ten as structured entries. The duplication is manual and
      nothing checks it.
    status: open
    issue: null
  - id: data-flow-mold-packages-pattern-mocs-without-their-recipe-pages
    raised_by: freeform-summary-to-galaxy-data-flow
    observed_in:
      mold: freeform-summary-to-galaxy-data-flow
      path: content/molds/freeform-summary-to-galaxy-data-flow/index.md
      revision: 3
      content_hash: 22613db4e710f58b29bfd2d6fe1d49f167afe451e09afefb2efd755b4a6e4da1
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: freeform-summary-to-galaxy-data-flow
      locator: content/molds/freeform-summary-to-galaxy-data-flow/index.md
      content_hash: 22613db4e710f58b29bfd2d6fe1d49f167afe451e09afefb2efd755b4a6e4da1
    kind: gap
    severity: major
    what: >-
      All four packaged pattern references are MOC index pages. Each is a list of wiki-links to
      operation and recipe pages — sync-collections-by-identifier, collection-cleanup-after-
      mapover-failure, tabular-to-collection-by-row, tabular-filter-by-column-value and roughly
      thirty others — and none of those pages is in the bundle, while the Mold's runtime notes
      forbid reading Foundry source at runtime. The collection MOC states its own role plainly:
      "the operation and recipe pages are the actionable references". So the Mold packages the
      index and withholds the content it indexes. Every collection-idiom choice in this run's
      brief was made from a one-line MOC description, with no corpus-observed recipe available to
      check the shape, the built-in tool id, or the failure modes against.
    expected: >-
      Package the recipe and operation pages the MOCs name, or at least the subset a data-flow
      brief can act on (the Cleanup, Identifiers, Structural Reshape and Bridges sections of
      galaxy-collection-patterns and the Bridges section of galaxy-tabular-patterns). If the full
      leaf set is too large for a bundle, say so in the Mold and state that MOC entries are naming
      hints only, so the brief records idiom selections as unverified rather than as
      corpus-grounded. Applies identically to the Nextflow and CWL data-flow Molds, which package
      the same MOCs. The refs manifest marks these `evidence: corpus-observed`, which as packaged
      is true of the map but not of anything the runtime can actually read.
    evidence: >-
      Bundle references/patterns holds exactly four files, all `pattern_kind: moc`. The brief at
      <run>/freeform-galaxy-data-flow.md selects the identifier-sync idiom for its central
      decision and has to state in its confidence section that the selection rests on a one-line
      index entry.
    status: open
    issue: null
  - id: data-flow-contract-silent-on-contradicting-the-interface-brief
    raised_by: freeform-summary-to-galaxy-data-flow
    observed_in:
      mold: freeform-summary-to-galaxy-data-flow
      path: content/molds/freeform-summary-to-galaxy-data-flow/index.md
      revision: 3
      content_hash: 22613db4e710f58b29bfd2d6fe1d49f167afe451e09afefb2efd755b4a6e4da1
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-data-flow-draft-contract
      locator: content/research/galaxy-data-flow-draft-contract/index.md
      content_hash: f8f75f5a4ff750bff5f682847466492dd8a9910b95b5078373567f84ce16fb54
    kind: gap
    severity: major
    what: >-
      The contract partitions ownership between data-flow, template and step implementation, and
      the Mold declares the interface brief as an input that "pins inputs, outputs, and labels".
      Neither says what to do when the data-flow analysis shows a pinned interface decision is not
      achievable under Galaxy semantics. That happened three times in this run: FastQC mapped over
      a sample_sheet:paired input fans out per read direction, so two outputs the interface
      declared as flat lists are nested; the outer axis of every mapped output is sample_sheet-
      shaped, not the declared `list`; and a linear fold-change parameter is declared against a
      DESeq2 column reported in log2. Whether the data-flow brief may correct the interface, must
      defer to it, or must only record the conflict is not stated anywhere in the bundle. The
      route taken here — ledger each one and state the correction in the brief — was invented.
    expected: >-
      State the rule in the contract's Boundary section: the data-flow draft may contradict an
      interface decision when Galaxy collection or map-over semantics make it unachievable, and
      must record each contradiction as an open-requirements entry naming the interface decision
      it invalidates, so the interface brief and the template cannot silently diverge. Worth
      pairing with a note in the open-requirements research page, whose `supersedes` mechanism
      covers a refuted justification on an existing entry but has nothing for a refuted claim in
      a sibling brief that was never ledgered.
    evidence: >-
      Interface brief section 3 declares outputs 1-8 as `list`/`list:paired` and input 7 as a
      linear fold change; the data-flow brief section 7 shows all three are unachievable as
      written and raises `fastqc-per-read-fanout-not-a-flat-list`,
      `mapped-outputs-carry-sample-sheet-outer-axis` and
      `fold-change-threshold-linear-vs-deseq2-log2fc` in the open-requirements ledger.
    status: open
    issue: null
  - id: sample-sheet-note-omits-sample-sheet-to-tabular-output-columns
    raised_by: freeform-summary-to-galaxy-data-flow
    observed_in:
      mold: freeform-summary-to-galaxy-data-flow
      path: content/molds/freeform-summary-to-galaxy-data-flow/index.md
      revision: 3
      content_hash: 22613db4e710f58b29bfd2d6fe1d49f167afe451e09afefb2efd755b4a6e4da1
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-sample-sheet-collections
      locator: content/research/galaxy-sample-sheet-collections/index.md
      content_hash: 98736fdcd0c4f4698af4e0cd4c539b33e3d0db872de1e10487556b21033030ec
    kind: gap
    severity: major
    what: >-
      The note's closing guidance is that carry-forward of sample-sheet metadata past map-over
      "must be explicit (re-attaching metadata via `__SAMPLE_SHEET_TO_TABULAR__` or rules DSL
      `add_column_from_sample_sheet_index`)", and the note is packaged into this Mold precisely to
      drive that re-attachment. But it describes the tool in one clause — "iterates and tab-joins
      for downstream tabular consumers" — and never states its output columns. Whether the element
      identifier is emitted as a column is the single fact the re-attachment turns on, because the
      identifier is the only key that survives map-over and therefore the only possible join key
      back to a downstream collection. The note recommends the mechanism without supplying what is
      needed to wire it.
    expected: >-
      State the tool's output schema in the "Tool-side access" section: whether the element
      identifier is emitted, in which column, and how the remaining columns order relative to
      `column_definitions`. Do the same for the rules-DSL alternative — which axis
      `add_column_from_sample_sheet_index` reads and whether it can apply to a collection that has
      already lost its `column_definitions`, since that is the case a reader arrives with. The
      sources list already cites lib/galaxy/tools/sample_sheet_to_tabular.xml, so the fact is one
      line away from where it is needed.
    evidence: >-
      This run's central data-flow decision (<run>/freeform-galaxy-data-flow.md section 4) is an
      identifier-keyed split that joins this tool's output to the featureCounts collection. It had
      to ship with open-requirements entry `sample-sheet-to-tabular-identifier-column-unverified`
      and a named substitute node, because the join key could not be confirmed from the bundle and
      the Mold's runtime notes forbid reading Galaxy or Foundry source.
    status: open
    issue: null
  - id: iwc-corpus-path-hard-coded-with-no-override
    raised_by: compare-against-iwc-exemplar
    observed_in:
      mold: compare-against-iwc-exemplar
      path: content/molds/compare-against-iwc-exemplar/index.md
      revision: 10
      content_hash: 5b198e9b87cf554cb01aceaa14a5aaa3ab24c6a35afc35dab7d09ce154cc3a77
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: compare-against-iwc-exemplar
      locator: content/molds/compare-against-iwc-exemplar/index.md
      content_hash: 5b198e9b87cf554cb01aceaa14a5aaa3ab24c6a35afc35dab7d09ce154cc3a77
    kind: defect
    severity: blocker
    what: >-
      The procedure's first and only prerequisite step is "Clone or pull and merge the IWC corpus
      (https://github.com/galaxyproject/iwc) to `~/.foundry/iwc`". That path is hard-coded, and the
      Mold names no environment variable, no configuration key, and no fallback for a machine where
      it cannot be created. On this machine it cannot: `mkdir ~/.foundry` fails with `Operation not
      permitted`, with and without the Bash sandbox. The Mold is the corpus-first check of every
      Galaxy-targeting pipeline, so an unsatisfiable corpus location makes the whole phase
      unrunnable, and the only reason this run produced a comparison at all is that the harness
      supplied a clone at a different path out of band. Nothing in the bundle describes that route,
      and an unattended run has no way to discover it.
    expected: >-
      Read the corpus location from an environment variable with `~/.foundry/iwc` as the default —
      the same shape `OMC_STATE_DIR` already has elsewhere in this toolchain — and state in the
      procedure that a caller-supplied corpus path is honoured and pinned rather than pulled.
      Two smaller corrections belong with it. State that when the corpus is supplied rather than
      cloned, the Mold must not `git pull` or write inside it, since a caller-owned checkout may be
      shared or read-only. And require the corpus commit to be recorded in `iwc-comparison-notes`
      as provenance in every case, not only the supplied one: a structural comparison is a claim
      about a moving corpus, and without the commit no later reader can tell whether a divergence
      this Mold reported has since been closed upstream. The output artifact description says
      nothing about provenance today.
    evidence: >-
      `mkdir ~/.foundry` returned `Operation not permitted` on darwin 25.5.0 under both the
      sandboxed and unsandboxed Bash tool. The phase ran against a caller-supplied shallow clone
      pinned at galaxyproject/iwc `fe41a79`, recorded in <run>/iwc-comparison-notes.md under
      "Corpus provenance" on this phase's own initiative, since no packaged instruction asks for it.
    status: open
    issue: null
  - id: iwc-exemplar-artifact-assumes-a-single-nearest-exemplar
    raised_by: compare-against-iwc-exemplar
    observed_in:
      mold: compare-against-iwc-exemplar
      path: content/molds/compare-against-iwc-exemplar/index.md
      revision: 10
      content_hash: 5b198e9b87cf554cb01aceaa14a5aaa3ab24c6a35afc35dab7d09ce154cc3a77
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: compare-against-iwc-exemplar
      locator: content/molds/compare-against-iwc-exemplar/index.md
      content_hash: 5b198e9b87cf554cb01aceaa14a5aaa3ab24c6a35afc35dab7d09ce154cc3a77
    kind: gap
    severity: major
    what: >-
      The Mold speaks of "nearest IWC exemplar(s)" in its summary, its procedure and its confidence
      table, but the gxformat2 artifact is specified strictly in the singular — one declared
      filename, "the nearest exemplar's relevant subgraph", "Once the nearest exemplar is chosen
      (High or Medium confidence), convert it". Nothing says what to emit when the honest answer is
      several exemplars covering disjoint parts of the subject. That is what happened here and it
      was not an edge case: IWC publishes the subject's journey as two workflows joined at the
      count-table boundary, so one exemplar covers the map-over head and a different one covers the
      differential-expression tail, and a third workflow in another domain was the only corpus
      source for one tool's output shape. Whether that is one file with several YAML documents,
      several files, or a single forced choice was guessed at.
    expected: >-
      State the multi-exemplar case in the "Nearest exemplar (gxformat2) view" section and settle
      the encoding. The shape that worked here and is worth specifying: one file at the declared
      filename, one YAML document per exemplar, each document headed by the abstract IWC workflow
      ID it came from, the steps covered, the steps dropped, and its own confidence level — so a
      Low-confidence cross-domain citation cannot be mistaken for a domain exemplar, which the
      section already warns about in prose but gives no structural way to express. Worth pairing
      with a sentence in the Feature Hierarchy noting that a corpus which splits the subject's
      scope across workflows is itself a first-class structural finding for the template tier, not
      a retrieval failure.
    evidence: >-
      Ranking produced transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting at High for the DE
      tail, transcriptomics/rnaseq-pe/rnaseq-pe at Medium for the map-over head, and
      epigenetics/cutandrun/cutandrun as tool-level-only evidence for the Cutadapt output shape.
      <run>/iwc-exemplar.gxwf.yml carries all three as three YAML documents at the one declared
      filename, an encoding invented in this run.
    status: open
    issue: null
  - id: bounded-subgraph-artifact-has-no-stated-validity-contract
    raised_by: compare-against-iwc-exemplar
    observed_in:
      mold: compare-against-iwc-exemplar
      path: content/molds/compare-against-iwc-exemplar/index.md
      revision: 10
      content_hash: 5b198e9b87cf554cb01aceaa14a5aaa3ab24c6a35afc35dab7d09ce154cc3a77
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: compare-against-iwc-exemplar
      locator: content/molds/compare-against-iwc-exemplar/index.md
      content_hash: 5b198e9b87cf554cb01aceaa14a5aaa3ab24c6a35afc35dab7d09ce154cc3a77
    kind: gap
    severity: major
    what: >-
      The artifact is described as a "Cleaned gxformat2 conversion (via convert --to format2
      --compact) of the nearest IWC exemplar's relevant subgraph", and separately as "bounded to
      the relevant subgraph, not the whole workflow". Those two requirements cannot both hold
      literally. `convert` emits the whole workflow; bounding it to a subgraph is hand surgery that
      necessarily leaves dangling `source:` references to removed steps and outputs, so the result
      is well-formed YAML but not a loadable gxformat2 workflow. The Mold never says which property
      matters, and the downstream consumer is named only as something that "pattern-matches
      against" the file — which does not distinguish a human-read reference from a parsed one. This
      run resolved it by treating the artifact as a reading aid, eliding non-structural tool_state
      and marking every elision inline, but a template Mold that tried to load the file would
      fail, and nothing warned it.
    expected: >-
      State the contract in the "Nearest exemplar (gxformat2) view" section: the artifact is a
      bounded reading aid, not a runnable or validatable workflow; dangling references to elided
      steps are expected; and elisions must be marked in place so a reader can tell a bound from
      an absence. If instead the file is meant to stay loadable, say that and specify how — keep
      every transitively referenced step, or rewrite dropped sources to workflow inputs. Either
      answer is workable; the absence of both means each run invents its own and the downstream
      Mold cannot rely on either. A one-line statement in the artifact description would settle it,
      since that is where a consumer looks.
    evidence: >-
      <run>/iwc-exemplar.gxwf.yml drops the visualization tail of one exemplar and roughly
      two-thirds of the other, leaving references to steps that are no longer present — for
      example the second document's `outputs[Counts Table].outputSource: _unlabeled_step_25/...`,
      whose producing step was dropped — and requiring several step `in:` entries to be commented
      out rather than resolved. It parses as three valid YAML documents and would not load as
      gxformat2. The convert CLI reference packaged in this bundle documents no subsetting option,
      so the bounding is necessarily out-of-band.
    status: open
    issue: null
  - id: gxwf-version-flag-reports-stale-version
    raised_by: compare-against-iwc-exemplar
    observed_in:
      mold: compare-against-iwc-exemplar
      path: content/molds/compare-against-iwc-exemplar/index.md
      revision: 10
      content_hash: 5b198e9b87cf554cb01aceaa14a5aaa3ab24c6a35afc35dab7d09ce154cc3a77
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: "@galaxy-tool-util/cli — gxwf --version"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: minor
    what: >-
      `gxwf --version` prints `1.0.0` regardless of the installed package version. The version
      actually installed here is `@galaxy-tool-util/cli@1.10.1`, confirmed by `npm ls -g`. Every
      Foundry Mold that requires gxwf pins a package version in `_required_tools.json` — this one
      pins `^1.8.1` — and none of those pins can be checked at runtime, because the only version
      the CLI reports is a constant that satisfies no pin and matches no release. The packaged
      availability check works around this by grepping `--help` for a subcommand name, which
      detects presence but says nothing about version.
    expected: >-
      Have `gxwf --version` report the package version from package.json. Until it does, a Mold
      that needs version-sensitive behaviour has no runtime signal, and a run that records its tool
      versions as provenance records a number that is always `1.0.0`.
    evidence: >-
      `gxwf --version` → `1.0.0`; `npm ls -g @galaxy-tool-util/cli` → `@galaxy-tool-util/cli@1.10.1`;
      this Mold's `_required_tools.json` pins `package_version: ^1.8.1` with
      `availability_check: gxwf --help | grep -q draft-validate`. Conversions in this phase
      succeeded, so the defect is in version reporting only, not in the tool's behaviour.
    status: open
    issue: null

  - id: draft-format-note-silent-on-step-addressing-by-label
    raised_by: freeform-summary-to-galaxy-template
    observed_in:
      mold: freeform-summary-to-galaxy-template
      path: content/molds/freeform-summary-to-galaxy-template/index.md
      revision: 6
      content_hash: 0a95f8c04465d48131ace523e0b27af53f1ade6e525a9173c5ff1e936ddc3366
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-workflow-draft-format
      locator: content/research/galaxy-workflow-draft-format/index.md
      content_hash: 93f3b231dd063bf059c550f18b2200c61909af0da9408e0d43351116e886bf7b
    kind: gap
    severity: major
    what: >-
      The note never says which step field a connection's `source:` resolves against. Its only
      example uses the map form, where the step key and its identity are the same string, so the
      question cannot arise there. In the list form a step may carry both `id` and `label`, and
      `gxwf draft-validate` resolves `source:` against the LABEL when one is present — an `id`
      that differs from the label is not addressable at all.
    expected: >-
      State the rule in the "Relaxations vs. gxformat2" section: a labelled step is addressed by
      its label, so in list form `id` and `label` must be the same string (which is what the IWC
      corpus does), or the label must be omitted. Adding a list-form example alongside the
      existing map-form sketch would carry the rule by demonstration.
    evidence: >-
      A first draft of 27 steps, each with a snake_case `id` and a separate human `label`, with
      every `source:` written against the id. `gxwf draft-validate` returned 47 topology errors,
      all of the form `references unknown step "<id>"`, while reporting the same steps in its
      diagnostic paths under their labels. Renaming every `id` to match its `label` and rewriting
      all 47 references cleared it to `draft valid` with no other change. Neither the note nor the
      packaged `draft-validate` command reference mentions the distinction.
    status: open
    issue: null

  - id: draft-format-note-example-writes-format-as-a-scalar
    raised_by: freeform-summary-to-galaxy-template
    observed_in:
      mold: freeform-summary-to-galaxy-template
      path: content/molds/freeform-summary-to-galaxy-template/index.md
      revision: 6
      content_hash: 0a95f8c04465d48131ace523e0b27af53f1ade6e525a9173c5ff1e936ddc3366
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-workflow-draft-format
      locator: content/research/galaxy-workflow-draft-format/index.md
      content_hash: 93f3b231dd063bf059c550f18b2200c61909af0da9408e0d43351116e886bf7b
    kind: defect
    severity: major
    what: >-
      The note's "Example (sketch)" declares a workflow input as `format: fastqsanger.gz` — a bare
      scalar. The draft schema the same Mold packages types `format` as `null | ReadonlyArray<string>`,
      so the sketch is not valid against the contract it illustrates. An author who copies the
      sketch, as it invites, gets a structure error.
    expected: >-
      Write `format:` as a list in the sketch (`format: [fastqsanger.gz]` or the block form), and
      say in the relaxations section that `format` is a list even for a single value. If the
      scalar form is in fact accepted by some other consumer, say which and why the draft schema
      rejects it.
    evidence: >-
      The draft was authored with `format: fastqsanger.gz`, `format: fasta` and `format: gtf`, all
      copied in form from the note's sketch. `gxwf draft-validate` returned one structure error
      against the whole `inputs` array; converting all three to single-element lists cleared it.
      Confirmed independently against the packaged
      `references/schemas/galaxy-workflow-draft.schema.json`, where the input variants type
      `format` as an array of strings only.
    status: open
    issue: null

  - id: draft-format-tiers-conflate-a-named-tool-with-a-named-tool-id
    raised_by: freeform-summary-to-galaxy-template
    observed_in:
      mold: freeform-summary-to-galaxy-template
      path: content/molds/freeform-summary-to-galaxy-template/index.md
      revision: 6
      content_hash: 0a95f8c04465d48131ace523e0b27af53f1ade6e525a9173c5ff1e936ddc3366
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-workflow-draft-format
      locator: content/research/galaxy-workflow-draft-format/index.md
      content_hash: 93f3b231dd063bf059c550f18b2200c61909af0da9408e0d43351116e886bf7b
    kind: gap
    severity: major
    what: >-
      The Identity-pinned tier admits "a source summary that names a specific `tool_id` with
      evidence", and the Mold's source-tendency paragraph relaxes that to "a free-form source that
      does name a specific tool/version with evidence hardens to the matching tier". Two
      paragraphs earlier the same note forbids pinning "on plausibility". A paper naming "FastQC"
      names a piece of software, not a Galaxy `tool_id`; reading the tier rule literally licenses
      writing a Tool Shed path from memory, which is exactly the plausibility pin the note
      forbids. The note never distinguishes the two kinds of naming, and they come apart on every
      free-form source.
    expected: >-
      Say explicitly that Identity-pinned requires a concrete `tool_id` STRING from evidence — a
      corpus workflow, a pattern page's worked example, or a source that quotes the Galaxy tool id
      — and that a source naming a tool by its software name, with no corpus or pattern hit for
      its wrapper, is Deferred with the software name recorded in `_plan_context`. The
      source-tendency paragraph should be reworded to match, since as written it points the other
      way.
    evidence: >-
      All six tools in this run's source are named by the paper. Five had a corpus-confirmed
      `tool_id` from the phase-4 exemplar and were Identity-pinned. FastQC had none: the nearest
      exemplar uses `iuc/falco` instead, so the only route to a FastQC `tool_id` was recall. That
      step was Deferred, against the plainest reading of the source-tendency sentence, on the
      strength of the plausibility prohibition. The two rules gave opposite answers for the same
      step and the note offers nothing to break the tie.
    status: open
    issue: null

  - id: draft-format-has-no-way-to-mark-a-region-provisional
    raised_by: freeform-summary-to-galaxy-template
    observed_in:
      mold: freeform-summary-to-galaxy-template
      path: content/molds/freeform-summary-to-galaxy-template/index.md
      revision: 6
      content_hash: 0a95f8c04465d48131ace523e0b27af53f1ade6e525a9173c5ff1e936ddc3366
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-workflow-draft-format
      locator: content/research/galaxy-workflow-draft-format/index.md
      content_hash: 93f3b231dd063bf059c550f18b2200c61909af0da9408e0d43351116e886bf7b
    kind: gap
    severity: minor
    what: >-
      The draft superset can mark a STEP as unresolved (`TODO`, `_plan_*`) but has no way to mark
      a REGION as provisional, or to record the named alternative a later phase should swap in.
      The `_plan_*` family is per-step and explicitly wrapper-tier, so a multi-step topology
      choice that is settled-but-unprecedented has nowhere durable to live.
    expected: >-
      Decide a home for it and add it to the note — a workflow-level annotation beside the
      `topology_repair` budget the open-requirements note already wants moved into the draft, or
      an explicit statement that a titled `comments:` frame plus an open-requirements entry IS the
      intended mechanism. Either answer is fine; the absence of one means each template Mold run
      invents its own.
    evidence: >-
      Phase 4 instructed this phase to build the sample-sheet condition split "as a delimited
      region with the phase-2 alternative one deletion away". Thirteen steps, no corpus precedent,
      settled topology. The delimitation was expressed three ways, none of them contractual: YAML
      comments (lost on any round-trip through Galaxy), a titled `comments:` frame (schema-legal
      and durable, but the packaged `galaxy-workflow-comments` note describes frames as narrative
      stage annotation, not risk marking), and a `type: markdown` comment holding the four-step
      swap procedure as prose. A reader of the draft alone has no typed signal that the region is
      provisional.
    status: open
    issue: null

  - id: template-mold-packages-pattern-mocs-without-their-recipe-pages
    raised_by: freeform-summary-to-galaxy-template
    observed_in:
      mold: freeform-summary-to-galaxy-template
      path: content/molds/freeform-summary-to-galaxy-template/index.md
      revision: 6
      content_hash: 0a95f8c04465d48131ace523e0b27af53f1ade6e525a9173c5ff1e936ddc3366
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: freeform-summary-to-galaxy-template
      locator: content/molds/freeform-summary-to-galaxy-template/index.md
      content_hash: 0a95f8c04465d48131ace523e0b27af53f1ade6e525a9173c5ff1e936ddc3366
    kind: gap
    severity: major
    what: >-
      The Mold packages four pattern references and all four are MOC index pages. Each names the
      recipe pages that carry the actual tool ids, port names and worked wiring, and none of those
      pages is in the bundle. The runtime notes forbid reading Foundry source, so a recipe named
      in a packaged MOC is unreachable at runtime — the reference resolves to a one-line
      description and a dead wiki-link.
    expected: >-
      Add the recipe pages the template actually reaches for to the manifest, or state in the
      Mold's procedure that the MOCs are for naming an idiom only and that port names and tool ids
      must come from the exemplar artifact or from `discover-shed-tool`. The first is better; the
      second at least stops the bundle promising what it cannot deliver. The data-flow Mold has
      the same defect, filed separately as
      `data-flow-mold-packages-pattern-mocs-without-their-recipe-pages` — different manifest, same
      correction, so fixing one does not fix the other.
    evidence: >-
      Three steps in this draft implement idioms the packaged MOCs name and cannot describe.
      `sync-collections-by-identifier` (collection MOC) is the pattern behind the three
      `__FILTER_FROM_FILE__` steps, but with no recipe page its input and output port names stayed
      `TODO_` sentinels. `tabular-filter-by-column-value` and `tabular-cut-and-reorder-columns`
      (tabular MOC) are the `Filter1` and `Cut1` steps; `Filter1`'s ports were recoverable only
      because the phase-4 exemplar happened to contain a worked instance, and `Cut1`'s were not
      recoverable at all and are recorded as assumed-by-convention in `_plan_context`. The phase-3
      brief flagged the same limitation prospectively in its section 9 evidence-class caveat; this
      is the concrete cost at the template tier.
    status: open
    issue: null
  - id: gxwf-drops-connectedvalue-under-format2-state-key
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: "@galaxy-tool-util/cli — gxwf tool-state validation (state: vs tool_state:)"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: major
    what: >-
      `gxwf validate` (and `draft-validate --concrete`) drops a
      `component_value: {__class__: ConnectedValue}` from a step's tool state before validating
      it when the state is written under the format2 key `state:`, but honours it under
      `tool_state:`. The same bytes under the two keys therefore get opposite verdicts. Any
      conditional parameter case whose generated `workflow_step_linked` branch marks
      `component_value` as required — every non-text case — then fails as
      "component_value: is missing", and the anyOf fallback reports the misleading
      "select_param_type: Expected \"text\", actual \"float\"" alongside it.
    expected: >-
      Treat `state:` and `tool_state:` identically in the tool-state validation path, so a
      connected value placeholder survives to validation under both keys. Failing that, the
      generated `workflow_step_linked` schema should not require `component_value` for a
      parameter that a step connection supplies, which is exactly the relaxation that
      distinguishes it from `workflow_step`. A one-line note in the `draft-validate` /
      `validate` docs would not be enough: the failure names a required field the author
      deliberately connected, which reads as an authoring error rather than a tool bug.
    evidence: >-
      gxwf 1.10.1. Minimal one-step workflow around
      `iuc/compose_text_param/compose_text_param@0.1.1`, second repeat component in the float
      case, `component_value: {__class__: ConnectedValue}`, connected via
      `components_1|param_type|component_value`. Under `state:` the tool-state check fails with
      the five diagnostics above; byte-identical content under `tool_state:` reports
      `tool_state: OK`. Failure does not depend on the connection being present, on
      `__index__`, or on `__current_case__`; substituting a literal float under `state:` passes,
      which is the wrong workaround to be nudged toward — it leaves a shadow default behind a
      connected parameter. The IWC exemplar this run compares against
      (`transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting` at fe41a79) passes its three
      compose_text_param steps precisely because `gxwf convert --to format2` emits `tool_state:`.
    status: open
    issue: null

  - id: advance-draft-mold-needs-a-tool-cache-it-never-declares
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: advance-galaxy-draft-step
      locator: content/molds/advance-galaxy-draft-step/index.md
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
    kind: gap
    severity: major
    what: >-
      The procedure names `galaxy-tool-cache list` as the way to resolve a stock tool's version,
      and the per-step validator it mandates only produces a real verdict when given
      `--cache-dir` with the step's tool cached. Neither the cache nor the `galaxy-tool-cache`
      binary appears anywhere in the bundle's contract: `_required_tools.json` declares `gxwf`
      alone, derived from the three `gxwf` commands the Mold cites, and no step of the procedure
      says to populate a cache. Without it `draft-validate --concrete` reports
      `skip_tool_not_found` for every step and the iteration's green is vacuous.
    expected: >-
      Declare `galaxy-tool-cache` as a required tool of this Mold (it ships in the same
      `@galaxy-tool-util/cli` package, so it costs no new install), and make cache population an
      explicit move in the procedure — `galaxy-tool-cache add <tool_id> --tool-version <v>` after
      wrapper resolution, with the resulting `--cache-dir` passed to `draft-validate --concrete`.
      Citing the command as a CLI reference rather than in prose would also let the required-tools
      deriver pick it up, which is why it is missing today.
    evidence: >-
      Iteration 1 of this run. `galaxy-tool-cache` is not on PATH after the documented
      `npm install -g @galaxy-tool-util/cli` (it exists in the package's bin directory but only
      `gxwf` was linked), and its default cache root `~/.galaxy/tool_info_cache` is not writable
      in this sandbox, so both had to be worked around before any concrete verdict was possible.
      Until the cache was populated by hand, `draft-validate --concrete` reported
      `Tool state: 0 ok, 0 fail, 1 skip` — which the procedure's "on green, return" would have
      accepted.
    status: open
    issue: null

  - id: implement-step-mold-says-state-without-saying-which-key
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: implement-galaxy-tool-step
      locator: content/molds/implement-galaxy-tool-step/index.md
      content_hash: 197be2d9b27741c3f258092f8392654369d1ca234d7fad492387ff69d656292b
    kind: gap
    severity: major
    what: >-
      The procedure says to "shape the step's `state` against
      `input_schemas.workflow_step_linked`" and speaks of `state` throughout, but the
      galaxy-workflow-draft schema it packages admits both `state` and `tool_state` on a step and
      the Mold never says which to write. The two are not interchangeable in practice: under
      gxwf 1.10.1 a connected non-text parameter validates under `tool_state:` and fails under
      `state:` (filed as `gxwf-drops-connectedvalue-under-format2-state-key`), so following the
      Mold's own wording is what produces the red verdict.
    expected: >-
      Name one key as the authoring convention and say it once — `tool_state:`, which is what
      `gxwf convert --to format2` emits and what every converted IWC exemplar a run compares
      against will therefore show — and note that a step mixing the two conventions within one
      draft is a readability cost, not a correctness one. This is worth fixing independently of
      the gxwf defect: the ambiguity is in the instruction, and it will still be there after the
      validator is corrected.
    status: open
    issue: null

  - id: gxwf-cannot-decode-collection-outputs-of-builtin-collection-operations
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: "@galaxy-tool-util/cli — gxwf tool fetch/decode for built-in collection operations"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: major
    what: >-
      `gxwf draft-validate --concrete` cannot bring the built-in `__FLATTEN__` into its tool
      cache. The fetch is reported as `toolshed fetch failed ... for __FLATTEN__`, but the
      accompanying dump is a decode failure against the tool schema, not a transport error:
      `["outputs"][0]` is decoded as the collection-output branch and `["structure"]` `is
      missing`. The tool's collection output carries no `structure`, and the schema requires
      one. The step is then reported `skip_tool_not_found`, so its tool state is never
      validated — and no amount of cache priming can fix it, because the tool cannot be
      decoded into the cache in the first place. The same shape is likely to hit the other
      collection-operation built-ins (`__UNZIP_COLLECTION__`, `__FILTER_FROM_FILE__`,
      `__APPLY_RULES__`), which are exactly the steps a Galaxy draft uses for collection
      plumbing.
    expected: >-
      Make `structure` optional on the collection-output branch of the tool schema (or supply a
      default for tools that declare a collection output without one), so built-in collection
      operations decode and cache like any other tool. Separately, distinguish a transport
      failure from a decode failure in the message: "toolshed fetch failed" sent this run
      looking at network and cache-priming for a fault that was neither.
    evidence: >-
      `gxwf draft-validate <run>/galaxy-workflow-draft.gxwf.yml --concrete --cache-dir <cache>`
      at iteration 2, gxwf 1.10.1. Verdict line `Tool state: 2 ok, 0 fail, 1 skip`, with
      `0 (__FLATTEN__) [skip_tool_not_found]  __FLATTEN__ not in cache (fetch failed)`. The two
      compose_text_param steps validated from the same cache in the same run, so the cache path
      and network were both working. The decode dump is also ~4 KB of expanded structural type
      text printed ahead of the verdict, which buries the one line that names the cause.
    status: open
    issue: null

  - id: advance-draft-mold-reresolves-a-wrapper-already-pinned-in-the-draft
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: advance-galaxy-draft-step
      locator: content/molds/advance-galaxy-draft-step/index.md
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
    kind: friction
    severity: minor
    what: >-
      The procedure's resolve-then-summarize sequence is written as if each iteration met its
      wrapper for the first time. It has no notion of a wrapper already resolved earlier in the
      same run: this draft uses one wrapper at five steps, and the identity-pinned branch still
      directs the iteration to confirm the pin via discover-shed-tool and then invoke
      summarize-galaxy-tool, both of which a sibling step's already-concrete
      `tool_id` + `tool_version` pair has settled. The Mold never describes what this iteration
      actually did, which was to take the sibling's pin and read the summary already in the
      cache.
    expected: >-
      Add a short third case to the resolve step: when another step of this same draft is
      already concrete on the same `tool_id`, adopt its `tool_version` and skip discovery,
      re-using the cached summary rather than re-summarizing. Say plainly that this is the
      intended move, so an iteration that takes it is following the procedure rather than
      departing from it. The Mold should also state that a per-run resolved-wrapper reuse is
      safe precisely because the pin is recorded in the artifact, not in operator memory.
    evidence: >-
      Iteration 2 of this run, step `Build first-contrast row predicate`. Iteration 1 had
      already resolved `iuc/compose_text_param` to 0.1.1 against the live Tool Shed for
      `Build adjusted p-value predicate`, and written the pin into the draft. Three more steps
      of the same draft await the same wrapper.
    status: open
    issue: null

  - id: advance-draft-mold-cites-draft-format-note-it-does-not-package
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: advance-galaxy-draft-step
      locator: content/molds/advance-galaxy-draft-step/index.md
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
    kind: gap
    severity: major
    what: >-
      The procedure's resolve step branches on the draft tiers and sends the reader to
      `galaxy-workflow-draft-format` for them, but the Mold does not package that note: the
      bundle's references are three gxwf CLI pages, `galaxy-tool-job-failure-reference`,
      `open-requirements-ledger`, and the tool-summary and draft JSON Schemas. The Runtime
      Notes then forbid reading Foundry source files at runtime, so the one document the
      procedure names for the distinction it asks the iteration to make is unreachable from
      inside the bundle. The draft JSON Schema does not carry the tiers — they are prose.
    expected: >-
      Package `content/research/galaxy-workflow-draft-format/index.md` in this Mold's
      references, as `freeform-summary-to-galaxy-template` already does. It is the note this
      Mold's central branch depends on, and the per-step loop reads a draft written against it
      on every iteration. If the intent is that the tiers be inferable from the draft alone,
      say that instead and drop the cross-reference.
    evidence: >-
      Iteration 3 of this run, step `Build log2 fold-change predicate`. The tier vocabulary was
      taken from the draft's own `doc:` strings (`Tier: Identity-pinned.` / `Tier: Resolved.`,
      a convention iteration 1 introduced) rather than from any packaged contract, and the same
      gap left the fate of the step's `_plan_context` provenance undecided — see
      `advance-draft-mold-silent-on-plan-provenance-at-concretion`.
    status: open
    issue: null

  - id: advance-draft-mold-silent-on-plan-provenance-at-concretion
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: advance-galaxy-draft-step
      locator: content/molds/advance-galaxy-draft-step/index.md
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
    kind: gap
    severity: minor
    what: >-
      Concretizing a step forces its `_plan_*` fields to be deleted — `draft-next-step` counts
      any surviving `_plan_*` as remaining work, so a finished step that keeps one is selected
      again on the next iteration and the harness loop cannot terminate. Those fields are also
      where the corpus provenance lives: the step implemented here carried a `_plan_context`
      citing the exemplar workflow and step that fixes its component text. The Mold says only
      that the implement phase "resolves the chosen step's remaining `TODO_*` / `_plan_*`
      slots", and never says whether that provenance should be carried into `doc:`, dropped,
      or recorded elsewhere. This run has been keeping it in a YAML comment above the step, a
      convention it invented, and nothing states whether `draft-extract` preserves comments
      when it re-serializes the concrete workflow (this iteration did not test that).
    expected: >-
      Say in the implement step what becomes of a concretized step's planning provenance: fold
      the load-bearing part of `_plan_context` into the step's `doc:` (which survives
      extraction and is visible to a Galaxy user), and drop the rest. State plainly that
      `_plan_*` must not survive concretion, and why — `draft-next-step` would re-select the
      step forever. If YAML comments are in fact preserved by `draft-extract`, say so and
      sanction the comment form; if they are not, say that too, so a run does not park
      provenance somewhere that silently disappears at loop endstate.
    evidence: >-
      Iteration 3, step `Build log2 fold-change predicate`. Its `_plan_context` cited
      `transcriptomics/rnaseq-de` at corpus fe41a79, step `_unlabeled_step_10`, as the source
      of the `abs(c3)>` component text; the step's post-implementation `doc:` and the YAML
      comment above it were written by hand to retain that, with no instruction either way.
      `gxwf draft-next-step` lists every `_plan_*` field of the selected step under `work`,
      which is what makes their removal mandatory rather than stylistic.
    status: open
    issue: null

  - id: advance-draft-mold-treats-a-skipped-tool-state-as-green
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: advance-galaxy-draft-step
      locator: content/molds/advance-galaxy-draft-step/index.md
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
    kind: gap
    severity: major
    what: >-
      The validate step is written as a two-valued gate — "On green, return; on red, route per
      the failure-routing rules" — but `draft-validate --concrete` reports three tool-state
      outcomes, `ok`, `fail`, and `skip`. A step reported `skip` had its tool state checked
      against nothing at all, yet the run prints `Concrete: OK` and the procedure's green branch
      accepts it. The Mold never says which of the two branches a skip belongs to, that a
      skipped step's state is unvalidated, or that some skips are permanent and cannot be
      cleared by priming the cache. An iteration reading only this Mold would return green on a
      draft whose collection-plumbing steps have never been validated, and no later phase would
      know which steps those were.
    expected: >-
      Make the accept criterion three-valued in the validate step: `fail` routes per the
      existing rules; `skip` is neither green nor red but an explicit hole — say that the
      iteration must name each skipped step in its report and hand-review that step's `state`
      against the wrapper, because nothing else will. Distinguish a skip that cache priming
      would clear (tool simply not yet added) from one that priming cannot clear (the tool
      cannot be decoded into the cache at all), and say that repeatedly re-priming the latter is
      wasted work. Carry the list of skipped steps into the loop endstate so terminal validation
      inherits it rather than rediscovering it.
    evidence: >-
      Iteration 4 of this run. `gxwf draft-validate <run>/galaxy-workflow-draft.gxwf.yml
      --concrete --cache-dir <cache>` returned `Concrete: OK` with
      `Tool state: 4 ok, 0 fail, 1 skip`, the skip being `0 (__FLATTEN__)
      [skip_tool_not_found]`. That step's `state` has gone unvalidated for four consecutive
      iterations and is accepted as green every time. Only a convention carried in this run's
      own harness brief — not anything in the Mold bundle — told the iteration to treat the skip
      as a permanent hole needing review by eye rather than as a cache-priming task to retry.
    status: open
    issue: null
  - id: gxwf-linked-step-schema-rejects-the-keys-its-own-converter-emits
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: "@galaxy-tool-util/cli — generated input_schemas.workflow_step_linked vs gxwf convert --to format2"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: major
    what: >-
      Three surfaces of the same CLI disagree about whether format2 `tool_state` may carry the
      `__`-prefixed bookkeeping keys Galaxy writes into conditionals and repeats. For
      `iuc/map_param_value/map_param_value` 0.2.0, `galaxy-tool-cache summarize` generates
      `input_schemas.workflow_step_linked` with `additionalProperties: false` on every
      conditional branch object (allowing only `type`, `input_param`, `mappings`) and on every
      `mappings` item (allowing only `from`, `to`). That schema rejects `__current_case__` and
      `__index__`. But `gxwf convert --to format2` of a real Galaxy workflow emits exactly those
      keys, and `gxwf draft-validate --concrete` accepts a state block containing them. The
      generated schema is therefore stricter than both the converter that produces format2 and
      the validator that checks it — and it is the one surface a Mold is told to author against.
      implement-galaxy-tool-step step 2 says to "shape the step's `state` against
      `input_schemas.workflow_step_linked`"; following that literally produces a state block
      that diverges from what Galaxy round-trips, and no gate reports the divergence, because
      the validator is the permissive one.
    expected: >-
      Make the three surfaces agree, and say which one is normative. Either the generated
      linked-step schema should admit the `__`-prefixed bookkeeping keys the converter emits
      (`__current_case__`, `__index__`, and the top-level `__page__` /
      `__rerun_remap_job_id__`), or `gxwf convert --to format2` should strip them and the
      validator should reject them. Until then, document in the schema-generation output which
      form is canonical for authoring, so a Mold binding a step against the published schema
      gets the same answer as the validator and the converter.
    evidence: >-
      Iteration 6 of this run, gxwf 1.10.1, step `Get featureCounts strandedness parameter`.
      Four state shapes were probed against `gxwf draft-validate <run>/galaxy-workflow-draft.gxwf.yml
      --concrete --cache-dir <cache>`: with the bookkeeping keys, without them, with the
      connected `input_param` omitted from the state block, and with the whole block moved from
      `tool_state:` to `state:`. All four returned `Tool state: 6 ok, 0 fail, 1 skip` — the
      validator discriminates none of them. `gxwf convert --to format2` of the corpus workflow
      `transcriptomics/rnaseq-pe/rnaseq-pe` at IWC fe41a79 emits, for the same step,
      `input_param_type: {type: text, __current_case__: 0, input_param: {__class__:
      ConnectedValue}, mappings: [{__index__: 0, from: ..., to: ...}, ...]}` and
      `unmapped: {on_unmapped: fail, __current_case__: 1}` — every key the generated schema
      forbids. The round-trip form was taken as authoritative for this step.
    status: open
    issue: null
  - id: shed-search-normalization-misses-abbreviation-expansion
    raised_by: discover-shed-tool
    observed_in:
      mold: discover-shed-tool
      path: content/molds/discover-shed-tool/index.md
      revision: 5
      content_hash: 05c99cbac2ec6a4d130b06204c15d0e76cf4fd0e00cdc8164dd606dac72809e2
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: discover-shed-tool
      locator: content/molds/discover-shed-tool/index.md
      content_hash: 05c99cbac2ec6a4d130b06204c15d0e76cf4fd0e00cdc8164dd606dac72809e2
    kind: gap
    severity: major
    what: >-
      The procedure's query-normalization recipe for a tool-id-shaped need is to strip any
      `owner/` prefix, split on `_` / `-` into space-separated words, and also try the bare
      significant word — and it states that "a `miss` is only honest after the name variants
      have been tried". Every variant that recipe generates fails for
      `iuc/map_param_value/map_param_value`, a tool that is published, current, and pinned by
      the IWC corpus. The recipe does not cover the one transform that works: expanding an
      abbreviated word in the id to the word the human tool name actually uses. Galaxy tool ids
      abbreviate routinely (`param`, `val`, `seq`, `align`, `qc`, `col`), so this is a recurring
      shape, not a one-off. An iteration following the packaged recipe literally would have
      declared `miss` and fallen through to author-galaxy-tool-wrapper — authoring a new wrapper
      for a tool that already exists, which is the most expensive possible wrong answer this
      Mold can produce.
    expected: >-
      Add abbreviation expansion to the §1 normalization list, with the common Galaxy
      abbreviations spelled out, and make the honest-miss condition explicit that it includes
      expanded variants. Better still, invert the last resort: before returning `miss`, search
      the significant words with the id token dropped entirely (here `parameter value` alone
      ranks the target first), and state that a `miss` is only honest when a name-shaped query
      has been tried, not merely an id-shaped one.
    evidence: >-
      Iteration 6 of this run, gxwf 1.10.1, resolving the wrapper for step
      `Get featureCounts strandedness parameter`. `gxwf tool-search` returned zero hits for
      `map param value` (the recipe's underscore split), zero for the same query scoped
      `--owner iuc`, and zero for the raw token `map_param_value`. The bare significant word
      `map` returned only unrelated tools (`tdrmapper`, `multi_fasta_glimmerhmm`,
      `glimmerhmm_predict`). Expanding `param` to `parameter` found it immediately and first:
      `map parameter value` scores 42.1 at rank 1, and `parameter value` scores 29.7 at rank 1.
      The tool's human name is "Map parameter value".
    status: open
    issue: null

  - id: stock-tool-version-resolution-is-circular-in-the-builtin-branch
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: advance-galaxy-draft-step
      locator: content/molds/advance-galaxy-draft-step/index.md
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
    kind: gap
    severity: major
    what: >-
      Procedure step 2's built-in/stock branch offers exactly two ways to get a stock tool's
      version — "read it from a populated cache via `galaxy-tool-cache list` or take a known pin
      from the step plan" — and then forbids the only remaining move: "never hand-guess a stock
      version". Both offered sources presuppose the version is already known. For a stock tool
      meeting the run for the first time, with `tool_version: TODO` in the draft and no step-plan
      pin, the cache is empty of it precisely because populating the cache requires an `add`
      that takes `--tool-version`. The branch is circular, and its stated prohibition closes the
      only exit. There is no third route to fall back on: the Tool Shed publishes no version
      listing for a bare stock id at all.
    expected: >-
      Say that a bare-id probe is the sanctioned route for a stock tool and why it is not a
      guess: `galaxy-tool-cache add <id> --tool-version <v>` either returns a summary whose own
      `id` and `version` fields confirm the pin, or fails loudly with a 404 from
      `/api/tools/<id>/versions/<v>`. There is no silent-wrong outcome, so the probe is
      self-verifying and the prohibition should be narrowed to "never record an unconfirmed
      version" rather than "never try one". Name the starting probe — Galaxy's built-in
      collection operations ship at `1.0.0` — and require that the confirming summary be read
      back before the version is written into the draft.
    evidence: >-
      Iteration 7 of phase 6 in this run, gxwf/galaxy-tool-cache 1.10.1, resolving
      `__SAMPLE_SHEET_TO_TABULAR__` for step `Project sample sheet to tabular`. `galaxy-tool-cache
      list` held only the two Tool Shed wrappers earlier iterations had cached. The step plan
      pinned identity but not version (`tool_version: TODO`). Bare `add __SAMPLE_SHEET_TO_TABULAR__`
      failed twice over — TRS versions 500, then `/versions/_default_` 404 — and reported
      "Failed to fetch tool", which reads as "not available" rather than "version unresolved".
      `add __SAMPLE_SHEET_TO_TABULAR__ --tool-version 1.0.0` succeeded immediately and returned
      the full summary; `--tool-version 0.1.0` 404'd. `GET /api/tools/__SAMPLE_SHEET_TO_TABULAR__/versions`
      (the non-TRS listing) returns "No route for" — checked directly, so no listing fallback
      exists. The cached summary then let `draft-validate --concrete` validate the step's state
      for real (7 ok, 0 fail, 1 skip), rather than skipping it.
    status: open
    issue: null

  - id: tool-cache-records-a-shed-path-for-a-bare-stock-tool-id
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: "galaxy-tool-util-ts / galaxy-tool-cache (@galaxy-tool-util/cli 1.10.1)"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: minor
    what: >-
      When `galaxy-tool-cache add` caches a stock Galaxy tool by its bare id, it writes the id
      into `index.json` with a Tool Shed repository path glued on the front, producing
      `toolshed.g2.bx.psu.edu/repos/__SAMPLE_SHEET_TO_TABULAR__` — an identifier that names
      nothing: there is no such shed repository, and no workflow may legally carry that tool id.
      The cached summary body alongside it is correct (`"id": "__SAMPLE_SHEET_TO_TABULAR__"`), so
      the damage is confined to the index, but the index is what `galaxy-tool-cache list` prints.
      That matters because `list` is the surface the advance-galaxy-draft-step procedure directs
      an author to read a stock tool's version off, and what it shows there cannot be pasted into
      a draft.
    expected: >-
      Preserve a bare stock tool id verbatim in the cache index, as the summary body already
      does. A tool id with no `owner/repo` path is not a shed-relative name and should not have
      `toolshed.g2.bx.psu.edu/repos/` prefixed to it.
    evidence: >-
      Iteration 7 of phase 6 in this run. After `galaxy-tool-cache add __SAMPLE_SHEET_TO_TABULAR__
      --tool-version 1.0.0 --cache-dir <run-scratch>/gxwf-cache`, `index.json` records
      `"tool_id": "toolshed.g2.bx.psu.edu/repos/__SAMPLE_SHEET_TO_TABULAR__"` against
      `"source_url": "https://toolshed.g2.bx.psu.edu/api/tools/__SAMPLE_SHEET_TO_TABULAR__/versions/1.0.0"`,
      while the summary file it points at carries `"id": "__SAMPLE_SHEET_TO_TABULAR__"`.
      `galaxy-tool-cache list` prints the prefixed form. Lookup at validation time is evidently
      keyed off the summary body rather than the index, since `draft-validate --concrete`
      resolved the step against the cache and reported it ok.
    status: open
    issue: null

  - id: collection-output-decode-is-a-flat-vs-nested-shape-mismatch-not-a-missing-field
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: "@galaxy-tool-util/cli — ParsedTool collection-output decoding vs the Tool Shed tool API"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: major
    what: >-
      This CORRECTS the root cause and the scope recorded in
      `gxwf-cannot-decode-collection-outputs-of-builtin-collection-operations`, which is filed against
      the same decode failure. Two things in that entry are wrong.
      (1) The collection output does NOT "carry no `structure`". The Tool Shed serializes a collection
      output FLAT — `collection_type`, `collection_type_source`, `collection_type_from_rules`,
      `structured_like` and `discover_datasets` all sit at the top level of the output object — while
      the decoder expects exactly those five fields nested inside a `structure` object. Every field the
      decoder wants is present in the payload; only the nesting differs. The correction that entry
      proposes — make `structure` optional, or default it — would therefore decode cutadapt's
      `out_pairs` with a NULL collection type when the API plainly said `"collection_type": "paired"`.
      That is worse than the current loud failure: downstream shape reasoning would silently lose the
      one fact the output exists to carry.
      (2) It is not a built-in phenomenon. The trigger is a `<collection>` output, wherever it occurs.
      `toolshed.g2.bx.psu.edu/repos/lparsons/cutadapt/cutadapt` — a mainstream IUC-maintained wrapper
      pinned by eight workflows in IWC at fe41a79 — fails identically at every version tried, because
      its paired-collection branch declares `<collection name="out_pairs" type="paired">`. So the
      blast radius is not "the collection-operation built-ins a draft uses for plumbing"; it is every
      tool with a collection output, which includes a large share of the wrappers real Galaxy
      workflows are built from. Any such step is permanently unvalidatable by
      `draft-validate --concrete`, and no cache priming can help.
    expected: >-
      Decode the flat form the Tool Shed actually emits: read `collection_type`,
      `collection_type_source`, `collection_type_from_rules`, `structured_like` and
      `discover_datasets` from the output object itself when no `structure` key is present, and
      populate `structure` from them. Do not make `structure` optional or default it to nulls — that
      discards `collection_type`, which is the only thing a caller needs the branch for. If the
      nested form is the intended contract, then the Tool Shed's
      `/api/tools/<trs-id>/versions/<v>` serializer is the side that must change, and the decoder
      should say which shape it received rather than printing the expected type.
    evidence: >-
      Iteration 8 of phase 6 in this run, gxwf/galaxy-tool-cache 1.10.1, resolving the Cutadapt
      wrapper for step `Quality-trim reads`. `galaxy-tool-cache add
      toolshed.g2.bx.psu.edu/repos/lparsons/cutadapt/cutadapt --tool-version 5.2+galaxy2` fails with
      `toolshed fetch failed ... for lparsons~cutadapt~cutadapt` and the same
      `["outputs"][0] ... ["structure"] is missing` dump as `__FLATTEN__`; 5.2+galaxy0, 4.9+galaxy1
      and 3.7+galaxy0 fail identically, so it is not version-specific. Fetching the same payload
      directly — `GET https://toolshed.g2.bx.psu.edu/api/tools/lparsons~cutadapt~cutadapt/versions/5.2+galaxy2`,
      200 — shows `outputs[0]` is `{"name": "out_pairs", "type": "collection", "collection_type":
      "paired", "collection_type_source": null, "collection_type_from_rules": null,
      "structured_like": null, "discover_datasets": ..., "hidden": ..., "label": ...}` with no
      `structure` key and nothing missing from it. `outputs[1]` (`split_output`) has the same shape;
      the thirteen `data` outputs decode fine. The eight IWC workflows at fe41a79 that pin this
      wrapper all pin 5.2+galaxy2 at changeset f6168dd17f82. Net effect on this run:
      `draft-validate --concrete` reports `7 ok, 0 fail, 2 skip`, both skips being collection-output
      tools, and the Cutadapt step's tool state had to be bound by hand against the wrapper XML and a
      `gxwf convert --to format2` round-trip of a corpus `.ga`.
    status: open
    issue: null

  - id: tool-search-default-page-returns-every-hit-three-times
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: "@galaxy-tool-util/cli — gxwf tool-search result de-duplication"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: minor
    what: >-
      With no `--max-results`, `gxwf tool-search` returns every hit three times — in the table
      rendering and in `--json` alike. Passing any explicit `--max-results` returns distinct hits, so
      the duplication is confined to the default page size. It is not cosmetic for a discovery
      procedure that triages by counting and comparing candidates: the default view of a search makes
      a field of twenty wrappers look like sixty, and the triage rule "multiple plausible hits ... →
      weak" reads a repeated single candidate as a cluster.
    expected: >-
      De-duplicate on `(repoName, repoOwnerUsername, toolId)` before returning, on the default page
      size as well as an explicit one — or, if the repetition encodes distinct revisions, surface the
      revision that distinguishes the rows instead of emitting rows that are byte-identical.
    evidence: >-
      Iteration 8 of phase 6 in this run, gxwf 1.10.1. `gxwf tool-search "cutadapt" --json` returns 50
      `trsToolId` values of which 20 are distinct, each repeated three times; the table rendering of
      the same query shows the same triples. `--max-results 5`, `--max-results 10` and
      `--max-results 20` each return exactly N values, all distinct.
    status: open
    issue: null

  - id: implement-step-mold-has-no-binding-path-when-no-tool-summary-exists
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: implement-galaxy-tool-step
      locator: content/molds/implement-galaxy-tool-step/index.md
      content_hash: 197be2d9b27741c3f258092f8392654369d1ca234d7fad492387ff69d656292b
    kind: gap
    severity: major
    what: >-
      Both Molds assume a tool summary always exists. advance-galaxy-draft-step's sequence step 3 is
      an unconditional "Invoke summarize-galaxy-tool on the resolved wrapper", and
      implement-galaxy-tool-step's sequence step 2 is an unconditional "Read the galaxy-tool-summary
      manifest". Neither says what to do when the summary cannot be produced at all. That is not a
      hypothetical: summarize-galaxy-tool can only summarize what `galaxy-tool-cache add` could
      decode, and a wrapper with a collection output cannot be decoded (see
      `collection-output-decode-is-a-flat-vs-nested-shape-mismatch-not-a-missing-field`). The Mold's
      one nearby escape hatch does not cover it either — "If `input_schemas` is `null`, consult
      `warnings[]`" presupposes a manifest with a warnings array, and here there is no manifest.
      The result is an author improvising the most consequential part of a step, its tool state, with
      no stated evidence standard, on the first genuinely complex wrapper of the run.
    expected: >-
      Give implement-galaxy-tool-step an explicit no-summary branch that names the fallback evidence
      in priority order and requires the step to record which one it used: (1) the wrapper XML at the
      pinned changeset, which is authoritative for parameter names, defaults and output filters;
      (2) `gxwf convert --to format2` over a corpus `.ga` that pins the same version, which is
      authoritative for the round-trip state shape; (3) nothing else. summarize-galaxy-tool already
      sanctions raw XML as supporting evidence ("Optional raw XML source for ambiguity checks"), so
      the material exists — it is the implement Mold that never mentions it. The branch should also
      require the step to carry a visible marker that its state was hand-bound and is therefore
      unchecked by `draft-validate`, since the verdict line will show it as a skip and a skip is
      indistinguishable from an absent step in the counts.
    evidence: >-
      Iteration 8 of phase 6 in this run, step `Quality-trim reads`, wrapper
      `lparsons/cutadapt/cutadapt` 5.2+galaxy2 at changeset f6168dd17f82. `galaxy-tool-cache add`
      failed at four different versions, so summarize-galaxy-tool could not run and no
      `galaxy-tool-summary.json` was ever produced. The step's `tool_state` — a `library`
      conditional at `__current_case__: 2` with six adapter repeats, plus `output_selector`, which
      gates the promoted `report` output through an XML `<filter>` — was bound entirely by hand from
      the tools-iuc wrapper XML at the pinned version and from `gxwf convert --to format2` over
      `epigenetics/cutandrun/cutandrun.ga` and
      `VGP-assembly-v2/post-curation-processing/Post_Curation.ga` at fe41a79. That route was
      inferred, not instructed. `draft-validate --concrete` then reported `7 ok, 0 fail, 2 skip` with
      this step among the skips, so nothing in the run checks the binding.
    status: open
    issue: null
  - id: discovery-schema-cannot-express-an-unresolved-alternate
    raised_by: discover-shed-tool
    observed_in:
      mold: discover-shed-tool
      path: content/molds/discover-shed-tool/index.md
      revision: 5
      content_hash: 05c99cbac2ec6a4d130b06204c15d0e76cf4fd0e00cdc8164dd606dac72809e2
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: schema
      label: galaxy-tool-discovery
      locator: package://@galaxy-foundry/gxwf-foundry#galaxyToolDiscoverySchema
      content_hash: c0940fc0265e8fd25955f4380f03e088b62d25155fe98e7f051a5a3957d2840c
    kind: gap
    severity: minor
    what: >-
      `alternates[]` reuses the full `ToolCandidate` shape, which requires `version` and
      `changeset_revision` (both `minLength: 1`) plus a numeric `score`. There is no way to record
      a candidate the discovery deliberately did NOT pin. The procedure asks for exactly that —
      "multiple plausible hits ... → `weak` with the leading candidate plus alternates", and the
      surrounding Molds hand this skill named alternatives in a step's `_plan_context` — but an
      alternate is only representable after it has been fully resolved to a changeset. So an author
      recording a runner-up faces two bad options: spend Tool Shed calls pinning a wrapper being
      rejected, or write placeholder strings. The second validates green. `"changeset_revision":
      "unresolved"` and `"score": 0` pass `validate-galaxy-tool-discovery` without a murmur, in a
      field the schema itself documents as "Selected Tool Shed Mercurial changeset revision for
      reproducible gxformat2 tool_shed_repository pinning" and one documented as "Higher is better".
      A machine-read pin contract should not be able to carry a fabricated pin.
    expected: >-
      Let an alternate be unresolved. Either make `version` and `changeset_revision` nullable on
      `ToolCandidate` and require them non-null only for the selected `candidate` (the existing
      `allOf` if/then on `status` is already the place to say so), or split a lighter
      `AlternateCandidate` shape carrying identity plus rationale and no pin fields. Add a
      `match_fields` enum value for a candidate surfaced from an upstream plan hint rather than from
      the lexical index — `matched_terms` already anticipates this case ("Empty only when the
      candidate came from a non-lexical hint") but `match_fields` has no corresponding value, so a
      plan-hint alternate has to claim lexical evidence it does not have or leave the array empty.
    evidence: >-
      Iteration 9 of phase 6 in this run, step `Read quality report`. The step was Deferred and its
      `_plan_context` named `iuc/falco` as the corpus-observed alternative to `devteam/fastqc`,
      asking that it be recorded if not taken. Recording it cost three extra Tool Shed calls
      (`tool-search falco`, `tool-versions`, `tool-revisions`) for a wrapper being rejected, and the
      score needed its own query because `iuc~falco~falco` does not appear in the `fastqc` result set
      at all. The first artifact written instead used `"changeset_revision": "unresolved"` and
      `"score": 0`; `foundry validate-galaxy-tool-discovery` reported `galaxy-tool-pin.json: valid`.
      It was corrected by hand afterwards, not by the gate. Notably the schema ref's own
      `verification` line in the cast provenance is "Run discover-shed-tool against known FastQC,
      ambiguous BWA-style, and no-hit queries and validate each emitted recommendation" — this is
      that FastQC query, and the alternates path is what it does not exercise.
    status: open
    issue: null
  - id: no-authority-for-scalar-value-encoding-in-tool-state
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: implement-galaxy-tool-step
      locator: content/molds/implement-galaxy-tool-step/index.md
      content_hash: 197be2d9b27741c3f258092f8392654369d1ca234d7fad492387ff69d656292b
    kind: gap
    severity: minor
    what: >-
      Nothing the Mold names settles how a scalar parameter's VALUE is encoded in the step's
      state block, and the two authorities an author can reach disagree. Procedure step 2 sends
      the author to `parsed_tool` for ports and datatypes and to
      `input_schemas.workflow_step_linked` for the state shape. For `Filter1`'s `header_lines`,
      `parsed_tool` declares `parameter_type: gx_integer`, `type: integer`, `value: 0` — read
      literally that says write the YAML integer `0`. Every real Galaxy workflow writes the
      string `'0'`. This is a different axis from the two disagreements already filed here:
      `implement-step-mold-says-state-without-saying-which-key` is about WHICH block key, and
      `gxwf-linked-step-schema-rejects-the-keys-its-own-converter-emits` is about `__`-prefixed
      bookkeeping keys. This one is about the scalar leaf value, and it is unaddressed by either
      correction. The Mold's guidance is not merely silent — following the one authority it
      names produces the form the corpus never uses.
    expected: >-
      Say in procedure step 2 that a wrapper's declared parameter type fixes the SEMANTICS of a
      state value, not its serialization, and that the authoring form for scalar leaves is what
      `gxwf convert --to format2` emits from a real workflow — strings for integer, float, and
      select parameters alike, because Galaxy's `tool_state` is JSON-string-encoded at the
      source. That is the same authority the run already had to fall back on for conditional
      and repeat bookkeeping, so naming it once covers all three axes. If the Foundry would
      rather not carry that rule, say instead that either encoding is accepted and that the
      choice is a consistency matter within a draft — but say which, because an author who
      reads only `parsed_tool` currently has no way to find out and no gate that will tell them.
    evidence: >-
      Phase 6 iteration 10 of this run, step `Select first-contrast samples` (`Filter1` 1.1.1).
      The cached summary declares `header_lines` as `gx_integer` with `value: 0`. A survey of
      all 21 `Filter1` states in the pinned IWC corpus found the value serialized as a string in
      every one ('0' x14, '1' x7) and as an integer in none; `gxwf convert --to format2` of
      `transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting` likewise emits `header_lines: "1"`.
      `gxwf draft-validate --concrete` was then run twice over otherwise byte-identical drafts,
      once with `header_lines: '0'` and once with `header_lines: 0`, against a cache holding the
      real `Filter1` summary. Both returned `draft valid` / `Concrete: OK` /
      `Tool state: 9 ok, 0 fail, 2 skip` — identical verdicts, so the gate decides nothing here
      and an author cannot resolve the question by trying it. The binding was settled by corpus
      frequency alone, which is exactly the guess the Mold should have removed.
    status: open
    issue: null

  - id: tool-summary-whens-order-is-the-current-case-index-and-nothing-says-so
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: schema
      label: galaxy-tool-summary conditional whens
      locator: package://@galaxy-foundry/gxwf-foundry#galaxyToolSummarySchema
      content_hash: bf77d0735410fa0028bc9bea08edaff3016580508bf27851fdcae65d038f2d94
    kind: gap
    severity: major
    what: >-
      A conditional's `__current_case__` is an INDEX, and nothing in the packaged summary
      schema, the Mold, or any packaged note says what it indexes. The summary gives each
      conditional a `test_parameter.options` array and a `whens` array with a `discriminator`
      per element. The index that `__current_case__` must carry is the position in `whens`
      (i.e. `<when>` document order), NOT the position of the matching value in `options`. The
      two are not interchangeable and this run hit a live divergence: in
      iuc/rgrnastar/rna_star 2.7.11b+galaxy1, `refGenomeSource[history].GTFconditional` lists
      its options as `without-gtf, with-gtf` but its whens as `with-gtf, without-gtf`, so
      `with-gtf` is case 0 while the dropdown shows it second. An author reading the field an
      author would naturally read gets 1. The bundle offers no way to know which array is
      authoritative, so the rule had to be re-established from outside the bundle -- by
      round-tripping a corpus `.ga` that happens to use the same tool and reading the wrapper
      XML's `<when>` order to confirm the correspondence.
    expected: >-
      State once, in the summary schema's own description of `whens` (and echo it in
      implement-galaxy-tool-step's state-authoring guidance), that `whens` is emitted in
      `<when>` document order and that a conditional's `__current_case__` is the zero-based
      index into that array -- explicitly warning that it may differ from the option order,
      with a one-line example. Better still, emit the index on each `whens` element so the
      author never has to count, which also makes the contract checkable rather than
      conventional. This is a third axis of the same family already filed here:
      `implement-step-mold-says-state-without-saying-which-key` is about which block key,
      `no-authority-for-scalar-value-encoding-in-tool-state` is about the scalar leaf value,
      and this one is about the bookkeeping key's VALUE. The `__index__` key on repeat entries
      has the same problem and the same fix.
    status: open
    issue: null

  - id: tool-summary-drops-output-filter-expressions
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: schema
      label: galaxy-tool-summary parsed_tool.outputs
      locator: package://@galaxy-foundry/gxwf-foundry#galaxyToolSummarySchema
      content_hash: bf77d0735410fa0028bc9bea08edaff3016580508bf27851fdcae65d038f2d94
    kind: gap
    severity: major
    what: >-
      `parsed_tool.outputs` carries name, label, hidden, type, format, format_source,
      metadata_source, discover_datasets, from_work_dir and precreate_directory -- but not the
      output's `<filter>` expression. A Galaxy output filter is what decides whether a declared
      output EXISTS for a given parameter branch, so the summary can enumerate ten outputs for
      a tool that will produce four, with nothing marking the difference. That is not cosmetic
      for this Mold: a step's `out:` list and every workflow output wired to it are only valid
      if the chosen tool state keeps those outputs alive, and the summary is the artifact
      procedure step 3 produces expressly so step 4 can bind the step against it. Concretely,
      this run's STAR step had to establish that `output_log` and `mapped_reads` survive
      `quantMode: '-'` and `outWigType: None`, because its plan makes a suppressed `output_log`
      a hard failure -- and the summary cannot answer that question at all. The answer
      (`reads_per_gene` and `transcriptome_mapped_reads` are filtered on quantMode; the two
      promoted outputs carry no filter) came only from reading the wrapper XML. Iteration 8 hit
      the same wall on lparsons/cutadapt, where the `report` and `out_pairs` filters are what
      keep two promoted outputs alive.
    expected: >-
      Carry the raw `<filter>` expression string on each output in `parsed_tool.outputs` (a
      nullable `filter` field is enough; the Mold does not need it evaluated, only visible),
      and have implement-galaxy-tool-step's procedure say that before finalizing a step's
      `out:` list the author must check each promoted output's filter against the state just
      bound. Without the field the correct instruction is unfollowable, because the evidence is
      not in the artifact the procedure hands the author.
    status: open
    issue: null

  - id: no-reachable-wrapper-xml-at-a-pinned-changeset
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: "Galaxy Tool Shed + galaxy-tool-cache: raw wrapper source retrieval"
      locator: https://toolshed.g2.bx.psu.edu
    kind: gap
    severity: major
    what: >-
      Several authoring paths depend on reading a wrapper's XML at the pinned changeset -- the
      correction already filed as `implement-step-mold-has-no-binding-path-when-no-tool-summary-exists`
      names it as fallback evidence "authoritative for parameter names, defaults and output
      filters", and it is the only source for output filters at all (see
      `tool-summary-drops-output-filter-expressions`). No declared tool can fetch it. The
      summary manifest sets `artifacts.raw_tool_source_path: null` for a toolshed-sourced tool;
      the Tool Shed's `repos/<owner>/<repo>/raw-file/<changeset>/<path>` endpoint answers 403;
      and `/api/tools/<trs-id>/versions/<v>/raw_tool_source` answers 404 with "No route". The
      only thing that worked this run was guessing the upstream repository layout from the
      repository record's `remote_repository_url` and fetching the file from GitHub at the
      default branch -- where the filename was `rg_rnaStar.xml`, not the `<repo>/<tool_id>.xml`
      the shed path implies, so the first two guesses 404'd. That fetch is also UNPINNED: it
      happened to match the pinned version here only because tools-iuc main still carried
      @TOOL_VERSION@ 2.7.11b / @VERSION_SUFFIX@ 1, which a later iteration on a lagging wrapper
      cannot count on.
    expected: >-
      Give `galaxy-tool-cache add` an option to retain the raw tool source it already
      downloads, and populate `artifacts.raw_tool_source_path` for toolshed sources rather than
      nulling it -- the bytes are in hand at fetch time, so this is retention, not new
      retrieval. Failing that, the Tool Shed should expose a supported raw-source route per
      (tool id, version) so the pin and the source agree. Until one of those exists, any Mold
      instruction of the form "bind against the wrapper XML at the pinned version" names an
      artifact the run has no sanctioned way to obtain, and authors will keep reaching an
      unpinned copy by guesswork.
    status: open
    issue: null
  - id: draft-format-sentinel-hint-cannot-express-a-conditional-nested-port
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-workflow-draft-format
      locator: content/research/galaxy-workflow-draft-format/index.md
      content_hash: 93f3b231dd063bf059c550f18b2200c61909af0da9408e0d43351116e886bf7b
    kind: gap
    severity: major
    what: >-
      The `TODO_<port>` input sentinel carries the port's semantic hint as a FLAT identifier, so
      it cannot express a port that lives inside a conditional -- and it reads as though no
      qualification were needed. `__FILTER_FROM_FILE__` has one top-level `input` and a
      `filter_source` that exists only inside the `how` conditional, but the template wrote both
      as siblings, `TODO_input` and `TODO_filter_source`. The concrete `in:` key for the second
      is `how|filter_source`; a literal reading of the sentinel yields a bare `filter_source:`,
      which connects nothing. `gxwf draft-next-step` then restates the flat hint verbatim --
      "assign the real wrapper input port name (semantic hint: 'filter_source')" -- so the loop's
      own work list repeats the wrong shape at the moment the author acts on it. The note does
      define qualified keys elsewhere (concrete steps in the same draft carry
      `select_data|rep_factorName_0|rep_factorLevel_0|countsFile`), so the shape is expressible;
      what is missing is any statement that the sentinel's hint is a NAME rather than a PATH,
      and that concretion may have to qualify it.
    expected: >-
      State in the sentinel section that a `TODO_<port>` hint names a parameter, not its address,
      and that the concrete `in:` key must be the tool's full parameter path -- `|`-qualified
      through every enclosing conditional and repeat. Better still, let the sentinel carry the
      path it already knows when the template resolves a nested port, so the hint and the answer
      have the same shape. A worked conditional-nested example next to the existing flat one
      would carry the rule by demonstration, as the map-form/list-form correction did.
    evidence: >-
      Established from the raw TRS payload `GET /api/tools/__FILTER_FROM_FILE__/versions/1.1.0`
      (200), where `filter_source` appears only under `how.whens[*].parameters` and never at the
      top level, and from `mgnify-amplicon-pipeline-v5-rrna-prediction` at IWC fe41a79 through
      `gxwf convert --to format2`, which writes `- id: how|filter_source`. The cost of the gap is
      that NOTHING catches the literal reading: this tool's outputs are collections, so it fails
      the cache decode and reports `skip_tool_not_found`, and a probe run of the same draft with
      the unqualified `filter_source:` key returned a byte-identical verdict to the correct one
      -- `draft valid`, `Concrete: OK`, `Tool state: 16 ok, 0 fail, 3 skip`. A silently
      disconnected input survives every gate in the run.
    status: open
    issue: null

  - id: no-evidence-route-for-an-output-content-property-the-declaration-cannot-carry
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: implement-galaxy-tool-step
      locator: content/molds/implement-galaxy-tool-step/index.md
      content_hash: 197be2d9b27741c3f258092f8392654369d1ca234d7fad492387ff69d656292b
    kind: gap
    severity: major
    what: >-
      Some step bindings depend on a property of an UPSTREAM output's CONTENT -- does this
      tabular output carry a header row, is the element identifier emitted as a column -- and no
      artifact the Mold names can answer that class of question. `parsed_tool.outputs` carries
      name, label, type, format and discovery rules; it carries nothing about the rows inside the
      file, and it never could, because the property is a fact about the wrapper's script rather
      than about its declaration. The Mold has no instruction for the case, so the author either
      guesses or invents a method. This is not the same wall as
      `implement-step-mold-has-no-binding-path-when-no-tool-summary-exists`: there the summary is
      merely absent, and the fallback list that entry proposes -- wrapper XML at the pinned
      changeset, then a corpus round-trip, then "nothing else" -- would still not answer this,
      because the XML's `<param>`/`<data>` declarations are silent on it and the round-trip
      actively misleads. The cost is silent: a wrong `header_lines` drops the first data row of
      every filtered table, or passes a header into a numeric comparison, and either way the
      workflow emits a plausible result.
    expected: >-
      Add to the binding procedure a named evidence route for output CONTENT properties, distinct
      from the one for parameter shape, and require the step to record which source settled it:
      the wrapper's `<test>` blocks -- `assert_contents` on the output in question is a positive
      statement about what the file actually contains, and the absence of a header assertion on
      one output beside its presence on a sibling is itself evidence -- and, when the wrapper is
      open-source, its script's write call. Say explicitly that a corpus round-trip is NOT
      authority for this class: a corpus workflow's binding is evidence about the table IT filters,
      which may be several steps removed from the producer, and copying it across is exactly the
      error this route exists to prevent. Related but separable: this run's phase-5 Mold recorded
      the question as "closable by a summarize-galaxy-tool pass", which is false for the same
      reason -- the pass it names cannot see the property.
    evidence: >-
      Phase 6 iteration 21 of this run, closing open-requirement
      `deseq2-result-table-header-presence-unverified` for `iuc/deseq2/deseq2` @ 2.11.40.8+galaxy4
      (changeset 05f9e54d7e81), which gates `header_lines` on the four downstream `Filter1` steps.
      The tool summary route was unavailable at all (the cache add fails on the collection-output
      decode). What settled it was outside every sanctioned source: `deseq2.R` writes the result
      table with `write.table(..., col.names = FALSE)` while writing the normalized counts one
      screen up with `col.names = NA` -- one tool, opposite answers for two tabular outputs, so no
      tool-level generalization is available either; and `deseq2.xml`'s test asserts `deseq_out`
      content starting at a data row with `has_n_lines`, while the `vst_out` and `counts_out`
      assertions in the same test block DO assert a sample-name header line. The round-trip is the
      trap: `gxwf convert --to format2` over IWC
      `transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting` at fe41a79 shows BOTH its `Filter1`
      steps binding `header_lines: "1"`, which reads as a direct answer and is not one -- that
      workflow manufactures a header with `tp_text_file_with_recurring_lines` + `tp_sed_tool` and
      concatenates it onto the DESeq2 output with `tp_cat` before filtering. Copying its binding
      into a workflow that filters `deseq_out` directly would have silently discarded the most
      significant gene from every result table.
    status: open
    issue: null
  - id: draft-format-cannot-express-a-cross-step-invariant
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-workflow-draft-format
      locator: content/research/galaxy-workflow-draft-format/index.md
      content_hash: 93f3b231dd063bf059c550f18b2200c61909af0da9408e0d43351116e886bf7b
    kind: gap
    severity: major
    what: >-
      Every annotation the draft format offers is scoped to ONE step (`TODO_*`, `_plan_*`, the
      tier vocabulary), so an invariant that binds two or more steps has nowhere to live except
      prose inside one of them -- and `advance-galaxy-draft-step` advances exactly one step per
      invocation, so it is also structurally unable to check one. The step's own `_plan_state`
      said it outright: "Bind both steps together so the factor name, level ordering, header flag
      and output_selector cannot drift apart between the two contrasts." That is an instruction
      with no mechanism behind it. The author is the only enforcement, and only if they happen to
      read a sibling step's plan prose while implementing a different step.
    expected: >-
      Give the format a typed, workflow-level way to declare that named steps must agree on named
      state paths -- a sibling of the `comments:` frame, or a `_plan_invariant` block naming the
      step ids and the `|`-qualified paths -- and have `draft-validate` check it once every named
      step is concrete. It is a cheap check: the paths are already addressable, and this run's
      loop concretized the two coupled steps 1 iteration apart. Failing that, state in the note
      that cross-step couplings are out of scope for the draft and must be carried as
      open-requirements entries, so a Mold run stops inventing its own mechanism.
    evidence: >-
      Two couplings in one workflow, both invisible to every gate. (1) The two DESeq2 nodes must
      stay key-for-key identical apart from one counts collection and one factor level; this was
      settled only by hand-diffing the two steps' parsed `in:`/`out:`/`tool_state` trees (5 `in:`
      keys, 3 `out:` ids, 27 state leaves). (2) Worse, `advanced_options.lfc_shrinkage_type: none`
      on BOTH DESeq2 nodes determines the result table's column count, and two DIFFERENT steps
      bake the resulting indices as opaque string literals -- `"c7<"` and `"abs(c3)>"` in two
      `compose_text_param` bridges. Any other shrinkage value drops the `stat` column, moves padj
      from c7 to c6, and `"c7<"` then names a column that does not exist. A three-step coupling
      expressed as a magic number in a text field. Both DESeq2 steps additionally report
      `skip_tool_not_found` (their `split_output` collection output fails the cache decode), so
      not even the tool-state gate looks at them.
    status: open
    issue: null
  - id: a-required-port-missing-from-in-carries-no-sentinel-and-no-gate
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: galaxy-workflow-draft-format
      locator: content/research/galaxy-workflow-draft-format/index.md
      content_hash: 93f3b231dd063bf059c550f18b2200c61909af0da9408e0d43351116e886bf7b
    kind: gap
    severity: major
    what: >-
      A `TODO_<port>` sentinel marks a port whose NAME is unresolved, but there is no way to mark
      a port that is simply ABSENT from `in:` -- and an absent port is invisible to
      `gxwf draft-next-step`, whose `work` list enumerates sentinels and `_plan_*` fields and so
      reports nothing at all. The obligation survives only as prose. Distinct from
      `draft-format-sentinel-hint-cannot-express-a-conditional-nested-port`, which is about a
      sentinel whose hint has the wrong SHAPE; here there is no sentinel to misread.
    expected: >-
      Require in the note that every port the step's plan calls for appears in `in:` -- as a
      `TODO_<port>` sentinel while unresolved -- so that "the plan named it" and "the work list
      reports it" cannot come apart. A cheap partial check for `draft-next-step`: when a step's
      `_plan_*` prose names a sibling step as its model, flag an `in:` key set that is a strict
      subset of that sibling's.
    evidence: >-
      `Differential expression: second contrast vs reference` declared three `in:` keys; the
      already-concrete sibling it was told to copy declared five. The two missing ones were
      `select_data|rep_factorName_0|rep_factorLevel_{0,1}|factorLevel`. `draft-next-step` reported
      only `TODO[tool_version]` plus four `_plan_*` strings; the sole trace of the two missing
      connections was the sentence "As for the first contrast, with `Counts for the second
      contrast level` at level index 0". Left unwired, both factor levels would fall back to the
      wrapper's empty default -- blank contrast titles on the diagnostic plots, and the silent
      re-opening of resolved entry `deseq2-factor-level-names-not-parameterized`, whose closure
      asserts that no level literal is baked into any step. Nothing would have failed: this run
      already measured that a disconnected input validates byte-identically green, and this step
      reports `skip_tool_not_found` besides.
    status: open
    issue: null
  - id: ledger-resolved-entries-are-never-read-back-at-binding-time
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: advance-galaxy-draft-step
      locator: content/molds/advance-galaxy-draft-step/index.md
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
    kind: gap
    severity: major
    what: >-
      The open-requirements ledger is the only artifact that carries a decision from one
      iteration of the per-step loop to a later one, but the Mold never reads it as an INPUT to
      binding. Its `Inputs` section declares the ledger carries "the run's open, resolved, and
      surrendered entries", and then the procedure's single ledger touchpoint is step 5, AFTER
      implement: "Inspect the open-requirements-ledger for a new `open` blocking entry ...
      appended against this step." New, open, post-hoc. A decision settled five iterations
      earlier lives in a `resolved` entry -- whose `note` field is where the ledger note itself
      says the closure reasoning goes -- and no procedure step ever routes an author back to it
      before they bind state. The packaged `open-requirements-ledger` note does not close the
      gap either: its "how downstream reads it" paragraph covers only
      repair-galaxy-draft-topology reading OPEN blocking entries, and its one line about
      resolved entries ("Resolving is not deleting -- a resolved entry stays in the ledger as
      the audit trail") frames them as provenance, not as an input.
    expected: >-
      Add a read to the procedure BEFORE implement: select the ledger entries whose `step` names
      the step being concretized -- regardless of status -- and treat a `resolved` entry's
      `note` as a binding already settled for this step, to be copied rather than re-derived.
      Entries carry a `step` field precisely so this lookup is a filter, not a search. Say
      explicitly that a resolved entry is authority at binding time and not merely an audit
      trail, and say the converse too: a Mold that settles a binding for a step it is not
      currently implementing must record it in that entry's `note`, because the draft itself has
      nowhere to put it (the sibling steps' `_plan_state` prose is deleted at their own
      concretion).
    evidence: >-
      This iteration concretized `Filter with p-adj threshold: first contrast`. Its whole
      remaining decision was `header_lines`, and the answer -- `'0'`, because the pinned
      iuc/deseq2 wrapper writes `deseq_out` with `col.names = FALSE` -- had been established
      two iterations earlier and recorded ONLY in the `note` of the now-`resolved` entry
      `deseq2-result-table-header-presence-unverified`. Following the procedure literally, that
      note is never opened: step 5 filters for new open entries and runs after the binding is
      already written. The default an author reaches for instead is the corpus, and the corpus
      is wrong here -- `rnaseq-de-filtering-plotting` binds `header_lines: "1"` because it
      MANUFACTURES a header (tp_text_file_with_recurring_lines -> tp_sed -> tp_cat) before
      filtering, which this workflow omits. Binding "1" against the headerless `deseq_out`
      discards row 1 of a table sorted by padj: the most significant gene, which for this
      paper is *SCF1*, the finding the workflow exists to reproduce. No error, no warning, and
      `draft-validate --concrete` returns identical green for '0' and '1'. Three more steps
      (the remaining significance filters) must copy that same value, each from an iteration
      that will not be told to look.
    status: open
    issue: null
  - id: draft-extract-destroys-yaml-comments
    raised_by: advance-galaxy-draft-step
    observed_in:
      mold: advance-galaxy-draft-step
      path: content/molds/advance-galaxy-draft-step/index.md
      revision: 4
      content_hash: c92d452f69a7b40bcab5fc8e63b03e1e2abf98f2a636abc4c93ad7f3c1d03082
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: cli-command
      label: gxwf draft-extract
      locator: content/cli/gxwf/draft-extract.md
      content_hash: f8e9350c96ce005ed5e38deefff6c993824c951784f17f79f7e60d1ad664d8d7
    kind: gap
    severity: major
    what: >-
      `gxwf draft-extract` re-serializes the workflow from a parsed data model, so every YAML
      `#` comment in the draft is destroyed. The note describes the command as three subtractive
      operations — drop drafty steps, strip `_plan_*`, promote `class` — and its Output and
      Gotchas sections say nothing about comments. Nothing in this Mold's bundle does either.
      The loss is total and silent: it is not reported in `--report-json` (which counts only
      dropped steps, dropped outputs and rewritten inputs) and the extracted file validates
      clean, so no gate in the pipeline can see it. It matters because a per-step draft loop has
      nowhere else to put the reasoning: `_plan_*` fields MUST be deleted at concretion or
      `draft-next-step` re-selects the step forever, and the gxformat2 `doc:` field is
      user-facing prose, not a place for a corpus citation or a cross-step warning. A run that
      parks that material in comments above each step — as this one did, and as the concrete
      steps of a template naturally invite — loses all of it at the exact moment the artifact
      becomes the one downstream Molds consume. The two surfaces that DO survive are `doc:` and
      the gxformat2 `comments:` block (frames/markdown), and the note names neither as the place
      to put anything load-bearing.
    expected: >-
      State in the note's Output section that `draft-extract` is a re-serialization, not a
      textual edit, and that YAML comments do not survive it; add it to Gotchas next to the
      existing "this is a transformation, not a validator" warning. Name the two surfaces that
      do survive — step `doc:` and the gxformat2 `comments:` block — and say that anything a
      later reader must not lose belongs there before the loop reaches endstate. Ideally
      `--report-json` should also count discarded comment lines, so the loss is at least
      visible in the sidecar; failing that, the note is the only place a run can learn it, and
      it has to say so. (The upstream fix — a comment-preserving round-trip — is a separate,
      larger ask; the note must describe today's behaviour either way.)
    evidence: >-
      Iteration 26 of this run, at loop endstate, gxwf 1.10.1. `gxwf draft-extract
      <run>/galaxy-workflow-draft.gxwf.yml -o <run>/galaxy-workflow.gxwf.yml --report-json
      <report>` exited 0 with `0 steps dropped, 0 outputs dropped, 0 input rewrites;
      class_after=GalaxyWorkflow`. Parsing both files and comparing: `steps`, `inputs`,
      `outputs` and the `comments:` block are deep-equal, and the only structural difference is
      `class`. The textual difference is 850 comment lines / 62,565 characters present in the
      draft and 0 in the extract — 106,359 bytes down to 41,862, about 59% of the file. What was
      in them: the per-step corpus citations this loop moved out of `_plan_*` at concretion, and
      the two cross-step invariant warnings this run established by hand — that
      `lfc_shrinkage_type: none` on both DESeq2 nodes is what makes the `c7<` and `abs(c3)>`
      predicates in four other steps refer to the right columns, and that `header_lines: '0'` on
      the four significance filters must NOT be the corpus's `'1'`. Neither invariant is
      expressible in the draft schema (see `draft-format-cannot-express-a-cross-step-invariant`)
      and neither is checked by any gate, so the comment was the only record, and the extract is
      where it stopped existing.
    status: open
    issue: null
  - id: paper-to-test-data-declares-no-test-data-refs-shape
    raised_by: paper-to-test-data
    observed_in:
      mold: paper-to-test-data
      path: content/molds/paper-to-test-data/index.md
      revision: 2
      content_hash: 863d7b503a77fa63449899f124d1669e181fdfd9a72f81f516a219cdc89cb6bb
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: paper-to-test-data
      locator: content/molds/paper-to-test-data/index.md
      content_hash: 863d7b503a77fa63449899f124d1669e181fdfd9a72f81f516a219cdc89cb6bb
    kind: gap
    severity: major
    what: >-
      The Mold emits `test-data-refs.json` and specifies no shape for it. The whole procedure is
      one sentence — "derive concrete workflow test inputs and expected outputs — resolvable
      URLs, file shapes, and expected hashes — emitted as `test-data-refs`" — and the bundle
      packages no schema, template, or worked example (`refs: []`, "Load Upfront" and "Load On
      Demand" both "None declared"). The whole JSON structure was invented at runtime: how to key
      an input against a workflow input label, how to express a collection input's elements and
      per-row metadata, where expected outputs live and how to separate an assertion the data
      supports from one it does not. Two downstream Molds in this pipeline consume the artifact
      (freeform-summary-to-galaxy-test-plan at phase 8, implement-galaxy-workflow-test at phase
      9) and neither can rely on any particular key existing. This is the same defect class
      already filed as `summarize-paper-declares-no-freeform-summary-shape`, at the other end of
      the same pipeline.
    expected: >-
      Package a `test-data-refs` skeleton or worked example, or state in the procedure the
      minimum keys every consumer can assume. At minimum: one entry per workflow input keyed by
      the input's label, carrying source URL, hash, datatype and (for collections) collection
      type, element identifiers and per-element metadata; expected outputs separated from
      inputs; and an explicit place to record which expected outputs the resolved data can
      actually produce versus which it cannot. The three sibling test-data Molds
      (`nextflow-to-test-data`, `cwl-to-test-data`, `find-test-data`) emit the same artifact id
      and should share whatever contract is written.
    evidence: >-
      Bundle carries SKILL.md plus _feedback/_provenance/_verify only. The Output section names
      the filename, the format (`json`) and a one-line description, and nothing else. The
      artifact this run wrote (<run>/test-data-refs.json) has seventeen top-level keys, none of
      which were prescribed.
    status: open
    issue: null

  - id: paper-to-test-data-silent-on-subsetting-deposited-data
    raised_by: paper-to-test-data
    observed_in:
      mold: paper-to-test-data
      path: content/molds/paper-to-test-data/index.md
      revision: 2
      content_hash: 863d7b503a77fa63449899f124d1669e181fdfd9a72f81f516a219cdc89cb6bb
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: paper-to-test-data
      locator: content/molds/paper-to-test-data/index.md
      content_hash: 863d7b503a77fa63449899f124d1669e181fdfd9a72f81f516a219cdc89cb6bb
    kind: gap
    severity: major
    what: >-
      The Mold is silent on subsetting, which is the central problem of deriving test data from a
      paper. A paper's deposited data is essentially always orders of magnitude too large to be a
      workflow test fixture — this run's six SRA runs are 22–30 M read pairs each, ~6.5 GB — so
      the Mold's real task is not "resolve URLs" but "resolve URLs AND decide a subset AND
      establish whether that subset still supports the paper's claims". The procedure's phrasing
      ("resolvable URLs, file shapes, and expected hashes") reads as though the deposited data is
      used as-is. Nothing tells the runtime to subset, how to choose a depth, how to keep the
      recipe deterministic, or how to decide whether the result is a fixture that reproduces the
      paper's finding or one that only exercises the workflow's shape. That last decision is
      exactly what the Foundry's own fixture rule makes mandatory — "a fixture must also be able
      to produce the outcome the scenario bound to it claims" (AGENTS.md) — and the Mold never
      routes to it. The subsetting strategy, the depth, the evidence standard and the tiering of
      assertions by whether the data supports them were all invented here.
    expected: >-
      Add a subsetting step to the procedure and state the decision it has to reach. Name the
      common strategies and when each applies (deterministic head subset; alignment-based region
      targeting; downsampling to a fixed seed), require the recipe be reproducible and the
      subset be hashed on its UNCOMPRESSED bytes, and require the output artifact to state
      explicitly whether the subset reproduces the paper's result or is shape-only — citing the
      fixture rule so the runtime knows an unmarked shape-only fixture is a defect and not a
      shortcut. Relatedly, "Required Tools: None declared. Procedure should not assume external
      CLIs are present" sits in tension with the same procedure demanding "expected hashes": a
      hash of a subsetted fixture cannot be produced without tooling, so either the tool
      expectation or the hash expectation should be stated honestly.
    evidence: >-
      This run reached a defensible answer only by going well outside anything the Mold
      describes: measuring SCF1-matching reads per 1 M-read block to show a head subset carries
      no positional bias, then aligning 200,000-pair subsets of all six runs to the pinned
      reference to measure that SCF1 retains 414/391 fragments in AR0382 against 1–4 in the two
      comparison conditions. Without that work the honest answer would have been "shape-only,
      probably"; the Mold gave no reason to do it and no standard to judge it against.
    status: open
    issue: null

  - id: galaxy-test-staging-drops-sample-sheet-column-definitions
    raised_by: paper-to-test-data
    observed_in:
      mold: paper-to-test-data
      path: content/molds/paper-to-test-data/index.md
      revision: 2
      content_hash: 863d7b503a77fa63449899f124d1669e181fdfd9a72f81f516a219cdc89cb6bb
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: Galaxy — test-data staging for sample_sheet collections (galaxy.tool_util.cwl.util / galaxy.tool_util.client.staging)
      locator: https://github.com/galaxyproject/galaxy
    kind: gap
    severity: major
    what: >-
      A sample_sheet collection staged from a Planemo/gxwf test `job:` block never carries
      collection-level `column_definitions`, although the API it posts to accepts them.
      `galactic_job_json`'s `replacement_collection()` passes only `rows` and `name` for a
      sample_sheet collection type, and `StagingInterface`'s `create_collection_func` has no
      `column_definitions` parameter to pass — while `CreateNewCollectionPayload` declares the
      field and `SampleSheetDatasetCollectionType.generate_elements` reads it. The result is
      that a workflow input declaring `column_definitions` is, under test, fed a collection that
      has none. Per-element `columns` survive, so most workflows still behave, and the two
      `column_definitions_compatible()` call sites are both in `DataCollectionToolParameter`
      option-building (UI dropdown filtering) which a test bypasses by supplying the HDCA by id
      — so the divergence is silent rather than caught. It becomes visible at
      `__SAMPLE_SHEET_TO_TABULAR__`, whose header line is emitted only
      `#if $include_headers and $input.collection.column_definitions`: a workflow using that
      tool with headers on produces a header in the UI and no header under test, and no gate
      reports the difference.
    expected: >-
      Accept `column_definitions` on a `Collection` entry in a test job block and thread it
      through `galactic_job_json` → `create_collection_func` → the collections API, alongside
      `rows`. Then a sample sheet staged for a test is the same object a user builds, and
      `validate_row` actually validates the fixture's rows against the declared columns instead
      of short-circuiting. Failing that, Galaxy should say plainly that test-staged sample
      sheets are column-definition-free, so workflow authors know that any behaviour gated on
      `column_definitions` is untestable.
    evidence: >-
      Read at galaxyproject/galaxy dev while resolving open-requirements entry
      `sample-sheet-input-test-fixture-expressibility` for this run:
      lib/galaxy/tool_util/cwl/util.py `replacement_collection()` (sample_sheet branch sets only
      `kwds["rows"]`); lib/galaxy/tool_util/client/staging.py `create_collection_func` signature
      `(element_identifiers, collection_type, rows=None, name=None)`;
      lib/galaxy/schema/schema.py `CreateNewCollectionPayload.column_definitions`;
      lib/galaxy/model/dataset_collections/types/sample_sheet.py `generate_elements`;
      lib/galaxy/model/dataset_collections/types/sample_sheet_util.py `validate_row` and
      `column_definitions_compatible`; lib/galaxy/tools/sample_sheet_to_tabular.xml. This run's
      workflow is unaffected only because its `Project sample sheet to tabular` step sets
      `include_headers: false`. This entry also supplies part of the evidence asked for by
      `testability-note-silent-on-sample-sheet-test-fixtures`, which remains open on its own
      subject.
    status: open
    issue: null

  - id: galaxy-unit-test-documents-a-sample-sheet-rows-shape-the-api-rejects
    raised_by: paper-to-test-data
    observed_in:
      mold: paper-to-test-data
      path: content/molds/paper-to-test-data/index.md
      revision: 2
      content_hash: 863d7b503a77fa63449899f124d1669e181fdfd9a72f81f516a219cdc89cb6bb
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: Galaxy — test/unit/tool_util/test_cwl_util.py sample_sheet rows tests
      locator: https://github.com/galaxyproject/galaxy
    kind: defect
    severity: minor
    what: >-
      `test_galactic_job_json_sample_sheet_collection_with_rows` asserts that `rows` round-trips
      as `{"el1": {"condition": "treatment"}, "el2": {"condition": "control"}}` — a mapping of
      column name to value. The real contract is a POSITIONAL LIST per row: `validate_row` in
      lib/galaxy/model/dataset_collections/types/sample_sheet_util.py rejects on
      `len(row) != len(column_definitions)` and then zips `row` against `column_definitions` in
      order. The unit test never catches this because it mocks `collection_create_func`, so
      nothing downstream of `galactic_job_json` is exercised. These unit tests are the most
      discoverable documentation of the job-block `rows` syntax, and the shape they document
      will not validate whenever the target collection has column definitions.
    expected: >-
      Change the two `rows` unit tests to the positional-list form
      (`{"el1": ["treatment"], "el2": ["control"]}`), matching
      lib/galaxy_test/workflow/collection_semantics_cat_sample_sheet.gxwf-tests.yml which
      already uses lists (`rows: {el1: [], el2: []}`). Better, add an integration test that
      stages a sample sheet WITH column definitions through the real API, which would have
      caught the divergence and would also cover the gap filed as
      `galaxy-test-staging-drops-sample-sheet-column-definitions`.
    evidence: >-
      Read at galaxyproject/galaxy dev: test/unit/tool_util/test_cwl_util.py lines ~192–219
      versus `validate_row` in
      lib/galaxy/model/dataset_collections/types/sample_sheet_util.py. Encountered while
      deriving the job block now recorded in <run>/test-data-refs.json; the dict form was
      adopted from the unit test first and corrected only after reading the validator.
    status: open
    issue: null
  - id: test-plan-mold-ignores-available-concrete-workflow
    raised_by: freeform-summary-to-galaxy-test-plan
    observed_in:
      mold: freeform-summary-to-galaxy-test-plan
      path: content/molds/freeform-summary-to-galaxy-test-plan/index.md
      revision: 2
      content_hash: 9c1d56e625a5dd26cb7a82082ac8f4eeb7db287a626bf00e73457c8dd491ad6d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: mold
      label: freeform-summary-to-galaxy-test-plan
      locator: content/molds/freeform-summary-to-galaxy-test-plan/index.md
      content_hash: 9c1d56e625a5dd26cb7a82082ac8f4eeb7db287a626bf00e73457c8dd491ad6d
    kind: gap
    severity: major
    what: >-
      The procedure's "Labels and fixtures are assumed, not bound" section instructs the Mold to
      bind assertions to interface-brief labels with `label_status: assumed` and
      `workflow.label_source: interface-brief`, and to record fixtures as `storage: unresolved` with
      `location: null`, on the stated premise that the concrete workflow and the resolved test-data
      refs "exist in the harness run-state by the time the plan is authored, but they are reconciled
      downstream rather than here". In this run both were supplied as phase-8 inputs and both were
      settled: `galaxy-workflow.gxwf.yml` (27 steps, 9 inputs, 16 outputs) carries the real labels,
      and `test-data-refs.json` carries resolved URLs, md5s and a measured verdict. Following the
      instruction would have meant writing `assumed` over labels read byte-for-byte from the
      workflow and `unresolved` over fixtures with pinned NCBI URLs and md5s — discarding verified
      information and handing implement-galaxy-workflow-test a reconciliation job already done. The
      Mold was deliberately disobeyed on both counts, and the deviation had to be argued inside the
      artifact rather than settled by the procedure.
    expected: >-
      Make the premise conditional rather than absolute. State that when a concrete workflow or a
      resolved test-data-refs artifact is available to the invocation, the plan binds to it and
      records `label_status: resolved` / `label_source: draft` and real fixture storage; the
      `assumed` / `unresolved` path is the fallback for the template-era case the section describes.
      Declaring `galaxy-workflow-gxformat2` and `test-data-refs` as optional consumed artifacts
      would make that explicit in the bundle instead of leaving it to the runtime to notice.
    evidence: >-
      <run>/galaxy-test-plan.yml `workflow.notes` records the deviation and its reasoning; every
      `label_status` in the plan is `resolved` and `workflow.label_source` is `draft`. The interface
      brief had itself drifted from the draft (open requirement
      `interface-brief-output-and-parameter-surface-drifted-from-draft`), so binding to the brief
      would have produced labels that do not exist in the workflow under test.
    status: open
    issue: null
  - id: test-plan-schema-has-no-evidence-class-for-run-measured-assertions
    raised_by: freeform-summary-to-galaxy-test-plan
    observed_in:
      mold: freeform-summary-to-galaxy-test-plan
      path: content/molds/freeform-summary-to-galaxy-test-plan/index.md
      revision: 2
      content_hash: 9c1d56e625a5dd26cb7a82082ac8f4eeb7db287a626bf00e73457c8dd491ad6d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: schema
      label: galaxy-workflow-test-plan
      locator: package://@galaxy-foundry/gxwf-foundry#galaxyWorkflowTestPlanSchema
      content_hash: 325e9cf1bb5e075fa818dc68868647de95389ca8df8581ee8cf1e2be61196e37
    kind: gap
    severity: minor
    what: >-
      `AssertionIntent.evidence` is a two-value enum, `test-evidence` or `intent`, described as
      "whether this assertion was translated from upstream test evidence or synthesized from
      intent". A third case is routine on the paper and interview paths and occurred throughout this
      run: an assertion neither translated from an upstream fixture nor merely synthesized, but
      MEASURED against the run's own resolved test data before the plan was written. The same gap
      exists at plan level, where `source.derived_from` has the same two values. Forced to pick,
      every such assertion is recorded `evidence: intent`, so a reviewer filtering on the field sees
      a uniformly speculative plan and systematically undercounts its grounding. `confidence: high`
      is the only signal left, and it means something different.
    expected: >-
      Add a third value — `run-measured` (or `fixture-measured`) — to `AssertionIntent.evidence` and
      to `source.derived_from`, meaning the value was obtained by measuring the run's own resolved
      fixtures rather than read from upstream tests or inferred from intent. Nothing else in the
      schema needs to change, and the existing two values keep their meaning.
    evidence: >-
      <run>/galaxy-test-plan.yml records the same complaint twice because the field cannot carry it:
      `source.notes` argues that `derived_from: intent` understates the plan's grounding, and
      warnings[] `fixture-measured-assertions-lack-an-evidence-class` states the consequence.
      Roughly twenty assertions in the plan rest on direct measurement of the resolved fixtures by
      phase 7 (per-sample SCF1 fragment counts, strandedness, mapping rate, annotation gene count)
      or by phase 8 (DESeq2 log2 fold change and adjusted p-value per contrast), and all of them are
      filed as `intent`.
    status: open
    issue: null
  - id: an-interrupted-phase-leaves-its-declared-artifact-written-but-unvalidated
    raised_by: freeform-summary-to-galaxy-test-plan
    observed_in:
      mold: freeform-summary-to-galaxy-test-plan
      path: content/molds/freeform-summary-to-galaxy-test-plan/index.md
      revision: 2
      content_hash: 9c1d56e625a5dd26cb7a82082ac8f4eeb7db287a626bf00e73457c8dd491ad6d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: foundry-feedback-ledger
      locator: content/research/foundry-feedback-ledger/index.md
      content_hash: null
    kind: gap
    severity: major
    what: >-
      A Mold that writes one large artifact and validates it at the end has a window in which the
      artifact exists on disk and nothing records whether it conforms to its schema. This run
      entered that window: phase 8's session was terminated after writing a 67 KB
      `galaxy-test-plan.yml` and before running `foundry validate-galaxy-workflow-test-plan`, and
      before its ledger pass. The run-lifecycle section addresses only run status — it says a hard
      interruption leaves the run `running`, "which is distinguishable from success without a
      recovery write" — and says nothing about the phase's declared OUTPUT. So `status: running` on
      a phase is ambiguous in a way that matters: the artifact may be absent, present and valid,
      present and invalid, or present and half-written, and the four look identical from the ledger.
      A resuming agent, or a next phase that simply reads the artifact because it is there, has
      nothing telling it the verify step never ran.
    expected: >-
      State in the run-lifecycle section that a phase left `running` may have written its declared
      output artifacts in an UNVALIDATED state, and that whoever resumes or succeeds it must re-run
      that phase's declared validation before treating those artifacts as input. Stronger: give the
      phase row a field the skill writes when its own validation passes — `artifacts_validated:
      true`, the exact analogue of `feedback_checked` and justified by the same argument the note
      already makes for it, that a run which looked and found nothing must be distinguishable from
      one that never looked.
    evidence: >-
      The terminated phase-8 session left `<run>/galaxy-test-plan.yml` complete and, as it turned
      out, schema-valid, but nothing on disk said so; the finishing session had to re-run the
      validator to find out. The same interruption also left the artifact citing a feedback entry id
      (`test-plan-mold-ignores-available-concrete-workflow`) that did not exist in this ledger,
      because the ledger pass is likewise an end-of-phase step — the artifact and the ledger were
      inconsistent with each other and only the ledger's `running` status hinted at it.
    status: open
    issue: null
  - id: cast-validation-fallback-resolves-an-unpinned-published-validator
    raised_by: freeform-summary-to-galaxy-test-plan
    observed_in:
      mold: freeform-summary-to-galaxy-test-plan
      path: content/molds/freeform-summary-to-galaxy-test-plan/index.md
      revision: 2
      content_hash: 9c1d56e625a5dd26cb7a82082ac8f4eeb7db287a626bf00e73457c8dd491ad6d
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: implementation
      label: cast-mold caster — generated Validation section
      locator: packages/build-cli/src/commands/cast-mold.ts
      content_hash: null
    kind: friction
    severity: minor
    what: >-
      The caster emits a Validation instruction of the form "run `foundry <cmd> <artifact>` from
      `@galaxy-foundry/gxwf-foundry`; if the command is not on PATH, run `npx --package
      @galaxy-foundry/gxwf-foundry foundry <cmd> <artifact>`". The fallback names no version, so it
      resolves whatever `latest` is on the registry at runtime, while the bundle already carries the
      schema it was cast against, verbatim, under `references/schemas/`. Those two can disagree, and
      when they do the fallback returns a confident green verdict against a contract the skill is
      not bound by. In this run the fallback was the only available route — the checkout had no
      installed dependencies and no package manager on PATH — and the published 0.1.2 schema had to
      be diffed against the bundled copy by hand to know the verdict meant anything, which the
      runtime notes ("do not read Foundry source files at runtime") discourage in the first place.
    expected: >-
      Record the validator package version in `_verify.json` at cast time and emit it in the
      fallback (`npx --package @galaxy-foundry/gxwf-foundry@<version> ...`), so the fallback
      validates against the same contract the bundle carries. Alternatively, or additionally, say in
      the generated Validation line that `references/schemas/<name>.schema.json` in the bundle is the
      binding contract and that a fallback verdict is only meaningful if the two agree.
    evidence: >-
      `_verify.json` for this Mold carries `validator_bin: foundry` and args, with no package
      version anywhere in the bundle. The published schema happened to be byte-identical to
      `references/schemas/galaxy-workflow-test-plan.schema.json` here, so the verdict stands, but
      nothing in the bundle establishes that and nothing would have flagged it had they diverged.
    status: open
    issue: null
  - id: tests-format-schema-rejects-two-shapes-the-iwc-corpus-uses
    raised_by: implement-galaxy-workflow-test
    observed_in:
      mold: implement-galaxy-workflow-test
      path: content/molds/implement-galaxy-workflow-test/index.md
      revision: 8
      content_hash: 966c486ccda2eb1e0f06afa674c5d6b85a7e7bfcc9b1545309e3ef8b1016f931
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: related-project
      label: galaxy-tool-util-ts — tests-format schema / gxwf validate-tests
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: major
    what: >-
      The tests-format schema rejects two shapes that production IWC workflow tests use and that
      Galaxy accepts at run time. (1) The inner element of a `list:paired` job input written as
      `class: Collection` + `type: paired` — the `Collection` `$def` is `additionalProperties:
      false` with only `collection_type`, so `type` is an unknown property and the whole job input
      fails `oneOf`. (2) A nested-collection output assertion whose outer `element_tests` entry
      carries `elements:` without a `class: Collection` discriminator — the schema's `if/then` on
      `class` routes it to the dataset-element model, which has no `elements` property. Run over
      the pinned IWC corpus (fe41a79) with `gxwf validate-tests` from @galaxy-tool-util/cli 1.10.1,
      31 of 122 committed `*-tests.yml` files fail, and the two clusters above account for the bulk
      of them. This makes the static gate unusable as a conformance check against the corpus the
      Foundry treats as normative, and it silently invalidates the corpus-derived recipes the
      Mold's own packaged notes teach.
    expected: >-
      Accept `type` as an alias for `collection_type` on a nested `Collection` in a job block, and
      allow `elements:` on a collection element assertion without requiring an explicit
      `class: Collection`, matching what the Galaxy job-block loader and the test-format runner
      actually accept. If the strict form is deliberate, the schema should say so and Galaxy's
      Pydantic models (galaxyproject/galaxy, the source these are generated from) plus the IWC
      corpus should be migrated together, rather than leaving a validator that fails a quarter of
      the published corpus.
    evidence: >-
      Probed directly. A two-file probe differing only in `type: paired` versus
      `collection_type: paired` gives 48 schema errors versus OK. Removing `class: Collection` from
      one outer `element_tests` entry of this run's own test file gives
      `/0/outputs/Trimmed reads/element_tests/AR0382_A: must NOT have additional properties`.
      Corpus sweep at fe41a79: 91 pass, 31 fail; failures cluster on `/0/job/<label>/class`
      (e.g. sars-cov-2-pe-illumina-wgs-variant-calling, three VGP Hi-C workflows,
      generic-variant-calling-wgs-pe) and on `/0/outputs/<label>/element_tests`
      (e.g. scrna-seq-fastq-to-matrix-10x-cellplex, metagenomic-raw-reads-amr-analysis).
    status: open
    issue: null
  - id: iwc-test-data-conventions-teaches-a-nesting-shape-the-schema-gate-rejects
    raised_by: implement-galaxy-workflow-test
    observed_in:
      mold: implement-galaxy-workflow-test
      path: content/molds/implement-galaxy-workflow-test/index.md
      revision: 8
      content_hash: 966c486ccda2eb1e0f06afa674c5d6b85a7e7bfcc9b1545309e3ef8b1016f931
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: iwc-test-data-conventions
      locator: content/research/iwc-test-data-conventions/index.md
      content_hash: 1921e939703444604440e0768caca6e6b834a73dcfa964f450fa081143e008d3
    kind: defect
    severity: major
    what: >-
      Section 2e states normatively "Note: outer `collection_type: list:paired`, inner
      `type: paired` (not `collection_type:`)", and section 2f's nested-output example omits the
      `class: Collection` discriminator on the outer `element_tests` entry. Both forms are faithful
      transcriptions of the IWC corpus and both are rejected by `references/schemas/
      tests-format.schema.json`, which the same Mold packages and names as the gate its output must
      pass. An agent that follows the note produces a file that fails the Mold's own step 4, with
      48 unhelpful `oneOf` errors pointing at the whole job input rather than at the offending key.
      The note is the only packaged guidance on these two shapes, so there is nothing else to fall
      back to.
    expected: >-
      Give both shapes in each section, marked: the corpus form (`type: paired`; bare `elements:`)
      and the schema-valid form (`collection_type: paired`; `class: Collection` then `elements:`),
      with a one-line note that the validator accepts only the latter and a pointer to the upstream
      entry `tests-format-schema-rejects-two-shapes-the-iwc-corpus-uses`. Until upstream converges,
      a Foundry-authored test file should use the schema-valid form, and the note should say so
      rather than leaving the reader to discover it from the validator.
    evidence: >-
      This run authored the reads input from section 2e verbatim; `gxwf validate-tests` 1.10.1
      returned 48 errors. Switching the single key `type` to `collection_type` returned OK. The
      same substitution is needed for the section 2f output form, confirmed by removing
      `class: Collection` from one element of the finished file.
    status: open
    issue: null
  - id: asserts-idioms-nested-element-tests-recipe-omits-the-class-discriminator
    raised_by: implement-galaxy-workflow-test
    observed_in:
      mold: implement-galaxy-workflow-test
      path: content/molds/implement-galaxy-workflow-test/index.md
      revision: 8
      content_hash: 966c486ccda2eb1e0f06afa674c5d6b85a7e7bfcc9b1545309e3ef8b1016f931
      foundry_head: 63a3f9cf9c97a637fe1628d298e524f35709e289
    subject:
      kind: research
      label: planemo-asserts-idioms
      locator: content/research/planemo-asserts-idioms/index.md
      content_hash: 81b8feeccd891642572e587a494d1e4d3a1b369a19d92947626230494fd92beb
    kind: gap
    severity: minor
    what: >-
      Section 5 is the note an agent reaches for when writing collection-output assertions, and its
      nested-collection rule — "outer `element_tests:` keyed by outer identifier; inner `elements:`
      (note plural, no `_tests` suffix on the inner)" — is incomplete in the one way that matters
      to the gate: the outer entry must also carry `class: Collection`, or the schema routes it to
      the dataset-element model and rejects `elements` as an additional property. The section also
      does not mention that an `asserts:` mapping cannot repeat a key, so any output needing two
      `has_text` probes has to use the list form (`- that: has_text`) — which this run needed on
      five outputs and which the section's own examples never show.
    expected: >-
      Add `class: Collection` to the nested example in section 5 and say why it is required. Add a
      line to section 4 or 9 stating that repeated assertion families require the `- that: <family>`
      list form, since the dict form silently loses all but the last occurrence in YAML.
    evidence: >-
      `Trimmed reads` in <run>/galaxy-workflow.gxwf-tests.yml is a `list:paired` output needing the
      nested form; without `class: Collection` on the outer entry `gxwf validate-tests` reports
      `must NOT have additional properties` at `/0/outputs/Trimmed reads/element_tests/AR0382_A`.
      Four of this run's outputs carry between four and six assertions of which two or more share a
      family, which the dict form cannot express.
    status: open
    issue: null
  - id: tonative-shape-sniff-skips-normalization-on-list-form-format2
    raised_by: validate-galaxy-workflow
    observed_in:
      mold: validate-galaxy-workflow
      path: content/molds/validate-galaxy-workflow/index.md
      revision: 5
      content_hash: 74e3743fc376f33c1eb37832ab4869c327d48f87e015c7a38597cf4c4ca939cd
      foundry_head: 79bf5c3ab98eec3a67948d8bdebb9dba7d8a2359
    subject:
      kind: related-project
      label: "@galaxy-tool-util/schema - toNative normalization guard (_isNormalizedFormat2)"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: major
    what: >-
      `toNative` mistakes a raw format2 workflow for an already-normalized one whenever `inputs:` and
      `steps:` are written in list form, skips normalization, and then aborts with an uncaught
      `TypeError: step.in is not iterable` as soon as a step's `in:` uses the mapping form gxformat2
      equally permits. `_isNormalizedFormat2` decides on three top-level facts that say nothing about
      per-step shape - `class === "GalaxyWorkflow"`, `Array.isArray(inputs)`, `Array.isArray(steps)` -
      so `normalizedFormat2`, whose `normalizeStepIn`/`normalizeStepOut` handle both spellings
      correctly, is never called and `_extractConnections` iterates a mapping. The defect is in
      `toNative`, not in the connection validator that surfaced it: `gxwf convert --to native` and
      `ensureNative` crash on the same input with no connection flag involved. For this run the
      consequence is that `gxwf validate --connections` - the only static gate covering connection
      types, collection algebra and map-over - could not be run at all.
    expected: >-
      Normalize unconditionally. `normalizedFormat2` is idempotent and performs plain object shaping
      rather than a schema decode, so the guard bought nothing; teaching it to inspect every step's
      `in`/`out` would cost more than normalizing and would leave the same class of bug waiting on the
      next field. Whatever the shape of the fix, a validator crashing on its own supported input
      format gives a harness no way to distinguish a broken tool from a broken workflow. Submitted
      upstream with a regression test as jmchilton/galaxy-tool-util-ts#179.
    evidence: >-
      First seen on gxwf 1.10.1 in this run's phase 10; reproduced unchanged on 1.12.0 (both
      `@galaxy-tool-util/cli` and `@galaxy-tool-util/schema`), the current release as of 2026-09-17,
      so it is not a stale-CLI artifact. Trigger is one corner of the shape matrix, established on
      20-line workflows with a single `Filter1@1.1.1` step: list `inputs:` + list `steps:` + mapping
      `in:` crashes; the same workflow with list-form `in:`, with map-form `inputs:`, or with map-form
      `steps:` converts cleanly, because those spellings fail the sniff and get normalized. The
      earlier claim in this entry that the flag breaks on every format2 workflow with a tool step was
      too broad - it breaks on the dialect this Foundry emits. `gxwf convert <run>/galaxy-workflow.gxwf.yml
      --to native` crashes identically with no `--connections`, at `runConvert` in
      `cli/dist/commands/convert.js:83`. Upstream's own suite never crosses the guard: every existing
      `toNative` test writes `inputs`/`steps` in map form. With the guard removed, all 4999
      `packages/schema` tests pass and the three new list-form cases go from crash to green. Workaround
      available to the Foundry without waiting on a release: emit `in:` in list form
      (`- id: <port>` / `source: <ref>`).
    status: filed
    issue: https://github.com/jmchilton/galaxy-tool-util-ts/pull/179
  - id: gxwf-strict-encoding-demands-the-state-key-its-validator-mishandles
    raised_by: validate-galaxy-workflow
    observed_in:
      mold: validate-galaxy-workflow
      path: content/molds/validate-galaxy-workflow/index.md
      revision: 5
      content_hash: 74e3743fc376f33c1eb37832ab4869c327d48f87e015c7a38597cf4c4ca939cd
      foundry_head: 79bf5c3ab98eec3a67948d8bdebb9dba7d8a2359
    subject:
      kind: related-project
      label: "@galaxy-tool-util/cli — gxwf validate --strict-encoding vs the tool-state validator"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: major
    what: >-
      gxwf's two validation paths demand opposite tool-state keys, so no format2 workflow can satisfy
      both. `--strict-encoding` rejects `tool_state:` on every step with `uses "tool_state" instead
      of "state" (format2 should use "state")` and exits 2, while the tool-state validator silently
      drops a `{__class__: ConnectedValue}` placeholder under `state:` and honours it only under
      `tool_state:` (this ledger's `gxwf-drops-connectedvalue-under-format2-state-key`). Moving to the
      key `--strict-encoding` wants reintroduces that bug; staying on the key that works means
      `--strict` can never be part of a Foundry gate. `gxwf convert --to format2` emits `tool_state:`,
      so gxwf's own converter produces output its own `--strict-encoding` rejects.
    expected: >-
      Reconcile the two paths as one decision. If `state:` is the canonical format2 key, fix the
      tool-state validator to honour ConnectedValue under it first, then keep the strict-encoding
      diagnostic. If `tool_state:` is to remain accepted, `--strict-encoding` should not flag it, and
      `gxwf convert --to format2` should emit whichever key the strict path blesses. Until then the
      diagnostic actively steers authors toward the broken key.
    evidence: >-
      gxwf 1.10.1. `gxwf validate <run>/galaxy-workflow.gxwf.yml --json --strict-structure
      --strict-encoding --cache-dir <cache>` exits 2 and emits the message for all 26 steps that
      carry tool state (step 0 has none); the same file under default strictness reports zero
      encoding errors and zero structure errors. The workflow uses `tool_state:` on 26 steps because
      phase 6 established at iteration level that `state:` breaks the three
      `iuc/compose_text_param/compose_text_param@0.1.1` steps. Companion to, not a duplicate of,
      `gxwf-drops-connectedvalue-under-format2-state-key`: that entry is the validator dropping the
      placeholder, this one is the strict gate requiring the key that triggers it. A maintainer
      fixing one need not touch the other; triage may merge them into a single upstream issue.
    status: open
    issue: null
  - id: gxwf-validate-json-output-is-not-machine-parseable
    raised_by: validate-galaxy-workflow
    observed_in:
      mold: validate-galaxy-workflow
      path: content/molds/validate-galaxy-workflow/index.md
      revision: 5
      content_hash: 74e3743fc376f33c1eb37832ab4869c327d48f87e015c7a38597cf4c4ca939cd
      foundry_head: 79bf5c3ab98eec3a67948d8bdebb9dba7d8a2359
    subject:
      kind: related-project
      label: "@galaxy-tool-util/cli — gxwf validate --json stdout contract"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: defect
    severity: major
    what: >-
      `gxwf validate --json` does not put JSON, and only JSON, on stdout, so the documented
      machine interface cannot be consumed by parsing stdout. Two separate breaches. (1) Every
      uncached tool that fails to decode prints a multi-line `toolshed fetch failed (...) for <id>:`
      block to stdout ahead of the JSON document - the full Effect schema type, roughly 55 lines for
      seven such tools - so `JSON.parse(stdout)` throws and the report has to be located by scanning
      backwards for a line that is exactly `{`. (2) Adding any strict flag makes gxwf abandon JSON
      entirely: exit 2, plain-text diagnostics on stderr, and stdout completely empty, although
      `--json` was passed.
    expected: >-
      With `--json`, write the report and nothing else to stdout, and route fetch/decode diagnostics
      to stderr. Carry the strict-mode findings inside the JSON report (the schema already has
      `structure_errors` and `encoding_errors` arrays for exactly this) and keep emitting it on the
      strict failure path, signalling the verdict through the exit code rather than by withholding
      the document. A harness cannot classify a failure it cannot parse.
    evidence: >-
      gxwf 1.10.1. Terminal validation of <run>/galaxy-workflow.gxwf.yml: stdout begins with the
      decode-failure text for `__FLATTEN__` and continues for six more tools before the report; the
      JSON body starts at line 57 of 273. The same command plus `--strict-structure
      --strict-encoding` returns exit 2 with 0 bytes on stdout and the 26 encoding messages on
      stderr. The packaged CLI reference states "JSON output should be treated as the preferred
      cast-skill interface", which is the contract being broken.
    status: open
    issue: null
  - id: gxwf-validate-never-checks-in-key-names
    raised_by: validate-galaxy-workflow
    observed_in:
      mold: validate-galaxy-workflow
      path: content/molds/validate-galaxy-workflow/index.md
      revision: 5
      content_hash: 74e3743fc376f33c1eb37832ab4869c327d48f87e015c7a38597cf4c4ca939cd
      foundry_head: 79bf5c3ab98eec3a67948d8bdebb9dba7d8a2359
    subject:
      kind: related-project
      label: "@galaxy-tool-util/cli — gxwf validate tool-state / step-input name checking"
      locator: https://github.com/jmchilton/galaxy-tool-util-ts
    kind: gap
    severity: major
    what: >-
      `gxwf validate` never checks that a step's `in:` keys name real parameters of the pinned tool,
      even when that tool is fully cached and its state validates. A step wiring a connection to a
      port the tool does not have is reported `tool_state: OK` and counted in the validated total.
      The same permissiveness runs the other way inside `tool_state:`: an unknown extra parameter is
      accepted, and a required parameter that is simply absent is accepted. What the validator
      actually checks is the type and value of the parameters that happen to be present. No
      strictness flag changes this - not `--strict-structure`, not `--strict-state`, not `--strict`.
      This matters most exactly where Galaxy's own syntax is easiest to get wrong: a conditional's
      nested port must be qualified (`how|filter_source`), and the unqualified spelling is silently
      accepted.
    expected: >-
      When the step's tool is resolved from the cache, check each `in:` key against the tool's input
      tree - including conditional qualification (`cond|param`), repeat indexing (`name_<n>|param`)
      and sections - and report an unknown port as an error, or at minimum under `--strict-state`.
      Report a missing required parameter the same way. A validator that reports `20 validated` while
      never having looked at a single port name overstates its own coverage to any harness reading
      the summary.
    evidence: >-
      gxwf 1.10.1, measured directly. A minimal format2 workflow with a single `Filter1@1.1.1` step,
      the tool present in the cache, and `in: {not_a_real_tool_input: tbl}` reports `Structural
      validation: OK`, `tool_state: OK`, `Tool state: 1 validated, 0 skipped`, identically under
      `--strict-structure --strict-state`. On <run>/galaxy-workflow.gxwf.yml, adding
      `bogus_param_probe: "xyz"` to a validated Filter1 step leaves the summary at 20 ok / 0 fail /
      7 skip, and deleting the required `cond` parameter also leaves it at 20 / 0 / 7; by contrast
      `header_lines: "notanint"` on the same step does fail, which is what the check does cover.
      This run's three `__FILTER_FROM_FILE__` steps depend on the qualified `how|filter_source`
      spelling and nothing in the toolchain would have caught the unqualified one.
    status: open
    issue: null
  - id: validate-mold-terminal-pass-has-no-cache-and-no-skip-vocabulary
    raised_by: validate-galaxy-workflow
    observed_in:
      mold: validate-galaxy-workflow
      path: content/molds/validate-galaxy-workflow/index.md
      revision: 5
      content_hash: 74e3743fc376f33c1eb37832ab4869c327d48f87e015c7a38597cf4c4ca939cd
      foundry_head: 79bf5c3ab98eec3a67948d8bdebb9dba7d8a2359
    subject:
      kind: mold
      label: validate-galaxy-workflow
      locator: content/molds/validate-galaxy-workflow/index.md
      content_hash: 74e3743fc376f33c1eb37832ab4869c327d48f87e015c7a38597cf4c4ca939cd
    kind: gap
    severity: major
    what: >-
      The Mold owns the run's last automated gate and its procedure never mentions the tool cache,
      never mentions skips, and gives the artifact a three-valued status - `pass`, `fail`,
      `not-run` - with no value for the outcome this gate actually produces. Two consequences, both
      hit in this run. (1) The invocation the procedure implies, `gxwf validate <file> --json`, omits
      `--cache-dir` and returns `Tool state: 0 validated, 27 skipped` with exit code 0: a green that
      checked nothing, on the terminal gate, with no warning anywhere in the bundle. (2) When the
      cache is supplied, seven steps still skip permanently, and the Mold offers no way to say so -
      `pass` overstates it, `not-run` understates it - and no route to discharge a skip, although one
      exists and is cheap. The Mold's one instruction on this, "A `not-run` status is never reported
      as a pass", guards the case that cannot happen quietly and not the one that can.
    expected: >-
      Three changes to the procedure. Declare the tool cache as part of the contract and make
      `--cache-dir` explicit in the invocation, stating that omitting it turns every tool step into
      a skip and still exits 0. Make the artifact's status carry coverage: either a fourth value for
      a pass with unvalidated steps, or a required `validated`/`skipped` split alongside `status`,
      so a downstream reader cannot cite the green without the number. And give the skip a discharge
      route rather than leaving it to eyeball review: a skipped step's tool can be fetched from a
      Galaxy instance at `GET /api/tools/<tool_id>?io_details=true` and its `in:` keys and tool_state
      parameter names checked against the real input tree, which is what this phase did by hand and
      what the procedure nowhere describes. Sibling of
      `advance-draft-mold-needs-a-tool-cache-it-never-declares` and
      `advance-draft-mold-treats-a-skipped-tool-state-as-green`, which found the same two holes in
      the per-step loop Mold; this is the terminal gate, where they cost more.
    evidence: >-
      This run, phase 10. Only the harness brief - not anything in the cast bundle - carried the
      `--cache-dir` requirement and the expected skip list. With the cache, `gxwf validate
      <run>/galaxy-workflow.gxwf.yml --json --cache-dir <cache>` returns `ok: 20, fail: 0, skip: 7`
      and exit 0; phase 6 iteration 26 recorded that the same command without it returns `0
      validated, 27 skipped`, also exit 0. The seven skips are `__FLATTEN__`, `lparsons/cutadapt`,
      three `__FILTER_FROM_FILE__` and two `iuc/deseq2`, none clearable by priming the cache (see
      `collection-output-decode-is-a-flat-vs-nested-shape-mismatch-not-a-missing-field`). All seven
      were discharged here against the live Galaxy tool API with zero mismatches, in one pass.
    status: open
    issue: null
  - id: validate-cli-note-recommends-a-flag-that-crashes-and-omits-the-cache-trap
    raised_by: validate-galaxy-workflow
    observed_in:
      mold: validate-galaxy-workflow
      path: content/molds/validate-galaxy-workflow/index.md
      revision: 5
      content_hash: 74e3743fc376f33c1eb37832ab4869c327d48f87e015c7a38597cf4c4ca939cd
      foundry_head: 79bf5c3ab98eec3a67948d8bdebb9dba7d8a2359
    subject:
      kind: cli-command
      label: gxwf validate
      locator: content/cli/gxwf/validate.md
      content_hash: 1bf4b4c50e9987f63e65ad36f6deefe6e399fe009280d4bee8533d6c451acf1e
    kind: defect
    severity: major
    what: >-
      The note's Gotchas section warns about the one flag that weakens validation visibly and is
      silent on the two ways it fails invisibly. It says `--no-tool-state` weakens validation, but
      never says that omitting `--cache-dir` has the same effect and worse - every tool step becomes
      a skip and the command still exits 0. `--cache-dir` appears only as a bare option line, "Tool
      cache directory", with nothing about what happens without it. Separately, the note actively
      recommends a flag that cannot run: "Use `--connections` when tool cache metadata is available
      and data-shape compatibility matters, especially around collections and map-over", and its
      Examples block lists `gxwf validate workflow.gxwf.yml --json --connections --strict` - a
      command that at 1.10.1 crashes on any format2 workflow with a tool step, and would abandon
      JSON even if it did not.
    expected: >-
      Add a Gotchas line making the cache trap explicit - without `--cache-dir` pointing at a
      populated cache, `validate` reports every tool step as skipped and still exits 0, so the split
      must be read rather than the exit code. Mark `--connections` as not working at 1.10.1 with a
      pointer to the upstream defect, and drop or flag the `--connections --strict` example so the
      note stops recommending a crash. Record which gxwf version the page was verified against; the
      page carries none today, and `gxwf --version` self-reports 1.0.0 regardless.
    evidence: >-
      gxwf 1.10.1, this run's phase 10. Following the note's own recommendation was the first thing
      attempted and it exited 1 with `TypeError: step.in is not iterable` and no report; see
      `tonative-shape-sniff-skips-normalization-on-list-form-format2`. The `--strict` half of the same
      example returns exit 2 with empty stdout; see
      `gxwf-validate-json-output-is-not-machine-parseable`. The cache trap is the failure phase 6
      iteration 26 hit on its first terminal validate: `0 validated, 27 skipped`, exit 0.
    status: open
    issue: null
foundry-run-manifestoptional-absentfoundry-run.yml

Runtime artifact initialized by the harness ([[foundry-run-manifest]]).

declared by
— (phase —)
consumed at
nothing downstream reads it
schema
none declared
sha256

not on disk.

freeform-galaxy-data-flowpresentfreeform-galaxy-data-flow.md

Reviewable Markdown brief: abstract operations, collection map/reduce choices, shape-changing placeholder steps, unresolved Galaxy tool needs, confidence, open questions.

declared by
freeform-summary-to-galaxy-data-flow (phase 3)
consumed at
4, 5, 8
schema
none declared
sha256
1b0bc5140d6c30f71f293daea2be3f740badc01fec661f3a1510e7721735ebd6
  • Galaxy data-flow brief — *C. auris* Scf1 RNA-seq differential expression
  • 1. What this phase settled
  • 2. Abstract operation graph
  • 2.1 Nodes
  • 2.2 Edges
  • 3. Collection map/reduce decisions
  • 4. The condition factor: how it reaches DESeq2 (settled)
  • 4.1 The problem, precisely
  • 4.2 The settled route
  • 4.3 The one residual
  • 4.4 Why this is independent of the DESeq2 realization
  • 4.5 Why this is robust to the interface's own fallback
  • 4.6 Rejected alternatives
  • 5. Shape-changing and placeholder transformations
  • 6. Unresolved tool needs
  • 7. Where the interface brief's declared shapes do not hold
  • 7.1 FastQC outputs are nested, not flat lists
  • 7.2 Trimmed reads may arrive as two parallel collections
  • 7.3 Two parameter gaps the wiring exposes
  • 8. Data-flow evidence bearing on entries this phase did not close
  • 9. Confidence
  • 10. Open questions
  • 11. Handoff notes
freeform-galaxy-interfacepresentfreeform-galaxy-interface.md

Reviewable Markdown brief: Galaxy workflow inputs, outputs, labels, collection shapes, checkpoint outputs, source-summary provenance, confidence, open questions.

declared by
freeform-summary-to-galaxy-interface (phase 2)
consumed at
3, 4, 5, 7, 8
schema
none declared
sha256
0d8c34f3296a4e720f814f5986c81d17b10c0628f479180177e0160baaaae3b0
  • Galaxy workflow interface brief — *C. auris* Scf1 RNA-seq differential expression
  • 1. Scope
  • 2. Workflow inputs
  • 2.1 Why `sample_sheet:paired` for the reads
  • 2.2 Reference data delivery
  • 2.3 Parameter surface
  • 3. Workflow outputs
  • 4. Provenance and confidence
  • 5. Open questions
freeform-summarypresentfreeform-summary.md

Methods, tools, sample data, references, and workflow intent extracted from a primary paper, normalized into the shared free-form source summary handoff.

declared by
summarize-paper (phase 1)
consumed at
2, 3, 5, 7, 8
schema
none declared
sha256
74cc140952e788633aa3b8c4b42f9364f001a0fe54b12bd6b824c8608e00a840
  • Free-form source summary — Santana et al. 2023, *Science*: Scf1, a *Candida auris*-specific adhesin
  • Source
  • Workflow intent
  • Pipeline A — RNA-seq differential expression
  • Steps, tools, parameters (all as stated in the supplement)
  • Library / sequencing facts
  • Reference data
  • Sample data (public, resolvable — strong test-data leads)
  • Contrasts and expected biological result (useful as workflow assertions)
  • Pipeline B — AtMT T-DNA insertion-site mapping (WGS)
  • Steps, tools, parameters
  • Library / sequencing facts
  • Sample data
  • Reference data
  • Other computational / analytical methods in the paper (not sequencing workflows)
  • Strains and other identifiers worth carrying forward
  • Assumptions carried forward
  • Open questions
galaxy-test-planpresentgalaxy-test-plan.yml

Reviewable Galaxy workflow test plan (see [[galaxy-workflow-test-plan]]): synthesized test cases with job inputs, expected outputs, assertion intent, fixture provenance, label assumptions, unresolved mappings, and omissions.

declared by
freeform-summary-to-galaxy-test-plan (phase 8)
consumed at
9
schema
galaxy-workflow-test-plan
sha256
a16f4db071a9d7746f186a9c21e5897190ff363977e9536081a557d8f2bfa3b4
# galaxy-test-plan — run auris-scf1 (paper-to-galaxy, phase 8)
# Produced by freeform-summary-to-galaxy-test-plan.
# Reviewable handoff for implement-galaxy-workflow-test. NOT a tests-format file.
plan_version: "1"

source:
  kind: freeform
  name: "Santana DJ et al. 2023, Science 381(6665):1461-1467 (PMID 37769084, doi 10.1126/science.adf8972) — Pipeline A, bulk RNA-seq differential expression"
  derived_from: intent
  notes: >-
    There is no upstream test evidence: the source is a paper, not a Nextflow or CWL pipeline with
    its own fixtures, so every assertion below is synthesized rather than translated. That said,
    `derived_from: intent` understates this plan's grounding. Phase 7 (paper-to-test-data) resolved
    the fixtures and then MEASURED the workflow's central claim against them — six 200,000-read-pair
    head subsets aligned to GCA_002759435.2 with bwa-mem and counted over the NCBI GTF's exon
    features. SCF1 (B9J08_001458) carries 414/391 fragments in the two AR0382 replicates against 2/1
    in AR0387 and 4/3 in tnSWI1. Assertions resting on that measurement are marked `confidence: high`
    even though the schema forces `evidence: intent`, because the only two enum values are
    `test-evidence` (upstream test fixtures, which do not exist here) and `intent`. See
    warnings[] `no-evidence-class-for-fixture-measured`.

    Phase 7's measurement used bwa-mem plus a strand-agnostic exon counter, NOT RNA STAR plus
    featureCounts with -s 2. It establishes that the signal survives subsetting; it does not predict
    the workflow's exact integers. No assertion in this plan pins an exact count. Count-derived
    assertions are expressed as thresholds with a stated margin.

    THE DIFFERENTIAL-EXPRESSION NUMBERS IN THIS PLAN COME FROM A DIFFERENT IMPLEMENTATION THAN THE
    WORKFLOW RUNS. Phase 7 measured counts only. The statistics quoted below — SCF1's log2 fold
    change and adjusted p-value per contrast, and the rank figures in omissions[] and unresolved[] —
    were measured during phase 8 with pydeseq2 over a strand-aware count matrix built from the same
    six fixture BAMs, as two separate two-level n=2 analyses with no LFC shrinkage, mirroring the
    workflow's two DESeq2 nodes. pydeseq2 is a faithful reimplementation, NOT the R DESeq2 the Galaxy
    wrapper runs, and its input matrix came from bwa-mem rather than RNA STAR plus featureCounts -s
    2. Exact values WILL differ. Measured: tnSWI1 vs AR0382 log2FC -6.93, padj 2.1e-31; AR0387 vs
    AR0382 log2FC -7.95, padj 4.1e-18. Read those as evidence that a threshold has margin, never as
    values to assert. What survives the implementation difference is margin — 31 and 18 orders of
    magnitude on padj against a 0.05 cut, and roughly 5 and 6 log2 units of headroom over the plan's
    `<= -2` floor. What does not survive it is rank or any exact value, which is why neither is
    asserted anywhere. See warnings[] `deseq2-statistics-measured-with-pydeseq2-not-r`.

workflow:
  title: "C. auris Scf1 RNA-seq differential expression (Santana et al. 2023)"
  label_source: draft
  notes: >-
    DELIBERATE DEVIATION FROM THE MOLD DEFAULT. freeform-summary-to-galaxy-test-plan instructs that
    labels come from the interface brief and be recorded `assumed` / `label_source: interface-brief`,
    on the premise that the concrete workflow does not exist yet. In this run it does: the harness
    supplied `galaxy-workflow.gxwf.yml` (27 steps, 9 inputs, 16 outputs, produced by phase 6) as a
    phase-8 input, and every input and output label below was read from it and matches byte-for-byte.
    Recording them as `assumed` would discard verified information and hand
    implement-galaxy-workflow-test a reconciliation job that is already done. Filed as
    `test-plan-mold-ignores-available-concrete-workflow` in foundry-feedback.ledger.yml.

    The interface brief has itself drifted from the draft (open requirement
    `interface-brief-output-and-parameter-surface-drifted-from-draft`), which is a second reason to
    bind to the workflow rather than the brief.

    Labels are the public API of this workflow. Four of them embed condition level names
    (`... tnSWI1 vs AR0382`) that are NOT derived from the `First/Second contrast condition level`
    parameters; changing those parameters makes the labels stale without breaking anything, and both
    test cases below therefore leave the three level parameters at their defaults.

test_cases:
  - id: scf1-both-contrasts-200k
    doc: >-
      End-to-end run over all six deposited runs of PRJNA904261, subset to 200,000 read pairs each,
      against the NCBI GCA_002759435.2 reference and annotation. Exercises the whole DAG — flatten,
      FastQC, Cutadapt, RNA STAR, featureCounts, the sample-sheet condition split, both DESeq2
      reductions, both two-link significance chains — and asserts the paper's central biological
      result: SCF1 (B9J08_001458) is significantly down-regulated in BOTH contrasts against the
      AR0382 reference level. This is the plan's primary case; it is the only one that makes
      biological claims.
    derived_from: intent
    provenance: >-
      Synthesized from the paper's Fig. 1D / Fig. S5A result statement and its |fold change| > 2 and
      adjusted p-value < 0.05 significance criteria (freeform-summary.md §"Contrasts and expected
      biological result"), then grounded on phase 7's direct measurement of the resolved fixtures
      (test-data-refs.json `verdict`, `expected_outputs.safe`, `expected_outputs.biological`).
    job_inputs:
      - workflow_label: RNA-seq reads (sample sheet)
        label_status: resolved
        description: >-
          Six paired-end runs, 2 x 50 bp NextSeq 2000, each subset to the first 200,000 read pairs.
          Element identifiers are the deposited ENA library_name values and are load-bearing: the
          condition split joins the featureCounts collection back to the sample sheet on them.
          Identifiers, in sample-sheet row order — AR0382_A (SRR22376032), AR0382_B (SRR22376031),
          AR0387_A (SRR22376030), AR0387_B (SRR22376029), AR0382_tnSWI1_A (SRR22376028),
          AR0382_tnSWI1_B (SRR22376027).
        collection_shape: sample_sheet:paired
        datatype: fastqsanger.gz
        fixture:
          storage: unresolved
          location: null
          checksum: null
          provenance: >-
            GENERATED AND HASHED BUT NOT HOSTED — this is the one blocking gap in the plan (open
            requirement `test-fixtures-not-hosted`; test-data-refs.json `unresolved`
            `fixture-hosting-not-done`). Regenerate deterministically by streaming each ENA FASTQ and
            taking `head -n 800000`, then verify each against the per-file `md5_uncompressed` in
            test-data-refs.json. Hash the UNCOMPRESSED .fastq: gzip output is not byte-stable across
            gzip versions, which is why phase 7 pinned uncompressed md5s and why a `hashes:` block on
            the .gz fixtures cannot be authored from what phase 7 supplies. ~61 MB compressed for the
            12 files, so the IWC idiom applies: host on Zenodo and reference by URL, substituting for
            `<FIXTURE_BASE>` in test-data-refs.json `planemo_test_job_block`.

            The `rows:` shape is settled and has a trap: it is a mapping of element identifier to a
            POSITIONAL LIST of column values ordered to match `column_definitions` — `AR0382_A:
            [AR0382, A]`. Galaxy's own non-paired unit tests in test_cwl_util.py write `rows` as a
            dict; that form never reaches `validate_row()` in those tests and will NOT validate here.
            Staging never passes `column_definitions`, which is harmless only because this workflow's
            `Project sample sheet to tabular` step sets `include_headers: false`. Do NOT set
            `decompress: true` on the reads — the input declares fastqsanger.gz and the .gz must
            survive.
      - workflow_label: Reference genome FASTA
        label_status: resolved
        description: >-
          C. auris B8441, NCBI assembly GCA_002759435.2 (Cand_auris_B8441_V2), used whole. At 12.4 Mb
          it needs no subsetting, so the gene ID space stays exactly the real one and RNA STAR builds
          its own index inside each of the six mapped jobs.
        collection_shape: null
        datatype: fasta
        fixture:
          storage: remote-url
          location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.fna.gz
          checksum: MD5:a008b270d3aaa04736a8bb7daf3f6dd5
          provenance: >-
            Pinned NCBI FTP URL, staged with `decompress: true`. Needs no hosting. The md5 is of the
            COMPRESSED file as served; see unresolved `hash-vs-decompress-ordering` before pairing it
            with a `hashes:` block on a `decompress: true` input.
      - workflow_label: Gene annotation GTF
        label_status: resolved
        description: >-
          NCBI's own GTF for the exact accession the paper names. The paper names no annotation at
          all, so this choice is load-bearing for both consumers (RNA STAR splice junctions,
          featureCounts gene assignment) and determines whether SCF1 appears as B9J08_001458 at all.
          Phase 7 downloaded and inspected it: true GTF (#gtf-version 2.2), `gene_id` values ARE the
          paper's locus tags, 5586 distinct gene_id values, SCF1 resolves as B9J08_001458, so the
          workflow's `gff_feature_attribute: gene_id` and `gff_feature_type: exon` bindings are
          correct as they stand.
        collection_shape: null
        datatype: gtf
        fixture:
          storage: remote-url
          location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.gtf.gz
          checksum: MD5:6e5b9528d48c0a8fc2c8588e6eeea929
          provenance: >-
            Pinned NCBI FTP URL, staged with `decompress: true`. Closes open requirement
            `featurecounts-annotation-source-unnamed`. Same `hashes:` caveat as the FASTA.
      - workflow_label: Reference condition level
        label_status: resolved
        description: The condition value both contrasts are taken against; must match a `condition` value in the sample sheet exactly. Left at the workflow default.
        collection_shape: null
        datatype: null
        fixture:
          storage: null
          location: null
          checksum: null
          provenance: "Workflow default `AR0382`, the paper's high-adhesion parent and the reference level of both contrasts."
      - workflow_label: First contrast condition level
        label_status: resolved
        description: Condition value for the first contrast. Left at the workflow default, because four output labels hard-code this level name and are not derived from this parameter.
        collection_shape: null
        datatype: null
        fixture:
          storage: null
          location: null
          checksum: null
          provenance: "Workflow default `tnSWI1`, the AR0382 tnSWI1 insertional mutant."
      - workflow_label: Second contrast condition level
        label_status: resolved
        description: Condition value for the second contrast. Left at the workflow default, same label-staleness reason as the first.
        collection_shape: null
        datatype: null
        fixture:
          storage: null
          location: null
          checksum: null
          provenance: "Workflow default `AR0387`, the low-adhesion clinical isolate."
      - workflow_label: Strandedness
        label_status: resolved
        description: >-
          Left at the workflow default `stranded - reverse`. No longer an inference from the library
          kit name: phase 7 measured it. Of R1 reads unambiguously overlapping a single annotated
          gene, 98.4% / 98.4% / 98.2% (AR0382_A / AR0387_A / tnSWI1_A) are antisense to the gene,
          which is reverse-stranded. Closes open requirement
          `rnaseq-strandedness-inferred-from-kit-name`.
        collection_shape: null
        datatype: null
        fixture:
          storage: null
          location: null
          checksum: null
          provenance: "Workflow default, confirmed by measurement in test-data-refs.json `inputs.Strandedness.evidence`."
      - workflow_label: Adjusted p-value threshold
        label_status: resolved
        description: Left at the workflow default 0.05, which is the paper's stated criterion. Applied to DESeq2 column c7.
        collection_shape: null
        datatype: null
        fixture:
          storage: null
          location: null
          checksum: null
          provenance: "Stated by the paper (adjusted p-value < 0.05)."
      - workflow_label: log2 fold change threshold
        label_status: resolved
        description: >-
          Left at the workflow default 1.0. A log2 FC of 1 is exactly the paper's |fold change| > 2
          criterion; the filter compares the raw c3 column with no conversion, so a linear 2.0 here
          would silently apply a 4-fold cut.
        collection_shape: null
        datatype: null
        fixture:
          storage: null
          location: null
          checksum: null
          provenance: "Paper's |fold change| > 2, expressed in log2 units per open requirement `fold-change-threshold-linear-vs-deseq2-log2fc`."
    expected_outputs:
      - workflow_label: "FastQC raw reads: text summary"
        label_status: resolved
        description: >-
          Twelve elements, not six — the reads are flattened per read direction before QC, so
          identifiers are <sample>_forward / <sample>_reverse. The assertable text checkpoint behind
          the HTML binary, and the cheapest proof that the flatten side-branch fanned out correctly
          and that the fixture is the depth it claims to be.
        output_kind: collection
        collection_shape: list
        assertion_intent:
          - family: has_text_matching
            intent: >-
              Each element's fastqc_data.txt reports exactly 200,000 sequences, proving the fixture
              is the declared head subset and that no read loss happened upstream of QC. Apply to all
              twelve elements via element_tests.
            expected_value: 'Total Sequences\s+200000'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_text
            intent: "Each element carries the Basic Statistics module, i.e. FastQC actually parsed the file rather than erroring into an empty report."
            expected_value: ">>Basic Statistics"
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_text
            intent: >-
              The twelve element identifiers are addressed explicitly through element_tests, which is
              itself the cardinality and identifier-space assertion: AR0382_A_forward,
              AR0382_A_reverse, AR0382_B_forward, AR0382_B_reverse, AR0387_A_forward,
              AR0387_A_reverse, AR0387_B_forward, AR0387_B_reverse, AR0382_tnSWI1_A_forward,
              AR0382_tnSWI1_A_reverse, AR0382_tnSWI1_B_forward, AR0382_tnSWI1_B_reverse. Do NOT add
              an `attributes: {collection_type: list}` assertion here — see unresolved
              `promoted-collection-type-string-unverified`.
            expected_value: 12
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: Cutadapt trimming report
        label_status: resolved
        description: >-
          Six elements on the sample axis. Deterministic text with exact input counts, so it is a
          strong and cheap checkpoint that map-over produced one job per sample row and that the
          paired shape survived into Cutadapt.
        output_kind: collection
        collection_shape: list
        assertion_intent:
          - family: has_text_matching
            intent: "Each of the six reports states that 200,000 read pairs were processed — proving the paired-collection map-over kept both mates together on the six-element sample axis."
            expected_value: 'Total read pairs processed:\s+200,000'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_text
            intent: >-
              Element_tests address the six sample identifiers explicitly (AR0382_A, AR0382_B,
              AR0387_A, AR0387_B, AR0382_tnSWI1_A, AR0382_tnSWI1_B), asserting that the identifier
              space did NOT pick up the _forward/_reverse suffixes of the flatten side-branch.
            expected_value: 6
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: Trimmed reads
        label_status: resolved
        description: >-
          Six paired elements. Asserted only for shape: that Cutadapt under
          `library.type: paired_collection` emitted one paired-inner collection per sample, so no
          re-pair node is needed ahead of RNA STAR. Content is not asserted — FASTQ quality strings
          are read-id dependent and the downstream counts table is the stronger checkpoint.
        output_kind: collection
        collection_shape: list:paired
        assertion_intent:
          - family: has_size
            intent: >-
              Each of the six outer elements has non-empty `forward` and `reverse` inner elements
              (nested element_tests: outer keyed by sample identifier, inner using `elements:`).
              Catches the failure mode where the paired structure collapses or one mate is dropped.
            expected_value: null
            tolerance:
              kind: none
              magnitude: null
              rationale: "Min-only bound; trimmed output size depends on adapter content and is not predicted here."
            element_identifier: null
            evidence: intent
            confidence: medium
      - workflow_label: STAR mapping summary
        label_status: resolved
        description: >-
          Log.final.out per sample — the assertable text behind the binary BAM, which is why the BAM
          itself is an intentional omission. Proves the six STAR jobs ran and that the history-FASTA
          index build produced a usable genome.
        output_kind: collection
        collection_shape: list
        assertion_intent:
          - family: has_text
            intent: "The log carries the uniquely-mapped-reads stanza at all, i.e. STAR completed rather than dying in genomeGenerate."
            expected_value: "Uniquely mapped reads %"
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: >-
              Uniquely mapped reads % is at least 80% in every element. Phase 7 measured 198.2-198.6k
              of 200k R1 reads primary-mapping with bwa-mem (>99%); STAR against the same reference
              will be lower but comfortably above 80%, so the threshold carries a wide margin. NOTE
              the full-line anchoring convention recorded in warnings[].
            expected_value: '\s*Uniquely mapped reads % \|\s+(?:8[0-9]|9[0-9]|100)\.[0-9]+%'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
      - workflow_label: Gene counts per sample
        label_status: resolved
        description: >-
          The strongest deterministic checkpoint in the workflow: exact per-gene integer counts,
          addressable by gene id, keyed by the six sample element identifiers — and the only place
          the SCF1 signal can be asserted independently of DESeq2. featureCounts on a fixed
          reference, annotation and strandedness is deterministic, so the IWC comparison notes are
          right that an existence-only probe here would be the smell the anti-patterns note names.
          This plan is stricter than the corpus default here without pinning exact integers, which
          phase 7 cannot supply because it measured with a different aligner.
        output_kind: collection
        collection_shape: list
        assertion_intent:
          - family: has_n_columns
            intent: "featureCounts `format: tabdel_short` emitted the two-column gene-id/count table the DESeq2 per-level ports consume."
            expected_value: 2
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_n_lines
            intent: >-
              One row per gene in the annotation. Phase 7 counted exactly 5586 distinct gene_id
              values in the NCBI GTF. The ±1 allowance is for the header row: the DESeq2 steps bind
              `header: true` on their counts inputs, implying the wrapper emits one, but that was not
              read directly from the featureCounts wrapper in this run.
            expected_value: 5586
            tolerance:
              kind: delta
              magnitude: 1
              rationale: "Header-row presence on featureCounts output_short is inferred from the DESeq2 steps' `header: true` binding, not verified."
            element_identifier: null
            evidence: intent
            confidence: medium
          - family: has_line_matching
            intent: >-
              THE ANNOTATION-SPACE GUARD. At least 5000 rows are a bare B9J08_ six-digit locus tag
              plus an integer count. This is the assertion that fails loudly if the annotation input
              is swapped for the FungiDB GFF3 (keyed on `ID`, different gene ID space) or if
              `gff_feature_attribute` drifts off `gene_id` — failure modes whose natural symptom is a
              complete, plausible counts table of the WRONG identifiers rather than an error.
            expected_value: 'B9J08_[0-9]{6}\t[0-9]+'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: "SCF1 is present as a row in every one of the six count tables, with an integer count."
            expected_value: 'B9J08_001458\t[0-9]+'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: >-
              SCF1 is highly expressed in the AR0382 replicates — at least 100 fragments. Phase 7
              measured 414 and 391, so the threshold sits about 4x below the measurement; it is a
              threshold, not the exact count, because phase 7 counted with bwa-mem plus a
              strand-agnostic counter rather than STAR plus featureCounts -s 2.
            expected_value: 'B9J08_001458\t[0-9]{3,}'
            tolerance: null
            element_identifier: AR0382_A
            evidence: intent
            confidence: medium
          - family: has_line_matching
            intent: "Same SCF1 >= 100 threshold on the second AR0382 replicate (measured 391)."
            expected_value: 'B9J08_001458\t[0-9]{3,}'
            tolerance: null
            element_identifier: AR0382_B
            evidence: intent
            confidence: medium
          - family: has_line_matching
            intent: >-
              SCF1 is at most 99 fragments in AR0387_A. Measured 2, so the threshold sits ~25x above
              the measurement. Together with the two AR0382 assertions this pins the paper's central
              separation at the counts table, upstream of DESeq2 and independent of any statistical
              model.
            expected_value: 'B9J08_001458\t[0-9]{1,2}'
            tolerance: null
            element_identifier: AR0387_A
            evidence: intent
            confidence: medium
          - family: has_line_matching
            intent: "SCF1 <= 99 in AR0387_B (measured 1)."
            expected_value: 'B9J08_001458\t[0-9]{1,2}'
            tolerance: null
            element_identifier: AR0387_B
            evidence: intent
            confidence: medium
          - family: has_line_matching
            intent: "SCF1 <= 99 in AR0382_tnSWI1_A (measured 4)."
            expected_value: 'B9J08_001458\t[0-9]{1,2}'
            tolerance: null
            element_identifier: AR0382_tnSWI1_A
            evidence: intent
            confidence: medium
          - family: has_line_matching
            intent: "SCF1 <= 99 in AR0382_tnSWI1_B (measured 3)."
            expected_value: 'B9J08_001458\t[0-9]{1,2}'
            tolerance: null
            element_identifier: AR0382_tnSWI1_B
            evidence: intent
            confidence: medium
      - workflow_label: featureCounts assignment summary
        label_status: resolved
        description: >-
          Assigned / Unassigned_NoFeatures / Unassigned_Ambiguous totals per sample — and the direct
          read-out of whether the Strandedness parameter reached featureCounts as the right -s value.
          The two assertions below are a paired strandedness regression detector: with -s 2 (correct)
          Assigned is ~180,000 and NoFeatures ~20,000; with -s 1 (inverted) the two swap, and each
          assertion fails on its own digit-count bound.
        output_kind: collection
        collection_shape: list
        assertion_intent:
          - family: has_line_matching
            intent: >-
              Assigned is between 150,000 and 199,999 fragments, i.e. >=75% of the ~198k fragments
              entering featureCounts. Phase 7 measured 90-91% of primary-mapped R1 falling inside an
              annotated exon (178.9k-181.4k assigned). Under an inverted -s 1 the Assigned total
              collapses to five digits and this fails.
            expected_value: 'Assigned\t1[5-9][0-9]{4}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
          - family: has_line_matching
            intent: >-
              Unassigned_NoFeatures is below 100,000, i.e. under half the fragments. Under an
              inverted -s 1 setting NoFeatures balloons to ~180,000 — six digits — and this fails.
              This is the empirical strandedness check the workflow input's own doc names.
            expected_value: 'Unassigned_NoFeatures\t[0-9]{1,5}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
      - workflow_label: "DESeq2 results: tnSWI1 vs AR0382"
        label_status: resolved
        description: >-
          Full raw result table for the first contrast, a single dataset (the workflow's only
          reduction happens here). HEADERLESS — deseq2.R writes it with `col.names = FALSE` — which
          is what makes `header_lines: '0'` correct on the two downstream Filter1 steps. The column
          contract is c1 GeneID, c2 baseMean, c3 log2FoldChange, c4 lfcSE, c5 stat, c6 pvalue,
          c7 padj, and it holds ONLY while `lfc_shrinkage_type: none`.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_n_columns
            intent: >-
              CROSS-STEP INVARIANT 1 DETECTOR. Seven columns. `get_result_output_columns()` drops the
              `stat` column under any shrinkage, which moves padj from c7 to c6 and makes the
              `c7<0.05` predicate — built in a different step — filter on nothing. If anyone changes
              `lfc_shrinkage_type` off `none` on either DESeq2 node, this assertion fails. Because
              the table is headerless, has_n_columns reads a data row here, which is exactly what is
              wanted.
            expected_value: 7
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: not_has_text
            intent: >-
              HEADER-ASYMMETRY ASSERTION. The word `baseMean` does not appear anywhere in this table.
              deseq_out is written with `col.names = FALSE`, so it carries no header row; this is the
              premise `header_lines: '0'` on the significance filters depends on, and the gene ID
              space (B9J08_ locus tags) guarantees the token cannot occur as data. If a wrapper
              upgrade ever starts emitting a header, this fails before the filters silently start
              passing a header line into a numeric comparison. The SAME assertion inverted applies to
              the normalized-counts tables below, which DO carry a header.
            expected_value: baseMean
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: "SCF1 is present as a full seven-field row in the raw result table."
            expected_value: 'B9J08_001458(\t[^\t]*){6}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: >-
              SCF1's log2FoldChange (c3) is negative — down in tnSWI1 relative to AR0382, the
              direction the paper reports. Magnitude is asserted only on the filtered table below,
              and the paper's ~29-fold figure is NOT asserted anywhere (see omissions).
            expected_value: 'B9J08_001458\t[^\t]+\t-[0-9.]+(\t[^\t]*){4}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: "DESeq2 results: AR0387 vs AR0382"
        label_status: resolved
        description: >-
          Full raw result table for the second contrast. Same shape, same header asymmetry, same
          column contract as the first. Its normalized counts are not interchangeable with the first
          contrast's — size factors are estimated over a different sample set.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_n_columns
            intent: "CROSS-STEP INVARIANT 1 DETECTOR on the second DESeq2 node. The invariant must hold on BOTH nodes, so it is asserted on both raw tables."
            expected_value: 7
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: not_has_text
            intent: "Headerless, as for the first contrast — the premise of `header_lines: '0'` on this contrast's two significance filters."
            expected_value: baseMean
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: "SCF1 is present as a full seven-field row."
            expected_value: 'B9J08_001458(\t[^\t]*){6}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: "SCF1's log2FoldChange is negative — down in AR0387 relative to AR0382, the paper's headline between-isolate result."
            expected_value: 'B9J08_001458\t[^\t]+\t-[0-9.]+(\t[^\t]*){4}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: "DESeq2 normalized counts: tnSWI1 vs AR0382"
        label_status: resolved
        description: >-
          Normalized counts for the first contrast's four samples. THE OPPOSITE HALF OF THE HEADER
          ASYMMETRY: deseq2.R writes this table with `col.names = NA`, so it DOES carry a header row
          — padded with a leading blank field so the header is the same width as the data rows. That
          is why sample-name text assertions are valid here and invalid on deseq_out. This output is
          also the only place the condition split's correctness becomes visible at workflow level,
          since every step of the split is hidden.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_line_matching
            intent: >-
              CROSS-STEP INVARIANT 2 DETECTOR (condition-split half). SCF1's row carries EXACTLY four
              numeric sample fields. If `header_lines` on `Select first-contrast samples` were ever
              changed from '0' to '1', Filter1 would pass row 1 of the sample metadata table
              (AR0382_A) through unconditionally regardless of the `c2=='tnSWI1'` predicate, that
              sample would enter the first-contrast counts sub-collection as well as the
              reference-level one, and this contrast would run on five sample columns instead of
              four. Asserting on a DATA row rather than on the header sidesteps the open question of
              whether the padded header counts as four or five fields.
            expected_value: 'B9J08_001458(\t[0-9.eE+-]+){4}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
          - family: has_text
            intent: >-
              The header names the four samples this contrast should see. Assert each of AR0382_A,
              AR0382_B, AR0382_tnSWI1_A and AR0382_tnSWI1_B as a separate has_text. This is the
              membership half of the condition-split check — the workflow's join key is the element
              identifier, and this is the only workflow-level output that shows which identifiers
              actually reached the reduction.
            expected_value: AR0382_tnSWI1_A
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
          - family: not_has_text
            intent: >-
              No AR0387 sample leaked into the first contrast. `AR0387` cannot occur as data — gene
              ids are B9J08_ locus tags — so this is an unambiguous negative on the split.
            expected_value: AR0387
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: "DESeq2 normalized counts: AR0387 vs AR0382"
        label_status: resolved
        description: >-
          Normalized counts for the second contrast's four samples. Headered, same as the first
          contrast, and the same condition-split visibility — here against a leak of the tnSWI1
          samples.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_line_matching
            intent: "CROSS-STEP INVARIANT 2 DETECTOR on the second contrast's split: SCF1's row carries exactly four numeric sample fields."
            expected_value: 'B9J08_001458(\t[0-9.eE+-]+){4}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
          - family: has_text
            intent: "The header names the four samples this contrast should see: AR0382_A, AR0382_B, AR0387_A and AR0387_B, each as a separate has_text."
            expected_value: AR0387_A
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
          - family: not_has_text
            intent: "No tnSWI1 sample leaked into the second contrast. `tnSWI1` cannot occur as data."
            expected_value: tnSWI1
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: "Significant genes: tnSWI1 vs AR0382"
        label_status: resolved
        description: >-
          Endpoint of the first contrast's two-link chain — rows of the raw table that cleared BOTH
          adjusted p-value < 0.05 (c7) and |log2 fold change| > 1 (abs(c3)). Membership in this table
          IS the paper's significance criterion, which is why the biological claim is asserted here
          rather than by re-deriving thresholds from the raw table.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_n_columns
            intent: "Seven columns — the filter chain preserved the raw table's shape rather than reshaping it, and invariant 1 still holds at the chain's endpoint."
            expected_value: 7
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: >-
              THE PAPER'S CENTRAL CLAIM, FIRST CONTRAST. SCF1 appears in the significant-genes table,
              which by construction means padj < 0.05 AND |log2FC| > 1. Phase 7 measured 414/391
              fragments in AR0382 against 4/3 in tnSWI1 — a separation DESeq2 will call significant
              at n=2 with a wide margin.
            expected_value: 'B9J08_001458(\t[^\t]*){6}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: >-
              SCF1's log2FoldChange is at most -2, i.e. at least 4-fold down. This is a threshold,
              deliberately far below the ~270-fold raw separation phase 7 measured and deliberately
              NOT the paper's ~29-fold figure, which this plan does not assert anywhere. The regex
              encodes "negative with an integer part of 2 or more". MEASURED for this contrast:
              log2FC -6.93 at padj 2.1e-31, so the -2 floor clears by ~5 log2 units and the 0.05 cut
              by 31 orders of magnitude. The measured value is NOT the assertion: it comes from
              pydeseq2 over bwa-mem counts, not from the wrapper's R DESeq2 over featureCounts -s 2
              output, and will differ in the last digits at least (warnings[]
              `deseq2-statistics-measured-with-pydeseq2-not-r`). The margin is what makes the
              threshold safe across that difference.
            expected_value: 'B9J08_001458\t[^\t]+\t-(?:[2-9]|[1-9][0-9]+)\.[0-9]+(\t[^\t]*){4}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: "Significant genes: AR0387 vs AR0382"
        label_status: resolved
        description: >-
          Endpoint of the second contrast's chain. The paper's headline between-isolate result: SCF1
          the most down-regulated gene between AR0382 and the low-adhesion clinical isolate AR0387.
          Measurement backs "most down-regulated" HERE — SCF1 ranks 1st by |log2FC| among the 35
          genes clearing both filters, and 1st among the 9 down-regulated ones — but not in the first
          contrast, where it ranks 2nd of 63 (2nd of 52 down-regulated). No rank is asserted on
          either table; see omissions[] for why, and warnings[] for the implementation caveat on
          those figures.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_n_columns
            intent: "Seven columns at the second chain's endpoint."
            expected_value: 7
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: >-
              THE PAPER'S CENTRAL CLAIM, SECOND CONTRAST. SCF1 cleared both filters. Phase 7 measured
              414/391 against 2/1 — the widest separation in the fixture.
            expected_value: 'B9J08_001458(\t[^\t]*){6}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_line_matching
            intent: >-
              SCF1's log2FoldChange is at most -2. Same threshold and same refusal to assert the
              paper's ~29-fold magnitude. MEASURED for this contrast: log2FC -7.95 at padj 4.1e-18,
              clearing the -2 floor by ~6 log2 units and the 0.05 cut by 18 orders of magnitude. Same
              implementation caveat as the first contrast — pydeseq2 over bwa-mem counts, not the
              wrapper's R DESeq2 — so the margin is asserted and the value is not.
            expected_value: 'B9J08_001458\t[^\t]+\t-(?:[2-9]|[1-9][0-9]+)\.[0-9]+(\t[^\t]*){4}'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: "DESeq2 diagnostic plots: tnSWI1 vs AR0382"
        label_status: resolved
        description: >-
          Multi-page PDF (dispersion estimates, PCA, MA plot, sample distance heatmap). Deliberately
          weak check: the PDF embeds a creation timestamp and the plots are the user-facing view of
          numbers already asserted exactly on the tables above.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_size
            intent: >-
              Non-trivially sized PDF (min only, ~10 KB), catching the "R rendered an empty device"
              failure mode. Recorded as a deliberate accepted shortcut per iwc-shortcuts-anti-patterns
              §2-3: the sibling tabular checkpoints carry the content, so a size band adds nothing a
              min bound does not.
            expected_value: 10000
            tolerance:
              kind: none
              magnitude: null
              rationale: "Min bound rather than a delta band; PDF size depends on the gene count that survives independent filtering and is not predicted here."
            element_identifier: null
            evidence: intent
            confidence: medium
      - workflow_label: "DESeq2 diagnostic plots: AR0387 vs AR0382"
        label_status: resolved
        description: Multi-page PDF for the second contrast. Same deliberate weak check.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_size
            intent: "Non-trivially sized PDF (min only, ~10 KB)."
            expected_value: 10000
            tolerance:
              kind: none
              magnitude: null
              rationale: "Min bound rather than a delta band, as for the first contrast."
            element_identifier: null
            evidence: intent
            confidence: medium

  - id: topology-smoke-25k
    doc: >-
      Fast structural smoke test over the same six samples at 25,000 read pairs each (~7 MB
      compressed total, small enough to commit in-repo). Makes NO biological claim: at that depth
      SCF1 falls to roughly 50 fragments in AR0382 and near zero elsewhere, per-gene depth drops to
      ~4.5 reads/gene, and the significance tables may legitimately be empty. What it does buy is a
      CI lane that is not hostage to a Zenodo fetch and that still exercises every structural thing
      that can silently break: the sample_sheet:paired input shape, the flatten side-branch, the
      six-element map-over identifier space, both condition-split joins, both DESeq2 reductions, and
      BOTH cross-step invariants. Every assertion below is a subset of case 1's, with the depth
      constant and the biological claims removed.
    derived_from: intent
    provenance: >-
      test-data-refs.json `subsetting.shape_only_alternative`, which records this depth as a
      committable option and states exactly what it loses. The fixtures at this depth were NOT
      generated by phase 7 — only the 200k set was — so they must be produced by the same
      deterministic recipe with `head -n 100000`.
    job_inputs:
      - workflow_label: RNA-seq reads (sample sheet)
        label_status: resolved
        description: >-
          The same six samples and the same six element identifiers as case 1, at 25,000 read pairs
          each. The identifier spine is unchanged, because it is the thing this case exists to
          exercise.
        collection_shape: sample_sheet:paired
        datatype: fastqsanger.gz
        fixture:
          storage: in-repo
          location: test-data/
          checksum: null
          provenance: >-
            NOT YET GENERATED. Produce with the same deterministic head-subset recipe as the 200k
            fixture but `head -n 100000` per mate, from the same six ENA runs; the source URLs and
            source md5s are in test-data-refs.json `inputs["RNA-seq reads (sample sheet)"].elements`.
            Record an uncompressed md5 per file, for the same gzip-instability reason. At ~7 MB total
            this is under the ~1 MB-per-file IWC in-repo threshold only marginally, so a reviewer may
            still ask for Zenodo hosting; see unresolved `smoke-fixture-not-generated`.
      - workflow_label: Reference genome FASTA
        label_status: resolved
        description: Identical to case 1 — the reference is used whole at both depths.
        collection_shape: null
        datatype: fasta
        fixture:
          storage: remote-url
          location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.fna.gz
          checksum: MD5:a008b270d3aaa04736a8bb7daf3f6dd5
          provenance: "Pinned NCBI FTP URL, staged with `decompress: true`. Same as case 1."
      - workflow_label: Gene annotation GTF
        label_status: resolved
        description: Identical to case 1 — the annotation is what fixes the gene ID space, so it must not differ between cases.
        collection_shape: null
        datatype: gtf
        fixture:
          storage: remote-url
          location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.gtf.gz
          checksum: MD5:6e5b9528d48c0a8fc2c8588e6eeea929
          provenance: "Pinned NCBI FTP URL, staged with `decompress: true`. Same as case 1."
      - workflow_label: Reference condition level
        label_status: resolved
        description: Workflow default AR0382, as case 1.
        collection_shape: null
        datatype: null
        fixture: {storage: null, location: null, checksum: null, provenance: "Workflow default."}
      - workflow_label: First contrast condition level
        label_status: resolved
        description: Workflow default tnSWI1, as case 1.
        collection_shape: null
        datatype: null
        fixture: {storage: null, location: null, checksum: null, provenance: "Workflow default."}
      - workflow_label: Second contrast condition level
        label_status: resolved
        description: Workflow default AR0387, as case 1.
        collection_shape: null
        datatype: null
        fixture: {storage: null, location: null, checksum: null, provenance: "Workflow default."}
      - workflow_label: Strandedness
        label_status: resolved
        description: Workflow default `stranded - reverse`, as case 1. Strandedness is a property of the library, not of the subset depth.
        collection_shape: null
        datatype: null
        fixture: {storage: null, location: null, checksum: null, provenance: "Workflow default, measured by phase 7."}
      - workflow_label: Adjusted p-value threshold
        label_status: resolved
        description: Workflow default 0.05, as case 1 — left alone so the filter chain is exercised in its shipped configuration even though no significance claim is made.
        collection_shape: null
        datatype: null
        fixture: {storage: null, location: null, checksum: null, provenance: "Workflow default."}
      - workflow_label: log2 fold change threshold
        label_status: resolved
        description: Workflow default 1.0, as case 1.
        collection_shape: null
        datatype: null
        fixture: {storage: null, location: null, checksum: null, provenance: "Workflow default."}
    expected_outputs:
      - workflow_label: "FastQC raw reads: text summary"
        label_status: resolved
        description: Twelve elements. The depth constant is the only thing that changes from case 1.
        output_kind: collection
        collection_shape: list
        assertion_intent:
          - family: has_text_matching
            intent: "Each of the twelve elements reports exactly 25,000 sequences, proving the reduced fixture is the depth it claims and that the flatten side-branch fanned out to twelve."
            expected_value: 'Total Sequences\s+25000'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: Gene counts per sample
        label_status: resolved
        description: >-
          Six elements. The row count and gene ID space are depth-independent — featureCounts emits
          one row per annotated gene regardless of coverage — so the annotation-space guard carries
          over unchanged. The SCF1 count thresholds do NOT carry over.
        output_kind: collection
        collection_shape: list
        assertion_intent:
          - family: has_n_columns
            intent: "Two-column gene-id/count table."
            expected_value: 2
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_n_lines
            intent: "5586 annotated genes, depth-independent. Same ±1 header allowance as case 1."
            expected_value: 5586
            tolerance:
              kind: delta
              magnitude: 1
              rationale: "Header-row presence on featureCounts output_short is inferred, not verified."
            element_identifier: null
            evidence: intent
            confidence: medium
          - family: has_line_matching
            intent: "The annotation-space guard: at least 5000 rows are a bare B9J08_ locus tag plus an integer count."
            expected_value: 'B9J08_[0-9]{6}\t[0-9]+'
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: "DESeq2 results: tnSWI1 vs AR0382"
        label_status: resolved
        description: Raw result table, first contrast. Carries both invariant detectors, which are structural and therefore depth-independent.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_n_columns
            intent: "CROSS-STEP INVARIANT 1 DETECTOR: seven columns, i.e. `lfc_shrinkage_type` is still `none` and c7 is still padj."
            expected_value: 7
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: not_has_text
            intent: "Headerless, the premise of `header_lines: '0'` on the significance filters."
            expected_value: baseMean
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: "DESeq2 results: AR0387 vs AR0382"
        label_status: resolved
        description: Raw result table, second contrast. Both invariant detectors again, because invariant 1 must hold on BOTH DESeq2 nodes.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_n_columns
            intent: "CROSS-STEP INVARIANT 1 DETECTOR on the second node."
            expected_value: 7
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: not_has_text
            intent: "Headerless."
            expected_value: baseMean
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
      - workflow_label: "DESeq2 normalized counts: tnSWI1 vs AR0382"
        label_status: resolved
        description: Headered table. Carries the condition-split half of invariant 2 without needing any biological signal.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_text
            intent: >-
              The header names exactly the four samples of this contrast — AR0382_A, AR0382_B,
              AR0382_tnSWI1_A, AR0382_tnSWI1_B — each asserted separately.
            expected_value: AR0382_tnSWI1_A
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
          - family: not_has_text
            intent: "CROSS-STEP INVARIANT 2 DETECTOR (condition-split half): no AR0387 sample leaked into the first contrast."
            expected_value: AR0387
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_n_columns
            intent: >-
              Five fields on the first line — one padded row-name column plus four samples. A
              `header_lines: '1'` regression on `Select first-contrast samples` would leak AR0382_A
              into this contrast and make it six. See unresolved
              `normalized-counts-header-width-unverified`: if the wrapper's header turns out NOT to
              carry the leading blank field, the correct n is 4, and implement-galaxy-workflow-test
              must settle this from the first real invocation.
            expected_value: 5
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
      - workflow_label: "DESeq2 normalized counts: AR0387 vs AR0382"
        label_status: resolved
        description: Headered table, second contrast. Same split check against a tnSWI1 leak.
        output_kind: dataset
        collection_shape: null
        assertion_intent:
          - family: has_text
            intent: "The header names exactly AR0382_A, AR0382_B, AR0387_A, AR0387_B, each asserted separately."
            expected_value: AR0387_A
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium
          - family: not_has_text
            intent: "CROSS-STEP INVARIANT 2 DETECTOR on the second split: no tnSWI1 sample leaked in."
            expected_value: tnSWI1
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: high
          - family: has_n_columns
            intent: "Five fields on the first line, same caveat as the first contrast."
            expected_value: 5
            tolerance: null
            element_identifier: null
            evidence: intent
            confidence: medium

unresolved:
  - kind: fixture
    description: >-
      The twelve 200,000-read-pair FASTQ fixtures for case 1 were generated and hashed by phase 7 but
      never published; they existed only in that session's scratchpad. `location` is null and
      `<FIXTURE_BASE>` in test-data-refs.json `planemo_test_job_block` has no value to substitute.
      Case 1 cannot run until they are regenerated and hosted.
    blocking: true
    suggested_resolution: >-
      Regenerate with test-data-refs.json `subsetting.recipe` (stream each ENA FASTQ, `head -n
      800000`), verify each against its `md5_uncompressed`, publish to Zenodo, substitute the record
      URL for `<FIXTURE_BASE>`. Mirrors open requirement `test-fixtures-not-hosted`.
  - kind: fixture
    description: >-
      Case 2's 25,000-read-pair fixtures do not exist at all. Phase 7 costed and characterized this
      depth (`subsetting.shape_only_alternative`) but generated only the 200k set.
    blocking: true
    suggested_resolution: >-
      Same recipe with `head -n 100000` per mate from the same six ENA runs. Decide in-repo versus
      Zenodo at that point: ~7 MB total is committable but a reviewer may still push back on twelve
      binary files in `test-data/`. If the answer is Zenodo, case 2 loses its main advantage over
      case 1 and should be reconsidered rather than kept for its own sake.
  - kind: fixture
    description: >-
      Whether a `hashes:` block on an input that also sets `decompress: true` validates the
      compressed bytes as fetched or the decompressed bytes was not established. The reference FASTA
      and GTF carry md5s of the COMPRESSED NCBI files and are staged with `decompress: true`, so a
      wrong assumption here fails the fetch with a hash mismatch rather than anything diagnostic.
    blocking: false
    suggested_resolution: >-
      Confirm against Galaxy's fetch handler before adding `hashes:` to those two inputs; if it
      cannot be confirmed cheaply, omit the hashes there (the URLs are stable NCBI FTP paths) and
      keep them on the read fixtures, where no decompression happens.
  - kind: fixture
    description: >-
      No `hashes:` block can be authored for the twelve read fixtures from what phase 7 supplies.
      Phase 7 pinned UNCOMPRESSED md5s deliberately, because gzip output is not byte-stable across
      gzip versions, but the staged file is the .gz and that is what a `hashes:` block would check.
      IWC convention puts a hash on every remote input `location:`, so this will be noticed in review.
    blocking: false
    suggested_resolution: >-
      Hash the published .gz files once, after hosting, and record those hashes in the test file —
      the uncompressed md5s stay in test-data-refs.json as the regeneration check, and the published
      .gz hashes become the fetch-integrity check. The two serve different purposes and both are
      needed.
  - kind: assertion
    description: >-
      Two of phase 7's biological claims are RANK claims and no tests-format assertion family can
      express a rank or a row position: "SCF1 is within the 10 smallest padj rows" and "SCF1's
      normalized count sits in the top 5% of the table". has_line_matching has no positional
      anchoring; has_n_lines cannot bound where a match occurs. The padj-rank claim is TRUE as
      measured — SCF1 is 5th of the 2393 genes carrying a padj in the first contrast and 4th of 1376
      in the second — so it sits here because it is INEXPRESSIBLE, not because it is doubted. It is
      also the weaker form of a claim this plan already asserts well: a rank of 5 has no margin
      against the implementation difference recorded in warnings[]
      `deseq2-statistics-measured-with-pydeseq2-not-r`, whereas the padj threshold it rests on clears
      by 31 and 18 orders of magnitude. Promoting it is not worth a workflow change.
    blocking: false
    suggested_resolution: >-
      Either accept membership in the significant-genes table as the proxy (what this plan does), or,
      if the rank claim must be tested, add a `Select first` step after each DESeq2 node and promote
      its output — which turns a rank into an ordinary membership assertion. That is a workflow
      change, not a test change, and should be weighed against the output-list clutter.
  - kind: assertion
    description: >-
      The `has_n_columns` value for the DESeq2 normalized-counts tables depends on whether
      `col.names = NA` produces a header padded with a leading blank field (making the header the
      same width as the data rows, n=5 for four samples) or an unpadded one (n=4). Galaxy's
      has_n_columns reads the FIRST line only, so the two differ. Case 1 sidesteps this by asserting
      on SCF1's data row instead; case 2 uses has_n_columns and carries the risk.
    blocking: false
    suggested_resolution: >-
      Settle from the first successful invocation — `planemo workflow_test_init --from_invocation`
      will show the real header — and correct case 2's n before committing the test file.
  - kind: assertion
    description: >-
      Whether the DESeq2 wrapper names normalized-counts columns by the input dataset ELEMENT
      IDENTIFIER (AR0382_A) or by some other dataset name is assumed, not verified. Every
      sample-name has_text and not_has_text assertion on the two normalized-counts outputs, and
      therefore the condition-split half of invariant 2, rests on it.
    blocking: false
    suggested_resolution: >-
      Confirm from the first invocation. If the wrapper uses a different naming scheme, the
      `not_has_text: AR0387` / `not_has_text: tnSWI1` negatives survive only if the condition token
      still appears in the column name; otherwise fall back to the SCF1-row field-count assertion as
      the sole split detector.
  - kind: collection-shape
    description: >-
      The promoted per-sample outputs (FastQC text/HTML, Cutadapt report, Trimmed reads, STAR BAM and
      summary, Gene counts, featureCounts summary) are recorded here with their BEHAVIOURAL shapes
      `list` and `list:paired`, but a tool mapped over a `sample_sheet:paired` collection produces a
      sample_sheet-shaped collection without `column_definitions`, whose declared collection_type
      string is not `list`. This is open requirement `mapped-outputs-carry-sample-sheet-outer-axis`,
      which was filed explicitly for this phase.
    blocking: false
    suggested_resolution: >-
      Do NOT author `attributes: {collection_type: ...}` assertions on any of those eight outputs —
      the plan asserts element identifiers instead, which is what actually matters and is
      shape-agnostic. Record the real type strings from the first invocation and only then decide
      whether a collection_type assertion is worth having.
  - kind: assertion
    description: >-
      featureCounts `output_short` header-row presence was not verified. The DESeq2 steps bind
      `header: true` on their counts inputs, which implies one, and the `has_n_lines` assertions
      carry a ±1 allowance for it, but the featureCounts wrapper itself was not read in this run.
    blocking: false
    suggested_resolution: "Read the pinned iuc/featurecounts 2.1.1+galaxy1 wrapper, or take the line count from the first invocation, and tighten the ±1 to an exact value."

omissions:
  - target: "FastQC raw reads: HTML report"
    reason: >-
      The sibling text summary is promoted precisely so it can carry the assertions, and it does —
      exact sequence counts and the Basic Statistics module per element. The HTML embeds the FastQC
      version banner and base64-encoded images, so any content assertion on it either restates the
      text summary or pins a version string. Asserting both would be duplication, not coverage.
    category: weak-output
  - target: "STAR alignments (BAM)"
    reason: >-
      BAM is a gzipped block format whose header carries @PG lines with full command lines and Galaxy
      job ids, so it is never byte-stable; and the workflow already promotes `STAR mapping summary`
      as the assertable text behind it, which is asserted. Per planemo-asserts-idioms the fallback
      would be has_size plus has_archive_member, which buys nothing the mapping summary does not
      already prove.
    category: weak-output
  - target: "The paper's ~29-fold SCF1 expression difference between AR0382 and AR0387"
    reason: >-
      REFUSED DELIBERATELY. Direct measurement of the deposited data disagrees by roughly an order of
      magnitude — ~270-fold on aligned fragments at fixture depth, ~158-fold by exact 31-mer matching
      over 5.2 M reads. Reference bias is ruled out as the explanation, because AR0387 IS the B8441
      reference strain. The 29-fold may be the RT-qPCR assay rather than the RNA-seq, a shrunk rather
      than raw log2FC, or a different normalization; nothing in this run distinguishes them. Open
      requirement `scf1-fold-change-magnitude-disagrees-with-paper`. Direction and significance are
      asserted; magnitude is asserted only as a loose `<= -2` floor.
    category: out-of-scope
  - target: "SCF1 is the single top gene ranked by |log2 fold change|"
    reason: >-
      REFUSED DELIBERATELY, AND NOW MEASURED FALSE AS A CROSS-CONTRAST CLAIM. The mechanism was
      predicted correctly: with `lfc_shrinkage_type: none` — which this workflow pins, for the
      column-contract reason that is invariant 1 — low-count genes take extreme unshrunk log2FC
      values and can outrank SCF1. Measurement confirms it, and also kills the fallback this rationale
      previously proposed. Phase 8 ran pydeseq2 on a strand-aware count matrix built from phase 7's
      six BAMs, as two separate two-level n=2 analyses with no shrinkage, mirroring the workflow's two
      DESeq2 nodes. SCF1 (B9J08_001458) came out: tnSWI1 vs AR0382 — log2FC -6.93, padj 2.1e-31, 5th
      of 2393 genes with a padj, 2nd of the 63 clearing padj<0.05 and |log2FC|>1 by |log2FC| (2nd of
      the 52 down-regulated ones); AR0387 vs AR0382 — log2FC -7.95, padj 4.1e-18, 4th of 1376 by padj,
      1st of 35 by |log2FC| (1st of 9 down-regulated). So NO rank-1 assertion holds in both contrasts:
      not by |log2FC| (2nd, then 1st), and not by adjusted p-value either (5th, then 4th). The earlier
      claim in this slot that "ranking by adjusted p-value is the safe form" was wrong and has been
      removed. Membership plus a threshold is the honest form and is what the plan asserts: SCF1
      appears in each significant-genes table — which by construction means padj < 0.05 and |log2FC| >
      1 — with log2FoldChange <= -2. That form is also the robust one. It clears by 31 and 18 orders
      of magnitude on padj and ~5 and ~6 log2 units on fold change, which is margin enough to survive
      the fact that these numbers come from pydeseq2 over bwa-mem counts rather than the wrapper's R
      DESeq2 over featureCounts -s 2 (warnings[] `deseq2-statistics-measured-with-pydeseq2-not-r`); a
      rank of 1, 2, 4 or 5 has no such margin and would be a flaky assertion even if tests-format
      could express it, which it cannot (see unresolved).
    category: other
  - target: "No significant dysregulation of the ALS or IFF/HYR adhesin families in the tnSWI1 contrast"
    reason: >-
      REFUSED DELIBERATELY. A negative claim over twelve named genes (B9J08_002582/_004498/_004112
      ALS; _004100/_004109/_004098/_004110/_001531/_004892/_001155/_004451/_000675 IFF-HYR). At
      200,000 read pairs most of those sit at counts where DESeq2 simply lacks power, so their
      absence from the significant table is a depth artefact rather than evidence for the paper's
      claim. Asserting it would be a fixture claiming an outcome it cannot produce — exactly the
      defect the Foundry's fixture rule names. It would need full-depth runs.
    category: out-of-scope
  - target: "The tnBCY1 vs AR0382 contrast (Fig. S2)"
    reason: >-
      No tnBCY1 run was ever deposited in PRJNA904261; only three conditions exist. There is no test
      input for it at any depth, and the workflow does not model it.
    category: out-of-scope
  - target: "Exact per-gene integer counts on `Gene counts per sample`"
    reason: >-
      The IWC comparison notes recommend being stricter than the corpus default here, and they are
      right in principle — featureCounts on a fixed reference, annotation and strandedness is
      deterministic. But phase 7 measured with bwa-mem plus a strand-agnostic exon counter, not with
      RNA STAR plus featureCounts -s 2, so it cannot supply the exact integers and inventing them
      would be a fixture asserting something nothing measured. The plan asserts thresholds with a
      stated margin instead, and the exact-count upgrade is an explicit follow-up once a real
      invocation exists.
    category: other
  - target: "Negative / failure-path test cases"
    reason: >-
      `expect_failure:` does not appear anywhere in the 115-file IWC test corpus. Error-path testing
      belongs in tool wrappers, not workflow tests. No such case is authored here.
    category: out-of-scope

warnings:
  - code: fixture-measured-assertions-lack-an-evidence-class
    message: >-
      Every assertion carries `evidence: intent` because the schema's only alternative is
      `test-evidence`, which means "translated from upstream test fixtures". A large share of these
      assertions are neither: they were measured directly against the resolved fixtures by phase 7.
      Confidence is raised to `high` where that is the case, but a reader filtering on `evidence`
      will see a uniformly synthesized plan and undercount its grounding.
    path: "test_cases[*].expected_outputs[*].assertion_intent[*].evidence"
  - code: full-line-anchoring-convention
    message: >-
      Every `has_line_matching` expected_value in this plan is written to match a WHOLE line, because
      Galaxy evaluates has_line_matching as `^(?:expression)$` under re.MULTILINE. The
      `has_text_matching` values are written for `re.search` over the whole content instead, and `^`
      would NOT be multiline there. Do not move an expression between the two families without
      rewriting it.
    path: "test_cases[*].expected_outputs[*].assertion_intent[*].expected_value"
  - code: iwc-notes-and-test-data-refs-disagree-on-count-strictness
    message: >-
      iwc-comparison-notes.md §5 recommends exact per-gene integer counts on `Gene counts per
      sample`; test-data-refs.json `unresolved.counts-measured-with-bwa-not-star` forbids pinning any
      exact count. Resolved in favour of the later and better-grounded artifact: thresholds now,
      exact counts as a recorded follow-up once a real invocation exists. Both positions are correct
      about different moments in the run.
    path: "test_cases[0].expected_outputs[4]"
  - code: deseq2-at-reduced-depth-untested
    message: >-
      Case 2 assumes DESeq2 completes at 25,000 read pairs per sample (~4.5 reads/gene average). It
      should — the gene count and sample count are unchanged and only the counts shrink — but no run
      at that depth has happened, and a DESeq2 failure there would surface as a whole-case failure
      rather than an assertion failure. Run case 2 once before relying on it as the fast CI lane.
    path: "test_cases[1]"
  - code: reference-level-filter-regression-undetectable
    message: >-
      Of the seven Filter1 steps that bind `header_lines: '0'`, this plan can detect a regression on
      only two — `Select first-contrast samples` and `Select second-contrast samples`. On `Select
      reference-level samples` the leaked row IS AR0382_A, which the `c2=='AR0382'` predicate keeps
      anyway, so the output is identical and no assertion can see it. On the four significance
      filters the leaked row is row 1 of a padj-sorted table, which passes both filters on its own
      merits, so the regression is again invisible at workflow level. See the report for the full
      accounting.
    path: "test_cases[*].expected_outputs[*]"
  - code: deseq2-statistics-measured-with-pydeseq2-not-r
    message: >-
      Every differential-expression number quoted in this plan — SCF1's log2 fold change and adjusted
      p-value per contrast, and every rank figure in omissions[] and unresolved[] — was produced by
      pydeseq2 over a strand-aware count matrix built from phase 7's six bwa-mem BAMs, as two separate
      two-level n=2 analyses with no LFC shrinkage. The workflow runs the Galaxy DESeq2 wrapper, i.e.
      R DESeq2, over RNA STAR plus featureCounts -s 2. pydeseq2 is a faithful reimplementation, but it
      is not the same implementation and the counts are not the same counts, so exact values will
      differ. NO number from that measurement may be turned into an equality assertion, and none is.
      What survives the difference is margin: padj 2.1e-31 and 4.1e-18 against a 0.05 threshold,
      log2FC -6.93 and -7.95 against a -2 floor. What does not survive it is rank — 5th and 4th by
      padj, 2nd and 1st by |log2FC| — or any exact value, which is why neither is asserted.
    path: "test_cases[*].expected_outputs[*].assertion_intent[*].intent"
  - code: no-corpus-precedent-for-sample-sheet-test-fixture
    message: >-
      No IWC test fixture in the 115-file corpus declares a `sample_sheet` collection or any
      `column_definitions`, so the job block for the reads input has no worked example to
      pattern-match against. Phase 7 established from Galaxy's source that it IS expressible and
      pinned the positional-list `rows:` shape; that reading is the plan's only authority here, and
      the first `planemo test` run is what confirms it.
    path: 'test_cases[*].job_inputs[0]'
galaxy-workflowpresentgalaxy-workflow.gxwf.yml

Concrete gxformat2 workflow (`class: GalaxyWorkflow`) extracted from the fully-concretized draft at loop endstate via [[draft-extract]]: drafty steps dropped, `_plan_*` planning fields stripped, class promoted. The runnable, testable artifact that downstream Molds ([[implement-galaxy-workflow-test]], [[validate-galaxy-workflow]], [[run-workflow-test]]) consume.

declared by
advance-galaxy-draft-step (phase 6)
consumed at
9
schema
none declared
sha256
1c75a7f7eb052ddb200fd008c90404cd977d6eddca00e3e2a5713711c14e6831

GalaxyWorkflow — 27 step(s), 0 still drafty, 9 input(s), 16 output(s).

steplabeltoolstateplan keys
Flatten reads for per-fastq QCFlatten reads for per-fastq QC__FLATTEN__resolved
Read quality reportRead quality reporttoolshed.g2.bx.psu.edu/repos/devteam/fastqc/fastqcresolved
Quality-trim readsQuality-trim readstoolshed.g2.bx.psu.edu/repos/lparsons/cutadapt/cutadaptresolved
Splice-aware alignmentSplice-aware alignmenttoolshed.g2.bx.psu.edu/repos/iuc/rgrnastar/rna_starresolved
Get featureCounts strandedness parameterGet featureCounts strandedness parametertoolshed.g2.bx.psu.edu/repos/iuc/map_param_value/map_param_valueresolved
Count reads per geneCount reads per genetoolshed.g2.bx.psu.edu/repos/iuc/featurecounts/featurecountsresolved
Project sample sheet to tabularProject sample sheet to tabular__SAMPLE_SHEET_TO_TABULAR__resolved
Build reference-level row predicateBuild reference-level row predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Select reference-level samplesSelect reference-level samplesFilter1resolved
Reference-level identifier listReference-level identifier listCut1resolved
Counts for the reference levelCounts for the reference level__FILTER_FROM_FILE__resolved
Build first-contrast row predicateBuild first-contrast row predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Select first-contrast samplesSelect first-contrast samplesFilter1resolved
First-contrast identifier listFirst-contrast identifier listCut1resolved
Counts for the first contrast levelCounts for the first contrast level__FILTER_FROM_FILE__resolved
Build second-contrast row predicateBuild second-contrast row predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Select second-contrast samplesSelect second-contrast samplesFilter1resolved
Second-contrast identifier listSecond-contrast identifier listCut1resolved
Counts for the second contrast levelCounts for the second contrast level__FILTER_FROM_FILE__resolved
Differential expression: first contrast vs referenceDifferential expression: first contrast vs referencetoolshed.g2.bx.psu.edu/repos/iuc/deseq2/deseq2resolved
Differential expression: second contrast vs referenceDifferential expression: second contrast vs referencetoolshed.g2.bx.psu.edu/repos/iuc/deseq2/deseq2resolved
Build adjusted p-value predicateBuild adjusted p-value predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Build log2 fold-change predicateBuild log2 fold-change predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Filter with p-adj threshold: first contrastFilter with p-adj threshold: first contrastFilter1resolved
Filter with log2 FC threshold: first contrastFilter with log2 FC threshold: first contrastFilter1resolved
Filter with p-adj threshold: second contrastFilter with p-adj threshold: second contrastFilter1resolved
Filter with log2 FC threshold: second contrastFilter with log2 FC threshold: second contrastFilter1resolved
galaxy-workflow-draftpresentgalaxy-workflow-draft.gxwf.yml

gxformat2 draft (see [[galaxy-workflow-draft-format]]): topology fully resolved (workflow inputs, outputs, step set, edges); tool_id / state / tool_shed_repository and wrapper-determined port names may be TODO with free-text _plan_state / _plan_context / _plan_in / _plan_out per step for later implementation Molds.

declared by
advance-galaxy-draft-step, freeform-summary-to-galaxy-template (phase 5, 6)
consumed at
6
schema
galaxy-workflow-draft
sha256
1b3152706ce1e30738ef6bae539a073ddbb2dd9d867a7f394ddce90d9dcbb5ee

GalaxyWorkflowDraft — 27 step(s), 0 still drafty, 9 input(s), 16 output(s).

steplabeltoolstateplan keys
Flatten reads for per-fastq QCFlatten reads for per-fastq QC__FLATTEN__resolved
Read quality reportRead quality reporttoolshed.g2.bx.psu.edu/repos/devteam/fastqc/fastqcresolved
Quality-trim readsQuality-trim readstoolshed.g2.bx.psu.edu/repos/lparsons/cutadapt/cutadaptresolved
Splice-aware alignmentSplice-aware alignmenttoolshed.g2.bx.psu.edu/repos/iuc/rgrnastar/rna_starresolved
Get featureCounts strandedness parameterGet featureCounts strandedness parametertoolshed.g2.bx.psu.edu/repos/iuc/map_param_value/map_param_valueresolved
Count reads per geneCount reads per genetoolshed.g2.bx.psu.edu/repos/iuc/featurecounts/featurecountsresolved
Project sample sheet to tabularProject sample sheet to tabular__SAMPLE_SHEET_TO_TABULAR__resolved
Build reference-level row predicateBuild reference-level row predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Select reference-level samplesSelect reference-level samplesFilter1resolved
Reference-level identifier listReference-level identifier listCut1resolved
Counts for the reference levelCounts for the reference level__FILTER_FROM_FILE__resolved
Build first-contrast row predicateBuild first-contrast row predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Select first-contrast samplesSelect first-contrast samplesFilter1resolved
First-contrast identifier listFirst-contrast identifier listCut1resolved
Counts for the first contrast levelCounts for the first contrast level__FILTER_FROM_FILE__resolved
Build second-contrast row predicateBuild second-contrast row predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Select second-contrast samplesSelect second-contrast samplesFilter1resolved
Second-contrast identifier listSecond-contrast identifier listCut1resolved
Counts for the second contrast levelCounts for the second contrast level__FILTER_FROM_FILE__resolved
Differential expression: first contrast vs referenceDifferential expression: first contrast vs referencetoolshed.g2.bx.psu.edu/repos/iuc/deseq2/deseq2resolved
Differential expression: second contrast vs referenceDifferential expression: second contrast vs referencetoolshed.g2.bx.psu.edu/repos/iuc/deseq2/deseq2resolved
Build adjusted p-value predicateBuild adjusted p-value predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Build log2 fold-change predicateBuild log2 fold-change predicatetoolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_paramresolved
Filter with p-adj threshold: first contrastFilter with p-adj threshold: first contrastFilter1resolved
Filter with log2 FC threshold: first contrastFilter with log2 FC threshold: first contrastFilter1resolved
Filter with p-adj threshold: second contrastFilter with p-adj threshold: second contrastFilter1resolved
Filter with log2 FC threshold: second contrastFilter with log2 FC threshold: second contrastFilter1resolved
galaxy-workflow-testpresentgalaxy-workflow.gxwf-tests.yml

Galaxy workflow test file (tests-format) with job inputs, expected outputs, assertions; passes static schema + label cross-check. Named as the workflow basename + `-tests.yml` so Planemo discovers it as the companion of `galaxy-workflow.gxwf.yml`.

declared by
implement-galaxy-workflow-test (phase 9)
consumed at
nothing downstream reads it
schema
none declared
sha256
2b3aea677dcae22115a55e6e2d56d33ab220527669a75be8afbb0891dca38476
# Galaxy workflow tests for galaxy-workflow.gxwf.yml
#
# Two cases:
#   1. scf1-both-contrasts-200k -- the biological case. Six ENA runs at 200,000 read pairs.
#   2. topology-smoke-25k       -- structural smoke test at 25,000 read pairs, no biology.
#
# READ FIXTURES. test-data/*.fastq.gz (200k) and test-data/smoke-25k/*.fastq.gz (25k) are
# deterministic head subsets of the six PRJNA904261 runs, produced by
#   curl -L <ENA url> | gzip -dc | head -n 800000   (200k pairs;  -n 100000 for 25k pairs)
# and gzipped with `gzip -n`. Each 200k .fastq was verified byte-for-byte against the
# UNCOMPRESSED md5 pinned in test-data-refs.json before compression -- all twelve matched --
# and that md5 is the regeneration check, because gzip output is not byte-stable across gzip
# versions. The 25k set was then derived locally as the first 100,000 lines of each VERIFIED
# 200k .fastq, which is identical to re-streaming the ENA source with `head -n 100000`. The
# SHA-1 values in the `hashes:` blocks below are of the .gz artifacts as produced here and
# are the fetch-integrity check. Both are needed and they check different things. If these
# fixtures are published (Zenodo), replace each `path:` with the published `location:` and
# keep the SHA-1 unchanged -- it pins exactly these bytes.
#
# REFERENCE. Genome FASTA and GTF are pinned NCBI FTP URLs staged with `decompress: true`.
# They carry NO `hashes:` block: test-data-refs.json pins md5s of the COMPRESSED files as
# served, and whether a `hashes:` block on a `decompress: true` input checks the fetched or
# the decompressed bytes was not established. Omitting is the resolution the test plan
# names for that case (unresolved `hash-vs-decompress-ordering`).
#
# SETTLED HERE, FROM THE PINNED WRAPPER SOURCES (both are the versions this workflow pins).
#   * featurecounts 2.1.1+galaxy1 builds output_short as `grep -v "^#" output | cut -f 1,7`,
#     which strips only the "# Program:" comment and keeps the header row, renamed to the
#     element identifier. output_short is therefore 1 header + one row per annotated gene.
#     The NCBI GTF for GCA_002759435.2 carries exactly 5586 distinct gene_id values, all of
#     them matching B9J08_[0-9]{6} -- hence has_n_lines n: 5587 and the annotation guard.
#   * deseq2 2.11.40.8+galaxy4's deseq2.R writes counts_out with `col.names = NA`, so its
#     header is padded with a leading blank field: 4 samples -> 5 tab-separated fields on
#     line 1. has_n_columns n: 5 on both normalized-counts tables is therefore verified,
#     not assumed, and the test plan's `normalized-counts-header-width-unverified` is closed.
#   * Both wrappers name count columns by $file.element_identifier, so the normalized-counts
#     header really does carry AR0382_A / AR0382_tnSWI1_A and the sample-name assertions
#     below are well founded (test plan's `normalized-counts-column-naming-assumed`, closed).
#
# HEADER ASYMMETRY. deseq2.R writes `deseq_out` with col.names = FALSE (NO header, which is
# why the significance filters bind header_lines: '0') and `counts_out` with col.names = NA
# (header present, padded with a leading blank field). The assertions below respect that:
# `not_has_text: baseMean` on the raw result tables, sample-name `has_text` on the
# normalized-counts tables.
#
# NOT ASSERTED, DELIBERATELY. No rank claim (SCF1 is 5th/4th by padj and 2nd/1st by
# |log2FC| in the two contrasts -- no rank-1 assertion is true in both, and tests-format
# cannot express a rank anyway). No exact DESeq2 value, no exact per-gene count, not the
# paper's ~29-fold magnitude, and not the negative ALS / IFF-HYR adhesin claim. See the
# test plan's omissions[] for each refusal.
- doc: >-
    End-to-end run over all six deposited runs of PRJNA904261 (SRP409192), each subset to
    the first 200,000 read pairs, against the NCBI GCA_002759435.2 (Cand_auris_B8441_V2)
    reference and its own GTF. Exercises the whole DAG - flatten side-branch, FastQC,
    Cutadapt, RNA STAR with a history reference, featureCounts, the sample-sheet condition
    split, both DESeq2 reductions and both two-link significance chains - and asserts the
    paper's central result: SCF1 (B9J08_001458) is significantly down-regulated in BOTH
    contrasts against the AR0382 reference level. Membership in each significant-genes
    table means padj < 0.05 AND |log2FC| > 1 by construction; the extra floor asserted
    here is log2FC <= -2. No rank and no measured value is asserted anywhere: the
    grounding measurements come from pydeseq2 over bwa-mem counts, not from this
    workflow's R DESeq2 over featureCounts -s 2, so only margins survive the difference.
  job:
    RNA-seq reads (sample sheet):
      class: Collection
      collection_type: sample_sheet:paired
      # rows is a mapping of element identifier -> POSITIONAL LIST of column values,
      # ordered to match the input's column_definitions (condition, replicate).
      # Not a dict of column-name -> value: validate_row() zips row against
      # column_definitions and rejects on length mismatch.
      rows:
        AR0382_A: [AR0382, A]
        AR0382_B: [AR0382, B]
        AR0387_A: [AR0387, A]
        AR0387_B: [AR0387, B]
        AR0382_tnSWI1_A: [tnSWI1, A]
        AR0382_tnSWI1_B: [tnSWI1, B]
      elements:
      - class: Collection
        collection_type: paired
        identifier: AR0382_A
        elements:
        - class: File
          identifier: forward
          path: test-data/AR0382_A_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 81847033e7a1bd1430efc68499b08ec2a3de619e
        - class: File
          identifier: reverse
          path: test-data/AR0382_A_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: bc59eb5c253f42454631792d59405a019c56dcfe
      - class: Collection
        collection_type: paired
        identifier: AR0382_B
        elements:
        - class: File
          identifier: forward
          path: test-data/AR0382_B_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: a11438bf4c886b081b887ca650642bf925014177
        - class: File
          identifier: reverse
          path: test-data/AR0382_B_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: ab93048b0e16a0729fb0e38ed206cf398ed798df
      - class: Collection
        collection_type: paired
        identifier: AR0387_A
        elements:
        - class: File
          identifier: forward
          path: test-data/AR0387_A_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 0bcb624b2347148ab740c81ebcb19894449c30b4
        - class: File
          identifier: reverse
          path: test-data/AR0387_A_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 4590862aba8082c42557f31cbe62c12d694de4b6
      - class: Collection
        collection_type: paired
        identifier: AR0387_B
        elements:
        - class: File
          identifier: forward
          path: test-data/AR0387_B_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: f0e1a1e97ab993d91baf38ae4ffdd3edf5028905
        - class: File
          identifier: reverse
          path: test-data/AR0387_B_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 7ccf13c15a7a0ac4571bd2b06796243050d4a17c
      - class: Collection
        collection_type: paired
        identifier: AR0382_tnSWI1_A
        elements:
        - class: File
          identifier: forward
          path: test-data/AR0382_tnSWI1_A_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 8c76f5caecf41f7a9e62689f5cbdac2e877c31c7
        - class: File
          identifier: reverse
          path: test-data/AR0382_tnSWI1_A_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 8d8d57b28c41d50a6f5dd13074440deeb801d1da
      - class: Collection
        collection_type: paired
        identifier: AR0382_tnSWI1_B
        elements:
        - class: File
          identifier: forward
          path: test-data/AR0382_tnSWI1_B_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: d666ddad269a39d17b736079f15d5fb98d980851
        - class: File
          identifier: reverse
          path: test-data/AR0382_tnSWI1_B_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 7a7d6932aed4f46604cd030b04eab9355106e5f9
    Reference genome FASTA:
      class: File
      location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.fna.gz
      filetype: fasta
      decompress: true
    Gene annotation GTF:
      class: File
      location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.gtf.gz
      filetype: gtf
      decompress: true
    Reference condition level: AR0382
    First contrast condition level: tnSWI1
    Second contrast condition level: AR0387
    Strandedness: 'stranded - reverse'
    Adjusted p-value threshold: 0.05
    log2 fold change threshold: 1.0
  outputs:
    'FastQC raw reads: text summary':
      class: Collection
      element_count: 12
      element_tests:
        AR0382_A_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0382_A_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0382_B_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0382_B_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0387_A_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0387_A_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0387_B_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0387_B_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0382_tnSWI1_A_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0382_tnSWI1_A_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0382_tnSWI1_B_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
        AR0382_tnSWI1_B_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+200000\b'
          - that: has_text
            text: '>>Basic Statistics'
    Cutadapt trimming report:
      class: Collection
      element_count: 6
      element_tests:
        AR0382_A:
          asserts:
          - that: has_text_matching
            expression: 'Total read pairs processed:\s+200,000'
        AR0382_B:
          asserts:
          - that: has_text_matching
            expression: 'Total read pairs processed:\s+200,000'
        AR0387_A:
          asserts:
          - that: has_text_matching
            expression: 'Total read pairs processed:\s+200,000'
        AR0387_B:
          asserts:
          - that: has_text_matching
            expression: 'Total read pairs processed:\s+200,000'
        AR0382_tnSWI1_A:
          asserts:
          - that: has_text_matching
            expression: 'Total read pairs processed:\s+200,000'
        AR0382_tnSWI1_B:
          asserts:
          - that: has_text_matching
            expression: 'Total read pairs processed:\s+200,000'
    Trimmed reads:
      class: Collection
      element_count: 6
      element_tests:
        AR0382_A:
          class: Collection
          elements:
            forward:
              asserts:
              - that: has_size
                min: 1000
            reverse:
              asserts:
              - that: has_size
                min: 1000
        AR0382_B:
          class: Collection
          elements:
            forward:
              asserts:
              - that: has_size
                min: 1000
            reverse:
              asserts:
              - that: has_size
                min: 1000
        AR0387_A:
          class: Collection
          elements:
            forward:
              asserts:
              - that: has_size
                min: 1000
            reverse:
              asserts:
              - that: has_size
                min: 1000
        AR0387_B:
          class: Collection
          elements:
            forward:
              asserts:
              - that: has_size
                min: 1000
            reverse:
              asserts:
              - that: has_size
                min: 1000
        AR0382_tnSWI1_A:
          class: Collection
          elements:
            forward:
              asserts:
              - that: has_size
                min: 1000
            reverse:
              asserts:
              - that: has_size
                min: 1000
        AR0382_tnSWI1_B:
          class: Collection
          elements:
            forward:
              asserts:
              - that: has_size
                min: 1000
            reverse:
              asserts:
              - that: has_size
                min: 1000
    STAR mapping summary:
      class: Collection
      element_count: 6
      element_tests:
        AR0382_A:
          asserts:
          - that: has_text
            text: 'Uniquely mapped reads %'
          - that: has_line_matching
            expression: '\s*Uniquely mapped reads % \|\s+(?:8[0-9]|9[0-9]|100)\.[0-9]+%'
        AR0382_B:
          asserts:
          - that: has_text
            text: 'Uniquely mapped reads %'
          - that: has_line_matching
            expression: '\s*Uniquely mapped reads % \|\s+(?:8[0-9]|9[0-9]|100)\.[0-9]+%'
        AR0387_A:
          asserts:
          - that: has_text
            text: 'Uniquely mapped reads %'
          - that: has_line_matching
            expression: '\s*Uniquely mapped reads % \|\s+(?:8[0-9]|9[0-9]|100)\.[0-9]+%'
        AR0387_B:
          asserts:
          - that: has_text
            text: 'Uniquely mapped reads %'
          - that: has_line_matching
            expression: '\s*Uniquely mapped reads % \|\s+(?:8[0-9]|9[0-9]|100)\.[0-9]+%'
        AR0382_tnSWI1_A:
          asserts:
          - that: has_text
            text: 'Uniquely mapped reads %'
          - that: has_line_matching
            expression: '\s*Uniquely mapped reads % \|\s+(?:8[0-9]|9[0-9]|100)\.[0-9]+%'
        AR0382_tnSWI1_B:
          asserts:
          - that: has_text
            text: 'Uniquely mapped reads %'
          - that: has_line_matching
            expression: '\s*Uniquely mapped reads % \|\s+(?:8[0-9]|9[0-9]|100)\.[0-9]+%'
    Gene counts per sample:
      class: Collection
      element_count: 6
      element_tests:
        AR0382_A:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]+'
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]{3,}'
        AR0382_B:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]+'
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]{3,}'
        AR0387_A:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]+'
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]{1,2}'
        AR0387_B:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]+'
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]{1,2}'
        AR0382_tnSWI1_A:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]+'
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]{1,2}'
        AR0382_tnSWI1_B:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]+'
          - that: has_line_matching
            expression: 'B9J08_001458\t[0-9]{1,2}'
    featureCounts assignment summary:
      class: Collection
      element_count: 6
      element_tests:
        AR0382_A:
          asserts:
          - that: has_line_matching
            expression: 'Assigned\t1[5-9][0-9]{4}'
          - that: has_line_matching
            expression: 'Unassigned_NoFeatures\t[0-9]{1,5}'
        AR0382_B:
          asserts:
          - that: has_line_matching
            expression: 'Assigned\t1[5-9][0-9]{4}'
          - that: has_line_matching
            expression: 'Unassigned_NoFeatures\t[0-9]{1,5}'
        AR0387_A:
          asserts:
          - that: has_line_matching
            expression: 'Assigned\t1[5-9][0-9]{4}'
          - that: has_line_matching
            expression: 'Unassigned_NoFeatures\t[0-9]{1,5}'
        AR0387_B:
          asserts:
          - that: has_line_matching
            expression: 'Assigned\t1[5-9][0-9]{4}'
          - that: has_line_matching
            expression: 'Unassigned_NoFeatures\t[0-9]{1,5}'
        AR0382_tnSWI1_A:
          asserts:
          - that: has_line_matching
            expression: 'Assigned\t1[5-9][0-9]{4}'
          - that: has_line_matching
            expression: 'Unassigned_NoFeatures\t[0-9]{1,5}'
        AR0382_tnSWI1_B:
          asserts:
          - that: has_line_matching
            expression: 'Assigned\t1[5-9][0-9]{4}'
          - that: has_line_matching
            expression: 'Unassigned_NoFeatures\t[0-9]{1,5}'
    'DESeq2 results: tnSWI1 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 7
      - that: not_has_text
        text: baseMean
      - that: has_line_matching
        expression: 'B9J08_001458(\t[^\t]*){6}'
      - that: has_line_matching
        expression: 'B9J08_001458\t[^\t]+\t-[0-9.]+(\t[^\t]*){4}'
    'DESeq2 results: AR0387 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 7
      - that: not_has_text
        text: baseMean
      - that: has_line_matching
        expression: 'B9J08_001458(\t[^\t]*){6}'
      - that: has_line_matching
        expression: 'B9J08_001458\t[^\t]+\t-[0-9.]+(\t[^\t]*){4}'
    'DESeq2 normalized counts: tnSWI1 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 5
      - that: has_line_matching
        expression: 'B9J08_001458(\t[0-9.eE+-]+){4}'
      - that: has_text
        text: AR0382_A
      - that: has_text
        text: AR0382_B
      - that: has_text
        text: AR0382_tnSWI1_A
      - that: has_text
        text: AR0382_tnSWI1_B
      - that: not_has_text
        text: AR0387
    'DESeq2 normalized counts: AR0387 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 5
      - that: has_line_matching
        expression: 'B9J08_001458(\t[0-9.eE+-]+){4}'
      - that: has_text
        text: AR0382_A
      - that: has_text
        text: AR0382_B
      - that: has_text
        text: AR0387_A
      - that: has_text
        text: AR0387_B
      - that: not_has_text
        text: tnSWI1
    'Significant genes: tnSWI1 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 7
      - that: has_line_matching
        expression: 'B9J08_001458(\t[^\t]*){6}'
      - that: has_line_matching
        expression: 'B9J08_001458\t[^\t]+\t-(?:[2-9]|[1-9][0-9]+)(?:\.[0-9]+)?(\t[^\t]*){4}'
    'Significant genes: AR0387 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 7
      - that: has_line_matching
        expression: 'B9J08_001458(\t[^\t]*){6}'
      - that: has_line_matching
        expression: 'B9J08_001458\t[^\t]+\t-(?:[2-9]|[1-9][0-9]+)(?:\.[0-9]+)?(\t[^\t]*){4}'
    'DESeq2 diagnostic plots: tnSWI1 vs AR0382':
      asserts:
      - that: has_size
        min: 10000
    'DESeq2 diagnostic plots: AR0387 vs AR0382':
      asserts:
      - that: has_size
        min: 10000
- doc: >-
    Fast structural smoke test over the same six samples at 25,000 read pairs each
    (~7 MB compressed total). Makes NO biological claim: at that depth SCF1 falls to
    roughly 50 fragments in AR0382 and near zero elsewhere, and the significance tables
    may legitimately be empty. What it does exercise is everything that can silently
    break without erroring - the sample_sheet:paired input shape, the flatten
    side-branch, the six-element map-over identifier space, both condition-split joins,
    both DESeq2 reductions, and both cross-step invariants. Every assertion here is a
    subset of case 1's, with the depth constant changed and the biological claims removed.
  job:
    RNA-seq reads (sample sheet):
      class: Collection
      collection_type: sample_sheet:paired
      # rows is a mapping of element identifier -> POSITIONAL LIST of column values,
      # ordered to match the input's column_definitions (condition, replicate).
      # Not a dict of column-name -> value: validate_row() zips row against
      # column_definitions and rejects on length mismatch.
      rows:
        AR0382_A: [AR0382, A]
        AR0382_B: [AR0382, B]
        AR0387_A: [AR0387, A]
        AR0387_B: [AR0387, B]
        AR0382_tnSWI1_A: [tnSWI1, A]
        AR0382_tnSWI1_B: [tnSWI1, B]
      elements:
      - class: Collection
        collection_type: paired
        identifier: AR0382_A
        elements:
        - class: File
          identifier: forward
          path: test-data/smoke-25k/AR0382_A_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 7686fa073a0089b0376bbe1d3b819b4ea14e3a67
        - class: File
          identifier: reverse
          path: test-data/smoke-25k/AR0382_A_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: f1749c97d907abffd4f5ee36139c22b15376d2af
      - class: Collection
        collection_type: paired
        identifier: AR0382_B
        elements:
        - class: File
          identifier: forward
          path: test-data/smoke-25k/AR0382_B_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: a120941043541873bd674fc8e5cab4bf7d8dcc37
        - class: File
          identifier: reverse
          path: test-data/smoke-25k/AR0382_B_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 8f315372a23e305fc5dc5ee04d31d8a771ccb8f3
      - class: Collection
        collection_type: paired
        identifier: AR0387_A
        elements:
        - class: File
          identifier: forward
          path: test-data/smoke-25k/AR0387_A_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: d8e170332a2ae20242b8f79c66f287b79fa2c56a
        - class: File
          identifier: reverse
          path: test-data/smoke-25k/AR0387_A_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 89a32f456f670a00719362b8d72a1a90110c7e94
      - class: Collection
        collection_type: paired
        identifier: AR0387_B
        elements:
        - class: File
          identifier: forward
          path: test-data/smoke-25k/AR0387_B_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 022d0851c4b94b919c62d6c1af64e3713bb8e138
        - class: File
          identifier: reverse
          path: test-data/smoke-25k/AR0387_B_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 28d347eddac1455d3df000846cffb4dc0a472c4d
      - class: Collection
        collection_type: paired
        identifier: AR0382_tnSWI1_A
        elements:
        - class: File
          identifier: forward
          path: test-data/smoke-25k/AR0382_tnSWI1_A_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 3b6109b5392bba5cb75f312fd7c078b67346013c
        - class: File
          identifier: reverse
          path: test-data/smoke-25k/AR0382_tnSWI1_A_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: e5599b61d6ec223f9cd88e7eb7e7877c6d7cf29c
      - class: Collection
        collection_type: paired
        identifier: AR0382_tnSWI1_B
        elements:
        - class: File
          identifier: forward
          path: test-data/smoke-25k/AR0382_tnSWI1_B_1.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: f803a1ee765bd35a575ab66f80133ad09e50a4ea
        - class: File
          identifier: reverse
          path: test-data/smoke-25k/AR0382_tnSWI1_B_2.fastq.gz
          filetype: fastqsanger.gz
          hashes:
          - hash_function: SHA-1
            hash_value: 97a623f6a006e90740a44d6391685ede0157acf4
    Reference genome FASTA:
      class: File
      location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.fna.gz
      filetype: fasta
      decompress: true
    Gene annotation GTF:
      class: File
      location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.gtf.gz
      filetype: gtf
      decompress: true
    Reference condition level: AR0382
    First contrast condition level: tnSWI1
    Second contrast condition level: AR0387
    Strandedness: 'stranded - reverse'
    Adjusted p-value threshold: 0.05
    log2 fold change threshold: 1.0
  outputs:
    'FastQC raw reads: text summary':
      class: Collection
      element_count: 12
      element_tests:
        AR0382_A_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0382_A_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0382_B_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0382_B_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0387_A_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0387_A_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0387_B_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0387_B_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0382_tnSWI1_A_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0382_tnSWI1_A_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0382_tnSWI1_B_forward:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
        AR0382_tnSWI1_B_reverse:
          asserts:
          - that: has_text_matching
            expression: 'Total Sequences\s+25000\b'
    Gene counts per sample:
      class: Collection
      element_count: 6
      element_tests:
        AR0382_A:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
        AR0382_B:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
        AR0387_A:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
        AR0387_B:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
        AR0382_tnSWI1_A:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
        AR0382_tnSWI1_B:
          asserts:
          - that: has_n_columns
            n: 2
          - that: has_n_lines
            n: 5587
            delta: 1
          - that: has_line_matching
            expression: 'B9J08_[0-9]{6}\t[0-9]+'
            min: 5000
    'DESeq2 results: tnSWI1 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 7
      - that: not_has_text
        text: baseMean
    'DESeq2 results: AR0387 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 7
      - that: not_has_text
        text: baseMean
    'DESeq2 normalized counts: tnSWI1 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 5
      - that: has_text
        text: AR0382_A
      - that: has_text
        text: AR0382_B
      - that: has_text
        text: AR0382_tnSWI1_A
      - that: has_text
        text: AR0382_tnSWI1_B
      - that: not_has_text
        text: AR0387
    'DESeq2 normalized counts: AR0387 vs AR0382':
      asserts:
      - that: has_n_columns
        n: 5
      - that: has_text
        text: AR0382_A
      - that: has_text
        text: AR0382_B
      - that: has_text
        text: AR0387_A
      - that: has_text
        text: AR0387_B
      - that: not_has_text
        text: tnSWI1
galaxy-workflow-validation-resultpresentgalaxy-workflow-validation-result.json

Terminal gxwf validation handoff: the exact command run, a pass/fail/not-run status, the classified workflow-level diagnostics, and the residual runtime risks static validation cannot settle.

declared by
validate-galaxy-workflow (phase 10)
consumed at
nothing downstream reads it
schema
none declared
sha256
45c44c23f1b2a5e03455359c1619f449f0f709b8160adeb871ec1b13ba04523e
status
pass
rationale
gxwf 1.12.0 exits 0 with 27 of 27 steps tool-state validated, 0 failures, 0 structure errors and 0 encoding errors under default strictness. The 7 steps the 1.10.1 run could not check now decode and validate, so the `pass-with-unvalidated-steps` qualifier no longer applies. Two holes remain and are recorded below: connection validation still does not run at all, and a stray `in:` key on a step that carries `tool_state` is still not caught.
steps_total
27
tool_state_validated
27
tool_state_skipped
0
tool_state_failed
0
structure_errors
0
encoding_errors_default_mode
0
connection_check_ran
false
cold_cache_behaviour
Cache dir created empty for this run; gxwf populated it with 13 unique tool entries during validation. The 7 previously undecodable tools (__FLATTEN__, __FILTER_FROM_FILE__, cutadapt, deseq2) now decode.
iwc-comparison-notespresentiwc-comparison-notes.md

Structural diff against the nearest IWC exemplar(s); guidance for the downstream *-summary-to-galaxy-template Mold before per-step authoring. Carries an inline, bounded gxformat2 excerpt of the nearest exemplar's relevant subgraph under a labeled section, cross-referencing the iwc-exemplar-gxformat2 sibling file.

declared by
compare-against-iwc-exemplar (phase 4)
consumed at
5, 8
schema
none declared
sha256
3278e8ce8f1dec171c97b74cb67113076557739e8938e12a3c9a17b1d93ac938
  • IWC exemplar comparison — *C. auris* Scf1 RNA-seq differential expression
  • Corpus provenance
  • 1. Ranking
  • 1.1 Why rank 1 is High
  • 1.2 Why rank 2 is Medium, not High
  • 1.3 The headline structural finding
  • 2. The log2FC question — settled by the corpus
  • 2.1 Inline excerpt — the parameter-to-filter bridge
  • 3. Structural divergences that matter for template authoring
  • 3.1 The condition split has no corpus precedent at all
  • 3.2 DESeq2 node arity — two nodes, settled
  • 3.3 FastQC fan-out — flatten, contra the data-flow brief
  • 3.4 Cutadapt emits one paired collection — no re-pair node
  • 3.5 RNA STAR from a history FASTA has no corpus precedent
  • 3.6 Strandedness is a mapped parameter, not a raw one
  • 4. Where the corpus confirms the design
  • 5. Test-fixture guidance (for phase 8)
  • 6. Findings routed by authoring surface
  • 7. Open-requirements ledger changes
iwc-exemplar-gxformat2presentiwc-exemplar.gxwf.yml

Cleaned gxformat2 conversion (via [[convert]] --to format2 --compact) of the nearest IWC exemplar's relevant subgraph — the concrete idiom the downstream template draft pattern-matches against. Bounded to the relevant subgraph, not the whole workflow. Absent when no nearest exemplar is found.

declared by
compare-against-iwc-exemplar (phase 4)
consumed at
5, 8
schema
none declared
sha256
13c0e9a85cee236f7d4f206a58931f63d9c67078b2b9564d4f4e3ca39282b2d6
# Nearest IWC exemplar(s) — bounded subgraphs
#
# Corpus:      https://github.com/galaxyproject/iwc
# Corpus HEAD: fe41a79 (galaxyproject/iwc main, shallow clone taken 2026-09-16)
# Produced by: gxwf convert <workflow>.ga --to format2 --compact, then bounded to the
#              relevant subgraph and stripped of tool_state keys that carry no structural
#              signal. Load-bearing tool_state is kept verbatim.
#
# THIS FILE IS A READING AID, NOT A RUNNABLE WORKFLOW. Steps have been removed and
# tool_state elided; every elision is marked with a `# [elided]` comment. For the full
# exemplars, convert the corpus files named in each document header.
#
# The subject workflow (C. auris Scf1 RNA-seq DE: FastQC -> Cutadapt -> RNA STAR ->
# featureCounts -> DESeq2 -> significance filter) spans a journey that IWC publishes as
# TWO workflows joined at the count-table boundary. Both halves are given below.
# See iwc-comparison-notes.md for the structural diff and the ranking.

---
# ============================================================================
# DOCUMENT 1 — PRIMARY EXEMPLAR (High confidence) for the DE tail (nodes H, I)
#
# IWC workflow ID: transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting
# Release:         0.12
# Steps covered:   Differential Analysis, the two compose_text_param parameter
#                  bridges, Annotate DESeq2 table, Filter with p-adj threshold,
#                  Filter with log2 FC threshold
# Steps dropped:   volcano plot, both heatmaps, the normalized-counts join/cut
#                  chain, and the recurring-header generator (visualization tail;
#                  the subject brief exposes no counterpart)
#
# This is the document that settles `fold-change-threshold-linear-vs-deseq2-log2fc`
# and `deseq2-contrast-realization-unsettled`.
# ============================================================================
class: GalaxyWorkflow
label: RNA-Seq Differential Expression Analysis with Visualization
doc: >-
  Identifies differentially expressed genes between exactly two experimental conditions
  from count tables. [elided: full doc string]
inputs:
  # NOTE: the condition grouping lives in the INTERFACE, as two pre-grouped `list`
  # collections of per-sample count tables. There is no metadata-driven split anywhere
  # in this workflow. Contrast with the subject's data-flow brief section 4.
  - id: Counts from changed condition
    type: collection
    collection_type: list
    optional: false
    doc: Counts from experimental condition or changed condition.
  - id: Counts from reference condition
    type: collection
    collection_type: list
    optional: false
    doc: Counts from reference condition or base condition.
  - id: Count files have header
    type: boolean
    optional: false
    doc: >-
      featureCounts count files have a header line; RNA-STAR count files do not.
  - id: Gene Annotaton
    type: data
    optional: false
    doc: The same annotation GTF used for mapping and counting
  - id: Adjusted p-value threshold
    type: float
    optional: false
    default: 0.05
  # *** THE log2FC ANSWER ***
  # IWC states the effect-size threshold in log2 units at the interface. It does NOT
  # expose a linear fold change and convert internally. Default 1.0 == linear 2-fold,
  # which is exactly the subject paper's |FC| > 2 criterion.
  - id: log2 fold change threshold
    type: float
    optional: false
    default: 1
    doc: >-
      log2 fold change threshold to filter for highly regulated genes.
      A log2 FC of 3 equals to an absolute fold change of 8 (2^3).
outputs:
  - id: DESeq2 Plots
    outputSource: Differential Analysis/plots
  - id: DESeq2 Normalized Counts
    outputSource: Differential Analysis/counts_out
  - id: Annotated DESeq2 results table
    outputSource: Annotate DESeq2 table/out_file1
  - id: Significantly differentially expressed genes
    outputSource: Filter with log2 FC threshold/out_file1
steps:
  # --- Node H equivalent: ONE DESeq2 job == ONE contrast ---------------------
  # `how: datasets_per_level` with a `rep_factorLevel` repeat: each level port takes a
  # COLLECTION of per-sample count files, reduced into the tool's multiple=true input.
  # Exactly two levels here, and `deseq_out` is a single dataset (proved downstream: it
  # feeds deg_annotate / tp_cat / Filter1, all single-dataset tools, and the sibling
  # -tests.yml asserts has_text_matching on it as a dataset, not as a collection).
  - id: Differential Analysis
    label: Differential Analysis
    tool_id: toolshed.g2.bx.psu.edu/repos/iuc/deseq2/deseq2/2.11.40.8+galaxy4
    tool_version: 2.11.40.8+galaxy4
    tool_shed_repository:
      changeset_revision: 05f9e54d7e81
      name: deseq2
      owner: iuc
      tool_shed: toolshed.g2.bx.psu.edu
    in:
      - id: header
        source: Count files have header
      - id: output_options|alpha_ma
        source: Adjusted p-value threshold
      - id: select_data|rep_factorName_0|rep_factorLevel_0|countsFile
        source: Counts from changed condition
      - id: select_data|rep_factorName_0|rep_factorLevel_1|countsFile
        source: Counts from reference condition
    out:
      - id: deseq_out
        hide: true
    tool_state:
      # [elided] advanced_options, batch_factors, tximport
      output_options:
        output_selector:
          - pdf
          - normCounts
        alpha_ma:
          __class__: ConnectedValue
      select_data:
        how: datasets_per_level
        __current_case__: 1
        rep_factorName:
          - __index__: 0
            factorName: DEFactor
            rep_factorLevel:
              - __index__: 0
                factorLevel: MainFactor
                countsFile:
                  __class__: ConnectedValue
              - __index__: 1
                factorLevel: BaseFactor
                countsFile:
                  __class__: ConnectedValue

  # --- Annotation: gene positions / biotype / symbol onto the results table --
  # Appends columns 8-13; columns 1-7 (GeneID, BaseMean, log2FC, StdErr, Wald, pval,
  # padj) are unchanged, so the c3 / c7 filter indices below hold with or without it.
  - id: _unlabeled_step_11
    tool_id: toolshed.g2.bx.psu.edu/repos/iuc/deg_annotate/deg_annotate/1.1.0+galaxy1
    tool_version: 1.1.0+galaxy1
    in:
      - id: annotation
        source: Gene Annotaton
      - id: input_table
        source: Differential Analysis/deseq_out
    out:
      - id: output
        hide: true
    tool_state:
      advanced_parameters:
        gff_feature_type: exon
        gff_feature_attribute: gene_id
        gff_transcript_attribute: transcript_id
        gff_attributes: gene_biotype, gene_name
      mode: degseq
      # [elided] chromInfo, ConnectedValue stubs

  - id: Annotate DESeq2 table
    label: Annotate DESeq2 table
    tool_id: toolshed.g2.bx.psu.edu/repos/bgruening/text_processing/tp_cat/9.11+galaxy0
    tool_version: 9.11+galaxy0
    in:
      # [elided] `inputs` comes from a generated single-line header dataset
      - id: queries_0|inputs2
        source: _unlabeled_step_11
    out:
      - id: out_file1
        change_datatype: tabular
        rename: Annotated DESeq2 results

  # --- PARAMETER -> FILTER EXPRESSION BRIDGE --------------------------------
  # `Filter1` takes its predicate as a TEXT parameter, so a numeric workflow input
  # cannot reach it directly. The corpus idiom is one compose_text_param step per
  # filter: a literal prefix plus the connected float. Two steps per threshold.
  - id: _unlabeled_step_8
    tool_id: toolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_param/0.1.1
    tool_version: 0.1.1
    in:
      - id: components_1|param_type|component_value
        source: Adjusted p-value threshold
    out:
      - id: out1
        hide: true
    tool_state:
      components:
        - __index__: 0
          param_type:
            select_param_type: text
            __current_case__: 0
            component_value: c7<          # c7 == P-adj
        - __index__: 1
          param_type:
            select_param_type: float
            __current_case__: 2
            component_value:
              __class__: ConnectedValue

  - id: _unlabeled_step_10
    tool_id: toolshed.g2.bx.psu.edu/repos/iuc/compose_text_param/compose_text_param/0.1.1
    tool_version: 0.1.1
    in:
      - id: components_1|param_type|component_value
        source: log2 fold change threshold
    out:
      - id: out1
        hide: true
    tool_state:
      components:
        - __index__: 0
          param_type:
            select_param_type: text
            __current_case__: 0
            component_value: abs(c3)>     # c3 == log2(FC); NO conversion applied
        - __index__: 1
          param_type:
            select_param_type: float
            __current_case__: 2
            component_value:
              __class__: ConnectedValue

  # --- Node I equivalent: TWO CHAINED Filter1 STEPS, not one ----------------
  - id: Filter with p-adj threshold
    label: Filter with p-adj threshold
    tool_id: Filter1
    tool_version: 1.1.1
    in:
      - id: cond
        source: _unlabeled_step_8/out1
      - id: input
        source: Annotate DESeq2 table/out_file1
    out:
      - id: out_file1
        hide: true
        rename: Genes filtered with adj p-value threshold
    tool_state:
      header_lines: "1"

  - id: Filter with log2 FC threshold
    label: Filter with log2 FC threshold
    tool_id: Filter1
    tool_version: 1.1.1
    in:
      - id: cond
        source: _unlabeled_step_10/out1
      - id: input
        source: Filter with p-adj threshold/out_file1
    out:
      - id: out_file1
        rename: Genes filtered with adj p-value and log2(FC) thresholds
    tool_state:
      header_lines: "1"
tags:
  - transcriptomics
  - RNAseq
license: MIT
release: "0.12"

---
# ============================================================================
# DOCUMENT 2 — SECONDARY EXEMPLAR (Medium confidence) for the map-over head
#              (nodes A-D)
#
# IWC workflow ID: transcriptomics/rnaseq-pe/rnaseq-pe
# Steps covered:   __FLATTEN__, fastp (the trimmer slot), RNA STAR, the
#                  map_param_value strandedness bridge, featureCounts, and the
#                  `More QC` subworkflow reduced to its Falco (FastQC-equivalent) step
# Steps dropped:   Cufflinks, StringTie, all coverage/bigwig generation, MultiQC,
#                  Picard / RSeQC / idxstats QC, the reference-genome text bridge
#                  (the subject brief has no counterpart for any of these)
#
# This is the document that settles `fastqc-per-read-fanout-not-a-flat-list` and
# supplies the strandedness-parameter idiom.
# ============================================================================
class: GalaxyWorkflow
label: "RNA-Seq Analysis: Paired-End Read Processing and Quantification"
inputs:
  - id: Collection paired FASTQ files
    type: collection
    collection_type: list:paired
    optional: false
    doc: Should be a list of paired-end RNA-seq fastqs
  # Built-in STAR index selected by name. `restrictOnConnections: true` narrows the
  # option list to the genomes STAR actually has indexed on the server.
  # The subject's C. auris B8441 has no such entry — see iwc-comparison-notes.md.
  - id: Reference genome
    type: string
    optional: false
    restrictOnConnections: true
  - id: GTF file of annotation
    type: data
    optional: false
  - id: Strandedness
    type: string
    optional: false
    restrictions:
      - stranded - forward
      - stranded - reverse
      - unstranded
outputs:
  - id: Mapped Reads
    outputSource: "STAR: map and count and coverage splitted/mapped_reads"
  - id: Counts Table
    outputSource: _unlabeled_step_25/Counts Table   # [elided] relabel step
steps:
  # --- *** THE FastQC FAN-OUT ANSWER *** ------------------------------------
  # IWC does NOT promote a nested collection from per-fastq QC. It flattens the
  # list:paired collection FIRST, with an explicit join_identifier, and runs the
  # single-dataset QC tool over the resulting flat `list`. Element identifiers
  # become <sample>_forward / <sample>_reverse.
  - id: _unlabeled_step_11
    tool_id: __FLATTEN__
    tool_version: 1.0.0
    in:
      - id: input
        source: Collection paired FASTQ files
    out:
      - id: output
        hide: true
    tool_state:
      input:
        __class__: ConnectedValue
      join_identifier: _

  # --- Trimmer slot. rnaseq-pe uses fastp, the subject's paper names Cutadapt. ---
  # Both emit ONE paired-inner collection when driven by a list:paired input:
  # fastp -> output_paired_coll; Cutadapt (lparsons/cutadapt, library.type
  # `paired_collection`) -> out_pairs + report. See the Cutadapt excerpt below.
  - id: remove adapters + bad quality bases
    label: remove adapters + bad quality bases
    tool_id: toolshed.g2.bx.psu.edu/repos/iuc/fastp/fastp/1.3.6+galaxy0
    tool_version: 1.3.6+galaxy0
    in:
      - id: single_paired|paired_input
        source: Collection paired FASTQ files
      # [elided] adapter_sequence1 / adapter_sequence2 from optional string inputs
    out:
      - id: output_paired_coll
        hide: true
      - id: report_json
        hide: true
    tool_state:
      filter_options:
        quality_filtering_options:
          disable_quality_filtering: false
          qualified_quality_phred: "30"
      # [elided] duplicated_reads, read_mod_options, output_options

  # --- Aligner. ENCODE long-RNA parameter set, INDEXED reference. -----------
  - id: "STAR: map and count and coverage splitted"
    label: "STAR: map and count and coverage splitted"
    tool_id: toolshed.g2.bx.psu.edu/repos/iuc/rgrnastar/rna_star/2.7.11b+galaxy1
    tool_version: 2.7.11b+galaxy1
    tool_shed_repository:
      changeset_revision: 55c9ac3aa8f4
      name: rgrnastar
      owner: iuc
      tool_shed: toolshed.g2.bx.psu.edu
    in:
      - id: refGenomeSource|GTFconditional|genomeDir
        source: Reference genome
      - id: refGenomeSource|GTFconditional|sjdbGTFfile
        source: GTF file of annotation
      - id: singlePaired|input
        source: remove adapters + bad quality bases/output_paired_coll
    out:
      - id: output_log            # Log.final.out — the assertable text behind the BAM
        hide: true
      - id: reads_per_gene
        hide: true
        rename: Reads per gene from STAR
      - id: mapped_reads
        rename: Mapped Reads
      # [elided] signal_unique_str1/2, signal_uniquemultiple_str1/2, splice_junctions
    tool_state:
      refGenomeSource:
        geneSource: indexed         # <-- every STAR step in IWC uses `indexed`
        __current_case__: 0
        GTFconditional:
          GTFselect: without-gtf-with-gtf
          __current_case__: 1
          genomeDir:
            __class__: ConnectedValue
          sjdbGTFfile:
            __class__: ConnectedValue
          sjdbGTFfeatureExon: exon
          sjdbOverhang: "100"
      # [elided] full ENCODE algo.params block (seed / align / junction settings)

  # --- User-facing strandedness string -> per-tool parameter value ----------
  # One map_param_value step per consumer (featureCounts, Cufflinks, StringTie).
  # Keeps ONE user-facing vocabulary while each wrapper gets its own encoding.
  - id: Get featureCounts strandedness parameter
    label: Get featureCounts strandedness parameter
    tool_id: toolshed.g2.bx.psu.edu/repos/iuc/map_param_value/map_param_value/0.2.0
    tool_version: 0.2.0
    in:
      - id: input_param_type|input_param
        source: Strandedness
    out:
      - id: output_param_text
    tool_state:
      # [elided] the value_map repeat: 'unstranded'->0, 'stranded - forward'->1,
      # 'stranded - reverse'->2
      output_param_type: text
      unmapped:
        on_unmapped: fail
        __current_case__: 1

  - id: _unlabeled_step_21
    tool_id: toolshed.g2.bx.psu.edu/repos/iuc/featurecounts/featurecounts/2.1.1+galaxy1
    tool_version: 2.1.1+galaxy1
    tool_shed_repository:
      changeset_revision: 37d067694d40
      name: featurecounts
      owner: iuc
      tool_shed: toolshed.g2.bx.psu.edu
    in:
      - id: alignment
        source: "STAR: map and count and coverage splitted/mapped_reads"
      - id: anno|reference_gene_sets
        source: GTF file of annotation
      - id: strand_specificity
        source: Get featureCounts strandedness parameter/output_param_text
      - id: when
        source: Use featureCounts for generating count tables
    out:
      - id: output_short           # subject interface output 7
        hide: true
      - id: output_summary         # subject interface output 8
        hide: true
    tool_state:
      anno:
        anno_select: history       # GTF from the history, same dataset as STAR's
        __current_case__: 2
        reference_gene_sets:
          __class__: ConnectedValue
        gff_feature_type: exon
        gff_feature_attribute: gene_id
        summarization_level: false
      format: tabdel_short
      pe_parameters:
        paired_end_status: PE_fragments
        __current_case__: 2
        exclude_chimerics: true
      # [elided] extended_parameters, read_filtering_parameters
    when: $(inputs.when)

  # --- QC subworkflow, reduced to the per-fastq QC step --------------------
  - id: More QC
    label: More QC
    run:
      class: GalaxyWorkflow
      label: RNA-seq-QC
      inputs:
        - id: FASTQ collection
          type: collection
          collection_type: list      # <-- FLAT, because of __FLATTEN__ upstream
          optional: false
      outputs:
        - id: Falco text output
          outputSource: _unlabeled_step_3/text_file
      steps:
        # Falco is a drop-in FastQC reimplementation; same single-dataset input,
        # same html_file + text_file output pair the subject promotes as outputs 1-2.
        - id: _unlabeled_step_3
          tool_id: toolshed.g2.bx.psu.edu/repos/iuc/falco/falco/1.3.2+galaxy0
          tool_version: 1.3.2+galaxy0
          in:
            - id: input_file
              source: FASTQ collection
          out:
            - id: html_file
              hide: true
            - id: text_file
              hide: true
      # [elided] gtftobed12, samtools view/idxstats, Picard MarkDuplicates,
      # RSeQC read_distribution and geneBody_coverage, and their three other inputs
    in:
      - id: FASTQ collection
        source: _unlabeled_step_11        # the FLATTEN output
      - id: STAR BAM
        source: "STAR: map and count and coverage splitted/mapped_reads"
      - id: reference_annotation_gtf
        source: GTF file of annotation
      - id: when
        source: Generate additional QC reports
    when: $(inputs.when)

---
# ============================================================================
# DOCUMENT 3 — TOOL-LEVEL EVIDENCE ONLY (cross-domain, NOT a domain exemplar)
#
# IWC workflow ID: epigenetics/cutandrun/cutandrun
# Steps covered:   the Cutadapt step only
#
# Cited for one fact the subject needs and the transcriptomics exemplars cannot
# supply, because neither uses Cutadapt: the output shape of the IUC Cutadapt
# wrapper when mapped over a list:paired collection. This settles
# `trimmed-reads-paired-reassembly-conditional`. CUT&RUN is a different domain
# and this document must not be read as a structural exemplar for anything else.
# ============================================================================
class: GalaxyWorkflow
label: CUT&RUN / CUT&TAG analysis
inputs:
  - id: PE fastq input
    type: collection
    collection_type: list:paired
    optional: false
steps:
  - id: Cutadapt (remove adapter + bad quality bases)
    label: Cutadapt (remove adapter + bad quality bases)
    tool_id: toolshed.g2.bx.psu.edu/repos/lparsons/cutadapt/cutadapt/5.2+galaxy2
    tool_version: 5.2+galaxy2
    in:
      - id: library|input_1
        source: PE fastq input
      # [elided] library|r1|adapters_0|... and library|r2|adapters2_0|... adapter
      #          sequences, wired from two text parameter inputs
    out:
      # ONE paired collection out, shape `input` (i.e. same as the input collection),
      # plus one report per element. No re-pair node is needed downstream.
      - id: out_pairs
      - id: report
        rename: cutadapt report
    tool_state:
      library:
        type: paired_collection
        __current_case__: 2
        input_1:
          __class__: ConnectedValue
        pair_adapters: false
        # [elided] r1 / r2 adapter repeats
open-requirements-ledgerpresentopen-requirements.ledger.yml

Carried obligations ledger re-emitted by this step: entries it appended or closed updated, every other entry passed through with its provenance intact.

declared by
advance-galaxy-draft-step, compare-against-iwc-exemplar, freeform-summary-to-galaxy-data-flow, freeform-summary-to-galaxy-interface, freeform-summary-to-galaxy-template (phase 2, 3, 4, 5, 6)
consumed at
2, 3, 4, 5, 6
schema
none declared
sha256
bfef2433c89340bf8fa0c661dbae698bef8dd949a87687fe4d7462c002abaa36
# open-requirements-ledger — run auris-scf1 (paper-to-galaxy)
# Started empty by freeform-summary-to-galaxy-interface, the first Mold of this run to carry it.
entries:
  - id: featurecounts-annotation-source-unnamed
    status: resolved
    raised_by: freeform-summary-to-galaxy-interface
    unmet: "gene annotation for workflow input `Gene annotation GTF`"
    missing: >-
      The paper names the genome assembly (GCA_002759435.2, C. auris B8441) but never names a
      GTF/GFF for it. featureCounts requires one and RNA STAR uses one for splice junctions.
      NCBI RefSeq GFF and FungiDB B8441 GFF differ in gene ID space and attribute keys
      (`gene_id` vs `ID`), which changes the featureCounts `-g` attribute, every downstream gene
      identifier, and whether SCF1 appears as `B9J08_001458` at all. FungiDB and CGOB are cited in
      the paper only for synteny inspection, not as the counting annotation. The interface fixes
      the datatype as `gtf`; a GFF3 source would require a different datatype or a conversion step.
    resolved_by: paper-to-test-data
    supersedes: null
    note: >-
      Largest single gap for reproducing this analysis. Whoever picks a source must record which
      one, because count values and gene IDs are not comparable across the two.

      CARRIED FORWARD by advance-galaxy-draft-step at `Count reads per gene`, which is now
      concrete and STILL DOES NOT CLOSE THIS. Both consumers are wired to the declared
      `Gene annotation GTF` input (RNA STAR `sjdbGTFfile`, featureCounts
      `anno|reference_gene_sets` under `anno_select: history`, case 2), so the PORTS are
      settled. What this step adds to the entry is a second dependent binding:
      `gff_feature_attribute: gene_id` (featureCounts `-g`), taken from the wrapper default and
      from corpus transcriptomics/rnaseq-pe/rnaseq-pe at fe41a79. That is correct for a GTF and
      wrong for a FungiDB B8441 GFF3, which keys on `ID` — and the failure mode is a complete,
      plausible counts table of the WRONG identifiers, not an error. `gff_feature_type: exon`
      has the same shape of exposure. So this entry now gates three things, not one: the input's
      datatype, the `-g` attribute, and the gene ID space every downstream result is addressed
      in. Naming the source settles all three at once; nothing else will.


      CLOSED by paper-to-test-data. The annotation is NCBI's own GTF for the exact accession the
      paper names: `GCA_002759435.2_Cand_auris_B8441_V2_genomic.gtf.gz` under
      `https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/`
      (md5 6e5b9528d48c0a8fc2c8588e6eeea929). It was downloaded and inspected, not merely cited.
      It settles all three things this entry gated, in the direction the workflow already assumes:
      it is a true GTF (`#gtf-version 2.2`), so the input's `gtf` datatype needs no conversion
      step; its `gene_id` values ARE the paper's locus tags, so `gff_feature_attribute: gene_id`
      is correct as bound and SCF1 resolves as `B9J08_001458` with no identifier translation
      (PEKT02000003.1:864995-867292, + strand, single exon, 2298 bp); and it carries 6057 `exon`
      features across 5586 distinct genes, so `gff_feature_type: exon` is correct too. The
      FungiDB-GFF3 hazard this entry described is avoided by not using FungiDB — nothing about
      that hazard was wrong, it simply does not arise for this source.

  - id: rnaseq-strandedness-inferred-from-kit-name
    status: resolved
    raised_by: freeform-summary-to-galaxy-interface
    unmet: "value for workflow parameter `featureCounts strandedness`"
    missing: >-
      The paper states only the library kit (Illumina Stranded Total RNA Prep with Ribo-Zero Plus).
      Reverse-stranded (dUTP) is inferred from that kit name and is not stated anywhere in the
      supplement. The interface exposes the parameter with default `reverse` so the inference is
      visible and changeable rather than buried in a step default.
    resolved_by: paper-to-test-data
    supersedes: null
    note: >-
      Empirically checkable without new information: the promoted output
      `featureCounts assignment summary` shows a large Unassigned_NoFeatures fraction when the
      setting is wrong. A test-plan or run phase can close this from evidence.


      CLOSED by paper-to-test-data, empirically rather than by argument. The entry itself named
      the check; it was performed a step earlier than expected, on alignments rather than on a
      featureCounts summary. 200,000 read pairs from each of AR0382_A, AR0387_A and tnSWI1_A were
      aligned to GCA_002759435.2 with bwa-mem and each R1 that unambiguously overlapped a single
      annotated gene was compared against that gene's strand. 98.4% / 98.4% / 98.2% map ANTISENSE.
      That is the dUTP reverse-stranded signature and it is not a close call. `stranded - reverse`
      → featureCounts `-s 2` is correct; the value is now measured, not inferred from the kit
      name. The `featureCounts assignment summary` check this entry proposed remains valid as a
      regression guard and is carried into the test plan as an assertion.

  - id: sample-sheet-condition-to-deseq2-factor-wiring
    status: resolved
    raised_by: freeform-summary-to-galaxy-interface
    unmet: "path from per-sample `condition` metadata to DESeq2's factor-level inputs"
    missing: >-
      The interface carries condition and replicate as `column_definitions` on a
      `sample_sheet:paired` reads input. Galaxy does not propagate `column_definitions` or per-row
      `columns` through map-over, so by the time featureCounts has produced counts the condition
      metadata is gone and an explicit step must reattach it
      (`__SAMPLE_SHEET_TO_TABULAR__` plus a filter/split, or the rules DSL) before DESeq2 can
      receive one multi-data input per factor level.
    resolved_by: freeform-summary-to-galaxy-data-flow
    supersedes: null
    note: >-
      Named fallback if the wiring proves unbuildable: replace the sample-sheet input with a
      `list:paired` reads collection plus a `data` input `Sample metadata table`
      (tabular: sample_id, condition, replicate) and split on that table. Taking the fallback
      changes workflow input 1 and must be reflected back into the interface brief, not applied
      silently in the template.
      Closed by the data-flow brief, section 4. The condition reaches DESeq2 by an
      identifier-keyed split taken off the workflow input, where the column metadata still lives:
      `__SAMPLE_SHEET_TO_TABULAR__` projects the sample sheet to a tabular (element identifier,
      condition, replicate); a row filter plus column projection yields one identifier list per
      factor level; `__FILTER_FROM_FILE__` filters the featureCounts collection to each level's
      identifiers; each per-level counts sub-collection reduces into one DESeq2 multi-data factor
      port. The map-over region and every promoted per-sample output are untouched, and the join
      key is the element identifier, which Galaxy preserves across map-over. The route is
      independent of how the DESeq2 contrasts are realized
      (`deseq2-contrast-realization-unsettled`) and survives the named fallback intact: under the
      fallback the tabular node simply disappears and the user-supplied metadata table lands in
      its place, with the filter and collection-split nodes unchanged. Two narrower successors
      carry what remains: `sample-sheet-to-tabular-identifier-column-unverified` (is the element
      identifier emitted as a column) and `deseq2-factor-level-names-not-parameterized` (where the
      per-level literal comes from). Rejected alternatives and why are recorded in the brief's
      section 4.6 — notably filtering by element-identifier regex, which is unsafe here because
      `AR0382_tnSWI1_A` contains the reference level's own identifier as a substring.

  - id: sample-sheet-input-test-fixture-expressibility
    status: resolved
    raised_by: freeform-summary-to-galaxy-interface
    unmet: "test-fixture form for the `sample_sheet:paired` workflow input"
    missing: >-
      Whether a `sample_sheet`-family workflow input — element identifiers plus per-row typed
      `columns` and collection-level `column_definitions` — can be expressed in a Planemo/IWC
      `-tests.yml` job block. If it cannot, the workflow's primary input is untestable as designed
      and the fallback in `sample-sheet-condition-to-deseq2-factor-wiring` becomes mandatory.
    resolved_by: paper-to-test-data
    supersedes: null
    note: >-
      Not verified in this phase; no fixture syntax for sample-sheet inputs appears in the
      references packaged with this Mold. Naturally closed by the test-plan phase.


      CLOSED by paper-to-test-data: YES, it is expressible, and the `list:paired` fallback is not
      required. Galaxy's job-block loader has an explicit branch for it —
      `lib/galaxy/tool_util/cwl/util.py`, `replacement_collection()`:
      `if collection_type.startswith("sample_sheet"): kwds["rows"] = value.get("rows")`, carried
      to the collections API by `lib/galaxy/tool_util/client/staging.py`. The nested
      `sample_sheet:paired` shape specifically is covered by
      `test/unit/tool_util/test_cwl_util.py::test_galactic_job_json_sample_sheet_paired_collection`,
      and an end-to-end worked example ships as
      `lib/galaxy_test/workflow/collection_semantics_cat_sample_sheet.gxwf-tests.yml`.

      Two things a test author must get right. (1) `rows` is a mapping of element identifier to a
      POSITIONAL LIST of column values ordered to match `column_definitions` — authority is
      `validate_row()` in
      `lib/galaxy/model/dataset_collections/types/sample_sheet_util.py`, which rejects on
      `len(row) != len(column_definitions)` then `zip(row, column_definitions)`. The dict form
      appearing in Galaxy's own non-paired unit tests never reaches that validator and will not
      work. (2) Neither `galactic_job_json` nor `staging.py` passes `column_definitions` when
      creating the collection, although the API payload supports it, so a test-staged sample sheet
      has collection-level `column_definitions: None`. This is NOT fatal: per-element `columns`
      are still populated from `rows`, and the two `column_definitions_compatible()` call sites
      are both in `DataCollectionToolParameter` option-building (UI dropdown filtering), which a
      test bypasses by supplying the HDCA by id. The one real consequence is that
      `__SAMPLE_SHEET_TO_TABULAR__` emits its header line only
      `#if $include_headers and $input.collection.column_definitions` — harmless here because
      `Project sample sheet to tabular` sets `include_headers: false`, but a later phase that
      flips it to true will see the header silently vanish under test while it appears in the UI.
      Concrete job block is in `test-data-refs.json` under `planemo_test_job_block`.

  - id: deseq2-contrast-realization-unsettled
    status: resolved
    raised_by: freeform-summary-to-galaxy-interface
    unmet: "step realization behind the two contrast outputs"
    missing: >-
      The paper reports two contrasts (tnSWI1 vs AR0382, AR0387 vs AR0382) and states the
      significance thresholds, but writes no design formula. One factor, three levels, n = 2 is
      inferred from the six deposited runs. Whether the two contrasts come from one DESeq2 run over
      a three-level factor or two runs over two-level factors depends on the chosen wrapper's
      contrast handling, which is not known at interface time.
    resolved_by: compare-against-iwc-exemplar
    supersedes: null
    note: >-
      The interface fixes only the output surface — two result tables and two filtered tables,
      labelled per contrast. Either realization satisfies it.
      Closed by the IWC exemplar comparison, section 3.2, against
      transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting at corpus fe41a79. Two DESeq2
      nodes, each with a two-level factor, the reference-level counts collection feeding both:
      H1 (AR0382, tnSWI1) and H2 (AR0382, AR0387). Evidence: the exemplar's DESeq2 step uses
      `select_data.how: datasets_per_level` with a `rep_factorName_0.rep_factorLevel` repeat of
      exactly two levels, each port consuming a whole collection through the wrapper's
      multiple=true input; its `deseq_out` is a single dataset, not a collection, proved by the
      single-dataset tools it feeds (deg_annotate, tp_cat, two Filter1 steps) and by the sibling
      -tests.yml asserting has_text_matching on the derived output as a dataset; and the
      workflow README bounds itself to "exactly 2 conditions with at least 2 replicates per
      condition". One DESeq2 job therefore yields one results table, so the interface's two
      distinctly labelled result tables require two jobs. A three-level factor is expressible
      via the repeat but would still emit one `deseq_out`, which cannot satisfy a two-output
      interface — so the output surface settles the arity regardless of the wrapper's
      three-level contrast behaviour. Upstream wiring is unaffected, exactly as the data-flow
      brief's section 4.4 predicted: nodes E, F1-F3 and G1-G3 are identical either way. One
      consequence for the interface: under two DESeq2 nodes, `DESeq2 normalized counts`
      (output 9) and `DESeq2 diagnostic plots` (output 14) are each produced twice while the
      interface declares one of each. Promote from one designated run and say which, or relabel
      per contrast; do not leave it implicit in the template.

  - id: reference-genome-delivery-shape-unverified
    status: open
    raised_by: freeform-summary-to-galaxy-interface
    unmet: "verified delivery shape for the C. auris B8441 reference genome"
    missing: >-
      The interface settles the genome as a history `fasta` dataset on portability grounds (a
      remote-URL fixture resolves on any server; a CVMFS built-in index does not), with RNA STAR
      building its index at run time. Whether usegalaxy.org carries a built-in index for
      GCA_002759435.2 was not checked, and this run's phase roster contains no reference-data Mold
      that owns the question.
    resolved_by: null
    supersedes: null
    note: >-
      Provisionally settled, not verified. A built-in index would be cheaper at run time but less
      portable for tests; revisiting it changes workflow input 2.

  - id: cutadapt-adapter-and-length-filter-unstated
    status: open
    raised_by: freeform-summary-to-galaxy-interface
    unmet: "Cutadapt adapter sequence and minimum-length filter"
    missing: >-
      The paper gives only "Cutadapt with a Phred cutoff score of 20". No adapter sequence appears
      anywhere in the supplement, and no minimum-length filter is stated. The interface reads the
      step as quality trimming only (`-q 20`, baked in) and exposes no adapter input.
    resolved_by: null
    supersedes: null
    note: >-
      If a later phase adds adapter trimming it is adding method the paper does not describe, and
      that addition has to be labelled as such rather than presented as a faithful port.
      STILL OPEN after phase 6 iteration 8 concretized `Quality-trim reads`, and deliberately so.
      The step pins lparsons/cutadapt 5.2+galaxy2 with all six adapter repeats written explicitly
      empty (adapters / front_adapters / anywhere_adapters on R1, the adapters2 / front_adapters2 /
      anywhere_adapters2 trio on R2) and `other_trimming_options.quality_cutoff: '20'`, which is
      the one Cutadapt parameter the paper actually states. The unstated minimum-length filter is
      now recorded concretely: the pinned wrapper's default is `filter_options.minimum_length: 1`,
      which the XML flags as a deliberate wrapper-side departure from cutadapt's own default of 0
      ("intentionally set to 1 ... to avoid hard to debug issues with downstream tools"). The step
      writes that 1 out explicitly so a later wrapper bump cannot move it silently -- but it is the
      WRAPPER's default, not the paper's parameter, and the corpus shows the value is genuinely
      chosen per workflow (cutandrun at fe41a79 uses 15, the VGP workflows use 1). Closing this
      entry requires a stated adapter and a stated length cut, neither of which exists in the
      source.

  - id: galaxy-tool-versions-unpinnable-from-source
    status: open
    raised_by: freeform-summary-to-galaxy-interface
    unmet: "tool versions matching the published analysis"
    missing: >-
      The paper gives no version for any Galaxy step (FastQC, Cutadapt, RNA STAR, featureCounts,
      DESeq2). Only non-Galaxy software is versioned (R 4.0.3, DescTools 0.99.49, survminer 0.4.9,
      Fiji 1.52, CellProfiler 3.1.9). The constructed workflow will pin its own versions and can
      reproduce the authors' method but never their software stack.
    resolved_by: null
    supersedes: null
    note: >-
      Unclosable from the source. Expected to be surrendered at the terminal and stated on the
      workflow itself, so a reader does not mistake the run for a version-faithful reproduction.

  - id: tnbcy1-contrast-not-carried
    status: open
    kind: dropped
    raised_by: freeform-summary-to-galaxy-interface
    units: "the tnBCY1 (B9J08_002818) vs AR0382 transcriptome comparison reported in Fig. S2"
    because: >-
      No tnBCY1 runs exist in BioProject PRJNA904261. Only three conditions were deposited
      (AR0382, AR0387, AR0382 tnSWI1), across six runs SRR22376027–SRR22376032.
    unmet: "one of the three transcriptome comparisons the paper reports"
    missing: >-
      The workflow can build only the two contrasts whose input data is public. A test that tries
      to reproduce Fig. S2 has no input.
    resolved_by: null
    supersedes: null
    note: >-
      Cut by data availability, not by a design decision. Adding a fourth condition later needs
      no interface change: the reads input is a sample sheet whose `condition` restrictions widen.

  - id: pipeline-b-tdna-mapping-not-carried
    status: open
    kind: dropped
    raised_by: freeform-summary-to-galaxy-interface
    units: >-
      the whole of Pipeline B — AtMT T-DNA insertion-site mapping: FastQC, Trimmomatic,
      BWA-MEM against linearized pTO128 (seed 50, band width 2), extractSoftClipped,
      BWA-MEM of the soft-clipped flanks against C. auris B8441 (5 computational steps,
      2 deposited runs SRR22376033–SRR22376034)
    because: >-
      Two blockers, both recorded in the source summary's open questions (4 and 5) and neither
      independently verified in this phase. (a) `extractSoftClipped` from SE-MEI
      (github.com/dpryan79/SE-MEI) has no known Galaxy Tool Shed wrapper, so the pipeline cannot be
      assembled from stock tools as written; substituting a samtools/awk soft-clip extraction would
      change the method. (b) The pTO128 (pPZP-NAT) plasmid reference has no public accession in the
      paper — it is cited to the prior AtMT method paper (ref. 52) — so step 3's reference is
      unresolvable from the publication alone.
    unmet: "the second of the paper's two sequencing analyses"
    missing: >-
      No T-DNA integration sites are produced by this run. The scope decision was taken by the
      harness before this phase.
    resolved_by: null
    supersedes: null
    note: >-
      Uncited in the strict sense: no Tool Shed search for extractSoftClipped and no Addgene or
      ref. 52 lookup for pTO128 was run in this phase; both reasons are carried from the source
      summary. Treat as a debt, not a finding. If a later phase's discovery step contradicts either
      reason, this entry's `because` is refuted and the entry must be reopened rather than left
      standing.

  - id: sample-sheet-to-tabular-identifier-column-unverified
    status: resolved
    raised_by: freeform-summary-to-galaxy-data-flow
    step: sample_metadata_table
    unmet: "confirmation that the sample-sheet-to-tabular bridge emits the element identifier"
    missing: >-
      The settled condition wiring joins the sample metadata table to the featureCounts collection
      on element identifier, so node E's output must carry that identifier as a column. The
      packaged note galaxy-sample-sheet-collections documents `__SAMPLE_SHEET_TO_TABULAR__` only as
      iterating elements and tab-joining "for downstream tabular consumers"; it does not state the
      output column set or ordering, and no other packaged reference covers it. Galaxy source was
      not readable from inside this run.
    resolved_by: advance-galaxy-draft-step
    supersedes: null
    note: >-
      RESOLVED, affirmatively, by the tool summary itself. `__SAMPLE_SHEET_TO_TABULAR__` v1.0.0
      caches cleanly from the Tool Shed API by bare id (unlike `__FLATTEN__`, whose collection
      output defeats gxwf's summary decoder), and its packaged help states the contract
      directly: "The first column is always the element identifier (sample name). The remaining
      columns match the metadata fields defined in the sample sheet." With the optional
      `include_headers` enabled the first header cell is literally `element_identifier`. So the
      identifier IS emitted, as column 1, and the ordering this region assumed throughout --
      (element identifier, condition, replicate), metadata in column_definitions order -- is
      confirmed. All six provisional column bindings downstream (three Filter1 predicates on c2,
      three Cut1 projections of c1) stand unchanged; the Apply Rules substitute node named in
      the data-flow brief section 4.2 is not needed and was not built. The step is pinned with
      `include_headers: false`, which also settles the three row filters' `header_lines` at 0 --
      that binding is no longer blocked on this entry. Evidence is the wrapper's own documented
      contract, not a corpus exemplar; `no-iwc-precedent-for-sample-sheet-workflow-input` is
      unaffected and stays open.

  - id: deseq2-factor-level-names-not-parameterized
    status: resolved
    raised_by: freeform-summary-to-galaxy-data-flow
    step: select_level_L
    unmet: "a source for the two non-reference condition level literals"
    missing: >-
      The condition split needs one literal condition value per factor level to filter the sample
      metadata table (AR0382, AR0387, tnSWI1). The interface exposes only `Reference condition
      level` (input 4, default AR0382). The two contrast levels have no parameter, so as the
      interface stands they would be baked into two filter steps — re-hard-coding into the steps
      the three-level design that the sample-sheet input was chosen to keep out of the public
      interface.
    resolved_by: freeform-summary-to-galaxy-template
    supersedes: null
    note: >-
      Two ways to settle it, both changing the interface's parameter surface rather than the
      topology: expose two more text parameters (one per contrast level) alongside the existing
      reference-level parameter, or accept the bake and state plainly in the interface that the
      three level names are fixed in the steps. Whichever is chosen must be reflected back into
      freeform-galaxy-interface.md section 2.3, not applied silently in the template.
      Closed by the template on the first of those two options. The draft exposes two further
      `text` workflow inputs, `First contrast condition level` (default tnSWI1) and
      `Second contrast condition level` (default AR0387), alongside the existing
      `Reference condition level` (default AR0382). Each feeds a `compose_text_param` step that
      builds the row predicate for its level, so no level literal is baked into any step.
      The deciding argument is symmetry: the reference level was already a parameter for exactly
      this reason, and baking the other two would have left the interface half-generalized while
      re-hard-coding into the steps the design the sample-sheet input was chosen to keep out of
      the interface. The cost is two inputs the phase-2 brief does not declare, which is one facet
      of the interface drift recorded in
      `interface-brief-output-and-parameter-surface-drifted-from-draft`; that entry carries the
      owed edit to freeform-galaxy-interface.md section 2.3, since this Mold does not own that
      artifact. One caveat the draft states on both new inputs: the contrast output labels
      (`... tnSWI1 vs AR0382`, `... AR0387 vs AR0382`) are the public API and are not derived
      from these parameters, so changing a level value makes the labels stale.

  - id: fold-change-threshold-linear-vs-deseq2-log2fc
    status: resolved
    raised_by: freeform-summary-to-galaxy-data-flow
    step: filter_significant
    unmet: "unit agreement between workflow input 7 and the DESeq2 result column it thresholds"
    missing: >-
      Workflow input 7 is `Minimum absolute fold change`, default 2.0, in linear units — the form
      the paper states (|fold change| > 2). DESeq2 reports log2FoldChange. The significance filter
      must therefore compare abs(log2FoldChange) > log2(threshold), which needs either a conversion
      the filter expression may not support or a restatement of the parameter. Nothing in the
      interface brief notes the mismatch.
    resolved_by: compare-against-iwc-exemplar
    supersedes: null
    note: >-
      Consequential if missed rather than merely untidy: comparing abs(log2FoldChange) > 2.0
      applies a 4-fold cut and silently fails to reproduce the paper's gene lists, while still
      producing a plausible-looking filtered table. Options are a log2 conversion node before the
      filter, a filter expression that computes the conversion inline, or restating input 7 as a
      log2 threshold (default 1.0) with its label changed — the last changes the interface.
      Closed by the IWC exemplar comparison, section 2, on the third of those three options.
      transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting at corpus fe41a79 exposes
      `log2 fold change threshold` as a float workflow input, default 1.0, documented as
      "A log2 FC of 3 equals to an absolute fold change of 8 (2^3)", and filters with the raw
      column and no conversion anywhere in the workflow: `abs(c3)>` concatenated with the
      connected float. There is no linear fold-change parameter, no log2() call and no
      conversion node in the corpus exemplar. So: restate interface input 7 as a log2 threshold
      with default 1.0 — which is exactly the paper's |fold change| > 2 — and change its label;
      teach the conversion in the doc string as the corpus does. Three implementation details
      come with the answer. (a) Column indices are c3 for log2FC and c7 for adjusted p-value;
      deg_annotate appends columns 8-13 and leaves 1-7 untouched, so the indices hold whether or
      not an annotation step is added. (b) Filter1's predicate is a text parameter, so a numeric
      workflow input cannot reach it directly: the corpus idiom is one iuc/compose_text_param
      step per threshold concatenating a literal prefix (`c7<`, `abs(c3)>`) with the connected
      float, making node I two Galaxy steps per filter rather than one. (c) The two thresholds
      are two chained Filter1 steps, p-adj then log2FC, each with header_lines "1", not one
      compound predicate. Restating input 7 changes the interface brief's section 2.3 and its
      output-3 table entry; that edit must be made there, not silently in the template.

  - id: fastqc-per-read-fanout-not-a-flat-list
    status: resolved
    raised_by: freeform-summary-to-galaxy-data-flow
    step: qc_raw_reads
    unmet: "agreement between the declared shape of outputs 1-2 and what map-over actually produces"
    missing: >-
      The interface declares `FastQC raw reads: text summary` and `: HTML report` as `list`
      collections. FastQC consumes a single dataset, so mapping it over a `sample_sheet:paired`
      input fans out over the inner paired axis as well — 12 jobs, and outputs nested one report
      per read direction per sample, not six flat elements. As declared, outputs 1 and 2 are not
      what the workflow computes.
    resolved_by: compare-against-iwc-exemplar
    supersedes: null
    note: >-
      The data-flow brief (section 7.1) recommends promoting the nested collection and correcting
      the interface: per-read-direction QC is what a reader wants from raw-read QC, and it
      preserves the element identifier space that every checkpoint assertion keys on. The
      alternative is an explicit flatten node, which satisfies the declared `list` but rewrites
      identifiers to a doubled vocabulary (`AR0382_A_forward`) for no analytical gain. Either way
      the test plan must know which, because it changes every assertion on outputs 1 and 2.
      Closed by the IWC exemplar comparison, section 3.3, in favour of the explicit flatten —
      the option the data-flow brief rated as having no analytical gain. The brief's preference
      was a reasonable call made with no corpus evidence available; the evidence now exists and
      points the other way. transcriptomics/rnaseq-pe/rnaseq-pe at corpus fe41a79 puts an
      explicit `__FLATTEN__` step with `join_identifier: _` between the list:paired reads input
      and the per-fastq QC tool, and the consuming QC subworkflow declares its input as
      `collection_type: list` — flat. This is published IWC convention rather than a workaround,
      and it satisfies the interface's declared `list` shape for outputs 1 and 2 as originally
      written, so no interface correction is needed. Consequences the test plan must key on: the
      identifier vocabulary on outputs 1 and 2 doubles to twelve (AR0382_A_forward,
      AR0382_A_reverse, and so on for each of the six samples), while outputs 3-8 keep the
      six-element sample identifier space untouched, because the flatten is a side branch off the
      workflow input and does not enter the map-over region. The data-flow brief's placeholder
      transformation 5.5 is therefore required, not conditional.

  - id: trimmed-reads-paired-reassembly-conditional
    status: resolved
    raised_by: freeform-summary-to-galaxy-data-flow
    step: trim_reads
    unmet: "the output collection shape of the paired-aware trimming step"
    missing: >-
      The interface declares `Trimmed reads` as `list:paired`. Whether the trimmer emits one
      paired-inner collection per sample or two parallel single-ended collections (R1, R2) is
      wrapper-dependent and was not resolvable in this phase, which pins no Tool Shed tools. If it
      is the latter, the design needs a re-pair node between trimming and alignment, or the
      alignment step must take two parallel collections in dot-product.
    resolved_by: compare-against-iwc-exemplar
    supersedes: null
    note: >-
      Conditional shape repair, not method: recorded in the data-flow brief as placeholder
      transformation 5.6 so the template does not assume one reading. Naturally closed by the IWC
      exemplar comparison or by tool discovery on the trimmer.
      Closed by the IWC exemplar comparison, section 3.4: one paired-inner collection per sample,
      no re-pair node. Neither transcriptomics exemplar uses Cutadapt — both use fastp — so the
      evidence comes from epigenetics/cutandrun/cutandrun at corpus fe41a79, cited for the
      wrapper's IO shape only and for nothing else, since CUT&RUN is a different domain.
      toolshed.g2.bx.psu.edu/repos/lparsons/cutadapt/cutadapt/5.2+galaxy2 driven from a
      list:paired collection with `library.type: paired_collection` declares outputs `out_pairs`
      (type `input`, i.e. the input collection's own shape) and `report` — one paired collection
      plus one report per element. Placeholder transformation 5.6 is therefore not needed and
      node C takes a single collection input, as the data-flow brief's primary reading assumed.
      The finding is consistent across both wrappers that could fill the trimmer slot:
      transcriptomics/rnaseq-pe/rnaseq-pe's fastp emits `output_paired_coll` the same way. The
      residual is version-scoped rather than structural — a future Cutadapt wrapper could change
      its output set, so the per-step loop should confirm the output name against whatever
      version it pins.

  - id: mapped-outputs-carry-sample-sheet-outer-axis
    status: open
    raised_by: freeform-summary-to-galaxy-data-flow
    unmet: "accurate declared collection types for the promoted per-sample outputs"
    missing: >-
      A tool mapped over a `sample_sheet`-family collection produces a `sample_sheet`-shaped output
      without `column_definitions` (packaged note galaxy-sample-sheet-collections, "Mapping
      rules"). The interface's output table calls outputs 1-8 `list` / `list:paired`. Behaviourally
      that is accurate — such a collection maps, reduces and filters exactly like a list — but the
      declared collection type string is not `list`, which matters to anything that type-checks,
      including workflow test assertions on collection type.
    resolved_by: null
    supersedes: null
    note: >-
      No topology consequence; the fix is either a corrected type column in the interface brief or
      an explicit statement that the promoted outputs are sample_sheet-shaped lists without column
      metadata. Flagged primarily for the test-plan phase, which is where a wrong collection_type
      assertion would surface as a confusing failure.


      ADVANCED, NOT CLOSED, by implement-galaxy-workflow-test. The test file authors no `attributes:
      {collection_type: ...}` assertion on any of the eight mapped outputs, as the test plan's `promoted-
      collection-type-string-unverified` instructs. What it asserts instead is `element_count` plus a named
      `element_tests` entry per identifier - twelve for the flattened FastQC outputs, six for every output on
      the sample axis - which is shape-agnostic and pins the thing that actually matters, the identifier space
      the condition split joins on. So no assertion in this run depends on the declared type string and nothing
      here forces the question. It stays open as an interface-brief accuracy item, and the real type strings
      should be read off the first successful invocation before anyone decides whether a collection_type
      assertion is worth having.


  - id: iwc-splits-rnaseq-de-at-the-count-table-boundary
    status: open
    raised_by: compare-against-iwc-exemplar
    unmet: "agreement between this run's workflow scope and published IWC practice"
    missing: >-
      IWC has no single workflow spanning FastQC to DESeq2. At corpus fe41a79 the journey is
      published as two workflows joined at the count-table boundary:
      transcriptomics/rnaseq-pe/rnaseq-pe ends at per-sample count tables, and
      transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting begins there, taking pre-grouped
      `list` collections of count tables as its workflow inputs. This run's design is one
      workflow spanning both halves, which is why it needs an in-workflow condition split at all
      — a workflow that starts from count tables can demand pre-grouped collections at its
      interface and never reconstruct grouping internally.
    resolved_by: null
    supersedes: null
    note: >-
      Not a defect in either design, and not a reason to re-scope: the paper describes one
      analysis and the run was asked to build it. Recorded because it is the root of the single
      largest divergence from corpus practice (see
      `no-iwc-precedent-for-sample-sheet-workflow-input`) and because any later
      mature-galaxy-workflow-for-iwc pass will meet it as a packaging question — a reviewer may
      ask why this is not two workflows. Whoever revisits it should weigh that the split also
      costs reusability of the DE half, which is presumably why IWC made it.

  - id: no-iwc-precedent-for-sample-sheet-workflow-input
    status: open
    raised_by: compare-against-iwc-exemplar
    step: sample_metadata_table
    unmet: "a worked corpus precedent for the sample-sheet condition split"
    missing: >-
      Searched across all of `workflows/` at corpus fe41a79: `sample_sheet` as a collection type
      appears in zero workflows; `__SAMPLE_SHEET_TO_TABULAR__` appears in zero workflows;
      `column_definitions` appears only as the literal `null` that newer Galaxy serializes onto
      ordinary collection inputs, never with a value. `__FILTER_FROM_FILE__` does appear, in six
      workflows, but none in transcriptomics and none for a condition split. So the second half
      of the data-flow brief's section 4.2 route is a real used Galaxy idiom while the
      sample-sheet half is unprecedented in published IWC practice, and no corpus fixture
      declares a sample-sheet collection either.
    resolved_by: null
    supersedes: null
    note: >-
      Absence of precedent is not refutation, and nothing in the corpus contradicts the route —
      the phase-3 reasoning stands on its own. What changes is the risk profile: this is the one
      region of the workflow with no worked example to pattern-match against, and two open
      entries ride on it (`sample-sheet-to-tabular-identifier-column-unverified`,
      `sample-sheet-input-test-fixture-expressibility`). Guidance for the template, from the
      comparison notes section 3.1: build the split as a clearly delimited region with the
      phase-2 fallback (a `data` input `Sample metadata table`) reachable by deleting one node,
      so that if either open entry resolves against the sample sheet the cost is a deletion
      rather than a redesign. Worth knowing for the record that IWC made the opposite trade
      deliberately — it pushes grouping into the interface as two pre-grouped count collections,
      which is the same shape the interface brief's section 2.1 rejected for hard-coding the
      design into the public API.

  - id: star-history-reference-wiring-has-no-corpus-precedent
    status: resolved
    raised_by: compare-against-iwc-exemplar
    step: align_reads
    unmet: "worked wiring for the RNA STAR history-reference conditional branch"
    missing: >-
      Every RNA STAR step in the corpus at fe41a79 — there are exactly two, in
      transcriptomics/rnaseq-pe/rnaseq-pe and transcriptomics/rnaseq-sr/rnaseq-sr — uses
      `refGenomeSource.geneSource: indexed`, selecting a built-in `genomeDir` through a
      `restrictOnConnections: true` string parameter with `sjdbGTFfile` supplied from the
      history. Their test jobs pass a plain genome string (`Reference genome: sacCer3`). The
      interface settles input 2 as a history FASTA, which is the right call for C. auris B8441
      since no public server indexes it, but that is a different `__current_case__` in the
      iuc/rgrnastar wrapper with a different set of required sub-parameters, and the corpus has
      no example of it.
    resolved_by: advance-galaxy-draft-step
    supersedes: null
    note: >-
      RESOLVED, affirmatively, from the wrapper rather than from the corpus — which is what the
      original note asked for. iuc/rgrnastar/rna_star @ 2.7.11b+galaxy1 caches and summarizes
      cleanly (all ten of its outputs are plain `data`, so the collection-output decode failure
      that blocks `__FLATTEN__` and lparsons/cutadapt does not apply here), and its schema
      answers every part of the question. `refGenomeSource.geneSource` publishes exactly two
      options, `indexed` and `history`; the HISTORY branch is real, is `__current_case__: 1`, and
      carries `genomeFastaFiles` (gx_data, formats fasta/fasta.gz, optional false),
      `genomeSAindexNbases` (integer, min 2 max 16, default 14), its own TWO-case
      `GTFconditional`, and `diploidconditional` (Yes=0 / No=1, default No). So the two TODO
      ports resolve to `refGenomeSource|genomeFastaFiles` and
      `refGenomeSource|GTFconditional|sjdbGTFfile`; the exemplar's
      `refGenomeSource|GTFconditional|genomeDir` exists only under `indexed` and is correctly
      absent. Under history, GTFconditional's option order (without-gtf, with-gtf) is the
      REVERSE of its `<when>` order (with-gtf, without-gtf), so `with-gtf` is case 0 — a case
      index that cannot be read off the dropdown. That the case index follows `<when>` document
      order is not assumed: `gxwf convert --to format2` over
      transcriptomics/rnaseq-pe/rnaseq-pe.ga at fe41a79 emits `geneSource: indexed` with
      `__current_case__: 0` and `GTFselect: without-gtf-with-gtf` with `__current_case__: 1`,
      matching the indexed branch's `<when>` order (with-gtf, without-gtf-with-gtf, without-gtf)
      and not its option order. Cross-read against rg_rnaStar.xml and macros.xml at tools-iuc
      main, which carries @TOOL_VERSION@ 2.7.11b / @VERSION_SUFFIX@ 1 — this exact version.
      Two consequences worth carrying: the history branch builds its index in-job via
      `STAR --runMode genomeGenerate` into tempstargenomedir, confirming the data-flow brief's
      no-separate-index-node call; and `output_log` and `mapped_reads` carry no `<filter>`
      whatsoever in `<outputs>`, so the promoted `STAR mapping summary` cannot be suppressed by
      any parameter choice here. The step now validates for real against the cache
      (`12 ok, 0 fail, 2 skip`), not by skip. `reference-genome-delivery-shape-unverified` is
      untouched and STAYS OPEN — this entry establishes that the history route WORKS, never
      that it is preferable to a built-in index nobody checked for.
      `featurecounts-annotation-source-unnamed` also stays open: the GTF PORT is now wired, the
      GTF SOURCE is still unnamed by the paper.

  - id: star-genome-length-drives-sa-index-parameter
    status: open
    raised_by: advance-galaxy-draft-step
    step: "Splice-aware alignment"
    unmet: "a measured length for the reference actually wired into `Reference genome FASTA`"
    missing: >-
      `genomeSAindexNbases` exists ONLY in the RNA STAR history branch — it configures the
      in-job `--runMode genomeGenerate` that this workflow's history-FASTA delivery choice
      introduces, so it is not a mapping parameter the paper's "default parameters" could have
      covered. The wrapper's own help carries STAR's formula verbatim: "For small genomes, the
      parameter --genomeSAindexNbases must be scaled down to min(14, log2(GenomeLength)/2 - 1)".
      The step binds `'10'`, from min(14, log2(12.5e6)/2 - 1) = min(14, 10.79). The 12.5 Mb is
      the GCA_002759435.2 assembly record's length as carried through this run's briefs; nothing
      in this run measured the dataset that will actually be wired.
    resolved_by: null
    supersedes: null
    note: >-
      Silent-degradation class, not a hard gate: the wrapper default of 14 is the documented
      seg-fault-at-mapping hazard for a genome this small (the wrapper's own test data uses 5),
      and a value too small merely costs search speed. Two ways this becomes wrong: a different
      reference is supplied at run time (a mammalian genome would want 14 back), or the B8441
      FASTA in hand differs materially in length from the assembly record. Settle it by reading
      the length of the actual dataset, or by reading STAR's own recommendation out of the
      promoted `STAR mapping summary` / job stderr on the first real run — STAR prints the
      recommended value when the supplied one is too large. Tied to
      `reference-genome-delivery-shape-unverified`: if that resolves to a built-in index, this
      parameter disappears with the branch.


      ADVANCED, NOT CLOSED, by paper-to-test-data. The entry asked for a measured length of the
      reference actually wired in. The reference this run will wire is now pinned — NCBI
      `GCA_002759435.2_Cand_auris_B8441_V2_genomic.fna.gz`, md5 a008b270d3aaa04736a8bb7daf3f6dd5 —
      and it was downloaded and measured: 12,365,959 bp across 15 contigs. min(14,
      log2(12365959)/2 - 1) = min(14, 10.78), so the bound `'10'` is correct for this dataset.
      What keeps the entry open is exactly the exposure it already named: the parameter is
      correct for the reference the TEST supplies, and a user who supplies a different genome at
      run time still gets a silently wrong value. The first real STAR run should confirm from the
      promoted `STAR mapping summary`.

  - id: deseq2-normalized-counts-and-plots-promoted-per-contrast
    status: resolved
    raised_by: freeform-summary-to-galaxy-template
    step: "Differential expression: first contrast vs reference"
    unmet: >-
      a decision on the two DESeq2 outputs the interface declares once and two DESeq2 jobs
      produce twice
    missing: >-
      `deseq2-contrast-realization-unsettled` closed on two DESeq2 nodes, which makes
      `DESeq2 normalized counts` (interface output 9) and `DESeq2 diagnostic plots` (interface
      output 14) each produced twice while the interface declares one of each. The IWC comparison
      flagged the consequence and instructed the template not to leave it implicit: promote from
      one designated run and say which, or relabel per contrast.
    resolved_by: freeform-summary-to-galaxy-template
    supersedes: null
    note: >-
      Settled by relabelling per contrast. The draft promotes four outputs where the interface
      declared two: `DESeq2 normalized counts: tnSWI1 vs AR0382`,
      `DESeq2 normalized counts: AR0387 vs AR0382`, `DESeq2 diagnostic plots: tnSWI1 vs AR0382`
      and `DESeq2 diagnostic plots: AR0387 vs AR0382`, taking the workflow to sixteen outputs.
      This is not a tie broken on taste. Each DESeq2 job sees only its own four samples, so its
      size factors, normalized values and PCA/dispersion plots are computed over that subset: the
      two normalized-count tables are different tables, not two copies of one. Promoting either
      and calling it "the" normalized counts would be wrong rather than merely arbitrary, and
      dropping one would discard a real artifact while leaving the surviving one silently
      contrast-scoped. Relabelling also makes the four outputs symmetric with the result and
      filtered-gene tables, which are already labelled per contrast.
      One consequence for the test plan: the paper's sanity check that SCF1 sits in the top 2.5%
      of AR0382 expression can be asserted against either normalized-counts table, because both
      carry the two AR0382 replicates. Assert it on one and say which.
      The owed edit to freeform-galaxy-interface.md section 3 rides on
      `interface-brief-output-and-parameter-surface-drifted-from-draft`; labels are the public API
      and this is a breaking change to be made once, before any test is written.

  - id: interface-brief-output-and-parameter-surface-drifted-from-draft
    status: open
    raised_by: freeform-summary-to-galaxy-template
    unmet: >-
      agreement between freeform-galaxy-interface.md and the settled draft's public input and
      output surface
    missing: >-
      Phases 4 and 5 changed the interface in four places that the phase-2 brief still states in
      its original form. The brief is the artifact the test plan reads for labels, and labels are
      the API, so the drift has to be visible rather than inferred by diffing two artifacts.
      (a) Interface input 7 `Minimum absolute fold change` (float, linear, default 2.0) is in the
      draft `log2 fold change threshold` (float, log2 units, default 1.0) — closed on corpus
      evidence by `fold-change-threshold-linear-vs-deseq2-log2fc`.
      (b) Interface input 5 `featureCounts strandedness` (free text, allowed values named only in
      prose) is in the draft `Strandedness` (text restricted to `stranded - forward` /
      `stranded - reverse` / `unstranded`, default `stranded - reverse`) feeding a
      `map_param_value` bridge with `on_unmapped: fail` — the corpus idiom adopted per the IWC
      comparison section 3.6.
      (c) Two inputs the brief does not declare at all: `First contrast condition level` and
      `Second contrast condition level` — see `deseq2-factor-level-names-not-parameterized`.
      (d) Fourteen declared outputs become sixteen — see
      `deseq2-normalized-counts-and-plots-promoted-per-contrast`.
    resolved_by: null
    supersedes: null
    note: >-
      Bookkeeping, not a design gap: every one of the four changes is itself settled and recorded,
      and the draft is the current statement of the interface. What is unmet is that no Mold in
      this run's roster owns freeform-galaxy-interface.md after phase 2, so the edits have no
      writer. Recorded here so the test-plan phase reads labels off the draft rather than off the
      stale brief, and so a later maturation pass knows which artifact was authoritative. Closable
      by editing the brief's sections 2.3 and 3, or by declaring the draft authoritative for the
      interface and saying so in the brief.

  - id: deseq2-result-table-header-presence-unverified
    status: resolved
    raised_by: freeform-summary-to-galaxy-template
    step: "Filter with p-adj threshold: first contrast"
    unmet: "the `header_lines` binding for the four Filter1 steps of the significance chains"
    missing: >-
      The corpus exemplar binds `header_lines: "1"` on both of its Filter1 steps, but it filters a
      table that has been through `deg_annotate` and then `tp_cat`, and the tp_cat step exists
      precisely to concatenate a separately generated single-line header onto the DESeq2 output.
      That strongly implies the raw `deseq_out` is headerless. This workflow omits the annotation
      pair — the paper names no annotation step and the interface's tool set is closed — so its
      Filter1 steps consume `deseq_out` directly and the corpus binding cannot be copied across.
      Whether the pinned iuc/deseq2 wrapper writes a header row on `deseq_out` is not established
      by anything this run read.
    resolved_by: advance-galaxy-draft-step
    supersedes: null
    note: >-
      Consequential if missed rather than untidy, which is why it is a ledger entry and not only a
      `_plan_state` line. Binding "1" against a headerless table silently drops the first gene
      row of every filtered result; binding "0" against a headed table passes the header line into
      a numeric comparison, where Filter1 discards it with a warning rather than an error. Either
      way the workflow produces a plausible table. The same binding applies to the three row
      filters in the condition-split region, where the unknown is the header behaviour of
      `__SAMPLE_SHEET_TO_TABULAR__` rather than of DESeq2 — a different producer, the same class
      of silent loss, tracked there by
      `sample-sheet-to-tabular-identifier-column-unverified`. Closable by a summarize-galaxy-tool
      pass on iuc/deseq2 during the per-step loop, or by inspecting the first real run's output.
      SETTLED in the per-step loop, at the step that pins the producer
      (`Differential expression: first contrast vs reference`, iuc/deseq2/deseq2 @
      2.11.40.8+galaxy4, changeset 05f9e54d7e81): `deseq_out` carries NO header row, so all four
      significance filters bind `header_lines: '0'` — the same value the three condition-split row
      filters already use, for an unrelated reason. Three independent lines of evidence, all at
      the pinned version. (1) The wrapper's own `deseq2.R` writes the result table with
      `write.table(out_df, file = opt$outfile, sep = "\t", quote = FALSE, row.names = FALSE,
      col.names = FALSE)` — `col.names = FALSE`, at both of the two call sites that write a result
      table (the single-contrast path and the `many_contrasts` loop). The same script writes
      `counts_out` with `col.names = NA`, so the normalized-counts table DOES carry a header:
      one tool, opposite answers for its two tabular outputs, which is exactly why this could not
      be settled by analogy. (2) `deseq2.xml`'s test assertions for `deseq_out` match a data row
      first (`FBgn0003360\t1933.9504…\t-2.8399…`) with `has_n_lines n="3999"` and assert no
      header, while `vst_out` and `counts_out` assertions in the same test block DO assert a
      sample-name header line. (3) The corpus chain explains its own "1": rnaseq-de MANUFACTURES
      the header it later skips — `tp_text_file_with_recurring_lines` emits the single line
      `GeneID__tc__Base mean__tc__log2(FC)…`, `tp_sed_tool` turns `__tc__` into tabs, and `tp_cat`
      (labelled `Annotate DESeq2 table`) concatenates it on top of the deg_annotate output; only
      then do its two Filter1 steps bind `header_lines: "1"`. That "1" is about the manufactured
      header, not about `deseq_out`, which confirms rather than contradicts the reading above.
      One rider the four filters inherit: the c3 = log2FC / c7 = padj column contract holds only
      while `lfc_shrinkage_type` is `none` — `get_result_output_columns()` drops the `stat` column
      under any shrinkage, moving padj to c6. The DESeq2 steps pin `none` explicitly for that
      reason.
  - id: cross-step-invariants-survive-only-in-the-draft
    status: open
    raised_by: advance-galaxy-draft-step
    step: "Differential expression: first contrast vs reference / second contrast vs reference; the four significance filters"
    unmet: "a durable record, in the runnable artifact, of the two cross-step couplings that keep the filter predicates pointing at the right columns"
    missing: >-
      Two couplings hold this workflow together and neither is expressible in gxformat2 nor
      checked by any gate. (1) `advanced_options.lfc_shrinkage_type: none` on BOTH DESeq2 nodes
      is what makes `deseq_out` a 7-column table, so `c3` = log2FoldChange and `c7` = padj — the
      columns the `compose_text_param` bridges hard-code as `"abs(c3)>"` and `"c7<"` in four
      OTHER steps. Under any other shrinkage the `stat` column is dropped, padj moves to c6, and
      `"c7<"` silently filters on nothing. (2) `header_lines: '0'` on all four significance
      filters, against a corpus exemplar that binds `'1'` — the exemplar manufactures the header
      it then skips, this workflow does not. `'1'` would discard row 1 of a padj-sorted table,
      i.e. *SCF1*, the paper's finding. Both were established by hand and recorded only in YAML
      comments above the steps in `galaxy-workflow-draft.gxwf.yml`. `gxwf draft-extract`
      re-serializes from a parsed model and destroys every comment, so
      `galaxy-workflow.gxwf.yml` — the artifact the test, validate, run and any later maturation
      phase actually consume — carries no trace of either. The values themselves are correct in
      the extract (verified at iteration 26); what is missing is any reason a later editor would
      know not to change them.
    resolved_by: null
    supersedes: null
    note: >-
      Not closable inside the per-step loop: the loop's only writable surface is the draft, and
      the draft's comments are exactly what the extract drops. Closable by (a) folding both
      couplings into the step `doc:` strings of the two DESeq2 nodes and the four filters — `doc:`
      survives extraction and is visible in the Galaxy editor — or (b) making the test plan assert
      the column contract directly (a test on `deseq_out` that the padj column is c7, or on a
      filtered table that its top row survived), which converts a silent invariant into a failing
      test. (a) is the cheap one and should be done before the draft is discarded; (b) is the one
      that would actually hold. Filed upstream as `draft-extract-destroys-yaml-comments` in
      `foundry-feedback.ledger.yml`; that is the tool-side gap, this entry is the run-side
      obligation it leaves behind.


      ADVANCED, NOT CLOSED, by freeform-summary-to-galaxy-test-plan. Route (b) was taken and it
      lands unevenly across the two invariants. INVARIANT 1 is now fully detectable: both test cases
      assert `has_n_columns: 7` on BOTH `DESeq2 results:` tables, and case 1 asserts it again on both
      `Significant genes:` tables, so any shrinkage setting other than `none` on either node drops
      the `stat` column and fails the assertion before the `c7<` predicate can silently filter on
      nothing. A `not_has_text: baseMean` assertion on both raw tables covers the headerless premise
      the same predicates rest on. INVARIANT 2 is only partly detectable, and the plan says so in
      warnings[] `reference-level-filter-regression-undetectable`: of the seven Filter1 steps binding
      `header_lines: '0'`, a regression is visible on two — `Select first-contrast samples` and
      `Select second-contrast samples`, where the leaked metadata row would add a fifth sample column
      to that contrast's normalized-counts table, which both cases assert against. On `Select
      reference-level samples` the leaked row IS `AR0382_A`, which the `c2=='AR0382'` predicate keeps
      anyway, so the output is byte-identical and no assertion can see it. On the four significance
      filters the leaked row is row 1 of a padj-sorted table, which passes both filters on its own
      merits, so the regression is invisible at workflow level. Route (a) — folding both couplings
      into the `doc:` strings of the two DESeq2 nodes and the seven filters, which survives
      `draft-extract` — was NOT done and remains the cheap half of this entry. It is now the only
      protection the three undetectable reference-level and significance filters have. Do it before
      `galaxy-workflow-draft.gxwf.yml` is discarded.


      FURTHER ADVANCED by implement-galaxy-workflow-test, on the invariant-2 half only. The normalized-counts
      header width is no longer an assumption: deseq2.R at the pinned 2.11.40.8+galaxy4 writes `counts_out` with
      `write.table(..., col.names = NA)`, which emits a header padded with a leading blank field, so a four-
      sample contrast has exactly five tab-separated fields on line 1. Both test cases now assert
      `has_n_columns: 5` on both normalized-counts tables, and case 1 keeps the SCF1 data-row field-count
      assertion alongside it. A `header_lines: '1'` regression on either contrast-selecting Filter1 leaks a
      fifth sample column and fails both. The three undetectable filters - `Select reference-level samples` and
      the four significance filters - are unchanged, and route (a), folding both couplings into the `doc:`
      strings that survive `draft-extract`, is still not done and is still the only protection they have.


      RE-VERIFIED AT THE TERMINAL GATE by validate-galaxy-workflow, still not closed. Both couplings
      hold in `galaxy-workflow.gxwf.yml` as assembled: `advanced_options.lfc_shrinkage_type: 'none'`
      on both DESeq2 nodes (neither selects `many_contrasts`/`split_output`; both are
      `select_data.how: datasets_per_level`), and `header_lines: '0'` as the string `'0'` on all
      seven Filter1 steps (8, 12, 16, 23, 24, 25, 26 — the Filter1 count is exactly seven). Two
      refinements to the entry's own statement of invariant 1, from reading the assembled artifact.
      The `"c7<"` predicate is hard-coded in ONE `compose_text_param` bridge, step 21, consumed by
      two Filter1 steps (23 and 25); the second bridge, step 22, is `"abs(c3)>"`, consumed by 24 and
      26. Since `stat` sits at c5, dropping it moves padj from c7 to c6 but leaves log2FoldChange at
      c3 — so the shrinkage coupling protects the padj predicate specifically, and the `abs(c3)>`
      bridge is not exposed to it. Route (a) remains not done and the terminal artifact confirms
      where the hole is: the four significance filters' `doc:` strings do name their `c7<` /
      `abs(c3)>` predicates and the three condition-split filters' docs do say "headerless", but
      NEITHER DESeq2 node's `doc:` mentions shrinkage at all, and no filter doc says why
      `header_lines` is `'0'`. An editor changing `lfc_shrinkage_type` has nothing in the runnable
      artifact to warn them. `galaxy-workflow-draft.gxwf.yml` still holds the rationale in the YAML
      comments `draft-extract` strips; do not discard it before the `doc:` strings are written.

  - id: scf1-fold-change-magnitude-disagrees-with-paper
    status: open
    raised_by: paper-to-test-data
    unmet: "reconciliation of the paper's ~29-fold SCF1 figure with the deposited read data"
    missing: >-
      The paper states a ~29-fold SCF1 expression difference between AR0382 and AR0387. Direct
      measurement of the deposited runs disagrees on magnitude by roughly an order of magnitude:
      aligned fragments over the SCF1 exon at 200,000 read pairs per sample give 414 and 391 in
      the two AR0382 replicates against 2 and 1 in AR0387 (~270-fold), and exact 31-mer matching
      against the SCF1 CDS over 5.2 M reads gives 1942.7 vs 12.3 per million (~158-fold).
      Direction and significance are not in doubt; only the ratio is. Candidates not
      distinguished here: the 29-fold may be the RT-qPCR assay rather than the RNA-seq, a shrunk
      rather than raw log2FC, or a different normalization. Reference bias is ruled out as the
      explanation — AR0387 IS the B8441 reference strain, so its SCF1 reads cannot be undercounted;
      if anything AR0382's divergent allele is undercounted, which makes the true ratio larger,
      not smaller.
    resolved_by: null
    supersedes: null
    note: >-
      Consequence for the test plan, and the reason this is filed rather than glossed: no
      assertion may pin the paper's 29-fold figure. `test-data-refs.json` asserts direction and
      significance only (`log2FoldChange < -2`, `padj < 0.05`), which both the paper and the data
      support. Settling this needs the paper's source data (Data S1) or the authors, not more
      compute.


      REINFORCED by freeform-summary-to-galaxy-test-plan, and one candidate explanation is now the
      leading one. The disagreement was previously measured only on raw counts and k-mer matches;
      phase 8 ran an actual differential-expression model — pydeseq2, two separate two-level n=2
      analyses with NO LFC shrinkage, mirroring the workflow's two DESeq2 nodes, on a strand-aware
      count matrix built from phase 7's six BAMs. AR0387 vs AR0382 gives SCF1 log2FC -7.95 at padj
      4.1e-18, i.e. ~247-fold, consistent with the ~270-fold raw fragment ratio and not with the
      paper's ~29-fold. tnSWI1 vs AR0382 gives log2FC -6.93 at padj 2.1e-31. Because that run applied
      no shrinkage, it does not rule out the "shrunk rather than raw log2FC" candidate — log2(29) =
      4.86 against an unshrunk 7.95 is exactly the direction and roughly the magnitude a shrinkage
      estimator would move a gene whose counts are near zero in one group, so that candidate is now
      the most plausible of the three and the RT-qPCR candidate is not excluded either. This matters
      for the workflow, not just the paper: the workflow pins `lfc_shrinkage_type: none`, so ITS
      log2FoldChange for SCF1 will be the unshrunk ~-8 and will never reproduce the paper's figure by
      construction. Do not treat that as a defect when the first real invocation lands. Caveat on all
      of the above: pydeseq2 is not the R DESeq2 the Galaxy wrapper runs, and the counts came from
      bwa-mem rather than RNA STAR plus featureCounts -s 2; see
      `deseq2-statistics-unverified-against-the-galaxy-wrapper`.

  - id: test-fixtures-not-hosted
    status: open
    raised_by: paper-to-test-data
    unmet: "a resolvable URL for each of the 12 read fixtures"
    missing: >-
      The 200,000-read-pair subsets were generated and hashed but not published. They exist only
      in the producing session's scratchpad and do not survive it. At ~61 MB compressed for 12
      files they are too large to commit alongside the workflow, so the IWC idiom applies: host
      them (Zenodo or equivalent) and reference them by URL. `test-data-refs.json` carries
      `<FIXTURE_BASE>` as the placeholder in `planemo_test_job_block`.
    resolved_by: null
    supersedes: null
    note: >-
      Not a research gap — the recipe is deterministic and the artifact pins it. Regenerate by
      streaming each ENA FASTQ and taking `head -n 800000`, then verify against the
      `md5_uncompressed` value recorded per file (hash the UNCOMPRESSED `.fastq`; gzip output is
      not byte-stable across gzip versions). The reference genome and annotation need no hosting:
      both are pinned to stable NCBI FTP URLs with md5s and are staged with `decompress: true`.


      ADVANCED, NOT CLOSED, by implement-galaxy-workflow-test. All twelve 200,000-read-pair fixtures were
      regenerated in this phase with the recorded recipe (`curl -L <ENA url> | gzip -dc | head -n 800000`) and
      every one verified byte-for-byte against its `md5_uncompressed` in test-data-refs.json before compression,
      so the recipe and the hashes are now proven reproducible rather than merely recorded. They live at
      `<run>/test-data/*.fastq.gz` (59 MB) and `galaxy-workflow.gxwf-tests.yml` addresses them as local `path:`
      entries, each with the SHA-1 of the `.gz` artifact as produced here. A twelve-file 25,000-read-pair set
      for the smoke case was derived from the verified 200k files by taking their first 100,000 lines, at
      `<run>/test-data/smoke-25k/` (6.7 MB). The test therefore runs against real data in this run directory and
      nowhere else: the fixtures are still unhosted and do not survive the run, which is the whole of what
      remains unmet. Publishing them is now a mechanical step - upload the exact `.gz` bytes, replace each
      `path:` with the published `location:`, and leave the SHA-1 values unchanged, because they pin those
      bytes. Keep the uncompressed md5s in test-data-refs.json as the regeneration check; the SHA-1s are the
      fetch-integrity check and the two answer different questions.


  - id: deseq2-statistics-unverified-against-the-galaxy-wrapper
    status: open
    raised_by: freeform-summary-to-galaxy-test-plan
    step: "Differential expression: first contrast vs reference / second contrast vs reference"
    unmet: "confirmation that the workflow's own DESeq2 node reproduces the differential-expression result the test plan asserts"
    missing: >-
      Every differential-expression figure this run has is from a different implementation than the
      workflow runs. Phase 8 measured SCF1's log2 fold change and adjusted p-value with pydeseq2 over
      a strand-aware count matrix built from phase 7's six bwa-mem BAMs, as two separate two-level
      n=2 analyses with no LFC shrinkage — tnSWI1 vs AR0382 log2FC -6.93 / padj 2.1e-31, AR0387 vs
      AR0382 log2FC -7.95 / padj 4.1e-18. The workflow runs the Galaxy DESeq2 wrapper, i.e. R DESeq2,
      over RNA STAR plus featureCounts -s 2. pydeseq2 is a faithful reimplementation, but it is not
      the same implementation and the counts are not the same counts, so the exact values will
      differ. What is unverified is not the biology — the separation is ~200-fold at the counts level
      and the padj margins are 18 and 31 orders of magnitude — but that the workflow's own node
      produces a row for B9J08_001458 clearing `padj < 0.05` and `|log2FC| > 1` into each significant
      table with `log2FoldChange <= -2`, which is what the test plan asserts, and that the raw table
      really is the 7-column shape both cases assert `has_n_columns: 7` on.
    resolved_by: null
    supersedes: null
    note: >-
      Closable by the first successful invocation and nothing cheaper; phase 11 (run-workflow-test)
      is where it lands, and phase 9 should not try to settle it by adding assertions. The plan is
      already written so that this can only bite as a margin question rather than a value question:
      no assertion anywhere pins an exact statistic, and the recorded measurements are stated as
      headroom over thresholds (~5 and ~6 log2 units over the -2 floor; 31 and 18 orders of magnitude
      over the 0.05 cut). Three things to check against the first invocation while it is open — the
      sign and magnitude of SCF1's c3 in both contrasts, that c7 is padj and not something else, and
      whether R DESeq2's independent filtering keeps SCF1 in the padj-carrying set at this depth as
      pydeseq2 did (2393 and 1376 genes retained). Recorded in the plan as warnings[]
      `deseq2-statistics-measured-with-pydeseq2-not-r`.

  - id: collection-algebra-never-statically-checked
    status: open
    raised_by: validate-galaxy-workflow
    step: whole workflow; acutely steps 0-3, 5, 10/14/18, 19-20
    unmet: any static confirmation that the workflow's collection shapes and map-over structure are compatible end to end
    missing: >-
      Terminal validation could not run the one check that covers this. `gxwf validate --connections`
      — the flag whose entire purpose is connection-type compatibility, collection algebra and
      map-over — aborts with `TypeError: step.in is not iterable` before emitting any report, on a
      format2 workflow written the way this one is: list-form `inputs:`/`steps:` with mapping-form
      `in:`, which gxwf's `toNative` mistakes for an already-normalized workflow and therefore never
      normalizes. Reproduced on a 20-line minimal case and unchanged on gxwf 1.12.0, so it is a gxwf
      defect and not a property of this workflow (filed as
      `tonative-shape-sniff-skips-normalization-on-list-form-format2`; fix submitted as
      jmchilton/galaxy-tool-util-ts#179). Two routes could still close this without an upstream
      release: rewriting the workflow's `in:` blocks in list form, or running `--connections` against
      a patched gxwf. What the default validation does cover is per-step tool state on the 20
      decodable steps, and nothing about how shapes compose.
      The map-over structure is the load-bearing part of this design and none of it is proven: the
      `sample_sheet:paired` input fanning through `__FLATTEN__` into FastQC's per-fastq axis; the
      paired collection entering Cutadapt at `library|input_1` and leaving as `out_pairs` into RNA
      STAR's `singlePaired|input`; three `__FILTER_FROM_FILE__` reductions over `Count reads per
      gene/output_short`; and the workflow's only true reduction, where each DESeq2 node's
      multiple=true `countsFile` ports consume whole 2-element collections and the per-sample axis
      collapses.
    resolved_by: null
    supersedes: null
    note: >-
      Not a defect claim against the workflow — every `in:` source resolves, every `in:` key name and
      tool_state parameter name was checked by hand against the real tool input trees for all 27
      steps (the 20 cached tools from the gxwf cache, the 7 undecodable ones fetched from
      `GET /api/tools/<id>?io_details=true`) with zero mismatches, and every select and boolean value
      on the previously unchecked steps is legal. What is unmet is that shape COMPOSITION has no
      static witness, and cannot get one until gxwf is fixed. Closable only by the first successful
      invocation. Phase 11 should read `GET /api/invocations/{id}/step_jobs_summary` per step rather
      than only the terminal invocation state: a shape mismatch surfaces as invocation message
      `collection_failed`, or as an unexpected job count on a mapped step, and a workflow can report
      `completed` while a mapped step ran the wrong number of jobs. Two adjacent exposures ride along
      and are settled by the same run: whether the target Galaxy supports the `sample_sheet:paired`
      input with `column_definitions` and `__SAMPLE_SHEET_TO_TABULAR__` at all (see
      `no-iwc-precedent-for-sample-sheet-workflow-input` and the filed
      `galaxy-test-staging-drops-sample-sheet-column-definitions`), which would fail at request-time
      validation or input materialization rather than in any tool's stderr; and the element
      identifiers `__FLATTEN__` produces, since step 0 carries no tool_state and `join_identifier`
      therefore takes the wrapper default, while the sample-sheet element identifiers are declared
      the spine of this workflow.
test-data-refspresenttest-data-refs.json

Resolved workflow test inputs and expected outputs derived from paper evidence (URLs, file shapes, expected hashes).

declared by
find-test-data, paper-to-test-data (phase 7)
consumed at
9
schema
none declared
sha256
9ed5c59287a4a2856798e06da958590defe73fbb14f6eef9e715ed7806e662d8
schema
test-data-refs
schema_version
1
produced_by
paper-to-test-data
run_slug
auris-scf1
branch_path_selected
paper-to-test-data
target_workflow
galaxy-workflow.gxwf.yml
generated
2026-09-16
planemo_test_job_block
# job: block for the workflow's -tests.yml. Substitute hosted URLs for <FIXTURE_BASE>. RNA-seq reads (sample sheet): class: Collection collection_type: sample_sheet:paired rows: AR0382_A: [AR0382, A] AR0382_B: [AR0382, B] AR0387_A: [AR0387, A] AR0387_B: [AR0387, B] AR0382_tnSWI1_A: [tnSWI1, A] AR0382_tnSWI1_B: [tnSWI1, B] elements: - identifier: AR0382_A class: Collection type: paired elements: - identifier: forward class: File location: <FIXTURE_BASE>/AR0382_A_1.fastq.gz filetype: fastqsanger.gz - identifier: reverse class: File location: <FIXTURE_BASE>/AR0382_A_2.fastq.gz filetype: fastqsanger.gz - identifier: AR0382_B class: Collection type: paired elements: - identifier: forward class: File location: <FIXTURE_BASE>/AR0382_B_1.fastq.gz filetype: fastqsanger.gz - identifier: reverse class: File location: <FIXTURE_BASE>/AR0382_B_2.fastq.gz filetype: fastqsanger.gz - identifier: AR0387_A class: Collection type: paired elements: - identifier: forward class: File location: <FIXTURE_BASE>/AR0387_A_1.fastq.gz filetype: fastqsanger.gz - identifier: reverse class: File location: <FIXTURE_BASE>/AR0387_A_2.fastq.gz filetype: fastqsanger.gz - identifier: AR0387_B class: Collection type: paired elements: - identifier: forward class: File location: <FIXTURE_BASE>/AR0387_B_1.fastq.gz filetype: fastqsanger.gz - identifier: reverse class: File location: <FIXTURE_BASE>/AR0387_B_2.fastq.gz filetype: fastqsanger.gz - identifier: AR0382_tnSWI1_A class: Collection type: paired elements: - identifier: forward class: File location: <FIXTURE_BASE>/AR0382_tnSWI1_A_1.fastq.gz filetype: fastqsanger.gz - identifier: reverse class: File location: <FIXTURE_BASE>/AR0382_tnSWI1_A_2.fastq.gz filetype: fastqsanger.gz - identifier: AR0382_tnSWI1_B class: Collection type: paired elements: - identifier: forward class: File location: <FIXTURE_BASE>/AR0382_tnSWI1_B_1.fastq.gz filetype: fastqsanger.gz - identifier: reverse class: File location: <FIXTURE_BASE>/AR0382_tnSWI1_B_2.fastq.gz filetype: fastqsanger.gz Reference genome FASTA: class: File location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.fna.gz filetype: fasta decompress: true Gene annotation GTF: class: File location: https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/GCA_002759435.2_Cand_auris_B8441_V2_genomic.gtf.gz filetype: gtf decompress: true Reference condition level: AR0382 First contrast condition level: tnSWI1 Second contrast condition level: AR0387 Strandedness: stranded - reverse Adjusted p-value threshold: 0.05 log2 fold change threshold: 1.0
planemo_test_job_notes
2 item(s)
unresolved
4 item(s)
workflow-debug-reportnot-yet-dueworkflow-debug-report.md

Failure-surface classification with captured job/invocation/collection/assertion evidence and a recommended next step or reference-gap follow-up.

declared by
debug-galaxy-workflow-output (phase 12)
consumed at
nothing downstream reads it
schema
none declared
sha256

not on disk.

workflow-test-resultmissingworkflow-test-result.json

Structured status plus captured evidence — Planemo result, invocation/history/workflow ids, artifact paths, Galaxy mode, and (on failure) the observed modality and next reference surface — for debug-galaxy-workflow-output. Also the faithful handoff when no test exists or none could be run.

declared by
run-workflow-test (phase 11)
consumed at
12
schema
none declared
sha256

not on disk.

Obligations

what the workflow still owes

15 open, 13 resolved, 0 surrendered, 2 deliberately dropped. 0 open entries are blocking.

collection-algebra-never-statically-checkedopenvalidate-galaxy-workflow → still open
unmet
any static confirmation that the workflow's collection shapes and map-over structure are compatible end to end

Terminal validation could not run the one check that covers this. `gxwf validate --connections` — the flag whose entire purpose is connection-type compatibility, collection algebra and map-over — aborts with `TypeError: step.in is not iterable` before emitting any report, on a format2 workflow written the way this one is: list-form `inputs:`/`steps:` with mapping-form `in:`, which gxwf's `toNative` mistakes for an already-normalized workflow and therefore never normalizes. Reproduced on a 20-line minimal case and unchanged on gxwf 1.12.0, so it is a gxwf defect and not a property of this workflow (filed as `tonative-shape-sniff-skips-normalization-on-list-form-format2`; fix submitted as jmchilton/galaxy-tool-util-ts#179). Two routes could still close this without an upstream release: rewriting the workflow's `in:` blocks in list form, or running `--connections` against a patched gxwf. What the default validation does cover is per-step tool state on the 20 decodable steps, and nothing about how shapes compose. The map-over structure is the load-bearing part of this design and none of it is proven: the `sample_sheet:paired` input fanning through `__FLATTEN__` into FastQC's per-fastq axis; the paired collection entering Cutadapt at `library|input_1` and leaving as `out_pairs` into RNA STAR's `singlePaired|input`; three `__FILTER_FROM_FILE__` reductions over `Count reads per gene/output_short`; and the workflow's only true reduction, where each DESeq2 node's multiple=true `countsFile` ports consume whole 2-element collections and the per-sample axis collapses.

Not a defect claim against the workflow — every `in:` source resolves, every `in:` key name and tool_state parameter name was checked by hand against the real tool input trees for all 27 steps (the 20 cached tools from the gxwf cache, the 7 undecodable ones fetched from `GET /api/tools/<id>?io_details=true`) with zero mismatches, and every select and boolean value on the previously unchecked steps is legal. What is unmet is that shape COMPOSITION has no static witness, and cannot get one until gxwf is fixed. Closable only by the first successful invocation. Phase 11 should read `GET /api/invocations/{id}/step_jobs_summary` per step rather than only the terminal invocation state: a shape mismatch surfaces as invocation message `collection_failed`, or as an unexpected job count on a mapped step, and a workflow can report `completed` while a mapped step ran the wrong number of jobs. Two adjacent exposures ride along and are settled by the same run: whether the target Galaxy supports the `sample_sheet:paired` input with `column_definitions` and `__SAMPLE_SHEET_TO_TABULAR__` at all (see `no-iwc-precedent-for-sample-sheet-workflow-input` and the filed `galaxy-test-staging-drops-sample-sheet-column-definitions`), which would fail at request-time validation or input materialization rather than in any tool's stderr; and the element identifiers `__FLATTEN__` produces, since step 0 carries no tool_state and `join_identifier` therefore takes the wrapper default, while the sample-sheet element identifiers are declared the spine of this workflow.

cross-step-invariants-survive-only-in-the-draftopenadvance-galaxy-draft-step → still open
unmet
a durable record, in the runnable artifact, of the two cross-step couplings that keep the filter predicates pointing at the right columns

Two couplings hold this workflow together and neither is expressible in gxformat2 nor checked by any gate. (1) `advanced_options.lfc_shrinkage_type: none` on BOTH DESeq2 nodes is what makes `deseq_out` a 7-column table, so `c3` = log2FoldChange and `c7` = padj — the columns the `compose_text_param` bridges hard-code as `"abs(c3)>"` and `"c7<"` in four OTHER steps. Under any other shrinkage the `stat` column is dropped, padj moves to c6, and `"c7<"` silently filters on nothing. (2) `header_lines: '0'` on all four significance filters, against a corpus exemplar that binds `'1'` — the exemplar manufactures the header it then skips, this workflow does not. `'1'` would discard row 1 of a padj-sorted table, i.e. *SCF1*, the paper's finding. Both were established by hand and recorded only in YAML comments above the steps in `galaxy-workflow-draft.gxwf.yml`. `gxwf draft-extract` re-serializes from a parsed model and destroys every comment, so `galaxy-workflow.gxwf.yml` — the artifact the test, validate, run and any later maturation phase actually consume — carries no trace of either. The values themselves are correct in the extract (verified at iteration 26); what is missing is any reason a later editor would know not to change them.

Not closable inside the per-step loop: the loop's only writable surface is the draft, and the draft's comments are exactly what the extract drops. Closable by (a) folding both couplings into the step `doc:` strings of the two DESeq2 nodes and the four filters — `doc:` survives extraction and is visible in the Galaxy editor — or (b) making the test plan assert the column contract directly (a test on `deseq_out` that the padj column is c7, or on a filtered table that its top row survived), which converts a silent invariant into a failing test. (a) is the cheap one and should be done before the draft is discarded; (b) is the one that would actually hold. Filed upstream as `draft-extract-destroys-yaml-comments` in `foundry-feedback.ledger.yml`; that is the tool-side gap, this entry is the run-side obligation it leaves behind. ADVANCED, NOT CLOSED, by freeform-summary-to-galaxy-test-plan. Route (b) was taken and it lands unevenly across the two invariants. INVARIANT 1 is now fully detectable: both test cases assert `has_n_columns: 7` on BOTH `DESeq2 results:` tables, and case 1 asserts it again on both `Significant genes:` tables, so any shrinkage setting other than `none` on either node drops the `stat` column and fails the assertion before the `c7<` predicate can silently filter on nothing. A `not_has_text: baseMean` assertion on both raw tables covers the headerless premise the same predicates rest on. INVARIANT 2 is only partly detectable, and the plan says so in warnings[] `reference-level-filter-regression-undetectable`: of the seven Filter1 steps binding `header_lines: '0'`, a regression is visible on two — `Select first-contrast samples` and `Select second-contrast samples`, where the leaked metadata row would add a fifth sample column to that contrast's normalized-counts table, which both cases assert against. On `Select reference-level samples` the leaked row IS `AR0382_A`, which the `c2=='AR0382'` predicate keeps anyway, so the output is byte-identical and no assertion can see it. On the four significance filters the leaked row is row 1 of a padj-sorted table, which passes both filters on its own merits, so the regression is invisible at workflow level. Route (a) — folding both couplings into the `doc:` strings of the two DESeq2 nodes and the seven filters, which survives `draft-extract` — was NOT done and remains the cheap half of this entry. It is now the only protection the three undetectable reference-level and significance filters have. Do it before `galaxy-workflow-draft.gxwf.yml` is discarded. FURTHER ADVANCED by implement-galaxy-workflow-test, on the invariant-2 half only. The normalized-counts header width is no longer an assumption: deseq2.R at the pinned 2.11.40.8+galaxy4 writes `counts_out` with `write.table(..., col.names = NA)`, which emits a header padded with a leading blank field, so a four- sample contrast has exactly five tab-separated fields on line 1. Both test cases now assert `has_n_columns: 5` on both normalized-counts tables, and case 1 keeps the SCF1 data-row field-count assertion alongside it. A `header_lines: '1'` regression on either contrast-selecting Filter1 leaks a fifth sample column and fails both. The three undetectable filters - `Select reference-level samples` and the four significance filters - are unchanged, and route (a), folding both couplings into the `doc:` strings that survive `draft-extract`, is still not done and is still the only protection they have. RE-VERIFIED AT THE TERMINAL GATE by validate-galaxy-workflow, still not closed. Both couplings hold in `galaxy-workflow.gxwf.yml` as assembled: `advanced_options.lfc_shrinkage_type: 'none'` on both DESeq2 nodes (neither selects `many_contrasts`/`split_output`; both are `select_data.how: datasets_per_level`), and `header_lines: '0'` as the string `'0'` on all seven Filter1 steps (8, 12, 16, 23, 24, 25, 26 — the Filter1 count is exactly seven). Two refinements to the entry's own statement of invariant 1, from reading the assembled artifact. The `"c7<"` predicate is hard-coded in ONE `compose_text_param` bridge, step 21, consumed by two Filter1 steps (23 and 25); the second bridge, step 22, is `"abs(c3)>"`, consumed by 24 and 26. Since `stat` sits at c5, dropping it moves padj from c7 to c6 but leaves log2FoldChange at c3 — so the shrinkage coupling protects the padj predicate specifically, and the `abs(c3)>` bridge is not exposed to it. Route (a) remains not done and the terminal artifact confirms where the hole is: the four significance filters' `doc:` strings do name their `c7<` / `abs(c3)>` predicates and the three condition-split filters' docs do say "headerless", but NEITHER DESeq2 node's `doc:` mentions shrinkage at all, and no filter doc says why `header_lines` is `'0'`. An editor changing `lfc_shrinkage_type` has nothing in the runnable artifact to warn them. `galaxy-workflow-draft.gxwf.yml` still holds the rationale in the YAML comments `draft-extract` strips; do not discard it before the `doc:` strings are written.

cutadapt-adapter-and-length-filter-unstatedopenfreeform-summary-to-galaxy-interface → still open
unmet
Cutadapt adapter sequence and minimum-length filter

The paper gives only "Cutadapt with a Phred cutoff score of 20". No adapter sequence appears anywhere in the supplement, and no minimum-length filter is stated. The interface reads the step as quality trimming only (`-q 20`, baked in) and exposes no adapter input.

If a later phase adds adapter trimming it is adding method the paper does not describe, and that addition has to be labelled as such rather than presented as a faithful port. STILL OPEN after phase 6 iteration 8 concretized `Quality-trim reads`, and deliberately so. The step pins lparsons/cutadapt 5.2+galaxy2 with all six adapter repeats written explicitly empty (adapters / front_adapters / anywhere_adapters on R1, the adapters2 / front_adapters2 / anywhere_adapters2 trio on R2) and `other_trimming_options.quality_cutoff: '20'`, which is the one Cutadapt parameter the paper actually states. The unstated minimum-length filter is now recorded concretely: the pinned wrapper's default is `filter_options.minimum_length: 1`, which the XML flags as a deliberate wrapper-side departure from cutadapt's own default of 0 ("intentionally set to 1 ... to avoid hard to debug issues with downstream tools"). The step writes that 1 out explicitly so a later wrapper bump cannot move it silently -- but it is the WRAPPER's default, not the paper's parameter, and the corpus shows the value is genuinely chosen per workflow (cutandrun at fe41a79 uses 15, the VGP workflows use 1). Closing this entry requires a stated adapter and a stated length cut, neither of which exists in the source.

deseq2-statistics-unverified-against-the-galaxy-wrapperopenfreeform-summary-to-galaxy-test-plan → still open
unmet
confirmation that the workflow's own DESeq2 node reproduces the differential-expression result the test plan asserts

Every differential-expression figure this run has is from a different implementation than the workflow runs. Phase 8 measured SCF1's log2 fold change and adjusted p-value with pydeseq2 over a strand-aware count matrix built from phase 7's six bwa-mem BAMs, as two separate two-level n=2 analyses with no LFC shrinkage — tnSWI1 vs AR0382 log2FC -6.93 / padj 2.1e-31, AR0387 vs AR0382 log2FC -7.95 / padj 4.1e-18. The workflow runs the Galaxy DESeq2 wrapper, i.e. R DESeq2, over RNA STAR plus featureCounts -s 2. pydeseq2 is a faithful reimplementation, but it is not the same implementation and the counts are not the same counts, so the exact values will differ. What is unverified is not the biology — the separation is ~200-fold at the counts level and the padj margins are 18 and 31 orders of magnitude — but that the workflow's own node produces a row for B9J08_001458 clearing `padj < 0.05` and `|log2FC| > 1` into each significant table with `log2FoldChange <= -2`, which is what the test plan asserts, and that the raw table really is the 7-column shape both cases assert `has_n_columns: 7` on.

Closable by the first successful invocation and nothing cheaper; phase 11 (run-workflow-test) is where it lands, and phase 9 should not try to settle it by adding assertions. The plan is already written so that this can only bite as a margin question rather than a value question: no assertion anywhere pins an exact statistic, and the recorded measurements are stated as headroom over thresholds (~5 and ~6 log2 units over the -2 floor; 31 and 18 orders of magnitude over the 0.05 cut). Three things to check against the first invocation while it is open — the sign and magnitude of SCF1's c3 in both contrasts, that c7 is padj and not something else, and whether R DESeq2's independent filtering keeps SCF1 in the padj-carrying set at this depth as pydeseq2 did (2393 and 1376 genes retained). Recorded in the plan as warnings[] `deseq2-statistics-measured-with-pydeseq2-not-r`.

galaxy-tool-versions-unpinnable-from-sourceopenfreeform-summary-to-galaxy-interface → still open
unmet
tool versions matching the published analysis

The paper gives no version for any Galaxy step (FastQC, Cutadapt, RNA STAR, featureCounts, DESeq2). Only non-Galaxy software is versioned (R 4.0.3, DescTools 0.99.49, survminer 0.4.9, Fiji 1.52, CellProfiler 3.1.9). The constructed workflow will pin its own versions and can reproduce the authors' method but never their software stack.

Unclosable from the source. Expected to be surrendered at the terminal and stated on the workflow itself, so a reader does not mistake the run for a version-faithful reproduction.

interface-brief-output-and-parameter-surface-drifted-from-draftopenfreeform-summary-to-galaxy-template → still open
unmet
agreement between freeform-galaxy-interface.md and the settled draft's public input and output surface

Phases 4 and 5 changed the interface in four places that the phase-2 brief still states in its original form. The brief is the artifact the test plan reads for labels, and labels are the API, so the drift has to be visible rather than inferred by diffing two artifacts. (a) Interface input 7 `Minimum absolute fold change` (float, linear, default 2.0) is in the draft `log2 fold change threshold` (float, log2 units, default 1.0) — closed on corpus evidence by `fold-change-threshold-linear-vs-deseq2-log2fc`. (b) Interface input 5 `featureCounts strandedness` (free text, allowed values named only in prose) is in the draft `Strandedness` (text restricted to `stranded - forward` / `stranded - reverse` / `unstranded`, default `stranded - reverse`) feeding a `map_param_value` bridge with `on_unmapped: fail` — the corpus idiom adopted per the IWC comparison section 3.6. (c) Two inputs the brief does not declare at all: `First contrast condition level` and `Second contrast condition level` — see `deseq2-factor-level-names-not-parameterized`. (d) Fourteen declared outputs become sixteen — see `deseq2-normalized-counts-and-plots-promoted-per-contrast`.

Bookkeeping, not a design gap: every one of the four changes is itself settled and recorded, and the draft is the current statement of the interface. What is unmet is that no Mold in this run's roster owns freeform-galaxy-interface.md after phase 2, so the edits have no writer. Recorded here so the test-plan phase reads labels off the draft rather than off the stale brief, and so a later maturation pass knows which artifact was authoritative. Closable by editing the brief's sections 2.3 and 3, or by declaring the draft authoritative for the interface and saying so in the brief.

iwc-splits-rnaseq-de-at-the-count-table-boundaryopencompare-against-iwc-exemplar → still open
unmet
agreement between this run's workflow scope and published IWC practice

IWC has no single workflow spanning FastQC to DESeq2. At corpus fe41a79 the journey is published as two workflows joined at the count-table boundary: transcriptomics/rnaseq-pe/rnaseq-pe ends at per-sample count tables, and transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting begins there, taking pre-grouped `list` collections of count tables as its workflow inputs. This run's design is one workflow spanning both halves, which is why it needs an in-workflow condition split at all — a workflow that starts from count tables can demand pre-grouped collections at its interface and never reconstruct grouping internally.

Not a defect in either design, and not a reason to re-scope: the paper describes one analysis and the run was asked to build it. Recorded because it is the root of the single largest divergence from corpus practice (see `no-iwc-precedent-for-sample-sheet-workflow-input`) and because any later mature-galaxy-workflow-for-iwc pass will meet it as a packaging question — a reviewer may ask why this is not two workflows. Whoever revisits it should weigh that the split also costs reusability of the DE half, which is presumably why IWC made it.

mapped-outputs-carry-sample-sheet-outer-axisopenfreeform-summary-to-galaxy-data-flow → still open
unmet
accurate declared collection types for the promoted per-sample outputs

A tool mapped over a `sample_sheet`-family collection produces a `sample_sheet`-shaped output without `column_definitions` (packaged note galaxy-sample-sheet-collections, "Mapping rules"). The interface's output table calls outputs 1-8 `list` / `list:paired`. Behaviourally that is accurate — such a collection maps, reduces and filters exactly like a list — but the declared collection type string is not `list`, which matters to anything that type-checks, including workflow test assertions on collection type.

No topology consequence; the fix is either a corrected type column in the interface brief or an explicit statement that the promoted outputs are sample_sheet-shaped lists without column metadata. Flagged primarily for the test-plan phase, which is where a wrong collection_type assertion would surface as a confusing failure. ADVANCED, NOT CLOSED, by implement-galaxy-workflow-test. The test file authors no `attributes: {collection_type: ...}` assertion on any of the eight mapped outputs, as the test plan's `promoted- collection-type-string-unverified` instructs. What it asserts instead is `element_count` plus a named `element_tests` entry per identifier - twelve for the flattened FastQC outputs, six for every output on the sample axis - which is shape-agnostic and pins the thing that actually matters, the identifier space the condition split joins on. So no assertion in this run depends on the declared type string and nothing here forces the question. It stays open as an interface-brief accuracy item, and the real type strings should be read off the first successful invocation before anyone decides whether a collection_type assertion is worth having.

no-iwc-precedent-for-sample-sheet-workflow-inputopencompare-against-iwc-exemplar → still open
unmet
a worked corpus precedent for the sample-sheet condition split

Searched across all of `workflows/` at corpus fe41a79: `sample_sheet` as a collection type appears in zero workflows; `__SAMPLE_SHEET_TO_TABULAR__` appears in zero workflows; `column_definitions` appears only as the literal `null` that newer Galaxy serializes onto ordinary collection inputs, never with a value. `__FILTER_FROM_FILE__` does appear, in six workflows, but none in transcriptomics and none for a condition split. So the second half of the data-flow brief's section 4.2 route is a real used Galaxy idiom while the sample-sheet half is unprecedented in published IWC practice, and no corpus fixture declares a sample-sheet collection either.

Absence of precedent is not refutation, and nothing in the corpus contradicts the route — the phase-3 reasoning stands on its own. What changes is the risk profile: this is the one region of the workflow with no worked example to pattern-match against, and two open entries ride on it (`sample-sheet-to-tabular-identifier-column-unverified`, `sample-sheet-input-test-fixture-expressibility`). Guidance for the template, from the comparison notes section 3.1: build the split as a clearly delimited region with the phase-2 fallback (a `data` input `Sample metadata table`) reachable by deleting one node, so that if either open entry resolves against the sample sheet the cost is a deletion rather than a redesign. Worth knowing for the record that IWC made the opposite trade deliberately — it pushes grouping into the interface as two pre-grouped count collections, which is the same shape the interface brief's section 2.1 rejected for hard-coding the design into the public API.

pipeline-b-tdna-mapping-not-carriedopendroppedfreeform-summary-to-galaxy-interface → still open
unmet
the second of the paper's two sequencing analyses

No T-DNA integration sites are produced by this run. The scope decision was taken by the harness before this phase.

units
the whole of Pipeline B — AtMT T-DNA insertion-site mapping: FastQC, Trimmomatic, BWA-MEM against linearized pTO128 (seed 50, band width 2), extractSoftClipped, BWA-MEM of the soft-clipped flanks against C. auris B8441 (5 computational steps, 2 deposited runs SRR22376033–SRR22376034)
because
Two blockers, both recorded in the source summary's open questions (4 and 5) and neither independently verified in this phase. (a) `extractSoftClipped` from SE-MEI (github.com/dpryan79/SE-MEI) has no known Galaxy Tool Shed wrapper, so the pipeline cannot be assembled from stock tools as written; substituting a samtools/awk soft-clip extraction would change the method. (b) The pTO128 (pPZP-NAT) plasmid reference has no public accession in the paper — it is cited to the prior AtMT method paper (ref. 52) — so step 3's reference is unresolvable from the publication alone.

Uncited in the strict sense: no Tool Shed search for extractSoftClipped and no Addgene or ref. 52 lookup for pTO128 was run in this phase; both reasons are carried from the source summary. Treat as a debt, not a finding. If a later phase's discovery step contradicts either reason, this entry's `because` is refuted and the entry must be reopened rather than left standing.

reference-genome-delivery-shape-unverifiedopenfreeform-summary-to-galaxy-interface → still open
unmet
verified delivery shape for the C. auris B8441 reference genome

The interface settles the genome as a history `fasta` dataset on portability grounds (a remote-URL fixture resolves on any server; a CVMFS built-in index does not), with RNA STAR building its index at run time. Whether usegalaxy.org carries a built-in index for GCA_002759435.2 was not checked, and this run's phase roster contains no reference-data Mold that owns the question.

Provisionally settled, not verified. A built-in index would be cheaper at run time but less portable for tests; revisiting it changes workflow input 2.

scf1-fold-change-magnitude-disagrees-with-paperopenpaper-to-test-data → still open
unmet
reconciliation of the paper's ~29-fold SCF1 figure with the deposited read data

The paper states a ~29-fold SCF1 expression difference between AR0382 and AR0387. Direct measurement of the deposited runs disagrees on magnitude by roughly an order of magnitude: aligned fragments over the SCF1 exon at 200,000 read pairs per sample give 414 and 391 in the two AR0382 replicates against 2 and 1 in AR0387 (~270-fold), and exact 31-mer matching against the SCF1 CDS over 5.2 M reads gives 1942.7 vs 12.3 per million (~158-fold). Direction and significance are not in doubt; only the ratio is. Candidates not distinguished here: the 29-fold may be the RT-qPCR assay rather than the RNA-seq, a shrunk rather than raw log2FC, or a different normalization. Reference bias is ruled out as the explanation — AR0387 IS the B8441 reference strain, so its SCF1 reads cannot be undercounted; if anything AR0382's divergent allele is undercounted, which makes the true ratio larger, not smaller.

Consequence for the test plan, and the reason this is filed rather than glossed: no assertion may pin the paper's 29-fold figure. `test-data-refs.json` asserts direction and significance only (`log2FoldChange < -2`, `padj < 0.05`), which both the paper and the data support. Settling this needs the paper's source data (Data S1) or the authors, not more compute. REINFORCED by freeform-summary-to-galaxy-test-plan, and one candidate explanation is now the leading one. The disagreement was previously measured only on raw counts and k-mer matches; phase 8 ran an actual differential-expression model — pydeseq2, two separate two-level n=2 analyses with NO LFC shrinkage, mirroring the workflow's two DESeq2 nodes, on a strand-aware count matrix built from phase 7's six BAMs. AR0387 vs AR0382 gives SCF1 log2FC -7.95 at padj 4.1e-18, i.e. ~247-fold, consistent with the ~270-fold raw fragment ratio and not with the paper's ~29-fold. tnSWI1 vs AR0382 gives log2FC -6.93 at padj 2.1e-31. Because that run applied no shrinkage, it does not rule out the "shrunk rather than raw log2FC" candidate — log2(29) = 4.86 against an unshrunk 7.95 is exactly the direction and roughly the magnitude a shrinkage estimator would move a gene whose counts are near zero in one group, so that candidate is now the most plausible of the three and the RT-qPCR candidate is not excluded either. This matters for the workflow, not just the paper: the workflow pins `lfc_shrinkage_type: none`, so ITS log2FoldChange for SCF1 will be the unshrunk ~-8 and will never reproduce the paper's figure by construction. Do not treat that as a defect when the first real invocation lands. Caveat on all of the above: pydeseq2 is not the R DESeq2 the Galaxy wrapper runs, and the counts came from bwa-mem rather than RNA STAR plus featureCounts -s 2; see `deseq2-statistics-unverified-against-the-galaxy-wrapper`.

star-genome-length-drives-sa-index-parameteropenadvance-galaxy-draft-step → still open
unmet
a measured length for the reference actually wired into `Reference genome FASTA`

`genomeSAindexNbases` exists ONLY in the RNA STAR history branch — it configures the in-job `--runMode genomeGenerate` that this workflow's history-FASTA delivery choice introduces, so it is not a mapping parameter the paper's "default parameters" could have covered. The wrapper's own help carries STAR's formula verbatim: "For small genomes, the parameter --genomeSAindexNbases must be scaled down to min(14, log2(GenomeLength)/2 - 1)". The step binds `'10'`, from min(14, log2(12.5e6)/2 - 1) = min(14, 10.79). The 12.5 Mb is the GCA_002759435.2 assembly record's length as carried through this run's briefs; nothing in this run measured the dataset that will actually be wired.

Silent-degradation class, not a hard gate: the wrapper default of 14 is the documented seg-fault-at-mapping hazard for a genome this small (the wrapper's own test data uses 5), and a value too small merely costs search speed. Two ways this becomes wrong: a different reference is supplied at run time (a mammalian genome would want 14 back), or the B8441 FASTA in hand differs materially in length from the assembly record. Settle it by reading the length of the actual dataset, or by reading STAR's own recommendation out of the promoted `STAR mapping summary` / job stderr on the first real run — STAR prints the recommended value when the supplied one is too large. Tied to `reference-genome-delivery-shape-unverified`: if that resolves to a built-in index, this parameter disappears with the branch. ADVANCED, NOT CLOSED, by paper-to-test-data. The entry asked for a measured length of the reference actually wired in. The reference this run will wire is now pinned — NCBI `GCA_002759435.2_Cand_auris_B8441_V2_genomic.fna.gz`, md5 a008b270d3aaa04736a8bb7daf3f6dd5 — and it was downloaded and measured: 12,365,959 bp across 15 contigs. min(14, log2(12365959)/2 - 1) = min(14, 10.78), so the bound `'10'` is correct for this dataset. What keeps the entry open is exactly the exposure it already named: the parameter is correct for the reference the TEST supplies, and a user who supplies a different genome at run time still gets a silently wrong value. The first real STAR run should confirm from the promoted `STAR mapping summary`.

test-fixtures-not-hostedopenpaper-to-test-data → still open
unmet
a resolvable URL for each of the 12 read fixtures

The 200,000-read-pair subsets were generated and hashed but not published. They exist only in the producing session's scratchpad and do not survive it. At ~61 MB compressed for 12 files they are too large to commit alongside the workflow, so the IWC idiom applies: host them (Zenodo or equivalent) and reference them by URL. `test-data-refs.json` carries `<FIXTURE_BASE>` as the placeholder in `planemo_test_job_block`.

Not a research gap — the recipe is deterministic and the artifact pins it. Regenerate by streaming each ENA FASTQ and taking `head -n 800000`, then verify against the `md5_uncompressed` value recorded per file (hash the UNCOMPRESSED `.fastq`; gzip output is not byte-stable across gzip versions). The reference genome and annotation need no hosting: both are pinned to stable NCBI FTP URLs with md5s and are staged with `decompress: true`. ADVANCED, NOT CLOSED, by implement-galaxy-workflow-test. All twelve 200,000-read-pair fixtures were regenerated in this phase with the recorded recipe (`curl -L <ENA url> | gzip -dc | head -n 800000`) and every one verified byte-for-byte against its `md5_uncompressed` in test-data-refs.json before compression, so the recipe and the hashes are now proven reproducible rather than merely recorded. They live at `<run>/test-data/*.fastq.gz` (59 MB) and `galaxy-workflow.gxwf-tests.yml` addresses them as local `path:` entries, each with the SHA-1 of the `.gz` artifact as produced here. A twelve-file 25,000-read-pair set for the smoke case was derived from the verified 200k files by taking their first 100,000 lines, at `<run>/test-data/smoke-25k/` (6.7 MB). The test therefore runs against real data in this run directory and nowhere else: the fixtures are still unhosted and do not survive the run, which is the whole of what remains unmet. Publishing them is now a mechanical step - upload the exact `.gz` bytes, replace each `path:` with the published `location:`, and leave the SHA-1 values unchanged, because they pin those bytes. Keep the uncompressed md5s in test-data-refs.json as the regeneration check; the SHA-1s are the fetch-integrity check and the two answer different questions.

tnbcy1-contrast-not-carriedopendroppedfreeform-summary-to-galaxy-interface → still open
unmet
one of the three transcriptome comparisons the paper reports

The workflow can build only the two contrasts whose input data is public. A test that tries to reproduce Fig. S2 has no input.

units
the tnBCY1 (B9J08_002818) vs AR0382 transcriptome comparison reported in Fig. S2
because
No tnBCY1 runs exist in BioProject PRJNA904261. Only three conditions were deposited (AR0382, AR0387, AR0382 tnSWI1), across six runs SRR22376027–SRR22376032.

Cut by data availability, not by a design decision. Adding a fourth condition later needs no interface change: the reads input is a sample sheet whose `condition` restrictions widen.

deseq2-contrast-realization-unsettledresolvedfreeform-summary-to-galaxy-interfacecompare-against-iwc-exemplar
unmet
step realization behind the two contrast outputs

The paper reports two contrasts (tnSWI1 vs AR0382, AR0387 vs AR0382) and states the significance thresholds, but writes no design formula. One factor, three levels, n = 2 is inferred from the six deposited runs. Whether the two contrasts come from one DESeq2 run over a three-level factor or two runs over two-level factors depends on the chosen wrapper's contrast handling, which is not known at interface time.

The interface fixes only the output surface — two result tables and two filtered tables, labelled per contrast. Either realization satisfies it. Closed by the IWC exemplar comparison, section 3.2, against transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting at corpus fe41a79. Two DESeq2 nodes, each with a two-level factor, the reference-level counts collection feeding both: H1 (AR0382, tnSWI1) and H2 (AR0382, AR0387). Evidence: the exemplar's DESeq2 step uses `select_data.how: datasets_per_level` with a `rep_factorName_0.rep_factorLevel` repeat of exactly two levels, each port consuming a whole collection through the wrapper's multiple=true input; its `deseq_out` is a single dataset, not a collection, proved by the single-dataset tools it feeds (deg_annotate, tp_cat, two Filter1 steps) and by the sibling -tests.yml asserting has_text_matching on the derived output as a dataset; and the workflow README bounds itself to "exactly 2 conditions with at least 2 replicates per condition". One DESeq2 job therefore yields one results table, so the interface's two distinctly labelled result tables require two jobs. A three-level factor is expressible via the repeat but would still emit one `deseq_out`, which cannot satisfy a two-output interface — so the output surface settles the arity regardless of the wrapper's three-level contrast behaviour. Upstream wiring is unaffected, exactly as the data-flow brief's section 4.4 predicted: nodes E, F1-F3 and G1-G3 are identical either way. One consequence for the interface: under two DESeq2 nodes, `DESeq2 normalized counts` (output 9) and `DESeq2 diagnostic plots` (output 14) are each produced twice while the interface declares one of each. Promote from one designated run and say which, or relabel per contrast; do not leave it implicit in the template.

deseq2-factor-level-names-not-parameterizedresolvedfreeform-summary-to-galaxy-data-flowfreeform-summary-to-galaxy-template
unmet
a source for the two non-reference condition level literals

The condition split needs one literal condition value per factor level to filter the sample metadata table (AR0382, AR0387, tnSWI1). The interface exposes only `Reference condition level` (input 4, default AR0382). The two contrast levels have no parameter, so as the interface stands they would be baked into two filter steps — re-hard-coding into the steps the three-level design that the sample-sheet input was chosen to keep out of the public interface.

Two ways to settle it, both changing the interface's parameter surface rather than the topology: expose two more text parameters (one per contrast level) alongside the existing reference-level parameter, or accept the bake and state plainly in the interface that the three level names are fixed in the steps. Whichever is chosen must be reflected back into freeform-galaxy-interface.md section 2.3, not applied silently in the template. Closed by the template on the first of those two options. The draft exposes two further `text` workflow inputs, `First contrast condition level` (default tnSWI1) and `Second contrast condition level` (default AR0387), alongside the existing `Reference condition level` (default AR0382). Each feeds a `compose_text_param` step that builds the row predicate for its level, so no level literal is baked into any step. The deciding argument is symmetry: the reference level was already a parameter for exactly this reason, and baking the other two would have left the interface half-generalized while re-hard-coding into the steps the design the sample-sheet input was chosen to keep out of the interface. The cost is two inputs the phase-2 brief does not declare, which is one facet of the interface drift recorded in `interface-brief-output-and-parameter-surface-drifted-from-draft`; that entry carries the owed edit to freeform-galaxy-interface.md section 2.3, since this Mold does not own that artifact. One caveat the draft states on both new inputs: the contrast output labels (`... tnSWI1 vs AR0382`, `... AR0387 vs AR0382`) are the public API and are not derived from these parameters, so changing a level value makes the labels stale.

deseq2-normalized-counts-and-plots-promoted-per-contrastresolvedfreeform-summary-to-galaxy-templatefreeform-summary-to-galaxy-template
unmet
a decision on the two DESeq2 outputs the interface declares once and two DESeq2 jobs produce twice

`deseq2-contrast-realization-unsettled` closed on two DESeq2 nodes, which makes `DESeq2 normalized counts` (interface output 9) and `DESeq2 diagnostic plots` (interface output 14) each produced twice while the interface declares one of each. The IWC comparison flagged the consequence and instructed the template not to leave it implicit: promote from one designated run and say which, or relabel per contrast.

Settled by relabelling per contrast. The draft promotes four outputs where the interface declared two: `DESeq2 normalized counts: tnSWI1 vs AR0382`, `DESeq2 normalized counts: AR0387 vs AR0382`, `DESeq2 diagnostic plots: tnSWI1 vs AR0382` and `DESeq2 diagnostic plots: AR0387 vs AR0382`, taking the workflow to sixteen outputs. This is not a tie broken on taste. Each DESeq2 job sees only its own four samples, so its size factors, normalized values and PCA/dispersion plots are computed over that subset: the two normalized-count tables are different tables, not two copies of one. Promoting either and calling it "the" normalized counts would be wrong rather than merely arbitrary, and dropping one would discard a real artifact while leaving the surviving one silently contrast-scoped. Relabelling also makes the four outputs symmetric with the result and filtered-gene tables, which are already labelled per contrast. One consequence for the test plan: the paper's sanity check that SCF1 sits in the top 2.5% of AR0382 expression can be asserted against either normalized-counts table, because both carry the two AR0382 replicates. Assert it on one and say which. The owed edit to freeform-galaxy-interface.md section 3 rides on `interface-brief-output-and-parameter-surface-drifted-from-draft`; labels are the public API and this is a breaking change to be made once, before any test is written.

deseq2-result-table-header-presence-unverifiedresolvedfreeform-summary-to-galaxy-templateadvance-galaxy-draft-step
unmet
the `header_lines` binding for the four Filter1 steps of the significance chains

The corpus exemplar binds `header_lines: "1"` on both of its Filter1 steps, but it filters a table that has been through `deg_annotate` and then `tp_cat`, and the tp_cat step exists precisely to concatenate a separately generated single-line header onto the DESeq2 output. That strongly implies the raw `deseq_out` is headerless. This workflow omits the annotation pair — the paper names no annotation step and the interface's tool set is closed — so its Filter1 steps consume `deseq_out` directly and the corpus binding cannot be copied across. Whether the pinned iuc/deseq2 wrapper writes a header row on `deseq_out` is not established by anything this run read.

Consequential if missed rather than untidy, which is why it is a ledger entry and not only a `_plan_state` line. Binding "1" against a headerless table silently drops the first gene row of every filtered result; binding "0" against a headed table passes the header line into a numeric comparison, where Filter1 discards it with a warning rather than an error. Either way the workflow produces a plausible table. The same binding applies to the three row filters in the condition-split region, where the unknown is the header behaviour of `__SAMPLE_SHEET_TO_TABULAR__` rather than of DESeq2 — a different producer, the same class of silent loss, tracked there by `sample-sheet-to-tabular-identifier-column-unverified`. Closable by a summarize-galaxy-tool pass on iuc/deseq2 during the per-step loop, or by inspecting the first real run's output. SETTLED in the per-step loop, at the step that pins the producer (`Differential expression: first contrast vs reference`, iuc/deseq2/deseq2 @ 2.11.40.8+galaxy4, changeset 05f9e54d7e81): `deseq_out` carries NO header row, so all four significance filters bind `header_lines: '0'` — the same value the three condition-split row filters already use, for an unrelated reason. Three independent lines of evidence, all at the pinned version. (1) The wrapper's own `deseq2.R` writes the result table with `write.table(out_df, file = opt$outfile, sep = "\t", quote = FALSE, row.names = FALSE, col.names = FALSE)` — `col.names = FALSE`, at both of the two call sites that write a result table (the single-contrast path and the `many_contrasts` loop). The same script writes `counts_out` with `col.names = NA`, so the normalized-counts table DOES carry a header: one tool, opposite answers for its two tabular outputs, which is exactly why this could not be settled by analogy. (2) `deseq2.xml`'s test assertions for `deseq_out` match a data row first (`FBgn0003360\t1933.9504…\t-2.8399…`) with `has_n_lines n="3999"` and assert no header, while `vst_out` and `counts_out` assertions in the same test block DO assert a sample-name header line. (3) The corpus chain explains its own "1": rnaseq-de MANUFACTURES the header it later skips — `tp_text_file_with_recurring_lines` emits the single line `GeneID__tc__Base mean__tc__log2(FC)…`, `tp_sed_tool` turns `__tc__` into tabs, and `tp_cat` (labelled `Annotate DESeq2 table`) concatenates it on top of the deg_annotate output; only then do its two Filter1 steps bind `header_lines: "1"`. That "1" is about the manufactured header, not about `deseq_out`, which confirms rather than contradicts the reading above. One rider the four filters inherit: the c3 = log2FC / c7 = padj column contract holds only while `lfc_shrinkage_type` is `none` — `get_result_output_columns()` drops the `stat` column under any shrinkage, moving padj to c6. The DESeq2 steps pin `none` explicitly for that reason.

fastqc-per-read-fanout-not-a-flat-listresolvedfreeform-summary-to-galaxy-data-flowcompare-against-iwc-exemplar
unmet
agreement between the declared shape of outputs 1-2 and what map-over actually produces

The interface declares `FastQC raw reads: text summary` and `: HTML report` as `list` collections. FastQC consumes a single dataset, so mapping it over a `sample_sheet:paired` input fans out over the inner paired axis as well — 12 jobs, and outputs nested one report per read direction per sample, not six flat elements. As declared, outputs 1 and 2 are not what the workflow computes.

The data-flow brief (section 7.1) recommends promoting the nested collection and correcting the interface: per-read-direction QC is what a reader wants from raw-read QC, and it preserves the element identifier space that every checkpoint assertion keys on. The alternative is an explicit flatten node, which satisfies the declared `list` but rewrites identifiers to a doubled vocabulary (`AR0382_A_forward`) for no analytical gain. Either way the test plan must know which, because it changes every assertion on outputs 1 and 2. Closed by the IWC exemplar comparison, section 3.3, in favour of the explicit flatten — the option the data-flow brief rated as having no analytical gain. The brief's preference was a reasonable call made with no corpus evidence available; the evidence now exists and points the other way. transcriptomics/rnaseq-pe/rnaseq-pe at corpus fe41a79 puts an explicit `__FLATTEN__` step with `join_identifier: _` between the list:paired reads input and the per-fastq QC tool, and the consuming QC subworkflow declares its input as `collection_type: list` — flat. This is published IWC convention rather than a workaround, and it satisfies the interface's declared `list` shape for outputs 1 and 2 as originally written, so no interface correction is needed. Consequences the test plan must key on: the identifier vocabulary on outputs 1 and 2 doubles to twelve (AR0382_A_forward, AR0382_A_reverse, and so on for each of the six samples), while outputs 3-8 keep the six-element sample identifier space untouched, because the flatten is a side branch off the workflow input and does not enter the map-over region. The data-flow brief's placeholder transformation 5.5 is therefore required, not conditional.

featurecounts-annotation-source-unnamedresolvedfreeform-summary-to-galaxy-interfacepaper-to-test-data
unmet
gene annotation for workflow input `Gene annotation GTF`

The paper names the genome assembly (GCA_002759435.2, C. auris B8441) but never names a GTF/GFF for it. featureCounts requires one and RNA STAR uses one for splice junctions. NCBI RefSeq GFF and FungiDB B8441 GFF differ in gene ID space and attribute keys (`gene_id` vs `ID`), which changes the featureCounts `-g` attribute, every downstream gene identifier, and whether SCF1 appears as `B9J08_001458` at all. FungiDB and CGOB are cited in the paper only for synteny inspection, not as the counting annotation. The interface fixes the datatype as `gtf`; a GFF3 source would require a different datatype or a conversion step.

Largest single gap for reproducing this analysis. Whoever picks a source must record which one, because count values and gene IDs are not comparable across the two. CARRIED FORWARD by advance-galaxy-draft-step at `Count reads per gene`, which is now concrete and STILL DOES NOT CLOSE THIS. Both consumers are wired to the declared `Gene annotation GTF` input (RNA STAR `sjdbGTFfile`, featureCounts `anno|reference_gene_sets` under `anno_select: history`, case 2), so the PORTS are settled. What this step adds to the entry is a second dependent binding: `gff_feature_attribute: gene_id` (featureCounts `-g`), taken from the wrapper default and from corpus transcriptomics/rnaseq-pe/rnaseq-pe at fe41a79. That is correct for a GTF and wrong for a FungiDB B8441 GFF3, which keys on `ID` — and the failure mode is a complete, plausible counts table of the WRONG identifiers, not an error. `gff_feature_type: exon` has the same shape of exposure. So this entry now gates three things, not one: the input's datatype, the `-g` attribute, and the gene ID space every downstream result is addressed in. Naming the source settles all three at once; nothing else will. CLOSED by paper-to-test-data. The annotation is NCBI's own GTF for the exact accession the paper names: `GCA_002759435.2_Cand_auris_B8441_V2_genomic.gtf.gz` under `https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/002/759/435/GCA_002759435.2_Cand_auris_B8441_V2/` (md5 6e5b9528d48c0a8fc2c8588e6eeea929). It was downloaded and inspected, not merely cited. It settles all three things this entry gated, in the direction the workflow already assumes: it is a true GTF (`#gtf-version 2.2`), so the input's `gtf` datatype needs no conversion step; its `gene_id` values ARE the paper's locus tags, so `gff_feature_attribute: gene_id` is correct as bound and SCF1 resolves as `B9J08_001458` with no identifier translation (PEKT02000003.1:864995-867292, + strand, single exon, 2298 bp); and it carries 6057 `exon` features across 5586 distinct genes, so `gff_feature_type: exon` is correct too. The FungiDB-GFF3 hazard this entry described is avoided by not using FungiDB — nothing about that hazard was wrong, it simply does not arise for this source.

fold-change-threshold-linear-vs-deseq2-log2fcresolvedfreeform-summary-to-galaxy-data-flowcompare-against-iwc-exemplar
unmet
unit agreement between workflow input 7 and the DESeq2 result column it thresholds

Workflow input 7 is `Minimum absolute fold change`, default 2.0, in linear units — the form the paper states (|fold change| > 2). DESeq2 reports log2FoldChange. The significance filter must therefore compare abs(log2FoldChange) > log2(threshold), which needs either a conversion the filter expression may not support or a restatement of the parameter. Nothing in the interface brief notes the mismatch.

Consequential if missed rather than merely untidy: comparing abs(log2FoldChange) > 2.0 applies a 4-fold cut and silently fails to reproduce the paper's gene lists, while still producing a plausible-looking filtered table. Options are a log2 conversion node before the filter, a filter expression that computes the conversion inline, or restating input 7 as a log2 threshold (default 1.0) with its label changed — the last changes the interface. Closed by the IWC exemplar comparison, section 2, on the third of those three options. transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting at corpus fe41a79 exposes `log2 fold change threshold` as a float workflow input, default 1.0, documented as "A log2 FC of 3 equals to an absolute fold change of 8 (2^3)", and filters with the raw column and no conversion anywhere in the workflow: `abs(c3)>` concatenated with the connected float. There is no linear fold-change parameter, no log2() call and no conversion node in the corpus exemplar. So: restate interface input 7 as a log2 threshold with default 1.0 — which is exactly the paper's |fold change| > 2 — and change its label; teach the conversion in the doc string as the corpus does. Three implementation details come with the answer. (a) Column indices are c3 for log2FC and c7 for adjusted p-value; deg_annotate appends columns 8-13 and leaves 1-7 untouched, so the indices hold whether or not an annotation step is added. (b) Filter1's predicate is a text parameter, so a numeric workflow input cannot reach it directly: the corpus idiom is one iuc/compose_text_param step per threshold concatenating a literal prefix (`c7<`, `abs(c3)>`) with the connected float, making node I two Galaxy steps per filter rather than one. (c) The two thresholds are two chained Filter1 steps, p-adj then log2FC, each with header_lines "1", not one compound predicate. Restating input 7 changes the interface brief's section 2.3 and its output-3 table entry; that edit must be made there, not silently in the template.

rnaseq-strandedness-inferred-from-kit-nameresolvedfreeform-summary-to-galaxy-interfacepaper-to-test-data
unmet
value for workflow parameter `featureCounts strandedness`

The paper states only the library kit (Illumina Stranded Total RNA Prep with Ribo-Zero Plus). Reverse-stranded (dUTP) is inferred from that kit name and is not stated anywhere in the supplement. The interface exposes the parameter with default `reverse` so the inference is visible and changeable rather than buried in a step default.

Empirically checkable without new information: the promoted output `featureCounts assignment summary` shows a large Unassigned_NoFeatures fraction when the setting is wrong. A test-plan or run phase can close this from evidence. CLOSED by paper-to-test-data, empirically rather than by argument. The entry itself named the check; it was performed a step earlier than expected, on alignments rather than on a featureCounts summary. 200,000 read pairs from each of AR0382_A, AR0387_A and tnSWI1_A were aligned to GCA_002759435.2 with bwa-mem and each R1 that unambiguously overlapped a single annotated gene was compared against that gene's strand. 98.4% / 98.4% / 98.2% map ANTISENSE. That is the dUTP reverse-stranded signature and it is not a close call. `stranded - reverse` → featureCounts `-s 2` is correct; the value is now measured, not inferred from the kit name. The `featureCounts assignment summary` check this entry proposed remains valid as a regression guard and is carried into the test plan as an assertion.

sample-sheet-condition-to-deseq2-factor-wiringresolvedfreeform-summary-to-galaxy-interfacefreeform-summary-to-galaxy-data-flow
unmet
path from per-sample `condition` metadata to DESeq2's factor-level inputs

The interface carries condition and replicate as `column_definitions` on a `sample_sheet:paired` reads input. Galaxy does not propagate `column_definitions` or per-row `columns` through map-over, so by the time featureCounts has produced counts the condition metadata is gone and an explicit step must reattach it (`__SAMPLE_SHEET_TO_TABULAR__` plus a filter/split, or the rules DSL) before DESeq2 can receive one multi-data input per factor level.

Named fallback if the wiring proves unbuildable: replace the sample-sheet input with a `list:paired` reads collection plus a `data` input `Sample metadata table` (tabular: sample_id, condition, replicate) and split on that table. Taking the fallback changes workflow input 1 and must be reflected back into the interface brief, not applied silently in the template. Closed by the data-flow brief, section 4. The condition reaches DESeq2 by an identifier-keyed split taken off the workflow input, where the column metadata still lives: `__SAMPLE_SHEET_TO_TABULAR__` projects the sample sheet to a tabular (element identifier, condition, replicate); a row filter plus column projection yields one identifier list per factor level; `__FILTER_FROM_FILE__` filters the featureCounts collection to each level's identifiers; each per-level counts sub-collection reduces into one DESeq2 multi-data factor port. The map-over region and every promoted per-sample output are untouched, and the join key is the element identifier, which Galaxy preserves across map-over. The route is independent of how the DESeq2 contrasts are realized (`deseq2-contrast-realization-unsettled`) and survives the named fallback intact: under the fallback the tabular node simply disappears and the user-supplied metadata table lands in its place, with the filter and collection-split nodes unchanged. Two narrower successors carry what remains: `sample-sheet-to-tabular-identifier-column-unverified` (is the element identifier emitted as a column) and `deseq2-factor-level-names-not-parameterized` (where the per-level literal comes from). Rejected alternatives and why are recorded in the brief's section 4.6 — notably filtering by element-identifier regex, which is unsafe here because `AR0382_tnSWI1_A` contains the reference level's own identifier as a substring.

sample-sheet-input-test-fixture-expressibilityresolvedfreeform-summary-to-galaxy-interfacepaper-to-test-data
unmet
test-fixture form for the `sample_sheet:paired` workflow input

Whether a `sample_sheet`-family workflow input — element identifiers plus per-row typed `columns` and collection-level `column_definitions` — can be expressed in a Planemo/IWC `-tests.yml` job block. If it cannot, the workflow's primary input is untestable as designed and the fallback in `sample-sheet-condition-to-deseq2-factor-wiring` becomes mandatory.

Not verified in this phase; no fixture syntax for sample-sheet inputs appears in the references packaged with this Mold. Naturally closed by the test-plan phase. CLOSED by paper-to-test-data: YES, it is expressible, and the `list:paired` fallback is not required. Galaxy's job-block loader has an explicit branch for it — `lib/galaxy/tool_util/cwl/util.py`, `replacement_collection()`: `if collection_type.startswith("sample_sheet"): kwds["rows"] = value.get("rows")`, carried to the collections API by `lib/galaxy/tool_util/client/staging.py`. The nested `sample_sheet:paired` shape specifically is covered by `test/unit/tool_util/test_cwl_util.py::test_galactic_job_json_sample_sheet_paired_collection`, and an end-to-end worked example ships as `lib/galaxy_test/workflow/collection_semantics_cat_sample_sheet.gxwf-tests.yml`. Two things a test author must get right. (1) `rows` is a mapping of element identifier to a POSITIONAL LIST of column values ordered to match `column_definitions` — authority is `validate_row()` in `lib/galaxy/model/dataset_collections/types/sample_sheet_util.py`, which rejects on `len(row) != len(column_definitions)` then `zip(row, column_definitions)`. The dict form appearing in Galaxy's own non-paired unit tests never reaches that validator and will not work. (2) Neither `galactic_job_json` nor `staging.py` passes `column_definitions` when creating the collection, although the API payload supports it, so a test-staged sample sheet has collection-level `column_definitions: None`. This is NOT fatal: per-element `columns` are still populated from `rows`, and the two `column_definitions_compatible()` call sites are both in `DataCollectionToolParameter` option-building (UI dropdown filtering), which a test bypasses by supplying the HDCA by id. The one real consequence is that `__SAMPLE_SHEET_TO_TABULAR__` emits its header line only `#if $include_headers and $input.collection.column_definitions` — harmless here because `Project sample sheet to tabular` sets `include_headers: false`, but a later phase that flips it to true will see the header silently vanish under test while it appears in the UI. Concrete job block is in `test-data-refs.json` under `planemo_test_job_block`.

sample-sheet-to-tabular-identifier-column-unverifiedresolvedfreeform-summary-to-galaxy-data-flowadvance-galaxy-draft-step
unmet
confirmation that the sample-sheet-to-tabular bridge emits the element identifier

The settled condition wiring joins the sample metadata table to the featureCounts collection on element identifier, so node E's output must carry that identifier as a column. The packaged note galaxy-sample-sheet-collections documents `__SAMPLE_SHEET_TO_TABULAR__` only as iterating elements and tab-joining "for downstream tabular consumers"; it does not state the output column set or ordering, and no other packaged reference covers it. Galaxy source was not readable from inside this run.

RESOLVED, affirmatively, by the tool summary itself. `__SAMPLE_SHEET_TO_TABULAR__` v1.0.0 caches cleanly from the Tool Shed API by bare id (unlike `__FLATTEN__`, whose collection output defeats gxwf's summary decoder), and its packaged help states the contract directly: "The first column is always the element identifier (sample name). The remaining columns match the metadata fields defined in the sample sheet." With the optional `include_headers` enabled the first header cell is literally `element_identifier`. So the identifier IS emitted, as column 1, and the ordering this region assumed throughout -- (element identifier, condition, replicate), metadata in column_definitions order -- is confirmed. All six provisional column bindings downstream (three Filter1 predicates on c2, three Cut1 projections of c1) stand unchanged; the Apply Rules substitute node named in the data-flow brief section 4.2 is not needed and was not built. The step is pinned with `include_headers: false`, which also settles the three row filters' `header_lines` at 0 -- that binding is no longer blocked on this entry. Evidence is the wrapper's own documented contract, not a corpus exemplar; `no-iwc-precedent-for-sample-sheet-workflow-input` is unaffected and stays open.

star-history-reference-wiring-has-no-corpus-precedentresolvedcompare-against-iwc-exemplaradvance-galaxy-draft-step
unmet
worked wiring for the RNA STAR history-reference conditional branch

Every RNA STAR step in the corpus at fe41a79 — there are exactly two, in transcriptomics/rnaseq-pe/rnaseq-pe and transcriptomics/rnaseq-sr/rnaseq-sr — uses `refGenomeSource.geneSource: indexed`, selecting a built-in `genomeDir` through a `restrictOnConnections: true` string parameter with `sjdbGTFfile` supplied from the history. Their test jobs pass a plain genome string (`Reference genome: sacCer3`). The interface settles input 2 as a history FASTA, which is the right call for C. auris B8441 since no public server indexes it, but that is a different `__current_case__` in the iuc/rgrnastar wrapper with a different set of required sub-parameters, and the corpus has no example of it.

RESOLVED, affirmatively, from the wrapper rather than from the corpus — which is what the original note asked for. iuc/rgrnastar/rna_star @ 2.7.11b+galaxy1 caches and summarizes cleanly (all ten of its outputs are plain `data`, so the collection-output decode failure that blocks `__FLATTEN__` and lparsons/cutadapt does not apply here), and its schema answers every part of the question. `refGenomeSource.geneSource` publishes exactly two options, `indexed` and `history`; the HISTORY branch is real, is `__current_case__: 1`, and carries `genomeFastaFiles` (gx_data, formats fasta/fasta.gz, optional false), `genomeSAindexNbases` (integer, min 2 max 16, default 14), its own TWO-case `GTFconditional`, and `diploidconditional` (Yes=0 / No=1, default No). So the two TODO ports resolve to `refGenomeSource|genomeFastaFiles` and `refGenomeSource|GTFconditional|sjdbGTFfile`; the exemplar's `refGenomeSource|GTFconditional|genomeDir` exists only under `indexed` and is correctly absent. Under history, GTFconditional's option order (without-gtf, with-gtf) is the REVERSE of its `<when>` order (with-gtf, without-gtf), so `with-gtf` is case 0 — a case index that cannot be read off the dropdown. That the case index follows `<when>` document order is not assumed: `gxwf convert --to format2` over transcriptomics/rnaseq-pe/rnaseq-pe.ga at fe41a79 emits `geneSource: indexed` with `__current_case__: 0` and `GTFselect: without-gtf-with-gtf` with `__current_case__: 1`, matching the indexed branch's `<when>` order (with-gtf, without-gtf-with-gtf, without-gtf) and not its option order. Cross-read against rg_rnaStar.xml and macros.xml at tools-iuc main, which carries @TOOL_VERSION@ 2.7.11b / @VERSION_SUFFIX@ 1 — this exact version. Two consequences worth carrying: the history branch builds its index in-job via `STAR --runMode genomeGenerate` into tempstargenomedir, confirming the data-flow brief's no-separate-index-node call; and `output_log` and `mapped_reads` carry no `<filter>` whatsoever in `<outputs>`, so the promoted `STAR mapping summary` cannot be suppressed by any parameter choice here. The step now validates for real against the cache (`12 ok, 0 fail, 2 skip`), not by skip. `reference-genome-delivery-shape-unverified` is untouched and STAYS OPEN — this entry establishes that the history route WORKS, never that it is preferable to a built-in index nobody checked for. `featurecounts-annotation-source-unnamed` also stays open: the GTF PORT is now wired, the GTF SOURCE is still unnamed by the paper.

trimmed-reads-paired-reassembly-conditionalresolvedfreeform-summary-to-galaxy-data-flowcompare-against-iwc-exemplar
unmet
the output collection shape of the paired-aware trimming step

The interface declares `Trimmed reads` as `list:paired`. Whether the trimmer emits one paired-inner collection per sample or two parallel single-ended collections (R1, R2) is wrapper-dependent and was not resolvable in this phase, which pins no Tool Shed tools. If it is the latter, the design needs a re-pair node between trimming and alignment, or the alignment step must take two parallel collections in dot-product.

Conditional shape repair, not method: recorded in the data-flow brief as placeholder transformation 5.6 so the template does not assume one reading. Naturally closed by the IWC exemplar comparison or by tool discovery on the trimmer. Closed by the IWC exemplar comparison, section 3.4: one paired-inner collection per sample, no re-pair node. Neither transcriptomics exemplar uses Cutadapt — both use fastp — so the evidence comes from epigenetics/cutandrun/cutandrun at corpus fe41a79, cited for the wrapper's IO shape only and for nothing else, since CUT&RUN is a different domain. toolshed.g2.bx.psu.edu/repos/lparsons/cutadapt/cutadapt/5.2+galaxy2 driven from a list:paired collection with `library.type: paired_collection` declares outputs `out_pairs` (type `input`, i.e. the input collection's own shape) and `report` — one paired collection plus one report per element. Placeholder transformation 5.6 is therefore not needed and node C takes a single collection input, as the data-flow brief's primary reading assumed. The finding is consistent across both wrappers that could fill the trimmer slot: transcriptomics/rnaseq-pe/rnaseq-pe's fastp emits `output_paired_coll` the same way. The residual is version-scoped rather than structural — a future Cutadapt wrapper could change its output set, so the per-step loop should confirm the output name against whatever version it pins.

Foundry feedback

what the run showed to be wrong with the Foundry

61 entries about the Foundry's own assets. Triage them with the report-foundry-run-feedback skill rather than filing from here.

iwc-corpus-path-hard-coded-with-no-overrideblockerdefectcompare-against-iwc-exemplar

The procedure's first and only prerequisite step is "Clone or pull and merge the IWC corpus (https://github.com/galaxyproject/iwc) to `~/.foundry/iwc`". That path is hard-coded, and the Mold names no environment variable, no configuration key, and no fallback for a machine where it cannot be created. On this machine it cannot: `mkdir ~/.foundry` fails with `Operation not permitted`, with and without the Bash sandbox. The Mold is the corpus-first check of every Galaxy-targeting pipeline, so an unsatisfiable corpus location makes the whole phase unrunnable, and the only reason this run produced a comparison at all is that the harness supplied a clone at a different path out of band. Nothing in the bundle describes that route, and an unattended run has no way to discover it.

expected
Read the corpus location from an environment variable with `~/.foundry/iwc` as the default — the same shape `OMC_STATE_DIR` already has elsewhere in this toolchain — and state in the procedure that a caller-supplied corpus path is honoured and pinned rather than pulled. Two smaller corrections belong with it. State that when the corpus is supplied rather than cloned, the Mold must not `git pull` or write inside it, since a caller-owned checkout may be shared or read-only. And require the corpus commit to be recorded in `iwc-comparison-notes` as provenance in every case, not only the supplied one: a structural comparison is a claim about a moving corpus, and without the commit no later reader can tell whether a divergence this Mold reported has since been closed upstream. The output artifact description says nothing about provenance today.
evidence
`mkdir ~/.foundry` returned `Operation not permitted` on darwin 25.5.0 under both the sandboxed and unsandboxed Bash tool. The phase ran against a caller-supplied shallow clone pinned at galaxyproject/iwc `fe41a79`, recorded in <run>/iwc-comparison-notes.md under "Corpus provenance" on this phase's own initiative, since no packaged instruction asks for it.
raised by
compare-against-iwc-exemplar (phase 4)
observed at
63a3f9cf9c97 rev 10
issue
not filed
a-required-port-missing-from-in-carries-no-sentinel-and-no-gatemajorgapgalaxy-workflow-draft-format

A `TODO_<port>` sentinel marks a port whose NAME is unresolved, but there is no way to mark a port that is simply ABSENT from `in:` -- and an absent port is invisible to `gxwf draft-next-step`, whose `work` list enumerates sentinels and `_plan_*` fields and so reports nothing at all. The obligation survives only as prose. Distinct from `draft-format-sentinel-hint-cannot-express-a-conditional-nested-port`, which is about a sentinel whose hint has the wrong SHAPE; here there is no sentinel to misread.

expected
Require in the note that every port the step's plan calls for appears in `in:` -- as a `TODO_<port>` sentinel while unresolved -- so that "the plan named it" and "the work list reports it" cannot come apart. A cheap partial check for `draft-next-step`: when a step's `_plan_*` prose names a sibling step as its model, flag an `in:` key set that is a strict subset of that sibling's.
evidence
`Differential expression: second contrast vs reference` declared three `in:` keys; the already-concrete sibling it was told to copy declared five. The two missing ones were `select_data|rep_factorName_0|rep_factorLevel_{0,1}|factorLevel`. `draft-next-step` reported only `TODO[tool_version]` plus four `_plan_*` strings; the sole trace of the two missing connections was the sentence "As for the first contrast, with `Counts for the second contrast level` at level index 0". Left unwired, both factor levels would fall back to the wrapper's empty default -- blank contrast titles on the diagnostic plots, and the silent re-opening of resolved entry `deseq2-factor-level-names-not-parameterized`, whose closure asserts that no level literal is baked into any step. Nothing would have failed: this run already measured that a disconnected input validates byte-identically green, and this step reports `skip_tool_not_found` besides.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
advance-draft-mold-cites-draft-format-note-it-does-not-packagemajorgapadvance-galaxy-draft-step

The procedure's resolve step branches on the draft tiers and sends the reader to `galaxy-workflow-draft-format` for them, but the Mold does not package that note: the bundle's references are three gxwf CLI pages, `galaxy-tool-job-failure-reference`, `open-requirements-ledger`, and the tool-summary and draft JSON Schemas. The Runtime Notes then forbid reading Foundry source files at runtime, so the one document the procedure names for the distinction it asks the iteration to make is unreachable from inside the bundle. The draft JSON Schema does not carry the tiers — they are prose.

expected
Package `content/research/galaxy-workflow-draft-format/index.md` in this Mold's references, as `freeform-summary-to-galaxy-template` already does. It is the note this Mold's central branch depends on, and the per-step loop reads a draft written against it on every iteration. If the intent is that the tiers be inferable from the draft alone, say that instead and drop the cross-reference.
evidence
Iteration 3 of this run, step `Build log2 fold-change predicate`. The tier vocabulary was taken from the draft's own `doc:` strings (`Tier: Identity-pinned.` / `Tier: Resolved.`, a convention iteration 1 introduced) rather than from any packaged contract, and the same gap left the fate of the step's `_plan_context` provenance undecided — see `advance-draft-mold-silent-on-plan-provenance-at-concretion`.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
advance-draft-mold-needs-a-tool-cache-it-never-declaresmajorgapadvance-galaxy-draft-step

The procedure names `galaxy-tool-cache list` as the way to resolve a stock tool's version, and the per-step validator it mandates only produces a real verdict when given `--cache-dir` with the step's tool cached. Neither the cache nor the `galaxy-tool-cache` binary appears anywhere in the bundle's contract: `_required_tools.json` declares `gxwf` alone, derived from the three `gxwf` commands the Mold cites, and no step of the procedure says to populate a cache. Without it `draft-validate --concrete` reports `skip_tool_not_found` for every step and the iteration's green is vacuous.

expected
Declare `galaxy-tool-cache` as a required tool of this Mold (it ships in the same `@galaxy-tool-util/cli` package, so it costs no new install), and make cache population an explicit move in the procedure — `galaxy-tool-cache add <tool_id> --tool-version <v>` after wrapper resolution, with the resulting `--cache-dir` passed to `draft-validate --concrete`. Citing the command as a CLI reference rather than in prose would also let the required-tools deriver pick it up, which is why it is missing today.
evidence
Iteration 1 of this run. `galaxy-tool-cache` is not on PATH after the documented `npm install -g @galaxy-tool-util/cli` (it exists in the package's bin directory but only `gxwf` was linked), and its default cache root `~/.galaxy/tool_info_cache` is not writable in this sandbox, so both had to be worked around before any concrete verdict was possible. Until the cache was populated by hand, `draft-validate --concrete` reported `Tool state: 0 ok, 0 fail, 1 skip` — which the procedure's "on green, return" would have accepted.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
advance-draft-mold-treats-a-skipped-tool-state-as-greenmajorgapadvance-galaxy-draft-step

The validate step is written as a two-valued gate — "On green, return; on red, route per the failure-routing rules" — but `draft-validate --concrete` reports three tool-state outcomes, `ok`, `fail`, and `skip`. A step reported `skip` had its tool state checked against nothing at all, yet the run prints `Concrete: OK` and the procedure's green branch accepts it. The Mold never says which of the two branches a skip belongs to, that a skipped step's state is unvalidated, or that some skips are permanent and cannot be cleared by priming the cache. An iteration reading only this Mold would return green on a draft whose collection-plumbing steps have never been validated, and no later phase would know which steps those were.

expected
Make the accept criterion three-valued in the validate step: `fail` routes per the existing rules; `skip` is neither green nor red but an explicit hole — say that the iteration must name each skipped step in its report and hand-review that step's `state` against the wrapper, because nothing else will. Distinguish a skip that cache priming would clear (tool simply not yet added) from one that priming cannot clear (the tool cannot be decoded into the cache at all), and say that repeatedly re-priming the latter is wasted work. Carry the list of skipped steps into the loop endstate so terminal validation inherits it rather than rediscovering it.
evidence
Iteration 4 of this run. `gxwf draft-validate <run>/galaxy-workflow-draft.gxwf.yml --concrete --cache-dir <cache>` returned `Concrete: OK` with `Tool state: 4 ok, 0 fail, 1 skip`, the skip being `0 (__FLATTEN__) [skip_tool_not_found]`. That step's `state` has gone unvalidated for four consecutive iterations and is accepted as green every time. Only a convention carried in this run's own harness brief — not anything in the Mold bundle — told the iteration to treat the skip as a permanent hole needing review by eye rather than as a cache-priming task to retry.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
an-interrupted-phase-leaves-its-declared-artifact-written-but-unvalidatedmajorgapfoundry-feedback-ledger

A Mold that writes one large artifact and validates it at the end has a window in which the artifact exists on disk and nothing records whether it conforms to its schema. This run entered that window: phase 8's session was terminated after writing a 67 KB `galaxy-test-plan.yml` and before running `foundry validate-galaxy-workflow-test-plan`, and before its ledger pass. The run-lifecycle section addresses only run status — it says a hard interruption leaves the run `running`, "which is distinguishable from success without a recovery write" — and says nothing about the phase's declared OUTPUT. So `status: running` on a phase is ambiguous in a way that matters: the artifact may be absent, present and valid, present and invalid, or present and half-written, and the four look identical from the ledger. A resuming agent, or a next phase that simply reads the artifact because it is there, has nothing telling it the verify step never ran.

expected
State in the run-lifecycle section that a phase left `running` may have written its declared output artifacts in an UNVALIDATED state, and that whoever resumes or succeeds it must re-run that phase's declared validation before treating those artifacts as input. Stronger: give the phase row a field the skill writes when its own validation passes — `artifacts_validated: true`, the exact analogue of `feedback_checked` and justified by the same argument the note already makes for it, that a run which looked and found nothing must be distinguishable from one that never looked.
evidence
The terminated phase-8 session left `<run>/galaxy-test-plan.yml` complete and, as it turned out, schema-valid, but nothing on disk said so; the finishing session had to re-run the validator to find out. The same interruption also left the artifact citing a feedback entry id (`test-plan-mold-ignores-available-concrete-workflow`) that did not exist in this ledger, because the ledger pass is likewise an end-of-phase step — the artifact and the ledger were inconsistent with each other and only the ledger's `running` status hinted at it.
raised by
freeform-summary-to-galaxy-test-plan (phase 8)
observed at
63a3f9cf9c97 rev 2
issue
not filed
bounded-subgraph-artifact-has-no-stated-validity-contractmajorgapcompare-against-iwc-exemplar

The artifact is described as a "Cleaned gxformat2 conversion (via convert --to format2 --compact) of the nearest IWC exemplar's relevant subgraph", and separately as "bounded to the relevant subgraph, not the whole workflow". Those two requirements cannot both hold literally. `convert` emits the whole workflow; bounding it to a subgraph is hand surgery that necessarily leaves dangling `source:` references to removed steps and outputs, so the result is well-formed YAML but not a loadable gxformat2 workflow. The Mold never says which property matters, and the downstream consumer is named only as something that "pattern-matches against" the file — which does not distinguish a human-read reference from a parsed one. This run resolved it by treating the artifact as a reading aid, eliding non-structural tool_state and marking every elision inline, but a template Mold that tried to load the file would fail, and nothing warned it.

expected
State the contract in the "Nearest exemplar (gxformat2) view" section: the artifact is a bounded reading aid, not a runnable or validatable workflow; dangling references to elided steps are expected; and elisions must be marked in place so a reader can tell a bound from an absence. If instead the file is meant to stay loadable, say that and specify how — keep every transitively referenced step, or rewrite dropped sources to workflow inputs. Either answer is workable; the absence of both means each run invents its own and the downstream Mold cannot rely on either. A one-line statement in the artifact description would settle it, since that is where a consumer looks.
evidence
<run>/iwc-exemplar.gxwf.yml drops the visualization tail of one exemplar and roughly two-thirds of the other, leaving references to steps that are no longer present — for example the second document's `outputs[Counts Table].outputSource: _unlabeled_step_25/...`, whose producing step was dropped — and requiring several step `in:` entries to be commented out rather than resolved. It parses as three valid YAML documents and would not load as gxformat2. The convert CLI reference packaged in this bundle documents no subsetting option, so the bounding is necessarily out-of-band.
raised by
compare-against-iwc-exemplar (phase 4)
observed at
63a3f9cf9c97 rev 10
issue
not filed
collection-output-decode-is-a-flat-vs-nested-shape-mismatch-not-a-missing-fieldmajordefect@galaxy-tool-util/cli — ParsedTool collection-output decoding vs the Tool Shed tool API

This CORRECTS the root cause and the scope recorded in `gxwf-cannot-decode-collection-outputs-of-builtin-collection-operations`, which is filed against the same decode failure. Two things in that entry are wrong. (1) The collection output does NOT "carry no `structure`". The Tool Shed serializes a collection output FLAT — `collection_type`, `collection_type_source`, `collection_type_from_rules`, `structured_like` and `discover_datasets` all sit at the top level of the output object — while the decoder expects exactly those five fields nested inside a `structure` object. Every field the decoder wants is present in the payload; only the nesting differs. The correction that entry proposes — make `structure` optional, or default it — would therefore decode cutadapt's `out_pairs` with a NULL collection type when the API plainly said `"collection_type": "paired"`. That is worse than the current loud failure: downstream shape reasoning would silently lose the one fact the output exists to carry. (2) It is not a built-in phenomenon. The trigger is a `<collection>` output, wherever it occurs. `toolshed.g2.bx.psu.edu/repos/lparsons/cutadapt/cutadapt` — a mainstream IUC-maintained wrapper pinned by eight workflows in IWC at fe41a79 — fails identically at every version tried, because its paired-collection branch declares `<collection name="out_pairs" type="paired">`. So the blast radius is not "the collection-operation built-ins a draft uses for plumbing"; it is every tool with a collection output, which includes a large share of the wrappers real Galaxy workflows are built from. Any such step is permanently unvalidatable by `draft-validate --concrete`, and no cache priming can help.

expected
Decode the flat form the Tool Shed actually emits: read `collection_type`, `collection_type_source`, `collection_type_from_rules`, `structured_like` and `discover_datasets` from the output object itself when no `structure` key is present, and populate `structure` from them. Do not make `structure` optional or default it to nulls — that discards `collection_type`, which is the only thing a caller needs the branch for. If the nested form is the intended contract, then the Tool Shed's `/api/tools/<trs-id>/versions/<v>` serializer is the side that must change, and the decoder should say which shape it received rather than printing the expected type.
evidence
Iteration 8 of phase 6 in this run, gxwf/galaxy-tool-cache 1.10.1, resolving the Cutadapt wrapper for step `Quality-trim reads`. `galaxy-tool-cache add toolshed.g2.bx.psu.edu/repos/lparsons/cutadapt/cutadapt --tool-version 5.2+galaxy2` fails with `toolshed fetch failed ... for lparsons~cutadapt~cutadapt` and the same `["outputs"][0] ... ["structure"] is missing` dump as `__FLATTEN__`; 5.2+galaxy0, 4.9+galaxy1 and 3.7+galaxy0 fail identically, so it is not version-specific. Fetching the same payload directly — `GET https://toolshed.g2.bx.psu.edu/api/tools/lparsons~cutadapt~cutadapt/versions/5.2+galaxy2`, 200 — shows `outputs[0]` is `{"name": "out_pairs", "type": "collection", "collection_type": "paired", "collection_type_source": null, "collection_type_from_rules": null, "structured_like": null, "discover_datasets": ..., "hidden": ..., "label": ...}` with no `structure` key and nothing missing from it. `outputs[1]` (`split_output`) has the same shape; the thirteen `data` outputs decode fine. The eight IWC workflows at fe41a79 that pin this wrapper all pin 5.2+galaxy2 at changeset f6168dd17f82. Net effect on this run: `draft-validate --concrete` reports `7 ok, 0 fail, 2 skip`, both skips being collection-output tools, and the Cutadapt step's tool state had to be bound by hand against the wrapper XML and a `gxwf convert --to format2` round-trip of a corpus `.ga`.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
data-flow-contract-silent-on-contradicting-the-interface-briefmajorgapgalaxy-data-flow-draft-contract

The contract partitions ownership between data-flow, template and step implementation, and the Mold declares the interface brief as an input that "pins inputs, outputs, and labels". Neither says what to do when the data-flow analysis shows a pinned interface decision is not achievable under Galaxy semantics. That happened three times in this run: FastQC mapped over a sample_sheet:paired input fans out per read direction, so two outputs the interface declared as flat lists are nested; the outer axis of every mapped output is sample_sheet- shaped, not the declared `list`; and a linear fold-change parameter is declared against a DESeq2 column reported in log2. Whether the data-flow brief may correct the interface, must defer to it, or must only record the conflict is not stated anywhere in the bundle. The route taken here — ledger each one and state the correction in the brief — was invented.

expected
State the rule in the contract's Boundary section: the data-flow draft may contradict an interface decision when Galaxy collection or map-over semantics make it unachievable, and must record each contradiction as an open-requirements entry naming the interface decision it invalidates, so the interface brief and the template cannot silently diverge. Worth pairing with a note in the open-requirements research page, whose `supersedes` mechanism covers a refuted justification on an existing entry but has nothing for a refuted claim in a sibling brief that was never ledgered.
evidence
Interface brief section 3 declares outputs 1-8 as `list`/`list:paired` and input 7 as a linear fold change; the data-flow brief section 7 shows all three are unachievable as written and raises `fastqc-per-read-fanout-not-a-flat-list`, `mapped-outputs-carry-sample-sheet-outer-axis` and `fold-change-threshold-linear-vs-deseq2-log2fc` in the open-requirements ledger.
raised by
freeform-summary-to-galaxy-data-flow (phase 3)
observed at
63a3f9cf9c97 rev 3
issue
not filed
data-flow-mold-packages-pattern-mocs-without-their-recipe-pagesmajorgapfreeform-summary-to-galaxy-data-flow

All four packaged pattern references are MOC index pages. Each is a list of wiki-links to operation and recipe pages — sync-collections-by-identifier, collection-cleanup-after- mapover-failure, tabular-to-collection-by-row, tabular-filter-by-column-value and roughly thirty others — and none of those pages is in the bundle, while the Mold's runtime notes forbid reading Foundry source at runtime. The collection MOC states its own role plainly: "the operation and recipe pages are the actionable references". So the Mold packages the index and withholds the content it indexes. Every collection-idiom choice in this run's brief was made from a one-line MOC description, with no corpus-observed recipe available to check the shape, the built-in tool id, or the failure modes against.

expected
Package the recipe and operation pages the MOCs name, or at least the subset a data-flow brief can act on (the Cleanup, Identifiers, Structural Reshape and Bridges sections of galaxy-collection-patterns and the Bridges section of galaxy-tabular-patterns). If the full leaf set is too large for a bundle, say so in the Mold and state that MOC entries are naming hints only, so the brief records idiom selections as unverified rather than as corpus-grounded. Applies identically to the Nextflow and CWL data-flow Molds, which package the same MOCs. The refs manifest marks these `evidence: corpus-observed`, which as packaged is true of the map but not of anything the runtime can actually read.
evidence
Bundle references/patterns holds exactly four files, all `pattern_kind: moc`. The brief at <run>/freeform-galaxy-data-flow.md selects the identifier-sync idiom for its central decision and has to state in its confidence section that the selection rests on a one-line index entry.
raised by
freeform-summary-to-galaxy-data-flow (phase 3)
observed at
63a3f9cf9c97 rev 3
issue
not filed
draft-extract-destroys-yaml-commentsmajorgapgxwf draft-extract

`gxwf draft-extract` re-serializes the workflow from a parsed data model, so every YAML `#` comment in the draft is destroyed. The note describes the command as three subtractive operations — drop drafty steps, strip `_plan_*`, promote `class` — and its Output and Gotchas sections say nothing about comments. Nothing in this Mold's bundle does either. The loss is total and silent: it is not reported in `--report-json` (which counts only dropped steps, dropped outputs and rewritten inputs) and the extracted file validates clean, so no gate in the pipeline can see it. It matters because a per-step draft loop has nowhere else to put the reasoning: `_plan_*` fields MUST be deleted at concretion or `draft-next-step` re-selects the step forever, and the gxformat2 `doc:` field is user-facing prose, not a place for a corpus citation or a cross-step warning. A run that parks that material in comments above each step — as this one did, and as the concrete steps of a template naturally invite — loses all of it at the exact moment the artifact becomes the one downstream Molds consume. The two surfaces that DO survive are `doc:` and the gxformat2 `comments:` block (frames/markdown), and the note names neither as the place to put anything load-bearing.

expected
State in the note's Output section that `draft-extract` is a re-serialization, not a textual edit, and that YAML comments do not survive it; add it to Gotchas next to the existing "this is a transformation, not a validator" warning. Name the two surfaces that do survive — step `doc:` and the gxformat2 `comments:` block — and say that anything a later reader must not lose belongs there before the loop reaches endstate. Ideally `--report-json` should also count discarded comment lines, so the loss is at least visible in the sidecar; failing that, the note is the only place a run can learn it, and it has to say so. (The upstream fix — a comment-preserving round-trip — is a separate, larger ask; the note must describe today's behaviour either way.)
evidence
Iteration 26 of this run, at loop endstate, gxwf 1.10.1. `gxwf draft-extract <run>/galaxy-workflow-draft.gxwf.yml -o <run>/galaxy-workflow.gxwf.yml --report-json <report>` exited 0 with `0 steps dropped, 0 outputs dropped, 0 input rewrites; class_after=GalaxyWorkflow`. Parsing both files and comparing: `steps`, `inputs`, `outputs` and the `comments:` block are deep-equal, and the only structural difference is `class`. The textual difference is 850 comment lines / 62,565 characters present in the draft and 0 in the extract — 106,359 bytes down to 41,862, about 59% of the file. What was in them: the per-step corpus citations this loop moved out of `_plan_*` at concretion, and the two cross-step invariant warnings this run established by hand — that `lfc_shrinkage_type: none` on both DESeq2 nodes is what makes the `c7<` and `abs(c3)>` predicates in four other steps refer to the right columns, and that `header_lines: '0'` on the four significance filters must NOT be the corpus's `'1'`. Neither invariant is expressible in the draft schema (see `draft-format-cannot-express-a-cross-step-invariant`) and neither is checked by any gate, so the comment was the only record, and the extract is where it stopped existing.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
draft-format-cannot-express-a-cross-step-invariantmajorgapgalaxy-workflow-draft-format

Every annotation the draft format offers is scoped to ONE step (`TODO_*`, `_plan_*`, the tier vocabulary), so an invariant that binds two or more steps has nowhere to live except prose inside one of them -- and `advance-galaxy-draft-step` advances exactly one step per invocation, so it is also structurally unable to check one. The step's own `_plan_state` said it outright: "Bind both steps together so the factor name, level ordering, header flag and output_selector cannot drift apart between the two contrasts." That is an instruction with no mechanism behind it. The author is the only enforcement, and only if they happen to read a sibling step's plan prose while implementing a different step.

expected
Give the format a typed, workflow-level way to declare that named steps must agree on named state paths -- a sibling of the `comments:` frame, or a `_plan_invariant` block naming the step ids and the `|`-qualified paths -- and have `draft-validate` check it once every named step is concrete. It is a cheap check: the paths are already addressable, and this run's loop concretized the two coupled steps 1 iteration apart. Failing that, state in the note that cross-step couplings are out of scope for the draft and must be carried as open-requirements entries, so a Mold run stops inventing its own mechanism.
evidence
Two couplings in one workflow, both invisible to every gate. (1) The two DESeq2 nodes must stay key-for-key identical apart from one counts collection and one factor level; this was settled only by hand-diffing the two steps' parsed `in:`/`out:`/`tool_state` trees (5 `in:` keys, 3 `out:` ids, 27 state leaves). (2) Worse, `advanced_options.lfc_shrinkage_type: none` on BOTH DESeq2 nodes determines the result table's column count, and two DIFFERENT steps bake the resulting indices as opaque string literals -- `"c7<"` and `"abs(c3)>"` in two `compose_text_param` bridges. Any other shrinkage value drops the `stat` column, moves padj from c7 to c6, and `"c7<"` then names a column that does not exist. A three-step coupling expressed as a magic number in a text field. Both DESeq2 steps additionally report `skip_tool_not_found` (their `split_output` collection output fails the cache decode), so not even the tool-state gate looks at them.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
draft-format-note-example-writes-format-as-a-scalarmajordefectgalaxy-workflow-draft-format

The note's "Example (sketch)" declares a workflow input as `format: fastqsanger.gz` — a bare scalar. The draft schema the same Mold packages types `format` as `null | ReadonlyArray<string>`, so the sketch is not valid against the contract it illustrates. An author who copies the sketch, as it invites, gets a structure error.

expected
Write `format:` as a list in the sketch (`format: [fastqsanger.gz]` or the block form), and say in the relaxations section that `format` is a list even for a single value. If the scalar form is in fact accepted by some other consumer, say which and why the draft schema rejects it.
evidence
The draft was authored with `format: fastqsanger.gz`, `format: fasta` and `format: gtf`, all copied in form from the note's sketch. `gxwf draft-validate` returned one structure error against the whole `inputs` array; converting all three to single-element lists cleared it. Confirmed independently against the packaged `references/schemas/galaxy-workflow-draft.schema.json`, where the input variants type `format` as an array of strings only.
raised by
freeform-summary-to-galaxy-template (phase 5)
observed at
63a3f9cf9c97 rev 6
issue
not filed
draft-format-note-silent-on-step-addressing-by-labelmajorgapgalaxy-workflow-draft-format

The note never says which step field a connection's `source:` resolves against. Its only example uses the map form, where the step key and its identity are the same string, so the question cannot arise there. In the list form a step may carry both `id` and `label`, and `gxwf draft-validate` resolves `source:` against the LABEL when one is present — an `id` that differs from the label is not addressable at all.

expected
State the rule in the "Relaxations vs. gxformat2" section: a labelled step is addressed by its label, so in list form `id` and `label` must be the same string (which is what the IWC corpus does), or the label must be omitted. Adding a list-form example alongside the existing map-form sketch would carry the rule by demonstration.
evidence
A first draft of 27 steps, each with a snake_case `id` and a separate human `label`, with every `source:` written against the id. `gxwf draft-validate` returned 47 topology errors, all of the form `references unknown step "<id>"`, while reporting the same steps in its diagnostic paths under their labels. Renaming every `id` to match its `label` and rewriting all 47 references cleared it to `draft valid` with no other change. Neither the note nor the packaged `draft-validate` command reference mentions the distinction.
raised by
freeform-summary-to-galaxy-template (phase 5)
observed at
63a3f9cf9c97 rev 6
issue
not filed
draft-format-sentinel-hint-cannot-express-a-conditional-nested-portmajorgapgalaxy-workflow-draft-format

The `TODO_<port>` input sentinel carries the port's semantic hint as a FLAT identifier, so it cannot express a port that lives inside a conditional -- and it reads as though no qualification were needed. `__FILTER_FROM_FILE__` has one top-level `input` and a `filter_source` that exists only inside the `how` conditional, but the template wrote both as siblings, `TODO_input` and `TODO_filter_source`. The concrete `in:` key for the second is `how|filter_source`; a literal reading of the sentinel yields a bare `filter_source:`, which connects nothing. `gxwf draft-next-step` then restates the flat hint verbatim -- "assign the real wrapper input port name (semantic hint: 'filter_source')" -- so the loop's own work list repeats the wrong shape at the moment the author acts on it. The note does define qualified keys elsewhere (concrete steps in the same draft carry `select_data|rep_factorName_0|rep_factorLevel_0|countsFile`), so the shape is expressible; what is missing is any statement that the sentinel's hint is a NAME rather than a PATH, and that concretion may have to qualify it.

expected
State in the sentinel section that a `TODO_<port>` hint names a parameter, not its address, and that the concrete `in:` key must be the tool's full parameter path -- `|`-qualified through every enclosing conditional and repeat. Better still, let the sentinel carry the path it already knows when the template resolves a nested port, so the hint and the answer have the same shape. A worked conditional-nested example next to the existing flat one would carry the rule by demonstration, as the map-form/list-form correction did.
evidence
Established from the raw TRS payload `GET /api/tools/__FILTER_FROM_FILE__/versions/1.1.0` (200), where `filter_source` appears only under `how.whens[*].parameters` and never at the top level, and from `mgnify-amplicon-pipeline-v5-rrna-prediction` at IWC fe41a79 through `gxwf convert --to format2`, which writes `- id: how|filter_source`. The cost of the gap is that NOTHING catches the literal reading: this tool's outputs are collections, so it fails the cache decode and reports `skip_tool_not_found`, and a probe run of the same draft with the unqualified `filter_source:` key returned a byte-identical verdict to the correct one -- `draft valid`, `Concrete: OK`, `Tool state: 16 ok, 0 fail, 3 skip`. A silently disconnected input survives every gate in the run.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
draft-format-tiers-conflate-a-named-tool-with-a-named-tool-idmajorgapgalaxy-workflow-draft-format

The Identity-pinned tier admits "a source summary that names a specific `tool_id` with evidence", and the Mold's source-tendency paragraph relaxes that to "a free-form source that does name a specific tool/version with evidence hardens to the matching tier". Two paragraphs earlier the same note forbids pinning "on plausibility". A paper naming "FastQC" names a piece of software, not a Galaxy `tool_id`; reading the tier rule literally licenses writing a Tool Shed path from memory, which is exactly the plausibility pin the note forbids. The note never distinguishes the two kinds of naming, and they come apart on every free-form source.

expected
Say explicitly that Identity-pinned requires a concrete `tool_id` STRING from evidence — a corpus workflow, a pattern page's worked example, or a source that quotes the Galaxy tool id — and that a source naming a tool by its software name, with no corpus or pattern hit for its wrapper, is Deferred with the software name recorded in `_plan_context`. The source-tendency paragraph should be reworded to match, since as written it points the other way.
evidence
All six tools in this run's source are named by the paper. Five had a corpus-confirmed `tool_id` from the phase-4 exemplar and were Identity-pinned. FastQC had none: the nearest exemplar uses `iuc/falco` instead, so the only route to a FastQC `tool_id` was recall. That step was Deferred, against the plainest reading of the source-tendency sentence, on the strength of the plausibility prohibition. The two rules gave opposite answers for the same step and the note offers nothing to break the tie.
raised by
freeform-summary-to-galaxy-template (phase 5)
observed at
63a3f9cf9c97 rev 6
issue
not filed
galaxy-test-staging-drops-sample-sheet-column-definitionsmajorgapGalaxy — test-data staging for sample_sheet collections (galaxy.tool_util.cwl.util / galaxy.tool_util.client.staging)

A sample_sheet collection staged from a Planemo/gxwf test `job:` block never carries collection-level `column_definitions`, although the API it posts to accepts them. `galactic_job_json`'s `replacement_collection()` passes only `rows` and `name` for a sample_sheet collection type, and `StagingInterface`'s `create_collection_func` has no `column_definitions` parameter to pass — while `CreateNewCollectionPayload` declares the field and `SampleSheetDatasetCollectionType.generate_elements` reads it. The result is that a workflow input declaring `column_definitions` is, under test, fed a collection that has none. Per-element `columns` survive, so most workflows still behave, and the two `column_definitions_compatible()` call sites are both in `DataCollectionToolParameter` option-building (UI dropdown filtering) which a test bypasses by supplying the HDCA by id — so the divergence is silent rather than caught. It becomes visible at `__SAMPLE_SHEET_TO_TABULAR__`, whose header line is emitted only `#if $include_headers and $input.collection.column_definitions`: a workflow using that tool with headers on produces a header in the UI and no header under test, and no gate reports the difference.

expected
Accept `column_definitions` on a `Collection` entry in a test job block and thread it through `galactic_job_json` → `create_collection_func` → the collections API, alongside `rows`. Then a sample sheet staged for a test is the same object a user builds, and `validate_row` actually validates the fixture's rows against the declared columns instead of short-circuiting. Failing that, Galaxy should say plainly that test-staged sample sheets are column-definition-free, so workflow authors know that any behaviour gated on `column_definitions` is untestable.
evidence
Read at galaxyproject/galaxy dev while resolving open-requirements entry `sample-sheet-input-test-fixture-expressibility` for this run: lib/galaxy/tool_util/cwl/util.py `replacement_collection()` (sample_sheet branch sets only `kwds["rows"]`); lib/galaxy/tool_util/client/staging.py `create_collection_func` signature `(element_identifiers, collection_type, rows=None, name=None)`; lib/galaxy/schema/schema.py `CreateNewCollectionPayload.column_definitions`; lib/galaxy/model/dataset_collections/types/sample_sheet.py `generate_elements`; lib/galaxy/model/dataset_collections/types/sample_sheet_util.py `validate_row` and `column_definitions_compatible`; lib/galaxy/tools/sample_sheet_to_tabular.xml. This run's workflow is unaffected only because its `Project sample sheet to tabular` step sets `include_headers: false`. This entry also supplies part of the evidence asked for by `testability-note-silent-on-sample-sheet-test-fixtures`, which remains open on its own subject.
raised by
paper-to-test-data (phase 7)
observed at
63a3f9cf9c97 rev 2
issue
not filed
gxwf-cannot-decode-collection-outputs-of-builtin-collection-operationsmajordefect@galaxy-tool-util/cli — gxwf tool fetch/decode for built-in collection operations

`gxwf draft-validate --concrete` cannot bring the built-in `__FLATTEN__` into its tool cache. The fetch is reported as `toolshed fetch failed ... for __FLATTEN__`, but the accompanying dump is a decode failure against the tool schema, not a transport error: `["outputs"][0]` is decoded as the collection-output branch and `["structure"]` `is missing`. The tool's collection output carries no `structure`, and the schema requires one. The step is then reported `skip_tool_not_found`, so its tool state is never validated — and no amount of cache priming can fix it, because the tool cannot be decoded into the cache in the first place. The same shape is likely to hit the other collection-operation built-ins (`__UNZIP_COLLECTION__`, `__FILTER_FROM_FILE__`, `__APPLY_RULES__`), which are exactly the steps a Galaxy draft uses for collection plumbing.

expected
Make `structure` optional on the collection-output branch of the tool schema (or supply a default for tools that declare a collection output without one), so built-in collection operations decode and cache like any other tool. Separately, distinguish a transport failure from a decode failure in the message: "toolshed fetch failed" sent this run looking at network and cache-priming for a fault that was neither.
evidence
`gxwf draft-validate <run>/galaxy-workflow-draft.gxwf.yml --concrete --cache-dir <cache>` at iteration 2, gxwf 1.10.1. Verdict line `Tool state: 2 ok, 0 fail, 1 skip`, with `0 (__FLATTEN__) [skip_tool_not_found] __FLATTEN__ not in cache (fetch failed)`. The two compose_text_param steps validated from the same cache in the same run, so the cache path and network were both working. The decode dump is also ~4 KB of expanded structural type text printed ahead of the verdict, which buries the one line that names the cause.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
gxwf-drops-connectedvalue-under-format2-state-keymajordefect@galaxy-tool-util/cli — gxwf tool-state validation (state: vs tool_state:)

`gxwf validate` (and `draft-validate --concrete`) drops a `component_value: {__class__: ConnectedValue}` from a step's tool state before validating it when the state is written under the format2 key `state:`, but honours it under `tool_state:`. The same bytes under the two keys therefore get opposite verdicts. Any conditional parameter case whose generated `workflow_step_linked` branch marks `component_value` as required — every non-text case — then fails as "component_value: is missing", and the anyOf fallback reports the misleading "select_param_type: Expected \"text\", actual \"float\"" alongside it.

expected
Treat `state:` and `tool_state:` identically in the tool-state validation path, so a connected value placeholder survives to validation under both keys. Failing that, the generated `workflow_step_linked` schema should not require `component_value` for a parameter that a step connection supplies, which is exactly the relaxation that distinguishes it from `workflow_step`. A one-line note in the `draft-validate` / `validate` docs would not be enough: the failure names a required field the author deliberately connected, which reads as an authoring error rather than a tool bug.
evidence
gxwf 1.10.1. Minimal one-step workflow around `iuc/compose_text_param/compose_text_param@0.1.1`, second repeat component in the float case, `component_value: {__class__: ConnectedValue}`, connected via `components_1|param_type|component_value`. Under `state:` the tool-state check fails with the five diagnostics above; byte-identical content under `tool_state:` reports `tool_state: OK`. Failure does not depend on the connection being present, on `__index__`, or on `__current_case__`; substituting a literal float under `state:` passes, which is the wrong workaround to be nudged toward — it leaves a shadow default behind a connected parameter. The IWC exemplar this run compares against (`transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting` at fe41a79) passes its three compose_text_param steps precisely because `gxwf convert --to format2` emits `tool_state:`.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
gxwf-linked-step-schema-rejects-the-keys-its-own-converter-emitsmajordefect@galaxy-tool-util/cli — generated input_schemas.workflow_step_linked vs gxwf convert --to format2

Three surfaces of the same CLI disagree about whether format2 `tool_state` may carry the `__`-prefixed bookkeeping keys Galaxy writes into conditionals and repeats. For `iuc/map_param_value/map_param_value` 0.2.0, `galaxy-tool-cache summarize` generates `input_schemas.workflow_step_linked` with `additionalProperties: false` on every conditional branch object (allowing only `type`, `input_param`, `mappings`) and on every `mappings` item (allowing only `from`, `to`). That schema rejects `__current_case__` and `__index__`. But `gxwf convert --to format2` of a real Galaxy workflow emits exactly those keys, and `gxwf draft-validate --concrete` accepts a state block containing them. The generated schema is therefore stricter than both the converter that produces format2 and the validator that checks it — and it is the one surface a Mold is told to author against. implement-galaxy-tool-step step 2 says to "shape the step's `state` against `input_schemas.workflow_step_linked`"; following that literally produces a state block that diverges from what Galaxy round-trips, and no gate reports the divergence, because the validator is the permissive one.

expected
Make the three surfaces agree, and say which one is normative. Either the generated linked-step schema should admit the `__`-prefixed bookkeeping keys the converter emits (`__current_case__`, `__index__`, and the top-level `__page__` / `__rerun_remap_job_id__`), or `gxwf convert --to format2` should strip them and the validator should reject them. Until then, document in the schema-generation output which form is canonical for authoring, so a Mold binding a step against the published schema gets the same answer as the validator and the converter.
evidence
Iteration 6 of this run, gxwf 1.10.1, step `Get featureCounts strandedness parameter`. Four state shapes were probed against `gxwf draft-validate <run>/galaxy-workflow-draft.gxwf.yml --concrete --cache-dir <cache>`: with the bookkeeping keys, without them, with the connected `input_param` omitted from the state block, and with the whole block moved from `tool_state:` to `state:`. All four returned `Tool state: 6 ok, 0 fail, 1 skip` — the validator discriminates none of them. `gxwf convert --to format2` of the corpus workflow `transcriptomics/rnaseq-pe/rnaseq-pe` at IWC fe41a79 emits, for the same step, `input_param_type: {type: text, __current_case__: 0, input_param: {__class__: ConnectedValue}, mappings: [{__index__: 0, from: ..., to: ...}, ...]}` and `unmapped: {on_unmapped: fail, __current_case__: 1}` — every key the generated schema forbids. The round-trip form was taken as authoritative for this step.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
gxwf-strict-encoding-demands-the-state-key-its-validator-mishandlesmajordefect@galaxy-tool-util/cli — gxwf validate --strict-encoding vs the tool-state validator

gxwf's two validation paths demand opposite tool-state keys, so no format2 workflow can satisfy both. `--strict-encoding` rejects `tool_state:` on every step with `uses "tool_state" instead of "state" (format2 should use "state")` and exits 2, while the tool-state validator silently drops a `{__class__: ConnectedValue}` placeholder under `state:` and honours it only under `tool_state:` (this ledger's `gxwf-drops-connectedvalue-under-format2-state-key`). Moving to the key `--strict-encoding` wants reintroduces that bug; staying on the key that works means `--strict` can never be part of a Foundry gate. `gxwf convert --to format2` emits `tool_state:`, so gxwf's own converter produces output its own `--strict-encoding` rejects.

expected
Reconcile the two paths as one decision. If `state:` is the canonical format2 key, fix the tool-state validator to honour ConnectedValue under it first, then keep the strict-encoding diagnostic. If `tool_state:` is to remain accepted, `--strict-encoding` should not flag it, and `gxwf convert --to format2` should emit whichever key the strict path blesses. Until then the diagnostic actively steers authors toward the broken key.
evidence
gxwf 1.10.1. `gxwf validate <run>/galaxy-workflow.gxwf.yml --json --strict-structure --strict-encoding --cache-dir <cache>` exits 2 and emits the message for all 26 steps that carry tool state (step 0 has none); the same file under default strictness reports zero encoding errors and zero structure errors. The workflow uses `tool_state:` on 26 steps because phase 6 established at iteration level that `state:` breaks the three `iuc/compose_text_param/compose_text_param@0.1.1` steps. Companion to, not a duplicate of, `gxwf-drops-connectedvalue-under-format2-state-key`: that entry is the validator dropping the placeholder, this one is the strict gate requiring the key that triggers it. A maintainer fixing one need not touch the other; triage may merge them into a single upstream issue.
raised by
validate-galaxy-workflow (phase 10)
observed at
79bf5c3ab98e rev 5
issue
not filed
gxwf-validate-json-output-is-not-machine-parseablemajordefect@galaxy-tool-util/cli — gxwf validate --json stdout contract

`gxwf validate --json` does not put JSON, and only JSON, on stdout, so the documented machine interface cannot be consumed by parsing stdout. Two separate breaches. (1) Every uncached tool that fails to decode prints a multi-line `toolshed fetch failed (...) for <id>:` block to stdout ahead of the JSON document - the full Effect schema type, roughly 55 lines for seven such tools - so `JSON.parse(stdout)` throws and the report has to be located by scanning backwards for a line that is exactly `{`. (2) Adding any strict flag makes gxwf abandon JSON entirely: exit 2, plain-text diagnostics on stderr, and stdout completely empty, although `--json` was passed.

expected
With `--json`, write the report and nothing else to stdout, and route fetch/decode diagnostics to stderr. Carry the strict-mode findings inside the JSON report (the schema already has `structure_errors` and `encoding_errors` arrays for exactly this) and keep emitting it on the strict failure path, signalling the verdict through the exit code rather than by withholding the document. A harness cannot classify a failure it cannot parse.
evidence
gxwf 1.10.1. Terminal validation of <run>/galaxy-workflow.gxwf.yml: stdout begins with the decode-failure text for `__FLATTEN__` and continues for six more tools before the report; the JSON body starts at line 57 of 273. The same command plus `--strict-structure --strict-encoding` returns exit 2 with 0 bytes on stdout and the 26 encoding messages on stderr. The packaged CLI reference states "JSON output should be treated as the preferred cast-skill interface", which is the contract being broken.
raised by
validate-galaxy-workflow (phase 10)
observed at
79bf5c3ab98e rev 5
issue
not filed
gxwf-validate-never-checks-in-key-namesmajorgap@galaxy-tool-util/cli — gxwf validate tool-state / step-input name checking

`gxwf validate` never checks that a step's `in:` keys name real parameters of the pinned tool, even when that tool is fully cached and its state validates. A step wiring a connection to a port the tool does not have is reported `tool_state: OK` and counted in the validated total. The same permissiveness runs the other way inside `tool_state:`: an unknown extra parameter is accepted, and a required parameter that is simply absent is accepted. What the validator actually checks is the type and value of the parameters that happen to be present. No strictness flag changes this - not `--strict-structure`, not `--strict-state`, not `--strict`. This matters most exactly where Galaxy's own syntax is easiest to get wrong: a conditional's nested port must be qualified (`how|filter_source`), and the unqualified spelling is silently accepted.

expected
When the step's tool is resolved from the cache, check each `in:` key against the tool's input tree - including conditional qualification (`cond|param`), repeat indexing (`name_<n>|param`) and sections - and report an unknown port as an error, or at minimum under `--strict-state`. Report a missing required parameter the same way. A validator that reports `20 validated` while never having looked at a single port name overstates its own coverage to any harness reading the summary.
evidence
gxwf 1.10.1, measured directly. A minimal format2 workflow with a single `Filter1@1.1.1` step, the tool present in the cache, and `in: {not_a_real_tool_input: tbl}` reports `Structural validation: OK`, `tool_state: OK`, `Tool state: 1 validated, 0 skipped`, identically under `--strict-structure --strict-state`. On <run>/galaxy-workflow.gxwf.yml, adding `bogus_param_probe: "xyz"` to a validated Filter1 step leaves the summary at 20 ok / 0 fail / 7 skip, and deleting the required `cond` parameter also leaves it at 20 / 0 / 7; by contrast `header_lines: "notanint"` on the same step does fail, which is what the check does cover. This run's three `__FILTER_FROM_FILE__` steps depend on the qualified `how|filter_source` spelling and nothing in the toolchain would have caught the unqualified one.
raised by
validate-galaxy-workflow (phase 10)
observed at
79bf5c3ab98e rev 5
issue
not filed
implement-step-mold-has-no-binding-path-when-no-tool-summary-existsmajorgapimplement-galaxy-tool-step

Both Molds assume a tool summary always exists. advance-galaxy-draft-step's sequence step 3 is an unconditional "Invoke summarize-galaxy-tool on the resolved wrapper", and implement-galaxy-tool-step's sequence step 2 is an unconditional "Read the galaxy-tool-summary manifest". Neither says what to do when the summary cannot be produced at all. That is not a hypothetical: summarize-galaxy-tool can only summarize what `galaxy-tool-cache add` could decode, and a wrapper with a collection output cannot be decoded (see `collection-output-decode-is-a-flat-vs-nested-shape-mismatch-not-a-missing-field`). The Mold's one nearby escape hatch does not cover it either — "If `input_schemas` is `null`, consult `warnings[]`" presupposes a manifest with a warnings array, and here there is no manifest. The result is an author improvising the most consequential part of a step, its tool state, with no stated evidence standard, on the first genuinely complex wrapper of the run.

expected
Give implement-galaxy-tool-step an explicit no-summary branch that names the fallback evidence in priority order and requires the step to record which one it used: (1) the wrapper XML at the pinned changeset, which is authoritative for parameter names, defaults and output filters; (2) `gxwf convert --to format2` over a corpus `.ga` that pins the same version, which is authoritative for the round-trip state shape; (3) nothing else. summarize-galaxy-tool already sanctions raw XML as supporting evidence ("Optional raw XML source for ambiguity checks"), so the material exists — it is the implement Mold that never mentions it. The branch should also require the step to carry a visible marker that its state was hand-bound and is therefore unchecked by `draft-validate`, since the verdict line will show it as a skip and a skip is indistinguishable from an absent step in the counts.
evidence
Iteration 8 of phase 6 in this run, step `Quality-trim reads`, wrapper `lparsons/cutadapt/cutadapt` 5.2+galaxy2 at changeset f6168dd17f82. `galaxy-tool-cache add` failed at four different versions, so summarize-galaxy-tool could not run and no `galaxy-tool-summary.json` was ever produced. The step's `tool_state` — a `library` conditional at `__current_case__: 2` with six adapter repeats, plus `output_selector`, which gates the promoted `report` output through an XML `<filter>` — was bound entirely by hand from the tools-iuc wrapper XML at the pinned version and from `gxwf convert --to format2` over `epigenetics/cutandrun/cutandrun.ga` and `VGP-assembly-v2/post-curation-processing/Post_Curation.ga` at fe41a79. That route was inferred, not instructed. `draft-validate --concrete` then reported `7 ok, 0 fail, 2 skip` with this step among the skips, so nothing in the run checks the binding.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
implement-step-mold-says-state-without-saying-which-keymajorgapimplement-galaxy-tool-step

The procedure says to "shape the step's `state` against `input_schemas.workflow_step_linked`" and speaks of `state` throughout, but the galaxy-workflow-draft schema it packages admits both `state` and `tool_state` on a step and the Mold never says which to write. The two are not interchangeable in practice: under gxwf 1.10.1 a connected non-text parameter validates under `tool_state:` and fails under `state:` (filed as `gxwf-drops-connectedvalue-under-format2-state-key`), so following the Mold's own wording is what produces the red verdict.

expected
Name one key as the authoring convention and say it once — `tool_state:`, which is what `gxwf convert --to format2` emits and what every converted IWC exemplar a run compares against will therefore show — and note that a step mixing the two conventions within one draft is a readability cost, not a correctness one. This is worth fixing independently of the gxwf defect: the ambiguity is in the instruction, and it will still be there after the validator is corrected.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
iwc-exemplar-artifact-assumes-a-single-nearest-exemplarmajorgapcompare-against-iwc-exemplar

The Mold speaks of "nearest IWC exemplar(s)" in its summary, its procedure and its confidence table, but the gxformat2 artifact is specified strictly in the singular — one declared filename, "the nearest exemplar's relevant subgraph", "Once the nearest exemplar is chosen (High or Medium confidence), convert it". Nothing says what to emit when the honest answer is several exemplars covering disjoint parts of the subject. That is what happened here and it was not an edge case: IWC publishes the subject's journey as two workflows joined at the count-table boundary, so one exemplar covers the map-over head and a different one covers the differential-expression tail, and a third workflow in another domain was the only corpus source for one tool's output shape. Whether that is one file with several YAML documents, several files, or a single forced choice was guessed at.

expected
State the multi-exemplar case in the "Nearest exemplar (gxformat2) view" section and settle the encoding. The shape that worked here and is worth specifying: one file at the declared filename, one YAML document per exemplar, each document headed by the abstract IWC workflow ID it came from, the steps covered, the steps dropped, and its own confidence level — so a Low-confidence cross-domain citation cannot be mistaken for a domain exemplar, which the section already warns about in prose but gives no structural way to express. Worth pairing with a sentence in the Feature Hierarchy noting that a corpus which splits the subject's scope across workflows is itself a first-class structural finding for the template tier, not a retrieval failure.
evidence
Ranking produced transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting at High for the DE tail, transcriptomics/rnaseq-pe/rnaseq-pe at Medium for the map-over head, and epigenetics/cutandrun/cutandrun as tool-level-only evidence for the Cutadapt output shape. <run>/iwc-exemplar.gxwf.yml carries all three as three YAML documents at the one declared filename, an encoding invented in this run.
raised by
compare-against-iwc-exemplar (phase 4)
observed at
63a3f9cf9c97 rev 10
issue
not filed
iwc-test-data-conventions-teaches-a-nesting-shape-the-schema-gate-rejectsmajordefectiwc-test-data-conventions

Section 2e states normatively "Note: outer `collection_type: list:paired`, inner `type: paired` (not `collection_type:`)", and section 2f's nested-output example omits the `class: Collection` discriminator on the outer `element_tests` entry. Both forms are faithful transcriptions of the IWC corpus and both are rejected by `references/schemas/ tests-format.schema.json`, which the same Mold packages and names as the gate its output must pass. An agent that follows the note produces a file that fails the Mold's own step 4, with 48 unhelpful `oneOf` errors pointing at the whole job input rather than at the offending key. The note is the only packaged guidance on these two shapes, so there is nothing else to fall back to.

expected
Give both shapes in each section, marked: the corpus form (`type: paired`; bare `elements:`) and the schema-valid form (`collection_type: paired`; `class: Collection` then `elements:`), with a one-line note that the validator accepts only the latter and a pointer to the upstream entry `tests-format-schema-rejects-two-shapes-the-iwc-corpus-uses`. Until upstream converges, a Foundry-authored test file should use the schema-valid form, and the note should say so rather than leaving the reader to discover it from the validator.
evidence
This run authored the reads input from section 2e verbatim; `gxwf validate-tests` 1.10.1 returned 48 errors. Switching the single key `type` to `collection_type` returned OK. The same substitution is needed for the section 2f output form, confirmed by removing `class: Collection` from one element of the finished file.
raised by
implement-galaxy-workflow-test (phase 9)
observed at
63a3f9cf9c97 rev 8
issue
not filed
ledger-resolved-entries-are-never-read-back-at-binding-timemajorgapadvance-galaxy-draft-step

The open-requirements ledger is the only artifact that carries a decision from one iteration of the per-step loop to a later one, but the Mold never reads it as an INPUT to binding. Its `Inputs` section declares the ledger carries "the run's open, resolved, and surrendered entries", and then the procedure's single ledger touchpoint is step 5, AFTER implement: "Inspect the open-requirements-ledger for a new `open` blocking entry ... appended against this step." New, open, post-hoc. A decision settled five iterations earlier lives in a `resolved` entry -- whose `note` field is where the ledger note itself says the closure reasoning goes -- and no procedure step ever routes an author back to it before they bind state. The packaged `open-requirements-ledger` note does not close the gap either: its "how downstream reads it" paragraph covers only repair-galaxy-draft-topology reading OPEN blocking entries, and its one line about resolved entries ("Resolving is not deleting -- a resolved entry stays in the ledger as the audit trail") frames them as provenance, not as an input.

expected
Add a read to the procedure BEFORE implement: select the ledger entries whose `step` names the step being concretized -- regardless of status -- and treat a `resolved` entry's `note` as a binding already settled for this step, to be copied rather than re-derived. Entries carry a `step` field precisely so this lookup is a filter, not a search. Say explicitly that a resolved entry is authority at binding time and not merely an audit trail, and say the converse too: a Mold that settles a binding for a step it is not currently implementing must record it in that entry's `note`, because the draft itself has nowhere to put it (the sibling steps' `_plan_state` prose is deleted at their own concretion).
evidence
This iteration concretized `Filter with p-adj threshold: first contrast`. Its whole remaining decision was `header_lines`, and the answer -- `'0'`, because the pinned iuc/deseq2 wrapper writes `deseq_out` with `col.names = FALSE` -- had been established two iterations earlier and recorded ONLY in the `note` of the now-`resolved` entry `deseq2-result-table-header-presence-unverified`. Following the procedure literally, that note is never opened: step 5 filters for new open entries and runs after the binding is already written. The default an author reaches for instead is the corpus, and the corpus is wrong here -- `rnaseq-de-filtering-plotting` binds `header_lines: "1"` because it MANUFACTURES a header (tp_text_file_with_recurring_lines -> tp_sed -> tp_cat) before filtering, which this workflow omits. Binding "1" against the headerless `deseq_out` discards row 1 of a table sorted by padj: the most significant gene, which for this paper is *SCF1*, the finding the workflow exists to reproduce. No error, no warning, and `draft-validate --concrete` returns identical green for '0' and '1'. Three more steps (the remaining significance filters) must copy that same value, each from an iteration that will not be told to look.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
no-evidence-route-for-an-output-content-property-the-declaration-cannot-carrymajorgapimplement-galaxy-tool-step

Some step bindings depend on a property of an UPSTREAM output's CONTENT -- does this tabular output carry a header row, is the element identifier emitted as a column -- and no artifact the Mold names can answer that class of question. `parsed_tool.outputs` carries name, label, type, format and discovery rules; it carries nothing about the rows inside the file, and it never could, because the property is a fact about the wrapper's script rather than about its declaration. The Mold has no instruction for the case, so the author either guesses or invents a method. This is not the same wall as `implement-step-mold-has-no-binding-path-when-no-tool-summary-exists`: there the summary is merely absent, and the fallback list that entry proposes -- wrapper XML at the pinned changeset, then a corpus round-trip, then "nothing else" -- would still not answer this, because the XML's `<param>`/`<data>` declarations are silent on it and the round-trip actively misleads. The cost is silent: a wrong `header_lines` drops the first data row of every filtered table, or passes a header into a numeric comparison, and either way the workflow emits a plausible result.

expected
Add to the binding procedure a named evidence route for output CONTENT properties, distinct from the one for parameter shape, and require the step to record which source settled it: the wrapper's `<test>` blocks -- `assert_contents` on the output in question is a positive statement about what the file actually contains, and the absence of a header assertion on one output beside its presence on a sibling is itself evidence -- and, when the wrapper is open-source, its script's write call. Say explicitly that a corpus round-trip is NOT authority for this class: a corpus workflow's binding is evidence about the table IT filters, which may be several steps removed from the producer, and copying it across is exactly the error this route exists to prevent. Related but separable: this run's phase-5 Mold recorded the question as "closable by a summarize-galaxy-tool pass", which is false for the same reason -- the pass it names cannot see the property.
evidence
Phase 6 iteration 21 of this run, closing open-requirement `deseq2-result-table-header-presence-unverified` for `iuc/deseq2/deseq2` @ 2.11.40.8+galaxy4 (changeset 05f9e54d7e81), which gates `header_lines` on the four downstream `Filter1` steps. The tool summary route was unavailable at all (the cache add fails on the collection-output decode). What settled it was outside every sanctioned source: `deseq2.R` writes the result table with `write.table(..., col.names = FALSE)` while writing the normalized counts one screen up with `col.names = NA` -- one tool, opposite answers for two tabular outputs, so no tool-level generalization is available either; and `deseq2.xml`'s test asserts `deseq_out` content starting at a data row with `has_n_lines`, while the `vst_out` and `counts_out` assertions in the same test block DO assert a sample-name header line. The round-trip is the trap: `gxwf convert --to format2` over IWC `transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting` at fe41a79 shows BOTH its `Filter1` steps binding `header_lines: "1"`, which reads as a direct answer and is not one -- that workflow manufactures a header with `tp_text_file_with_recurring_lines` + `tp_sed_tool` and concatenates it onto the DESeq2 output with `tp_cat` before filtering. Copying its binding into a workflow that filters `deseq_out` directly would have silently discarded the most significant gene from every result table.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
no-reachable-wrapper-xml-at-a-pinned-changesetmajorgapGalaxy Tool Shed + galaxy-tool-cache: raw wrapper source retrieval

Several authoring paths depend on reading a wrapper's XML at the pinned changeset -- the correction already filed as `implement-step-mold-has-no-binding-path-when-no-tool-summary-exists` names it as fallback evidence "authoritative for parameter names, defaults and output filters", and it is the only source for output filters at all (see `tool-summary-drops-output-filter-expressions`). No declared tool can fetch it. The summary manifest sets `artifacts.raw_tool_source_path: null` for a toolshed-sourced tool; the Tool Shed's `repos/<owner>/<repo>/raw-file/<changeset>/<path>` endpoint answers 403; and `/api/tools/<trs-id>/versions/<v>/raw_tool_source` answers 404 with "No route". The only thing that worked this run was guessing the upstream repository layout from the repository record's `remote_repository_url` and fetching the file from GitHub at the default branch -- where the filename was `rg_rnaStar.xml`, not the `<repo>/<tool_id>.xml` the shed path implies, so the first two guesses 404'd. That fetch is also UNPINNED: it happened to match the pinned version here only because tools-iuc main still carried @TOOL_VERSION@ 2.7.11b / @VERSION_SUFFIX@ 1, which a later iteration on a lagging wrapper cannot count on.

expected
Give `galaxy-tool-cache add` an option to retain the raw tool source it already downloads, and populate `artifacts.raw_tool_source_path` for toolshed sources rather than nulling it -- the bytes are in hand at fetch time, so this is retention, not new retrieval. Failing that, the Tool Shed should expose a supported raw-source route per (tool id, version) so the pin and the source agree. Until one of those exists, any Mold instruction of the form "bind against the wrapper XML at the pinned version" names an artifact the run has no sanctioned way to obtain, and authors will keep reaching an unpinned copy by guesswork.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
paper-to-galaxy-path-has-no-reference-data-ownermajorgappaper-to-galaxy

The Nextflow path has a dedicated Mold for deciding the Galaxy-side shape of external reference data (nextflow-summary-to-galaxy-reference-data, listed among the producers of the open-requirements ledger). The paper path has no equivalent, and this run's phase roster contains no reference-data phase. The source paper nonetheless pins a specific non-model assembly (GCA_002759435.2, C. auris B8441) and requires an annotation file, so the decision between a built-in index, a data-table string, and a portable history dataset had to be made somewhere. The interface Mold made it, unprompted: nothing in its procedure or references assigns reference-data shape to the interface tier.

expected
Either add a reference-data phase to the paper-to-galaxy pipeline (a freeform sibling of nextflow-summary-to-galaxy-reference-data, since the input is a narrative accession rather than an igenomes-style parameter), or state explicitly in freeform-summary-to-galaxy-interface that reference-data delivery shape is its call and package the guidance that decision needs. Today the decision is unowned, which means it gets made by whichever Mold notices, with no shared basis.
evidence
Run roster phases 1-12 in this ledger contain no `*-to-galaxy-reference-data` step. The interface brief settles the genome as a history fasta on portability grounds and carries open-requirements entry `reference-genome-delivery-shape-unverified` recording that no built-in-index check was performed, because no phase owns performing one. The subject's content hash is null: this Mold's runtime notes forbid reading Foundry source, so the pipeline note could not be hashed from inside the run.
raised by
freeform-summary-to-galaxy-interface (phase 2)
observed at
63a3f9cf9c97 rev 3
issue
not filed
paper-to-test-data-declares-no-test-data-refs-shapemajorgappaper-to-test-data

The Mold emits `test-data-refs.json` and specifies no shape for it. The whole procedure is one sentence — "derive concrete workflow test inputs and expected outputs — resolvable URLs, file shapes, and expected hashes — emitted as `test-data-refs`" — and the bundle packages no schema, template, or worked example (`refs: []`, "Load Upfront" and "Load On Demand" both "None declared"). The whole JSON structure was invented at runtime: how to key an input against a workflow input label, how to express a collection input's elements and per-row metadata, where expected outputs live and how to separate an assertion the data supports from one it does not. Two downstream Molds in this pipeline consume the artifact (freeform-summary-to-galaxy-test-plan at phase 8, implement-galaxy-workflow-test at phase 9) and neither can rely on any particular key existing. This is the same defect class already filed as `summarize-paper-declares-no-freeform-summary-shape`, at the other end of the same pipeline.

expected
Package a `test-data-refs` skeleton or worked example, or state in the procedure the minimum keys every consumer can assume. At minimum: one entry per workflow input keyed by the input's label, carrying source URL, hash, datatype and (for collections) collection type, element identifiers and per-element metadata; expected outputs separated from inputs; and an explicit place to record which expected outputs the resolved data can actually produce versus which it cannot. The three sibling test-data Molds (`nextflow-to-test-data`, `cwl-to-test-data`, `find-test-data`) emit the same artifact id and should share whatever contract is written.
evidence
Bundle carries SKILL.md plus _feedback/_provenance/_verify only. The Output section names the filename, the format (`json`) and a one-line description, and nothing else. The artifact this run wrote (<run>/test-data-refs.json) has seventeen top-level keys, none of which were prescribed.
raised by
paper-to-test-data (phase 7)
observed at
63a3f9cf9c97 rev 2
issue
not filed
paper-to-test-data-silent-on-subsetting-deposited-datamajorgappaper-to-test-data

The Mold is silent on subsetting, which is the central problem of deriving test data from a paper. A paper's deposited data is essentially always orders of magnitude too large to be a workflow test fixture — this run's six SRA runs are 22–30 M read pairs each, ~6.5 GB — so the Mold's real task is not "resolve URLs" but "resolve URLs AND decide a subset AND establish whether that subset still supports the paper's claims". The procedure's phrasing ("resolvable URLs, file shapes, and expected hashes") reads as though the deposited data is used as-is. Nothing tells the runtime to subset, how to choose a depth, how to keep the recipe deterministic, or how to decide whether the result is a fixture that reproduces the paper's finding or one that only exercises the workflow's shape. That last decision is exactly what the Foundry's own fixture rule makes mandatory — "a fixture must also be able to produce the outcome the scenario bound to it claims" (AGENTS.md) — and the Mold never routes to it. The subsetting strategy, the depth, the evidence standard and the tiering of assertions by whether the data supports them were all invented here.

expected
Add a subsetting step to the procedure and state the decision it has to reach. Name the common strategies and when each applies (deterministic head subset; alignment-based region targeting; downsampling to a fixed seed), require the recipe be reproducible and the subset be hashed on its UNCOMPRESSED bytes, and require the output artifact to state explicitly whether the subset reproduces the paper's result or is shape-only — citing the fixture rule so the runtime knows an unmarked shape-only fixture is a defect and not a shortcut. Relatedly, "Required Tools: None declared. Procedure should not assume external CLIs are present" sits in tension with the same procedure demanding "expected hashes": a hash of a subsetted fixture cannot be produced without tooling, so either the tool expectation or the hash expectation should be stated honestly.
evidence
This run reached a defensible answer only by going well outside anything the Mold describes: measuring SCF1-matching reads per 1 M-read block to show a head subset carries no positional bias, then aligning 200,000-pair subsets of all six runs to the pinned reference to measure that SCF1 retains 414/391 fragments in AR0382 against 1–4 in the two comparison conditions. Without that work the honest answer would have been "shape-only, probably"; the Mold gave no reason to do it and no standard to judge it against.
raised by
paper-to-test-data (phase 7)
observed at
63a3f9cf9c97 rev 2
issue
not filed
sample-sheet-note-omits-sample-sheet-to-tabular-output-columnsmajorgapgalaxy-sample-sheet-collections

The note's closing guidance is that carry-forward of sample-sheet metadata past map-over "must be explicit (re-attaching metadata via `__SAMPLE_SHEET_TO_TABULAR__` or rules DSL `add_column_from_sample_sheet_index`)", and the note is packaged into this Mold precisely to drive that re-attachment. But it describes the tool in one clause — "iterates and tab-joins for downstream tabular consumers" — and never states its output columns. Whether the element identifier is emitted as a column is the single fact the re-attachment turns on, because the identifier is the only key that survives map-over and therefore the only possible join key back to a downstream collection. The note recommends the mechanism without supplying what is needed to wire it.

expected
State the tool's output schema in the "Tool-side access" section: whether the element identifier is emitted, in which column, and how the remaining columns order relative to `column_definitions`. Do the same for the rules-DSL alternative — which axis `add_column_from_sample_sheet_index` reads and whether it can apply to a collection that has already lost its `column_definitions`, since that is the case a reader arrives with. The sources list already cites lib/galaxy/tools/sample_sheet_to_tabular.xml, so the fact is one line away from where it is needed.
evidence
This run's central data-flow decision (<run>/freeform-galaxy-data-flow.md section 4) is an identifier-keyed split that joins this tool's output to the featureCounts collection. It had to ship with open-requirements entry `sample-sheet-to-tabular-identifier-column-unverified` and a named substitute node, because the join key could not be confirmed from the bundle and the Mold's runtime notes forbid reading Galaxy or Foundry source.
raised by
freeform-summary-to-galaxy-data-flow (phase 3)
observed at
63a3f9cf9c97 rev 3
issue
not filed
shed-search-normalization-misses-abbreviation-expansionmajorgapdiscover-shed-tool

The procedure's query-normalization recipe for a tool-id-shaped need is to strip any `owner/` prefix, split on `_` / `-` into space-separated words, and also try the bare significant word — and it states that "a `miss` is only honest after the name variants have been tried". Every variant that recipe generates fails for `iuc/map_param_value/map_param_value`, a tool that is published, current, and pinned by the IWC corpus. The recipe does not cover the one transform that works: expanding an abbreviated word in the id to the word the human tool name actually uses. Galaxy tool ids abbreviate routinely (`param`, `val`, `seq`, `align`, `qc`, `col`), so this is a recurring shape, not a one-off. An iteration following the packaged recipe literally would have declared `miss` and fallen through to author-galaxy-tool-wrapper — authoring a new wrapper for a tool that already exists, which is the most expensive possible wrong answer this Mold can produce.

expected
Add abbreviation expansion to the §1 normalization list, with the common Galaxy abbreviations spelled out, and make the honest-miss condition explicit that it includes expanded variants. Better still, invert the last resort: before returning `miss`, search the significant words with the id token dropped entirely (here `parameter value` alone ranks the target first), and state that a `miss` is only honest when a name-shaped query has been tried, not merely an id-shaped one.
evidence
Iteration 6 of this run, gxwf 1.10.1, resolving the wrapper for step `Get featureCounts strandedness parameter`. `gxwf tool-search` returned zero hits for `map param value` (the recipe's underscore split), zero for the same query scoped `--owner iuc`, and zero for the raw token `map_param_value`. The bare significant word `map` returned only unrelated tools (`tdrmapper`, `multi_fasta_glimmerhmm`, `glimmerhmm_predict`). Expanding `param` to `parameter` found it immediately and first: `map parameter value` scores 42.1 at rank 1, and `parameter value` scores 29.7 at rank 1. The tool's human name is "Map parameter value".
raised by
discover-shed-tool
observed at
63a3f9cf9c97 rev 5
issue
not filed
stock-tool-version-resolution-is-circular-in-the-builtin-branchmajorgapadvance-galaxy-draft-step

Procedure step 2's built-in/stock branch offers exactly two ways to get a stock tool's version — "read it from a populated cache via `galaxy-tool-cache list` or take a known pin from the step plan" — and then forbids the only remaining move: "never hand-guess a stock version". Both offered sources presuppose the version is already known. For a stock tool meeting the run for the first time, with `tool_version: TODO` in the draft and no step-plan pin, the cache is empty of it precisely because populating the cache requires an `add` that takes `--tool-version`. The branch is circular, and its stated prohibition closes the only exit. There is no third route to fall back on: the Tool Shed publishes no version listing for a bare stock id at all.

expected
Say that a bare-id probe is the sanctioned route for a stock tool and why it is not a guess: `galaxy-tool-cache add <id> --tool-version <v>` either returns a summary whose own `id` and `version` fields confirm the pin, or fails loudly with a 404 from `/api/tools/<id>/versions/<v>`. There is no silent-wrong outcome, so the probe is self-verifying and the prohibition should be narrowed to "never record an unconfirmed version" rather than "never try one". Name the starting probe — Galaxy's built-in collection operations ship at `1.0.0` — and require that the confirming summary be read back before the version is written into the draft.
evidence
Iteration 7 of phase 6 in this run, gxwf/galaxy-tool-cache 1.10.1, resolving `__SAMPLE_SHEET_TO_TABULAR__` for step `Project sample sheet to tabular`. `galaxy-tool-cache list` held only the two Tool Shed wrappers earlier iterations had cached. The step plan pinned identity but not version (`tool_version: TODO`). Bare `add __SAMPLE_SHEET_TO_TABULAR__` failed twice over — TRS versions 500, then `/versions/_default_` 404 — and reported "Failed to fetch tool", which reads as "not available" rather than "version unresolved". `add __SAMPLE_SHEET_TO_TABULAR__ --tool-version 1.0.0` succeeded immediately and returned the full summary; `--tool-version 0.1.0` 404'd. `GET /api/tools/__SAMPLE_SHEET_TO_TABULAR__/versions` (the non-TRS listing) returns "No route for" — checked directly, so no listing fallback exists. The cached summary then let `draft-validate --concrete` validate the step's state for real (7 ok, 0 fail, 1 skip), rather than skipping it.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
summarize-paper-declares-no-freeform-summary-shapemajorgapsummarize-paper

The Mold produces the shared freeform-summary handoff but specifies no shape for it. The procedure says only "a free-form Markdown summary capturing the workflow's steps, tools, parameters, and sample/reference-data leads", the cast bundle packages no template, example, or reference, and the runtime notes forbid reading Foundry source at runtime. Four downstream Molds consume this artifact (freeform-summary-to-galaxy-interface, -data-flow, -template, -test-plan), so the section structure they can rely on was guessed at, not settled by the instructions.

expected
Package a freeform-summary skeleton or worked example in the cast bundle, or state in the procedure the minimum sections every consumer can assume (source identity, candidate workflow scope, per-step tool/parameter table, reference data, sample data accessions, assumptions, open questions). interview-to-freeform-summary emits the same artifact id and should share whatever contract is written.
evidence
Bundle carries SKILL.md plus _feedback/_provenance/_verify only; refs is empty and "Load Upfront"/"Load On Demand" are both "None declared". Output section names the filename and format but no structure.
raised by
summarize-paper (phase 1)
observed at
63a3f9cf9c97 rev 2
issue
not filed
summarize-paper-silent-on-supplementary-methods-retrievalmajorgapsummarize-paper

The Mold's entire job is to extract methods from a paper, but it declares no required tools and describes no procedure for actually obtaining the text. For this run the article's main text contained no Methods section at all: every computational detail lived only in a supplementary PDF. The publisher page returned HTTP 403, the article is not in the Europe PMC open-access set, and a single fetch of the PMC article page yielded a tool list with no parameters and no pipeline. The retrieval route that worked was improvised and is not described anywhere in the bundle.

expected
State in the procedure that Science/Nature-style papers carry their Methods in supplementary materials and that the main text alone is not a sufficient source, and package a reference covering the retrieval routes worth trying in order (NCBI eutils efetch db=pmc for JATS full text, the supplementary-material file list inside that XML, Europe PMC, preprint and institutional-repository copies), plus the instruction to report inaccessible Methods as an open question rather than summarize around them.
evidence
Main-text JATS body ended at a bare "Supplementary Material" label; the full methods, tool parameters, and reference-assembly accession came only from the supplement PDF listed in that XML. Artifact open question 10 records the same for the next reader.
raised by
summarize-paper (phase 1)
observed at
63a3f9cf9c97 rev 2
issue
not filed
template-mold-packages-pattern-mocs-without-their-recipe-pagesmajorgapfreeform-summary-to-galaxy-template

The Mold packages four pattern references and all four are MOC index pages. Each names the recipe pages that carry the actual tool ids, port names and worked wiring, and none of those pages is in the bundle. The runtime notes forbid reading Foundry source, so a recipe named in a packaged MOC is unreachable at runtime — the reference resolves to a one-line description and a dead wiki-link.

expected
Add the recipe pages the template actually reaches for to the manifest, or state in the Mold's procedure that the MOCs are for naming an idiom only and that port names and tool ids must come from the exemplar artifact or from `discover-shed-tool`. The first is better; the second at least stops the bundle promising what it cannot deliver. The data-flow Mold has the same defect, filed separately as `data-flow-mold-packages-pattern-mocs-without-their-recipe-pages` — different manifest, same correction, so fixing one does not fix the other.
evidence
Three steps in this draft implement idioms the packaged MOCs name and cannot describe. `sync-collections-by-identifier` (collection MOC) is the pattern behind the three `__FILTER_FROM_FILE__` steps, but with no recipe page its input and output port names stayed `TODO_` sentinels. `tabular-filter-by-column-value` and `tabular-cut-and-reorder-columns` (tabular MOC) are the `Filter1` and `Cut1` steps; `Filter1`'s ports were recoverable only because the phase-4 exemplar happened to contain a worked instance, and `Cut1`'s were not recoverable at all and are recorded as assumed-by-convention in `_plan_context`. The phase-3 brief flagged the same limitation prospectively in its section 9 evidence-class caveat; this is the concrete cost at the template tier.
raised by
freeform-summary-to-galaxy-template (phase 5)
observed at
63a3f9cf9c97 rev 6
issue
not filed
test-plan-mold-ignores-available-concrete-workflowmajorgapfreeform-summary-to-galaxy-test-plan

The procedure's "Labels and fixtures are assumed, not bound" section instructs the Mold to bind assertions to interface-brief labels with `label_status: assumed` and `workflow.label_source: interface-brief`, and to record fixtures as `storage: unresolved` with `location: null`, on the stated premise that the concrete workflow and the resolved test-data refs "exist in the harness run-state by the time the plan is authored, but they are reconciled downstream rather than here". In this run both were supplied as phase-8 inputs and both were settled: `galaxy-workflow.gxwf.yml` (27 steps, 9 inputs, 16 outputs) carries the real labels, and `test-data-refs.json` carries resolved URLs, md5s and a measured verdict. Following the instruction would have meant writing `assumed` over labels read byte-for-byte from the workflow and `unresolved` over fixtures with pinned NCBI URLs and md5s — discarding verified information and handing implement-galaxy-workflow-test a reconciliation job already done. The Mold was deliberately disobeyed on both counts, and the deviation had to be argued inside the artifact rather than settled by the procedure.

expected
Make the premise conditional rather than absolute. State that when a concrete workflow or a resolved test-data-refs artifact is available to the invocation, the plan binds to it and records `label_status: resolved` / `label_source: draft` and real fixture storage; the `assumed` / `unresolved` path is the fallback for the template-era case the section describes. Declaring `galaxy-workflow-gxformat2` and `test-data-refs` as optional consumed artifacts would make that explicit in the bundle instead of leaving it to the runtime to notice.
evidence
<run>/galaxy-test-plan.yml `workflow.notes` records the deviation and its reasoning; every `label_status` in the plan is `resolved` and `workflow.label_source` is `draft`. The interface brief had itself drifted from the draft (open requirement `interface-brief-output-and-parameter-surface-drifted-from-draft`), so binding to the brief would have produced labels that do not exist in the workflow under test.
raised by
freeform-summary-to-galaxy-test-plan (phase 8)
observed at
63a3f9cf9c97 rev 2
issue
not filed
testability-note-silent-on-sample-sheet-test-fixturesmajorgapgalaxy-workflow-testability-design

Section 5 ("Design inputs with fixtures in mind") instructs the designer to match workflow input collection types to realistic fixture shapes and to choose labels readable as test `job:` keys, and its evidence covers `list`, `list:paired`, data, string, boolean and int inputs. It says nothing about the `sample_sheet` family, even though the sibling note galaxy-sample-sheet-collections is packaged in the same bundle and is what pushes the designer toward that shape. Nothing in the bundle states whether a `sample_sheet:paired` workflow input — element identifiers plus per-row typed `columns` plus collection-level `column_definitions` — can be expressed in a Planemo/IWC `-tests.yml` job block at all. This run's most consequential interface decision, the primary reads input, had to be made without knowing whether the same run's own test phase can express it.

expected
Add a rule and evidence to section 5 covering sample_sheet-family inputs: either a corpus or Planemo-schema citation showing the job-block fixture syntax for `column_definitions` and per-row `columns`, or an explicit statement that no such fixture form exists yet and that a sample_sheet input therefore trades testability for metadata carriage. Either answer settles the decision; the absence of both leaves it a guess. The sample_sheet note's "Edges to flag" list would be the natural second home for the same statement.
evidence
Bundle packages galaxy-sample-sheet-collections (which documents the workflow YAML `column_definitions` form) and galaxy-workflow-testability-design (which documents fixture design) with no overlap on this point; the note that would cover it, iwc-test-data-conventions, is cross-referenced from section 5 but is not packaged. The interface brief at <run>/freeform-galaxy-interface.md settles on `sample_sheet:paired` and carries open-requirements entry `sample-sheet-input-test-fixture-expressibility` plus a named `list:paired` fallback precisely because the question could not be answered.
raised by
freeform-summary-to-galaxy-interface (phase 2)
observed at
63a3f9cf9c97 rev 3
issue
not filed
tests-format-schema-rejects-two-shapes-the-iwc-corpus-usesmajordefectgalaxy-tool-util-ts — tests-format schema / gxwf validate-tests

The tests-format schema rejects two shapes that production IWC workflow tests use and that Galaxy accepts at run time. (1) The inner element of a `list:paired` job input written as `class: Collection` + `type: paired` — the `Collection` `$def` is `additionalProperties: false` with only `collection_type`, so `type` is an unknown property and the whole job input fails `oneOf`. (2) A nested-collection output assertion whose outer `element_tests` entry carries `elements:` without a `class: Collection` discriminator — the schema's `if/then` on `class` routes it to the dataset-element model, which has no `elements` property. Run over the pinned IWC corpus (fe41a79) with `gxwf validate-tests` from @galaxy-tool-util/cli 1.10.1, 31 of 122 committed `*-tests.yml` files fail, and the two clusters above account for the bulk of them. This makes the static gate unusable as a conformance check against the corpus the Foundry treats as normative, and it silently invalidates the corpus-derived recipes the Mold's own packaged notes teach.

expected
Accept `type` as an alias for `collection_type` on a nested `Collection` in a job block, and allow `elements:` on a collection element assertion without requiring an explicit `class: Collection`, matching what the Galaxy job-block loader and the test-format runner actually accept. If the strict form is deliberate, the schema should say so and Galaxy's Pydantic models (galaxyproject/galaxy, the source these are generated from) plus the IWC corpus should be migrated together, rather than leaving a validator that fails a quarter of the published corpus.
evidence
Probed directly. A two-file probe differing only in `type: paired` versus `collection_type: paired` gives 48 schema errors versus OK. Removing `class: Collection` from one outer `element_tests` entry of this run's own test file gives `/0/outputs/Trimmed reads/element_tests/AR0382_A: must NOT have additional properties`. Corpus sweep at fe41a79: 91 pass, 31 fail; failures cluster on `/0/job/<label>/class` (e.g. sars-cov-2-pe-illumina-wgs-variant-calling, three VGP Hi-C workflows, generic-variant-calling-wgs-pe) and on `/0/outputs/<label>/element_tests` (e.g. scrna-seq-fastq-to-matrix-10x-cellplex, metagenomic-raw-reads-amr-analysis).
raised by
implement-galaxy-workflow-test (phase 9)
observed at
63a3f9cf9c97 rev 8
issue
not filed
tonative-shape-sniff-skips-normalization-on-list-form-format2majordefect@galaxy-tool-util/schema - toNative normalization guard (_isNormalizedFormat2)

`toNative` mistakes a raw format2 workflow for an already-normalized one whenever `inputs:` and `steps:` are written in list form, skips normalization, and then aborts with an uncaught `TypeError: step.in is not iterable` as soon as a step's `in:` uses the mapping form gxformat2 equally permits. `_isNormalizedFormat2` decides on three top-level facts that say nothing about per-step shape - `class === "GalaxyWorkflow"`, `Array.isArray(inputs)`, `Array.isArray(steps)` - so `normalizedFormat2`, whose `normalizeStepIn`/`normalizeStepOut` handle both spellings correctly, is never called and `_extractConnections` iterates a mapping. The defect is in `toNative`, not in the connection validator that surfaced it: `gxwf convert --to native` and `ensureNative` crash on the same input with no connection flag involved. For this run the consequence is that `gxwf validate --connections` - the only static gate covering connection types, collection algebra and map-over - could not be run at all.

expected
Normalize unconditionally. `normalizedFormat2` is idempotent and performs plain object shaping rather than a schema decode, so the guard bought nothing; teaching it to inspect every step's `in`/`out` would cost more than normalizing and would leave the same class of bug waiting on the next field. Whatever the shape of the fix, a validator crashing on its own supported input format gives a harness no way to distinguish a broken tool from a broken workflow. Submitted upstream with a regression test as jmchilton/galaxy-tool-util-ts#179.
evidence
First seen on gxwf 1.10.1 in this run's phase 10; reproduced unchanged on 1.12.0 (both `@galaxy-tool-util/cli` and `@galaxy-tool-util/schema`), the current release as of 2026-09-17, so it is not a stale-CLI artifact. Trigger is one corner of the shape matrix, established on 20-line workflows with a single `Filter1@1.1.1` step: list `inputs:` + list `steps:` + mapping `in:` crashes; the same workflow with list-form `in:`, with map-form `inputs:`, or with map-form `steps:` converts cleanly, because those spellings fail the sniff and get normalized. The earlier claim in this entry that the flag breaks on every format2 workflow with a tool step was too broad - it breaks on the dialect this Foundry emits. `gxwf convert <run>/galaxy-workflow.gxwf.yml --to native` crashes identically with no `--connections`, at `runConvert` in `cli/dist/commands/convert.js:83`. Upstream's own suite never crosses the guard: every existing `toNative` test writes `inputs`/`steps` in map form. With the guard removed, all 4999 `packages/schema` tests pass and the three new list-form cases go from crash to green. Workaround available to the Foundry without waiting on a release: emit `in:` in list form (`- id: <port>` / `source: <ref>`).
raised by
validate-galaxy-workflow (phase 10)
observed at
79bf5c3ab98e rev 5
issue
https://github.com/jmchilton/galaxy-tool-util-ts/pull/179
tool-summary-drops-output-filter-expressionsmajorgapgalaxy-tool-summary parsed_tool.outputs

`parsed_tool.outputs` carries name, label, hidden, type, format, format_source, metadata_source, discover_datasets, from_work_dir and precreate_directory -- but not the output's `<filter>` expression. A Galaxy output filter is what decides whether a declared output EXISTS for a given parameter branch, so the summary can enumerate ten outputs for a tool that will produce four, with nothing marking the difference. That is not cosmetic for this Mold: a step's `out:` list and every workflow output wired to it are only valid if the chosen tool state keeps those outputs alive, and the summary is the artifact procedure step 3 produces expressly so step 4 can bind the step against it. Concretely, this run's STAR step had to establish that `output_log` and `mapped_reads` survive `quantMode: '-'` and `outWigType: None`, because its plan makes a suppressed `output_log` a hard failure -- and the summary cannot answer that question at all. The answer (`reads_per_gene` and `transcriptome_mapped_reads` are filtered on quantMode; the two promoted outputs carry no filter) came only from reading the wrapper XML. Iteration 8 hit the same wall on lparsons/cutadapt, where the `report` and `out_pairs` filters are what keep two promoted outputs alive.

expected
Carry the raw `<filter>` expression string on each output in `parsed_tool.outputs` (a nullable `filter` field is enough; the Mold does not need it evaluated, only visible), and have implement-galaxy-tool-step's procedure say that before finalizing a step's `out:` list the author must check each promoted output's filter against the state just bound. Without the field the correct instruction is unfollowable, because the evidence is not in the artifact the procedure hands the author.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
tool-summary-whens-order-is-the-current-case-index-and-nothing-says-somajorgapgalaxy-tool-summary conditional whens

A conditional's `__current_case__` is an INDEX, and nothing in the packaged summary schema, the Mold, or any packaged note says what it indexes. The summary gives each conditional a `test_parameter.options` array and a `whens` array with a `discriminator` per element. The index that `__current_case__` must carry is the position in `whens` (i.e. `<when>` document order), NOT the position of the matching value in `options`. The two are not interchangeable and this run hit a live divergence: in iuc/rgrnastar/rna_star 2.7.11b+galaxy1, `refGenomeSource[history].GTFconditional` lists its options as `without-gtf, with-gtf` but its whens as `with-gtf, without-gtf`, so `with-gtf` is case 0 while the dropdown shows it second. An author reading the field an author would naturally read gets 1. The bundle offers no way to know which array is authoritative, so the rule had to be re-established from outside the bundle -- by round-tripping a corpus `.ga` that happens to use the same tool and reading the wrapper XML's `<when>` order to confirm the correspondence.

expected
State once, in the summary schema's own description of `whens` (and echo it in implement-galaxy-tool-step's state-authoring guidance), that `whens` is emitted in `<when>` document order and that a conditional's `__current_case__` is the zero-based index into that array -- explicitly warning that it may differ from the option order, with a one-line example. Better still, emit the index on each `whens` element so the author never has to count, which also makes the contract checkable rather than conventional. This is a third axis of the same family already filed here: `implement-step-mold-says-state-without-saying-which-key` is about which block key, `no-authority-for-scalar-value-encoding-in-tool-state` is about the scalar leaf value, and this one is about the bookkeeping key's VALUE. The `__index__` key on repeat entries has the same problem and the same fix.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
validate-cli-note-recommends-a-flag-that-crashes-and-omits-the-cache-trapmajordefectgxwf validate

The note's Gotchas section warns about the one flag that weakens validation visibly and is silent on the two ways it fails invisibly. It says `--no-tool-state` weakens validation, but never says that omitting `--cache-dir` has the same effect and worse - every tool step becomes a skip and the command still exits 0. `--cache-dir` appears only as a bare option line, "Tool cache directory", with nothing about what happens without it. Separately, the note actively recommends a flag that cannot run: "Use `--connections` when tool cache metadata is available and data-shape compatibility matters, especially around collections and map-over", and its Examples block lists `gxwf validate workflow.gxwf.yml --json --connections --strict` - a command that at 1.10.1 crashes on any format2 workflow with a tool step, and would abandon JSON even if it did not.

expected
Add a Gotchas line making the cache trap explicit - without `--cache-dir` pointing at a populated cache, `validate` reports every tool step as skipped and still exits 0, so the split must be read rather than the exit code. Mark `--connections` as not working at 1.10.1 with a pointer to the upstream defect, and drop or flag the `--connections --strict` example so the note stops recommending a crash. Record which gxwf version the page was verified against; the page carries none today, and `gxwf --version` self-reports 1.0.0 regardless.
evidence
gxwf 1.10.1, this run's phase 10. Following the note's own recommendation was the first thing attempted and it exited 1 with `TypeError: step.in is not iterable` and no report; see `tonative-shape-sniff-skips-normalization-on-list-form-format2`. The `--strict` half of the same example returns exit 2 with empty stdout; see `gxwf-validate-json-output-is-not-machine-parseable`. The cache trap is the failure phase 6 iteration 26 hit on its first terminal validate: `0 validated, 27 skipped`, exit 0.
raised by
validate-galaxy-workflow (phase 10)
observed at
79bf5c3ab98e rev 5
issue
not filed
validate-mold-terminal-pass-has-no-cache-and-no-skip-vocabularymajorgapvalidate-galaxy-workflow

The Mold owns the run's last automated gate and its procedure never mentions the tool cache, never mentions skips, and gives the artifact a three-valued status - `pass`, `fail`, `not-run` - with no value for the outcome this gate actually produces. Two consequences, both hit in this run. (1) The invocation the procedure implies, `gxwf validate <file> --json`, omits `--cache-dir` and returns `Tool state: 0 validated, 27 skipped` with exit code 0: a green that checked nothing, on the terminal gate, with no warning anywhere in the bundle. (2) When the cache is supplied, seven steps still skip permanently, and the Mold offers no way to say so - `pass` overstates it, `not-run` understates it - and no route to discharge a skip, although one exists and is cheap. The Mold's one instruction on this, "A `not-run` status is never reported as a pass", guards the case that cannot happen quietly and not the one that can.

expected
Three changes to the procedure. Declare the tool cache as part of the contract and make `--cache-dir` explicit in the invocation, stating that omitting it turns every tool step into a skip and still exits 0. Make the artifact's status carry coverage: either a fourth value for a pass with unvalidated steps, or a required `validated`/`skipped` split alongside `status`, so a downstream reader cannot cite the green without the number. And give the skip a discharge route rather than leaving it to eyeball review: a skipped step's tool can be fetched from a Galaxy instance at `GET /api/tools/<tool_id>?io_details=true` and its `in:` keys and tool_state parameter names checked against the real input tree, which is what this phase did by hand and what the procedure nowhere describes. Sibling of `advance-draft-mold-needs-a-tool-cache-it-never-declares` and `advance-draft-mold-treats-a-skipped-tool-state-as-green`, which found the same two holes in the per-step loop Mold; this is the terminal gate, where they cost more.
evidence
This run, phase 10. Only the harness brief - not anything in the cast bundle - carried the `--cache-dir` requirement and the expected skip list. With the cache, `gxwf validate <run>/galaxy-workflow.gxwf.yml --json --cache-dir <cache>` returns `ok: 20, fail: 0, skip: 7` and exit 0; phase 6 iteration 26 recorded that the same command without it returns `0 validated, 27 skipped`, also exit 0. The seven skips are `__FLATTEN__`, `lparsons/cutadapt`, three `__FILTER_FROM_FILE__` and two `iuc/deseq2`, none clearable by priming the cache (see `collection-output-decode-is-a-flat-vs-nested-shape-mismatch-not-a-missing-field`). All seven were discharged here against the live Galaxy tool API with zero mismatches, in one pass.
raised by
validate-galaxy-workflow (phase 10)
observed at
79bf5c3ab98e rev 5
issue
not filed
advance-draft-mold-reresolves-a-wrapper-already-pinned-in-the-draftminorfrictionadvance-galaxy-draft-step

The procedure's resolve-then-summarize sequence is written as if each iteration met its wrapper for the first time. It has no notion of a wrapper already resolved earlier in the same run: this draft uses one wrapper at five steps, and the identity-pinned branch still directs the iteration to confirm the pin via discover-shed-tool and then invoke summarize-galaxy-tool, both of which a sibling step's already-concrete `tool_id` + `tool_version` pair has settled. The Mold never describes what this iteration actually did, which was to take the sibling's pin and read the summary already in the cache.

expected
Add a short third case to the resolve step: when another step of this same draft is already concrete on the same `tool_id`, adopt its `tool_version` and skip discovery, re-using the cached summary rather than re-summarizing. Say plainly that this is the intended move, so an iteration that takes it is following the procedure rather than departing from it. The Mold should also state that a per-run resolved-wrapper reuse is safe precisely because the pin is recorded in the artifact, not in operator memory.
evidence
Iteration 2 of this run, step `Build first-contrast row predicate`. Iteration 1 had already resolved `iuc/compose_text_param` to 0.1.1 against the live Tool Shed for `Build adjusted p-value predicate`, and written the pin into the draft. Three more steps of the same draft await the same wrapper.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
advance-draft-mold-silent-on-plan-provenance-at-concretionminorgapadvance-galaxy-draft-step

Concretizing a step forces its `_plan_*` fields to be deleted — `draft-next-step` counts any surviving `_plan_*` as remaining work, so a finished step that keeps one is selected again on the next iteration and the harness loop cannot terminate. Those fields are also where the corpus provenance lives: the step implemented here carried a `_plan_context` citing the exemplar workflow and step that fixes its component text. The Mold says only that the implement phase "resolves the chosen step's remaining `TODO_*` / `_plan_*` slots", and never says whether that provenance should be carried into `doc:`, dropped, or recorded elsewhere. This run has been keeping it in a YAML comment above the step, a convention it invented, and nothing states whether `draft-extract` preserves comments when it re-serializes the concrete workflow (this iteration did not test that).

expected
Say in the implement step what becomes of a concretized step's planning provenance: fold the load-bearing part of `_plan_context` into the step's `doc:` (which survives extraction and is visible to a Galaxy user), and drop the rest. State plainly that `_plan_*` must not survive concretion, and why — `draft-next-step` would re-select the step forever. If YAML comments are in fact preserved by `draft-extract`, say so and sanction the comment form; if they are not, say that too, so a run does not park provenance somewhere that silently disappears at loop endstate.
evidence
Iteration 3, step `Build log2 fold-change predicate`. Its `_plan_context` cited `transcriptomics/rnaseq-de` at corpus fe41a79, step `_unlabeled_step_10`, as the source of the `abs(c3)>` component text; the step's post-implementation `doc:` and the YAML comment above it were written by hand to retain that, with no instruction either way. `gxwf draft-next-step` lists every `_plan_*` field of the selected step under `work`, which is what makes their removal mandatory rather than stylistic.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
asserts-idioms-nested-element-tests-recipe-omits-the-class-discriminatorminorgapplanemo-asserts-idioms

Section 5 is the note an agent reaches for when writing collection-output assertions, and its nested-collection rule — "outer `element_tests:` keyed by outer identifier; inner `elements:` (note plural, no `_tests` suffix on the inner)" — is incomplete in the one way that matters to the gate: the outer entry must also carry `class: Collection`, or the schema routes it to the dataset-element model and rejects `elements` as an additional property. The section also does not mention that an `asserts:` mapping cannot repeat a key, so any output needing two `has_text` probes has to use the list form (`- that: has_text`) — which this run needed on five outputs and which the section's own examples never show.

expected
Add `class: Collection` to the nested example in section 5 and say why it is required. Add a line to section 4 or 9 stating that repeated assertion families require the `- that: <family>` list form, since the dict form silently loses all but the last occurrence in YAML.
evidence
`Trimmed reads` in <run>/galaxy-workflow.gxwf-tests.yml is a `list:paired` output needing the nested form; without `class: Collection` on the outer entry `gxwf validate-tests` reports `must NOT have additional properties` at `/0/outputs/Trimmed reads/element_tests/AR0382_A`. Four of this run's outputs carry between four and six assertions of which two or more share a family, which the dict form cannot express.
raised by
implement-galaxy-workflow-test (phase 9)
observed at
63a3f9cf9c97 rev 8
issue
not filed
cast-validation-fallback-resolves-an-unpinned-published-validatorminorfrictioncast-mold caster — generated Validation section

The caster emits a Validation instruction of the form "run `foundry <cmd> <artifact>` from `@galaxy-foundry/gxwf-foundry`; if the command is not on PATH, run `npx --package @galaxy-foundry/gxwf-foundry foundry <cmd> <artifact>`". The fallback names no version, so it resolves whatever `latest` is on the registry at runtime, while the bundle already carries the schema it was cast against, verbatim, under `references/schemas/`. Those two can disagree, and when they do the fallback returns a confident green verdict against a contract the skill is not bound by. In this run the fallback was the only available route — the checkout had no installed dependencies and no package manager on PATH — and the published 0.1.2 schema had to be diffed against the bundled copy by hand to know the verdict meant anything, which the runtime notes ("do not read Foundry source files at runtime") discourage in the first place.

expected
Record the validator package version in `_verify.json` at cast time and emit it in the fallback (`npx --package @galaxy-foundry/gxwf-foundry@<version> ...`), so the fallback validates against the same contract the bundle carries. Alternatively, or additionally, say in the generated Validation line that `references/schemas/<name>.schema.json` in the bundle is the binding contract and that a fallback verdict is only meaningful if the two agree.
evidence
`_verify.json` for this Mold carries `validator_bin: foundry` and args, with no package version anywhere in the bundle. The published schema happened to be byte-identical to `references/schemas/galaxy-workflow-test-plan.schema.json` here, so the verdict stands, but nothing in the bundle establishes that and nothing would have flagged it had they diverged.
raised by
freeform-summary-to-galaxy-test-plan (phase 8)
observed at
63a3f9cf9c97 rev 2
issue
not filed
design-briefs-duplicate-open-questions-into-ledger-entriesminorfrictionopen-requirements-ledger

The Mold's output contract requires an "open questions" section in the Markdown brief and a ledger of open entries, with no rule for which destination an unresolved choice belongs in. Every unresolved item in this run was genuinely both an obligation a later Mold must discharge and a thing a human reviewer should see, so all ten were written twice, in two different shapes, and cross-referenced by entry id by hand to stop them drifting. Confirming a real-run instance of something the note already lists under Open work ("Reconcile the design-tier briefs' free-text 'open questions' sections with the ledger"); filed as corroboration rather than as a new finding.

expected
Pick one source of truth and say so. The cheapest version that would have helped here: make the ledger authoritative and specify the brief's open-questions section as a rendered index of the entries it raised, one line each keyed by entry id, rather than an independently authored list.
evidence
Interface brief section 5 carries ten numbered open questions, each ending in the entry id it mirrors; the ledger carries the same ten as structured entries. The duplication is manual and nothing checks it.
raised by
freeform-summary-to-galaxy-interface (phase 2)
observed at
63a3f9cf9c97 rev 3
issue
not filed
discovery-schema-cannot-express-an-unresolved-alternateminorgapgalaxy-tool-discovery

`alternates[]` reuses the full `ToolCandidate` shape, which requires `version` and `changeset_revision` (both `minLength: 1`) plus a numeric `score`. There is no way to record a candidate the discovery deliberately did NOT pin. The procedure asks for exactly that — "multiple plausible hits ... → `weak` with the leading candidate plus alternates", and the surrounding Molds hand this skill named alternatives in a step's `_plan_context` — but an alternate is only representable after it has been fully resolved to a changeset. So an author recording a runner-up faces two bad options: spend Tool Shed calls pinning a wrapper being rejected, or write placeholder strings. The second validates green. `"changeset_revision": "unresolved"` and `"score": 0` pass `validate-galaxy-tool-discovery` without a murmur, in a field the schema itself documents as "Selected Tool Shed Mercurial changeset revision for reproducible gxformat2 tool_shed_repository pinning" and one documented as "Higher is better". A machine-read pin contract should not be able to carry a fabricated pin.

expected
Let an alternate be unresolved. Either make `version` and `changeset_revision` nullable on `ToolCandidate` and require them non-null only for the selected `candidate` (the existing `allOf` if/then on `status` is already the place to say so), or split a lighter `AlternateCandidate` shape carrying identity plus rationale and no pin fields. Add a `match_fields` enum value for a candidate surfaced from an upstream plan hint rather than from the lexical index — `matched_terms` already anticipates this case ("Empty only when the candidate came from a non-lexical hint") but `match_fields` has no corresponding value, so a plan-hint alternate has to claim lexical evidence it does not have or leave the array empty.
evidence
Iteration 9 of phase 6 in this run, step `Read quality report`. The step was Deferred and its `_plan_context` named `iuc/falco` as the corpus-observed alternative to `devteam/fastqc`, asking that it be recorded if not taken. Recording it cost three extra Tool Shed calls (`tool-search falco`, `tool-versions`, `tool-revisions`) for a wrapper being rejected, and the score needed its own query because `iuc~falco~falco` does not appear in the `fastqc` result set at all. The first artifact written instead used `"changeset_revision": "unresolved"` and `"score": 0`; `foundry validate-galaxy-tool-discovery` reported `galaxy-tool-pin.json: valid`. It was corrected by hand afterwards, not by the gate. Notably the schema ref's own `verification` line in the cast provenance is "Run discover-shed-tool against known FastQC, ambiguous BWA-style, and no-hit queries and validate each emitted recommendation" — this is that FastQC query, and the alternates path is what it does not exercise.
raised by
discover-shed-tool
observed at
63a3f9cf9c97 rev 5
issue
not filed
draft-format-has-no-way-to-mark-a-region-provisionalminorgapgalaxy-workflow-draft-format

The draft superset can mark a STEP as unresolved (`TODO`, `_plan_*`) but has no way to mark a REGION as provisional, or to record the named alternative a later phase should swap in. The `_plan_*` family is per-step and explicitly wrapper-tier, so a multi-step topology choice that is settled-but-unprecedented has nowhere durable to live.

expected
Decide a home for it and add it to the note — a workflow-level annotation beside the `topology_repair` budget the open-requirements note already wants moved into the draft, or an explicit statement that a titled `comments:` frame plus an open-requirements entry IS the intended mechanism. Either answer is fine; the absence of one means each template Mold run invents its own.
evidence
Phase 4 instructed this phase to build the sample-sheet condition split "as a delimited region with the phase-2 alternative one deletion away". Thirteen steps, no corpus precedent, settled topology. The delimitation was expressed three ways, none of them contractual: YAML comments (lost on any round-trip through Galaxy), a titled `comments:` frame (schema-legal and durable, but the packaged `galaxy-workflow-comments` note describes frames as narrative stage annotation, not risk marking), and a `type: markdown` comment holding the four-step swap procedure as prose. A reader of the draft alone has no typed signal that the region is provisional.
raised by
freeform-summary-to-galaxy-template (phase 5)
observed at
63a3f9cf9c97 rev 6
issue
not filed
galaxy-unit-test-documents-a-sample-sheet-rows-shape-the-api-rejectsminordefectGalaxy — test/unit/tool_util/test_cwl_util.py sample_sheet rows tests

`test_galactic_job_json_sample_sheet_collection_with_rows` asserts that `rows` round-trips as `{"el1": {"condition": "treatment"}, "el2": {"condition": "control"}}` — a mapping of column name to value. The real contract is a POSITIONAL LIST per row: `validate_row` in lib/galaxy/model/dataset_collections/types/sample_sheet_util.py rejects on `len(row) != len(column_definitions)` and then zips `row` against `column_definitions` in order. The unit test never catches this because it mocks `collection_create_func`, so nothing downstream of `galactic_job_json` is exercised. These unit tests are the most discoverable documentation of the job-block `rows` syntax, and the shape they document will not validate whenever the target collection has column definitions.

expected
Change the two `rows` unit tests to the positional-list form (`{"el1": ["treatment"], "el2": ["control"]}`), matching lib/galaxy_test/workflow/collection_semantics_cat_sample_sheet.gxwf-tests.yml which already uses lists (`rows: {el1: [], el2: []}`). Better, add an integration test that stages a sample sheet WITH column definitions through the real API, which would have caught the divergence and would also cover the gap filed as `galaxy-test-staging-drops-sample-sheet-column-definitions`.
evidence
Read at galaxyproject/galaxy dev: test/unit/tool_util/test_cwl_util.py lines ~192–219 versus `validate_row` in lib/galaxy/model/dataset_collections/types/sample_sheet_util.py. Encountered while deriving the job block now recorded in <run>/test-data-refs.json; the dict form was adopted from the unit test first and corrected only after reading the validator.
raised by
paper-to-test-data (phase 7)
observed at
63a3f9cf9c97 rev 2
issue
not filed
gxwf-version-flag-reports-stale-versionminordefect@galaxy-tool-util/cli — gxwf --version

`gxwf --version` prints `1.0.0` regardless of the installed package version. The version actually installed here is `@galaxy-tool-util/cli@1.10.1`, confirmed by `npm ls -g`. Every Foundry Mold that requires gxwf pins a package version in `_required_tools.json` — this one pins `^1.8.1` — and none of those pins can be checked at runtime, because the only version the CLI reports is a constant that satisfies no pin and matches no release. The packaged availability check works around this by grepping `--help` for a subcommand name, which detects presence but says nothing about version.

expected
Have `gxwf --version` report the package version from package.json. Until it does, a Mold that needs version-sensitive behaviour has no runtime signal, and a run that records its tool versions as provenance records a number that is always `1.0.0`.
evidence
`gxwf --version` → `1.0.0`; `npm ls -g @galaxy-tool-util/cli` → `@galaxy-tool-util/cli@1.10.1`; this Mold's `_required_tools.json` pins `package_version: ^1.8.1` with `availability_check: gxwf --help | grep -q draft-validate`. Conversions in this phase succeeded, so the defect is in version reporting only, not in the tool's behaviour.
raised by
compare-against-iwc-exemplar (phase 4)
observed at
63a3f9cf9c97 rev 10
issue
not filed
interface-mold-gives-no-rule-for-parameter-exposureminorgapfreeform-summary-to-galaxy-interface

The procedure names inputs, outputs, labels, collection shapes and checkpoints as the things to settle, but gives no rule for which source-stated parameters become exposed typed workflow inputs and which are baked into step defaults. The only nearby guidance is one line in the packaged testability note ("Keep typed parameters explicit when tests need to set them"), which answers the test-facing half and not the design half. For a paper source the distinction is sharp and recurring: a value the paper pins is settled and can be baked, while a value inferred from a kit name or an accession list is exactly what a reviewer must be able to see and change. That rule was invented here, not read.

expected
State the rule in the procedure: expose as a typed workflow parameter any value the source leaves unstated or that this brief infers, so the inference is visible and overridable, plus any value a test must set; bake values the source pins. Applies equally to the Nextflow and CWL interface Molds, whose sources pin parameters far more often than a paper does.
evidence
Interface brief section 2.3 has to justify its own exposure policy from scratch: strandedness exposed because inferred, Cutadapt `-q 20` and RNA STAR defaults baked because stated, significance thresholds exposed because stated but test-relevant. Three different justifications for one decision class, none of them sourced from the bundle.
raised by
freeform-summary-to-galaxy-interface (phase 2)
observed at
63a3f9cf9c97 rev 3
issue
not filed
no-authority-for-scalar-value-encoding-in-tool-stateminorgapimplement-galaxy-tool-step

Nothing the Mold names settles how a scalar parameter's VALUE is encoded in the step's state block, and the two authorities an author can reach disagree. Procedure step 2 sends the author to `parsed_tool` for ports and datatypes and to `input_schemas.workflow_step_linked` for the state shape. For `Filter1`'s `header_lines`, `parsed_tool` declares `parameter_type: gx_integer`, `type: integer`, `value: 0` — read literally that says write the YAML integer `0`. Every real Galaxy workflow writes the string `'0'`. This is a different axis from the two disagreements already filed here: `implement-step-mold-says-state-without-saying-which-key` is about WHICH block key, and `gxwf-linked-step-schema-rejects-the-keys-its-own-converter-emits` is about `__`-prefixed bookkeeping keys. This one is about the scalar leaf value, and it is unaddressed by either correction. The Mold's guidance is not merely silent — following the one authority it names produces the form the corpus never uses.

expected
Say in procedure step 2 that a wrapper's declared parameter type fixes the SEMANTICS of a state value, not its serialization, and that the authoring form for scalar leaves is what `gxwf convert --to format2` emits from a real workflow — strings for integer, float, and select parameters alike, because Galaxy's `tool_state` is JSON-string-encoded at the source. That is the same authority the run already had to fall back on for conditional and repeat bookkeeping, so naming it once covers all three axes. If the Foundry would rather not carry that rule, say instead that either encoding is accepted and that the choice is a consistency matter within a draft — but say which, because an author who reads only `parsed_tool` currently has no way to find out and no gate that will tell them.
evidence
Phase 6 iteration 10 of this run, step `Select first-contrast samples` (`Filter1` 1.1.1). The cached summary declares `header_lines` as `gx_integer` with `value: 0`. A survey of all 21 `Filter1` states in the pinned IWC corpus found the value serialized as a string in every one ('0' x14, '1' x7) and as an integer in none; `gxwf convert --to format2` of `transcriptomics/rnaseq-de/rnaseq-de-filtering-plotting` likewise emits `header_lines: "1"`. `gxwf draft-validate --concrete` was then run twice over otherwise byte-identical drafts, once with `header_lines: '0'` and once with `header_lines: 0`, against a cache holding the real `Filter1` summary. Both returned `draft valid` / `Concrete: OK` / `Tool state: 9 ok, 0 fail, 2 skip` — identical verdicts, so the gate decides nothing here and an author cannot resolve the question by trying it. The binding was settled by corpus frequency alone, which is exactly the guess the Mold should have removed.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
test-plan-schema-has-no-evidence-class-for-run-measured-assertionsminorgapgalaxy-workflow-test-plan

`AssertionIntent.evidence` is a two-value enum, `test-evidence` or `intent`, described as "whether this assertion was translated from upstream test evidence or synthesized from intent". A third case is routine on the paper and interview paths and occurred throughout this run: an assertion neither translated from an upstream fixture nor merely synthesized, but MEASURED against the run's own resolved test data before the plan was written. The same gap exists at plan level, where `source.derived_from` has the same two values. Forced to pick, every such assertion is recorded `evidence: intent`, so a reviewer filtering on the field sees a uniformly speculative plan and systematically undercounts its grounding. `confidence: high` is the only signal left, and it means something different.

expected
Add a third value — `run-measured` (or `fixture-measured`) — to `AssertionIntent.evidence` and to `source.derived_from`, meaning the value was obtained by measuring the run's own resolved fixtures rather than read from upstream tests or inferred from intent. Nothing else in the schema needs to change, and the existing two values keep their meaning.
evidence
<run>/galaxy-test-plan.yml records the same complaint twice because the field cannot carry it: `source.notes` argues that `derived_from: intent` understates the plan's grounding, and warnings[] `fixture-measured-assertions-lack-an-evidence-class` states the consequence. Roughly twenty assertions in the plan rest on direct measurement of the resolved fixtures by phase 7 (per-sample SCF1 fragment counts, strandedness, mapping rate, annotation gene count) or by phase 8 (DESeq2 log2 fold change and adjusted p-value per contrast), and all of them are filed as `intent`.
raised by
freeform-summary-to-galaxy-test-plan (phase 8)
observed at
63a3f9cf9c97 rev 2
issue
not filed
tool-cache-records-a-shed-path-for-a-bare-stock-tool-idminordefectgalaxy-tool-util-ts / galaxy-tool-cache (@galaxy-tool-util/cli 1.10.1)

When `galaxy-tool-cache add` caches a stock Galaxy tool by its bare id, it writes the id into `index.json` with a Tool Shed repository path glued on the front, producing `toolshed.g2.bx.psu.edu/repos/__SAMPLE_SHEET_TO_TABULAR__` — an identifier that names nothing: there is no such shed repository, and no workflow may legally carry that tool id. The cached summary body alongside it is correct (`"id": "__SAMPLE_SHEET_TO_TABULAR__"`), so the damage is confined to the index, but the index is what `galaxy-tool-cache list` prints. That matters because `list` is the surface the advance-galaxy-draft-step procedure directs an author to read a stock tool's version off, and what it shows there cannot be pasted into a draft.

expected
Preserve a bare stock tool id verbatim in the cache index, as the summary body already does. A tool id with no `owner/repo` path is not a shed-relative name and should not have `toolshed.g2.bx.psu.edu/repos/` prefixed to it.
evidence
Iteration 7 of phase 6 in this run. After `galaxy-tool-cache add __SAMPLE_SHEET_TO_TABULAR__ --tool-version 1.0.0 --cache-dir <run-scratch>/gxwf-cache`, `index.json` records `"tool_id": "toolshed.g2.bx.psu.edu/repos/__SAMPLE_SHEET_TO_TABULAR__"` against `"source_url": "https://toolshed.g2.bx.psu.edu/api/tools/__SAMPLE_SHEET_TO_TABULAR__/versions/1.0.0"`, while the summary file it points at carries `"id": "__SAMPLE_SHEET_TO_TABULAR__"`. `galaxy-tool-cache list` prints the prefixed form. Lookup at validation time is evidently keyed off the summary body rather than the index, since `draft-validate --concrete` resolved the step against the cache and reported it ok.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed
tool-search-default-page-returns-every-hit-three-timesminordefect@galaxy-tool-util/cli — gxwf tool-search result de-duplication

With no `--max-results`, `gxwf tool-search` returns every hit three times — in the table rendering and in `--json` alike. Passing any explicit `--max-results` returns distinct hits, so the duplication is confined to the default page size. It is not cosmetic for a discovery procedure that triages by counting and comparing candidates: the default view of a search makes a field of twenty wrappers look like sixty, and the triage rule "multiple plausible hits ... → weak" reads a repeated single candidate as a cluster.

expected
De-duplicate on `(repoName, repoOwnerUsername, toolId)` before returning, on the default page size as well as an explicit one — or, if the repetition encodes distinct revisions, surface the revision that distinguishes the rows instead of emitting rows that are byte-identical.
evidence
Iteration 8 of phase 6 in this run, gxwf 1.10.1. `gxwf tool-search "cutadapt" --json` returns 50 `trsToolId` values of which 20 are distinct, each repeated three times; the table rendering of the same query shows the same triples. `--max-results 5`, `--max-results 10` and `--max-results 20` each return exactly N values, all distinct.
raised by
advance-galaxy-draft-step (phase 6)
observed at
63a3f9cf9c97 rev 4
issue
not filed

Timeline

how the draft grew

no checkpoint history — re-run with --checkpoint to get a per-phase and per-iteration record.

Everything else

files no Mold declares

13 of 23 files map to a declared artifact. The rest are below. 3 path(s) were ignored entirely.

sub-mold1 — declared by a Mold a phase invokes internally rather than by a phase itself
pathsizemodifieddeclared by
galaxy-tool-pin.json2.5 KB2026-09-16 17:59discover-shed-tool
tool-output7 — output of a tool the run drove, not an artifact any Mold declares
pathsizemodifieddeclared by
planemo-biocontainers.log369.4 KB2026-09-18 16:41
planemo-smoke.attempt1.log317.3 KB2026-09-16 22:09
planemo-smoke.html340.4 KB2026-09-17 14:26
planemo-smoke.json14.8 KB2026-09-17 14:26
planemo-smoke.log290.7 KB2026-09-17 14:26
tool_test_output.html340.4 KB2026-09-18 16:41
tool_test_output.json14.8 KB2026-09-18 16:41
narrative2 — written about the run rather than by it
pathsizemodifieddeclared by
CASE_STUDY.md15.1 KB2026-09-18 16:36
README.md3.9 KB2026-09-18 16:39
directory1 — summarized, never walked
pathsizemodifieddeclared by
test-data24 file(s)2026-09-18 16:38