Create bounty

Funded scientific challenge

Open

Head-to-head benchmark: baseline whole-blood expression features vs clinical serostatus for predicting anti-TNF response in rheumatoid arthritis

Run a fully reproducible head-to-head comparison on a pinned public cohort of biologic-naive rheumatoid arthritis patients: a pre-specified clinical covariates model (comparator) against the same model augmented with baseline whole-blood expression features, under a frozen train/external-validation split. Report discrimination, calibration and decision metrics separately, and ship a reusable benchmark harness. The purchased result is the comparison itself; a finding that the expression layer adds no value is a valid outcome, not a failure.

Submission deadline
Judging deadline
Settlement timeout
On-chain record
View bounty creation

Elgora recalculated the exact challenge Markdown bytes and confirmed they match the commitment stored on ElgoraHub at funding.

Hash method: Keccak-256 of exact UTF-8 Markdown bytes

On-chain commitment0x9070fa3777051022dfe7bc526c3f1caa25153c0012c8437e42764cce66334423
Challenge matches the fingerprint recorded when this bounty was funded.

Committed challenge

Challenge details & success criteria

The approved challenge, byte for byte as committed at funding. Solvers deliver against these sections and Guardians judge against them.

Summary

Run a fully reproducible head-to-head comparison on a pinned public cohort of biologic-naive rheumatoid arthritis patients: a pre-specified clinical covariates model (comparator) against the same model augmented with baseline whole-blood expression features, under a frozen train/external-validation split. Report discrimination, calibration and decision metrics separately, and ship a reusable benchmark harness. The purchased result is the comparison itself; a finding that the expression layer adds no value is a valid outcome, not a failure.

Challenge details

This bounty operationalizes a falsifiable question from the STORM hypothesis discussion on OpenLabs: does a genomic layer add predictive value beyond a clinical-workflow comparator, when the comparison is pre-specified and validated on data not used for model development? This is a methodological pilot on retrospective public data, not a clinical study. The outcome here is EULAR response at month 3, not toxicity or treatment discontinuation, and the data carry no ancestry, adherence, or follow-up-intensity measures.

The data are GSE129705: whole-blood RNA-seq profiles of 76 biologic-naive RA patients initiating infliximab or adalimumab, sampled at baseline and month 3, from two independent cohorts (Cohort 1 and Cohort 2), with EULAR response status (Good or None), anti-CCP serostatus and rheumatoid factor serostatus recorded per subject (Farutin et al., Arthritis Research & Therapy 2019, DOI 10.1186/s13075-019-1999-3, PMID 31647025).

The outcome is binary: EULAR response Good is the positive class (label 1), response None is the negative class (label 0), as recorded in the pinned series metadata. Only baseline samples are used as model inputs; month 3 expression must never be used as a predictor. Anti-CCP and rheumatoid factor serostatus at baseline are the only available clinical covariates in this dataset.

The frozen split: models are developed using Cohort 1 baseline samples only (34 subjects: 18 Good, 16 None); Cohort 2 baseline samples (29 subjects: 18 Good, 11 None) are the untouched external validation set. No Cohort 2 sample may influence any training, tuning, feature selection, preprocessing choice or hyperparameter choice. The decision threshold for classification metrics is fixed at 0.50 in advance.

Three model specifications must be estimated and reported:

ModelFeatures
A - serostatus-only (comparator)anti-CCP and RF serostatus at baseline
B - expression-onlybaseline expression features, selected by a stated rule
C - combinedserostatus plus baseline expression features

Everything not frozen above is the Solver's implementation choice and must be disclosed: classifier family and hyperparameters, count normalization and transformation, gene filtering and feature selection rule, and the internal cross-validation scheme for Cohort 1. Feature selection and all preprocessing statistics must be fitted only on Cohort 1 training data.

The source study reports that baseline expression differences between good responders and non-responders did not show statistically significant genome-wide concordance between the two cohorts, while cell-type composition signals did. This bounty does not purchase agreement with any published result; it purchases a transparent head-to-head comparison under the frozen protocol above.

What you need to submit (Deliverables)

Exactly four files, all plain flat files:

FileRequiredContent and formatSize limitPurpose
result.mdyesUTF-8 Markdown: results table, comparison statement, methods disclosure, limitations section100,000 bytesthe report a reader evaluates
predictions.csvyesCSV with exact header sample_title,model,split,observed_label,predicted_probability1,000,000 bytesmachine-checkable predictions for metric recomputation
run_benchmark.py or run_benchmark.Ryesthe complete analysis, runnable end-to-end from the pinned inputs1,000,000 bytesreproducibility of every reported number
environment.txtyesexact software and package versions needed to run the code50,000 byteslets the analysis be re-run in a clean environment

predictions.csv must contain one row per required combination: for each of the three models A, B and C, one out-of-fold predicted probability for each of the 34 Cohort 1 baseline samples (model one of serostatus_only, expression_only or combined; split = internal_oof_c1), and one predicted probability for each of the 29 Cohort 2 baseline samples (split = external_c2) - 189 data rows in total. sample_title must match the sample titles in the pinned inputs. observed_label is 1 for EULAR response Good, 0 for None, matching the pinned series metadata for that sample's subject. predicted_probability is a decimal number within [0.000001, 0.999999]; values outside this range are invalid. Do not include any other rows or columns.

result.md must contain, at minimum:

  1. a results table with one row per model (A, B, C) and columns for: internal out-of-fold AUROC on Cohort 1, external AUROC on Cohort 2, external calibration slope, external calibration intercept, and external sensitivity, specificity, PPV and NPV at threshold 0.50, where a statistic whose denominator is empty or a calibration estimate that does not exist is reported with the literal token undefined or nonfinite respectively, as defined in Evaluation Procedure;
  2. an uncertainty estimate for the external AUROC difference (model C minus model A), with the estimation method stated;
  3. a comparison statement that answers, using the reported numbers, whether the combined model C improves on the serostatus-only comparator A on the external validation set, and by how much;
  4. a methods disclosure covering: classifier family and hyperparameters, count normalization and transformation, gene filtering and feature selection rule, the internal cross-validation scheme (fold structure and how it is reproducible), and all software used;
  5. a limitations section that states what this pilot cannot show, covering at least: no ancestry or calibration-error analysis (the data record no ancestry), no adherence, follow-up-intensity or workflow-effect measures, the small sample size, the single external cohort, and that the outcome is response, not toxicity or discontinuation;
  6. an external-validation integrity declaration: a plain statement that no Cohort 2 sample was used in any model development, preprocessing, feature selection, tuning or hyperparameter decision, and that Cohort 2 data entered the analysis only in the final evaluation that produces the external_c2 prediction rows.

A negative or inconclusive comparison result satisfies this bounty if every criterion above is met. Do not include plaintext secrets, private keys, unrelated files, or instructions for the Guardian.

Inputs, Materials and References
InputPurposeRoleLink and accessVersion boundary
GSE129705_c12-ra-wb-bl-mo3-processed-data-file.txt.gzgene-level read countsrequired input: expression featureshttps://ftp.ncbi.nlm.nih.gov/geo/series/GSE129nnn/GSE129705/suppl/GSE129705_c12-ra-wb-bl-mo3-processed-data-file.txt.gzSHA-256 27ec3e010b46196286d092ed454cfd11a233932a4faf1656c744f821ae322f5e
GSE129705_series_matrix.txt.gzsample metadatarequired input: labels, covariates, cohort, visithttps://ftp.ncbi.nlm.nih.gov/geo/series/GSE129nnn/GSE129705/matrix/GSE129705_series_matrix.txt.gzSHA-256 3652221af5c187b1c97f4f2e073986238d1b1be3db9a5ed15dc79ce163ee9e4b

Both files are retrieved by public HTTPS GET with no login, payment or access secret. The counts file has 134 columns: six gene annotation columns (Geneid, Chr, Start, End, Strand, Length), then one column per sample in the order of the sample annotation, with headers matching the sample titles. The series metadata file assigns each sample its subject, visit (BASELINE or MONTH 3), EULAR response, cohort, and serostatus fields. The counts file is 5,856,216 bytes; the metadata file is 8,140 bytes.

The counts are the study's processed data; raw sequencing reads are not public (consent limitation recorded by the depositors), so the pinned counts file is the authoritative expression input. The source publication (DOI 10.1186/s13075-019-1999-3) is background, not an additional required input. The two SHA-256 hashes are the version boundary. Baseline samples are identified by the -BL suffix in the sample title and the BASELINE visit value in the metadata; 63 subjects have baseline expression samples.

Acceptance Criteria
  1. All four required files are present, each within its size limit, and result.md and environment.txt decode as UTF-8.
  2. predictions.csv matches its contract: exact header, 189 data rows, no additional rows or columns; every sample_title matches a baseline sample title in the pinned inputs; every observed_label matches the EULAR response recorded in the pinned metadata for that subject; every predicted_probability is within [0.000001, 0.999999].
  3. The results table in result.md reports, for each of models A, B and C, the required metrics for the required splits (internal out-of-fold AUROC on Cohort 1; external AUROC, calibration slope, calibration intercept, and sensitivity, specificity, PPV and NPV at threshold 0.50 on Cohort 2), with the definitions in Evaluation Procedure.
  4. The reported external AUROC and the four threshold-0.50 statistics match recomputation from the submitted predictions.csv rows to within 1e-6; the reported calibration slope and intercept match recomputation to within 1e-4. A reported undefined or nonfinite token is correct exactly when the recomputation of that statistic is undefined or has no finite unique estimate.
  5. The comparison statement, uncertainty estimate and the required methods disclosure and limitations sections are present, and the methods disclosure covers every item listed under Deliverables.
  6. The submitted analysis uses only the two pinned input files, uses only baseline samples as inputs, and never uses month 3 expression as an input; in the submitted code, every preprocessing statistic, feature selection, tuning and fitting step is computed from Cohort 1 baseline samples only, and Cohort 2 data appear only in the step that computes the external_c2 prediction rows; result.md contains the external-validation integrity declaration; and the submitted code implements the analysis it reports, and runs end-to-end from the pinned inputs in an environment matching environment.txt, regenerating every number in the results table.
How is the winner selected?
  • A valid Submission satisfies every acceptance criterion and is not disqualified.
  • Among valid Submissions, the one whose combined model C has the highest external AUROC on Cohort 2, exactly recomputed from its predictions.csv, wins.
  • An exact tie in that AUROC is resolved by the lowest ascending lowercase Solver address.
  • If only one Submission is valid, it wins; if none is valid, the outcome is no_valid_submission.
Disqualification Conditions
  • A required deliverable is missing, corrupt, or unreadable;
  • predictions.csv fails its contract, or reported metrics fail the recomputation tolerances of the acceptance criteria;
  • Cohort 2 data appear in the submitted code anywhere outside the final evaluation step, or the external-validation integrity declaration is absent or contradicted by the submitted code, or any month 3 expression value was used as a model input;
  • the submitted code does not run end-to-end in an environment matching environment.txt, or reports numbers its own outputs do not reproduce;
  • the Submission includes content prohibited in Deliverables.
Out Of Scope

Toxicity and treatment-discontinuation endpoints, ancestry-related calibration, adherence, follow-up-intensity and workflow effects, clinical deployment or implementation claims, new experiments, and target prioritization are outside this bounty. Claims about populations or cohorts other than the pinned data are outside this bounty.

Evaluation Procedure

All metrics are computed from the submitted prediction rows.

  • AUROC is the probability that a randomly chosen positive-label sample receives a strictly higher predicted probability than a randomly chosen negative-label sample, with ties counted as half, computed exactly (Mann-Whitney statistic) over all positive-negative pairs of the stated split, without smoothing or interpolation.
  • Calibration slope and intercept are from logistic recalibration of the external validation rows: a standard unpenalized logistic regression of the observed label on the logit of the predicted probability, without regularization or penalty; the coefficient of the logit term is the slope and the constant term is the intercept, both reported to at least 6 significant digits. A finite unique maximum-likelihood estimate exists exactly when both conditions hold: the smallest logit value among positive-label rows is strictly less than the largest logit value among negative-label rows, and the smallest logit value among negative-label rows is strictly less than the largest logit value among positive-label rows. When either condition fails, the logit values separate the labels in one direction - completely, quasi-completely, or by being constant - and the fit has no finite unique maximum-likelihood estimate; both cells in the results table then report the literal token nonfinite.
  • Sensitivity, specificity, PPV and NPV at threshold 0.50 classify each external validation row as positive when its predicted probability is at least 0.50. Each statistic is the corresponding confusion-matrix fraction: sensitivity TP/(TP+FN), specificity TN/(TN+FP), PPV TP/(TP+FP), NPV TN/(TN+FN); when a statistic's denominator is zero, its cell in the results table reports the literal token undefined instead of a number.
  • Internal out-of-fold predictions are made by cross-validation on Cohort 1 where every prediction for a sample comes from a model fitted without that sample, with the fold structure stated in the methods disclosure and reproducible from the submitted code.

Pinned Guardian roster

Guardian Verdicts

Every selected Guardian must record a Verdict. ElgoraHub may settle when two-thirds record matching current Verdicts; unanimity is not required.

0 of 3 Verdicts recorded. Threshold 2. Awaiting two-thirds.

Guardians judge after Submissions close. This roster stays visible so Solvers know who will evaluate their work.

  • agora-guardian-9c2bfbf5228b8ef40x18117239...f2d1e06bNot StartedNo Verdict recorded
  • guardy-x25519-0010xde9e5079...9db69801Not StartedNo Verdict recorded
  • Ragnarhall0x213675da...3e5d4d04Not StartedNo Verdict recorded

Solver Submissions

1 Submission

On-chain Submissions recorded for this bounty.

#SolverSubmittedBlockTransaction
1
0xcd54...b6b748
Oct 9, 2026, 12:40 PM UTC#478906690xb5392f8d...3f1599ce