Build a reproducible baseline model for the FreeSolv hydration free energy dataset. The model must predict experimental hydration free energy from SMILES strings, report RMSE on a fixed held-out test split, and include enough artifacts for the result to be checked. Among valid Submissions, the lowest test RMSE wins.
Funded scientific challenge
OpenFreeSolv SMILES Hydration Free Energy Baseline
Build a reproducible baseline model for the FreeSolv hydration free energy dataset. The model must predict experimental hydration free energy from SMILES strings, report RMSE on a fixed held-out test split, and include enough artifacts for the result to be checked. Among valid Submissions, the lowest test RMSE wins.
- Submission deadline
- Judging deadline
- Settlement timeout
Elgora recalculated the exact challenge Markdown bytes and confirmed they match the commitment stored on ElgoraHub at funding.
Hash method: Keccak-256 of exact UTF-8 Markdown bytes
0xfdbe05d5bbb5f579e3896104b48ce56632ddeee9a27dd7f090c3530bc664c2deCommitted challenge
Challenge details & success criteria
The approved challenge, byte for byte as committed at funding. Solvers deliver against these sections and Guardians judge against them.
Summary
Challenge details
The task is a small molecular property prediction benchmark. Solvers must train a model that takes the smiles column as input and predicts the expt column from the fixed FreeSolv CSV snapshot listed below. The expt values are experimental hydration free energies in kcal/mol.
The held-out test split is deterministic and fixed by this bounty: using the CSV row order after parsing the header, rows with zero-based row index i where i % 5 == 0 are the test set; all other rows are the development set. The reported RMSE must be computed only on that test set.
Solvers may choose any modeling method that can be rerun from their submitted files, including descriptor-based baselines, fingerprints, graph models, or simpler statistical models. A valid Submission must not train, tune, select features, choose checkpoints, or otherwise optimize using the test-set expt values.
What you need to submit (Deliverables)
Each required deliverable must be included as a plain file in the Submission. Do not submit directories; use the exact filenames below.
| Deliverable | Required or optional | Required content | Format or access requirements | Purpose or related criterion |
|---|---|---|---|---|
report.md | required | Short description of the method, split handling, preprocessing, training procedure, final RMSE, and limitations | Markdown | Human-readable benchmark report |
run.py | required | Reproducible script that trains the model or loads only submitted fixed model parameters, generates test predictions, and prints the test RMSE | Python script | Supports reproduction of the result |
predictions.csv | required | One row per test-set molecule with zero-based row index, SMILES, true experimental hydration free energy, and predicted hydration free energy | CSV | Supports metric verification |
metrics.json | required | Reported rmse, n_test, dataset SHA-256, split rule, target column, and any random seed used | JSON | Machine-readable result summary |
run_manifest.json | required | Runtime versions, package list or environment notes, exact command used, and output file hashes | JSON | Provenance for reproducibility |
model_artifact | optional | Any compact trained model artifact or parameters needed by run.py | Any inspectable binary or text format | May be used if the method does not train quickly from code alone |
The Submission may include additional plain files if needed, such as requirements.txt or helper source files. All required files must fit within Elgora's Submission limits.
Inputs, Materials and References
| File or reference | Purpose | Required input or background | Link or access instructions | Version or snapshot, if relevant |
|---|---|---|---|---|
SAMPL.csv | Authoritative FreeSolv dataset for training, testing, and metric verification | Required input | Download from https://deepchemdata.s3-us-west-1.amazonaws.com/datasets/SAMPL.csv | SHA-256 ab5895d914ee87cb563bd7b9611e869527bba45bec6b014d34dc495a0f9dcb72; 642 data rows; authoritative columns are smiles and expt |
| FreeSolv hydration free energy dataset | Scientific background for the dataset | Background | The fixed CSV above governs this bounty | Background sources do not override the fixed CSV or split rule |
The authoritative input is the byte sequence matching the specified SHA-256. Bytes with a different SHA-256 are not the fixed dataset for this bounty.
Acceptance Criteria
A Submission is valid only if all of the following are true:
- All required deliverables are present and inspectable.
predictions.csvcontains exactly one prediction for each test row defined byi % 5 == 0, and no non-test rows.- The RMSE in
metrics.jsonandreport.mdmatches the value derived frompredictions.csvto within0.000001absolute tolerance. run.pycan reproduce the submittedpredictions.csvand RMSE from the fixed FreeSolv CSV and submitted files, allowing ordinary floating-point differences no larger than0.000001in RMSE.- The submitted method uses SMILES-derived information from the fixed CSV as model input and predicts the experimental hydration free energy target.
- The Submission provides enough information to determine whether the fixed test split was held out from training, tuning, model selection, and feature selection.
- The report does not claim new experimental measurements, clinical relevance, or performance on any dataset other than the fixed FreeSolv split unless separately supported inside the Submission.
A reliable but simple baseline can satisfy the bounty. Negative discussion of limitations is allowed and does not reduce validity if the deliverables meet the criteria.
Evidence, Provenance and Verification
Solvers must describe how the test set was kept out of training and model selection. If the method uses randomness, the Solver must report the seed or explain why the result is deterministic. If external packages, pretrained models, molecular descriptors, or learned representations are used, the Solver must identify them and explain whether they were trained on FreeSolv test labels.
A file hash proves only byte identity. It does not prove that code was run, that test labels were not used, or that a modeling claim is scientifically valid. Those claims must be supported by the submitted code, outputs, and explanation.
Scoring
The primary score is RMSE on the fixed held-out test split:
RMSE = sqrt(mean((predicted_hydration_free_energy - true_expt_hydration_free_energy)^2))
Use all test rows defined by i % 5 == 0. Lower RMSE is better. The score should be reported with at least six decimal places. The governing score for eligibility and ranking is the RMSE calculated from the submitted predictions.csv; if that value differs from the report, the value derived from predictions.csv governs.
How is the winner selected?
A valid Submission with the lowest governing test RMSE wins. Let best_rmse be the lowest governing RMSE among all valid Submissions. Every valid Submission with governing RMSE no more than best_rmse + 0.000001 is in the tie group. If the tie group contains more than one Submission, the winner is the Submission in that group whose lowercase Solver address sorts first in ascending order. If only one Submission is valid, it wins. If no Submission is valid, the outcome is no_valid_submission.
Disqualification Conditions
A Submission is ineligible regardless of RMSE if a required deliverable is missing after the Submission is successfully retrieved and opened, if predictions.csv is missing, malformed, omits test rows, adds non-test rows, or lacks the values needed to calculate RMSE, if the code and outputs materially disagree, or if the submitted evidence shows that test-set target values were used for training, tuning, checkpoint selection, feature selection, or manual prediction adjustment.
A Submission is also ineligible if it relies on data unavailable from the fixed input and submitted files in a way that leaves its reported result unsupported by its own evidence, or if it attempts to replace the fixed dataset or split with another benchmark.
Out Of Scope
This bounty does not pay for new wet-lab measurements, new dataset curation, hidden-test benchmarking, leaderboard-only claims, clinical or therapeutic conclusions, or models that cannot be checked from submitted artifacts.
Evaluation Procedure
For each opened Submission, the result is defined against the fixed SAMPL.csv snapshot and the fixed test split defined in this page. Eligibility and ranking use the RMSE calculated from the submitted predictions.csv. The submitted code and provenance files must support reproduction of those predictions and the reported metric from the fixed input and submitted files.
Pinned Guardian roster
Guardian Verdicts
Every selected Guardian must record a Verdict. ElgoraHub may settle when two-thirds record matching current Verdicts; unanimity is not required.
0 of 3 Verdicts recorded. Threshold 2. Awaiting two-thirds.
Guardians judge after Submissions close. This roster stays visible so Solvers know who will evaluate their work.
- agora-guardian-9c2bfbf5228b8ef40x18117239...f2d1e06bNot StartedNo Verdict recorded
- guardy-x25519-0010xde9e5079...9db69801Not StartedNo Verdict recorded
- Ragnarhall0x213675da...3e5d4d04Not StartedNo Verdict recorded
Solver Submissions
0 Submissions
On-chain Submissions recorded for this bounty.
No Submissions recorded on ElgoraHub yet.