Evaluation¶
Debug phase¶
Participants can submit their containers for a sanity check on training cases to see if the output on the platform matches the local one.
Test phase, qualification¶
Models are ranked by C-index on the hidden test cohort. Those above C-index 0.65 advance to the interpretability stage; below that threshold, the prognostic signal is too weak for the explanation to mean much.
Automatic faithfulness¶
Automatic evaluation of interpretability faithfulness will be executed via perturbation analysis: regions the attribution map marks as highly relevant are progressively removed, and the change in the model's output is compared against perturbing random regions. A larger gap means a more faithful explanation.
AI expert panel¶
Manuscripts, code, metrics, predictions, attribution maps, and reports are reviewed by at least 3 experts each, on Likert-scale questionnaires covering relevance and applicability for hypothesis-free scientific discovery, clarity, scientific plausibility, meaningful novelty, usefulness for understanding model behaviour, transferability to other pathology problems, and risk of hallucination.
Pathology expert panel¶
Practising pathologists and trainees independently assess a subset of cases per submission, at least two experts each, on clinical meaningfulness, alignment with the image evidence, usefulness for understanding and questioning the prediction, and support for clinical understanding.
Final ranking¶
A weighted sum of the aggregated scores.