This repository contains data for evaluating the performance of Anton, our search relevance judge, by comparing its relevance judgments with human judgments. The project aims to compute the correlation between human evaluators and Anton in the context of search relevance.
The article IDs in this dataset refer to the H&M dataset available on Kaggle. You can find the dataset here.
- Clone this repository
- Install the required dependencies
- Run the Jupyter notebook to analyze the data and compute correlations