I am interested in the outcome of the percentage of # rephrased sample contamination in Real-world datasets. Regarding large training datasets, you mentioned using a subset. Could you provide the parameters(like data_dir = "python" similar to those shown in the 'Pre-process' section of the READme.md file?
I am interested in the outcome of the percentage of # rephrased sample contamination in Real-world datasets. Regarding large training datasets, you mentioned using a subset. Could you provide the parameters(like data_dir = "python" similar to those shown in the 'Pre-process' section of the READme.md file?