您好,我想问一下您选取的MVSA-Single数据集为何与官网数据集规模不一致,为何不使用官方的总数为4511的数据集呢?
您论文中提到:
Image-text classification. For FOOD-101, following the previous work (Kiela et al., 2019), there are 60101 image-text pairs in the training set, 5000 image-text pairs in the validation set, and 21695 image-text pairs in the test set. For MVSA, we conduct the division strategy presents in (Kiela et al., 2019). There are 1555 image-text pairs in the training set. The validation set contains 518 image-text pairs, and the test set contains 519 image-text pairs.
但我在您提到的文章中并没有找到相应的”划分依据“,请问您此处的引用是否正确?
您好,我想问一下您选取的MVSA-Single数据集为何与官网数据集规模不一致,为何不使用官方的总数为4511的数据集呢?
您论文中提到:
Image-text classification. For FOOD-101, following the previous work (Kiela et al., 2019), there are 60101 image-text pairs in the training set, 5000 image-text pairs in the validation set, and 21695 image-text pairs in the test set. For MVSA, we conduct the division strategy presents in (Kiela et al., 2019). There are 1555 image-text pairs in the training set. The validation set contains 518 image-text pairs, and the test set contains 519 image-text pairs.但我在您提到的文章中并没有找到相应的”划分依据“,请问您此处的引用是否正确?