Hi,
I have downloaded Gigaword dataset from havardnlp, and there are 4 directories: train, Giga, DUC2004, and DUC2003.
I use valid.title.filter.txt, valid.article.filter.txt, train.title.txt and train.article.txt in train as the tgt/src of valid/train data, and use two txt files in Giga as test data. However, the valid data has wrong size (189651) after preprocess. The weird thing is that when I run preprocess.py, the result shows "(0 and 0 ignored due to length == 0 or > )".
Would you know any method to fix this?
Thanks a lot!
Hi,
I have downloaded Gigaword dataset from havardnlp, and there are 4 directories: train, Giga, DUC2004, and DUC2003.
I use valid.title.filter.txt, valid.article.filter.txt, train.title.txt and train.article.txt in train as the tgt/src of valid/train data, and use two txt files in Giga as test data. However, the valid data has wrong size (189651) after preprocess. The weird thing is that when I run preprocess.py, the result shows "(0 and 0 ignored due to length == 0 or > )".
Would you know any method to fix this?
Thanks a lot!