Skip to content

Gigaword Validation Data #25

Description

@stylelohan

Hi,

I have downloaded Gigaword dataset from havardnlp, and there are 4 directories: train, Giga, DUC2004, and DUC2003.

I use valid.title.filter.txt, valid.article.filter.txt, train.title.txt and train.article.txt in train as the tgt/src of valid/train data, and use two txt files in Giga as test data. However, the valid data has wrong size (189651) after preprocess. The weird thing is that when I run preprocess.py, the result shows "(0 and 0 ignored due to length == 0 or > )".

Would you know any method to fix this?
Thanks a lot!

Activity

  1. changed the title [-]Gigaword Valid Data[/-] [+]Gigaword Validation Data[/+] on Aug 20, 2019
  2. JustinLin610 commented on Oct 1, 2019

    @JustinLin610
    Collaborator

    This problem may stem from empty lines in your dataset. Check if you have this issue.

  3. jiahuanluo commented on Oct 15, 2019

    @jiahuanluo

    @JustinLin610 Thanks for your job.
    I wonder how to split the data into validation set and test set. There are 18,691 lines in the valid.article.filter.txt.
    How could I get the 8k validation set and the 2k test set identical to that of your paper?
    Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions