Skip to content

Any plan to open-source the IF-caps-Pro dataset? #2

Description

@melo1998

Hi, thanks a lot for this excellent work and for releasing the code/models.

While going through the paper and the training data description (Table A.1 /
training data summary), I noticed that the IF-caps-Pro dataset is used as a
key large-scale source for T2A and TV2A training. The scale and caption quality
of this dataset look very high, and it seems to play an important role in the
reported performance.

I have a few questions regarding this dataset:

  1. Is there any plan to open-source IF-caps-Pro? If so, is there an
    estimated timeline?
  2. If a full release is not planned, would it be possible to release a
    subset, or the caption annotations / metadata (e.g. clip IDs,
    timestamps, captions) so that the community can reconstruct it from the
    original audio sources?
  3. Could you share more details about how IF-caps-Pro was constructed
    (source datasets, captioning pipeline, filtering criteria)? This would help
    reproduce the results even if the data itself cannot be released.

Understanding the licensing or availability of this dataset would be very
helpful for reproducing the results and for follow-up research.

Thanks again for your time and for the great contribution!

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions