The datasets of Crowd Labels with Text Contents of instances available.
We have just uploaded the QUIZ datasets, the documentations, source codes for data processing and the RTE dataset are under construction.
Datasets used for our following papers:
In this paper, the text contents are used by LLM to generate the labels.
- Jiyi Li, "A Comparative Study on Annotation Quality of Crowdsourcing and LLM via Label Aggregation", Proceedings of the 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2024), pp. 6525-6529, Apr. 2024.
In this paper, only crowd labels are used for proposing the label aggregation method.
- Jiyi Li, Yukino Baba, Hisashi Kashima, "Hyper Questions: Unsupervised Targeting of a Few Experts in Crowdsourcing", the 26th ACM International Conference on Information and Knowledge Management (CIKM 2017), pp.1069-1078, Nov. 2017.
-
Chinese (CHI): the meaning of Chinese vocabularies.
-
English (ENG): the most analogically similar word pair to a word pair.
-
Information Technology (ITM, ITMANAGE): the basic knowledge of information technology.
-
Medicine (MED): about medicine efficacy and side effects.
-
Pokemon (POK): the Japanese name of a Pokemon with English name.
-
Science (SCI): intermediate knowledge of chemistry and physics.
If you use this dataset, please cite the following papers.
@inproceedings{CrowdLLMLabel,
author={Li, Jiyi},
booktitle={ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
title={A Comparative Study on Annotation Quality of Crowdsourcing and LLm Via Label Aggregation},
year={2024},
volume={},
number={},
pages={6525-6529},
keywords={Crowdsourcing;Annotations;Quality control;Benchmark testing;Signal processing;Chatbots;Reliability;Crowdsourcing;Label Aggregation;Large Language Model},
doi={10.1109/ICASSP48485.2024.10447803}
}
@inproceedings{HyperQuestion,
author = {Li, Jiyi and Baba, Yukino and Kashima, Hisashi},
title = {Hyper Questions: Unsupervised Targeting of a Few Experts in Crowdsourcing},
year = {2017},
isbn = {9781450349185},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3132847.3132971},
doi = {10.1145/3132847.3132971},
booktitle = {Proceedings of the 2017 ACM on Conference on Information and Knowledge Management},
pages = {1069–1078},
numpages = {10},
keywords = {crowdsourcing, answer aggregation, hyper question, heterogeneous-answer multiple-choice questions},
location = {Singapore, Singapore},
series = {CIKM '17}
}
Creative Commons CC BY 4.0