As documented, Kindly currently uses the twitter-roberta-base-offensive dataset, which is licensed without any restrictions, but refers to the original license for additional information.
This dataset is used in this study, therein referred as Offensive Language Identification Dataset (OLID) with its associated homepage belonging to Shervin Malmasi. It includes a download link, but it's a dead end.
Another reference to the same study, yields a code repository with the actual dataset, but is missing a license.
I am contacting both leads to request an explicit inclusion of an open license.
Cc: @amreenp7, @nathanfletcher
As documented, Kindly currently uses the twitter-roberta-base-offensive dataset, which is licensed without any restrictions, but refers to the original license for additional information.
This dataset is used in this study, therein referred as Offensive Language Identification Dataset (OLID) with its associated homepage belonging to Shervin Malmasi. It includes a download link, but it's a dead end.
Another reference to the same study, yields a code repository with the actual dataset, but is missing a license.
I am contacting both leads to request an explicit inclusion of an open license.
Cc: @amreenp7, @nathanfletcher