Libre AI Datasets #15
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
See:
https://spacecruft.org/deepcrayon/parrot-datasets
Like the models, many of the training data sets that claim to be "open source" aren't actually using open source licenses.
Find useful truly "open source" data sets.
Note which notable large data sets claim to be "open source", but don't have open source license.
Might not suck:
"RedPajama is a clean-room, fully open-source implementation of the LLaMa dataset."
https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T
Lists these licenses for what they scraped:
License
Please refer to the licenses of the data subsets you use.
Common Crawl is definitely not open source:
https://commoncrawl.org/terms-of-use
2023-11-15 10:54:42 -07:00
Poster
Collaborator
https://huggingface.co/datasets/c4
Under the "odc-by" license:
https://opendatacommons.org/licenses/by/
https://opendatacommons.org/licenses/by/1-0/
https://opendatacommons.org/licenses/by/summary/
OSI doesn't have the license listed.
FSF has comments on the Open Data DB license, from the same organization:
"Open Database license (#ODbl)
This is a free and copyleft license meant for data. It is incompatible with the GNU GPL. Please don't use it for software or documentation, since it is incompatible with the GNU GPL and with the GNU FDL. It makes inconvenient requirements about signing contracts which try to create an effect like copyleft for data that is not copyrightable, so we don't recommend using it; however, there is no reason to avoid using data released this way."
I don't see that any outside group has reviewed/approved odc-by as libre. Punt.
"Books" from the pile:
https://huggingface.co/datasets/the_pile_books3#licensing-information
"Defunct: Dataset "the_pile_books3" is defunct and no longer accessible due to reported copyright infringement."
pg19 is project gutenberg. Project Gutenberg is based on public domain works, so is libre.
https://huggingface.co/datasets/pg19
The set is licensed Apache 2.0, so is good to use. Just old texts though, not programming...
Fastest way to download pg19 is ala:
The compilers of The Pile have a good table about datasets they included.
https://arxiv.org/abs/2101.00027
Probably best to use the Pile, but exclude non-libre sections.
Stackoverflow may actually be free to use, this is in their terms:
"From time to time, Stack Overflow may make available compilations of all the Subscriber Content on the public Network (the “Creative Commons Data Dump”). The Creative Commons Data Dump is licensed under the CC BY-SA license. By downloading the Creative Commons Data Dump, you agree to be bound by the terms of that license."
https://stackoverflow.com/legal/terms-of-service/public
Creative Commons licensing terms (CC BY-SA 4.0)
https://archive.org/download/stackexchange/
https://archive.org/download/stackexchange/stackexchange_archive.torrent
Also remove non-libre os questions, etc. from stackoverflow and others.
Perhaps rust cargo mirror for dataset.
Huggingface no_robots is non-libre NC.
AI2 ImpACT License, more shit. Non libre.
I see "The Stack" created and used by BigCode has GPL, AGPL, and CC by SA-4.0 files in it created by a company I owned and one I own now.
Built with lots of libre data sets:
https://huggingface.co/TigerResearch
See:
https://www.kaggle.com/datasets
https://huggingface.co/glaiveai