corpora
Here are 175 public repositories matching this topic...
A collection of corpora for named entity recognition (NER) and entity recognition tasks. These annotated datasets cover a variety of languages, domains and entity types.
-
Updated
Jul 2, 2026 - Python
Data repository for pretrained NLP models and NLP corpora.
-
Updated
Mar 16, 2018 - Python
A collaborative catalog of NLP resources for Indic languages
-
Updated
Dec 14, 2024
微信公众号语料库
-
Updated
Jan 7, 2019
Official source for spanish Language Models and resources made @ BSC-TEMU within the "Plan de las Tecnologías del Lenguaje" (Plan-TL).
-
Updated
Jul 27, 2023 - Python
A web-based engine for creating and annotating textual corpora
-
Updated
Aug 26, 2023 - PHP
CrossNER: Evaluating Cross-Domain Named Entity Recognition (AAAI-2021)
-
Updated
Jan 5, 2021 - Python
Unannotated Spanish 3 Billion Words Corpora
-
Updated
Oct 20, 2022 - Python
Automatic categorization of documents, consists in assigning a category to a text based on the information it contains. We'll follow different approach of Supervised Machine Learning.
-
Updated
Jan 1, 2019 - Python
An advanced, extensible web front-end for the Manatee-open corpus search engine
-
Updated
Aug 20, 2026 - TypeScript
An R package for dynamic exploration of text collections
-
Updated
Mar 6, 2026 - R
[NLPCC 2023] CCAE: A Corpus of Chinese-based Asian Englishes
-
Updated
Dec 6, 2023 - Python
Named Entity Recognition for biomedical entities
-
Updated
Jan 11, 2023 - Python
Tools for filtering and cleaning parallel and monolingual corpora for machine translation and other natural language processing tasks.
-
Updated
Dec 19, 2023 - PHP
A comprehensive list of annotated training datasets classified by use case.
-
Updated
Jul 8, 2022
Add this topic to your repo
To associate your repository with the corpora topic, visit your repo's landing page and select "manage topics."