# Mozilla Data Collective > Mozilla Data Collective is rebuilding the AI data ecosystem with communities at the centre. Access over 600 high-quality global datasets, built by and for the community in a transparent and ethical way. > Mozilla Data Collective is the data platform for human agency and fair value exchange. People should be able to choose where their datasets show up, and they should be able to define what it looks like to benefit; whether that’s swapping data for tool access, or for expertise, donating it to the public, or asking for fair compensation. > Data providers can share data openly, using existing licenses like Creative Commons, or by building their own license. Datasets can be open for everyone, or just for some types of downloaders. Uploaders can set custom constraints, ask for exchange, compensation or recognition. Downloaders who access datasets are fully authenticated, and held in legally binding contracts, and we have a number of dataset protection features. ## Datasets - [Datasets](/datasets): Mozilla Data Collective has 600+ ethically sourced datasets shared by over 150 organizations. > Each dataset is identified by a unique slug: /datasets/ > Uploaders provide their set of terms and conditions for accessing the dataset alongside the license to the dataset itself. Downloaders must agree to these conditions before they are able to download the dataset. > Some datasets require the downloader to share their email address to request access before the uploader grants access to the dataset. ## Uploading - [Uploads](/profile/submissions/create): Account holders can submit requests to become uploaders on Mozilla Data Collective by submitting a form explaining the type of data that they want to publish and their goals in sharing their data. To request to become an uploader, account holders should visit their profile page and submit a request. > Once approved as an uploader on the platform, which generally takes less than a week, uploaders can create a dataset listing by sharing a `.tar.gz` file and filling out an accompanying datasheet. - [How to Upload](https://community.mozilladatacollective.com/uploading-your-dataset-to-the-mozilla-data-collective-platform/) ## Organizations - [Organizations](/organization/): Organizations who upload to Mozilla Data Collective can create organizational profiles that show all of their available datasets. ## Legal Documents - [Terms of Service](/terms): Use of Mozilla Data Collective is governed by different sets of terms for uploader organizations and downloaders. - [Privacy](/privacy): The Mozilla Data Collective privacy policy explains how we use information gathered by the platform. ## API - [API](/api-reference): The Mozilla Data Collective API allows you to programatically download datasets. It provides a REST API and [Python SDK] ([datacollective](https://pypi.org/project/datacollective/)). The API is in beta. Authentication uses Bearer tokens via the Authorization header. The base URL is https://mozilladatacollective.com/api and downloads are limited to 30 per day per organization. Users must agree to dataset terms via the web interface before downloading and presigned URLs are returned for direct storage downloads (valid 12 hours). ### API Documentation - [API Reference Overview](https://mozilladatacollective.com/api-reference): Landing page with quickstart steps for API access - [REST API Docs](https://mozilladatacollective.com/api-reference/docs): Full endpoint reference including GET /datasets/:datasetId, POST /datasets/:datasetId/download, authentication, rate limiting, and error handling ### Python SDK - [datacollective Python SDK Documentation](https://mozilla-data-collective.github.io/datacollective-python/): Installation, configuration, usage guide for download_dataset, load_dataset, and programmatic uploads - [datacollective on PyPI](https://pypi.org/project/datacollective/): Python package installation - [SDK Source Code](https://github.com/Mozilla-Data-Collective/datacollective-python): GitHub repository ### Optional - [Schema-Based Loading](https://mozilla-data-collective.github.io/datacollective-python/schema_documentation/): How `schema.yaml` files drive dataset loading into DataFrames - [ASR Loader](https://mozilla-data-collective.github.io/datacollective-python/loaders/asr/): ASR-specific dataset loading - [TTS Loader](https://mozilla-data-collective.github.io/datacollective-python/loaders/tts/): TTS-specific dataset loading - [Programmatic Uploads](https://mozilla-data-collective.github.io/datacollective-python/upload/): Creating submissions and uploading datasets via the SDK ## Relationship to Common Voice > In 2025, our Founder and CEO E.M. Lewis-Jong, was leading Common Voice (the world’s largest public participation speech dataset) at Mozilla Foundation, and was looking for a release platform that would give Common Voice communities more choice: choice of license, features that undergirded stronger control, and a radically anti-extractivist form of value exchange. > The team couldn’t find that platform, so in September we built Mozilla Data Collective. Common Voice was Mozilla Data Collective’s first community user, piloting the platform for its own datasets. ## Relationship to Mozilla Foundation > Mozilla Data Collective was incubated at the Mozilla Foundation. In April 2026, we spun out a UK entity specifically focused on the Mozilla Data Collective platform. Common Voice continues to be stewarded by Mozilla Foundation. ## Optional - [More information](https://community.mozilladatacollective.com/): The Mozilla Data Collective community blog. - [How Mozilla Data Collective came to be](https://community.mozilladatacollective.com/about/): About Mozilla Data Collective as an organization - [Guides for using Mozilla Data Collective](https://community.mozilladatacollective.com/tag/guides/) ### Social media You can find Mozilla Data Collective at the following social media sites: - [LinkedIn](https://www.linkedin.com/company/mozilla-data-collective) - [Reddit](https://www.reddit.com/r/MozillaDataCollective/) - [Discord](https://discord.gg/cs9tJPqQB5) ## Datasets currently available on the platform The following is a list of datasets currently available on the platform and the data they contain: ## Datasets - [CoVoST 2 English - Persian](https://mozilladatacollective.com/datasets/cmpr9hex201bhnv07aoztund6): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Persian. - [KerOfis toponymic database Breton-French](https://mozilladatacollective.com/datasets/cmpr2sd43014knv07mpsae5yf): KerOfis is a database of toponyms. It aims to provide the correct forms of Breton place names. - [CoVoST 2 English - Turkish](https://mozilladatacollective.com/datasets/cmpqxcsb900wbnu07wlxrag4b): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Turkish. - [CoVoST 2 English - Chinese (China)](https://mozilladatacollective.com/datasets/cmpqxcoy400zinv07bmrdx59x): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Chinese (China). - [CoVoST 2 English - Tamil](https://mozilladatacollective.com/datasets/cmpqxclu100w7nu07890m5s45): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Tamil. - [CoVoST 2 English - Slovenian](https://mozilladatacollective.com/datasets/cmpqxciw600w3nu07pagrlqka): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Slovenian. - [CoVoST 2 English - Latvian](https://mozilladatacollective.com/datasets/cmpqxcg1z00zanv07m3qeb196): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Latvian. - [CoVoST 2 English - Japanese](https://mozilladatacollective.com/datasets/cmpqxcd0300vznu07ynpc8du5): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Japanese. - [CoVoST 2 English - Welsh](https://mozilladatacollective.com/datasets/cmpq9bnz000esnu073lop6tud): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Welsh. - [CoVoST 2 English - German](https://mozilladatacollective.com/datasets/cmpq9bjzy00f7nv07qff3brpb): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in German. - [CoVoST 2 English - Estonian](https://mozilladatacollective.com/datasets/cmpq9bfw000eonu07uwrpnygg): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Estonian. - [CoVoST 2 English - Indonesian](https://mozilladatacollective.com/datasets/cmpq9bamw00eknu07hsiurgso): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Indonesian. - [Friulian TTS - female voice](https://mozilladatacollective.com/datasets/cmp3rmip800ozo30799h7wp6g): A single-speaker read speech dataset in Friulian. The dataset contains ~5 hours of pre-segmented utterances, recorded by an anonymous ~30-years-old female Friulian native speaker. - [CoVoST 2 English - Catalan](https://mozilladatacollective.com/datasets/cmpps4tzv0019nu07p7h9zl0e): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Catalan. - [Mada-French Parallel Corpus 1.0](https://mozilladatacollective.com/datasets/cmmamvzrz04qtmk077j1k99vt): This dataset comprises a parallel corpus of Mada–French literary text translations totalling 2,154 lines. It is designed to support the benchmarking, training and evaluation of machine translation models for Mada, a language spoken in Cameroon. - [Jombang Dialect-Javanese TTS](https://mozilladatacollective.com/datasets/cmpo53l5b000pku07cem453ac): The Jombang dialect is part of the Arekan dialect (East Javanese) in East Java Province, Indonesia. However, this dialect has a unique position because it is the meeting point between the cultural influences between "Mataraman" (Solo-Yogyakarta) and "Arekan" (Surabaya-Malang) across Yogyakarta, Central Java, and East Java Province, Indonesia. - [Tequila Zongolica Nahuatl Audio](https://mozilladatacollective.com/datasets/cmpo51atc000mnu07idjb21pw): Audio corpus of Orizaba-Zongolica Nahuatl language (Glottocode:oriz1235) with a total duration of approximately **122:02:38** (hours:mins:secs). - [AmaWar](https://mozilladatacollective.com/datasets/cmpdyqe6u00oznu07c3scvwz0): Bitext from the online AmaWar dictionary of the Tamazight dialect of Ait Warain spoken in northeastern Morocco. Contains sentences, stories, and poems in Tamazight written in the Neo-Tifinagh script along with their translations into Modern Standard Arabic. - [Duala-TTS-Dataset](https://mozilladatacollective.com/datasets/cmpmpf0jw021bnu0743hu7763): This dataset comprises 1,521 high-quality audio recordings of read speech produced by a single Duala speaker over several sessions. Duala (ISO 639-3: dua), also known as Douala, is a Bantu language of the Niger-Congo family spoken primarily in the Littoral Region of Cameroon, notably in the city of Douala and its surrounding areas. - [Ngemba-ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmpmpdwwh0217nu07zomtwz7o): Ngemba_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Ngemba language, also referred to in the literature as Ghomala-Ouest (Breton and Bikia Fohtung 1991). Ngemba is a Grassfields Bantu language spoken in the West Region of Cameroon and is rarely represented in existing standard grammatical descriptions, computational resources or lexicographical tools. - [Otomí (Hñähñu) TTS Voz Masculina](https://mozilladatacollective.com/datasets/cmo0cro1g00hlmr07oichasyk): 4.5 horas de habla leída alineada con texto. Estos texto vienen de un solo libro. - [Cross Agency Federal Violations by Company](https://mozilladatacollective.com/datasets/cmpfuh0p0004rmm070vm3ggnx): Multi-Agency Federal Enforcement: U.S. Companies Cited by 2+ Federal Agencies. What separates a company that had a one-off OSHA citation from one with systemic compliance problems? - [Buscador de Obligaciones de Transparencia de la Plataforma Nacional de Transparencia](https://mozilladatacollective.com/datasets/cmpfmnia3000dnx0776qa5xp6): El dataset comprende una extracción parcial del buscador de obligaciones de transparencia de la Plataforma Nacional de Transparencia (PNT) de México, basado en el Sistema de Portales de Obligaciones de Transparencia (SIPOT). Fue extraído sistemáticamente por el equipo Amnesia durante 2025 ante el riesgo de pérdida asociado a la extinción del Instituto Nacional de Transparencia, Acceso a la Información Pública y Protección de Datos Personales (INAI). - [Ewondo-French Parallel Corpus](https://mozilladatacollective.com/datasets/cmpedjr9y00wqnv07styloz0d): This dataset is a parallel corpus of Ewondo and French texts. The text was obtained by transcribing raw audio files recorded in Yaoundé in the 1980s. - [CoVoST 2 Estonian - English](https://mozilladatacollective.com/datasets/cmp77wrdm02wdmp071z00ldk3): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Estonian audio (8 hours) and the translations in English. - [Diboum-ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmpe4sp6100rznu07p61gmppd): Diboum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Diboum variety of Basaa (ISO 639-3: bas), a Bantu language of Cameroon. Diboum is a localised and socially embedded speech form that is rarely represented in standard grammatical descriptions or lexicographical resources. - [Multilingual Audio Visual Speech (Lip-Reading) Dataset](https://mozilladatacollective.com/datasets/cmpe4rs9800r2nv07qpq8972c): Community-sourced dataset of anonymised mouth-only video with separate audio track file and transcribed text in the format used for training and evaluating lip-reading (visual speech recognition) models. 7 different languages from 16 different anonymous speakers in a noisy environment. - [StackOverflow Knowledge Graph - Bash](https://mozilladatacollective.com/datasets/cmpdyogk700ohnv07aasx1q4f): N-Triples/RDF export of the StackOverflow knowledge graph for the Bash programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - C](https://mozilladatacollective.com/datasets/cmpdyoaa300ovnu07ruxpomwb): N-Triples/RDF export of the StackOverflow knowledge graph for the C programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Cobol](https://mozilladatacollective.com/datasets/cmpdyo45100o9nv07np0uqmpm): N-Triples/RDF export of the StackOverflow knowledge graph for the Cobol programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - C#](https://mozilladatacollective.com/datasets/cmpdynyzu00ornu074r3i2ndm): N-Triples/RDF export of the StackOverflow knowledge graph for the C# programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - CSS](https://mozilladatacollective.com/datasets/cmpdyntcn00o5nv07wwkgz5x7): N-Triples/RDF export of the StackOverflow knowledge graph for the CSS programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Dart](https://mozilladatacollective.com/datasets/cmpdynmf900o1nv07s9a3fopl): N-Triples/RDF export of the StackOverflow knowledge graph for the Dart programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Elixir](https://mozilladatacollective.com/datasets/cmpdyng3300nxnv07p9nlq3na): N-Triples/RDF export of the StackOverflow knowledge graph for the Elixir programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Erlang](https://mozilladatacollective.com/datasets/cmpdynb1g00onnu0783qgrvdr): N-Triples/RDF export of the StackOverflow knowledge graph for the Erlang programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Fortran](https://mozilladatacollective.com/datasets/cmpdyn3lc00ojnu07qfq4ff3h): N-Triples/RDF export of the StackOverflow knowledge graph for the Fortran programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - F#](https://mozilladatacollective.com/datasets/cmpdymy7o00ofnu0743zb36hs): N-Triples/RDF export of the StackOverflow knowledge graph for the F# programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - GO](https://mozilladatacollective.com/datasets/cmpdymqte00o9nu07mzkxwxfb): N-Triples/RDF export of the StackOverflow knowledge graph for the GO programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Haskell](https://mozilladatacollective.com/datasets/cmpdymk8t00ntnv07vw6h2zhr): N-Triples/RDF export of the StackOverflow knowledge graph for the Haskell programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Html](https://mozilladatacollective.com/datasets/cmpdymcul00npnv07amwfj0zq): N-Triples/RDF export of the StackOverflow knowledge graph for the Html programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Java](https://mozilladatacollective.com/datasets/cmpdym5uw00o1nu071kqpz3vw): N-Triples/RDF export of the StackOverflow knowledge graph for the Java programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Javascript](https://mozilladatacollective.com/datasets/cmpdylxq900nlnv07sobvszik): N-Triples/RDF export of the StackOverflow knowledge graph for the Javascript programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Julia](https://mozilladatacollective.com/datasets/cmpdylr4p00nxnu07j6vnqwl6): N-Triples/RDF export of the StackOverflow knowledge graph for the Julia programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Kotlin](https://mozilladatacollective.com/datasets/cmpdyl8p800ntnu07od7r1u51): N-Triples/RDF export of the StackOverflow knowledge graph for the Kotlin programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Lisp](https://mozilladatacollective.com/datasets/cmpdyl2wk00npnu0727nrkvfk): N-Triples/RDF export of the StackOverflow knowledge graph for the Lisp programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - LUA](https://mozilladatacollective.com/datasets/cmpdykw5h00nlnu079ty86xbw): N-Triples/RDF export of the StackOverflow knowledge graph for the LUA programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Matlab](https://mozilladatacollective.com/datasets/cmpdyknrh00nhnv07ohqk1hrc): N-Triples/RDF export of the StackOverflow knowledge graph for the Matlab programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Objective-C](https://mozilladatacollective.com/datasets/cmpdykii200ndnv07wkw14oqa): N-Triples/RDF export of the StackOverflow knowledge graph for the Objective-C programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Perl](https://mozilladatacollective.com/datasets/cmpdyk9yd00n9nv073odu03go): N-Triples/RDF export of the StackOverflow knowledge graph for the Perl programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - PHP](https://mozilladatacollective.com/datasets/cmpdyk3ap00n5nv078lcz7evy): N-Triples/RDF export of the StackOverflow knowledge graph for the PHP programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Powershell](https://mozilladatacollective.com/datasets/cmpdyjvuj00nhnu07sbmm4giq): N-Triples/RDF export of the StackOverflow knowledge graph for the Powershell programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Prolog](https://mozilladatacollective.com/datasets/cmpdyjnp900n1nv079ccqu2x3): N-Triples/RDF export of the StackOverflow knowledge graph for the Prolog programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Ruby](https://mozilladatacollective.com/datasets/cmpdyjegk00ndnu07rq7ai8g7): N-Triples/RDF export of the StackOverflow knowledge graph for the Ruby programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Rust](https://mozilladatacollective.com/datasets/cmpdyi6wl00mvnv077tm2mqdf): N-Triples/RDF export of the StackOverflow knowledge graph for the Rust programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Scala](https://mozilladatacollective.com/datasets/cmpdyhww200n5nu07ez4y7ppw): N-Triples/RDF export of the StackOverflow knowledge graph for the Scala programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Shell](https://mozilladatacollective.com/datasets/cmpdyhlyw00mrnv07va1jwwko): N-Triples/RDF export of the StackOverflow knowledge graph for the Shell programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Vb.Net](https://mozilladatacollective.com/datasets/cmpdyes6u00mnnv07wyjjzpgg): N-Triples/RDF export of the StackOverflow knowledge graph for the Vb.Net programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - .Net](https://mozilladatacollective.com/datasets/cmpdyeljq00mvnu07y0s26kw5): N-Triples/RDF export of the StackOverflow knowledge graph for the .Net programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Abap](https://mozilladatacollective.com/datasets/cmpdyefi900mrnu076a7ge28i): N-Triples/RDF export of the StackOverflow knowledge graph for the Abap programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Assembly](https://mozilladatacollective.com/datasets/cmpdye7xc00mjnv07ug7zhu9k): N-Triples/RDF export of the StackOverflow knowledge graph for the Assembly programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Python](https://mozilladatacollective.com/datasets/cmpdye1q800mnnu07toknnvps): N-Triples/RDF export of the StackOverflow knowledge graph for the Python programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Typescript](https://mozilladatacollective.com/datasets/cmpdydrp600mfnv07yake2j40): N-Triples/RDF export of the StackOverflow knowledge graph for the Typescript programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - Swift](https://mozilladatacollective.com/datasets/cmpdydl4400mbnv07rzc2sf56): N-Triples/RDF export of the StackOverflow knowledge graph for the Swift programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - R](https://mozilladatacollective.com/datasets/cmpdydc4o00mfnu0761rbiv86): N-Triples/RDF export of the StackOverflow knowledge graph for the R programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - SQL](https://mozilladatacollective.com/datasets/cmpdyd5r500mbnu07ntygcv16): N-Triples/RDF export of the StackOverflow knowledge graph for the SQL programming language. The archive contains the schema file plus language-specific RDF files. - [StackOverflow Knowledge Graph - C++](https://mozilladatacollective.com/datasets/cmpdycv7k00m7nu07co3jmxt8): N-Triples/RDF export of the StackOverflow knowledge graph for the C++ programming language. The archive contains the schema file plus language-specific RDF files. - [Khmer ASR Cultural Dataset (Version 3 - Part 5)](https://mozilladatacollective.com/datasets/cmpdy19nn00llnu07y94pzeu4): 87.11 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 9 seconds. - [Khmer ASR Cultural Dataset (Version 3 - Part 2)](https://mozilladatacollective.com/datasets/cmpdy14ka00lhnu07yzk3ixyn): 92.53 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 9 seconds. - [Khmer ASR Cultural Dataset (Version 3 - Part 1)](https://mozilladatacollective.com/datasets/cmpdy0si400ldnu07rs6zwql4): 134.60 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. - [Khmer ASR Cultural Dataset (Version 3 - Part 7)](https://mozilladatacollective.com/datasets/cmpdy0icy00l9nu07zo5hl1m3): 81.01 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. - [Khmer ASR Cultural Dataset (Version 3 - Part 6)](https://mozilladatacollective.com/datasets/cmpdy0e2c00lxnv071292ocxf): 45.22 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 6 seconds. - [Khmer ASR Cultural Dataset (Version 3 - Part 3)](https://mozilladatacollective.com/datasets/cmpdy0a9h00l5nu078u4y91x1): 81.18 hours of manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. - [Khmer ASR Cultural Dataset (Version 3 - Part 4)](https://mozilladatacollective.com/datasets/cmpdy07op00l1nu07lucrntkj): 64.47 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 7 seconds. - [LegendNER-ID (Javanese)](https://mozilladatacollective.com/datasets/cmpcp7cl800ydmg077wte52mv): The JavLegends-NER dataset was developed to address the scarcity of labeled linguistic resources for Indonesian regional languages, specifically Javanese. It focuses on the domain of folklore and legends, which contains unique linguistic structures and cultural entities. - [TTS Balinese Language](https://mozilladatacollective.com/datasets/cmmm2ru5r003nmd07p53h9wdw): The Balinese TTS dataset is created and narrated by native Balinese speakers with code-mixing in Indonesian. This dataset is designed to showcase the use of the Balinese language in everyday contexts, covering topics such as family, social interactions, and routine community activities. - [Synthetic Ladino Parallel Corpus](https://mozilladatacollective.com/datasets/cmpbmhj4i0067nw07tk46v2jp): This dataset contains over 20 million synthetic parallel sentence pairs for Ladino (Judeo-Spanish) paired with English (5.7M pairs), Spanish (10.3M pairs), and Turkish (4.6M pairs). The data was generated using rule-based Spanish-Ladino translation methods to support the preservation and digital development of this endangered language. - [CoVoST 2 French - English](https://mozilladatacollective.com/datasets/cmpbfyxlc002nmj07k67e3ok2): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the French audio (276 hours) and the translations in English. - [Multispeaker Hindi ASR Dataset](https://mozilladatacollective.com/datasets/cmp8w3owe03phmp076hciar56): This dataset consists of Hindi audio recordings paired with their corresponding text transcriptions. It includes a variety of speech samples that may cover different speakers, accents, speaking styles, and recording conditions, reflecting real-world audio diversity. - [Gujarati News and Blogs Corpus](https://mozilladatacollective.com/datasets/cmp8vqj5h03m0o007qu438s4a): This dataset consists of a collection of Gujarati news articles and blog posts gathered from various online sources. It covers a wide range of topics, including current affairs, politics, lifestyle, culture, and general interest content, providing diverse linguistic patterns and writing styles. - [CoVoST 2 Russian - English](https://mozilladatacollective.com/datasets/cmp787jb102whmp07hbogpvaw): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Russian audio (35 hours) and the translations in English. - [CoVoST 2 Mongolian - English](https://mozilladatacollective.com/datasets/cmp77ufsd02w9mp07q65muzr3): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Mongolian audio (7 hours) and the translations in English. - [CoVoST 2 Turkish - English](https://mozilladatacollective.com/datasets/cmp77tlx802w5mp07dwv1rsq0): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Turkish audio (6 hours) and the translations in English. - [CoVoST 2 Latvian - English](https://mozilladatacollective.com/datasets/cmp77t2kg02weo007s6hkkzec): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Latvian audio (4 hours) and the translations in English. - [TWB Voice 1.0 - Kanuri](https://mozilladatacollective.com/datasets/cmp77rkiv02wao007n5ut2oxr): TWB Voice 1.0 - Kanuri is the Kanuri language portion of the TWB Voice 1.0 multilingual speech corpus, created by CLEAR Global (formerly Translators without Borders). It contains approximately 52 hours of read speech recorded by native Kanuri speakers through the TWB Voice platform. - [CoVoST 2 Tamil - English](https://mozilladatacollective.com/datasets/cmp753p4902t8o007nejqkhlo): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Tamil audio (2 hours) and the translations in English. - [Public GenAI Vulnerability Disclosures](https://mozilladatacollective.com/datasets/cmowvw8oq011hnn07re4l1vi3): 0DIN, the 0Day Investigative Network, was founded by Mozilla in 2024 to reward responsible researchers for their efforts in securing GenAI models. This dataset is a weekly export of the public 0DIN disclosures published at the 0DIN disclosures page, in accordance with the 0DIN Research Terms and Disclosure Policy. - [TWB Voice 1.0 - Hausa](https://mozilladatacollective.com/datasets/cmp74k47s02u2mp07zr9zwhpt): TWB Voice 1.0 - Hausa is the Hausa language portion of the TWB Voice 1.0 multilingual speech corpus, created by CLEAR Global (formerly Translators without Borders). It contains approximately 58 hours of read speech recorded by native Hausa speakers through the TWB Voice platform. - [CoVoST 2 Japanese - English](https://mozilladatacollective.com/datasets/cmp73hbl502swo007eb7ealfq): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Japanese audio (2 hours) and the translations in English. - [CoVoST 2 Dutch - English](https://mozilladatacollective.com/datasets/cmp73gzl502twmp07g6f1bkrx): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Dutch audio (9 hours) and the translations in English. - [CoVoST 2 Portuguese - English](https://mozilladatacollective.com/datasets/cmp72jw8l02sgo007cvzf77lu): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Portuguese audio (17 hours) and the translations in English. - [CoVoST 2 Spanish - English](https://mozilladatacollective.com/datasets/cmp704ive02qymp075orb8ok4): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Spanish audio (143 hours) and the translations in English. - [Khmer ASR Cultural Dataset (V2)](https://mozilladatacollective.com/datasets/cml9h5vgc01bxmn075sjeftek): 106.53 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.42 seconds with the standard deviation of 3.39. - [Khmer ASR Cultural Dataset](https://mozilladatacollective.com/datasets/cmkcy8in2004umo0775mye43g): 37.62 hours manually curated speech-text pairs by native speakers in Khmer language about Cambodian cultural topics. On average, each recording is 8.29 seconds with the standard deviation of 3.87. - [śmigiel - machine generated text detection](https://mozilladatacollective.com/datasets/cmp6qb2al02inmp07dlajrjke): Consisting of 64 538 human-written and machine-generated texts in Polish from various domains, ŚMIGIEL is a comprehensive resource for training and benchmarking Machine-Generated Text (MGT) detection systems focusing on Polish language. The dataset was originally created to for needs of the Shared Task 1 at PolEval 2025 (see it here: http://poleval.pl/tasks/task1) organized by The Linguistic Engineering (LE) Group (learn more at https://zil.ipipan.waw.pl/), part of the Department of Artificial Intelligence at the Institute of Computer Science, Polish Academy of Sciences [IPI PAN (official site: https://ipipan.waw.pl/). - [CoVoST 2 Persian - English](https://mozilladatacollective.com/datasets/cmp5kw3o301hcmp07d8dgejzx): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Persian audio (50 hours) and the translations in English. - [CoVoST 2 Chinese (China) - English](https://mozilladatacollective.com/datasets/cmp5kv8j101hvo007hjamlqet): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Chinese (China) audio (24 hours) and the translations in English. - [CoVoST 2 Italian - English](https://mozilladatacollective.com/datasets/cmp5k3fsy01gmmp07w2kaous5): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Italian audio (67 hours) and the translations in English. - [TWB Voice 1.0 - Shuwa Arabic](https://mozilladatacollective.com/datasets/cmp5ht3jx01dno007i1opx2rj): TWB Voice 1.0 - Shuwa Arabic is the Shuwa Arabic language portion of the TWB Voice 1.0 multilingual speech corpus, created by CLEAR Global (formerly Translators without Borders). It contains approximately 15 hours of read speech recorded by native Shuwa Arabic speakers through the TWB Voice platform. - [CoVoST 2 Indonesian - English](https://mozilladatacollective.com/datasets/cmp5hjood01comp07fxt1fb74): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Indonesian audio (2 hours) and the translations in English. - [CommonLID](https://mozilladatacollective.com/datasets/cmp5c60at015po007bbql6h3s): CommonLID is a community-created language identification (LID) benchmark. CommonLID consists of web text manually annotated for the language that it is written in. - [CoVoST 2 English - Arabic](https://mozilladatacollective.com/datasets/cmp5b9e10014mo007qkcmck75): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio (436 hours) and the translations in Arabic. - [Bangladesh Traffic Signs Dataset](https://mozilladatacollective.com/datasets/cmp4htii700glmp07syfvo010): This dataset consists of traffic sign images collected from Bangladesh, capturing both Bengali signage commonly found on roads. It includes multiple categories such as regulatory, warning, and informational signs, reflecting real-world driving environments. - [Pakistan Traffic Signs Dataset](https://mozilladatacollective.com/datasets/cmp4htnf000gpmp07shg6l5v9): The Pakistan Traffic Signs Dataset is a curated collection of traffic sign images captured under diverse real-world road conditions across Pakistan. The dataset contains multiple categories of regulatory, warning, and informational traffic signs, collected from urban, suburban, and highway environments. - [India Traffic Signs Dataset](https://mozilladatacollective.com/datasets/cmp4htkri00hao007kbrbew2h): This dataset consists of traffic sign images collected from different regions of India, including signs in English and regional languages. It covers a wide range of categories such as regulatory, warning, and informational signs, reflecting real-world road environments. - [ReRooted: Speech Corpus of Testimonials from Armenian Refugees and Immigrants](https://mozilladatacollective.com/datasets/cmp4ijosh00h9mp078ggnaatm): ReRooted is an open-access online YouTube corpus of interviews with Armenian refugees and immigrants. As of now, the online corpus has over 80hrs of recordings on YouTube, alongside subtitles in Armenian. - [Less is More Corpus](https://mozilladatacollective.com/datasets/cmp2hgs2j00nsmp07q9s6g3m0): This repository contains the dataset associated with the paper "Less is More? The Role of Demographic Author Information in Emotion Classification of Ambiguous Text". - [VoxForge - Turkish](https://mozilladatacollective.com/datasets/cmp2h2som00kzno07n8lwesrr): 3 hours (1717 utterances) of read speech of Türkçe (Turkish), collected via the VoxForge project. - [VoxForge - Albanian](https://mozilladatacollective.com/datasets/cmp2h2jnt00n8mp07lkr11xru): 30 minutes (330 utterances) of read speech of Shqip (Albanian), collected via the VoxForge project. - [VoxForge - Ukrainian](https://mozilladatacollective.com/datasets/cmp2h28f400n4mp07aeka38ca): 1 hour (390 utterances) of read speech of Українська (Ukrainian), collected via the VoxForge project. - [VoxForge - Russian](https://mozilladatacollective.com/datasets/cmp2h1zvg00n0mp07wrjxow3l): 15.5 hours (6412 utterances) of read speech of Русский (Russian), collected via the VoxForge project. - [VoxForge - Hebrew](https://mozilladatacollective.com/datasets/cmp2h1qmo00kvno072hav0upy): 32 minutes (307 utterances) of read speech of ‭עברית‬ (Hebrew), collected via the VoxForge project. - [VoxForge - Persian](https://mozilladatacollective.com/datasets/cmp2h1iop00krno07itbxve77): 16 minutes (157 utterances) of read speech of ‭فارسی‬ (Persian), collected via the VoxForge project. - [VoxForge - French](https://mozilladatacollective.com/datasets/cmp2h0vtl00mwmp07ra6g2pre): 39 hours (23752 utterances) of read speech of Français (French), collected via the VoxForge project. - [VoxForge - Croatian](https://mozilladatacollective.com/datasets/cmp2h0m5p00msmp07okqm2fzz): 14 minutes (127 utterances) of read speech of Hrvatski (Croatian), collected via the VoxForge project. - [VoxForge - Portuguese](https://mozilladatacollective.com/datasets/cmp2gp36800m9mp072hodigy6): 4.5 hours (4571 utterances) of read speech of Português (Portuguese), collected via the VoxForge project. - [VoxForge - Dutch](https://mozilladatacollective.com/datasets/cmp2goc8g00m5mp07dst6o64k): 10.5 hours (8440 utterances) of read speech of Nederlands (Dutch), collected via the VoxForge project. - [VoxForge - Italian](https://mozilladatacollective.com/datasets/cmp2ghayo00lzmp07eh8fssxm): 20 hours (10643 utterances) of read speech of Italiano (Italian), collected via the VoxForge project. - [VoxForge - Spanish](https://mozilladatacollective.com/datasets/cmp2ggvbx00lvmp07i0xqlntr): 53 hours (23537 utterances) of read speech of Español (Spanish), collected via the VoxForge project. - [Efik-TTS-Dataset](https://mozilladatacollective.com/datasets/cmp1ma8ty007eo607klisho1a): This dataset comprises audio recordings of Efik speech aligned with textual transcriptions. The dataset is structured into 10 folders, each containing audio files and a corresponding audio-text mapping file. - [Teke-Laali-TTS-Dataset](https://mozilladatacollective.com/datasets/cmj8g2kw902fwmb07hub8puq8): The dataset contains paired audio and text resources for Teke-Laali, a Bantu language spoken in the Congo. It consists of seven folders containing a total of 9,069 audio clips from raw audio recordings, with a total duration of 7:01:50.126 (HH:MM:SS.mmm). - [Laari-TTS-Dataset](https://mozilladatacollective.com/datasets/cmj324gbx00p4ny078pi26kfz): The dataset contains audio and text resources on Laari, a Bantu language spoken in the Congo. The resources, which are suitable for TTS tasks and possibly ASR tasks, consist of the following: - 6,311 audio clips totalling 241 minutes and 44.97 seconds; - an audio mapping file with 5,321 lines, each beginning with the name of an audio file, followed by a tab and then the corresponding text excerpt; - two raw audio files totalling 120 minutes and 54.90 seconds; - two long audio files with their original, non-split transcription files, for a total duration of 120 minutes and 41.90 seconds. - [Bomitaba-TTS-Dataset](https://mozilladatacollective.com/datasets/cmj2rze7r00j5ny07uhs85go2): The dataset comprises three components: audio clips, an audio mapping file, and raw audio of Bomitaba, a Bantu language spoken in the Congo. Each audio clip is paired with its corresponding transcription. - [Suundi-TTS-Dataset](https://mozilladatacollective.com/datasets/cmj1qhr5n001wnv07hqcatq25): The dataset consists of paired audio and text data on Suundi (sdj), a language spoken in Congo. The audio corpus consists of 4,187 clips read by one speaker totaling 188 min 22.68 sec. - [Mbosi-TTS-Dataset](https://mozilladatacollective.com/datasets/cmj1gdg1s00vrnu07nmlrax7g): The dataset consists of paired audio and text data on Mbosi (mdw), a language spoken in Congo. The audio corpus consists of 2,575 clips read by one speaker totaling 275 min 48.35 sec. - [Beembe-TTS-Dataset](https://mozilladatacollective.com/datasets/cmj1gd6j400uvnw07ylungxjz): The dataset consists of paired audio and text data on Beembe (beq), a language spoken in Congo. The audio corpus consists of 6,933 clips read by one speaker totaling 275 min 48.35 sec. - [Yaka-TTS-Dataset](https://mozilladatacollective.com/datasets/cmj0c3cpe003enw07zxujgdt1): Paired audio and text data on Yaka (also known as West Teke), a language spoken in Congo. The audio corpus consists of 7,648 clips read by one speaker for a total duration of 344 min 40.48 sec. - [Kituba-TTS-Dataset](https://mozilladatacollective.com/datasets/cmj0az3yz002enw07okqee8jr): Paired audio and text data on Kituba (mkw), a language spoken in Congo. The audio corpus consists of 8,302 clips read by one speaker, totalling 350 min 11.98 sec. - [Bulu_ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmnopvxxr00t6mf07nm4cp4qs): ALCAM-Bulu-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Bulu language, a Bantu language spoken in southern Cameroon. The dataset comprises three closely aligned components: (i) a structured datasheet containing carefully selected example sentences and lexical entries reflecting usage in the Bulu language; (ii) high-quality audio recordings of these sentences and lexical items, produced by a native speaker; and (iii) an explicit audio–sentence mapping file enabling precise alignment between the textual and acoustic data. - [Basaa-ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmind910n0096nx07gme9v3wj): This dataset comprises a datasheet of lexical entries in Basaa, accompanied by illustrative sentences, word-by-word glosses, and corresponding translations in French. Each entry is enriched with aligned audio recordings, making the resource suitable for linguistic analysis and speech technology development. - [Bati-MultiDialectalASR-Dataset](https://mozilladatacollective.com/datasets/cmj8fdiyv02egmb07m2wfdltm): This dataset contains paired audio and text resources for three Bati dialects (Kelleng, Mbougue, and Nyambat), which belong to the Yambasa group of Bantu languages found in Cameroon. It contains 13,344 audio clips totalling 6 hours, 8 minutes and 12.286 seconds and 44 audio/text mapping files totalling 13,346 lines. - [Igbo-TTS-Dataset](https://mozilladatacollective.com/datasets/cmp0f5u2n02a8mp0771ryoly8): This dataset comprises audio recordings of Igbo speech aligned with textual transcriptions. The dataset is structured into 17 folders, each containing audio files and a corresponding audio-text mapping file. - [Adamawa Fulfulde-TTS-Dataset](https://mozilladatacollective.com/datasets/cmp0f3tib02a4mp07ch5tr808): This dataset comprises 1,303 high-quality audio recordings of read speech produced by a single Adamawa Fulfulde speaker over several months. Adamawa Fulfulde (ISO 639-3: fub), also known as Fula Adamawa, is a language of the Niger-Congo family spoken in the Adamawa region of Cameroon and in adjacent areas of Chad and Nigeria. - [Spoken-Congolese-French-Dataset](https://mozilladatacollective.com/datasets/cmk1bdl0q39v6mb07gee6udp2): The dataset consists of paired audio and text resources on spoken French from the Republic of the Congo. The audio files were extracted from longer recordings of semi-guided interviews conducted in Brazzaville, and orthographic transcriptions were added. - [Bamun-French Parallel Corpus 2.0](https://mozilladatacollective.com/datasets/cmn6ay1xg016jnv07drtmw7qo): This dataset is an extended and updated version of the 'Bamun-French Parallel Corpus 1.1' that is published on the Mozilla Data Collective platform. It is a parallel corpus of 4,444 lines in Bamun and French suitable for machine translation tasks. - [Naija-TTS-Dataset](https://mozilladatacollective.com/datasets/cmnykldcz010knu0737o8bgh9): This dataset comprises audio recordings of Nigerian Pidgin English speech aligned with textual transcriptions. The dataset is structured into 16 folders, each containing audio files and a corresponding audio-text mapping file. - [Bamun-TTS-Dataset](https://mozilladatacollective.com/datasets/cmnhjbnjp0115mh07kiha0rei): This dataset comprises audio recordings of Bamun (Shupamem) speech aligned with textual transcriptions. The dataset is structured into 24 folders totalling 4h 30m 25s, each containing audio files and a corresponding audio-text mapping file. - [Hausa-TTS-Dataset](https://mozilladatacollective.com/datasets/cmnopto3q00t0mf07v2dtc0ej): This dataset comprises audio recordings of Hausa speech aligned with textual transcriptions. The dataset is structured into 19 folders, each containing audio files and a corresponding audio-text mapping file. - [Yoruba-TTS-Dataset](https://mozilladatacollective.com/datasets/cmo1nlaah0071mk077mw0qhpv): This dataset comprises audio recordings of Yoruba speech aligned with textual transcriptions. The dataset is structured into 17 folders, each containing audio files and a corresponding audio-text mapping file. - [isiXhosa-TTS-Dataset](https://mozilladatacollective.com/datasets/cmo4gtixz00kwny07hayfsk8s): This dataset comprises audio recordings of isiXhosa speech aligned with textual transcriptions. The dataset is structured into 24 folders, each containing audio files and a corresponding audio-text mapping file. - [Tiv-TTS-Dataset](https://mozilladatacollective.com/datasets/cmo4nmfam00nxny07rssox2tj): This dataset comprises audio recordings of Tiv speech aligned with textual transcriptions. The dataset is structured into 14 folders, each containing audio files and a corresponding audio-text mapping file. - [CoVoST 2 Catalan - English](https://mozilladatacollective.com/datasets/cmoz2e4oo01n7nt079rbo308t): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Catalan audio (157 hours) and the translations in English. - [CoVoST 2 Welsh - English](https://mozilladatacollective.com/datasets/cmoyrppho01kvnt075pt9pj5x): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Welsh audio (3 hours) and the translations in English. - [CoVoST 2 Arabic-English](https://mozilladatacollective.com/datasets/cmoyfkei701eynt07ut2uoiif): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the Arabic audio and the translations in English. - [VoxForge - German](https://mozilladatacollective.com/datasets/cmoxa7ndg01e1ny07zpwgtu22): 32 hours (24685 utterances) of read speech of Deutsch (German), collected via the VoxForge project. - [VoxForge - Greek](https://mozilladatacollective.com/datasets/cmowwx2gq012hnn073izoljfz): 4 hours (1377 utterances) of read speech of Ελληνικά (Greek), collected via the VoxForge project. - [VoxForge - Catalan](https://mozilladatacollective.com/datasets/cmowvzwts011lnn078f3719jg): 39 minutes of Catalan read speech, collected as part of the VoxForge project. - [VoxForge - Bulgarian](https://mozilladatacollective.com/datasets/cmovv4qrj00hcmk07yuaxyvzv): 1 hour of read speech and transcriptions in Bulgarian. - [CoVoST 2 German-English](https://mozilladatacollective.com/datasets/cmou2fdyx015fl307bux4c4gi): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the German audio and the translations in English. - [CoVoST 2 English-Slovenian](https://mozilladatacollective.com/datasets/cmosu28090085ny07x7d8ezoq): CoVoST 2 is a large-scale multilingual speech to text translation corpus based on Mozilla Common Voice 4.0. This segment of the corpus contains the English audio and the translations in Slovenian. - [TWB Parallel Sentence kits - Congo Swahili (25k)](https://mozilladatacollective.com/datasets/cmosl07v400w9nu07g3puif2t): The Congo Swahili portion of CLEAR Global's Gamayun Language Data Kits — 25,305 parallel French–Congo Swahili sentences in total, distributed across three independent kit sizes: 5,000 (`kit5k`), 10,000 (`kit10k`), and 10,305 (`kit15k`, partial delivery of the medium-kit target). French source sentences were drawn from the Tatoeba repository using a selection algorithm that ensures representation of the most frequently used words in French; Congo Swahili translations were produced by professionals and volunteers of the Translators without Borders (now CLEAR Global) translator community. - [TWB Parallel Sentence kits - Swahili (5k)](https://mozilladatacollective.com/datasets/cmoskxn8k00vtmj07ubxtu0f2): The Swahili portion of CLEAR Global's Gamayun Language Data Kits — 5,000 parallel English–Swahili sentences (the `kit5k` mini-kit). English source sentences were drawn from the Tatoeba repository using a selection algorithm that ensures representation of the most frequently used words in English; Swahili translations were produced by professionals and volunteers of the Translators without Borders (now CLEAR Global) translator community. - [TWB Parallel Sentence kits - Rohingya (5k)](https://mozilladatacollective.com/datasets/cmoskwsac00w3nu07b1nydlfb): The Rohingya portion of CLEAR Global's Gamayun Language Data Kits — 5,000 parallel English–Rohingya sentences (the `kit5k` mini-kit). English source sentences were drawn from the Tatoeba repository using a selection algorithm that ensures representation of the most frequently used words in English; Rohingya translations were produced by professionals and volunteers of the Translators without Borders (now CLEAR Global) translator community. - [TWB Parallel Sentence kits - Lingala (5k)](https://mozilladatacollective.com/datasets/cmosknxap00vlmj07kf6mugba): The Lingala portion of CLEAR Global's Gamayun Language Data Kits — 5,000 parallel French–Lingala sentences (the `kit5k` mini-kit). French source sentences were drawn from the Tatoeba repository using a selection algorithm that ensures representation of the most frequently used words in French; Lingala translations were produced by professionals and volunteers of the Translators without Borders (now CLEAR Global) translator community. - [TWB Parallel Sentence kits - Nande (15k)](https://mozilladatacollective.com/datasets/cmoskn32a00vhmj07prz6k5ng): The Nande portion of CLEAR Global's Gamayun Language Data Kits — 15,000 parallel French–Nande sentences in total, distributed across two independent kit sizes: 5,000 (`kit5k`) and 10,000 (`kit10k`). French source sentences were drawn from the Tatoeba repository using a selection algorithm that ensures representation of the most frequently used words in French; Nande translations were produced by professionals and volunteers of the Translators without Borders (now CLEAR Global) translator community. - [TWB Parallel Sentence kits - Tigrinya (5k)](https://mozilladatacollective.com/datasets/cmoskmbpj00vxnu07w8lu7rrk): The Tigrinya portion of CLEAR Global's Gamayun Language Data Kits — 5,000 parallel English–Tigrinya sentences (the `kit5k` mini-kit). English source sentences were drawn from the Tatoeba repository using a selection algorithm that ensures representation of the most frequently used words in English language; Tigrinya translations were produced by professionals and volunteers of the Translators without Borders (now CLEAR Global) translator community. - [TWB Parallel Sentence kits - Kanuri (5k)](https://mozilladatacollective.com/datasets/cmoskkop100vjnu0775hclbd1): The Kanuri portion of CLEAR Global's Gamayun Language Data Kits — 5,000 parallel English–Kanuri sentences (the `kit5k` mini-kit). English source sentences were drawn from the Tatoeba repository using a selection algorithm that ensures representation of the most frequently used words in English; Kanuri translations were produced by professionals and volunteers of the Translators without Borders (now CLEAR Global) translator community. - [Yezoum_ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmndjfw9d000qmj070dece40f): Yezoum_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Yezoum variety of the widely designated 'Ewondo language', or sometimes of the macro linguistic group known as Beti or Beti-Fang. Yezoum is a localised and socially embedded speech form that is rarely represented in standard grammatical descriptions or lexicographical resources. - [The LJSpeech Dataset](https://mozilladatacollective.com/datasets/cmonjhxee01ako007kohpbg34): This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books. A transcription is provided for each clip. - [TWB Parallel Sentence kits - Hausa (30k)](https://mozilladatacollective.com/datasets/cmom43ixg00hao00731e0o0jg): The Hausa portion of CLEAR Global's Gamayun Language Data Kits — 30,000 parallel English–Hausa sentences in total, distributed across three independent kit sizes: 5,000 (`kit5k`), 10,000 (`kit10k`), and 15,000 (`kit15k`). English source sentences were drawn from the Tatoeba repository using a selection algorithm that ensures representation of the most frequently used words in English; Hausa translations were produced by professionals and volunteers of the Translators without Borders (now CLEAR Global) translator community. - [TTS Bugis - Barru Dialect: Language and Identity](https://mozilladatacollective.com/datasets/cmom41aj500h6o007zkt02p8n): The Bugis language is one of the major regional languages spoken in South Sulawesi, Indonesia, particularly in areas such as Bone, Wajo, Soppeng, Sidrap, Sinjai, and Pare-pare. It consists of several dialects, including Bone, Wajo, Soppeng and Barru, with the Barru dialect showing distinctive lexical and phonological features. - [BAHANA-Manggarai TTS](https://mozilladatacollective.com/datasets/cmom3z1mc00h2o0071zxeq1ur): This dataset supports an underrepresented language in technology by contributing to the development of speech recognition for the Manggarai language. Manggarai is part of the Austronesian language family and is primarily spoken on the island of Flores, Indonesia. - [BAHANA-Betawi TTS](https://mozilladatacollective.com/datasets/cmom3xala00gyo007hc0evymi): BAHANA-Betawi TTS is a Betawi language dataset that represents the language dynamics of communities around Indonesia's urban administrative centers. This dataset consists of Betawi variations found in West Java Province and Betawi dialects around the center of Jakarta Province, resulting in strong contact between Indonesian and English. - [Tamazight Open Speech Dataset](https://mozilladatacollective.com/datasets/cmok1w0j002jcmr075bsof72y): This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice recordings and corresponding text transcripts in Standard Moroccan Amazigh (ⵜⴰⵎⴰⵣⵉⵖⵜ ⵜⴰⵏⴰⵡⴰⵢⵜ ⵜⴰⵎⵓⵔⴰⴽⵓⵛⵜ) intended for training Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models. - [Italian TTS - female voice](https://mozilladatacollective.com/datasets/cmoiuyem401j5mr07s0jx8rqr): A single-speaker read speech dataset in Italian. The dataset contains ~10 hours of pre-segmented utterances, recorded by an anonymous 66-years-old female Italian speaker. - [Read Speech in Kenyan Swahili (6h)](https://mozilladatacollective.com/datasets/cmocxuhkn00y4md07d8bllj2x): A single-speaker read speech dataset in Kenyan Swahili, produced as part of CLEAR Global's Gamayun Language Data Kits initiative. The dataset contains 4,700 pre-segmented utterances (~6 hours, 21,852 seconds) recorded by an anonymous male Kenyan speaker. - [Imágenes de Señalamientos en México](https://mozilladatacollective.com/datasets/cmo1kqs2a004anr07sxn2mtf5): Una colección de imágenes anotadas de señales de tránsito y otras señales viales en México. Cada imagen cuenta con anotaciones estructuradas que incluyen bounding boxes y etiquetas de clasificación, lo que lo hace adecuado para tareas de visión por computadora como detección de objetos, clasificación de señales y análisis de escenas viales. - [Yoruba-English Code-Switching (YECS) Corpus](https://mozilladatacollective.com/datasets/cmo09pqp300gbnx07xcl42los): The Yoruba-English Code-Switching (YECS) Corpus is a comprehensive, ~120-hour dataset designed to capture the natural linguistic phenomenon of intra-sentential code-mixing. Curated by the LynguaTech Innovative Foundation (LyngualLabs), this dataset provides nearly 100,000 validated audio-text pairs recorded by 140 demographically diverse bilingual speakers in Nigeria. - [RFE/RL Belarusian News Text Corpus](https://mozilladatacollective.com/datasets/cmobujn7b00aenw07wcn5dntu): This dataset serves as a comprehensive longitudinal news corpus for the Belarusian language, sourced from Radio Svaboda (svaboda.org), the Belarusian service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from October 1997 to March 2026, the corpus contains 338,937 unique articles, totaling over 134 million tokens. - [RFE/RL Afghan Dari News Text Corpus](https://mozilladatacollective.com/datasets/cmobui7g400cbmd07kmfe7oj4): This dataset serves as a comprehensive longitudinal news corpus for the Afghan Dari language, sourced from Radio Azadi (da.azadiradio.com), the Afghan service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from November 2000 to March 2026, the corpus contains 212,870 unique articles, totaling over 50 million tokens. - [RFE/RL Afghan Pashto News Text Corpus](https://mozilladatacollective.com/datasets/cmobugzgs00c7md0783t69yd7): This dataset serves as a comprehensive longitudinal news corpus for the Pashto language, sourced from Radio Azadi (pa.azadiradio.com), the Afghan service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from January 2000 to March 2026, the corpus contains 204,817 unique articles, totaling over 16 million tokens. - [RFE/RL Armenian News Text Corpus](https://mozilladatacollective.com/datasets/cmobufn0500a5nw07u01ivqa6): This dataset serves as a comprehensive longitudinal news corpus for the Armenian language, sourced from Radio Azatutyun (azatutyun.am), the Armenian service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from June 1999 to March 2026, the corpus contains 232,925 unique articles, totaling over 54 million tokens across both Armenian and English texts. - [Akoose-ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmj0b1ymo002pnu076g7oseat): This dataset comprises a datasheet of Akoose (bss) lexical entries collected from the 'Western-Bakossi' subgroup. Each entry is accompanied by illustrative sentences, French translations and a word-by-word breakdown of the Akoose sentences, as well as an equivalent breakdown in English. - [Ewondo-Yanda-ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmiv2rbv901mwmf07iyhv4pva): This dataset comprises a datasheet of Ewondo (Ewo) lexical entries collected in the Yanda subgroup. Each entry is accompanied by illustrative sentences, word-by-word glosses and French translations. - [Ewondo_Mbida-Mbani_ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmk1bbz7o39v2mb07jjupvzud): This dataset comprises a datasheet of Ewondo (Ewo) lexical entries collected in the speech area known as Mbida Mbani. Each entry is accompanied by illustrative sentences, word-by-word glosses and French translations. - [Ewondo_Fong_ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmkmoepzx018wnw07cnkupls3): Ewondo_Fong_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Fong variety of the widely designated 'Ewondo language'. Fong is a localised and socially embedded speech form that is rarely represented in standard grammatical descriptions or lexicographical resources. - [Mvele_ALCAM-MultimodalDataset](https://mozilladatacollective.com/datasets/cmmyyd6jq00k5mf07tjn4xh3j): Mvele_ALCAM-MultimodalDataset is a richly curated, multimodal linguistic dataset dedicated to the documentation and technological enhancement of the Mvele variety of the widely designated 'Ewondo language'. Mvele is a localised and socially embedded speech form that is rarely represented in standard grammatical descriptions or lexicographical resources. - [Sentence translation difficulty in Spanish - BOUQuET](https://mozilladatacollective.com/datasets/cmngbf1tt0050nn07i49aebnk): This dataset is a collection of sentences in Spanish from the BOUQuET benchmark (total 1990 sentences) which have been annotated with sentence translation difficulty scores on a Likert scale. The annotators are speakers of three Indigenous languages of the Mexico and Guatemala and scored the sentences as part of the work on translating the benchmark into their languages. - [Sylheti Text Corpus by Haque Publishers](https://mozilladatacollective.com/datasets/cmo8rj0f60258mg0783ls9gjb): The Sylheti Text Corpus is a professionally assembled linguistic dataset comprising approximately 524,000 tokens of the Sylheti language. This collection focuses on drama scripts, which preserve cultural folklore, social nuances, and everyday idiomatic expressions. - [Rohingya Literature Corpus](https://mozilladatacollective.com/datasets/cmo8q6f9a024mmg0708ccv04j): The Rohingya Literature Corpus is an exceptionally rare linguistic resource comprising approximately 613,500 tokens. This dataset is uniquely characterized by its use of the Myanmar (Burmese) script to represent the Rohingya language, which is very difficult to find in available linguistic resources. - [Noakhalian (নোয়াখাইল্লা) Text Corpus](https://mozilladatacollective.com/datasets/cmo8oicwk020dnt077o5dqhwm): The Noakhalian (নোয়াখাইল্লা) Text Corpus is a systematically curated linguistic resource comprising approximately 504,500 tokens of the Noakhalian language variety. This dataset is primarily composed of drama scripts, a multifaceted domain that provides rich insights into the phonetic, morphological, and sociolinguistic nuances of the Greater Noakhali region. - [Adamawa Fulfulde-French Parallel Corpus of Narratives 1.2](https://mozilladatacollective.com/datasets/cml5asbhf009sme079y6sa9hm): This dataset is the second iteration of the 'Adamawa Fulfulde–French Parallel Corpus of Narratives 1.2'. It comprises a parallel corpus of Adamawa Fulfulde–French literary texts, most of which are narratives. - [Rangpuri (অংপুরি Ôṅgpuri) Text Corpus](https://mozilladatacollective.com/datasets/cmo3f6wem001wny0792yg26as): The Rangpuri (অংপুরি Ôṅgpuri) Text Corpus is a professionally curated collection of approximately 501,500 tokens representing the linguistic and cultural heritage of the Rangpur Division in Bangladesh and neighboring regions. The dataset includes a diverse range of genres such as poetry, folklore, and drama scripts, bringing together everyday social themes, cultural expressions, and oral traditions in one resource. - [Chittagonian (চাটগাঁইয়া, saṭgãia) Text Corpus](https://mozilladatacollective.com/datasets/cmo3f5ydj001umo07y8579zna): The Chittagonian (চাটগাঁইয়া, saṭgãia) Text Corpus is a curated collection of approximately 690,000 tokens reflecting the distinct linguistic and cultural identity of the Greater Chittagong region in Bangladesh. This dataset features a large collection of drama scripts, a unique domain that captures folklore, everyday social themes, and traditional cultural expressions within a single narrative framework. - [Speech Corpus of English Learners from Mexico](https://mozilladatacollective.com/datasets/cmo29suhm00gjo2078rhmqn3p): A corpus of read speech by learners of English living in Mexico. The current version represents 8 speakers and makes up nearly 8 hours of recorded speech. - [Highland Puebla Nahuatl Spoken Image Descriptions](https://mozilladatacollective.com/datasets/cmo29siht00gfo207oj2jrnrd): 100 culturally-salient images from the Sierra Norte of Puebla, Mexico, with spoken descriptions in Highland Puebla Nahuatl. - [Archivo GELED: Muestra general de audios del cuicateco](https://mozilladatacollective.com/datasets/cmo29qkrl00g5o207hash6gk2): Corpus de 3 horas de audio transcrito de diferentes comunidades de habla cuicateca. Respecto de la clasificación del INALI, la muestra cubre de manera más extensa la variante del cuicateco del centro (San Juan Tepeuxila, Santos Reyes Pápalo, San Lorenzo Pápalo), seguida de la del norte (San Andrés Teotilálpam) y en menor medida del oriente (Colonia Constitución). - [IsiZulu Second Language Learner Speech Corpus](https://mozilladatacollective.com/datasets/cmo1y8vrv006ol207lc86hc13): This corpus is specifically designed to assist in evaluating the performance of pronunciation feedback tools for second language learning. The corpus is comprised of gold standard recordings from isiZulu teachers (2,493 recordings) and recordings from isiZulu L2 learners that have been annotated by isiZulu teachers for phonemic and tonal pronunciation errors (9,639 recordings). - [Modern Greek Dictionary](https://mozilladatacollective.com/datasets/cmo1sklpo00d0mk07taoepe1a): This dataset is a structured digital export of the Triantafyllides Modern Greek Dictionary (Λεξικό της Κοινής Νεοελληνικής — Dictionary of Standard Modern Greek), sourced from the (Greek Language Portal: https://www.greek-language.gr/greekLang/modern_greek/tools/lexica/triantafyllides/). It contains 46,745 entries covering the full Modern Greek lexicon. - [ERT Press](https://mozilladatacollective.com/datasets/cmo1sg74a00cwmk07q6nin2of): This dataset is a structured digital collection of press releases and news articles sourced from the official press platform of the Hellenic Broadcasting Corporation (https://press.ert.gr). It contains 18,979 entries representing the official communication and archival record of the national Greek public broadcaster. - [Ladino-Spanish Lexical Resources](https://mozilladatacollective.com/datasets/cmo1qb62200anmk07ls9feuh5): Ladino-Spanish Lexical Resources is a collection of four lexical files for Ladino (Judeo-Spanish) and Spanish, compiled by Col·lectivaT and the Sephardic Center of Istanbul for use in a rule-based Spanish-to-Ladino machine translation system. The files include a Spanish–Ladino phrase dictionary (136 entries) digitized from the printed dictionary "Diksionaryo de Ladino a Espanyol" by Güler, Portal i Tinoco, a Spanish–Ladino word list (3,884 entries), a list of irregular Spanish verbs (1,299 entries), and a Spanish–Ladino conjugated verb pairs list (2,379 entries). - [Ladino: Una Fraza al Diya](https://mozilladatacollective.com/datasets/cmo1krloc004zmk07fon30uqs): "Una fraza al diya" (A Phrase a Day) is a Ladino language learning dataset prepared by Karen Sarhon of the Sephardic Center of Istanbul (SKAD). It consists of 307 sentences in Ladino (Judeo-Spanish) with parallel translations in Spanish, Turkish, and English. - [Şalom Ladino Corpus](https://mozilladatacollective.com/datasets/cmo1ks4zv004enr07la1rkr9x): Şalom Ladino Articles is a monolingual text corpus in Ladino (Judeo-Spanish), compiled from 397 articles published in the Judeo-Espanyol section of Şalom newspaper. The corpus contains 176,843 words, provided as a single segmented and shuffled plain-text file. - [Bulu-TTS-Dataset 1.0](https://mozilladatacollective.com/datasets/cml9iik7d01efmn07miuf8yof): This dataset comprises denoised audio recordings of read speech from a single Bulu male speaker. Bulu is a Bantu language spoken in Cameroon. - [Ewondo-TTS-Dataset](https://mozilladatacollective.com/datasets/cml16fpkn009lnt07ht6k406o): This dataset comprises high-quality audio recordings of read speech from a single female speaker of Ewondo, a Bantu language spoken in Cameroon. The dataset also contains audio/sentence mapping files, making it suitable for TTS tasks on Ewondo. - [Lingala-TTS-Dataset](https://mozilladatacollective.com/datasets/cmm23jslb003znq07vdska54l): The dataset contains audio and text resources in Lingala, a Bantu language spoken in the Republic of the Congo (also known as 'Congo Brazzaville') and the Democratic Republic of the Congo (DRC). These resources are suitable for TTS and ASR tasks and consist of the following: - 8,572 audio clips totalling 4 hours, 25 minutes and 54 seconds; - an audio mapping file containing 8,572 lines. - [Zacatlán Tepetzintla Nahuatl ASR Dataset](https://mozilladatacollective.com/datasets/cmls27zfd0043ma07mxvsz8zg): An ASR dataset of Zacatlán-Ahuacatlán-Tepetzintla (Western Sierra Puebla) Nahuatl, ISO 639-3 nhi. This is a derivative work of the Zacatlán Tepetzintla Nahuatl Audio and Transcriptions datasets. - [Kanuri Books Corpus](https://mozilladatacollective.com/datasets/cmo1hv2fm001xmk07ity2g0xw): A text corpus of 10,281 randomized sentences (90,706 words) extracted from books by Kanuri authors Dr. Baba Kura Alkali Gazali, Lawan Dalama, Kaka Gana Abba, and Lawan Hassan. - [LibriVox Italian TTS Female Voice](https://mozilladatacollective.com/datasets/cmo0qtw7v003knt07u6yupncc): 4 hours of sentence-aligned speech/text from "Le avventure di Pinocchio" by Carlo Collodi, on LibriVox, containing 2,175 utterances and 41,642 tokens. - [LibriVox Czech TTS Female Voice](https://mozilladatacollective.com/datasets/cmo0jfvnw00p1nx070preklt5): 2 hours of sentence-aligned speech/text from "Krysař" by Viktor Dyk, on LibriVox, containing over 1,500 sentences and nearly 16,000 words. - [UK Sort Codes - ASR Evaluation](https://mozilladatacollective.com/datasets/cmo0h34b600n8nx07khjct4iy): This dataset consists of 1,000 UK bank sort codes read out aloud by a single male speaker of UK English from the Midlands. Sort codes are used to identify bank branches for making money transfers in the United Kingdom. - [Awal Tamazight Dataset](https://mozilladatacollective.com/datasets/cmo07ixww00ddmr07enqrrpfw): This dataset is a compilation of Tamazight (zgh) language resources created by CIEMEN as part of the Awal project (https://awaldigital.org/), with funding from the Municipality of Barcelona and the Government of Catalonia. It includes 1,002 monolingual sentences from a Tamazight language learning material, and over 417,000 parallel sentence pairs spanning multiple language pairs: English–Tamazight, French–Tamazight, Catalan–Tamazight, Spanish–Tamazight, and Arabic–Tamazight. - [RFE/RL Serbian, Bosnian, and Montenegrin (Balkan) News Text Corpus](https://mozilladatacollective.com/datasets/cmo06ikpt00chnx07xmaj8rxe): This dataset serves as a comprehensive longitudinal news corpus for the Serbian, Bosnian, and Montenegrin languages, sourced from Radio Slobodna Evropa (slobodnaevropa.org), the Balkan service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from December 2003 to March 2026, the corpus contains 389,883 unique articles, totaling over 24 million tokens. - [RFE/RL Bulgarian News Text Corpus](https://mozilladatacollective.com/datasets/cmnz26xn200hgnr07m1x2po5v): This dataset serves as a comprehensive longitudinal news corpus for the Bulgarian language, sourced from Radio Svobodna Evropa (svobodnaevropa.bg), the Bulgarian service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from January 2019 to March 2026, the corpus contains 26,753 unique articles, totaling over 8.3 million tokens. - [RFE/RL Azerbaijani News Text Corpus](https://mozilladatacollective.com/datasets/cmnz269cd00hcnr075v74km69): This dataset serves as a comprehensive longitudinal news corpus for the Azerbaijani language, sourced from Radio Azadlıq (azadliq.org), the Azerbaijani service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from January 2005 to March 2026, the corpus contains 239,110 unique articles, totaling over 37 million tokens across both Azerbaijani and Russian texts. - [RFE/RL Belarusian News Text Corpus](https://mozilladatacollective.com/datasets/cmnz260wu00h8nr07h3uf4499): This dataset serves as a comprehensive longitudinal news corpus for the Belarusian language, sourced from Radio Svaboda (svaboda.org), the Belarusian service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from October 1997 to March 2026, the corpus contains 338,937 unique articles, totaling over 134 million tokens. - [RFE/RL Macedonian News Text Corpus](https://mozilladatacollective.com/datasets/cmnz0epgs00e4nx071mq5eiz1): This dataset serves as a comprehensive longitudinal news corpus for the Macedonian language, sourced from Radio Slobodna Evropa (slobodnaevropa.mk), the Macedonian service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from May 2002 to March 2026, the corpus contains 204,934 unique articles, totaling over 46 million tokens. - [LibriVox Croatian TTS Male Voice](https://mozilladatacollective.com/datasets/cmnyx788r00d3nr07u04smvap): 4 hours of sentence-aligned speech/text from "Priče iz Davnine" by Ivana Brlić Mažuranić (1874 - 1938) on LibriVox, containing over 2,000 sentences and 31,000 words. - [RFE/RL Romanian (Moldova) News Text Corpus](https://mozilladatacollective.com/datasets/cmnyq2hwx005qnx07lakjchhb): This dataset serves as a comprehensive longitudinal news corpus for the Romanian language focused on Moldova, sourced from Europa Liberă Moldova (moldova.europalibera.org), the Moldovan service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from October 2002 to March 2026, the corpus contains 244,404 unique articles, totaling over 63 million tokens across Romanian, Russian, and English texts. - [RFE/RL Tajik News Text Corpus](https://mozilladatacollective.com/datasets/cmnypmsyf005enx07bc7nr16l): This dataset serves as a comprehensive longitudinal news corpus for the Tajik language, sourced from Radio Ozodi (ozodi.org), the Tajik service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from June 2000 to March 2026, the corpus contains 166,721 unique articles, totaling over 20 million tokens across Tajik and Russian texts. - [Punjabi 10 Hours TTS](https://mozilladatacollective.com/datasets/cmnypcx5p004jnr07k6ptxwcq): The Punjabi TTS Dataset (Shahmukhi) is a high-quality speech corpus containing approximately 10 hours of read speech in Punjabi written in the Shahmukhi script. It is designed to support text-to-speech development, speech synthesis research, pronunciation modeling, and broader language technology work for Punjabi in its Perso-Arabic writing tradition. - [RFE/RL Turkmen News Text Corpus](https://mozilladatacollective.com/datasets/cmnym5gkr0033nx07oz0gg63f): This dataset serves as a comprehensive longitudinal news corpus for the Turkmen language, sourced from Azatlyk Radiosy (azathabar.com), the Turkmen service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from June 2009 to March 2026, the corpus contains 64,949 unique articles, totaling over 16.5 million tokens across Turkmen and Russian texts. - [RFE/RL Kyrgyz News Text Corpus](https://mozilladatacollective.com/datasets/cmnylnboa001gnr076jrr8vhp): This dataset serves as a comprehensive longitudinal news corpus for the Kyrgyz language, sourced from Radio Azattyk (azattyk.org), the Kyrgyz service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from October 2002 to March 2026, the corpus contains 352,396 unique articles, totaling over 79 million tokens across Kyrgyz, Russian, and English texts. - [RFE/RL Georgian News Text Corpus](https://mozilladatacollective.com/datasets/cmnyln0rl0012nx07wd8cnyis): This dataset serves as a comprehensive longitudinal news corpus for the Georgian language, sourced from Radio Tavisupleba (radiotavisupleba.ge), the Georgian service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from December 2001 to March 2026, the corpus contains 238,829 unique articles, totaling over 38 million tokens. - [RFE/RL Kazakh News Text Corpus](https://mozilladatacollective.com/datasets/cmnylmd5c000ynx07qonxot0v): This dataset serves as a comprehensive longitudinal news corpus for the Kazakh language, sourced from Radio Azattyq (azattyq.org), the Kazakh service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from May 2003 to March 2026, the corpus contains 139,065 unique articles, totaling over 35 million tokens. - [RFE/RL Crimean Tatar News Text Corpus](https://mozilladatacollective.com/datasets/cmnykmjx6010xny079bqway5r): This dataset serves as a comprehensive longitudinal news corpus for the Crimean Tatar language, sourced from Qırım.Aqiqat (ktat.krymr.com), the Crimean Tatar-language broadcaster of Radio Free Europe/Radio Liberty's Crimea.Realities project. Spanning from March 2014 to March 2026, the corpus contains 32,684 unique articles, totaling over 7.5 million tokens. - [RFE/RL Chechen News Text Corpus](https://mozilladatacollective.com/datasets/cmnykm7yi010rny0737lq2fth): This dataset serves as a comprehensive longitudinal news corpus for the Chechen language, sourced from Radio Marsho (radiomarsho.com), the Chechen-language broadcaster of the North Caucasus Service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from April 2006 to March 2026, the corpus contains 30,192 unique articles, totaling over 8.4 million tokens. - [RFE/RL Hungarian News Text Corpus](https://mozilladatacollective.com/datasets/cmnxk95so00a6ny07vtscd9a9): This dataset serves as a complete historical news corpus for the Hungarian language, sourced from Szabad Európa (szabadeuropa.hu), the Hungarian service of Radio Free Europe/Radio Liberty (RFE/RL). Because the RFE/RL Hungarian service officially ceased operations on November 21, 2025, this dataset represents a complete, finalized archive of its modern iteration. - [RFE/RL Ukrainian (Crimea) News Text Corpus](https://mozilladatacollective.com/datasets/cmnxk7u8900a2ny076pyqzama): This dataset serves as a comprehensive longitudinal news corpus for the Ukrainian language focused on the Crimean region, sourced from Krym.Realii (ua.krymr.com), the Crimean project of the Ukrainian Service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from December 2012 to March 2026, the corpus contains 162,119 unique articles, totaling over 54 million tokens. - [RFE/RL Pashto (Pakistani) News Text Corpus](https://mozilladatacollective.com/datasets/cmnxk6nou00a5nu07kid0mewy): This dataset serves as a comprehensive longitudinal news corpus for the Pashto language targeting the Pakistan-Afghanistan border region, sourced from Radio Mashaal (mashaalradio.com), a broadcaster of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from October 2010 to March 2026, the corpus contains 52,729 unique articles, totaling over 14.5 million tokens. - [RFE/RL Ukrainian News Text Corpus](https://mozilladatacollective.com/datasets/cmnxk58mo009sny0794yybczp): This dataset serves as a comprehensive longitudinal news corpus for the Ukrainian language, sourced from Radio Svoboda (radiosvoboda.org), the main Ukrainian service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning over three decades from January 1995 to March 2026, the corpus contains 504,805 unique articles, totaling over 171 million tokens across Ukrainian and Russian texts. - [Synthetic Text Corpus for African Language ASR](https://mozilladatacollective.com/datasets/cmnxi09v2008ony078mfiehri): This dataset contains 13,488 synthetic sentences across 10 African languages (Bambara, Chichewa, Hausa, Kanuri, Luo, Nande, Somali, Twi, Wolof, Yoruba) generated using large language models (GPT-4o, GPT-4.5, Claude 3.5 Sonnet, Claude 3.7 Sonnet). Each sentence has been evaluated by human linguists on readability and naturalness (1-7 scale), translation adequacy and accuracy (1-7 scale), grammatical correctness, word validity, and presence of notable errors. - [Kaler Kantho Bengali Newspaper Corpus](https://mozilladatacollective.com/datasets/cmnxhzcgy0094nu07arwxhq53): The Kaler Kantho Bengali Newspaper Corpus is a large-scale text dataset containing over 10 million tokens collected from the digital archives of Kaler Kantho, a major daily newspaper in Bangladesh. It represents modern Bengali journalistic writing across domains such as national politics, international affairs, social issues, and cultural content. - [Marma Text Corpus](https://mozilladatacollective.com/datasets/cmnxhxm1i0090nu071ba8o4sg): This dataset contains 5,675 sentences in the Marma language (ISO 639-3: rmz), a Tibeto-Burman language spoken primarily by the Marma people in Bangladesh and Myanmar. Each entry includes the original sentence and its normalized form, along with the source of the text. - [Prothom Alo Bengali Newspaper Corpus](https://mozilladatacollective.com/datasets/cmnxhw57z008kny07uy0kj8a5): The Prothom Alo Bengali Newspaper Corpus is a large-scale text dataset containing over 10 million tokens from the archives of Prothom Alo, one of Bangladesh’s most widely read newspapers. It represents modern Bengali journalistic writing across domains such as national and international news, social issues, culture, and literary content. - [RFE/RL Uzbek News Text Corpus](https://mozilladatacollective.com/datasets/cmnx7y68r033znn07g6g8bosq): This dataset serves as a comprehensive longitudinal news corpus for the Uzbek language, sourced from Radio Ozodlik (ozodlik.org), the Uzbek service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from January 2002 to March 2026, the corpus contains 166,316 unique articles, totaling over 31 million tokens. - [RFE/RL Romanian (Romania) News Text Corpus](https://mozilladatacollective.com/datasets/cmnx7twx3033xml07rctvp2gx): This dataset serves as a comprehensive longitudinal news corpus for the Romanian language focused on Romania, sourced from Europa Liberă România (romania.europalibera.org), the Romanian service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from January 2013 to March 2026, the corpus contains 34,765 unique articles, totaling over 17.6 million tokens. - [Hindi 10 Million Text Corpus](https://mozilladatacollective.com/datasets/cmnx6rab3032gml07xu2xkwym): The Hindi Ten Million Corpus is a curated Hindi text collection of around 10 million tokens drawn from multiple authors. It includes both literary texts and personal articles, making it useful for research in Hindi NLP, corpus linguistics, stylistic analysis, and digital humanities. - [The Daily Jugantor Bengali Language Corpus](https://mozilladatacollective.com/datasets/cmnx6375n031gnn07sm6juaz4): The Daily Jugantor Bengali Language Corpus is a monolingual Bengali text collection containing approximately 10.6 million words sourced from Daily Jugantor, a widely read Bengali news publication. The corpus reflects contemporary written Bengali as used in journalistic reporting and public communication. - [Bangor Miami Spanish-English Corpus](https://mozilladatacollective.com/datasets/cmmfulo4r018bnz07py4q9t09): The Bangor Miami Corpus of Spanish-English bilingual speech, containing around 240,000 words over 35 hours of recorded audio conversations. The dataset includes the audios, transcriptions and glosses in CHAT format, and word-level analyses of the transcriptions in .tsv files. - [CV Korean Test 25.0 - Noise-Augmented (SCAI)](https://mozilladatacollective.com/datasets/cmnprpfot01khnz07xdjosybq): This dataset is a noise-augmented version of the test split (test.csv) from the Mozilla Common Voice Scripted Speech 25.0 – Korean dataset, designed to support research in automatic speech recognition (ASR), particularly in noisy environments. The original clean speech dataset is already hosted on Mozilla Data Collective. - [IBT Torwali Literature Corpus](https://mozilladatacollective.com/datasets/cmnoqtoq900ugmf077kiuny8y): The IBT Torwali Literature Corpus by is a text dataset of about 233,000 tokens from multiple literary and cultural domains, including poetry, folktales, biographies, educational materials, religious translations, and other literary texts. It contains 21 UTF-8 encoded text files and is useful for language preservation and NLP research. - [Tamil Time Aligned Speech Dataset](https://mozilladatacollective.com/datasets/cmnmyptri02glo107p5cx5por): The Tamil Time-Aligned Speech Dataset is a curated 5-hour speech corpus consisting of Tamil audio recordings paired with precise time-aligned transcriptions. The dataset is designed to support a wide range of speech and language technology tasks, including automatic speech recognition, forced alignment, speech segmentation, subtitle generation, and timestamp-aware linguistic analysis. - [ViQua² — Visual Question-answering about Quantities](https://mozilladatacollective.com/datasets/cmnmy7tmf02k6ml07m7vvcxzr): This dataset provides 116 images of varying quantities of a range of different objects, each with an accompanying question about the quantity in the image and a numeric answer. For example, an image of a table with multiple fruits and vegetables, with the question "How many carrots?". - [Saraiki 10 Hours TTS Dataset](https://mozilladatacollective.com/datasets/cmnggqr8z0082mh07vbsbm6t5): he Saraiki TTS Dataset – 10 Hours is a curated speech dataset developed to support research and development in text-to-speech (TTS) and related speech technologies for the Saraiki language. The dataset contains approximately 10 hours of audio recordings with corresponding text transcripts, prepared for high-quality speech synthesis and language technology applications. - [Territórios Digitais](https://mozilladatacollective.com/datasets/cmnhi28fc00xpmh07fwall3wn): This dataset was developed in accordance with ethical research practices commonly applied in social science and participatory research. All participants were informed about the purpose of the study, the nature of their participation, how the data would be used, and any potential risks involved. - [Chuvash TTS](https://mozilladatacollective.com/datasets/cmnhhi0by00zknn07edrnd82e): Chuvash TTS is a speech dataset sourced from the Turkic_TTS GitHub repository. It comprises 4 hours and 8 minutes of news article text from chuvash.org and 1 hour and 1 minute of recorded digits, all read by a single female speaker at a rapid tempo. - [RFE/RL Persian News Text Corpus](https://mozilladatacollective.com/datasets/cmnhgmkgg00wimh07wc0v3s49): This dataset serves as a comprehensive longitudinal news corpus for the Persian language, sourced from Radio Farda (radiofarda.com), the Persian service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from January 2001 to March 2026, the corpus contains 350,485 unique articles, totaling over 51 million tokens. - [TidyVoiceX2_ASV](https://mozilladatacollective.com/datasets/cmkv32i5e02tumg07j79d3c35): This dataset is designed for speaker verification using the Mozilla Common Voice corpus, covering approximately 40 additional languages beyond those included in TidyVoiceX_ASV. It comprises recordings from different speakers, each of whom appears in multiple languages. - [Kannada Time Aligned Speech Corpus](https://mozilladatacollective.com/datasets/cmnggocgg007ymh079k30st39): The Kannada Time-Aligned Speech Corpus is a 5-hour speech dataset containing Kannada audio with corresponding time-aligned transcriptions. It is designed to support speech technology and research tasks such as automatic speech recognition, forced alignment, speech segmentation, pronunciation modeling, and spoken language analysis. - [Common Voice Spontaneous Speech 3.0 - Serian Bidayuh](https://mozilladatacollective.com/datasets/cmndf7vbn002gma07np6nc35o): A collection of spontaneous responses to questions in Serian Bidayuh (sdo). - [Common Voice Scripted Speech 25.0 - Pashto](https://mozilladatacollective.com/datasets/cmndf6mgs001lnz07bf9t3skp): A collection of read speech recordings in Pashto (پښتو). - [Common Voice Scripted Speech 25.0 - English](https://mozilladatacollective.com/datasets/cmndapwry02jnmh07dyo46mot): A collection of read speech recordings in English (English). - [Common Voice Scripted Speech 25.0 - Catalan](https://mozilladatacollective.com/datasets/cmnd4la5a02fwmh074t1fx5y9): A collection of read speech recordings in Catalan (català). - [English Hausa Parallel Corpus](https://mozilladatacollective.com/datasets/cmn3ht40i00eami07lgydmrgg): This English–Hausa Parallel Corpus is a curated bilingual dataset of 5,000 aligned sentence pairs, translated from English into Hausa and organized into a clean sentence-level format to ensure reliable alignment. The dataset is designed to support machine translation training and evaluation, bilingual lexicon development, and broader linguistic and natural language processing (NLP) research for Hausa, including data-driven language technology development. - [Common Voice Scripted Speech 25.0 - Kinyarwanda](https://mozilladatacollective.com/datasets/cmn60xnfi00wnnv07028xltoz): A collection of read speech recordings in Kinyarwanda (Ikinyarwanda). - [Common Voice Scripted Speech 25.0 - French](https://mozilladatacollective.com/datasets/cmn5zugst00w3nv07upovf2bg): A collection of read speech recordings in French (Français). - [Common Voice Scripted Speech 25.0 - Spanish](https://mozilladatacollective.com/datasets/cmn4z1n52000knv07h01532dd): A collection of read speech recordings in Spanish (Español). - [Araina Text Corpus (Occitan Aranese)](https://mozilladatacollective.com/datasets/cmn4xrzyx00ednz074sa8dwkp): This text corpus includes sentences from three sources. Public domain literary texts translated by Antòni Nogués. - [Common Voice Scripted Speech 25.0 - Belarusian](https://mozilladatacollective.com/datasets/cmn4xg3a900d3nu075gnh4jpt): A collection of read speech recordings in Belarusian (Беларуская). - [Corpus de llenguatge ofensiu en català](https://mozilladatacollective.com/datasets/cmn4s1j5d0091nu07e1hgzwgn): This dataset consists of sentences tagged as offensive-language in the version 25.0 release of Mozilla Common Voice in Catalan. The sentences are provided with the aim that they promote the development of offensive language detection in Catalan. - [Common Voice Scripted Speech 25.0 - German](https://mozilladatacollective.com/datasets/cmn4rsdh6009unz07jdn2ol9p): A collection of read speech recordings in German (Deutsch). - [Common Voice Scripted Speech 25.0 - Esperanto](https://mozilladatacollective.com/datasets/cmn4o8691005pnu07fxmq06px): A collection of read speech recordings in Esperanto (Esperanto). - [Oro_Word](https://mozilladatacollective.com/datasets/cmn4l2u61001nnu07wtrkoqe4): This dataset contains word-level recordings in Afaan Oromoo collected from native speakers to support the development of open-source speech technologies. The dataset is designed for training and evaluating automatic speech recognition (ASR) and text-to-speech (TTS) systems. - [INEL Kalmyk Speech Corpus](https://mozilladatacollective.com/datasets/cmn4kxlaj001xnz07a5yugnew): This dataset is a machine-learning-ready subset of the INEL Kalmyk Corpus (Version 1.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. - [INEL Nganasan Speech Corpus](https://mozilladatacollective.com/datasets/cmn4kxhp8001tnz07jyfvf2qt): This dataset is a machine-learning-ready subset of the INEL Nganasan Corpus (Version 1.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. - [INEL Evenki Speech Corpus](https://mozilladatacollective.com/datasets/cmn4kxexu001pnz07kn6wr985): This dataset is a machine-learning-ready subset of the INEL Evenki Corpus (Version 2.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. - [INEL Dolgan Speech Corpus](https://mozilladatacollective.com/datasets/cmn4kqzzt0013nu07caxllg3t): This dataset is a machine-learning-ready subset of the INEL Dolgan Corpus (Version 2.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. - [INEL Kamas Speech Corpus](https://mozilladatacollective.com/datasets/cmn4kq2fl001dnz078d6n7r9a): This dataset is a machine-learning-ready subset of the INEL Kamas Corpus (Version 2.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. - [INEL Selkup Speech Corpus](https://mozilladatacollective.com/datasets/cmn4kpt540019nz07sapke78o): This dataset is a machine-learning-ready subset of the INEL Selkup Corpus (Version 2.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. - [INEL Enets Speech Corpus](https://mozilladatacollective.com/datasets/cmn4knp2w0011nz07e6y641qy): This dataset is a machine-learning-ready subset of the INEL Enets Corpus (Version 1.1), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the highly detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. - [INEL Nenets Speech Corpus](https://mozilladatacollective.com/datasets/cmn4kmw70000nnu07vwcwfdi5): This dataset is a machine-learning-ready subset of the INEL Nenets Corpus (Version 1.0), processed specifically for Automatic Speech Recognition (ASR) / Speech-to-Text (STT) training. It translates the detailed EXMARaLDA XML annotations into the standard tabular layout utilized by Mozilla Common Voice. - [Heroes English-Spanish Dubbed Movie Speech Corpus](https://mozilladatacollective.com/datasets/cmn3fheeo00cjmb078qc8xr62): Heroes corpus contains mapped bilingual (English and Spanish) speech segments from the TV series Heroes. It contains 7000 single speaker speech segments extracted from the original and Spanish dubbed version of 21 episodes. - [Common Voice Scripted Speech 25.0 - Bengali](https://mozilladatacollective.com/datasets/cmn3ipo8b00ejmi079e8upl2k): A collection of read speech recordings in Bengali (বাংলা). - [Common Voice Scripted Speech 25.0 - Chinese (China)](https://mozilladatacollective.com/datasets/cmn3iaztg00e4mb070uvufz7q): A collection of read speech recordings in Chinese (China) (汉语(中国大陆)). - [Persian Literature Corpus by Najwai Sukhan](https://mozilladatacollective.com/datasets/cmn3hs01700e0mb07n3y7brmz): The Persian Literature Corpus by Najwai Sukhan is a curated collection of Persian (Farsi) literary and educational texts created for research, computational use, and cultural preservation. It contains about 1.26 million tokens across 20 complete works spanning classical literature, poetry, modern prose, educational writing, philosophy, translations, and culturally rooted creative texts. - [Common Voice Scripted Speech 25.0 - Swahili](https://mozilladatacollective.com/datasets/cmn3ailbd008nmb07mjyu3xro): A collection of read speech recordings in Swahili (Kiswahili). - [Common Voice Scripted Speech 25.0 - Kabyle](https://mozilladatacollective.com/datasets/cmn38spwm005vmi07bejigyo6): A collection of read speech recordings in Kabyle (Taqbaylit). - [Common Voice Scripted Speech 25.0 - Basque](https://mozilladatacollective.com/datasets/cmn2hwe0d01n8mm07wug9r5he): A collection of read speech recordings in Basque (Euskara). - [Common Voice Scripted Speech 25.0 - Japanese](https://mozilladatacollective.com/datasets/cmn2hm68r01n4mm071qux43yu): A collection of read speech recordings in Japanese (日本語). - [Common Voice Scripted Speech 25.0 - Luganda](https://mozilladatacollective.com/datasets/cmn2hjqe001n0mm07lbfcq6bp): A collection of read speech recordings in Luganda (Luganda). - [Common Voice Scripted Speech 25.0 - Czech](https://mozilladatacollective.com/datasets/cmn2h5zd801h3o1075tita1ap): A collection of read speech recordings in Czech (Čeština). - [Common Voice Scripted Speech 25.0 - Urdu](https://mozilladatacollective.com/datasets/cmn2h58bw01mwmm07t3ypteqz): A collection of read speech recordings in Urdu (اردو). - [Common Voice Scripted Speech 25.0 - Georgian](https://mozilladatacollective.com/datasets/cmn2h4m7901gzo1072qn7zoes): A collection of read speech recordings in Georgian (ქართული). - [Common Voice Scripted Speech 25.0 - Thai](https://mozilladatacollective.com/datasets/cmn2h1svx01gvo1074l8g2a27): A collection of read speech recordings in Thai (ไทย). - [Common Voice Scripted Speech 25.0 - Russian](https://mozilladatacollective.com/datasets/cmn2h1dg201gro107lpynbbd6): A collection of read speech recordings in Russian (Русский). - [Common Voice Scripted Speech 25.0 - Italian](https://mozilladatacollective.com/datasets/cmn2h0yei01msmm07u8z5vu87): A collection of read speech recordings in Italian (Italiano). - [Common Voice Scripted Speech 25.0 - Galician](https://mozilladatacollective.com/datasets/cmn2h0nw001momm07xxarkyd4): A collection of read speech recordings in Galician (Galego). - [Common Voice Scripted Speech 25.0 - Latvian](https://mozilladatacollective.com/datasets/cmn2gmdle01gmo1071lmpg5sj): A collection of read speech recordings in Latvian (Latviešu). - [Common Voice Scripted Speech 25.0 - Persian](https://mozilladatacollective.com/datasets/cmn2gho8i01gio107ckfuqzxo): A collection of read speech recordings in Persian (فارسی). - [Common Voice Scripted Speech 25.0 - Tamil](https://mozilladatacollective.com/datasets/cmn2gfvyp01geo107izoftfki): A collection of read speech recordings in Tamil (தமிழ்). - [Common Voice Scripted Speech 25.0 - Uyghur](https://mozilladatacollective.com/datasets/cmn2gfexo01mamm07lcv4ako1): A collection of read speech recordings in Uyghur (ئۇيغۇرچە). - [Common Voice Scripted Speech 25.0 - Kabardian](https://mozilladatacollective.com/datasets/cmn2gc4hz01gao107ak92w86f): A collection of read speech recordings in Kabardian (Адыгэбзэ (Къэбэрдей)). - [Common Voice Scripted Speech 25.0 - Frisian](https://mozilladatacollective.com/datasets/cmn2gaxv301g6o107fdzr4kct): A collection of read speech recordings in Frisian (Frysk). - [Common Voice Scripted Speech 25.0 - Welsh](https://mozilladatacollective.com/datasets/cmn2g9w1h01m6mm07lq2w14dd): A collection of read speech recordings in Welsh (Cymraeg). - [Common Voice Scripted Speech 25.0 - Central Kurdish](https://mozilladatacollective.com/datasets/cmn2g9npx01g2o107hentahj9): A collection of read speech recordings in Central Kurdish (کوردیی ناوەندی). - [Common Voice Scripted Speech 25.0 - Hungarian](https://mozilladatacollective.com/datasets/cmn2g9aoi01fyo107xhdrwb5d): A collection of read speech recordings in Hungarian (Magyar). - [Common Voice Scripted Speech 25.0 - Chinese (Hong Kong)](https://mozilladatacollective.com/datasets/cmn2g8zqd01m2mm07prcmehku): A collection of read speech recordings in Chinese (Hong Kong) (中文(香港)). - [Common Voice Scripted Speech 25.0 - Meadow Mari](https://mozilladatacollective.com/datasets/cmn2g8gyb01fuo107z5w28cqc): A collection of read speech recordings in Meadow Mari (Марий). - [Common Voice Scripted Speech 25.0 - Arabic](https://mozilladatacollective.com/datasets/cmn2g7uu701fqo1072r5na25l): A collection of read speech recordings in Arabic (العربية). - [Common Voice Scripted Speech 25.0 - Dutch](https://mozilladatacollective.com/datasets/cmn2g7nu901fmo107a1ydn0n5): A collection of read speech recordings in Dutch (Nederlands). - [Common Voice Scripted Speech 25.0 - Chinese (Taiwan)](https://mozilladatacollective.com/datasets/cmn2g7eaj01fio10769r1m96n): A collection of read speech recordings in Chinese (Taiwan) (華語(台灣)). - [Common Voice Scripted Speech 25.0 - Kidaw'ida](https://mozilladatacollective.com/datasets/cmn2e92b801limm07whtrn9te): A collection of read speech recordings in Kidaw'ida (dav). - [Common Voice Scripted Speech 25.0 - Macedonian](https://mozilladatacollective.com/datasets/cmn2e8yb101lemm07j3flmvjs): A collection of read speech recordings in Macedonian (Македонски). - [Common Voice Scripted Speech 25.0 - Kyrgyz](https://mozilladatacollective.com/datasets/cmn2e8urv01lamm07vgcmkqnt): A collection of read speech recordings in Kyrgyz (Кыргызча). - [Common Voice Scripted Speech 25.0 - Romanian](https://mozilladatacollective.com/datasets/cmn2e8rmi01l6mm07vxurptse): A collection of read speech recordings in Romanian (Română). - [Common Voice Scripted Speech 25.0 - Slovak](https://mozilladatacollective.com/datasets/cmn2e8ojy01l2mm07giwrvaqf): A collection of read speech recordings in Slovak (Slovenčina). - [Common Voice Scripted Speech 25.0 - Armenian](https://mozilladatacollective.com/datasets/cmn2e8k9z01kymm07yqqy4bk1): A collection of read speech recordings in Armenian (Հայերեն). - [Common Voice Scripted Speech 25.0 - Swedish](https://mozilladatacollective.com/datasets/cmn2e8gr301evo1079gujuzqr): A collection of read speech recordings in Swedish (Svenska). - [Common Voice Scripted Speech 25.0 - Dhivehi](https://mozilladatacollective.com/datasets/cmn2e8dkd01ero107fgwmo0qz): A collection of read speech recordings in Dhivehi (ދިވެހި). - [Common Voice Scripted Speech 25.0 - Indonesian](https://mozilladatacollective.com/datasets/cmn2e8ats01eno107glwgoasv): A collection of read speech recordings in Indonesian (Bahasa Indonesia). - [Common Voice Scripted Speech 25.0 - Estonian](https://mozilladatacollective.com/datasets/cmn2e880l01kumm07i9upoz99): A collection of read speech recordings in Estonian (eesti). - [Common Voice Scripted Speech 25.0 - Kalenjin](https://mozilladatacollective.com/datasets/cmn2e84ge01kqmm07esyp21xq): A collection of read speech recordings in Kalenjin (kln). - [Common Voice Scripted Speech 25.0 - Adyghe](https://mozilladatacollective.com/datasets/cmn2e80ea01kmmm07lzsoe5z9): A collection of read speech recordings in Adyghe (Адыгабзэ). - [Common Voice Scripted Speech 25.0 - Kurmanji Kurdish](https://mozilladatacollective.com/datasets/cmn2e7wkd01kimm070ofrk3ie): A collection of read speech recordings in Kurmanji Kurdish (Kurdî (Kurmancî)). - [Common Voice Scripted Speech 25.0 - Dholuo](https://mozilladatacollective.com/datasets/cmn2e7tbt01kemm07drju0bcf): A collection of read speech recordings in Dholuo (luo). - [Common Voice Scripted Speech 25.0 - Ukrainian](https://mozilladatacollective.com/datasets/cmn2e7qgt01kamm07oftersjt): A collection of read speech recordings in Ukrainian (Українська). - [Common Voice Scripted Speech 25.0 - Mongolian](https://mozilladatacollective.com/datasets/cmn2e7nxs01k6mm07you99zve): A collection of read speech recordings in Mongolian (Монгол хэл). - [Common Voice Scripted Speech 25.0 - Turkish](https://mozilladatacollective.com/datasets/cmn2e7kbl01k2mm07gm5n1bc9): A collection of read speech recordings in Turkish (Türkçe). - [Common Voice Scripted Speech 25.0 - Maltese](https://mozilladatacollective.com/datasets/cmn2cyjd001jmmm07meob0o2j): A collection of read speech recordings in Maltese (Malti). - [Common Voice Scripted Speech 25.0 - Toki Pona](https://mozilladatacollective.com/datasets/cmn2cyfr901jimm070o431als): A collection of read speech recordings in Toki Pona (toki pona). - [Common Voice Scripted Speech 25.0 - Taiwanese (Minnan)](https://mozilladatacollective.com/datasets/cmn2cyd8901jemm0738nubysq): A collection of read speech recordings in Taiwanese (Minnan) (台語). - [Common Voice Scripted Speech 25.0 - Finnish](https://mozilladatacollective.com/datasets/cmn2cyal501jamm07q2dnsy5x): A collection of read speech recordings in Finnish (suomi). - [Common Voice Scripted Speech 25.0 - Slovenian](https://mozilladatacollective.com/datasets/cmn2cy7z701j6mm07axskhd0a): A collection of read speech recordings in Slovenian (slovenščina). - [Common Voice Scripted Speech 25.0 - Sakha](https://mozilladatacollective.com/datasets/cmn2cy58m01j2mm07v4dx049d): A collection of read speech recordings in Sakha (Саха тыла). - [Common Voice Scripted Speech 25.0 - Dagbani](https://mozilladatacollective.com/datasets/cmn2cy2su01iymm07xfr6ul2b): A collection of read speech recordings in Dagbani (Dagbanli). - [Common Voice Scripted Speech 25.0 - Hindi](https://mozilladatacollective.com/datasets/cmn2cxzy701iumm077t5ayw0e): A collection of read speech recordings in Hindi (हिंदी). - [Common Voice Scripted Speech 25.0 - Laz](https://mozilladatacollective.com/datasets/cmn2cxxeu01iqmm07qww9uflb): A collection of read speech recordings in Laz (Lazuri). - [Common Voice Scripted Speech 25.0 - Marathi](https://mozilladatacollective.com/datasets/cmn2cxubn01ebo107piypnoix): A collection of read speech recordings in Marathi (मराठी). - [Common Voice Scripted Speech 25.0 - Palula](https://mozilladatacollective.com/datasets/cmn2cxqwz01e7o107b9tl06ao): A collection of read speech recordings in Palula (پالولا). - [Common Voice Scripted Speech 25.0 - Latgalian](https://mozilladatacollective.com/datasets/cmn2cxmpv01e3o107b7ryplbt): A collection of read speech recordings in Latgalian (Latgalīšu). - [Common Voice Scripted Speech 25.0 - Chuvash](https://mozilladatacollective.com/datasets/cmn2cxjsd01dzo107sraxorft): A collection of read speech recordings in Chuvash (Чӑвашла). - [Common Voice Scripted Speech 25.0 - Puno Quechua](https://mozilladatacollective.com/datasets/cmn2cxfs801dvo107r5rkzmvx): A collection of read speech recordings in Puno Quechua (Punu qhichwa). - [Common Voice Scripted Speech 25.0 - Lithuanian](https://mozilladatacollective.com/datasets/cmn2cxca301dro107ucz5j8ey): A collection of read speech recordings in Lithuanian (Lietuvių). - [Common Voice Scripted Speech 25.0 - Greek](https://mozilladatacollective.com/datasets/cmn2cx91x01dno10754vxfu3b): A collection of read speech recordings in Greek (Ελληνικά). - [Common Voice Scripted Speech 25.0 - Breton](https://mozilladatacollective.com/datasets/cmn2cx1ko01djo1076muxlj24): A collection of read speech recordings in Breton (Brezhoneg). - [Common Voice Scripted Speech 25.0 - Odia](https://mozilladatacollective.com/datasets/cmn2cww3s01dfo107etbjbif1): A collection of read speech recordings in Odia (ଓଡ଼ିଆ). - [Common Voice Scripted Speech 25.0 - Tatar](https://mozilladatacollective.com/datasets/cmn2cwsv501dbo1071091borr): A collection of read speech recordings in Tatar (Татар). - [Common Voice Scripted Speech 25.0 - Sindhi](https://mozilladatacollective.com/datasets/cmn2cwq6k01d7o107ykckzpyr): A collection of read speech recordings in Sindhi (sd). - [Common Voice Scripted Speech 25.0 - Baoule](https://mozilladatacollective.com/datasets/cmn2cqa8r01immm074kdpi5v8): A collection of read speech recordings in Baoule (bci). - [Common Voice Scripted Speech 25.0 - Romansh Sursilvan](https://mozilladatacollective.com/datasets/cmn2cq76201iimm07avtwokjf): A collection of read speech recordings in Romansh Sursilvan (romontsch sursilvan). - [Common Voice Scripted Speech 25.0 - Wakhi](https://mozilladatacollective.com/datasets/cmn2cq4j601iemm0765824vbi): A collection of read speech recordings in Wakhi (Wakhi (Wuk̃hikwor)). - [Common Voice Scripted Speech 25.0 - Gawri](https://mozilladatacollective.com/datasets/cmn2cq22i01iamm07xoo8zcnw): A collection of read speech recordings in Gawri (گاؤری). - [Common Voice Scripted Speech 25.0 - Southern Pastaza Quechua](https://mozilladatacollective.com/datasets/cmn2cpzd901i6mm075p7u9e1b): A collection of read speech recordings in Southern Pastaza Quechua (qup). - [Common Voice Scripted Speech 25.0 - Batanga](https://mozilladatacollective.com/datasets/cmn2cpwp901i2mm07hpqh5mre): A collection of read speech recordings in Batanga (bnm). - [Common Voice Scripted Speech 25.0 - Danish](https://mozilladatacollective.com/datasets/cmn2cptsh01hymm07mulngxv0): A collection of read speech recordings in Danish (Dansk). - [Common Voice Scripted Speech 25.0 - Kohistani Shina](https://mozilladatacollective.com/datasets/cmn2cpm5101humm07n6c947br): A collection of read speech recordings in Kohistani Shina (کوہستانی شینا). - [Common Voice Scripted Speech 25.0 - Balti](https://mozilladatacollective.com/datasets/cmn2cpjc301hqmm072p9pm8c1): A collection of read speech recordings in Balti (بلتی). - [Common Voice Scripted Speech 25.0 - Svan](https://mozilladatacollective.com/datasets/cmn2cpgcr01hmmm0762lczqjs): A collection of read speech recordings in Svan (ლუშნუ). - [Common Voice Scripted Speech 25.0 - Aragonese](https://mozilladatacollective.com/datasets/cmn2cpd2m01himm07yj9w1lxn): A collection of read speech recordings in Aragonese (Aragonés). - [Common Voice Scripted Speech 25.0 - Irish](https://mozilladatacollective.com/datasets/cmn2cp9uy01hemm07jfogi1zf): A collection of read speech recordings in Irish (Gaeilge). - [Common Voice Scripted Speech 25.0 - Ibibio](https://mozilladatacollective.com/datasets/cmn2cp74s01hamm07s74uypfk): A collection of read speech recordings in Ibibio (ibb). - [Common Voice Scripted Speech 25.0 - Igbo](https://mozilladatacollective.com/datasets/cmn2cp3yv01h6mm07x6tl0t1i): A collection of read speech recordings in Igbo (Ásụ̀sụ́ Ìgbò). - [Common Voice Scripted Speech 25.0 - Torwali](https://mozilladatacollective.com/datasets/cmn2cp1b101h2mm07vy11w8td): A collection of read speech recordings in Torwali (توروالی). - [Common Voice Scripted Speech 25.0 - Khowar](https://mozilladatacollective.com/datasets/cmn2coyno01gymm072evfrx7b): A collection of read speech recordings in Khowar (کھوار). - [Common Voice Scripted Speech 25.0 - Dargwa](https://mozilladatacollective.com/datasets/cmn2cosov01gumm070n74b5i6): A collection of read speech recordings in Dargwa (Дарган). - [Common Voice Scripted Speech 25.0 - Bulgarian](https://mozilladatacollective.com/datasets/cmn2coplj01gqmm07dbibr68y): A collection of read speech recordings in Bulgarian (Български). - [Common Voice Scripted Speech 25.0 - Ormuri](https://mozilladatacollective.com/datasets/cmn2comsw01gmmm07m07dewn1): A collection of read speech recordings in Ormuri (اُرمړي). - [Common Voice Scripted Speech 25.0 - Vietnamese](https://mozilladatacollective.com/datasets/cmn2cojzk01gimm07zqlmypmn): A collection of read speech recordings in Vietnamese (Việt). - [Common Voice Scripted Speech 25.0 - Indus Kohistani](https://mozilladatacollective.com/datasets/cmn2cogvj01gemm074y9z3z5c): A collection of read speech recordings in Indus Kohistani (اباسیْن کوستَیں). - [Common Voice Scripted Speech 25.0 - Adamawa Fulfulde](https://mozilladatacollective.com/datasets/cmn2cijyg01gamm07434y1o8l): A collection of read speech recordings in Adamawa Fulfulde (fub). - [Common Voice Scripted Speech 25.0 - Kichwa](https://mozilladatacollective.com/datasets/cmn2cigqs01d3o107mzljz8gp): A collection of read speech recordings in Kichwa (qvi). - [Common Voice Scripted Speech 25.0 - Adja](https://mozilladatacollective.com/datasets/cmn2cidwc01g6mm07bv38rtqx): A collection of read speech recordings in Adja (ajg). - [Common Voice Scripted Speech 25.0 - Mingrelian](https://mozilladatacollective.com/datasets/cmn2ci35o01czo107gtywv73c): A collection of read speech recordings in Mingrelian (მარგალური). - [Common Voice Scripted Speech 25.0 - Ghomala](https://mozilladatacollective.com/datasets/cmn2chz1n01cvo107cain4z31): A collection of read speech recordings in Ghomala (bbj). - [Common Voice Scripted Speech 25.0 - Eastern Balochi](https://mozilladatacollective.com/datasets/cmn2chw3k01cro1072gkohebe): A collection of read speech recordings in Eastern Balochi (بلوچی). - [Common Voice Scripted Speech 25.0 - Cornish](https://mozilladatacollective.com/datasets/cmn2chs8a01cno107aydgie03): A collection of read speech recordings in Cornish (Kernowek). - [Common Voice Scripted Speech 25.0 - Quechua Pasco Santa Ana de Tusi](https://mozilladatacollective.com/datasets/cmn2cho3b01cjo1079ybiidtg): A collection of read speech recordings in Quechua Pasco Santa Ana de Tusi (qxt). - [Common Voice Scripted Speech 25.0 - Quechua Jauja Wanka](https://mozilladatacollective.com/datasets/cmn2chlg201cfo10794fnwn6p): A collection of read speech recordings in Quechua Jauja Wanka (qxw). - [Common Voice Scripted Speech 25.0 - Medumba](https://mozilladatacollective.com/datasets/cmn2chivm01cbo107vqvbgn2i): A collection of read speech recordings in Medumba (byv). - [Common Voice Scripted Speech 25.0 - Baatonum](https://mozilladatacollective.com/datasets/cmn2chcee01c7o10782n6ozvp): A collection of read speech recordings in Baatonum (Baatɔnum). - [Common Voice Scripted Speech 25.0 - Pahari-Pothwari](https://mozilladatacollective.com/datasets/cmn2ch9gl01c3o107pqn1noey): A collection of read speech recordings in Pahari-Pothwari (phr). - [Common Voice Scripted Speech 25.0 - Duala](https://mozilladatacollective.com/datasets/cmn2ch6p301bzo107o6gib3e7): A collection of read speech recordings in Duala (dua). - [Common Voice Scripted Speech 25.0 - Sakizaya](https://mozilladatacollective.com/datasets/cmn2ch3pj01bvo107xchtu05n): A collection of read speech recordings in Sakizaya (Sakizaya). - [Common Voice Scripted Speech 25.0 - Ewondo](https://mozilladatacollective.com/datasets/cmn2ch06r01bro107450c5m72): A collection of read speech recordings in Ewondo (ewo). - [Common Voice Scripted Speech 25.0 - Paiwan](https://mozilladatacollective.com/datasets/cmn2cgx9p01bno1073smxze18): A collection of read speech recordings in Paiwan (pwn). - [Common Voice Scripted Speech 25.0 - Nigerian Pidgin English](https://mozilladatacollective.com/datasets/cmn2cgr3101g2mm07mt1zagmz): A collection of read speech recordings in Nigerian Pidgin English (pcm). - [Common Voice Scripted Speech 25.0 - Wadiyara Koli](https://mozilladatacollective.com/datasets/cmn2ca7oj01fwmm07piqekyg7): A collection of read speech recordings in Wadiyara Koli (kxp). - [Common Voice Scripted Speech 25.0 - Loja Highland Kichwa](https://mozilladatacollective.com/datasets/cmn2ca4jd01fsmm07v6aorg7u): A collection of read speech recordings in Loja Highland Kichwa (qvj). - [Common Voice Scripted Speech 25.0 - Tupuri](https://mozilladatacollective.com/datasets/cmn2ca0qt01fomm07ivn5e89r): A collection of read speech recordings in Tupuri (t'pur). - [Common Voice Scripted Speech 25.0 - Oadki](https://mozilladatacollective.com/datasets/cmn2c9xqu01fkmm07rzk6m11b): A collection of read speech recordings in Oadki (odk). - [Common Voice Scripted Speech 25.0 - Bankon](https://mozilladatacollective.com/datasets/cmn2c9uwa01fgmm073lt5l0i0): A collection of read speech recordings in Bankon (Ɓàŋkón). - [Common Voice Scripted Speech 25.0 - Loarki](https://mozilladatacollective.com/datasets/cmn2c9ron01fcmm07ej0e597h): A collection of read speech recordings in Loarki (lrk). - [Common Voice Scripted Speech 25.0 - Yadgha](https://mozilladatacollective.com/datasets/cmn2c9ouw01f8mm0764l0eb27): A collection of read speech recordings in Yadgha (ydg). - [Common Voice Scripted Speech 25.0 - Nüpode Huitoto](https://mozilladatacollective.com/datasets/cmn2c9jar01bio107d52mdzdl): A collection of read speech recordings in Nüpode Huitoto (hux). - [Common Voice Scripted Speech 25.0 - Central Puebla Nahuatl](https://mozilladatacollective.com/datasets/cmn2c9fhc01beo1074z081p8t): A collection of read speech recordings in Central Puebla Nahuatl (Nahuat). - [Common Voice Scripted Speech 25.0 - Yaqui](https://mozilladatacollective.com/datasets/cmn2c9cm601bao1071vlhmoqj): A collection of read speech recordings in Yaqui (Jiak noki). - [Common Voice Scripted Speech 25.0 - Bamun](https://mozilladatacollective.com/datasets/cmn2c99q101b6o1079w6l6r5l): A collection of read speech recordings in Bamun (Shüpamom). - [Common Voice Scripted Speech 25.0 - Bunun](https://mozilladatacollective.com/datasets/cmn2c96vl01b2o107ys2d3rim): A collection of read speech recordings in Bunun (Bunun). - [Common Voice Scripted Speech 25.0 - Huarijio](https://mozilladatacollective.com/datasets/cmn2c942j01ayo107i14q7o9p): A collection of read speech recordings in Huarijio (Makurawe). - [Common Voice Scripted Speech 25.0 - Kom](https://mozilladatacollective.com/datasets/cmn2c90xf01auo107fh76jofh): A collection of read speech recordings in Kom (bkm). - [Common Voice Scripted Speech 25.0 - Orizaba Nahuatl](https://mozilladatacollective.com/datasets/cmn2c8xa101aqo1074qep48bl): A collection of read speech recordings in Orizaba Nahuatl (Nauatl). - [Common Voice Scripted Speech 25.0 - Fe’efe’e](https://mozilladatacollective.com/datasets/cmn2bk50x01alo107rql5nty3): A collection of read speech recordings in Fe’efe’e (fmp). - [Common Voice Scripted Speech 25.0 - Basaa](https://mozilladatacollective.com/datasets/cmn2bk0jn01f2mm07dv6v1jbt): A collection of read speech recordings in Basaa (Basaa). - [Common Voice Scripted Speech 25.0 - Atayal](https://mozilladatacollective.com/datasets/cmn2bjxwf01eymm071zgnlf93): A collection of read speech recordings in Atayal (Atayal). - [Common Voice Scripted Speech 25.0 - Serbian](https://mozilladatacollective.com/datasets/cmn2bjv8101aho107xj4uoklm): A collection of read speech recordings in Serbian (Српски). - [Common Voice Scripted Speech 25.0 - Tush](https://mozilladatacollective.com/datasets/cmn2bjso001ado1076ox3kg0t): A collection of read speech recordings in Tush (ვაჲღეჼ). - [Common Voice Scripted Speech 25.0 - Quechua Yauyos](https://mozilladatacollective.com/datasets/cmn2bjq6a01a9o107ifab182z): A collection of read speech recordings in Quechua Yauyos (qux). - [Common Voice Scripted Speech 25.0 - Bakoko](https://mozilladatacollective.com/datasets/cmn2bjje101eumm07avcbr9fk): A collection of read speech recordings in Bakoko (bkh). - [Common Voice Scripted Speech 25.0 - Gawarbaiti](https://mozilladatacollective.com/datasets/cmn2bjg7201eqmm0716x5ugzj): A collection of read speech recordings in Gawarbaiti (گوَرباتی). - [Common Voice Scripted Speech 25.0 - Quechua Arequipa-La Unión](https://mozilladatacollective.com/datasets/cmn2bjcjx01a5o1078dzypsfb): A collection of read speech recordings in Quechua Arequipa-La Unión (qxu). - [Common Voice Scripted Speech 25.0 - Malayalam](https://mozilladatacollective.com/datasets/cmn2awrzn019zo1070mgq51jm): A collection of read speech recordings in Malayalam (മലയാളം). - [Common Voice Scripted Speech 25.0 - Dameli](https://mozilladatacollective.com/datasets/cmn2awoti01elmm07gfwlnqs8): A collection of read speech recordings in Dameli (dml). - [Common Voice Scripted Speech 25.0 - Bamvele](https://mozilladatacollective.com/datasets/cmn2awma9019vo1077myulutp): A collection of read speech recordings in Bamvele (beb). - [Common Voice Scripted Speech 25.0 - Bafut](https://mozilladatacollective.com/datasets/cmn2awjf501ehmm07zqozqv3q): A collection of read speech recordings in Bafut (bfd). - [Common Voice Scripted Speech 25.0 - Kachhi](https://mozilladatacollective.com/datasets/cmn2avy1001edmm07bpvopkl3): A collection of read speech recordings in Kachhi (gjk). - [Common Voice Scripted Speech 25.0 - Tigre](https://mozilladatacollective.com/datasets/cmn2avv9x01e9mm07cfm9tqve): A collection of read speech recordings in Tigre (ትግረ). - [Common Voice Scripted Speech 25.0 - Fang](https://mozilladatacollective.com/datasets/cmn2avq1d01e5mm07tfbbfmjq): A collection of read speech recordings in Fang (fan). - [Common Voice Scripted Speech 25.0 - Parkari Koli](https://mozilladatacollective.com/datasets/cmn2avl52019ro107fand91bl): A collection of read speech recordings in Parkari Koli (kvx). - [Common Voice Scripted Speech 25.0 - Brushaski](https://mozilladatacollective.com/datasets/cmn2avimf019no107bb37vfx8): A collection of read speech recordings in Brushaski (Mishaski). - [Common Voice Scripted Speech 25.0 - Ngomba](https://mozilladatacollective.com/datasets/cmn2avfo201e1mm07ihajuhcn): A collection of read speech recordings in Ngomba (jgo). - [Common Voice Scripted Speech 25.0 - Sindhi Bhil](https://mozilladatacollective.com/datasets/cmn2avcqr01dxmm07fgghrfq6): A collection of read speech recordings in Sindhi Bhil (sbn). - [Common Voice Scripted Speech 25.0 - Western Highland Purepecha](https://mozilladatacollective.com/datasets/cmn2ava20019jo107uax2avuc): A collection of read speech recordings in Western Highland Purepecha (pua). - [Common Voice Scripted Speech 25.0 - Mina](https://mozilladatacollective.com/datasets/cmn2av7jh01dtmm07xogduo3x): A collection of read speech recordings in Mina (gej). - [Common Voice Scripted Speech 25.0 - Tuki](https://mozilladatacollective.com/datasets/cmn2av5d9019fo10727qxuymi): A collection of read speech recordings in Tuki (Tukí). - [Common Voice Scripted Speech 25.0 - Manx](https://mozilladatacollective.com/datasets/cmn2av2qr01dpmm078edm64ux): A collection of read speech recordings in Manx (gv). - [Common Voice Scripted Speech 25.0 - Rukai](https://mozilladatacollective.com/datasets/cmn2av09s01dlmm07el2llge3): A collection of read speech recordings in Rukai (Drekay). - [Common Voice Scripted Speech 25.0 - Sansi](https://mozilladatacollective.com/datasets/cmn2auxhh01dhmm072b74oup1): A collection of read speech recordings in Sansi (ssi). - [Common Voice Scripted Speech 25.0 - Borgu Fulfulde](https://mozilladatacollective.com/datasets/cmn2auup5019bo107zvq9bitg): A collection of read speech recordings in Borgu Fulfulde (fue). - [Common Voice Scripted Speech 25.0 - Gujari](https://mozilladatacollective.com/datasets/cmn2aus9q01ddmm07rsn6o1jj): A collection of read speech recordings in Gujari (گوجری). - [Common Voice Scripted Speech 25.0 - Quechua Corongo Ancash](https://mozilladatacollective.com/datasets/cmn2aupcc01d9mm07qdphmygg): A collection of read speech recordings in Quechua Corongo Ancash (qwa). - [Common Voice Scripted Speech 25.0 - Teutila Cuicatec](https://mozilladatacollective.com/datasets/cmn2aumhf0197o107rr09hs4w): A collection of read speech recordings in Teutila Cuicatec (Dbaku). - [Common Voice Scripted Speech 25.0 - Bulu](https://mozilladatacollective.com/datasets/cmn2aok00018vo107c7he8w00): A collection of read speech recordings in Bulu (bum). - [Common Voice Scripted Speech 25.0 - Asheninka South Ucayali](https://mozilladatacollective.com/datasets/cmn2aoh0g01d3mm07de71ivsg): A collection of read speech recordings in Asheninka South Ucayali (cpy). - [Common Voice Scripted Speech 25.0 - Central Tarahumara](https://mozilladatacollective.com/datasets/cmn2aodxp01czmm078t6zlp2s): A collection of read speech recordings in Central Tarahumara (Rarámuri). - [Common Voice Scripted Speech 25.0 - Quechua Sihuas Ancash](https://mozilladatacollective.com/datasets/cmn2ao6ae01cvmm07kjt76a5e): A collection of read speech recordings in Quechua Sihuas Ancash (qws). - [Common Voice Scripted Speech 25.0 - Kateviri](https://mozilladatacollective.com/datasets/cmn2ao1wu01crmm07nlzyjq1w): A collection of read speech recordings in Kateviri (کتہ وری). - [Common Voice Scripted Speech 25.0 - Guidar](https://mozilladatacollective.com/datasets/cmn2anyzw01cnmm07lir99pol): A collection of read speech recordings in Guidar (gid). - [Common Voice Scripted Speech 25.0 - Goaria](https://mozilladatacollective.com/datasets/cmn2anvy601cjmm07jzgf5vdh): A collection of read speech recordings in Goaria (gig). - [Common Voice Scripted Speech 25.0 - Copainalá Zoque](https://mozilladatacollective.com/datasets/cmn2ansfl01cfmm07ivk4w1rf): A collection of read speech recordings in Copainalá Zoque (zoc). - [Common Voice Scripted Speech 25.0 - Marwari](https://mozilladatacollective.com/datasets/cmn2anp8p01cbmm07v1u46jig): A collection of read speech recordings in Marwari (mve). - [Common Voice Scripted Speech 25.0 - Jaqaru](https://mozilladatacollective.com/datasets/cmn2anmfw01c7mm0734ie8efh): A collection of read speech recordings in Jaqaru (jqr). - [Common Voice Scripted Speech 25.0 - Shina](https://mozilladatacollective.com/datasets/cmn2anjcv01c3mm07qmeksg2p): A collection of read speech recordings in Shina (شِینا). - [Common Voice Scripted Speech 25.0 - Seediq](https://mozilladatacollective.com/datasets/cmn2angd301bzmm0701c0wims): A collection of read speech recordings in Seediq (trv). - [Common Voice Scripted Speech 25.0 - Bateri](https://mozilladatacollective.com/datasets/cmn2an7at018ro1076m25xok8): A collection of read speech recordings in Bateri (btv). - [Common Voice Scripted Speech 25.0 - Losso](https://mozilladatacollective.com/datasets/cmn2amch2018no1077o0iwot5): A collection of read speech recordings in Losso (nmz). - [Common Voice Scripted Speech 25.0 - Northern Hindko](https://mozilladatacollective.com/datasets/cmn2alr7p018jo1071gyxr951): A collection of read speech recordings in Northern Hindko (شمالی ہندکو). - [Common Voice Scripted Speech 25.0 - Guiziga](https://mozilladatacollective.com/datasets/cmn2alntw018fo10716d0qef6): A collection of read speech recordings in Guiziga (giz). - [Common Voice Scripted Speech 25.0 - Korean](https://mozilladatacollective.com/datasets/cmn2ale8p018bo107nbvos0f7): A collection of read speech recordings in Korean (한국어). - [Common Voice Scripted Speech 25.0 - Dawoodi](https://mozilladatacollective.com/datasets/cmn2al6jd01bvmm07tfcmt9k0): A collection of read speech recordings in Dawoodi (داوُدِݵ). - [Common Voice Scripted Speech 25.0 - Quechua Santiago del Estero](https://mozilladatacollective.com/datasets/cmn2akyt40187o107ihyanqz5): A collection of read speech recordings in Quechua Santiago del Estero (qus). - [Common Voice Scripted Speech 25.0 - Seri](https://mozilladatacollective.com/datasets/cmn2akuj101brmm0750ppoekz): A collection of read speech recordings in Seri (sei). - [Common Voice Scripted Speech 25.0 - Sorbian, Upper](https://mozilladatacollective.com/datasets/cmn2aepaq01bkmm073dkd515d): A collection of read speech recordings in Sorbian, Upper (Hornjoserbšćina). - [Common Voice Scripted Speech 25.0 - Kalasha](https://mozilladatacollective.com/datasets/cmn2aelvm01bgmm07kvcwqxpw): A collection of read speech recordings in Kalasha (Kalasha). - [Common Voice Scripted Speech 25.0 - Quechua Chiquián](https://mozilladatacollective.com/datasets/cmn2aeie901bcmm073fn2gl34): A collection of read speech recordings in Quechua Chiquián (qxa). - [Common Voice Scripted Speech 25.0 - Brahui](https://mozilladatacollective.com/datasets/cmn2aefkq01b8mm07dblhdy6q): A collection of read speech recordings in Brahui (براہوئی). - [Common Voice Scripted Speech 25.0 - Huautla Mazatec](https://mozilladatacollective.com/datasets/cmn2aeca0017uo107pcvp1zlu): A collection of read speech recordings in Huautla Mazatec (mau). - [Common Voice Scripted Speech 25.0 - Khetrani](https://mozilladatacollective.com/datasets/cmn2ae9xk01b4mm07c3333tkk): A collection of read speech recordings in Khetrani (کھیترانی). - [Common Voice Scripted Speech 25.0 - Tunen](https://mozilladatacollective.com/datasets/cmn2ae7fu01b0mm07588vrse1): A collection of read speech recordings in Tunen (tvu). - [Common Voice Scripted Speech 25.0 - Eton](https://mozilladatacollective.com/datasets/cmn2ae4ke01awmm075jsuvk3j): A collection of read speech recordings in Eton (eto). - [Common Voice Scripted Speech 25.0 - Quechua Cajatambo](https://mozilladatacollective.com/datasets/cmn2ae1x001asmm073je0m7tj): A collection of read speech recordings in Quechua Cajatambo (qvl). - [Common Voice Scripted Speech 25.0 - Kalkoti](https://mozilladatacollective.com/datasets/cmn2adzcv01aomm07mwmyaeoq): A collection of read speech recordings in Kalkoti (کلکوٹی). - [Common Voice Scripted Speech 25.0 - Albanian](https://mozilladatacollective.com/datasets/cmn29zkso01aimm07wb1ar40j): A collection of read speech recordings in Albanian (Shqip). - [Common Voice Scripted Speech 25.0 - Dhatki](https://mozilladatacollective.com/datasets/cmn29z2pb01aemm072es1wj06): A collection of read speech recordings in Dhatki (mki). - [Common Voice Scripted Speech 25.0 - Asheninka Perene](https://mozilladatacollective.com/datasets/cmn29z01q01aamm07oyzrkbnb): A collection of read speech recordings in Asheninka Perene (prq). - [Common Voice Scripted Speech 25.0 - Quechua Yanahuanca](https://mozilladatacollective.com/datasets/cmn29yxas01a6mm07ykfdt77p): A collection of read speech recordings in Quechua Yanahuanca (qur). - [Common Voice Scripted Speech 25.0 - Quechua Ambo-Pasco](https://mozilladatacollective.com/datasets/cmn29yuqh01a2mm07t68iibpo): A collection of read speech recordings in Quechua Ambo-Pasco (qva). - [Common Voice Scripted Speech 25.0 - Mokpwe](https://mozilladatacollective.com/datasets/cmn29ys5q019ymm07txqe23ap): A collection of read speech recordings in Mokpwe (bri). - [Common Voice Scripted Speech 25.0 - Nyungwe](https://mozilladatacollective.com/datasets/cmn29yppi019umm075uy16q04): A collection of read speech recordings in Nyungwe (nyu). - [Common Voice Scripted Speech 25.0 - Matses](https://mozilladatacollective.com/datasets/cmn29ymwn019qmm07nrqputof): A collection of read speech recordings in Matses (mcf). - [Common Voice Scripted Speech 25.0 - Hazargi](https://mozilladatacollective.com/datasets/cmn29yk9b019mmm07sixvc3j4): A collection of read speech recordings in Hazargi (haz). - [Common Voice Scripted Speech 25.0 - Kotokoli](https://mozilladatacollective.com/datasets/cmn29yhin019imm075uejkq7p): A collection of read speech recordings in Kotokoli (kdh). - [Common Voice Scripted Speech 25.0 - Tepeuxila Cuicatec](https://mozilladatacollective.com/datasets/cmn29vw2k019emm07n9jrgbn0): A collection of read speech recordings in Tepeuxila Cuicatec (cux). - [Common Voice Scripted Speech 25.0 - Yoruba](https://mozilladatacollective.com/datasets/cmn29vsoh019amm07d95id0mo): A collection of read speech recordings in Yoruba (Yòrùbá). - [Common Voice Scripted Speech 25.0 - Assamese](https://mozilladatacollective.com/datasets/cmn29vpx50196mm07mm08x8xm): A collection of read speech recordings in Assamese (অসমীয়া). - [Common Voice Scripted Speech 25.0 - Turkmen](https://mozilladatacollective.com/datasets/cmn29vm4b0192mm07fpfluszr): A collection of read speech recordings in Turkmen (Türkmençe). - [Common Voice Scripted Speech 25.0 - Mengambo](https://mozilladatacollective.com/datasets/cmn29vjg6017po107eo73vgqf): A collection of read speech recordings in Mengambo (bce). - [Common Voice Scripted Speech 25.0 - Hebrew](https://mozilladatacollective.com/datasets/cmn29vgka017lo107v8ebc8r1): A collection of read speech recordings in Hebrew (עברית). - [Common Voice Scripted Speech 25.0 - Ushojo](https://mozilladatacollective.com/datasets/cmn29vdo7017ho107gocd0uf2): A collection of read speech recordings in Ushojo (ush). - [Common Voice Scripted Speech 25.0 - Ekoti](https://mozilladatacollective.com/datasets/cmn29v9kx017do107ce0r76ku): A collection of read speech recordings in Ekoti (eko). - [Common Voice Scripted Speech 25.0 - Romansh Vallader](https://mozilladatacollective.com/datasets/cmn29u34u0179o107folbxhma): A collection of read speech recordings in Romansh Vallader (Rumantsch vallader). - [Common Voice Scripted Speech 25.0 - Ligurian](https://mozilladatacollective.com/datasets/cmn29tvj90175o107f982bzce): A collection of read speech recordings in Ligurian (Ligure). - [Common Voice Scripted Speech 25.0 - Punjabi](https://mozilladatacollective.com/datasets/cmn29trq60171o107e4l3tzlo): A collection of read speech recordings in Punjabi (ਪੰਜਾਬੀ). - [Common Voice Scripted Speech 25.0 - Setswana](https://mozilladatacollective.com/datasets/cmn29tnd0016xo107zjauonue): A collection of read speech recordings in Setswana (Setswana). - [Common Voice Scripted Speech 25.0 - Kazakh](https://mozilladatacollective.com/datasets/cmn29sufc018ymm071hvsk595): A collection of read speech recordings in Kazakh (Қазақ тілі). - [Common Voice Scripted Speech 25.0 - Malay](https://mozilladatacollective.com/datasets/cmn29sqde018umm071bb6u1ot): A collection of read speech recordings in Malay (Bahasa Melayu). - [Common Voice Scripted Speech 25.0 - Sardinian](https://mozilladatacollective.com/datasets/cmn29sn1b018qmm07kagz65ws): A collection of read speech recordings in Sardinian (Sardu). - [Common Voice Scripted Speech 25.0 - Cantonese](https://mozilladatacollective.com/datasets/cmn29rqn9016to107eniyak65): A collection of read speech recordings in Cantonese (粵語). - [Common Voice Scripted Speech 25.0 - Erzya](https://mozilladatacollective.com/datasets/cmn29m1vc016co107za0i4zp0): A collection of read speech recordings in Erzya (Эрзянь кель). - [Common Voice Scripted Speech 25.0 - Telugu](https://mozilladatacollective.com/datasets/cmn29lt270168o107nhemmxkh): A collection of read speech recordings in Telugu (తెలుగు). - [Common Voice Scripted Speech 25.0 - Amharic](https://mozilladatacollective.com/datasets/cmn29lq6f0164o10748yd3o7w): A collection of read speech recordings in Amharic (አማርኛ). - [Common Voice Scripted Speech 25.0 - Zaza](https://mozilladatacollective.com/datasets/cmn29lne80160o107scm9ustt): A collection of read speech recordings in Zaza (Kurdkî (Zazakî)). - [Common Voice Scripted Speech 25.0 - Norwegian Bokmål](https://mozilladatacollective.com/datasets/cmn29lkh2018kmm07ywneb3o0): A collection of read speech recordings in Norwegian Bokmål (Norsk (bokmål)). - [Common Voice Scripted Speech 25.0 - Ebrie](https://mozilladatacollective.com/datasets/cmn29lghi018gmm07rxencde0): A collection of read speech recordings in Ebrie (ebr). - [Common Voice Scripted Speech 25.0 - Yiddish](https://mozilladatacollective.com/datasets/cmn29l9yi018cmm07n4mdklvd): A collection of read speech recordings in Yiddish (אידיש). - [Common Voice Scripted Speech 25.0 - Tamazight](https://mozilladatacollective.com/datasets/cmn29l72y015wo107538tz82c): A collection of read speech recordings in Tamazight (ⵜⴰⵎⴰⵣⵉⵖⵜ). - [Common Voice Scripted Speech 25.0 - Nepali](https://mozilladatacollective.com/datasets/cmn29l48i015so107ibwqdewu): A collection of read speech recordings in Nepali (नेपाली). - [Common Voice Scripted Speech 25.0 - Norwegian Nynorsk](https://mozilladatacollective.com/datasets/cmn29l17i015oo107fb396rv4): A collection of read speech recordings in Norwegian Nynorsk (Norsk (nynorsk)). - [Common Voice Scripted Speech 25.0 - Azerbaijani](https://mozilladatacollective.com/datasets/cmn29hqvk015ko107fblsr5ay): A collection of read speech recordings in Azerbaijani (Azərbaycanca). - [Common Voice Scripted Speech 25.0 - Afrikaans](https://mozilladatacollective.com/datasets/cmn29hngy0188mm07kzspi4d4): A collection of read speech recordings in Afrikaans (Afrikaans). - [Common Voice Scripted Speech 25.0 - Alsatian](https://mozilladatacollective.com/datasets/cmn29hk3a0184mm072ledt03r): A collection of read speech recordings in Alsatian (Elsassisch). - [Common Voice Scripted Speech 25.0 - Ossetian](https://mozilladatacollective.com/datasets/cmn29hfds0180mm07xsk6my6x): A collection of read speech recordings in Ossetian (Ирон). - [Common Voice Scripted Speech 25.0 - Santali (Ol Chiki)](https://mozilladatacollective.com/datasets/cmn29h9y2015go107c2gfl7l2): A collection of read speech recordings in Santali (Ol Chiki) (ᱥᱟᱱᱛᱟᱲᱤ (ᱚᱞ ᱪᱤᱠᱤ)). - [Common Voice Scripted Speech 25.0 - Tajik](https://mozilladatacollective.com/datasets/cmn29h5x9017vmm079ptonmpm): A collection of read speech recordings in Tajik (Тоҷикӣ). - [Common Voice Scripted Speech 25.0 - Tigrinya](https://mozilladatacollective.com/datasets/cmn29h20x017nmm07t01j5evj): A collection of read speech recordings in Tigrinya (ትግርኛ). - [Common Voice Scripted Speech 25.0 - Western Sierra Puebla Nahuatl](https://mozilladatacollective.com/datasets/cmn29gycw0158o1074yca3w5x): A collection of read speech recordings in Western Sierra Puebla Nahuatl (nhi). - [Common Voice Scripted Speech 25.0 - Lao](https://mozilladatacollective.com/datasets/cmn29gv5j0154o107vfidhfbg): A collection of read speech recordings in Lao (ພາສາລາວ). - [Common Voice Scripted Speech 25.0 - Croatian](https://mozilladatacollective.com/datasets/cmn29gs96017jmm0705ypjod7): A collection of read speech recordings in Croatian (Hrvatski). - [Common Voice Scripted Speech 25.0 - Sorbian, Lower](https://mozilladatacollective.com/datasets/cmn29gp8o017fmm07l18jdkdf): A collection of read speech recordings in Sorbian, Lower (Dolnoserbšćina). - [Common Voice Scripted Speech 25.0 - Uzbek](https://mozilladatacollective.com/datasets/cmn29fa4y0150o1073fqww73p): A collection of read speech recordings in Uzbek (O‘zbek). - [Common Voice Scripted Speech 25.0 - Portuguese](https://mozilladatacollective.com/datasets/cmn29f4cb017bmm07pd9yd8mw): A collection of read speech recordings in Portuguese (Português). - [Common Voice Scripted Speech 25.0 - Abkhaz](https://mozilladatacollective.com/datasets/cmn29f1jp014to1077s6w9o6g): A collection of read speech recordings in Abkhaz (Аԥсшәа). - [Common Voice Scripted Speech 25.0 - Bashkir](https://mozilladatacollective.com/datasets/cmn29exhf014po107d6mpl5ec): A collection of read speech recordings in Bashkir (Башҡорт). - [Common Voice Scripted Speech 25.0 - Polish](https://mozilladatacollective.com/datasets/cmn27nz69015hmm0720txf781): A collection of read speech recordings in Polish (polski). - [Common Voice Scripted Speech 25.0 - Hill Mari](https://mozilladatacollective.com/datasets/cmn1qh20m00zymm07mka0fu7i): A collection of read speech recordings in Hill Mari (Кырык мары). - [Common Voice Scripted Speech 25.0 - Guarani](https://mozilladatacollective.com/datasets/cmn1qfbos00zumm079uvzxgtb): A collection of read speech recordings in Guarani (Guarani). - [Common Voice Scripted Speech 25.0 - Interlingua](https://mozilladatacollective.com/datasets/cmn1qf7ns00y3o107oa3cxsut): A collection of read speech recordings in Interlingua (Interlingua). - [Common Voice Scripted Speech 25.0 - Ngiembon](https://mozilladatacollective.com/datasets/cmn1qf3al00xzo107byg0pine): A collection of read speech recordings in Ngiembon (nnh). - [Common Voice Scripted Speech 25.0 - Bafia](https://mozilladatacollective.com/datasets/cmn1qeyjg00xvo107ica6bh2f): A collection of read speech recordings in Bafia (ksf). - [Common Voice Scripted Speech 25.0 - Chokwe](https://mozilladatacollective.com/datasets/cmn1qeu5500xro107r3og0kg7): A collection of read speech recordings in Chokwe (cjk). - [Common Voice Scripted Speech 25.0 - Occitan](https://mozilladatacollective.com/datasets/cmn1qeqny00xno107p0rrn9lq): A collection of read speech recordings in Occitan (occitan). - [Common Voice Scripted Speech 25.0 - Hausa](https://mozilladatacollective.com/datasets/cmn1qen4q00xjo107gln14ztz): A collection of read speech recordings in Hausa (Hausa). - [Common Voice Scripted Speech 25.0 - Mada](https://mozilladatacollective.com/datasets/cmn1qejfy00zmmm07m68oodlv): A collection of read speech recordings in Mada (mxu). - [Common Voice Scripted Speech 25.0 - Gurgula](https://mozilladatacollective.com/datasets/cmn1qdzix00xdo107oun0cxj8): A collection of read speech recordings in Gurgula (ggg). - [Common Voice Scripted Speech 25.0 - Nuasue](https://mozilladatacollective.com/datasets/cmn1qdhpu00zimm07szibarzh): A collection of read speech recordings in Nuasue (yav). - [Common Voice Scripted Speech 25.0 - Mbo](https://mozilladatacollective.com/datasets/cmn1qc3ct00zemm07h05b4qls): A collection of read speech recordings in Mbo (mbo). - [Common Voice Scripted Speech 25.0 - Kirombo](https://mozilladatacollective.com/datasets/cmn1qbzuw00x9o1071mhurgt4): A collection of read speech recordings in Kirombo (rof). - [Common Voice Scripted Speech 25.0 - Kunabembe](https://mozilladatacollective.com/datasets/cmn1qbwf000x5o107frlc1ukd): A collection of read speech recordings in Kunabembe (mgg). - [Common Voice Scripted Speech 25.0 - Mungaka](https://mozilladatacollective.com/datasets/cmn1qbs8000x1o107crsgsqv3): A collection of read speech recordings in Mungaka (mhk). - [Common Voice Scripted Speech 25.0 - Massa](https://mozilladatacollective.com/datasets/cmn1qbo6w00wxo107nxpj67j5): A collection of read speech recordings in Massa (mcn). - [Common Voice Scripted Speech 25.0 - Kwasio](https://mozilladatacollective.com/datasets/cmn1qbjqy00wto10779vs86bz): A collection of read speech recordings in Kwasio (nmg). - [Common Voice Scripted Speech 25.0 - Mpiemo](https://mozilladatacollective.com/datasets/cmn1qb5gg00zamm07v7vv9an3): A collection of read speech recordings in Mpiemo (mcx). - [Common Voice Scripted Speech 25.0 - Tshiluba](https://mozilladatacollective.com/datasets/cmn1qaqkp00wmo107jncj3y1j): A collection of read speech recordings in Tshiluba (lua). - [Common Voice Scripted Speech 25.0 - Northwest Gbaya](https://mozilladatacollective.com/datasets/cmn1qamc600wio1079l8jwbkf): A collection of read speech recordings in Northwest Gbaya (gya). - [Common Voice Scripted Speech 25.0 - Tlingit](https://mozilladatacollective.com/datasets/cmn1qahdl00weo107s7r9ceds): A collection of read speech recordings in Tlingit (tli). - [Common Voice Scripted Speech 25.0 - Cameroon Pidgin](https://mozilladatacollective.com/datasets/cmn1qa0u300z4mm07egfgo1k4): A collection of read speech recordings in Cameroon Pidgin (wes). - [Common Voice Scripted Speech 25.0 - Kihemba](https://mozilladatacollective.com/datasets/cmn1q9gn100z0mm07fee16bjb): A collection of read speech recordings in Kihemba (hem). - [Common Voice Scripted Speech 25.0 - Mbum](https://mozilladatacollective.com/datasets/cmn1q93fk00ywmm07917rri91): A collection of read speech recordings in Mbum (mdd). - [Common Voice Scripted Speech 25.0 - Ouldémé](https://mozilladatacollective.com/datasets/cmn1q8p8e00w6o107ff2k89o8): A collection of read speech recordings in Ouldémé (udl). - [Common Voice Scripted Speech 25.0 - Mundang](https://mozilladatacollective.com/datasets/cmn1q88st00ysmm071vj0faa1): A collection of read speech recordings in Mundang (mua). - [Common Voice Scripted Speech 25.0 - Ngombale](https://mozilladatacollective.com/datasets/cmn1q7ug300w2o107mkunl96o): A collection of read speech recordings in Ngombale (nla). - [Common Voice Scripted Speech 25.0 - Lassi](https://mozilladatacollective.com/datasets/cmn1q7dpq00yomm074hxklb6h): A collection of read speech recordings in Lassi (لاسي). - [Common Voice Scripted Speech 25.0 - Hakha Chin](https://mozilladatacollective.com/datasets/cmn1q6t6u00ykmm07okk16ed2): A collection of read speech recordings in Hakha Chin (Laiholh (Hakha)). - [Common Voice Scripted Speech 25.0 - Moussey](https://mozilladatacollective.com/datasets/cmn1q5vro00ygmm07omw8p4b5): A collection of read speech recordings in Moussey (mse). - [Common Voice Scripted Speech 25.0 - Iñupiaq](https://mozilladatacollective.com/datasets/cmn1q5h2g00ycmm074ucll31q): A collection of read speech recordings in Iñupiaq (ipk). - [Common Voice Scripted Speech 25.0 - Central Alaskan Yup’ik](https://mozilladatacollective.com/datasets/cmn1q4qsg00vyo107sjh2vufw): A collection of read speech recordings in Central Alaskan Yup’ik (esu). - [Common Voice Scripted Speech 25.0 - Saraiki](https://mozilladatacollective.com/datasets/cmn1q4nr800vuo107m8htiiy4): A collection of read speech recordings in Saraiki (سرائیکی). - [Common Voice Scripted Speech 25.0 - Musgum](https://mozilladatacollective.com/datasets/cmn1q4jwf00y8mm07bs03sz3l): A collection of read speech recordings in Musgum (mug). - [Common Voice Scripted Speech 25.0 - Asturian](https://mozilladatacollective.com/datasets/cmn1q4g7i00y4mm07eb1yeh11): A collection of read speech recordings in Asturian (Asturianu). - [Common Voice Scripted Speech 25.0 - Quechua Chanka](https://mozilladatacollective.com/datasets/cmn1q4bb400vno1075pvgxlib): A collection of read speech recordings in Quechua Chanka (Quechua Chanka). - [Common Voice Scripted Speech 25.0 - Icelandic](https://mozilladatacollective.com/datasets/cmn1q47o300vjo1078y9lpkff): A collection of read speech recordings in Icelandic (Íslenska). - [Common Voice Scripted Speech 25.0 - Twi](https://mozilladatacollective.com/datasets/cmn1q42wn00y0mm07dfoe454l): A collection of read speech recordings in Twi (Twi). - [Common Voice Scripted Speech 25.0 - Moksha](https://mozilladatacollective.com/datasets/cmn1q3wj800vfo107qflke6e7): A collection of read speech recordings in Moksha (Мокшень кяль). - [Common Voice Scripted Speech 25.0 - Dioula](https://mozilladatacollective.com/datasets/cmn1q3sgr00xwmm07t7te56k4): A collection of read speech recordings in Dioula (Dioula ye). - [Common Voice Scripted Speech 25.0 - Votic](https://mozilladatacollective.com/datasets/cmn1q3p3l00xsmm076b72jjou): A collection of read speech recordings in Votic (vad̕d̕a). - [Common Voice Scripted Speech 25.0 - Northern Sotho](https://mozilladatacollective.com/datasets/cmn1q2jfh00xomm072a2mqrcj): A collection of read speech recordings in Northern Sotho (Sesotho sa Leboa). - [Common Voice Scripted Speech 25.0 - Xhosa](https://mozilladatacollective.com/datasets/cmn1q21i400xkmm0707onm8fb): A collection of read speech recordings in Xhosa (IsiXhosa). - [Common Voice Scripted Speech 25.0 - Southern Sotho](https://mozilladatacollective.com/datasets/cmn1q1rhn00vbo107b64tb5b8): A collection of read speech recordings in Southern Sotho (Sesotho sa Borwa). - [Common Voice Scripted Speech 25.0 - Siswati](https://mozilladatacollective.com/datasets/cmn1q1f3q00xgmm07150pr650): A collection of read speech recordings in Siswati (Siswati). - [Common Voice Scripted Speech 25.0 - Aromanian](https://mozilladatacollective.com/datasets/cmn1q11oz00v7o107a25yjpsp): A collection of read speech recordings in Aromanian (rup). - [Common Voice Scripted Speech 25.0 - Zulu](https://mozilladatacollective.com/datasets/cmn1q0ng600xcmm07hhwtnuq7): A collection of read speech recordings in Zulu (Zulu). - [Common Voice Scripted Speech 25.0 - Tshivenda](https://mozilladatacollective.com/datasets/cmn1pzszy00x8mm07fkoydxpq): A collection of read speech recordings in Tshivenda (Tshivenḓa). - [Common Voice Scripted Speech 25.0 - Haitian](https://mozilladatacollective.com/datasets/cmn1pz91w00v3o107hknri5xy): A collection of read speech recordings in Haitian (Ayisyen). - [Common Voice Scripted Speech 25.0 - Xitsonga](https://mozilladatacollective.com/datasets/cmn1pyucy00x4mm071setoehx): A collection of read speech recordings in Xitsonga (Xitsonga). - [Common Voice Scripted Speech 25.0 - IsiNdebele (South)](https://mozilladatacollective.com/datasets/cmn1pyaq600uzo1079yby8vs2): A collection of read speech recordings in IsiNdebele (South) (IsiNdebele (Sewula)). - [Common Voice Spontaneous Speech 3.0 - English](https://mozilladatacollective.com/datasets/cmn1pv5hi00uto1072y1074y7): A collection of spontaneous responses to questions in English (English). - [Common Voice Spontaneous Speech 3.0 - Puno Quechua](https://mozilladatacollective.com/datasets/cmn1pujk200uno107g6el5r9y): A collection of spontaneous responses to questions in Puno Quechua (Punu qhichwa). - [Common Voice Spontaneous Speech 3.0 - Scots](https://mozilladatacollective.com/datasets/cmn1ptsol00wsmm070mh2j79o): A collection of spontaneous responses to questions in Scots (sco). - [Common Voice Spontaneous Speech 3.0 - Michoacán Mazahua](https://mozilladatacollective.com/datasets/cmn1ptaru00ufo107w2z6gzub): A collection of spontaneous responses to questions in Michoacán Mazahua (Jñatjo). - [Common Voice Spontaneous Speech 3.0 - Gorani](https://mozilladatacollective.com/datasets/cmn1psdr200ubo107t8sf2ro5): A collection of spontaneous responses to questions in Gorani (hac). - [Common Voice Spontaneous Speech 3.0 - Betawi](https://mozilladatacollective.com/datasets/cmn1prwpx00wlmm07aq3out8f): A collection of spontaneous responses to questions in Betawi (bew). - [Common Voice Spontaneous Speech 3.0 - Cypriot Greek](https://mozilladatacollective.com/datasets/cmn1prgv200wdmm074tv16ejk): A collection of spontaneous responses to questions in Cypriot Greek (Κυπριακά Ελληνικά). - [Common Voice Spontaneous Speech 3.0 - Papantla Totonac](https://mozilladatacollective.com/datasets/cmn1pqa0900w9mm074zrq7tbr): A collection of spontaneous responses to questions in Papantla Totonac (top). - [Common Voice Spontaneous Speech 3.0 - Pashto](https://mozilladatacollective.com/datasets/cmn1ppxp700w5mm079ze83bhb): A collection of spontaneous responses to questions in Pashto (پښتو). - [Common Voice Spontaneous Speech 3.0 - Toba Qom](https://mozilladatacollective.com/datasets/cmn1ppks200u7o1078xxw5izl): A collection of spontaneous responses to questions in Toba Qom (tob). - [Common Voice Spontaneous Speech 3.0 - Kabardian](https://mozilladatacollective.com/datasets/cmn1pp8ln00w1mm07xega82ar): A collection of spontaneous responses to questions in Kabardian (Адыгэбзэ (Къэбэрдей)). - [Common Voice Spontaneous Speech 3.0 - Basaa](https://mozilladatacollective.com/datasets/cmn1pow3f00vtmm07dre8danl): A collection of spontaneous responses to questions in Basaa (Basaa). - [Common Voice Spontaneous Speech 3.0 - Adyghe](https://mozilladatacollective.com/datasets/cmn1polzs00u3o1074eqe1d2w): A collection of spontaneous responses to questions in Adyghe (Адыгабзэ). - [Common Voice Spontaneous Speech 3.0 - Ushojo](https://mozilladatacollective.com/datasets/cmn1poblg00vnmm075qaox16n): A collection of spontaneous responses to questions in Ushojo (ush). - [Common Voice Spontaneous Speech 3.0 - Galician](https://mozilladatacollective.com/datasets/cmn1pnmci00vjmm07lum8cjfd): A collection of spontaneous responses to questions in Galician (Galego). - [Common Voice Spontaneous Speech 3.0 - Russian](https://mozilladatacollective.com/datasets/cmn1pnb4n00vfmm07eqydzilq): A collection of spontaneous responses to questions in Russian (Русский). - [Common Voice Spontaneous Speech 3.0 - Ligurian](https://mozilladatacollective.com/datasets/cmn1pn1qk00tvo1070q0jzudy): A collection of spontaneous responses to questions in Ligurian (Ligure). - [Common Voice Spontaneous Speech 3.0 - Arvanitika](https://mozilladatacollective.com/datasets/cmn1pmsl400tro107gce75iqt): A collection of spontaneous responses to questions in Arvanitika (aat). - [Common Voice Spontaneous Speech 3.0 - Sena](https://mozilladatacollective.com/datasets/cmn1pmgxb00vbmm07a9dgloq9): A collection of spontaneous responses to questions in Sena (seh). - [Common Voice Spontaneous Speech 3.0 - Breton](https://mozilladatacollective.com/datasets/cmn1pm3tk00v7mm07wkhdv5p9): A collection of spontaneous responses to questions in Breton (Brezhoneg). - [Common Voice Spontaneous Speech 3.0 - Catalan](https://mozilladatacollective.com/datasets/cmn1plu3z00uzmm07cacsbro3): A collection of spontaneous responses to questions in Catalan (català). - [Common Voice Spontaneous Speech 3.0 - Latvian](https://mozilladatacollective.com/datasets/cmn1pllvd00tlo10743ce16nh): A collection of spontaneous responses to questions in Latvian (Latviešu). - [Common Voice Spontaneous Speech 3.0 - Turkish](https://mozilladatacollective.com/datasets/cmn1pleap00tho107nvvwnbyj): A collection of spontaneous responses to questions in Turkish (Türkçe). - [Common Voice Spontaneous Speech 3.0 - Welsh](https://mozilladatacollective.com/datasets/cmn1pl37b00uvmm07woxp33ef): A collection of spontaneous responses to questions in Welsh (Cymraeg). - [Common Voice Spontaneous Speech 3.0 - Nubi](https://mozilladatacollective.com/datasets/cmmytprab00ifmf074pd0bvn3): A collection of spontaneous responses to questions in Nubi (kcn). - [Common Voice Spontaneous Speech 3.0 - Thur](https://mozilladatacollective.com/datasets/cmmytpom600ibmf07abftw5m0): A collection of spontaneous responses to questions in Thur (lth). - [Common Voice Spontaneous Speech 3.0 - Konzo](https://mozilladatacollective.com/datasets/cmmytnyhe00i3mf07l1q1gwx1): A collection of spontaneous responses to questions in Konzo (koo). - [Common Voice Spontaneous Speech 3.0 - Sabah Malay](https://mozilladatacollective.com/datasets/cmmytmz7k00fxnz07yj3lastx): A collection of spontaneous responses to questions in Sabah Malay (msi). - [Common Voice Spontaneous Speech 3.0 - Rutoro](https://mozilladatacollective.com/datasets/cmmytm3ip00humf07dmc8r7lo): A collection of spontaneous responses to questions in Rutoro (ttj). - [Common Voice Spontaneous Speech 3.0 - Lendu](https://mozilladatacollective.com/datasets/cmmytlaj900fpnz074nf5ei09): A collection of spontaneous responses to questions in Lendu (led). - [Common Voice Spontaneous Speech 3.0 - Amba](https://mozilladatacollective.com/datasets/cmmytkjq600hqmf07m6ychfex): A collection of spontaneous responses to questions in Amba (rwm). - [Common Voice Spontaneous Speech 3.0 - Bukusu](https://mozilladatacollective.com/datasets/cmmytjs2c00hmmf07towm00mv): A collection of spontaneous responses to questions in Bukusu (bxk). - [Common Voice Spontaneous Speech 3.0 - Kenyi](https://mozilladatacollective.com/datasets/cmmytiprz00himf07y1a1qhvu): A collection of spontaneous responses to questions in Kenyi (lke). - [Common Voice Spontaneous Speech 3.0 - Western Penan](https://mozilladatacollective.com/datasets/cmmythoy000fenz07az97qz0h): A collection of spontaneous responses to questions in Western Penan (pne). - [Common Voice Spontaneous Speech 3.0 - Eastern Min](https://mozilladatacollective.com/datasets/cmmytgpsh00hemf07itahf0oc): A collection of spontaneous responses to questions in Eastern Min (cdo). - [Common Voice Spontaneous Speech 3.0 - Bahasa Malay](https://mozilladatacollective.com/datasets/cmmytgn7200fanz071knkp7gb): A collection of spontaneous responses to questions in Bahasa Malay (Bahasa Melayu). - [Common Voice Spontaneous Speech 3.0 - Alsatian](https://mozilladatacollective.com/datasets/cmmytgkr900f6nz073umxgzk3): A collection of spontaneous responses to questions in Alsatian (Elsassisch). - [Common Voice Spontaneous Speech 3.0 - French](https://mozilladatacollective.com/datasets/cmmytgij900f2nz07xm0wyzrd): A collection of spontaneous responses to questions in French (Français). - [Common Voice Spontaneous Speech 3.0 - German](https://mozilladatacollective.com/datasets/cmmytggb100hamf07z06g4iae): A collection of spontaneous responses to questions in German (Deutsch). - [Common Voice Spontaneous Speech 3.0 - Gheg Albanian](https://mozilladatacollective.com/datasets/cmmytg2r000eynz07epeq9u30): A collection of spontaneous responses to questions in Gheg Albanian (aln). - [Common Voice Spontaneous Speech 3.0 - Mixteco Yucuhiti](https://mozilladatacollective.com/datasets/cmmytg00z00h6mf076gsilniw): A collection of spontaneous responses to questions in Mixteco Yucuhiti (meh). - [Common Voice Spontaneous Speech 3.0 - Sabah Bisaya](https://mozilladatacollective.com/datasets/cmmytfxan00eunz07ormig9ra): A collection of spontaneous responses to questions in Sabah Bisaya (bsy). - [Common Voice Spontaneous Speech 3.0 - Heng Hua](https://mozilladatacollective.com/datasets/cmmytftf900eqnz0760kydlnm): A collection of spontaneous responses to questions in Heng Hua (cpx). - [Common Voice Spontaneous Speech 3.0 - Kuku](https://mozilladatacollective.com/datasets/cmmytfqkp00emnz072015yasf): A collection of spontaneous responses to questions in Kuku (ukv). - [Common Voice Spontaneous Speech 3.0 - Chiga](https://mozilladatacollective.com/datasets/cmmyte4zw00einz07mucw81ok): A collection of spontaneous responses to questions in Chiga (cgg). - [Common Voice Spontaneous Speech 3.0 - Sa’ban](https://mozilladatacollective.com/datasets/cmmytb9jf00gvmf077g36y3r7): A collection of spontaneous responses to questions in Sa’ban (snv). - [Common Voice Spontaneous Speech 3.0 - Kenyah](https://mozilladatacollective.com/datasets/cmmyt67yw00grmf07uggmsoyk): A collection of spontaneous responses to questions in Kenyah (xkl). - [Common Voice Spontaneous Speech 3.0 - Melanau](https://mozilladatacollective.com/datasets/cmmyt5c2400ecnz07hk5mn5o7): A collection of spontaneous responses to questions in Melanau (mel). - [Common Voice Spontaneous Speech 3.0 - Wixárika](https://mozilladatacollective.com/datasets/cmmyt1ceg00gnmf07ynjd0avk): A collection of spontaneous responses to questions in Wixárika (hch). - [Common Voice Spontaneous Speech 3.0 - Kelabit](https://mozilladatacollective.com/datasets/cmmyt0mhv00e8nz07j2yjk8g7): A collection of spontaneous responses to questions in Kelabit (kzi). - [Common Voice Spontaneous Speech 3.0 - Aragonese](https://mozilladatacollective.com/datasets/cmmyswoay00e4nz07d005aacz): A collection of spontaneous responses to questions in Aragonese (Aragonés). - [Common Voice Spontaneous Speech 3.0 - Tashlhiyt](https://mozilladatacollective.com/datasets/cmmysvqya00gjmf07wtgzt1nd): A collection of spontaneous responses to questions in Tashlhiyt (shi). - [Common Voice Spontaneous Speech 3.0 - Manx](https://mozilladatacollective.com/datasets/cmmyss5a200gcmf07e4hffezp): A collection of spontaneous responses to questions in Manx (gv). - [Common Voice Spontaneous Speech 3.0 - Tudaga](https://mozilladatacollective.com/datasets/cmmysq4af00g8mf07gkpqtwx5): A collection of spontaneous responses to questions in Tudaga (tuq). - [Common Voice Spontaneous Speech 3.0 - Sundanese](https://mozilladatacollective.com/datasets/cmmysp2sf00e0nz0767ofkj4r): A collection of spontaneous responses to questions in Sundanese (Basa Sunda). - [Common Voice Spontaneous Speech 3.0 - Spanish](https://mozilladatacollective.com/datasets/cmmysoeaj00g4mf074e6j5cw5): A collection of spontaneous responses to questions in Spanish (Español). - [Common Voice Spontaneous Speech 3.0 - Esperanto](https://mozilladatacollective.com/datasets/cmmysniwj00g0mf07o2efljbv): A collection of spontaneous responses to questions in Esperanto (Esperanto). - [Common Voice Spontaneous Speech 3.0 - Georgian](https://mozilladatacollective.com/datasets/cmmysmqds00fwmf07e72ap8dg): A collection of spontaneous responses to questions in Georgian (ქართული). - [Common Voice Spontaneous Speech 3.0 - Rakhine](https://mozilladatacollective.com/datasets/cmmyslwhj00fsmf07a1iwom0w): A collection of spontaneous responses to questions in Rakhine (rki). - [Common Voice Spontaneous Speech 3.0 - Bashkir](https://mozilladatacollective.com/datasets/cmmysgpd600fomf07qipsogyg): A collection of spontaneous responses to questions in Bashkir (Башҡорт). - [Common Voice Spontaneous Speech 3.0 - Javanese](https://mozilladatacollective.com/datasets/cmmysfmtu00dwnz078kvqu1gp): A collection of spontaneous responses to questions in Javanese (basa jawa). - [Common Voice Spontaneous Speech 3.0 - Sinhala](https://mozilladatacollective.com/datasets/cmmysetwg00fkmf07v0akjit6): A collection of spontaneous responses to questions in Sinhala (සිංහල). - [Common Voice Spontaneous Speech 3.0 - Dutch](https://mozilladatacollective.com/datasets/cmmyse72100dsnz07i4e2cijq): A collection of spontaneous responses to questions in Dutch (Nederlands). - [Common Voice Spontaneous Speech 3.0 - Shona](https://mozilladatacollective.com/datasets/cmmysd3sh00fgmf07lioiqs3z): A collection of spontaneous responses to questions in Shona (sn). - [Common Voice Spontaneous Speech 3.0 - Bodo](https://mozilladatacollective.com/datasets/cmmysc65y00fcmf075nhrqs9k): A collection of spontaneous responses to questions in Bodo (brx). - [Common Voice Spontaneous Speech 3.0 - Thai](https://mozilladatacollective.com/datasets/cmmysas5x00f8mf07lbevjqbb): A collection of spontaneous responses to questions in Thai (ไทย). - [Common Voice Spontaneous Speech 3.0 - Frisian](https://mozilladatacollective.com/datasets/cmmys9qyv00f4mf0794yqc63c): A collection of spontaneous responses to questions in Frisian (Frysk). - [Common Voice Spontaneous Speech 3.0 - Croatian](https://mozilladatacollective.com/datasets/cmmy9phqr0064nz07v3w1pfww): A collection of spontaneous responses to questions in Croatian (Hrvatski). - [Common Voice Spontaneous Speech 3.0 - Danish](https://mozilladatacollective.com/datasets/cmmy9neso0067mf07gf9u0rdu): A collection of spontaneous responses to questions in Danish (Dansk). - [Common Voice Spontaneous Speech 3.0 - Ruuli](https://mozilladatacollective.com/datasets/cmmy9n21i0060nz07x66x631h): A collection of spontaneous responses to questions in Ruuli (ruc). - [Common Voice Spontaneous Speech 3.0 - Irish](https://mozilladatacollective.com/datasets/cmmy33q7l0018mf0721oocrh7): A collection of spontaneous responses to questions in Irish (Gaeilge). - [Istorima](https://mozilladatacollective.com/datasets/cmmxib8ic00v6nw07bjrci8vj): Dataset Language: Greek Dataset Info: This dataset consists of oral history content collected from the Istorima archive, including transcribed interviews and associated metadata. The material reflects personal narratives and life stories, primarily in Greek, covering a wide range of social, cultural, and historical topics. - [UP - DSP - Philippine Languages Database (UP-DSP-PLD)](https://mozilladatacollective.com/datasets/cmmxhw46c00tqnw07xyr94zjk): This dataset contains multilingual, text and speech pairs for ten Philippine languages namely Filipino, English, Cebuano, Kapampangan, Hiligaynon, Ilokano, Bikolano, Waray, and Tausug. The dataset contains over 454 hours of recordings, covering multiple domains in news, medical, education, tourism and spontaneous speech. - [Urdu Multi-Speaker TTS Dataset](https://mozilladatacollective.com/datasets/cmmvykcrs0050ny07vkwww5gi): This dataset is an Urdu text-to-speech corpus designed for speech technology development and related computational research. It contains approximately 10 hours of speech from 3 speakers, including 2 male and 1 female speaker. - [BECO Brahui Literature Corpus](https://mozilladatacollective.com/datasets/cmmus79780030ms07vyi4dsr5): This Brahui literary corpus contains short stories, novels, and other creative literary works, representing a broad range of narrative styles and themes within Brahui literature. The texts reflect both classical and contemporary writing, offering insight into cultural expression and linguistic variation in Brahui. - [ddd-kenya-somali-68hrs-asr-part1](https://mozilladatacollective.com/datasets/cmmng8btl000yl807k8qtx891): This dataset, curated by Digital Divide Data (DDD), provides high-quality audio recordings and corresponding text transcriptions for the Somali (som) language. The collection includes thousands of unique utterances per language to support diverse acoustic modeling. - [Malayalam Time-Aligned Speech Corpus](https://mozilladatacollective.com/datasets/cmmno795h009hml07dh7uefvp): This dataset is a speaker-organized Malayalam speech corpus consisting of 100 audio recordings and 100 corresponding transcription files in .srt format. The transcriptions are time-aligned and include timestamps matched to the audio. - [ddd-kenya-somali-68hrs-asr-part3](https://mozilladatacollective.com/datasets/cmmniydvi0047ml07nt6z5xud): This dataset, curated by Digital Divide Data (DDD), provides high-quality audio recordings and corresponding text transcriptions for the Somali (som) language. The collection includes thousands of unique utterances per language to support diverse acoustic modeling. - [ddd-kenya-somali-68hrs-asr-part2](https://mozilladatacollective.com/datasets/cmmnixx2d0043ml07zbq1i7oi): This dataset, curated by Digital Divide Data (DDD), provides high-quality audio recordings and corresponding text transcriptions for the Somali (som) language. The collection includes thousands of unique utterances per language to support diverse acoustic modeling. - [ddd-kenya-luhya-70hrs-asr](https://mozilladatacollective.com/datasets/cmkh2stj6004mmj0751rn8vq6): A 70-hour subset of Luhya speech data collected by Digital Divide Data in Kenya. The dataset includes recorded sentences from native speakers and is intended to support research and development in Automatic Speech Recognition for low-resource African languages. - [TODa: Tamazight Open Dataset](https://mozilladatacollective.com/datasets/cmmm7rvm200b9md07h3pv8uae): Welcome to the Tamazight Open Dataset (TODa), a groundbreaking open-source project dedicated to preserving and advancing the Tamazight language. With its extensive collection of linguistic data, TODa stands as a pioneering collaborative project for Tamazight <=> Englis translation, specifically designed for Natural Language Processing applications. - [Kokoro Speech Dataset](https://mozilladatacollective.com/datasets/cmmknsho4014wmf087kvq5rc6): Kokoro Speech Dataset is a public domain Japanese speech dataset. It contains 43,253 short audio clips of a single speaker reading 14 novel books. - [Sundanese TTS](https://mozilladatacollective.com/datasets/cmmj6vyb902ownz07j4k7cunj): The Sundanese TTS dataset represents the Sundanese language using the Priangan Sundanese dialect as the standard Sundanese in West Java province, Indonesia, reflecting both traditional forms and modern variations in everyday communication practices. This dataset can be utilized for linguistic research, cultural documentation, sociolinguistic studies, and the development of regional language technologies involving code-mixing with Indonesian. - [Elkhani Hazargi Literature Corpus](https://mozilladatacollective.com/datasets/cmmdtenxt0050mh0792d10knv): The Hazargi Literature Corpus (Keblagh e Azergi) is a monolingual literary dataset for documenting and supporting computational research on Hazargi (Hazaragi), an eastern Persian (Dari) dialect spoken by Hazara communities in Afghanistan and the diaspora. It contains 12 digitized works (prose, poetry, folklore, drama) converted from Word into UTF-8 normalized plain text while preserving original orthography and dialectal features. - [Dari Literature Corpus by Anjuman e Adabi Nayestan](https://mozilladatacollective.com/datasets/cmmdpikpq003imh077foix53d): The Dari Literature Corpus (Anjuman e Adabi Nayestan) is a curated collection of written Dari (Afghan Persian) literary texts totaling about 1 million tokens. It includes prose, poetry, folklore-inspired narratives, and other culturally significant writings from both contemporary and classical traditions. - [IBT Torwali Wordlist](https://mozilladatacollective.com/datasets/cmmdpbs5z003emh07yvbymzu5): The IBT Torwali Wordlist contains approximately 20,000 unique entries in Torwali (ISO 639-3: trw), an under-documented Indo-Aryan language spoken in northern Pakistan. The dataset comprises standardized lexical entries covering core vocabulary, function words, and culturally salient terms, with consistent orthography and normalization suitable for linguistic and computational use. - [Bangor Siarad Welsh-English Corpus](https://mozilladatacollective.com/datasets/cmmccg5qi00enmu07w9wjpnrn): The Siarad Welsh-English corpus, containing around 450,000 words, 84% Welsh, 4% English, 13% indeterminate (the relevant word appears in the dictionaries of both main languages). The dataset includes the audios, transcriptions and glosses in CHAT format, and word-level analyses of the transcriptions in .tsv files. - [Bangor Patagonia Welsh-Spanish Corpus](https://mozilladatacollective.com/datasets/cmmccfc5000efmu07ommi3zfr): The Patagonia Welsh-Spanish corpus contains around 195,000 words: 78% Welsh, 17% Spanish, 5% indeterminate (i.e. the relevant word appears in the dictionaries of both main languages). - [Saraiki-English Parallel Corpus](https://mozilladatacollective.com/datasets/cmmaphscg04t2mk07i1f8yc0q): This English–Saraiki Parallel Corpus is a curated bilingual dataset of 51,447 aligned sentence pairs (about 0.89 million words in total), translated from English into Saraiki by Kaleem Art Press and cleaned into a consistent sentence-level format for reliable alignment; it is designed to support machine translation training and evaluation, bilingual lexicon and terminology work, and broader linguistic and NLP research for Saraiki, including data-driven language technology development. - [Jhoke Publisher Multan’s Saraiki Newspaper Corpus](https://mozilladatacollective.com/datasets/cmmao7dc504oamh0710j4wau1): Jhoke Publishers Multan’s Saraiki Newspaper Corpus is a curated text dataset with about 1.25M tokens (1,258K) of Saraiki content collected from Daily Jhoke Saraiki (Multan, Pakistan) and Jhoke Publishers (Multan, Pakistan). Daily Jhoke Multan (ݙین٘ھ وار جھوک ملتان) is a Saraiki newspaper and publishing house based in Multan. - [Javanese TTS of Banyumasan Dialect](https://mozilladatacollective.com/datasets/cmmamtaf104nemh07xa9e7sdx): This dataset comprises speech data produced by a speaker of the Banyumasan dialect of Javanese (locally known as Ngapak), Central Java Province, Indonesia. All datasets use the informal register (Ngoko) and include various topics. - [Dolgan Folklore Text Corpus](https://mozilladatacollective.com/datasets/cmm0n4ro2000hnq079tknw6gv): This corpus contains a curated collection of 19 Dolgan fairy tales (15,618 words) digitized from a 2000 academic volume published in Novosibirsk. The Dolgans are the northernmost Turkic-speaking people, and their language is highly endangered. - [Finnish Public Domain 20th Century Literature Text Corpus](https://mozilladatacollective.com/datasets/cmm5078n50168mk07v64792sf): This corpus contains a curated collection of public domain literature from Finland, featuring works by authors who died between 1901 and 1955. The dataset captures the literary landscape of early 20th-century Finland and includes independent texts in both of the country's official languages: Finnish (fi) and Swedish (sv). - [Thorsten-Voice-44kHz-Full](https://mozilladatacollective.com/datasets/cmm4de9w500ntmh073nx14k7p): TV-44kHz-Full is a high-quality German speech dataset containing approximately 40 hours of transcribed recordings (38,000+ files) by Thorsten Müller, a single native male speaker. It combines multiple Thorsten-Voice subsets (neutral, emotional, and Hessian dialect) in original 44.1 kHz sampling rate. - [Thorsten-Voice Dataset 2023.09 Hessisch](https://mozilladatacollective.com/datasets/cmm4ddcsx00npmh07hxmlax2b): Thorsten-Voice Dataset 2023.09 (Hessisch) is a German regional dialect speech dataset based on standard German text pronounced in the Hessian dialect (“Hessisch”), primarily spoken in the southern region of the German state of Hessen. The recordings were made by Thorsten Müller and audio-optimized by Dominik Kreutz. - [Thorsten-Voice Dataset 2022.10](https://mozilladatacollective.com/datasets/cmm4b8y6100nomk07w678zb9d): Thorsten-Voice Dataset 2022.10 is a high-quality German neutral speech dataset recorded by Thorsten Müller and audio-optimized by Dominik Kreutz. It contains 12,450 phrases with more than 11 hours of clean speech audio. - [Thorsten-Voice Dataset 2021.06 Emotional](https://mozilladatacollective.com/datasets/cmm4b8f7700mgmh07cha3549n): Thorsten-Voice Dataset 2021.06 (emotional) is a German emotional speech dataset recorded by Thorsten Müller and audio-optimized by Dominik Kreutz. It contains 2,400 recordings representing eight distinct emotions. - [Daily Expressions in Highland Puebla Nahuatl](https://mozilladatacollective.com/datasets/cmm3vxr6r00bpmk07b7ayxdgx): A corpus of more than 1,000 common expressions in Highland Puebla Nahuatl, annotated for child-directedness and code-switching. 80% of the phrases are accompanied by Spanish translations. - [Cuentos en Mam leídos en voz alta](https://mozilladatacollective.com/datasets/cmm3oy19w004umh071icuxao8): Una colección de cuentos (audio y texto) en la lengua Mam. 40 cuentos, un total de 1 hora 23 minutos de audio con 958 oraciones (7,441 palabras) de texto, del Currículo Nacional Base de Guatemala. - [Cuentos en Kʼicheʼ leídos en voz alta](https://mozilladatacollective.com/datasets/cmm3j5gyb00l9nt073masycl3): Una colección de cuentos (audio y texto) en la lengua Kʼicheʼ. 1 hora 51 minutos de audio con 726 oraciones (8,283 palabras) de texto, del Currículo Nacional Base de Guatemala. - [Finance Sentences - North American Spanish](https://mozilladatacollective.com/datasets/cmm3gntdp00jymn07zjetwkyy): This is a public domain corpus of North American Spanish sentences in the finance domain. The corpus was collected in the second half of 2023 in aid of the Mozilla Common Voice project. - [CorCenCC: Corpws Cenedlaethol Cymraeg Cyfoes](https://mozilladatacollective.com/datasets/cmm3gotrz00j8nt07bsjz2znh): The CorCenCC corpus contains over 11 million words (circa 14.4m tokens) from written, spoken and electronic (online, digital texts) Welsh language sources, taken from a range of genres, language varieties (regional and social) and contexts. The contributors to CorCenCC are representative of the over half a million Welsh speakers in the country. - [Thorsten-Voice Dataset 2021.02](https://mozilladatacollective.com/datasets/cmm2fwi490024nt0779vv9o0d): Thorsten-Voice Dataset 2021.02 is a high-quality German neutral speech dataset recorded by Thorsten Müller and audio-optimized by Dominik Kreutz. It contains 22,668 phrases with more than 23 hours of clean speech audio. - [Persian VOA Corpus 2003-2008](https://mozilladatacollective.com/datasets/cmm27p8oy00auo3073kg3f1gj): VOA news articles in Persian (Farsi) from 2003 to 2008. Each entry begins with the original URL, the date of publication, and the headline, in the following format: # File: www... - [Polish Public Domain 20th Century Literature Text Corpus](https://mozilladatacollective.com/datasets/cmm0nm2ua000eo007qs4r3m8q): This corpus contains a curated collection of 54 iconic Polish literary works, including major novels, sprawling multi-volume historical epics, and documentary prose from the late 19th and early 20th centuries. The dataset features the complete canonical works of literary titans such as Władysław Reymont, Stefan Żeromski, Henryk Sienkiewicz, Bolesław Prus, Józef Ignacy Kraszewski, Eliza Orzeszkowa, Tadeusz Dołęga-Mostowicz, and Zofia Nałkowska. - [GeoLogicQA: An LLM Benchmark for Logical Reasoning in Georgian](https://mozilladatacollective.com/datasets/cmm0n37lm000dnq07vpctdtc9): GeoLogicQA is a manually-curated logical and inferential reasoning dataset for the Georgian language (a Kartvelian language). Designed to evaluate deep language understanding, the dataset bypasses simple pattern recognition in favor of multi-step deduction, reading comprehension, and arithmetic problem-solving. - [Tatar Folklore Text Corpus](https://mozilladatacollective.com/datasets/cml8gixh60087o407lfgoumgu): This corpus contains a curated collection of Tatar folklore texts, including fairy tales, proverbs, short songs (quatrains), and legends. To ensure linguistic relevance for modern NLP tasks, the content was selected from 5 volumes of the 13-volume Tatar Folklore series, explicitly excluding archaic genres (such as Dastannar) to focus on contemporary language. - [Bojonegoro Javanese TTS](https://mozilladatacollective.com/datasets/cmltnzkug0012mh07k1obic7v): The Bojonegoro Javanese TTS is a speech synthesis dataset containing more than 8 hours of audio recordings, recorded by native speakers of the Javanese language, specifically the Bojonegoro dialect from East Java, Indonesia. This dataset represents the Bojonegoro dialect of Javanese as well as variations of the Aneman dialect. - [ATLAS Cross-Lingual Transfer Matrix](https://mozilladatacollective.com/datasets/cmlth9lrp000ams07yjdscjgu): This matrix is helpful for determining what languages to train a language model with. Given a Target Language, we hope to optimize the performance for (shown as rows), we estimate how beneficial it is to train with each Source Language (shown in the columns). - [Trabajo de Campo - Huave](https://mozilladatacollective.com/datasets/cmll3ucql00mfmn07x1oeiygu): Un corpus de audio anotado de la región de San Mateo del Mar, Oaxaca, una lengua de comunidades originarias de México. El corpus contiene monólogos y diálogos de la comunidad de San Mateo del Mar. - [Kyrgyz Folklore Text Corpus](https://mozilladatacollective.com/datasets/cmlqoukmi000hnr07cprdmxsc): This corpus contains a curated collection of Kyrgyz folklore texts, including fairy tales, magical tales, tales of everyday life, proverbs, sayings, and aphorisms. The content was digitized from 5 academic volumes published in Bishkek between 2016 and 2017, sourced from the electronic collections of the Central Scientific Library of the National Academy of Sciences of the Kyrgyz Republic. - [Finweb-Edu-Chinese-v2.2](https://mozilladatacollective.com/datasets/cmlqmsm0d00pinx072m0mm3aw): Fineweb-Edu-Chinese v2.2 is the updated Fineweb-derived dataset of refined Chinese educational web content. It enhances content quality and expands education-stage coverage, fitting education-focused LLM training & educational AI tools. - [World Factbook (JSON)](https://mozilladatacollective.com/datasets/cmll4w0ln00nemn07axe7mj09): This dataset contains the full text of the CIA World Factbook converted into machine-readable JSON. It covers over 260 world entities, organized hierarchically by region (e.g., Africa, Europe). - [Manggarai Language for NLP](https://mozilladatacollective.com/datasets/cmll54ryz00ljl60764ivc60m): This dataset is a specialized linguistic collection designed to support the development of computational resources for Manggarai, a low-resource Austronesian language spoken primarily on the island of Flores, Indonesia. The dataset bridges the gap between written and spoken language by providing a synchronized collection of textual prompts and their corresponding high-quality audio recordings. - [Eastern Balochi Literature Corpus](https://mozilladatacollective.com/datasets/cmll4tg5200lcl6070hpil2b9): The Eastern Balochi Literature Corpus by Balochi Academy is a curated and cleaned collection of literary texts written in Eastern Balochi. The corpus represents a wide range of genres including poetry, folklore, novels, short stories, translations, and cultural writings. - [ABC-Draco](https://mozilladatacollective.com/datasets/cmll4i2n400l2l607uze0glc2): ABC-Draco is a GLTF Draco conversion of the NYU ABC-Dataset. The collection contains 751,407 dense point clouds with normals; using lossy Draco compression, the total archive size of 50GB is approximately two orders of magnitude smaller than the original, which makes this version useful in more resource-constrained environments. - [Gojri Literature Corpus](https://mozilladatacollective.com/datasets/cmlgz1vdt001omg07rv25gwnu): The Gojri Literature Corpus ontains approximately 60,821 tokens of Gojri (Gujari) text drawn from poetry, short stories, narrative prose, and question–answer literary books. It reflects creative writing and traditional cultural expression, including social themes, folklore, and community knowledge, and supports linguistic research, NLP tasks (e.g., text analysis and language modeling), and the documentation and preservation of Gojri language and literary heritage. - [Zacatlán Tepetzintla Nahuatl Transcriptions](https://mozilladatacollective.com/datasets/cmlct0jzu01s4nv07023lv3m3): This corpus contains the most up-to-date version of the ongoing transcription effort corresponding to the "Zacatlan Tepetzintla Nahuatl Audio Corpus", also available on MDC. At present, approximately 13 hours of audio have been transcribed using the Transcriber software. - [Khowar Literature Corpus by FLI](https://mozilladatacollective.com/datasets/cmlgys7fj001knx077yktc0pd): The Khowar Literature Corpus by FLI is a curated multi-genre textual dataset consisting of 12 UTF-8 encoded text files with a total of 108K tokens. It includes literary works, poetry, folklore narratives, magazine editions, translated books, articles, reports, and official documents, and supports corpus linguistics, low-resource NLP, digital humanities research, and the preservation of Khowar linguistic and cultural heritage. - [Khowar Word List](https://mozilladatacollective.com/datasets/cmlgxqdl80019mg07p0197u76): The Khowar Word List and Alphabet Corpus is a curated lexical dataset of 22K tokens, organized into two UTF-8 encoded text files. It includes the full Khowar script (letters and language-specific characters) along with an extensive word list, and supports lexicography, morphological analysis, language documentation, and low-resource NLP tasks for Khowar. - [Kohistani Shina Word List](https://mozilladatacollective.com/datasets/cmlgxhwha0015mg07ufma156h): The Kohistani Shina Dictionary and Word List Corpus is a curated lexical dataset consisting of a single UTF-8 encoded text file with 154K tokens. It contains Kohistani Shina vocabulary entries with definitions, grammatical notes, and cross-linguistic references, primarily mapping Kohistani Shina words to Urdu. - [Brahui Research Work Corpus](https://mozilladatacollective.com/datasets/cmlgx1kdz000nmg0744z7bbxl): This corpus consists of research works, academic theses, and scholarly papers, comprising approximately 185,000 tokens. It covers a range of academic topics and formal registers, reflecting standardized writing practices and disciplinary conventions. - [Talar (تلار) Barahui Magazine Corpus](https://mozilladatacollective.com/datasets/cmlgx0s7j000jmg07pr4o635t): The corpus consists of approximately 150,000 words collected from Talar, a monthly Brahui-language magazine. The corpus includes a range of written genres such as editorials, essays, fiction, poetry, and socio-cultural commentary, reflecting contemporary Brahui usage. - [Western Balochi Literature Cropus](https://mozilladatacollective.com/datasets/cmlgv2ucp0005nx07wwocbpux): This dataset is a curated literary corpus of General/Western Balochi (Rakhshani) prepared by the Balochistan Educational and Cultural Organization (BECO), bringing together digitized UTF-8 texts across genres such as poetry, creative literature, folklore-based writing, research articles, academic theses, translations, and other written materials. It reflects authentic Balochi usage from traditional and modern sources and is intended to support language documentation, linguistic research, digital humanities, and NLP development for an under-resourced Iranian language. - [NAWA-E-WATAN Balochi Newspaper Corpus](https://mozilladatacollective.com/datasets/cmlgrqom000jrnx07zywfpblb): The NAWA-E-WATAN Balochi Newspaper Corpus is a large-scale collection of contemporary Balochi journalistic text comprising approximately ~1.02 million tokens. The corpus reflects modern written Balochi as used in daily news reporting, political coverage, social issues, editorials, and public discourse. - [Gawri (گاؤری) Magazine Corpus](https://mozilladatacollective.com/datasets/cmlgmqqul009lny07rhsa7aey): The Gawri (گاؤری) Magazine Corpus is a curated collection of monthly Gawri (گاؤری) magazine text drawn from a periodical magazine, totaling approximately 67,724 tokens. It reflects contemporary community writing across recurring magazine sections and provides a natural sample of edited, publishable Gawri prose and poetyr. - [TTS Javanese - Ngapak Dialect](https://mozilladatacollective.com/datasets/cmlgmf58l0096nx07tahttd6y): This dataset captures the vibrant and dynamic linguistic variety found along the North Coast (Pantura) of Central Java Province, Indonesia. Unlike the inland varieties of Javanese which are heavily stratified, this dialect reflects the spirit of coastal communities. - [Jember Javanese Spontaneous Speech Corpus](https://mozilladatacollective.com/datasets/cmlgm5a94008kny07nz2intus): The Jember Javanese Spontaneous Speech Corpus is a spoken dataset of approximately 10 hours of audio collected from native Javanese speakers in Jember Regency, East Java, Indonesia. The corpus represents the Jember dialect of Javanese as well as the Pandhalungan variety. - [Zacatlán Tepetzintla Nahuatl Audio](https://mozilladatacollective.com/datasets/cmlcqxjwl01t8mm07wz7c08bz): This corpus contains 578 audio recordings, comprising 114 hours, from 31 speakers of Zacatlán-Ahuacatlán-Tepetzintla Nahuatl (alternatively Nahuatl de la Sierra Oeste de Puebla "Western Sierra Puebla Nahuatl", Glottocode zaca1241) from the municipalities of Zacatlán and Tepetzintla. The transcription of this audio is an ongoing effort, with periodic releases of transcriptions in a separate MDC dataset (https://datacollective.mozillafoundation.org/datasets/cmlct0jzu01s4nv07023lv3m3), "Zacatlan Tepetzintla Nahuatl Transcriptions.". - [openbook.gr v1.0](https://mozilladatacollective.com/datasets/cmkwuye570028mo074bppsmiw): This dataset provides a comprehensive Corpus of Greek Digital Books systematically aggregated from OpenBook.gr. Since its inception in 2010, the OpenBook platform has functioned as a central hub for the Greek open-access movement. - [Greek PhD Theses Corpus v1.0](https://mozilladatacollective.com/datasets/cmkwvpu7s0032mo07jpk20pj1): The Greek PhD Theses Corpus is a large-scale, AI-ready text dataset consisting of 55,423 Greek doctoral dissertations produced between 1975 and 2025. It represents the most comprehensive and technically homogenized collection of Greek PhD-level academic writing assembled to date. - [TTS Sasak Language](https://mozilladatacollective.com/datasets/cml9i39tx018so40750re64v6): This dataset is compiled based on a series of questions related to daily life, such as routine activities, social interactions, personal experiences, and common customs of the Sasak people. Answers to these questions were delivered both orally and in writing using the informal Sasak language, as used in everyday communication. - [Betawi TTS of Cultural Language (BEKAL)](https://mozilladatacollective.com/datasets/cml9hmuis017yo407k0p4i0t4): Betawi TTS of Cultural Language (BEKAL) is a dataset that represents the Betawi language as a living and evolving language within the urban context of Jakarta, reflecting both traditional forms and modern variations that emerge in everyday communicative practices. This dataset can be utilized for linguistic research, cultural documentation, urban sociolinguistic studies, and the development of language technologies based on regional languages with Indonesian code-mixing. - [TTS-Tolaki](https://mozilladatacollective.com/datasets/cml6ywgg0007xmn07ppq469gt): The Tolaki language is the predominant language in Southeast Sulawesi Province, Indonesia. It is spoken across the regencies of Kolaka, North Kolaka, Konawe, North Konawe, South Konawe, and East Konawe, as well as in several areas within the city of Kendari. - [Mandar Spontaneous Speech](https://mozilladatacollective.com/datasets/cml5e30pd00eskr072e6a4rrh): Mandar Spontaneous Speech is a representative dataset of the Mandar language and contains a variety of dialects, particularly those used in Majene and Polewali Mandar. It also includes Mandar–Indonesian code-switching varieties that reflect both traditional and modern forms of speech. - [TTS Central Javanese](https://mozilladatacollective.com/datasets/cml5bn4k900aame07u0rwidcg): The Central Javanese dialect is a variety of the Javanese language that is politically regarded as “Standard Javanese.” It is generally spoken in Semarang, Solo, and Yogyakarta, Indonesia. Central Java and Yogyakarta are considered the historical centers of the Mataram Kingdom in Java. - [TTS Javanese-Lumajang Dialect](https://mozilladatacollective.com/datasets/cml5bgysg00bhkr07g23kewke): The Lumajang dialect is a unique variation of the Javanese language spoken across Lumajang Regency in East Java, Indonesia. Locally known as “Arekan”, this dialect emerged from a cultural blend of Javanese and Madurese influences. - [Bamun-French Parallel Corpus 1.1](https://mozilladatacollective.com/datasets/cml16bz8v008pnt07scfh66p5): This dataset is the second iteration of the 'Bamun–French parallel corpus' that was initially published on the Mozilla Data Collective platform. It is a parallel corpus of Bamun (Shupament) and French texts. - [TTS Muna Dataset](https://mozilladatacollective.com/datasets/cml161p08007ynt07iilell66): The Muna language, locally known as Wuna, is an Austronesian language spoken by approximately 300,000 people in Southeast Sulawesi, primarily on Muna Island, large parts of West Muna, and the western coastal areas of Buton and Central Buton, Indonesia. According to data from Ethnologue and the Kemendikbud Language Map, this language has an extensive distribution but faces sustainability challenges, currently holding a Threatened status as it is increasingly less common for it to be actively passed down to the youngest generation. - [Hawrami Kurdish TTS dataset 1.0](https://mozilladatacollective.com/datasets/cml0wtouz026wno075aq7v201): This dataset contains high-quality single-speaker audio recordings in Hawrami Kurdish (Hewrami, ISO 639-3:hac), also known as the Gorani language, intended for building Text-to-Speech (TTS) and Automatic Speech recognition (ASR) systems. The dataset comprises 5 hours and 15 minutes of aligned audio and text data. - [Common Voice 7.0 - Single Word Target Segment](https://mozilladatacollective.com/datasets/cmkzhp64p00wlno07elrmt20y): This dataset contains the numbers 0 to 9 and the words "yes" and "no" in 34 languages. It contains 84 validated hours of speech and 142 hours in total. - [Sermon-Malaysian-English](https://mozilladatacollective.com/datasets/cmkq08sqg005gmg07hshr5o5n): This is a 7 minute mp4 file of a sermon. There is also a txt file of the original sermon text, and an srt file generated with Descript, that adds time stamps at 5 second intervals. - [Reading Recommendations List](https://mozilladatacollective.com/datasets/cmkpzexin005gnq07f2ujqtar): A reading recommendations list of 249 fiction (fantasy, literary, mystery) books read by a single individual in her early-thirties. The data was pulled from the individual's StoryGraph account. - [chinese-cosmopedia](https://mozilladatacollective.com/datasets/cmkpon5es03j2mg07j2kppzn7): A large-scale high-quality Chinese text dataset developed by OpenCSG, containing ~15 million entries (≈60B tokens) covering multi-domain content (encyclopedia, education, etc.). Cleaned and deduplicated to remove low-quality content, it is optimized for large language model pretraining, text generation, and other Chinese NLP downstream tasks, compatible with mainstream toolchains (Hugging Face Datasets, PyTorch). - [smoltalk-chinese](https://mozilladatacollective.com/datasets/cmkplyrln03jknw07u3iy35rm): SmolTalk-Chinese is a high-quality multi-task dataset for Chinese conversational scenarios, covering 19 task types (e.g., advice-seeking, code generation, daily chat). The full dataset is accessible via www.opencsg.com. - [Common Voice v24 English - en-AU subset for Everything Open 2026](https://mozilladatacollective.com/datasets/cmko7havo02f5nw07rbwwhowe): This is a subset of Common Voice v24 English filtered for Australian-clustered accents. It is designed to be used in conjunction with the hands-on Tutorial delivered at Everything Open 2026 in Canberra, Australia. - [Informes de Actividades InfoCDMX (Ponencia Laura Enríquez)](https://mozilladatacollective.com/datasets/cmkmnb80w0177nw07iswxo3re): Descripción: El conjunto de datos se compone de los informes anuales de actividades y resultados del InfoCDMX, los cuales documentan el desempeño del organismo garante, el pleno y las ponencias de las personas comisionadas. Estos documentos son instrumentos fundamentales de rendición de cuentas que detallan la gestión institucional, la actividad cuasi-jurisdiccional (resolución de recursos de revisión y denuncias), y las acciones de promoción de derechos. - [Ficha de Documentación de Datos: Resoluciones InfoNL (Ponencia F. Guajardo)](https://mozilladatacollective.com/datasets/cmkmn6o6e0165mg07s9h542vt): Este conjunto de datos documenta la actividad resolutiva de la ponencia del Consejero Francisco Guajardo Martínez dentro del órgano garante de transparencia de Nuevo León. Cubre un periodo significativo de gestión (identificado preliminarmente entre 2018 y 2025), reflejando las disputas entre ciudadanos (solicitantes de información) y sujetos obligados (gobierno). - [Compar:IA conversations](https://mozilladatacollective.com/datasets/cmklbq9qt006mnw077t9lqh89): The compar:IA dataset is a large-scale collection of real user conversations generated on the compar:IA platform, a public conversational AI comparison service developed within the French Ministry of Culture. The platform allows users to interact with two conversational AI models side by side and compare their answers in a blind setting. - [RFE/RL Tatar-Bashkir News Text Corpus](https://mozilladatacollective.com/datasets/cmkh3nwms0052nv07997y38zx): This dataset serves as a comprehensive longitudinal news corpus for the Tatar and Bashkir languages, sourced from Azatliq Radiosi (azatliq.org), the Tatar-Bashkir service of Radio Free Europe/Radio Liberty (RFE/RL). Spanning from December 2001 to December 2025, the corpus contains over 105,000 unique articles. - [DhoNam: Dholuo Speech dataset](https://mozilladatacollective.com/datasets/cmjepxo6t08nmmk07iauvua6v): DhoNam: Dholuo Speech dataset is a speech corpus designed to supercharge Automatic Speech Recognition (ASR) and other speech technologies for Dholuo, one of Kenya’s major indigenous languages. This dataset contains native-speaker audio recordings collected through a platform where users read aloud a displayed sentence. - [English–Punjabi (Shahmukhi) Parallel Sentences Corpus (Mediamen Archives)](https://mozilladatacollective.com/datasets/cmkh9rso90076nv076jwxgjv3): This parallel sentences corpus containing 30,405 aligned sentence pairs with a total of approximately 0.62 million tokens, curated from the archival materials of Mediamen (Advertising Agency). The sentences were professionally translated from English into Punjabi (Shahmukhi) and are intended to support machine translation, linguistic research, and Punjabi language technology development, particularly for real-world and contemporary language use. - [HESEIA Sentence Bias Dataset](https://mozilladatacollective.com/datasets/cmkh3vrob005pmj07vgm55iha): This repository contains a dataset collected during the teacher training course HESEIA Sentence Bias (Tools for Exploring Biases and Artificial Intelligence). organized by Vía Libre, the Ministry of Education, and FAMAF-UNC. - [Effect AI Scripted Speech 1.0 - English](https://mozilladatacollective.com/datasets/cmkfm9fbl00nto0070sdcrak2): An open dataset of scripted English sentences, recorded by speakers using the Effect AI platform for the development, training, and evaluation of speech recognition and language technologies. - [DataTrust Africa: Speech Corpus of Public Radio Recordings from Northern Uganda](https://mozilladatacollective.com/datasets/cmkfm6xtw00k2nv07oakesnix): This is an open-access corpus of short clips of public radio content from Mega 100 FM, Q FM, Radio Pacis and Radio Rupiny in Northern Uganda. As of now, the online corpus has over 350 clips of recordings in English. - [Corpus of Panjebar Semangat Javanese-Language Magazine](https://mozilladatacollective.com/datasets/cmk719jvs02z1nt07tjswt52s): This dataset is a TXT-format collection compiled from three years of popular articles published in the Javanese-language weekly magazine Panjebar Semangat. It compiles widely read, non-academic Javanese texts reflecting contemporary themes and language use. - [SI-NLI](https://mozilladatacollective.com/datasets/cmk70opha02yxnt072patcjnd): SI-NLI (Slovene Natural Language Inference Dataset) contains 5,937 human-created Slovene sentence pairs (premise and hypothesis) that are manually labeled with the labels "entailment", "contradiction", and "neutral". We created the dataset using sentences that appear in the Slovenian reference corpus ccKres (http://hdl.handle.net/11356/1034). - [Sindh Line Publishers](https://mozilladatacollective.com/datasets/cmk1bfsez3htsmk07xfzaf9oe): The corpus contains 1.029 million tokens from the Sindh Line a Sindhi Newspaper published from the year 2024-2025. The text consists of the complete newspaper content including headlines, editorials, finance news and advertisements. - [Vallader Newspaper Corpus](https://mozilladatacollective.com/datasets/cmk40plhc001lmd07kwixjls0): 6.2 million tokens in the Vallader variety of Romansh from the daily newspaper ”La Quotidiana”. - [Hussain Faizy Indus Kohistani Corpus](https://mozilladatacollective.com/datasets/cmixehxof005dnr07yp80kluh): The Indus Kohistani corpus contains around 500k tokens of folktales, stories, poetry, biographies, and conversational texts, all transcribed with a consistent community orthography. Reviewed by native speakers, the corpus offers a representative snapshot of the language’s vocabulary and grammar for linguistic and computational research. - [Multilingual Humanitarian Response Eval (MHRE)](https://mozilladatacollective.com/datasets/cmixmkzhk003vnw07a7d2pgns): This multilingual humanitarian dataset contains 655 annotated datapoints evaluating AI chatbot safety and quality in migration and asylum scenarios across four language pairs (English–Farsi (Iranian Persian), Arabic, Kurdish (Sorani), Pashto). Built from 120 expert prompts, it includes outputs from GPT-4o, Gemini 2.5 Flash, and Mistral Small. - [KyrgyzLLM-Bench: Kyrgyz LLM Evaluation Dataset](https://mozilladatacollective.com/datasets/cmj0e4p03003mnu077qbbn17z): KyrgyzLLM-Bench is a comprehensive suite purpose-built to evaluate LLMs’ deep understanding and reasoning in Kyrgyz. It combines natively authored benchmarks with carefully translated and post-edited international tasks to provide broad and culturally grounded coverage. - [Flemishguy 1.0](https://mozilladatacollective.com/datasets/cmiupbboq01i4mf071clmd43n): Text to speech dataset for Dutch, male speaker, approximately 1 hour of read speech. - [Central Kurdish TTS dataset 1.0](https://mozilladatacollective.com/datasets/cmj77njd701ljmb07m97pw1p3): This dataset contains high-quality single-speaker audio recordings in Central Kurdish (ckb), intended for building Text-to-Speech (TTS) and Automatic Speech recognition (ASR) systems. The dataset comprises 2 hours and 18 minutes of aligned audio and text data. - [AI on the Frontline: Evaluating Large Language Models in Real-World Conflict Resolution](https://mozilladatacollective.com/datasets/cmj8mjbas02lmmk07glipgg79): Findings of an experimental evaluation to assess how leading free-access LLMs perform when asked to respond to realistic conflict resolution scenarios. - [Improving AI Conflict Resolution Capacities: A Prompts-Based Evaluation](https://mozilladatacollective.com/datasets/cmj8mkwot02ltmb07a8pyg3ne): Findings of a follow-up study assessing how leading free-access LLMs perform when adding instructions directing them to apply basic conflict-resolution practices. - [Darkman 1.0](https://mozilladatacollective.com/datasets/cmiuparkl01fpnv077wvg3hhm): Text to speech dataset for Polish, male speaker, approximately 2 hours of read speech. - [Denis 1.0](https://mozilladatacollective.com/datasets/cmiup9seu01flnv076fexaqp9): Text to speech dataset for Russian, male speaker, approximately 2 hours of read speech. - [Future-proofing Gbagyi: A community centered approach](https://mozilladatacollective.com/datasets/cmit3jhj000nlnv073a6zduck): This dataset comprises 360 audio recordings of the Gbagyi language, comprising approximately 7 hours 52 minutes of speech data, with paired transcripts. - [Gosia 1.0](https://mozilladatacollective.com/datasets/cmiugoh1u01c9nv07x7i2t25c): Text to speech dataset for Polish, female speaker, approximately 2 hours of read speech. - [Cadu 1.0](https://mozilladatacollective.com/datasets/cmiuimd7e01fimf072njn3mbb): Text to speech dataset for Brazilian Portuguese, male speaker, approximately 1.5 hours of read speech. - [Pim 1.0](https://mozilladatacollective.com/datasets/cmiugo0h501exmf073drr56rr): Text to speech dataset for Dutch, male speaker, approximately 2 hours of read speech. - [Jeff 1.0](https://mozilladatacollective.com/datasets/cmiupaf3e01i0mf07qkqxbzzx): Text to speech dataset for Brazilian Portuguese, male speaker, approximately 1.5 hours of read speech. - [Nathalie 1.0](https://mozilladatacollective.com/datasets/cmiugnlxi01c5nv07yt3kxwhx): Text to speech dataset for Dutch, female speaker, approximately 1 hour of read speech. - [Lili 1.0](https://mozilladatacollective.com/datasets/cmiup5wqw01hlmf074qy07b80): Text to speech dataset for Slovak, female speaker, approximately 2 hours of read speech. - [Chitwan 1.0](https://mozilladatacollective.com/datasets/cmiugmupp01etmf07h89hfpir): Text to speech dataset for Nepali, male speaker, approximately 1 hour of read speech. - [Dmitri 1.0](https://mozilladatacollective.com/datasets/cmiup9hrx01hsmf074vi5iqgz): Text to speech dataset for Russian, male speaker, approximately 2 hours of read speech. - [Faber 1.0](https://mozilladatacollective.com/datasets/cmiupazq801ftnv079bo1zu4h): Text to speech dataset for Brazilian Portuguese, male speaker, approximately 1.5 hours of read speech. - [Mihai 1.0](https://mozilladatacollective.com/datasets/cmiupa1t801hwmf07xcawi3ve): Text to speech dataset for Romanian, male speaker, approximately 2 hours of read speech. - [Ronnie 1.0](https://mozilladatacollective.com/datasets/cmiujt8r401d2nv07avt6446s): Text to speech dataset for Dutch, male speaker, approximately 2 hours of read speech. - [Tugão 1.0](https://mozilladatacollective.com/datasets/cmiui8v3101cpnv074qwggdsj): Text to speech dataset for Portuguese, male speaker, approximately 1.5 hours of read speech. - [Aim Foundation Dari Literature Corpus](https://mozilladatacollective.com/datasets/cmiq9ulwl004ao207pglxpszv): This corpus is a collection of more than seven hundred thousand tokens of Dari language. The corpus contains work of literature including poems, stories, novels, fictional and non-fictional and different articles. - [Mozilla Common Voice Spontaneous Speech ASR Shared Task Test Data](https://mozilladatacollective.com/datasets/cminc35no007no707hql26lzk): A bundle of the held-out test data for the Mozilla Common Voice Spontaneous Speech ASR shared task. - [TidyVoiceX_ASV](https://mozilladatacollective.com/datasets/cmihtsewu023so207xot1iqqw): This dataset is designed for speaker verification using the Mozilla Common Voice corpus across 40 languages. It includes approximately 5,000 speakers who each have recordings in more than one language. - [Keblagh-e-Azergi Hazargi literature corpus](https://mozilladatacollective.com/datasets/cmiq9mdsg004sns07lrp0rypp): This corpus is a collection of more than one hundred thousand tokens of Hazargi language. The corpus contains work of literature, poems, folk and short stories and dramas. - [Sindh Sujag Newspaper Corpus](https://mozilladatacollective.com/datasets/cmiqbnn0x005uns07c3evoh0k): The corpus contains approximately 1.2 million tokens from the Sindh Sujag Newspaper Agency published between 2024 and 2025. It includes complete newspaper content such as headlines, editorials, finance news, and advertisements. - [Rana Printers Urdu Literature Corpus](https://mozilladatacollective.com/datasets/cmiq9pmmx0050ns07s0vgvobm): This corpus comprises 1.68 million tokens of high-quality Urdu text collected over the past decade through Rana Printers. It includes a diverse range of literary genres such as stories, short stories, novels, fiction, non-fiction, poetry, and historical works. - [Kaleem Art Press Urdu Literature Corpus](https://mozilladatacollective.com/datasets/cmiq9o5rn004wns07zry0skak): This corpus is a collection of 1.44 million tokens of Urdu language . The data was produced under the Kaleem Art Press over the last fifteen years . - [Kaleem Art Press Saraiki Literature Corpus](https://mozilladatacollective.com/datasets/cmiq9nlul0042o207d0gnxv5p): This corpus contains approximately one million tokens of Saraiki text curated over the past ten years by Kaleem Art Press. It features a wide range of literary genres, including stories, short stories, novels, fiction, non-fiction, travelogues, poetry, biographies, and historical writings. - [Atyap Afwan_: Preserving Tyap Through Community-Driven Speech Data](https://mozilladatacollective.com/datasets/cmiodqzui00wwnx07j549ek4c): This dataset contains 98 recordings (≈1.16 hours) of everyday Tyap speech from 10 community speakers, each paired with detailed transcripts and English translations. - [Ehugbo TTS: biblical text to speech dataset in Ehugbo Language](https://mozilladatacollective.com/datasets/cmihqro9h0238o207fgg5cmf6): This dataset contains audio recordings of Bible verses in Ehugbo, a dialect of Igbo (a Niger-Congo language spoken in Nigeria). It contains 312 audio recordings of biblical text-to-speech data comprising 1 hour and 30 seconds of speech data. - [Everyday Interactions in Ibọnọ and Obolo Languages](https://mozilladatacollective.com/datasets/cmin6i6da001so707uwggcn7a): This dataset offers 11.3 hours of natural everyday speech in Ibọnọ and Obolo, captured from 20 adult speakers across 120 recordings, each paired with a clean transcript and metadata. - [Anjuman-e-Katib Farsi/Persian Literature Corpus](https://mozilladatacollective.com/datasets/cmiq9p8e70046o2073yogyn73): This corpus is a collection of more than one million tokens of Farsi/Persian language. The corpus contains work of literature including novels, fictional and non-fictional, poems and many more. - [Documenting Ekpeye Folktales and Preserving Cultural Heritage](https://mozilladatacollective.com/datasets/cmiohwz2t011hnx07urwsx55i): This dataset presents 21 video-recorded Ekpeye folktales (1h28m) narrated by two community elders, each paired with transcripts and English translations that include narrative summaries. It offers a rich multimodal resource for speech, video, storytelling, and cultural heritage research, as well as training multilingual and multimodal AI systems. - [Joe 1.0](https://mozilladatacollective.com/datasets/cmid3cyuy00gnnv07mo54373n): Text to speech dataset for English, male speaker, approximately 1 hour of read speech. - [Rumantsch Grischun Newspaper Corpus](https://mozilladatacollective.com/datasets/cmifzczuo00w6o207wgmqb97m): 6.1 million tokens in the Rumantsch Grischun variety of Romansh from the daily newspaper “La Quotidiana”. - [Podcast Homostoria (Indonesia)](https://mozilladatacollective.com/datasets/cmiepnyu1001jo207hzslvb5m): This dataset features discussions on modern media—including film, podcasts, and social media—and its connection to local customs and traditions. The conversations are primarily in Indonesian, with frequent code-switching between English and Javanese. - [Imre 1.0](https://mozilladatacollective.com/datasets/cmid4wbrc00hvnv07e448rwv3): Text to speech dataset for Hungarian, male speaker, approximately 1.5 hours of read speech. - [Putèr Newspaper Corpus](https://mozilladatacollective.com/datasets/cmig0fljr00yrmd075gqlj2qo): 1.3 million tokens in the Putèr variety of Romansh from the daily newspaper “La Quotidiana”. - [Sutsilvan Newspaper Corpus](https://mozilladatacollective.com/datasets/cmig0f00m00ynmd07kqzscviy): 1.3 million tokens in the Sutsilvan variety of Romansh from the daily newspaper “La Quotidiana”. - [KyrgyzNER: Human-Annotated NER Dataset for Kyrgyz](https://mozilladatacollective.com/datasets/cmihefmde01vxmd07djfchz51): KyrgyzNER is the first manually annotated Named-Entity Recognition (NER) dataset for the Kyrgyz language. It consists of 1,499 news articles (10,900 sentences, 140k tokens) from the 24.kg news portal, annotated with 39,075 entity mentions across 27 classes. - [Kerstin 1.0](https://mozilladatacollective.com/datasets/cmi7mgbam000bnx074097g2yg): Text to speech dataset for German, female speaker, approximately 2 hours of read speech. - [Berta 1.0](https://mozilladatacollective.com/datasets/cmid4vszk00hrnv07ljzu2xsa): Text to speech dataset for Hungarian, female speaker, approximately 1 hour of read speech. - [Dave 1.0](https://mozilladatacollective.com/datasets/cmid4ir9100bgnu07du03svuc): Text to speech dataset for Spanish, male speaker, approximately 1.5 hours of read speech. - [Speech Data Collection for The Nupe Language](https://mozilladatacollective.com/datasets/cmihkoth10246md07dtojxehg): This dataset contains audio recordings of the Nupe language. It features 1,583 audio recordings comprising 2 hours, 40 minutes, and 32 seconds of speech data, with paired transcripts. - [Kathleen 1.0](https://mozilladatacollective.com/datasets/cmid4ekjn00b8nu07xonhj0x8): Text to speech dataset for English, female speaker, approximately 1 hour of read speech. - [Sursilvan Newspaper Corpus](https://mozilladatacollective.com/datasets/cmig0eqj700x6o207tp8pb5za): 14.6 million tokens in the Sursilvan variety of Romansh from the daily newspaper “La Quotidiana”. - [Anna 1.0](https://mozilladatacollective.com/datasets/cmid4jner00hnnv072tkwlbps): Text to speech dataset for Hungarian, female speaker, approximately 1.5 hours of read speech. - [Kaleem Magazine Urdu Corpus](https://mozilladatacollective.com/datasets/cmi39jopd08k5mn0745umscb3): This corpus is a collection of around 1.4 million tokens of Urdu language. The data was extracted from the archives of a famous Urdu magazine "Kaleem" published weekly from last 30 years. - [Dimitar 1.0](https://mozilladatacollective.com/datasets/cmhpahaib00d8mk07ely2m8wh): Text to speech dataset for Bulgarian, male speaker, approximately 2 hours of read speech. - [Mediamen Punjabi Literature Corpus](https://mozilladatacollective.com/datasets/cmhrlxa1100h3mn07l6tpvqwz): This corpus is a collection of one million tokens of Western Punjabi language. The data was produced under the Mediamen publishing agency over the last ten years. - [Saraiki Quarterly Magazine Wasson Wehray Corpus](https://mozilladatacollective.com/datasets/cmhp3uclf00ammk07nnd1m0bj): This corpus contains 11,79,200 tokens, from a Saraiki Quarterly Magazine "Wasson Wehray". - [Punjabi Literature Corpus](https://mozilladatacollective.com/datasets/cmhp42gmu00aumk07w7k22tbp): This corpus contains 10,39,430 tokens of Punjabi Shahmukhi script. - [FUB-Narratives](https://mozilladatacollective.com/datasets/cmhvzlidq0326mn07hk4do3pj): This dataset contains literary texts derived from oral Fulfulde Adamawa (fub) performances. The texts are of various genres, including narratives, hymns, riddles and poems. - [Jazab Sindhi Newspaper Corpus](https://mozilladatacollective.com/datasets/cmhrus0ty00lpkx07slh9mqez): The corpus contains 1.07 million tokens from the Jazab a Sindhi Newspaper published from the year 2023-2025. The text consists of the complete newspaper content including headlines, editorials, finance news and advertisements. - [Chishti Sons Punjabi Literature Corpus](https://mozilladatacollective.com/datasets/cmi39dqwq08jnmn0739vleijb): This corpus is a collection of more than one million tokens of Western Punjabi language. The data was produced under the Chishti Sons publishing agency. - [Tamir Sindhi News Corpus](https://mozilladatacollective.com/datasets/cmhruaw9j00llkx07uxgm81a7): The corpus contains 1.1 million tokens from the Tamir Sindhi Newspaper published from the year 2022-2025. The text consists of the complete newspaper content including headlines, editorials, finance news and advertisements. - [Baloch Publishers Saraiki Literature Corpus](https://mozilladatacollective.com/datasets/cmi39h0xr08jzmn07hij4fz40): This corpus is a collection of one million tokens of Saraiki language. The data was produced under the Baloch Publishers over the last ten years. - [Speech Corpus of Armenian Question-Answer Dialogues](https://mozilladatacollective.com/datasets/cmhqr666h009gmn07fpp0egby): A collection of question-answer dialogues in Western and Eastern Armenian. - [Mozilla Common Voice Spontaneous Speech ASR Shared Task Train/Dev Data](https://mozilladatacollective.com/datasets/cmfzu8u8wa555eq8onrk334h4): This datasheet is for the bundle of Mozilla Common Voice spontaneous speech datasets to be used in the Shared Task on Spontaneous Speech. - [Saraiki Literature Corpus](https://mozilladatacollective.com/datasets/cmhnlbtx401ihmr07co7x8ptw): This contains multiple Saraiki Language books of Stories, Short Stories, Novel, Travelogue, Sentences and collection of articles. - [Urdu Literature Corpus](https://mozilladatacollective.com/datasets/cmhp3jurx00afmk07z5z4fktu): This corpus contains 16,17,074 tokens of multiple Urdu literature books published by Bismillah Graphics Publishers. - [Tetelancingo Nahuatl](https://mozilladatacollective.com/datasets/cmhkl8z2a007rnr07p9bm5kmz): Audio del Nahuatl de la Sierra Oeste de Puebla, transcrito, ortográficamente normalizado, traducido, y etiquetado. - [Podcast Hari Minggoean (Indonesia)](https://mozilladatacollective.com/datasets/cmhlvswz200fio007c7szypw8): This dataset is derived from the "Hari Minggoean" podcast, featuring over ten hours of recorded speech from a single, consistent speaker. The content, tailored for a young Indonesian audience, is presented in Indonesian (Bahasa Indonesia) characterized by code-switching with English and a discernible Javanese accent. - [Multilingual Religious Parallel Corpus (Kaleem Art Press)](https://mozilladatacollective.com/datasets/cmk1bhogs3htwmk07wo7o9p6y): This dataset is a multilingual parallel sentences corpus containing 6,465 aligned sentence units with approximately 0.98 million words, curated from Kaleem Art Press archives. It includes parallel religious text data in Arabic, Urdu, Saraiki (standard and dialectal), Punjabi (Shahmukhi), and English, supporting research in machine translation, comparative linguistics, digital humanities, and low-resource language studies. - [Balochi Academy Text Corpus](https://mozilladatacollective.com/datasets/cmk1bb49f3htomk0760xg4cb6): This corpus contains approximately 500k tokens of text from novels, poetry, articles, riddles, and proverbs, covering both literary and traditional genres. It is intended for linguistic research, NLP tasks (e.g., language modeling and text analysis), and cultural documentation. - [Mada Narratives](https://mozilladatacollective.com/datasets/cmk19uruw39urmb07alqf493e): This dataset contains 17 transcribed oral narratives in Mada (mxu), a language belonging to the Afro-Asiatic family that is spoken in Cameroon. The texts, derived from audio recordings of oral literature, reflect natural spoken discourse. - [Surmiran Newspaper Corpus](https://mozilladatacollective.com/datasets/cmjhe0xap09gamb078g9loi3q): 2.9 million tokens in the Surmiran variety of Romansh from the daily newspaper “La Quotidiana”. - [Archivo de la Comisionada María de los Ángeles Guzmán García (COTAI Nuevo León / InfoNL)](https://mozilladatacollective.com/datasets/cmjcc6g9z06c7mk07yolcdyjr): Este archivo preserva la memoria institucional y académica de la gestión de la Dra. María de los Ángeles Guzmán García como Comisionada de la Comisión de Transparencia y Acceso a la Información del Estado de Nuevo León (COTAI / INFONL) durante el periodo 2018-2025. - [Mozilla Common Voice Text Language Identification dataset](https://mozilladatacollective.com/datasets/cmj8ddapc02c8mb07l6wyr882): A dataset for text-based language identification of 19 Million sentences from over 300 languages taken from Mozilla Common Voice scripted (v23) and spontaneous (v1) speech projects.