Mova corpus 0.2
Open data Mova is developed and tested on. Every part says where it comes from, what was changed, its license and how Mova uses it. Nothing here restricts the purpose of use.
Download corpus-0.2.tar.gz (8.8M)SHA-256: corpus-0.2.tar.gz.sha256. Each folder in the archive has a README with its full description.
Parts
| Part | What | License |
|---|---|---|
tales/ | 42 books of fairy tales, fables and children's stories from Project Gutenberg, one text file per book; Gutenberg header and footer removed, text unchanged | public domain in the USA (check your country) |
tale-summaries/ | one-line summaries of 51,070 paragraphs of those books | Apache-2.0 OR MIT |
tale-questions/ | 12,251 short questions; each answer is at most three words copied verbatim from the paragraph | Apache-2.0 OR MIT |
treebanks/ | Universal Dependencies annotation: Mova silver (348 sentences) and bronze (326) sets, and UD English ESLSpok with Mova's dialect rules applied (2,320 sentences) | silver CC BY 4.0; bronze Apache-2.0 OR MIT; ESLSpok CC BY-SA 4.0 |
math-scripts/ | 2,680 word problems from SVAMP and GSM8K (training splits) with step-by-step world scripts that compute the gold answer | problems MIT (original sets); scripts Apache-2.0 OR MIT |
How Mova uses it
The tales, summaries and questions train and test reading: Mova finds the sentence with the answer and extracts the span, with explanations. The treebanks test the English parser and its grammar seeds. The math scripts teach Mova to write the world of a word problem itself. Code and instructions: ElamurAI/mova-free, folders ai/, tests/, train/.