🐍 Elamur Research

Mova corpus 0.2

Open data Mova is developed and tested on. Every part says where it comes from, what was changed, its license and how Mova uses it. Nothing here restricts the purpose of use.

Download corpus-0.2.tar.gz (8.8M)

SHA-256: corpus-0.2.tar.gz.sha256. Each folder in the archive has a README with its full description.

Parts

PartWhatLicense
tales/42 books of fairy tales, fables and children's stories from Project Gutenberg, one text file per book; Gutenberg header and footer removed, text unchangedpublic domain in the USA (check your country)
tale-summaries/one-line summaries of 51,070 paragraphs of those booksApache-2.0 OR MIT
tale-questions/12,251 short questions; each answer is at most three words copied verbatim from the paragraphApache-2.0 OR MIT
treebanks/Universal Dependencies annotation: Mova silver (348 sentences) and bronze (326) sets, and UD English ESLSpok with Mova's dialect rules applied (2,320 sentences)silver CC BY 4.0; bronze Apache-2.0 OR MIT; ESLSpok CC BY-SA 4.0
math-scripts/2,680 word problems from SVAMP and GSM8K (training splits) with step-by-step world scripts that compute the gold answerproblems MIT (original sets); scripts Apache-2.0 OR MIT

How Mova uses it

The tales, summaries and questions train and test reading: Mova finds the sentence with the answer and extracts the span, with explanations. The treebanks test the English parser and its grammar seeds. The math scripts teach Mova to write the world of a word problem itself. Code and instructions: ElamurAI/mova-free, folders ai/, tests/, train/.