Skip to content
lilernoYour workspace

Source transparency

Language data needs provenance, licensing, and human-readable context.

Lilerno's vocabulary pipeline records source identifiers and keeps a catalog of suitable language-data providers. Configured source families include a manual core seed, PanLex, Kaikki/Wiktionary extracts, Wikidata Lexemes, Tatoeba, and UniMorph, each with different licensing and review needs.

Configured source roles

  • Manual core seed: initial project-authored coverage and schema validation.
  • PanLex: multilingual translation candidates under its published license.
  • Kaikki/Wiktionary and Wikidata Lexemes: dictionary and structured lexeme candidates.
  • Tatoeba: sentence candidates that require per-item license tracking.
  • UniMorph: morphology candidates under source-specific open-data terms.

Publication standard

A configured source is not automatically proof that every displayed item came from it. Public resources should identify the actual record provenance, preserve required attribution, and separate imported facts from project-authored explanations. Unverified generated sentences are not treated as authoritative examples.

Content reviewed and updated 2026-07-25.