@3sln/donki-lle

Language-learning elements: custom elements for interactive lessons, made so a language model can build a lesson as a small HTML page and a chat interface can render it in an iframe. No build step and no dependencies: load one script from a CDN and write the tags.

<script type="module" src="https://cdn.jsdelivr.net/npm/@3sln/donki-lle@0/index.js"></script>

<donki-lle-tts lang="es-MX">¿Dónde está la biblioteca?</donki-lle-tts>
<donki-lle-stroke text="你好"></donki-lle-stroke>
ElementWhat it's for
<donki-lle-tts>A phrase that reads itself aloud, with a neural voice where the device has none
<donki-lle-stroke>Stroke-order writing practice: Latin letters, hangul, kana, hanzi, kanji
<donki-lle-deck-share>A Donki flashcard deck (or one shard of one) to add, download or copy
<donki-lle-send-results>Sends the learner's answers back into the chat, or copies them where it can't

Every element works with touch, pen and mouse, has 44 px touch targets, follows the reader's light or dark preference, and can be themed (see Theming). Each is also its own module (…/src/tts.js, …/src/stroke.js, …) for a page that wants only one.

For how to teach with these, see the language teaching guide.


<donki-lle-tts>#

<donki-lle-tts lang="fr-FR">Je voudrais un café</donki-lle-tts>
<donki-lle-tts lang="ja-JP" text="ありがとう" hide-text rate="0.8"></donki-lle-tts>

A pill-shaped button showing the phrase; tapping it speaks. Tapping again stops.

Attribute
langBCP 47 tag (es-MX), a comma-separated list, or a JSON array, in order of preference. Name the variety you mean.
textThe phrase. If absent, the element's own text is used, which is the easier form to write.
hide-textA round speaker button with no text, for listening exercises where the text would give the answer away.
pacelearner (default): a little slower than natural, with short pauses between clauses. natural: conversational speed. slow: slower, with longer pauses, for a phrase heard for the first time.
rate, pitchA speed multiplier (0.1–10) that replaces the pace's, and pitch (0–2).
enginesystem or neural to force one; normally left out.

Which voice speaks it:

  1. An installed system voice for the language.
  2. A stand-in the reader already accepted.
  3. A neural voice, run in the browser. Used when the device has no voice for the language at all: Android WebViews have no speech synthesis, and many desktops have voices for only a few languages. The download starts as soon as the element appears, except on a data-saver or slow connection, where it waits for a tap. The button shows download progress in a ring, and a tap during the download explains the wait and plays once the voice is ready.
  4. A related system voice (es-ES for es-MX), offered in a small dialog: once, for the session, or always.
  5. Otherwise, a note saying no voice is available, with a link on how to install one.

The neural voices run on sherpa-onnx, compiled to WebAssembly without a built-in voice (vendor/sherpa/). Each voice is its own npm package, loaded only when a phrase needs it:

Voice packageSpeaksModelDownload
@3sln/donki-lle-voice-yueCantonese (yue, zh-HK, zh-MO)VITS trained by xiaomaiiwn, 8-bit~38 MB
@3sln/donki-lle-voice-zhMandarin (zh, cmn, zh-CN, zh-TW)MeloTTS, 8-bit~61 MB
@3sln/donki-lle-voice-multi (+ -multi-2)Japanese, Korean, English, Spanish, French, German, Russian, Arabic, Hindi, Vietnamese and 21 moreSupertonic 3, 8-bit~145 MB

Everything arrives by import(): the runtime is one module with its WebAssembly inside, and each voice's files are modules holding base64 chunks. That is deliberate. A chat interface's iframe typically lets a page load scripts from a few CDNs and nothing else, so no fetch(), .wasm or .onnx URL would get through. Synthesis runs in a Web Worker where one can start, so the page stays responsive. Every voice is checked in CI by a speech recognizer, which has to hear the right language and the right words.

Pacing is done by the voice, not by slowing the audio, which would lower its pitch. The pace sets the model's own speaking speed, and each clause is synthesized separately with real silence between them. Each voice has a natural-speed calibration and a floor below which it isn't slowed further, both measured with a speech recognizer. The Cantonese model, for example, is brisk at its default speed (5.5 syllables a second) and loses words below 0.75 of it, so it is slowed only that far and gets its extra space from the pauses. System voices get the pace's speed, and pause at punctuation by themselves.

Voices are downloaded once per device. The element first tries to mount a hidden speech frame: donki.3sln.com/lle/speech/, a page with no UI that runs the engine and keeps voice files in its IndexedDB. It talks to the page only by postMessage, and only in terms of a voice's name and the text to say. Where the frame is refused (a chat's iframe usually forbids other sites' frames), the element notices within a moment and runs the engine itself, with the same cache in the page's own IndexedDB. Either way, the next lesson starts from the device instead of the network.

Browsers partition a framed site's storage by the site embedding it, so each site that hosts lessons keeps its own copy. Safari may also clear a frame's storage after a period without visits. configureSherpa({ frame: false }) skips the frame; frame: '<url>' points it at your own copy of the page.

Languages none of these cover (Afrikaans, Catalan, Swahili, Tagalog, Yoruba and others) fall back to Meta's MMS-TTS through transformers.js. That path fetches its model from Hugging Face, so it works on a normal page but not inside a strict chat iframe.

import { configureSherpa, configureNeural, registerEngine } from '@3sln/donki-lle';
configureSherpa({ cdn: 'https://my.cdn/npm/' });   // host the voice packages yourself
configureSherpa({ enabled: false });               // no sherpa voices
configureNeural({ enabled: false });               // no transformers.js fallback
registerEngine({                                   // your own engine, tried first
  id: 'mine', label: 'My voice',
  supports: (tag) => /^eu\b/i.test(tag),
  async load(tag, onProgress) { /* … */ return { synthesize: async (text) => ({ samples, rate }) }; },
});

<donki-lle-stroke>#

<donki-lle-stroke text="我"></donki-lle-stroke>                     <!-- Chinese -->
<donki-lle-stroke text="あめ" lang="ja"></donki-lle-stroke>         <!-- kana/kanji, Japanese order -->
<donki-lle-stroke text="안녕"></donki-lle-stroke>                   <!-- Korean -->
<donki-lle-stroke text="cat" mode="guided"></donki-lle-stroke>      <!-- Latin, traced -->
<donki-lle-stroke text="我" prefilled="3"></donki-lle-stroke>       <!-- finish the character -->

A square writing box. The learner draws each stroke in order. Each stroke is judged on where it starts, which way it runs, where it ends, and its shape. The judge forgives a wobble, a stroke a bit long or short, or one a little off-centre, but not the wrong place, the wrong direction or the wrong stroke. A miss says what was wrong ("Right place, other way round", "That's stroke 3. Stroke 2 comes first"). Repeated misses on one stroke escalate:

MissesHelp
1the start point is marked
2the stroke is shown faintly to trace
3 (demo-after)the stroke is demonstrated, then left to trace
5it is drawn for them, and practice moves on

The Hint button climbs the same ladder on request, and Show me animates the rest of the character. Several characters are practised one after another, with a strip showing progress.

Attribute
textOne or more characters. Spaces and punctuation are skipped.
langja for Japanese stroke order of kanji (inherited from an ancestor's lang too).
modeguided (default): the character is shown faintly. recall: a blank box, written from memory.
prefilledStrokes already drawn: a count (3) or indices (0,2). For "finish the character". Applies to the first character.
leniencystrict, normal (default), lenient.
demo-afterMisses before a demonstration (default 3).
gridcross, star (米), lines (handwriting ruling; the default for Latin), none.
json, srcStroke data inline or by URL, for anything not built in. A <script type="application/json"> child also works.
name, requiredIt's a form control. Inside a <form>, a named box submits its result as JSON (see Results): the finished result, or how far the learner got. So FormData and ordinary form posts carry the writing along with the other answers. required blocks submission until the writing is finished, and resetting the form starts the box again.

Any other content is shown as a caption under the box.

Built-in stroke data:

  • Latin letters A–Z and a–z, digits, and accented letters, composed from base letter plus mark (é, ñ, ü, ç, å…). These are the manuscript forms taught in schools, drawn to exercise-book ruling. This is for children learning to write and for adults learning the Latin alphabet.
  • Korean: every hangul syllable and jamo, composed from the jamo's strokes.
  • Chinese: about 9,500 characters from Make Me a Hanzi, plus colloquial Cantonese characters (哋 咗 啲 喺 嚟 佢 冇 …), which no open dataset covers. They're composed from real strokes of their parts, each part taken from a character that already has it in that position, and fitted together.
  • Japanese: kana and about 6,700 kanji from KanjiVG.

The Chinese and Japanese tables ship in the package as small shards of 16 code points (typically 5–20 KB), each an ES module. They're loaded with import() relative to the module's URL when a character needs them, so they work anywhere a script can load. configureStrokeTable({ baseUrl }) serves them from elsewhere.

Your own stroke data, for Cyrillic, Greek, Devanagari, a symbol, anything:

<donki-lle-stroke>
  <script type="application/json">
    { "character": "г", "size": 100,
      "strokes": [ [[30, 20], [70, 20]], [[30, 20], [30, 80]] ] }
  </script>
  г — the Cyrillic letter ghe
</donki-lle-stroke>

strokes lists the strokes in drawing order. Each is its centre line as [x, y] points in the drawing direction, in a size × size box with y pointing down. A stroke can also be { "path": "M… C…" }, an SVG path, when curves are easier to write that way. "prefilled": 1 in the data works like the attribute. Make Me a Hanzi JSON and KanjiVG SVG are also accepted by src as-is.

When the text is finished, it emits donki-lle-result (see Results). element.strokeData exposes the current character's normalized data.

<donki-lle-deck-share>#

<donki-lle-deck-share>
  <script type="application/json">
    { "type": "deck", "urn": "urn:donki:course:spanish-a1", "shard": "week-01",
      "name": "Spanish A1", "updatedAt": 1767225600000,
      "cards": [ { "urn": "urn:donki:course:spanish-a1:hola", "updatedAt": 1767225600000,
                   "face": "@tts(hola){\"lang\":\"es-MX\"}", "prompts": [{ "type": "anki", "answer": "hello" }] } ] }
  </script>
</donki-lle-deck-share>

Shows the deck's name, card count and shard, plus a list of what's in it, and three ways to take it:

  • Add to Donki opens Donki's import preview with the deck loaded.
  • Download saves a .donki.json file.
  • Copy copies the JSON, to paste into Donki's importer.

The deck is validated first, and problems are listed in plain words. The format is at donki.3sln.com/docs. app overrides where "Add to Donki" goes. json and src work as for the stroke element.

<donki-lle-send-results>#

<form id="quiz">
  <label>“library” in Spanish: <input name="library"></label>
  <donki-lle-stroke text="书"></donki-lle-stroke>
</form>
<donki-lle-send-results scope="#quiz" intro="Quiz, unit 3:"></donki-lle-send-results>

A button that collects everything the learner did inside scope (the whole page by default) and sends it into the chat as a message from the learner. It collects:

  • every named form field, so a quiz can be a plain HTML form;
  • the latest result of every donki-lle element, including unfinished ones ("Started writing 书 but didn't finish…"), labelled with the element's name if it has one.

The message is a readable summary followed by the same data as JSON:

Quiz, unit 3:
- library: la biblioteca
- Wrote 书 (4 strokes): 1 mistake.

```json
{ "answers": { "library": "la biblioteca" }, "results": [ { "kind": "writing", … } ] }
```
Attribute
scopeCSS selector of the container to collect from.
introThe message's first line.
labelThe button text. By default it reads "Send my results", or "Copy my results for the chat" where sending isn't possible.

How a page talks back to the chat#

sendToChat(text) (exported, and used by the button) tries in order:

  1. A sender the page configured: configureBridge({ send: async (text, data) => … }).
  2. window.sendPrompt(text), where the host provides one: Claude's inline widgets do.
  3. An MCP Apps host: detected with a ui/initialize handshake to the parent frame; the message goes as ui/message.
  4. The clipboard: the learner is told to paste it into the chat, and shown the text in case copying is blocked too.

A cancelable donki-lle-send event is dispatched first. Call preventDefault() on it to deliver the message your own way.

Results#

Interactive elements dispatch a bubbling donki-lle-result event when the learner finishes, and keep the latest as element.result. For writing practice it looks like this:

{ "kind": "writing", "text": "你好", "mistakes": 2, "completed": true,
  "characters": [ { "character": "你", "strokes": 7, "mistakes": 2, "missesByStroke": [0,0,2,0,0,0,0],
                    "shownStrokes": [], "hints": 0, "summary": "你 (7 strokes): 2 mistakes" }, … ],
  "summary": "Wrote 你好, 2 mistakes in total — …" }

Theming#

Colours come from custom properties with neutral defaults for light and dark:

:root {
  --lle-accent: #3d63dd;  --lle-accent-ink: #fff;
  --lle-surface: #fff;    --lle-surface-2: #f2f4f8;
  --lle-border: #d5dae3;  --lle-text: #18202e;  --lle-muted: #5b6577;
  --lle-good: #1d8a4e;    --lle-warn: #b36b00;  --lle-bad: #c9372c;
  --lle-radius: 12px;     --lle-font: inherit;
  --lle-stroke-size: 320px;   /* the writing box's largest size */
}

In a chat interface's iframe#

Which interfaces can render the elements, as far as we know (the teaching guide has the details and a test for any other):

  • Claude artifacts: yes. Scripts load from jsDelivr, the speech frame can't be embedded (voices load in the artifact instead), and downloads are blocked, so a deck leaves by Add to Donki or Copy.
  • Gemini's web_code_canvas: no. The elements don't load and files can't be handed over. Teach in text there, with decks as .donki.json code blocks.
  • Load the module from jsDelivr (above) or unpkg. Pin a major version (@0) so a lesson keeps working.
  • Everything loads by import() from the CDN (stroke shards, the speech runtime and the voices), because a chat's iframe usually allows scripts and nothing else. The neural voices also need WebAssembly, which a host's content security policy may forbid; where it does, speech falls back to the device's own voices.
  • Keep a page to one activity or one short quiz; it will usually be read on a phone.

Licences#

The code is MIT. The Chinese stroke data is from Make Me a Hanzi, under the Arphic Public License, which is included with it. The Japanese stroke data is from KanjiVG, © Ulrich Apel, CC BY-SA 3.0, also included with it.

For a language model: this page as Markdown.