Status: prototype. Wired for
highlight.jstoday; sketches the same shape for Mermaid. Seesrc/highlight.ts,src/highlight-hljs.ts, andsrc/highlight-lazy.test.ts.The demo site (
docs/index.html) exercises both paths live: its bundle is built with code splitting, so the highlight.js grammars are fetched as a separate chunk on the first fenced code block, and mermaid is fetched via an import map on the first ```mermaid fence. Status chips on the page show each load happening.
renderMarkdownUnsafe() (the "generator") is a synchronous string→HTML function. Its
only heavy runtime dependency is highlight.js — the core plus a dozen language
grammars. Because render-blocks.ts imported highlight.ts, which statically
imported highlight.js, the whole grammar payload landed in every consumer's
main bundle, even one that never renders a fenced code block.
Measured with esbuild (--bundle --minify --format=esm) on dist/:
| Bundle | Size | Contains highlight.js? |
|---|---|---|
Main entry (index.js), before |
~164 KB | yes (unconditionally) |
Main entry (index.js), after |
~88 KB | no |
highlighters/highlightjs chunk (lazy) |
~71 KB | yes (fetched on demand) |
So ~71 KB — the grammars — now loads only when a host asks for highlighting.
The figures above are minified-only, whole-payload sizes. The number a consumer
actually pays over the wire is gzipped, and — because the peer dependencies
(dompurify, highlight.js, katex, mermaid, shiki, entities) are the host's to
bundle — it excludes them. Measured that way (--bundle --minify --format=esm --platform=browser, peers external, gzipped), the main entry is ~31.5 KB
gzipped, and the emoji data subpath is the next largest at ~16.2 KB.
This is no longer a hand-run figure. scripts/check-bundle-size.mts
(npm run size) measures the main entry and every key subpath against the
committed budgets in scripts/bundle-size-budget.json, and the CI size job
fails the build if any entry exceeds its budget. The emoji table — bundled, not
external — is caught by that byte budget: dragging it into the core entry blows
the . budget outright.
A peer dependency is trickier: because the host bundles it, esbuild keeps a
leaked import 'shiki' as an ~30-byte external statement, well inside the ~5%
headroom, while the consumer's real chunk silently grows ~70 KB. So the gate
also reads esbuild's metafile and fails if the main entry statically imports
any peer — a leaked grammar or math payload reddens the PR on the import
itself, not on the bytes. (A lazy import() of a peer on its own subpath is the
intended pattern and is left alone.) When an increase is intentional, bump the
budget deliberately with npm run size:update and commit the JSON.
This reuses the exact pattern the sanitizer already uses (the sanitizerBackend
config field + the @copse/streaming-markdown/sanitizers/dompurify subpath):
highlight.ts(core) carries no highlight.js code. It keeps only the cheap string work — language aliases, theKNOWN_LANGUAGESset, andfenceCodeClass— and a config-injected slot: thecodeHighlighterfield, read internally viagetCodeHighlighter()and applied around each render bywithConfig. With no highlighter configured,highlightFenceCodereturns escaped plain text.highlight-hljs.ts(backend) is the only module that imports highlight.js. It registers the grammars and exportshighlightjsHighlighter(the value) andloadHighlightjs()(which returns thatCodeHighlighter), to pass viaMarkdownConfig.codeHighlighter. It lives behind the@copse/streaming-markdown/highlighters/highlightjssubpath, so a bundler drops it unless the host references that entry.
fenceCodeClass('ts') returns hljs lang-typescript before the backend loads
and the identical string after. That stability matters for streaming: the core
renders a code fence as plain escaped text with the final class immediately, and a
later re-render (once the grammar chunk arrives) only swaps the interior to token
spans — the <pre><code class> element never churns. KNOWN_LANGUAGES in the core
must stay in sync with the grammars the backend registers.
Eager (highlighting from first paint — pulls the chunk into your bundle):
import { renderMarkdown } from '@copse/streaming-markdown'
import { highlightjsHighlighter } from '@copse/streaming-markdown/highlighters/highlightjs'
renderMarkdown(md, { codeHighlighter: highlightjsHighlighter })Lazy (the grammars are a separate chunk, fetched only when this runs):
// e.g. on the first fenced block seen, or during an idle callback
const { loadHighlightjs } = await import('@copse/streaming-markdown/highlighters/highlightjs')
const codeHighlighter = await loadHighlightjs() // returns the highlighter
// re-render the message with { codeHighlighter } in config so already-rendered
// fences upgrade from plain → highlighted
rerender({ codeHighlighter })Until either runs, code fences render as safe, escaped plain text with the correct
hljs lang-* class.
@copse/streaming-markdown/highlighters/shiki is a second highlighter backend
for the same codeHighlighter config slot. It sits between the two patterns
above: like mermaid, shiki is an optional peer dependency reached only
through variable-specifier dynamic imports (the package builds without it, and
zero shiki bytes can land in the main entry — or in the subpath chunk itself);
like highlight.js, the backend highlights synchronously once ready.
Because shiki can only initialize asynchronously, loadShiki() is the load
seam: it awaits shiki's fine-grained core (shiki/core +
createHighlighterCore with the JavaScript regex engine — no oniguruma WASM,
no bundled-registry entry) plus the grammar and theme modules, then registers a
backend that highlights synchronously against the loaded instance. Until it
resolves, fences render as escaped plain text with the same stable
core-resolved hljs lang-* class, and a re-render upgrades them in place —
identical UX to lazy highlight.js. Token colors are emitted as classes (not
shiki's inline style attributes, which the sink sanitizer strips); see the
Shiki section of EXTENDING.md for the styling contract
and shikiThemeCss().
Mermaid is not bundled by this package: mermaid-source.ts is pure string
preparation, and the generator only emits inert scaffolding
(<div class="mermaid-diagram mermaid-diagram--pending"><pre class="mermaid">…).
The heavy mermaid library is host-injected and rendered after sanitization, so
it is already lazy by construction. What the package was missing is an official
hook so every host stops hand-rolling the "find pending diagrams, load mermaid,
inject SVG, retry on the aggressive source candidate" dance.
The same registry shape as highlighting now ships for it:
mermaid.ts(core) carries no mermaid code. It holds the async renderer seam (thediagramRendererconfig field, consumed byhydrate()/hydratePendingDiagrams'srendereroption — not the synchronous render), theDiagramRendererinterface, andhydratePendingDiagrams(root, opts?)— which finds everymermaid-diagram--pendingcontainer underroot, tries the gentle then aggressivemermaidSourceCandidates()until one renders, and flips the container to--rendered(SVG injected) or--error.mermaid-mermaidjs.ts(backend) is the only module that referencesmermaid. It lives behind@copse/streaming-markdown/diagrams/mermaidand exportsmermaidDiagramRendererandloadMermaid()(which returns the renderer value — pass it viaMarkdownConfig.diagramRendereror the hydraterendereroption).mermaidis an optional peer dependency — the host installs it, the package never bundles it. The dynamic import uses a variable specifier so the package builds and type-checks even when the peer isn't installed.
import { hydratePendingDiagrams } from '@copse/streaming-markdown'
const { loadMermaid } = await import('@copse/streaming-markdown/diagrams/mermaid')
const renderer = await loadMermaid() // returns the backend; library loads lazily
await hydratePendingDiagrams(messageEl, { renderer }) // pending → rendered SVGTrust boundary: mermaid SVG is produced by the trusted library after the sink
sanitizer runs and is injected without re-sanitization (matching the existing
design invariant). A safety-conscious host can pass hydratePendingDiagrams(root, { transformSvg }) to run the SVG through its own sanitizer first.
The mechanism is covered by src/mermaid-lazy.test.ts with a stub renderer (the
real mermaid library needs a browser DOM it can't get in jsdom); the backend module
is the thin adapter over mermaid.render.
Math (#70) mirrors mermaid exactly. The generator emits inert scaffolding for
every math form — ```math fences, $$ … $$ / \[ … \] display
blocks, and $…$ / \(…\) inline spans
(math-block--pending / math-inline--pending, escaped TeX source inside) —
so the KaTeX payload is never needed at render time. The prose delimiter
grammar (everything except the always-on ```math fence) is itself
opt-in via mathSyntax (#78): until { mathSyntax: true } turns it on (the
null default defers to a process-wide renderer registration), $…$-style text
stays ordinary prose and output is byte-identical to a math-free build, so hosts
that never opt in pay nothing at all:
math.ts(core) carries no KaTeX code: the async renderer seam (themathRendererconfig field, consumed byhydrate()/hydratePendingMath'srendereroption — not the synchronous render), theMathRendererinterface, andhydratePendingMath(root, opts?)— which finds every pending block/span underroot, renders it (display mode for blocks, inline for spans), and flips the element to--rendered(HTML injected) or--error(escaped source kept visible).math-katex.ts(backend) is the only module that referenceskatex. It lives behind@copse/streaming-markdown/math/katexand exportskatexMathRendererandloadKatex()(which returns the renderer value — pass it viaMarkdownConfig.mathRenderer, withmathSyntax: truefor the prose grammar, or the hydraterendereroption).katexis an optional peer dependency, imported through a variable-specifier dynamic import so the package builds and type-checks without it.
import { hydratePendingMath } from '@copse/streaming-markdown'
const { loadKatex } = await import('@copse/streaming-markdown/math/katex')
const renderer = await loadKatex() // returns the backend; library loads lazily
// render with { mathSyntax: true } to turn the prose grammar on, then:
await hydratePendingMath(messageEl, { renderer }) // pending → rendered KaTeX HTMLThe host also loads KaTeX's stylesheet/fonts (katex/dist/katex.min.css).
Trust boundary: KaTeX HTML is injected after the sink sanitizer and not
re-sanitized, matching the mermaid invariant; hydratePendingMath(root, { transformHtml }) is the seam for a host that wants to (and the required hook
under Trusted Types enforcement).
Emoji shortcodes (#86) reuse the split for a data payload rather than a
library. The optional pass and its GitHub/gemoji-aligned map live in
src/emoji-shortcodes.ts / src/emoji-shortcode-map.ts behind
@copse/streaming-markdown/inline/emoji; the ~50 KB alias table is pulled into a
bundle only when a host imports that entry, so a consumer that never opts in pays
zero bytes for it (src/emoji-shortcodes.test.ts asserts this with an esbuild
bundle of the main entry). Unlike the highlighter/mermaid/KaTeX backends there is
no optional peer dependency — the pass is pure string work built on the public
inlinePasses config field, so it needs no registry in the core at all.
import { renderMarkdown } from '@copse/streaming-markdown'
import { emojiInlinePass } from '@copse/streaming-markdown/inline/emoji'
renderMarkdown(md, { inlinePasses: [emojiInlinePass] }) // `:smile:` → 😄; see EXTENDING.md#custom-inline-syntax-inline-passesChunky token arrival shows as chunky updates: an LLM transport delivers text in
bursts, and the incremental DOM emitter renders exactly what it is given, when it
is given it. The optional input smoother steadies that reveal — it sits
between the host's chunk arrival and renderer.update() and releases the growing
string a few characters per frame (via requestAnimationFrame) instead of in raw
token bursts.
It lives behind @copse/streaming-markdown/smoothing and is never re-exported
from the main entry, so a host that doesn't import it pays zero bytes and the
default emitter behaviour is byte-for-byte unchanged (the same "zero bytes in the
main bundle" contract as the highlighter/mermaid/KaTeX backends above —
src/smoothing.test.ts asserts an esbuild bundle of
the main entry contains none of the smoother's code).
import { StreamingMarkdownRenderer } from '@copse/streaming-markdown'
import { createInputSmoother } from '@copse/streaming-markdown/smoothing'
const renderer = new StreamingMarkdownRenderer(host)
const smoother = createInputSmoother({
update: (text) => renderer.update(text), // the sink for each revealed prefix
cadence: 'adaptive', // follow the stream's own rate (default 'fixed')
})
for await (const fullTextSoFar of stream) smoother.push(fullTextSoFar)
smoother.finish(() => showFinalRender()) // stream end: drain the rest, then settle
smoother.dispose() // tear down (cancels any pending frame)push(text) takes the full accumulated message so far (the same argument
renderer.update takes), so every value the smoother releases is a prefix of
the target. That is exactly the streaming contract the pending-state machinery
and DOM morph already converge on, so smoothing composes with them for free
rather than fighting the morph.
'fixed'(the default, for compatibility) walks the revealed prefix at a constantcharsPerSecond(default 600). It is predictable, but it does not follow the stream: a transport slower than the rate still shows as bursts (each chunk drains in a frame or two, then the reveal waits for the next), and one faster than the rate falls further and further behind.'adaptive'— recommended for LLM output — reveals at a velocity that tracks the arrival rate, running aboutlagMs(default 120) behind it. The velocity is low-pass filtered, so a steady stream reveals steadily, a burst speeds the reveal up over a few frames rather than in one jump, and the lag stays bounded however fast the model is. After a frame gap long enough to mean the page was hidden, it catches up at once instead of replaying text.
Measured on a transport delivering 12 characters every 50ms, 'fixed' leaves
the text frozen in roughly two frames of three; 'adaptive' advances it in
nearly every frame, 4–5 characters at a time.
finish(onSettled?)reveals whatever is still pending over a short drain (~60ms of lag, whatever the cadence), then callsonSettledonce — the place to swap in a final, at-rest render. It settles synchronously when nothing is pending, smoothing is off, or the page is hidden. Apushafterfinishmeans the stream resumed and drops the pending callback.flush()releases everything immediately (and settles a pendingfinish). Use it when completion must carry no delay at all.
A revealed prefix never ends right after markdown punctuation or whitespace
(` * _ ~ [ ] ( ) | # > ! - + = . : < \ and digits): the cut moves forward
past that run — a bounded few characters — so a frame always ends just after a
word character. Otherwise the half-arrived syntax would show raw for a frame
before the renderer could tell what it becomes: the first two backticks of a
closing fence, the || of a table separator, a 1. before its list item, the
brackets of [text] before (url) arrives, or a trailing newline briefly
opening an empty line. A cut also never splits a surrogate pair.
Pass initial with the text already on screen when a host rebuilds a message
element partway through a stream: the reveal starts from its end instead of
replaying the whole message.
Two designs were on the table (see #84): throttle the input string fed to
update(), or animate the output (CSS/opacity transitions on newly-added
nodes). We ship input smoothing only — it is lower-risk and framework- and
theme-agnostic, and it reuses the renderer's existing convergence guarantees
instead of introducing a second animation system that could fight the DOM morph.
CSS entrance animations are deliberately left to the host: a host that wants
nodes to fade in can style the emitter's class hooks (.stream-pending,
.stream-complete) with its own transitions — that is a presentation concern the
package does not own.
- Convergence. After
flush(), or oncefinish()settles, the sink has seen the full text, so the rendered DOM equals a single un-smoothedupdate(fullText)— asserted against a direct full-string render in the tests, for both cadences. prefers-reduced-motion. Honoured by default: when the environment reports reduced motion, smoothing is disabled andpushpasses straight through immediately (respectReducedMotion: falseoverrides;disabled: trueforces pass-through unconditionally).smoother.enabledreports which mode is active.- No completion lag on demand.
flush()releases the whole pending target at once and stops the loop;finish()drains it within a few frames. - Never recursive. A frame scheduler that calls back synchronously (some test shims do) cannot pace anything; the smoother detects it and releases text as it is pushed instead of recursing.
- Environment-guarded.
requestAnimationFrame/cancelAnimationFrame,performance.now, andmatchMediaare all read through defaulted, injectable seams (requestFrame/cancelFrame/now/matchMediaoptions), so the smoother runs — and is unit-tested with a fake clock — under node/jsdom where those globals may be absent.