#3389·solid

Perf tracker: Solid primitives vs the hand-rolled `@yak/solid` runtime

Author: ryansolidCreated Sep 12, 2026Updated Sep 17, 2026
Labelsenhancement

DigitecGalaxus/next-yak#644 rewrote the @yak/solid styled() runtime and made it 12× faster on SSR and ~2× on hydrate/mount over its own base. Reading the PR against our code, almost every piece of the rewrite is a reimplementation of a Solid primitive that was too slow or too generic to use directly:

yak hand-rolls instead of because
serializeElement ssrElement needs a finished props object; walks getter-backed merged props again; emits style=""/class="" (#3382)
copyProps / proxyProps merge / omit proxy layers when any source is $PROXY; omit has no predicate/prefix form; plain path leaks $SOURCES (#3384)
createElementRenderer (template + getNextElement + spread + insert) Dynamic / dynamic() per-instance memo for a constant tag; no static class in the template; no runHydrationEvents; tag-only namespace (#3386)
shared withTheme proxy merge(props, { theme }) one merge object per instance is measurable
once() on the server createMemo server memos are real nodes held until the deferred dispose (#3385)

This issue tracks closing those gaps so a library can call Solid and get the same numbers. Baseline harness: ryansolid/yak-bench — the 14 css-in-js-bench workloads, @yak/solid at the PR's base and head, both on 2.0.0-rc.8, with next-yak/React as the reference lane. Acceptance for every item below: a variant of the yak runtime using our primitive lands within noise of the hand-rolled one in that harness.

Where the remaining time goes

Solid 2.0.0-rc.8, @yak/solid at PR head. Full tables: SSR · Chromium.

Static styled intrinsics are fast (SSR 25× React, hydrate/mount 2–3.5× React). The cases still behind React are the ones that go through our generic paths — styled(Component) with {...rest} spreads, dynamic $props, imported primitives:

case SSR PR/React hydrate PR/React mount PR/React
tabs (styled(Component) + spread) 1.10× 0.45× 0.63×
multifile-composition 1.05× 0.46× 0.64×
product-grid (400 tiles × 11 el) 4.15× 0.46× 0.81×
realistic-button 6.10× 0.72× 1.02×
dyn-fair (CSS var per element) 10.06× 0.81× 1.04×
btn-variant vs compose-1 (dynamic vs static props, same element) 2.7M vs 10M inst/s 30.4 vs 15.7 ms 22.6 vs 9.3 ms

Checklist

  • spread() → one render effect per element (client.ts ~809, the // TODO: make this better). Three reactive nodes per element today; the dynamic-props path costs 2× the static one on hydrate and mount, and spread is the largest share. (#3388) — landed in #3419: ref folded into the attribute effect, children stay owned (two nodes with children, one without), sources array; compilers emit the array form for element spreads (#3423, universal #3424). Browser hydrate on the styled-element shape went from 1.47× behind yak's hand-rolled path to 1.00×.
  • Constant-tag Dynamic/dynamic(): when source() is a string that can't change, skip the factory + instance memos, allow a static class baked into a cached template, and take the compiled-JSX element path. Same fix carries runHydrationEvents and namespace correctness. (#3387, #3386) — landed as dynamic(source, { static }) + isStatic(o, key) in #3471: no memo per instance, the tag goes straight to ssrElement / the compiled element path, and a library decides per instance from the prop's descriptor (isStatic sees through merge/omit layers), so a literal as="a" at the call site takes the static path while a reactive as keeps the memo. #3386 part 1 (runHydrationEvents) fixed in #3396, part 2 (namespace: dynamic() honors xmlns) in #3436; <Dynamic> deprecated in favor of dynamic(). The compiler-lowered <element tag> (#3429) stays open as a later question — the runtime path is fast enough that it is a spread-shape question, not a memo one.
  • Server merge/omit fast path when every source is a plain object (the rc.8 spread fast path 4e730a9, but on the server), and a predicate / $-prefix form of omit so "drop all $ keys" doesn't require enumerating first. — landed in #3454 (merge/omit always return O(1) lazy views; an omit over a merge flattens to filtered leaf entries; omit(props, key => …) predicate form; truthful getOwnPropertyDescriptor through every layer), #3470 (the view protocol consumers walk moves behind solid-js/internal), and #3475: a view's resolved key table is built by enumeration or after 16 reads, never on the first read, and ssrElement walks a view's entries instead — profiled on the Kobalte-shaped chain, a third of SSR time was the table code and its garbage. Same-process A/B: SSR chain 8.2× → 5.8× the compiled floor. A follow-up that chained omit's folded hidden-key lists instead of copying them (#3487) was closed unmerged: against a correctly built next, no fold form beats slice()+push on instruction count, and its TurboFan gains appeared only on an artificial depth-7 chain — the allocation it targeted is real but costs no time. #3497 then took the other route: an omit over a merge holds the merge record (SOURCE_MERGE) instead of one filtered leaf view per flattened source — one record per layer, and the nested walk is by function call with a MISSING sentinel (one pass per read, no trap hop, the outer view builds the table in one pass). ssrElement gets sourceOwners: keys and owning objects for any source in one walk. Kobalte-shaped depth-7 chain: build+consume +62%, ownKeys(defaults 100) ×4.9, the nine merge-* construction benches +9–20%; yak-bench composition cases move from 0.8× to parity (0.95–0.99×) with the merged runtime. Original design note follows. CPU profile of the tabs SSR case (styled(Component) + omit + spread): merge 18%, omit 10%, yak's copyProps 11%, GC 16% (mostly from the same allocations); Solid's rendering proper (ssrElement, escape, resolveSSRNode, ssr) ~7%; component machinery ~2%. Each Tab builds a chain of three or four Proxies and every consumer enumerates through the traps. Direction: omit carries its skip-set on the $SOURCES brand and ssrElement/spread/merge walk branded views directly, never through traps; the Proxy stays only for direct property reads.
  • ssrElement with skip rules / extra sources, or a lower-level serializer, so a library doesn't have to materialize getter-bearing props just to have them walked again (yak measured −18.9% HTTP throughput on that path). — landed in #3418: ssrElement(tag, sources[], children, needsId, skip?), single walk, later wins, winner read once. On dynamic-prop shapes ssrElement is now within 3–23% of yak's serializeElement (was 2–3× behind); the fully static shapes remain 3× behind because yak string-concats those at build time (not our gap). #3486: plain string/number/null/boolean and finished-node children concatenate in place without resolveSSRNode/ssr(), and a uniform source array carries one kind instead of an array of them — element-dense SSR (dyn-translate) +21% throughput, −27% bytes/instance; gap to yak's direct writer 1.76× → 1.46×.
  • Hydration per-element overhead: hydrate is ~1.7× mount for 1000 static buttons (15.7 vs 9.3 ms at 4× throttle). Profile gatherHydratable / registry build and per-element getNextElement, and stripTextSeparators + [...childNodes] in claimInitial on large flat lists (interacts with #3383 separators). — the styled-element shape closed with #3419 (1.47× behind yak's path → 1.00×), and yak reports the whole PR's hydrate on rc.8 "inside the floor" with rc.8's spread fast path alone worth ~12 points to their runtime. The composition cases (tabs, multifile: 0.45× React on rc.8's predecessor) have not been re-measured since; re-run on rc.9 before profiling further. — Re-measured after #3509 (4× throttle, median of 15 fresh pages): hydrate tabs 0.44× React, multifile-composition 0.46×, polymorphic-chain 0.53× — unchanged; #3419 fixed the styled-element shape, not composition. And on all three, Solid's hydrate is slower than Solid's mount (54.9 vs 47.6 ms on tabs) while React's hydrate is 25% cheaper than its mount. Chromium CPU profile of tabs hydrate vs mount, sourcemap-attributed, per page: hydration saves ~1.5 ms of DOM writes (setAttribute, className, clone/append) and spends ~3.6 ms claiming. Hydrate-only costs: gatherHydratable 0.85 ms (per-[_hk] closest("[data-fid]") + contains to skip frame interiors — a DOM-ancestor walk per element even when the page has no frames), claimInitial 0.8 ms ([...parent.childNodes] iterator spread on every hydrating insert(), then a second stripTextSeparators pass), getNextElement 0.4 ms (string key + Map get/delete + WeakSet add), clearSnapshots 0.47 ms (delete source._x._snapshotValue over every source captured during hydration — delete on a fixed-shape object; assignment of undefined is equivalent and is what the store branch already does), isHydrating(node) 0.29 ms (node.isConnected per call), _x extension allocated per hydrated source to hold its snapshot (GC 1.8 vs 1.4 ms). The first three plus clearSnapshots are ~2.1 of the ~3.7 ms/page excess and are local fixes in client.ts / core.ts.
  • Cheap "props plus one key" for the theme case (merge(props, { theme }) allocates a merge object per instance; Object.create(props) was slower). — since #3454 a merge() is one record + one Proxy with no descriptor copy, the same shape as the shared-handler view yak settled on (withTheme / viewTraps, 279317a6). Measured at parity on next with the merged-prim overlay (__PRIM_THEME__ 1.03× geomean, inside the ±8% lane noise); the yak-side deletion waits on rc.9.
  • Re-run the harness after each item (gated on an rc.9 so yak targets published packages); when a primitive is within noise of yak's version, open the corresponding follow-up on @yak/solid to delete the duplicate. — blocked on rc.9 (#3399). DigitecGalaxus/next-yak#644 merged 2026-09-15 on rc.8, so none of #3454 / #3471 / the table follow-up is in what it measures. Its final experiment table (rc.8 vs rc.6) is the list to re-run: see the status comment below.

Not our gap

Definition-time chain flattening, static-class collection, choosing the renderer when the module loads, caching template parts: that is yak doing at runtime what our compiler does at build time. No Solid change removes it.

Landed alongside (correctness surfaced by the same audit)

  • #3383 <!--!$--> separators decided on resolved values → #3394 (+ #3430 recovers the ~2% SSR cost of the walker split)
  • #3382 empty style=""/class="" → #3395
  • #3384 merge() stale $SOURCES → #3401
  • #3385 renderToString deferred dispose → #3422
  • #3386 (1) runHydrationEvents after dynamic() string tag → #3396

Versus React (the part that matters more than parity with yak's PR)

Element-dense cases: Solid 1.5–4× faster than React on SSR, hydrate and mount. Deep composition (compose-*, button-variants*): 5.3–6.8× React on SSR after #3497 (yak's PR: 4.2–5.5×). Still behind React on SSR — and identically so across every yak lane, so none of the props-plumbing work above touched it: polymorphic-chain 0.29×, tabs 0.50×, multifile-composition 0.51×. That residue is Solid's component layer on the server, not merge/omit. Attributed (--cpu-prof --no-turbo-inlining, sourcemapped, polymorphic-chain / tabs SSR): compiled component code 24–26% (the props literals with accessors the compiler emits at every call site — #3511, which also finds the non-identifier keys are emitted computed, get ["aria-disabled"](), a 45% cost by itself), merge/omit views 33% on the chain, GC 13–15%, reactive graph / owner / child-id 16%; Solid's serializer is faster than React's renderer. Server memo() identity was read and rejected: _$memo is id-load-bearing on the server (#3033, #2976) — its owner slot is the retry-stable scope a deferred hole re-runs under, so it can't be transparent or a bare id burn. Hydrate half: re-measured, still 0.44–0.53× React, and slower than our own mount — see the hydration item above.