ProGPU validates its clean-room SkiaSharp shim against the public ECMA-335
metadata in the official SkiaSharp NuGet package. The lock in
eng/skiasharp-api-baseline.json pins the package URI, SHA-512, target-framework
reference assembly, namespace, and monotonic regression budget.
The current contract is SkiaSharp 4.151.0, using
ref/net10.0/SkiaSharp.dll. This advances the previous implementation record
from Skia m148 to the current stable SkiaSharp package without consulting or
copying its implementation source.
Run the complete gate with:
./eng/progpu-verify-skiasharp-api.shThe gate verifies the official package hash, extracts only its public reference
metadata, self-tests the canonical metadata reader, builds the ProGPU shim, and
writes deterministic JSON and Markdown reports under
artifacts/skiasharp-api/. CI fails if exact matches decrease or missing entries
increase.
API equality is necessary but not sufficient. Every implementation slice must also include independent behavioral tests, Svg.Skia/Avalonia.Skia compatibility evidence where applicable, and matched Release benchmarks for native SkiaSharp and ProGPU. Rendering work must preserve ProGPU's WebGPU ownership, quality, device-loss, bounded-resource, and allocation contracts.
The matched benchmark runner is
eng/progpu-run-skiasharp-benchmarks.sh. It compiles identical source against
official SkiaSharp and ProGPU, alternates process order, verifies semantic
checksums, and preserves raw median/p95 timing and allocation distributions plus
environment metadata. Its scheduled workflow runs on macOS, Linux, and Windows;
small timing deltas on shared runners remain informational until calibrated on
dedicated hardware.
The initial local Release run on an Apple M3 Pro, .NET 10.0.5, macOS 26.4.1, using three alternating process pairs and 72 measured samples per backend, produced the following diagnostic baseline:
| Workload | Native median ns/op | ProGPU median ns/op | ProGPU/native | Native B/op | ProGPU B/op |
|---|---|---|---|---|---|
| point arithmetic | 2.076 | 2.158 | 1.039 | 0 | 0 |
| matrix map point | 8.713 | 4.547 | 0.522 | 0 | 0 |
| path builder, detach, and bounds | 808.542 | 3,284.499 | 4.062 | 168 | 3,520 |
These figures identify path construction/ownership as the first measured CPU and allocation hotspot. They are not a cross-platform performance claim; raw distributions and environment records remain in generated artifacts, and the path work requires matched profiling plus equivalent before/after runs.
The current pinned comparison records 4,222 official entries, 5,219 ProGPU entries, all 4,222 exact matches, zero missing entries, and 997 ProGPU-only entries. The matching/missing budget is now locked at full official coverage. ProGPU-only entries are audited and removed when accidental; explicitly documented extension seams remain outside the official parity claim.
The continuation branch regenerated the original 97-entry gap from the pinned
official package at v0.1.0-preview.34 (39b53dbb) before implementation.
Public metadata closure does not by itself complete rendering compatibility:
GPU-visible families still require original retained WebGPU implementations
and quality/performance tests, while unsupported platform codecs must fail
explicitly rather than silently emit another format.
The remaining mask-filter and forwarding-canvas slice is a clean-room design based on public contracts and independently observable behavior. No implementation source from SkiaSharp or another renderer is used.
Primary sources consulted:
- Skia
SkMaskFilter: mask filters transform coverage before compositing; Gaussian sigma must be positive and may be transformed by the current matrix. - Skia
SkCanvasandSkOverdrawCanvas: canvas state is a matrix/clip stack, while overdraw records every touched pixel rather than final source color. - Direct2D Gaussian blur: a separable GPU blur uses transparent soft borders and a conservative three-sigma radius.
- Direct2D effects overview and custom effects: retained effect graphs compose GPU transforms and explicitly expand input rectangles for non-local sampling.
- Win2D GaussianBlurEffect: retained effect nodes expose bounds/invalidation and may cache their output.
- WebRender: retained display-list rendering keeps scene preparation separate from GPU raster/composition.
- Vello and its image-filter status: GPU compute is the intended parallel execution boundary, while filter semantics remain an explicit renderer concern.
- Parley and HarfBuzz shaping: shaped glyph IDs, positions, and layout are reusable CPU results and are not recomputed by a coverage effect.
- DirectWrite glyph runs and Direct2D/DirectWrite integration: text layout remains independent from the renderer that consumes the retained glyph run.
Adopted: immutable filter snapshots, conservative three-sigma bounds, retained command replay, transparent out-of-bounds sampling, and GPU effect composition. Adapted: mask filters route through ProGPU's existing WebGPU save-layer/effect graph so paths, glyphs, images, and custom visuals share one backend. Rejected: CPU pixel fallbacks, moving Unicode or OpenType shaping to the GPU, unbounded per-frame filter allocation, and source-shaped ports of another engine's implementation.
- Close metadata-only value, enum, descriptor, and ownership contracts that do not require GPU initialization.
- Complete bitmap, pixmap, image, codec, stream, and color-space contracts with explicit CPU/GPU ownership and no accidental readback or upload.
- Complete paths, regions, paint, text, picture, document, and canvas behavior over reusable ProGPU primitives.
- Complete shaders, filters, blenders, masks, vertices, atlas, surface, and GPU context contracts through retained WebGPU pipelines and embedded shaders.
- Prove source-level Avalonia.Skia substitution, close the full Svg.Skia corpus, and enforce representative CPU, GPU, frame-time, and memory advantages over the official runtime on supported platforms.
Primary public contracts:
- https://www.nuget.org/packages/SkiaSharp/4.151.0
- https://learn.microsoft.com/dotnet/api/skiasharp
- https://www.w3.org/TR/SVG2/
- https://www.w3.org/TR/webgpu/
All 4,222 public entries in the pinned SkiaSharp 4.151 reference metadata now match exactly, including declaring type, inheritance, fields, signatures, parameter names, layout, ownership hooks, obsolete payloads, and nullable metadata. The regression gate is ratcheted to zero missing entries.
SKMaskFilter now retains immutable blur, table, gamma, clip, and shader
coverage descriptions; conversions and fast paint bounds are fixed O(1),
while table construction is bounded O(256). Overdraw color filters retain
their six-color palette and clamp positive coverage counts to the last color.
No factory initializes WebGPU or reads pixels.
At draw time, a typed retained-brush marker intercepts only commands that carry a mask filter. Ordinary commands perform two marker checks and never invoke the mask delegate. Filtered commands render their retained geometry or glyph run to an offscreen texture and reuse the existing WebGPU image-effect graph for separable blur, alpha tables, shader masks, and solid/inner/outer composition. Overdraw palette mapping uses a dedicated 16x16 WebGPU compute pass with one texture read and write per texel and a fixed 96-byte six-color uniform. Pixel tests cover Gaussian falloff and exact zero/one/saturated overdraw counts; the shader resource audit enforces embedding and complexity documentation.
SKNoDrawCanvas, SKNWayCanvas, and SKOverdrawCanvas share a typed retained
command-forwarding seam in DrawingContext. Every newly recorded command is
forwarded immediately, including its packed buffer slices and retained resource
leases. Fan-out costs O(T * (B + R)) for T targets, referenced packed data
B, and retained resources R; ordinary canvases pay one direct delegate-null
check per recorded command. Overdraw forwards additive 1/255 coverage rather
than source paint color. Focused tests cover immutable tables, conversion and
bounds behavior, immediate target removal/fan-out, and additive coverage
commands.
SKFontMetrics now uses the official sequential flags-plus-fifteen-floats ABI.
The four nullable decoration metrics are represented by validity bits and
inline values, preserving null semantics in a fixed 64-byte value with no heap
storage. SKRawRunBuffer<T> now uses the official readonly pointer/length
layout. Builder arrays are allocated directly in pinned managed storage and
remain owned by the builder, so glyph, position, text, and cluster spans stay
valid across compacting collections without GCHandle, copying, or per-access
allocation.
Focused tests verify the exact field types and size, nullable metric behavior,
raw span lengths and snapshots, compacting-GC stability, and exactly zero
managed bytes across 10,000 warmed position reads. The matched benchmark suite
includes the same public raw-buffer access workload for official SkiaSharp and
ProGPU. Three alternating clean Release pairs at c78266ec retained matching
checksums and measured 4.614 ns/op for official SkiaSharp versus 4.701 ns/op for
ProGPU (1.019 ratio). One-time sample setup amortized to 0.002 versus 0.004
B/op; the warmed access loop itself remains allocation-free. This is neutral
timer-floor evidence, not a performance-win claim. The exact metadata gate
advances from 4,183 to 4,186 matches, reduces
missing entries from 39 to 36, and removes six accidental ProGPU-only metadata
entries.
The WebP frame value, static encoder surface, SVG canvas type shape, managed
stream fork/duplicate declarations, and memory-stream native-disposal hook now
match the pinned public metadata. WebP frames borrow pixmaps directly; the
bitmap constructor reuses PeekPixels, while the image constructor performs
the explicit image-to-raster readback requested by that API and retains its
pixel owner through the pixmap. Static and animated WebP encoding currently
return null or false without writing because the reviewed dependency-free
platform layer does not yet expose a WebP encoder on every supported target.
This is an explicit capability failure and never emits PNG/JPEG bytes under a
WebP contract.
Focused tests cover frame layout and mutation, borrowed pixels, zero-byte failure behavior, and the non-static/non-constructible SVG helper shape. The exact metadata gate advances from 4,157 to 4,183 matches, reduces missing entries from 65 to 39, and removes one accidental ProGPU-only metadata entry.
Thirty-two official protected ownership hooks now appear on their declaring
SkiaSharp types while retaining ProGPU's single SKNativeObject lifetime
engine. The declarations cover managed/read/write streams, bitmap and codec
wrappers, color and image filters, color spaces, drawables, font styles, paint,
paths, pictures, surfaces, and text blobs. They delegate to the existing
idempotent base implementation; no native handle model, allocation, rendering
path, or GPU initialization was added. The protected SKDrawable(bool owns)
constructor now preserves borrowed ownership without adding a public adapter.
Independent reflection and lifetime tests verify every declaring type, virtual override shape, borrowed drawable ownership, and post-disposal handle state. The exact metadata gate advances from 4,125 to 4,157 matches and ratchets missing entries from 97 to 65 without increasing the 1,019 documented ProGPU-only entries.
The current checkpoint closes 62 official metadata gaps without importing a
native ownership model. SKData, SKFont, SKRegion, its three iterators, and
SKPixmap now participate in the shared SKObject lifetime contract. Data
subsets still share one pinned reference-counted store, release callbacks still
run once after the final view, and the protected empty singleton remains usable
after public disposal. Region iterators retain their bounded snapshots and
pixmap disposal resets only its borrowed CPU view; none of these operations
initializes WebGPU or takes ownership of caller memory.
Public parameter names, optional metadata, and overloads now match the pinned
4.151 reference for color values, sampling values, color spaces, pixmaps,
regions, discrete path effects, and the remaining image-filter factories.
CreateEmpty maps to the existing transparent GPU shader-filter path, while
the legacy crop factories map to the existing input graph plus retained crop
rectangle. Object creation and signature adapters are fixed O(1) work;
SKData final release is fixed work plus its caller-owned callback, and region
iterator snapshots remain O(R) time/storage for R normalized rectangles.
The public contract was derived only from the pinned official NuGet reference
metadata. Independent focused tests cover exact parameter names, transparent
and crop graph state, shared data ownership, disposed pixmap views, region
operations/iterators, and font behavior. Legacy path iterators and the path-
operation builder are extensible SKObject instances with exact disposal
overrides, while disposed temporary paths continue to preserve geometry already
owned by retained commands. The shared SKObject disposal declarations,
read-only public-disposal policy, and matrix equality parameter metadata also
match the pinned contract. The exact metadata gate advances from 4,063 to 4,125
matches and ratchets missing entries from 159 to 97. No shader,
rendering algorithm, text-shaping boundary, cache policy, or GPU submission path
changed, so the prior cross-engine rendering research and matched performance
evidence remain applicable.
SKRuntimeEffect, its shader/color-filter/blender builders, uniform and child
collections, stack-only uniform values, and typed child values now match the
official SkiaSharp 4.151.0 public metadata. The clean-room parser validates a
top-level main function, records scalar/vector/matrix uniform layout in source
order, separates shader, color-filter, and blender children, and snapshots both
uniform bytes and child references into immutable effect instances. Construction
and lookup are CPU-only; parsing is O(S) for S source characters, uniform
snapshots are O(U) time/storage for U bytes, and child snapshots are O(C)
for C children. Invalid names, sizes, kinds, and sources fail explicitly.
The matched Release benchmark preserves an exact native/ProGPU packing checksum
for float, float2, and float4 uniforms. A preliminary nine-sample Apple M3 Pro
run measured 641.042 ns/op and 584.928 managed bytes for ProGPU versus
2,770.375 ns/op and 968.744 bytes for native (0.231 time ratio). This is a
contract checkpoint rather than the final renderer claim: SkSL-to-WGSL lowering,
child sampling, runtime color-filter execution, and destination-aware blender
execution remain in the active GPU slice and will receive matched three-pair
and Instruments evidence before release.
The design follows the public Skia runtime-effect contract, WGSL, and WebGPU execution and resource models. ProGPU adopts immutable compiled programs and typed byte-packed uniforms, adapts execution to retained WebGPU pipelines, and rejects native source-code reuse, runtime reflection, and CPU pixel fallback.
SKFontVariationAxis, SKFontVariationPositionCoordinate,
SKFontPaletteOverride, and the stack-only SKFontArguments now match the
official 4.151.0 sequential value contracts, mutable properties, readonly
accessors, typed equality, hashing, and operators. SKTypeface now uses the
official SKObject ownership hierarchy and exposes allocation-free span APIs
for variation axes and current positions. Variation cloning maps four-byte axis
tags directly onto ProGPU's existing immutable OpenType instances; unknown axes
are ignored, omitted axes use their defaults, and user coordinates are clamped
and normalized by the existing fvar/avar engine. Array properties allocate
only their documented result, while warmed span queries are O(A) with zero
managed allocation for A axes. Typeface cloning is O(A + R) for R
requested coordinates and preserves distinct font-instance identity for shaping
and retained glyph/cache keys. A thread-safe immutable last-position entry
turns repeated clones into bounded O(R) comparison plus one required wrapper,
without weakening the existing 32-instance normalized-coordinate cache. It
remains CPU-only and cannot initialize WebGPU.
Three alternating Apple M3 Pro Release process pairs, 72 samples per backend,
retained exact semantic checksums. Span queries measured 11.850 ns/op for
ProGPU versus 594.391 ns/op for native (0.020 ratio), both at zero managed
allocation. Repeated clones measured 137.959 versus 31,021.542 ns/op
(0.004 ratio) and 88 versus 112 managed bytes per clone. The value-only
contract measured 3.658 versus 3.654 ns/op with zero allocation, neutral at
timer resolution. Matched Xcode Time Profiler captures measured queries at
11.819 versus 580.759 ns/op and clones at 134.521 versus 31,392.875
ns/op. Allocations plus VM Tracker retained zero bytes per query and 88 versus
112 managed bytes per clone while preserving the same ordering. Metal System
Trace completed both exact-binary workloads with zero target-process command
buffer submissions and no MTLDevice.currentAllocatedSize samples, confirming
that font instance selection does not initialize a GPU.
The clean-room design used the public
SkiaSharp font-arguments contract,
Skia font-argument model,
OpenType fvar and
avar contracts,
DirectWrite axis selection,
Core Text variation descriptors,
and HarfBuzz variation settings.
ProGPU adopts their shared immutable axis/value instance model and the rule that
unspecified axes resolve to defaults. It adapts that model to bounded managed
instance caches and retained WebGPU glyph resources. It rejects per-draw font
mutation and GPU shaping: Skia's
text architecture,
Parley, and
WebRender
all reinforce reusable CPU shaping/layout followed by cached glyph preparation
and GPU composition. Color-palette clone behavior remains an explicit follow-up;
the descriptor contract is present but no incomplete palette renderer is
advertised by this checkpoint.
GRGlFramebufferInfo, GRGlTextureInfo, and GRMtlTextureInfo now match the
official 4.151.0 value surfaces and sequential ABI layouts. OpenGL framebuffer
and texture descriptors retain their unsigned object identifiers and formats
inline, with protection state normalized into the final byte field. The Metal
descriptor retains one native texture handle. Official constructor and property
names, overloads, typed equality, object equality, hashing, operators, readonly
accessors, and the two declared IEquatable<T> interfaces are preserved; the
former accidental aliases and optional-parameter signature have been removed.
These structs are CPU-only borrowed-handle metadata. Construction, mutation,
comparison, and hashing are fixed O(1) work, allocate nothing, do not claim
ownership of the referenced native resource, and cannot initialize GL, Metal,
or WebGPU. ProGPU rendering continues through its typed WebGPU resource model;
these compatibility values do not introduce a second renderer. Independent
tests verify private field order/types, byte protection normalization, official
parameter names, pointer identity, and complete value behavior. Three
alternating Apple M3 Pro Release process pairs retained exact checksums and zero
managed allocations at 0.996 ProGPU/native (2.401 versus 2.410 ns/op).
Matched Time Profiler captures measured 2.373 versus 2.385 ns/op and matched
Allocations captures measured zero bytes per operation. The clean-room contract
uses the public
framebuffer API,
OpenGL texture API,
Metal texture API,
OpenGL framebuffer model,
and Metal resource ownership model.
GRVkAlloc, GRVkImageInfo, GRVkYcbcrComponents, and
GRVkYcbcrConversionInfo now match the complete 4.151.0 metadata contract,
including sequential nested field layouts, byte-backed Boolean transport,
readonly accessors, value equality, hashing, and operators. The obsolete
GrVkYcbcrConversionInfo spelling is retained as one inline current-value
wrapper with exact conversion operators and an intentionally inert obsolete
FormatFeatures property. Allocation metadata includes device memory, size,
offset, flags, backend memory, and its hidden transport byte in official order;
image metadata carries allocation, tiling/layout/format/usage, sample and mip
counts, queue ownership, protection, YCbCr conversion, and sharing mode.
These values describe caller-owned Vulkan resources without creating, mapping,
destroying, or submitting them. All getters, setters, and comparisons are
allocation-free fixed O(1) CPU work and cannot initialize Vulkan or WebGPU.
Field-wise equality is aggressively inlined so both equal values and a
last-field mismatch avoid boxing and reflection while preserving every public
field's observable contribution. Independent tests inspect every private field
type/order and cover full mutation, nested equality, byte normalization, and
legacy/current conversion. Three alternating Apple M3 Pro Release process
pairs retained exact checksums and zero managed allocations at 0.976
ProGPU/native (2.844 versus 2.915 ns/op). Matched Time Profiler and
Allocations captures ranged from 2.784–2.879 for ProGPU and
2.808–2.813 ns/op for native, straddling at sub-nanosecond timer resolution,
with zero bytes per operation. The clean-room contract uses the public
allocation API,
image API,
YCbCr API,
and Vulkan's
sampler-conversion structure
and image-view rules.
GRD3DTextureResourceInfo and GRBackendState now match their complete
4.151.0 contracts. The extensible disposable descriptor retains the borrowed
D3D resource pointer, resource state, DXGI format, mip count, sample count,
quality pattern, and protection flag. Disposal follows the official observable
contract: every call dispatches through the protected virtual hook and leaves
the caller-owned resource metadata intact. The unsigned flags enum preserves
exact None and all-bits All values.
Construction allocates only the required descriptor object; all subsequent
property reads/writes and disposal dispatch are fixed O(1) CPU work with no
incremental allocation, COM call, resource transition, device creation, or
WebGPU initialization. Independent tests cover defaults, all mutable values,
post-disposal retention, repeated virtual dispatch, enum width, and flags.
Three alternating Apple M3 Pro Release process pairs retained exact checksums
at 0.985 ProGPU/native (1.506 versus 1.528 ns/op) with the same amortized
0.00048 bytes per operation from the one required descriptor object per
100,000-operation sample. Matched Time Profiler/Allocations captures measured
1.495–1.506 versus 1.516–1.524 ns/op with identical allocation. The
clean-room contract uses the public
D3D resource-info API,
backend-state API,
and Microsoft's
D3D12 resource-state model.
GRBackendTexture and GRBackendRenderTarget now derive from the official
SKObject ownership base and match the public GL, Vulkan, Metal, and Direct3D
constructor, backend, dimensions, size, rectangle, validity, mip, sample,
stencil, GL-query, and protected-disposal contracts. The wrappers retain only
typed descriptor metadata and one synthetic managed wrapper identity; they
never create, upload, transition, submit, or destroy the caller-owned native
resource. The existing ProGPU GpuTexture constructors remain explicit Dawn
extensions so Avalonia, LibreWPF, LibreWinForms, and media composition can
share a typed zero-copy WebGPU texture without reflection or an intermediate
pixel copy.
All property and descriptor queries are fixed O(1) CPU work and allocate
nothing after wrapper construction. GL TryGet overloads fail closed with a
default descriptor for non-GL backends. Disposal invalidates only the wrapper
and does not dispose the supplied D3D descriptor or WebGPU/native texture.
Independent tests cover the exact non-sealed hierarchy, declared protected
overrides, constructor parameter names, backend classification, immutable
geometry, GL success/failure, mip/sample propagation, invalidation, and
borrowed ownership across all five backend identities.
Three alternating Apple M3 Pro Release process pairs retained the exact native
checksum. The combined metadata query measured 2.340 ns/op for ProGPU versus
14.863 ns/op for native (0.157 ratio), with amortized construction at
0.001 versus 0.002 bytes per operation. Matched Xcode Time Profiler runs
measured 2.352 versus 14.737 ns/op; matched Allocations runs measured
2.329 versus 14.540 ns/op with the same bounded construction allocation.
The clean-room contract uses the public
backend texture,
backend render target,
GL texture query,
GL framebuffer query,
Vulkan external-memory rules,
D3D12 resource ownership,
and WebGPU object ownership.
SKPMColor now matches the complete 4.151.0 public metadata contract. Scalar
premultiply and unpremultiply are allocation-free fixed-work operations; array
overloads allocate exactly one result array and process N colors in O(N)
time with O(1) auxiliary storage. The implementation retains the official
platform-native N32 packing (RGBA on Apple targets and BGRA on the official
Windows/Linux assets), rounded divide-by-255 premultiplication, and a generated
read-only 8.24 reciprocal table for deterministic unpremultiplication without
per-channel division. It is CPU-only and cannot initialize WebGPU.
Independent tests cover packed identity, logical channels, formatting,
operators, allocation ownership, transparent input, and component bounds. The
matched benchmark exhaustively checks every alpha/component pair and separately
measures scalar and 64-element array overloads against the official package.
On the recorded Apple M3 Pro Release run, all four semantic checksums and
managed allocations matched exactly. ProGPU/native median ratios were 1.014
for scalar premultiply, 1.121 for scalar unpremultiply, 1.080 for the
64-element premultiply array, and 1.183 for the unpremultiply array. These
small but repeatable remaining CPU gaps are retained as optimization work; this
checkpoint establishes parity without claiming a performance win.
The design used the public
SkiaSharp contract
and Skia's documented
premultiplied color and
unpremultiply scale
contracts. No foreign implementation code, source layout, or helper structure
was incorporated.
SKFourByteTag now matches all 18 entries in its 4.151.0 metadata contract.
The four-byte readonly value uses OpenType's big-endian display order, preserves
packed uint identity, pads non-empty short tags with trailing spaces, truncates
long tags, and preserves native zero identity for null or empty input. Character
construction narrows each UTF-16 code unit to its low byte, matching the
observable API behavior without validating font-table policy at this value
boundary.
Construction, parsing, equality, hashing, and conversions are allocation-free
fixed-work operations. Formatting allocates only its four-character result.
Matched Release checksums cover string/span parsing, construction, conversion,
and formatting. Across three alternating Apple M3 Pro process pairs, value
operations measured 1.127 ProGPU/native and formatting measured 0.127, with
32 versus 280 managed bytes per formatted tag. These local figures are evidence for the
slice, not a cross-platform claim. The clean-room design follows the
OpenType Tag data type
and the public
SkiaSharp parsing contract.
SKSwizzle now matches all six entries in the 4.151.0 public metadata
contract. The reusable PixelChannelSwizzler core operates on tightly packed
four-byte pixels in O(N) time and O(1) auxiliary storage, supports bounded
overlapping copies, and never initializes WebGPU. On ARM64, copy and in-place
paths use fixed 32-bit and 16-bit byte reversals followed by a mask select,
avoiding both an intermediate buffer and table-lookup stalls. Other targets use the
portable hardware-accelerated Vector128 shuffle and scalar tails.
Independent tests cover in-place and copy overloads, pointer entry points,
count clamping, overlap direction, stable replay allocations, and incomplete
trailing pixels. Valid complete-pixel inputs match the official behavior. The
span-only overload deliberately preserves an incomplete trailing pixel rather
than allowing the official wrapper's observable out-of-bounds native access;
this is a memory-safety improvement outside the documented complete-pixel
contract. Three alternating Apple M3 Pro Release process pairs retained equal
managed allocations and exact semantic checksums. Copy measured 0.962
ProGPU/native and in-place measured 1.213; the latter remains an explicit CPU
optimization target. Matched Time Profiler and Allocations traces from the same
Release binaries retained stable checksums and 0.824/0.412 managed bytes per
operation for both implementations. The raw distributions, trace bundles, and
exported sample tables remain diagnostic evidence rather than a cross-platform
claim.
The design follows the public
SkiaSharp swizzle contract
and Skia's documented
RGBA/BGRA transform.
SkiaSharpVersion now matches all four entries in its 4.151.0 metadata
contract. The clean-room shim reports the observed 151.0 native and minimum
compatibility levels and succeeds in both throwing and non-throwing check modes
because ProGPU supplies the complete implementation without loading a separate
native Skia binary. Both properties share one immutable process-wide Version,
so repeated queries are allocation-free fixed O(1) operations.
Independent tests cover exact version values, compatibility modes, stable
identity, and one million allocation-free queries. Three alternating Apple M3
Pro Release process pairs produced exact semantic checksums; ProGPU measured
0.066 of native time and 0 versus 32 managed bytes per operation. The
clean-room behavior follows the public
SkiaSharpVersion contract
and retains no native-library discovery or loader side effects.
SkiaExtensions now matches all 18 entries in the 4.151.0 metadata contract,
replacing the former non-official SKGlExtensions identity. Pixel-geometry
classification, byte and bit-shift sizes, alpha compatibility, and OpenGL sized
formats cover all 29 declared color types. Unknown declared formats retain
their documented zero values, while out-of-range enum values fail with the
official colorType argument boundary. SKImageInfo now delegates to the same
single format-size contract instead of retaining a second mapping.
Every valid query is allocation-free fixed O(1) CPU work and cannot initialize
WebGPU. Independent tests exhaust the color-type and alpha-type matrices,
geometry categories, GL mappings, invalid enums, and one million stable queries.
The source-built Avalonia.Skia projects continue to compile for net8 and net10
against the official extension identity. Three alternating Apple M3 Pro Release
process pairs produced exact checksums and zero allocations; ProGPU measured
0.683 of native time for the combined workload. Matched Time Profiler and
Allocations captures from the same binaries preserved that ordering, exact
checksums, and zero managed bytes per operation. The clean-room contract uses
the public
SkiaExtensions API,
Skia color-type documentation, and
Khronos sized internal formats.
StringUtilities now matches all ten entries in the 4.151.0 metadata contract.
UTF-8, little-endian UTF-16, and little-endian UTF-32 conversion use replacement
fallbacks, return exactly one owned byte array or string, expose bounded array,
span, slice, and pointer decode overloads, and reject glyph-ID or out-of-range
encodings before conversion. Encoding is O(C + B) and decoding is O(B + C)
for C UTF-16 code units and B encoded bytes, with only the caller-owned
result allocation and no WebGPU initialization.
GetUnicodeCharacterCode validates exactly one complete Unicode scalar and
returns it allocation-free for every supported UTF encoding. This intentionally
corrects the official 4.151 wrapper's observable short-buffer failure for
ordinary UTF-8/UTF-16 characters while retaining the documented API contract;
incomplete surrogates and multiple scalars fail before returning partial data.
Independent tests cover exact byte forms, supplementary scalars, replacement
fallbacks, pointer/slice boundaries, null/empty ownership, invalid encodings,
and glyph-ID rejection. Three alternating Apple M3 Pro Release process pairs
produced exact checksums for matched workloads: roundtrip conversion measured
0.960 ProGPU/native with equal 290.651 managed bytes per operation, while
the scalar query measured 0.041 and 0 versus 256 bytes. The clean-room
Matched Time Profiler and Allocations traces from the same Release binaries
retained the checksum, allocation, and timing ordering. The clean-room design
follows the public
StringUtilities contract,
Unicode encoding forms,
and .NET Encoding contract.
SKColorSpacePrimaries now matches all 35 entries in the 4.151.0 metadata
contract. Eight mutable inline floats retain the red, green, blue, and white
chromaticities; constructors and the public Values snapshot preserve caller
ownership. Conversion solves one homogeneous 3x3 primary matrix and applies
Bradford chromatic adaptation into the ICC D50 profile-connection space. The
general conversion is fixed O(1) CPU work with no heap allocation or WebGPU
initialization. Degenerate matrices and non-finite or out-of-unit coordinates
fail transactionally with an empty result.
The common sRGB and Display P3/D65 combinations use immutable matrices computed
from the same public chromaticities and D50 model. This keeps the dominant path
at fixed comparisons plus one inline struct copy without weakening arbitrary
gamut support. Independent tests cover value ownership, every mutable scalar,
equality, invalid and degenerate inputs, a zero-y boundary primary, and sRGB/P3
conversion. Three alternating Apple M3 Pro Release process pairs retained exact
semantic checksums and zero managed allocations: the common-gamut workload
measured 0.172 ProGPU/native (6.001 versus 34.830 ns/op). Matched Time
Profiler and Allocations captures from the same binaries measured 5.904–
6.022 versus 33.791–33.861 ns/op with the same checksum and zero bytes per
operation. The clean-room design follows the public
SkiaSharp primaries contract,
Skia public color-space contract,
and the
ICC.1:2022 D50/Bradford model.
SKCodecFrameInfo now matches the official sequential layout as well as its
existing public value behavior. Its two Boolean properties use normalized
one-byte storage in the declared native field order, preserving the compact
codec interop contract without exposing the storage fields. All eight public
properties, equality, hashing, and operators remain allocation-free fixed
O(1) CPU operations and do not initialize a decoder or WebGPU.
Independent tests inspect the compiled private layout, verify byte rather than
managed-Boolean storage, and exercise property normalization and full-value
equality. Three alternating Apple M3 Pro Release process pairs retained exact
checksums and zero allocations at 1.001 ProGPU/native (1.255 versus 1.254
ns/op), which is performance-neutral at timer resolution. Matched Time Profiler
and Allocations captures retained exact checksums and zero bytes per operation;
their medians ranged from 1.220–1.291 ns/op for ProGPU and 1.196–1.226
ns/op for native. The clean-room contract follows the public
SKCodecFrameInfo API
and the official package's ECMA-335 sequential field metadata.
SKJpegEncoderOptions, SKPngEncoderOptions, and SKDocumentXpsOptions now
match their official sequential layouts in addition to retaining their existing
public value contracts. JPEG keeps its three value fields followed by zeroed
metadata pointer/length/origin transport slots; PNG keeps its filter and level
followed by three zeroed native pointer slots; XPS uses one float and a
normalized byte-backed Boolean. The private transport fields are never exposed,
dereferenced, or used to add an external encoder dependency.
Construction, property access, equality, and hashing remain fixed O(1) CPU
work with zero allocation and no codec or WebGPU initialization. Independent
tests inspect the compiled private field order/types and verify all public
values. Three alternating Apple M3 Pro Release process pairs retained exact
checksums and zero allocations at 1.024 ProGPU/native (1.261 versus 1.232
ns/op), within timer noise for the combined value workload. Matched Time
Profiler and Allocations captures retained the exact checksum and zero bytes;
ProGPU measured 1.169–1.217 versus native 1.190–1.221 ns/op. The
clean-room contract follows the public
JPEG options API,
PNG options API,
XPS options API,
and the pinned package's ECMA-335 sequential field metadata.
The point, size, rectangle, color, and rounded-rectangle families now match 41
additional entries in the official 4.151.0 reference contract. This checkpoint
preserves the existing fixed-work value algorithms while aligning official
parameter metadata, adds allocation-free ReadOnlySpan<char> color parsing,
and makes SKRoundRect participate in the official SKObject ownership and
idempotent-disposal hierarchy. Primitive arithmetic and parsing remain O(1)
CPU work with no WebGPU initialization; rounded-rectangle construction owns one
bounded four-corner array and one managed handle, with no native resource.
The clean-room contract was derived from the public
SkiaSharp primitive API documentation,
SKColor parsing API,
SKRoundRect API,
and the pinned package's ECMA-335 public reference metadata. Independent tests
cover all newly aligned parameter names, span parsing output and steady-state
allocation, and the SKObject handle lifetime. Repeatable matched workloads
exercise point arithmetic, span parsing, and rounded-rectangle construction and
disposal against the official package. Three alternating Apple M3 Pro Release
process pairs retained exact checksums: canonical span parsing measured 0.491
ProGPU/native (11.358 versus 23.147 ns/op) with zero allocation, while
rounded-rectangle lifetime measured 0.535 (44.656 versus 83.425 ns/op)
with 120 versus 80 managed bytes per owned instance. The extra 40 bytes are
the managed handle/lifetime state required by the official SKObject contract.
Matched Time Profiler captures measured parsing at 11.185 versus 22.316
ns/op and rounded-rectangle lifetime at 44.827 versus 86.215 ns/op;
Allocations retained the same zero/120 versus zero/80 byte ordering. Metal
System Trace exported zero target command-buffer, device-allocation, and Metal
resource-allocation rows for both CPU-only binaries.
The pinned Svg.Skia 03f64b67badfca9fca216dc25896d0c0ee04e7b7
validation improved with the transformed-stroke slice: native W3C reported 530
passed and 3 skipped; the ProGPU raw lane reported 486 passed, 44 reviewed known
differences, and 3 skipped after filters-overview-02-b reached native image
parity; the resvg lane reported 927 passed and 37 intentional skips; and the
remaining suite passed 1,147 of 1,147 tests. The repository's parity verifier
accepted the complete reviewed-difference inventory.
SKTypeface and SKPathBuilder now match 16 additional entries in the
official 4.151.0 contract. Typeface cloning combines a collection face,
variation coordinates, a CPAL palette, and caller overrides into one immutable
font instance. ProGPU.Text owns the package-neutral FontPaletteOverride and
TtfFont.WithColorPalette primitive so WinUI, WPF, WinForms, and the SkiaSharp
shim share the same color-glyph path. The first selected non-default palette is
O(B + A + P) time and O(B + P) storage for font bytes B, variation axes
A, and palette entries P; non-color fonts and the default palette reuse the
font in O(1). Repeated variation instances continue to use the bounded
32-entry normalized-coordinate cache. SKPathBuilder changes in this slice are
metadata-only delegates and preserve its retained analytic geometry behavior.
The clean-room design follows the public
SKFontArguments contract,
SKTypeface clone contract,
and the authoritative
OpenType CPAL table
and COLR table
formats. It adopts CPAL's base-zero palette and entry indices, contiguous BGRA
records, unpremultiplied sRGB values, and palette-zero fallback; it adapts them
to immutable linear-float render colors and rejects out-of-range override
entries without mutating the source typeface. Independent tests cover combined
collection/variation/palette arguments, non-color reuse, official legacy
parameter names, and path-builder ownership metadata. The matched
font-arguments-clone workload uses the same Inter variable-font bytes and
semantic checksum in both binaries.
Three alternating Apple M3 Pro Release process pairs measured the combined
arguments clone at 0.005 ProGPU/native (156.791 versus 30,604.417 ns/op)
and 88 versus 112 managed bytes per operation. This workload exercises the
bounded repeated-instance path; a first non-default CPAL materialization is
reported separately because it necessarily copies font storage and palette
records. Matched Time Profiler captures measured 168.188 versus 30,561.271
ns/op; Allocations retained 88 versus 112 bytes. Metal System Trace exported
zero target command-buffer, device-allocation, and resource-allocation rows for
both CPU-only binaries.
The official SKGraphics, SKTraceMemoryDump, GRGlBackendState, and
SKBlender.CreateArithmetic contracts close 41 additional 4.151.0 metadata
entries. Cache budgets use atomic process-wide values; setters return the prior
budget, reads are fixed O(1), and the compatibility counters and dump callbacks
do not initialize WebGPU. Purge entry points are safe idempotent boundaries for
the shim's process caches. GRGlBackendState preserves the official 16-bit
OpenGL invalidation mask exactly, while ProGPU's WebGPU backend continues to use
its typed resource ownership instead of interpreting GL state bits.
Independent tests cover every state-mask group, atomic budget round trips,
negative-budget rejection, cache accounting, and protected memory-dump
callbacks. The repeatable graphics-cache-controls workload performs two
atomic setter/getter pairs per operation with identical native and ProGPU
checksums and zero managed allocation. The design follows the public
SKGraphics API,
SKTraceMemoryDump API,
and the pinned package's ECMA-335 enum and method metadata.
Three alternating Apple M3 Pro Release process pairs measured 0.131
ProGPU/native (2.373 versus 18.165 ns/op), with zero allocation. Matched
Time Profiler captures measured 2.408 versus 17.892 ns/op; Allocations
retained zero bytes per operation. Metal System Trace exported zero target
command-buffer, device-allocation, and resource-allocation rows for both
CPU-only binaries.
PlatformConfiguration, IPlatformLock, PlatformLock, and
SKAutoCoInitialize close 26 additional entries in the official 4.151.0
metadata contract. Runtime flags use the platform and process-architecture
information supplied by .NET, while the mutable Linux flavor remains an atomic
process-wide compatibility setting. The default lock is a typed
ReaderWriterLockSlim adapter supporting read, upgradeable-read, write, and
recursive entry without reflection or per-entry allocation. Lock entry and
exit are fixed O(1) work when uncontended and use the runtime lock's bounded
per-instance state; contention has scheduler-dependent wait time. On Windows,
SKAutoCoInitialize balances each successful multithreaded-apartment
initialization, including S_FALSE, with exactly one CoUninitialize call.
Other platforms use the same idempotent object lifetime without loading a
Windows library or initializing WebGPU.
The clean-room design follows the public
RuntimeInformation contract,
ReaderWriterLockSlim contract,
CoInitializeEx contract,
and CoUninitialize balance rule.
Independent tests cover platform-flag consistency, factory replacement,
recursive read, upgradeable/read/write modes, one million steady-state lock
pairs with zero managed allocation, and idempotent COM lifetime behavior.
Three alternating Apple M3 Pro Release process pairs retained the exact native
checksum. The read-lock pair measured 1.194 ProGPU/native (9.241 versus
7.737 ns/op); both harnesses reported only the same amortized 0.0012 B/op
one-time measurement overhead. Matched Time Profiler captures measured 9.194
versus 7.757 ns/op, while Allocations captures measured 9.171 versus
7.907 ns/op with the same checksum and allocation result. Metal System Trace
exported zero target command-buffer, device-allocation, and resource-allocation
rows for both CPU-only binaries.
SKSurface and SKSurfaceReleaseDelegate now close all 65 missing entries in
their official 4.151.0 contracts. Surfaces participate in the shared
SKObject lifetime, snapshot immutable surface properties, retain the typed
GRRecordingContext, and expose the complete raster, recording-context,
backend-texture, backend-render-target, sample-count, origin, color-space, and
mipmap overload families. Caller WebGPU textures and external pixel pointers
remain borrowed and zero-copy. External release callbacks run exactly once.
Ordinary WebGPU surfaces allocate no CPU mirror until PeekPixels is requested;
the first peek performs one explicit readback and retains a stable pointer,
while later GPU flushes update that view. A bounded snapshot performs a direct
texture-to-texture region copy and stays GPU-backed until an explicit readback.
Wrapped-surface creation is O(1) CPU work and storage. Rendering remains
O(C + P) for retained commands C and affected pixels P. A snapshot uses
O(1) command-encoding work, O(P) GPU bandwidth, and one destination texture;
readback and first peek use O(P) transfer/conversion work. Null surfaces own no
texture or GPU context, do not initialize WebGPU, and discard retained commands
at flush. Independent tests cover the
official ownership hierarchy, surface-property isolation, typed contexts,
every overload family through the metadata verifier, stable lazy CPU views,
one-shot release callbacks, null surfaces, bounded GPU snapshots, and existing
backend target/origin/readback behavior.
The clean-room architecture review used Skia's public surface contract and canvas/surface model, Direct2D's device-dependent render-target model, Win2D's incremental offscreen target contract, WebRender's display-list, scene, frame, and GPU submission split, and Vello's explicit wgpu scene-to-texture pipeline. ProGPU adopts explicit device ownership, retained commands, incremental target contents, immutable snapshots, and GPU-native copies; it rejects API-specific GL/Vulkan/Metal handle interpretation in favor of typed WebGPU resources. The required text-stack review also covered Skia's text architecture, DirectWrite's layout/render separation, and HarfBuzz's buffer shaping contract. Those CPU-reusable shaping/layout boundaries remain unchanged by this surface slice.
Three alternating Apple M3 Pro Release process pairs retained the exact native
checksum for 32-by-32 bounded snapshots from a stable 64-by-64 surface. Native
raster copy-on-write measured 453.540 ns/op and 120.8 B/op; ProGPU's current
explicit WebGPU copy-and-submit path measured 65,455.835 ns/op and 442
B/op. Removing per-snapshot native label marshalling reduced the ProGPU managed
cost from 666.4 to 442 B/op. Matched Time Profiler captures measured
437.705 versus 67,390.415 ns/op, and Allocations captures measured
2,217.085 versus 95,860.205 ns/op with the same byte counts. Metal System
Trace correctly reported no native raster work and recorded 3,275 ProGPU
command-buffer rows, 4,204 current-allocation rows, and 175 resource-allocation
rows. This is a documented
performance blocker for the final parity release: repeated immutable snapshots
still need deferred/batched submission and shared copy-on-write texture
ownership before ProGPU can meet the goal's matched native latency and
allocation criterion.
SKImage, SKImageRasterReleaseDelegate, and
SKImageTextureReleaseDelegate now close all 65 missing entries in their
official 4.151.0 contracts. The complete factory surface covers raster
creation, immutable pixel copies, caller-owned pixmaps, encoded data and files,
pictures, borrowed and adopted backend textures, recording contexts, color and
alpha metadata, release callbacks, raster/texture conversion, filter
application, shaders, and subsets. Caller pixel and texture release callbacks
run exactly once with their original pointer/context. Encoded images retain an
independent encoded snapshot. Raster PeekPixels materializes one stable
pinned CPU view; a GPU-backed image does not silently claim a CPU pointer.
Contained subsets are immutable O(1) texture-region views. One atomic
reference retains the source texture storage and the view composes bounded
CPU-pixel and GPU-texture origins; creating, nesting, or disposing a subset
performs no pixel copy, command encoding, queue submission, or GPU allocation.
The final owner releases an adopted texture and invokes its borrowed-texture
release callback exactly once, so a subset remains valid after its parent is
disposed. Raster provenance remains observable as
raster and texture provenance remains observable as texture; texture-backed
subsets require the matching recording context.
Region materialization is deferred to the operation that requires an independent
resource. Same-context texture conversion and retained image drawing issue one
typed base-level WebGPU rectangle copy, and texture conversion can generate
mip levels afterward. A CPU read requests only the view rectangle; immutable
raster-backed views copy directly from their retained row-stride storage, while
GPU-only views use one bounded readback texture. Cross-context conversion uses
one explicit tight upload because WebGPU resources cannot be copied between
devices. Filter application runs through ProGPU's retained WebGPU filter graph
and clips its output to the caller's expected device bounds. Creation and
wrapping validation are O(1) apart from required pixel ownership; view
creation is O(1) time/storage and one managed wrapper; materialization and
cross-device transfers are O(P) bandwidth and storage for P view pixels.
Independent tests cover stride-aware immutable copies, stable raster views, encoded ownership, exact-once raster/texture callbacks, borrowed versus adopted textures, shared and nested region views, parent-before-child disposal, contained GPU rectangle materialization, invalid subsets, mip generation, and filtered output bounds. The focused image/surface contract selection passes 87 tests. The metadata verifier at this image checkpoint reported 4,222 official entries, 4,933 candidate entries, 3,756 exact matches, 466 missing entries, and 1,177 documented extensions. The isolated package gate also produced the runtime and Avalonia 11/12 integration packages in a fresh feed, then restored and built the package-only Avalonia consumer with zero warnings or errors.
The clean-room architecture uses Skia's public image contract, image factory contract, and filter-bounds model, WebGPU's texture-copy validation and ordering model, Direct2D's source-rectangle bitmap model, Win2D's CanvasBitmap contract, WebRender's external-image and frame split, and Vello's explicit wgpu scene-to-texture pipeline. The Skia/SkParagraph, DirectWrite/Direct2D, Win2D, WebRender, Vello/Parley, and HarfBuzz shaping/layout review recorded by the surface checkpoint remains unchanged: image ownership does not move Unicode/OpenType shaping onto the GPU.
Three alternating Apple M3 Pro Release process pairs retained the exact native
checksum for 32-by-32 subsets of a stable 64-by-64 image. The final shared-view
implementation measured 399.790 ns/op and 402.64 managed B/op versus native
raster copy-on-write at 675.210 ns/op and 106.08 B/op (0.592 latency
ratio). Relative to the previous ProGPU immediate-copy result, this reduces
median latency from 38,778.335 ns/op by 99.0% and managed allocation from
722.08 B/op by 44.2%. ProGPU's remaining managed-byte difference is its
visible managed image/view ownership while the native counter excludes Skia's
native object allocation, so no total-memory advantage is inferred.
Matched final-binary Time Profiler, Allocations plus VM Tracker, and Metal
System Trace captures all completed. For the same workload, Xcode's persistent
native heap plus anonymous VM fell from 165,785,280 to 110,526,736 bytes,
and total native heap bytes fell from 728,675,488 to 196,383,680 bytes.
The former Metal trace exported 6,429 command-buffer submission rows, 4,509
currentAllocatedSize rows, and 268 resource-allocation rows; the final trace
contains no modeled target Metal track because subset creation no longer
records or submits GPU work. These whole-process Instruments numbers include
runtime/device startup and are correlated evidence rather than per-operation
allocation claims. Before/final raw traces, TOCs, exported tables, and exact-run
JSON are retained under artifacts/performance/skiasharp-image-api-instruments
and artifacts/performance/skiasharp-image-subset-zero-copy-instruments.
The benchmark workflow now installs the same Linux Vulkan prerequisites as the
main build and resolves the packaged RID-native WebGPU directory on Linux,
macOS, and Windows. This fixes the prior Ubuntu libwgpu_native loader failure
without skipping the GPU workload or relaxing comparison evidence.
The source-built Avalonia 12 WriteableBitmapImpl creates an immutable image
with SKImage.FromPixels(info, address, rowBytes) whenever its writable pixel
version changes, then reuses that image across draws. ProGPU now copies common
RGBA, sRGBA, and BGRA rows directly into one tight immutable portable snapshot
and uploads that same snapshot to WebGPU. The former temporary SKBitmap
wrapper and its second row walk are gone; arbitrary supported formats keep the
conservative conversion fallback. Snapshot work remains O(P) time and
storage for P pixels because the public pointer is caller-owned and the image
must remain immutable after the writable framebuffer changes.
Whole images drawn in the same WebGPU device now cross the retained-command
boundary through IProGpuContextTextureLeaseSource. The first draw records one
bounded lifetime lease and every subsequent draw in that context reuses the
same GpuTexture, texture view, and bindable identity. Disposal of the public
SKImage releases its ownership but cannot destroy the texture while a
deferred context or picture still holds a lease. Subsets, cross-device images,
and mipmap generation retain their normalized materialization paths. This makes
ordinary same-device recording O(C) command work for C draws with one GPU
texture and one lease, rather than O(C * P) texture allocation and copy
bandwidth.
The clean-room design follows Skia's public immutable image contract, WebGPU's texture ownership and copy model, Direct2D's device-context bitmap drawing contract, Win2D's CanvasBitmap contract, WebRender's external-image frame split, and Vello's explicit wgpu scene-to-texture pipeline. ProGPU adopts immutable CPU ownership at the public pointer boundary and typed same-device leases at the deferred GPU boundary; it rejects borrowed pointer lifetime assumptions, per-draw GPU copies, reflection, and backend-specific public handles. Text shaping remains unchanged at the reusable CPU-result boundary established by SkParagraph, DirectWrite, Parley, and HarfBuzz.
On the Apple M3 Pro Release baseline, the 16-by-16 Avalonia snapshot workload
improved from 13,356.445 to 10,934.730 ns/op and from 1,568 to 1,424
managed B/op with the exact native checksum. The new 1,000-draw retained-picture
workload isolates reuse of that immutable image: replacing one GPU texture copy
per draw with one lifetime lease reduced ProGPU from 69,164.500 to 608.354
ns/draw and from 2,831.500 to 2,486.000 managed B/draw. Native measured
48.479 ns/draw and 2 managed B/draw because its retained command storage is
native and outside the managed counter. The remaining ProGPU command-storage
and snapshot gaps are explicit optimization targets; these shared-machine
figures establish the direction and do not claim final cross-platform parity.
Matched final-binary macOS profiling compared exact pre-lease commit
1c60239b with exact candidate 79d86548 on the same Apple M3 Pro, macOS
26.4.1, and .NET 10.0.5 workload. Time Profiler measured 327,088.874 versus
816.041 median ns/draw; Allocations plus VM Tracker measured 82,988.745
versus 929.165; Metal System Trace measured 49,527.290 versus 797.290;
and EventPipe measured 60,615.875 versus 627.041. EventPipe retained the
exact checksum while managed allocation fell from 2,831 to 2,486 B/draw
(12.2%). Profiler overhead perturbs the absolute latency, so the ordinary
Release process numbers above remain the throughput result and these matched
captures provide causal evidence.
The Metal capture reduced target resource-allocation rows from 188 to 53
and target application command-buffer submission rows from 5,627 to zero.
The baseline target stack contains WebGPU copy_texture_to_texture; the
candidate target stack does not. Both captures reported zero Metal
command-buffer errors, compiler spills, and hang risks. Completion and
currentAllocatedSize row counts include process/device sampling and are not
interpreted as bytes or per-draw totals. The Allocations template did not
export a native retained-byte table on this Xcode version, so no unsupported
native-memory claim is made. Compact results are recorded here; the 221 MiB of
raw trace and EventPipe data, temporary publishes, packages, and exact-baseline
worktree were removed after the audit.
SKCanvas now closes all 45 missing entries in its official 4.151.0 owner
contract plus the two missing readonly matrix-parameter attributes. It derives
from SKObject, owns one stable compatibility handle, and clears that handle
through the shared idempotent lifetime. Official parameter names, optional
values, and compile-time-obsolete text overloads now match the reference
metadata. Rectangle and path clips use the official non-antialiased default;
explicit antialias choices continue through the same typed retained API.
Bitmap, image, surface, lattice, nine-patch, picture, primitive, and text
overloads remain thin routes into the existing retained WebGPU command graph.
An empty saved clip scope is now removed transactionally on restore instead of
retaining a large general push/pop command pair. This peephole is fixed O(1)
time and storage and is valid only when no command was recorded after the push;
a scope containing drawing retains its balanced push, content, and pop. After
one capacity warmup, 100,000 empty save/clip/restore cycles allocate exactly
zero managed bytes and leave no commands. Drawn clips remain O(C) retained
storage for commands C; lattice construction remains O((X + 1)(Y + 1))
patch work for X and Y divider counts and submits those patches through one
retained image source rather than uploading once per patch.
The clean-room design follows Skia's public canvas and lattice contract, Direct2D's device-context bitmap contract, Win2D's retained offscreen drawing model, WebRender's display-list, spatial-tree, clip-tree, and frame split, and Vello's wgpu scene-to-texture architecture. ProGPU adopts retained draw routing, separate transform/clip state, one image source per lattice, and GPU submission after scene recording; it rejects immediate CPU rasterization and API-specific native-handle branches. The required text review used Skia's text architecture, DirectWrite's layout/render separation, and HarfBuzz's buffer shaping contract. Canvas overload alignment therefore leaves reusable shaping and glyph placement on the existing CPU-result boundary and changes only retained draw routing.
The isolated package gate produced all runtime and Avalonia 11/12 integration packages in a fresh feed, then restored and built the package-only Avalonia consumer with zero warnings or errors.
The exact-checksum Apple M3 Pro Release workload performs 10,000
save/scale/concat/clip/restore cycles per sample. Before empty-scope elision,
ProGPU measured 3,839.419 ns/op and 6,979.893 B/op. Afterward it measured
679.500 ns/op and 0.7792 amortized B/op, versus native 213.525 ns/op and
0.1752 B/op: an 82.3% ProGPU latency reduction and more than 99.98% allocation
reduction, while the remaining direct-run latency is still a documented
optimization target. Matched Time Profiler captures measured 190.623 versus
195.177 ns/op; Allocations captures 184.444 versus 198.840 ns/op with
the same managed allocation counts; Metal System Trace captures 194.098
versus 203.783 ns/op. Both Metal traces export zero target command-buffer,
device-allocation, and resource-allocation rows, confirming this state-only
path does not initialize WebGPU. Raw traces, TOCs, exported Metal tables, and
exact-run JSON are retained under
artifacts/performance/skiasharp-canvas-api-instruments.
SKShader now closes all 33 entries that remained missing from its official
4.151.0 owner contract. The public bitmap, image, and picture factories use the
official src, tmx, tmy, and tile parameter names; float-color gradient
factories consistently expose colorspace; and compose/filter wrappers expose
the official shaderA, shaderB, and filter names. This is metadata parity
over the existing original retained implementation, not a native Skia call or
source port.
Color, gradient, picture, image, local-matrix, color-filter, noise, and composed
shader nodes keep immutable ownership. Gradient colors and offsets are
converted and clamped once in O(S) time and O(S) retained storage for S
stops. Every ToBrush call returns an independent compact stop array so caller
mutation cannot alter the shader. Linear-color spaces select scRGB-linear
interpolation; tile modes and the inverse local matrix survive through linear,
radial, two-point conical, and sweep gradients. Image shaders continue to own
one retained texture snapshot with explicit nearest/linear/mipmap/cubic
sampling rather than uploading once per tile. Actual gradient evaluation,
tiled texture sampling, composition, and post-filter work remain in ProGPU's
retained WebGPU render/compute paths; factory construction is intentionally
CPU-only and does not initialize WebGPU.
The clean-room design follows Skia's public shader contract and gradient degeneracy rules, Direct2D's solid, gradient, image, and bitmap brush model, Win2D's color-space-aware linear gradient contract, and WebGPU's immutable samplers, addressing, filtering, and external-texture model. ProGPU adopts immutable factory state, explicit interpolation and addressing, and deferred GPU evaluation; it rejects render-target-bound public resources, per-tile uploads, CPU raster fallbacks, and backend-specific public handles. The required Skia/SkParagraph, DirectWrite/Direct2D, Win2D, WebRender, Vello/Parley, and HarfBuzz review recorded by the canvas checkpoint remains the shaping/layout boundary: this shader-only slice does not change text shaping, glyph caching, or CPU layout reuse.
The exact-checksum Apple M3 Pro Release workload creates and disposes linear,
radial, sweep, and two-point conical float-color gradients with three stops,
different tile modes, one sRGB color space, and one local matrix.
ProGPU measured 850.979 ns/op and 1,448 managed B/op versus native
2,353.083 ns/op and 416 managed B/op (0.362 latency ratio). The extra
managed bytes are ProGPU's visible immutable stop/closure ownership, whereas
the native harness does not count Skia's native allocations; reducing the
managed representation remains an optimization target and no total-memory
advantage is claimed from this counter alone. Matched Time Profiler captures
measured 853.521 versus 2,389.021 ns/op, Allocations captures 838.396
versus 5,591.000 ns/op with the same 1,448 versus 416 managed B/op, and
Metal System Trace captures 846.980 versus 2,414.667 ns/op. Both Metal
traces export zero target command-buffer, current-allocation-size, and resource
allocation rows, confirming factory construction remains CPU-only. Raw traces,
TOCs, exported Metal tables, and exact-run JSON are retained under
artifacts/performance/skiasharp-shader-api-instruments.
The GRContext cluster now closes 75 official 4.151.0 metadata entries across
the direct recording context, its options, GL interface, Vulkan extensions,
typed GL/Vulkan/Metal/Direct3D descriptors, procedure-address delegates, and
their disposal contracts. Backend descriptors are CPU-only borrowed-handle
DTOs. Their disposal never releases caller-owned API objects, while
GRGlInterface and GRVkExtensions own only their managed compatibility
handles and immutable extension metadata.
Every public factory maps to ProGPU's process-wide typed WgpuContext; the
foreign GL, Vulkan, Metal, or Direct3D descriptor selects a compatibility entry
point but is never exposed as ProGPU's device ownership. A GRContext wrapper
does not own that shared WebGPU device. Abandonment is local and idempotent, so
abandoning or disposing one wrapper cannot invalidate another wrapper or an
Avalonia/WinUI/WPF/WinForms host sharing the device. Flush and asynchronous
Submit poll the queue without an idle wait because ProGPU submits recorded
render/compute work at the owning surface/compositor boundary; synchronous
submission uses the existing device wait. Reset is an O(1) state-coherency
acknowledgement because WebGPU tracks explicit immutable pipeline and bind-group
state rather than a mutable GL state vector.
The compatibility cache budget is an atomic O(1) wrapper value. Usage reports
the exact process-device shader-module, bind-group-layout, pipeline-layout,
render-pipeline, and compute-pipeline counts and reports zero bytes when the
backend cannot attribute shared GPU residency to one wrapper. Purging processes
the context's deferred resource-release queue but never destroys leased shared
pipelines or another presentation context's atlases. The memory dump therefore
reports bounded counts, the configured limit, and the WebGPU backend without
inventing per-wrapper native allocation totals.
The clean-room design follows Skia's public direct-context submission, abandonment, and cache contract, WebGPU's device/queue timeline and completion semantics, and Direct3D 12's explicit command-list, queue, and fence ownership. It adopts explicit submission, shared-device lifetime, device-loss observation, and bounded deferred cleanup; it rejects fake native-backend ownership, unconditional idle waits, and eviction of live cross-host resources. The required Skia/SkParagraph, DirectWrite/Direct2D, Win2D, WebRender, Vello/Parley, and HarfBuzz review recorded above remains unchanged because this slice does not alter scene compilation, shaping, layout, or glyph residency.
The exact-checksum Apple M3 Pro Release workload constructs and reads every
official GRContextOptions property 100,000 times per sample. Native measured
8.389 ns/op and 32 B/op; ProGPU measured 8.252 ns/op and 32 B/op
(0.984 latency ratio). Matched Time Profiler captures measured 7.906 versus
8.135 ns/op, Allocations captures 7.906 versus 7.820 ns/op, and both
retain exactly 32 managed B/op. Metal System Trace captures measured 8.115
versus 8.357 ns/op. The ProGPU trace exports zero target command-buffer,
current-allocation-size, and resource-allocation rows. The native Metal trace
and TOC were retained, but exporting its individual Metal tables reports an
Instruments run error, so no unsupported native row-count claim is made. Raw
traces, TOCs, available exported tables, and exact-run JSON are retained under
artifacts/performance/skiasharp-gr-context-api-instruments.
The 43 legacy SKPath mutation overloads now carry the official
Obsolete("Use SKPathBuilder instead.") contract without changing their
existing clean-room behavior. The attribute is advisory rather than an error,
so source compatibility remains intact while new callers receive the same
migration signal as the official 4.151.0 surface. An independent metadata test
enumerates every declared public obsolete method, fixes the count at 43, and
verifies the exact message and non-error policy.
This is a metadata-only closure over ProGPU's already validated CPU path view.
Path mutation remains retained, CPU-only O(1) work per line/curve operation
and O(N) storage for N segments; it does not initialize WebGPU, flatten
analytic arcs, or change renderer cache keys. The original clean-room path and
builder checkpoints above continue to define topology, ownership, conic,
iterator, transform, serialization, and performance behavior. Because no
algorithm, allocation path, shader, or rendered output changed, the matched
performance and Instruments evidence for those checkpoints remains applicable;
this slice introduces no executable hot-path work to benchmark.
The public migration policy was derived solely from the pinned NuGet reference
metadata and the official
SKPathBuilder API contract.
No implementation source was consulted. The required cross-engine rendering
review remains unchanged because this checkpoint neither changes scene/path
compilation nor text shaping, caching, or GPU submission.
The pinned official SkiaSharp 4.151.0 comparison now reports 4,222 exact
matches of 4,222 reference entries and zero missing entries. The final slice
closes nullable/obsolete metadata, managed disposal, WebP frame/encoder, pinned
raw text-run buffers, SKMaskFilter, SKNoDrawCanvas, SKNWayCanvas, and
SKOverdrawCanvas contracts. Metadata equality is the contract-ledger result;
behavior and performance remain independently gated.
Mask filters retain immutable blur, alpha-table/gamma/clip, and shader
descriptions. Ordinary draw commands remain on the existing direct retained
path. A typed marker activates interception only for filtered brushes; the
source command renders once into a bounded offscreen target and the existing
WebGPU image-filter graph performs separable blur, alpha lookup, or DstIn
shader masking. Overdraw uses a dedicated 16-by-16 WebGPU compute shader and a
96-byte six-color uniform, mapping transparent input to transparent output,
counts one through five to their palette entries, and saturated counts to the
last entry. No CPU readback, external codec, reflection, or per-pixel managed
loop is introduced.
The clean-room design used the public
SkMaskFilter and
SkCanvas contracts,
Direct2D Gaussian blur,
Win2D Gaussian blur,
WebRender's retained frame architecture,
Vello's wgpu renderer,
Parley's reusable layout model, and
HarfBuzz shaping. ProGPU
adopts retained filter descriptions, bounded GPU intermediates, and explicit
compute/composite stages; it rejects copied engine structure, CPU bitmap
fallback, per-frame reflection, and changes to reusable shaping/layout output.
Focused mask/forwarding/shader-resource tests pass, including GPU blur-tail and
overdraw pixel checks. The complete macOS core suite passes 3,167/3,167 and the
headless suite passes 225/225. Three alternating matched Release process pairs
preserve every semantic checksum. The final Apple M3 Pro run records retained
canvas routing at 738.425 versus native 207.625 ns/op (3.557), path build
and bounds at 3,793.146 versus 711.500 ns/op (5.331), and bounded surface
snapshot at 66,053.955 versus 482.085 ns/op (137.017). These gaps remain
explicit optimization work; full metadata closure does not claim an overall
performance win.
The filtered-command marker is now gated behind the presence of an interceptor,
so ordinary framework-neutral command lists do not pay two type tests. Matched
macOS Instruments runs against exact pre-change commit 65cc9641 retained the
same checksum and 0.788 managed B/op: Time Profiler measured 169.313 before
and 165.688 ns/op after, Allocations measured 171.219 and 166.627 ns/op,
and Metal System Trace measured 173.121 and 166.933 ns/op. Target Metal
command-buffer submissions, current device allocation, and resource-allocation
exports are empty in both runs, as expected for state-only recording. Raw
traces, TOCs, table exports, and exact-run JSON are retained under
artifacts/performance/skiasharp-interceptor-instruments.
This continuation closes the three explicit Preview.35 performance slices without changing the complete 4,222-of-4,222 official 4.151.0 metadata ledger.
Common SKPathBuilder move, line, quadratic, cubic, and close operations now
write one pooled contiguous command stream. Bounds are maintained
incrementally, immutable detach transfers ownership in O(1), and the public
PathGeometry graph materializes only when requested. Complex conic,
analytic-arc, add-path, transform, reverse, and iterator paths retain the typed
geometry implementation. Construction is CPU-only O(N) time and storage for
N commands, bounds are O(1), and storage retention is bounded to one
thread-local array of at most 1,024 commands; larger arrays return to the
shared pool.
Surface snapshots now create one immutable full-surface WebGPU texture per content generation. Bounded images are constant-time shared views with composed origins and reference-counted lifetime. The next surface command invalidates only the cache reference; returned images retain the old generation and preserve the immutable snapshot contract. A generation performs one GPU texture copy and requires copy-source, copy-destination, and texture-binding usage; repeated snapshots allocate no texture, submit no copy, and perform no CPU readback. Borrowed externally mutable targets remain uncached. Cold cross-context, raster, encoded-data, and release-callback state is held lazily, so ordinary views do not allocate unrelated locks or maps.
The clean-room design uses Skia's public
SkPathBuilder,
SkPath, and
SkSurface::makeImageSnapshot
contracts; Direct2D's
path geometry model;
Win2D's
offscreen target model;
WebGPU's
texture usage, lifetime, and texel-copy rules;
WebRender's
retained display-list architecture;
and Vello's
compute-centric renderer. ProGPU adopts
immutable generations, explicit GPU ownership, lazy typed materialization, and
retained-resource reuse. It rejects copied source structure, per-view GPU
copies, CPU readback, unbounded exact-position caches, and GPU initialization
in path construction. SkParagraph, Parley, DirectWrite, and HarfBuzz were also
reviewed at the architecture boundary; this slice does not alter shaping or
line layout, so their reusable CPU result boundary remains unchanged.
Three alternating exact-checksum Apple M3 Pro Release process pairs at commit
c989623c produced these medians:
| Workload | Native SkiaSharp | ProGPU | Ratio | Managed B/op, native/ProGPU |
|---|---|---|---|---|
| retained canvas state routing | 206.148 ns | 215.754 ns | 1.047 | 0.175 / 0.788 |
| path build and bounds | 766.542 ns | 512.875 ns | 0.669 | 168 / 224 |
| bounded surface snapshot | 411.052 ns | 223.758 ns | 0.544 | 104.168 / 104.890 |
The former canvas ratio was a Tier-0 measurement artifact. Thirty-two full
warmups stabilize dynamic PGO before sampling; the steady route is within 4.7%
of native and remains below one managed byte per operation. Against exact
Preview.35, packed path construction fell from 2,868.104 to 534.312 ns/op
and from 3,520 to 224 managed B/op. It is faster than the native 764.459
ns/op result, while native's 168 managed B/op excludes its native
allocations; no unsupported total-memory comparison is made.
Matched macOS profiling compares exact Preview.35 product commit 561a5bd2
with exact product commit c989623c; the baseline harness contains only the
dependency-reference and operation-count changes needed to run the same case.
Time Profiler measured snapshots at 34,701.302 versus 114.981 ns/op.
Allocations plus VM Tracker measured 78,493.290 versus 124.304 ns/op and
512.834 versus 104.890 managed B/op; raw native tables are retained because
xctrace does not expose an allocation-table export schema for this template.
EventPipe measured 34,994.175 versus 147.485 ns/op and attributes the
baseline to per-call SKSurface.Snapshot, GpuTexture.Allocate, and queue
submission, while the candidate samples the shared-view/reference-count path.
A bounded 100-operation Metal System Trace avoids an unusable multi-gigabyte
baseline while exercising the same path. Baseline/candidate exports contain
4,090/131 command-buffer-submission rows and 198/121 resource allocation or
deallocation rows. The candidate creates the surface backing texture and one
labelled immutable snapshot generation rather than a texture per view. Raw
traces, TOCs, exported tables, exact-run JSON, EventPipe, and Speedscope files
are retained under artifacts/performance/skiasharp-surface-c989623c; packed
path evidence is under
artifacts/performance/skiasharp-packed-path-0cba9fb9.
The three Preview.35 blockers are closed. Residual matched CPU ratios above
native remain separately visible: platform-lock read 1.139, PM-color array
unpremultiply 1.165, string round-trip 1.140, and in-place 4 KiB swizzle
1.094. Image-subset and gradient-factory managed representations also remain
larger where their latency is faster. These are future optimization slices,
not blockers for this release boundary.
Product commit 4be7dbb1 closes the Preview.36 gradient-factory allocation
item without changing the complete 4,222-of-4,222 SkiaSharp 4.151.0 metadata
ledger or the existing WebGPU gradient renderer. SKShader now stores one
typed payload plus a compact kind instead of retaining eight nullable payload
references. Linear, radial, sweep, and two-point-conical gradients use original
typed descriptors rather than closure-backed brush factories. The descriptors
pack spread/interpolation options into one byte and preserve the exact local
matrix, geometry, color-space selection, and immutable stop snapshot.
The common zero-to-three-stop path uses a bounded per-thread last-input lookup.
It reuses compact immutable stop storage only while the exact source array,
positions array, and values remain unchanged. Caller mutation creates a new
snapshot, so existing shaders cannot observe later input changes. More than
three stops always receive an independent owned array. The lookup retains at
most one three-element SKColor input and one three-element SKColorF input,
their optional positions and immutable snapshots, plus one matrix result per
thread. Factory validation and snapshotting are O(S) time for S stops;
unchanged common inputs are bounded O(S) comparisons with S <= 3 and no
stop-array allocation. Overflow storage and the public ToBrush ownership
boundary remain O(S) time and storage. Matrix inversion is fixed O(1) work
and is reused only for an exact unchanged matrix. No reflection, unbounded
cache, CPU raster fallback, GPU initialization, or backend-specific public
handle is introduced.
The clean-room review used Skia's public
SkGradientShader contract,
Skia's shaped-text architecture,
Direct2D brushes,
Direct2D/DirectWrite separation,
Win2D linear gradients,
the WebGPU specification,
WebRender's retained-frame model,
Vello's wgpu renderer,
Parley's reusable layout model, and
HarfBuzz shaping. ProGPU
adopts immutable retained parameters, explicit interpolation/addressing, typed
ownership, and deferred GPU evaluation. It adapts those contracts to one
framework-neutral descriptor shared by Avalonia, WinUI, WPF, and WinForms. It
rejects copied implementation structure, render-target-bound factory objects,
per-tile uploads, source-array aliasing, and moving Unicode shaping or line
layout onto this shader path. Actual gradient sampling and compositing remain
in the existing WebGPU pipeline.
Three alternating exact-checksum Apple M3 Pro Release process pairs compare
exact Preview.36 commit 7a94fb3c with the candidate implementation. The
median of run medians fell from 369.452 to 218.160 ns/op (40.95%), while
managed allocation fell from 1,480 to 472 B/op (68.11%). At 2,000 warmup
passes the same binaries measured 243.584 versus 137.938 ns/op, confirming
the ordering after final dynamic PGO. The stabilized official SkiaSharp
4.151.0 differential measured native 1,283.516 versus ProGPU 198.848
ns/op with the same checksum. Managed allocation was 416 versus 472 B/op;
the native counter excludes Skia's native heap work, so no unsupported total
memory comparison is made. The baseline harness changes are limited to the
direct backend reference required by a clean worktree and increasing this
case from 1,000 to 16,000 operations to clear the sub-millisecond timer-noise
floor.
Matched macOS Time Profiler captures measured exact Preview.36 at 329.578
ns/op and the candidate at 182.654 ns/op. Allocations plus VM Tracker measured
329.314 versus 184.297 ns/op and the same 1,480 versus 472 managed B/op.
EventPipe measured 327.648 versus 189.569 ns/op. Metal System Trace measured
330.810 versus 183.625 ns/op; both exported zero target command-buffer,
current-allocation-size, and resource-allocation rows, confirming construction
is CPU-only. Raw process JSON, Instruments traces and TOCs, Metal table exports,
EventPipe traces, and Speedscope conversions are retained under
artifacts/performance/skiasharp-gradient-typed-final.
Mutation, overflow ownership, colorspace, tile-mode, local-matrix, degeneracy,
paint-alpha, transformed-picture, and GPU pixel-coverage tests pass. The full
macOS core suite passes 3,237/3,237 and the headless suite passes 225/225. The
official API metadata gate still reports reference=4222, matching=4222,
missing=0, and extra=997 ProGPU extensions.
SKRoundRect now keeps its fixed four SKPoint corner radii in an inline value
buffer instead of allocating a second managed array for every instance. The
copy constructor transfers those four values directly, uniform initialization
uses four bounded stores, and internal SKPath/canvas consumers borrow a
ReadOnlySpan<SKPoint>. The public Radii property still returns a fresh
caller-owned four-point array, preserving the official ownership boundary.
Construction, copying, normalization, classification, and radius access remain
fixed O(1) CPU work and storage. They do not initialize WebGPU, allocate a
native geometry object, or change retained path topology.
This clean-room change was designed from Skia's public
SkRRect contract, the official
SKRoundRect API,
GetRadii,
and
SetRectRadii
contracts, plus the official .NET
InlineArrayAttribute
and C# inline-array specification.
ProGPU adopts the observable four-corner value and ownership contracts and
adapts them to its typed CPU-only geometry model. It rejects source-array
aliasing, an unbounded cache, native allocation, reflection, and GPU setup for
metadata operations. No foreign implementation source was consulted. This is
an object-storage change rather than a renderer, text, scene, or GPU-pipeline
change, so the existing cross-engine rendering architecture review remains
unchanged.
Three alternating Apple M3 Pro Release process pairs compared exact
Preview.37 commit d510dd5c with the final candidate after 2,000 dynamic-PGO
warmups and 192 samples per binary. The exact semantic checksum remained
13947687467187634243. Aggregate median construction/disposal latency fell
from 31.0687 to 25.1271 ns/op, a 19.12% latency reduction or 23.65%
more operations per second. Managed allocation fell from 120 to 88 B/op,
a 26.67% reduction. Scheduler interruptions dominate the raw p95 values
(111.9750 versus 112.3459 ns/op), so no tail-latency improvement is
claimed. A separate three-pair differential against official SkiaSharp
4.151.0 measured 63.417 versus 25.038 ns/op with the same checksum.
Official managed allocation was 80 B/op versus ProGPU's 88 B/op, but that
counter excludes Skia's native SkRRect allocation, so it is not treated as a
total-memory comparison. The complete default-warmup matrix also preserved all
checksums and measured this case at 90.442 versus 35.606 ns/op.
Matched macOS 26.4.1 profiling used the same 400-million-construction Release
workload for Preview.37 and the candidate. Time Profiler and Allocations plus
VM Tracker each sampled 12 seconds. EventPipe sampled-thread-time plus verbose
GC attributed 99.23%/99.64% exclusive CPU to the benchmark body; the
baseline's SpanHelpers.ClearWithoutReferences entry (0.23%) disappeared
from the candidate hot list. Three-second Metal System Trace captures exported
zero target application-encoder and zero target Metal-driver rows for both
binaries, confirming the value path remains CPU-only. Raw tracing artifacts
were removed after extracting these summaries to recover local disk space, as
requested; reproducible benchmark JSON and Markdown remain under
artifacts/performance/skiasharp-roundrect-inline.
The focused rounded-rectangle/path/canvas tests pass, including a 10,000-object
allocation guard requiring at most 96 managed B/op. The complete macOS core
suite passes 3,238/3,238 and the headless suite passes 225/225. The official
SkiaSharp metadata gate reports reference=4222, matching=4222, missing=0,
and extra=998; the one-entry extension-count movement is compiler-emitted
nullable metadata redistribution, not a new public member.
SKImage subsets now share one root-invariant TextureStorage containing the
texture, pixel format, alpha format, color space, portable pixel snapshot, row
width, texture-backed classification, ownership callbacks, and atomic lifetime.
Each immutable view retains only that storage reference, its width and height,
one composed origin pair, and lazily created view-specific state. Info remains
an official value-returning boundary and is reconstructed from those fields.
Nested subsets compose checked origins in O(1) time, never copy pixels, never
upload another texture, and remain valid after their parent view is disposed.
The compare/exchange retain loop deliberately preserves the existing
no-resurrection ownership rule; a tempting increment-and-rollback shortcut was
rejected because concurrent retainers could observe the rollback after final
release.
This clean-room design used Skia's public
SkImage immutability and subset contract,
the WebGPU texture-view model,
Direct2D source rectangles,
Win2D image source rectangles,
WebRender's retained-scene model,
and Vello's wgpu renderer. ProGPU adopts
immutable shared backing storage, cheap typed views, explicit ownership, and
deferred source-rectangle evaluation. It rejects copied implementation
structure, per-view pixel buffers, texture duplication, reflection, and an
unbounded view cache. The text boundary was reviewed against
Skia's shaping architecture,
Parley, and
HarfBuzz; this storage
change does not move shaping, layout, glyph caching, or renderer work. No
foreign implementation source was consulted.
Three alternating Apple M3 Pro Release process pairs compared exact
Preview.38 commit 65f86cf4 with implementation commit c7046673 after 2,000
dynamic-PGO warmups and 192 samples. The exact checksum remained
15041971963811491075. Aggregate median subset latency fell from 29.312 to
26.450 ns/op (9.76%), or 10.82% more operations per second. Managed
allocation fell from 105.694 to 65.693 B/op (37.85%). Scheduler
interruptions dominate the raw p95 values, so no tail-latency improvement is
claimed. A stabilized official SkiaSharp 4.151.0 differential measured
339.927 versus 26.450 ns/op with the same checksum. Official managed
allocation was 104.021 B/op versus ProGPU's 65.693 B/op, but that counter
excludes Skia's native heap, so it is not treated as a total-memory comparison.
Matched macOS 26.4.1 profiling isolated one long-lived immutable source image
from source upload by overriding only the benchmark operation count. Each Time
Profiler and EventPipe process constructed 400 million subset views. Time
Profiler measured exact Preview.38 at 36.429 and the candidate at 32.553
ns/op (10.64% lower); EventPipe measured 40.414 versus 35.315 ns/op
(12.62% lower), with about 96% exclusive sampled CPU in the intended
benchmark body. The same sustained path allocated exactly 104 versus 64
managed B/op (38.46% lower). A bounded 40-million-view Allocations plus VM
Tracker lane measured 37.198 versus 33.065 ns/op. Its native heap and VM
totals include runtime/device startup and did not show a retained regression;
they are not used as evidence for the managed object-size claim.
The matched 40-million-view Metal System Trace lane measured 37.882 versus
33.152 ns/op. Both binaries exported exactly 26 target resource-allocation
rows, 42 current-allocation-size intervals, the same 1,196,032-byte peak,
and zero target command-buffer submissions, errors, compiler spills, or hangs.
Thus source creation remains the only GPU work and view count does not multiply
GPU resources. Raw Instruments traces, table exports, EventPipe traces, and
temporary exact-baseline binaries were removed after these summaries were
extracted. Reproducible benchmark distributions remain under
artifacts/performance/skiasharp-image-view-c7046673.
A 10,000-view focused allocation guard requires at most 72 managed B/view;
image, surface, pixmap, and effect ownership tests pass. The complete macOS
core suite passes 3,239/3,239 and the headless suite passes 225/225. The
official API metadata gate remains reference=4222, matching=4222,
missing=0, and extra=998; this implementation and its benchmark operation
override add no public API.
The shared 32-bit pixel channel swizzler now uses .NET 10's hardware-native
128-bit byte-table shuffle whenever Vector128 acceleration is available. Its
constant indices are all in [0, 15], so ShuffleNative can lower directly to
the architecture's native table instruction without the normalization required
by the portable Shuffle operation. This replaces the former Apple ARM64
sequence of two element reversals plus a bitwise select. Four vectors are
unrolled per iteration, remaining complete vectors use the same primitive, and
the existing scalar loop preserves incomplete trailing bytes. Forward copy,
in-place conversion, count clamping, backward overlap handling, and every
public SKSwizzle overload retain their existing contracts.
For P complete pixels the algorithm is O(P) time and O(1) auxiliary
storage. It performs one load, one native byte shuffle, and one store per
16-byte vector on the common path. It allocates no managed memory, initializes
no WebGPU device, submits no GPU command, and does not change alpha or the
middle two bytes of any pixel. A 4,099-byte regression covers every vector
block plus an incomplete tail; the focused suite passes 6/6.
The clean-room design used only public contracts and independently measured
behavior: Skia's public
SkSwapRB
RGBA/BGRA contract; Direct2D's
BGRA/RGBA format guidance;
Win2D's raw
CanvasBitmap.GetPixelBytes
default-BGRA contract; WebRender's
swizzling architecture;
Vello's wgpu renderer boundary; and the
.NET 10
Vector128.ShuffleNative
contract. ProGPU adopts explicit format boundaries, avoids conversion when the
existing caller already has the target format, and uses one portable native
SIMD primitive when a CPU-visible buffer must observably change. It rejects a
GPU/shader substitute for the public mutating CPU API, per-platform duplicate
loops, runtime reflection, copied implementation structure, and hidden buffer
allocation. No foreign implementation source was copied or adapted.
The mandatory text-boundary review used Skia's
Shaped Text, DirectWrite
glyph runs,
Parley's shared layout resources, and
HarfBuzz's
shaping output.
This byte-format transform remains below those reusable shaping/layout results
and changes no font, glyph, atlas, subpixel, or fallback state.
Three interleaved Apple M3 Pro Release process pairs compared exact
Preview.39 merge 3efcf9e5 with product commit bfc6c62a, using 128 complete
warmups and 192 samples per process. Across 576 samples per side the exact
checksum remained 12185046443090060243, median latency fell from 90.375 to
79.942 ns/op (11.55% lower), and throughput rose 13.05%; both sides
allocated exactly 0 managed B/op. Scheduler interference dominates the raw
p95 distribution, so no tail-latency claim is made. An exploratory matched
official SkiaSharp 4.151.0 process set measured 85.442 versus 81.804 ns/op,
so the candidate was 4.26% faster with the same checksum and allocation.
Matched long-running macOS profiling used 20 million 4-KiB operations per
sample and checksum 895921851728446851. Time Profiler measured 93.264
versus 46.231 ns/op (50.43% lower). Allocations plus VM Tracker measured
97.338 versus 49.173 ns/op (49.48% lower) and retained zero managed
bytes per operation. Persistent native heap plus anonymous VM was effectively
unchanged at 107,098,816 versus 107,111,136 bytes; the 12-KiB difference
is startup noise, while candidate total allocation bytes were lower.
EventPipe measured 92.400 versus 85.772 ns/op and attributed 95.70%/95.18%
exclusive sampled time to the intended swizzle body. Metal System Trace
measured 97.363 versus 47.386 ns/op (51.33% lower); both traces exported
zero target command-buffer submissions, command-buffer errors, compiler
spills, hangs, Metal resource allocations, and currentAllocatedSize rows.
Raw distributions and compact profiler target results are retained under
artifacts/performance/skiasharp-swizzle-native-shuffle. After extracting the
summaries, 433 MiB of raw Instruments/EventPipe data and 102 MiB of exact
baseline build state were deleted. No task-owned .trace, .nettrace,
Speedscope, Xcode scratch, or temporary worktree remains.
SKRuntimeEffectUniforms now publishes its current uniform byte storage as an
immutable retained snapshot. A later write or reset clones the buffer only
when that storage has already been published, preserving every older shader,
color-filter, or blender instance without copying on every construction.
Effects with no child slots reuse the empty child array, and the common
identity instance no longer stores a 36-byte SKMatrix; a compact derived
instance carries the matrix only for a non-identity transform. Public API and
observable mutation isolation remain unchanged.
Snapshot publication is O(1). The first mutation after publication is
O(U) time and storage for U uniform bytes; later mutations before another
snapshot are O(1). Existing child capture remains O(C) for C child slots.
The path is CPU-only and performs no WebGPU initialization, upload, or command
submission.
The clean-room design used Skia's public Runtime Effects and SkSL contract; Direct2D's resource-format boundary; Win2D's premultiplied-alpha contract; WebRender's retained blob-image architecture; Vello's typed retained renderer boundary; DirectWrite glyph runs; and HarfBuzz's reusable shaping outputs. ProGPU adopts immutable retained payloads, copy-on-write ownership, and compact identity state. It rejects writable shared snapshots, reflection, CPU-pixel fallbacks, copied foreign implementation structure, moving text shaping into this layer, and GPU initialization for this CPU ownership API.
Three interleaved Apple M3 Pro Release process pairs compared exact Preview.40
merge 3dbf79b8 with product commit fa77c4ad, using 128 warmups and 192
samples per process. Across 576 samples per side, the exact checksum remained
1721237190835759209; median latency fell from 220.8291 to 140.1958
ns/op (36.51% lower), throughput rose 57.51%, and managed allocation fell
from 544 to 360 B/op (33.82% lower). Scheduler interference dominates
the tail, so no P95 improvement is claimed. An exploratory official SkiaSharp
4.151.0 comparison produced the same checksum; its managed counter excludes
native allocation and is not used for a total-memory claim.
Matched macOS profiling retained the same allocation result. Time Profiler
measured 206.865 versus 175.498 ns/op; Allocations plus VM Tracker measured
259.502 versus 177.112; EventPipe sampled-thread-time measured 212.559
versus 166.342; and Metal System Trace measured 224.460 versus 172.181.
EventPipe attributed 38.54% exclusive baseline samples to ToShader; that
frame left the candidate top 15 after the compact path became inlineable. Both
Metal traces exported zero target command-buffer submissions and zero
MTLDevice.currentAllocatedSize rows.
Focused runtime-effect tests pass 7/7, including snapshot mutation isolation,
transformed-matrix fidelity, and a 400-B/op allocation ceiling. The full core
suite passes 3,242/3,242, the headless suite passes 225/225, and the XAML
compiler suite passes 307/307. The official Skia API gate remains
reference=4222, matching=4222, missing=0, and extra=998; documentation
and package-manifest gates pass. Distributions, compact profiler results, the
complete research record, and reproduction protocol are retained under
artifacts/performance/skiasharp-runtime-effect-cow. After extraction,
approximately 8.6 GiB of task-owned raw EventPipe/Instruments data and
temporary exact-baseline state were deleted; no task-owned trace, scratch, path
marker, or worktree remains.
SKPath now defers PathGeometry construction until an empty path is actually
materialized or mutated. The packed constructor no longer allocates an empty
geometry only to discard it before retaining PackedPathData. Packed bounds,
detach ownership, analytic verbs, fill rules, close/current-point behavior,
and later geometry materialization are unchanged.
Empty construction, packed detach, and packed bounds remain O(1) beyond the
retained command stream. First materialization remains O(N) time and storage
for N commands. The path stays CPU-only and adds no tessellation, WebGPU
initialization, upload, or command submission.
The clean-room design used Skia's public
SkPathBuilder contract;
Direct2D's
ID2D1GeometrySink;
Win2D's
path/figure behavior;
WebRender's
retained scene-building boundary;
Vello's GPU scene model and
compact encoding;
Skia shaped text;
DirectWrite glyph runs;
Parley's shared layout resources; and
HarfBuzz's shaped output.
ProGPU adopts lazy retained storage and preserves analytic commands until an
explicit materialization boundary. It rejects eager tessellation, reflection,
GPU setup, copied foreign implementation structure, and text reshaping.
Three interleaved long-running process pairs compared exact merged main
3c3c46b8 with product commit c8213c55, using 128 warmups, 192 samples per
process, and 10,000 path builds per sample. Managed allocation fell from 224
to 136 B/op (39.29%), while the exact checksum remained
8402956917441101891. Median latency differed by only 0.40%, inside
process/frequency noise, so no process-pair latency claim is made. A matched
official SkiaSharp 4.151.0 set measured approximately 725.5 versus 529.3
ns/op and 168 versus 136 managed B/op; native Skia allocations are outside
that managed counter.
Matched Time Profiler measured 523.428 versus 406.363 ns/op (22.36%
lower), Allocations plus VM Tracker measured 418.704 versus 406.572, and
Metal System Trace measured 425.999 versus 414.783. EventPipe whole-process
timing was 427.227 versus 440.719, but its exclusive Detach samples fell
from 0.27% to 0.10%; no EventPipe throughput claim is made. Both Metal
traces exported zero target command-buffer submissions and zero
MTLDevice.currentAllocatedSize rows.
The focused path suite passes 93/93 and tightens the packed-detach ceiling to
192 managed bytes for both small and 256-segment paths. Core passes
3,242/3,242, headless passes 225/225, and the XAML compiler passes 307/307.
Official API metadata remains 4,222/4,222 required with zero missing; docs and
package manifests pass. Distributions, profiler target results, and research
are retained under
artifacts/performance/skiasharp-path-lazy-geometry. After extraction, 272 MiB
of raw profiler data and 102 MiB of exact-baseline build state were deleted; no
task-owned trace, scratch directory, or worktree remains.
SKCanvas now stores saved state, pushed scopes, and layer frames in lazy typed
value buffers rather than eagerly allocating generic stack wrapper objects.
Active clips are derived from the already authoritative pushed-scope stack and
materialized as full RenderCommand values only when a save-layer snapshot
needs them. This removes the former second copy of every active clip command.
Popped reference-containing entries are cleared immediately, arbitrary nesting
still grows geometrically, clip/layer order remains LIFO, and the public
one-based save-count contract is unchanged. Bitmap flushes retain live clip
semantics by temporarily borrowing the active commands, clearing consumed draw
state, replaying only those pushes into the reused context, and rebasing their
typed scope indices. This keeps later draws and save-layer snapshots valid
without restoring duplicate per-clip storage to the normal recording path.
Save, push, and pop are amortized O(1); an occasional capacity growth is
O(D) time/storage for depth D. A layer snapshot is O(S + C) time and
O(C) output storage for S active scopes and C clips. Warm state cycling is
allocation-free. The change is CPU-only: it does not initialize WebGPU, alter a
retained draw command, change raster quality, or move shaping/layout work.
The clean-room design used Skia's public
SkCanvas save/restore contract;
Direct2D's
PushAxisAlignedClip
LIFO nesting contract; Win2D's
CanvasDrawingSession
stateful drawing boundary; WebRender's
retained display-list architecture;
Vello's typed GPU scene; Parley's
reusable layout model; and HarfBuzz's
shaping output contract.
ProGPU adopts compact typed storage, strict nested ownership, and deferred GPU
evaluation. It rejects copied foreign implementation structure, reflection,
duplicate full-command storage, GPU bookkeeping for CPU state, and reshaping
text during canvas save/restore.
Three interleaved Apple M3 Pro Release process pairs compared exact unpublished
Preview.41 tag commit 19867237 with product commit b1a30c1c, using 128
warmups and 192 samples per process. Across 576 samples per side, the semantic
checksum remained 17022205643649352006; median latency fell from 177.5375
to 109.2416 ns/op (38.47%), throughput rose 62.52%, and P95 fell from
226.1542 to 121.9334 ns/op (46.08%). Cold one-cycle managed allocation
fell from 7,880 to 4,472 bytes (43.25%). Official SkiaSharp 4.151.0
measured 190.4417 ns/op with the same checksum; its managed counter excludes
native allocations and is not used for a total-memory comparison.
Matched final-binary profiling measured Preview.41 versus candidate at
220.872/105.889 ns/op in Time Profiler, 217.826/109.157 in Allocations
plus VM Tracker, 231.644/110.841 in EventPipe sampled thread time, and
231.903/107.858 in Metal System Trace. EventPipe's baseline duplicate
active-clip copy frames disappear from the candidate. Both Metal captures
report zero target resources, submissions, waits, errors, spills, hangs, and
currentAllocatedSize rows. The 60,176-byte persistent native/VM delta is
startup/JIT noise and is not treated as a memory improvement.
Focused canvas/state tests pass 111/111, including active-clip replay, bitmap
flush rebasing, nested save counts, zero-allocation warm cycling, and a cold
allocation ceiling that rejects duplicate clip-command storage. The complete
core suite passes 3,249/3,249, headless passes 225/225, and the XAML compiler
passes 307/307.
Official API metadata remains 4,222/4,222 required with zero missing and 998
documented extensions; shader-resource, docs, and package-manifest gates pass.
Compact distributions and profiler summaries are retained under
artifacts/performance/skiasharp-canvas-state-routing. Raw Instruments,
EventPipe, Xcode scratch, preliminary captures, and the exact-baseline worktree
were removed after extraction, reclaiming roughly 1.2 GiB of task-owned data.
The Avalonia.Skia-first recording tranche removes optional media-effect state
from every RenderCommand, pools the mutable recorder's large command array,
and stores immutable pictures as typed core, text, texture, uncommon-command,
and deduplicated-transform arrays. The existing public RenderCommand[]
inspection boundary remains available through one lazy immutable
materialization; ordinary compositor, Skia playback, serialization, operation
counting, and byte accounting consume the compact typed view directly. Large
pooled scratch arrays are cleared and returned after a snapshot, while stable
small retained contexts keep their exact trimmed capacity.
Positioned SKTextBlob runs now convert their owned SKPoint[] positions to
the renderer's Vector2[] representation lazily on the first retained draw.
The converted array is atomically published on the owning run and reused by
every later draw. Blob construction therefore keeps its previous allocation
profile, while recording is no longer O(G) allocation per draw for G
glyphs. Rotation/scale runs retain their separate per-glyph transform path.
The Avalonia canvas path also reuses package-private solid brushes and pens
while relevant SKPaint state is unchanged; public mutable conversion results
remain independent, and a paint mutation publishes a new retained resource so
earlier commands stay immutable. Ordinary rectangles no longer materialize a
path merely to reject a special-shader route. Antialiased image draws reuse one
immutable unit-rectangle edge clip and place it through the retained command
transform instead of allocating a four-segment geometry per draw.
Recorder growth is amortized O(1) per command and O(C) for a capacity
change. Immutable snapshot construction is O(C) average with O(C) pooled
scratch for C commands; its open-addressed transform table is kept below a
0.5 load factor, with O(C²) only under adversarial matrix-hash collisions.
Replay expansion is allocation-free O(1) per command. Retained storage is
O(C + A + T + X + U) for compact commands, uncommon payloads, text payloads,
texture payloads, and unique transforms. No raster quality, DPI/subpixel
policy, scene invalidation, WebGPU submission, or resource-lifetime boundary
changes in this CPU recording tranche. Solid-paint reuse, the ordinary
rectangle shader check, and unit-clip placement are allocation-free O(1).
The clean-room design used these primary contracts and architecture records:
- Skia
SkPicturepermits recorded operation count to differ from canvas calls and defines approximate storage without charging large referenced objects. ProGPU adopts immutable retained playback and accurate owned-storage accounting, but not Skia's private op encoding or implementation structure. - Direct2D
ID2D1CommandListrecords replayable commands, references bitmap resources, and stores drawing state by value. Win2D'sCanvasCommandListlikewise separates recording from later drawing/effect use. ProGPU adapts this ownership split to typed WebGPU resources and reference-counted leases. - DirectWrite's Direct2D text integration explicitly identifies cached glyph positions in reusable text layouts as a performance advantage. Skia's shaped-text model, HarfBuzz's shaping output and shape-plan caching, and Parley's shared layout/scratch contexts all keep shaping/layout results reusable. ProGPU therefore caches only the representation conversion and never reshapes during drawing.
- WebRender's retained display-list architecture and Vello's typed GPU scene inform the separation between compact CPU scene encoding and later parallel GPU work. The WebGPU render-bundle specification reinforces immutable replay, but bundles are rejected for this layer because ProGPU commands still need current DPI, atlas-generation, effect, clip, and device-loss validation before encoding a render pass.
The implementation rejects copied foreign source/layout, reflection, boxed per-frame adapters, hiding retained allocations in native memory, eager glyph position conversion, unbounded exact-position caches, and moving Unicode or OpenType shaping onto the GPU.
Three alternating Apple M3 Pro Release process pairs compared exact Preview.44
84a86f68 with product commit d22fcef3, using 64 warmups and 96 samples per
process. Across 288 samples per side, the mixed Avalonia-shaped picture retained
checksum 2454466986173768955; median latency fell from 7,543.213 to
4,457.357 ns/op (40.91%), throughput rose 69.23%, and managed allocation
fell from 35,344 to 1,627 B/op (95.40%). Scheduler and GC interference
dominate the tail, so P95 is recorded in the artifact but not used as the
primary claim. Immutable-image picture recording separately fell from 2,486
to 146 managed B/op (94.13%), and the inline command value fell from 816
to 576 bytes before compact picture packing.
Final integration commit e75723db also makes the existing explicit
DrawingContext.EnsureCommandCapacity reservation contract persistent across
Clear, while organically grown large transient command buffers remain pooled
and bounded. The matched picture benchmark does not call that reservation API;
the focused MotionMark allocation regression passes repeatedly with the final
behavior.
Official SkiaSharp 4.151.0 remains faster for the mixed wrapper workload at a
pooled 396.078 ns/op median and 10 managed B/op with the same checksum.
That counter excludes native Skia picture allocation, so it is neither a total
memory comparison nor evidence that ProGPU has reached native latency.
Matched macOS Allocations plus VM Tracker, Time Profiler, and Metal System
Trace launches each completed the same four warmups plus eight 16,384-operation
samples. Instrumented latency measured Preview.44/candidate at
19,748.180/3,428.551, 19,587.496/3,421.159, and
19,956.059/3,226.458 ns/op respectively; managed allocation was
35,281/1,537 B/op throughout. Allocations reported total
heap-plus-anonymous-VM allocation falling from 2,441,702,816 to
327,050,480 bytes, while bounded pool retention raised persistent storage by
7.37 MB, so no persistent-footprint improvement is claimed. The Metal pair
was resource-identical: 42 resources totaling 3,227,648 bytes, maximum
MTLDevice.currentAllocatedSize 1,589,248 bytes, zero target submissions,
and zero waits, errors, spills, or hangs.
Matched EventPipe measured 23,176.839 versus 2,460.077 ns/op. Preview.44's
exclusive command-list growth/copy, rectangle geometry, and paint-conversion
frames leave the candidate hot list; remaining samples center on GC polling,
reference clearing, compact snapshot construction, and retained-array
allocation.
The complete core suite passes 3,268/3,268, headless passes 225/225, Avalonia
renderer contracts pass 86/86, and the XAML compiler suite passes. The unchanged
Avalonia.Skia 12.0.5 source project builds with zero warnings and errors. The
official API gate remains reference=4222, matching=4222, missing=0, and
extra=998; documentation and package-manifest gates pass. Distributions,
compact profiler summaries, research, and reproduction details are retained in
artifacts/performance/skiasharp-avalonia-hotpaths-final. More than 3.4 GiB of
raw Instruments/EventPipe data, XML exports, Xcode scratch, exploratory runs,
and incomplete captures were deleted after extraction; no raw trace remains.
Preview.46 preserves the official 4,222/4,222 metadata match and prioritizes the call shapes used by source-built Avalonia.Skia 12.0.5. Path boolean results are retained as typed deferred geometry and evaluated by the WebGPU path rasterizer. Picture SaveLayer operations retain immutable commands and leases, then prepare their effect/layer textures before ordered replay so a nested offscreen submission cannot split or prematurely release the enclosing image stream. Common blur, shadow, table-filter, dash, rounded-rectangle, and retained-layer state is compacted without CPU raster fallback, reflection, or recording-time WebGPU initialization.
The exact implementation head passed all 15 PR checks. Focused compositor and Skia compatibility coverage passes 380/380; the macOS CI lane passes 3,236/3,236 core and 225/225 headless tests. The source-built Avalonia Composition workload measures 80.619 versus 78.305 frames/s and 6,466.76 versus 7,485.97 managed bytes/frame for ProGPU versus official Skia. Allocations/VM Tracker reports 198,790,144 versus 205,265,120 persistent heap-plus-anonymous-VM bytes, with zero command-buffer errors, spills, hangs, or hang risks.
The evidence is deliberately mixed. ProGPU P95 is 22.616 ms versus 17.021 ms;
managed retained heap, first active physical footprint, native heap, and
IOAccelerator VM are higher. The matched microbenchmark also keeps surface
readback/composition, immutable-image and mixed-picture recording, path
combination, and SaveLayer recording on the remaining optimization ledger.
The compact matched evidence is under
artifacts/performance/skiasharp-avalonia-canvas-image-hotpaths-final, and the
clean-room architecture/research record is
docs/AVALONIA_SKIA_PAINT_EFFECT_RESEARCH.md. All raw .trace, .gcdump,
heap-dump, XML-export, and Xcode scratch artifacts were deleted after summary
extraction.
Preview.47 preserves the official 4,222/4,222 metadata match and continues to
prioritize public call shapes used by source-built Avalonia.Skia 12.0.5. An
original ordered 32-bit token stream plus typed records compacts common
picture operations while retaining exact full records for uncommon commands.
Consecutive immutable-image draws reuse their context-owned texture;
native-format Disallow readback copies directly from the reusable WebGPU
staging buffer into caller rows; map polling no longer imposes a fixed
one-millisecond sleep; and the common single-run text builder avoids list
mutation. Rounded rectangles and retained visuals also use compact analytic
records. No CPU renderer, eager CPU image mirror, reflection, external media
dependency, or foreign command encoding was added.
Against the preceding source-equivalent endpoint, repeated immutable-image readback improves 91.8%, direct surface readback 85.1%, and conversion readback 85.9%. Mixed picture allocation falls from 1,627 to 424 B/op, common layer recording reaches 8,180 B/op, and focused positioned text measures 268.375 ns/op and 89 B/op versus official SkiaSharp at 289.270 ns/op and 136 B/op. Official CPU-raster surfaces remain much faster for synchronous readback, so no universal performance-superiority claim is made.
The exact PR #84 head passed all 15 CI checks, including three operating-system
build/test lanes, portable/mobile packaging, official API metadata, native
Dawn, source-built Avalonia contracts, SVG image parity, and matched benchmark
lanes. Local final gates pass 3,305 core tests, 225 headless tests, 28 Avalonia
compositor tests, 287 Avalonia text tests including the focused corpus, and the
patched Avalonia 12.0.5 ControlCatalog source build. Matched Xcode Allocations,
Time Profiler, and Metal System Trace retain the exact composition checksum and
992 B/op; persistent heap plus anonymous VM changes by 0.019%, with zero waits,
spills, hangs, or command-buffer errors. Raw trace, ETLX, XML-export, and Xcode
scratch artifacts were deleted after compact evidence was retained. Full
methods, complexity, research sources, distributions, and rejected experiments
are in docs/AVALONIA_SKIA_RETAINED_COMMAND_STREAM_RESEARCH.md.
Preview.48 preserves the official 4,222/4,222 metadata match and corrects the shared retained stroke pipeline exercised by SkiaSharp and Avalonia.Skia. Source-local pen provenance now survives recording, append, retained picture replay, archive round-trip, CPU/GPU transform selection, opacity-mask and special-shader routes, and GPU hit testing. Conformal scale is applied once; anisotropic and sheared normal strokes transform their local outline exactly; and zero-width hairlines plus fixed positive widths expand in framebuffer or device space. Caps, joins, miters, dashes, and reflected transforms retain the same stroke mode. Special image, picture, composed, and color-filter shaders no longer discard hairline-only coverage.
Indexed polyline and spline recording no longer allocates eager
PathGeometry/segment graphs. Direct polyline compilation is bounded O(N),
and spline replay restores transform-adaptive sampling instead of forcing 100
segments at every scale. The exact PR #87 head passed all 16 CI checks,
including official metadata, native/ProGPU SVG image parity, matched CPU
benchmarks on three operating systems, source-built Avalonia contracts,
portable/mobile packaging, and native Dawn. Local final gates pass 3,569 core,
240 headless, and 185 focused stroke/hairline/hit-test cases. Algorithms,
quality bounds, complexity, and primary research sources are recorded in
docs/STROKE_TRANSFORM_RESEARCH.md.
Preview.62 preserves the official 4,222/4,222 metadata match and optimizes the
Avalonia-shaped bounded SaveLayer containing one analytic rounded rectangle.
The exact PushClip, DrawRoundedRect, PopClip replay is retained in typed
fields with inline transforms, and sequential layer recording reuses at most
one cleared transient context. Nested layers cannot borrow an active context;
unmatched command shapes retain the existing general compact path. Construction
and retained storage are O(1) for the specialized shape, replay is three
allocation-free indexed expansions, and the canvas-local reuse bound is
independent of sequential layer count.
The final alternating three-process Release matrix preserves all 62 semantic
checksums. avalonia-layer-recording improves from 3,847.625 to 2,673.188
ns/op median (-30.5%), from 14,958.313 to 4,333.313 ns/op p95 (-71.0%), and
from 8,189 to 6,131 managed B/op (-25.1%). Matched 50,000-operation Xcode
Allocations/VM Tracker, Time Profiler, and Metal System Trace captures retain
the exact checksum and reduce measured managed allocation from 3,309 to 1,205
B/op (-63.6%); persistent heap plus anonymous VM changes by +0.16%, with zero
target Metal resources, submissions, waits, spills, hangs, or command-buffer
errors. The managed/native audit finds no renderer delta because both scene
compilers consume the same expanded commands. Full research, rejected
alternatives, distributions, validation counts, and reproduction evidence are
in docs/AVALONIA_SKIA_RETAINED_COMMAND_STREAM_RESEARCH.md.