Skip to content

Add an experimental Dawn backend (generated Cython bindings) - #839

Draft
hmaarrfk wants to merge 8 commits into
pygfx:mainfrom
hmaarrfk:dawn
Draft

hmaarrfk wants to merge 8 commits into
pygfx:mainfrom
hmaarrfk:dawn

Conversation

@hmaarrfk

@hmaarrfk hmaarrfk commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Note

This PR was written with the help of an AI assistant (Claude).

dawn

pyodide demo: https://www.markharfouche.com/wgpu-py/

claude

Draft. Adds an experimental, optional Dawn backend (wgpu.backends.dawn) next to the wgpu-native backend. It is one Cython extension module that runs natively (against Dawn's libwebgpu_dawn) and in Pyodide (against Emdawnwebgpu, Dawn's webgpu.h on top of the browser's WebGPU). It implements the same public classes as wgpu/_classes.py, so existing code and tests can select it, either with import wgpu.backends.dawn or with WGPUPY_BACKEND=dawn. In Pyodide it is selected automatically if it is installed. wgpu-native stays the default natively, and wgpu-py installs and works without Dawn.

xref: #838. The Pyodide support is adapted from the cffi-based #840; this PR keeps Cython as the binding technology. The backend-agnostic fixes and tests that #840 gained since have been ported here too.

Live demo: https://www.markharfouche.com/wgpu-py/

Binding strategy: generated Cython

The goal was the lowest Python→Dawn call overhead, with bindings that are generated rather than written by hand.

  • Generated: codegen/dawn_patcher.py parses two vendored headers of the same Dawn release (v20260929.200159) with pycparser, using the same cleaning step as the wgpu-native headers:

    • wgpu/resources/dawn_webgpu.h: Dawn's native webgpu.h.
    • wgpu/resources/emdawn_webgpu.h: Emdawnwebgpu's webgpu.h, which is a subset of the native one.

    From them it generates:

    • backends/dawn/_webgpu.pxd. Its main block comes from Emdawnwebgpu's header, so it is valid for both targets (207 functions, 91 structs, enums, flags, callbacks and the WGPU_*_INIT initializers). The codegen checks that these declarations and the enum values are identical in the native header.
    • A separate, clearly marked block for the few native-only declarations that the backend uses (NATIVE_ONLY in the codegen: the logging callback, wgpuDeviceTick, and the native surface sources). They come from the generated backends/dawn/dawn_native_only.h, which just includes webgpu.h natively. With Emscripten, it defines them (copied verbatim from Dawn's header) with no-op functions, so that the same _api.pyx compiles. The backend does not call them in the browser.
    • backends/dawn/_mappings.py: enum mappings.

    The codegen also checks _api.pyx against the base API and lists anything not implemented in the report (currently none).

  • Hand-written, like wgpu_native/_api.py: backends/dawn/_api.pyx mirrors the wgpu-native backend class by class.

    • The classes are ordinary Python classes that subclass wgpu.classes.*, but Cython compiles their methods.
    • Each object keeps its native pointer in a tiny extension type, so methods such as draw or set_bind_group reach the pointer without any Python attribute lookups.
    • Descriptors are C structs filled from Dawn's *_INIT defaults, so the C compiler checks every field against the header.
    • The few differences in the browser are marked with _IS_EMSCRIPTEN (see below).
  • Build and ABI: tools/build_dawn.py builds the extension.

    • Natively, it builds against an installed Dawn, with Cython and a C compiler. A C++ compiler is also used for one small helper that needs Dawn's C++ API (adapter enumeration, see below); it is skipped if Dawn's C++ headers are missing. On CPython ≥ 3.12 it builds against the Stable ABI (abi3), and the Limited API has no measurable cost here.
    • For Pyodide, see below.
    • The wheel build hook can also build it, opt-in with WGPU_PY_BUILD_DAWN=1.

Pyodide

How it works (the approach is from #840):

  • Emdawnwebgpu consists of C++ (webgpu.cpp) and an Emscripten JS library. A Pyodide extension is an Emscripten side module, and side modules cannot carry JS libraries.
  • So, at build time, tools/gen_emdawn_glue.py links a throw-away program with --use-port=emdawnwebgpu using Pyodide's Emscripten, and extracts the expanded JS. That JS ships in the wheel as emdawn_glue.js.
  • webgpu.cpp and emdawn_glue.cpp are compiled into the extension. At import, an EM_JS function evals the glue in Pyodide's module scope and registers its functions in wasmImports, where the extension's (lazy) wgpu* import stubs find them.
  • The glue generator also patches one bug in Emdawnwebgpu: wgpuBufferGetMappedRange returns uninitialized memory instead of the buffer contents (HEAPU8.fill(0, data, mapped.byteLength) passes a size where an end index is expected).
  • The wheel is built with WGPU_PY_BUILD_NOARCH=1 WGPU_PY_BUILD_DAWN=1 pyodide build --exports pyinit. It is one wgpu wheel (about 520 KB) with the extension and no wgpu-native. The pinned Emdawnwebgpu package is downloaded and sha256-checked.

Async model:

  • Natively (unchanged): callbacks run when wgpuInstanceProcessEvents is called (AllowProcessEvents), so they always run on a thread we control.
    • sync_wait() blocks on wgpuInstanceWaitAny (TimedWaitAny) with the GIL released, so it wakes up as soon as the future completes.
    • await polls wgpuInstanceProcessEvents.
    • .then() is driven by a small pump that schedules process_events() on the event loop's thread.
    • I kept this instead of Dawn backend via cffi that runs natively and in Pyodide (Emdawnwebgpu) #840's polling, because the blocking wait has no polling latency.
  • In the browser, callbacks are AllowSpontaneous: the JS event loop calls them when the JS promise resolves.
    • await is plain asyncio.
    • sync_wait() suspends the Python stack with JSPI (pyodide.ffi.run_sync), waiting on the promise's async event. This needs a runtime with JSPI and code that is run via pyodide.runPythonAsync(); otherwise it raises and points to the async API.
    • request_adapter/request_device do not block.
  • Validation errors: they are still raised at the call site when they are reported during the call (Dawn in Node.js does that). In Chrome they arrive later, via the uncapturederror event. They are then logged from the event loop, instead of being raised by a later, unrelated call.
  • Mapped buffers: Emdawnwebgpu copies each accessed range between JS and the wasm heap, and the browser does not allow overlapping ranges. So in the browser the backend fetches the whole mapped range once. Natively, copies use wgpuBufferReadMappedRange/WriteMappedRange.
  • Canvas: the surface is a <canvas> element (via a CSS selector), and the browser presents it. This is verified in headless Chrome: the triangle example renders to the canvas, and a screenshot is uploaded by CI.

Adapter enumeration and naming

pygfx selects an adapter by name: PYGFX_WGPU_ADAPTER_NAME=llvmpipe is matched against adapter.summary from enumerate_adapters_sync(). Its screenshot tests are enabled only if adapter.info["vendor"] == "llvmpipe". With the Dawn backend neither worked:

  • Enumeration: webgpu.h cannot enumerate adapters, and requestAdapter returns Dawn's "best" adapter. forceFallbackAdapter only matches SwiftShader, so lavapipe was never listed when a GPU is present.
    • The backend now uses Dawn's exported C++ API (dawn::native::Instance::EnumerateAdapters) in a small native-only helper, backends/dawn/dawn_native_extras.cpp. It lists for example the Intel GPU and llvmpipe (Vulkan).
    • Dawn's Null backend is left out.
    • Without Dawn's C++ headers, it falls back to requesting adapters, now also with fallback and compatibility options.
  • Vendor and description:
    • Dawn reports a normalized vendor ("mesa" for lavapipe, "intel") and, for Vulkan, "<driver name>: <driver info>" as the description.
    • wgpu-native reports the driver name as the vendor ("llvmpipe") and the driver info as the description.
    • The Dawn backend now matches wgpu-native, and keeps Dawn's vendor as info["vendor_name"]. This is a deliberate deviation from the spec meaning of vendor, to make the backends interchangeable. Reviewers may prefer otherwise.
  • Int enums: like wgpu-native, ints are now accepted for enum arguments, as raw webgpu.h values. pygfx uses cull_mode=0.

Result, run locally with pygfx main, lavapipe, PYGFX_WGPU_ADAPTER_NAME=llvmpipe and a GPU present:

  • The screenshot comparisons now run with Dawn.
  • examples/tests: Dawn and wgpu-native give identical results, 155 passed and 27 failed. 40 of 53 screenshot comparisons pass on both, and the same 13 fail on both. The failures are missing optional dependencies in my environment, plus screenshot differences that occur with wgpu-native too. I assume those are due to my Mesa/LLVM version, but I did not verify that.
  • pygfx's tests/ also give identical results for both backends (354 passed, 15 failed).

Evidence

1. Raw binding micro-benchmark (tools/binding_benchmark/)

  • Setup: Dawn Null backend, Python 3.13, Intel Core Ultra 7 270K; times are ns per call, best of 7.
  • All variants call the same function, wgpuBufferGetSize, plus an "encode+submit" sequence of 11 calls.
  • The "method" column is the realistic case: obj.get_size(). For cffi that means a Python wrapper class, which is how wgpu-py works today. For the compiled variants it is a native method.
binding bound func obj.get_size() compiled func behind a Python wrapper class encode+submit (11 calls)
cffi ABI (dlopen, current approach) 46.6 67.3 67.3 1.64 µs
cffi API / out-of-line (abi3) 30.6 52.7 52.7 1.39 µs
Cython 3 (full API) 12.9 13.3 36.0 0.83 µs
Cython 3 (Limited API, abi3) 13.6 13.4 36.8 0.92 µs
nanobind (stable ABI) 14.6 15.8 37.6 0.97 µs
pybind11 29.0 42.4 49.6 1.29 µs
pure C loop (Dawn's own cost) 0.7 – – 0.67 µs
  • Fastest: Cython and nanobind are effectively tied, at about 4–5× less overhead than cffi ABI mode. Cython's Limited API build costs nothing here.
  • Where the wrapper class lives: keeping a plain Python wrapper class in front of a fast binding gives back most of the gain (about 36 ns instead of 13 ns). So to be fastest, the class methods themselves have to be compiled.
  • Class layout: tools/binding_benchmark/bench_cls.py compares three ways to do that. The layout used here, a Python API class combined with an extension base that holds the pointer, costs 13.6 ns, or 14.1 ns with abi3. A plain compiled class that stores the pointer as an int costs 18.9 ns, or 24.4 ns with abi3.
  • Why Cython over nanobind: the bindings are C, not C++. The Python-like syntax lets the existing _api.py logic be ported almost line by line. Its Limited API build is just as fast.

2. Through the public wgpu-py API, natively (python tools/bench_backends.py --all)

  • Same machine as above, both backends on the Intel iGPU via Vulkan; times are ns per call, best of 5.
  • The remaining per-call time for the Dawn backend is mostly Dawn's own validation. For example, a raw DispatchWorkgroups call costs about 180 ns on its own.
call wgpu-native (cffi ABI) dawn (Cython)
render_pass.draw(3) 322 46
render_pass.set_bind_group(0, bg) 472 115
render_pass.set_viewport(...) 339 70
compute_pass.dispatch_workgroups(1) 472 247
queue.write_buffer(64 bytes) 2,412 2,528
device.create_bind_group(2 entries) 6,536 2,293
frame: encoder + pass + 10 draws + submit 50,378 37,300

3. Through the public wgpu-py API, in Pyodide (tools/bench_backends.py run with tools/dawn_pyodide)

  • Same machine, Pyodide 314.0.7; times are ns per call, best of 2 runs (each best of 5).
  • The Dawn backend of this PR (Cython) is compared with the cffi (API mode) Dawn backend of Dawn backend via cffi that runs natively and in Pyodide (Emdawnwebgpu) #840, with the same Emdawnwebgpu.
  • Node.js 26 uses the webgpu npm package 0.6.1 (Dawn) on lavapipe. Chrome is headless with SwiftShader.
call Node: cffi (#840) Node: Cython Chrome: cffi (#840) Chrome: Cython
render_pass.draw(3) 704 249 463 58
render_pass.set_bind_group(0, bg) 1,250 469 729 90
render_pass.set_viewport(...) 1,181 307 889 156
compute_pass.dispatch_workgroups(1) 1,179 449 671 73
queue.write_buffer(64 bytes) 3,220 1,108 2,305 460
device.create_bind_group(2 entries) 22,350 10,538 16,120 4,940
frame: encoder + pass + 10 draws + submit 93,368 67,199 72,150 59,300
  • In Pyodide, Cython has about 2–9× less per-call overhead than cffi.
  • The frame numbers depend mostly on the GPU and on waiting, and they vary between runs. In Node, the frame times ranged from 67–105 µs with Cython and 93–187 µs with cffi.
  • CI runners are slower but show the same pattern, for example draw 806 ns in Node, 154 ns in Chrome, and 128 ns natively on lavapipe.
  • The native CI benchmark now also runs wgpu-native, on the same lavapipe adapter. For example, draw was 914 ns with wgpu-native versus 129 ns with Dawn, and create_bind_group 33.6 µs versus 13.4 µs.

Trade-offs

  • Contributors who work on the Dawn backend need Cython and a C (and C++) compiler; for Pyodide, also pyodide-build and its Emscripten. The rest of wgpu-py still needs no compiler.
  • cffi API mode (Dawn backend via cffi that runs natively and in Pyodide (Emdawnwebgpu) #840) would allow reusing the current _api.py almost unchanged. It is about 4× slower per call than Cython natively, and about 2–9× slower in Pyodide.
  • The backend is a port of _api.py rather than shared code. Struct field use is checked by the Cython and C compilers instead of by codegen annotations.
  • Dawn is not thread-safe by default, so promises never touch Dawn from a background thread (see the async model above).
  • The Emdawnwebgpu glue depends on the Emscripten version of Pyodide and on the Dawn release. It is regenerated at every build, and the glue patch fails loudly if Emdawnwebgpu changes.

Version constraints

  • Dawn v20260929.200159 everywhere: the vendored headers, the conda package dawn 20260929.200159, and Emdawnwebgpu (emdawnwebgpu_pkg-v20260929.200159.zip, pinned with sha256 in tools/build_dawn.py). Updating Dawn means updating all three together and rerunning the codegen.
  • Native: Cython ≥ 3.1. abi3 on CPython ≥ 3.12; CI tests Python 3.12 and 3.14. Dawn is linux-64 only for now.
  • Pyodide: Pyodide 314.0.7 (Python 3.14, Emscripten 5.0.3), pyodide-build 0.39.1, built with --exports pyinit. The wheel is for that Pyodide ABI (cp314-pyemscripten_2026_0_wasm32).
  • Sync API in the browser: requires JSPI and runPythonAsync. Tested in Node.js 26 and in Chrome as installed on ubuntu-latest.

How Dawn is obtained

  • Dawn is being packaged for conda-forge in Add Google Dawn -- a webgpu implementation conda-forge/staged-recipes#35015.
  • Until that lands, CI uses the temporary mark.harfouche channel: -c mark.harfouche -c conda-forge dawn. The package is dawn 20260929.200159, linux-64 only for now.
  • The package contains libwebgpu_dawn, webgpu.h, webgpu_cpp.h, Dawn's C++ headers and a CMake config.
  • For Pyodide, nothing needs to be installed: the Emdawnwebgpu package from Dawn's GitHub release is downloaded at build time.

Status

Tests (CI, lavapipe / SwiftShader):

where selection result
native, Python 3.12 and 3.14 Dawn tests plus the generic tests that apply 172 passed, 2 skipped
native, full suite (informational) pytest tests 193 passed, 14 failed, 11 errors
Pyodide in Node.js (webgpu npm on lavapipe) 11 test files (PYODIDE_TESTS in the workflow) 123 passed, 3 skipped
Pyodide in headless Chrome (SwiftShader) same, minus 5 tests that expect a validation error at the call site 118 passed, 3 skipped
  • Ported from Dawn backend via cffi that runs natively and in Pyodide (Emdawnwebgpu) #840 into test_dawn_backend.py: encoder errors raised at finish(), set_bind_group with dynamic offsets, and a render bundle with a depth_stencil_format.
  • Also ported from Dawn backend via cffi that runs natively and in Pyodide (Emdawnwebgpu) #840:
    • The base GPUCompilationMessage/GPUCompilationInfo in wgpu/_classes.py now store their values.
    • The codegen struct check parses the backend being patched, instead of always parsing wgpu-native.
    • The codegen report lists unimplemented methods.
    • The shared files (backends/__init__.py, conftest.py, tools/bench_backends.py) are identical in both PRs.
    • auto.py differs only in how each PR detects its compiled part in Pyodide. Here it checks for the extension file, because importing the package would register the backend.
  • The Pyodide runs also include the example scripts in tools/dawn_pyodide: compute async/sync, and a triangle rendered offscreen and (in Chrome) to a <canvas>.
  • The native full-suite failures are wgpu-native-specific or real Dawn differences:
    • naga error-message formats
    • native-only features such as immediates and pipeline statistics
    • the poller and repr
    • Dawn's hard limit of 4 bind groups
    • test shaders that Tint rejects as invalid WGSL
  • Modules that import wgpu.backends.wgpu_native directly are not collected when Dawn is the active backend.
  • The invalid test shaders are fixed separately, in the independent draft Add @interpolate(flat) to integer varyings in test shaders #841 (@interpolate(flat) on integer varyings). This PR does not duplicate that change.

Works (verified):

  • Adapter and device requests, including features and limits; adapter info; adapter enumeration (natively all adapters, in the browser the browser's adapter).
  • Buffers: create, write_buffer, read_buffer, map read/write (including overlapping and unaligned access in the browser).
  • WGSL shaders. SPIR-V natively only; the browser has no SPIR-V.
  • Compute pipelines and passes.
  • Render pipelines and passes:
    • render to texture and read back pixels
    • vertex and index buffers
    • indirect and indexed draws
    • render bundles
    • depth/stencil
    • occlusion queries
  • Textures, views, samplers, copies.
  • Error raising (natively, and in Node.js), error scopes, device.lost.
  • async/await and .then(); sync_wait() in Pyodide via JSPI.
  • Canvas presentation in the browser (headless Chrome).
  • pygfx's examples and tests on lavapipe, with results identical to wgpu-native (run locally; see above).

Not implemented or not tested yet:

  • GLSL shaders (Dawn does not support them).
  • Pipeline statistics queries.
  • Texture view swizzle.
  • Native canvas/surface presentation: the code is ported but untested.
  • In the browser:
    • validation errors are logged, not raised by the failing call
    • the sync API needs JSPI
    • test_api.py, test_async.py and test_canvas.py are not run in Pyodide: they need subprocesses, trio, or a rendercanvas install in the runner
  • macOS and Windows: the code is written to be portable, but no Dawn packages exist for them yet.

CI

.github/workflows/dawn.yml is kept separate from the main CI:

  • Native (unchanged approach): micromamba with the channels mark.harfouche then conda-forge, on ubuntu-latest with lavapipe, for Python 3.12 and 3.14.
    • Builds the extension and runs the Dawn tests plus the generic tests that apply, and the example scripts.
    • Reports the full suite and the benchmark; those steps are informational.
  • Codegen: checks that the generated files are up to date.
  • Pyodide, modeled on Dawn backend via cffi that runs natively and in Pyodide (Emdawnwebgpu) #840:
    • One job builds the wheel.
    • One job runs the scripts and tests in Node.js (webgpu npm, lavapipe).
    • One job runs them in headless Chrome (SwiftShader), with Pyodide from the CDN.
    • Both run the benchmark (informational). The native job runs it for Dawn and wgpu-native.
    • The Chrome job uploads a pyodide-demo artifact (canvas screenshot, plus a static site with the wheel and runner.html).
  • Pages (fork only): on pushes to dawn, the demo is deployed to https://www.markharfouche.com/wgpu-py/.

CI runs:

Other changes:

  • backends/auto.py respects an already-registered backend and the WGPUPY_BACKEND variable. In Pyodide, it prefers the Dawn backend if its extension is installed.
  • Registering a second, different backend now raises an error. Before, it silently replaced the first one.
  • tests/renderutils.py no longer imports wgpu_native.
  • tools/binding_benchmark/ exists to support this comparison and can be dropped before merging.

https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG

<details><summary>Claude's draft</summary>

Vendor Dawn's webgpu.h (v20260929.200159, from the conda package) as
wgpu/resources/dawn_webgpu.h and extend the codegen to generate, from it:

* backends/dawn/_webgpu.pxd: the complete Cython declaration of the C API
  (handles, enums, flags, structs, callbacks, 276 functions, constants and
  the WGPU_*_INIT struct initializers), parsed with pycparser after the
  same cleaning step that is used for the wgpu-native headers. It also
  contains generated helpers to convert WGPULimits from/to a dict.
* backends/dawn/_mappings.py: enum string <-> int maps for Dawn.

webgpu.h is used rather than dawn.json because it is what the conda
package ships, it is the exact ABI that is compiled against, and the
existing hparser already digests it.

The codegen also checks the Cython backend (_api.pyx) against the base
API and lists methods that are not implemented in the report.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary>

Add wgpu.backends.dawn, a backend based on Google's Dawn, alongside the
wgpu-native backend. It implements the same public classes (from
wgpu._classes), so user code and the test suite can select it.

The backend is a single Cython module (_api.pyx) whose classes mirror the
wgpu-native backend: regular Python classes that subclass the public API,
but whose methods are compiled and call Dawn's webgpu.h directly. Each
object stores its native pointer in a small extension type, so hot-path
methods (draw, set_bind_group, ...) need no Python attribute lookups.
Descriptors are filled as C structs on the stack (or in a small arena),
using Dawn's WGPU_*_INIT initializers. Validation errors, which Dawn
reports synchronously via the uncaptured-error callback, are raised as
exceptions at the call site by checking a C flag after each call.
Promises wrap Dawn futures: sync_wait() uses wgpuInstanceWaitAny with the
GIL released, awaiting polls wgpuInstanceProcessEvents, and then() is
driven by a pump that schedules process_events() on the event loop.

The extension is optional: build it with tools/build_dawn.py against an
installed Dawn (e.g. from conda). On CPython >= 3.12 it is built against
the Stable ABI. Select it with `import wgpu.backends.dawn` or with
WGPUPY_BACKEND=dawn; the auto backend now also respects an already
registered backend, and registering a second, different backend raises.

Not implemented yet: GLSL shaders, pipeline statistics queries,
texture view swizzle, and canvas/surface presentation is untested.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary>

* .github/workflows/dawn.yml: separate from the main CI. Installs Dawn,
  Cython and a C compiler with micromamba (channels mark.harfouche, then
  conda-forge), builds the Dawn backend, and runs the Dawn tests plus the
  generic tests that apply, on lavapipe. Also reports on the full test
  suite and runs the API benchmark (informational). Runs on pushes to
  main/dawn and on PRs.
* tools/binding_benchmark: micro-benchmark comparing cffi ABI, cffi API,
  Cython (full and Limited API), nanobind and pybind11 for calling the
  same Dawn functions, plus the class layout used by the backend.
* tools/bench_backends.py: per-call overhead through the public wgpu-py
  API, for the wgpu-native and Dawn backends.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
hmaarrfk added a commit to hmaarrfk/wgpu-py that referenced this pull request Oct 1, 2026
<details><summary>Claude's draft</summary>

Adopt the backend selection changes of the Cython Dawn backend PR (pygfx#839),
with identical content, so that both PRs can share them:

* auto.py uses a backend that is already registered (e.g. after
  `import wgpu.backends.dawn`), so that `wgpu.gpu.request_adapter_sync()`
  no longer loads (and silently switches to) wgpu-native in that case. It
  also honours the WGPUPY_BACKEND env var.
* _register_backend() raises if a second, different backend is registered.
* conftest.py loads the backend from WGPUPY_BACKEND before collection.
* tools/bench_backends.py: the per-call overhead benchmark from pygfx#839.

auto.py additionally selects the Dawn backend in Pyodide when the compiled
wgpu_dawn package is installed.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
hmaarrfk added a commit to hmaarrfk/wgpu-py that referenced this pull request Oct 1, 2026
<details><summary>Claude's draft</summary>

Performance, keeping cffi (API mode) as the binding layer:

* build_cffi.py post-processes the C code that cffi generates, so that the
  wrappers of wgpuRenderPassEncoder*/wgpuComputePassEncoder*/
  wgpuRenderBundleEncoder* functions keep the GIL and skip the errno
  save/restore. These calls are cheap, never block, and never call back into
  Python. All other functions still release the GIL.
* Hot pass encoder methods (draw, set_pipeline, set_bind_group,
  set_vertex_buffer, set_viewport, dispatch_workgroups, ...) call lib
  directly: Dawn defers encoder errors to finish(), so no error-checking
  wrapper is needed. set_bind_group() without dynamic offsets does no
  allocation. Objects are no longer kept alive from Python in passes and
  bundles; Dawn holds its own references.
* The error-checking wrapper for other calls is now a call plus a list
  truthiness test (was a thread-local capture stack). Validation errors that
  Dawn reports during a call are still raised at the call site natively, and
  in Pyodide when the WebGPU implementation reports them during the call.
* Sync waits block in wgpuInstanceWaitAny() (GIL released) on the promise's
  future, instead of polling with sleeps.
* create_bind_group() writes entries directly into a C array, and
  write_buffer() avoids an address round trip.

API, matching the Cython backend's tests (tests/test_dawn_backend.py, copied
unchanged from pygfx#839, passes):

* push_error_scope() / pop_error_scope_async().
* get_compilation_info_async() returns real messages. The base
  GPUCompilationMessage/GPUCompilationInfo now store their values.
* The device-lost promise (_get_lost_async).
* promise.then() works natively: a light thread schedules process_events()
  in the event loop while promises are pending.
* Released objects raise RuntimeError instead of cffi's TypeError.
* process_events is exported from wgpu.backends.dawn.

Adapters:

* enumerate_adapters_sync() returns all adapters natively, for all
  backends, including CPU adapters like lavapipe. webgpu.h can only request
  the "best" adapter, so wgpu_dawn now includes a small C++ helper that
  calls dawn::native::Instance::EnumerateAdapters (the native build is now
  C++20).
* On Vulkan, info["vendor"] is the driver name, as with wgpu-native (Dawn
  puts it at the start of the description). This lets downstream code that
  detects lavapipe with info["vendor"] == "llvmpipe" work unchanged.

Fixes: create_render_bundle_encoder() with a depth_stencil_format.

Codegen: the struct-check validation now parses the backend being patched
(it always parsed the wgpu-native backend), and the report lists API
methods that a backend does not implement.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary>

The same _api.pyx now also compiles for Pyodide, against Emdawnwebgpu
(Dawn's webgpu.h on top of the browser's WebGPU), so one backend runs
natively and in the browser. Adapted from the cffi-based work in pygfx#840.

Codegen:
* Vendor Emdawnwebgpu's webgpu.h (same Dawn release, v20260929.200159) as
  wgpu/resources/emdawn_webgpu.h. It is a subset of Dawn's native header.
* The main block of _webgpu.pxd is now generated from it, so it is valid for
  both targets. The few native-only declarations that the backend uses
  (logging callback, wgpuDeviceTick, native surface sources) are listed in
  NATIVE_ONLY and declared in a separate block, from the generated
  dawn_native_only.h. That header includes webgpu.h natively, and with
  Emscripten defines them (copied from Dawn's header) with no-op functions.
* The codegen checks that the common declarations and enum values are the
  same in both headers.

Build (tools/build_dawn.py): when cross-compiling for Emscripten (e.g. with
`pyodide build`), download the pinned Emdawnwebgpu package (sha256-checked),
generate the JS glue with gen_emdawn_glue.py (from pygfx#840: link a throw-away
program with --use-port=emdawnwebgpu, extract the expanded JS library), and
compile Emdawnwebgpu's webgpu.cpp plus emdawn_glue.cpp into the extension.
At import, the extension evals the glue in Pyodide's module scope and puts
its functions in wasmImports. The glue generator also fixes a bug in
Emdawnwebgpu's wgpuBufferGetMappedRange (uninitialized memory instead of
the buffer contents). The hatch hook builds it with WGPU_PY_BUILD_DAWN=1 and
adds Cython/setuptools as build dependencies; a Pyodide wheel is built with
`WGPU_PY_BUILD_NOARCH=1 WGPU_PY_BUILD_DAWN=1 pyodide build --exports pyinit`.

Backend (_api.pyx), with _IS_EMSCRIPTEN:
* Callbacks use AllowSpontaneous, so the JS event loop calls them. Natively
  the model is unchanged (AllowProcessEvents, TimedWaitAny for sync waits,
  an event pump for then()).
* sync_wait() suspends with JSPI (pyodide.ffi.run_sync) and waits for the
  promise's async event; awaiting is plain asyncio.
* request_adapter/request_device do not block in the browser.
* Validation errors are still raised at the call site when reported during
  the call (Dawn in Node.js does that); otherwise (Chrome) they are logged
  from the event loop instead of being raised by a later, unrelated call.
* Surfaces are <canvas> elements (CSS selector); the browser presents.
* Mapped ranges: copying is done with wgpuBufferRead/WriteMappedRange
  natively. In the browser, the whole mapped range is fetched once, because
  Emdawnwebgpu copies each accessed range and the browser does not allow
  overlapping ranges.
* enumerate_adapters returns the browser's adapter.

In Pyodide, the Dawn backend is selected automatically if it is built.
tools/dawn_pyodide has runners for Node.js (webgpu npm package) and
headless Chrome, and example scripts that also run natively.

In headless Chrome, the tests that expect a validation error to be raised
by the failing call do not pass, because the browser reports it later.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary>

pygfx selects an adapter by name (PYGFX_WGPU_ADAPTER_NAME=llvmpipe matches
against adapter.summary in enumerate_adapters_sync()), and enables its
screenshot tests only if adapter.info["vendor"] == "llvmpipe". With the
Dawn backend neither worked:

* enumerate_adapters_sync() requested one adapter per backend and power
  preference. Dawn always returns the "best" one, so a CPU adapter like
  lavapipe was never listed when there is a GPU (forceFallbackAdapter only
  matches SwiftShader). webgpu.h has no way to enumerate adapters, so this
  uses Dawn's C++ API (dawn::native::Instance::EnumerateAdapters) in a small
  native-only helper, dawn_native_extras.cpp. It is compiled if Dawn's C++
  headers are available (they are in the conda package); otherwise the
  previous request-based enumeration is used, which now also tries fallback
  adapters and the compatibility level for OpenGL. Dawn's Null backend is
  left out.
* Dawn reports a normalized vendor ("mesa" for lavapipe, "intel") and, for
  Vulkan, "<driver name>: <driver info>" as the description. wgpu-native
  reports the driver name as vendor ("llvmpipe") and the driver info as
  description. The Dawn backend now does the same, and keeps Dawn's vendor
  as info["vendor_name"].

Also accept ints for enum arguments, passed through as the raw webgpu.h
value, like the wgpu-native backend (pygfx uses cull_mode=0).

With this, pygfx's example tests on lavapipe give the same results with the
Dawn backend as with wgpu-native, including the screenshot comparisons.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary>

Modeled on the workflow of pygfx#840, without the GitHub Pages deployment:

* pyodide-build: build the wgpu wheel with the Dawn extension for Pyodide
  (Pyodide 314.0.7, pyodide-build 0.39.1, Pyodide's Emscripten).
* pyodide-node: run the example scripts and the wgpu-py tests that apply in
  Pyodide in Node.js, with navigator.gpu from the `webgpu` npm package
  (Dawn) on lavapipe. Also run the call-overhead benchmark.
* pyodide-chrome: the same in headless Chrome (SwiftShader), with Pyodide
  from the CDN. Uploads a screenshot of the canvas and the demo page (a
  static site with the wheel) as an artifact.

The native jobs now install a C++ compiler (for dawn_native_extras.cpp),
print the enumerated adapters, and also run the example scripts.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
Comment thread tools/bench_backends.py
@@ -0,0 +1,226 @@
"""
Benchmark the per-call overhead of the wgpu-py API for a given backend.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Have you ran this? I'm curious to the results :)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The long Claude vomit claims it did.

Cython is fastest from what Claude told me.

I did run my application with dawn and it seemed to be working with all this.

I think…

I kinda did this whole effort while working on other things

<details><summary>Claude's draft</summary>

Port what pygfx#840 gained that is not specific to cffi:

* wgpu/_classes.py: the base GPUCompilationMessage and GPUCompilationInfo
  now store their values (same change as in pygfx#840), so the Dawn backend's
  subclasses no longer duplicate that code.
* codegen: the struct-check validation parses the backend being patched,
  instead of always parsing the wgpu-native backend, and the report lists API
  methods that a (Python) backend does not implement.
* tests/test_dawn_backend.py: the tests that pygfx#840 added in
  test_dawn_cffi.py: encoder errors raised at finish(), set_bind_group with
  dynamic offsets (list, numpy, start/length), and a render bundle with a
  depth_stencil_format (pygfx#839 did not have pygfx#840's bug there; the test now
  also executes the bundle in a pass with a depth attachment). Also check
  that enumerated adapters are unique and can create a device.

The dynamic-offsets test found a bug in this backend: set_bind_group()
tested `if dynamic_offsets_data:`, which drops a numpy array [0] (and raises
for longer arrays). It now checks the length.

auto.py: in Pyodide, detect the Dawn extension by its file instead of with
find_spec(), which imported (and so registered) the backend as a side effect.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary>

* tools/dawn_pyodide/demo: a demo page (adapted from pygfx#840's): it installs
  the wgpu wheel with the Cython Dawn backend in Pyodide, prints the adapter
  info, runs a compute shader on 1M floats and checks the result, and renders
  a rotating triangle to a <canvas>, through the browser's WebGPU.
* New `pages` job (pushes to `dawn` on the fork only, after the Pyodide jobs
  pass): deploys that demo to the fork's GitHub Pages. It also builds the
  demo of the cffi variant from the dawn-pyodide branch (pygfx#840) and puts it
  under /cffi/, so both stay reachable on one site. If that build fails, a
  placeholder is deployed at /cffi/ and the main demo still deploys.
* The native benchmark step now downloads wgpu-native and runs
  `bench_backends.py --all`, for a Dawn vs wgpu-native table (informational).
  The Node and Chrome benchmark steps were already there.
* Deselect test_encoder_errors_raise_at_finish in Chrome: like the other
  call-site error tests, the browser reports the error asynchronously.

Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>

Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
@hmaarrfk

hmaarrfk commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

I want to set expectations on this front. For me, this was mostly an experiment on how perhaps we could execute a rapid transition to a different backend, trying to reduce the "fear" of the "large change" and to show how it could be achieved quickly, so that we may get to the fun parts (for me) of making use of the new features afforded by Dawn.

Shared buffers for display and computation are really exciting for me.

Some benchmarks really excite me
call wgpu-native (cffi ABI) dawn (Cython)
render_pass.draw(3) 322 46
render_pass.set_bind_group(0, bg) 472 115
render_pass.set_viewport(...) 339 70
compute_pass.dispatch_workgroups(1) 472 247
queue.write_buffer(64 bytes) 2,412 2,528
device.create_bind_group(2 entries) 6,536 2,293
frame: encoder + pass + 10 draws + submit 50,378 37,300

However, I realize that there nothing more frustrating than "talking to claude through a human intermediary".

I don't intend for the next few iterations to be "human written" by opening a conventional code editor and typing.

I do agree that claude's organization can use some work, but I would fix it, by vibing.

If for any reason this means that this PR should just be closed, or remain dormant as an idea, I'm more than happy to back away from it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants