Repository navigation
Conversation
<details><summary>Claude's draft</summary> Vendor Dawn's webgpu.h (v20260929.200159, from the conda package) as wgpu/resources/dawn_webgpu.h and extend the codegen to generate, from it: * backends/dawn/_webgpu.pxd: the complete Cython declaration of the C API (handles, enums, flags, structs, callbacks, 276 functions, constants and the WGPU_*_INIT struct initializers), parsed with pycparser after the same cleaning step that is used for the wgpu-native headers. It also contains generated helpers to convert WGPULimits from/to a dict. * backends/dawn/_mappings.py: enum string <-> int maps for Dawn. webgpu.h is used rather than dawn.json because it is what the conda package ships, it is the exact ABI that is compiled against, and the existing hparser already digests it. The codegen also checks the Cython backend (_api.pyx) against the base API and lists methods that are not implemented in the report. Resume this Claude session: ``` cd /home/mark/git/feedstock/staged-recipes claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651 ``` </details> Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary> Add wgpu.backends.dawn, a backend based on Google's Dawn, alongside the wgpu-native backend. It implements the same public classes (from wgpu._classes), so user code and the test suite can select it. The backend is a single Cython module (_api.pyx) whose classes mirror the wgpu-native backend: regular Python classes that subclass the public API, but whose methods are compiled and call Dawn's webgpu.h directly. Each object stores its native pointer in a small extension type, so hot-path methods (draw, set_bind_group, ...) need no Python attribute lookups. Descriptors are filled as C structs on the stack (or in a small arena), using Dawn's WGPU_*_INIT initializers. Validation errors, which Dawn reports synchronously via the uncaptured-error callback, are raised as exceptions at the call site by checking a C flag after each call. Promises wrap Dawn futures: sync_wait() uses wgpuInstanceWaitAny with the GIL released, awaiting polls wgpuInstanceProcessEvents, and then() is driven by a pump that schedules process_events() on the event loop. The extension is optional: build it with tools/build_dawn.py against an installed Dawn (e.g. from conda). On CPython >= 3.12 it is built against the Stable ABI. Select it with `import wgpu.backends.dawn` or with WGPUPY_BACKEND=dawn; the auto backend now also respects an already registered backend, and registering a second, different backend raises. Not implemented yet: GLSL shaders, pipeline statistics queries, texture view swizzle, and canvas/surface presentation is untested. Resume this Claude session: ``` cd /home/mark/git/feedstock/staged-recipes claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651 ``` </details> Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary> * .github/workflows/dawn.yml: separate from the main CI. Installs Dawn, Cython and a C compiler with micromamba (channels mark.harfouche, then conda-forge), builds the Dawn backend, and runs the Dawn tests plus the generic tests that apply, on lavapipe. Also reports on the full test suite and runs the API benchmark (informational). Runs on pushes to main/dawn and on PRs. * tools/binding_benchmark: micro-benchmark comparing cffi ABI, cffi API, Cython (full and Limited API), nanobind and pybind11 for calling the same Dawn functions, plus the class layout used by the backend. * tools/bench_backends.py: per-call overhead through the public wgpu-py API, for the wgpu-native and Dawn backends. Resume this Claude session: ``` cd /home/mark/git/feedstock/staged-recipes claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651 ``` </details> Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary> Adopt the backend selection changes of the Cython Dawn backend PR (pygfx#839), with identical content, so that both PRs can share them: * auto.py uses a backend that is already registered (e.g. after `import wgpu.backends.dawn`), so that `wgpu.gpu.request_adapter_sync()` no longer loads (and silently switches to) wgpu-native in that case. It also honours the WGPUPY_BACKEND env var. * _register_backend() raises if a second, different backend is registered. * conftest.py loads the backend from WGPUPY_BACKEND before collection. * tools/bench_backends.py: the per-call overhead benchmark from pygfx#839. auto.py additionally selects the Dawn backend in Pyodide when the compiled wgpu_dawn package is installed. Resume this Claude session: ``` cd /home/mark/git/feedstock/staged-recipes claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651 ``` </details> Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary> Performance, keeping cffi (API mode) as the binding layer: * build_cffi.py post-processes the C code that cffi generates, so that the wrappers of wgpuRenderPassEncoder*/wgpuComputePassEncoder*/ wgpuRenderBundleEncoder* functions keep the GIL and skip the errno save/restore. These calls are cheap, never block, and never call back into Python. All other functions still release the GIL. * Hot pass encoder methods (draw, set_pipeline, set_bind_group, set_vertex_buffer, set_viewport, dispatch_workgroups, ...) call lib directly: Dawn defers encoder errors to finish(), so no error-checking wrapper is needed. set_bind_group() without dynamic offsets does no allocation. Objects are no longer kept alive from Python in passes and bundles; Dawn holds its own references. * The error-checking wrapper for other calls is now a call plus a list truthiness test (was a thread-local capture stack). Validation errors that Dawn reports during a call are still raised at the call site natively, and in Pyodide when the WebGPU implementation reports them during the call. * Sync waits block in wgpuInstanceWaitAny() (GIL released) on the promise's future, instead of polling with sleeps. * create_bind_group() writes entries directly into a C array, and write_buffer() avoids an address round trip. API, matching the Cython backend's tests (tests/test_dawn_backend.py, copied unchanged from pygfx#839, passes): * push_error_scope() / pop_error_scope_async(). * get_compilation_info_async() returns real messages. The base GPUCompilationMessage/GPUCompilationInfo now store their values. * The device-lost promise (_get_lost_async). * promise.then() works natively: a light thread schedules process_events() in the event loop while promises are pending. * Released objects raise RuntimeError instead of cffi's TypeError. * process_events is exported from wgpu.backends.dawn. Adapters: * enumerate_adapters_sync() returns all adapters natively, for all backends, including CPU adapters like lavapipe. webgpu.h can only request the "best" adapter, so wgpu_dawn now includes a small C++ helper that calls dawn::native::Instance::EnumerateAdapters (the native build is now C++20). * On Vulkan, info["vendor"] is the driver name, as with wgpu-native (Dawn puts it at the start of the description). This lets downstream code that detects lavapipe with info["vendor"] == "llvmpipe" work unchanged. Fixes: create_render_bundle_encoder() with a depth_stencil_format. Codegen: the struct-check validation now parses the backend being patched (it always parsed the wgpu-native backend), and the report lists API methods that a backend does not implement. Resume this Claude session: ``` cd /home/mark/git/feedstock/staged-recipes claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651 ``` </details> Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary> The same _api.pyx now also compiles for Pyodide, against Emdawnwebgpu (Dawn's webgpu.h on top of the browser's WebGPU), so one backend runs natively and in the browser. Adapted from the cffi-based work in pygfx#840. Codegen: * Vendor Emdawnwebgpu's webgpu.h (same Dawn release, v20260929.200159) as wgpu/resources/emdawn_webgpu.h. It is a subset of Dawn's native header. * The main block of _webgpu.pxd is now generated from it, so it is valid for both targets. The few native-only declarations that the backend uses (logging callback, wgpuDeviceTick, native surface sources) are listed in NATIVE_ONLY and declared in a separate block, from the generated dawn_native_only.h. That header includes webgpu.h natively, and with Emscripten defines them (copied from Dawn's header) with no-op functions. * The codegen checks that the common declarations and enum values are the same in both headers. Build (tools/build_dawn.py): when cross-compiling for Emscripten (e.g. with `pyodide build`), download the pinned Emdawnwebgpu package (sha256-checked), generate the JS glue with gen_emdawn_glue.py (from pygfx#840: link a throw-away program with --use-port=emdawnwebgpu, extract the expanded JS library), and compile Emdawnwebgpu's webgpu.cpp plus emdawn_glue.cpp into the extension. At import, the extension evals the glue in Pyodide's module scope and puts its functions in wasmImports. The glue generator also fixes a bug in Emdawnwebgpu's wgpuBufferGetMappedRange (uninitialized memory instead of the buffer contents). The hatch hook builds it with WGPU_PY_BUILD_DAWN=1 and adds Cython/setuptools as build dependencies; a Pyodide wheel is built with `WGPU_PY_BUILD_NOARCH=1 WGPU_PY_BUILD_DAWN=1 pyodide build --exports pyinit`. Backend (_api.pyx), with _IS_EMSCRIPTEN: * Callbacks use AllowSpontaneous, so the JS event loop calls them. Natively the model is unchanged (AllowProcessEvents, TimedWaitAny for sync waits, an event pump for then()). * sync_wait() suspends with JSPI (pyodide.ffi.run_sync) and waits for the promise's async event; awaiting is plain asyncio. * request_adapter/request_device do not block in the browser. * Validation errors are still raised at the call site when reported during the call (Dawn in Node.js does that); otherwise (Chrome) they are logged from the event loop instead of being raised by a later, unrelated call. * Surfaces are <canvas> elements (CSS selector); the browser presents. * Mapped ranges: copying is done with wgpuBufferRead/WriteMappedRange natively. In the browser, the whole mapped range is fetched once, because Emdawnwebgpu copies each accessed range and the browser does not allow overlapping ranges. * enumerate_adapters returns the browser's adapter. In Pyodide, the Dawn backend is selected automatically if it is built. tools/dawn_pyodide has runners for Node.js (webgpu npm package) and headless Chrome, and example scripts that also run natively. In headless Chrome, the tests that expect a validation error to be raised by the failing call do not pass, because the browser reports it later. Resume this Claude session: ``` cd /home/mark/git/feedstock/staged-recipes claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651 ``` </details> Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary>
pygfx selects an adapter by name (PYGFX_WGPU_ADAPTER_NAME=llvmpipe matches
against adapter.summary in enumerate_adapters_sync()), and enables its
screenshot tests only if adapter.info["vendor"] == "llvmpipe". With the
Dawn backend neither worked:
* enumerate_adapters_sync() requested one adapter per backend and power
preference. Dawn always returns the "best" one, so a CPU adapter like
lavapipe was never listed when there is a GPU (forceFallbackAdapter only
matches SwiftShader). webgpu.h has no way to enumerate adapters, so this
uses Dawn's C++ API (dawn::native::Instance::EnumerateAdapters) in a small
native-only helper, dawn_native_extras.cpp. It is compiled if Dawn's C++
headers are available (they are in the conda package); otherwise the
previous request-based enumeration is used, which now also tries fallback
adapters and the compatibility level for OpenGL. Dawn's Null backend is
left out.
* Dawn reports a normalized vendor ("mesa" for lavapipe, "intel") and, for
Vulkan, "<driver name>: <driver info>" as the description. wgpu-native
reports the driver name as vendor ("llvmpipe") and the driver info as
description. The Dawn backend now does the same, and keeps Dawn's vendor
as info["vendor_name"].
Also accept ints for enum arguments, passed through as the raw webgpu.h
value, like the wgpu-native backend (pygfx uses cull_mode=0).
With this, pygfx's example tests on lavapipe give the same results with the
Dawn backend as with wgpu-native, including the screenshot comparisons.
Resume this Claude session:
```
cd /home/mark/git/feedstock/staged-recipes
claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651
```
</details>
Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary> Modeled on the workflow of pygfx#840, without the GitHub Pages deployment: * pyodide-build: build the wgpu wheel with the Dawn extension for Pyodide (Pyodide 314.0.7, pyodide-build 0.39.1, Pyodide's Emscripten). * pyodide-node: run the example scripts and the wgpu-py tests that apply in Pyodide in Node.js, with navigator.gpu from the `webgpu` npm package (Dawn) on lavapipe. Also run the call-overhead benchmark. * pyodide-chrome: the same in headless Chrome (SwiftShader), with Pyodide from the CDN. Uploads a screenshot of the canvas and the demo page (a static site with the wheel) as an artifact. The native jobs now install a C++ compiler (for dawn_native_extras.cpp), print the enumerated adapters, and also run the example scripts. Resume this Claude session: ``` cd /home/mark/git/feedstock/staged-recipes claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651 ``` </details> Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
| @@ -0,0 +1,226 @@ | |||
| """ | |||
| Benchmark the per-call overhead of the wgpu-py API for a given backend. | |||
There was a problem hiding this comment.
Have you ran this? I'm curious to the results :)
There was a problem hiding this comment.
The long Claude vomit claims it did.
Cython is fastest from what Claude told me.
I did run my application with dawn and it seemed to be working with all this.
I think…
I kinda did this whole effort while working on other things
<details><summary>Claude's draft</summary> Port what pygfx#840 gained that is not specific to cffi: * wgpu/_classes.py: the base GPUCompilationMessage and GPUCompilationInfo now store their values (same change as in pygfx#840), so the Dawn backend's subclasses no longer duplicate that code. * codegen: the struct-check validation parses the backend being patched, instead of always parsing the wgpu-native backend, and the report lists API methods that a (Python) backend does not implement. * tests/test_dawn_backend.py: the tests that pygfx#840 added in test_dawn_cffi.py: encoder errors raised at finish(), set_bind_group with dynamic offsets (list, numpy, start/length), and a render bundle with a depth_stencil_format (pygfx#839 did not have pygfx#840's bug there; the test now also executes the bundle in a pass with a depth attachment). Also check that enumerated adapters are unique and can create a device. The dynamic-offsets test found a bug in this backend: set_bind_group() tested `if dynamic_offsets_data:`, which drops a numpy array [0] (and raises for longer arrays). It now checks the length. auto.py: in Pyodide, detect the Dawn extension by its file instead of with find_spec(), which imported (and so registered) the backend as a side effect. Resume this Claude session: ``` cd /home/mark/git/feedstock/staged-recipes claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651 ``` </details> Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
<details><summary>Claude's draft</summary> * tools/dawn_pyodide/demo: a demo page (adapted from pygfx#840's): it installs the wgpu wheel with the Cython Dawn backend in Pyodide, prints the adapter info, runs a compute shader on 1M floats and checks the result, and renders a rotating triangle to a <canvas>, through the browser's WebGPU. * New `pages` job (pushes to `dawn` on the fork only, after the Pyodide jobs pass): deploys that demo to the fork's GitHub Pages. It also builds the demo of the cffi variant from the dawn-pyodide branch (pygfx#840) and puts it under /cffi/, so both stay reachable on one site. If that build fails, a placeholder is deployed at /cffi/ and the main demo still deploys. * The native benchmark step now downloads wgpu-native and runs `bench_backends.py --all`, for a Dawn vs wgpu-native table (informational). The Node and Chrome benchmark steps were already there. * Deselect test_encoder_errors_raise_at_finish in Chrome: like the other call-site error tests, the browser reports the error asynchronously. Resume this Claude session: ``` cd /home/mark/git/feedstock/staged-recipes claude --resume 2cab6db9-a6ea-4976-a0e6-f4153fe5a651 ``` </details> Claude-Session: https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG
|
I want to set expectations on this front. For me, this was mostly an experiment on how perhaps we could execute a rapid transition to a different backend, trying to reduce the "fear" of the "large change" and to show how it could be achieved quickly, so that we may get to the fun parts (for me) of making use of the new features afforded by Dawn. Shared buffers for display and computation are really exciting for me. Some benchmarks really excite me
However, I realize that there nothing more frustrating than "talking to claude through a human intermediary". I don't intend for the next few iterations to be "human written" by opening a conventional code editor and typing. I do agree that claude's organization can use some work, but I would fix it, by vibing. If for any reason this means that this PR should just be closed, or remain dormant as an idea, I'm more than happy to back away from it. |
Note
This PR was written with the help of an AI assistant (Claude).
dawn
pyodide demo: https://www.markharfouche.com/wgpu-py/
claude
Draft. Adds an experimental, optional Dawn backend (
wgpu.backends.dawn) next to the wgpu-native backend. It is one Cython extension module that runs natively (against Dawn'slibwebgpu_dawn) and in Pyodide (against Emdawnwebgpu, Dawn'swebgpu.hon top of the browser's WebGPU). It implements the same public classes aswgpu/_classes.py, so existing code and tests can select it, either withimport wgpu.backends.dawnor withWGPUPY_BACKEND=dawn. In Pyodide it is selected automatically if it is installed. wgpu-native stays the default natively, and wgpu-py installs and works without Dawn.xref: #838. The Pyodide support is adapted from the cffi-based #840; this PR keeps Cython as the binding technology. The backend-agnostic fixes and tests that #840 gained since have been ported here too.
Live demo: https://www.markharfouche.com/wgpu-py/
Binding strategy: generated Cython
The goal was the lowest Python→Dawn call overhead, with bindings that are generated rather than written by hand.
Generated:
codegen/dawn_patcher.pyparses two vendored headers of the same Dawn release (v20260929.200159) with pycparser, using the same cleaning step as the wgpu-native headers:wgpu/resources/dawn_webgpu.h: Dawn's nativewebgpu.h.wgpu/resources/emdawn_webgpu.h: Emdawnwebgpu'swebgpu.h, which is a subset of the native one.From them it generates:
backends/dawn/_webgpu.pxd. Its main block comes from Emdawnwebgpu's header, so it is valid for both targets (207 functions, 91 structs, enums, flags, callbacks and theWGPU_*_INITinitializers). The codegen checks that these declarations and the enum values are identical in the native header.NATIVE_ONLYin the codegen: the logging callback,wgpuDeviceTick, and the native surface sources). They come from the generatedbackends/dawn/dawn_native_only.h, which just includeswebgpu.hnatively. With Emscripten, it defines them (copied verbatim from Dawn's header) with no-op functions, so that the same_api.pyxcompiles. The backend does not call them in the browser.backends/dawn/_mappings.py: enum mappings.The codegen also checks
_api.pyxagainst the base API and lists anything not implemented in the report (currently none).Hand-written, like
wgpu_native/_api.py:backends/dawn/_api.pyxmirrors the wgpu-native backend class by class.wgpu.classes.*, but Cython compiles their methods.draworset_bind_groupreach the pointer without any Python attribute lookups.*_INITdefaults, so the C compiler checks every field against the header._IS_EMSCRIPTEN(see below).Build and ABI:
tools/build_dawn.pybuilds the extension.WGPU_PY_BUILD_DAWN=1.Pyodide
How it works (the approach is from #840):
webgpu.cpp) and an Emscripten JS library. A Pyodide extension is an Emscripten side module, and side modules cannot carry JS libraries.tools/gen_emdawn_glue.pylinks a throw-away program with--use-port=emdawnwebgpuusing Pyodide's Emscripten, and extracts the expanded JS. That JS ships in the wheel asemdawn_glue.js.webgpu.cppandemdawn_glue.cppare compiled into the extension. At import, anEM_JSfunction evals the glue in Pyodide's module scope and registers its functions inwasmImports, where the extension's (lazy)wgpu*import stubs find them.wgpuBufferGetMappedRangereturns uninitialized memory instead of the buffer contents (HEAPU8.fill(0, data, mapped.byteLength)passes a size where an end index is expected).WGPU_PY_BUILD_NOARCH=1 WGPU_PY_BUILD_DAWN=1 pyodide build --exports pyinit. It is onewgpuwheel (about 520 KB) with the extension and no wgpu-native. The pinned Emdawnwebgpu package is downloaded and sha256-checked.Async model:
wgpuInstanceProcessEventsis called (AllowProcessEvents), so they always run on a thread we control.sync_wait()blocks onwgpuInstanceWaitAny(TimedWaitAny) with the GIL released, so it wakes up as soon as the future completes.awaitpollswgpuInstanceProcessEvents..then()is driven by a small pump that schedulesprocess_events()on the event loop's thread.AllowSpontaneous: the JS event loop calls them when the JS promise resolves.awaitis plain asyncio.sync_wait()suspends the Python stack with JSPI (pyodide.ffi.run_sync), waiting on the promise's async event. This needs a runtime with JSPI and code that is run viapyodide.runPythonAsync(); otherwise it raises and points to the async API.request_adapter/request_devicedo not block.uncapturederrorevent. They are then logged from the event loop, instead of being raised by a later, unrelated call.wgpuBufferReadMappedRange/WriteMappedRange.<canvas>element (via a CSS selector), and the browser presents it. This is verified in headless Chrome: the triangle example renders to the canvas, and a screenshot is uploaded by CI.Adapter enumeration and naming
pygfx selects an adapter by name:
PYGFX_WGPU_ADAPTER_NAME=llvmpipeis matched againstadapter.summaryfromenumerate_adapters_sync(). Its screenshot tests are enabled only ifadapter.info["vendor"] == "llvmpipe". With the Dawn backend neither worked:webgpu.hcannot enumerate adapters, andrequestAdapterreturns Dawn's "best" adapter.forceFallbackAdapteronly matches SwiftShader, so lavapipe was never listed when a GPU is present.dawn::native::Instance::EnumerateAdapters) in a small native-only helper,backends/dawn/dawn_native_extras.cpp. It lists for example the Intel GPU and llvmpipe (Vulkan)."mesa"for lavapipe,"intel") and, for Vulkan,"<driver name>: <driver info>"as the description."llvmpipe") and the driver info as the description.info["vendor_name"]. This is a deliberate deviation from the spec meaning ofvendor, to make the backends interchangeable. Reviewers may prefer otherwise.webgpu.hvalues. pygfx usescull_mode=0.Result, run locally with pygfx
main, lavapipe,PYGFX_WGPU_ADAPTER_NAME=llvmpipeand a GPU present:examples/tests: Dawn and wgpu-native give identical results, 155 passed and 27 failed. 40 of 53 screenshot comparisons pass on both, and the same 13 fail on both. The failures are missing optional dependencies in my environment, plus screenshot differences that occur with wgpu-native too. I assume those are due to my Mesa/LLVM version, but I did not verify that.tests/also give identical results for both backends (354 passed, 15 failed).Evidence
1. Raw binding micro-benchmark (
tools/binding_benchmark/)wgpuBufferGetSize, plus an "encode+submit" sequence of 11 calls.obj.get_size(). For cffi that means a Python wrapper class, which is how wgpu-py works today. For the compiled variants it is a native method.obj.get_size()tools/binding_benchmark/bench_cls.pycompares three ways to do that. The layout used here, a Python API class combined with an extension base that holds the pointer, costs 13.6 ns, or 14.1 ns with abi3. A plain compiled class that stores the pointer as an int costs 18.9 ns, or 24.4 ns with abi3._api.pylogic be ported almost line by line. Its Limited API build is just as fast.2. Through the public wgpu-py API, natively (
python tools/bench_backends.py --all)DispatchWorkgroupscall costs about 180 ns on its own.render_pass.draw(3)render_pass.set_bind_group(0, bg)render_pass.set_viewport(...)compute_pass.dispatch_workgroups(1)queue.write_buffer(64 bytes)device.create_bind_group(2 entries)3. Through the public wgpu-py API, in Pyodide (
tools/bench_backends.pyrun withtools/dawn_pyodide)webgpunpm package 0.6.1 (Dawn) on lavapipe. Chrome is headless with SwiftShader.render_pass.draw(3)render_pass.set_bind_group(0, bg)render_pass.set_viewport(...)compute_pass.dispatch_workgroups(1)queue.write_buffer(64 bytes)device.create_bind_group(2 entries)draw806 ns in Node, 154 ns in Chrome, and 128 ns natively on lavapipe.drawwas 914 ns with wgpu-native versus 129 ns with Dawn, andcreate_bind_group33.6 µs versus 13.4 µs.Trade-offs
_api.pyalmost unchanged. It is about 4× slower per call than Cython natively, and about 2–9× slower in Pyodide._api.pyrather than shared code. Struct field use is checked by the Cython and C compilers instead of by codegen annotations.Version constraints
dawn 20260929.200159, and Emdawnwebgpu (emdawnwebgpu_pkg-v20260929.200159.zip, pinned with sha256 intools/build_dawn.py). Updating Dawn means updating all three together and rerunning the codegen.--exports pyinit. The wheel is for that Pyodide ABI (cp314-pyemscripten_2026_0_wasm32).runPythonAsync. Tested in Node.js 26 and in Chrome as installed onubuntu-latest.How Dawn is obtained
mark.harfouchechannel:-c mark.harfouche -c conda-forge dawn. The package isdawn 20260929.200159, linux-64 only for now.libwebgpu_dawn,webgpu.h,webgpu_cpp.h, Dawn's C++ headers and a CMake config.Status
Tests (CI, lavapipe / SwiftShader):
pytest testswebgpunpm on lavapipe)PYODIDE_TESTSin the workflow)test_dawn_backend.py: encoder errors raised atfinish(),set_bind_groupwith dynamic offsets, and a render bundle with adepth_stencil_format.[0]was dropped.depth_stencil_formatbug; the test now also executes the bundle in a pass with a depth attachment.GPUCompilationMessage/GPUCompilationInfoinwgpu/_classes.pynow store their values.backends/__init__.py,conftest.py,tools/bench_backends.py) are identical in both PRs.auto.pydiffers only in how each PR detects its compiled part in Pyodide. Here it checks for the extension file, because importing the package would register the backend.tools/dawn_pyodide: compute async/sync, and a triangle rendered offscreen and (in Chrome) to a<canvas>.immediatesand pipeline statisticsreprwgpu.backends.wgpu_nativedirectly are not collected when Dawn is the active backend.@interpolate(flat)on integer varyings). This PR does not duplicate that change.Works (verified):
write_buffer,read_buffer, map read/write (including overlapping and unaligned access in the browser).device.lost..then();sync_wait()in Pyodide via JSPI.Not implemented or not tested yet:
test_api.py,test_async.pyandtest_canvas.pyare not run in Pyodide: they need subprocesses, trio, or a rendercanvas install in the runnerCI
.github/workflows/dawn.ymlis kept separate from the main CI:mark.harfouchethenconda-forge, onubuntu-latestwith lavapipe, for Python 3.12 and 3.14.webgpunpm, lavapipe).pyodide-demoartifact (canvas screenshot, plus a static site with the wheel andrunner.html).dawn, the demo is deployed to https://www.markharfouche.com/wgpu-py/./cffi/.dawn-pyodide. Until that job is removed, the last deployment wins.CI runs:
Other changes:
backends/auto.pyrespects an already-registered backend and theWGPUPY_BACKENDvariable. In Pyodide, it prefers the Dawn backend if its extension is installed.tests/renderutils.pyno longer importswgpu_native.tools/binding_benchmark/exists to support this comparison and can be dropped before merging.https://claude.ai/code/session_01QMLpZTQYYCu2K7EaWNnkEG