Repository navigation
Backend: Browser (remote) #40
Description
Activity
I've been doing some research to image file formats in almarklein/ppng#2. Some conclusions relevant to this topic:
- Compressing with jpeg is a lot faster than png, and produces smaller blobs (and easy to tune size). But we already know this.
- With png compression, the size of the blob depends heavily on how well the image can be compressed (i.e. the content of the image).
- The above actually holds for lossless formats (like jpg) as well, but to a lesser degree.
- Compression with (lossy) WebP can produce smaller blobs than jpeg, but its much slower.
- When sending png, dropping the filter and lowering the compression help speed a lot, with relatively small effect on the file size.
One viable idea might be to zlib-compress the image on the GPU (people have done this) and then sending as png. The downside is that if the image has high entropy, it does not compress well. So if possible, I'd rather move towards jpeg/mpeg compression.
I got a tip from @hmaarrfk to align the buffer to 4096 when downloading from the GPU, which ought to result in a faster download. I have yet to try this.
I got a tip from @hmaarrfk to align the buffer to 4096 when downloading from the GPU, which ought to result in a faster download. I have yet to try this.
I had a look at this, but we don't allocate the data for the buffer to copy to. The wgpu API does that:
src_ptr = libf.wgpuBufferGetMappedRange(self._internal, offset, size) src_address = int(ffi.cast("intptr_t", src_ptr))
Interestingly, the
src_addressalways seems aligned to 4096 😉I also tried padding the buffer size to make it a multiple of 4096, but I did not observe a difference in performance.
I spent some time investigating possible ways to improve the speed by which we can download data from the GPU.
From what I read, the steps that we do are the correct ones:
- Allocate (or re-use) a buffer large enough to contain the pixel data.
- Use
command_encoder.copy_texture_to_buffer()to copy the data over. - Map the buffer (async op).
- Wait for it to be mapped ⏳
- Read the mapped data, and copy to a new array (the mapped data is released when the buffer is unmapped).
- Unmap the buffer.
Re-using the target texture and the buffer to copy to, helps significantly. I also found that one data-copy can be avoided in some cases (when bytes-per-row is not a multiple of 256). And Using numpy to copy the data to avoid a for-loop and enabling allocating an empty array (instead of zeros).
Together these bring the FPS for full screen apps from 24-ish to 30-ish (on a 4k screen).
I also found a much bigger way to improve performance. The part where we wait for the buffer to be mapped is (apparently) a significant portion of the total spent time. As an experiment, I mapped the buffer, but then wait for it to be mapped in the next frame, and then read data from it. FPS goes up to 50+ 🚀 Sweet! However, this by itself is not a viable approach, because in cases like the notebook, where the network can be the bottleneck, being one frame behind may become noticeable. What it means is that we need some form of async / callbacks here, if we want to have this kind of fps for rendering via a bitmap. And we need some more work on this front. Nevertheless, this is promising!
Reacted by Mark Harfouche, Korijn van Golen and Clay DugoI also tried doing the
copy_texture_to_bufferin the Pygfx flusher, using the command_encoder that's already doing the flushing, to avoid creating a new encoder and doing an extraqueue.submit().@apasarkar and I have been making some progress on JPEG encoding, I think we'll have a webgpu jpeg encoder soon!
Some notes and things to think about:
- By jpeg encoding on the GPU we also significantly reduce the amount of data that has to be downloaded from the GPU.
- Caveat, the number of bytes that represent a jpeg image are not consistent between frames. AFAIK, in
wgpuyou can't download an arbitrary sized buffer from the GPU to the CPU? - Naive idea to get around this: for a given quality level, set the max number of bytes allowed based on the most complex expected frame based on testing the size for a wide variety of plots?
- Another idea: chunk it into multiple large buffers, download the buffers and see if a stop byte is at the end of any buffer, if so don't download the later chunks.
- Caveat, the number of bytes that represent a jpeg image are not consistent between frames. AFAIK, in
- On the client end, I'm using
anywidgetwhich is amazing. Started prototyping webgpu canvases with anywidget. This will give us the best flexibility for use in the browser across as many notebook platforms as possible, including marimo which has native anywidget support.- Questions: What is the fastest way to send binary data from the server, get that on the js end, and upload it to the GPU for jpeg decoding?
- Some general ideas here maybe: https://toji.dev/webgpu-best-practices/buffer-uploads.html
- there are some jpeg decoders people have written on github if we need ideas: https://github.com/kushalkolar/Compeg, https://github.com/tamani-coding/webgpu-computing-jpeg
- By jpeg encoding on the GPU we also significantly reduce the amount of data that has to be downloaded from the GPU.
FAIK, in wgpu you can't download an arbitrary sized buffer from the GPU to the CPU?
Yeah, you can download any size, except for some alignment constraints (nbytes must be a multiple of 8 or 16). But you have to specify it on the CPU, which (I assume) does not know the actual size.
You could try encoding white noise and see how large that becomes as a max. That could provide a sensible upper bound, but would be nice if images that encode well are also fast.
Downloading in multiple steps does result in extra wait time, but this may not be so bad, since you can start sending the first chunk over the network while the GPU is downloading the next chunk. Steaming! Definitely need benchmarks to test what works well and what not.
Questions: What is the fastest way to send binary data from the server, get that on the js end, and upload it to the GPU for jpeg decoding?
Binary websocket. Then probably
queue.write_buffer. Ordevice.createBuffer(..., mapped_at_creation=True).Mmm, if we stream the encoded data in chunks, the client does not know the size until the last chunk has arrived. In that case it would be good if there is a big-enough buffer on the GPU, and just send chunks to it with
queue.write_buffer.On the GPU, after encoding, what does the data look like? Is it stored in a buffer or texture? Does it already naturally come in chunks of some sort? I'm half wondering if the client can start processing chunks once the first arrives 😉
Thanks for the thoughts! Will keep updating.
On the GPU, after encoding, what does the data look like? Is it stored in a buffer or texture?
JPEG encoded output is just raw bytes, so I imagine the file output would just be from a storage buffer.
Meanwhile, almost have DCT working in wgsl 🥳 . Need to debug why the basis weights are not in the right place, keeping track of
[row, col]->[x, y]in wgsl is tedious 😅 . Thank goodness we have numpy, signal processing with a low level language is always tedious, last time I did this was in C and ARM assembly; at least wgsl is easier to debug because sometimes ARM assembly will just spew out random unicode errors 🤦EDIT: apprently I can't write a basic for loop, I've been spoiled by writing loops with python's
enumerate, anditertools. AndtextureLoad()does not raise if you try to load a texture from an index that is beyond the size of that dimension 🤦and we have DCT in wgsl 🥳 🚀
prototyping here: https://github.com/kushalkolar/webgpu-rfb-prototype
Reacted by Sébastien BoisgéraultMade some more progress on this and finally started templating. Wow texture fetches are expensive, 35% faster if we just dump the DCT basis as a
const.ok now I'm at ~6ms for DCT of luma, I need to see if this can be optimized further by doing more invocations per workgroup and having larger workgroups than just 8x8 blocks. The way that you arrange your workgroups and invocations makes a huge difference!
Strange thing is, it takes ~6ms on my 8 year old AMD RX 570 GPU which was midrange at the time, but it takes ~34ms on my Nvidia RTX 3080. I wonder if it prefers different layouting of workgroups vs invocations or if I should update the drivers. Really bizarre because my nvidia does a much better job at all the viz with pygfx.
@almarklein if you see any obvious ways to improve the shader let me know! @hmaarrfk heard you're also interested in this
are you sure your timing setup is measuring only the GPU work and not any overhead?
are you sure your timing setup is measuring only the GPU work and not any overhead?
I'm using device._poll_wait() which I saw in the pygfx compute shader object:
https://github.com/kushalkolar/webgpu-rfb-prototype/blob/main/jpeg_encode.py#L79-L104
Also I should start combining this with pygfx. Any pointer on how I can feed the final rendered texture from pygfx into this?
i guess the only thing i have to say, is that i'm not sure the "JPEG" path is the best one. I think that JPEG has its place, but "video streaming" is essentially "improved JPEG".
JPEG is a "step forward" i believe, but to get performance, over "remote" networks, you may have to go to "video" streaming algorithms.
I've found it difficult to have "video streaming" through PyAV.
Generally, the issue is that you want to "shape the stream" correctly, where the metadata comes in first, then the "video stream" can be "interrupted" randomly. And I never got that right.
i guess the only thing i have to say, is that i'm not sure the "JPEG" path is the best one. I think that JPEG has its place, but "video streaming" is essentially "improved JPEG".
JPEG is a "step forward" i believe, but to get performance, over "remote" networks, you may have to go to "video" streaming algorithms.
I've found it difficult to have "video streaming" through PyAV.
Generally, the issue is that you want to "shape the stream" correctly, where the metadata comes in first, then the "video stream" can be "interrupted" randomly. And I never got that right.
Yea we eventually want to do a video codec-like method after we have implemented jpeg.
The best thing I've found for gpu accelerated video decoding with random frame index access is decord, no longer maintained but it was really good! https://github.com/dmlc/decord
Interesting about decore... neat find. I'm going to have to see what kind of tricks they use.
I generally don't think the "backend browser" will need random seeks since the main audiance is a "human" if you "wait 2-3 seconds" for the stream to catch a "I" frame.
i guess the only thing i have to say, is that i'm not sure the "JPEG" path is the best one.
I agree that a video stream would produce better results, since it can compress temporally as well. But it's also quite a bit more complex, so IMO it makes sense get things moving with a jpeg-like solution, and improve from there.
Strange thing is, it takes ~6ms on my 8 year old AMD RX 570 GPU which was midrange at the time, but it takes ~34ms on my Nvidia RTX 3080.
I found benchmarking to be quite tricky. I found that if you submit the workload n times (say 100 or 1000), and do the
_poll_wait()after that, you get reasonable results that actually go up if the shader does more work, and down if it does less. See this example:Any pointer on how I can feed the final rendered texture from pygfx into this?
In current main that'd be:
texture = renderer._blender.get_texture("color") # or to make sure you include all effects texture = renderer._blender.get_texture(renderer._name_of_texture_with_effects or "color")
I had a look at the code, and it looks good to me:
- constants for arrays so that theses don't take up registers 👌
- unrolling the loop that does texture fetches 👌
The loop further down probably can be a regular loop; the unrolling is mostly for allowing the compiler to combine texture fetches.
Can we use texture samplers in compute shaders? If we can, you could try
textureSampleLevelwith the integer offset argument. That gives the compiler more static hints on what you're sampling.The loop further down probably can be a regular loop; the unrolling is mostly for allowing the compiler to combine texture fetches.
Ah ok, sometimes I don't see a performance difference with unrolling them. Been playing around with all combinations to get a feel of how wgsl performs and create a mental model.
Can we use texture samplers in compute shaders?
Yes we can use texture samplers in compute shaders.
If we can, you could try
textureSampleLevelwith the integer offset argument. That gives the compiler more static hints on what you're sampling.What kinds of hints? If we are just grabbing the value at a given coord I assumed there would be no difference but I have no idea.
I've been trying to reduce the number of operations, first I tried combining the RGB -> luma weights along with the DCT coefficients and quantization factor into a single "quantized dct with rgba weights" constant of shape
[64, 8, 8, 4]. This actually increased the compute time from 6.2 ms to 14 ms! I guess having huge constants (4096vec4fitems in this case), is not ideal.Looking into how fast DCT algorithms are implemented and came across this doc from nvidia, see page 10: https://developer.download.nvidia.com/assets/cuda/files/dct8x8.pdf
We can exploit the symmetry in the DCT coefficients to reduce the number of multiplications. A lot of papers seem to site this paper: Practical_fast_1_D_DCT_algorithms_with_1.pdf
Hopefully can reduce the compute time by half.
So the factorization seems to give results that look right, diff of Frobenius norms is ~0.05 for 512x512 image which is acceptable I think for this purpose. Might be less numerically stable or their weighting term does something different. But the differences in the result are small and only in the less important basis anyways.
Performance is amazing though, reducing all those multiplications!. And the nvidia results finally make sense.
AMD RX 570, 4k mean: 0.919 ms median: 0.916 ms nvidia RTX 3080, 4k mean: 0.544 ms median: 0.535 ms

Would be interesting to have a generic browser backend. A bit like
jupyter_rfbbut without Jupyter. It may also make it easier to test / benchmark new techniques.Overview in short
The simplified and naive approach:
More detailed steps and notes on performance
See also pygfx/wgpu-py#378
We can also use a custom present-method, so that the image can be encoded. That broadens the options quite a bot. Assuming rendering with wgpu here:
canvas._rc_present_xx(), or a bit in both.srcof an<img>object to the base64 encoded png/jpg. This is surprisingly fast.There are multiple paths to go through these steps. What's tricky is that a choice in one step affects the other steps, so
we cannot simply pick "the best" from each step and stack them together.
MP4 has nice properties in being pretty good in a situations with low/variable bandwidth, and it very naturally improves the resul over time for static images.
Something like WebCRT can possibly be used to implement some of these steps.
Some unsorted thoughts and links
Open questions