Skip to content

Backend: Browser (remote) #40

Description

@almarklein

Would be interesting to have a generic browser backend. A bit like jupyter_rfb but without Jupyter. It may also make it easier to test / benchmark new techniques.

Overview in short

The simplified and naive approach:

  • Context must present as image.
  • Canvas backend sends image over network.
  • Browser presents the image.

More detailed steps and notes on performance

                                                                                     
  ┌───────────────┐                                                                  
  │Compress/reduce│      download                                                    
  │   on GPU      │─────┐                                       ┌──────────────────┐ 
  └───────────────┘     │                                       │Present via       │ 
                        │                                       │ <img> or <canvas>│ 
                        ▼                                       │                  │ 
                     ┌───────────────┐              ┌──────┐    └──────────────────┘ 
                     │Compress/encode│              │Decode│            ▲            
                     │   on CPU      │ ───────────► │      │ ───────────┘            
                     └───────────────┘              └──────┘                         

See also pygfx/wgpu-py#378

We can also use a custom present-method, so that the image can be encoded. That broadens the options quite a bot. Assuming rendering with wgpu here:

  • Encoding the image on the GPU (optional):
    • The purpose is to reduce the image size (in bytes), so the next steps are faster.
    • To make this step fast, we have to stick to encodings that parallelize well.
    • There are compressed texture formats in wgpu.
    • Can use YUV (yuv420 results in 37.5% the size compared to rgba)
    • JPEG encoding with variable quality
    • Diff images
    • mp4 encoding
  • Downloading from GPU into RAM:
    • The bottleneck is the amount of bytes.
    • For small images, as typical in a notebook, this step does not cost a lot of time.
    • For larger images, going towards full screen, the impact becomes very significant (when downloading raw rgba data)
    • A lot of time is spent on waiting for the copy-buffer to become mapped. An async approach would matter a lot.
    • Also see the little benchmark in Refactor canvas context to allow presenting as image wgpu-py#586
  • Encoding the image on the CPU:
    • Can be done by the context, or in canvas._rc_present_xx(), or a bit in both.
    • Let's assume we have a binary websocket, but if we don't than base64 encoding is part of this step too.
    • jupyter_rfb does png and jpg encoding.
    • Making the image size smaller will make the next step faster.
    • But the encoding also greatly affects the presentation in the browser.
  • Sending over the network, io:
    • The speed of this step is pretty good if you're on localhost.
    • Otherwise, the speed can vary a lot, depending on where the server and browser are.
    • Speed can vary over time too.
    • It would be nice to measure this speed and use it to control the encoding steps above.
  • Presenting in the browser:
    • This step includes any necessary decoding. Browsers have builtin support for some decoding mechanisms.
    • jupyter_rfb sets the src of an <img> object to the base64 encoded png/jpg. This is surprisingly fast.
    • Using a canvas. IIRC drawing an image to a canvas is pretty slow.
    • Using WebGL / WebGPU.

There are multiple paths to go through these steps. What's tricky is that a choice in one step affects the other steps, so
we cannot simply pick "the best" from each step and stack them together.

MP4 has nice properties in being pretty good in a situations with low/variable bandwidth, and it very naturally improves the resul over time for static images.

Something like WebCRT can possibly be used to implement some of these steps.

Some unsorted thoughts and links

  • Partial jpg encoding/decoding on the GPU: https://github.com/negge/jpeg_gpu
  • Can also run deflate on the GPU (GDeflate), so basically encode to png.
  • Simply down-sampling the image at the GPU will reduce the size by a factor 4, which affects all steps. For moving images this is barely noticeably (because by default our images are 2x larger for aa). We could even go for a factor of 9 or 16 in low-latency cases.
  • Jpeg encoding is a lot faster and smaller than png.
  • Formats like webp and avif have good compression, which help with the file size, but they are rather slow.
  • Advanced solutions like webCRT are cool, but maybe heavy on the dependencies.

Open questions

  • Can we jpeg-compress on the GPU?
  • How good are compressed texture formats?
  • How do game streaming platforms do this?

Activity

  1. almarklein commented on Feb 12, 2025

    @almarklein
    MemberAuthor

    I've been doing some research to image file formats in almarklein/ppng#2. Some conclusions relevant to this topic:

    • Compressing with jpeg is a lot faster than png, and produces smaller blobs (and easy to tune size). But we already know this.
    • With png compression, the size of the blob depends heavily on how well the image can be compressed (i.e. the content of the image).
    • The above actually holds for lossless formats (like jpg) as well, but to a lesser degree.
    • Compression with (lossy) WebP can produce smaller blobs than jpeg, but its much slower.
    • When sending png, dropping the filter and lowering the compression help speed a lot, with relatively small effect on the file size.

    One viable idea might be to zlib-compress the image on the GPU (people have done this) and then sending as png. The downside is that if the image has high entropy, it does not compress well. So if possible, I'd rather move towards jpeg/mpeg compression.

  2. almarklein commented on Feb 12, 2025

    @almarklein
    MemberAuthor

    I got a tip from @hmaarrfk to align the buffer to 4096 when downloading from the GPU, which ought to result in a faster download. I have yet to try this.

  3. almarklein commented on Mar 11, 2025

    @almarklein
    MemberAuthor

    I got a tip from @hmaarrfk to align the buffer to 4096 when downloading from the GPU, which ought to result in a faster download. I have yet to try this.

    I had a look at this, but we don't allocate the data for the buffer to copy to. The wgpu API does that:

            src_ptr = libf.wgpuBufferGetMappedRange(self._internal, offset, size)
            src_address = int(ffi.cast("intptr_t", src_ptr))

    Interestingly, the src_address always seems aligned to 4096 😉

    I also tried padding the buffer size to make it a multiple of 4096, but I did not observe a difference in performance.

  4. almarklein commented on Mar 11, 2025

    @almarklein
    MemberAuthor

    I spent some time investigating possible ways to improve the speed by which we can download data from the GPU.

    From what I read, the steps that we do are the correct ones:

    • Allocate (or re-use) a buffer large enough to contain the pixel data.
    • Use command_encoder.copy_texture_to_buffer() to copy the data over.
    • Map the buffer (async op).
    • Wait for it to be mapped ⏳
    • Read the mapped data, and copy to a new array (the mapped data is released when the buffer is unmapped).
    • Unmap the buffer.

    Re-using the target texture and the buffer to copy to, helps significantly. I also found that one data-copy can be avoided in some cases (when bytes-per-row is not a multiple of 256). And Using numpy to copy the data to avoid a for-loop and enabling allocating an empty array (instead of zeros).

    Together these bring the FPS for full screen apps from 24-ish to 30-ish (on a 4k screen).

    I also found a much bigger way to improve performance. The part where we wait for the buffer to be mapped is (apparently) a significant portion of the total spent time. As an experiment, I mapped the buffer, but then wait for it to be mapped in the next frame, and then read data from it. FPS goes up to 50+ 🚀 Sweet! However, this by itself is not a viable approach, because in cases like the notebook, where the network can be the bottleneck, being one frame behind may become noticeable. What it means is that we need some form of async / callbacks here, if we want to have this kind of fps for rendering via a bitmap. And we need some more work on this front. Nevertheless, this is promising!

  5. almarklein commented on Mar 11, 2025

    @almarklein
    MemberAuthor

    I also tried doing the copy_texture_to_buffer in the Pygfx flusher, using the command_encoder that's already doing the flushing, to avoid creating a new encoder and doing an extra queue.submit().

  6. kushalkolar commented on Jun 11, 2025

    @kushalkolar
    Contributor

    @apasarkar and I have been making some progress on JPEG encoding, I think we'll have a webgpu jpeg encoder soon!

    Some notes and things to think about:

    • By jpeg encoding on the GPU we also significantly reduce the amount of data that has to be downloaded from the GPU.
      • Caveat, the number of bytes that represent a jpeg image are not consistent between frames. AFAIK, in wgpu you can't download an arbitrary sized buffer from the GPU to the CPU?
      • Naive idea to get around this: for a given quality level, set the max number of bytes allowed based on the most complex expected frame based on testing the size for a wide variety of plots?
      • Another idea: chunk it into multiple large buffers, download the buffers and see if a stop byte is at the end of any buffer, if so don't download the later chunks.
    • On the client end, I'm using anywidget which is amazing. Started prototyping webgpu canvases with anywidget. This will give us the best flexibility for use in the browser across as many notebook platforms as possible, including marimo which has native anywidget support.
  7. almarklein commented on Jun 12, 2025

    @almarklein
    MemberAuthor

    FAIK, in wgpu you can't download an arbitrary sized buffer from the GPU to the CPU?

    Yeah, you can download any size, except for some alignment constraints (nbytes must be a multiple of 8 or 16). But you have to specify it on the CPU, which (I assume) does not know the actual size.

    You could try encoding white noise and see how large that becomes as a max. That could provide a sensible upper bound, but would be nice if images that encode well are also fast.

    Downloading in multiple steps does result in extra wait time, but this may not be so bad, since you can start sending the first chunk over the network while the GPU is downloading the next chunk. Steaming! Definitely need benchmarks to test what works well and what not.

    Questions: What is the fastest way to send binary data from the server, get that on the js end, and upload it to the GPU for jpeg decoding?

    Binary websocket. Then probably queue.write_buffer. Or device.createBuffer(..., mapped_at_creation=True).

    Mmm, if we stream the encoded data in chunks, the client does not know the size until the last chunk has arrived. In that case it would be good if there is a big-enough buffer on the GPU, and just send chunks to it with queue.write_buffer.

    On the GPU, after encoding, what does the data look like? Is it stored in a buffer or texture? Does it already naturally come in chunks of some sort? I'm half wondering if the client can start processing chunks once the first arrives 😉

  8. kushalkolar commented on Jun 12, 2025

    @kushalkolar
    Contributor

    Thanks for the thoughts! Will keep updating.

    On the GPU, after encoding, what does the data look like? Is it stored in a buffer or texture?

    JPEG encoded output is just raw bytes, so I imagine the file output would just be from a storage buffer.

    Meanwhile, almost have DCT working in wgsl 🥳 . Need to debug why the basis weights are not in the right place, keeping track of [row, col] -> [x, y] in wgsl is tedious 😅 . Thank goodness we have numpy, signal processing with a low level language is always tedious, last time I did this was in C and ARM assembly; at least wgsl is easier to debug because sometimes ARM assembly will just spew out random unicode errors 🤦

    EDIT: apprently I can't write a basic for loop, I've been spoiled by writing loops with python's enumerate, and itertools. And textureLoad() does not raise if you try to load a texture from an index that is beyond the size of that dimension 🤦

    and we have DCT in wgsl 🥳 🚀

    Image

    prototyping here: https://github.com/kushalkolar/webgpu-rfb-prototype

  9. kushalkolar commented on Jul 13, 2025

    @kushalkolar
    Contributor

    Made some more progress on this and finally started templating. Wow texture fetches are expensive, 35% faster if we just dump the DCT basis as a const.

  10. kushalkolar commented on Jul 15, 2025

    @kushalkolar
    Contributor

    ok now I'm at ~6ms for DCT of luma, I need to see if this can be optimized further by doing more invocations per workgroup and having larger workgroups than just 8x8 blocks. The way that you arrange your workgroups and invocations makes a huge difference!

    https://github.com/kushalkolar/webgpu-rfb-prototype/blob/b6bb454ceea2de0cfad238c46ed16ae6ce91f753/luma_dct_8-1-64.wgsl

    Strange thing is, it takes ~6ms on my 8 year old AMD RX 570 GPU which was midrange at the time, but it takes ~34ms on my Nvidia RTX 3080. I wonder if it prefers different layouting of workgroups vs invocations or if I should update the drivers. Really bizarre because my nvidia does a much better job at all the viz with pygfx.

    @almarklein if you see any obvious ways to improve the shader let me know! @hmaarrfk heard you're also interested in this

  11. Korijn commented on Jul 15, 2025

    @Korijn
    Contributor

    are you sure your timing setup is measuring only the GPU work and not any overhead?

  12. kushalkolar commented on Jul 15, 2025

    @kushalkolar
    Contributor

    are you sure your timing setup is measuring only the GPU work and not any overhead?

    I'm using device._poll_wait() which I saw in the pygfx compute shader object:

    https://github.com/kushalkolar/webgpu-rfb-prototype/blob/main/jpeg_encode.py#L79-L104

  13. kushalkolar commented on Jul 15, 2025

    @kushalkolar
    Contributor

    Also I should start combining this with pygfx. Any pointer on how I can feed the final rendered texture from pygfx into this?

  14. hmaarrfk commented on Jul 15, 2025

    @hmaarrfk
    Contributor

    i guess the only thing i have to say, is that i'm not sure the "JPEG" path is the best one. I think that JPEG has its place, but "video streaming" is essentially "improved JPEG".

    JPEG is a "step forward" i believe, but to get performance, over "remote" networks, you may have to go to "video" streaming algorithms.

    I've found it difficult to have "video streaming" through PyAV.

    Generally, the issue is that you want to "shape the stream" correctly, where the metadata comes in first, then the "video stream" can be "interrupted" randomly. And I never got that right.

  15. kushalkolar commented on Jul 15, 2025

    @kushalkolar
    Contributor

    i guess the only thing i have to say, is that i'm not sure the "JPEG" path is the best one. I think that JPEG has its place, but "video streaming" is essentially "improved JPEG".

    JPEG is a "step forward" i believe, but to get performance, over "remote" networks, you may have to go to "video" streaming algorithms.

    I've found it difficult to have "video streaming" through PyAV.

    Generally, the issue is that you want to "shape the stream" correctly, where the metadata comes in first, then the "video stream" can be "interrupted" randomly. And I never got that right.

    Yea we eventually want to do a video codec-like method after we have implemented jpeg.

    The best thing I've found for gpu accelerated video decoding with random frame index access is decord, no longer maintained but it was really good! https://github.com/dmlc/decord

  16. hmaarrfk commented on Jul 15, 2025

    @hmaarrfk
    Contributor

    Interesting about decore... neat find. I'm going to have to see what kind of tricks they use.

    I generally don't think the "backend browser" will need random seeks since the main audiance is a "human" if you "wait 2-3 seconds" for the stream to catch a "I" frame.

  17. almarklein commented on Jul 16, 2025

    @almarklein
    MemberAuthor

    i guess the only thing i have to say, is that i'm not sure the "JPEG" path is the best one.

    I agree that a video stream would produce better results, since it can compress temporally as well. But it's also quite a bit more complex, so IMO it makes sense get things moving with a jpeg-like solution, and improve from there.

    Strange thing is, it takes ~6ms on my 8 year old AMD RX 570 GPU which was midrange at the time, but it takes ~34ms on my Nvidia RTX 3080.

    I found benchmarking to be quite tricky. I found that if you submit the workload n times (say 100 or 1000), and do the _poll_wait() after that, you get reasonable results that actually go up if the shader does more work, and down if it does less. See this example:

    https://github.com/almarklein/ppaa-experiments/blob/bc593bb5625f2f457fbf68c997db8d6185b3609e/scripts/renderer_wgsl.py#L142-L157

    Any pointer on how I can feed the final rendered texture from pygfx into this?

    In current main that'd be:

    texture = renderer._blender.get_texture("color")
    # or to make sure you include all effects
    texture = renderer._blender.get_texture(renderer._name_of_texture_with_effects or "color")
  18. almarklein commented on Jul 16, 2025

    @almarklein
    MemberAuthor

    I had a look at the code, and it looks good to me:

    • constants for arrays so that theses don't take up registers 👌
    • unrolling the loop that does texture fetches 👌

    The loop further down probably can be a regular loop; the unrolling is mostly for allowing the compiler to combine texture fetches.

    Can we use texture samplers in compute shaders? If we can, you could try textureSampleLevel with the integer offset argument. That gives the compiler more static hints on what you're sampling.

  19. kushalkolar commented on Jul 16, 2025

    @kushalkolar
    Contributor

    The loop further down probably can be a regular loop; the unrolling is mostly for allowing the compiler to combine texture fetches.

    Ah ok, sometimes I don't see a performance difference with unrolling them. Been playing around with all combinations to get a feel of how wgsl performs and create a mental model.

    Can we use texture samplers in compute shaders?

    Yes we can use texture samplers in compute shaders.

    If we can, you could try textureSampleLevel with the integer offset argument. That gives the compiler more static hints on what you're sampling.

    What kinds of hints? If we are just grabbing the value at a given coord I assumed there would be no difference but I have no idea.

    I've been trying to reduce the number of operations, first I tried combining the RGB -> luma weights along with the DCT coefficients and quantization factor into a single "quantized dct with rgba weights" constant of shape [64, 8, 8, 4]. This actually increased the compute time from 6.2 ms to 14 ms! I guess having huge constants (4096 vec4f items in this case), is not ideal.

    Looking into how fast DCT algorithms are implemented and came across this doc from nvidia, see page 10: https://developer.download.nvidia.com/assets/cuda/files/dct8x8.pdf

    We can exploit the symmetry in the DCT coefficients to reduce the number of multiplications. A lot of papers seem to site this paper: Practical_fast_1_D_DCT_algorithms_with_1.pdf

    Hopefully can reduce the compute time by half.

  20. kushalkolar commented on Jul 16, 2025

    @kushalkolar
    Contributor

    So the factorization seems to give results that look right, diff of Frobenius norms is ~0.05 for 512x512 image which is acceptable I think for this purpose. Might be less numerically stable or their weighting term does something different. But the differences in the result are small and only in the less important basis anyways.

    Performance is amazing though, reducing all those multiplications!. And the nvidia results finally make sense.

    AMD RX 570, 4k
    	mean: 0.919 ms
    	median: 0.916 ms
    
    nvidia RTX 3080, 4k
    	mean: 0.544 ms
    	median: 0.535 ms
    
  21. almarklein commented on Jun 24, 2026

    @almarklein
    MemberAuthor

    Closed by #202, let's track improvements in #215

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions