perf(viewer): share one mtime-keyed cache for the trends CSV and dir listing - #555
Open
omkargaikwad23 wants to merge 3 commits into
Open
perf(viewer): share one mtime-keyed cache for the trends CSV and dir listing#555omkargaikwad23 wants to merge 3 commits into
omkargaikwad23 wants to merge 3 commits into
Conversation
| return None | ||
|
|
||
| logging.info(f"Loaded {len(df)} rows from trends cache.") | ||
| _TRENDS_DF_CACHE = (mtime, df) |
| d for d in os.listdir(results_dir) | ||
| if os.path.isdir(os.path.join(results_dir, d)) | ||
| ] | ||
| _RUN_DIRS_CACHE = (mtime, dirs) |
omkargaikwad23
marked this pull request as ready for review
August 9, 2026 18:43
omkargaikwad23
requested review from
IsmailMehdi,
helloeve and
prernakakkar-google
as code owners
August 9, 2026 18:43
omkargaikwad23
force-pushed
the
perf/shared-trends-caches
branch
from
August 20, 2026 08:12
fef7b53 to
c2dc17f
Compare
State is serialized to the browser on every interaction, so keeping the whole parsed trends cache in `eval_summaries` made every click upload megabytes. Clicks that touched no data at all — switching to the Dataset Quality tab — could stall long enough to look like a hang. Parse into a module-level cache keyed on the trends cache mtime instead. Click payload drops to ~1.2 KB and repeat parses go from ~122ms to ~0.07ms. The cached list is shared across requests in a worker, so the sort now builds a new list rather than sorting in place.
generate_d3_chart serialized every column of the trends cache into each sandboxed iframe, including ai_summary — a ~1.7 KB LLM paragraph per run that no chart reads. Four charts on 300 runs shipped 2.56 MB of HTML. Project to the columns the chart actually uses: 131 KB, ~19x smaller. job_id is kept because chart.js reads it for the hover tooltip.
…listing trends_cache.csv was reparsed on every render by the Status tab, the Charts tab and the run detail view, and the results directory was relisted three times per page load. On the GCS FUSE mount each listing costs one stat per run directory, so both scaled with the number of runs and ran on interactions that needed neither. Add load_trends_df and list_run_dirs, keyed on mtime, and route every caller through them. load_summaries now builds on load_trends_df so the CSV is parsed once rather than twice. The frame is now shared, so the Charts filter chain uses assign() rather than an in-place column write, which would otherwise have welded product_dataset onto the cached object.
omkargaikwad23
force-pushed
the
perf/shared-trends-caches
branch
from
August 20, 2026 08:23
c2dc17f to
6e43752
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
trends_cache.csv was reparsed on every render by the Status tab, the
Charts tab and the run detail view, and the results directory was
relisted three times per page load. On the GCS FUSE mount each listing
costs one stat per run directory, so both scaled with the number of runs
and ran on interactions that needed neither.
Add load_trends_df and list_run_dirs, keyed on mtime, and route every
caller through them. load_summaries now builds on load_trends_df so the
CSV is parsed once rather than twice.
The frame is now shared, so the Charts filter chain uses assign() rather
than an in-place column write, which would otherwise have welded
product_dataset onto the cached object.
Stack created with GitHub Stacks CLI • Give Feedback 💬