Skip to content

Add parser cache - #950

Open
hugo-vrijswijk wants to merge 1 commit into
mainfrom
feature/parser-and-validation-cache
Open

hugo-vrijswijk wants to merge 1 commit into
mainfrom
feature/parser-and-validation-cache

Conversation

@hugo-vrijswijk

@hugo-vrijswijk hugo-vrijswijk commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Adds a parse and compileParsed to QueryCompiler.

Add CachingQueryCompiler and QueryCache, which cache the parse result by document text, including parse failures. Default: 1024 documents, LRU eviction. Both are configurable, and callers can supply their own QueryCache.

ParsedDocument is a plain data representation of the parse result, with Codec instances to store it in an external store.

Add docs for setup and configuration.

@hugo-vrijswijk
hugo-vrijswijk force-pushed the feature/parser-and-validation-cache branch from a97019b to e1425f7 Compare September 15, 2026 15:12

@milessabin milessabin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm happy with the general idea, but I think the implementation here is too specific. Maybe move the concrete implementation out of core? It should be possible to back the cache with caffeine on the JVM, for instance, or Redis.

Also, using the raw document text as a cache key is problematic IME. Normalizing it first (ie. running it through the minimzer) and hashing is a better bet.

But I think this all belongs in separate, or even external modules ... there are many different choices you might make.

@hugo-vrijswijk

Copy link
Copy Markdown
Contributor Author

Good points, I initially leaned towards a serializable format too. But because of all the references to the schema and validations. The validation results could be dropped, but that is quite a large part of what is cached in this PR. Actual pluggable remote caches would be a big win.

For normalization, I agree this would be great. But the (current) QueryMinimizer first parses the document to its AST, and then minimizes from the AST. Which kind of defeats the point of the cache if every document still has to be parsed first. I can't really see a why to normalize without parsing. Dropping all whitespace would cause conflicts in different documents that get the same normalized output. For example, query { a b } would get the same cache hit as query { ab } while being very different documents.

I think different cache hits on different whitespace is fine. The cost is just a single reparse per document, and most applications will be built to send the same documents in most cases.

What we could do instead is only cache the parsed AST result as long as it is a Success/Warning/Failure (InternalError would need to serialize Throwable), on the (hash of) the document string, and don't cache any validation. If we provide a io.circe.Codec[Result[Ast.Document]] that would make remote caches easy to plug in to.

For a separate module, do you mean the in-memory implementation to be separate? Or the caching interface? Something like grackle-cache?

@hugo-vrijswijk
hugo-vrijswijk force-pushed the feature/parser-and-validation-cache branch 2 times, most recently from 0e68e24 to 51f1b05 Compare October 7, 2026 14:53
@hugo-vrijswijk hugo-vrijswijk changed the title Add parser and validation cache Add parser cache Oct 7, 2026
@hugo-vrijswijk
hugo-vrijswijk force-pushed the feature/parser-and-validation-cache branch from 51f1b05 to 50a5f51 Compare October 7, 2026 14:58
@hugo-vrijswijk

Copy link
Copy Markdown
Contributor Author

@milessabin I've simplified this PR quite a bit to add a cache that only caches parsing. Either for successes or failures. I've also done some performance checks with it, and though it speeds up parsing a bit with in-memory (the validation from before is pretty fast), remote caches are almost never worth it. Regardless, it is probably best to keep support for it open.

Caching is still in the core module. It could go in a separate module, but it is pretty small altogether so I'm not sure it is worth it.

Some rough benchmarks below. "Remote" is a valkey server running on the same machine. Times are in µs per request:

Document No cache Local hit Local miss Remote hit Remote miss
tiny 9.7 5.0 (1.9x) 10.5 321 (33x slower) 485
smallInline 55.6 32.8 (1.7x) 56.4 327 (5.9x slower) 560
smallVars 62.7 35.4 (1.8x) 63.1 305 (4.9x slower) 575
mediumFragments 1,010 776 (1.3x) 900 1,590 (1.6x slower) 1,935
large 20,383 20,210 (1.0x) 20,301 23,758 23,500
introspection 324 232 (1.4x) 327 694 (2.1x slower) 1,244
malformed 23.7 0.35 (67x) 24.7 228 (9.6x slower) 500

Adds a `parse` and `compileParsed` to `QueryCompiler`.

Add `CachingQueryCompiler` and `QueryCache`, which cache the `parse` result by document text, including parse failures. Default: 1024 documents, LRU eviction. Both are configurable, and callers can supply their own `QueryCache`.

`ParsedDocument` is a plain data representation of the parse result, with `Codec` instances to store it in an external store.

Add docs for setup and configuration.
@hugo-vrijswijk
hugo-vrijswijk force-pushed the feature/parser-and-validation-cache branch from 50a5f51 to c196504 Compare October 7, 2026 15:43

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants