✨ Rewrite self-referencing urls outside HTML - #17
Conversation
7e63190 to
af3bb7a
Compare
af3bb7a to
4d4eb77
Compare
There was a problem hiding this comment.
If we're going to make a blunt force mechanism like this the default, then we need to make sure we know how to have an escape hatch. Not saying it has to exist with this PR, but we should have a good idea of what the escape hatch would look like
- how do we handle externally escaping urls in arbitrary text?
- how do we opt out of certain files
This slurps every text document into memory which for files that have no urls, could be very wasteful and potentially crash-inducing. Can the rebase be made streaming? I reckon with a plain generator function, it can be made to stream the replacement rather than slurp n' scan.
`--base` only reached html documents: everything else was streamed to disk byte for byte, so absolute urls in the text a site serves — `llms.txt`, the markdown twins of its pages, `feed.xml` — kept the crawl origin. A deployment had to tell the application its own public url a second way. Textual bodies (`text/*`, json, xml, javascript and any `+json`/`+xml` type) now have the crawl origin substituted for the base. There is no document to walk in text, so `rebaseText` is a substitution rather than a rewrite of known url-bearing attributes, but it maps the origin through the same `rebase` the html pass uses and so honors the base's path. Bodies that are not text are still streamed untouched, so nothing binary gets decoded and re-encoded. The html pass walked `link[href]`, `[src]` and `[content]`, which left `<a href>` behind; anchors are now rewritten too. They are not followed: the sitemap says what the site is made of, and downloading every link would crawl past it. Closes #13
4d4eb77 to
be302e9
Compare
Reading each text body into a string to scan it undid, for text, the one good property the byte-for-byte branch had: memory that does not depend on the size of the document. With the default concurrency of 75 that is 75 whole documents in flight, and css and js bundles are not small. `RebaseTextStream` is a TransformStream that does the substitution as the body arrives, after the shape of `TextLineStream` in effectionx's jsonl-store: the crawl origin is a fixed string, so it holds back the longest partial match at the end of each chunk and lets `flush` emit whatever is left. Nothing else is buffered. The text branch now writes the same way the byte-for-byte branch does, and inherits the same behavior on failure: a body that dies midway can leave a truncated file where slurping left none. Verified on a 12mb document: byte-identical output to the slurping version (same sha256) at half the peak rss, 190mb down to 89mb.
4d33cdc to
27afcfa
Compare
Yes. Done. Good catch.
Something like an ignore pattern? or an ignore file? |
|
Yes, I think we need a staticalize configuration soon. |
Motivation
Closes #13. Now that #16 has merged this is a plain PR against
mainrather than a stacked one.--baseonly reached html documents. Everything else was streamed to disk byte for byte, so absolute urls in the text a site serves kept the crawl origin:The Effection site serves its agent-facing documents as text —
llms.txt,/AGENTS.md, a.mdtwin of every guide and readme,feed.xml— so today the application is told its own public url through a build-time environment variable, a second source of truth next to--base.Approach
Two decisions the issue left open, both settled with @taras:
On by default, not opt-in. The point is that
--baseis the only place a deployment says where it lives; a flag is one more thing to forget in a workflow.All textual types.
text/*,application/json,application/xml,application/javascript, and any+jsonor+xmlsuffix — soapplication/rss+xml,atom+xmlandld+jsonare covered. That includes css and js, where an absolute url is just as wrong. Anything else still streams untouched, so no binary body is decoded and re-encoded (there is a test with a png for that).There is no document to walk in a text body, so this is a substitution of the crawl origin rather than a rewrite of known url-bearing attributes. It maps the origin through
rebasefrom #16, so it honors the base's path for free — one expression of the mapping, not two.<a href>. The html pass walkedlink[href],[src]and[content], which left anchors behind. They are now rewritten — through the samerebaseas every other attribute — but deliberately not followed: the sitemap says what the site is made of, and downloading every link would crawl past it. Tested both ways.Streaming, per @cowboyd's review
The first pass read each body into a string to scan it, which for text undid the one good property the byte-for-byte branch had: memory that does not depend on the size of the document. At the default concurrency of 75 that is 75 whole documents in flight, and css and js bundles are not small.
RebaseTextStreamnow does the substitution as the body arrives, in the shape ofTextLineStreamfrom effectionx'sjsonl-store. The crawl origin is a fixed string, so the transform holds back the longest possible partial match at the end of each chunk and letsflushemit whatever is left:@effectionx/stream-helpershaslines(), which solves the same shape in Effection space, but it hands the leftover back as the stream's close value (Remainder<TReturn>) — right for a line fragment that is not a line, wrong here where the held-back tail is ordinary output bound for disk. Staying in web streams also means no Effection↔web conversion: the whole thing is onepipeThroughchain intoDeno.writeFile, and the transform is the only new code.The text branch now writes the same way the byte-for-byte branch does, and inherits its behavior on failure: a body that dies midway can leave a truncated file, where slurping left none. Taking the consequence rather than adding a temp-file-and-rename dance, which would belong on both branches or neither.
Still open: the escape hatch
Not in this PR, per your "doesn't have to exist with this PR" — but the shape I'd propose:
--no-rewrite-text, a coarse kill switch back to byte-for-byte. Cheap, and it makes the new default reversible from day one, which is the substance of the concern.--no-rewrite <glob>, repeatable, matched on the source pathname. That is "opt out of certain files", precisely. Worth its own issue.On externally escaping urls in arbitrary text I want to check I read you right. If you mean urls pointing outside the site, those are already untouched: we match the crawl origin exactly and nothing else. If you mean a document that legitimately mentions the crawl origin and must not be rewritten, then I think in-band escaping is not worth building — any marker is one we invent, the document's generator has to cooperate, and we would have to strip it again — and path-level opt-out is the mechanism. Which did you mean?
Verification
The repro from #12, with
--base https://frontside.com/effection:On a 12mb markdown document with urls scattered across real 64k chunk boundaries, the streaming output is byte-identical to the slurping version — same
sha256, same 13,880,405 bytes — at half the peak rss, 190mb down to 89mb. The slurp figure scales with document size times concurrency; streaming is flat.15 new tests, 24 green in all:
test/rebase.test.ts(new): 4 overrebase— protocol/host/port, base path prefix, a trailing slash meaning no path, query and fragment preserved. 6 overRebaseTextStream, including matches split across chunk boundaries at sizes 1, 2, 3, 7 and either side of the needle length, a body that ends mid-match, and an empty body. The boundary carry is the part that can break silently, so it is tested directly rather than only through whole files.test/staticalize.test.ts: anchors rewritten, anchor targets not downloaded, text bodies rewritten across plain/markdown/rss/json/css, text bodies carrying the base path, and a binary body left byte-identical.