Skip to content

✨ Rewrite self-referencing urls outside HTML - #17

Merged
taras merged 2 commits into
mainfrom
rewrite-text-urls
Sep 26, 2026
Merged

taras merged 2 commits into
mainfrom
rewrite-text-urls

Conversation

@taras

@taras taras commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

Motivation

Closes #13. Now that #16 has merged this is a plain PR against main rather than a stacked one.

--base only reached html documents. Everything else was streamed to disk byte for byte, so absolute urls in the text a site serves kept the crawl origin:

out/llms.txt     [API]: http://localhost:8321/api.md
out/AGENTS.md    see http://localhost:8321/api.md
out/index.html   <a href="http://localhost:8321/AGENTS.md">contract</a>

The Effection site serves its agent-facing documents as text — llms.txt, /AGENTS.md, a .md twin of every guide and readme, feed.xml — so today the application is told its own public url through a build-time environment variable, a second source of truth next to --base.

Approach

Two decisions the issue left open, both settled with @taras:

On by default, not opt-in. The point is that --base is the only place a deployment says where it lives; a flag is one more thing to forget in a workflow.

All textual types. text/*, application/json, application/xml, application/javascript, and any +json or +xml suffix — so application/rss+xml, atom+xml and ld+json are covered. That includes css and js, where an absolute url is just as wrong. Anything else still streams untouched, so no binary body is decoded and re-encoded (there is a test with a png for that).

There is no document to walk in a text body, so this is a substitution of the crawl origin rather than a rewrite of known url-bearing attributes. It maps the origin through rebase from #16, so it honors the base's path for free — one expression of the mapping, not two.

<a href>. The html pass walked link[href], [src] and [content], which left anchors behind. They are now rewritten — through the same rebase as every other attribute — but deliberately not followed: the sitemap says what the site is made of, and downloading every link would crawl past it. Tested both ways.

Streaming, per @cowboyd's review

The first pass read each body into a string to scan it, which for text undid the one good property the byte-for-byte branch had: memory that does not depend on the size of the document. At the default concurrency of 75 that is 75 whole documents in flight, and css and js bundles are not small.

RebaseTextStream now does the substitution as the body arrives, in the shape of TextLineStream from effectionx's jsonl-store. The crawl origin is a fixed string, so the transform holds back the longest possible partial match at the end of each chunk and lets flush emit whatever is left:

export class RebaseTextStream extends TransformStream<string, string> {
  #carry = "";

  constructor(host: URL, base: URL) {
    let needle = host.origin;
    let prefix = `${rebase(new URL(needle), base)}`.replace(/\/$/, "");
    let keep = needle.length - 1;
    super({
      transform: (chunk, controller) => { /* enqueue all but the held-back tail */ },
      flush: (controller) => { /* emit the tail */ },
    });
  }
}

@effectionx/stream-helpers has lines(), which solves the same shape in Effection space, but it hands the leftover back as the stream's close value (Remainder<TReturn>) — right for a line fragment that is not a line, wrong here where the held-back tail is ordinary output bound for disk. Staying in web streams also means no Effection↔web conversion: the whole thing is one pipeThrough chain into Deno.writeFile, and the transform is the only new code.

The text branch now writes the same way the byte-for-byte branch does, and inherits its behavior on failure: a body that dies midway can leave a truncated file, where slurping left none. Taking the consequence rather than adding a temp-file-and-rename dance, which would belong on both branches or neither.

Still open: the escape hatch

Not in this PR, per your "doesn't have to exist with this PR" — but the shape I'd propose:

  • --no-rewrite-text, a coarse kill switch back to byte-for-byte. Cheap, and it makes the new default reversible from day one, which is the substance of the concern.
  • --no-rewrite <glob>, repeatable, matched on the source pathname. That is "opt out of certain files", precisely. Worth its own issue.
  • Narrowing by content type is probably unnecessary if globs exist — paths are more intuitive than media types.

On externally escaping urls in arbitrary text I want to check I read you right. If you mean urls pointing outside the site, those are already untouched: we match the crawl origin exactly and nothing else. If you mean a document that legitimately mentions the crawl origin and must not be rewritten, then I think in-band escaping is not worth building — any marker is one we invent, the document's generator has to cooperate, and we would have to strip it again — and path-level opt-out is the mechanism. Which did you mean?

Verification

The repro from #12, with --base https://frontside.com/effection:

out/index.html   <link rel="alternate" href="https://frontside.com/effection/llms.txt">
                 <a href="https://frontside.com/effection/AGENTS.md">contract</a>
out/llms.txt     [API]: https://frontside.com/effection/api.md
out/AGENTS.md    see https://frontside.com/effection/api.md

On a 12mb markdown document with urls scattered across real 64k chunk boundaries, the streaming output is byte-identical to the slurping version — same sha256, same 13,880,405 bytes — at half the peak rss, 190mb down to 89mb. The slurp figure scales with document size times concurrency; streaming is flat.

15 new tests, 24 green in all:

  • test/rebase.test.ts (new): 4 over rebase — protocol/host/port, base path prefix, a trailing slash meaning no path, query and fragment preserved. 6 over RebaseTextStream, including matches split across chunk boundaries at sizes 1, 2, 3, 7 and either side of the needle length, a body that ends mid-match, and an empty body. The boundary carry is the part that can break silently, so it is tested directly rather than only through whole files.
  • test/staticalize.test.ts: anchors rewritten, anchor targets not downloaded, text bodies rewritten across plain/markdown/rss/json/css, text bodies carrying the base path, and a binary body left byte-identical.

@taras
taras changed the base branch from main to base-path September 26, 2026 12:54
@taras
taras added this pull request to stack #18 September 26, 2026 13:06
@taras
taras requested a review from cowboyd September 26, 2026 13:39

@cowboyd cowboyd left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we're going to make a blunt force mechanism like this the default, then we need to make sure we know how to have an escape hatch. Not saying it has to exist with this PR, but we should have a good idea of what the escape hatch would look like

  • how do we handle externally escaping urls in arbitrary text?
  • how do we opt out of certain files

This slurps every text document into memory which for files that have no urls, could be very wasteful and potentially crash-inducing. Can the rebase be made streaming? I reckon with a plain generator function, it can be made to stream the replacement rather than slurp n' scan.

Base automatically changed from base-path to main September 26, 2026 14:37
`--base` only reached html documents: everything else was streamed to disk
byte for byte, so absolute urls in the text a site serves — `llms.txt`, the
markdown twins of its pages, `feed.xml` — kept the crawl origin. A
deployment had to tell the application its own public url a second way.

Textual bodies (`text/*`, json, xml, javascript and any `+json`/`+xml`
type) now have the crawl origin substituted for the base. There is no
document to walk in text, so `rebaseText` is a substitution rather than a
rewrite of known url-bearing attributes, but it maps the origin through the
same `rebase` the html pass uses and so honors the base's path. Bodies that
are not text are still streamed untouched, so nothing binary gets decoded
and re-encoded.

The html pass walked `link[href]`, `[src]` and `[content]`, which left
`<a href>` behind; anchors are now rewritten too. They are not followed:
the sitemap says what the site is made of, and downloading every link
would crawl past it.

Closes #13
Reading each text body into a string to scan it undid, for text, the one
good property the byte-for-byte branch had: memory that does not depend on
the size of the document. With the default concurrency of 75 that is 75
whole documents in flight, and css and js bundles are not small.

`RebaseTextStream` is a TransformStream that does the substitution as the
body arrives, after the shape of `TextLineStream` in effectionx's
jsonl-store: the crawl origin is a fixed string, so it holds back the
longest partial match at the end of each chunk and lets `flush` emit
whatever is left. Nothing else is buffered.

The text branch now writes the same way the byte-for-byte branch does, and
inherits the same behavior on failure: a body that dies midway can leave a
truncated file where slurping left none.

Verified on a 12mb document: byte-identical output to the slurping version
(same sha256) at half the peak rss, 190mb down to 89mb.
@taras

taras commented Sep 26, 2026

Copy link
Copy Markdown
Member Author

Can the rebase be made streaming?

Yes. Done. Good catch.

If we're going to make a blunt force mechanism like this the default, then we need to make sure we know how to have an escape hatch.

Something like an ignore pattern? or an ignore file?

@taras
taras merged commit 8ea23d6 into main Sep 26, 2026
1 check passed
@taras
taras requested a review from cowboyd September 26, 2026 15:50
@cowboyd

cowboyd commented Sep 26, 2026

Copy link
Copy Markdown
Member

Yes, I think we need a staticalize configuration soon.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Rewrite self-referencing urls outside HTML

2 participants