From 2d981737bb4f61d298060379253ef8deefba40e6 Mon Sep 17 00:00:00 2001 From: Will Kahn-Greene Date: Wed, 23 Sep 2026 10:12:33 -0400 Subject: [PATCH 1/6] docs(guarantees): rewrite as design principles Replaces the status-tracking shape with principles, per _plans/048. Each entry is a statement with Why, Rules out, and (only for a deliberate trade-off) Accepts. There are no statuses, enforcement notes, or history, so the file changes only when a principle changes. Gaps between a principle and the code are issues: #184, #185, #186. Ids and labels are unchanged. The no-markfluence.yaml trade-off is now an Accepts under S8 and L4. L5 is reworded to what it always meant: the page keeps its meaning, and the Markdown is a fixed point. The byte-level measurement behind that moves to docs/confluence/storage-format.md. The "how each kind is verified" table is dropped, since it describes test practice rather than a principle. Refs #191 --- docs/confluence/storage-format.md | 16 + docs/guarantees.md | 1014 +++++++++++------------------ 2 files changed, 388 insertions(+), 642 deletions(-) diff --git a/docs/confluence/storage-format.md b/docs/confluence/storage-format.md index 267a18c..c928ca2 100644 --- a/docs/confluence/storage-format.md +++ b/docs/confluence/storage-format.md @@ -34,6 +34,22 @@ would churn ids forever. So the guarantee is **semantic**, not byte-for-byte, which is also the converter's stated design target. +### A page from the editor does not survive an export and republish byte-for-byte + +**Verified 2026-09-05**, exporting a live page the editor had written and +publishing the Markdown back. The stored storage changed in two ways, neither of +which changes what renders: + +- The editor writes a list item as `
  • text

  • `; the converter writes + `
  • text
  • `. +- The editor's TOC macro carries `ac:local-id`, `ac:macro-id` and `data-layout` + attributes; the converter's canonical form has none of them. + +This is why L5 (`roundtrip-from-confluence`, [../guarantees.md](../guarantees.md)) +promises the page keeps its *meaning* rather than its bytes. The Markdown side +is stricter: after one cycle it stops changing, which +`TestRoundTripMarkdownIsAFixedPoint` checks over every `storage2md` case. + ### Confluence strips HTML comments on write **Verified 2026-09-13.** A comment does not survive the write at all — it is not diff --git a/docs/guarantees.md b/docs/guarantees.md index 6a93417..cdec3ef 100644 --- a/docs/guarantees.md +++ b/docs/guarantees.md @@ -1,126 +1,81 @@ -# What markfluence guarantees +# Design principles -These are the properties that markfluence holds itself to. You cite them. A -change that breaks one needs an argument. If a new feature cannot hold one, the -problem is probably in the design of the feature, and not in the guarantee. +These are the principles that guide design decisions in markfluence. Cite them +by id. A change that breaks one needs an argument. If a new feature cannot hold +one, the problem is probably in the design of the feature, and not in the +principle. The [confluence/](confluence/) directory records what we know about -*Confluence*. This document is its counterpart: it gives claims about -**markfluence**. +*Confluence*. This document is its counterpart: it states what **markfluence** +aims to do. -Each guarantee is about one thing on purpose, so you can argue about it -separately. +This document does not say whether the code holds each principle today. A gap +between a principle and the code is a GitHub issue that names the principle by +id. Thus this document changes when a principle changes, and at no other time. ## How to read an entry -Every guarantee has a status. Every entry in [confluence/](confluence/) gives -its provenance for the same reason: "we promise this" and "we intend this" need -different levels of trust. +Each entry has an id, a label, and a statement, and then these parts: -- **Holds**: it is true now. The entry names the thing that enforces it. -- **Partial**: it is true on some paths. The entry names the paths where it - fails. -- **Aspirational**: it is not true yet. The entry names what would make it true. -- **Vacuous**: nothing uses it yet. It is not tested, so it is not proved. +- **Why:** the reason for the principle. +- **Rules out:** the designs that the principle forbids. +- **Accepts:** a trade-off that we chose on purpose. Only some entries have + one. An Accepts records a decision. It never records a bug. ## Rules for changes to this document -**Identifiers are permanent.** Never change the number of a guarantee, and never -use a number again. Each guarantee also has a **label**, such as -`no-read-outside-root`. The label helps you to read and cite the guarantee. In -a heading, "S2 (no-read-outside-root)" gives enough information, and you do not -need a full sentence to explain it. +**Ids and labels are permanent.** Never change the id or the label of a +principle, and never use an id again. Plans, pull requests, and code comments +cite them, and a changed or reused id makes an old reference wrong with no +warning. If the id and the label ever disagree, the id is correct. -Labels are also permanent, for the same reason as the ids. Persons cite them, -and a changed label makes an old reference wrong with no warning. If the id and -the label ever disagree, the id is correct. If a guarantee stops being useful, -mark it as retired and give the reason. Do not delete it. Plans and pull -requests cite guarantees by id, and an id used again makes an old reference -wrong with no warning. - -**A change cannot lower a status without a statement.** A change from Holds to -Partial is a decision. Write it in the commit message and in this file. Do not -let it be a side effect that somebody finds later. +You can reword a statement when the principle itself changes. If a principle +stops being useful, mark it as retired and give the reason. Do not delete it. *Retired: "Nothing in Confluence is deleted." This was a fact about the code, and not a principle. It would become false on the day that the first prune -feature shipped. S4, S5, and S6 replace it. They stay true with that feature, -and they control how it works.* +feature shipped. S4, S5, and S6 replace it.* ## Safety -If a safety guarantee fails, the result is damage, and not only a wrong answer. - -| | label | guarantee | status | -|---|---|---|---| -| **S1** | `no-write-outside-root` | markfluence writes no file outside the root. | Holds | -| **S2** | `no-read-outside-root` | markfluence reads no file outside the root. | Holds | -| **S3** | `no-overwrite-without-force` | markfluence does not overwrite a file that exists, unless you give `--force`. | Holds | -| **S4** | `no-removal-as-side-effect` | markfluence removes nothing as a side effect. A removal occurs only when it is the stated purpose of a command. | Vacuous | -| **S5** | `remove-only-ours` | markfluence removes only what markfluence created. | Vacuous | -| **S6** | `removal-is-previewable` | A command that removes things says what it will remove before it removes them, and it obeys `--dry-run`. | Vacuous | -| **S7** | `no-partial-create` | If `create` cannot publish a file, it leaves no page behind. | Partial | -| **S8** | `no-overwrite-of-a-moved-page` | markfluence does not overwrite a page that has a newer version than the base of the local copy, unless you give `--force`. | Partial | - -**S1**: `attachfile.Resolve` enforces it. It refuses a path that goes outside -the root. It does not clip the path. - -**S2** now holds for all 3 reads that `_plans/025` names. - -`root.FS` enforces it for the image leaf (`internal/convert/images.go`). -`root.FS` is an `os.Root` that the documentation root bounds. markfluence -refuses a path that escapes by its text before it asks `os.Root`. `os.Root` -also refuses an escape that only it can see: an intermediate directory that is -a symlink. The lexical comparison of `withinRoot` did not find that case. The -same leaf also refuses every symlink with `os.Lstat`, also a symlink that -resolves inside the root. - -Link and anchor resolution needs no clamp at all. `internal/linkindex.Build` -goes *down* from the root one time. Thus nothing outside the root can be in the -index, and markfluence never opens a file outside the root for this purpose. -The guarantee holds by construction, and not by a check, as `_plans/025` -describes. See [Non-goals](#symlinks). - -markfluence reads a `parent:` path in frontmatter in the same way as the image -leaf (`cmd/create.resolveParent`, through `root.FS`). But the failure is -different on purpose. A parent that escapes, or that is a symlink, is a **hard -error**. It is not an unresolved case that markfluence reports, as a link is. -The parent is important: a publish under the wrong parent, or under no parent -with no message, is worse than no publish (`_plans/026` commit 6). - -**S3**: `export` enforces it. It does a check of the destination. It skips the -Markdown file and each attachment unless you give `--force`. - -### S8 is about the page, and S3 is about the file - -**S3** (`no-overwrite-without-force`) protects a *file* on disk, so it is about -the filesystem. Nothing protected a *page* that exists, and that is the side -where the work of another person is. #149 was about that gap. Thus S8 is a new -guarantee, and not a wider S3. - -S8 uses a **merge base**: the version that this copy came from. The base is -different for each copy, because the page cannot know where a local copy came -from. Thus `create`, `update`, and `export` record it locally -(`internal/actionlog`). `update` refuses when the live version is newer than -the base. - -S8 is **Partial**. The gap is that protection starts at a later time, and it is -not a defect. A file with no recorded base has nothing to compare. It publishes -with no message, as markfluence did before S8 existed. A file gets protection -the first time that you publish or export it. Thus you need no setup step and -no adopt command. - -But the first run after S8 shipped gave no protection to any file that already -existed. A new clone also starts with no base, because the log is local to each -checkout and is not committed. `update` reports the count of files with no -check one time in each run. - -`--force` overrides S8, and that is the purpose of the flag. It is not a hole. -When a repository is the source of truth, you publish with `--force`. Then an -edit in the Confluence UI is drift that you overwrite, and not work that you -protect. - -### An overwrite and a removal are not the same risk +If a safety principle fails, the result is damage, and not only a wrong answer. + +### S1 `no-write-outside-root` + +markfluence writes no file outside the root. + +**Why:** Data from the server decides where some files go. An attachment +records the path of its source, and `..` is correct in a source path. Thus a +recorded path can point outside the destination. + +**Rules out:** A clip of a path that escapes. A clip writes *something*, with a +name that nobody chose. markfluence refuses the path. + +### S2 `no-read-outside-root` + +markfluence reads no file outside the root. + +**Why:** A publish sends what it reads to Confluence. A reference in a +Markdown file must not publish a file from outside the project. + +**Rules out:** + +- A check of the text of a path only. A directory in the path that is a + symlink can escape while the text stays inside the root. +- A `parent:` path outside the root that publishes as though it had no parent. + This is an error. A publish under the wrong parent, or under no parent, is + worse than no publish. + +### S3 `no-overwrite-without-force` + +markfluence does not overwrite a file that exists, unless you give `--force`. + +**Why:** A file on disk can have edits that exist nowhere else. + +**Rules out:** An export that replaces a file that you edited, because you ran +it again. + +### An overwrite and a removal are different risks S3 covers overwrites, and S4 to S6 cover removals, because the two risks have different shapes. @@ -131,85 +86,85 @@ gives your consent. A removal is **out of scope**. "Publish this file" or "export this page" does not include a removal. Thus the risk has no limit: which things, and how many. -There is also no defined way to give consent for it. "Not without `--force`" is -the wrong shape for a removal. The correct shape is that a removal does not -occur unless you asked for a removal. - -### Why S4 to S6 exist before any command removes things - -No command removes a local file (there is no `os.Remove` in the tree). No -command deletes anything in Confluence. Thus all 3 guarantees are vacuous and -not tested. We wrote them anyway, for two reasons. Without them, a guarantee -would stop being true one day with no warning. Also, we can already see two -places where a removal will be necessary: - -- **Orphaned attachments.** The identity of an attachment comes from the file - name of its asset (L3). Thus a rename of an asset leaves the old attachment - behind on the page. #99 tracks a future `attachment-prune` command to remove - those attachments. -- **`export --clean`.** Subtree export will need it. Then a new export does not - leave stale files for pages that somebody deleted in Confluence. - -**S5 is the guarantee with force, and the mechanism for it exists.** -`client.AttachmentMeta.Managed` is true when an attachment has the -`markfluence: ` prefix in its comment. It is false for an attachment that a -person uploaded by hand. Now, only `attachment-list` reports it. It lets a -prune remove markfluence attachments that nothing references, and never touch -a file that a person attached by hand. - -### S7 and the stub that create leaves behind - -**S7** is **Partial**. The exact limit is important, because the gap is not -the part that looks dangerous. - -`create` has 3 phases: - -1. Preflight does a check of every file. -2. Reserve creates a page with no content for each file, and writes its - `page_id` to the file. -3. Publish converts each file and fills in its page. - -If preflight refuses a file, the batch stops and creates nothing. Since #127, -this includes **every defect that the converter can find in the files on -disk**. Preflight converts each file and keeps the error. Thus a document that -the converter refuses never gets to the reserve phase. An example is two assets -that need the same attachment name. - -Before #127, such a document got to the reserve phase, and that is why this -guarantee was necessary. The author got a page with no content and a `page_id` -that they did not ask for. A second run then refused, because a page already -had that id. - -3 gaps remain, and only the first one is fully remote: - -- **A server failure or a network failure during publish.** The stub stays, and - its `page_id` is already in the frontmatter. Thus a plain `markfluence update` - completes the publish. -- **An attachment that markfluence cannot read.** `client.SyncAttachments` - opens every asset to calculate its checksum and to upload it. The converter - never opens an asset: it only calls `Lstat`. Thus an unreadable image fails - in the publish phase, locally, after the stub exists. Examples are an image - with mode `000`, or a file that somebody replaced between the two steps. - Preflight does not do this check, on purpose. It would copy the read that the - upload does anyway. It would also race the filesystem, so the check can pass - and the upload can still fail. -- **A frontmatter file that markfluence cannot write.** `reserveOne` calls - `os.WriteFile` *after* `CreatePage`. Thus a read-only `.md` file leaves a - stub whose id is **not** in the file. This is the one case where - `markfluence update` cannot continue the work, because nothing on disk names - the page. For this reason, `failKeepingPage` keeps the id and the URL in the - result. The output of the run is the only remaining record. - -The first gap is a decision, and not a thing that we forgot. `_plans/026` -accepted it as the cost of a reservation of every id before any conversion. -That reservation stopped link resolution from depending on the sequence of -creation. In all 3 cases, a page exists that the command reported as failed. -Thus the guarantee does not hold as written. - -To remove the gap, markfluence would have to delete the stub. This guarantee -cannot approve that change alone. The change would make **S4** and **S5** -not vacuous. S4 says that a removal occurs only when it is the stated purpose -of a command. +"Not without `--force`" is the wrong shape for a removal. The correct shape is +that a removal does not occur unless you asked for a removal. + +### S4 `no-removal-as-side-effect` + +markfluence removes nothing as a side effect. A removal occurs only when it is +the stated purpose of a command. + +**Why:** See above. Nobody who asks for a publish or an export expects to lose +anything. + +**Rules out:** A publish that deletes the attachments that no file references +now. An export that deletes the files of pages that are gone. Those removals +come as their own commands: #99 and #129. + +### S5 `remove-only-ours` + +markfluence removes only what markfluence created. + +**Why:** A person can attach a file to a page by hand, or put a file in a +directory. markfluence cannot know why it is there. + +**Rules out:** A prune that removes an attachment that markfluence did not +upload. + +### S6 `removal-is-previewable` + +A command that removes things says what it will remove before it removes them, +and it obeys `--dry-run`. + +**Why:** markfluence cannot undo a removal. + +**Rules out:** A command that removes things and has no `--dry-run`. + +### S7 `no-partial-create` + +If `create` cannot publish a file, it leaves no page behind. + +**Why:** An empty page, and a `page_id` that the author did not ask for, must +be undone by hand. A second run refuses, because the id is already in use. + +**Rules out:** A check that finds a defect in a file on disk after `create` +made a page. Every defect that the converter can find must stop the batch +before the first page exists. + +**Accepts:** `create` reserves every page before it publishes any, so that link +resolution does not depend on the sequence of creation. Thus some failures +during publish leave an empty page: a server or network failure, or an asset +that markfluence cannot read when it uploads. The `page_id` of that page is +already in the file, so `markfluence update` completes the work. A delete of +the empty page would be a removal as a side effect, which S4 forbids. + +### S8 `no-overwrite-of-a-moved-page` + +markfluence does not overwrite a page that has a newer version than the base of +the local copy, unless you give `--force`. + +**Why:** S3 protects a *file*. S8 protects a *page*, and the page is where the +work of other persons is. A page that changed after your copy came from it has +edits that your copy does not have. + +**Rules out:** + +- An `update` that publishes over a page that somebody edited after your last + publish or export. +- A base that the page records. The base is different for each copy, and the + page cannot know where a copy came from. + +**Accepts:** + +- Protection needs a base, and the base is local. markfluence records it in the + action log under a `markfluence.yaml`, when you publish or export a file. A + file with no base publishes with no check. This includes every file in a new + clone, because the log is not committed. +- Without a `markfluence.yaml`, there is no action log and thus no protection. + markfluence does not create a project file for you (#139). +- `--force` overrides S8, and that is the purpose of the flag. When a + repository is the source of truth, an edit in the Confluence UI is drift that + you overwrite, and not work that you protect. ## Laws @@ -221,455 +176,257 @@ The laws are algebraic properties of the 3 mappings that markfluence does: We state each law so that a property test can generate trees and assert it. -| | label | guarantee | status | -|---|---|---|---| -| **L1** | `resolve-what-was-named` | A reference resolves to the file that it names, or to nothing. | Holds | -| **L2** | `invocation-independent` | How a reference resolves, and the name of an attachment, depend only on the files on disk. They do not depend on the working directory, or on the other files in the same command. | Holds | -| **L3** | `identity-from-asset-location` | The identity of an attachment depends only on the file name of the asset. A move of the asset keeps the same attachment. | Holds | -| **L4** | `publish-is-idempotent` | A publish of a file that did not change makes no change in Confluence. | Holds | -| **L5** | `roundtrip-from-confluence` | If you export a page and then publish it again with no edits, the page does not change. | Partial | -| **L6** | `roundtrip-from-disk` | If you publish a file and then export it, the result is Markdown that publishes to the same page. | Partial | -| **L7** | `output-is-valid-markdown` | Everything that markfluence writes to disk is Markdown that renders. | Holds | -| **L8** | `no-layout-inference` | markfluence never gets the identity or the hierarchy of a page from the layout on disk. | Holds | -| **L9** | `declared-metadata-is-asserted` | markfluence makes a metadata field that a file declares true of the page. A field that the file does not declare leaves the page alone. | Partial | - -**L1** is about correctness, and not about how many files match. A lookup by -base name used to resolve to exactly one file, but not always the file that the -reference named. That is how a link to `sub/dup.md` got to `./dup.md`. -`internal/linkindex` now resolves by path. Thus a base name cannot match the -wrong file (`_plans/026` commit 5). - -**L2** is narrow on purpose. A flag such as the `--title` flag of `create` -changes what markfluence publishes, and that is its purpose. Thus the law -controls resolution and naming only. - -In that scope, the law does not let the root come from the working directory. It -also does not let the root come from the *set* of arguments. Otherwise, the same -file would get a different name for each batch. `internal/project` finds the -root when it goes up from the directory of each file. The working directory and -the other files in the command do not change the result (`_plans/026` commits 1 -to 4). - -A `space:` or `page_width:` default for the whole project in `markfluence.yaml` -(#100) is in that scope, and it helps L2. A committed file on disk declares the -value. markfluence finds that file by the same walk, which does not depend on -the working directory. Thus two persons in different directories resolve it the -same way. It is better for L2 than the `--space` flag that it replaces, because -a flag is invocation state by definition. The status does not change. - -A `pages:` entry (#139) takes the same argument further. That is why the path -keys are **lexical and relative to the root**, and markfluence does not resolve -them. If markfluence resolved a symlink, the meaning of a key would depend on -the layout of a checkout. Then the same repository could publish to different -pages on two machines, and L2 forbids exactly that. - -It is also why `update` lost `--title`, `--page-id`, and `--page-width`. Page -metadata now lives only in files on disk. Thus the page that a file publishes -to does not depend on how you ran the command. The exception that L2 gives to -flags is narrower than it was. - -One thing is outside L2, and it is better to say so now than to find it later. -The action log (#149) is **local to each checkout and not committed**. Thus two -persons who run the same command on the same tree can get different -*behavior*. One of them has a base for a file and gets a refusal for a moved -page. The other has no base and publishes it. - -That is not L2 as written. L2 controls how a reference resolves and the name of -an attachment, and neither of those changes. But it is against the purpose of -the law. It is also the reason that markfluence never keeps -`pagedoc.UserCache` on disk. - -The difference cannot be avoided. A merge base is different for each copy by -definition. The alternative is to commit the log, and that would help only the -one arrangement that does not need it. The published result does not change: -with the same files and the same page, a run that publishes sends the same -bytes. - -**L3** is why a move of a page or of an asset costs nothing. Two places in the -code enforce it: - -- `convert.AttachmentFilename` (`internal/convert/attachname.go`) gives an - attachment the base name of its file, and nothing else. -- `planAttachments` (`internal/client/client.go`) matches each local file to - an attachment on the page by that name. When the name agrees but the recorded - `Source` path does not, it updates the attachment to record the new path. It - does not create a second attachment. - -Thus a move keeps the attachment, and a rename of the file creates a new -attachment. The old attachment stays on the page, because markfluence never -deletes an attachment (see S4 to S6). - -The statement of L3 was corrected. It said that the identity depends on the -*location* of the asset. That was true when the name was the whole path, -encoded (`_plans/026` commit 4). It became false when the name became the base -name (`_plans/029`). The label is permanent (see *Rules for changes to this -document*), so `identity-from-asset-location` does not change, although the -guarantee is now about the file name. The status does not change. - -A history note: an earlier version of this entry said that a move of an asset -changes its identity, and that a fix would need names from the content of the -file. Neither is true now. The name never had to build the tree again on -export, because the attachment comment holds the path. - -**L4** had the status Holds when it was not true. It is better to record the -correction than to change it with no statement. The skip in `update` used the -mtime of the file, and git does not keep mtimes. Thus a clone, pull, checkout, -`touch`, file copy, or backup restore looked the same as an edit. Each one made -markfluence publish a file that did not change. - -On a new CI checkout, the mtime of every file is the time of the clone. Thus -the whole tree was published again on every run. That is a counterexample, and -not an edge case. The status should have been **Partial** from the first day. - -L4 holds now, by construction and not by approximation. `update` calculates a -sha of exactly what the body `PUT` would send: the resolved title and the -rendered body. It compares that sha with the sha that the last publish -recorded. When they agree, it sends no request (#149, `internal/actionlog`). -Nothing reads a clock. - -The scope of L4 is the **body**. markfluence still uploads an attachment whose -bytes changed. It still asserts a width or a label that the file declares, -because each one has its own pass. That is why the result is "no change", and -not "no request at all". - -A publish of a body that did not change would give the page a new version. It -would send a notification to every watcher. It would also fill the history of -the page with versions that are all the same. - -**L5** and **L6** stay **Partial**. This file said that #59 would decide them. -#59 showed that they cannot have the status Holds as they are written. - -Multi-page export now exists, and the earlier note said that it was missing. -It has attachment placement from provenance, it mirrors directories, and it -writes a tree whose `parent:` paths let it publish into new pages -(`_plans/029`). - -The attachment part of the round trip is also correct now, because of a -different change than we expected. An attachment gets the base name of its -file. Thus when export places the images of a page, their attachments do not -move (`_plans/029` §"The thing 025 got wrong"). - -The problem is the text of the law: *"publish it again with no edits, the page -does not change"*. We measured it on 2026-09-05 with a live page. An export and -a new publish change the stored storage format in two ways that have no relation -to the content: - -- The editor of Confluence writes `
  • text

  • `. The converter writes - `
  • text
  • `. -- A TOC macro has `ac:local-id`, `ac:macro-id`, and `data-layout` attributes. - The canonical form of the converter does not have them. - -Both forms render the same way, and neither form loses anything. That is the -stated design target of the converter: equivalence of meaning, and not -byte-for-byte equivalence. Thus the law as written asks for a thing that -markfluence does not do, on purpose. The honest reading is that L5 and L6 have -the wrong shape, and not that markfluence fails them. A change to their text is -a separate decision, and this document does not make it. - -A property test now proves one thing (`TestRoundTripMarkdownIsAFixedPoint`, -over every `storage2md` case and not a list that a person keeps). **Markdown is -a fixed point.** Export a page, publish that Markdown again, and export again. -The Markdown is the same. After a page goes through markfluence one time, it -stops changing. - -That test found real drift on its first run: a hard break got one more leading -space on every cycle. That is the argument for the test. It also answers the -note that this file added when #125 showed that L5 had no property test at all. - -We expect two exceptions to the fixed point, and both become stable on the -second cycle. A native Confluence attachment is not managed, so the first new -publish writes a new comment on it. Also, a mention from the editor of -Confluence has an `ri:local-id` that markfluence does not write (#91). Thus the -first new publish removes it. We found that this does no harm. A mention with -only the account id resolves to the same person -([links-and-anchors.md](confluence/links-and-anchors.md)). - -**#91 makes the gap in L5 and L6 smaller, but it does not remove it.** A -mention was the most frequent thing that an `` could be: 80% of all -use. Before #91, a mention stayed correct through the round trip only as raw -storage format. It now converts in both directions. Thus the *readable* part of -the round trip covers the case that most real pages have. The laws stay -**Partial** for the reasons above: the text asks for byte-for-byte -equivalence, and the converter does not have that target, on purpose. - -**L9** is what makes a new frontmatter field safe. Without it, each new field -has two bad defaults. Assume that markfluence always asserts the field. Then a -run that never mentioned the field reverts a page that a person set up by hand, -with no message. If markfluence never asserts the field, you cannot use it to -manage anything. The declaration of the field is the signal. - -L9 is **Partial**, and the gap is `page_width`, not `labels`. `update` obeys -the law exactly. It asserts a `page_width` line, and an absent line leaves the -live width alone. But `create` asserts a *default* width of `max` for a file -that declares no width. Thus in `create`, an absent field does not leave the -page alone. - -`labels` holds in both verbs. If the field is absent, markfluence makes no -label request at all. That is stronger than "no write", and a test in -`cmd/update` pins it. - -`page_status` (#168) holds in both verbs in the same way as `labels`, with the -same "no request at all" test. It has one difference from `labels`, and it is -better to write it down than to let a reader guess it. A file can set and -change a status, but a file **cannot clear** a status. - -`labels: []` can mean "remove them all", because an empty sequence has its own -spelling, and that spelling is different from a null scalar. A scalar field has -no such spelling. `page_status:`, `page_status: ~`, and `page_status: null` all -read as `""`, and that is exactly what an unfinished edit looks like. Thus -markfluence refuses an empty value. That is a missing *declaration*, and not a -declaration that markfluence does not assert, so the law is not weaker. To -clear a status, use the UI, until somebody asks for a spelling. - -The law now has **no exception**. It had one until `fix` was removed (#151). -`fix` changed the *file* to agree with the page. Thus in `fix`, the page was -the authority, and markfluence filled in an absent field. - -To `update`, an absent `labels` key meant "leave the page alone". To `fix`, it -meant "use the labels that the page has". A reader had to remember that -difference. Now every verb that writes goes in one direction, and the law -describes all of them. - -Thus you adopt a page that a person labeled or resized by hand with a manual -step. Look at the page, and then edit the file. `page-info` shows the labels, -the width, and the page status. `read` and `export` write all three. #154 -(`markfluence diff`) is the correct way to see what is different. Nothing -writes the file for you, on purpose. - -The status does not change. The default width of `create` makes L9 partial, -and the removal of `fix` did not change that. +### L1 `resolve-what-was-named` + +A reference resolves to the file that it names, or to nothing. + +**Why:** A near match publishes a link to the wrong page, and nothing reports +it. + +**Rules out:** A lookup by base name. That is how a link to `sub/dup.md` once +got to `./dup.md`. + +### L2 `invocation-independent` + +How a reference resolves, and the name of an attachment, depend only on the +files on disk. They do not depend on the working directory, or on the other +files in the same command. + +**Why:** Two persons, or a person and CI, publish the same tree. They must get +the same result. + +**Rules out:** + +- A root that comes from the working directory, or from the set of arguments. + Otherwise the same file gets a different name in each batch. +- A `pages:` key that markfluence resolves through a symlink. Then the meaning + of a key depends on the layout of a checkout. +- Page metadata from flags. That is why `update` has no `--title`, + `--page-id`, or `--page-width`. +- A cache on disk that changes what markfluence writes. + +The law controls resolution and naming only. A flag such as the `--title` flag +of `create` changes what markfluence publishes, and that is its purpose. + +**Accepts:** Two checkouts of the same tree can *behave* differently, because +the action log is local and not committed. One person has a base and gets a +refusal for a moved page. The other has no base and publishes. A merge base is +different for each copy by definition. The published bytes do not differ: with +the same files and the same page, a run that publishes sends the same bytes. + +### L3 `identity-from-asset-location` + +The identity of an attachment depends only on the file name of the asset. A +move of the asset keeps the same attachment. + +The label is older than the statement. The identity once came from the +location, and the label is permanent. + +**Why:** A move of a page or an asset in the tree is a usual operation. If the +name moved with the file, each move would upload a new attachment and leave +the old one on the page. + +**Rules out:** An attachment name that comes from the path of the asset. + +**Accepts:** + +- Two assets in one document with the same file name cannot both publish. R2 + reports it. +- A rename of an asset leaves the old attachment on the page, because + markfluence removes nothing as a side effect (S4). #99 is where a prune will + arrive. + +### L4 `publish-is-idempotent` + +A publish of a file that did not change makes no change in Confluence. + +**Why:** A publish of a body that did not change gives the page a new version. +It sends a notification to every watcher, and it fills the history of the page +with versions that are all the same. + +**Rules out:** A decision about "changed" from the mtime of the file. git does +not keep mtimes, so a clone, a checkout, or a CI run looks the same as an edit. +The decision must come from the content that markfluence would send. + +The scope is the **body**. markfluence still uploads an attachment whose bytes +changed, and it still asserts a width or labels that the file declares. Thus +the result is "no change", and not "no request at all". + +**Accepts:** The same dependency as S8. The comparison needs a record of the +last publish. That record exists only under a `markfluence.yaml`, and only +after the first publish or export of the file. Without it, `update` publishes +the body again. + +### L5 `roundtrip-from-confluence` + +If you export a page and then publish it again with no edits, the page keeps +its meaning. A second export gives the same Markdown as the first. + +**Why:** Persons edit exported files and publish them. A round trip that loses +content makes export a trap. + +**Rules out:** A conversion that drops a construct that it cannot map to +Markdown. markfluence keeps the storage format as raw markup instead, and +publishes it again unchanged. + +**Accepts:** The stored storage format can change. The Confluence editor writes +some forms that the converter does not, and both forms render the same way +([storage-format.md](confluence/storage-format.md)). The design target of the +converter is equivalence of meaning, and not byte-for-byte equivalence. + +### L6 `roundtrip-from-disk` + +If you publish a file and then export it, the result is Markdown that publishes +to the same page. + +**Why:** An export of a page that markfluence published must not be a +different document. + +**Rules out:** A published form that the reverse conversion cannot read back. +If the converter writes a construct, the reverse conversion must recognize it. + +**Accepts:** The same as L5. The exported Markdown can differ in spelling from +the file that you published, but it publishes to the same page. + +### L7 `output-is-valid-markdown` + +Everything that markfluence writes to disk is Markdown that renders. + +**Why:** Persons and tools read these files, and markfluence publishes them +again. + +**Rules out:** Output that only markfluence can read, such as a link +destination that is not encoded and thus does not parse. + +L7 is separate from C2 because each one has a different external +specification. A file with bad frontmatter can still render as Markdown. + +### L8 `no-layout-inference` + +markfluence never gets the identity or the hierarchy of a page from the layout +on disk. + +**Why:** A move or a rename in the tree is a usual operation. It must not +change which page a file publishes to, or where that page is. `page_id` gives +the identity, and `parent:` gives the hierarchy. + +**Rules out:** A directory that becomes a parent page, or a file name that +selects a page. The file names that `export` writes are for persons only. + +### L9 `declared-metadata-is-asserted` + +markfluence makes a metadata field that a file declares true of the page. A +field that the file does not declare leaves the page alone. + +**Why:** Without this rule, each new field has two bad defaults. If +markfluence always asserts the field, a run that never mentioned it reverts a +page that a person set up by hand, with no message. If markfluence never +asserts it, you cannot use the field to manage anything. The declaration is the +signal. + +**Rules out:** + +- A request about a field that the file does not declare. That includes a + read. +- A verb that writes the file from the page. Every verb that writes goes in one + direction: from the file to the page. To adopt a page that a person set up by + hand, look at the page (`page-info`, `read`, `diff`), and then edit the file. + +**Accepts:** A file cannot clear a scalar field such as `page_status`. An empty +value looks exactly like an unfinished edit, so markfluence refuses it. Use +the Confluence UI to clear a status. `labels: []` can mean "remove them all", +because an empty sequence has its own spelling. ## Conformance -| | label | guarantee | status | -|---|---|---|---| -| **C1** | `preview-compatible-resolution` | A reference resolves the same way that a Markdown preview resolves it, the GitHub preview included. | Holds | -| **C2** | `frontmatter-is-valid-yaml` | The frontmatter that markfluence writes parses as YAML, and it reads back as the values that markfluence wrote. A value is a single-line scalar, or a sequence in either YAML style whose elements are all single-line scalars. | Holds | - -**C1** is not an internal property. It is agreement with an external -specification. It always held for images, which resolve relative to the page. -Links now resolve the same way. Inside markfluence they are relative to the -root, but markfluence composes them from the directory of the page that -references them, as a preview does. They do not resolve by base name in one -directory (`_plans/026` commit 5). - -C1 is separate from L1 because we could, in principle, give up C1. -markfluence could use its own resolution rules and document them. We could not -give up L1. - -**C2** was false until #130, and only an outside tool could see that it was -false. `internal/frontmatter` was a flat-key parser that we wrote by hand. It -split each line at the first `:`. Thus it read its own -`title: Deploy Runbook: Part 2` correctly, and every real YAML parser refused -it. The colon was one member of a class. markfluence wrote booleans, numbers, -nulls, flow collections, and the reserved indicators with no quotes. It read -them back as the wrong type, or not at all. - -C2 holds now for two reasons. `goccy/go-yaml` parses and writes the block. The -writer also **does a check of its own output**, and it does not trust it. It -writes with the style that goccy chooses, and it reads the result again. When -the two disagree, it writes a double-quoted scalar instead. - -That fallback is necessary, and not only an extra safety. goccy drops a tab. It -writes a value that starts with `? ` as a document that it then cannot parse. -It writes `.inf` and `.nan` with no quotes, and every conforming reader then -sees a float. - -A predicate that lists those shapes would be incomplete, because we found them -only by experiment. A check is better than a prediction, and the check stays -correct if goccy gets worse. - -The check compares the **node kind** that it reads again, and also the text. -That is what makes the `.inf` case work. An earlier version compared only the -text, and it passed. The reader changes every scalar to its token. Thus -markfluence read `.inf` back as `.inf`, and the round trip looked correct. But -the file said "float" to every other tool. A comparison of text compares the -wrong thing. - -markfluence enforces the single-line rule when it **reads**. It refuses a -scalar whose source is on more than one line. It must do this, because -markfluence writes a key that it did not change from the node that the parser -made. When goccy writes a parsed node again, the result is not always the same -as the source. - -We measured this with the pinned version of goccy. goccy writes a `|` block or -a `>` block again as a block, and then the node-kind whitelist of the reader -refuses it. Thus a write would make a file that markfluence cannot read. In -`create`, that occurs only after markfluence made the page. - -A plain scalar on more than one line, and a quoted scalar on more than one -line, are less dangerous. goccy writes both again on one line. The result -parses, but markfluence changes the file of the author with no message. That -must not occur while markfluence sets a different field. - -Sequences made this rule larger, but not weaker. A value can now be a list, in -either YAML style. The rule is now: "a *sequence* can be on more than one line, -but every element must be a single-line scalar". Thus a block list is correct, -and a flow list on more than one line is correct. An element that continues on -the next line is not correct. - -The element check trims its origin at both ends. A leading newline in the -origin of an element is structure (the item started on a new line), and not -content. - -The self-check of the writer had to get a second context. The reason for that -is the nearest thing to a repeat of #130. The scalar check tests a value as a -*mapping* value, and two shapes pass it and then break a list. In a flow -sequence with no quotes, `x,y` becomes two elements, and `has]bracket` ends -the sequence. - -Block style has a different trap. There, `? q` is the explicit-key indicator -of YAML, and it parses as a mapping. Thus the check runs in the style that -markfluence will write. Flow is the stricter context, so a check in flow style -is always *safe*. The style parameter gives the minimum of quotes. It protects -against one fault: a check in block style for a write in flow style. - -markfluence now writes both styles, because a rewrite keeps the style that it -found. Some sets are large enough to be a block list. The flow spelling of -such a set is a single line that nobody can read. - -markfluence keeps the style, but not the format. It writes block items at its -own indent, and not at the indent of the author. Valid YAML in the correct -style is the guarantee. Byte-for-byte preservation is not. - -One result is important, because a code comment used the old rule as its -reason. `Normalize` removes blank lines with a **textual** filter. The comment -said that this was safe "because nothing this package emits spans more than one -line". A block list that markfluence passes through makes that false. - -The correct statement is that no value that markfluence can write has a -*meaningful* blank line. A blank line between two list items has no effect in -YAML. The shape where a blank line is content is a `|` block, and markfluence -refuses that shape when it reads. - -C2 has two limits, and we state them. The quotes come from goccy, but **the -types come from markfluence**. markfluence writes `page_id` and `parent` as -YAML integers and nulls, because `page_id: "123"` is valid YAML that says the -wrong thing. That is a rule that we wrote by hand, and it applies only to two -keys whose sets of values are closed. - -Also, the check is a *self*-check. It proves that goccy can read again what -goccy wrote. It does not prove that a different implementation can. There is -no second YAML library in `go.mod` to decide that, on purpose. If a real tool -reports a difference, that is an issue to correct, and the frontmatter that -fails is the evidence. - -C2 is separate from **L7** (`output-is-valid-markdown`), because each one has -a different external specification. A file with bad frontmatter still renders -as Markdown. GitHub shows it, and the YAML extension of VSCode was the tool -that found the fault. If L7 included C2, its status would not tell you which -specification failed. +### C1 `preview-compatible-resolution` + +A reference resolves the same way that a Markdown preview resolves it, the +GitHub preview included. + +**Why:** Authors check their documents in a preview. If markfluence resolves a +reference differently, the preview shows the wrong thing. + +**Rules out:** A link that resolves by base name in one directory, or relative +to anything other than the file that contains it. + +C1 is separate from L1. We could give up C1, use our own resolution rules, and +document them. We could not give up L1. + +### C2 `frontmatter-is-valid-yaml` + +The frontmatter that markfluence writes parses as YAML, and it reads back as +the values that markfluence wrote. A value is a single-line scalar, or a +sequence in either YAML style whose elements are all single-line scalars. + +**Why:** Other tools read these files: editors, GitHub, and YAML linters. A +file that only markfluence reads correctly is broken. + +**Rules out:** + +- A parser that is not a YAML parser, such as one that splits each line at the + first `:`. +- A value that a conforming reader gets as a different type, such as `.inf` as + a float or `page_id: "123"` as a string. +- A write of one field that changes the text of another field. + +**Accepts:** + +- markfluence checks its output with its own YAML library only. That proves + that the library can read what it wrote, and not that every implementation + can. A difference that a real tool reports is an issue to fix. +- A rewrite keeps the YAML style of a list, but not its indent. ## Reporting These are not invariants. A publish of a dead link can be acceptable. A publish of a dead link **with no message** is not acceptable. The obligation is to tell -the user. Thus R1 can be false while nothing calculates a wrong answer. - -| | label | guarantee | status | -|---|---|---|---| -| **R1** | `report-unresolved-references` | markfluence reports every reference that it could not resolve. | Holds | -| **R2** | `report-unplaceable-attachments` | markfluence reports every attachment that it could not name or place. | Holds | - -**R1** was false by design, and a document said so. The README said that an -unresolved link was "published as-is, which on Confluence is a dead relative -link. There is no warning for this." - -An unresolved image already used the `warnings` list. A same-tree `.md` link -that does not resolve was the next thing to go into it (`_plans/026` commit 5). -The status was Partial, and not Holds. That was a minimum warning with a -mechanism that already existed. It was not the dedicated diagnostic that -`_plans/025` described and left for later. That diagnostic tells you *why* a -reference failed, and it lets you audit a tree with no publish. - -That dedicated diagnostic now exists (#42). A doc-link target that is missing, -or that resolves outside the documentation root, is Broken. It replaces the -published element, as a broken image already does. Some targets exist but -have no `page_id` yet, or have a `#fragment` that matches no heading. Such a -target gives a warning, and markfluence does not publish it with no message. - -Every message gives the source line that it came from, when markfluence can -find it. A link or image with no visible text has no `*ast.Text` to go to. -Thus markfluence reports the message with no line prefix, and not with a wrong -line (the documented `ok=false` case of `nodeLine`). `check` adds the audit of -a tree with no publish. It gives the same diagnostics, and it never touches -Confluence. It only reads the filesystem. - -R1 covers the two types of reference that markfluence tries to resolve: -doc-links (`.md` siblings) and images. A relative link to a local file that is -not `.md` and not an image, such as a PDF, gets no existence check at all. That -was true before #42 and after it. - -`rewriteDocLink` tries to resolve only an href that ends in `.md`, so -markfluence never tries to resolve that case, by design. markfluence uploads -only images, and a relative href to any other file would be dead in all cases. -That case is outside the claim of R1, and not a part of R1 that fails. Thus it -does not stop the status Holds. - -**R2** covers naming and placement. The note is wider, but the label is not. -Labels are permanent (see *Rules for changes to this document*). Thus -`report-unplaceable-attachments` does not change, although the guarantee now -starts one step earlier in the pipeline. - -The step is new. Since `_plans/029`, an attachment gets the base name of its -file. Thus two assets in one document can need the same name. There is no -correct way to publish that, because a name is unique on a page. The converter -refuses the file, and it names both paths and both lines. - -`attachment-upload` refuses the same collision in a batch. `check` reports it -offline as a Broken entry, in the same list as a dead link. In both cases, the -fix is to rename a file. An attachment that markfluence cannot name -is as unusable as an attachment that it cannot place. The obligation to report -it is the same. +the user. -## Non-goals +### R1 `report-unresolved-references` -These are decisions about what markfluence will not do. They are in this -document because they control future work in the same way as a guarantee. A -request that needs a reversal of one is a design discussion, and not a bug -report. +markfluence reports every reference that it could not resolve. -### Symlinks +**Why:** Otherwise a reader finds the dead link on the published page, too +late. -**markfluence does not follow symlinks.** This includes a symlink that goes -outside the project, and a symlink that stays inside it. There is one rule, -because two rules made a system where the same symlink worked for an image and -failed for a link. +**Rules out:** A publish of an unresolved link or image with no message. -There are 3 points of enforcement, from the most work to the least work: +The scope is the references that markfluence tries to resolve: links to `.md` +files and images. markfluence uploads only images, so it does not check a +relative link to any other file, such as a PDF. -| where | mechanism | cost | -|---|---|---| -| the link and anchor index | `filepath.WalkDir`, which reports a symlinked directory and does not descend it | free | -| the read of a leaf, such as an image | `os.Lstat`, and a refusal of anything that is not a regular file | one call | -| an escape through an intermediate symlinked directory | `os.Root` bounded to the root, which refuses it | `internal/attachfile` already uses this pattern | +### R2 `report-unplaceable-attachments` -**Verified 2026-08-28.** `WalkDir` from `docs/` over a tree that has -`docs/escape → ../outside`: +markfluence reports every attachment that it could not name or place. -``` -dir docs/ -SYMLINK (not descended) docs/escape ← outside/out.md never enumerated -dir docs/sub/ -file docs/sub/in.md -``` +The label is older than "name". It is permanent. -This is `os.Root` bounded to `docs/`, for the escape case that an `Lstat` on a -leaf cannot see: +**Why:** An attachment that markfluence cannot name is as unusable as one that +it cannot place. -| path | `os.Root` | -|---|---| -| no symlink | allowed | -| relative symlink that stays inside the root | allowed | -| relative symlink that escapes the root | refused: `path escapes from parent` | -| absolute symlink, any target | refused | +**Rules out:** A publish that drops one of two assets that need the same name. +A name is unique on a page, so markfluence refuses the file and names both +paths. + +## Non-goals + +These are decisions about what markfluence will not do. A request that needs a +reversal of one is a design discussion, and not a bug report. + +### Symlinks + +**markfluence does not follow symlinks.** This includes a symlink that goes +outside the project, and a symlink that stays inside it. There is one rule, +because two rules made a system where the same symlink worked for an image and +failed for a link. This rule has two good results. **The origin of the root stops being important.** markfluence addresses paths -relative to the open root handle. Thus a tree in a checkout that is a symlink -works. On macOS, `/tmp` is `/private/tmp`, and home directories are frequently -links. None of that needs a special case. +relative to the open root. Thus a tree in a checkout that is a symlink works. +On macOS, `/tmp` is `/private/tmp`, and home directories are frequently links. +None of that needs a special case. **markfluence does not support an asset directory that you share with a symlink**, on purpose. Persons use a symlink to let many pages use one asset @@ -678,39 +435,12 @@ that gives the same result directly. Also, git symlinks do not work on Windows without `core.symlinks` and developer mode. Thus a repository that uses them cannot be cloned everywhere. -This is how **S2** stopped being lexical. The converter does a check of each -image with `root.FS.Lstat` (`internal/convert/images.go`). `root.FS` is an -`os.Root` bounded to the documentation root (`internal/project/project.go`). -It refuses an escape through a symlinked directory, and the `Lstat` refuses a -symlinked leaf. The lexical `convert.withinRoot` clamp, which could not see a -symlinked directory, no longer exists. - -One limit remains. The check goes through `os.Root`, but the upload does not. -The converter gives the upload an ordinary path, `filepath.Join(root.Dir, -rootRel)`, and `fileChecksum` and the upload call `os.Open` on it -(`internal/client/client.go`). If somebody replaces a directory with a symlink -between the check and the upload, the uploaded bytes can come from outside the -root. That is a race, and a static layout of files cannot cause it. The S7 -note already accepts a file that somebody replaces between the two steps, for -a different reason. - -## When markfluence cannot hold a guarantee +## When markfluence cannot hold a principle markfluence does its best. If that is not possible, it gives an error that names the problem and a safe next step. It never gives a partial success with no message. -This is a policy, and not a guarantee. A property test cannot prove it, and it -applies to all of the guarantees above. It is also the reason that S1 refuses -an attachment path that goes outside the root, and does not clip it. A clip -would write *something*, with a name that nobody chose. - -## How we do a check of each kind - -| kind | how we do the check | -|---|---| -| Safety | adversarial tests: tries to go outside the root, files that already exist, and a preflight that fails and must create nothing | -| Laws | property tests: generate trees and assert the equation | -| Conformance | C1: fixtures, compared with what a Markdown preview renders. C2: the writer does a check of its own output when it runs, and there is a round-trip test and a fuzz target. Agreement with *other* YAML implementations is review judgement | -| Reporting | example tests that assert that a specific message appears | -| Policy | review judgement | +This is a policy, and not a principle. A property test cannot prove it, and it +applies to all of the principles above. It is also the reason that S1 refuses +a path, and does not clip it. From 1ace7024816c46882b05adf03e16691042fc3c5c Mon Sep 17 00:00:00 2001 From: Will Kahn-Greene Date: Wed, 23 Sep 2026 10:12:33 -0400 Subject: [PATCH 2/6] docs: stop citing principle statuses docs/guarantees.md no longer carries a status, so references to Holds and Partial in CLAUDE.md, the README, and three code comments now point at nothing. CLAUDE.md's description of the file also had stale id ranges (S1-S7, L1-L8, C1). Refs #191 --- CLAUDE.md | 10 +++++----- README.md | 2 +- cmd/create/create.go | 2 +- cmd/export/actionlog.go | 2 +- internal/actionlog/actionlog.go | 2 +- 5 files changed, 9 insertions(+), 9 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index bb2b3c0..0a4b65e 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -56,7 +56,7 @@ Module `github.com/mozilla/markfluence` (`go 1.25`). `main.go` is a shim to `cmd **Before changing anything that talks to Confluence, read [docs/confluence/](docs/confluence/)** — what we established by experiment, since Atlassian documents little of it. Two traps recorded there have each already produced a confident wrong conclusion: `body-format=view` is not what the browser renders, and `body.storage` proves only what was stored, never what takes effect. -**[docs/guarantees.md](docs/guarantees.md) holds the properties markfluence holds itself to** — safety (S1-S7), laws (L1-L8), conformance (C1), reporting (R1-R2). Each carries a status, because several are aspirational rather than true today: a spec or PR cites them by id to say what it changes. The ids are permanent and never reused, and a change that downgrades a status says so in the commit message and in that file rather than letting it be noticed later. +**[docs/guarantees.md](docs/guarantees.md) holds the design principles markfluence holds itself to** — safety (S1-S8), laws (L1-L9), conformance (C1-C2), reporting (R1-R2). Each is a statement with its **Why**, what it **Rules out**, and what it **Accepts** (a deliberate trade-off, never a bug); a spec or PR cites them by id to say what it changes. The file deliberately carries **no status, no enforcement notes and no history** — those went stale faster than they could be maintained (#191) — so a gap between a principle and the code is a GitHub issue naming the id, never an edit to the file, which changes only when a principle does. The ids and labels are permanent and never reused. ### Layout @@ -76,7 +76,7 @@ Module `github.com/mozilla/markfluence` (`go 1.25`). `main.go` is a shim to `cmd - `schema/` — the published JSON Schema (`json-output/v1.json`) *and* the `schema` Go package that embeds it (`V1`). The Go file lives beside the schema because `go:embed` cannot reach outside its own directory, and the schema stays at a top-level path a non-Go consumer can browse, mirroring its own `$id`. `internal/schematest` validates against the embed rather than reading the file, which is what makes "what ships" and "what the tests checked" the same bytes — do not reintroduce a disk read or a second copy. The version number is **not** restated here: `jsonout.SchemaVersion` and the document's own `schema_version` const are the two copies, tied together by a test in `cmd/schema`. The envelope's and the error object's top-level **`warnings`** are the one field no command fills: `jsonout.NewEnvelope`/`EmitError` drain a package-level collector (`AddWarning`), because the only thing in it is raised during credential resolution — below any command, before either document exists. - `internal/project` — the documentation root: `Discover` (walk up from a directory looking for `markfluence.yaml`), `FromPath` (`--root`), `Resolve`, and `Cache`, which consults itself at every level of the walk so a batch spanning a subtree pays for the walk — and `os.OpenRoot` — once rather than per directory (the quadratic cost `_plans/025` measured). A `Root` carries `Dir`, `File`, `Config` and an `os.Root` that refuses an escape even through a symlink partway down. `Discover` is called from **two starting points for two reasons** — once per invocation from the working directory to locate `.env`, and once per Markdown file from its own directory to bound its reads and name its attachments — which is why it returns a type rather than a string; the two diverge legitimately, so a multi-root batch is allowed and nothing refuses it ([docs/root-model.md](docs/root-model.md)). `config.go` reads the project file's **settings** (#100): `space` and `page_width`, resolving **flag > frontmatter > project file**, which is *not* the credentials chain and must never be conflated with it. Three things about it are load-bearing. It is read through `frontmatter.Dialect.ReadMapping` rather than a second parser, since every rule there was found by probing goccy and a second copy would be a second set of the same bugs. An **unknown top-level key is fatal**, and that is the point rather than a cost — a silently ignored `spce: ENG` is wrong for every file at once, and a file written for a newer markfluence holds keys this binary would ignore, so there is no schema version and this must not be loosened; an empty or comment-only file stays valid, being what ships and what `export` plants. And loading happens in `open()`, the single place a `Root` is built from a marker hit, so `Discover`/`Cache`/`FromPath` cannot disagree that a file which cannot be understood **is not a valid marker**: the walk does not continue upward and does not fall back to the starting directory, because the root decides every attachment name and guessing at it is worse than stopping. `ConfigError`/`IsConfigError`/`RootError` exist so a caller reports that as a local defect (`VALIDATION`) rather than under `resolving the documentation root` as I/O. `pages.go` holds the **`pages:`** key (#139) and `SetPageEntry`, which records one file's metadata there: read-modify-write **once per page, not once per run**, because `create` writes each file's frontmatter as that page is published and a run that dies partway has to leave every already-created page recorded (D10). It **retries once** on a concurrent write, because the read-modify-write is not serialized and this file is shared by every page in the project: before the manifest each page's metadata went into its own file, so two concurrent `create`s could not collide, and they now can — A reads, B reads, A writes, B writes, and A's entry is gone while A's page exists. Optimistic rather than locked, matching `client.SetContentProperty`'s retry-once, since a lock file brings stale-lock handling for a verb a person invokes by hand. `beforeReplace` is a test hook for the give-up path, the `SetRetryLogger` arrangement. It also verifies twice — `frontmatter.SetNested` re-reads its own output, then `parseConfig` (split out of `loadConfig` for this) re-runs the *loader's* rules — since a tool that corrupts the file it is recording success in is the worst version of the feature; a test pins that a file which would not load afterwards is left byte-identical. A root with no project file refuses rather than creating one (#5's question, not a `create`'s to answer silently). An `Entry` is `{Fields, Lists}`, the same two maps `frontmatter.MarkdownFile` carries, which is the design rather than a convenience — `pagewidth` and `labels` both reach `client`, which holds a `*Cache`, so a typed validated entry would need a broken cycle or a second copy of every field's rules, and with two maps `labels.Declared(e.Lists, e.Fields)` works unchanged. `entryFields` is the manifest's schema and the only place it is written down; it mirrors frontmatter's fields deliberately, since adding one there and not here would make a field expressible in a file and not in an entry. Load checks **structure** — a mapping of mappings, legal paths, known field *names*, right shapes — and never a field's **value**, because #139 requires a semantically bad entry to be reported only when its file is one of the arguments, and this package has no idea which files the command was given. An unknown field *name* is the exception and is fatal at load, being the same typo class as an unknown setting. `NormalizePageKey` is lexical (L2 forbids a key whose meaning depends on the checkout's layout), and an escaping key or two keys normalizing to one are load-time errors naming both spellings. `Config.Pages` is **nil when there is no `pages:` key and empty-non-nil for `pages: {}`**, which is how a command tells "has not chosen the manifest" from "has, and has registered nothing". What it validates is **structure only**: `internal/pagewidth` cannot be imported here (`pagewidth` → `client` → `project`), so a width's vocabulary is checked by `pagewidth.Declared` where it already runs and by `check`'s offline lint. `Config` deliberately holds **no `url` or token**, and the reason is sharper than "those are credentials": basic auth goes to whatever host the resolved URL names, so a committed, walked-up file naming one would decide where `CONFLUENCE_TOKEN` is sent — a worse version of the `.env` hole #136 records. - `internal/pagemeta` — `Resolve`, the one merge of a file's page metadata from the two places it may live: its own frontmatter and a `pages:` entry in the project file (#139). A package because `update`, `create`, `check` **and `internal/linkindex`** all need it, and a per-command copy is how two commands come to publish one file to two different pages. It imports `frontmatter` and `project` and nothing else, which is also why it validates no *value*: `pagewidth` and `labels` are unreachable from here, and the commands that need them already call them on the maps it returns. **Frontmatter and an entry are two spellings of one level, not two levels of a precedence chain** — when both speak the rule is not "the higher wins" but a grading: `page_id`/`space`/`parent` are coordinates and a disagreement fails the file (for `create`, the batch, since it preflights everything), while `title`/`page_width`/`labels` are visible and recoverable, so they warn and frontmatter wins. Agreement is silent, which is what makes migration incremental. Three things took a second pass and should not be flattened: `Source` (what `--json` reports as `metadata_source`) and `Managed` are computed from **different predicates** — the first answers "who contributed metadata", the second "should `update` act on this file", which is true when an entry exists or a `page_id` is named, so `a.md: {}` is a claim that must fail for want of an id rather than be skipped; contributed metadata is counted only over fields `project.IsPageField` knows, or a file carrying only keys markfluence *preserves but does not understand* (`reviewers:`, pinned by a frontmatter test) reads as claimed and a whole tree of them fails; and `labels` is compared as a **set**, since Confluence has no label order and a reordering cannot reach the page. A blank value is not a disagreement — every null spelling already reads as `""`. `KeyFor` is the one place a file's path becomes a manifest key, used on both sides because a mismatch is a silent skip rather than an error. `Origin` is the same question per field — which location supplied *this* value — recorded as `Resolve` grades them, because `diff` reports it beside every frontmatter difference and "the title differs" is otherwise ambiguous about which file to edit; unlike `MetadataSource()` it does **not** collapse `FromBoth`, which is the most useful of the three there (correcting a field two locations supply means editing two files). Computed here rather than by the caller for the reason the package exists: a second copy of the precedence rules is a second copy whatever it is used for. -- `internal/actionlog` — the append-only record of what markfluence published, one log per project root at `/.markfluence/log.jsonl` (#149). `create`, `update` and `export` each append a line as a page completes, carrying the page version they left behind and `Sum`'s hash of what the body `PUT` sent; the last successful line for a file is that copy's **merge base**. It exists because nothing else can tell "the page differs because I have edits" from "the page differs because somebody published first" — that needs what *this copy* was derived from, which is a per-copy fact no page-side state can hold. Three things about it are load-bearing. **It is not committed**, and the reason is structural: a shared repository is itself a declaration that the repository is the source of truth, which is the arrangement where `update --force` is the answer and no base is consulted — so the log serves a local copy with the source of truth in Confluence, where per-checkout state is the right shape. The directory ignores itself (a planted `.gitignore` holding `*`, never overwritten) rather than editing a `.gitignore` markfluence does not own. **Nothing in it may fail a command**: a missing, unreadable, corrupt or half-written log degrades the check that reads it and never the run, which is why `read` skips a line it cannot parse instead of erroring and why a failed `Append` is a warning at the call site — the page is published by then, so failing would report that it was not. And **`Sum` covers exactly what the body `PUT` sends**, the resolved title and the rendered body: the title because `convert.ConfluencePage` carries none (a render-only hash would skip the publish of a file whose only change was its title), and *not* page width, labels or attachments, each of which has its own pass that runs whether or not the body is republished — folding them in would bump the page version for a change that never touched the body. Its parts are length-prefixed, which is part of the persisted format: changing the framing invalidates every recorded base. A root with `File == ""` gets no log at all (`For` returns nil, and every method is nil-safe), matching #139's rule that a root with no project file refuses rather than creating one. `Cache` hands out one `Log` per root so a batch reads each log once, and a batch spanning roots writes to several. `update` reads it for **two orthogonal checks, both only when `--force` is absent** (`--force` means always PUT, and no logic may suppress the request): the logged **`page_version`** against the live one decides **divergence** and *refuses* the file with `CodeConflict` — deliberately not qualified by the sha, since the case where you have no local edits is the worse one, publishing their work away with the bytes they started from — and the logged **`publish_sha256`** against this run's decides **idempotence** and skips the *body* `PUT` alone, leaving the attachment, width and label passes to run. An entry naming a different `page_id` is discarded rather than compared (a retarget would otherwise read as "the page moved 40 versions"), and the two fields degrade independently, so an `export` line written before its sha pass still refuses a moved page. A body-unchanged skip **records a line too**, which is load-bearing: once the sha does the skipping most runs skip, and a publish-only log would never keep a base current in a tree that is already published. Guarantee: **S8** (`no-overwrite-of-a-moved-page`, Partial — a file with no base publishes), and **L4** now holds by construction rather than by the mtime approximation that replaced it. +- `internal/actionlog` — the append-only record of what markfluence published, one log per project root at `/.markfluence/log.jsonl` (#149). `create`, `update` and `export` each append a line as a page completes, carrying the page version they left behind and `Sum`'s hash of what the body `PUT` sent; the last successful line for a file is that copy's **merge base**. It exists because nothing else can tell "the page differs because I have edits" from "the page differs because somebody published first" — that needs what *this copy* was derived from, which is a per-copy fact no page-side state can hold. Three things about it are load-bearing. **It is not committed**, and the reason is structural: a shared repository is itself a declaration that the repository is the source of truth, which is the arrangement where `update --force` is the answer and no base is consulted — so the log serves a local copy with the source of truth in Confluence, where per-checkout state is the right shape. The directory ignores itself (a planted `.gitignore` holding `*`, never overwritten) rather than editing a `.gitignore` markfluence does not own. **Nothing in it may fail a command**: a missing, unreadable, corrupt or half-written log degrades the check that reads it and never the run, which is why `read` skips a line it cannot parse instead of erroring and why a failed `Append` is a warning at the call site — the page is published by then, so failing would report that it was not. And **`Sum` covers exactly what the body `PUT` sends**, the resolved title and the rendered body: the title because `convert.ConfluencePage` carries none (a render-only hash would skip the publish of a file whose only change was its title), and *not* page width, labels or attachments, each of which has its own pass that runs whether or not the body is republished — folding them in would bump the page version for a change that never touched the body. Its parts are length-prefixed, which is part of the persisted format: changing the framing invalidates every recorded base. A root with `File == ""` gets no log at all (`For` returns nil, and every method is nil-safe), matching #139's rule that a root with no project file refuses rather than creating one. `Cache` hands out one `Log` per root so a batch reads each log once, and a batch spanning roots writes to several. `update` reads it for **two orthogonal checks, both only when `--force` is absent** (`--force` means always PUT, and no logic may suppress the request): the logged **`page_version`** against the live one decides **divergence** and *refuses* the file with `CodeConflict` — deliberately not qualified by the sha, since the case where you have no local edits is the worse one, publishing their work away with the bytes they started from — and the logged **`publish_sha256`** against this run's decides **idempotence** and skips the *body* `PUT` alone, leaving the attachment, width and label passes to run. An entry naming a different `page_id` is discarded rather than compared (a retarget would otherwise read as "the page moved 40 versions"), and the two fields degrade independently, so an `export` line written before its sha pass still refuses a moved page. A body-unchanged skip **records a line too**, which is load-bearing: once the sha does the skipping most runs skip, and a publish-only log would never keep a base current in a tree that is already published. Principles: **S8** (`no-overwrite-of-a-moved-page`) and **L4** (`publish-is-idempotent`), decided by content rather than by mtime; both accept that a file with no base, or a root with no project file, gets neither. - `internal/pageslug` — `Slug`/`For`/`Filename`: a title to a filename-safe slug. A package rather than a helper because `export`, `read` and `attachment-download` all place attachments under a page's own directory and must agree. It lowercases (so case-variant titles collide and can be caught) and drops `/` (so no title can inject a path separator); it is lossy, and no readable slug can avoid being, so the caller decides what a collision means. Known limit: NFD and NFC spellings of one title are different Go strings but one filename on APFS, so that pair is not disambiguated. - `internal/pagedoc` — a fetched page as a Markdown document: `Render` (frontmatter + converted body), `Frontmatter` (which since #168 also emits `page_status`, widening `RenderFrontmatter`'s positional signature and adding an unconditional `GET /content/{id}/state` per page — a third best-effort read beside the width and label ones, so `export --space --depth all` pays it for every page in the space), and the lookups the converter can't do for itself — `Sources`/`SourcesFrom` (attachment name → recorded source path) and `PageLinks` (the page an `` points at → its URL). **One conversion, parameterized by a `Placement`**: where the page's file sits, where its unrecorded attachments go, what `parent:` says, and the attachment listing the caller already has. `read`, `export` and `attachment-download` all go through `Options`/`AttachmentDirFor` rather than assembling their own, so they cannot drift by accident — only by argument. For a page at the top level of what is being written, which is what `read` prints and what a single-page export writes, `read` and `export` are byte-identical; deeper in a tree they differ in exactly the position-dependent parts (a sourced attachment's `../` prefix, a `-` suffix a sibling forced, and `parent:`), because `read` has no tree to be positioned in. It needs a client (page width, attachment list, title lookups), which is why it isn't in `internal/convert` — that package is deliberately client-free, and it's why `StorageToMarkdown` takes those maps rather than fetching them. Every one of them is best-effort in the same shape: no references in the body means no request at all, and a lookup that fails is omitted rather than fatal (an omitted page link renders as raw storage, not as a link with no destination). It also owns **`UserCache`**, the user-name lookup both mention directions share (#91): a per-run, **cross-page** cache, threaded in from the caller the way `project.Cache`/`linkindex.Cache` are, because the obvious structure is wrong — `PageLinks` builds its space-id map per page and `Options` is built per page, so a user map written that way would re-resolve the same twelve people on every page of a 200-page export. It **remembers misses**, or a page mentioning deactivated people costs a request each, every page, to learn the same failures. Not persisted to disk, and the reason is **L2**: output must depend only on the files on disk, not on what a cache happens to hold. `MentionWarnings` is the forward direction's use of it, and sharing the cache is what makes that warning affordable — publishing needs no display names at all, so that lookup exists purely to report an id that names nobody. `PageLinks` resolves a space id **once per space key**, not once per link, and refuses to search site-wide when it can't scope a title to a space — a same-titled page in the wrong space is a wrong answer, which is worse than the passthrough a miss produces. - `internal/attachfile` — `Resolve` (where an attachment goes under a destination root, **including the traversal clamp**) and `Write` (download it there, honoring force/dry-run). Shared by `attachment-download` and `export`; the clamp must never exist in two copies. @@ -84,9 +84,9 @@ Module `github.com/mozilla/markfluence` (`go 1.25`). `main.go` is a shim to `cmd - `internal/pageref` — `Resolve`, the single page-argument resolver: a numeric id, a Confluence page **or folder** URL (`pagePathRE` matches both `/pages/` and `/folder/`, since `children` takes a folder and a folder URL is what a browser hands you — the id is all it returns, so a command that can only use a page reports its own not-found), or a `.md` file that **declares** a `page_id` — in its own frontmatter or in a `pages:` entry for it (#139), resolved through `pagemeta` after discovering the root from the file's own directory (stat'd first, so `123.md` is a file). Every command taking a page uses it, which is why the manifest lookup lives here rather than at the seven call sites: without it the page argument meant one thing to `update` and another to `page-info`/`read`/`children`/`export`/`attachment-*`, so a file `update` could publish could not be named to any of them. A **disagreement** between the two locations is fatal here — it is the question being asked — while a **malformed** `markfluence.yaml` is not: a project file this resolver never consults must not make `page-info 123` fail, and the commands that bound reads by the root report it themselves. `message.go` also owns the wording for the two ways a *frontmatter* `page_id` is wrong — `NotFoundMessage` (caller supplies the remedy, which differs per command) and `NotNumericMessage` — because `create` and `update` report both and `check` reports the non-numeric one, and a reader should recognize the same problem across all of them. They return strings, not errors: `create` wraps the text in its typed `pageIDFailure` (which also carries the `--json` fields), the others want a plain error. Anything checking a `page_id` before a request uses `IsDigits`, since the API answers a non-numeric id with a 400 whose body says nothing useful. - `internal/client` — `ConfluenceClient` over `net/http` with basic auth. Built from a `Config` (site URL, cloud ID, username, token) via `New`; it carries **two bases**: `BaseURL()` is where requests go (the gateway when a cloud ID is set) and `SiteURL()` is always the site. Anything a reader sees uses `SiteURL()` — printed page URLs and, critically, the `baseURL` handed to `convert.MdToConfluence`, since rewritten links are published *into* the page. Pages are Confluence **v2**; attachment writes and the user lookup are **v1** (`/wiki/rest/api/...`). A **folder** — the Cloud content type that can parent a page — has its own v2 route, `GetFolderOrNil` against `/wiki/api/v2/folders/{id}`, because every v2 *page* route answers a folder id with 404; enumerating children, if it is ever added, must be v1, since v2 cannot list inside a folder at all and its page-children route silently omits folders ([docs/confluence/folders.md](docs/confluence/folders.md)). Page **status** — the title lozenge — is v1 only and lives in `state.go` (`PageState`/`AvailableStates`/`SetPageState`, plus `StateVocabulary`, which keeps the space's statuses and the caller's own custom ones in separate fields because only the first are valid for a file); v2 carries no state field on a page in any form and there is no expansion that adds one, so a lozenge is one extra request per page, always. `AvailableStates` must be asked about the page the status is going on — its answer varies by caller *and* page, and it needs edit permission on that page. There is deliberately **no `ClearPageState`**: the `DELETE` route exists, but nothing can reach it until a clearing spelling does, and an unused write method is a loaded gun. Typed `HTTPError`, per-attempt context timeouts, centralized retry/backoff in `send`. `HTTPError.Error()` appends a **hint** for the three auth failures whose status misleads, matched on the response *body* rather than deduced from the status and always **appended** to it, never replacing it. The one that matters: **a rejected credential is a 404 on every v2 route**, so a revoked token used to make `read` answer `page ... not found` about a page that exists. `RejectedCredential` tells it apart by the fact that every genuine v2 404 *names* what it could not find and the auth one does not, `notFound` gates the three `…OrNil` helpers on it so they stop reading it as "absent", and `jsonout.CodeFor` checks it before the status switch so `--json` reports `AUTH` rather than `NOT_FOUND`. **Two error types on the request path, and one predicate for them**: an `*HTTPError` once a response has a status, an unexported `requestError` when there is none (a transport failure, a request that would not build, a body that would not decode), and `FromRequest` answers whether an error is either. That is what lets a caller tell a server failure from a local one — `jsonout.CodeOr(err, fallback)` is the whole point of it, since `CodeFor` alone reports every non-`HTTPError` as `NETWORK` and so turns `no title given` into a network problem (#133). The rule is deliberately scoped to the request: `DownloadAttachment` writing to the caller's writer, `uploadAttachment` opening the caller's file, and `Resolve` reading the environment stay untyped, because tagging them would misreport an unreadable file as a network failure. The wrapper carries no message of its own, so `Error()` is the inner text verbatim and nothing a reader sees changed. A 403 that is not one of the two measured credential phrasings gets no hint, because that is what a genuine permission denial looks like ([docs/confluence/api.md](docs/confluence/api.md#scopes)). **Retry rules**: 429 for any method; 502/503/504 for idempotent methods; **any other 5xx only when the response carries `Retry-After`** — that is how a 500 becomes retryable, and it is why `parseRetryAfter` reports the header's *presence* apart from its delay (`Retry-After: 0` means "retry now", not "no header"). The exponential delay is jittered, a server-supplied `Retry-After` never is. Decisions go to a package-level hook (`SetRetryLogger`, set once in `root.go` beside `ui.SetDebug`) and fire whichever way they went, because `internal/client` prints nothing and a silent twelve-minute retry storm is otherwise indistinguishable from a hang. **A versioned PUT is not as idempotent as its method**: `SetContentProperty` retry-once on top (recovers a lost create-POST response) and `UpdatePage`'s `updateLanded` both exist for the same reason — a write whose response was lost gets re-sent, and the re-sent version is refused. `updateLanded` requires version *and* title *and* body to match what was sent, since a concurrent edit could have produced the version alone and claiming success over someone else's content is worse than a false failure ([docs/confluence/api.md](docs/confluence/api.md)). `SyncAttachments` (skip/update by a SHA-256 recorded in the attachment's comment, alongside the source path so `read` recovers image paths exactly; only the current comment form is parsed — an attachment stamped by a markfluence predating a comment-format change reads as unmanaged and is re-uploaded once, the same as any hand-uploaded file — except that a *recorded path disagreeing with the local source* is an update even when the checksum matches, so a mangled path repairs itself instead of surviving every later publish; a comment with no source recorded at all is not a disagreement. Every text part of the upload form must go through `writeTextField`, never `multipart.Writer.WriteField`, which emits no charset and gets decoded as Latin-1), `_links.next` pagination. **Four pagination schemes, and picking the wrong one truncates silently.** v1 *child/attachment* collections page through the generic `listV1` helper by `start`/`limit` offset, never `_links.next` (absent when the results fit one page, so it cannot terminate a loop); `ListAttachments`, `ListChildPages`, and `ListChildFolders` all go through it. v2 collections page through `listV2` by the cursor in `_links.next` (whose loop is `walkV2`, shared rather than copied so a counting caller can stream — a second implementation of v2 paging is how one of them comes to terminate on a short page), which is a `/wiki`-prefixed absolute path `resolveNext` handles unchanged; `ListContentProperties` and `SearchPagesByTitle` share it. **`/wiki/rest/api/search` is neither**: it ignores `start` outright, its `next` is context-relative so it needs the `/wiki` prefix `resolveNext` does not add, a short page does *not* mean the end, and `totalSize` can be nonzero against an empty `results` — so `searchCQL` terminates only on a missing `next` and nothing may branch on `totalSize` ([docs/confluence/search.md](docs/confluence/search.md)). `searchCQLBounded` adds a row bound under it (`SearchCQL` is that call with no bound, which is why `find` is unaffected): it asks for `max+1` and reports the surplus as `more`, since `totalSize` cannot supply a count. **`/wiki/rest/api/space` is a fifth, and the one that punishes the obvious choice**: it pages by `start`/`limit` offset exactly as the child collections do, and **a short page is not the end** — asked for 250 from `start=0` it answered 200, and `start=200` then answered 250 more, against 525 spaces. `listV1` stops on that short page, so `WalkSpaceOperations` has its own loop terminating on an **empty** page, advancing by rows *returned* rather than by the limit asked for, bounded by `maxSpacePages` since an empty page is the only end signal offset paging has here. The first version of the probe that found this trusted the short page and reported 200 spaces with total confidence ([users.md](docs/confluence/users.md)). **`/wiki/rest/api/search/user` is a fourth**, and the one that looks most like an existing scheme while not being it: it pages by `start`/`limit` offset exactly as `listV1` does, so `listV1` is the obvious home for it and is a trap — the route **caps a page at 100 rows while echoing back whatever limit was requested** (101, 250 and 500 all answer 100), and `listV1` asks for `v1PageSize = 250` and reads a short page as the end of the collection, so it would truncate every result set past 100 with no error at all. `SearchUsers` lives in its own `users.go` with `userPageSize = 100` and the measurement beside it for that reason, and `TestPageCapDoesNotTruncate` is the regression. It also carries `maxUserPages`, `searchCQLBounded`'s guard for the same hazard reached a different way: a short page is the *only* end signal offset paging here has, so a server that clamped `start` — or ignored it the way `/wiki/rest/api/search` ignores it outright — would return a full page forever and an unbounded walk would collect rows until it ran out of memory. Its `totalSize` is a *third* kind of wrong: not absent like v1's and not an estimate like `/search`'s, but the row count of the page just fetched, so `limit=3` answers 3 and `limit=500` answers 100 against 304 real matches. `user.go` holds the two identity routes (`CurrentUser`, `UserInfo` — both `read:confluence-user`, both seeing a deactivated account the directory cannot) and `WalkSpaceOperations`; `space.go` holds `GetSpace` (one v1 request answering identity, the caller's own operations, description, labels and the homepage *with its title*), `SpaceStateSettings` (space-admin only, so a 403 that is not a rejected credential is `(nil, nil)` rather than an error) and `WalkSpacePages`. Both space routes decode the space `id` as a `json.Number`: v1 reports it as a **number** where every v2 route reports a string, and `homepage.id` in the same response is a string. Full text goes through `SearchText`/`SearchRawCQL`, which return the cleaned `SearchMatch` the way `FindByTitle` returns `TitleMatch` — and **every field of a match comes from the row's `content` object**, because the row-level `title` is HTML-escaped *and* wrapped in `@@@hl@@@` markers where `content.title` is neither. The `excerpt` exists only at row level, so `cleanExcerpt` strips those markers, unescapes once, and collapses to one line — in the client, so the human and `--json` paths cannot disagree about it. `excerpt=highlight` is passed explicitly and **re-attached when following the cursor** (the `next` link carries `cql` and `limit` but not `excerpt`, and `doJSON` appends params with a bare `?`); an unrecognized value there yields an empty excerpt with a 200, so a rename by Atlassian degrades to no excerpts rather than an error. A row with no `content` object is skipped and **counted** — `type = space` answers with hundreds of them, and a silent skip would report a successful empty result. A bare v1 child row already carries `webui`, `status`, and `extensions.position`, so child listing needs no `expand`. `DownloadAttachment` goes through `send` (inheriting retry/backoff) against `_links.download`; **never** add a `CheckRedirect` that forwards headers — it would leak site credentials to Atlassian's media host, which neither needs nor wants them. `config.go` holds `Resolve` and the `.env` reader, plus the **`.env` permission warning** (#136): a `.env` reachable by anyone but its owner (`mode.Perm()&0o077`) *and* containing `CONFLUENCE_TOKEN` earns a warning naming the file, its mode, and the `chmod`. Both halves matter — a `.env` holding only the URL and username leaks nothing, and a warning that fires on a file with no secret in it is how one becomes something people scroll past. It stats rather than lstats (a link's own `0777` would cry wolf over a `0600` target), lives in `loadDotenv` because that is the one function both the discovered `.env` and `--env-file` pass through, and reaches the reader through `SetSecurityWarner` for the same reason `SetRetryLogger` exists — wired to `cmd/root.go`'s `reportSecurityWarning`, which prints it (human mode) *and* records it via `jsonout.AddWarning`, since stderr under `--json` is a schema-validated document with no room for a stray line. A group/world-*writable* `.env` with no token in it is knowingly **not** covered, though `CONFLUENCE_URL` resolves from there too and rewriting it would redirect the token: see #136. Why each of these is shaped this way, with the evidence: [docs/confluence/api.md](docs/confluence/api.md) and [attachments.md](docs/confluence/attachments.md). - `internal/convert` — the converter (the crux). `MdToConfluence(md *frontmatter.MarkdownFile, root *project.Root, index *linkindex.Index, baseURL, spaceKey string) (*ConfluencePage, error)`. `root` bounds which images and parent references may be read (S1/S2) and is what an image's recorded `Source` is relative to; `index` is the tree-wide link/anchor index for `root` (`internal/linkindex.Build`), built once and shared across every file converted under it rather than rebuilt per conversion — both are discovered/built by the caller (`internal/project`/`internal/linkindex`), which is why this package stays client-free. It parses with goldmark (GFM) and renders through a custom `storageRenderer` registered at priority 100 (below the default HTML=1000 and table=500 renderers) that emits Confluence storage format. `shield.go` renames raw `ac:`/`ri:` tags to colon-free sentinels around the goldmark step so pasted storage passes through; `callouts.go` is an AST transformer + blockquote renderer for GitHub alerts; `mention.go` owns the user-mention mapping in both directions (#91): a mention is 80% of all `` usage, and it converts to `[@Display Name](https://home.atlassian.com/people/{accountId})`. Three things decide its shape, each measured rather than reasoned. **The URL is Atlassian Home, not the site** — Confluence's own renderer still emits `{site}/wiki/people/{id}`, which no longer resolves usefully in a browser, so a mention in Markdown names *no site*, needs nothing from configuration, and is therefore recognisable by `check` with no client at all. **Matching is on the path, ignoring host and query**, because several spellings of one target circulate (the Home URL, the modal's `?cloudId=` copy, the `/o/{orgId}` redirect, both Confluence forms, root-relative) and none of `cloudId`/`ref`/the org segment identifies the person — only the id does, and `ri:user` stores nothing else. **The `@` on the link text is the marker**, load-bearing rather than decoration: the URL cannot tell "mention this person" from "link to their profile", so without it anyone writing the second would silently get the first. The account id is **not** pattern-validated (two shapes are live on one instance, so a pattern tight enough for one rejects the other) and `ri:local-id` is never emitted (a mention carrying only the id resolves to the same person, verified via ADF). `MentionMarkdown` is the shared builder for a mention's whole Markdown line, exported because `user-find` (#143) prints exactly it and a second copy there would be a second place to get the `@` marker and name escaping right. `ConfluencePage.Mentions` reports the ids the *forward* direction emitted so the caller can warn about one that names nobody — the `Attachments` arrangement, and necessary because Confluence accepts any id and renders `@Unlicensed user` rather than failing, and the profile URL 200s either way. An unresolvable mention still renders as a link, `[@Unlicensed user](…)` — that wording mirrors Confluence because the only ids reaching it are the ones the page labels that way: a **deactivated account resolves normally** and keeps its name (measured across every mention on a real page — 18 of them, six departed, all 200, returning e.g. `Mark Reid (Deactivated)`), so a departed colleague never takes that branch. Name resolution is `pagedoc.UserCache`, a per-run cross-page cache, and the tri-state is the part to preserve: `client.LookupUser` separates a name from `ErrNoSuchUser` from an unaskable question, `StorageOptions.UserNames` carries that as name / `""` / absent, and only a **confirmed** absence renders the placeholder. Flattening those would write a fabricated name over a real one the moment a VPN dropped mid-export, across a whole tree, into a file that then looks authoritative — which is also why the cache remembers a 404 but not a timeout (one is an answer, the other is not) and why `MentionWarnings` warns only about a confirmed absence; `aclink.go` is the *inverse* direction's one element with enough shape to need its own file — ``, which the editor writes for every internal link and `MdToConfluence` never emits, so nothing in the regression suite covers it. One rule decides its whole mapping: **convert when the Markdown republishes to a link resolving to the same target, pass the storage through when it would not** — so a page link and a space link convert, while a mention and a space link convert, and an attachment link (only images are uploaded, so a relative href would be dead) and an unresolvable target stay raw, which the shield republishes byte-identical. A page target is a **title, never an id**, so `PageLinkTargets` reports what needs resolving and `StorageOptions.PageLinks` carries the answers back. An `ac:anchor` is **percent-encoded** where `confluenceSlug` output is not: decode it before matching a heading, leave it encoded inside a URL. A same-page anchor recovers its heading from the document rather than inverting the slug, which is impossible — `confluenceSlug` turns both a space and a hyphen into `-`. The survey the mapping rests on, and the `xml.HTMLAutoClose` trap that made `` crash the parser outright (#88), are in [docs/confluence/links-and-anchors.md](docs/confluence/links-and-anchors.md); Plain text used as a Markdown link's text goes through `escapeLinkText` (`\`, `[`, `]`), applied to the *raw* sources only — a page title, a space key, an anchor, a display name — and via `inlineTextForLink` to a body whose every descendant is a text node. Never to already-rendered output: an `ac:link-body` holding markup has been converted to Markdown already, and escaping it yields a literal `\*\*bold\*\*`. Both directions are tested, because a fix at either extreme passes one and fails the other. `attachname.go` owns the source-path→attachment-name mapping, which is now the path's **base name** and nothing else (#59/`_plans/029`): the name is the attachment's identity, so an encoded path moved the name every time the file moved and orphaned the old attachment, and the path is recorded in the comment anyway. The mapping is therefore lossy, and what the bijection used to buy is an explicit refusal — two assets in one document whose base names agree return a typed `NameCollisionError` from `MdToConfluence`, which is a *failure* and not a `Broken` entry, since nothing blocks a publish on `Broken`. `check` catches that error and reports it as `Broken` anyway, because there it is a document defect like a dead link rather than a converter failure. A stored name is never interpreted in the other direction either: `sourceFor` reads the recorded path or uses the name verbatim. What names Confluence accepts is in [docs/confluence/attachments.md](docs/confluence/attachments.md); `destination.go` owns the **other** codec, destination↔path (`decodeDestination`/`encodeDestination`), shared by images *and* doc links — a Markdown destination is a URL, so decode inbound (**before** `withinRoot`, or an encoded `..%2F` slips the clamp) and encode outbound in `storage_to_md.go` (or `export` emits Markdown that no longer parses, and `sourceFor`'s absolute-path refusal is undone by the next read); an undecodable destination is a literal `%` in a filename, not an error; the reasoning is in [docs/confluence/links-and-anchors.md](docs/confluence/links-and-anchors.md); `images.go` (resolution stays page-relative like GitHub; the documentation root — cwd — bounds what may be published, and an image above it is `IMAGE BROKEN`), `links.go` (GitHub/Confluence slugs, doc-link + anchor rewriting against `internal/linkindex`'s tree-wide index; `resolveDocKey` resolves a destination to the index's root-relative key and reports `escapes` — a purely lexical check on the *query* side, since the index itself needs no clamp: an escaping key can never be in it, built by walking downward from root). A doc-link target is one of four severities, #42: missing entirely or escaping root is **Broken** (`LINK BROKEN: … (not found|outside the documentation root)`) and replaces the whole `` element — tags and visible text alike — with that literal message, matching `images.go`'s precedent for a missing image (`renderLink` needs a small per-node flag, `linkBrokenText`, since goldmark still invokes a container node's renderer on the matching leaving call regardless of `WalkSkipChildren` on entering, and there is no `` to write in the broken case); existing on disk with no `page_id` yet is unchanged — a **warning**, the normal state of an unpublished tree; a `#fragment` matching no heading on an otherwise-resolving target also **warns**, gated on `linkindex.Index.FileExists` so a missing/escaping target isn't double-reported. `tables.go` (the `` tag, stamped with `data-layout="align-start"` so tables auto-size and left-align — this must stay if column widths are ever emitted, or a `` silently induces a layout; plus cells: an AST transformer consumes a leading `` comment in a cell and `renderTableCell` emits it as `data-highlight-colour`; `storage_to_md.go`'s `cellTexts` reverses this, reading `data-highlight-colour` back into a `bg:` marker (`cellBGNames`, the reverse of `tables.go`'s swatch map — a hex outside the 21 swatches round-trips as the literal hex, and where two names share a hex the British spelling wins, matching Confluence's own `-colour`), and a column's GFM alignment becomes a `

    ` wrapper **inside** the cell — never the `align` attribute the GFM renderer would emit, which is the one form Confluence discards. Only center and right are emitted: Confluence has no explicit left, so `:---` publishes bare and `read` recovers it as `---`. Since alignment is per-paragraph there and per-column in GFM, `columnSeparators` in `storage_to_md.go` takes each column's most common declared alignment (ties to the first seen) and drops the rest. Rows still fall through to the GFM renderer. A multi-line cell is one `

    ` per line — Enter in the editor starts a new `

    `, it does not insert a `
    ` — so `renderCellLines` in `storage_to_md.go` joins sibling `

    ` children with a literal `
    ` rather than nothing: a GFM table row is exactly one physical line, so a real newline isn't an option, and the same substitution catches a bare mid-line `
    ` (Shift+Enter) that would otherwise render as the two-space hard break valid in ordinary block content but not inside a table row. A `