Skip to content

Support use as a Python library - #14

Merged
ewels merged 5 commits into
mainfrom
claude/python-library-packaging-yo7ut7
Aug 9, 2026
Merged

ewels merged 5 commits into
mainfrom
claude/python-library-packaging-yo7ut7

Conversation

@ewels

@ewels ewels commented Aug 9, 2026

Copy link
Copy Markdown
Owner

Makes nf-docs usable as an importable Python package, alongside the CLI. The CLI stays the primary interface and its behaviour is unchanged.

Why this was cheap

The core was already side-effect-clean, so most of the work was packaging and documenting what existed rather than restructuring anything:

  • Every rich call, sys.exit() and print() already lived in cli.py / tailwind.py. Core modules only use logging.
  • Progress was already a decoupled callback protocol, not wired to rich.
  • The test suite already imported the package directly.

The real gaps were that nothing was exported, there was no py.typed, and ~150 lines of reusable output policy were trapped inside the generate click command.

The API

import nf_docs

pipeline = nf_docs.extract("./my_pipeline")          # -> Pipeline
markdown = nf_docs.render(pipeline, "markdown")      # -> str
files    = nf_docs.generate("./my_pipeline", output_format="html", output="site/")  # -> list[Path]

generate() always writes files and returns their paths. The CLI's habit of streaming json/yaml to stdout stays in cli.py, since that's a command-line concern rather than an API one.

Changes

  • nf_docs/api.py — the three functions above, wrapping PipelineExtractor and the renderers.
  • nf_docs/output.py — source and output path policy (format aliases, single-module auto-detection, default filenames and directories), lifted out of cli.py so both entry points share one implementation. cli.py also extracts via api.extract(), so the two stay in sync by construction. It imports nothing from the rest of the package.
  • nf_docs/__init__.py — re-exports the facade, models, renderers, progress types and exceptions, with an explicit __all__.
  • py.typed — so the existing type hints reach downstream users. Verified present in the built wheel.
  • PipelineExtractor(config=...) — library callers get NfDocsConfig() defaults rather than the user's ~/.config/nf-docs/config.yaml, which keeps programmatic use reproducible. The CLI passes load_config() explicitly to preserve its behaviour. The get_config()/reset_config() globals had zero callers left afterwards and were removed.
  • docs/python-api.md — new page, added to the nav.

Verification

  • 326 tests pass. tests/test_cli.py is unmodified — that's the guardrail for "the CLI didn't move".
  • CLI and library output verified byte-identical across all five formats in every output mode: explicit file, explicit directory, default location, stdout, and single-module.
  • ruff check, ruff format --check and ty check clean.
  • Every code example on the new docs page was executed (this caught a real bug: PipelineCache().clear() needs a Path, not a str).

Not verified: the real Language Server path. The sandbox this was developed in blocks GitHub, so the language-server JAR could not be downloaded and no end-to-end extraction against a live LSP was possible. Everything upstream and downstream of that call was exercised with the LSP stubbed. Worth a manual nf-docs generate against a real pipeline before merging.

Known issue, not addressed here

Most NfDocsConfig fields are documented but never consulted during extraction — include_hidden_params, max_readme_length, strip_readme_badges, exclude_patterns, ignore_input_prefixes and default_format. The behaviours mostly exist but are hardcoded (badge-stripping always runs and never reads strip_readme_badges; hidden is parsed and stored but nothing filters on it; the LSP excludes are hardcoded). Only ignore_config_prefixes takes effect. tests/test_config.py only round-trips the values, so the suite doesn't catch it.

The new docs page documents just ignore_config_prefixes so the library doesn't advertise inert options. Wiring up the rest is a separate change.

🤖 Generated with Claude Code

https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H


Generated by Claude Code

claude added 5 commits August 8, 2026 19:02
nf-docs stays a CLI tool first, but the core was already side-effect-clean
(all rich output and sys.exit calls live in cli.py), so exposing a supported
library API was mostly a matter of packaging and documenting what exists.

Adds a small facade over the existing classes:

  import nf_docs
  pipeline = nf_docs.extract("./my_pipeline")
  markdown = nf_docs.render(pipeline, "markdown")
  files    = nf_docs.generate("./my_pipeline", output_format="html", output="site/")

extract() and render() mirror PipelineExtractor/get_renderer; generate()
always writes files and returns their paths. The CLI's habit of streaming
json/yaml to stdout stays in cli.py, since that's a command-line concern.

Output path policy (format aliases, single-module auto-detection, default
filenames) moves out of cli.py into nf_docs.output so both entry points share
one implementation. cli.py becomes a presentation layer over it and is
otherwise unchanged - tests/test_cli.py passes untouched, and CLI and library
output was verified byte-identical across all five formats.

PipelineExtractor now takes an optional config= argument. Library callers get
NfDocsConfig() defaults rather than the user's ~/.config/nf-docs/config.yaml,
so programmatic use is reproducible; the CLI passes load_config() explicitly
to keep its behaviour. The global get_config()/reset_config() remain.

Also adds a py.typed marker, an explicit __all__ covering the facade, models,
renderers, progress types and exceptions, and a "Python API" docs page.

Note: most NfDocsConfig fields (include_hidden_params, max_readme_length,
strip_readme_badges, exclude_patterns, ignore_input_prefixes, default_format)
are not actually consulted during extraction - the behaviours exist but are
hardcoded. Only ignore_config_prefixes takes effect. The docs page documents
just that one so the library doesn't advertise inert options; wiring up the
rest is left for a separate change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
rich is already an nf-docs dependency, so the progress-bar example costs
readers no extra install. Passing None for completed/total leaves rich's
task values alone, so one callback covers both the countable and
indeterminate extraction phases.

Also removed em dashes, tricolons and passive constructions from the
Python API page.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
Cleanup pass over the library API change. No behaviour change - CLI and
library output verified byte-identical across every format and output mode.

output.py now owns all source/output path decisions, not just the
single-file case:

- resolve_data_file_path() and default_output_dir() replace the
  <pipeline>/docs and pipeline.{ext} rules that were computed separately
  in both api.generate() and cli.generate()
- format_extension() replaces a third copy of the format->extension
  mapping ('json' if fmt == 'json' else 'yaml'), which silently labelled
  any future non-json format .yaml
- cli.py now uses DIRECTORY_FORMATS instead of spelling
  ("markdown", "html", "table") and its inverse inline four times
- resolve_single_file_output() split into a pure resolve_single_file_path()
  plus a streams_by_default() predicate. It no longer renders as a side
  effect, so the "None means stdout" sentinel is gone along with the
  special case in api.generate() that undid it. output.py now imports
  nothing from the package, keeping it a cheap leaf module.

Format aliases had drifted into two tables: get_renderer() had its own
"md" entry while output.FORMAT_ALIASES had another, and api.render()
never normalised at all. get_renderer() now resolves through
normalize_format(), so aliases live in one place.

cli.generate() extracts via api.extract() rather than rebuilding the
same workspace/target_file derivation, and api.generate() resolves the
source once via a shared _extract_resolved() instead of resolving twice.

Deleted get_config()/reset_config() and the module-level _config global.
Injecting config at PipelineExtractor left them with zero callers in
src/, and leaving an ambient accessor around invites reintroducing the
non-hermetic behaviour the API promises against.

__init__.py now resolves everything except the data models lazily via
PEP 562 __getattr__. Re-exporting eagerly had pulled httpx, jinja2,
markdown, pygments and yaml into every `import nf_docs`:

  import nf_docs   264 ms -> 90 ms   (sys.modules 386 -> 151)

CLI startup is unaffected either way, since cli.py already imported that
whole tree; the cost fell entirely on library consumers who only wanted
the models or the version.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
The lazy PEP 562 __getattr__ saved ~170 ms on `import nf_docs`, but that
is not worth the maintenance cost: every new public name had to be
registered in both _LAZY_EXPORTS and a TYPE_CHECKING block, and the
mechanism is non-obvious to anyone adding one.

nf-docs is not a library you import in a hot path. It is imported to
parse pipeline docs, which takes seconds waiting on the Language Server,
so the import time is noise. A plain list of imports is easier to read
and harder to get wrong.

The rest of the cleanup pass stands - only __init__.py reverts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
Four issues, all introduced by this PR, each reproduced before and after.

The cache ignored configuration. Its key was version + workspace path +
file contents, so a CLI run (which loads the user's config file) and a
library call (which uses defaults) shared an entry for the same unchanged
pipeline and returned each other's config-filtered results. Reproduced in
both directions. NfDocsConfig.cache_key() now folds into the content
hash, and PipelineCache.get/set take the config.

Folding into the existing hash rather than adding a filename component
keeps _cleanup_old_caches working unchanged. The trade-off is that
alternating two configs on one pipeline evicts each entry in turn instead
of keeping both - a re-extraction, where sharing an entry gave wrong
answers.

load_config() had moved outside the CLI's try block, so it no longer sat
inside the extractor's `except Exception: logger.warning(...)`. A config
file holding valid YAML of the wrong shape (a top-level list) escaped as
an uncaught AttributeError and dumped a traceback; `inspect` was
unaffected, only `generate` regressed. load_config() now checks the
parsed value is a mapping, which fixes it for every caller rather than
just that call site.

resolve_source() didn't check the path exists. The CLI gets that from
click's exists=True, but the new public API had nothing, so
extract("./my_pipline") returned an empty Pipeline named after the typo
with no error at all. (It did not start the Language Server - there's an
early return on an empty file scan.)

Tests wrote into the developer's real ~/.cache/nf-docs/: 32 entries per
pytest run, containing empty extractions because the Language Server is
mocked. An autouse conftest fixture now points XDG_CACHE_HOME and
XDG_CONFIG_HOME at tmp dirs.

Regression tests added for all four. CLI and library output re-verified
byte-identical across every format and output mode.

Not addressed: resolve_single_file_path() treats a not-yet-created output
directory as a file, because is_dir() is false for it. That predates this
PR - the same check was in the original _resolve_single_file_output - so
it belongs in its own change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
@ewels
ewels merged commit e7d5b1c into main Aug 9, 2026
8 checks passed
@ewels
ewels deleted the claude/python-library-packaging-yo7ut7 branch August 9, 2026 09:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants