Support use as a Python library - #14
Merged
Merged
Conversation
nf-docs stays a CLI tool first, but the core was already side-effect-clean
(all rich output and sys.exit calls live in cli.py), so exposing a supported
library API was mostly a matter of packaging and documenting what exists.
Adds a small facade over the existing classes:
import nf_docs
pipeline = nf_docs.extract("./my_pipeline")
markdown = nf_docs.render(pipeline, "markdown")
files = nf_docs.generate("./my_pipeline", output_format="html", output="site/")
extract() and render() mirror PipelineExtractor/get_renderer; generate()
always writes files and returns their paths. The CLI's habit of streaming
json/yaml to stdout stays in cli.py, since that's a command-line concern.
Output path policy (format aliases, single-module auto-detection, default
filenames) moves out of cli.py into nf_docs.output so both entry points share
one implementation. cli.py becomes a presentation layer over it and is
otherwise unchanged - tests/test_cli.py passes untouched, and CLI and library
output was verified byte-identical across all five formats.
PipelineExtractor now takes an optional config= argument. Library callers get
NfDocsConfig() defaults rather than the user's ~/.config/nf-docs/config.yaml,
so programmatic use is reproducible; the CLI passes load_config() explicitly
to keep its behaviour. The global get_config()/reset_config() remain.
Also adds a py.typed marker, an explicit __all__ covering the facade, models,
renderers, progress types and exceptions, and a "Python API" docs page.
Note: most NfDocsConfig fields (include_hidden_params, max_readme_length,
strip_readme_badges, exclude_patterns, ignore_input_prefixes, default_format)
are not actually consulted during extraction - the behaviours exist but are
hardcoded. Only ignore_config_prefixes takes effect. The docs page documents
just that one so the library doesn't advertise inert options; wiring up the
rest is left for a separate change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
rich is already an nf-docs dependency, so the progress-bar example costs readers no extra install. Passing None for completed/total leaves rich's task values alone, so one callback covers both the countable and indeterminate extraction phases. Also removed em dashes, tricolons and passive constructions from the Python API page. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
Cleanup pass over the library API change. No behaviour change - CLI and
library output verified byte-identical across every format and output mode.
output.py now owns all source/output path decisions, not just the
single-file case:
- resolve_data_file_path() and default_output_dir() replace the
<pipeline>/docs and pipeline.{ext} rules that were computed separately
in both api.generate() and cli.generate()
- format_extension() replaces a third copy of the format->extension
mapping ('json' if fmt == 'json' else 'yaml'), which silently labelled
any future non-json format .yaml
- cli.py now uses DIRECTORY_FORMATS instead of spelling
("markdown", "html", "table") and its inverse inline four times
- resolve_single_file_output() split into a pure resolve_single_file_path()
plus a streams_by_default() predicate. It no longer renders as a side
effect, so the "None means stdout" sentinel is gone along with the
special case in api.generate() that undid it. output.py now imports
nothing from the package, keeping it a cheap leaf module.
Format aliases had drifted into two tables: get_renderer() had its own
"md" entry while output.FORMAT_ALIASES had another, and api.render()
never normalised at all. get_renderer() now resolves through
normalize_format(), so aliases live in one place.
cli.generate() extracts via api.extract() rather than rebuilding the
same workspace/target_file derivation, and api.generate() resolves the
source once via a shared _extract_resolved() instead of resolving twice.
Deleted get_config()/reset_config() and the module-level _config global.
Injecting config at PipelineExtractor left them with zero callers in
src/, and leaving an ambient accessor around invites reintroducing the
non-hermetic behaviour the API promises against.
__init__.py now resolves everything except the data models lazily via
PEP 562 __getattr__. Re-exporting eagerly had pulled httpx, jinja2,
markdown, pygments and yaml into every `import nf_docs`:
import nf_docs 264 ms -> 90 ms (sys.modules 386 -> 151)
CLI startup is unaffected either way, since cli.py already imported that
whole tree; the cost fell entirely on library consumers who only wanted
the models or the version.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
The lazy PEP 562 __getattr__ saved ~170 ms on `import nf_docs`, but that is not worth the maintenance cost: every new public name had to be registered in both _LAZY_EXPORTS and a TYPE_CHECKING block, and the mechanism is non-obvious to anyone adding one. nf-docs is not a library you import in a hot path. It is imported to parse pipeline docs, which takes seconds waiting on the Language Server, so the import time is noise. A plain list of imports is easier to read and harder to get wrong. The rest of the cleanup pass stands - only __init__.py reverts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
Four issues, all introduced by this PR, each reproduced before and after.
The cache ignored configuration. Its key was version + workspace path +
file contents, so a CLI run (which loads the user's config file) and a
library call (which uses defaults) shared an entry for the same unchanged
pipeline and returned each other's config-filtered results. Reproduced in
both directions. NfDocsConfig.cache_key() now folds into the content
hash, and PipelineCache.get/set take the config.
Folding into the existing hash rather than adding a filename component
keeps _cleanup_old_caches working unchanged. The trade-off is that
alternating two configs on one pipeline evicts each entry in turn instead
of keeping both - a re-extraction, where sharing an entry gave wrong
answers.
load_config() had moved outside the CLI's try block, so it no longer sat
inside the extractor's `except Exception: logger.warning(...)`. A config
file holding valid YAML of the wrong shape (a top-level list) escaped as
an uncaught AttributeError and dumped a traceback; `inspect` was
unaffected, only `generate` regressed. load_config() now checks the
parsed value is a mapping, which fixes it for every caller rather than
just that call site.
resolve_source() didn't check the path exists. The CLI gets that from
click's exists=True, but the new public API had nothing, so
extract("./my_pipline") returned an empty Pipeline named after the typo
with no error at all. (It did not start the Language Server - there's an
early return on an empty file scan.)
Tests wrote into the developer's real ~/.cache/nf-docs/: 32 entries per
pytest run, containing empty extractions because the Language Server is
mocked. An autouse conftest fixture now points XDG_CACHE_HOME and
XDG_CONFIG_HOME at tmp dirs.
Regression tests added for all four. CLI and library output re-verified
byte-identical across every format and output mode.
Not addressed: resolve_single_file_path() treats a not-yet-created output
directory as a file, because is_dir() is false for it. That predates this
PR - the same check was in the original _resolve_single_file_output - so
it belongs in its own change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
This was referenced Aug 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Makes
nf-docsusable as an importable Python package, alongside the CLI. The CLI stays the primary interface and its behaviour is unchanged.Why this was cheap
The core was already side-effect-clean, so most of the work was packaging and documenting what existed rather than restructuring anything:
richcall,sys.exit()andprint()already lived incli.py/tailwind.py. Core modules only uselogging.rich.The real gaps were that nothing was exported, there was no
py.typed, and ~150 lines of reusable output policy were trapped inside thegenerateclick command.The API
generate()always writes files and returns their paths. The CLI's habit of streaming json/yaml to stdout stays incli.py, since that's a command-line concern rather than an API one.Changes
nf_docs/api.py— the three functions above, wrappingPipelineExtractorand the renderers.nf_docs/output.py— source and output path policy (format aliases, single-module auto-detection, default filenames and directories), lifted out ofcli.pyso both entry points share one implementation.cli.pyalso extracts viaapi.extract(), so the two stay in sync by construction. It imports nothing from the rest of the package.nf_docs/__init__.py— re-exports the facade, models, renderers, progress types and exceptions, with an explicit__all__.py.typed— so the existing type hints reach downstream users. Verified present in the built wheel.PipelineExtractor(config=...)— library callers getNfDocsConfig()defaults rather than the user's~/.config/nf-docs/config.yaml, which keeps programmatic use reproducible. The CLI passesload_config()explicitly to preserve its behaviour. Theget_config()/reset_config()globals had zero callers left afterwards and were removed.docs/python-api.md— new page, added to the nav.Verification
tests/test_cli.pyis unmodified — that's the guardrail for "the CLI didn't move".ruff check,ruff format --checkandty checkclean.PipelineCache().clear()needs aPath, not astr).Not verified: the real Language Server path. The sandbox this was developed in blocks GitHub, so the language-server JAR could not be downloaded and no end-to-end extraction against a live LSP was possible. Everything upstream and downstream of that call was exercised with the LSP stubbed. Worth a manual
nf-docs generateagainst a real pipeline before merging.Known issue, not addressed here
Most
NfDocsConfigfields are documented but never consulted during extraction —include_hidden_params,max_readme_length,strip_readme_badges,exclude_patterns,ignore_input_prefixesanddefault_format. The behaviours mostly exist but are hardcoded (badge-stripping always runs and never readsstrip_readme_badges;hiddenis parsed and stored but nothing filters on it; the LSP excludes are hardcoded). Onlyignore_config_prefixestakes effect.tests/test_config.pyonly round-trips the values, so the suite doesn't catch it.The new docs page documents just
ignore_config_prefixesso the library doesn't advertise inert options. Wiring up the rest is a separate change.🤖 Generated with Claude Code
https://claude.ai/code/session_013qwciLzAxtZMwu3uWnA24H
Generated by Claude Code