From unfamiliar XML to analysis-ready R tables.
xmlrectr is an ambitious attempt at a universal XML rectangler for R.
It is not trying to be a universal XML parser. Excellent XML parsers already exist. The problem addressed here is different: how do you turn hierarchical XML into a table or tibble that is genuinely useful for analysis with the tidyverse, base R, Arrow, and other tools built around tabular data?
That problem becomes especially difficult when you are simply handed an XML file and know little or nothing about it. You may have no schema, no documentation, no knowledge of the XML vocabulary, and no predefined extraction rules. xmlrectr can inspect the actual document structure and, in many cases, construct a workable analyst-friendly tibble automatically.
When you do know the XML structure, have documentation or an XSD, or know exactly what you want to expose, the package becomes more explicit rather than less useful. You can review structural proposals, provide row and identifier choices, select and rename fields, use schema evidence, save reusable profiles, and obtain a rectangle closely aligned with your analytical purpose.
The package is deliberately generic. It contains no special rules for MARC, FHIR, EAD, ONIX, GPX, UBL, or any other XML vocabulary.
At present, xmlrectr can be installed from GitHub.
The recommended method is pak:
install.packages("pak")
pak::pak("larry77/xmlrectr")Alternatively:
install.packages("remotes")
remotes::install_github("larry77/xmlrectr")Because xmlrectr contains native C code, the GitHub version is compiled from source and requires a working build toolchain and libxml2 development files.
Windows build requirements
Windows is a first-class supported platform. Use the Rtools version matching your R installation and install it in the default location.
For the currently supported R series:
- R 4.6.x: Rtools45
- R 4.5.x: Rtools45
- R 4.4.x: Rtools44
With a normal R/Rtools installation, manual PATH configuration should not usually be necessary.
If compilation fails, first check the toolchain:
install.packages("pkgbuild")
pkgbuild::has_build_tools(debug = TRUE)Avoid mixing Rtools with unrelated MSYS2/MinGW/Strawberry Perl toolchains on PATH, as this can produce difficult-to-diagnose linking problems.
The GitHub Actions workflow compiles and tests the package on windows-latest.
Linux build requirements
On Debian/Ubuntu:
sudo apt install libxml2-dev pkg-configThen install xmlrectr from R using pak or remotes.
macOS build requirements
With Homebrew:
brew install libxml2 pkg-config
export PKG_CONFIG_PATH="$(brew --prefix libxml2)/lib/pkgconfig:$PKG_CONFIG_PATH"Then install xmlrectr from R.
For a local checkout:
pak::local_install(".")XML is hierarchical. Analytical data is usually rectangular.
A naive "flatten everything" strategy can easily:
- lose repeated values;
- multiply independent repeated branches into accidental Cartesian products;
- confuse elements that share names but belong to different namespaces;
- collapse document order that carries information;
- propagate parent text or context to the wrong descendants;
- invent associations that are not actually present in the source;
- create list-columns or highly irregular structures that are awkward for ordinary R analysis.
There is also an unavoidable conceptual limit: an arbitrary XML document does not have one mathematically unique tabular interpretation. Domain knowledge can always improve a rectangle.
xmlrectr therefore does not claim to infer the intended semantics of every XML document. Instead, it aims to make a useful, conservative structural interpretation when knowledge is scarce, while exposing progressively more control when the user knows more.
The core design is:
unknown XML
|
+--> automatic analyst-oriented rectangle
|
+--> structural proposal
|
v
human review
|
v
reusable profile
|
v
compiled rectangle
XSD information can contribute evidence, but it is advisory rather than blindly authoritative.
A major goal of xmlrectr is to automate both the analytical problem and the computational problem, while keeping both layers tunable.
For exploration, rectangle_xml_analyst() can start from the XML file itself:
library(xmlrectr)
file <- system.file("extdata", "orders.xml", package = "xmlrectr")
analyst <- rectangle_xml_analyst(file)
analystThe result is one self-contained tibble with:
- universal
xml_*provenance and entity columns; - atomic analyst-facing columns;
- no list-columns;
- explicit entity and parent-entity identities;
- repeated branches kept as separate observations rather than silently multiplied;
- namespace-aware source information;
- inferred scalar values where the structure supports them.
This is the deliberately ambitious part of the package: starting from unfamiliar XML and attempting to produce something an R analyst can immediately inspect and work with.
For one-off exploration, this may be all you need.
For repeated or production use, once you understand the structure you will usually want to move to an explicit profile so that the intended rectangle becomes a stable, reviewable contract.
Once a rectangle is defined, execution has its own set of choices: sequential or parallel processing, worker count, chunk size, task size, memory ownership, and whether a large document should be streamed instead of fully materialised.
Those choices are intentionally separated from the analytical meaning of the rectangle.
For the normal rectangling APIs, one argument can ask xmlrectr to choose whether parallel execution is worthwhile:
out <- rectangle_xml(
file,
profile,
parallel = "auto"
)Automatic execution planning is based on the observed structural workload and available resources, not on XML vocabulary names, filenames, or rules tuned to the validation corpus. The package does not apply one fixed worker/chunk recipe to every document.
Advanced controls remain available, but they are optional. The point of the architecture is that you should not need to turn every screw before getting useful work done on a large, nested or unfamiliar XML file.
xmlrectr is designed to support a continuum of prior knowledge.
Start with the automatic analyst table:
analyst <- rectangle_xml_analyst(file)This is the quickest route from an unfamiliar document to an R tibble.
Ask the package for a proposal:
proposal <- propose_xml_profile(file)
review_xml_proposal(proposal, "rows")
review_xml_proposal(proposal, "ids")
review_xml_proposal(proposal, "fields")The proposal is evidence, not an executable command. It lets the package inspect the XML and suggest plausible structural choices without pretending that software can know your analytical intent.
Define it explicitly:
profile <- xml_profile(
rows = "order",
id = "id"
)
out <- rectangle_xml(file, profile)A profile can also select and rename fields, specify types, provide namespace information, and choose the desired layout.
You can save the reviewed decision:
write_xml_profile(profile, "orders-profile.json")
profile2 <- read_xml_profile("orders-profile.json")and reuse it across files belonging to the same XML family.
For repeated processing, compile the profile once:
spec <- compile_xml_profile(profile, file)
out <- rectangle_xml(
file,
spec,
parallel = "auto"
)This is where domain knowledge pays off: when you know the XML structure and the analytical question, the package can produce a rectangle much more closely aligned with your desiderata than any fully automatic method could infer.
If an XSD is available, xmlrectr can use it as additional structural evidence:
xml <- system.file("extdata", "types.xml", package = "xmlrectr")
xsd <- system.file("extdata", "types.xsd", package = "xmlrectr")
proposal <- propose_xml_profile(xml, xsd = xsd)
review_xml_proposal(proposal, "xsd")inspect_xsd() can expose useful declaration, occurrence, required-attribute and scalar-type information.
The important design choice is that XSD is advisory. A schema describes valid document structure, but it does not necessarily tell an analyst what should constitute a row, which repeated structures should become separate entities, or which fields are relevant to a particular analysis.
So the workflow remains:
XML structure + optional XSD + user knowledge
|
v
reviewed profile
|
v
analytical rectangle
The automatic analyst representation is designed as one self-contained atomic table per XML document.
Its structural contract is conservative:
xml_entityandxml_entity_ididentify analytical entities;- entity IDs are nonblank and unique;
- parent IDs refer to entities in the same table;
- repeated branches remain separate rows;
- independent repetitions are not multiplied into accidental Cartesian products;
- analyst columns are atomic rather than list-columns;
- all-
NAanalyst columns are avoided; - source attributes and text remain structurally accounted for;
- namespace information and source paths remain available through
xml_*provenance columns; - derived text is not blindly propagated into descendants;
- order-sensitive sibling structures are preserved rather than being folded into invented key-value associations.
The underlying canonical representation is even more explicit. It records document order, node identity, parentage, depth, namespaces, attributes, text/CDATA and other retained XML node types.
That canonical layer is the loss-aware structural foundation on which higher-level rectangles are built.
The ambition to work with arbitrary XML is useful only if the implementation can cope with XML as it exists in practice: large documents, deep nesting, repeated records, namespaces, irregular branches and substantial structural overhead.
A large part of the development of xmlrectr has therefore focused on algorithmic cost, memory behaviour, workload decomposition and semantic parity.
The package deliberately separates two questions:
- What should the rectangle mean?
- How should that rectangle be computed efficiently on this particular XML document?
Changing the computational engine must never change the first answer.
Performance-critical canonical reading and structural/indexing operations have native C implementations using libxml2.
The R implementation remains the semantic reference. Compiled code is used to accelerate structural bottlenecks, not to introduce a second set of rectangling semantics.
This distinction matters: generic XML rectangling repeatedly performs structural operations for which interpreted R can become expensive on large trees. Moving those bottlenecks into compiled code makes the generic design practical without introducing vocabulary-specific shortcuts.
For XML that should not be represented as one complete in-memory canonical table, xmlrectr provides a streaming path based on complete record subtrees.
batches <- list()
stats <- xml_stream_rectangle(
"large.xml",
spec,
callback = function(batch) {
batches[[length(batches) + 1L]] <<- batch
},
parallel = "auto"
)The streaming architecture has an important correctness boundary:
- SAX parsing and record-boundary detection remain coordinator-side;
- XML is never split at arbitrary byte positions;
- workers receive complete, self-contained record material;
- live parser state and
xml2/libxml external pointers are never passed between processes; - source order and record identities remain controlled by the coordinator;
- duplicate source-ID checks and output callbacks remain coordinator-side.
The point is not merely to "use less RAM". The package tries to keep parser state, record ownership, output ordering and rectangling semantics cleanly separated.
Large results can be written without first collecting the entire rectangle into one R object:
rectangle_xml_csv(
"large.xml",
spec,
output = "large.csv",
parallel = "auto"
)For typed analytical output:
rectangle_xml_parquet(
"large.xml",
spec,
output_dir = "large-parquet",
parallel = "auto"
)Parquet support requires the optional arrow package.
The same high-level execution controls are used by the main in-memory, streaming, CSV and Parquet interfaces. Parallel execution is therefore part of the normal rectangling architecture rather than a separate workflow bolted onto one output format.
Parallelism is an execution choice of the same rectangling operation, not a separate family of user-facing functions.
# Exact sequential path
seq_out <- rectangle_xml(
file,
spec,
parallel = FALSE
)
# Request parallel execution with tuned defaults
par_out <- rectangle_xml(
file,
spec,
parallel = TRUE
)
# Let the engine decide whether parallel work is worthwhile
auto_out <- rectangle_xml(
file,
spec,
parallel = "auto"
)The sequential implementation is the semantic oracle. Parallel execution is required to preserve the same result:
identical(seq_out, par_out)The three modes have deliberately simple meanings:
parallel = FALSEuses the exact sequential path and does not require the parallel stack;parallel = TRUErequests parallel execution using the package's tuned defaults;parallel = "auto"asks the engine to decide, from the observed workload, whether parallel execution is justified.
Lower-level *_parallel() functions exist for compatibility, testing and diagnostics. They are not the intended everyday API.
The parallel architecture follows a few strict rules:
- Only independent complete XML record subtrees are parallelised.
- XML is never divided by arbitrary byte ranges.
- Live
xml2/libxml external pointers are never sent to workers. - SAX parsing and record-boundary detection remain coordinator-side.
- Workers receive self-contained record material.
- Source order, identifiers and sequential semantics must remain exact.
- Coordinator-side responsibilities such as duplicate-ID checks and callbacks remain coordinator-side.
- The caller's pre-existing Future plan is restored after execution.
These constraints are less flashy than a benchmark chart, but they are fundamental to making parallel execution a trustworthy implementation detail rather than a second semantics.
Parallel XML rectangling is not simply a matter of running:
workers <- parallel::detectCores()and dividing a file into equal pieces.
A useful execution plan depends on the actual structural work available:
- how many independent records exist;
- how large their canonical subtrees are;
- whether tasks are coarse enough to amortise process and scheduling overhead;
- how much working memory can safely be admitted at once;
- how many cores are available;
- whether throughput or bounded memory is the more important constraint.
For this reason, xmlrectr does not use filenames, XML vocabulary names, or rules learned specifically from the validation corpus to decide whether to parallelise.
The automatic policy is structural.
Benchmarks showed that increasing the number of workers beyond four can still reduce elapsed time, but efficiency falls and memory pressure increases.
The automatic/default worker policy therefore normally uses up to four workers, while preserving a core for the system where possible.
This is not a hard maximum.
Advanced users can explicitly request more workers:
rectangle_xml(
file,
spec,
parallel = TRUE,
workers = 8
)The default is intended to be a balanced choice, not a claim that four workers are universally optimal.
The current in-memory auto policy asks whether there is enough canonical record work per worker to justify process-level parallelism.
The policy includes a work floor of approximately:
25,000 canonical record nodes per worker
together with sufficient record/subtree structure, approximately:
at least 128 records per worker
or
a median record subtree of at least 1,000 canonical nodes
These are engineering defaults derived from broad structural benchmarking. They are not XML-vocabulary rules, and they should not be read as eternal constants or promises of a particular speedup.
There is also a cheap impossibility check. If the entire canonical table has fewer than roughly:
workers * 25,000
nodes, the workload cannot satisfy the per-worker work floor. Auto mode can then remain sequential immediately instead of performing a more expensive record-span analysis.
This fast path is important. An early version of automatic planning was semantically correct but could make small sequential jobs noticeably slower simply because planning repeated structural work that the sequential path would perform anyway. The current design reuses validation and rejects obviously too-small workloads before doing that extra work.
In other words, auto mode is designed not only to find parallel opportunities, but also to get out of the way when parallelism would be pointless.
xmlrectr retains two parallel execution strategies because throughput and memory pressure are different optimisation problems.
With parallel_chunks, workers own independent vectorised outer chunks.
Conceptually:
chunk 1 ---> worker 1
chunk 2 ---> worker 2
chunk 3 ---> worker 3
chunk 4 ---> worker 4
This strategy is designed primarily to:
- maximise throughput;
- keep scheduling straightforward;
- exploit coarse vectorised work;
- perform well when memory duplication is acceptable.
For in-memory automatic execution, this is generally the preferred strategy.
A useful scheduling model is approximately one active owned outer chunk per worker.
With shared_chunk, one bounded outer chunk is shared and subdivided into multiple vectorised tasks coordinated through mori.
Conceptually:
bounded shared outer chunk
/ | | \
task task task task
| | | |
workers process coarse ranges
Its purpose is to reduce input-memory duplication while still preserving parallel work.
This is particularly attractive for streaming, CSV and Parquet workflows, where bounded memory is part of the reason for choosing the execution mode in the first place.
Automatic streaming/output execution therefore prefers shared_chunk when the required stack is available. If mori is unavailable, automatic execution can fall back to parallel_chunks.
The trade-off is intentional:
parallel_chunksis primarily throughput-oriented;shared_chunkis primarily memory-oriented.
Neither strategy dominates the other on every machine and workload.
The advanced API exposes:
workers
strategy
chunk_records
task_records
The last two parameters solve different problems.
chunk_records controls the size of the outer batch admitted at once.
It therefore influences:
- memory pressure;
- how much vectorised work is available;
- how many independently owned chunks can be active.
For the current streaming defaults:
parallel_chunks:
chunk_records = 512
For shared_chunk:
chunk_records = min(1024, max(256, workers * 256))
These are tuned defaults, not XML-specific rules.
Within a shared outer chunk, task_records controls how finely the work is subdivided for scheduling.
Benchmarks found a fairly broad useful plateau around 4 to 8 tasks per worker. The balanced automatic setting is therefore approximately 4 tasks per worker.
That gives workers enough independent work for load balancing without producing a large number of tiny tasks whose scheduling cost dominates useful computation.
So:
chunk_records -> controls admitted working-set size / memory
task_records -> controls scheduling granularity inside that work
Keeping these concepts separate is especially important for shared_chunk.
Performance results are included here because the defaults were not chosen by intuition alone.
They should nevertheless be interpreted carefully:
Never compare raw elapsed times from different machines as though they belong to one benchmark series.
Absolute timings depend on processor, memory subsystem, operating system, R build, package versions and background load. The useful quantities are same-machine sequential/parallel speedup, worker efficiency, memory behaviour and exact semantic parity.
The following results are representative development measurements on one machine, referred to during development as einstein. They document why the current defaults exist; they are not runtime guarantees.
A representative synthetic workload was approximately 16.266 MiB.
Sequential execution:
193.977 s
Parallel results:
| Workers | parallel_chunks |
Speedup | shared_chunk |
Speedup |
|---|---|---|---|---|
| 2 | 121.585 s | 1.60x | 107.070 s | 1.81x |
| 4 | 63.340 s | 3.06x | 68.869 s | 2.82x |
| 11 | 46.051 s | 4.21x | 55.525 s | 3.49x |
Several conclusions follow.
First, more than four workers can improve elapsed time. Four is therefore not a hard ceiling.
Second, scaling efficiency declines substantially at high worker counts. The extra processes are doing useful work, but the cost of coordination, memory traffic and finite task parallelism becomes increasingly important.
Third, four workers gave a strong compromise between speedup, efficiency and memory pressure. That is why automatic/default worker selection is normally capped there unless the user explicitly chooses otherwise.
On the same approximate 16.266 MiB workload with four workers:
| Strategy | Peak PSS |
|---|---|
parallel_chunks |
about 3618 MiB |
shared_chunk |
about 3142 MiB |
In that experiment, shared_chunk reduced peak proportional set size by roughly 13% and private memory by roughly 15%.
Depending on phase, median PSS could fall by considerably more.
That reduction is meaningful, even though shared_chunk can be slower on some workloads. It is the empirical reason the memory-oriented strategy remains part of the package rather than being removed in favour of the single fastest throughput result.
These measurements do not imply:
- that four workers are always fastest;
- that
shared_chunkalways saves exactly 13% memory; - that a 16 MiB XML file will take anything close to the timings above;
- that file size alone determines whether parallelism helps.
The actual structural workload matters more than the byte size of the XML file.
The final automatic policy was also checked against real XML on the same development machine.
Representative P6.1 smoke timings were:
| XML | Sequential | Auto | Auto decision |
|---|---|---|---|
| UBL | 0.960 s | 0.874 s | sequential |
| Maven | 1.111 s | 1.126 s | sequential |
| EAD | 33.756 s | 30.973 s | parallel |
The important observation for UBL and Maven is not the tiny timing difference. It is that automatic planning did not impose a material penalty on small jobs that should remain sequential.
The EAD document crossed the structural threshold and was sent to the parallel engine.
That particular EAD run should not be used to argue either that parallelism is spectacular or that it is useless. It happened to lie relatively close to the crossover region on that machine. Larger synthetic workloads demonstrated much stronger same-machine speedups.
Most users should stop at:
rectangle_xml(file, spec, parallel = "auto")The following controls exist for benchmarking, unusually constrained machines, or specialist tuning:
rectangle_xml(
file,
spec,
parallel = TRUE,
workers = 4,
strategy = "shared_chunk",
chunk_records = 1024,
task_records = 64
)Useful questions for an advanced tuning exercise are:
- Is the workload CPU-bound or memory-bound?
- Are there enough independent records to keep additional workers busy?
- Are records very uneven in size?
- Does peak memory, rather than elapsed time, constrain the run?
- Are tasks large enough to amortise process scheduling?
- Is the document already close to the sequential/parallel crossover?
Do not tune from filenames or XML vocabulary names.
Do not assume that tiny chunk/task values used in semantic stress tests are production recommendations.
And do not optimise one specific XML file at the expense of the generic structural rules.
For performance work, compare strategies on the same machine and the same XML.
A simple reproducible pattern is:
spec <- compile_xml_profile(profile, file)
t_seq <- system.time(
seq_out <- rectangle_xml(
file,
spec,
parallel = FALSE
)
)
t_auto <- system.time(
auto_out <- rectangle_xml(
file,
spec,
parallel = "auto"
)
)
t_forced <- system.time(
forced_out <- rectangle_xml(
file,
spec,
parallel = TRUE
)
)
stopifnot(
identical(seq_out, auto_out),
identical(seq_out, forced_out)
)
rbind(
sequential = t_seq,
auto = t_auto,
forced_parallel = t_forced
)For serious benchmarking:
- run multiple repetitions;
- compare within-machine speedups rather than raw times from different platforms;
- inspect memory as well as elapsed time when that matters;
- keep semantic checks in the benchmark harness;
- distinguish performance benchmarks from tests whose only purpose is to stress correctness.
Small XML documents may correctly be faster sequentially because process startup and scheduling have a cost. Larger, sufficiently coarse record workloads are where parallel execution can pay off.
This is why parallel = "auto" exists: parallelism is a tool, not a goal in itself.
Some of the harshest parallel tests deliberately used settings that would make poor production defaults.
For example, the 30-file forced-parallel semantic stress run used approximately:
workers = 2
chunk_records = 64
task_records = 8
Those small values were chosen to force many scheduling boundaries and expose correctness problems.
They were not selected for speed.
The result was:
- 30/30 exact in-memory sequential/parallel parity;
- 30/30 exact streaming sequential/parallel parity.
This distinction is important when reading the project's benchmark history. Different experiments answer different questions:
- semantic stress tests ask whether alternative execution paths produce exactly the same result;
- throughput benchmarks ask how much elapsed time can be reduced;
- memory benchmarks ask what the working-set trade-offs are;
- scheduler/chunk experiments tune workload granularity;
- auto-policy experiments ask whether the engine chooses sensibly between sequential and parallel execution.
Conclusions from one class should not be casually transferred to another.
Before conversion into the R package, the frozen validated engine was exercised against a deliberately diverse corpus of 30 real-world XML documents.
The validation included:
- the focused unit/regression suite;
- exact in-memory sequential/forced-parallel parity;
- exact streaming sequential/forced-parallel parity;
- exact parity through the unified
parallel = "auto"interface; - synthetic scaling experiments;
- memory experiments;
- worker-count, chunk-size and task-size tuning.
In the final 30-file automatic-policy run:
- 30/30 documents preserved exact semantic parity;
- 29 smaller workloads stayed sequential;
- only the sufficiently large/coarse EAD workload went parallel;
- the planning overhead on sequential auto decisions was essentially eliminated.
Crucially, the engine was not modified with vocabulary-specific rules, filename-specific exceptions or thresholds tuned to make these files pass.
Current 30-file real-world validation corpus
The corpus covers very different XML domains and structures:
- UBL invoice
- FHIR patient
- GPX route
- KML places
- METS metadata
- MODS records
- JUnit report
- Nmap scan
- DocBook book
- PubMed articles
- BLAST result
- BioSample record
- RDF vocabulary
- GraphML graph
- RSS feed
- Atom feed
- XLIFF 2.0
- XLIFF 1.2
- MusicXML score
- OpenStreetMap
- SVG drawing
- SDMX Generic
- SDMX Structure
- OOXML shared strings
- Maven project
- Android manifest
- EAD finding aid
- SoapUI project
- TEI person data
- ONIX books
The purpose of this corpus is structural diversity, not optimisation for these particular vocabularies. General structural rules are preferred over corpus-specific special cases.
This validation does not mean that every arbitrary XML document has one objectively correct analyst table. It means that the package's generic structural rules and execution engine have been exercised across a broad set of real-world XML shapes without resorting to vocabulary-specific parsers.
The engineering conclusions behind the current defaults can be summarised compactly:
- Sequential execution is the semantic oracle.
- Automatic workers normally stop at four, but explicit higher counts are allowed.
- More workers can improve raw elapsed time, with declining efficiency.
parallel_chunksis the primary throughput-oriented strategy.shared_chunkis the primary memory-oriented strategy.shared_chunkhas shown materially lower peak/private memory in representative tests.- Outer
chunk_recordsand innertask_recordssolve different problems and are tuned separately. - Shared execution benefits from a modest number of coarse tasks; excessive tiny tasks waste scheduling time.
- Auto mode estimates structural work, not filenames, XML types or file size alone.
- Fast auto rejection is important so small documents pay essentially no planning penalty.
- No performance optimisation is accepted merely because it helps the existing validation corpus.
- Any tuning change must preserve exact sequential parity before its speed or memory result matters.
The public API is simple because this machinery sits underneath it, not because the machinery does not exist.
xmlrectr expects well-formed XML.
Malformed XML is detected and reported as an error or warning where appropriate. Repairing broken XML is intentionally outside the scope of the package: xmlrectr rectangles XML; it does not try to guess how a malformed source document should be rewritten.
Likewise, XSD inspection is intended to provide useful schema evidence for rectangling. xmlrectr is not a complete XSD validation or repair framework.
This README is intended to be a self-contained introduction. You should not need to install the package or open a vignette merely to understand what xmlrectr is trying to do.
For readers who want more detail, the repository also contains technical material that can be read directly on GitHub:
- Getting started with unknown XML
- Large XML, streaming and parallel execution
- Architecture and validation
- Engineering history and development notes
The technical articles are intentionally more detailed than this README. They are the right place for readers interested in scheduler design, bounded-memory execution, native acceleration, benchmark interpretation, worker/chunk/task tuning, validation boundaries and the engineering decisions behind the simple public interface.
A few principles define the project:
- Generic before vocabulary-specific. Structural rules should work across XML vocabularies.
- Sequential semantics are the reference. Performance work must not change the rectangle.
- Automation should be inspectable. Proposals are reviewable and profiles are explicit.
- XSD is evidence, not analytical truth.
- Do not invent relationships. Independent repetitions should not become Cartesian products.
- Preserve provenance. Analyst-friendly output should remain traceable to the XML structure.
- Performance matters on real data. Streaming, native acceleration and parallel execution are part of the architecture, not afterthoughts.
- Advanced machinery should not make routine use complicated. Sensible structural and computational decisions should be available without manual tuning.
The intended experience is simple even though the implementation underneath is not:
Give
xmlrectran XML file. If you know nothing about it, start exploring immediately. If you know more, tell the package what you know. If the workload is large, let the execution engine do the heavy lifting.
xmlrectr aims to be a universal rectangler, not a universal semantic interpreter.
It is designed to take arbitrary well-formed XML and produce useful R-oriented tabular representations without requiring a vocabulary-specific parser. It can exploit schema information and user knowledge when they exist, but it does not require them for exploratory use.
That is a deliberately ambitious target. The package cannot know the domain meaning of every XML vocabulary, but it can do a great deal of the structural and computational work required to move from hierarchical XML to an analyst-friendly table.
That is the problem xmlrectr is built to solve.