Explainable alcohol-label screening for faster, safer human review.
Label Compass compares submitted alcohol label artwork with declared application
data. It uses locally bundled browser OCR and deterministic, beverage-specific
rules to return evidence-backed PASS, needs correction, or REVIEW findings.
No API key, account, database, or third-party artwork upload is required.
Open Label Compass · Run the two-minute sample · Test the 10-label corpus · Run locally · Review evaluator evidence
Important
Label Compass is decision support—not a Certificate of Label Approval (COLA), legal advice, TTB endorsement, or a replacement for an authorized reviewer. “PASS” means only that the submitted evidence satisfied the eight checks implemented by this prototype.
- Open Label Compass.
- Download the bundled
old-tom-ocr-test.pngandmanifest.csvfiles. - Choose Detailed, upload the PNG, and select Import CSV / JSON applications.
- Import the manifest and select Run AI verification.
- Inspect the evidence, official rule links, raw OCR, and CSV/JSON exports.
With visual attestations left unset, the expected result is 6 PASS, 0 corrections, and 2 REVIEW. The two review findings are intentional: wording recognition cannot prove typography or physical label placement, so the product asks a person instead of inventing certainty.
For a larger end-to-end check, use the
10-label generated control corpus. It ships
five expected-pass and five one-field expected-failure proof sheets, an
importable manifest, exact hashes, and live-browser reproduction steps. The
fictional synthetic labels are regression evidence, not TTB approvals or a
legal gold set.
| Finding status | Meaning | Human action |
|---|---|---|
| PASS | Submitted evidence agrees with the declared value for that implemented check. | Confirm the evidence as part of the authorized review process. |
| Needs correction | The artwork contains a reliable contradiction or a confirmed omission. | Correct the application or artwork and run it again. |
| REVIEW | OCR, typography, completeness, or placement evidence is not strong enough for automation. | Inspect the label manually; do not treat REVIEW as noncompliance. |
The overall result uses the most conservative finding: any correction becomes
Label needs correction; otherwise any REVIEW becomes Send to manual
review; only an all-pass report becomes Label fields verified.
Label Compass produces one explainable finding for each of these eight areas:
- Brand name
- Class or type designation
- Alcohol content
- Net contents
- Producer, bottler, or importer name and address
- Country of origin for imports
- Government Health Warning
- Beverage-specific mandatory-information placement
Separate rule paths cover distilled spirits, wine, and malt beverages. Simple view shows only the operational workflow and action items; Detailed view adds the batch queue, confidence, recognized text, rule citations, and exports.
label image in the browser
→ safe color-preserving canvas preparation
→ optional two-panel gutter detection and vertical reflow
→ local Tesseract WebAssembly OCR
primary structured read + bounded sparse fallback
→ structured field extraction
→ deterministic beverage-specific comparisons
→ PASS / needs correction / REVIEW
→ evidence and official-rule links for a person
The AI component is optical character recognition—not a general-purpose LLM. Deterministic rules make every outcome reproducible and keep the system from quietly converting uncertain text into a compliance conclusion.
Requirements: Node.js 22.13 or newer and pnpm 11.
pnpm install
pnpm devOpen the local URL printed by the development server. For the native Next.js
path used by Vercel, run pnpm dev:vercel instead.
Run the complete verification suite:
pnpm test
pnpm lint
pnpm build:vercel| Need | Evidence |
|---|---|
| Production application | Label Compass on Vercel |
| Source code repository | github.com/geoffdls/label-compass |
| Engineering approach and trade-offs | docs/APPROACH.md |
| Automated, browser, load, and accessibility tests | docs/TEST-PLAN.md |
| Final release evidence | docs/RELEASE-VALIDATION.md |
| Generated 5-pass/5-fail OCR corpus | docs/GENERATED-LABEL-CORPUS.md |
| Import-ready generated corpus manifest | Repository · Deployment |
| Real-label OCR campaign | docs/REAL-LABEL-CORPUS.md |
| Third-party runtime inventory | THIRD_PARTY_NOTICES.md |
| Requirement mapping | Requirement traceability |
| Verified source snapshot | Bundle · SHA-256 · Metadata |
The repository is intentionally evidence-heavy, but the default product path remains the three-step upload, application-data, and review workflow shown above.
Verify the downloadable source snapshot
The source artifact is a complete, cloneable Git snapshot of the final human-authored source state. It includes the application, tests, documentation, lockfiles, and deployment configuration; it is not presented as the project's complete development history. The deployed artifact commit adds only the generated bundle, checksum, and source metadata, which avoids an impossible recursive bundle. After downloading the bundle and checksum into one folder:
shasum -a 256 -c labelverify-source.bundle.sha256
git clone labelverify-source.bundle labelverify
git -C labelverify bundle verify ../labelverify-source.bundleFull implementation capabilities
- Reads JPEG, PNG, and WebP label images with a locally bundled browser OCR model—no API key and no third-party artwork upload.
- Downscales OCR input in color to at most a 2,400-pixel longest edge, paints transparent artwork onto white so alpha does not erase dark label text, and encodes the prepared result as a lossless PNG.
- Detects a clean near-center gutter in a landscape two-panel proof sheet using image pixels alone, then stacks the left and right panels vertically on a white canvas capped at a 3,200-pixel longest edge. Single-panel, portrait, undersized-gap, and materially off-center layouts remain unsplit.
- Uses Tesseract page-segmentation mode 3 for a reflowed proof sheet. If the class/type, alcohol, net, producer, warning, or required imported-country field is absent—or the primary brand resembles an operator statement—mode 11 provides an independent second read. Evidence from both passes is retained; credible disagreement cannot be hidden by selecting a matching pass, and a better non-operator brand may replace an operator line. Ordinary artwork uses mode 11 directly; the submitted application never influences OCR layout or candidate selection.
- Bounds declared and extracted comparison fields at 512 characters, raw OCR
at 100,000 characters, individual OCR lines at 1,024 characters, and
normalized OCR input at 2,048 logical lines. Crossing either OCR-wide limit
forces every otherwise-passing OCR-dependent check to
REVIEW; an explicit human placement attestation remains independent. - Compares eight regulated content and evidence checks:
- Brand name
- Class/type designation
- Alcohol content
- Net contents
- Name and address of producer/bottler/importer
- Country of origin for imports
- Government Health Warning
- Beverage-specific mandatory-information placement
- Applies separate rules for distilled spirits, wine, and malt beverages.
- Uses fuzzy matching for brand text, beverage-specific producer/importer checks, category-aware ABV tolerance, and volume unit conversion.
- Retains all relevant class, alcohol, volume, and origin candidates—including same-line and later-label-face candidates—so an early matching statement cannot conceal contradictory text elsewhere on the artwork.
- Checks the Government Warning’s prescribed text and all-caps prefix, then asks a reviewer to attest to the visual rules OCR cannot prove: bold prefix, non-bold remainder, continuity, separation, contrasting background, applicable minimum type size, and the corresponding 40, 25, or 12 characters-per-inch limit, plus front/back/side label location and firm affixation when the warning label is not integral to the container.
- Requires a separate reviewer placement attestation with a readable,
category-specific checklist. Spirits guidance covers same-field-of-vision,
the narrowly qualified distinctive-liquor-bottle exception, and paired proof,
alcohol, net, and specialty-designation statements. Wine guidance defines the
brand label; distinguishes designation, appellation, foreign-blend, alcohol,
and net-content rules; and avoids importing spirits/malt co-location rules.
Malt guidance identifies eligible affixed/direct-applied locations, excluded
surfaces and loose packaging, molded-text limits, the keg exception, and
metric/alcohol/designation pairing. An unrecorded attestation routes only
this finding to
REVIEW. - Routes low-confidence images and unmeasurable typography to
REVIEWinstead of inventing a confident answer. - Treats missing fields in uploaded photographs as field-specific incomplete
evidence without weakening a clearly read contradiction elsewhere. Unseen
fields route to
REVIEWunless a reviewer confirms the upload contains the complete, legible artwork set; confirmed omissions can then be corrections. - Accepts up to 300 images and filename-matched CSV/JSON application manifests.
- Preflights manifests at 1 MB, 300 rows, 64 fields per row, 32 nesting levels, and 4,096 JSON containers before application parsing.
- Exports completed results as CSV or structured JSON.
- Defaults to a Simple view with the essential single-label workflow and action items. Detailed view adds batch tools, full reviewer guidance, all evidence panels, official citations, confidence, exports, and raw OCR. The choice changes presentation only and never changes a verification result.
- Keeps the interface focused on the operational upload, application-data, verification, evidence, and export workflow.
The interface uses a restrained dark theme designed for sustained review work and for evaluators who benefit from larger, higher-contrast controls:
- 17 px base text, 16 px form inputs, generous line spacing, and no functional interface text below 12 px on desktop.
- At least 44×44 px interactive targets in the tested review and report flows, with primary actions and inputs at 50–52 px high.
- A 3 px keyboard focus indicator, explicit invalid-field focus, sticky-header scroll offsets, and reduced-motion support.
- A labeled, keyboard-operable Simple/Detailed selector with 48 px targets; changing views preserves the selected label, form values, attestations, and report.
- Text and icons accompany every PASS, correction, and REVIEW state; color is never the only signal. Finding reasons and long filenames wrap instead of being silently truncated.
- Responsive reflow verified at 1440×900, 390×844, and 320×800 without page horizontal overflow. The explicit result status remains visible on narrow screens.
- Measured dark-theme contrast includes 17.53:1 primary text, at least 7.2:1
muted text, and 3.38:1 functional control boundaries. A
prefers-contrast: moremode strengthens muted text and boundaries further.
image in browser
→ Canvas resize + white alpha compositing while preserving color
→ pixel-only two-panel gutter detection and optional vertical reflow
→ bundled Tesseract WebAssembly OCR (PSM 3 primary / PSM 11 sparse fallback)
→ semantic field extraction
→ field-specific comparators
→ beverage-specific TTB requirement rules
→ worst-of aggregation (FAIL > REVIEW > PASS)
→ evidence-rich report
The architecture deliberately separates extraction from compliance:
lib/extraction.tsturns OCR text into structured candidate fields.lib/compliance.tscompares those candidates against the declared application and produces a typed report.lib/ocr-evidence.tspreserves bounded, independent evidence from both OCR passes so one read cannot silently hide a contradiction in the other.app/LabelVerifier.tsxowns the reviewer workflow, queue, manifest import, local OCR worker, and exports.lib/ocr-layout.tsdetects substantial side-by-side panels without consulting application values.lib/upload-safety.tsperforms encoded-dimension and manifest row preflight checks before expensive decoding or parsing.tests/covers comparators, conditional beverage rules, extraction, manifest parsing, upload safety, safe report serialization, response security headers, proof-sheet layout detection, the generated ten-label oracle, and production server rendering.
See docs/APPROACH.md for the decision record,
docs/TEST-PLAN.md for validation coverage,
docs/GENERATED-LABEL-CORPUS.md for the
reproducible five-pass/five-fail synthetic OCR suite,
docs/REAL-LABEL-CORPUS.md for the reproducible
Google-indexed image campaign, and
docs/RELEASE-VALIDATION.md for the final release
record.
| Requirement | Implementation |
|---|---|
| Fast per-label response | The generated 1800×2400 OCR fixture exercises the actual bundled model, with measured local timings recorded in docs/TEST-PLAN.md and the final production timing rechecked during release handoff |
| Non-technical reviewer UX | Large high-contrast controls, visible keyboard focus, 320 px reflow, plain-language outcomes, a focused three-step workflow, and field-level evidence |
| 200–300-label batch | 300-image queue and cap, CSV/JSON filename matching, accurate invalid/error accounting, 300-record export tests, a 300-item orchestration stress run, and a 200-image compiled real-OCR run at 2.29 seconds per label |
| Imperfect images | Color-preserving downscale, white alpha compositing, lossless PNG encoding, conservative two-panel vertical reflow, structured PSM 3 plus sparse PSM 11 fallback, field-specific coverage tracking, reviewer completeness attestation, and REVIEW routing for unresolved evidence |
| Government Warning screening | Exact normalized wording and capitalization, all-caps prefix, 0.5% applicability threshold, low-confidence review boundary, and explicit reviewer attestation for every non-text visual rule |
| Fuzzy brand matching | Case/punctuation/whitespace-insensitive edit similarity |
| Standalone prototype | No COLA integration, authentication, database, or external model credentials |
| Deployed URL + source | Native Next.js deployment on Vercel, Cloudflare-compatible vinext build support, public GitHub repository, and verified source bundle |
Official rule grounding and conditional logic
The implementation was checked against current official guidance:
- Distilled Spirits Labeling
- Distilled Spirits Alcohol Content
- Distilled Spirits Net Contents
- Distilled Spirits Name and Address
- Wine Labeling
- Wine Net Contents
- Wine Alcohol Content
- Wine Name and Address
- Malt Beverage Class and Type
- Malt Beverage Net Contents
- Malt Beverage Mandatory Label Information
- Government Health Warning
- 27 CFR 5.63 — Spirits label placement
- 27 CFR 5.205 — Distinctive-liquor-bottle exception
- 27 CFR 5.156 — Spirits specialty designations
- Wine Brand Label
- Wine Appellations of Origin
- 27 CFR 4.32–4.37 — Wine mandatory-information rules
- 27 CFR 7.51 — Malt labels and keg exception
- 27 CFR 7.61 — Malt-beverage label placement
- 27 CFR 7.70 — Malt net-content co-placement
- 27 CFR 7.141 — Malt class/type designation
- 27 CFR 16.21–16.22 — Warning placement and presentation
- 27 CFR 5.65 — Spirits alcohol content
- 27 CFR 5.70 — Spirits net contents
- 27 CFR 7.65 — Malt-beverage alcohol content
- 27 CFR 7.145 — Malt designations below 0.5%
Important conditional logic:
- Distilled spirits require a numerical alcohol-content statement.
- Distilled-spirits brand name, class/type, and alcohol content must appear in the same field of vision, except when an approved distinctive-liquor-bottle authorization permits other placement because the design precludes it; paired optional alcohol/net statements and specialty designation components have additional co-placement rules. For wine, brand and complete designation belong on the brand label; required appellation and conditional foreign-blend text have their own conjunction/placement rules; ABV may appear on any label; and net-content marking rules are wine-specific. Malt information can be divided among eligible labels or permitted direct markings, subject to excluded surfaces, molded-text limits, keg conditions, and statement-pairing rules. These physical facts require explicit reviewer evidence and are not inferred from OCR or the complete-artwork attestation.
- If optional proof appears on distilled spirits, it must equal twice the printed alcohol-by-volume percentage. Additional contradictory percentage statements are rejected.
- Wine over 14% requires a numerical statement. For wine from 7–14%,
table wineorlight winemay serve in place of a numerical statement. - Malt-beverage ABV is federally mandatory in certain formulation circumstances and may also be affected by State law, so a missing declaration without formula context routes to review.
- Truthful supplementary alcohol by weight may accompany the required ABV on
spirits and malt beverages when it is spelled out;
ABWis rejected. Printed malt alcohol values at or above 0.5% are limited to one decimal place, while values below 0.5% may use at most two decimal places. - Part 4 wine applications are bounded to the inclusive 7–24% range handled by this prototype.
Low alcoholandreduced alcoholmalt claims require less than 2.5% ABV.Non-alcoholicrequires the exact adjacent statement using either0.5 percentor.5%and routes to visual review;0.0%requires a truthfulalcohol freeclaim. Sub-0.5% products use the applicablemalt beverage,cereal beverage, ornear beerdesignation, withnear beertypography left for reviewer confirmation.- Wine and distilled-spirits net contents must include metric units; malt-beverage net contents must include U.S. standard units, with the other system permitted only in addition. When both systems appear, their quantities must agree within the applicable conversion rounding; repeated quantities may not conflict.
- The required standard-fill form is milliliters below 1 liter and liters at 1 liter or more. Wine does not accept centiliters as either the required or an additional quantity. For distilled spirits, a centiliter-only statement does not satisfy the requirement, but consistent supplementary metric equivalents (including cL or a second L/mL expression) may accompany the required milliliter/liter statement.
- An optional wine U.S. equivalent must be stated only in fluid ounces and use the prescribed precision: one decimal place below 100 fluid ounces and a whole number at 100 fluid ounces or more. Optional metric/U.S. quantities are checked for conversion consistency rather than accepted as decorative text.
- Distilled spirits are screened against the current authorized metric standards of fill, including the May 2026 additions.
- Ordinary wine containers below 4 liters are screened against the current standard fills: 50, 100, 180, 187, 200, 250, 300, 330, 355, 360, 375, 473, 500, 550, 568, 600, 620, 700, 720, and 750 mL; 1, 1.5, 1.8, 2.25, and 3 liters. From 4 through 17 liters, whole-liter fills are screened separately; saké and containers of 18 liters or more are left to their applicable rules.
- Malt-beverage quantities use volume-dependent U.S. formats: fluid ounces or
a pint fraction below one pint; the exact named unit at one pint, quart, or
gallon; the applicable larger-unit fraction or compound form between those
sizes; and gallons plus fractions above one gallon. Slash and mixed fractions
are parsed explicitly and must be in lowest terms. For example,
16 FL. OZ.alone does not replace1 Pint, and20 FL. OZ.alone does not replace1 Pint 4 FL. OZ.or an authorized quart fraction. - Malt-beverage alcohol declarations using a range, minimum, maximum,
not less than, ornot more thanare flagged instead of being accepted as a single percentage. - Permitted wine alcohol ranges are evaluated for width, application
containment, and class-appropriate protected boundaries. Specific values use
the applicable tolerance—generally ±1.5 percentage points at 14% or less and
±1 point above 14%—without crossing a protected boundary.
ABVis flagged as an unauthorized label abbreviation. - Brand, class/type, and name/address fields can auto-pass only after punctuation-tolerant normalized equality. High-similarity token or location substitutions route to review instead of receiving a false pass.
- Imported products require an authorized importer/agent role; optional
Imported forwording does not replaceImported by. Domestic wine requiresBottled byat 4 liters or less andPacked byabove 4 liters. - Exact combined importer/U.S.-bottler phrases are recognized. Imported-bulk
pathways that show only a plausible domestic processor or bottler identity
route to
REVIEWwhen the application schema cannot establish the pathway, instead of receiving a false failure or pass. - Country extraction retains disambiguating regional context, and explicit additional foreign-style health warnings are flagged.
- Name-and-address screening is beverage-specific: imported products require an importer or authorized U.S.-agent statement; domestic wine requires an authorized bottler/packer statement; distilled spirits require an authorized operation statement; and domestic malt beverages may identify the responsible brewer without a generic role prefix.
- Country of origin is evaluated for imports.
- A foreign-origin statement on an application marked domestic is flagged.
- Alcohol, proof, and net-content parsing rejects non-finite, negative, out-of-range, over-precise, contradictory, and prohibited range/maximum/ minimum forms instead of allowing malformed numeric text to reach a PASS.
- The Government Warning applies to beverages containing at least 0.5% alcohol by volume and must preserve the prescribed wording and capitalization as well as the format rules in 27 CFR part 16.
- Artwork stays in browser memory for the active session.
- The OCR JavaScript, WebAssembly core, and English model are vendored under
public/vendor/tesseract; the runtime does not depend on a public CDN. - Uploads are limited to JPEG/PNG/WebP and 10 MB each.
- File contents must carry a structurally valid JPEG, PNG, or WebP signature; filename extensions and browser MIME declarations are not trusted alone.
- Encoded dimensions are checked before decode, and decoded artwork is checked again against 10,000 pixels per edge and 25 megapixels to reduce decompression-bomb risk.
- Manifests are limited to 1 MB and 300 records before full parsing. CSV input is capped at 64 columns; JSON requires object rows and is capped at 64 properties per row, 32 nesting levels, and 4,096 containers.
- Batch size is capped at 300.
- Application and extracted comparison fields are limited to 512 characters; raw OCR is truncated at 100,000 characters and normalized line input at 1,024 characters, with at most 2,048 logical lines processed. An OCR-wide limit forces manual review. Oversized evidence is rejected predictably and truncated in the report rather than echoed back in full.
- Spreadsheet exports neutralize formula triggers even when
=,+,-, or@follows leading whitespace or control characters, then apply RFC 4180 quoting. - Application HTML responses set CSP, frame protection, MIME sniffing protection, restrictive referrer and permissions policies, plus same-origin opener and resource policies.
- Vendored OCR licensing and attributions are recorded in
THIRD_PARTY_NOTICES.md, with the corresponding license/notices beside the runtime files. - No authentication or persistence is included because both are explicit prototype non-goals.
Dependency security posture
Security-sensitive framework and tool versions are pinned in package.json
and pnpm-lock.yaml: Next.js and eslint-config-next 16.2.12, React,
React DOM, and react-server-dom-webpack 19.2.8, Vite 8.2.0, the Cloudflare
Vite plugin 1.49.0, and Wrangler 4.116.0. Narrow workspace overrides replace
Next.js's older transitive PostCSS and Sharp resolutions with PostCSS 8.5.25
and Sharp 0.35.3. Additional exact overrides resolve Undici 7.29.0,
fast-uri 3.1.5, and brace-expansion 5.0.9 in the affected development-tool
paths. The installed graph also resolves ws 8.21.0, esbuild 0.28.1, and Sharp
0.35.2/0.35.3 where required by their respective consumers.
The production-only pnpm audit --prod --audit-level=high reports no known
vulnerabilities. The full pnpm audit --audit-level=high exits clean under one
explicit, documented metadata exception for
GHSA-mh99-v99m-4gvg:
the registry range still treats every brace-expansion release below 5.0.8 as
affected, while the official API-compatible 1.x backport in 1.1.18 includes
the CVE-2026-14257 total-expansion-length bound. Forcing 5.x into minimatch 3
breaks its CommonJS contract, so the project pins 1.1.18 and records the
advisory exception transparently rather than hiding a broken toolchain or
carrying a local source patch. The exception should be removed once registry
metadata recognizes the backport.
These pnpm audits cover the installed package graph, not the separately
vendored Tesseract.js 5.1.1 browser runtime, WebAssembly core, or language
model. Those files are inventoried and licensed in
THIRD_PARTY_NOTICES.md. The worker is terminated
when the application unmounts; upgrading the vendored OCR major version
remains a separately benchmarked regression task rather than an untested
release-time substitution.
- This is not legal advice or an approval system.
- OCR can miss small, curved, reflective, occluded, or stylized text.
- One queue item represents one verification decision. The prototype does not
yet aggregate separate front/back/side photographs into one report; upload a
complete flat-artwork image when available. Partial views intentionally
return
REVIEWfor unseen fields rather than false corrections. A reviewer can explicitly attest that one upload is complete and legible so a genuine omitted required statement is eligible for a correction. - Browser OCR can read wording but cannot reliably prove visual warning format.
Unattested presentation returns
REVIEW; a reviewer may explicitly confirm the bold prefix, non-bold remainder, continuity, separation, contrast, applicable minimum type size, corresponding 40, 25, or 12 characters-per-inch limit, qualifying front/back/side label location, and firm affixation when that label is not integral to the container. - OCR also cannot establish which container face or designated label carries a
statement. The separate beverage-specific placement control records
confirmed compliant,placement issue found, ornot yet confirmed; the last state remainsREVIEW. Confirming complete, legible artwork does not answer the placement question. - The prototype compares text and values; it does not automatically measure physical type size or characters per inch. Those checks depend on the documented reviewer attestation.
- OCR batch processing is intentionally sequential to bound browser memory. The 300-item seeded development fixture used during testing validates queue orchestration but is intentionally omitted from the reviewer interface and is not an OCR-throughput benchmark. A separate 200-image compiled run did exercise the real bundled OCR worker end to end at 2.29 seconds per label with zero errors. A production service would still use a server-side work queue with controlled concurrency for sustained batches.
- The generated ten-label corpus demonstrates deterministic five-pass/five-fail behavior and exercises live browser OCR, but it is synthetic and is not a statistically valid measure of performance on real COLA submissions.
- Before federal deployment, use a human-adjudicated holdout of real, appropriately handled label submissions and complete security, privacy, accessibility, and model-risk reviews.
The next production increment would replace the local OCR extractor with a benchmarked vision model behind a government-approved endpoint, retain the deterministic comparator layer, add pixel-level warning typography measurement, and introduce an auditable reviewer queue. The decision boundary should remain the same: confident match, definite mismatch, or human review.
Prototype submission. Source is provided for portfolio and evaluation review.
The project source is not open source and remains all rights reserved; see
LICENSE.md. Third-party components retain their own terms as
listed in THIRD_PARTY_NOTICES.md.

