Document intelligence / Sustainability
Emissions you can trace back to the bill.
A document pipeline that reads scanned utility invoices for a multi-site operator, derives emissions from published factors and keeps every number connected to the page it came from.
Talk about a similar challenge- Template
- Usage
- Factor
- Emissions
- As printed
- Read from the bill
- Standardized
- Converted to common units
- Published factor
- Pinned to its vintage
- Unreadable value
- Left blank, not guessed
- Setting
- A multi-site operator
- Challenge
- Emissions data locked inside thousands of scanned invoices
- Built
- A document pipeline, a lineage dashboard and a review workflow
- Decision owner
- The sustainability team and its reviewers
The starting point
The data existed. It was just trapped in PDFs.
A multi-site operator receives thousands of utility invoices a year, most of them scanned documents inside an accounts-payable platform. To meet climate-disclosure reporting requirements, the sustainability team needed the energy used on each bill: electricity, natural gas, propane and fuel oil. Water is out of scope. The alternative was re-keying everything by hand.
The inputs are plain. Accounts payable drops a work-list spreadsheet on a schedule. The pipeline compares it with a ledger, retrieves only the invoices it has not seen, and works from there. Weekly fuel files arrive separately and are reported as their own section.
A decision about the model
We began with document AI, then chose explicit templates.
The first approach used a trainable document-AI service. It struggled for a specific reason: each bill arrives wrapped in accounts-payable cover pages, and the model learned those wrapper pages instead of the layouts of the bills inside. A model trained on the wrong thing performs confidently on the wrong thing.
We moved to an explicit design. The pipeline recognizes each vendor’s bill layout from a template written in plain configuration. More than a hundred vendor layouts are covered, many with several services on one bill. There is no language model in the extraction path. For this problem, a readable rule that an engineer and a reviewer can both inspect was a better tool than a clever model.
We bring the same question to any AI project: does this step need a model at all?
The pipeline
From a work-list to a number, one bill at a time.
The pipeline separates the cover pages from the real bill, splits invoices that cover several departments and reads the bill pages with optical character recognition. It selects the matching vendor template and extracts usage in the bill’s own units, such as kilowatt-hours, therms and gallons.
Extracted dollars are reconciled against the accounts-payable amount as a check. Dollars never feed the emissions calculation. Usage does. Emissions are then derived from published EPA emission factors, pinned to a stated vintage, in a step that can be re-run without re-reading the bills.
A dashboard shows results by site, division, vendor and month. Every figure has a lineage view that shows the source page next to the extracted value, the conversion and the factor behind the result. The export is built from the same data, so the export always matches the dashboard.
Principles
Absent is better than wrong.
An emissions inventory is only useful if people can trust its gaps as much as its numbers. So the pipeline does not estimate. If a value cannot be read, it stays blank. If a machine-read value exceeds a plausibility cap, it is withheld for review. Values that passed only a lighter check are labelled as such.
Reviewers can correct a value from the dashboard. A correction is a record, not an edit: it overrides the machine value, keeps the original and notes where the confirmation came from, whether a client reviewer or one of our own verification passes.
Proving it
Ground truth written by people, checked on every change.
More than a hundred and fifty bills were read by eye and their expected values typed by hand, never copied from the OCR output or from accounts-payable data, so the tests do not inherit the system’s mistakes. Automated checks run on every change and compare the pipeline’s results with the hand-typed values across the whole set.
Audits found real problems. One was a double count: a gas supplier and a distributor both appeared for the same gas. The fix shipped with a regression test so it cannot return unnoticed.
Delivery
Built inside the client’s private cloud.
The pipeline and dashboard run on Azure inside a network the client owns, with no public access. Sign-in uses the client’s identity provider with role-based access, and the application stores no database passwords. Infrastructure is defined as code and documented.
The system runs in the client’s production environment. Releases move between environments through an automated, documented process.
The connected workflow
Follow the work, from start to decision.
- 01
Receive the work-list
Compare it with a ledger and retrieve only the invoices not yet processed.
- 02
Separate and read
Set aside cover pages, split multi-department invoices and read the bill pages.
- 03
Extract with a template
Match the vendor layout and extract usage in the bill’s own units.
- 04
Check and derive
Reconcile against the payable amount, then derive emissions from pinned published factors.
- 05
Review and report
Inspect each figure with its source page, correct what is wrong and export from the same data.
An engineering detail that matters
One long bill can stop a whole run.
Reading a scanned document uses more memory the more pages it has. A single unusually long bill stalled a batch before anyone noticed. The pipeline now reads each bill in a fresh process, applies a time limit that grows with its page count and sets aside any bill that still fails, so the rest continue.
A pipeline that handles thousands of documents has to assume that some will misbehave, and keep the failure local and visible.
Emissions use location-based Scope 2 accounting. A small share of bills cannot be read; they appear as visible gaps rather than estimates, and the methodology has not yet had third-party assurance.