Case studies/AI / Health-report prototype

AI / Health-report prototype

From a lab report to a plain-language plan.

A prototype that turns a member’s lab-report PDF into an educational plan, with extraction checks, retrieved research citations and guardrails that keep the clinician in charge.

Talk about a similar challenge
HEALTH REPORT PROTOTYPEEducational, with citations
A member’s lab report (PDF)
VALUES, CHECKEDValidated
  • Marker oneIn range
  • Marker twoWorth discussing
  • Marker threeIn range

A plain-language plan

Practical food and lifestyle suggestions, each tied to a research passage.

Passage 1Passage 2
Published
Setting
A health-testing company
Challenge
Members received numbers with no plain-language next step
Built
A prototype: PDF extraction, cited retrieval and a guarded assistant
Stage
A working prototype, demonstrated to the client’s product team

The starting point

People got numbers. They wanted to know what to do.

A health-testing company gives each member a laboratory report and a separate lifestyle report. Both are detailed and both are hard to act on. The question members ask is the one the documents do not answer: what does this mean for me, and what should I do about it?

We built a prototype that reads a member’s own documents and drafts a plain-language, educational plan, with the research behind each suggestion attached. It was designed from the start to support a conversation with a clinician and never to replace one.

Reading a messy document

Most of the early work was reading the PDF correctly.

Lab reports vary in layout, and a wrong number would spoil everything that follows. The prototype has two extraction routes: rule-based extractors with a section detector, a normalizer and a validator, and a route in which a model reads the PDF directly. Both feed the same checks for plausible ranges and units.

The list of markers and scoring rules is kept in configuration, so it can change without a new software release. We tested the extraction on a handful of real sample reports: enough to check the approach, not enough to measure accuracy.

Grounding the advice

Every suggestion should point to something a person can read.

A language model drafts the assessment from the validated values and the configured rules. A retrieval service finds relevant passages in a curated body of research, and the plan and the chat assistant cite them. The assistant is instructed to answer only from the passages it is given and to decline questions outside its scope.

The assessment organizes findings into six areas such as cardiovascular health, blood-sugar regulation and inflammation, with a three-month plan, a retest suggestion and a printable summary. The model drafts the risk narrative. Nothing in the prototype has been clinically validated, and the interface says so.

Guardrails

Decide what the system should not say.

The instructions to the model are explicit: it is not a doctor, it does not suggest medications, it keeps to diet and lifestyle, and it sends the member back to their provider. The retrieval assistant is told not to guess beyond its sources.

Client feedback shaped the product. After early demonstrations we stopped showing age as a risk factor and removed anything that read as medical advice. A prototype is the cheapest place to find those calls.

The connected workflow

Follow the work, from start to decision.

  1. 01

    Upload the report

    The member provides their own lab-report PDF.

  2. 02

    Extract and validate

    Read the values, normalize units and check them against plausible ranges.

  3. 03

    Retrieve the evidence

    Find relevant passages in a curated body of research.

  4. 04

    Draft with citations

    Write an educational plan, with each suggestion tied to its source.

  5. 05

    Take it to a clinician

    Print the summary and use it as the start of a conversation.

An engineering detail that matters

When the model fails, the app should fail loudly.

Models sometimes return text that does not parse or omits a required field. The prototype demands strict structured output, strips stray formatting and checks every required field. If the result still cannot be trusted, it returns an error. It does not fall back to a simpler rule-based answer.

That was deliberate. For health information, a plausible-looking substitute can mislead more than a visible failure. Because a language model drafts the narrative, two runs on the same report can also read differently. The model can be swapped for another without changing that rule.

Start with the workflow that needs to change.

Tell us where the work gets stuck. We’ll come back with questions and a practical first step.

Talk about your project