Reading the document is the hard part

Sep 30, 2026 | ai-reliability-engineering

Reading Time: 6 minutes
featured 4

If you want to know what actually went wrong in an AI project, do not read the architecture document. Read the prompt.

Prompts accumulate scar tissue. Every defensive clause in one is a record of a specific day something came out wrong. Here is a fragment from a document extraction step in a system my team at Aditi built for a staffing client. It reads a recruiter’s offer document and pulls out the parameters needed to generate the employment paperwork.

- Extract EVERY role listed in the document, do NOT skip, omit, or leave out any single item
- Extract ONLY the roles explicitly listed, do NOT add, invent, or infer additional roles
- Preserve the EXACT count: if the document lists 8 roles, return exactly 8 items
- Do NOT split a single role into multiple items (even if it contains "and")
- Do NOT merge multiple roles into one
- Each extracted item must correspond 1-to-1 with a role as written
- Before returning, verify your array length matches the number of roles

Seven instructions. All of them guarding one property: the number of items out equals the number of items in.

Nobody writes seven instructions for one property on the first attempt. Each of those lines is a Tuesday.

The estimate I got wrong

When this work was scoped, I treated extraction as the easy part. Reading a document and pulling out fields is what these models are unambiguously good at, and the interesting problems were clearly elsewhere: the arithmetic, the template rendering, the real-time progress reporting.

I was wrong by a wide margin. Extraction consumed more engineering time than anything else in the system, and the reason is not that the model is weak. It is that I had scoped it as a document problem when it is a human variability problem.

A person doing this job by hand does it in under a minute without thinking. That is the trap. They are not reading the document, they are recognising it, against fifteen years of having seen documents like it. All the work is in the recognition, and none of it is written down anywhere.

I have changed how I scope these. The first question is no longer “what fields do you need out”. It is “show me the five weirdest examples of this document you have”, and the estimate comes from those, not from the clean one in the specification.

Four things a human does not notice doing

Money is written in a local dialect of numbers

The contract amount arrives as “20 lakh”, or “20 lakhs”, or “20L”, or “2 crore”, or with a currency symbol, or spelled out. An Indian reader converts all of those without any conscious step. The downstream calculation needs an integer.

So the prompt carries the conversion table explicitly, with worked examples in both directions, and the instruction to output a pure number and nothing else. Not because the model cannot do the arithmetic, but because left to itself it will helpfully return a formatted string, and formatted strings are how you get a currency symbol into a calculation three steps later.

The lesson generalises past this domain. Anywhere a quantity is written by a human for another human, there is a local dialect, and the model will faithfully reproduce the dialect unless told to normalise. Ask for the shape you need, not the value you want.

Addresses have a shape, and the shape is the content

This one cost us a redraft of the extraction contract:

CRITICAL: If address has multiple lines, preserve the multi-line format
using \n between lines. Do NOT combine into single line.

An address flattened to one line is still correct, still contains every word, and looks wrong on a printed letter in a way that everybody notices and nobody can articulate. It reads as machine-generated, which for a document formalising somebody’s employment is precisely the impression to avoid.

The line breaks are semantic. They are not formatting to be normalised away, they carry information about what is a building and what is a locality. That is not obvious until you see the output on a page.

The field that decides everything is one word long

There is a field in the extracted structure that names which employee group the candidate belongs to. It is one token. It is also the routing key for the entire downstream process: it selects the template set, decides which letters get generated at all, and names the output folder.

Get it wrong and nothing errors. Every step succeeds. You get a complete, well-formatted, internally consistent set of documents on the wrong contract type.

That is the worst failure mode a system can have, because it is silent and it looks like success. A crash is a gift by comparison.

Worth being honest about how the code handles it today: the group defaults when the value is missing, in three separate places. That was a pragmatic call under delivery pressure and it is the thing in this system I am least comfortable with, because a default on a routing key converts a loud failure into a quiet one. The correct behaviour is to refuse and ask a human. It is on the list.

Some defaults are uncomfortable and someone has to choose them

The letters use gendered salutations, so the extraction infers gender from the salutation or pronouns in the source document, and when it cannot:

If unclear, use "Male"

I do not like that line. It is in the codebase, it is defensible on the specific distribution of names this client hires from, and it is still a system making an assumption about a person from their name.

I am including it here rather than quietly leaving it out, because a post that shows only the decisions I am pleased with is marketing. The honest position is that this is a known compromise with a known failure mode, the failure is visible on the face of the document rather than buried, and a human reviews every letter before it goes out. That is the mitigation. It is not the same as it being right.

What holds this together

The extraction does not run into a void. Two things make it recoverable.

Every extraction is written to disk as JSON, named and timestamped, next to the documents it produced. Not logged. Stored, as an artifact, alongside the output.

That means the question “why does this letter say that” always has a mechanical answer. You open the extraction for that run and read what the model saw. Without it, every dispute becomes an argument about what probably happened, and with a probabilistic component in the pipeline, that argument has no end.

This is the single highest-value thing in the system relative to what it cost. It is a file write.

A human reviews the output before it goes out. Which is the reason a defensible-but-imperfect default is acceptable, and would not be if the letters went straight to the candidate.

I want to name this properly, because it gets called the wrong thing. This is not prompt engineering. Prompt engineering is a tips-and-tricks genre about phrasing. What is described above is AI reliability engineering: bounding the output contract, normalising input dialects, making silent failures loud, and keeping an audit artifact for every probabilistic decision.

The distinction matters because it determines what you build. If you think you have a phrasing problem, you rewrite the prompt. If you know you have a reliability problem, you build the audit trail, and then the prompt rewrites are a maintenance activity rather than the strategy.

What travels

Any system where a model reads something a human wrote, for another human, will meet all four of these.

Normalise the dialect at the boundary. Quantities, dates, names and money all have local written forms. Specify the output shape, do not hope for it.

Structure can be semantic. Line breaks, ordering and grouping sometimes carry meaning that survives being flattened but does not survive being read.

Find the one-word field that routes everything, and make it impossible to default. There is always one. It is the field that fails silently.

Store every extraction as an artifact. The moment you cannot reconstruct what the model saw, every future disagreement is unresolvable.

I scoped this as the easy part because a person does it in under a minute. That was the mistake. A person does it in under a minute by not doing most of it.