Skip to content

What a build looks like

Six phases, two of which are designed as places to stop. Typically six to ten weeks from scoping to a pilot group, depending mostly on how tidy the documents are and how fast access is granted.

01

Scoping

About a week
We do
We sit with the people who currently answer the questions and work out which questions actually get asked, how often, and what a wrong answer costs. Then we look at a sample of the documents.
You do
An hour each with two or three people who field the questions today. Read access to a representative sample of documents, not all of them.
Comes out
A written scope, a fixed price, and an honest read on whether your documents are in a state where this will work.

Sensible stopping point one. If the documents are contradictory or badly structured, we say so here and you fix that first. Some clients never come back and that is a correct outcome.

02

Evaluation set

A few days
We do
We build the test set before the system: 50 to 100 real questions with known correct answers and the source document for each.
You do
Confirmation that the answers are right. This is the highest-value hour anyone at your end spends on the project.
Comes out
A measurable definition of "working", agreed before anyone can be tempted to move the goalposts.
03

Ingestion and indexing

One to three weeks
We do
Connect the sources, parse the documents, decide chunking against your actual structure, handle permissions, build the index.
You do
Access to the real repositories, and someone who can answer questions about the odd documents. There are always odd documents.
Comes out
A searchable index and a report on what did not parse cleanly, which is always something.
04

Retrieval tuning

One to two weeks
We do
Hybrid search, reranking, thresholds. Run the evaluation set, look at what fails, change one thing, run it again.
You do
Nothing much. This is the least visible and most determinative part of the whole project.
Comes out
Measured accuracy against the evaluation set, with the failures listed rather than averaged away.

Sensible stopping point two. If accuracy is not good enough here, more prompt engineering will not save it. Better to stop and fix the corpus than ship something people learn not to trust.

05

Interface and rollout

One to two weeks
We do
Put it where people already are — Teams, Slack, an intranet page — rather than making them visit a new tool. Wire logging.
You do
A decision on where it lives, and a pilot group who will actually use it.
Comes out
Something in front of real users, with the question log running from day one.
06

Aftercare

Ongoing, optional
We do
Re-run the evaluation set on a schedule, review the questions that returned nothing, re-index as documents change.
You do
Someone who owns it internally. Without that it decays, and no support contract prevents that.
Comes out
A system that is as good in month twelve as it was in month one, which is not the default.

The thing that delays every one of these

Access. Not the technical work — getting read permissions to the actual repositories, signed off by whoever owns them. On projects that run late, this is the reason about four times out of five.

Start that conversation the same week you start scoping, even though it feels premature. It is not.

Start with scoping

One week, fixed price, and it ends with a straight answer about whether to proceed. Tell us roughly what the documents are and who keeps getting interrupted to answer questions about them.

Talk about scoping