Every engagement is different. The shape below is our default, not a template we force a problem into. We adapt it in the first days based on what is riskiest, and we keep adapting as we learn.

What does not change is the sequence: understand before proposing, measure before building, and show progress every week.

The arc of an engagement

Intro call. Thirty minutes, free. You describe the problem. We tell you what we think, including when we think you should not build it or should not hire us. If there is a fit, we talk about budget early, because budget decides what scope is possible and pretending otherwise wastes everybody's time.

Discovery. Two weeks. We learn your problem, your data, and your constraints properly. Discovery is analysis, not construction.

Report, presentation, proposal. At the end of discovery you receive a written report, a live walkthrough with questions answered in the room, and a proposal for what we would do next.

Delivery. Weekly iterations with visible output. The first thing we build is the evaluation and the baseline, before any of the system they are meant to judge.

Handover. Documentation, runbooks, and knowledge transfer, so your team can run what we built. A full handover is always available. Ongoing support is available when you want it, and never required to keep the system working.

The sentence

Before anything else, we want one sentence, written down and agreed by the people who will pay for the work:

This system should change this specific workflow by this measurable amount.

It sounds trivial. It is the hardest artefact in the engagement, and it decides more outcomes than any technical choice that follows.

Projects that cannot produce this sentence tend to fail slowly and expensively. Not because the technology disappoints, but because nobody agreed what winning meant, so every result is arguable and no result is conclusive. A system gets built, it works about as well as such systems work, and eighteen months later there is still no answer to the question of whether it was worth it.

We ask for the sentence in week one of discovery. If nobody in the organisation can write it, that finding goes in the report, in those words, because it is more useful than anything else we could tell you. It usually means the problem is not yet a project, and it is far cheaper to discover at that point than after a team has been assembled around it.

Everything downstream hangs off the sentence. The evaluation measures the amount. The baseline says where you are starting from. The kill criteria say what result would mean the current plan is not going to get you there.

Discovery: two weeks

Discovery is the most concrete thing we sell, so here is exactly what happens.

Kickoff

We start with a kickoff meeting and we ask you to invite everyone relevant to the project, not only the people who will manage it. The engineers who own the systems. The people who do today, by hand, whatever this is meant to help with. Whoever will decide.

The goal is a high level shared understanding: what needs to be done, why now, who owns which part, what has already been tried, and what would count as success. We take notes and we share them.

Week one: understand

Week one is information gathering, on two tracks at once.

The business track: a series of interviews with the people we met at kickoff and anyone they point us to. What decision does this system support? Who acts on its output? What does an error cost? What is the value at stake if it works? What has been attempted before, and why did it stop? This is also where we work towards the sentence.

The technical track: how the relevant systems fit together, where the data lives, how it is produced, and what shape it is in. We start the access conversation immediately, because data access is the most common cause of a slow discovery and it usually involves people who were not in the kickoff.

We also write down our assumptions as we collect them, and rank them by two questions: how badly would we be hurt if this is wrong, and how confident are we? The assumptions that are both risky and uncertain set the agenda for week two.

Week two: analyse

Week two is where the material gets examined rather than described.

With access in hand we profile the actual data: distributions, coverage, label quality, missingness, and how far the historical record resembles what a live system would see. We read records rather than summary statistics, including the ugly ones. We also pull the interview notes together and reconcile them, because the account of a process given by the people who manage it and the account given by the people who perform it are rarely the same account, and the gap between them is usually where the real difficulty lives.

In parallel we sketch what a production architecture would have to look like, and cost it out, roughly but honestly, in both money and latency.

We write throwaway code during this. Notebooks, queries, quick probes to check whether a signal is there at all. That work exists to inform our judgement. It is not a deliverable and we do not present it as one. Two weeks is enough to form a well evidenced opinion; it is not enough to build anything you should trust, and a discovery that hands over a prototype is usually a discovery that spent its time on the wrong thing.

Understanding keeps being refined throughout. Interviews often continue into week two, because the data raises questions only a person can answer, and because the questions we can ask after seeing the data are much better than the ones we could ask before it.

When access does not arrive in time, we say so plainly and we mark what remains unverified rather than quietly assuming it.

What you receive

Three things.

1. A written report. What we found. What state the data is actually in. Whether the thing is feasible, and on what evidence. The recommended approach. How we would measure success, including the evaluation we would build, the metric that maps to your decision, and the threshold that would count as working. The risks we can see. And an explicit list of what we could not verify, and why.

2. A presentation and walkthrough. We present the report to the stakeholders from the kickoff and answer questions live, so the findings land with the people who have to act on them.

3. A proposal. Scoped options for what to do next, with the effort each would take and what each would buy you.

The report is yours unconditionally. If you decide to take it to a different supplier, or to your own team, or to nobody, that is a legitimate outcome and the document is written to be useful in that case. We would rather write a report you can act on without us than one that only works as a sales document.

Measurement comes before building

The first thing we build in an engagement is not the system. It is the evaluation set and the baseline.

The evaluation set is assembled deliberately rather than sampled for convenience: stratified so that a headline number cannot hide a segment where the system is useless, held out along whatever boundary prevents leakage, and weighted towards the rare and awkward cases that cost the most when they go wrong. The baseline is the simplest thing that could work, often several of them, including the performance of whatever you do today.

Together they establish the two numbers everything else is judged against: where you are now, and what would count as better.

How long this takes depends on the problem, the state of the data, and how quickly your side can answer questions. We do not quote a fixed duration for it, because a number invented before we have seen your labelling would be a guess dressed as a commitment. What we commit to is the order. Nothing gets optimised before there is something to optimise against.

Occasionally an engagement can skip this, when earlier work already left a trustworthy evaluation in place. That is rare, and we check the evaluation before we trust it.

Research and Evaluation covers how this is actually done.

Killing plans, not engagements

Every phase of work carries a stated kill criterion: the result, or the elapsed time, that would mean this particular plan is not going to reach the sentence.

It matters that this kills a plan and not a relationship. We work as an agile team. When an approach stops looking viable we change the approach, and the criterion exists so that decision gets made in week three rather than month six. Most of the time it fires quietly and produces a better plan. Occasionally it fires and the honest conclusion is that the whole thing should stop, and we will say that too.

Work without a stopping rule expands until the budget runs out, and the hardest experiment to abandon is always the one you have already spent three weeks on. Naming the criterion in advance, while nobody is invested, is what makes it possible to act on later, when everybody is.

Who does the work

The people you meet in discovery are the people who do the work.

The proposal names them. If somebody joins the team or leaves it, we tell you before it happens, not after you notice that the writing style changed. There is no bench of juniors learning on your project behind a senior who appeared once to sell it.

We staff to the problem rather than to a template, so the size of the team moves with what the work needs. What does not move is the seniority.

And when a problem needs a skill we do not have, we say so, and where we can we point you at somebody who does. Recommending ourselves for work we would be mediocre at is a bad trade in both directions. It costs you a project and it costs us the reference.


That is the arc, from the first call to the point where delivery starts. The next page covers the part that lasts longest: the shapes an engagement can take, what a week looks like, and what each side owes the other.

Next: Working Together

Want to see this run on your problem?

Book a free 30-minute call. Bring the problem, and we will work out the shape of an engagement together.

No pitch deck, no obligation. If AI is the wrong answer for your problem, we'll tell you that too.