
A pilot is a small, time-limited client project that tests one business workflow before a wider rollout. For a Claude or OpenAI pilot, agree what information goes in, what the model will do, what comes out and who checks the result before you quote a price.
Also agree who owns the account, which tests the output must pass, what the client receives at handover and when support ends. Acceptance criteria are the checks the result must pass before the client accepts the pilot. Handover is the transfer of the workflow, instructions, access and operating knowledge to the client.
This method applies whether the pilot uses Claude, ChatGPT or the OpenAI API. The technical work changes by product, but the client still needs an agreed task, price, test and handover.
1. Name the exact job the pilot will test
“Implement AI for our sales team” is not a pilot scope. It names a department and a technology, but it does not define the work.
A scoped question sounds like this:
Can the client’s account managers use an AI-assisted workflow to turn approved discovery notes into a first proposal outline, while preserving source references and preventing unsupported commitments?
That question identifies the users, input, output and central risk. It also makes a pilot testable. The consultant can prepare representative notes, define prohibited claims, run a set of cases and ask designated reviewers to accept or reject the outlines.
Choose a workflow with a repeatable input and a reviewable output. Avoid using the first pilot to automate an irreversible action, resolve a high-stakes judgment or connect every client system at once.
2. Write down the input, action, output and check
A short table shows exactly what the consultant will and will not deliver.
Illustrative pilot: proposal-outline assistant
This is a fictional example. The invented amounts and effort illustrate the pricing method. They do not represent market rates or a client result.
| Element | Pilot definition |
|---|---|
| Input | Ten sanitized discovery-note packs, the approved service catalogue, proposal structure and prohibited-claim list |
| Action | Extract stated needs, map only approved services, identify missing information and draft an outline with source references |
| Output | A proposal outline and a list of questions; no final proposal and no external sending |
| Check | A sales lead scores factual support, service fit, missing questions, prohibited claims and editing effort |
The table keeps pricing, legal terms, document design and customer relationship management (CRM) integration outside the project unless they are listed. It also requires the draft to pass the agreed checks instead of passing because it reads well.
3. Decide access and ownership before building
Ask who owns every part of the pilot:
- Which client account, workspace or API project will be used?
- Who authorizes access to the source material?
- Which user roles can run or change the workflow?
- Where will prompts, configuration, test cases and logs be stored?
- Who can revoke access when the engagement ends?
- Which assets will the client receive at handover?
Do not design a client workflow that depends indefinitely on the consultant’s personal account or personal API key. A temporary prototype may begin in a controlled consultant environment when the agreement permits it, but the proposal must state that fact, the data allowed there and the route to client ownership.
Provider data controls differ by product and plan. OpenAI’s current API documentation says API data is not used to train models unless the customer opts in, describes default abuse-monitoring retention, and lists additional retention controls available to eligible customers. Anthropic separately documents commercial-product privacy information and Claude Enterprise retention controls. Check the selected product, contract and client policy rather than copying a consumer-product assumption into the scope. (OpenAI API data controls; Anthropic commercial-customer privacy; Claude Enterprise retention).
The consultant should record the decision and its owner. The consultant should not unilaterally approve the client’s security, legal or retention position.
4. Agree the tests before the final prompt
Acceptance criteria tell the client and consultant which checks the output must pass.
Develop the prompt on a small working set that the team is allowed to reuse. Once the prompt, configuration and rubric are stable, freeze them. Only then open a separate, untouched acceptance set of ten note packs. Each acceptance pack contains known facts, missing information and at least one tempting but unsupported conclusion. The sales lead reviews each output against a rubric.
A simple rubric might require:
- every proposed service to point to an approved catalogue entry;
- every client fact to point to the supplied notes;
- missing commercial details to become questions; assumptions are prohibited;
- zero prohibited commitments;
- the outline to follow the approved structure;
- the reviewer to record material edits.
Set the pass rule before the final run. For example, the client might require all prohibited-claim checks to pass and at least eight of ten first-pass outlines to meet the other criteria. That threshold is illustrative; the client must choose a standard appropriate to the consequence of error.
Retain the first-pass output and score before anyone edits it. Then record the human-edited result and review effort separately. The edited version shows whether the workflow is usable; it must not replace the first-pass record or make model quality look better than it was. If the team changes the prompt after seeing an acceptance case, that case becomes development material and a new untouched case is needed for acceptance.
OpenAI provides an Evals facility for defining test data and grading criteria, while Anthropic recommends structured test records and verification tools for longer workflows. You do not need a sophisticated evaluation platform for every small pilot, but you do need fixed cases, explicit criteria and retained results. (OpenAI Evals; Anthropic prompting and verification guidance).
5. State what a later production system would still need
A pilot may prove that the workflow can produce acceptable outlines on controlled source packs. It does not by itself prove that the workflow will remain reliable across every account manager, document format, model update or integration failure.
State what production would add:
- identity and role design;
- approved data connections;
- monitoring and incident ownership;
- usage and cost controls;
- a change process for instructions and source material;
- regression tests when a model or workflow changes;
- support hours and response expectations;
- records required by the client’s policy.
OpenAI notes that model output is variable and recommends pinned model versions and evals for consistency-sensitive applications. Anthropic advises testing applications against replacement models before a model retirement. These facts support change testing in production planning. They cannot guarantee identical output. (OpenAI API compatibility guidance; Anthropic model deprecations).
6. Price the work from effort, cost and risk
There is no honest universal price for a “Claude pilot” or “OpenAI pilot.” A workshop using sanitized documents and a custom application connected to internal systems are different services.
Use a transparent calculation:
Pilot price = delivery effort + direct costs + contingency for defined uncertainty + commercial margin
Break delivery effort into activities:
| Activity | Estimated effort |
|---|---|
| Discovery and workflow definition | 10 hours |
| Access and data-preparation coordination | 6 hours |
| Prototype configuration or development | 18 hours |
| Evaluation-set and rubric creation | 10 hours |
| Pilot runs and repairs | 12 hours |
| User review session and training | 6 hours |
| Handover pack | 6 hours |
| Project management | 8 hours |
| Illustrative total | 76 hours |
Suppose the consultant’s internally calculated delivery rate is $100 per hour. This figure is invented for the example; it is not a recommended or observed market rate.
- Delivery effort: 76 × 100 = **7,600**
- Direct pilot costs budget: $400
- Defined contingency: 10% of delivery effort = $760
- Illustrative cost-and-contingency base = $7,600 + $400 + 760 = **8,760**
The consultant would then apply the margin needed for the business and present the client with a fixed price or staged commercial structure. If a 20% margin on selling price were chosen, the calculation would be 8, 760 ÷ 0.80 = **10,950**. The calculation demonstrates the arithmetic. It does not recommend a market price.
Explain the assumptions beside the price. If the client later supplies sensitive data, requests an integration or changes the acceptance set, record a scope change. Price the added work instead of hiding it inside “prompt tuning.”
Resolve a disputed result before it becomes a pricing dispute
Assume a fictional pilot reviews 20 policy questions. The agreed acceptance rule says a case passes only when the answer cites the correct section, preserves every condition and escalates an unresolved exception. The first output gets 16 cases right. A reviewer corrects three more; one remains unresolved.
Record the two numbers separately:
| Measure | Result | Why it matters |
|---|---|---|
| First-pass accepted | 16 of 20 | Shows what the workflow produced before correction |
| Accepted after reviewer correction | 19 of 20 | Shows the delivered draft after human effort |
| Unresolved | 1 of 20 | Prevents the final result from hiding a remaining gap |
If the client expected “19 correct,” the log settles whether that meant first-pass output or reviewed delivery. The proposal should define the measure before the work begins.
Now add a change request: halfway through the pilot, the client asks the workflow to write approved answers into the service desk. That request adds authentication, permissions, duplicate prevention, failure recovery and audit logging. Record it as a new requirement with its effect on effort, risk, tests and handover. Do not absorb it into the original fixed price simply because the drafting step already works.
A useful change record has six fields: request, business reason, added deliverable, added acceptance cases, price or schedule effect, and approver. This makes scope negotiation inspectable instead of personal.
7. Make exclusions specific
A useful pilot proposal says what is outside scope in terms the client can recognize.
For the illustrative assistant, exclusions could include:
- creating final commercial or legal terms;
- sending proposals to customers;
- integrating the CRM;
- processing unsanitized customer data;
- supporting languages absent from the evaluation set;
- production uptime or response-time commitments;
- ongoing model and prompt maintenance after the agreed support window.
“Anything not mentioned is excluded” is technically broad and operationally weak. Name the likely misunderstandings.
8. Give the client everything needed to run the workflow
A pilot is incomplete if only the consultant knows how it works.
The handover package should contain:
- Workflow definition: purpose, user, permitted inputs, outputs and exclusions.
- Configuration: prompts, schemas, code or platform settings that the agreement allows the client to own.
- Evaluation pack: development cases, the untouched acceptance set, first-pass scores, edited results, review effort and known failures.
- Operating guide: how to run, review, stop and escalate.
- Access record: accounts, roles, secrets owner and revocation steps without exposing secret values.
- Change log: versions used, decisions made and open issues.
- Support boundary: support dates, response channel, included fixes and the route for new work.
Run the handover as a test. Ask the client owner to operate the workflow from the guide while the consultant observes. Record where the owner has to ask for missing information, then repair the guide.
9. Decide what happens after the pilot
The pilot should end with a client decision. Expansion is a separate decision.
There are four legitimate outcomes:
- Stop: the workflow is not valuable or acceptable.
- Revise: the question is valuable, but the design or data is not ready.
- Adopt as a controlled manual workflow: users operate it with explicit review and limited scope.
- Plan production implementation: the evidence supports investing in integration, controls and operations.
The final report should show first-pass acceptance results, separately logged human edits and reviewer effort, failures, unresolved risks and the client owner’s decision. A polished demonstration is not a fifth outcome.
A proposal structure you can use
A first proposal can be concise if it completes the explanation:
- Client problem and pilot question
- Users, inputs and permitted data
- Workflow and deliverables
- Acceptance set, rubric and decision rule
- Client and consultant responsibilities
- Access, ownership and handover
- Timeline and checkpoints
- Price, assumptions and direct-cost treatment
- Exclusions and change control
- Post-pilot options and support boundary
This structure gives the buyer enough information to assess the work and gives the consultant a defensible boundary.
Poorna Reddy writes practical guidance for consultants turning AI prototypes into work a client can accept and operate. Follow Poorna on LinkedIn and visit Timo for the next worked example.
