How to hire an AI agent agency without buying a demo
Step-by-step hiring process for an AI agent agency: brief, shortlist, pilot boundary, acceptance test, and ownership transfer.
Written by Northstar
Northstar is an AI agent systems studio. Alex leads engineering and product systems; Jordan leads operations and workflow fit. We ship production agents inside tools teams already use.
Alex Morgan · LinkedIn · Northstar
On this page
Direct answer
Hire an AI agent agency by asking it to prove one workflow, not one demo.
Start with a one-page brief, keep the shortlist small enough to compare on the same evidence, run enough discovery to pressure-test the workflow, and sign only when the pilot names the acceptance test, the gate rules, the logging rules, and the operating model after launch.
OpenAI's practical guide is useful here because it says agents belong where deterministic and rule-based approaches fall short.
If the workflow is still a form, queue, or rules-engine problem, do not turn agency selection into an agent sale.
Why the market is noisy
Gartner warns that agent washing is real.
It says many products are being rebranded as agentic without substantial agentic capability, and it argues that many use cases positioned as agentic do not actually require agentic implementations.
That is why a buyer should not trust logos, awards, or pitch polish as proof.
Use vendor directories to build a longlist.
Do not use them as evidence that a team can ship production work.
Write the brief first
The one-page brief is the first real filter.
It should describe the actual workflow in the order it happens today.
It should also show the vendor where the risk lives.
| Brief field | Why it matters |
|---|---|
| Recurring work, step by step | Prevents a vendor from guessing at the workflow shape |
| Tools already in use | Forces scoped integration into real systems |
| Weekly and peak volume | Makes load and cost concrete |
| Exceptions that always go to a human | Defines the boundary of autonomy |
| Actions that must never run without review | Protects refunds, sends, deletions, and other irreversible steps |
| What "working" means | Gives the pilot an acceptance test |
| Constraints such as data residency or compliance | Prevents late scope surprises |
| Budget envelope and decision date | Keeps the process bounded |
A vendor that starts pitching before reading this brief is selling a template.
That is not a hiring process.
It is a sales process.
Shortlist on evidence
Keep the shortlist small enough that every finalist answers the same questions.
Do not enlarge the list just to feel safe.
Too many finalists dilute the comparison and make the brief worse.
Ask each vendor for the same inputs.
You are looking for process proof, not decoration.
| Discovery area | Ask this | A good answer looks like |
|---|---|---|
| Workflow mapping | How do you map the workflow before writing code? | They ask about steps, tools, exceptions, and ownership before they talk about models |
| Gates | Which actions would you gate behind human approval? | They can name irreversible actions and explain the approval boundary |
| Observability | Where do logs and traces live, and who can read them? | They can show how the team reconstructs a run |
| Cost | What will model usage cost at our volume? | They can estimate usage without hand-waving |
| Ownership | What is in the handoff package? | They can name docs, runbooks, eval sets, training, and the support or handoff path |
OpenAI says to build agents where deterministic and rule-based approaches fall short.
That means the shortlist should favor vendors who can explain why the workflow needs agentic behavior at all.
If they cannot explain that, the fit is weak.
Buy a fixed-scope pilot
Do not buy a broad program before the first workflow passes.
Buy a pilot with a written boundary.
IBM frames deployment as a real-environment problem.
Its guidance calls out consistent development, testing, and production environments, plus the need to connect agents to business systems.
Dataiku is even more explicit.
It defines production-ready agents as systems operating in a live environment with real users, real data, and real consequences, backed by guardrails, monitoring, access controls, rollback, and defined success metrics.
The pilot contract should name the following items.
| Pilot item | What to require |
|---|---|
| Acceptance test | Inputs, expected outputs, pass threshold, and the judge |
| Gate rules | Which actions require approval and who approves them |
| Logging rules | What gets recorded, where it is stored, and who can review it |
| Cost responsibility | Who pays the model bill and what the cap is |
| Explicit exclusions | What is not included in the first pilot |
| Handoff package | Docs, runbook, eval set, and training |
| Support window | How long the vendor stays on call after go-live |
| Ownership | Who owns code, prompts, and evaluation assets |
The operating model after launch can be a clean handoff or managed support.
The point is to name it before anyone signs.
If the vendor says "we will figure that out during the build," the scope is not fixed.
That is a red flag.
Judge the pilot like an operator
The pilot is not a presentation.
It is a small production-like test.
The buyer should care about the process as much as the output.
- Do approvals route to the right person.
- Does the system stop cleanly when it hits a risky action.
- Can the team retrieve logs and traces after the run.
- Can the vendor explain cost at real volume.
- Can the buyer operate the workflow without the vendor standing over it.
NIST's AI Risk Management Framework is useful here because it frames AI as a trustworthiness and risk-management problem, not a novelty contest.
OWASP's 2026 agentic Top 10 is also useful because it gives buyers current security language for autonomous and agentic systems.
Use both to keep the pilot grounded in control, not theater.
Red flags
| What you hear | What it usually means |
|---|---|
| "Let us show you a demo first" | The vendor is selling before understanding the workflow |
| "Unlimited agents" | The scope is probably vague |
| "Approvals slow things down" | They may not have a serious gate model |
| "Documentation comes after rollout" | Knowledge will stay trapped with the vendor |
| Broad admin credentials "just for the pilot" | The vendor is asking for more access than the task needs |
| "We can handle that later" | The pilot is drifting away from fixed scope |
Do not confuse friendliness with reliability.
Do not confuse confidence with ownership.
Do not confuse a polished meeting with a production-ready system.
How Northstar fits
Use the hiring page to get to a signed pilot.
Use the comparison page if you need a bake-off.
Use the criteria page if you need a rubric before the shortlist.
If you want a structured next step, start from solutions and bring one real workflow, the systems it touches, and the actions that must stay gated.
Related reading:
FAQ
Not always. A good one-page brief and enough discovery to pressure-test workflow, controls, and ownership are enough for many teams. Use an RFP only when your procurement process needs it.
