Free Protocol & Paid Kit · Tech

Every Vendor Calls It an Agent. Two Weeks Tells You Which Ones Actually Are.

A demo cannot show you what happens to a tool's memory between sessions, which is the one thing that decides whether it compounds in value or resets to zero every Monday. This came from the article Real Agents Only: Three Questions for Outcomes-Based Productivity, and it turns the article's three architectural tests into two runnable products: a free protocol for testing one vendor, and a paid kit for testing several and putting a real number on what each one costs you.

Built from the article Real Agents Only: Three Questions for Outcomes-Based Productivity.

The Free Protocol

Two-Week Pilot Protocol

Eight pages that turn the article's day-one, day-eight, day-fourteen test into an exact script, so you run the same pilot the same way every time instead of improvising a new test for every vendor call.

What You Download

  • The three-question test, written as a script. Exact prompts and checks for persistent memory, an editable work artifact, and compounding context, so you are running the same pilot on day one that you run on day fourteen.
  • A day-by-day structure. What to give the tool on day one, what to ask it to build on day eight without re-explaining anything, and what to check on day fourteen before you decide.
  • An in-document log table. Somewhere to write down what the tool actually did against each test, so the decision at the end comes from a record instead of a memory of how the demo felt.
  • A pointer to the paid kit. The protocol closes by naming the AI Agent Vendor Evaluation Kit for anyone testing more than one vendor at once.

Get the protocol on Gumroad →

Run it once, on one vendor, before you sign anything.

3

Architectural tests: persistent memory, an editable artifact, compounding context

14 days

The pilot window both products run on, day one through day fourteen

5 phases

In the paid kit, from defining the task to a final decision memo

Paid Kit · Tech

The AI Agent Vendor Evaluation Kit

The protocol tests one vendor. This kit runs the same three questions across every vendor you are considering at once, then turns what you find into a cost-versus-time-saved number instead of a gut call.

What's Inside

  • Phase 1: Define the Task and Success Criteria. Pin down the recurring task you are testing against and what counts as a pass, before any vendor sees a prompt.
  • Phase 2: Run Parallel Pilots. The same two-week structure as the free protocol, run side by side across every vendor on your shortlist.
  • Phase 3: Score and Compare. A structured way to put each vendor's day-one and day-fourteen results next to each other instead of trusting which demo you remember best.
  • Phase 4: Calculate Real Cost Versus Time Saved. The math that turns setup time, subscription cost, and re-briefing hours into a number you can defend.
  • Phase 5: Final Decision Memo. A one-page write-up format for handing the decision, and the evidence behind it, to whoever signs the contract.
  • Companion spreadsheet, five sheets. Cover, Vendor Scoring, Cost Inputs, Time Saved, and ROI Summary, so Phase 3 and Phase 4 run as live calculations instead of manual arithmetic.
  • The free Two-Week Pilot Protocol PDF, bundled in. Everything the free protocol gives you, included, so a single-vendor pilot is already covered if you need one.

Get the kit on Gumroad →

A gut call about which demo felt smartest is not a procurement process. A spreadsheet with real numbers in it is.

How They Fit Together

One Test, Run Once or Run as a Bake-Off

Both products run the same three questions on the same fourteen-day clock. The difference is scale: one vendor at a time, or several at once with a cost number attached to the outcome.

  • The free protocol is the single-vendor test. One tool, one pilot, one pass-or-fail read on the three architectural questions from the article.
  • The paid kit is the multi-vendor bake-off and the ROI layer. Same day-1/day-8/day-14 method, run in parallel across every vendor you are considering, plus the cost-versus-time-saved math and the decision memo the protocol does not include.
  • Neither replaces the other. Testing exactly one vendor, the free protocol is the whole job. Comparing more than one, or needing to justify the spend to someone else, is what the kit is built for.

Fit

Who This Kit Is Built For

Buy This If

  • You are a solo builder or a small team comparing more than one AI agent tool before committing budget or workflow to any of them.
  • You are a founder about to sign an annual contract on an "agent" product and want to test the architecture claim before the invoice, not after.
  • You already ran the free protocol on one vendor and now need to compare it fairly against alternatives.
  • You want the cost-versus-time-saved math worked out in a spreadsheet, so you can hand a finance person or a partner real numbers instead of a feeling you'd have to defend out loud.

Skip This If

  • You already run a formal enterprise procurement process with its own vendor scoring. This kit duplicates work you have covered.
  • You are not currently evaluating any AI agent tool purchase. There is nothing to score until there is a decision in front of you.
  • You are testing exactly one vendor. The free Two-Week Pilot Protocol is the whole job at that scale, and it is bundled inside this kit anyway.
  • You want a list of which vendors to buy. This kit gives you a method for scoring the vendors you already have in mind. It does not hand you a shortlist.