WORKSHOP · STRATEGY · WS·07

From Conversation to Capability

Turning discovery-call transcripts into a deployable AI roadmap. A four-stage pipeline that decomposes interviews into atoms, clusters tasks across people, and outputs an auditable hours build-up the CFO can sign. 24 slides, 30 minutes.

24 SLIDES · ~30 MIN · PLAYBOOK OVERVIEW
Key Takeaways
  • The transcript is the raw ore. The pipeline is the refinery.
  • Pick one process and one person. Not 'do AI strategy.' Pick one workflow and one human who runs it. Record one 45-minute conversation. The pipeline starts with a single honest interview, not a roadmap.
  • Turning discovery-call transcripts into a deployable AI roadmap. A four-stage pipeline that decomposes interviews into atoms, clusters tasks across people, and outputs an auditable hours build-up the CFO can sign. 24 slides, 30 minutes.
WORKSHOP · STRATEGY

From Conversation to Capability

Turning discovery-call transcripts into a deployable AI roadmap.

BuildClub Academy
01 · THESIS

The transcript
is the raw ore.

The pipeline is the refinery.
A 45-minute discovery call decomposed
into machine-readable units of work —
then reassembled into an auditable map
of where AI agents can be deployed.

The Thesis
02 · WHY THIS MATTERS

Three reasons reading transcripts by hand fails at executive scale.

1
Loss
A transcript contains numbers, tools, frustrations, and risks that leak out when humans skim and summarize. The artifacts the CFO will actually fund are the ones a fast reader rounds away first.
2
Seams
The highest-value insight lives between interviews, not in any one of them. One head cannot hold them all. The hour you can defend is the hour that shows up in two different conversations.
3
Proof
Evidence unlocks budget. 'Your team wastes time on spreadsheets' does not. Auditable hours and named clusters do — because the CFO can ask which person, which week, which tool.

Conversations are lossy. Insight hides between them. Executives fund evidence, not vibes.

03 · THE CEO QUESTION

You are not deciding 'do AI.' You are deciding where AI absorbs work.

Where
Which specific clusters of work — not departments — can an agent absorb today? The unit of analysis is a task pattern, not an org-chart box.
How big
How many recoverable hours sit in each cluster, and how confident is that number? An hour with provenance is fundable; an hour without is folklore.
Who decides
Which tasks are rote, which are rules-based, which require human judgment? This is the single field that gates every automation decision downstream.
What's next
Which three to five agents do you build first, and what is the order? The roadmap is a ranked queue, not a wish list.
04 · ARCHITECTURE

Four stages turn one conversation into one defensible roadmap.

1
2
3
4
1
Ingest
Fetch · dedupe · classify · resolve speakers · anchor in the system of record. Before any analysis runs, every transcript is attributed to named people in a named meeting.
2
Analyze
Seven sequential passes per interview, starting with PASS 0 — Resolve (the entity-resolution gate covered in WS-08 Company Brain). After that, six analytical passes ask one question each — process, friction, knowledge, opportunity, strategic, quantified. Shallow summaries are designed out of the system.
3
Synthesize
One overpass across all interviews. Cluster the atoms, corroborate the claims, rank the recoverable-hours catalog. Insight that lives between conversations becomes visible here, not before.
4
Deploy
Auditable hours build-up, ranked task-cluster catalog, candidate agents. The output is a roadmap a CFO can sign and an engineering lead can scope.

Three storage layers, one job each: Index (system of record), Content (verbatim + analysis, paired), Logic (versioned prompts and schemas). The pipeline is boring on purpose.

05 · STAGE ONE — INGEST

Before any analysis: dedupe, classify, resolve speakers, anchor.

1
2
3
4
5
1
Fetch
A nightly cron pulls the latest transcribed calls from the recording service. No human in the loop; no missed days.
2
Dedupe
Cross-reference the index; skip already-processed calls and empty stubs from failed recordings. The dataset stays clean by construction.
3
Classify
Tag each transcript by meeting type, client, topics, phase. The classification decides which downstream analysis runs — and which prompt version is used.
4
Resolve
Map 'Speaker 1 / Speaker 2' to named individuals using attendees, self-intros, and a confidence floor. Anything below the floor is queued for human review, not silently guessed.
5
Anchor
Match or create the parent meeting row; build bidirectional relations that make every transcript, atom, and quote navigable from any direction.

Attribution is the foundation. Every extracted task and number traces back to a named person and role — not 'Speaker 3.'

06 · STAGE TWO — ANALYZE

One transcript. Seven sequential passes. Seven different questions.

Pass 0 · Resolve
The entity-resolution gate. Every mention of a person, system, or company maps to a canonical ID — or the pipeline pauses for human clarification. Nothing downstream runs until this clears. Covered in depth in WS-08 · The Resolution Layer.
Pass 1 · Process Mapping
What actually happens — workflows, steps, tools, handoffs, triggers, timing. Extractive only: no inference allowed at this layer.
Pass 2 · Friction & Pain
Where it hurts — frustrations, workarounds, manual effort, emotional language. Emotion is signal. The verbatim is the data.
Pass 3 · Knowledge & Risk
What is fragile — key-person dependencies, undocumented processes, single points of failure. The risk register the business does not have.
Pass 4 · Opportunity
What they want — explicit requests, success criteria, constraints, what they have already tried. The wish list, scoped to what was actually said.
Pass 5 · Strategic Interp.
What it means — change-readiness, politics, dynamics. The one and only interpretive pass. Deliberately isolated so its judgment cannot contaminate the others.
Pass 6 · Quantified Baseline
The numbers — times, volumes, dollars, headcount. Raw material for the business case. The Triple lives here: frequency, volume, time-per-unit.
07 · WHY MULTI-PASS

Asking one big question gets a shallow summary. Asking six small ones — after PASS 0 — gets depth.

One generic pass
Reads the transcript once. Produces a shallow summary that smooths over frustrations and rounds numbers into adjectives. Misses the things that matter most: emotional language, specific dollar figures, named risks. The naive prompt always ends here.
Six analytical passes (after Pass 0)
Once entities are resolved, the analyzer reads the transcript six more times, each pass hunting for one category. Catches what a generic summary would average away. Same logic as a human review board assigning different reviewers different lenses — but cheaper, faster, and consistent.
Cognitive blending — bad
Mixing extractive and interpretive work in one prompt produces confident-sounding analysis nobody can defend. The model hides the seam between what was said and what it concluded. Trust falls when the CFO probes.
Cognitive separation — good
Each pass enforces a different evidence standard. Extractive passes are forbidden from inferring; the interpretive pass is the only one allowed to. Surface the seams; they become audit points, not bugs.

Cognitive separation beats cognitive blending — for models and for humans.

08 · EXTRACTION PRINCIPLES

Two non-negotiable rules govern the analysis prompt.

Preserve specificity
'Sarah spends 2–3 hours every Monday on the AP reconciliation' — not 'significant time spent on reporting.' Exact numbers, named people, real tool names. Specificity is what survives translation into a business case.
Quote generously
'Everyone knows it — it's just not documented' — not 'institutional knowledge gaps exist.' The verbatim carries emotion, hedging, and politics. Those become the punchy lines that land with leadership weeks later.
Generalization — banned
'Significant time on manual work' is rejected by the schema validator. The model cannot soften, summarize, or genericize. If a sentence cannot point to a verbatim, it does not enter the pipeline.
Fidelity over brevity
We optimize the analysis for fidelity, not brevity. A human never reads the per-interview output — the next stage does. Length is free; lost specificity is not recoverable downstream.

The prompt is engineered against the room's instinct to summarize. Summarization is the enemy of an auditable hour.

09 · THE AGGREGATION PROBLEM

Prose is for humans. Prose cannot be summed.

1
AP Analyst
'I reconcile the vendor file against the GL every Monday morning.' One person, one process, one verbatim. The pipeline records this exactly.
2
Controller
'My month-end close pain is tying the sub-ledger back to the trial balance.' Different role, different language, same task type. Prose alone won't show that.
3
FP&A Manager
'I spend Fridays matching forecast actuals to what hit the books.' Third name, third workflow, same underlying pattern: reconciliation. The data layer is what makes the pattern visible.

Three people, one cluster — but only if the data layer says so. At organizational scale, 'a human will spot the pattern' is not a plan. You need a structured layer.

10 · THE INNOVATION

Append a machine-readable atom layer to every prose analysis.

Task atoms
One per discrete, repeatable task. Fields: task_type · judgment_level · hours/week · tools · actor role. The clustering key sits here — same task_type across people defines a cluster.
Metric atoms
One per quantified claim. Fields: value · unit_class · frequency · annualized · confidence. The math you can defend in a board meeting lives here.
Schema as a gate
Schema-validated as a hard gate — incomplete output is rejected. Every atom carries a stable ID like T-kim-01 or M-ryan-03. Every ID traces back to a named speaker and a verbatim quote.
Paired forever
Prose and atoms live in the same file, never separated. The atoms are the math layer; the prose is the audit trail. Lose one and the other becomes worthless.

In addition to the six analytical prose passes (after PASS 0 — Resolve), the analysis prompt emits a final structured block — a JSON layer of atoms designed to be aggregated by machines, not humans.

11 · TASK ATOMS

What a task atom looks like — and why each field exists.

task_type
Controlled vocabulary. The clustering key. Examples: reconciliation, data-entry-transcription, report-building, portal-checking. Free text would defeat the cluster — controlled vocab makes it possible.
actor_role
Named person plus role. Never 'Speaker 3.' Provenance must survive all the way into the spreadsheet, or the downstream consumer cannot defend the number.
tools
Real system names — the Excel file, the ERP module, the portal. Lets you spot tool-driven clusters (every reconciliation that touches the same legacy portal, for example).
hours_per_week
Time estimate, tagged by how it was known: stated, volume-derived, or inferred. The confidence stratum is part of the data, not a footnote.
judgment_level
rote · rules-based · judgment-heavy. The single field that decides where an agent can go and where it cannot. Every other field supports this one.
12 · METRIC ATOMS

unit_class is the field that keeps you from adding dollars to hours.

value
The raw number as spoken: 2.5, 17, 4,000. No unit conversion at extraction time. The interpretation layer never blurs into the data layer.
unit_class
time · money · volume · count · percent. The system sums within a class — never across. The single most common AI-extraction bug is silently summing dissimilar units; this field prevents it.
frequency
How often. Daily, weekly, monthly, annual. The bridge from a one-off mention to a per-year figure. Without it, you cannot annualize defensibly.
annualized
The math, shown. value × frequency, expanded — so a reviewer can see how the headline number was built. Black-box totals lose budget meetings.
confidence
How well-supported the number is. Stated > volume-derived > inferred. Drives a credibility floor; low-confidence metrics never make it into the headline.
13 · MENTAL MODEL

The Task Cluster Method — decompose, cluster, deploy.

1
2
3
1
Decompose
Break each job into atomic tasks. Stop thinking in job titles. Start thinking in repeatable units of work. 'AP Analyst' is not the unit; 'reconcile vendor file to GL on Monday morning' is.
2
Cluster
Group the same task_type across people, regardless of department. A cluster is where six different humans do the same kind of work. Departments hide clusters; the atom layer reveals them.
3
Deploy
Build one agent per high-value cluster. The cluster, not the role, is the unit of automation. An agent that does reconciliation for six people beats six pilots that each touch one person.

Decompose the job into tasks. Cluster the tasks across people. Deploy an agent against the cluster.

14 · JUDGMENT_LEVEL

One field decides which tasks become agents and which stay with humans.

Rote — agent candidate
Same input, same output, every time. The agent absorbs the whole task. No exception path needed. Reconciliation between two clean systems is the canonical example.
Rules-based — agent + review
Predictable logic, but multiple branches. The agent handles the path; a human reviews exceptions. Vendor portal checking with conditional triggers lives here.
Judgment-heavy — human-led
Context, negotiation, taste. The agent prepares the work; the human decides. Variance memos and supplier disputes belong on this side of the line.
Genuinely creative — no agent
Strategy, relationships, ethics. The agent has no business here yet — protect it. Naming this category explicitly is what stops scope creep.

'We augment people. We do not replace them.' — operationalized as a database field, not a slogan.

15 · STAGE THREE — SYNTHESIZE

The Overpass — one pass that reads every interview at once.

Reasoning sections
Built from the verbatim prose. Cross-person corroboration, severity re-ranking, ownership-gap detection, and the quotes that land. The AI's interpretation layer — but constrained to the evidence pool.
Deterministic sections
Computed from the atoms, not reasoned about. Auditable recoverable-hours build-up, ranked task-cluster catalog, coordination ledger. Math, not opinion.
Per-interview ≠ engagement-wide
The per-interview analysis runs many times. The Overpass runs once — across every interview in the engagement at the same time. Concatenation gives you a pile. Synthesis gives you understanding.
AI reasons, atoms compute
We let the AI reason in prose where judgment matters, and compute from atoms where math matters. Mixing the two is the failure mode the Overpass design avoids.

Two halves, two evidence standards. The reasoning side hires the AI; the deterministic side fires it.

16 · WHAT ONLY THE OVERPASS CAN SEE

Three findings that exist only when every interview is held in one head.

1
Corroboration
Two people describing the same reality from different seats — one's 70% gut estimate corroborated by another's 60% data figure — becomes a confidence signal no single interview produces. The number is no longer one person's opinion.
2
Severity re-rank
The loudest complaint is rarely the biggest risk. The Overpass demotes noisy-but-survivable pains beneath quiet-but-existential ones. Volume in a transcript is not the same as importance to the business.
3
Seam detection
The highest-value finding usually lives in handoffs both teams assume the other owns. Nobody says it out loud — it emerges only in the overlap. Seams are where AI's biggest leverage hides.

Insight lives between the interviews, not within them. The Overpass is the only pass that can see the between.

17 · CFO AUDITABILITY

An AI number a CFO cannot audit is worthless.

1
Provenance
Every hour is tagged by how it was known: stated, volume-derived, or inferred. The system reports them separately — the 1,820-hour line and the 'inferred only' bucket never get summed into a single headline.
2
Honesty
Tasks without a time basis are not dropped. They go into a named follow-up queue. The headline number is explicitly a floor — not the answer, the lower bound of the answer.
3
Traceability
Every aggregate traces to atom IDs. Every ID traces to a named speaker and a verbatim quote. Line-by-line audit, end to end. The CFO can interrogate any cell in the spreadsheet.

'The question is never do you trust the AI? — it's here's the line item, here's who said it, here's the assumption.'

18 · FEEDBACK LOOP

The pipeline taught us how to interview. The output improved the input.

1
Frequency
'How often do you do this?' → weekly, monthly, daily. The easy question. Interviewers ask it without prompting.
2
Volume
'How many at a time?' → 12 invoices, 200 line items. The mid-difficulty question. Captured maybe half the time.
3
Time-per-unit — the missing question
'How long does one take you?' → the question that was asked zero times across the entire first engagement. Asking it would have rescued ~17 of 19 unquantified clusters.

The Triple — capture all three or it's not an hour. Empirical audit across 22 task clusters proved the discipline.

19 · FAILURE MODES

The five places a pipeline like this breaks — and how to harden each.

1
2
3
4
5
1
Empty stubs
Failed recordings poison the dataset. Detect zero-duration and null-host transcripts at ingest; skip them. One bad stub propagated downstream contaminates the cluster counts.
2
Speaker drift
Anonymous speaker labels destroy provenance. Enforce a confidence floor; route low-confidence calls for human review rather than silently guessing a name. 'Speaker 3' is never an answer.
3
Prompt creep
Analysts add fields; outputs balloon; the schema drifts. Version prompts in git with a change log and a kill switch to roll back. Prompts are code — treat them that way.
4
Silent malformed atoms
A bad analysis stored as if it were good is the worst possible outcome. Schema-validate atoms as a hard gate. Reject, do not store. Better an empty row than a corrupt one.
5
Hours without provenance
A headline number nobody can defend kills the roadmap. Tag every hour stated/derived/inferred and report them separately. Mixed buckets look impressive and fund nothing.

Every one of these came from a real failed run. None was theoretical.

20 · CEO DECISION FRAMEWORK

Five questions to ask before approving any AI deployment plan.

1
2
3
4
5
1
Is the unit of analysis a task, not a department?
If your roadmap is organized by team, you are still doing org-chart AI. Reorganize around task clusters. Departments are where AI dies.
2
Is every hour traceable to a person and a quote?
If the answer is 'the model said so,' reject the number. Demand atom-level provenance — name plus verbatim. No exceptions.
3
Is judgment_level a column in the plan?
If you cannot point to which tasks are rote, rules-based, or judgment-heavy, you have not done the work. That column gates every automation downstream.
4
Is there a follow-up queue for unquantified work?
A roadmap that hides its gaps is propaganda. A roadmap that names them is a plan. Honesty about what you don't know is the new differentiator.
5
Is the build order ranked by hours × automatability?
Without an explicit ranking rule, you will build agents for the loudest cluster, not the biggest one. Loud is not the same as large.

Five binary questions. If any answer is 'no,' the plan is not yet a plan.

21 · EXAMPLE OUTPUT

What the deployment roadmap looks like when it lands on your desk.

Reconciliation (AP/GL) — 1,820 hrs · stated
Rules-based · High confidence · Build order #1. The cluster where six different roles run the same AP-to-GL reconciliation each week. Highest hours, highest confidence, cleanest agent shape.
Vendor portal checking — 1,140 hrs · stated
Rote · High confidence · Build order #2. Anonymous bot work on third-party portals. Rote enough for an agent; high enough hours to fund the build twice over.
Invoice data entry — 920 hrs · derived
Rote · High confidence · Build order #3. Volume-derived hours, not stated — but the math is fully expanded in the build-up. CFO can audit every line.
Monthly variance memos — 680 hrs · stated
Rules-based · Medium confidence · Build order #4. Predictable structure but real judgment calls on outliers. Agent prepares the memo; controller signs it.
Forecast assumptions — ~unquantified · 3 clusters
Judgment-heavy · Follow-up queue. Three clusters surfaced; hours not yet defendable. Explicitly tracked as unquantified — not pretended away.
22 · MONDAY-MORNING ACTIONS

Five moves you can start before lunch on Monday.

Monday Morning
  1. 1
    Pick one process and one person. Not 'do AI strategy.' Pick one workflow and one human who runs it. Record one 45-minute conversation. The pipeline starts with a single honest interview, not a roadmap.
  2. 2
    Capture the Triple in every interview. Frequency × volume × time-per-unit. Train your team on this one rule. It rescues most unquantified work and turns folklore into a fundable hour.
  3. 3
    Separate content, index, and logic. Decide where transcripts live, where metadata lives, where prompts live. Three locations, not one. The boring half of the architecture is the half that keeps it from drifting.
  4. 4
    Define your judgment_level vocabulary. Decide today what counts as rote, rules-based, and judgment-heavy in your business. Write it down. Without this, every automation decision becomes a debate.
  5. 5
    Demand provenance from any vendor. No hours figure without an atom ID, a speaker name, and a verbatim quote. That single rule filters 80% of vendor pitches before they get past the cover slide.
CLOSING

Bottom-up evidence.
Auditable hours.
Agents deployed against
clusters — not departments.


1 — The transcript is the raw ore. The pipeline is the refinery.
2 — Prose for humans. Atoms for math. Keep both.
3 — Reason where judgment matters. Compute where math matters.
4 — An AI number a CFO can't audit is worthless.
5 — Decompose. Cluster. Deploy.

The Thesis
WS·07 · Questions CEOs Ask

Frequently Asked Questions

What is the core idea of From Conversation to Capability?
The transcript is the raw ore. The pipeline is the refinery.
What should a CEO do Monday morning after reading From Conversation to Capability?
Start here: Pick one process and one person. Not 'do AI strategy.' Pick one workflow and one human who runs it. Record one 45-minute conversation. The pipeline starts with a single honest interview, not a roadmap; Capture the Triple in every interview. Frequency × volume × time-per-unit. Train your team on this one rule. It rescues most unquantified work and turns folklore into a fundable hour; Separate content, index, and logic. Decide where transcripts live, where metadata lives, where prompts live. Three locations, not one. The boring half of the architecture is the half that keeps it from drifting.
What are the steps in From Conversation to Capability?
1) Ingest; 2) Analyze; 3) Synthesize; 4) Deploy; 1) Fetch.
Where do the claims in From Conversation to Capability come from?
The playbook cites Internal workshop source material; WS-08 Company Brain · The Resolution Layer.
00 / 24