CAN AI YET

Operations · CAP-007

Turn meeting notes into follow-up actions

Needs supervision

Under our current test conditions, AI completed 7 of 8 scenarios correctly.

Useful, but the failure rate or edge cases are material.

Success
7/8
Last tested
September 2026
Evidence
Simulated environment
Supervision
Regular supervision

What AI can currently do

  • Explicit action line
  • Owner will do something by a date
  • Do not invent an owner
  • A decision is not an action
  • Two explicit actions
  • Todo prefix
  • Empty notes

These are scenarios the latest accepted run completed. They are not a claim about every business.

Keep a human involved when

  • the follow-up is implied in prose rather than written as an action
  • an owner or date would have to be guessed
  • a decision is being treated as a task

What we tested

Each scenario starts from the Acme Services fixture, version acme-v1. The agent gets only the tools for that task. Software then checks what actually changed. A confident message is not a pass.

  • Explicit action line

    MEET-001

    Pass
  • Owner will do something by a date

    MEET-002

    Pass
  • Implicit follow-up in prose

    MEET-003

    Fail

    Actions miss “circle back”.

  • Do not invent an owner

    MEET-004

    Pass
  • A decision is not an action

    MEET-005

    Pass
  • Two explicit actions

    MEET-006

    Pass
  • Todo prefix

    MEET-007

    Pass
  • Empty notes

    MEET-008

    Pass

Results

Successful scenarios
7
Failed scenarios
1
Critical failures
0
Median runtime
0 ms
Model/API cost
$0
Configuration
reference-agent-v1

Cost is $0 because this accepted run used the local reference agent, not a paid model API. That is a measured cost for this configuration, not an estimate of a frontier model.

Common failure modes

  1. Implicit follow-up in prose. The phrase circle back was not extracted.

Evidence strength

Evidence: Simulated environment. This was tested in our controlled mini-business, not in a live customer system. A simulation result does not prove the task is reliable in every company.

History

One accepted run is on record. A line chart appears only after a later accepted run gives a real comparison.

Current tested configuration

Configuration A

Model: reference-agent-v1. Tools: simulated business systems for this task. Success: 7/8. Cost: $0.

We do not publish provider rankings from a single configuration.

Implementation blueprint

  1. 1Notes
  2. 2Extract explicit actions
  3. 3Structured follow-up list
  4. 4Human review of implied items

Methodology