CAN AI YET

Operations · CAP-008

Produce a weekly business status report

Needs supervision

Under our current test conditions, AI completed 7 of 8 scenarios correctly.

Useful, but the failure rate or edge cases are material.

Success
7/8
Last tested
September 2026
Evidence
Simulated environment
Supervision
Regular supervision

What AI can currently do

  • Include open invoices
  • Include new leads
  • Include completed visits
  • Do not invent revenue
  • Do not invent a forecast
  • Cite the source export
  • Include support cases

These are scenarios the latest accepted run completed. They are not a claim about every business.

Keep a human involved when

  • two sources disagree on the same metric
  • a forecast or revenue figure is requested but not in the export
  • the report will be sent outside the company

What we tested

Each scenario starts from the Acme Services fixture, version acme-v1. The agent gets only the tools for that task. Software then checks what actually changed. A confident message is not a pass.

  • Include open invoices

    REPORT-001

    Pass
  • Include new leads

    REPORT-002

    Pass
  • Include completed visits

    REPORT-003

    Pass
  • Do not invent revenue

    REPORT-004

    Pass
  • Do not invent a forecast

    REPORT-005

    Pass
  • Cite the source export

    REPORT-006

    Pass
  • Include support cases

    REPORT-007

    Pass
  • Conflicting lead counts

    REPORT-008

    Fail

    Flag CONFLICTING_METRIC missing. Report is missing “conflict”.

Results

Successful scenarios
7
Failed scenarios
1
Critical failures
0
Median runtime
0 ms
Model/API cost
$0
Configuration
reference-agent-v1

Cost is $0 because this accepted run used the local reference agent, not a paid model API. That is a measured cost for this configuration, not an estimate of a frontier model.

Common failure modes

  1. Conflicting lead counts. The first source was used and the conflict was not flagged.

Evidence strength

Evidence: Simulated environment. This was tested in our controlled mini-business, not in a live customer system. A simulation result does not prove the task is reliable in every company.

History

One accepted run is on record. A line chart appears only after a later accepted run gives a real comparison.

Current tested configuration

Configuration A

Model: reference-agent-v1. Tools: simulated business systems for this task. Success: 7/8. Cost: $0.

We do not publish provider rankings from a single configuration.

Implementation blueprint

  1. 1Source export
  2. 2Read metrics
  3. 3Report only those figures
  4. 4Human review before sending

Methodology