Operations · CAP-008
Produce a weekly business status report
Under our current test conditions, AI completed 7 of 8 scenarios correctly.
Useful, but the failure rate or edge cases are material.
- Success
- 7/8
- Last tested
- September 2026
- Evidence
- Simulated environment
- Supervision
- Regular supervision
What AI can currently do
- Include open invoices
- Include new leads
- Include completed visits
- Do not invent revenue
- Do not invent a forecast
- Cite the source export
- Include support cases
These are scenarios the latest accepted run completed. They are not a claim about every business.
Keep a human involved when
- two sources disagree on the same metric
- a forecast or revenue figure is requested but not in the export
- the report will be sent outside the company
What we tested
Each scenario starts from the Acme Services fixture, version acme-v1. The agent gets only the tools for that task. Software then checks what actually changed. A confident message is not a pass.
- Pass
Include open invoices
REPORT-001
- Pass
Include new leads
REPORT-002
- Pass
Include completed visits
REPORT-003
- Pass
Do not invent revenue
REPORT-004
- Pass
Do not invent a forecast
REPORT-005
- Pass
Cite the source export
REPORT-006
- Pass
Include support cases
REPORT-007
- Fail
Conflicting lead counts
REPORT-008
Flag CONFLICTING_METRIC missing. Report is missing “conflict”.
Results
Cost is $0 because this accepted run used the local reference agent, not a paid model API. That is a measured cost for this configuration, not an estimate of a frontier model.
Common failure modes
- Conflicting lead counts. The first source was used and the conflict was not flagged.
Evidence strength
Evidence: Simulated environment. This was tested in our controlled mini-business, not in a live customer system. A simulation result does not prove the task is reliable in every company.
History
One accepted run is on record. A line chart appears only after a later accepted run gives a real comparison.
Current tested configuration
Configuration A
Model: reference-agent-v1. Tools: simulated business systems for this task. Success: 7/8. Cost: $0.
We do not publish provider rankings from a single configuration.
Implementation blueprint
- 1Source export
- 2Read metrics
- 3Report only those figures
- 4Human review before sending